12 Sep
|
Utho Cloud
|
Noida
Cloud Support Engineer Level 2 / Level 3
Designation :Cloud Support Engineer (L2 / L3)
Experience- L2: 3-5 years | L3: 5-8 years
Location: Noida, Uttar Pradesh (work from office)
Interview : F2F only No virtual Round
Engagement :Full-time, 247 rotational shifts
POSITION OVERVIEW
Utho operates an Indian public cloud platform delivering compute, block and object storage, networking and Managed Kubernetes
Service (MKS). We are hiring experienced Cloud Support Engineers to own the technical resolution of escalated customer issues. Level 2 is accountable for diagnosis and resolution across the stack; Level 3 is the final technical authority for Support, leading major incidents and root cause analysis with Engineering.
KEY RESPONSIBILITIES
1. Ownership
Level 2 Diagnosis & Resolution
- Own escalated tickets end-to-end and ensure timely resolution within defined SLAs.
- Keep customers informed throughout the troubleshooting and resolution process.
Level 3 Technical Authority
- Act as the final technical escalation point for the Support team.
- Lead major incidents and drive resolution of critical technical issues.
- Own Root Cause Analysis (RCA) and work with Engineering teams to implement permanent fixes.
1. Compute
Level 2 Diagnosis & Resolution
- Diagnose VM boot failures, guest kernel panics, Virtio device faults, CPU steal, and I/O contention at the hypervisor level.
Level 3 Technical Authority
- Resolve host and fleet-wide compute issues involving NUMA, CPU pinning, memory pressure, and live migration failures.
- Review and approve corrective actions on production hypervisors.
1. Storage
Level 2 Diagnosis & Resolution
- Resolve stuck attach and detach states, block-device path failures, filesystem issues, and mount recovery.
- Perform and validate storage restore and recovery activities.
Level 3 – Technical Authority
- Diagnose replication lag, split-brain scenarios, and data-consistency issues.
- Own capacity forecasting and define recovery and remediation runbooks.
1. Network
Level 2 – Diagnosis & Resolution
- Troubleshoot load balancer, NAT, conntrack, MTU, and fragmentation issues.
- Perform packet captures, TLS handshake analysis, and latency troubleshooting.
Level 3 – Technical Authority
- Handle complex overlay-network, asymmetric-routing, and cross-datacenter issues.
- Work with the Network team to drive permanent fixes and validate solutions at scale.
1. Kubernetes
Level 2 – Diagnosis & Resolution
- Resolve Kubernetes node NotReady issues, CNI and CSI attach loops, kubelet and containerd faults.
- Troubleshoot pod eviction, cgroup-related issues, and node-pool scaling problems.
Level 3 – Technical Authority
- Own control-plane and etcd health, certificate rotation, cluster upgrades, and capacity management.
- Advise customers on production Kubernetes architecture and best practices.
1. Continuous Improvement
Level 2 – Diagnosis & Resolution
- Create and maintain technical runbooks and Knowledge Base articles.
- Develop scripts for repeatable diagnostics and recurring technical issues.
Level 3 – Technical Authority
- Build automation and internal tooling to eliminate recurring issues.
- Mentor L1 and L2 engineers and review their RCAs to improve technical quality and resolution effectiveness.
SKILLS REQUIRED
- Linux — advanced administration on Ubuntu, Debian and RHEL/CentOS: systemd, journalctl, boot and rescue recovery, fstab and
mount repair, LVM, ext4 and XFS; kernel and sysctl tuning, OOM and kernel-panic analysis, profiling with iostat, vmstat, sar, perf and strace.
- Cloud and virtualization — hands-on with a public cloud or virtualization platform: VM lifecycle, snapshots and images, block and
object storage, VPC and IP management, firewall rules, load balancers and DNS. KVM, QEMU and libvirt required at L3.
- Networking — TCP/IP and subnetting, DNS, static and energetic routing, NAT and conntrack, VLAN and overlay networking, iptables or
nftables chain analysis, load balancing, TLS handshake debugging, tcpdump and MTU or latency analysis.
- Kubernetes — kubectl debugging, pod and node lifecycle, ingress and LoadBalancer services, PV, PVC and CSI binding, kubelet and
containerd log analysis; control-plane and etcd troubleshooting, certificate rotation and cluster upgrades (L3).
- Automation and process — Bash scripting (Python required at L3), REST APIs with curl and jq, Git, Ansible or Terraform, monitoring
and alert tuning with Prometheus and Grafana or an equivalent stack, ITIL incident and change management, formal RCA authoring
QUALIFICATIONS
- Education — Bachelor’s degree in Computer Science, IT, Electronics or an equivalent technical discipline.
- Experience — L2: 2–5 years supporting cloud or virtualization infrastructure in production. L3: 5–8 years, including major incident
leadership and demonstrable RCA ownership.
- Availability — willingness to work 247 rotational shifts including weekends, public holidays and on-call rotations.
Please share your Cv on (phone hidden)/and whatsapp also.
📌 Walk-in || Cloud Support Engineer/ (Noida)
🏢 Utho Cloud
📍 Noida