Sr Engineer, Site Reliability (Hyderabad)

Sr Engineer, Site Reliability (Hyderabad)

02 Oct
|
TMUS Global Solutions
|
Hyderabad

02 Oct

TMUS Global Solutions

Hyderabad

What Youll Do:
- Resolve escalated incidents across Kubernetes, API Proxy, WAF, DBs, and infra platforms.
- Design and improve runbooks, automating manual steps wherever possible.
- Lead and contribute to building self-healing systems and self-service tooling for users.
- Analyze incident trends, propose improvements in monitoring, capacity, and reliability.
- Collaborate with engineering teams on deployment, upgrades, and performance optimization.
- Conduct postmortems, document RCA, and ensure learning is captured.
- Mentor and coach Engineer(s)

What Youll Bring:
- 7+ years in SRE/DevOps/Systems Engineering as Senior or Principal Engineer
- Robust hands-on experience with Kubernetes, container orchestration, and API management.
- Working knowledge of WAFs, networking security, and database technologies (SQL/NoSQL).
- Proficient in automation and scripting (Python, Go, Ansible, Terraform, etc.).
- Strong observability/monitoring experience.
- Experience with CI/CD pipelines, GitOps, and infrastructure as code.
- Solid problem-solving and collaboration skills





Must Have Skills:
Advanced Incident Troubleshooting & Resolution:
- Expectation: Diagnose and resolve escalated incidents that Engineer(s) cannot handle, often across multiple layers (infrastructure, application, network).
- Example: For an API outage, identify if the root cause is in Kubernetes pod networking, API gateway misconfig, or backend DB latency and apply fixes.

Kubernetes & Container Orchestration Expertise:
- Expectation: Comfortable with deployments, scaling, networking, and debugging cluster-level issues.
- Example: Troubleshoot why pods are pending by checking node capacity, taints/tolerations, and cluster autoscaler logs.

Automation & Scripting (Python, Go, Bash, Ansible, Terraform)
- Expectation: Write scripts and automation to reduce manual toil, enhance monitoring, and improve incident resolution speed.
- Example: Develop a Python script to automatically collect pod and system logs when a

📌 Sr Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr engineer, site reliability (hyderabad) / hyderabad