02 Oct
|
TMUS Global Solutions
|
Hyderabad
02 Oct
TMUS Global Solutions
Hyderabad
What Youll Do:
- Resolve escalated incidents across Kubernetes, API Proxy, WAF, DBs, and infra platforms.
- Design and improve runbooks, automating manual steps wherever possible.
- Lead and contribute to building self-healing systems and self-service tooling for users.
- Analyze incident trends, propose improvements in monitoring, capacity, and reliability.
- Collaborate with engineering teams on deployment, upgrades, and performance optimization.
- Conduct postmortems, document RCA, and ensure learning is captured.
- Mentor and coach Engineer(s)
What Youll Bring:
- 7+ years in SRE/DevOps/Systems Engineering as Senior or Principal Engineer
- Robust hands-on experience with Kubernetes, container orchestration, and API management.
- Working knowledge of WAFs, networking security, and database technologies (SQL/NoSQL).
- Proficient in automation and scripting (Python, Go, Ansible, Terraform, etc.).
- Strong observability/monitoring experience.
- Experience with CI/CD pipelines, GitOps, and infrastructure as code.
- Solid problem-solving and collaboration skills
Must Have Skills:
Advanced Incident Troubleshooting & Resolution:
- Expectation: Diagnose and resolve escalated incidents that Engineer(s) cannot handle, often across multiple layers (infrastructure, application, network).
- Example: For an API outage, identify if the root cause is in Kubernetes pod networking, API gateway misconfig, or backend DB latency and apply fixes.
Kubernetes & Container Orchestration Expertise:
- Expectation: Comfortable with deployments, scaling, networking, and debugging cluster-level issues.
- Example: Troubleshoot why pods are pending by checking node capacity, taints/tolerations, and cluster autoscaler logs.
Automation & Scripting (Python, Go, Bash, Ansible, Terraform)
- Expectation: Write scripts and automation to reduce manual toil, enhance monitoring, and improve incident resolution speed.
- Example: Develop a Python script to automatically collect pod and system logs when a
📌 Sr Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad