03 Oct
|
TMUS Global Solutions
|
India
03 Oct
TMUS Global Solutions
India
What Youll Do:
Lead resolution of high-severity/complex incidents across hybrid infrastructure.
Architect and implement automation frameworks, self-healing workflows, and AI-driven ops.
Define SRE best practices, reliability SLIs/SLOs/SLAs, and operational standards.
Partner with application and platform engineering teams to improve resilience.
Drive observability maturity: predictive monitoring, anomaly detection, automated RCA.
Own continuous improvement of Engineer(s)/Sr Engineer(s) runbooks and automation pipelines.
Provide technical leadership, mentor junior SREs, and conduct training.
Identify current technologies, tools, and processes that elevate operational excellence.
What Youll Bring:
10+ years in SRE/DevOps/Systems/Platform Engineering as Principal or Staff engineer.
Deep expertise in Kubernetes, distributed systems, and multi-cloud infrastructure.
Robust knowledge of security, WAFs, and networking at scale.
Advanced automation and programming (Python, Go, Terraform, Ansible).
Experience applying AI/ML to operations (AIOps platforms,
anomaly detection, predictive scaling).
Robust incident command and leadership skills during outages.
Proven track record of driving automation-first operations transformations.
Must Have Skills:
Incident Command & Complex Troubleshooting:
Expectation: Take leadership during high-severity outages, orchestrating technical response across teams.
Example: Lead a Sev-1 bridge call where multiple microservices are failing due to cascading Kubernetes issues; coordinate DB, infra, network, security and app teams to isolate the problem.
Deep Kubernetes & Distributed Systems Expertise:
Expectation: Design, troubleshoot, and optimize complex Kubernetes clusters and multi-region deployments
Example: Diagnose why inter-cluster communication in a service mesh is causing intermittent API failures and propose architectural fixes.
Automation Framework Design (Infra & Ops):
Expectation: Architect autom
📌 Principal Engineer, Site Reliability Hyderabad (India)
🏢 TMUS Global Solutions
📍 India