Principal Engineer, Site Reliability (Hyderabad)

Principal Engineer, Site Reliability (Hyderabad)

23 Aug
|
TMUS Global Solutions
|
Hyderabad

23 Aug

TMUS Global Solutions

Hyderabad

What Youll Do:

- Lead resolution of high-severity/complex incidents across hybrid infrastructure.
- Architect and implement automation frameworks, self-healing workflows, and AI-driven ops.
- Define SRE best practices, reliability SLIs/SLOs/SLAs, and operational standards.
- Partner with application and platform engineering teams to improve resilience.
- Drive observability maturity: predictive monitoring, anomaly detection, automated RCA.
- Own continuous improvement of Engineer(s)/Sr Engineer(s) runbooks and automation pipelines.
- Provide technical leadership, mentor junior SREs, and conduct training.
- Identify new technologies, tools, and processes that elevate operational excellence.

What Youll Bring:

- 10+ years in SRE/DevOps/Systems/Platform Engineering as Principal or Staff engineer.
- Deep expertise in Kubernetes, distributed systems, and multi-cloud infrastructure.
- Strong knowledge of security, WAFs, and networking at scale.
- Advanced automation and programming (Python, Go, Terraform, Ansible).
- Experience applying AI/ML to operations (AIOps platforms, anomaly detection, predictive scaling).
- Solid incident command and leadership skills during outages.
- Proven track record of driving automation-first operations transformations.

Must Have Skills:

Incident Command & Complex Troubleshooting:

- Expectation: Take leadership during high-severity outages, orchestrating technical response across teams.
- Example:



Lead a Sev-1 bridge call where multiple microservices are failing due to cascading Kubernetes issues; coordinate DB, infra, network, security and app teams to isolate the problem.

Deep Kubernetes & Distributed Systems Expertise:

- Expectation: Design, troubleshoot, and optimize complex Kubernetes clusters and multi-region deployments
- Example: Diagnose why inter-cluster communication in a service mesh is causing intermittent API failures and propose architectural fixes.

Automation Framework Design (Infra & Ops):

- Expectation: Architect automation platforms to reduce manual toil, enable self-service, and support auto-remediation.
- Example: Build an Ansible/Terraform-based automation pipeline that provisions, configures, and tests new app environments with zero manual steps.

Observability Strategy & Advanced Monitoring:

- Expectation: Define enterprise-wide observability standards (SLIs/SLOs/SLAs), implement anomaly detection, and predictive monitoring.
- Example: Roll out a metrics-based SLO framework for all API services with automated burn-rate alerts in Prometheus.

Database & Application Performance Engineering:

- Expectation: Tune databases, caching layers, and app performance to handle scale.
- Example: Identify DB query patterns that degrade API performance and recommend schema/index optimizations.

Cross-Domain SME Knowledge (Networking, Storage, APIs):

- Expectation: Act as a go-to expert across infrastructure layers.

📌 Principal Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal engineer, site reliability (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: principal engineer, site reliability (hyderabad) / hyderabad