Sr Engineer, Site Reliability (Hyderabad)

Sr Engineer, Site Reliability (Hyderabad)

30 Jul
|
TMUS Global Solutions
|
Hyderabad

30 Jul

TMUS Global Solutions

Hyderabad

About T-Mobile:
T-MobileUS, Inc. (NASDAQ: TMUS), headquartered in Bellevue, Washington, is Americas supercharged Un-carrier, connecting millions through its strong nationwide network and flagship brands,T-Mobileand Metro by T-Mobile. Customersbenefitfrom an unmatched combination of value, quality, and exceptional service experience.
TMUS Global Solutions:
TMUS Global Solutionsis a world-class technology powerhouse accelerating the companys global digital transformation. With a culture built on growth, inclusivity, and global collaboration, the teams here drive innovation at scale, powered by bold thinking.

Job Overview:

At T-Mobile, we dont just build technology we empower people. We believe in investing in YOU your growth, your impact, and your future. Were unstoppable when individuals like you come together to solve bold challenges, inspire innovation, and build platforms that serve millions.

As a Senior Site Reliability Engineer, youll join a world-class engineering team focused on building and scaling intelligent infrastructure for LLM-based applications, AI services, and enterprise-scale backend systems. Youll contribute to the design and implementation of observability, automation, and incident response strategies that ensure our platforms are high-performing, reliable, and cost-effective. Youll play a key role in driving operational excellence, supporting platform scalability, and collaborating across engineering and architecture teams. This role provides growth opportunities to influence large-scale architecture and AI/ML reliability.

What Youll Do:

- Implement and maintain observability, monitoring, and alerting systems for AI platforms and mission-critical backend services serving the T-Mobile Supply Chain domain.
- Collaborate in designing telemetry pipelines, logging infrastructure, and metrics dashboards using tools such as Splunk, Prometheus, Grafana, and OpenTelemetry.
- Contribute to the development of SLOs, SLIs, and real-time health indicators across platform services and APIs.




- Participate in on-call rotations and lead resolution of high-impact incidents, including root cause analysis and postmortem reporting.
- Work closely with platform engineering teams to enforce governance, compliance, and security standards in production environments.
- Improve deployment pipelines, CI/CD workflows, and infrastructure automation (e.g. GitLab).
- Tune and scale infrastructure components such as Kafka, HAProxy, RMQ, databases, and distributed APIs.
- Support capacity planning, cost analysis, and system tuning to optimize platform performance.
- Advocate for automation-first operations, reducing toil through scripting and system reliability tooling.
- Contribute to documentation, runbooks, and knowledge sharing across SRE and engineering teams.
- Mentor junior engineers and participate in the culture of technical rigor and continuous improvement.

What Youll Bring:

- Bachelors degree in Computer Science, Engineering, or a related field (Masters preferred).
- 7+ years of experience in SRE, DevOps, or operations engineering in cloud-based environments.

Engineering & Operations:

- CI/CD for integrations and infrastructure
- Automated testing (integration/regression)
- Performance & scalability engineering
- SRE fundamentals (SLIs, SLOs, RCA)
- Platform cost visibility & optimization
- Security-first design & compliance readiness
- Hands-on experience with monitoring, alerting, and incident response in distributed systems.
- Strong coding/scripting ability in Python, Java, or shell scripting languages like Bash or PowerShell.
- Proficiency in CI/CD pipelines, GitLab workflows
- Strong working knowledge of SQL and NoSQL databases, including Oracle DB and MongoDB.
- Working knowledge of AI/ML systems, APIs,



and modern LLM tooling is a strong plus.
- Familiarity with observability tools such as Splunk, Grafana, Prometheus.
- Experience with Kubernetes, container orchestration, and hybrid/multi-cloud deployments (Azure, AWS, GCP, OCI).
- Demonstrated ability to work in fast-paced, incident-driven environments with high stakes and uptime requirements.

Preferred Qualifications:

- Experience supporting AI workloads, model inference systems, or LLM-enabled platforms.
- Exposure to AIOpsor related ML platform observability and reliability practices.
- Experience in highly regulated finance industry with understanding of compliance and audit controls.
- Familiarity with platforms such as OpenAI or AI Gateway patterns.
- Background in building secure, zero-downtime platforms with enterprise-scale SLAs.

Must Have Skills:

- Strong grasp of SRE best practices including SLOs, SLIs, postmortems, and chaos engineering.
- Proficiency in diagnosing system bottlenecks across infrastructure, application, and network layers.
- Experience driving automation across observability, configuration, and deployment domains.
- Experience integrating with OFSLLor 3rd party finance systems
- Hands-on experience of OCI, OFSLL maintenance and support
- Robust communicator and collaborator in cross-functional technical teams.
- Curiosity-driven mindset with a passion for learning emerging AI technologies and systems reliability.

Why Join T-Mobile India:

At T-Mobile India, you wont just contribute to world-class technologyyoull help build it. Youll work with global leaders, solve complex system challenges, and build platforms that redefine how technology powers customer experience.

Were more than just a telecom companywere a technology powerhouse leading the way in AI, data, and digital innovation. And we do it all with heart, grit, and a passion for empowering people.

Join us and shape the future of intelligent platforms that serve millions at the scale and speed of T-Mobile.

📌 Sr Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr engineer, site reliability (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: sr engineer, site reliability (hyderabad) / hyderabad