04 Aug
|
TMUS Global Solutions
|
Hyderabad
04 Aug
TMUS Global Solutions
Hyderabad
ABOUT THE ROLE :
The Manager, Site Reliability Engineering PMT HR Systems leads the 14-person India SRE team responsible for the operational reliability and continuous improvement of T-Mobile's HR application portfolio. This portfolio spans 100+ applications supporting T-Mobile's workforce - including Workday, UKG Pro, Equifax, Fidelity, Fieldglass, and the broader perks, compensation, learning, and HR service platforms.
This is a player-managing role at the front line of T-Mobile's HR Systems operational team in Hyderabad. You will own day-to-day operational health for the portfolio, develop and coach a team of SREs in a greenfield Hyderabad hub, and partner with US-based Principal SREs and the Phase 2 Lead on architecture, automation, and the long-term migration from SaaS to custom-built solutions. You are not a delivery-only manager - you bring engineering judgment to incident reviews, automation priorities, and capability decisions.
WHAT YOU'LL DO:
- Lead, coach, and grow an 10-person India team (10 SRE + 1 Domain Architect) supporting T-Mobile's HR application portfolio (Workday, UKG, Learning, and adjacent HR platforms).
- Own operational metrics for the portfolio: SLA compliance, incident MTTR, change success rate, runbook coverage, automation adoption.
- Run the follow-the-sun on-call rotation for the HR application portfolio (IST shifts); coordinate handoffs with US-based SREs and Software Developers (DevOps).
- Partner with US-based Domain Architects on architecture decisions and the SaaS-to-custom migration roadmap; partner with US Software Developers (DevOps) on escalated incidents and platform tooling.
- Build the team's culture and engineering bar - recruit, ramp, mentor, and performance-manage SRE engineers in a greenfield Hyderabad office.
- Drive automation and toil reduction initiatives across the portfolio; set quarterly automation goals and hold the team accountable to them.
- Represent the team in cross-functional forums: portfolio reviews with HR Systems leadership (James Johnson portfolio), incident reviews, vendor governance, and global SRE leadership.
- Contribute hands-on to high-priority incidents, architecture reviews, and automation work where appropriate (player-coach mode).
- Establish and maintain runbook discipline, monitoring coverage, and post-incident learning processes.
- Leverage AI tools including Claude to accelerate team workflows, troubleshooting, and documentation.
WHAT YOU'LL BRING:
- 8+ years of experience in SRE, software engineering, or technical operations, with at least 3 years in a people-management capacity.
- Demonstrated experience building and leading SRE or operations teams - ideally in a greenfield or rapidly scaling environment.
- Solid technical foundation: you have written production code, debugged complex multi-system integration failures, and understand modern SRE practices (SLOs, error budgets,
blameless postmortems).
- Hands-on experience with at least one enterprise HR or SaaS platform (Workday, ServiceNow, Salesforce, or similar).
- Track record of driving automation and toil reduction with measurable results.
- Excellent written and verbal communication skills; comfort partnering across US/India time zones.
- Experience hiring, coaching, and managing performance in a multi-vendor or insourcing context is a strong plus.
- Demonstrated experience using AI tools in your own engineering work and driving AI adoption across a technical team - including setting standards for agent-based workflows, evaluating team-built agents, and measuring AI's impact on team productivity and toil reduction.
REQUIRED TECHNICAL SKILLS:
- AI & agent tooling: Working proficiency in AI coding/automation assistants (Claude, Copilot, or equivalent); able to evaluate team-built agents, set quality and governance standards, and drive adoption of agent-based workflows across the team. Expects every engineer on the team to use AI tooling as a default productivity multiplier.
- Working proficiency in at least one scripting/programming language (Python, PowerShell, Bash, or similar) - sufficient to read team code, contribute fixes, and judge quality.
- Familiarity with REST APIs, JSON/XML, and integration debugging patterns.
- Working knowledge of at least one monitoring/observability platform (Splunk, Datadog, AppDynamics, or similar).
- SQL proficiency for data validation and operational reporting.
- Familiarity with ticketing
📌 Manager, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad