19 Aug
|
TMUS Global Solutions
|
Hyderabad
19 Aug
TMUS Global Solutions
Hyderabad
Job Overview:
At T-Mobile, we dont just build technology we empower people. We believe in investing in YOU your growth, your leadership, and your impact. Were unstoppable when driven individuals like you come together to solve bold challenges, inspire innovation, and create systems that power the future.
As a Principal Site Reliability Engineer, youll be part of a world-class engineering team building and scaling intelligent infrastructure to support LLM-based applications, AI gateways, and enterprise-scale platform services across finance, credit, collections, and document systemsincluding integrations with lending platforms such as OFSLL. Youll lead architecture and execution of AI observability frameworks, incident response systems, and operational excellence standards for high-performance, high-availability environments.
Your leadership will directly shape AI/ML infrastructure reliability across platforms and support rapid innovation, secure operations, and high-scale system availability while championing modern SRE best practices across platform engineering.
What Youll Do;
- Design, implement, and scale observability and incident response frameworks for AI infrastructure and mission-critical backend services.
- Architect telemetry, tracing, and logging pipelines across LLM-based applications, inference APIs, and platform services
- Define and enforce SLAs, alerting standards, and real-time dashboards for platform health, latency, throughput, and cost insights.
- Drive real-time monitoring and incident detection across hybrid cloud environments
- Lead root cause analysis and resolution of high-severity incidents tied to system availability, model performance, API latency, or scaling issues.
- Partner with AI Architecture,
Platform Engineering, and Security to enforce governance, compliance, and policy guardrails for production systems.
- Drive automation-first reliability practices, leveraging GitLab CI/CD pipelines, and deployment automation.
- Guide performance tuning, capacity planning, and cloud cost optimization strategies.
- Mentor Sr. SREs and engineers, and foster a culture of technical excellence, continuous learning, and system ownership.
- Support AIOps, model evaluation, and audit-readiness for enterprise-grade reliability standards.
What Youll Bring:
- Bachelors or Masters degree in Computer Science, Engineering, or a related field (preferred).
- 10+ years in software operations, infrastructure, or SRE roles in distributed systems environments.
- Strong experience in incident management, customer-facing platform operations, and performance engineering.
- Hands-on experience with AI/ML infrastructure (OpenAI APIs, AI Gateways, AI Agents).
- Expertise in Python, Java, along with scripting (e.g., Bash, PowerShell, YML).
- Strong experience in tune and scale infrastructure components such as Kafka, HAProxy, RMQ, Oracle DB, MongoDB, and distributed APIs.
- Strong knowledge of CI/CD, GitLab pipelines, software SDLC workflows, and deployment automation.
- Proven success managing observability pipelines using Splunk, Prometheus,
Grafana, OpenTelemetry or similar tools.
- Familiarity with cloud-native platforms and infrastructure (OCI, Azure, AWS or GCP).
Preferred Qualifications:
- Experience managing AI observability (token usage, golden set accuracy, latency scoring, throughput).
- Experience with Oracle Financial Services applications beyond OFSLL
- Deep expertise in Oracle Financial Services Lending & Leasing (OFSLL) architecture, configuration, integrations, and platform capabilities
- Knowledge of AIOps, model validation, policy enforcement, and secure API access controls.
- Background in enterprise AI/ML platforms, especially in financial services.
- Hands-on experience with Oracle DB, MongoDB, and building data monitoring pipelines.
- Familiarity with governance frameworks, compliance audits, and AI risk management best practices.
- Background in building secure, zero-downtime platforms with enterprise-scale SLAs.
Must Have Skills:
- Deep expertise in SRE methodologies, including SLOs, SLIs, alerting, incident response, and chaos engineering.
- Solid technical debugging skills across infrastructure, cloud, data, network and application layers.
- Advanced knowledge of Kubernetes, container orchestration, and large-scale deployment patterns.
- Proficiency in observability, automation, and site reliability engineering tooling.
- Proven ability to design and manage high-availability, scalable platforms with security and performance in mind.
- Effective communicator and collaborator across product, engineering, and operations teams.
- Comfortable leading initiatives and mentoring engineers in a fast-paced, evolving environment.
📌 Principal Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad