18 Aug
|
TMUS Global Solutions
|
Hyderabad
18 Aug
TMUS Global Solutions
Hyderabad
Role Overview
At T-Mobile, we don't just build technology we empower people. We believe in investing in YOU your growth, your leadership, and your long-term impact. We're unstoppable when driven individuals come together to solve bold challenges, inspire innovation, and create platforms that power the future.
This role ensures the reliability and resilience of digital infrastructure, enabling efficient software development and deployment. It focuses on automating processes and reducing manual effort to prevent operational incidents, improve system performance, and enhance overall operational efficiency.
As part of the Enterprise Procurement Engineering organization, this role supports a large-scale enterprise device management, procurement, and application delivery ecosystem that enables critical operations.
The Sr Site Reliability Engineer leverages automation, CI/CD practices, scripting, observability, and incident management expertise to improve reliability, scalability, and operational efficiency across a complex, cloud-native technology environment. This work directly impacts organizational stability and customer experience by ensuring the availability, performance, and reliability of critical systems.
Key Responsibilities
Automate processes to accelerate software development and deployment while minimizing manual interventions including CI/CD pipelines on Azure DevOps or GitLab, deployment automation, and endpoint management workflows.
Design, build, and enhance automation solutions on Azure that improve operational efficiency, deployment consistency, and service reliability across complex enterprise environments.
Manage and scale Kubernetes-based workloads; own cluster reliability, resource optimization, deployment strategies, and container orchestration practices.
Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime and improve operational stability.
Conduct root cause analysis and collaborate with problem management teams to prevent incident recurrence and continuously improve system operations.
Leverage Python, Bash, and related scripting/programming tools to improve system robustness, deployment quality, and operational efficiency.
Implement and maintain observability solutions metrics, dashboards, alerting, and distributed tracing using Splunk, Grafana, Prometheus, or equivalent platforms to drive proactive incident detection.
Implement and support secure automation practices, including secrets management, credential lifecycle automation, and integration with approved enterprise security platforms (e.g., CyberArk, HashiCorp Vault).
Partner with product, engineering, cybersecurity, and operations teams to design and implement scalable deployment and automation solutions.
Support modernization initiatives involving mobile device management platforms, application deployment automation, platform migrations, and operational process improvements.
Continuously learn new skills and technologies to adapt to changing environments and drive innovation.
Perform other duties and projects as assigned.
Qualifications
Education & Experience
Bachelor's degree in Computer Science, Engineering, or a related technical field with 5+ years of relevant experience; or an advanced degree with 1+ year; or an equivalent combination of education and experience.
69 years of experience in Site Reliability Engineering (SRE), DevOps, platform engineering, or software development environments.
69 years of experience troubleshooting production and customer-related issues while supporting business-critical systems.
69 years of hands-on experience developing automation and software solutions using Python, Bash, or similar scripting/programming languages.
Hard Requirements (All Required)
Azure: Proven, hands-on experience designing and operating cloud infrastructure on Microsoft Azure including Azure DevOps, AKS, Azure Monitor, and core IaaS/PaaS services.
Kubernetes: Experience managing production-grade Kubernetes clusters workload orchestration, resource management, scaling, health, and troubleshooting.
SRE / DevOps:
Strong command of SRE principles SLOs, SLAs, error budgets, on-call, incident response, blameless post-mortems, root cause analysis, and continuous service improvement.
CI/CD: Experience designing, implementing, and maintaining CI/CD pipelines using GitLab, Azure DevOps, or comparable platforms.
Scripting: Proficiency in Python and/or Bash for automation, tooling, and operational workflows.
Observability (Interchangeable Platforms)
Candidates must have meaningful experience with observability tools and practices. Proficiency with any combination of the following is accepted:
Splunk, Grafana, Prometheus, Datadog, Dynatrace, Current Relic, Azure Monitor, AWS CloudWatch, or equivalent platforms.
Building dashboards, setting up alerting rules, and configuring telemetry pipelines.
Distributed tracing, log aggregation, and metrics-based SLO monitoring.
Additional Qualifications
Experience with Infrastructure as Code (Terraform, Ansible, or equivalent) for configuration management and infrastructure provisioning.
Experience with cloud platforms beyond Azure AWS and/or GCP is a plus.
Experience integrating enterprise secrets management solutions such as CyberArk or HashiCorp Vault.
Experience applying cybersecurity best practices and secure engineering principles across software delivery.
Strong analytical and problem-solving skills with the ability to diagnose and resolve complex operational and production issues.
Ability to collaborate effectively with product, engineering, cybersecurity, operations, and support teams.
Excellent verbal and written communication skills, with the ability to explain technical concepts to both technical and non-technical audiences.
Demonstrated ability to balance operational support, project delivery, modernization, and continuous improvement.
Experience mentoring peers, sharing technical knowledge, and adapting quickly to evolving technologies and business requirements.
Licenses & Certifications
Preferred (not required):
Microsoft Certified: Azure DevOps Engineer Expert
Microsoft Certified: Azure Solutions Architect Expert
Certified Kubernetes Administrator (CKA)
AWS Certified DevOps Engineer
Google Cloud Certified Qualified DevOps Engineer
📌 Sr Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad