16 Aug
|
TMUS Global Solutions
|
Hyderabad
16 Aug
TMUS Global Solutions
Hyderabad
Job Description:
The System Reliability Engineer (SRE) improves and protects the software and systems behind all of T-Mobile's IT services, including management of scalability, availability, latency, performance, security, and capacity, and delivering of software faster, better, and cheaper. From designing & maintaining CICD Pipelines to building the next generation of T-Mobile applications on cloud native platforms, the SRE's enable great customer experience and product innovation by continuous improvement of operational support.
This role supports the business operations related to payment processing, transaction authorization, and payment gateway integrations, ensuring seamless and secure payment flows for customers completing purchases and financial transactions. The position involves managing backend services and APIs to facilitate these processes efficiently.
In addition, the role leverages AI/ML-driven monitoring, anomaly detection, and predictive analytics to proactively identify issues before they impact customers, optimize system performance, and support self-healing capabilities. SREs implement intelligent automation across CI/CD pipelines, infrastructure management, and incident response, reducing manual effort and accelerating resolution times. They also explore ML-based automation frameworks to continuously enhance resiliency, scalability, and efficiency of T-Mobiles IT ecosystem.
we are a team that encourages innovation and advocates an agile and open approach, truly working and playing in the Un-carrier way!
Job Responsibilities:
- Utilizes fluent knowledge and skill in emerging DevOps-centric automation tools and technologies for CICD, configuration management, etc.
for non-prod environments.
- Performs environment management, automated server provisioning, pipeline configuration (VMs).
- Delivers software to improve the availability, scalability, latency, and efficiency of T-Mobiles services.
- Creates, manages, and uses dashboard for continuous monitoring and health check of applications, and the underlying infrastructure, improve the quality of services using the monitoring feedback for non-production workplace.
- Contributes in future improvement of software delivery processes and operations, e.g., cloud enablement, use of microservices with containerization.
- Relationship and People Management: Mentors/guides other Systems Reliability Engineers and vendor resources as needed.
- Applies AI/ML-driven monitoring and anomaly detection to proactively identify performance bottlenecks, predict outages, and optimize system health.
- Automates operational workflows using machine learning models and intelligent runbooks to reduce incident response time and manual intervention.
- Develops predictive analytics solutions for capacity planning, scaling, and fault-tolerant systems.
Education and Work Experience:
- Bachelor's Degree [Required]
- Master's/Advanced Degree [Preferred]
- 4-7 years Relevant experience
- Experience working in an Agile and DevOps environment [Required]
- Experience in one or more of: Java, Python, Go, or scripting experience in Shell [Required]
- Experience in Continuous Integration/Continuous Delivery tools, such as Jenkins, Cloudbees, Gitlab etc., and other automation tools [Required]
- Experience in working in a team that manages API/backend services [Required]
- Experience with DevOps tools, such as Ansible, Chef, Puppet, etc. Experience in Docker, Kubernetes, etc. [Required]
- Experience in APM tool, like AppDynamics, logging tool, like Splunk [Required]
- Experience working in a cloud environment (public/private) [Required]
- Experience in migrating to cloud or cloud native environments [Preferred]
- Experience with Cassandra database management, including cluster operations, data modeling, and performance tuning [Preferred]
- Experience with AI/ML frameworks for automation and system monitoring [Preferred]
- Hands-on experience with predictive modeling, anomaly detection, and ML-driven automation pipelines [Preferred]
Knowledge, Skills and Abilities:
- Strong analytical and communication skills to drive adoption of performance improvements across teams [Required]
- Deep understanding of JVM internals, memory management, garbage collection, threading, and profiling tools (JProfiler, VisualVM, Flight Recorder) [Preferred]
- Ability to design and execute realistic load/capacity models and synthetic workloads that simulate production traffic [Preferred]
- AI/ML Model Development and Deployment [Preferred]
- Predictive Analytics for Reliability Engineering [Preferred]
📌 Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad