31 Jul
|
TMUS Global Solutions
|
Hyderabad
31 Jul
TMUS Global Solutions
Hyderabad
About T-Mobile:
T-Mobile US, Inc. (NASDAQ: TMUS) is America’s supercharged Un-carrier, powered by an award-winning 5G network that connects more people in more places than ever before. With its unique value proposition of best network, best value, and best experiences, T-Mobile is redefining connectivity, fueling competition, and driving the next wave of innovation in wireless and beyond. Headquartered in Bellevue, Washington, T-Mobile provides services through its subsidiaries and operates its flagship brands, T-Mobile, Metro by T-Mobile, and Mint Mobile.
TMUS Global Solutions:
TMUS Global Solutions is a world-class technology organization accelerating T-Mobile’s global digital transformation. Our teams combine engineering talent, technology expertise, and collaborative ways of working to build secure, scalable solutions that improve customer and employee experiences. We foster innovation, agility, transparency, and strong enterprise partnerships to deliver measurable business outcomes. About the Role:
This role is responsible for developing, operating, and continuously improving production software systems while ensuring reliability, scalability, security, and operational excellence across critical enterprise platforms.
The engineer contributes throughout the software development lifecycle, from application development and deployment automation to monitoring, observability, incident response, and continuous improvement. Working closely with platform, infrastructure, security, and product engineering teams, this role emphasizes ownership of software in production, automation of operational processes, and technical leadership for the reliability, operability, and lifecycle management of production platforms.
Success in this role requires solid software engineering fundamentals combined with expertise in reliability engineering, cloud-native operations, automation, and observability practices. We pride ourselves on encouraging a culture of innovation, agile ways of working, and transparency in all we do. Join us in embodying the spirit of the Un-carrier and making a tangible impact. What You'll Do:
Design, develop, and maintain backend services, APIs, automation tools, and operational platforms.
Build and maintain CI/CD pipelines supporting rapid and reliable software delivery.
Develop infrastructure automation solutions using Infrastructure-as-Code and cloud-native tooling.
Implement monitoring, observability, logging, alerting, and operational dashboards to improve platform visibility.
Partner with engineering teams to improve service reliability, scalability, and operational readiness.
Participate in incident management, root cause analysis, and reliability improvement initiatives.
Design and implement automated testing, deployment, and recovery mechanisms.
Support cloud-native platforms and containerized workloads running in Kubernetes environments.
Collaborate across engineering organizations to establish operational best practices and engineering standards.
Create and maintain operational documentation, runbooks, and knowledge-sharing materials.
Support deployment, monitoring, observability, security, and operational reliability of MCP-enabled services and AI integration platforms. What You'll Bring:
7+ years of software engineering experience building and operating production systems.
Bachelor’s degree in Computer Science, Software Engineering, Information Systems, or a related field, or equivalent practical experience.
Strong proficiency in Python, Java, or similar programming languages.
Experience developing APIs, services, automation tools, or cloud-native applications.
Experience with CI/CD platforms, source control workflows, and automated testing practices.
Experience with monitoring, logging, observability platforms, and operational tooling.
Experience deploying and operating applications in public cloud environments.
Demonstrated experience leading reliability initiatives, operational improvements, and cross-functional engineering efforts.
Strong troubleshooting, debugging, and incident response skills.
Must Have
Skills
Python, Java, or similar programming languages
Backend services, APIs, and automation engineering
CI/CD, source control, and automated testing
Public cloud application deployment and operations
Observability, incident response, and reliability engineering Nice-to-Have:
Experience with Kubernetes, containers, and cloud-native infrastructure.
Experience with Infrastructure-as-Code platforms such as Terraform or similar technologies.
Familiarity with reliability engineering principles, including SLIs, SLOs, and error budgets.
Experience supporting AI-enabled services, model-serving platforms, or LLM-based applications.
Experience building internal operational platforms and developer productivity solutions.
Experience supporting AI platform services, Model Context Protocol (MCP) integrations, API gateways, or agent-based application architectures.
📌 Sr Engineer Software - Platform & Reliability Engineering-28034] (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad