31 Jul
|
TMUS Global Solutions
|
Hyderabad
31 Jul
TMUS Global Solutions
Hyderabad
About TMUS Global Solutions
T-Mobile is America’s supercharged Un-carrier, challenging conventions and setting new standards in wireless. With the nation’s largest and fastest 5G network, T-Mobile delivers advanced connectivity and unmatched value to millions across the U.S. We’re unwaveringly obsessed with providing the best possible service experience, driven by a spirit of disruption that fuels competition and innovation in wireless and beyond.
Job Description
What You'll Do:
- Design, develop, and maintain backend services, APIs, automation tools, and operational platforms.
- Build and maintain CI/CD pipelines supporting rapid and reliable software delivery.
- Develop infrastructure automation solutions using Infrastructure-as-Code and cloud-native tooling.
- Implement monitoring, observability, logging, alerting, and operational dashboards to improve platform visibility.
- Partner with engineering teams to improve service reliability, scalability, and operational readiness.
- Participate in incident management, root cause analysis, and reliability improvement initiatives.
- Design and implement automated testing, deployment, and recovery mechanisms.
- Support cloud-native platforms and containerized workloads running in Kubernetes environments.
- Collaborate across engineering organizations to establish operational best practices and engineering standards.
- Create and maintain operational documentation, runbooks, and knowledge-sharing materials.
- Support deployment, monitoring, observability, security, and operational reliability of MCP-enabled services and AI integration platforms.
What You'll Bring:
- 7+ years of software engineering experience building and operating production systems.
- Bachelors degree in Computer Science, Software Engineering, Information Systems, or a related field, or equivalent practical experience.
- Robust proficiency in Python, Java, or similar programming languages.
- Experience developing APIs, services, automation tools, or cloud-native applications.
- Experience with CI/CD platforms, source control workflows, and automated testing practices.
- Experience with monitoring, logging, observability platforms, and operational tooling.
- Experience deploying and operating applications in public cloud environments.
- Demonstrated experience leading reliability initiatives, operational improvements, and cross-functional engineering efforts.
- Strong troubleshooting, debugging, and incident response skills.
Must Have Skills:
- Python, Java, or similar programming languages
- Backend services, APIs, and automation engineering
- CI/CD, source control, and automated testing
- Public cloud application deployment and operations
- Observability, incident response, and reliability engineering
Nice-to-Have:
- Experience with Kubernetes, containers, and cloud-native infrastructure.
- Experience with Infrastructure-as-Code platforms such as Terraform or similar technologies.
- Familiarity with reliability engineering principles, including SLIs, SLOs, and error budgets.
- Experience supporting AI-enabled services, model-serving platforms, or LLM-based applications.
- Experience building internal operational platforms and developer productivity solutions.
- Experience supporting AI platform services, Model Context Protocol (MCP) integrations, API gateways, or agent-based application architectures.
📌 Sr Engineer ,Software (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad