23 Aug
|
TMUS Global Solutions
|
Hyderabad
23 Aug
TMUS Global Solutions
Hyderabad
About the Team:
The MPM (Enterprise Catalog) Platform team powers T-Mobiles enterprise product catalog ecosystem, enabling product modeling, pricing, promotions, orchestration, APIs, and omnichannel experiences across digital and assisted channels.
The platform serves as a foundational capability supporting enterprise commerce, product lifecycle management, integrations, and downstream fulfillment systems. Our engineering teams focus on scalability, resiliency, observability, automation, and operational excellence while enabling rapid innovation across the enterprise.
About the Role:
- As a Site Reliability Engineer on the MPM Platform team, you will contribute to enhancing system reliability, resilience, and operational efficiency for enterprise catalog platforms and APIs supporting T-Mobiles commerce ecosystem.
- You will work hands-on to automate operational processes, improve observability, reduce manual effort, prevent operational incidents, and support the stability of critical enterprise services powering customer and business experiences across channels.
- This role is ideal for engineers with a strong operational mindset who are eager to grow their expertise in reliability engineering, cloud-native platforms, distributed systems, and modern DevOps practices while working alongside experienced platform and engineering teams.
- We pride ourselves on fostering a culture of innovation, collaboration, transparency, and continuous improvement. Join us in helping modernize and strengthen the enterprise platforms that power T-Mobiles digital transformation journey.
What Youll Do:
- Support high-availability enterprise catalog platforms and APIs with strong uptime and reliability targets.
- Monitor platform health, validate integrations, troubleshoot API and service failures, and support incident management activities.
- Automate operational workflows using Java, Python, Bash, or similar technologies to improve efficiency and reduce manual effort.
- Participate in root cause analysis (RCA), problem management, and continuous improvement initiatives to prevent recurring incidents.
- Support observability and monitoring using tools such as Splunk, Grafana, AppDynamics, Dynatrace, and AWS CloudWatch.
- Assist with CI/CD pipeline maintenance and deployment automation to improve software delivery reliability and speed.
- Contribute to telemetry improvements, dashboard creation, alert tuning, and operational metrics reporting.
- Collaborate with engineering, platform, and architecture teams to troubleshoot customer-impacting issues and improve platform resiliency.
- Support production readiness activities, release validations, and operational governance processes.
- Create and maintain operational runbooks, troubleshooting guides, and technical documentation.
- Participate in on-call support rotations and operational response activities.
What Youll Bring:
- Bachelors degree in Computer Science, Engineering, or related technical field.
- 35 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or enterprise application support.
- Experience supporting enterprise APIs, distributed systems, and cloud-native applications.
- Strong development and scripting skills using Java, Python, Bash, or similar technologies.
- Understanding of system reliability, resiliency, observability,
and incident management principles.
- Experience with monitoring and observability platforms such as Splunk, Grafana, AppDynamics, Dynatrace, or CloudWatch.
- Familiarity with CI/CD pipelines and deployment automation practices.
- Basic understanding of cloud-native platforms and containerized environments.
- Solid troubleshooting, analytical, and problem-solving skills.
- Effective communication and collaboration skills across engineering and product teams.
- Ability to adapt and thrive in fast-paced Agile environments.
- Hands-on experience with SQL (PostgreSQL, MySQL) and NoSQL (MongoDB or similar) databases for backend systems, troubleshooting, and performance optimization
Must-Have Skills:
- Java and/or Python development experience
- Bash scripting for operational automation
- Monitoring and observability platforms (Splunk, Grafana, AppDynamics, Dynatrace)
- Incident management and root cause analysis
- API troubleshooting and operational support
- Linux system administration fundamentals
- CI/CD pipeline familiarity
- Cloud platform fundamentals (AWS preferred)
Nice-to-Have:
- Spring Boot and microservices architecture
- Kubernetes and container orchestration
- Kafka or event-driven systems
- Infrastructure as Code (Terraform, Helm)
- TM Forum/Open API standards
- Experience with enterprise catalog or commerce ecosystems
- Familiarity with Agile tools such as JIRA and Confluence
- AWS or Kubernetes certifications
Why Join Us:
At T-Mobile, we move fast, solve complex enterprise challenges, and continuously innovate to deliver exceptional customer experiences. As part of the MPM Platform team, you will help build and support mission-critical enterprise platforms that enable the future of digital commerce and enterprise transformation at scale.
📌 Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad