Sr Engineer, Site Reliability (Hyderabad)

Sr Engineer, Site Reliability (Hyderabad)

30 Jul
|
TMUS Global Solutions
|
Hyderabad

30 Jul

TMUS Global Solutions

Hyderabad

ABOUT THE ROLE:

The Senior Site Reliability Engineer owns the availability, performance, and scalability of production systems at TMUS Global Solutions. You will drive the engineering discipline that bridges software delivery and operations eliminating toil through automation, raising the reliability bar through rigorous SLO/SLI frameworks, and embedding security and observability into every layer of the stack. You will be a force multiplier for the team: setting technical direction, leading incident response, and mentoring engineers to build a culture of ownership and continuous improvement.

WHAT YOU'LL DELIVER:

- Own end-to-end CI/CD pipeline design and reliability from commit to production using GitLab CI/CD and GitHub Actions, cutting deployment lead times and failure rates.
- Define and enforce SLOs, SLIs, and error budgets; translate reliability targets into engineering decisions that directly impact customer experience.
- Drive incident command for P1/P2 events: lead root-cause analysis, author blameless postmortems, and ship permanent fixes that move the needle on MTTR.
- Build production-grade observability distributed tracing, metrics pipelines, and structured logging using Open Telemetry, Prometheus/Grafana, Splunk, and AppDynamics.
- Architect and maintain infrastructure as code (Terraform, Ansible) for cloud-native environments on AWS RDS, DynamoDBs, OpenSearch, Redshift etc with reproducibility and drift detection built in.
- Harden the platform end-to-end: IAM least-privilege, network segmentation, zero-touch secrets management (Vault / AWS Secrets Manager), and container image security scanning.




- Champion GitOps and progressive delivery patterns (GitlabCD, feature flags, canary / blue-green deployments) to reduce blast radius and accelerate safe releases.
- Reduce operational toil by ?20% per quarter through targeted automation, self-healing runbooks, and platform abstractions that empower development teams.
- Mentor junior and mid-level SREs; establish engineering standards, runbooks, and on-call practices that scale with the team.

WHAT YOU'LL BRING:

- 6+ years of progressive experience in SRE, DevOps, or cloud infrastructure engineering, with demonstrated ownership of production systems at scale.
- Deep Kubernetes expertise: helm charts, workload tuning, RBAC, network policies, and service mesh (Istio / Linked).
- Hands-on Terraform and Ansible at production scale modules, state management, drift remediation, and CI-integrated plan/apply workflows.
- Proven observability engineering: you have built telemetry pipelines, authored meaningful dashboards, and used data to justify reliability investments. Experience with Open Telemetry, Prometheus, Grafana, Splunk, and SignalFX.
- Demonstrated cloud security experience across multiple layers network (VPC, security groups, WAF, firewall), identity (IAM, OIDC, SAML), OS hardening,



and application-level controls (certificates, API gateway, mTLS).
- Strong scripting and automation skills (Python, Bash, Go) to eliminate toil and build internal tooling.
- Experience with GitOps workflows and progressive delivery tools (GitlabCD); comfort with canary, blue-green, and feature-flag release strategies.
- Track record of leading incident response at P1/P2 severity on-call experience, structured postmortems, and measurable MTTR improvements.
- Solid communication skills: able to translate reliability risks into business impact and influence engineering decisions across teams.
- Bachelor's degree in computer science, Software Engineering, or related field; master's preferred. AWS certifications (Solutions Architect, DevOps Engineer) a plus

MUST-HAVE SKILLS:

- Cloud Platform:AWS (EKS, ECS, Lambda, RDS, CloudWatch, IAM)
- Container Orchestration:Kubernetes production-grade cluster operations, RBAC, network policies
- IaC & Configuration:Terraform (required), Ansible; GitOps-first approach
- CI/CD & SCM:GitLab CI/CD, GitHub Actions; pipeline-as-code, reusable templates
- Observability:Open Telemetry, Prometheus, Grafana, Splunk, SignalFX tracing, metrics, logs
- Security:Multi-layer: IAM, VPC/security groups, WAF, certificates, secrets management (Vault / AWS SM), container scanning
- Reliability Practice:SLO/SLI definition, error budget management, blameless postmortems, chaos/fault injection
- Progressive Delivery:ArgoCD / Flux, canary & blue-green deployments, feature flags
- Scripting:Python, Bash, Go for automation and tooling

📌 Sr Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr engineer, site reliability (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: sr engineer, site reliability (hyderabad) / hyderabad