01 Aug
|
TMUS Global Solutions
|
Hyderabad
01 Aug
TMUS Global Solutions
Hyderabad
What Youll Do:
- Resolve escalated incidents across Kubernetes, API Proxy, WAF, DBs, and infra platforms.
- Design and improve runbooks, automating manual steps wherever possible.
- Lead and contribute to building self-healing systems and self-service tooling for users.
- Analyze incident trends, propose improvements in monitoring, capacity, and reliability.
- Collaborate with engineering teams on deployment, upgrades, and performance optimization.
- Conduct postmortems, document RCA, and ensure learning is captured.
- Mentor and coach Engineer(s)
What Youll Bring:
- 7+ years in SRE/DevOps/Systems Engineering as Senior or Principal Engineer
- Strong hands-on experience with Kubernetes, container orchestration, and API management.
- Working knowledge of WAFs, networking security, and database technologies (SQL/NoSQL).
- Proficient in automation and scripting (Python, Go, Ansible, Terraform, etc.).
- Strong observability/monitoring experience.
- Experience with CI/CD pipelines, GitOps, and infrastructure as code.
- Solid problem-solving and collaboration skills
Must Have Skills:
Advanced Incident Troubleshooting & Resolution:
- Expectation: Diagnose and resolve escalated incidents that Engineer(s) cannot handle, often across multiple layers (infrastructure, application, network).
- Example: For an API outage, identify if the root cause is in Kubernetes pod networking, API gateway misconfig, or backend DB latency and apply fixes.
Kubernetes & Container Orchestration Expertise:
- Expectation: Comfortable with deployments, scaling, networking, and debugging cluster-level issues.
- Example:
Troubleshoot why pods are pending by checking node capacity, taints/tolerations, and cluster autoscaler logs.
Automation & Scripting (Python, Go, Bash, Ansible, Terraform)
- Expectation: Write scripts and automation to reduce manual toil, enhance monitoring, and improve incident resolution speed.
- Example: Develop a Python script to automatically collect pod and system logs when a service crashes.
Observability & Monitoring Tooling:
- Expectation: Deep understanding of monitoring, alerting, tracing, and logging systems.
- Example: Build Prometheus alert rules to detect DB query spikes; configure Grafana dashboards for API latency.
CI/CD & Infrastructure as Code (IaC):
- Expectation: Familiarity with GitOps workflows, CI/CD pipelines, and infrastructure provisioning.
- Example: Enhance Jenkins pipeline to add automated smoke tests before promoting Kubernetes deployments.
Database Troubleshooting (SQL & NoSQL):
- Expectation: Identify performance bottlenecks, connection issues, and basic tuning opportunities.
- Example: Run queries to detect slow-running SQL statements causing latency in an application.
Incident Management & RCA:
- Expectation: Act as incident commander for escalated issues, lead bridge calls, and produce Root Cause Analyses.
- Example: After a WAF misconfiguration causes downtime, lead the investigation, document the timeline, and propose preventive actions.
Mentorship & Runbook Improvement:
- Expectation: Coach Engineer(s), refine runbooks, and introduce recent automated workflows.
- Example: Update a runbook to add automated Kubernetes log collection instead of manual steps.
Nice-to-Have:
Cloud Platform Engineering (AWS, Azure, GCP):
- Expectation: Hands-on skills in provisioning, scaling, and securing cloud workloads.
- Example: Diagnose why an AWS ALB is misrouting traffic after a deployment.
Security & WAF Management:
- Expectation: Understand WAF rules, common attacks (SQL injection, XSS), and how to apply fixes.
- Example: Investigate false positives in WAF logs and adjust rule sets with security teams.
Capacity & Performance Engineering:
- Expectation: Anticipate scaling needs, tune resource utilization, and propose optimizations.
- Example: Identify that a Kubernetes deployment is CPU-throttled and adjust HPA (Horizontal Pod Autoscaler) configs.
Automation Platform Integration (AIOps, ChatOps):
- Expectation: Integrate AI/ML-powered tools for anomaly detection and auto-remediation.
- Example: Implement a ChatOps bot that runs predefined Kubernetes troubleshooting commands in Slack.
Cross-Platform Expertise (Hybrid Infra):
- Expectation: Experience supporting both on-prem and cloud environments seamlessly.
- Example: Compare latency patterns between on-prem DBs and cloud-hosted APIs to identify bottlenecks.
📌 Sr Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad