Senior Site Reliability Engineer – L3
Senior SRE Engineer
? Location: Hyderabad, India
? Work Type: Full time | On-site / Hybrid
? Experience: 5–8 Years
-
Role Overview
Own infrastructure architecture, scalability, and reliability. Set technical direction, define SRE practice, and mentor engineers. Build and operate AI-scale infrastructure powering real-time healthcare workflows.
Key Responsibilities
Architecture & Platform
• Design scalable, secure, cost-efficient AWS/GCP architectures for AI and healthcare workloads
• Lead IaC standards and reusable modules org-wide; own GitOps practices
• Architect optimized Kubernetes (EKS/GKE) for HA, scale, and GPU workload orchestration
Reliability Engineering
• Define and drive reliability targets (SLOs/SLIs); establish SRE practice
• Lead incident management, RCA, and blameless post-mortems for critical systems
• Partner on infrastructure hardening and compliance (SOC 2, HIPAA)
AI-First Engineering
• Lead AI/ML inference and GPU infrastructure scaling in production
• Evaluate and introduce AI tooling, platforms,
and automation best practices
• Contribute to LLM deployment pipelines and AI systems reliability
Leadership & SDLC
• Mentor L1/L2 engineers; influence cross-team infrastructure decisions
• Drive DevSecOps maturity across the SDLC (policy-as-code, SBOM, supply-chain security)
• Evaluate and introduce new tooling and platforms
Requirements
• 5–8 years infrastructure/platform/DevOps/SRE
• Deep expertise in AWS & GCP large-scale production systems
• Advanced Kubernetes (cluster design, multi-tenant ops, resource quotas)
• AI/ML inference or GPU infrastructure scaling (production track record)
• Strong Terraform / equiv. and GitOps practices
• Python, Go, or Bash scripting and automation
• Track record in scale, reliability, and cost efficiency
• Lead incident response and on-call practices
Nice-to-Have
• LLM Gateway architecture and optimization
• Zero-trust networking and advanced security/compliance (SOC 2, HIPAA)