Site Reliability Engineering Lead (India)

Site Reliability Engineering Lead (India)

02 Sep
|
CS Tech AI
|
India

02 Sep

CS Tech AI

India

Location: Chennai
Experience: 5-8+ years
Open Positions: 01

Role Summary

Lead reliability engineering for business-critical manufacturing platforms, owning SLOs, observability, CI/CD, infrastructure automation, incident management and production readiness. Lead AI/Agentic-AI reliability covering model/agent observability, AIOps, drift monitoring, workflow guardrails and GenAI-assisted incident response.

Key Responsibilities

- Define SLI/SLO/SLA and error budgets; align release velocity with reliability targets.
- Build observability across metrics, logs and traces using Splunk, Datadog, Prometheus and Grafana.
- Create actionable alerts and dashboards that reduce alert fatigue.
- Develop Jenkins CI/CD pipelines and evaluate GitOps/Argo CD.
- Manage GCP infrastructure with Terraform; operate Docker/Kubernetes workloads.
- Conduct load/stress testing and production-readiness reviews.
- Own incident response, on-call, MTTA/MTTR, postmortems, runbooks and improvement.
- Monitor model-serving, agent orchestration and RAG systems for availability, latency, task success and drift.
- Define AI SLOs and safeguards including auditing, human-in-the-loop escalation and runaway-agent circuit breakers.
- Implement AIOps for anomaly detection, predictive alerts, automated triage and GenAI-assisted RCA/runbooks.
- Reduce operational toil through automation; collaborate across platform, manufacturing, data/AI and IT teams; mentor engineers.
- Provide weekly reliability, incident and KPI reporting.

Required Qualifications

- 5–8+ years in SRE, DevOps or platform engineering.
- Strong SLI/SLO, error-budget, observability and incident-management expertise.
- Strong GCP experience including networking, compute, managed services and access.
- Production experience with Kubernetes, Docker and Terraform.
- Jenkins CI/CD experience; GitOps/Argo CD exposure desirable.
- Splunk plus Datadog, Prometheus or Grafana experience.
- Python and/or Java scripting capability.
- Load/stress testing, capacity analysis and production-readiness experience.




- PagerDuty, Opsgenie or equivalent on-call tooling knowledge.
- Exposure to AI/ML reliability, MLOps, AIOps or Agentic-AI operations.

AI & Agentic-AI Reliability Experience

- Vertex AI: model monitoring, pipelines and agent-building services.
- LangChain, LangGraph or comparable agent orchestration frameworks.
- RAG, vector search and retrieval-system observability.
- Model performance and data-drift monitoring.
- Inference latency, model availability and agent task-completion observability.
- GenAI-assisted incident response, runbook retrieval, incident copilots and RCA summarisation.

Behavioural & Leadership Competencies

- Technical leadership and architectural decision-making.
- Analytical troubleshooting of distributed and AI-system failures.
- Transparent communication of reliability, risks and progress to technical/executive audiences.
- Cross-functional collaboration and SRE mentoring.
- Operational ownership through incident closure, corrective actions and documentation.

Key Deliverables

- SLO/error-budget framework and integrated observability dashboards.
- Controlled CI/CD and Terraform-based infrastructure automation.
- Incident-management process, on-call model, runbooks and postmortems.
- Load/stress testing evidence for production readiness.
- Initial AI/Agentic-AI observability and AIOps/GenAI incident-response proof of concept.
- Modular onboarding architecture and regular reliability/KPI reporting.

Initial Success Measures

- SLOs operational for at least three critical platforms within three months.
- Initial model-monitoring or AIOps capability operational within three months.
- Critical-incident MTTA below 15 minutes with continuous MTTR improvement.
- Operational toil at or below 50% per sprint, with remaining capacity for automation.

Additional Expectations

- Flexibility for onsite collaboration, milestones and on-call participation.
- Approved engineering workstation with GCP and AI/ML tooling access.
- Strong documentation and modular architecture to support handover and scale-up.

Send your CVs to [email protected]

📌 Site Reliability Engineering Lead (India)
🏢 CS Tech AI
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineering lead (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineering lead (india) / india