02 Sep
|
ADCC Global
|
Chennai
02 Sep
ADCC Global
Chennai
01
Role Summary
Lead reliability engineering for business-critical manufacturing platforms, owning SLOs, observability, CI/CD, infrastructure automation, incident management and production readiness. Lead AI/Agentic-AI reliability covering model/agent observability, AIOps, drift monitoring, workflow guardrails and GenAI-assisted incident response.
Key Responsibilities
- Define SLI/SLO/SLA and error budgets; align release velocity with reliability targets.
- Build observability across metrics, logs and traces using Splunk, Datadog, Prometheus and Grafana.
- Create actionable alerts and dashboards that reduce alert fatigue.
- Develop Jenkins CI/CD pipelines and evaluate GitOps/Argo CD.
- Manage GCP infrastructure with Terraform; operate Docker/Kubernetes workloads.
- Conduct load/stress testing and production-readiness reviews.
- Own incident response, on-call, MTTA/MTTR, postmortems, runbooks and improvement.
- Monitor model-serving,
agent orchestration and RAG systems for availability, latency, task success and drift.
- Define AI SLOs and safeguards including auditing, human-in-the-loop escalation and runaway-agent circuit breakers.
- Implement AIOps for anomaly detection, predictive alerts, automated triage and GenAI-assisted RCA/runbooks.
- Reduce operational toil through automation; collaborate across platform, manufacturing, data/AI and IT teams; mentor engineers.
- Provide weekly reliability, incident and KPI reporting. <
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineering Lead (Chennai)
🏢 ADCC Global
📍 Chennai