Sr Reliability Engineer- Ford (Hyderabad)

Sr Reliability Engineer- Ford (Hyderabad)

24 Sep
|
TCS_India
|
Hyderabad

24 Sep

TCS_India

Hyderabad

mandatory

SN Required Information Details

1 Role SRE (Site Reliability Engineer) - Lead

Engineer

2 Required Technical Skill Set Site Reliability Engineering (SRE),

Google Cloud Platform (GCP),

Kubernetes, Docker, Terraform,

Jenkins, Splunk, Datadog,

Prometheus, Grafana, CI/CD,

Infrastructure as Code (IaC), Incident

Management, Observability, Python,

AI/ML Reliability, MLOps, Vertex AI,

AIOps, LangChain, LangGraph, RAG

Systems

3 No of Requirements 1

4 Desired Experience Range 05-08 years

5 Location of Requirement Chennai/ Pune/Hyderabad

Desired Competencies (Technical/Behavioral Competency)

Must-Have

(Ideally should not be more than 8-10)

 Strong hands-on experience in Site Reliability Engineering (SRE) and modern DevOps practices.

 Experience in Google Cloud Platform (GCP) environment setup, deployment, and operations.

 Strong expertise in Kubernetes and Docker container orchestration.

 Experience developing and managing CI/CD pipelines using Jenkins.

 Strong knowledge of Terraform and Infrastructure as Code (IaC).

 Experience with Splunk and enterprise observability platforms.

 Hands-on experience with Datadog, Prometheus, or Grafana monitoring solutions.

 Experience defining and managing SLIs, SLOs, SLAs, and Error Budgets.

 Strong background in Incident Management, On-call Operations, MTTA/MTTR tracking,

and Blameless Postmortems.

 Experience in Python scripting and automation.

 Understanding of AI/ML platform reliability, MLOps, and model monitoring.

 Knowledge of Vertex AI, AI observability, model drift monitoring, and AIOps concepts.

 Strong analytical, troubleshooting, and root-cause analysis skills.

 Excellent stakeholder communication and leadership skills.

Good-to-Have Behavioral Attributes:





 Ability to lead a reliability workstream independently.

 Robust decision-making and ownership mindset.

 Excellent presentation and communication skills.

 Experience working with Manufacturing Engineering platforms.

 Experience mentoring SRE, DevOps, or platform engineering teams.

 Exposure to Agentic AI, LangChain, LangGraph, and RAG architectures.

 Experience in production-scale AI/ML deployments.

 Familiarity with GitOps and ArgoCD.

 Ability to work in a fast-paced enterprise environment with

CST 2017 JD Ver 2.0

executive visibility.

 Strong customer and business-oriented mindset.

SN Responsibility of / Expectations from the Role

 Define and maintain SLOs, SLAs, SLIs, and Error Budgets for critical Manufacturing Engineering platforms.

 Design and implement enterprise-grade observability solutions across applications and infrastructure.

 Build and maintain monitoring dashboards and actionable alerting mechanisms.

 Develop and support Jenkins-based CI/CD pipelines for reliable and scalable deployments.

 Manage Infrastructure as Code using Terraform across GCP environments.

 Deploy, administer, and optimize Kubernetes-based production platforms.

 Conduct load testing, stress testing, and production readiness assessments.

 Own major incident management processes and on-call operations.

 Develop runbooks, incident response procedures,



and postmortem documentation.

 Implement reliability and observability frameworks for AI/ML applications and Agentic AI systems.

 Monitor AI model performance, model drift, data drift, and inference reliability.

 Implement AIOps capabilities including anomaly detection, predictive alerting, and incident auto-triage.

 Support GenAI-assisted incident response and automated RCA initiatives.

 Collaborate with Manufacturing Engineering, AI/ML Engineering, and Enterprise IT teams.

 Drive automation initiatives to reduce operational toil and improve system reliability.

 Mentor engineering teams on SRE best practices and operational excellence

Details For Candidate Briefing**

(It should NOT contain any confidential information or references)

About Client

Global Auto OEM

Domain: Manufacturing Engineering, Cloud Engineering, Platform Engineering, Site Reliability

Engineering (SRE), AI/ML Engineering, Digital Manufacturing

USP of the Role and Project:

 Opportunity to establish enterprise-scale SRE practices for global manufacturing platforms.

 Lead AI and Agentic AI reliability initiatives in a cutting-edge manufacturing environment.

 Work on Google Cloud Platform, Kubernetes, Terraform, and enterprise observability solutions.

 Exposure to AI/ML Reliability Engineering, MLOps, and AIOps.

 Opportunity to design enterprise monitoring, incident management, and automation frameworks.

 Work closely with global manufacturing engineering and AI teams.

 High visibility role with leadership engagement.

CST 2017 JD Ver 2.0

 Opportunity to shape future-state reliability and automation strategy for Manufacturing

Engineering platforms

📌 Sr Reliability Engineer- Ford (Hyderabad)
🏢 TCS_India
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr reliability engineer- ford (hyderabad) / hyderabad