AI SRE/ AI Site Reliability Engineer (Bengaluru)

AI SRE/ AI Site Reliability Engineer (Bengaluru)

19 Aug
|
Tata Consultancy Services
|
Bengaluru

19 Aug

Tata Consultancy Services

Bengaluru

Key Responsibilities

- Manage and support infrastructure powering AI/ML and Generative AI applications.
- Design and implement scalable, highly available, and secure platform solutions.
- Build automation to reduce operational toil and improve platform reliability.
- Develop and maintain Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
- Operate Kubernetes-based environments, container platforms, and cloud services.
- Establish and monitor SLIs, SLOs, SLAs, observability dashboards, and alerting.
- Lead incident response, root cause analysis (RCA), and reliability improvement initiatives.
- Perform capacity planning, performance optimization, and cost management.
- Support GPU-based compute environments, AI model serving, and data pipelines.
- Implement disaster recovery, backup, resiliency, and security controls.
- Collaborate with engineering and security teams to deploy and operate AI services safely.
- Maintain operational documentation, runbooks, and best practices.
- Participate in on-call support and production incident management.

Required Skills & Experience

- 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform Engineering.
- Solid programming/scripting skills in Python, Go, Java, or similar languages.
- Hands-on experience with Docker,



Kubernetes, and container orchestration.
- Experience with AWS, Azure, or Google Cloud Platform (GCP).
- Strong knowledge of Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
- Experience with Observability and Monitoring tools such as Grafana, Prometheus, Loki, ELK/EFK, Datadog, OpenTelemetry.
- Knowledge of networking fundamentals: TCP/IP, DNS, Load Balancing, Routing.
- Experience in incident management, RCA, performance tuning, and system scaling.
- Strong understanding of security, compliance, governance, and production support.
- Excellent communication and cross-functional collaboration skills.

Preferred Skills

- Experience supporting Generative AI, LLM, MLOps, ModelOps, or AI Platforms.
- Knowledge of GPU clusters, HPC environments, Slurm, Kubernetes GPU scheduling.
- Experience with Kafka, Spark, Flink, and distributed data processing frameworks.
- Familiarity with databases such as Snowflake, Redis, SQL, PostgreSQL.
- Understanding of Embeddings, Fine-Tuning, RAG, Vector Databases, Model Serving.
- Experience with Canary Deployments, Blue-Green Deployments, Chaos Engineering.
- Exposure to financial services or highly regulated environments.

📌 AI SRE/ AI Site Reliability Engineer (Bengaluru)
🏢 Tata Consultancy Services
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai sre/ ai site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: ai sre/ ai site reliability engineer (bengaluru) / bengaluru