20 Aug
|
Tata Consultancy Services
|
Bengaluru
20 Aug
Tata Consultancy Services
Bengaluru
Key Responsibilities
Manage and support infrastructure powering AI/ML and Generative AI applications.
Design and implement scalable, highly available, and secure platform solutions.
Build automation to reduce operational toil and improve platform reliability.
Develop and maintain Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
Operate Kubernetes-based environments, container platforms, and cloud services.
Establish and monitor SLIs, SLOs, SLAs, observability dashboards, and alerting.
Lead incident response, root cause analysis (RCA), and reliability improvement initiatives.
Perform capacity planning, performance optimization, and cost management.
Support GPU-based compute environments, AI model serving, and data pipelines.
Implement disaster recovery, backup, resiliency, and security controls.
Collaborate with engineering and security teams to deploy and operate AI services safely.
Maintain operational documentation, runbooks, and best practices.
Participate in on-call support and production incident management.
Required Skills & Experience
5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform Engineering.
Solid programming/scripting skills in Python, Go, Java, or similar languages.
Hands-on experience with Docker,
Kubernetes, and container orchestration.
Experience with AWS, Azure, or Google Cloud Platform (GCP).
Robust knowledge of Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
Experience with Observability and Monitoring tools such as Grafana, Prometheus, Loki, ELK/EFK, Datadog, OpenTelemetry.
Knowledge of networking fundamentals: TCP/IP, DNS, Load Balancing, Routing.
Experience in incident management, RCA, performance tuning, and system scaling.
Solid understanding of security, compliance, governance, and production support.
Excellent communication and cross-functional collaboration skills.
Preferred Skills
Experience supporting Generative AI, LLM, MLOps, ModelOps, or AI Platforms.
Knowledge of GPU clusters, HPC environments, Slurm, Kubernetes GPU scheduling.
Experience with Kafka, Spark, Flink, and distributed data processing frameworks.
Familiarity with databases such as Snowflake, Redis, SQL, PostgreSQL.
Understanding of Embeddings, Fine-Tuning, RAG, Vector Databases, Model Serving.
Experience with Canary Deployments, Blue-Green Deployments, Chaos Engineering.
Exposure to financial services or highly regulated environments.
📌 Ai Sre/ Ai Site Reliability Engineer Bengaluru
🏢 Tata Consultancy Services
📍 Bengaluru