19 Aug
|
Tata Consultancy Services
|
Bengaluru
19 Aug
Tata Consultancy Services
Bengaluru
Key Responsibilities
- Manage and support infrastructure powering AI/ML and Generative AI applications.
- Design and implement scalable, highly available, and secure platform solutions.
- Build automation to reduce operational toil and improve platform reliability.
- Develop and maintain Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
- Operate Kubernetes-based environments, container platforms, and cloud services.
- Establish and monitor SLIs, SLOs, SLAs, observability dashboards, and alerting.
- Lead incident response, root cause analysis (RCA), and reliability improvement initiatives.
- Perform capacity planning, performance optimization, and cost management.
- Support GPU-based compute environments, AI model serving, and data pipelines.
- Implement disaster recovery, backup, resiliency, and security controls.
- Collaborate with engineering and security teams to deploy and operate AI services safely.
- Maintain operational documentation, runbooks, and best practices.
- Participate in on-call support and production incident management.
Required Skills & Experience
- 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform Engineering.
- Solid programming/scripting skills in Python, Go, Java, or similar languages.
- Hands-on experience with Docker,
Kubernetes, and container orchestration.
- Experience with AWS, Azure, or Google Cloud Platform (GCP).
- Strong knowledge of Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
- Experience with Observability and Monitoring tools such as Grafana, Prometheus, Loki, ELK/EFK, Datadog, OpenTelemetry.
- Knowledge of networking fundamentals: TCP/IP, DNS, Load Balancing, Routing.
- Experience in incident management, RCA, performance tuning, and system scaling.
- Strong understanding of security, compliance, governance, and production support.
- Excellent communication and cross-functional collaboration skills.
Preferred Skills
- Experience supporting Generative AI, LLM, MLOps, ModelOps, or AI Platforms.
- Knowledge of GPU clusters, HPC environments, Slurm, Kubernetes GPU scheduling.
- Experience with Kafka, Spark, Flink, and distributed data processing frameworks.
- Familiarity with databases such as Snowflake, Redis, SQL, PostgreSQL.
- Understanding of Embeddings, Fine-Tuning, RAG, Vector Databases, Model Serving.
- Experience with Canary Deployments, Blue-Green Deployments, Chaos Engineering.
- Exposure to financial services or highly regulated environments.
📌 AI SRE/ AI Site Reliability Engineer (Bengaluru)
🏢 Tata Consultancy Services
📍 Bengaluru