09 Sep
|
Tata Consultancy Services
|
Bengaluru
09 Sep
Tata Consultancy Services
Bengaluru
Dear Professionals
Greetings from Tata consultancy Services,
Job Title : Artificial Intelligence ( AI) / Site Reliability Engineer ( SRE)
Experiernce:6 to 10 Years
Location: Bengaluru
Mode of Work : Work from Office
Mode of Interview: Virtual
If you are Interested in the above opportunity kindly share your updated resume to
[email protected] immediately with the details below (Mandatory)
Name
Contact No.
Email id
Skillset
Total exp
Relevant Exp
Fulltime highest qualification (Year of completion with percentage scored):
Current organization details (Payroll company):
Current CTC
Expected CTC
Notice period
Current location
Preffered Location
Any gaps between your education or career (If yes pls specify the duration):
Will you be able to join within 30/45 days? (Yes/NO)
Key Responsibilities
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firms Technology principles and drives efficiency and consistency, controls, security and solid governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses.
This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastrucuture engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.
As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale,
while balancing cost, security, and compliance constraints.
The ideal candidate will have strong hands-on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large-scale API Gateway environments etc. Knowledge of AIML and hands-on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team-based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements.
Skills
- Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
- Design and build automation for core platform capabilities, reducing manual toil
- Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
- Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
- Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
- Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
- Optimize cost vs. performance tradeoffs in large-scale compute environments
- Production experience in SRE / Infrastructure / ops for large-scale systems
- Strong programming/scripting skills (Python, Go, Java, or equivalent)
- Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
- Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
- Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
📌 Artificial Intelligence / Site Reliability Engineer (Bengaluru)
🏢 Tata Consultancy Services
📍 Bengaluru