Machine Learning Operations Site Reliability Engineer (Bengaluru)

Machine Learning Operations Site Reliability Engineer (Bengaluru)

13 Sep
|
Persistent Systems
|
Bengaluru

13 Sep

Persistent Systems

Bengaluru

About Position:

We are seeking a skilled Machine Learning Operations Site Reliability Engineer (SRE) to join our growing AI and Engineering team. In this role, you will be responsible for designing, implementing, and managing robust MLOps and LLMOps platforms that support the complete machine learning lifecycle, from model development to production deployment, monitoring, and optimization.
- Role: Machine Learning Operations Site Reliability Engineer
- Location: Bangalore
- Experience: 5 to 8 Years
- Job Type: Full time Employment

What You'll Do:
- Design, implement, and maintain enterprise-grade MLOps and LLMOps pipelines for AI and machine learning workloads.
- Build secure, scalable, and reproducible model development, testing, deployment, and monitoring frameworks.
- Support the operationalization of machine learning models and Large Language Model (LLM)-based applications.
- Develop and manage Agentic AI and Retrieval-Augmented Generation (RAG) applications in production environments.
- Implement CI/CD and CT (Continuous Training)



pipelines for machine learning models and AI solutions.
- Establish observability frameworks for monitoring model accuracy, latency, drift, performance, availability, and reliability.
- Conduct model benchmarking, load testing, and performance optimization activities.
- Implement security controls, governance standards, and compliance requirements across AI platforms.
- Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
- Support application lifecycle management across development, testing, staging, and production environments.
- Monitor platform health and proactively identify potential risks, bottlenecks, and reliability issues.
- Troubleshoot production incidents and participate in root cause analysis activities.
- Create and maintain technical documentation, architecture guidelines, operational runbooks, code samples, and implementation blueprints.
- Collaborate with Data Scientists, ML Engi

📌 Machine Learning Operations Site Reliability Engineer (Bengaluru)
🏢 Persistent Systems
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: machine learning operations site reliability engineer (bengaluru) / bengaluru