09 Sep
|
Accolite Digital India
|
Hyderabad
09 Sep
Accolite Digital India
Hyderabad
Role Overview
We are looking for a Senior MLOps + DevOps Engineer (6+ years) to architect, build, and scale AI/ML platforms in an on-prem enterprise environment.
This role requires end-to-end ownership of ML systems, infrastructure, CI/CD, and production reliability, enabling scalable deployment of machine learning and GenAI solutions.
Key Responsibilities
1. 1. Platform Architecture & Ownership
- Design and own end-to-end ML platform architecture (data → training → deployment → monitoring)
- Define and enforce best practices for scalable and secure ML systems
- Standardize MLOps + DevOps frameworks and processes
-AI/ML Platform Architecture deployed and should be able to explain it clearly.
-Ability to design scalable ML/data platforms end-to-end
-Explain real-world integration patterns using Kafka, especially on-premise setups.
1. 2. Model Deployment & Serving
- Deploy and manage ML/LLM models on GPU-based on-prem infrastructure
- Optimize inference performance (latency, throughput, batching)
- Implement model versioning, A/B testing, and rollback strategies
1. 3. CI/CD & Automation
- Design and implement CI/CD pipelines for ML models, APIs, and data workflows
- Enable automated testing, deployment, and release management
1. 4. Infrastructure & Containerization
- Manage Linux-based (RHEL preferred) on-prem infrastructure
- Containerize applications using Docker
- Deploy and orchestrate workloads using Kubernetes / OpenShift
- Operate within restricted or air-gapped environments
-Understanding of OpenShift AI ecosystem
1. 5. Data & System Integration
- Build pipelines integrating structured databases and high-volume logs/streaming data
- Support batch and real-time inference architectures
1. 6. Monitoring, Observability & Reliability
- Implement end-to-end observability (model + infra)
- Use tools like Prometheus, Grafana, ELK stack
- Ensure high availability, SLA adherence, and incident response
1. 7. GenAI & Advanced ML Systems
- Deploy RAG pipelines and vector databases
- Manage LLM serving frameworks
- Work with agent orchestration frameworks
1. 8. Leadership & Collaboration
- Mentor engineers on MLOps and DevOps best practices
- Collaborate with cross-functional teams
- Drive design reviews and production readiness
Required Skills:
- Strong Python and scripting (Bash)
- Deep understanding of ML lifecycle and productionization
- Experience deploying ML/LLM systems in production
- Linux, Docker, Kubernetes/OpenShift
- CI/CD tools (Jenkins/GitLab CI)
- SQL and data pipeline experience
-Explain real-world integration patterns using Kafka, especially on-premise setups.
-Should be able to clearly explain Gunicorn
-Data ingestion pipelines (real-time and batch)
-AI/ML Platform Architecture deployed and should be able to explain it clearly.
-Understanding of OpenShift AI ecosystem
-Ability to design scalable ML/data platforms end-to-end
Valuable to Have
- GPU optimization knowledge
- MLflow / Kubeflow
- Terraform / Ansible
- Experience in on-prem or restricted environments
Experience
- 6+ years in MLOps / DevOps / Platform Engineering
- Proven experience scaling production ML systems
Ideal Candidate
A hands-on platform architect who can operate across ML systems and infrastructure, driving automation, scalability, and reliability.
About Bounteous
Bounteous is a global AI Services firm where agentic engineering and human experience converge to deliver transformative business outcomes for the enterprise. We help organizations design, build, and scale AI-driven products, platforms, and processes.
With more than 5,000 team members worldwide, Bounteous helps organizations take AI from experimentation to enterprise scale.
Learn more at www.bounteous.com
📌 Senior MLOps + DevOps Engineer (Hyderabad)
🏢 Accolite Digital India
📍 Hyderabad