Job Responsibilities:
Design, build, and operate end-to-end ML systems - data ingestion, training pipelines, deployment, monitoring, and retraining
Develop and deploy batch and real-time inference services (REST APIs) for production ML workloads
Package, convert, and migrate ML code across environments - from platforms such as Databricks to containerized AWS-based deployments (including serverless functions)
Build CI/CD pipelines and workflow orchestration for ML systems, including versioning, promotion, environment isolation, and rollbacks
Monitor models and services in production, implement drift detection, and drive retraining and re-deployment
Partner with data scientists and engineers to build models and support broader initiatives such as data collection projects
Skills:
Must have:
2–4 years of hands-on experience in software/data/ML engineering in production environments
Very robust Python - clean, production-quality code (not just notebooks); sharp problem-solving and the aptitude to pick up MLOps practices quickly
Experience building and deploying APIs/services (FastAPI, Flask, or similar) and working knowledge of AWS (EC2, S3, Lambda) for hosting and serving
Understanding of the end-to-end ML lifecycle - training vs inference pipelines, deployment, and monitoring
Working knowledge of SQL and familiarity with PySpark; exposure to Databricks or comparable platforms, with the ability to read, refactor, and convert code for other environments
CI/CD and version-control fundamentals (Git, testing, rollback); familiarity with Docker; strong ownership and comfort with a broad, evolving scope and global stakeholders
Prior hands-on MLOps tooling experience (MLflow, model registries, drift detection, ML observability)
Experience supporting GenAI or LLM workloads operationally (model serving, inference pipelines, cost/performance tuning)
Eligibility:
Master's or Bachelor's degree in Computer Science, Engineering, Math, Statistics, or a related field
2–4 years of rele
📌 ML Ops Engineer (India)
🏢 EXL
📍 India