Job Responsibilities:
- Design, build, and operate end-to-end ML systems - data ingestion, training pipelines, deployment, monitoring, and retraining
- Develop and deploy batch and real-time inference services (REST APIs) for production ML workloads
- Package, convert, and migrate ML code across environments - from platforms such as Databricks to containerized AWS-based deployments (including serverless functions)
- Build CI/CD pipelines and workflow orchestration for ML systems, including versioning, promotion, environment isolation, and rollbacks
- Monitor models and services in production, implement drift detection, and drive retraining and re-deployment
- Partner with data scientists and engineers to build models and support broader initiatives such as data collection projects
Skills:
Must have:
- 24 years of hands-on experience in software/data/ML engineering in production environments
- Very robust Python - clean, production-quality code (not just notebooks); sharp problem-solving and the aptitude to pick up MLOps practices quickly
- Experience building and deploying APIs/services (FastAPI, Flask, or similar)
and working knowledge of AWS (EC2, S3, Lambda) for hosting and serving
- Understanding of the end-to-end ML lifecycle - training vs inference pipelines, deployment, and monitoring
- Working knowledge of SQL and familiarity with PySpark; exposure to Databricks or comparable platforms, with the ability to read, refactor, and convert code for other environments
- CI/CD and version-control fundamentals (Git, testing, rollback); familiarity with Docker; strong ownership and comfort with a broad, evolving scope and global stakeholders
- Prior hands-on MLOps tooling experience (MLflow, model registries, drift detection, ML observability)
- Experience supporting GenAI or LLM workloads operationally (model serving, inference pipelines, cost/performance tuning)
Eligibility:
- Master's or Bachelor's degree in Computer Science, Engineering, Math, Statistics, or a related field
- 2–4 years of relevant hands-on experience; candidates who can join immediately will be prioritized
📌 ML Ops Engineer (Gurugram)
🏢 EXL
📍 Gurugram