Principal MLOps Engineer (Tamil Nadu)

Principal MLOps Engineer (Tamil Nadu)

09 Aug
|
NRM Analytix
|
Tamil Nadu

09 Aug

NRM Analytix

Tamil Nadu

Principal MLOps Enginee r

Location: Remote

Working Hours: 2:00 PM–10:00 PM IST

Employment Type: Contract or Full-Time

Role Overview

We are seeking a Principal MLOps Engineer to architect and build a next-generation enterprise Machine Learning Operations platform. This role will focus on developing secure, scalable, and automated ML infrastructure using AWS and Amazon SageMaker.

You will enable data science and engineering teams to train, deploy, monitor, and scale machine learning models across batch and real-time environments. The ideal candidate will have strong experience in cloud infrastructure, CI/CD automation, observability, infrastructure as code, and production-grade ML systems.

Key Responsibilities

MLOps Platform Architecture

- Architect and build scalable MLOps infrastructure using AWS and Amazon SageMaker.
- Support both real-time model hosting and batch transformation workloads.
- Implement multi-model endpoints and auto-scaling strategies to optimize performance, resource utilization, and cloud costs.
- Design secure AWS environments using VPC, PrivateLink, IAM roles, security controls, and KMS encryption.
- Develop reusable infrastructure patterns, deployment blueprints, and containerized components for standardized model promotion across environments.

CI/CD and Pipeline Automation

- Build and maintain automated pipelines for model training, testing, validation, and deployment.
- Develop Infrastructure as Code solutions using Terraform.
- Design and manage multi-stage Azure DevOps pipelines with approval gates for controlled model promotion.
- Integrate model registries to support model versioning, reproducibility, governance, and approval workflows.
- Diagnose and resolve issues related to training jobs, model deployments, pipeline execution, and registry updates.




- Automate the promotion of models across development, staging, and production environments.

Monitoring, Observability, and Governance

- Implement data quality and model drift monitoring using Amazon SageMaker Model Monitor and Amazon CloudWatch.
- Configure custom rules for detecting data drift, model performance degradation, and data quality issues.
- Build unified monitoring and observability dashboards using Grafana Cloud and Amazon CloudWatch.
- Establish proactive alerting for endpoint unavailability, high latency, error spikes, timeouts, deployment failures, pipeline failures, and drift detection events.
- Maintain operational visibility across model endpoints, infrastructure, pipelines, and supporting services.
- Track data lineage and maintain audit trails in alignment with enterprise security, compliance, and governance requirements.
- Support incident investigation and rapid resolution of production MLOps issues.

Required Qualifications

- At least 3 years of hands-on experience managing production workloads using AWS and Amazon SageMaker.
- Strong experience with AWS services, including Amazon SageMaker, Amazon S3, AWS Lambda, IAM, VPC, CloudWatch, and KMS.
- Proven proficiency with Terraform and Infrastructure as Code practices.
- Strong programming skills in Python.
- Extensive experience with Docker, containerization, and production-grade deployment patterns.




- Hands-on experience integrating data orchestration and processing tools such as AWS Glue, Apache Spark, or Apache Airflow with ML workflows.
- Experience designing, building, and maintaining Azure DevOps pipelines.
- Practical experience implementing multi-stage pipelines and approval gates for model deployment and promotion.
- Experience with Grafana Cloud and CloudWatch for monitoring infrastructure and model health.
- Strong understanding of monitoring and alerting for endpoint downtime, latency, error rates, timeouts, and model performance degradation.
- Experience troubleshooting ML training jobs, deployments, CI/CD pipelines, and production incidents.
- Strong understanding of security, access control, data lineage, auditability, and governance within cloud-based ML platforms.

Preferred Qualifications

- Experience designing enterprise-scale MLOps platforms.
- Experience with real-time inference, batch processing, multi-model endpoints, and autoscaling.
- Familiarity with model registries, model approval workflows, and ML lifecycle governance.
- Experience building reusable containers and deployment frameworks for data science teams.
- Solid understanding of cloud cost optimization and high-availability architecture.
- Excellent communication and collaboration skills, with the ability to work effectively with data scientists, software engineers, DevOps teams, and security stakeholders.

What You’ll Do in This Role You will play a key role in establishing a reliable and scalable ML platform that enables teams to move models from experimentation to production efficiently, securely, and consistently. This is an opportunity to shape the architecture, automation, governance, and observability standards for enterprise machine learning operations.

📌 Principal MLOps Engineer (Tamil Nadu)
🏢 NRM Analytix
📍 Tamil Nadu

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal mlops engineer (tamil nadu) / tamil nadu

Subscribe to this job alert:

Get the latest job offers by email for: principal mlops engineer (tamil nadu) / tamil nadu