23 Aug
|
Cloudxtreme
|
Bengaluru
23 Aug
Cloudxtreme
Bengaluru
- MLOps + FullStack :
- Contract to Hire
- Experience : 6 -12 Years
- Location : PAN INDIA
: Senior MLOps & Full Stack Systems Engineer
Job Summary
We are seeking a versatile and high-impact Senior MLOps & Full Stack Systems Engineer to bridge the gap between production machine learning infrastructure and user-facing enterprise platforms. In this hybrid engineering role, you will be responsible for designing, building, and operating end-to-end ML delivery pipelines while engineering the full-stack software interfaces (React frontend + Java/Python backend services) required to operationalize, monitor, and interact with production AI/ML models.
The ideal candidate possesses a unique blend of core software engineering rigor and machine learning platform operations. You will build automated CI/CD and Continuous Training (CT) pipelines, manage model registries and feature stores, deploy low-latency model inference microservices, and build full-stack interactive dashboards and admin portals using React (TypeScript) and Java (Spring Boot) / Python (FastAPI/Flask).
Key Roles & Responsibilities
1. MLOps & Platform Architecture
- Design, build, and maintain automated ML lifecycle pipelines spanning data ingestion, preprocessing, model training, validation, packaging, and continuous deployment (CI/CD/CT).
- Implement and manage Model Registries, Experiment Tracking, and Feature Stores using tools such as MLflow, Kubeflow, Feast, or Weights & Biases.
- Package, optimize, and deploy ML models into production as containerized microservices (using Triton Inference Server, TorchServe, TF Serving, or FastAPI).
- Implement real-time model monitoring, drift detection (data drift, concept drift), latency tracking, and automated retraining workflows using tools like Evidently AI, Prometheus, and Grafana.
2. Backend & Systems Engineering (Java / Python)
- Architect and implement robust, high-throughput microservices using Java (Spring Boot / Micronaut) or Python (FastAPI / Flask / AsyncIO) to serve as orchestration layers between frontend applications and ML inference engines.
- Design and implement secure RESTful APIs, gRPC services, and event-driven data streaming pipelines using Apache Kafka, RabbitMQ, or AWS Kinesis.
- Optimize backend services for low-latency scoring, batch inference jobs, asynchronous request queuing, and distributed data caching (Redis).
- Manage data persistence and access patterns across relational (PostgreSQL, MySQL) and NoSQL (MongoDB, DynamoDB, Vector DBs) data stores.
3. Frontend & UI Engineering (React.js)
- Design and develop responsive, modern web applications, internal tools, and administrative control panels using React.js, TypeScript, and modern UI libraries (Tailwind CSS, MUI).
- Build interactive model governance dashboards, telemetry visualizations, data labeling portals,
and human-in-the-loop (HITL) review interfaces.
- Implement state management using Redux Toolkit, React Context, or Zustand, and handle asynchronous data fetching/caching using TanStack Query (React Query).
- Integrate complex data visualization libraries (e.g., D3.js, Chart.js, Recharts) to render live inference metrics, feature importance plots, and operational telemetry.
4. Infrastructure, Cloud & Container Orchestration
- Deploy and orchestrate containerized workloads and ML pipelines on Kubernetes (EKS / GKE / AKS) using Docker, Helm, and service meshes (Istio).
- Implement Infrastructure as Code (IaC) using Terraform or CloudFormation to automate setting provisioning across AWS, GCP, or Azure.
- Configure autoscaling policies for inference endpoints based on GPU/CPU utilization and custom queue metrics.
- Enforce security best practices, including RBAC, secret management (HashiCorp Vault / AWS Secrets Manager), and API gateway security (OAuth 2.0 / JWT).
5. Quality, Governance & Cross-Functional Collaboration
- Enforce end-to-end testing standards: Unit/Integration tests for frontend (Jest, React Testing Library), backend (JUnit, PyTest), and data/model validation (Great Expectations).
- Collaborate closely with Data Scientists, ML Researchers, Product Owners, and DevOps teams to translate prototype algorithms into reliable, scalable production systems.
- Lead architecture reviews, maintain comprehensive documentation, and mentor team members in full-stack and MLOps best practices.
📌 MLOps + Full Stack (Bengaluru)
🏢 Cloudxtreme
📍 Bengaluru