13 Sep
|
Sama AI
|
Coimbatore
13 Sep
Sama AI
Coimbatore
Job Title: MLOps / AI Infrastructure Engineer
Experience: 5–8 Years
Job Type: Full-time
Job Summary
We are looking for an experienced MLOps / AI Infrastructure Engineer with strong expertise in Kubernetes, AWS, cloud-native infrastructure, GPU workloads, and production ML/GenAI model serving.
The ideal candidate will have hands-on experience building and operating scalable AI platforms, Kubernetes-based infrastructure, and production-grade model-serving environments.
Key Responsibilities & Requirements
- 5–8 years of experience in Software Engineering, Platform Engineering, DevOps, SRE, MLOps, or related infrastructure roles.
- Solid hands-on experience with Kubernetes, including development of Kubernetes Operators and Custom Resource Definitions (CRDs) using Kubebuilder, Operator SDK, or equivalent frameworks.
- Strong experience designing and operating cloud-native infrastructure on AWS, particularly:
- Amazon EKS
- EC2
- S3
- ECR
- IAM
- VPC
- CloudWatch
- Experience deploying and operating Machine Learning, Deep Learning, or Generative AI models in production.
- Hands-on experience running and troubleshooting GPU-accelerated workloads on Kubernetes, including GPU scheduling, utilization, memory constraints, and performance optimization.
- Experience with model-serving technologies such as:
- KServe
- NVIDIA Triton Inference Server
- vLLM
- Ray Serve
- TorchServe
- or equivalent frameworks.
- Strong programming skills in Go or Python,
with experience developing production-grade APIs, controllers, or distributed backend services.
- Experience with Docker/containers, Helm, CI/CD, Infrastructure as Code (IaC), and observability tools such as Prometheus, OpenTelemetry, and Grafana.
- Strong understanding of Linux, networking, storage, security, and distributed systems.
- Excellent debugging, problem-solving, communication, and cross-functional collaboration skills.
Preferred Qualifications
- Experience building an MLOps platform, AI Infrastructure Platform, Internal Developer Platform (IDP), or Kubernetes-based enterprise product.
- Experience with CUDA C/C or Triton for GPU kernels or performance-critical development.
- Experience with LLM serving, distributed inference, batching, quantization, and inference performance optimization.
- Experience with NVIDIA GPU Operator, MIG, GPU time-slicing, Dynamic Resource Allocation (DRA), or similar GPU-management technologies.
- Familiarity with AWS Inferentia, AWS Trainium, Amazon SageMaker, or Amazon Bedrock.
- Experience operating AI platforms in hybrid-cloud, on-premises, air-gapped, or multi-tenant environments.
Ideal Candidate The ideal candidate is a hands-on engineer who can work across Kubernetes AWS MLOps GPU infrastructure AI/LLM serving, and who is comfortable building reliable, scalable, production-grade AI infrastructure.
Pay: ₹923,856.93 - ₹3,000,000.00 per year
Work Location: In person
📌 ML Ops (Coimbatore)
🏢 Sama AI
📍 Coimbatore