- Openings: 1
- Exp: 7+years
- Mumbai
- Full-Time
We seek a seasoned AI/ML Engineer specializing in Natural Language Processing (NLP) and Transformer architectures to design, optimize, and deploy cutting-edge language models (e.g., BERT, GPT, T5) for real-world applications. You’ll own the end-to-end ML lifecycle - from research to rapid inferencing - leveraging distributed computing, container orchestration (Kubernetes), and microservices to serve models at scale.
Key Responsibilities
Model Development & Optimization
- Research, fine-tune, and deploy Transformer-based models (e.g., LLMs, encoder-decoder architectures) for tasks like text generation, summarization, and entity recognition.
- Optimize models for low-latency inferencing (e.g., quantization, distillation, ONNX runtime).
- Architect and deploy low-latency inference pipelines, optimizing for throughput (10K+ RPS) using Triton, vLLM, or TensorRT.
- Implement distributed training (PyTorch FSDP, Horovod) across GPU clusters (AWS/GCP/Azure).
- Implement distributed model serving (multi-GPU, multi-node) with Kubernetes/Knative and autoscaling.
- Pioneer hybrid cloud strategies (e.g., on-prem GPU clusters + cloud burst) for cost-effective training/inference.
ML Pipeline & Infrastructure
- Architect scalable ML pipelines (feature store, model registry, CI/CD) using MLflow, Kubeflow, or TFX.
- Containerize models (Docker)
and orchestrate via Kubernetes for high-availability deployments.
- Design microservices (FastAPI, gRPC) to expose models as APIs with <100ms latency.
Performance & Observability
- Monitor model performance (drift, accuracy) using Prometheus/Grafana.
- Implement A/B testing and shadow deployments for zero-downtime updates.
Collaboration & Leadership
- Mentor junior engineers and collaborate with Data Scientists, DevOps, and Product Teams.
- Translate business requirements into SOTA NLP solutions with measurable ROI.
Required Skills & Qualifications
Technical Expertise:
- 7+ years in NLP, with 3+ years focused on Transformer models (e.g., BERT, RoBERTa, GPT variants).
- Mastery of PyTorch/TensorFlow, Hugging Face transformers, and distributed training (DDP, DeepSpeed), distributed systems (Ray, Horovod, Spark).
- Production experience with Kubernetes, Docker, and cloud ML services (SageMaker, Vertex AI).
- Strong Python skills (asyncio, multiprocessing) and familiarity with Go/Rust for high-performance services.
Deployment & Ops:
- Proven track record of deploying low-latency model APIs (FastAPI, Triton Inference Server).
- Knowledge of ML observability tools (Weights & Biases, Evidently).
Can't find your role? Email us your resume at
[email protected]
📌 Senior AI/ML Engineer (India)
🏢 Stigasoft
📍 India