Staff Engineer – AI Data Platform | ML Infrastructure & Evaluation
Location: Viman Nagar, Pune (Hybrid 3 days onsite)
Experience: 10+ years
We are looking for a Staff Engineer to lead the architecture and technical direction of an AI Data Platform focused on ML infrastructure, model evaluation, observability, and production reliability.
Role Overview
This role is focused on designing and owning the long-term architecture for ML evaluation and observability systems that measure model quality, identify regressions, and monitor AI performance in production. You will work closely with AI/ML platform and computer vision teams and provide technical direction across the engineering organization.
Key Responsibilities
- Own the architecture and technical roadmap for a three-tier ML evaluation framework covering model-level benchmarking, end-to-end pipeline evaluation, and longitudinal performance analysis.
- Define versioned test-set governance and regression-detection architecture to identify model and pipeline regressions before production.
- Lead the technical direction for multimodal AI debugging and evaluation infrastructure.
- Own production model observability, including tracing, performance monitoring, escalation tracking, alerting, and production health signals.
- Define instrumentation standards across training and inference pipelines so model behavior can be measured end to end.
- Establish approaches for model drift detection, false-positive management, confidence scoring, and model-quality monitoring.
- Drive cost, latency,
and accuracy trade-offs for production ML workloads, including inference and model escalation strategies.
- Act as a technical escalation point for model-quality issues and lead root-cause investigations.
- Mentor ML infrastructure engineers and lead architecture/design reviews and engineering standards without formal people-management responsibility.
What We’re Looking For
- 10+ years of software engineering experience with strong production-quality Python and system design experience.
- Strong hands-on experience in ML infrastructure, MLOps, ML evaluation, model monitoring, observability, or related areas.
- Proven experience designing or owning production ML evaluation or regression-detection systems.
- Experience establishing model observability and production monitoring standards.
- Strong understanding of the ML lifecycle from training and evaluation through deployment and monitoring.
- Experience with tools such as MLflow, Weights & Biases/Weave, Datadog, or equivalent ML evaluation and observability platforms.
- Experience leading architecture or technical direction across engineering teams.
Positive to Have
- Computer Vision, Video AI, VLM or multimodal AI experience.
- Experience with IoT, smart-home, connected-device, robotics, or other real-world AI products.
- Experience with model serving, inference optimization, model routing, or production ML cost optimization.
- AWS ML experience or AWS Machine Learning certification.
📌 Staff ML Engineer (Pune)
🏢 Peoplefy
📍 Pune