We are Hiring: Senior ML Infrastructure Engineer
Vadodara, Gujarat | Wastefull Insights
Wastefull Insights builds AI-powered robotic sorting systems (Gantry Robot, Vision Box, Kuda.AI) solving one of India's toughest problems - sorting messy, mixed waste streams at scale. Backed by NVIDIA, NASSCOM, iCreate & the Govt. of Gujarat.
About the Role
As we grow from a handful of live deployments to many Materials Recovery Facilities - each running a growing number of vision systems and sorting robots - the hardest problem stops being “can we build an accurate model” and becomes “can we reliably manage data, training, and deployment at scale, as we keep adding sites.”
We are looking for a Senior ML Infrastructure Engineer to own this end-to-end: our AI datasets, annotation pipeline, model training and retraining, and the cloud and on-site infrastructure all of it runs on. You'll be the person responsible for architecting these systems - on AWS or a comparable public cloud - to be scalable from the ground up, not patched together as we grow. This is a systems and infrastructure role, not a model-research role - the goal is one person who can be trusted to own AI data, models, and infrastructure end-to-end as we scale.
Key Responsibilities
- Own how data moves from cameras, sensors, and connected devices across our sites into one reliable, organized system - nothing lost, nothing duplicated, nothing mixed up between clients
- Design and build the underlying architecture this runs on - this is a build-it role, not a use-it role
- Take real, hands-on ownership of our AWS (or equivalent public cloud) environment - storage, compute, and data movement - architecting it properly rather than assembling it ad hoc
- Be comfortable working with very large volumes of real-world data - think lakhs of images - and build systems that stay organized and fast as that volume keeps growing
- Own the annotation and data-quality process end-to-end, and make sure it scales smoothly as data volume grows well beyond what one person can manage by hand
- Build a system where models get retrained and updated automatically and reliably - not someone manually kicking off a training run every time
- Own the full model lifecycle - training, evaluation, retraining, and getting new or improved models into production smoothly
- Own deployment and monitoring of models across a growing number of in office devices at multiple client locations - always knowing which version is running where
- Design all of this to scale properly from the start - architecture meant to handle significantly more devices, sites, and data than exists today, not something that needs to be rebuilt at every stage of growth
- Own the data layer feeding the dashboard - ensuring operational data flows from every site into it reliably and close to real-time
- Work directly with our dashboard developer to define clean APIs and data contracts, so the data reaching the dashboard is well-structured and easy to build on - not something they have to untangle themselves
- Work daily with the CV/ML team,
robotics team, and dashboard developer to keep data, models, and product in sync
- Document architecture, retraining processes, and deployment procedures so the system doesn't depend on any one person's memory
- Coordinate with the CEO on infrastructure priorities, scaling plans, and site rollout timelines
What You Bring
- Extensive, hands-on experience in computer vision and AI - real depth in building, training, and deploying vision models, not adjacent or generalist ML exposure
- Experience deploying or working with LLMs and VLMs (Large Language Models / Vision-Language Models) in real production settings, not just experimenting with hosted APIs
- Experience working with real-time edge devices and GPUs - deploying and optimizing models to run directly on hardware in the field, with real constraints on latency and compute
- Experience working at the API level - designing, building, or integrating APIs so different systems (data pipeline, models, dashboard, devices) can talk to each other cleanly, not just working self-contained within one system
- Experience with streaming or high-throughput data movement (e.g. Kafka or similar) - comfortable handling continuous data arriving from many sources at once, not just batch jobs that run periodically
- Experience with containerized deployment (e.g. Docker) and running or orchestrating services consistently across both cloud and edge environments (e.g. Kubernetes or similar)
- Experience serving models in production - packaging and running inference efficiently, whether in the cloud or directly on edge hardware
- Experience setting up CI/CD pipelines - automating testing and deployment so changes to code or models go out safely and consistently, not manually
- Experience setting up monitoring and alerting for production systems - knowing when something has failed before a client has to tell you
- Comfortable working across a mix of storage systems - large media/object storage, structured databases, and systems built for high-volume device or sensor data
- Experience using Infrastructure as Code (e.g. Terraform or similar) to build and manage cloud environments in a repeatable, version-controlled way - not manually configuring things through a console
- Experience with configuration and secrets management - keeping credentials and settings secure and consistent across multiple environments and devices
- Comfortable designing for fault tolerance and cost-efficient resource use - so one device or component failing doesn't take down a whole site, and cloud costs stay sensible for a growing startup, not an enterprise budget
- 3–5 years of experience owning production data or ML systems end-to-end - not just analyzing data, but being responsible for datasets, models, and infrastructure that run continuously in the real world
- Strong, hands-on experience with AWS or a comparable public cloud (Azure, GCP) - real experience architecting and operating cloud infrastructure in production, not just using managed notebooks or occasional cloud tools
- Experience working on a live deployment that collects data - images, video, or sensor streams - continuously and in real-time from the field, not just working with static, pre-collected datasets
- A track record of designing scalable architecture - systems built (or rebuilt) to handle significantly more data, devices, or locations than they started with, and the ability to explain those design decisions clearly
- Strong Python and SQL, and real experience handling large-scale, messy, real-world data - comfortable working across lakhs of images or records, not just clean, sample-sized datasets
- Experience automating processes that used to be manual - e.g. turning a one-off training run or deployment step into something that happens reliably on its own
- Experience managing or monitoring systems running across multiple locations or devices at once, and knowing quickly when something, somewhere, isn't working
- Ability to work with messy, real-world industrial data and design systems that hold up outside a clean, single-site pilot
- Strong documentation instincts - comfortable being the person who builds systems others can rely on without hand-holding
Nice to Have
- Experience with device communication protocols used in IoT/edge fleets (e.g. MQTT or similar)
- Understanding of data security and access practices, particularly when handling data across multiple clients or geographies
- Experience keeping data from different clients or business units cleanly separated within the same system
- Experience deploying or monitoring software running on physical, on-site hardware (rather than only cloud servers)
- Prior experience at an early-stage startup, owning infrastructure with limited resources
- Exposure to robotics, industrial automation, recycling, or sustainability domains
KPIs / Success Metrics
- How quickly a new or improved model can go from ready to safely running across all active sites
- How clearly we can see, at any time, what's running where - and how fast an issue at any single site gets noticed
- How little manual intervention is needed to keep data flowing smoothly from every site into the dashboard
- System reliability holding steady, or improving, as we add more sites and devices
Why Join Us
- Own infrastructure that is genuinely load-bearing - not a side project, the backbone of how our AI scales
- Be part of a growing deep-tech startup backed by NVIDIA, NASSCOM, and the Government of Gujarat
- See your systems run real robots sorting real waste at real industrial sites, not just in a notebook
- Direct, visible impact on how far and how fast Wastefull Insights can scale
[email protected]
📌 ML Infrastructure Engineer (Vadodara)
🏢 Wastefull Insights
📍 Vadodara