14 Aug
|
Neurodrift
|
India
Fully remote · Immediate joiners only, non-negotiable
NeuroDrift builds production voice AI for enterprise contact centres. Our platform handles 300+ production calls a day and we have shipped over 350 voice AI deployments. Models we pick and tune sit in a live call path with a real customer on the other end, where a 200ms regression is something people hear.
We are hiring one senior, hands-on engineer to own the model layer: choosing models, serving them ourselves, fine-tuning them, and proving they hold up under load.
This is a deep individual contributor role. You write code every day and you are measured on what runs in production, not on managing people.
What you'll do
- Benchmark open-weight and hosted models against real workloads on quality, latency, throughput and cost per call, and make the call on what ships
- Deploy and serve models yourself on GPU using vLLM, Triton, TGI or similar, containerized, including into a client's own cloud account
- Fine-tune and adapt models, including LoRA and QLoRA and supervised fine-tuning, and build the data pipelines behind it
- Build evaluation harnesses and datasets that catch regressions before a customer does
- Tune inference for production: quantization, batching, KV cache, concurrency and GPU utilisation
- Get all of it into a live call path alongside our voice and product engineers
Who this is for
- 3 to 7 years engineering, still writing code daily, solid in Python
- You have deployed and served open-weight models in production yourself, not only called a hosted API
- You have run structured evaluations, built your own eval sets, and made model decisions from data rather than vibes
- You have fine-tuned a model and shipped the result
- Comfortable on GPU infrastructure: memory limits, quantization, throughput tuning, cost
- You deploy and operate your own services on AWS with Docker and CI/CD
Nice to have
- Speech models: ASR and TTS, self-hosted or containerized
- Real-time or streaming inference, where latency budgets are tight and non-negotiable
- Multi-provider orchestration with fallback across model vendors
- Distillation, quantization-aware work, or serving optimisation beyond the defaults
How we work
- Claude Code from day one. We are an AI-native team and expect you to ship faster with AI tooling, not work around it
- Small teams, weekly releases, no layers between you and the outcome
- Startup pace. This is a build, not a maintenance role The setup
- Fully remote
- Primarily IST, with availability for US-time client calls
- Immediate joiners only. Non-negotiable
- Compensation is set against a startup band and against what you can build, not years on a CV. If you are coming from a large-enterprise package, this will not match it
- This is for builders. If you have moved into diagrams and delegation, this is not the one.
📌 Senior AI Engineer, LLM Deployment and Evaluation (India)
🏢 Neurodrift
📍 India