24 Sep
|
Invyte.ai
|
India
This role sits at the intersection of multimodal AI, real-time voice, vision-language models, mobile-use agents, memory, on-device intelligence, and production-grade evaluation systems
. You will help build an AI system that can understand the user’s physical context through camera, microphones, device state and memory - then turn that context into useful action.
- Minimum 3+ years of qualified experience.
- Agentic systems
- multi-step task execution, planning and decomposition, state management
- Apply quantization (PTQ and QAT), pruning, and architecture search to hit per-product size, latency, and power budgets and efficient in working with SLM’s
- Applied vision-language models
- prompt architecture, structured output, context management, and visual understanding of interfaces, design and execute distillation strategies.
- Scale simulation and scaffolding environments for agentic RL: code execution sandboxes, computer use environments,
tool-calling harnesses, and Mobile Use Agents.
- Evaluation infrastructure
- task suites, automated scoring, regression harnesses, and success / latency / cost measurement
- Data pipelines
- capture, schema design, labelling, privacy-safe handling, and dataset curation for training
- Real-time voice
- ASR and TTS integration, streaming interaction, and end-to-end latency engineering
- Deep understanding of reinforcement learning: policy optimization, reward design, exploration, and the interplay between environment design and agent behavior.
- Model adaptation
- fine-tuning workflows, dataset construction, and running models under tight compute and memory budgets
- Partner with our mobile and hardware engineers to move capability from the cloud onto the device
📌 Applied Scientist (Multimodal AI) (India)
🏢 Invyte.ai
📍 India