14 Aug
|
Computer Futures
|
India
14 Aug
Computer Futures
India
About the Role
We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI systems that run efficiently on smartphones, wearables, and other resource-constrained devices.
This role sits at the intersection of machine learning research, systems optimization, and production engineering. You will design compact model architectures, develop advanced compression techniques, and optimize inference pipelines that enable real-time speech AI experiences directly on end-user devices.
You will work across a broad range of voice technologies, including automatic speech recognition (ASR), text-to-speech (TTS), speech translation, speech-to-speech systems, and neural audio codecs.
The ideal candidate combines robust research credentials with hands-on implementation skills and has a deep understanding of efficient deep learning, model optimization, and hardware-aware machine learning.
What You'll Do
Research and Model Development
- Drive research in efficient machine learning and edge AI for speech and audio applications.
- Design compact model architectures capable of operating under strict latency, memory, and power constraints.
- Develop and improve state-of-the-art approaches for:
- Knowledge distillation
- Model pruning
- Quantization
- Low-rank adaptation and compression
- Hardware-aware architectures
- Efficient training and inference techniques
- Contribute to speech-to-speech, speech recognition, speech translation, text-to-speech, and audio generation systems.
Inference Optimization
- Build highly optimized inference pipelines for mobile and embedded hardware.
- Improve performance across CPUs, GPUs, NPUs, and other acceleration hardware.
- Optimize:
- Memory utilization
- Operator execution
- Kernel performance
- Scheduling strategies
- Caching mechanisms
- End-to-end system latency
- Integrate models with production runtimes and deployment frameworks.
Performance Evaluation
- Develop rigorous benchmarking methodologies for edge AI systems.
- Measure and improve:
- Real-time factor
- End-to-end latency
- Time-to-first-audio
- Model size
- Peak memory consumption
- Power efficiency
- Thermal behavior
- Speech quality and accuracy
- Validate performance directly on target devices rather than relying solely on simulator environments.
Cross-Functional Collaboration
- Partner with machine learning researchers, mobile engineers, and systems engineers to bring research into production.
- Translate research prototypes into scalable products and customer-facing technologies.
- Communicate findings through internal documentation, technical publications, conference papers, and open-source contributions where appropriate.
Required Qualifications
- Master's degree, Ph.D., or equivalent industry experience in:
- Machine Learning
- Speech Processing
- Computer Science
- Computer Engineering
- Efficient Deep Learning
- Related quantitative disciplines
- Demonstrated expertise in model compression, effective inference, or edge AI through research publications, production systems, or both.
- Strong understanding of one or more of the following:
- Quantization
- Knowledge distillation
- Pruning
- Low-rank methods
- Efficient neural architectures
- Hardware-aware optimization
- Strong software engineering skills with:
- PyTorch or JAX
- C/C++ or equivalent systems-level programming experience
- Experience optimizing neural networks for resource-constrained hardware.
- Practical knowledge of:
- Mobile CPUs
- GPUs
- NPUs
- Memory systems
- Numerical precision tradeoffs
- Ability to make informed tradeoffs between model quality, latency, memory footprint, power consumption, and deployment portability.
- Professional proficiency in English.
- Ability to thrive in a fast-moving, research-driven environment.
Preferred Qualifications
- Experience working with:
- Automatic Speech Recognition (ASR)
- Text-to-Speech (TTS)
- Speech Translation
- Speech-to-Speech Models
- Audio-Language Models
- Neural Audio Codecs
- Familiarity with deployment frameworks such as:
- Core ML
- ExecuTorch
- ONNX Runtime
- LiteRT / TensorFlow Lite
- TensorRT
- llama.cpp
- Experience with acceleration technologies including:
- Metal
- Vulkan
- CUDA
- QNN
- XNNPACK
- Custom kernels and operator fusion
- Knowledge of:
- INT8 quantization
- INT4 quantization
- Quantization-aware training
- Mixed-precision execution
- Hardware-aware model design
- Experience building streaming and low-latency audio systems.
- Experience deploying machine learning models to:
- Mobile devices
- Wearables
- Embedded systems
- Experience training, distilling, or evaluating large-scale foundation models using distributed GPU infrastructure.
- Previous experience in industrial research labs, AI startups, or leading technology companies.
- Publication record at relevant conferences or meaningful open-source contributions
📌 Edge AI Researcher (Speech & Audio Models) (India)
🏢 Computer Futures
📍 India