16 Sep
|
Axiom Global Technologies
|
India
16 Sep
Axiom Global Technologies
India
Role & responsibilities
Multimodal AI Engineer (voice and vision)
Owns streaming speech in and out of the interview and the real-time computer vision for identity verification and proctoring, including vision-language models.
Must have
- Streaming speech (recognition, synthesis, or both) shipped in production under a latency budget, with real numbers for time to first audio and end-of-utterance latency.
- At least one vision model trained or fine-tuned on data they labeled or curated themselves, evaluated against the incumbent, and shipped.
- Strong PyTorch.
- Audio fundamentals: voice activity detection, endpointing, resampling, and feature extraction.
- Vision fundamentals: face detection and embedding (InsightFace or equivalent), gaze estimation, object detection (Ultralytics or YOLO), and OpenCV.
- Vision-language models in production, or strong working knowledge of prompting and batching them for structured output.
- Model export and optimization for constrained hardware: ONNX Runtime plus at least one of TensorRT, INT8 calibration, or CPU-targeted inference.
- Evaluation discipline: held-out sets, confusion matrices, calibrated thresholds, and honest reporting when a new model is worse.
- Python, Linux, and SQL at a skilled level. Comfortable inside a Kubernetes-deployed service.
- Fluent written and spoken English. Works autonomously with clean, reviewable pull requests.
Nice to have
- Fine-tuning speech recognition on accented or domain audio, or training text-to-speech voices.
- Speaker diarization, active speaker detection, or audio-visual sync.
- Liveness, anti-spoofing, or deepfake and virtual-camera detection.
- 3D reconstruction, depth estimation, or feature matching.
📌 Multimodal AI Engineer (voice and vision) (India)
🏢 Axiom Global Technologies
📍 India