02 Aug
|
Zoom Communications
|
Gurugram
02 Aug
Zoom Communications
Gurugram
Role Overview: Computer Vision Lead will head the vision and multimodal AI charter behind our next-generation sports AI products. This is a hands-on leadership role: you will own the technical direction of our real-time sports understanding stack, set the architecture and research agenda, and build and mentor a high-performing team of Computer Vision and AI Engineers. You will be accountable for taking models from research to reliable, low-latency production systems experienced by millions of fans.
Key Responsibility Areas: 1.
Technical
Leadership & Team Building ● Own the technical vision, architecture, and roadmap for Computer Vision and multimodal AI across the product portfolio. ● Lead, mentor, and grow a team of Computer Vision and AI Engineers; drive hiring, onboarding, and capability development. ● Set engineering standards for experimentation, code quality, documentation, reproducibility, and model governance. ● Translate product goals into technical milestones, effort estimates, and delivery plans; own execution end to end. ● Make build-vs-buy and architecture trade-off decisions across model, infrastructure, and vendor choices.
- AI Product Development & Language Intelligence ● Define the strategy for AI-powered language products that enhance sports content creation and fan engagement. ● Lead development of automated live commentary systems using Large Language Models (LLMs), multimodal AI, and speech technologies. ● Architect intelligent pipelines that fuse vision, audio, and contextual match data into real-time insights and narratives. ● Evaluate and adopt emerging AI architectures to keep product capabilities ahead of the market. 3.
Computer
Vision & Model Development ● Direct the design, training,
and optimization of computer vision models for live sports analytics. ● Guide algorithm development for image and video understanding, including player, ball, and object tracking. ● Set architecture direction across CNNs, Vision Transformers (ViTs), YOLO, Faster R-CNN, Mask R-CNN, and vision-language models. ● Oversee object detection, image classification, segmentation, pose estimation, OCR, facial recognition, and event detection workstreams. ● Establish data strategy: annotation pipelines, augmentation, dataset quality, and data engineering best practices. ● Define evaluation frameworks and benchmarks for accuracy, latency, and robustness; drive continuous improvement against them.
- Deployment, Performance Optimization & Production Engineering ● Own the deployment architecture for models across cloud, edge devices, and GPU-enabled environments. ● Architect scalable inference pipelines with high throughput and low latency for real-time applications. ● Drive deployment practices using Docker, ONNX, TensorRT, FastAPI, Kubernetes, and edge GPU platforms. ● Ensure robust integration into production systems through APIs, microservices, and modern software engineering practices. ● Establish monitoring, observability, and MLOps practices for production reliability and cost efficiency.
- Innovation, Collaboration & Research ● Partner with Product, Design, Data Science,
and Software Engineering leadership to shape the product roadmap. ● Lead applied research in Computer Vision, Multimodal AI, Generative AI, and Sports Analytics; identify what is worth productionizing. ● Represent the team in cross-functional and leadership forums; communicate technical trade-offs to non-technical stakeholders. ● Champion engineering excellence through architecture reviews, design reviews, and code reviews.
Ideal candidate will have: ● 8–12 years of hands-on experience in Computer Vision, Machine Learning, or Applied AI, including 2+ years leading or mentoring engineering teams. ● B.E./B.Tech./M.Tech. in Computer Science, Artificial Intelligence, Electronics, Robotics, or a related discipline; PhD is a plus. ● Proven track record of shipping computer vision systems to production at scale, ideally in real-time or low-latency environments. ● Deep fundamentals in Machine Learning, Deep Learning, Computer Vision, NLP, and Multimodal AI. ● Expertise in Python with strong software engineering and system design skills. ● Deep hands-on experience with PyTorch and/or TensorFlow. ● Robust command of transformer architectures and vision-language models. ● Practical, production-grade experience across: ○ Object Detection ○ Pose Estimation ○ Player & Object Tracking ○ Image Classification ○ Image Segmentation ○ OCR ○ Automatic Speech Recognition (ASR) ● Experience with model optimization and deployment (ONNX, TensorRT, quantization, edge GPU) and MLOps tooling. ● Ability to balance research ambition with delivery discipline, and to communicate clearly across technical and business audiences.
📌 Computer Vision Lead (Gurugram)
🏢 Zoom Communications
📍 Gurugram