31 Jul
|
Zoom Communications
|
Gurugram
31 Jul
Zoom Communications
Gurugram
Role Overview
Computer Vision Lead will head the vision and multimodal AI charter behind our next-generation sports AI products. This is a hands-on leadership role: you will own the technical direction of our real-time sports understanding stack, set the architecture and research agenda, and build and mentor a high-performing team of Computer Vision and AI Engineers. You will be accountable for taking models from research to reliable, low-latency production systems experienced by millions of fans.
Key Responsibility Areas
1.
Technical
Leadership & Team Building
● Own the technical vision, architecture, and roadmap for Computer Vision and multimodal AI across the product portfolio.
● Lead, mentor, and grow a team of Computer Vision and AI Engineers; drive hiring, onboarding, and capability development.
● Set engineering standards for experimentation, code quality, documentation, reproducibility, and model governance.
● Translate product goals into technical milestones, effort estimates, and delivery plans; own execution end to end.
● Make build-vs-buy and architecture trade-off decisions across model, infrastructure, and vendor choices.
- AI Product Development & Language Intelligence
● Define the strategy for AI-powered language products that enhance sports content creation and fan engagement.
● Lead development of automated live commentary systems using Large Language Models (LLMs), multimodal AI, and speech technologies.
● Architect intelligent pipelines that fuse vision, audio, and contextual match data into real-time insights and narratives.
● Evaluate and adopt emerging AI architectures to keep product capabilities ahead of the market. 3.
Computer
Vision & Model Development
● Direct the design, training,
and optimization of computer vision models for live sports analytics.
● Guide algorithm development for image and video understanding, including player, ball, and object tracking.
● Set architecture direction across CNNs, Vision Transformers (ViTs), YOLO, Faster R-CNN, Mask R-CNN, and vision-language models.
● Oversee object detection, image classification, segmentation, pose estimation, OCR, facial recognition, and event detection workstreams.
● Establish data strategy: annotation pipelines, augmentation, dataset quality, and data engineering best practices.
● Define evaluation frameworks and benchmarks for accuracy, latency, and robustness; drive continuous improvement against them.
- Deployment, Performance Optimization & Production Engineering
● Own the deployment architecture for models across cloud, edge devices, and GPU-enabled environments.
● Architect scalable inference pipelines with high throughput and low latency for real-time applications.
● Drive deployment practices using Docker, ONNX, TensorRT, FastAPI, Kubernetes, and edge GPU platforms.
● Ensure robust integration into production systems through APIs, microservices, and up-to-date software engineering practices.
● Establish monitoring, observability, and MLOps practices for production reliability and cost efficiency.
- Innovation, Collaboration & Research
● Partner with Product, Design, Data Science,
and Software Engineering leadership to shape the product roadmap.
● Lead applied research in Computer Vision, Multimodal AI, Generative AI, and Sports Analytics; identify what is worth productionizing.
● Represent the team in cross-functional and leadership forums; communicate technical trade-offs to non-technical stakeholders.
● Champion engineering excellence through architecture reviews, design reviews, and code reviews. Ideal candidate will have:
● 8–12 years of hands-on experience in Computer Vision, Machine Learning, or Applied AI, including 2+ years leading or mentoring engineering teams.
● B.E./B.Tech./M.Tech. in Computer Science, Artificial Intelligence, Electronics, Robotics, or a related discipline; PhD is a plus.
● Proven track record of shipping computer vision systems to production at scale, ideally in real-time or low-latency environments.
● Deep fundamentals in Machine Learning, Deep Learning, Computer Vision, NLP, and Multimodal AI.
● Expertise in Python with strong software engineering and system design skills.
● Deep hands-on experience with PyTorch and/or TensorFlow.
● Strong command of transformer architectures and vision-language models.
● Practical, production-grade experience across:
○ Object Detection
○ Pose Estimation
○ Player & Object Tracking
○ Image Classification
○ Image Segmentation
○ OCR
○ Automatic Speech Recognition (ASR)
● Experience with model optimization and deployment (ONNX, TensorRT, quantization, edge GPU) and MLOps tooling.
● Ability to balance research ambition with delivery discipline, and to communicate clearly across technical and business audiences.
📌 Computer Vision Lead (Gurugram)
🏢 Zoom Communications
📍 Gurugram