Research Engineer - Egocentric Perception Models (Bengaluru)

Research Engineer - Egocentric Perception Models (Bengaluru)

15 Aug
|
Humyn Labs
|
Bengaluru

15 Aug

Humyn Labs

Bengaluru

About Humyn Labs

Humyn Labs builds the intelligence layer for physical-world AI — systems that perceive, reason, and act in real environments. Our work sits at the intersection of egocentric video understanding, embodied AI, robotics perception, and voice-driven interaction. We move fast, obsess over data quality, and ship at scale.

Humyn Labs converts human action - across sound, sight, movement, and touch - into high-quality multi-modal data signals for physical AI. Operating across 20+ countries in India, southeast Asia, Latin America, and the Middle East: the real-world environments where physical AI deploys, not the labs where it is built.

Our data isn't just collected; it's evaluated, defended, and production-ready. Because before AI can be trusted, its training data must be.

THE CHANCE

Our auto-annotation stack currently runs on public off-the-shelf models — for head pose, hand keypoints, depth, and object detection and segmentation. Every one of them was trained on data that does not look like ours, and every one of them has a documented failure mode on our captures: wide-FOV fisheye rigs, gloved hands, heavy occlusion, motion blur, ego-motion, real workshops instead of clean scenes.

We have measured those failures precisely. We are hiring you to remove them — by training our own models, on our own data, and putting them in production behind a gate. This is not a research role adjacent to the pipeline. You own the model layer of the pipeline.

WHAT YOU'LL OWN

- Metric 3D hand pose. Build the in-house hand model that is stable in the world frame, and decide the representation (MANO topology, 6D vs axis-angle) rather than inheriting it.
- Egocentric pose that survives the camera.



We have narrowed the suspects (rectification residual, FOV, shutter type) and currently route between several public SLAM / VIO systems per clip. Replace routing-by-heuristic with a model that makes any capture device deliverable — including learned despiking and drift correction.
- Ground truth where we have none. Today we validate pose without ground truth (loop-closure drift, static-window jitter). Stand up real GT — fiducial-based (AprilTag / ArUco), gravity-aligned via IMU — under a hard constraint: our subjects are real workers doing their own jobs, so any protocol must add zero operator burden.
- Our own depth model. Public depth models do not survive our rigs. We need a depth model of our own: metric, stereo-native, valid across the full frame on wide FOV, and cheap enough to run on every video hour we deliver. Own the training data, the architecture call, and the accuracy-versus-throughput trade.
- Learned QC instead of thresholds. Our gates are hand-tuned numbers: epipolar residual, depth QC metrics, MCAP QA checks. Train error-detection models that flag a bad label before it reaches a customer, per clip, and quantify their catch rate against known-bad deliveries.
- Training and serving, both. Train on AWS and ship on AWS: batch GPU orchestration, throughput tuning, daily delivery SLAs.



Cost per labeled video hour is one of your numbers.

WHAT WE'RE LOOKING FOR

- MS / PhD in computer vision, robotics or ML — or equivalent published research output.
- You have trained a perception model that beat an off-the-shelf baseline on a real domain and shipped it. Data curation, loss design and eval decisions were yours.
- Deep hands-on work in at least two of: 3D hand / human pose reconstruction, stereo or monocular depth, VIO / SLAM, open-vocabulary detection and segmentation, video VLMs.
- Multi-view geometry is fluent, not looked up: rectification, epipolar residual, Q matrix and depth sign conventions, triangulation, world vs camera frame.
- You can read a calibration report and say whether the problem is the camera or the model — and be right.
- Strong Python and PyTorch; you own training and inference code, not notebooks.
- Production inference experience on AWS: containerization, batch pipelines, orchestration.
- 4+ years on problems in this space.

NICE TO HAVE

- Publications in egocentric vision, 3D hand pose, VLA / VLM models, or robot learning.
- IMU-synced multimodal data, gravity alignment, camera–IMU extrinsics.
- Robotics data formats: MCAP / Foxglove, ROS, LeRobot, RLDS.
- Worked with Ego4D, EgoExo4D, Open X-Embodiment or comparable large egocentric corpora.
- Open-source contributions to vision or robotics projects.

FIRST 60 DAYS

- Day 30. Reproduce our label QC numbers end to end and tell us which of the four label types is costing us the most, with evidence.
- Day 60. One in-house model beating its public baseline on our eval set, on that label type.

📌 Research Engineer - Egocentric Perception Models (Bengaluru)
🏢 Humyn Labs
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: research engineer - egocentric perception models (bengaluru) / bengaluru