About shift shift is a flexible contribution platform where eligible contributors can record short videos of everyday tasks to help train robotics and AI systems. It records how physical work actually happens and turns it into training-ready data. shift is trusted by 25000+ workers, across 15+ countries and paid out over 15M USD to contributors and annotators. shift has been featured in leading publications like Forbes, BBC, Business Insider, CBS, The Verge, Recent York Post, Entrepreneur, Morning Brew, Gizmodo.
Our mission is simple: make participation in the AI economy accessible, adaptable, and fair.
Join the shift.
About the role
You are the gold-standard authority for shift's egocentric training data. You own the golden sets, SOPs, and calibration system that annotators, reviewers, and an AI-first pipeline are all measured against and you partner with engineering to stand up new annotation workflows in quality control, hand tracking, object identification, and action labelling. As shift scales from thousands to millions of contributors, you build the quality system that lets output grow faster than headcount without letting the 99% accuracy bar slip.
What you'll do
Own the quality system
- Define and maintain gold-standard benchmark ("golden") sets for new and existing tasks, with explicit acceptance criteria and frame-level tolerance thresholds
- Author, version-control, and continuously refine annotation SOPs and rubrics based on disagreement patterns, audit findings, and the evolving client spec
- Design and run Inter-Annotator Agreement measurement (Cohen's/Fleiss' kappa, tolerance-based agreement for temporal boundaries) and produce regular quality reports
- Analyze disagreement trends to root cause and recommend data-driven process improvements
- Build and maintain edge-case libraries and example banks; adjudicate escalated edge cases and turn rulings into guideline updates
- Conduct precision-based QC audits and validate pilot batches before production scaling
Scale the people and the operation
- Stand up the training and certification pipeline that brings new annotators and reviewers to the quality bar quickly and consistently — onboarding, structured calibration sessions, targeted retraining
- Build and lead the reviewer team as it grows; manage performance with clear standards, feedback, and a fair improvement/exit process
- Establish operational metrics and reporting (acceptance/rejection, rework, quality-adjusted productivity) and drive week-over-week improvement
- Run capacity planning across competing annotation demands; allocate reviewer attention to the highest-impact work
- Own quality's contribution to unit economics: improve cost-per-accepted-hour while protecting the bar; inform the in-house vs. vendor mix with quality data
Build the workflows
- Design human-in-the-loop workflows where models pre-label and humans review, correct, and escalate — so throughput grows faster than headcount
- Partner with engineering and product to launch new annotation workflows: quality control (good vs bad tasks), hand tracking (keypoints, left/right attribution, hand-object contact states), object identification (bounding boxes/masks, class ontology, naming consistency across frames), and action labelling (controlled vocabulary, segment boundaries)
- Specify label schemas, annotation-tool requirements, and QC dashboards; translate research and client needs into clear instructions, rubrics, and SLAs
What we're looking for - must have
- 7+ years of hands-on video annotation, including 3+ years (5+ preferred) with egocentric/first-person video for robotics or embodied AI — ideally at or for a leading robotics company or project
- 3+ years in a lead or QA capacity: calibrating annotators, adjudicating disagreements, and owning guidelines that others work to
- Track record standing up 0→1 annotation programs and validating them through pilot to production
- Deep command of quality methodology: IAA frameworks (kappa/alpha), gold-set creation, sampling strategies, tolerance thresholds — you can explain what a 99% (versus 95%) accuracy bar means in practice and design the system that holds it
- Track record authoring (not just following) annotation SOPs and rubrics; able to translate ambiguous specs into precise, actionable guideline language
- Hands-on with 2+ professional annotation tools — CVAT, Labelbox, Encord, V7, Label Studio, or equivalent — including configuring label schemas and QA workflows, not just annotating in them
- Experience integrating model-based annotation into human workflows: auditing auto-generated labels and designing human-in-the-loop review that raises throughput without sacrificing quality
- Outstanding written English — your rubrics train annotators and your labels are commands a model learns from
- Strong analytical skills with Excel/Google Sheets for quality reporting (Python preferred)
- Data-driven, process-oriented mindset with strong ownership; startup-ready — comfortable with ambiguity, zero-to-one builds, and contributing when it matters most
- Working understanding of ML and why annotation quality drives model performance
Nice to have
- Robotics, mechanical/mechatronics, or computer-vision engineering background; tier-1 institution or AI/robotics startup experience
- Experience managing distributed/global annotator workforces or external vendors
- Multi-modal annotation exposure: 3D/depth, joint pose, grasp outcome classification
- Experience training or fine-tuning autolabeling models, or partnering closely with the ML teams that do
- Built annotation tooling or partnered tightly with a tooling team
Why it matters
Your ground truth becomes the benchmark and the model is measured against it. This is the role that sets the standard — and builds the system that holds it at 100,000-contributor scale.
📌 Data annotation lead (India)
🏢 Shift
📍 India