04 Sep
|
Beyond Border Consultants
|
India
04 Sep
Beyond Border Consultants
India
About the Opportunity
A fast-scaling AI data annotation and model evaluation startup operating at the intersection of Generative AI and Human-in-the-Loop Systems, we enable global tech firms to build safer, more reliable LLMs through precision evaluation, RLHF alignment, and structured data curation. Our team works directly with top AI labs and enterprise clients to design scalable feedback loops, refine model outputs, and train reinforcement learning signals that drive human-aligned AI behavior.
Role & Responsibilities
Design and execute human evaluation frameworks to assess LLM outputs across safety, helpfulness, truthfulness, and fluency dimensions.
Lead RLHF (Reinforcement Learning from Human Feedback) annotation campaigns—curating preference pairs, ranking responses, and labeling reward signals.
Train and manage annotation teams on complex AI evaluation guidelines; QA check outputs for consistency, bias, and edge-case coverage.
Collaborate with ML engineers to translate qualitative judgments into numerical reward models and calibration datasets.
Build evaluation benchmarks and scoring rubrics tailored to domain-specific use cases (e.g., coding assistants, chatbots, legal/medical AI).
Document workflow improvements and contribute to annotation platform tooling enhancements for higher throughput and quality.
Skills & Qualifications
Must-Have
Python
LLM evaluation frameworks (e.g., HELM, BIG-bench, MT-Bench)
Annotation platforms (e.g., Label Studio, Doccano, Scale AI, Labelbox)
RLHF workflows (preference ranking, reward modeling, trajectory sampling)
Data labeling QA & inter-annotator agreement (Kappa, Krippendorff’s Alpha)
Prompt engineering for evaluation tasks
Structured data annotation (JSON, CSV, classification, NER, entity linking)
Model output analysis (bias detection, hallucination scoring, toxicity flags)
Preferred
Experience with Hugging Face Transformers or LangChain pipelines
Familiarity with RLHF training loops in frameworks like TRL or DeepSpeed
Background in linguistics, NLP, or AI ethics
Perks & Culture Highlights
On-site team-oriented workspace in a modern tech hub with full-stack AI team access
Direct exposure to cutting-edge LLM evaluation challenges and global AI clients
Chance to shape industry standards in human-aligned AI through real-world impact
📌 Senior Ai Evaluation & Rlhf Specialist Ai Data Annotation New Delhi (India)
🏢 Beyond Border Consultants
📍 India