06 Aug
|
Deccan AI Experts
|
Maharashtra
06 Aug
Deccan AI Experts
Maharashtra
About Us
Deccan AI Experts is a pioneering AI company founded by IIT Bombay and IIM Ahmedabad alumni, with a strong founding team from IITs, NITs, and BITS. We specialize in high-quality human-curated data, AI-first operations, and advanced AI evaluation systems. About the Role
We are seeking a Reviewer Calibration Lead (Freelancer) to support advanced AI evaluation initiatives focused on reviewer calibration, annotation quality, evaluation consistency, audit governance, and AI benchmarking.
In this role, you will lead reviewer calibration activities, develop gold-standard evaluation datasets, measure inter-rater agreement, conduct quality audits, and coach reviewers to improve consistency and accuracy. Your expertise will help improve AI systems by ensuring high-quality human evaluations and standardized review processes.
This position is ideal for professionals with experience in AI data annotation, quality assurance, reviewer management, operations excellence, trust & safety, content moderation, or evaluation program management. Responsibilities
Create deliverables addressing real-world reviewer calibration and quality assurance scenarios.
Annotate and evaluate AI-generated reviewer assessments, audit reports, quality analyses, and calibration documentation.
Assess AI outputs for evaluation consistency, reviewer accuracy, adherence to guidelines, and quality standards.
Develop and maintain gold-standard (gold-set) datasets to establish benchmarking criteria, reference answers, and evaluation standards for reviewer performance.
Measure and improve inter-rater reliability (IRR) using agreement metrics such as Cohen's Kappa, Fleiss' Kappa, Krippendorff's Alpha, or percentage agreement, and identify areas requiring calibration.
Plan and conduct calibration sessions to align reviewers on evaluation guidelines, scoring methodologies, edge cases, and quality expectations.
Develop and maintain error taxonomies by categorizing reviewer errors, annotation inconsistencies, policy interpretation issues, and quality defects.
Perform audit sampling using statistical or risk-based sampling techniques to evaluate reviewer performance and quality compliance.
Deliver reviewer coaching by providing constructive feedback, identifying skill gaps, recommending corrective actions, and improving reviewer accuracy and consistency.
Analyze reviewer performance trends, calibration outcomes, and audit findings to identify continuous improvement opportunities.
Provide structured feedback to improve AI evaluation frameworks, reviewer guidelines, and quality governance.
Review peer-developed deliverables to maintain quality and consistency standards. Requirements
Bachelor's degree in Business Administration, Operations Management, Quality Management, Computer Science, Data Science, Psychology, Linguistics, or a related field.
5+ years of hands-on experience in Quality Assurance, AI Data Annotation, Reviewer Operations, Trust & Safety, Content Moderation, Audit Management, or AI Evaluation.
Creating and maintaining gold-set datasets for reviewer benchmarking and evaluation consistency.
Measuring and improving inter-rater reliability (IRR) using recognized statistical agreement methods.
Planning and facilitating calibration sessions to standardize reviewer interpretations and scoring.
Developing error taxonomies to classify annotation errors, reviewer inconsistencies, and quality issues.
Conducting audit sampling to evaluate reviewer performance and compliance with quality standards.
Providing reviewer coaching, mentoring, performance feedback, and continuous quality improvement.
Proficiency in Microsoft Excel, Google Sheets, Power BI, Tableau, or similar reporting tools.
Familiarity with annotation platforms, AI evaluation tools, or reviewer management systems is preferred.
Robust analytical thinking, leadership skills, coaching ability, and exceptional attention to detail.
Excellent written and verbal English communication skills.
Ability to critically evaluate reviewer performance and AI-generated outputs for consistency, quality, and policy adherence.
Ability to work independently in a remote, fast-paced environment.
Preferred
Qualifications
Experience leading reviewer teams, annotation operations, trust & safety programs, or AI benchmarking initiatives.
Experience designing quality scorecards, reviewer certification programs, onboarding curricula, or evaluation rubrics.
Familiarity with RLHF, human-in-the-loop evaluation, LLM benchmarking, content moderation, or AI safety evaluation.
Knowledge of SQL, Python, or analytics tools for quality reporting and reviewer performance analysis is an advantage.
Professional certifications such as Lean Six Sigma, ASQ Certified Quality Auditor (CQA), COPC Certification, or quality management credentials are highly desirable.
Why Join
Us
Competitive hourly pay: ₹1,500 - 2,200/hour
Fully remote with flexible working hours.
Opportunity to contribute to cutting-edge AI quality assurance and benchmarking initiatives.
Exposure to advanced AI systems focused on human evaluation, reviewer calibration, and operational excellence.
Flexible project-based opportunities with global teams.
Work on next-generation AI solutions supporting reliable, high-quality, and scalable AI evaluation systems.
📌 Reviewer Calibration Lead (Freelancer) (Maharashtra)
🏢 Deccan AI Experts
📍 Maharashtra