Expert Contributor - Frontier Benchmarks (Task-Based) (Hyderabad)

Expert Contributor - Frontier Benchmarks (Task-Based) (Hyderabad)

06 Aug
|
Evalixa AI
|
Hyderabad

06 Aug

Evalixa AI

Hyderabad

Expert Contributor – Frontier Benchmarks (Task-Based)

Company: Evalixa AI Private Limited

Engagement Type: Freelance Expert Contributor

Work Mode: Remote

Location: India

Schedule: Flexible and project-based To apply, email your details to [email protected]

About Evalixa AI

Evalixa AI Private Limited is an AI benchmarking and evaluation company focused on developing high-quality datasets, challenging evaluation tasks, and reliable assessment frameworks for frontier AI models.

We design complex evaluations that identify model limitations that standard benchmarks may not capture. Our work covers reasoning, coding, instruction-following, reliability, problem-solving, and long-horizon execution.

About the Prospect

We are looking for Freelance Expert Contributors to support the research, development, and evaluation of frontier AI benchmarks.

Contributors will create, review, validate, and improve complex benchmarking tasks designed to measure the capabilities and limitations of advanced AI models.

This is an independent, project-based freelance opportunity. Work will be assigned according to project availability, contributor expertise, performance, and quality requirements.

Key Responsibilities

- Design original and challenging benchmarks and evaluation tasks for frontier AI models.
- Identify capability gaps, recurring failure modes, shortcut behaviours, and under-evaluated problem areas.
- Develop accurate task specifications, deterministic test environments, hidden tests, scoring criteria, rubrics, and metadata.
- Evaluate model outputs and identify reasoning, coding, reliability, and instruction-following failures.
- Ensure that tasks are original, difficult, solvable, reproducible, and aligned with the intended evaluation criteria.
- Review task submissions and provide clear, actionable feedback.
- Revise tasks based on quality-review feedback and project requirements.
- Maintain consistency, accuracy, and confidentiality across all assigned work.
- Follow project documentation, submission requirements, and communicated deadlines.
- Participate in required training sessions, assessments, or project meetings.

Required Qualifications

- Strong academic, research, or industry background in computer science, machine learning, artificial intelligence, software engineering, data science, or a related field.
- Strong programming and analytical problem-solving skills.
- Proficiency in Python or another commonly used programming language.
- Familiarity with large language models, AI evaluation, benchmark datasets, or model testing.
- Ability to convert open-ended research questions into precise tasks and measurable evaluation criteria.
- Strong attention to detail when assessing model behaviour,



task validity, test quality, and evaluation integrity.
- Ability to understand technical documentation and follow detailed project instructions.
- Strong written communication and technical documentation skills.
- Ability to work independently and complete assignments within the required timelines.

Preferred Qualifications

- Master’s degree or Ph.D. in computer science, machine learning, NLP, data science, or a related discipline. Candidates with equivalent research or industry experience will also be considered.
- Experience developing benchmarks for coding models, reasoning systems, tool-using models, computer-use agents, or long-horizon tasks.
- Familiarity with sandboxed execution, CI systems, containers, hidden tests, rubrics, and structured evaluation pipelines.
- Experience creating datasets, coding tasks, evaluation environments, or automated test systems.
- Research publications, benchmark contributions, technical reports, GitHub projects, or relevant open-source work.
- Understanding of dataset governance, privacy, safety, licensing, contamination, and responsible AI evaluation.

Engagement Benefits

- Flexible remote working schedule.
- Practical exposure to frontier AI benchmarking and model evaluation.
- Opportunities to work across coding, machine learning, security, scientific computing, system administration, and other technical domains.
- Access to project documentation and training for selected contributors.
- Opportunities to qualify for additional projects or reviewer responsibilities based on performance.
- Experience contributing to technically challenging AI evaluation projects.

Selection Process

1. Application review: We review the candidate’s resume, academic background, technical skills, portfolio, and relevant experience.
2. Initial screening: Shortlisted applicants attend a discussion about their technical background, availability, and understanding of AI benchmarking.
3. Trial assessment: Candidates complete a limited technical assessment or trial task.
4. Quality review: The submission is evaluated for correctness, originality, technical quality, reasoning, and compliance with project requirements.
5. Onboarding: Successful candidates complete the required documentation, verification, confidentiality process, and project training.





Successfully completing the selection process does not guarantee continuous assignments or a minimum volume of work. Task availability depends on active projects, client requirements, contributor performance, and business needs.

Application Requirements

Applicants should submit

- An updated resume or CV.
- LinkedIn profile.
- GitHub profile, technical portfolio, publication list, or relevant work samples, where available.
- Educational and experience details.
- Availability and computer specifications when requested.

Freelance Relationship Selected contributors will work as independent freelancers and not as permanent employees of Evalixa AI Private Limited.

This engagement does not guarantee fixed working hours, continuous assignments, internships, placements, permanent employment, or employee benefits.

Ending the Freelance Engagement

Evalixa AI may issue a seven-day written notice if a contributor’s quality, productivity, activity, reliability, or overall performance does not meet project requirements.

Reasons may include

- Consistently poor work quality.
- Failure to meet project requirements or quality standards.
- Low output or repeated failure to complete assigned tasks.
- Continued inactivity or lack of communication.
- Repeatedly missing deadlines, meetings, or required revisions.
- Failure to follow project instructions or internal procedures.
- Unprofessional or unreliable participation.

The contributor will be removed from the engagement after the seven-day notice period. During this period, the contributor may be required to complete pending work, revisions, or handover responsibilities.

Applicant Privacy

Evalixa AI will use information submitted by applicants for recruitment, candidate evaluation, identity verification, background screening, fraud prevention, project administration, security, and legal compliance.

Applicant information will be accessible only to authorised personnel and relevant service providers or clients where necessary for these purposes. Information may be retained to meet legal, security, audit, contractual, or client requirements.

Applicants may request access to, correction of, or deletion of their personal information, subject to applicable legal and record-retention requirements. Applicants should not submit sensitive information unless specifically requested for a legitimate screening, compliance, or onboarding purpose.

Equal Opportunity

Evalixa AI Private Limited follows fair and merit-based selection practices. Selection and continued task allocation depend on technical capability, assessment performance, submission quality, reliability, project availability, and business requirements.

📌 Expert Contributor - Frontier Benchmarks (Task-Based) (Hyderabad)
🏢 Evalixa AI
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: expert contributor - frontier benchmarks (task-based) (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: expert contributor - frontier benchmarks (task-based) (hyderabad) / hyderabad