Company: Evalixa AI Private Limited
Engagement Type: Freelance / Independent Contractor
Work Mode: Remote
Location: India
Schedule: Flexible and project-based To Apply: Send your resume and relevant details to
[email protected]
About Evalixa AI
Evalixa AI Private Limited is an AI benchmarking and evaluation company focused on developing challenging benchmarks, high-quality datasets, and reliable evaluation systems for frontier AI models.
We work on real-world technical environments designed to evaluate how advanced AI systems perform across machine learning, data analysis, software engineering, reasoning, and complex problem-solving tasks.
Role Overview
We are looking for experienced Data Analysts for MLE-Bench to contribute to benchmark-driven evaluation projects focused on real-world machine learning systems.
In this role, you’ll work hands-on with production-like datasets, ML outputs, evaluation metrics, and analytical workflows to evaluate, diagnose, and understand the performance of advanced AI systems.
The ideal candidate is comfortable working at the intersection of data analysis and machine learning, with strong analytical skills and the ability to work with real-world datasets and ML evaluation workflows.
What You’ll Work On
- Analyze structured and unstructured datasets generated from ML training, inference, and evaluation pipelines.
- Define, calculate, and validate metrics used to evaluate model performance and behaviour.
- Investigate data distributions, model outputs, failure modes, anomalies, and edge cases relevant to benchmark tasks.
- Write and run Python and SQL code for data analysis and evaluation workflows.
- Clean, transform, and validate complex datasets.
- Validate data quality, consistency, completeness, and correctness across datasets and experiments.
- Analyze and compare model outputs to identify meaningful performance differences.
- Create reproducible and well-documented analytical workflows.
- Evaluate AI-generated analytical solutions for correctness and completeness.
- Identify incorrect assumptions, analytical errors, and misleading conclusions.
- Review datasets, queries, metrics, code, and analytical results.
- Collaborate with ML engineers and researchers to develop challenging, real-world MLE-Bench evaluation scenarios.
- Requirements
- 4+ years of experience as a Data Analyst, Analytics Engineer, or similar analytical role.
- Strong proficiency in Python for data analysis.
- Solid experience with SQL and relational datasets.
- Experience with Python data libraries such as Pandas, NumPy, or similar tools.
- Experience analyzing ML outputs and evaluation metrics.
- Solid understanding of statistics and analytical reasoning.
- Ability to work with large and complex datasets.
- Experience cleaning, transforming, validating, and analyzing real-world data.
- Ability to identify patterns, anomalies, inconsistencies, and edge cases.
- Strong debugging and problem-solving skills.
- Ability to write clean, readable, and reproducible analytical code.
- Good written and verbal English communication skills.
- Preferred Qualifications
- Good understanding of machine learning fundamentals.
- Familiarity with model training, inference, and evaluation workflows.
- Experience with Jupyter Notebook or similar analytical environments.
- Familiarity with PyTorch, TensorFlow, Scikit-learn, or similar ML frameworks.
- Experience with data visualization and analytical reporting.
- Familiarity with Git and software development workflows.
- Experience working with CSV, JSON, Parquet, APIs, and other common data formats.
- Familiarity with Linux and command-line environments.
- Experience with AI benchmarking, LLM evaluation, or AI agents is an advantage.
What We Look For
We’re looking for analysts who enjoy going deeper into data and understanding why a model or system behaves the way it does.
You should be comfortable receiving an unfamiliar dataset, understanding the problem, exploring the data, writing Python or SQL, validating metrics, identifying unusual behaviour, and reaching conclusions that can be supported by the data.
Attention to detail is important. MLE-Bench work may involve finding subtle errors, edge cases, data inconsistencies, or incorrect conclusions that are not immediately obvious.
Freelancing With Evalixa AI
- Fully remote work.
- Flexible and project-based engagement.
- Work on challenging MLE-Bench and AI evaluation projects.
- Work with real-world ML datasets, model outputs, and evaluation pipelines.
- Opportunity to contribute to the evaluation of advanced AI systems.
- Collaborate with engineers and researchers on technically challenging benchmark problems.
Offer Details
- Engagement Type: Freelance / Independent Contractor
- Work Mode: Remote
- Location: India
- Schedule: Flexible and project-based
- Duration: Based on project requirements
- Work Allocation: Based on project availability, expertise, performance, and quality
- Compensation: Project/task-specific compensation will be communicated before assignment.
Evaluation Process
Selected candidates may complete a technical evaluation covering:
- Python for data analysis
- SQL
- Statistics and analytical reasoning
- Data cleaning and transformation
- ML outputs and evaluation metrics
- Data quality validation
- Identifying anomalies and edge cases
- Practical data analysis and problem-solving
- Candidates who successfully complete the evaluation may be onboarded as Freelance MLE-Bench Data Analysts for relevant Evalixa AI projects.
📌 Data Analyst - MLE Bench (India)
🏢 Evalixa AI
📍 India