29 Aug
|
GET On Tarmack
|
Hyderabad
29 Aug
GET On Tarmack
Hyderabad
Employment Type: Contractor
Start Date: Immediately
Engagement Length: 8 weeks
Years of Experience: 3+ years of experience
Commitment Required: Full-time. 40 hours per week with at least 4 hours PST overlap
Pay Rate: (Average Handling Time per Task: 3 hours) Pay per task, $30/task, effective ~$10/hr
About the Client
Our client is one of the worlds fastest-growing AI companies, accelerating the advancement
and deployment of powerful AI systems.
It helps customers in two ways: Working with the worlds leading AI labs to advance frontier
model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality,
STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve
mission-critical priorities for companies.
About the Role
We are building an evaluation benchmark for frontier AI browsing agents. Your job is to design
research problems that a state-of-the-art AI cannot solve, even with full web access and multiple
attempts. This is not a subject matter expert role, nor is it a content-writing role. It is
investigative research.
You will start from a verifiable fact, work backwards to construct a question that makes that fact
extremely hard to locate, and then prove your work with a complete, auditable evidence trail.
What You Will Produce
- A natural-language research question with a short, stable, objectively verifiable answer
- Clues,
each independently checkable, spanning multiple fact types - including dates,
people, places, organisations, works, events, records, and quantities - with specific
constraints
- A validation record showing the obvious searches you ran and the results they
returned
Minimum Qualifications
- Demonstrated open-web research ability, including locating primary records and
navigating government and institutional databases, archives, registries, and PDF
documents
- Precision with sourcing. You cite exact pages, tables, and sectionsnot just
homepages
- Comfort researching unfamiliar subjects from scratch
- Native or near-native written English
- High tolerance for structured documentation. The evidence trail is the majority of the
work
- Experience with LLM evaluation, red-teaming, or benchmark construction
- Experience in one or more of the following domains:
1. Reference librarianship, archival research, or special collections
2. Investigative journalism or qualified fact-checking
3. OSINT, due diligence, KYC, or investigative research
4. Patent, prior-art, or legal-discovery search
5. Genealogy and records research
6. Competitive quizzing or puzzle-hunt construction
Nice to Have
- Experience with LLM evaluation, red-teaming, or benchmark construction
- Familiarity with JSON and structured data delivery formats
📌 Web Research Specialist - LLM Evaluation (Hyderabad)
🏢 GET On Tarmack
📍 Hyderabad