09 Sep
|
TELUS Digital
|
India
09 Sep
TELUS Digital
India
Role Overview
TELUS Digital is sourcing elite AI Evaluation Experts and Benchmark Engineers to support next-generation Recursive Self-Improvement (RSI) pipelines for frontier AI models. This unified role bridges two high-priority project tracks:
1. RSI Trace Diagnostics & Evaluation: Auditing end-to-end multi-turn agentic execution traces, isolating pipeline failure root causes (distinguishing model capability limits from harness bugs, environment drift, or leaky evaluation metrics), and assessing benchmark discrimination.
2. RSI Training Ready-Made Dataset Engineering: Authoring self-contained, hermetic agent task packages (Docker containers, ground-truth reference solutions, pytest verifiers) and generating complete execution trajectory run data across high-complexity technical domains.
Mandatory Criteria Role Type: Freelancer (AI Community)
Location: 100% Remote (United States, Europe, India, and Philippines preferred- global applicants are welcome)
- KEY POOLING NOTICE: This application is strictly for our Global Talent Pool. By applying, you agree to join our vetted candidate roster. Profiles will be activated and contacted as matching project demand is confirmed. Please apply only if you are comfortable joining a pool and exercising patience. Feedback and progress updates may not be provided during the initial pooling phase.
MANDATORY Core Technical Stack: Python, Shell/Bash, Docker Containerization, Claude Code / OpenAI Codex / Model Context Protocol (MCP), Pytest / Evaluation Harnesses. This is an ultra-niche engineering role. Candidates must pass a strict technical gateway. Non-coding prompt engineers, generic QA annotators, or non-technical profiles will be automatically filtered out.
- Language Requirement: English Full Professional Proficiency (C1/C2) or Native/Bilingual Proficiency is mandatory for precise technical reporting, bug attribution notes, and documentation.
- Hands-on Containerization: Proven experience spinning up, configuring,
and debugging local Docker container environments.
- Agentic Tooling: Daily operational proficiency with agentic coding CLI tools (Claude Code, OpenAI Codex, Cursor) and Model Context Protocol (MCP).
- Code & Scripting: Fluent in writing and reading Python (pytest, AST parsing, data processing) and Shell/Bash automation scripts.
Key Responsibilities (Dual Scope) Track 1: Execution Trace Diagnostics & Pipeline Auditing
- Audit full multi-turn agent rollout trajectories to diagnose exact root causes of failure (attributing errors to defective task instructions, metric computation bugs, incorrect reference artifacts, environment drift, or data leakage).
- Evaluate benchmark item discrimination across model capability tiers (ensuring tasks avoid 0% floor or 100% ceiling saturation and separate monotonically).
- Inspect visible selfcheck verification loops, reward functions, and hidden evaluation test suites in sandboxed setups.
Track 2: Agent Task Package & Dataset Engineering
- Author production-grade, hermetic agent task packages containing deterministic Docker setups (Dockerfile + pinned dependencies), explicit task specifications, and automated verifier/pytest suites.
- Generate complete trajectory run data (tool calls, terminal logs, CoT reasoning steps, and final artifacts) for RSI training sets.
- Validate task package reproducibility locally using digest-pinned base images and strict network/environment isolation.
Domain Specialization Requirements (At Least One Required) Candidates must demonstrate deep hands-on expertise in at least one of the following four scenario domains:
1.
AI Models, Software & Agents: SWE-bench/Terminal-Bench tasks, multi-agent tool orchestration, repository refactoring, MCP tool servers, API integration.
2. Physical Sciences & Engineering: Computational physics, robotics/VLA simulation pipelines, circuit/embedded control diagnostics, control theory, mathematical reasoning.
3. Optimization Planning & Control: Operations research, linear programming, multi-agent scheduling, supply chain routing, state-space control logic.
4. Finance & Quantitative Research: Quantitative backtesting, risk model validation, time-series forecasting, market data pipeline debugging, financial RAG.
About Company: TELUS Digital is the customer experience transformation partner to the world's most admired brands. Our diverse team weaves data, technology, and human ingenuity to deliver differentiated customer journeys, drive operational effectiveness, and scale AI solutions with meaningful value and positive impact.
We craft real-world solutions in the moments that matter, from customer acquisition to lifelong loyalty. Enabled by our global reach - spanning 78,000 experts in 33 countries - and deep industry expertise, we help over 600 organizations make the customer experience feel effortless.
Our solutions span Data & AI, Digital Experience & IT, CX Management and Trust & Safety. At the core of our innovation is Fuel iX™, an enterprise-grade generative AI platform that helps clients safely access and optimize leading LLMs to scale their own AI from pilot to production.
Equal Opportunity Employer
At TELUS Digital, we are proud to be an equal opportunity employer and are committed to creating a diverse and inclusive workplace. All aspects of employment, including the decision to hire and promote, are based on applicant's qualifications, merits, competence and performance without regard to any characteristic related to diversity.
📌 RSI Evaluation & Agent Training Data Engineer (India)
🏢 TELUS Digital
📍 India