Company: WillWare Technologies
Role: Freelance- LLM -Engineering Expert
Experience: 10yrs
Work Mode: Remote
Job Details:
- Commitments Required: 40 hours per week with 4 hours of overlap with PST.
- Engagement type: Contractor
- Engagement Length: upto 24 weeks
- Rate Range: $500 Per Task
Role Overview
We are seeking experienced AI Evaluation Engineers (Engineering Simulation & Design) to author and validate "model-breaking," simulation-based engineering design problems to train and evaluate state-of-the-art AI agents. Operating across major engineering disciplines—including Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics—you will create complex, multi-constraint tasks where AI agents must interpret requirements, navigate trade-offs, configure open-source simulation tools, diagnose failures, and iterate toward valid solutions. You will analyze agent execution logs, expose systemic reasoning gaps, and build automated, objective graders to elevate frontier model performance.
Job Requirements:
- Education & Expertise:
Master’s degree or PhD in Electrical, Mechanical, Aerospace, with 10+ years of hands-on engineering design experience.
- Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package (e.g., ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, Gmsh) combined with strong Python scripting skills.
- AI Evaluation & Failure Diagnostics: Hands-on experience with contemporary LLMs/coding agents and evaluation concepts (pass@k, failure-mode analysis, nondeterministic behavior), with the ability to audit trajectory logs and isolate core reasoning/tool-use failures.
- Domain Rigor & Precision: Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
- Availability & Commitment: Talent must have weekend on-call availability (part-time engagement is acceptable).
- Techni
📌 Freelance- LLM (India)
🏢 Turing
📍 India