Task Development Engineer
Posted 12hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Task Development Engineer creating difficult AI evaluation tasks for METR, a nonprofit assessing AI capabilities and risks. Improving task quality, scoring, and evaluation infrastructure.
Responsibilities:
- Develop difficult, novel tasks for AI models that remain challenging as model time horizons grow
- Perform quality assurance for existing tasks, verifying solvability and that models receive only the necessary information
- Baseline tasks within areas of expertise when helpful
- Score task completions from AIs or human baseliners
- Improve task development infrastructure and workflows
- Contribute to evaluations of frontier AI systems supporting METR’s Time Horizons methodology
Requirements:
- Several years of experience working on complex software engineering projects and codebases
- Experience building hard, ideally agent-based, AI evaluations
- Experience with evaluations such as RE-Bench, HCAST, SWE-bench Verified, Cybench, or GPQA
- Ideally experience using the Inspect framework
- High attention to detail, including spotting misspecifications and ambiguity
- Familiarity with METR infrastructure, including Hawk, is a plus
- Familiarity with the methodology behind METR’s Time Horizons work is a plus
- Availability for 20–40 hours per week
- Minimum of 1 hour of overlap with the Pacific Coast Time workday
Benefits:
- Flexible schedule determined by you
- Remote work worldwide
- Flexible working hours of 20–40 hours per week
- Individuals who contribute >80 hours will be acknowledged in the final research output (if desired)



















