Task Development Engineer

Posted 12hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Task Development Engineer creating difficult AI evaluation tasks for METR, a nonprofit assessing AI capabilities and risks. Improving task quality, scoring, and evaluation infrastructure.

Responsibilities:

  • Develop difficult, novel tasks for AI models that remain challenging as model time horizons grow
  • Perform quality assurance for existing tasks, verifying solvability and that models receive only the necessary information
  • Baseline tasks within areas of expertise when helpful
  • Score task completions from AIs or human baseliners
  • Improve task development infrastructure and workflows
  • Contribute to evaluations of frontier AI systems supporting METR’s Time Horizons methodology

Requirements:

  • Several years of experience working on complex software engineering projects and codebases
  • Experience building hard, ideally agent-based, AI evaluations
  • Experience with evaluations such as RE-Bench, HCAST, SWE-bench Verified, Cybench, or GPQA
  • Ideally experience using the Inspect framework
  • High attention to detail, including spotting misspecifications and ambiguity
  • Familiarity with METR infrastructure, including Hawk, is a plus
  • Familiarity with the methodology behind METR’s Time Horizons work is a plus
  • Availability for 20–40 hours per week
  • Minimum of 1 hour of overlap with the Pacific Coast Time workday

Benefits:

  • Flexible schedule determined by you
  • Remote work worldwide
  • Flexible working hours of 20–40 hours per week
  • Individuals who contribute >80 hours will be acknowledged in the final research output (if desired)