Task Development Engineer

METR

Completely RemoteContractInformation Technology
Posted Today

Job description

About the Company

METR is a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment. We are a mission-driven organization focused on producing high-quality, trustworthy science to help policymakers and civil society understand AI risks.

Responsibilities

  • Developing difficult, novel tasks for models that remain challenging as model time horizons grow
  • Performing quality assurance for existing tasks to ensure they are solvable and correctly specified
  • Baselining and scoring tasks within specific domains of expertise
  • Improving task development infrastructure and workflows to increase efficiency and quality

Requirements

  • Several years of software engineering experience working on complex projects and codebases
  • Experience building hard, ideally agent-based, AI evaluations (e.g., RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA)
  • High attention to detail regarding misspecifications and ambiguity

Preferred Qualifications

  • Experience using the Inspect framework
  • Prior experience with METR infrastructure, specifically Hawk
  • Familiarity with the methodology behind Time Horizons work

Skills & tools

PythonAILLM

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 01Software engineering experience
  2. 02AI evaluation experience
  3. 03High attention to detail
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches