Mercor Research Fellowship – APEX
Posted 21hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Research fellow building and validating APEX benchmarks for Mercor, an AI data and enterprise evaluation company. Designing evaluations, testing frontier models, and publishing results.
Responsibilities:
- Propose and scope a new benchmark or evaluation technique in an under-covered APEX domain or a meaningfully harder version of an existing one
- Design task specifications and grading rubrics with Mercor’s vetted domain experts
- Build and validate benchmarks through task pilots, scoring calibration, and stress-testing for contamination and gameable shortcuts
- Run frontier models against benchmarks and analyze failure modes
- Publish results as a paper, open dataset, APEX leaderboard, or methodology adopted internally
- Partner with Mercor’s research and engineering teams to incorporate findings into APEX’s public benchmark family
- Work directly with the APEX research team and access enterprise evaluation problems from Fortune 500 and frontier-lab partners
Requirements:
- Genuine interest in evaluation as a research discipline
- Background in CS, ML, statistics, or an adjacent field such as measurement, psychometrics, HCI, or social science
- Specific, well-scoped idea for a benchmark or evaluation technique to build
- Comfortable working in a startup environment with fast iteration and less hand-holding than an academic lab
- Able to commit at least 20 hours/week for the duration of the fellowship
- Minimum commitment of 30 hours/week; full-time preferred
- Bonus: experience with agentic evaluation, RL environments, or domain expertise in law, finance, medicine, or a scientific field
- Must submit a specific benchmark or evaluation-methodology proposal
Benefits:
- Unlimited API credits
- Dedicated budget for GPU compute
- Paid expert/human-data time
- Weekly 1:1 mentorship with a member of the APEX research team
- Regular access to the broader research organization
- Access to frontier model APIs
- Access to Mercor’s internal evaluation infrastructure
- Access to real enterprise evaluation problems from Mercor’s customers, where appropriate
- Optional desk in Mercor’s San Francisco office
- Introductions to Mercor’s network of researchers across frontier labs and academia
- Standout fellows considered for a full-time offer on the APEX research team at the end of the fellowship


















