Member of Technical Staff, Frontier AI
Posted 59mins ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Frontier AI technical professional owning research evaluation, ML-oriented data design, and failure analysis. Improving deployed model and agent performance through evidence-based iteration.
Responsibilities:
- Own research and evaluation initiatives from problem framing through data design, quality calibration, and signal validation
- Define rigorous approaches for determining whether experimental results provide reliable and defensible research signal
- Analyse model and system failures to identify root causes, edge cases, and opportunities for improvement
- Evaluate datasets, experiments, and conclusions against appropriate quality thresholds
- Act as a quality gate when signal strength, data integrity, or supporting evidence is insufficient
- Design ML-oriented data systems, including task definitions, annotation schemas, rubrics, incentives, and supporting pipelines
- Structure data and evaluation workflows around downstream model-performance objectives
- Translate ambiguous real-world behaviour into measurable evaluation frameworks and new data categories
- Identify evaluation or dataset coverage gaps and recommend additional investment or iteration
- Develop quality-assurance processes that maintain consistent research standards
- Investigate model and system behaviour to identify recurring weaknesses and performance limitations
- Iterate on evaluations, datasets, feedback loops, and quality standards
- Use experimental findings to guide improvements in model or agent performance
- Determine when research directions should be expanded, revised, paused, or discontinued based on evidence
- Maintain a systems-level perspective focused on end-to-end AI performance
- Collaborate with researchers, domain experts, operators, and cross-functional teams during project kickoff, calibration, and iteration
- Communicate research findings, trade-offs, limitations, and signal strength to technical and non-technical stakeholders
- Translate research progress into evidence-grounded narratives
- Support alignment between experimental work and real-world system requirements
- Contribute technical judgement in ambiguous, high-impact research environments
Requirements:
- Experienced technical professional
- Strong professional judgement regarding research signal quality
- Experience designing ML-oriented datasets, evaluation frameworks, annotation systems, rubrics, or QA processes
- Ability to translate complex and ambiguous real-world system behaviour into structured research and evaluation opportunities
- Strong ownership mindset and comfort making decisions in uncertain or rapidly evolving environments
- Excellent written and verbal communication skills
- Ability to explain technical trade-offs, limitations, evidence quality, and research findings clearly
- Proven experience working directly with researchers, technical experts, or domain specialists during project calibration and iteration
- Systems-level understanding of model, agent, or AI-system performance
- Experience with reinforcement-learning environments, simulators, or feedback-driven training systems is advantageous
- Experience improving agentic systems or AI systems operating within real-world workflows is beneficial
- Prior work within applied research or production environments with direct impact on deployed systems is advantageous
- Experience designing evaluations for complex or real-world tasks is strongly valued
- Familiarity with expert incentive design or high-stakes technical research programmes is beneficial
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
Benefits:
- Fully remote work
- Full-time engagement
- Remote consulting opportunity
- Compensation of $600,000–$2,000,000/year
















