AI Safety Expert, English, Dutch

Posted 2hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI safety experts red teaming conversational models for Mercor, which partners with AI labs to train frontier systems. Generating vulnerability data, attack cases, and reproducible safety reports.

Responsibilities:

  • Red team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
  • Apply taxonomies, benchmarks, and playbooks to maintain consistent testing
  • Produce reproducible reports, datasets, and attack cases for customers
  • Identify vulnerabilities that automated tests miss
  • Expand evaluation coverage and reduce production surprises
  • Collaborate with leading researchers and contribute to training and enhancing frontier AI systems

Requirements:

  • Native fluency in English and Dutch
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to probe AI systems adversarially and push systems to breaking points
  • Experience using frameworks or benchmarks for structured testing
  • Ability to explain risks clearly to technical and non-technical stakeholders
  • Adaptability across projects and customers
  • Adversarial ML experience, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction is a nice-to-have
  • Cybersecurity experience, including penetration testing, exploit development, or reverse engineering is a nice-to-have
  • Socio-technical risk experience, including harassment/disinformation probing, abuse analysis, or conversational AI testing is a nice-to-have
  • Creative probing experience in psychology, acting, or writing is a nice-to-have
  • Must be able to work as an independent contractor
  • H1-B and STEM OPT candidates are not supported

Benefits:

  • Fully remote role
  • Flexible schedule; work on your own schedule
  • Weekly payments via Stripe or Wise based on services rendered
  • Projects may be extended, shortened, or concluded early depending on needs and performance
  • Higher-sensitivity projects are optional
  • Clear content guidelines and wellness resources
  • Reasonable accommodations upon request
  • Competitive pay
  • Referral opportunity earning up to $250 per successful referral, subject to limits