AI Safety Expert, English – Dutch
Posted 6hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI safety red teamer probing conversational AI models for Mercor’s human-data AI training projects. Testing vulnerabilities, generating reproducible datasets, and strengthening frontier AI systems.
Responsibilities:
- Red team conversational AI models and agents
- Test jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Apply taxonomies, benchmarks, and playbooks to maintain consistent testing
- Produce reproducible reports, datasets, and actionable attack cases
- Probe AI outputs involving sensitive topics such as bias, misinformation, and harmful behaviors
- Uncover vulnerabilities automated tests miss
- Expand evaluation coverage and reduce production surprises
- Strengthen customer AI systems through adversarial testing
Requirements:
- Native fluency in English and Dutch required
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
- Ability to probe AI systems adversarially, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Ability to annotate failures, classify vulnerabilities, and flag systemic risks
- Ability to follow taxonomies, benchmarks, and playbooks
- Ability to produce reproducible reports, datasets, and attack cases
- Ability to explain risks clearly to technical and non-technical stakeholders
- Adaptability across projects and customers
- H1-B and STEM OPT candidates are not supported
- Independent contractor engagement
Benefits:
- Fully remote role
- Flexible own schedule
- Weekly payments via Stripe or Wise
- Projects may be extended, shortened, or concluded early depending on needs and performance
- Higher-sensitivity project participation is optional
- Clear guidelines and wellness resources for sensitive content
- Up to $250 referral payment per successful referral
- Competitive pay
- Collaboration with leading researchers
- Reasonable accommodations upon request







