AI Safety Expert, English, Danish
Posted 16hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI red team expert attacking conversational models for Mercor, which partners with AI labs and enterprises to train frontier systems. Generating reproducible vulnerability data, attack cases, and safety evaluations for customers.
Responsibilities:
- Red-team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Apply taxonomies, benchmarks, and playbooks to keep testing consistent
- Produce reproducible reports, datasets, and attack cases that customers can act on
- Uncover vulnerabilities automated tests miss
- Expand evaluation coverage and reduce production surprises
- Strengthen customer AI systems and help make them more robust, safe, and trustworthy
- Work on projects training and enhancing frontier AI systems
Requirements:
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
- Ability to probe AI systems adversarially and push them to breaking points
- Ability to use frameworks or benchmarks for structured testing
- Ability to explain risks clearly to technical and non-technical stakeholders
- Ability to adapt across projects and customers
- Fluent English and Danish required
- Native fluency in English and Danish required
- Independent contractor engagement
- H1-B and STEM OPT candidates are not supported
Benefits:
- Fully remote role
- Flexible schedule / work can be completed on your own schedule
- Higher-sensitivity projects are optional
- Clear guidelines and wellness resources for sensitive content
- Weekly payments via Stripe or Wise
- Competitive pay
- Opportunity to collaborate with leading researchers
- Referral payments of up to $250 per successful referral



















