AI Safety Expert – English, Danish
Posted 8hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI safety red teamer probing frontier AI models for Mercor. Testing jailbreaks, vulnerabilities, and harmful behaviors to generate safer human-data-driven AI systems.
Responsibilities:
- Red-team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Apply taxonomies, benchmarks, and playbooks to maintain consistent testing
- Produce reproducible reports, datasets, and attack cases for customers
- Probe AI systems to uncover vulnerabilities missed by automated tests
- Expand evaluation coverage and reduce production surprises
- Help Mercor customers make AI systems safer, more robust, and trustworthy
Requirements:
- Native fluency in English and Danish is required
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
- Ability to test AI systems adversarially and push them to breaking points
- Experience using frameworks or benchmarks for structured testing
- Ability to explain risks clearly to technical and non-technical stakeholders
- Adaptability across projects and customers
- Nice-to-have specialties include adversarial ML, cybersecurity, socio-technical risk, or creative probing
- Must work as an independent contractor
- H1-B and STEM OPT candidates are not supported
Benefits:
- Fully remote role
- Flexible schedule / work on your own schedule
- Weekly payments via Stripe or Wise based on services rendered
- Higher-sensitivity projects are optional
- Clear guidelines and wellness resources for sensitive-content projects
- Competitive pay
- Opportunity to collaborate with leading researchers
- Referral payments of up to $250 per successful referral












