AI Safety Expert – English, Danish

Posted 2hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI safety red team expert probing Mercor’s frontier AI models for vulnerabilities. Generating reproducible attack data that strengthens customer AI safety and robustness.

Responsibilities:

  • Red-team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
  • Apply taxonomies, benchmarks, and playbooks to keep testing consistent
  • Produce reproducible reports, datasets, and attack cases for customers
  • Identify vulnerabilities that automated tests miss
  • Expand evaluation coverage and reduce production surprises
  • Help customers strengthen the safety, robustness, and trustworthiness of AI systems
  • Work on projects training and enhancing frontier AI systems

Requirements:

  • Native fluency in English and Danish is required
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to probe AI systems adversarially and push systems to breaking points
  • Experience using frameworks or benchmarks for structured testing
  • Ability to explain risks clearly to technical and non-technical stakeholders
  • Ability to adapt across projects and customers
  • Nice-to-have: adversarial ML experience, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction
  • Nice-to-have: cybersecurity experience, including penetration testing, exploit development, or reverse engineering
  • Nice-to-have: socio-technical risk experience, including harassment/disinformation probing, abuse analysis, or conversational AI testing
  • Nice-to-have: psychology, acting, or writing experience for unconventional adversarial thinking
  • Must be an independent contractor
  • Candidates must not require H1-B or STEM OPT support

Benefits:

  • Fully remote role
  • Flexible schedule; work can be completed on your own schedule
  • Weekly payments via Stripe or Wise based on services rendered
  • Projects may be extended, shortened, or concluded early depending on needs and performance
  • Participation in higher-sensitivity projects is optional
  • Clear guidelines and wellness resources for sensitive-content work
  • Competitive pay
  • Referral payments of up to $250 for each successful referral