AI Safety Expert, English, Finnish

Posted 6hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI red-teaming experts probing frontier models for Mercor’s human-data AI training projects. Generating vulnerability annotations, attack cases, datasets, and reproducible safety reports.

Responsibilities:

  • Red-team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
  • Apply taxonomies, benchmarks, and playbooks to maintain consistent testing
  • Produce reproducible reports, datasets, and attack cases for customers
  • Review AI outputs involving sensitive topics such as bias, misinformation, or harmful behaviors
  • Identify vulnerabilities that automated tests miss
  • Expand evaluation coverage and reduce production surprises
  • Collaborate with leading researchers and contribute to training and enhancing frontier AI systems

Requirements:

  • Fluent/native fluency in English and Finnish
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to probe systems adversarially and push them to breaking points
  • Ability to use frameworks or benchmarks for structured testing
  • Ability to explain risks clearly to technical and non-technical stakeholders
  • Ability to adapt across projects and customers
  • Nice-to-have specialties: adversarial ML, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, or unconventional adversarial writing
  • Must be an independent contractor
  • H1-B and STEM OPT candidates are not supported

Benefits:

  • Fully remote role
  • Flexible schedule; work can be completed on your own schedule
  • Weekly payments via Stripe or Wise
  • Opportunity to build experience in human data-driven AI red teaming
  • Direct role in making AI systems more robust, safe, and trustworthy
  • Competitive payment
  • Collaboration with leading researchers
  • Reasonable accommodations upon request
  • Referral payments of up to $250 per successful referral