AI Safety Expert – English, Finnish

Posted 2hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI safety red teamer probing conversational AI models for Mercor’s human-data AI projects. Generating reproducible vulnerability data that strengthens frontier AI systems.

Responsibilities:

  • Red team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Annotate failures, classify vulnerabilities, and flag systemic risks
  • Follow taxonomies, benchmarks, and playbooks to keep testing consistent
  • Produce reproducible reports, datasets, and attack cases for customers
  • Uncover vulnerabilities automated tests miss
  • Expand evaluation coverage and reduce production surprises
  • Collaborate with leading researchers and work on projects training and enhancing AI systems

Requirements:

  • Fluent/native English and Finnish required
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to probe AI systems adversarially and push systems to breaking points
  • Experience using frameworks or benchmarks for structured testing
  • Ability to explain risks clearly to technical and non-technical stakeholders
  • Adaptability across projects and customers
  • Independent contractor engagement
  • Must not be an H1-B or STEM OPT candidate
  • Nice-to-have: adversarial ML experience, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction
  • Nice-to-have: cybersecurity experience, including penetration testing, exploit development, or reverse engineering
  • Nice-to-have: socio-technical risk experience, including harassment/disinformation probing, abuse analysis, or conversational AI testing
  • Nice-to-have: psychology, acting, or writing experience for unconventional adversarial thinking

Benefits:

  • Fully remote role
  • Flexible own schedule
  • Weekly payments via Stripe or Wise
  • Wellness resources and clear guidelines for higher-sensitivity projects
  • Opportunity to build experience in human data-driven AI red teaming
  • Direct role in making AI systems more robust, safe, and trustworthy
  • Competitive pay
  • Referral bonus of up to $250 per successful referral