AI Safety Expert, English, Finnish

Posted 15hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI red team expert testing Mercor’s frontier AI models through jailbreaks, prompt injections, and adversarial data generation. Producing reproducible vulnerability reports that improve AI safety.

Responsibilities:

  • Red team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Annotate failures, classify vulnerabilities, and flag systemic risks
  • Apply taxonomies, benchmarks, and playbooks to maintain consistent testing
  • Produce reproducible reports, datasets, and attack cases
  • Probe sensitive AI outputs involving bias, misinformation, and harmful behaviors
  • Expand evaluation coverage and identify vulnerabilities automated tests miss
  • Collaborate with leading researchers and contribute to training and enhancing frontier AI systems

Requirements:

  • Native fluency in English and Finnish
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to probe AI systems adversarially and push systems to breaking points
  • Experience using frameworks or benchmarks for structured testing
  • Ability to explain risks clearly to technical and non-technical stakeholders
  • Ability to adapt across projects and customers
  • Independent contractor status
  • Must not require H1-B or STEM OPT support
  • Nice-to-have specialties: adversarial ML, cybersecurity, socio-technical risk, or creative probing

Benefits:

  • Fully remote work
  • Flexible own schedule
  • Weekly payments via Stripe or Wise
  • Competitive pay
  • Reasonable accommodations upon request
  • Optional participation in higher-sensitivity projects
  • Clear guidelines and wellness resources for higher-sensitivity projects
  • Referral payments of up to $250 per successful referral