AI Safety Expert – English, Portuguese

Posted 45mins ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI safety red teamer probing conversational models for jailbreaks, prompt injections, and systemic risks. Producing human-data artifacts that help Mercor customers build safer frontier AI.

Responsibilities:

  • Red-team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Annotate failures, classify vulnerabilities, and flag systemic risks
  • Apply taxonomies, benchmarks, and playbooks to keep testing consistent
  • Produce reproducible reports, datasets, and attack cases
  • Probe AI outputs involving bias, misinformation, and harmful behaviors
  • Help expand evaluation coverage and strengthen customer AI systems

Requirements:

  • Native fluency in English and Portuguese (global, excluding Brazilian Portuguese)
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to probe AI systems adversarially and push systems to breaking points
  • Experience using frameworks or benchmarks for structured testing
  • Ability to explain risks clearly to technical and non-technical stakeholders
  • Ability to adapt across projects and customers
  • Independent contractor status
  • Must not be an H-1B or STEM OPT candidate
  • Nice-to-have specialties: adversarial ML, cybersecurity, socio-technical risk, or creative probing

Benefits:

  • Fully remote role
  • Flexible own schedule
  • Weekly payments via Stripe or Wise
  • Higher-sensitivity project participation is optional
  • Clear content guidelines and wellness resources
  • Reasonable accommodations upon request
  • Opportunity to build experience in human data-driven AI red teaming
  • Collaboration with leading researchers
  • Referral payments of up to $180 per successful referral