AI Safety Expert – English, Marathi

Posted 18hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI safety red teamer probing conversational AI models for Mercor, which partners with AI labs and enterprises to train frontier systems. Generating reproducible vulnerability data and attack cases.

Responsibilities:

  • Red-team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Annotate failures, classify vulnerabilities, and flag systemic risks
  • Follow taxonomies, benchmarks, and playbooks to keep testing consistent
  • Produce reproducible reports, datasets, and attack cases
  • Review AI outputs involving sensitive topics such as bias, misinformation, and harmful behaviors
  • Expand evaluation coverage and uncover vulnerabilities missed by automated tests
  • Collaborate on projects training and enhancing frontier AI systems for Mercor's customers

Requirements:

  • Fluent/native fluency in English and Marathi required
  • Strong judgment about language and content; ability to assess whether AI responses are accurate, complete, and appropriate, and explain why
  • Rigorous attention to subtle errors, inconsistencies, and gaps
  • Ability to follow guidelines and quality standards consistently
  • Ability to explain reasoning clearly to technical and non-technical audiences
  • Adaptability across projects, task types, and customers
  • Independent contractor status
  • Must not be an H-1B or STEM OPT candidate
  • Nice-to-have specialties: adversarial ML, cybersecurity, socio-technical risk, or creative probing

Benefits:

  • Fully remote role
  • Flexible schedule; work on your own schedule
  • Weekly payments via Stripe or Wise
  • Higher-sensitivity project participation is optional
  • Clear guidelines and wellness resources for higher-sensitivity projects
  • Competitive pay
  • Referral opportunity earning up to $90 per successful referral (referral limits apply)
  • Reasonable accommodations upon request