AI Safety Experts – English, Bengali
Posted 45mins ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI safety red team expert probing Mercor’s partner AI models for vulnerabilities. Generating reproducible adversarial data, reports, and attack cases to improve model safety.
Responsibilities:
- Red-team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Apply taxonomies, benchmarks, and playbooks to keep testing consistent
- Produce reproducible reports, datasets, and attack cases for customers
- Review AI outputs involving sensitive topics such as bias, misinformation, or harmful behaviors
- Uncover vulnerabilities that automated tests miss
- Expand evaluation coverage and reduce production surprises
- Strengthen customer AI systems through adversarial testing
Requirements:
- Native fluency in English and Bengali required
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
- Ability to probe systems adversarially and push them to breaking points
- Experience using frameworks, taxonomies, benchmarks, or playbooks
- Ability to explain risks clearly to technical and non-technical stakeholders
- Ability to adapt across projects and customers
- Independent contractor engagement
- Must not be an H1-B or STEM OPT candidate
- Nice-to-have specialties: adversarial ML, cybersecurity, socio-technical risk, creative probing
- Adversarial ML experience with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction is a plus
- Cybersecurity experience with penetration testing, exploit development, or reverse engineering is a plus
- Socio-technical experience with harassment/disinformation probing, abuse analysis, or conversational AI testing is a plus
- Creative probing experience in psychology, acting, or writing is a plus
Benefits:
- Fully remote role
- Flexible schedule / work on your own schedule
- Weekly payments via Stripe or Wise
- Projects may be extended, shortened, or concluded early depending on needs and performance
- Higher-sensitivity project participation is optional
- Clear guidelines and wellness resources for sensitive-content work
- Competitive pay
- Collaboration with leading researchers
- Referral opportunity earning up to $90 per successful referral













