AI Safety Expert – English, Kannada
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI safety red teamer probing conversational AI models for Mercor, which partners with leading AI labs to train frontier systems. Creating vulnerability data, attack cases, and reproducible reports.
Responsibilities:
- Red-team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Review AI outputs involving sensitive topics such as bias, misinformation, or harmful behaviors
- Annotate failures, classify vulnerabilities, and flag systemic risks
- Apply taxonomies, benchmarks, and playbooks to keep testing consistent
- Produce reproducible reports, datasets, and attack cases
- Uncover vulnerabilities automated tests miss
- Expand evaluation coverage and reduce production surprises
- Help Mercor customers make AI systems more robust, safe, and trustworthy
Requirements:
- Native fluency in English and Kannada is required
- Strong judgment about language and content
- Ability to assess whether AI responses are accurate, complete, and appropriate
- Ability to explain reasoning clearly to technical and non-technical audiences
- Rigorous attention to subtle errors, inconsistencies, and gaps
- Ability to follow guidelines, taxonomies, benchmarks, playbooks, and quality standards consistently
- Adaptability across projects, task types, and customers
- Independent contractor status
- H1-B and STEM OPT candidates are not supported
- Nice-to-have specialties: adversarial ML, cybersecurity, socio-technical risk, or creative probing
Benefits:
- Fully remote role
- Flexible schedule / work on your own schedule
- Weekly payments via Stripe or Wise
- Competitive pay
- Wellness resources and clear guidelines for higher-sensitivity projects
- Reasonable accommodations upon request
- Referral opportunity: earn up to $90 per successful referral






