AI Safety Experts – English, Malay

Posted 16hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI red teaming expert probing frontier models for Mercor, which partners with AI labs and enterprises to train advanced models. Generating reproducible attack data and vulnerability reports to improve AI safety and robustness.

Responsibilities:

  • Red team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
  • Apply taxonomies, benchmarks, and playbooks to keep testing consistent
  • Produce reproducible reports, datasets, and attack cases for customers
  • Probe AI systems to uncover vulnerabilities automated tests miss
  • Expand evaluation coverage and reduce production surprises
  • Support Mercor customers in improving the safety and robustness of their AI systems

Requirements:

  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to push systems to breaking points through adversarial testing
  • Experience using frameworks or benchmarks rather than random hacks
  • Ability to explain risks clearly to technical and non-technical stakeholders
  • Adaptability across projects and customers
  • Native fluency in English and Malay
  • Adversarial ML experience such as jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction is a nice-to-have
  • Cybersecurity specialties such as penetration testing, exploit development, or reverse engineering are a nice-to-have
  • Socio-technical risk specialties such as harassment/disinformation probing, abuse analysis, or conversational AI testing are a nice-to-have
  • Creative probing experience in psychology, acting, or writing is a nice-to-have
  • Must not be an H1-B or STEM OPT candidate

Benefits:

  • Fully remote role
  • Flexible own schedule
  • Higher-sensitivity project participation is optional
  • Clear guidelines and wellness resources for sensitive-content projects
  • Weekly payments via Stripe or Wise
  • Opportunity to build experience in human data-driven AI red teaming
  • Direct role in making AI systems more robust, safe, and trustworthy
  • Competitive pay
  • Collaboration with leading researchers