AI Policy Generalist
Posted 5hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI Policy Generalist evaluating AI model behavior for Handshake’s human-data business. Applying customer policies to ambiguous conversations and improving evaluation frameworks.
Responsibilities:
- Learn customer policies, definitions, taxonomies, and evaluation rubrics
- Evaluate user requests and AI model responses within full conversation context
- Classify cases according to the most applicable policy category
- Distinguish closely related labels, severity levels, and policy boundaries
- Select defensible classifications for ambiguous cases
- Write concise, evidence-based rationales citing policy language and conversation details
- Identify policy gaps, contradictions, unclear definitions, and emerging edge cases
- Raise questions when guidance does not resolve a case
- Participate in calibration discussions with evaluators, project leads, policy teams, and researchers
- Apply customer policies consistently without substituting personal beliefs
- Maintain accuracy and attention to detail across repeated evaluations
- Incorporate feedback and apply clarified guidance
- Help improve evaluation frameworks, examples, decision rules, and quality standards
- Move between projects covering different policy domains and customer needs
Requirements:
- Must be based in the United States
- Must be authorized to work lawfully in the United States for Handshake
- Must be able to work Monday–Friday from 8AM–5PM PT as core hours
- Ability to engage carefully, responsibly, and sustainably with sensitive material, including sexual content, emotional distress, self-harm, suicide, violence, weapons, abuse, exploitation, and discrimination
- Sound judgment and consistent work quality
- Ability to learn unfamiliar subject matter quickly
- Ability to communicate clearly and precisely in writing
- Ability to apply detailed policies, definitions, taxonomies, rubrics, and evaluation frameworks
- Ability to distinguish closely related labels, severity levels, and policy boundaries
- Ability to write concise, evidence-based rationales
- Ability to participate in calibration discussions and incorporate feedback
- Prior AI evaluation experience is helpful but not required
- Strong candidates may come from quality assurance, research, editing, law, teaching, operations, trust and safety, content moderation, social science, policy, investigations, compliance, or customer support
Benefits:
- W-2 employment classification
- Monday–Friday schedule
- Core hours of 8AM–5PM PT
- Remote work arrangement














