AI Quality & Evaluation Lead

Posted 18hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI evaluation lead building quality rubrics and validating judges for Samsung Health’s generated coaching. Samsung Food connects food, health, and home through personalized digital experiences.

Responsibilities:

  • Build the experience-quality evaluation layer for AI-generated coaching in Samsung Health
  • Author a failure taxonomy for coaching outputs
  • Hand-grade at least 150 traces with open-coded notes, including at least 40 sparse-data synthetic profiles
  • Categorize failures into 5–10 categories and quantify their frequency
  • Freeze at least 50 traces as a holdout set
  • Create binary pass/fail criteria for each taxonomy category
  • Write annotation guidelines with pass and fail examples
  • Double-code at least 30 traces, resolve disagreements, and maintain a guideline revision log
  • Create judge prompts for every criterion
  • Validate the LLM judge using true-positive and true-negative rates on development and holdout sets
  • Run weekly readouts on helpfulness, relevance, and tone
  • Document a revalidation routine triggered by model or prompt changes and quarterly regardless
  • Produce a complete playbook covering grading, taxonomy updates, rubric revisions, judge revalidation, and weekly readouts
  • Conduct a handover test enabling the internal owner to rerun validation and a weekly readout independently
  • Obtain Head of Product sign-off on the taxonomy and rubric
  • Collaborate with and hand over the method to an internal owner

Requirements:

  • Experience running the full evaluation loop at least once on conversational or generated-text output
  • Experience with error analysis on real traces
  • Experience building a failure taxonomy
  • Experience creating binary criteria and annotation guidelines
  • Experience validating an LLM judge against personal labels
  • Background in conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, RLHF, applied linguistics, or product management
  • Ability to identify previously unnamed failure modes by reading output
  • Ability to act as arbiter and document overrulings and rationale
  • Understanding that rubrics are discovered through grading
  • Ability to explain limitations of judge agreement rates
  • Ability to write unambiguous annotation guidelines
  • Comfort working in notebooks and spreadsheets
  • No production coding required
  • Nutrition, weight management, or behaviour change domain expertise is not required
  • Must be comfortable being responsible for own taxes and having no paid time off

Benefits:

  • Remote-first work arrangement
  • Independent contractor setup
  • Flexible remote work
  • Global team of more than 100 people in over 30 countries