AI Evaluation Specialist

Posted 2hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI Evaluation Specialist assessing AI-generated outputs, reasoning, and tool use for 24-MAG LLC’s remote consulting platform. Providing rubric-based quality evaluations and actionable written feedback.

Responsibilities:

  • Evaluate AI-generated outputs against detailed rubrics, guidelines, and defined quality standards
  • Assess responses for accuracy, relevance, completeness, reasoning quality, and adherence to instructions
  • Apply consistent and impartial judgement across large volumes of evaluation examples
  • Identify outputs that fail to satisfy important quality or task requirements
  • Maintain reliable assessment standards across repeated evaluation workflows
  • Identify reasoning gaps, logic errors, inconsistencies, unsupported conclusions, and tool-use failures
  • Analyse where AI-generated outputs diverge from expected reasoning or quality standards
  • Document recurring model weaknesses and opportunities for improvement
  • Produce clear, concise, and actionable written feedback on strengths and areas for improvement
  • Explain the reasoning behind evaluation decisions and quality scores
  • Maintain detailed, transparent, traceable, and reproducible assessment documentation
  • Participate in rubric interpretation and ambiguous-case discussions
  • Help refine assessment criteria as AI models and project requirements develop
  • Contribute insights supporting process optimisation and evaluation best practices
  • Collaborate with other reviewers to improve alignment and reliability across evaluation workflows

Requirements:

  • Experience in grading, quality assurance, editorial review, assessment, annotation, or another field requiring careful analysis and detailed feedback
  • Advanced, regular use of AI assistants such as ChatGPT, Claude, or comparable tools for professional work and productivity
  • Strong ability to synthesise complex information and communicate conclusions clearly in writing
  • Experience with process improvement, rubric development, operational quality assessment, or structured evaluation workflows is advantageous
  • Strong critical-thinking skills with particular emphasis on consistency, integrity, and fairness
  • High attention to detail and comfort reviewing large volumes of similar examples
  • Ability to work independently while maintaining consistent evaluation quality
  • Collaborative approach to discussing ambiguous cases and refining shared assessment standards
  • Excellent written English and professional documentation skills
  • Based in the United States, Canada, United Kingdom, Ireland, Australia, or New Zealand
  • Authorised to undertake contract work in the relevant country
  • No prior formal experience in AI research or model training is required

Benefits:

  • Part-time independent contractor engagement
  • Fully remote work
  • Flexible project scope, workload, timing, and duration depending on project requirements