AI Evaluation Specialist
Posted 2hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI Evaluation Specialist assessing AI-generated outputs, reasoning, and tool use for 24-MAG LLC’s remote consulting platform. Providing rubric-based quality evaluations and actionable written feedback.
Responsibilities:
- Evaluate AI-generated outputs against detailed rubrics, guidelines, and defined quality standards
- Assess responses for accuracy, relevance, completeness, reasoning quality, and adherence to instructions
- Apply consistent and impartial judgement across large volumes of evaluation examples
- Identify outputs that fail to satisfy important quality or task requirements
- Maintain reliable assessment standards across repeated evaluation workflows
- Identify reasoning gaps, logic errors, inconsistencies, unsupported conclusions, and tool-use failures
- Analyse where AI-generated outputs diverge from expected reasoning or quality standards
- Document recurring model weaknesses and opportunities for improvement
- Produce clear, concise, and actionable written feedback on strengths and areas for improvement
- Explain the reasoning behind evaluation decisions and quality scores
- Maintain detailed, transparent, traceable, and reproducible assessment documentation
- Participate in rubric interpretation and ambiguous-case discussions
- Help refine assessment criteria as AI models and project requirements develop
- Contribute insights supporting process optimisation and evaluation best practices
- Collaborate with other reviewers to improve alignment and reliability across evaluation workflows
Requirements:
- Experience in grading, quality assurance, editorial review, assessment, annotation, or another field requiring careful analysis and detailed feedback
- Advanced, regular use of AI assistants such as ChatGPT, Claude, or comparable tools for professional work and productivity
- Strong ability to synthesise complex information and communicate conclusions clearly in writing
- Experience with process improvement, rubric development, operational quality assessment, or structured evaluation workflows is advantageous
- Strong critical-thinking skills with particular emphasis on consistency, integrity, and fairness
- High attention to detail and comfort reviewing large volumes of similar examples
- Ability to work independently while maintaining consistent evaluation quality
- Collaborative approach to discussing ambiguous cases and refining shared assessment standards
- Excellent written English and professional documentation skills
- Based in the United States, Canada, United Kingdom, Ireland, Australia, or New Zealand
- Authorised to undertake contract work in the relevant country
- No prior formal experience in AI research or model training is required
Benefits:
- Part-time independent contractor engagement
- Fully remote work
- Flexible project scope, workload, timing, and duration depending on project requirements









