AI Interaction Evaluator – Codex, Claude Code

Posted 5hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior engineers evaluating AI coding-agent interactions for G2i. Assessing reasoning, explanations, outputs, and engineering judgment across Codex, Claude Code, and Cursor.

Responsibilities:

  • Evaluate AI-generated coding interactions end-to-end
  • Judge whether outputs are useful, correct at a high level, and aligned with how a strong engineer would think
  • Assess the quality of explanations and reasoning, not just code
  • Distinguish between different levels of response quality
  • Provide clear, opinionated feedback on what worked, what did not, and what felt off or misleading
  • Help define what great looks like when interacting with tools like Cursor
  • Evaluate whether AI coding agents' responses, preambles, reasoning, and outputs demonstrate strong engineering judgment
  • Participate in a take-home evaluation exercise and one behavioral interview

Requirements:

  • Highly experienced software engineer at Senior+ level
  • Staff- or Principal-level engineer, or equivalent experience
  • Strong background in TypeScript/JavaScript or Python
  • Hands-on experience using OpenAI Codex, Claude Code, and Cursor
  • Deep familiarity with modern AI-assisted development workflows
  • Ability to evaluate code without fully executing or deeply reviewing every line
  • Comfortable giving direct, opinionated feedback
  • High standards for engineering quality
  • Comfortable making subjective but rigorous judgments

Benefits:

  • Possible extension beyond early May
  • Flexible part-time engagement of approximately 10–20 hours per week