AI Interaction Evaluator – Codex, Claude Code
Posted 5hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior engineers evaluating AI coding-agent interactions for G2i. Assessing reasoning, explanations, outputs, and engineering judgment across Codex, Claude Code, and Cursor.
Responsibilities:
- Evaluate AI-generated coding interactions end-to-end
- Judge whether outputs are useful, correct at a high level, and aligned with how a strong engineer would think
- Assess the quality of explanations and reasoning, not just code
- Distinguish between different levels of response quality
- Provide clear, opinionated feedback on what worked, what did not, and what felt off or misleading
- Help define what great looks like when interacting with tools like Cursor
- Evaluate whether AI coding agents' responses, preambles, reasoning, and outputs demonstrate strong engineering judgment
- Participate in a take-home evaluation exercise and one behavioral interview
Requirements:
- Highly experienced software engineer at Senior+ level
- Staff- or Principal-level engineer, or equivalent experience
- Strong background in TypeScript/JavaScript or Python
- Hands-on experience using OpenAI Codex, Claude Code, and Cursor
- Deep familiarity with modern AI-assisted development workflows
- Ability to evaluate code without fully executing or deeply reviewing every line
- Comfortable giving direct, opinionated feedback
- High standards for engineering quality
- Comfortable making subjective but rigorous judgments
Benefits:
- Possible extension beyond early May
- Flexible part-time engagement of approximately 10–20 hours per week
















