Senior Software Engineer – Open Source, SWE-Bench Evaluation
Posted 5hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Software engineers reviewing open-source GitHub tasks, pull requests, and tests for Anyone AI. Evaluating reproducibility, quality, complexity, and benchmark suitability.
Responsibilities:
- Review software engineering tasks derived from real GitHub issues and pull requests
- Assess whether problem statements and success criteria are clear and complete
- Evaluate unit tests for correctness, coverage, and robustness
- Identify flaky tests, missing dependencies, version conflicts, and environment issues
- Determine whether tasks can be reliably reproduced across environments
- Assess the real-world difficulty and complexity of each task
- Provide clear recommendations on whether tasks should be accepted, improved, or excluded
- Review real-world bug fixes and feature implementations, repository setup, dependency management, multi-file and cross-module changes, and technical quality
Requirements:
- 3+ years of professional software engineering experience
- Strong experience working with large, multi-file codebases
- Experience reviewing pull requests, debugging issues, and maintaining production code
- Strong understanding of unit testing and test coverage
- Ability to evaluate whether tests correctly validate a solution without unnecessarily restricting implementation approaches
- Experience with dependency management, environment setup, and reproducibility
- Strong understanding of Git and GitHub-based development workflows
- Ability to analyze complex technical problems and provide clear written feedback
- Contributions to or maintenance of open-source projects (nice to have)
- Experience with SWE-Bench, SWE-Bench Verified, or similar coding benchmarks (nice to have)
- Experience with major Python open-source projects such as Django, Flask, scikit-learn, SymPy, matplotlib, requests, or pytest (nice to have)
- Experience with Docker, CI/CD, pip, conda, or dependency pinning (nice to have)
- Knowledge of test fixtures, test isolation, or property-based testing (nice to have)
- Experience designing technical assessments or reviewing coding challenges (nice to have)
- Experience with AI/ML evaluation, data curation, RLHF, or benchmark development (nice to have)
Benefits:
- Remote work
- Part-time, project-based consulting engagement

















