Public summary
This is a freelance opportunity to work on evaluating AI coding agents by creating and refining tasks within realistic developer environments. The role involves building virtual coding scenarios, designing evaluation criteria, and writing tests to assess AI solutions accuracy. Applicants should have over 5 years of software development experience with proficiency in Python (FastAPI), JavaScript/TypeScript (React), and related technologies, as well as experience in functional and integration testing. English proficiency at B2 level or higher is required. The work is project-based with flexible scheduling and payment up to $50 per hour depending on experience and pace.
Location and work setup
- Location
- Stuttgart
- Remote status
- Remote
- German requirement signal
- No German Required Detected
- Detected job language
- English
Salary
USD 50.00 - 50.00 hour
Responsibilities
Create realistic virtual developer environments representing a company's codebase and context. Design AI evaluation tasks with clear definitions of success. Write and refine tests that accept all valid AI-generated solutions while rejecting incorrect ones. Analyze AI performance and iterate on tasks based on quality assurance feedback to ensure fairness and robustness.
Qualifications
Minimum 5 years of professional software development experience. Strong knowledge of Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis. Proven experience in writing functional and integration tests. English language proficiency at B2 level or above. Ability to critically evaluate AI-generated code and create challenging development tasks that differentiate AI capabilities.