24-MAG
US - New York - New York
View Company Profile /
<< Go Back
**We are sharing a specialised full-time consulting opportunity for experienced software engineers with strong Python development, debugging, version-control, technical documentation, and AI-assisted coding experience.**
This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will design, implement, and review realistic multi-step software engineering tasks that test the capabilities of AI coding agents across Python development, environment setup, tooling, debugging, and technical problem-solving.
**Key Responsibilities**
**Software Engineering Task Design**
* Create realistic, multi-step software engineering challenges based on practical development workflows
* Design technically demanding problems that require implementation, debugging, environment configuration, and analytical reasoning
* Define clear requirements, constraints, expected outputs, and acceptance criteria
* Ensure tasks assess genuine software engineering capability rather than superficial code generation
**Reference Solution Development**
* Build complete and verifiable reference solutions in Python
* Create the supporting setup, dependencies, tests, and validation checks required for each task
* Write clean, readable, and maintainable code
* Confirm that solutions run reliably within the intended technical environment
* Document implementation decisions and expected behaviour clearly
**AI Coding Agent Evaluation**
* Use AI coding assistants and agent-based development tools within practical engineering workflows
* Evaluate how frontier models approach complex coding and debugging tasks
* Identify implementation errors, unsupported assumptions, inefficient approaches, and incomplete solutions
* Analyse where AI agents succeed, struggle, or exploit unintended shortcuts
* Document failure patterns and provide evidence supporting evaluation conclusions
**Peer Review \& Task Refinement**
* Review tasks and reference solutions created by other software engineering specialists
* Assess clarity, correctness, difficulty, reproducibility, and technical fairness
* Identify ambiguous instructions, hidden assumptions, grading gaps, and environment issues
* Provide actionable feedback that improves task quality and benchmark reliability
* Collaborate closely with researchers and fellow task authors
**Ideal Profile**
**Strong candidates may have:**
* At least 1 year of experience in software engineering, research engineering, or a related coding-intensive role
* Strong hands-on Python scripting, implementation, and debugging skills
* Experience developing clean, readable, and maintainable software
* Everyday fluency with Git, IDEs, repositories, and standard software development workflows
* Comfort configuring environments, dependencies, tooling, and validation processes
* Strong technical writing and documentation skills
* Ability to work independently through ambiguous and open-ended engineering problems
* Reliable availability for approximately 35 hours per week
**Educational Background**
* An MSc or PhD in computer science, software engineering, another STEM discipline, or a related technical field is highly relevant
* Equivalent practical experience in a research-intensive or engineering-intensive role may also be considered
* Academic or professional work involving significant coding, data analysis, or technical experimentation may strengthen an application
* Open-source contributions, technical projects, publications, or substantial software development work may also be valuable
**Nice to Have**
* Experience using AI coding assistants, prompt engineering methods, or agent-based workflows
* Previous work in AI training, model evaluation, or benchmark development
* Background authoring technical tasks, reference solutions, or grading criteria
* Familiarity with automated testing, CI/CD workflows, containers, or reproducible environments
* Experience reviewing code or technical assignments created by other engineers
* Knowledge of agentic AI systems and multi-step coding evaluations
* Strong ability to identify edge cases, unintended shortcuts, and subtle implementation issues
* Experience collaborating with AI research or evaluation teams
**Why This Opportunity**
* Apply practical software engineering expertise to frontier AI evaluation
* Design realistic coding tasks grounded in professional development workflows
* Help researchers understand where advanced AI coding agents succeed and fail
* Work across Python implementation, debugging, environment setup, and benchmark development
* Collaborate closely with AI researchers and experienced software engineers
* Participate in a structured full-time remote role with competitive hourly compensation
**Contract Details**
* Full-time W-2 contingent employment opportunity
* Fully remote within the United States
* Expected commitment of approximately 35 hours per week
* Competitive rates between $55--$85 per hour depending on expertise and project scope
* Individual tasks may require one to two days of focused engineering work
* Work may include task design, Python development, reference-solution creation, AI agent evaluation, peer review, and technical documentation
* Engagement scope and duration may evolve according to project requirements and performance
**About the Platform**
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX