24-MAG
US - New York - New York
View Company Profile /
<< Go Back
**We are sharing a specialised full-time consulting opportunity for machine learning engineers and research practitioners with hands-on experience training, evaluating, and experimenting with ML models end to end.**
This role supports the development of advanced agentic evaluation benchmarks for frontier AI systems. Selected professionals will transform real machine learning research ideas into rigorous multi-step tasks, implement and run experiments, analyse training behaviour, and evaluate where model-generated solutions fall short of technically correct results.
**Key Responsibilities**
**Machine Learning Task Design**
* Turn practical ML research ideas into well-defined, multi-step evaluation tasks
* Develop assignments involving model training, experimental modifications, and performance analysis
* Define clear technical requirements, expected outputs, and success criteria
* Ensure tasks assess genuine implementation and experimental reasoning rather than superficial library usage
**Experiment Implementation \& Execution**
* Implement reference solutions using Python, scripts, and notebook environments
* Configure and run model-training experiments from setup through final evaluation
* Modify model components, training procedures, reward functions, or experimental parameters
* Validate code, dependencies, datasets, intermediate outputs, and final results
* Document complete workflows so experiments can be reproduced independently
**Model Evaluation \& Analysis**
* Review how frontier AI models approach complex machine learning tasks
* Assess implementation quality, experimental methodology, and technical conclusions
* Identify coding errors, unsupported assumptions, weak experimental controls, and misleading interpretations
* Determine whether reported improvements are supported by the observed results
* Explain clearly where and why a model-generated solution fails
**Reinforcement Learning Experiments**
* Develop selected tasks involving reinforcement learning fundamentals
* Evaluate reward-function changes, policy-training behaviour, and experimental outcomes
* Assess whether proposed modifications produce the intended training effect
* Identify instability, unintended incentives, or incorrect interpretations of RL results
**Research Collaboration**
* Work closely with researchers, task authors, and fellow machine learning specialists
* Compare evaluation decisions to maintain consistent and rigorous benchmark standards
* Refine task instructions, reference solutions, and grading criteria based on testing outcomes
* Share recurring model failure patterns and opportunities for stronger benchmark coverage
**Ideal Profile**
**Strong candidates may have:**
* At least 1 year of experience in machine learning research, research engineering, or a comparable technical role
* Hands-on experience training and evaluating ML models through complete experimental workflows
* Strong understanding of experiment setup, execution, analysis, and reproducibility
* Familiarity with large language model capabilities, limitations, and evaluation techniques
* Working proficiency in Python and Git
* Comfort using both scripting and notebook-based environments
* Strong technical writing, analytical reasoning, and attention to detail
* Ability to work independently through ambiguous, open-ended research problems
* Reliable availability for approximately 35 hours per week
**Educational Background**
* A master's degree or PhD in machine learning, computer science, artificial intelligence, engineering, mathematics, or another relevant STEM discipline is highly relevant
* Equivalent practical experience in a research-intensive machine learning role may also be considered
* Academic or professional work involving model training, experimentation, or ML systems may strengthen an application
* Publications, open-source contributions, technical reports, or substantial research projects may also be valuable
**Nice to Have**
* Understanding of reinforcement learning concepts, including reward functions and policy training
* Experience in AI training, model evaluation, or benchmark development
* Background authoring technical tasks, reference solutions, or grading rubrics
* Familiarity with agentic AI systems and multi-step model evaluations
* Experience diagnosing model-training failures or unexpected experimental behaviour
* Knowledge of experimental design, ablation studies, and performance comparison
* Experience reviewing code, notebooks, or research analyses prepared by other practitioners
* Familiarity with reproducible ML environments and collaborative Git workflows
**Why This Opportunity**
* Apply practical machine learning research expertise to frontier AI evaluation
* Design realistic tasks grounded in end-to-end model experimentation
* Help improve how AI systems approach implementation, training, and analytical reasoning
* Work across Python, ML evaluation, reinforcement learning, and reproducible research
* Collaborate closely with AI researchers and machine learning specialists
* Participate in a structured full-time remote role with competitive hourly compensation
**Contract Details**
* Full-time W-2 contingent employment opportunity
* Fully remote within the United States
* Expected commitment of approximately 35 hours per week
* Competitive rates between $55--$85 per hour depending on expertise and project scope
* Individual tasks may require one to two days of focused implementation and experimental work
* Work may include task design, model training, experiment execution, notebook development, AI output evaluation, and technical reporting
* Engagement scope and duration may evolve according to project requirements and performance
**About the Platform**
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX