We are sharing a specialised full-time consulting opportunity for experienced data scientists and quantitative analysts with strong expertise in statistical analysis, data cleaning, method comparison, reproducible research, and evidence-based reporting. This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will create realistic data-analysis challenges, develop reproducible reference notebooks, evaluate model-generated analyses, and identify where statistical reasoning, interpretation, or reporting falls short of professional standards. Key Responsibilities Data Analysis Task Design Create realistic analytical tasks based on professional data science and quantitative research workflows Develop assignments involving messy data, anomaly detection, correlation analysis, hypothesis testing, and method comparison Design complex, multi-step problems requiring statistical judgment and careful interpretation Ensure tasks include realistic constraints, datasets, assumptions, and decision-making objectives Reproducible Notebook Development Complete reference analyses using Jupyter Notebook or Google Colab Build clear and reproducible workflows using Python, pandas, NumPy, and related libraries Document data-cleaning decisions, calculations, statistical methods, and analytical conclusions Validate intermediate results, spot checks, visualisations, and final recommendations Statistical Method Comparison Design fair comparisons between analytical models, algorithms, or statistical approaches Evaluate performance using appropriate metrics, manual checks, and sensitivity analyses Identify methodological trade-offs, limitations, and sources of uncertainty Produce recommendations supported by transparent quantitative evidence AI Model Evaluation Review model-generated analyses for statistical accuracy, methodological rigour, and sound interpretation Verify whether calculations, correlations, hypotheses, and conclusions are supported by the data Identify coding errors, unsupported assumptions, misleading summaries, and analytical shortcuts Explain where and why model outputs fail to meet professional data-analysis standards Research Collaboration Work closely with researchers, task authors, and fellow quantitative specialists Compare evaluation decisions to maintain consistent benchmark standards Refine tasks, reference notebooks, and grading criteria based on testing outcomes Document recurring model weaknesses and opportunities for stronger evaluation coverage Ideal Profile Strong candidates may have: At least 1 year of experience in data science, quantitative analysis, research engineering, or another research-intensive analytical role Deep hands-on experience with data cleaning, statistical correlation, hypothesis testing, and interpretation Strong proficiency in Python, including pandas, NumPy, or comparable analytical libraries Experience using Jupyter Notebook or Google Colab for analysis and reporting Working familiarity with Git and reproducible analytical workflows Ability to communicate complex quantitative findings clearly to technical and non-technical decision-makers Strong attention to detail and confidence working through ambiguous, open-ended problems Reliable availability for approximately 35 hours per week Educational Background A master's degree or PhD in statistics, data science, mathematics, economics, computer science, engineering, or another quantitative discipline is highly relevant Equivalent practical experience in a research-heavy analytical field may also be considered Academic or professional research involving statistical modelling, experimentation, or large-scale data analysis may strengthen an application Publications, technical reports, open-source work, or impactful analytical projects may also be valuable Nice to Have Experience in AI training, model evaluation, or benchmark development Background authoring analytical tasks, reference solutions, or grading rubrics Familiarity with anomaly detection, experimental design, or comparative model evaluation Experience conducting manual spot checks and validating automated analyses Knowledge of statistical modelling, machine learning, or scientific computing Familiarity with agentic AI systems and multi-step model evaluations Experience reviewing notebooks, code, or analyses prepared by other professionals Strong ability to identify subtle statistical errors and unsupported conclusions Why This Opportunity Apply advanced data science and quantitative analysis expertise to frontier AI evaluation Design realistic tasks grounded in professional analytical workflows Help improve how AI systems reason through statistics, data quality, and method comparison Work across Python, reproducible notebooks, model evaluation, and evidence-based reporting Collaborate closely with researchers and other quantitative specialists Participate in a structured full-time remote role with competitive hourly compensation Contract Details Full-time W-2 contingent employment opportunity Fully remote within the United States Expected commitment of approximately 35 hours per week Competitive rates between $55–$85 per hour depending on expertise and project scope Individual tasks may require one to two days of focused analysis and implementation Work may include task design, data cleaning, statistical analysis, notebook development, AI output evaluation, and technical reporting Engagement scope and duration may evolve according to project requirements and performance About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy .
Remote | Business Intelligence Analyst — $75–$110/hour
24 Mag
Remote | Finance Research Analyst — $65–$95/hour
24 Mag
Senior HR Technology Analyst - Hourly Compensation
Generalmotors
Enrollment Data & Storytelling Analyst (Graduate Student Hourly)
University of Wisconsin-Madison Login
Control Systems Analyst II (Off-hours)
GrayMatter Systems
AI Quality Analyst - Flexible Hours
Innodata Inc.