We are sharing a specialised part-time consulting opportunity for experienced Software Engineers with hands-on open-source contribution or maintainer experience and strong expertise in repository-level code review, testing, debugging, and software quality evaluation. This role focuses on reviewing software-engineering benchmark tasks for correctness, reproducibility, and grading integrity. Selected engineers will audit repository-level assignments, reference patches, test harnesses, containerised environments, and evaluation logic while identifying technical flaws, unintended shortcuts, and weaknesses in task design. Key Responsibilities Repository-Level Code Review Review software engineering tasks built around real code repositories Assess whether task requirements are technically clear, complete, and reproducible Evaluate repository state, dependencies, configuration, and expected behaviour Identify ambiguities or implementation issues that could affect task validity Apply practical engineering judgement to realistic codebase-level problems Reference Patch Auditing Review reference patches for correctness and completeness Determine whether proposed solutions appropriately address the underlying software issue Identify unintended behavioural changes, incomplete fixes, or unsupported assumptions Compare reference implementations against task requirements and expected outcomes Assess whether alternative valid implementations are treated fairly Test Harness & Grading Review Audit test runners and automated evaluation logic Assess whether tests accurately measure the intended behaviour Identify missing coverage, brittle assertions, or grading inconsistencies Verify that evaluation criteria appropriately distinguish correct from incorrect solutions Review benchmark tasks for reliable and repeatable scoring Reproducibility & Environment Validation Evaluate whether tasks can be reproduced consistently across clean environments Review dependency installation, build processes, configuration, and runtime requirements Assess Docker-based isolation and containerised execution Identify environmental dependencies or hidden assumptions affecting reproducibility Verify that tasks execute reliably under their intended setup Benchmark Integrity Identify potential answer leakage, unintended shortcuts, or reward-hacking opportunities Evaluate whether benchmark structure exposes information that makes tasks artificially easy Review task and grading design for loopholes or exploitable behaviours Assess whether successful completion genuinely demonstrates the intended engineering capability Recommend improvements where benchmark integrity is compromised Software Testing & Debugging Investigate failing or inconsistent benchmark tasks Review stack traces, logs, test failures, and repository behaviour Identify root causes of technical issues Distinguish task defects from legitimate implementation failures Assess whether debugging and validation processes follow sound engineering practices Open-Source Engineering Apply experience from contributing to or maintaining open-source software Evaluate repository conventions, contribution patterns, and realistic development workflows Review patches with the perspective of an experienced contributor or maintainer Assess whether proposed changes would meet reasonable code-review expectations Apply practical judgement derived from real-world pull request and repository experience Multi-Language Code Evaluation Review software written in Python Evaluate tasks involving at least one additional ecosystem such as Java, Go, TypeScript, or C++ Assess code structure, tests, implementation choices, and repository conventions across languages Identify language-specific implementation or testing issues Apply consistent engineering standards across different technology stacks Rubric-Based Evaluation Assess benchmark tasks against structured technical criteria Provide clear written explanations supporting evaluation decisions Reference specific code, tests, patches, or execution behaviour when identifying issues Apply grading standards consistently across assignments Distinguish substantive benchmark defects from minor implementation differences Ideal Profile 3+ years of professional software engineering experience Demonstrated open-source contribution or maintainer experience , such as merged pull requests, committer responsibilities, or maintainer roles Strong ability to review repository-level software changes Experience auditing reference patches, test runners, and automated test suites Comfortable evaluating Docker-based isolation and reproducible development environments Strong understanding of software testing, debugging, and code-review practices Ability to identify answer leakage, reward hacking, or other benchmark-integrity issues Strong proficiency in Python Professional fluency in at least one additional language such as Java, Go, TypeScript, or C++ Familiarity with SWE-Bench Verified or similar repository-level software engineering benchmarks is preferred Maintainer or contributor history with established Python open-source projects is highly valued Previous code-review, software evaluation, or task-grading experience is advantageous Strong written communication and ability to provide precise technical feedback Engagement Details Part-time independent contractor engagement Fully remote within the United States Flexible scheduling based on project requirements Compensation: $60–$80/hour Work focuses on repository-level software evaluation, reference-patch review, testing, reproducibility, benchmark integrity, and technical quality assessment Projects may be extended, shortened, or concluded based on project needs and performance Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party H1-B and STEM OPT support is unavailable for this engagement About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy .
Remote | Open Source Code Evaluation Engineer — Up to $90/hour
24-MAG
Remote | Software Engineer (AI-Assisted Development) — $60–$80/hour
24-MAG
Software Engineer - Hardware Diagnostics - New Grad
Nexthop Systems Inc
Senior Software Engineer – VEEComm
Generalmotors
Senior Project Engineer - Structural
Martin/Martin, Inc.
Structural Project Engineer
EwingCole, Inc.