24-MAG logo

Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

Hiring from
United States
Work type
Remote
Posted
Is this job info correct?
Show job description

We are sharing a specialised part-time consulting opportunity for experienced software engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research.

Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help research teams understand how advanced models perform throughout realistic software-development workflows.

Key Responsibilities

Code Curation & Solution Development

  • Curate high-quality code examples for model training and benchmarking
  • Develop precise solutions to software-engineering tasks
  • Correct and improve code across multiple programming languages
  • Work with Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go
  • Maintain strong standards for correctness and maintainability

AI-Generated Code Evaluation

  • Evaluate AI-generated code for technical correctness
  • Assess solutions for efficiency, scalability, and reliability
  • Identify implementation weaknesses and recurring error patterns
  • Review code quality against professional engineering standards
  • Provide structured rationales supporting evaluation decisions

Verification & Automated Assessment

  • Build agents that assess code quality
  • Design mechanisms for automatically verifying software solutions
  • Identify recurring model-generated coding errors
  • Develop reliable checks for engineering tasks
  • Support reproducible evaluation across repeated assignments

Software Engineering Lifecycle Evaluation

  • Evaluate model capabilities across the software-development lifecycle
  • Assess reasoning around prototyping and architecture design
  • Review API design and production implementation decisions
  • Evaluate launch, experimentation, monitoring, and maintenance scenarios
  • Identify areas where models struggle with real-world engineering workflows

Research & Benchmark Collaboration

  • Collaborate with research and cross-functional technical teams
  • Contribute to datasets used for training and benchmarking
  • Help define engineering evaluation strategies
  • Compare model performance against professional engineering expectations
  • Support iterative improvements to coding-focused evaluation systems

Ideal Profile

  • 3+ years of professional software-engineering experience
  • Strong full-stack development capabilities
  • Experience building scalable, production-grade software
  • Strong understanding of software architecture and system design
  • Deep knowledge of development, debugging, and code-quality assessment
  • Experience reviewing and improving complex software implementations
  • Proficiency in one or more of Python, JavaScript, Java, C++, Rust, or related languages
  • ReactJS, C, or Go experience may also be relevant to project assignments
  • Strong understanding of API design and production implementation
  • Familiarity with software monitoring and operational maintenance
  • Ability to reason across the complete software-engineering lifecycle
  • Strong analytical and problem-solving capabilities
  • Excellent written and verbal communication skills
  • Ability to provide clear, structured evaluation rationales
  • Comfortable collaborating remotely with research and technical teams

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote
  • Candidates must be based in the United States, Canada, or eligible Western European (WEU) countries
  • Source examples of WEU locations include Austria, Belgium, France, and Germany
  • Minimum commitment: 10 hours per week
  • Flexible workload of up to 40 hours per week
  • Initial project duration is approximately 1 month
  • Extension may be available depending on performance and project fit
  • No medical or paid-leave benefits are included under the contractor arrangement
  • Application process takes approximately 15–30 minutes
  • Completion of an AI video interview is required
  • Compensation is not specified in the source materials
  • Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Similar jobs

Apply for this job