Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Archetype Ai logo

World Model Evaluation Lead

Archetype Ai
Posted May 27, 2026, 10:02 PM UTC
🇺🇸United States🏠Remote📁Data & Analytics
Is this job info correct?

About A rchetype AI Archetype AI is developing the world's first AI platform to bring AI into the real world. Formed by an exceptionally high-caliber team from Google, Archetype AI is building a foundation model for the physical world, a real-time multimodal LLM for real life, transforming real-world data into valuable insights and knowledge that people will be able to interact with naturally. It will help people in their real lives, not just online, because it understands the real-time physical environment and everything that happens in it. Supported by deep tech venture funds in Silicon Valley, Archetype AI is currently at the Series A stage and is progressing rapidly to develop technology for their next stage. This presents a unique and once-in-a-lifetime opportunity to be part of an exciting AI team at the beginning of their journey, located in the heart of Silicon Valley. Our team is headquartered in San Mateo, California, with team members throughout the US and Europe. We are actively growing, so if you are an exceptional candidate excited to work on the cutting edge of physical AI and don’t see a role that exactly fits you below you can contact us directly with your resume via jobs archetypeai io. About the Role Archetype AI is seeking a hands-on Evaluation Lead to build and assess model performance for physical AI. You will design and implement advanced evaluation techniques for assessing the strengths and weaknesses of real-world AI models, and build and scale evaluation frameworks to rapidly test and generate reports on model performance. Responsibilities include partnering closely with research and engineering teams to develop evaluation methodologies, analytically assessing and improving test datasets, uncovering model weaknesses or risks, and tracking competitive industry benchmarks. This is a high-impact role for someone who thrives in a fast-paced AI environment and wants to directly influence our path as we scale our AI technologies and business. Core Responsibilities Drive Benchmarking & Evaluation Design and implement rigorous evaluation methodologies and benchmarks for measuring model effectiveness, reliability, alignment, and safety Lead evaluation of model performance, ranging from offline experiments to full production model testing Build & Scale Evaluation Frameworks Design and oversee the pipelines, dashboards, and tools that automate model evaluation Design and oversee tools for A/B model testing, regression testing, and production model performance Lead Evaluation Strategy Develop and implement strategies for evaluating physical AI models that can scale across a broad range of real-world use cases, sensor types, and edge cases Plan, run, and oversee evaluations, across internal teams and external customers Drive edge case discovery, red-teaming, safety, privacy, and risk evaluation - feeding back knowledge to key stakeholders in research and engineering teams Key Requirements Extensive expertise in evaluating AI and machine learning models, ideally in physical AI or a related AI field Experience in designing, implementing, and refining evaluation metrics Deep understanding of machine learning, AI, and generative models Excellent python and software engineering skills Experience designing and building scaleable data pipelines and evaluation tools Experience collaborating closely with key stakeholders from research, engineering, and product teams Strong communication and documentation skills, with a bias for creating detailed evaluation reports that help drive model performance Startup-ready mindset with the ability to thrive in high-velocity, high-ambiguity environments Minimum Qualifications Extensive expertise in evaluating AI and machine learning models, ideally in physical AI or a related AI field Experience in designing, implementing, and refining evaluation metrics Deep understanding of machine learning, AI, and generative models Excellent python and software engineering skills Experience designing and building scaleable data pipelines and evaluation tools Experience collaborating closely with key stakeholders from research, engineering, and product teams Strong communication and documentation skills, with a bias for creating detailed evaluation reports that help drive model performance Startup-ready mindset with the ability to thrive in high-velocity, high-ambiguity environments What We Would Love To See Experience evaluating real-world, real-time algorithms Experience evaluating a broad range of sensor types, such as cameras, LIDAR, physical sensors, RF sensors, and beyond A strong scientific approach to evaluation and understanding model performance Experience in evaluating production algorithms Experience building and curating data campaigns to create extensive test datasets Experience managing internal teams and/or external vendors

Similar jobs

Similar jobs

Takeda logo

Senior Director, Search and Evaluation Lead, Global Oncology

Takeda

🇺🇸United StatesYesterday
PCI GS logo

Data and Evaluation Lead

PCI GS

🇺🇸United States1 weeks ago
VA

Requirement, Evaluation, and Design (RED) Team Lead

ValidaTek

🇺🇸United States1 weeks ago
elly logo

AI Evaluation Lead

elly

🇺🇸United States1 weeks ago
Motional logo

Principal Engineer Tech Lead, Embodied AI & Off-Board Performance Evaluation

Motional

🇺🇸United States3 weeks ago
Fintentional Ai logo

AI Evaluation Lead

Fintentional Ai

🇺🇸United StatesJun 25, 2026, 9:09 AM UTC