Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Thinking Machines Lab logo

Software Engineer, ML Infra

Thinking Machines Lab
Posted 3 hours ago
📦Relocation support🛂Visa sponsorship
🇺🇸United States
💰$350.0K–$475.0K
📁
Engineering & Development
Is this job info correct?

About Thinking Machines The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. About the Role We're hiring a Software Engineer to sit at the day-to-day interface between research and infrastructure. This is a generalist role with broad scope: you'll be one of the people always in the room for infra decisions on the ML systems side, with a clear enough view of upcoming compute needs to see support burden coming before it arrives. You'll also be one of the people who babysits hero runs — the ones at 2am when a 4k-GPU job hits a weird Xid and someone needs to decide, quickly and correctly, whether to drain the node, restart the job, or escalate to NVIDIA before the run loses checkpoints. That kind of judgment, built across the kernel, the network, the scheduler, and the application layer, is the core of the job. What You'll Do Debug across the full stack — kernel, NCCL, scheduler, application, and telemetry — often in the same afternoon, to find root causes that don't show up in any single layer Provide embedded, hands-on support during hero runs and major incidents, staying with a problem until it's genuinely resolved Serve as the front door for researchers when something's broken and it isn't obvious who owns it Lead postmortems and build the tooling that prevents the next incident — acting as a force multiplier, not just a responder Mentor other engineers into this kind of cross-stack breadth Skills & Qualifications Minimum Qualifications Credible, hands-on competence in 4 or more of the following: Linux kernel, networking, GPUs / CUDA, distributed systems runtimes, storage, compilers / language runtimes, observability internals — at this level, this required range is the qualifying signal Comfort operating without a clearly defined owner, and the judgment to know when to dig in yourself versus when to escalate Preferred Qualifications Track record of being the person other engineers escalate to, across 2+ companies Shipped meaningful contributions in 3+ distinct technical stacks Track record of leading major incidents where the root cause was non-obvious Researcher-facing comfort: can talk to a researcher about their workload without making them feel dumb, and can tell them no when the right answer is no Experience operating at the scale of frontier training or inference clusters, and appetite for owning the hardest, least-defined problems in that stack Logistics Location: This role is based in San Francisco, CA. Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 – $475,000 USD, plus equity. Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together. Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar jobs

Similar jobs

Allen Institute logo

Senior Software Engineer - AI / ML Platforms and Infrastructure

Allen Institute

🇺🇸United StatesJul 21, 2026, 9:20 AM UTC
Preference Model logo

Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model

🌍Canada, United StatesJul 16, 2026, 9:14 PM UTC
Merck logo

Scientist, Engineering

Merck

🇺🇸United States2 hours ago
Merck logo

Vence Collar Tech Lead

Merck

🌍United Kingdom, United States2 hours ago
Schreiberfoods logo

Controls Engineering Intern (Summer 2027)

Schreiberfoods

🇺🇸United States3 hours ago
Schreiberfoods logo

Senior Manufacturing Engineer

Schreiberfoods

🇺🇸United States3 hours ago