Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
ST

Founding GPU Kernel Engineer

SF Tensor
Posted 1 weeks ago
📦Relocation support
🇺🇸United States
💰$285.0K–$315.0K📁Engineering & Development
Is this job info correct?

About SF Tensor At The San Francisco Tensor Company, we believe the future of AI and high-performance computing depends on rethinking the entire software and infrastructure stack. Today's developers face bottlenecks across hardware, cloud, and code optimization that slow progress before ideas can reach their full potential. Our mission is to remove those barriers and make compute faster, cheaper, and universally portable. We are building a Kernel Optimizer that automatically transforms code into its most efficient form, combined with Tensor Cloud for adaptive, cross-cloud compute and Emma Lang, a new programming language for high-performance, hardware-aware computation. Together, these technologies reinvent the foundations of AI and HPC. SF Tensor is proudly backed by Susa Ventures and Y Combinator, as well as a group of angels including Max Mullen and Paul Graham as well as founders and executives of NeuraLink, Notion and AMD. We are partnering with researchers, engineers, and organizations who share our belief that the next breakthroughs in AI require breakthroughs in compute. About the Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between hardware and software. Someone who thinks in warps, occupancy, and memory hierarchies, and can squeeze every last FLOP out of a GPU. Your job is to go deeper than anyone else. You'll hand-tune kernels to figure out what's actually possible on the hardware, and then turn that knowledge into compiler optimization passes that help every model we compile. What You'll Do Write and hand-optimize GPU kernels for ML workloads (matmuls, attention, normalization, etc.) to set the performance ceilings Profile at the microarchitectural level: look into SM utilization, warp stalls, memory bank conflicts, register pressure, instruction throughput Debug performance issues by digging deep into things like clock speeds, thermal throttling, driver behavior, hardware errata Turn your hand-optimization insights into automated compiler passes (working closely with our compiler team) Develop performance models that predict how kernels will behave across different GPU architectures Build tools and methods for systematic kernel optimization Work with NVIDIA, AMD, and emerging AI accelerators - understand the common parts and what's vendor-specific What We're Looking For Deep expertise in GPU architecture Proven track record of hand-writing kernels that match or beat vendor libraries (cuBLAS, cuDNN, CUTLASS) Strong skills with low-level profiling tools: Nsight Compute, Nsight Systems, rocprof, or equivalents Experience reading and reasoning about PTX/SASS or GPU assembly Solid systems programming in C++ and CUDA (or ROCm/HIP) Good understanding of how high-level ML operations map to hardware execution Experience with distributed training systems: collective ops like all-reduce and all-gather, NCCL/RCCL, multi-node communication patterns Nice to Have HPC background: experience with large-scale scientific computing, MPI, or work in supercomputing Background in electrical engineering, computer architecture, or hardware design Driver development experience (NVIDIA, AMD, or other accelerators) Experience with MLIR, LLVM, or compiler backends Deep knowledge of distributed ML training: gradient accumulation, activation checkpointing, pipeline/tensor parallelism, ZeRO-style optimizations Familiarity with custom accelerators: TPUs, Trainium, Inferentia, or similar Knowledge of high-speed interconnects: NVLink, NVSwitch, InfiniBand, RoCE Publications or contributions in GPU optimization, HPC, or ML systems Experience at NVIDIA, AMD, a national lab, or an AI hardware/infrastructure company Why Join Us This role is for someone who wants to know why things are fast or slow on the hardware. You'll have a direct impact on the performance of large-scale AI training, tackling problems that need real depth. If you've ever been annoyed that your hard-won optimization knowledge is stuck in your head and not baked into a compiler, here's your shot to change that. We believe in the power of in-person collaboration to solve the hardest problems and foster a strong team culture. We offer relocation assistance and look forward to you joining us in our San Francisco office. The base salary range for this full-time position is $285,000 - $315,000 + bonus + equity + benefits.

Similar jobs

Similar jobs

Etched logo

Applied AI Engineer, Kernel Performance

Etched

🇺🇸United States1 weeks ago
River AI Inc. logo

Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)

River AI Inc.

🇺🇸United StatesJun 29, 2026, 2:06 AM UTC
San Francisco Tensor Company logo

Founding GPU Kernel Engineer

San Francisco Tensor Company

🇺🇸United StatesMay 28, 2026, 3:58 AM UTC
Zyphra logo

Research Engineer - AI Performance & Kernel Optimization

Zyphra

🇺🇸United StatesMay 28, 2026, 3:33 AM UTC
Etched logo

Kernel Driver Software Engineer

Etched

🇺🇸United StatesMay 27, 2026, 10:39 PM UTC
Thinking Machines Lab logo

Research Engineer, Infrastructure, Kernels

Thinking Machines Lab

🇺🇸United StatesMay 27, 2026, 8:53 PM UTC