Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Realm Labs logo

Software Engineer, ML Infrastructure

Realm Labs
Posted May 27, 2026, 9:28 PM UTC
🛂Visa sponsorship
🇺🇸United States
📁Engineering & Development
Is this job info correct?

Role Overview We are hiring a Founding ML Infrastructure Engineer to own the end-to-end deployment, optimization, and operation of our suits of models in production . This is a core founding role focused on building and operating production-grade LLM systems . You will apply deep knowledge of model internals to deploy, optimize, and run modern LLMs at scale , owning performance end-to-end across latency, throughput, and reliability . You will design and operate the full ML serving stack from model artifacts to GPU execution, and work closely with Product and ML teams to ensure our models can support high QPS, strict SLAs, and production correctness . This role is ideal for someone who deeply understands how LLMs work internally , but chooses to specialize in making them fast, stable, and production-ready . About Realm Labs Realm Labs is an AI trust and security startup. We help enterprises detect, debug, and prevent AI’s misbehaviors in production. We are backed by top VCs and serve some of the most iconic global enterprises. Key Responsibilities Own the end-to-end LLM inference stack , including: Model loading and execution GPU utilization and memory efficiency Runtime performance tuning Production deployment and scaling Design and operate high-performance LLM serving systems using technologies such as: vLLM, TensorRT / TensorRT-LLM, Triton Inference Server, SGLang Optimize inference across: Latency Throughput (QPS) GPU memory footprint Cost efficiency Work hands-on with PyTorch and TensorFlow models , including: Model graph understanding Attention mechanisms, KV cache behavior, batching strategies Precision tradeoffs (FP16, BF16, INT8, etc.) Build and maintain production-grade GPU services : Multi-model serving Autoscaling strategies Fault isolation and graceful degradation Collaborate with application and platform teams to: Define serving APIs Ensure correctness and safety of outputs Debug production issues end-to-end Build a reproducible model training and versioning system for customer deployments Establish best practices for: Model versioning Rollouts and rollbacks Performance benchmarking Production validation Expected Qualifications 5+ years of professional experience in ML infrastructure, systems engineering, or production ML roles. Strong software engineering fundamentals; ability to write robust, maintainable production code . Deep hands-on experience with LLM inference infrastructure , including: PyTorch (required) TensorFlow (working knowledge) Proven experience with GPU inference optimization , including: TensorRT / TensorRT-LLM vLLM Triton Inference Server SGLang or similar serving runtimes Strong understanding of LLM internals , such as: Transformer architectures Attention and KV caching Batching, streaming, and token-level generation Experience running ML systems in production with high traffic and SLAs Comfortable working in Linux-based, cloud production environments Preferred Qualifications Experience deploying LLMs on Kubernetes and GPU clusters. Familiarity with CUDA, NCCL , or low-level GPU performance concepts. Experience with: Model sharding and parallelism strategies Multi-GPU inference Streaming inference systems Knowledge of observability for ML systems (metrics, latency breakdowns, GPU monitoring). Experience working at startups or owning systems with minimal abstraction layers. Additional Information This is a founding, high-ownership role with direct impact on core product capabilities. You will be expected to build, run, and own systems end-to-end . The role may include limited on-call responsibilities aligned with production ownership. Compensation & Benefits Market aligned compensation and benefits Founding engineer equity ( Equity is a significant component of this role and will be discussed) Medical, Dental, Vision, Life insurance, 401-K, In-office lunch etc. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and candidate. But if we make you an offer, we will make all reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. Compensation The base pay range for this role is $180,000 – $250,000 per year.

Similar jobs

Similar jobs

Alleninstitute logo

Senior Software Engineer - AI / ML Platforms and Infrastructure

Alleninstitute

🇺🇸United States1 weeks ago
Spellbrush logo

HPC/ML Infrastructure Engineer

Spellbrush

🌍Japan, United StatesJun 15, 2026, 8:22 AM UTC
Gray Swan logo

Software Engineer, Infrastructure

Gray Swan

🇺🇸United States5 hours ago
Gray Swan logo

Software Engineer

Gray Swan

🇺🇸United States5 hours ago
The Fountain Group logo

Manufacturing Engineer

The Fountain Group

🇺🇸United States5 hours ago
BeaconFire Inc. logo

Junior AI Developer- Python

BeaconFire Inc.

🇺🇸United States5 hours ago