Harness / AI Orchestration - Own the Core of the Platform
Tokyo, Japan (Kanda Ogawacho) · Full Remote (Japan) · Hybrid encouraged · Full-time
About the Company
I'm recruiting on behalf of my client, an AI research and product development company on a mission to build Japan's first general-purpose AI. Rather than offering one massive model uniformly to everyone, they're building a "ubiquitous AGI" that grows and adapts to each individual user. Alongside this, they're building global consumer products that leverage that technology. They recently closed one of Japan's largest funding rounds to date, and through a capital and business alliance with a major partner, they're accelerating both their research and product development.
About the Role
It's not the model itself but the harness around it — the orchestration layer, primitives, agent loops, tool calling, and sandboxing — that determines how an AI product actually performs in the real world. This role exists to turn a great model into a great product experience.
This is a Staff-level engineering role with ownership of the entire harness/orchestration layer behind the company's AI product platform. You'll help define the contract between backend and orchestration, design reusable primitives, and build the agent-execution foundation — setting technical standards for a platform built to serve millions of users globally. You'll also judge how much of the open-source harness/agent ecosystem to adopt, deciding what to build, what to borrow, and what to leave behind.
Responsibilities
- Lead architecture and technology selection for the orchestration/harness layer; define the contract between backend and orchestration and design an extensible, maintainable foundation
- Design and share reusable primitives (service engine, scale-out, tool calling/discovery, sandbox) as a platform product teams can freely compose
- Evolve product-relevant primitives (agent spawning, dynamic tool loading, long-running tasks, conversation session state) in a product-driven way
- Ensure scalability and operability for a large concurrent user base — resource isolation, usage controls, failover, horizontal scaling — while staying cost-efficient at scale
- Validate technical decisions through prototypes and quantitative ROI assessment before committing
- Drive standardization and automation of CI/CD, dev processes, and shared infrastructure org-wide
- Mentor engineers and elevate technical standards and culture through code review
Tech Stack
Go / Python / TypeScript / Rust / AWS / Datadog · LLM orchestration / agent frameworks / MCP (Model Context Protocol) / sandboxing / distributed systems
Tools: GitHub / Slack / Notion / Figma · AI tools: Claude Code / Cursor / ChatGPT / Gemini / Grok
Required Skills
- 8+ years of overall software development experience
- Deep expertise in designing, developing, and operating distributed systems at large scale
- Experience integrating LLMs into products (designing/implementing orchestration layers, agent loops, tool calling)
- Development experience in performance-critical languages such as Go, Rust, or Python
- Strong system design skills accounting for large-scale concurrency, scalability, and high availability
- Leadership skills to drive technical consensus across complex stakeholders and deliver high-uncertainty projects
- Ability to instill AI-native quality standards, holding AI-generated code to the same or higher bar as human-written code
- Business-level or higher Japanese and English conversational ability
Nice-to-Have
- Experience designing/building agent frameworks, LLM orchestration platforms, or inference pipelines
- Experience implementing MCP, tool discovery, or sandboxed execution environments
- Experience building on or extending open-source agent/harness assets
- OSS contributions or upstream project involvement
- Experience leading re-architecture or migration of systems processing billions of requests
- Experience leading AI-native engineering practices (spec-driven design, agentic workflows)
Work Style
- Full remote work from anywhere in Japan welcome; in-office attendance recommended Mondays/Fridays, with mandatory monthly All-Hands (travel expenses covered)
- Discretionary work system for specialists, or full flextime (standard 8-hr day, 45 hrs fixed overtime included, additional OT paid separately)
- 3 months trial period
Benefits
- Full flextime / discretionary work system, hybrid flexibility
- Relocation support for moving to Tokyo, device of choice, gadget subsidy
- Company-wide AI tool adoption support
- Book purchase & online course (MOOCs) subsidies, overseas conference support, English/Japanese learning support, access to select academic lab programs
- Team meal & offsite support, monthly All-Hands, free office drinks/snacks
- Egg freezing support (planned), student loan repayment support (planned), defined contribution pension
- Babysitting/sick child care subsidies, family care leave, parental leave with flexible hours, optional 1:1 check-ins during leave
Compensation
Negotiable, with stock option program available.
Reach out and I'll walk you through the details.