TR
Hiring from
Probably Worldwide
Work type
Remote
Posted
Sep 27, 2026
Is this job info correct?
You'll work solving real and complex AI problems alongside engineering, product, and client teams. We build AI agent systems that run in production and hold up there—bridging the gap between an LLM prototype and a reliable, mission-critical service.

About the role
Most of this work is not about the model—it is about everything around it: getting an agent to reach real systems safely, recover from tool failures, respond fast, and run cost-effectively. You will take full ownership of client systems, making key architectural calls and building high-reliability backend services.
Key focus areas include:
  • Designing request classification, model planning, tool permissions, and multi-turn state management.
  • Connecting agents to internal APIs, databases, and third-party tools via typed, versioned contracts, using Python async and streaming (SSE/WebSockets).
  • Building golden datasets, offline scoring, traces, and online experiments—ensuring performance is driven by metrics (p95 latency, cost per task, failure rates), not intuition.
  • Enforcing scoped permissions, confirmation gates for irreversible actions, and defense against prompt injection attempts.
What we're looking for
  • Extensive Python backend experience, with async as an everyday tool and strong experience designing service contracts.
  • Proven track record taking at least one LLM system to production and maintaining it.
  • Hands-on work with tool calling, multi-step workflows, and handling failure paths (timeouts, retries, bad arguments, hallucinated parameters).
  • Standard reliability engineering applied to non-deterministic runtimes: circuit breakers, backoff strategies, caching, and concurrency limits.
  • Discipline to build systematic measurement layers (evals, tracing) before deploying features.
Nice to have
  • Production experience with streaming architectures (SSE/WebSockets) and frontend contract design.
  • Hands-on use of agent frameworks (LangGraph, LangChain) or protocols like Model Context Protocol (MCP).
  • Experience with LLM tracing/eval tools, vector databases, hybrid search, and chunking/ingestion pipelines.
  • Quantifiable success in reducing LLM latency or token costs.
  • Core backend stack: PostgreSQL, Redis, Docker, AWS, CI/CD, OpenTelemetry, pytest.

Similar jobs

Apply for this job