Our client is building an enterprise LLM gateway platform - an OpenAI-compatible API layer that routes traffic to multiple model providers while cutting inference cost, enforcing policy, and giving customers full visibility into what their AI workloads actually cost. It serves engineering teams that run LLM workloads in production under real budget, compliance, and governance constraints.
We are looking for a Senior Full-Stack Engineer who is comfortable owning features end to end - from the PostgreSQL schema and the FastAPI endpoint through to the frontend screen the customer actually clicks. This is a single combined role, not a backend seat with occasional frontend work: most features in this product span the gateway middleware, the management API, and the dashboard, and you will be expected to ship all three layers.
Two Areas Will Be Your Primary Focus
Payments and billing - building the commercial layer on top of the platform's usage metering: subscription plans, usage-based billing, invoicing, credits, and entitlement enforcement.
Gateway guardrails and telemetry - building the policy enforcement and observability layer that makes the gateway safe, measurable, and auditable for enterprise customers.
You will join a small, senior, distributed team where engineers own their features through to production, take part in architectural decisions, and have real influence over how the platform evolves.
Real LLM infrastructure work - middleware in the request path of live model traffic, where a good optimization directly saves customers money
Measurable impact: every feature is expressed in tokens saved, cost avoided, or requests correctly handled
Substantial ownership of two significant domains - billing and guardrails - rather than incremental work on someone else's design
Multi-provider LLM work across the major model vendors, plus prompt optimization and cost/quality evaluation
Modern stack with architectural freedom: Python 3+, FastAPI, React, TypeScript, Kubernetes, Terraform
Responsibilities
Team Composition:
Own features end to end across the gateway middleware, backend API, and frontend - including schema design, API contracts, UI, tests, and rollout.
Build and operate the subscription and usage-based billing system, including payment provider integration, webhook reliability, usage-to-invoice mapping, and entitlement enforcement.
Design and implement guardrails that inspect and enforce policy on LLM prompts and completions, with per-tenant configuration and full auditability.
Extend the telemetry pipeline and analytics aggregations that power customer-facing cost and usage reporting, and add tracing, metrics, and alerting across services.
Write middleware that runs in the live request path, treating latency, failure modes, and graceful degradation as first-class design concerns.
Design multi-tenant data models with correct isolation between organizations, workspaces, and keys.
Write meaningful automated tests for the code you ship.
Participate in architecture decisions, code reviews, and planning; document decisions for a distributed team.
Handle sensitive data responsibly - secret encryption at rest, least-privilege access, and audit logging.
Qualifications
4+ years of professional full-stack experience, with genuine depth on both sides - you have owned production backend services and shipped non-trivial frontend features.
Strong Python skills and experience building production APIs with FastAPI (or Django REST / Flask, with a willingness to work in FastAPI), including async programming.
Strong React and TypeScript skills - comfortable with component-driven UI, SPA state management, routing, and integrating REST APIs including streamed responses.
Payment system integration experience - Stripe strongly preferred, or a comparable provider (Paddle, Adyen, Braintree, Chargebee, Recurly). You have dealt with webhooks, idempotency, subscription lifecycles, proration, refunds, or usage-based/metered billing in production.
Solid PostgreSQL skills: schema design, migrations, indexing, and query optimization - including aggregations over large event/log tables.
Experience with multi-tenant architectures, authentication and authorization, and tenant data isolation.
Comfortable with Docker and running services in Kubernetes; able to read a Helm chart and debug a failing pod.
Experience with Redis or a comparable cache, and with designing for cache correctness and invalidation.
Experience using AI-powered development tools (Cursor, VS Code with Copilot, or similar) and LLMs for research and problem-solving.
Strong ownership and autonomy: able to take an ambiguous problem, propose an approach, and deliver it with minimal oversight.
Clear written and spoken English for collaboration and documentation.