AC

Artificial Intelligence Engineer

Hiring from
Pakistan
Work type
Remote
Posted
Oct 2, 2026
Is this job info correct?

Full-Stack AI Engineer, Creative Intelligence

AI Case Leads · Remote · Full-time or long-term contract


Before you apply: this role has three must-haves (full-stack TypeScript, PostgreSQL and backend engineering, and production AI you have measured). If you do not have all three, it is not a fit. Listing a technology in a skills section is not enough; we look for where you used it in real, shipped work.


What we are building

We are a legal-intake advertising operation. We built an internal platform that watches every ad we run and measures it against one number: cost per signed case.

A multimodal model watches each video and returns a timestamped transcript and a scene breakdown. We map viewer retention onto those scenes, tag each creative on a structured genome (hook pattern, spokesperson, claim type, emotional arc, offer framing), and join it all to signed cases from our intake CRM.

On top of that sits what we call the Brain. It diagnoses why a creative is failing and points to the part of the script where viewers leave. It reviews a brief scene by scene, proposes the next iteration, and records a dated prediction that a nightly job later settles against what actually happened.

It is real, it is running on live spend, and we do not yet know how good it is. The job is to make it good, and to build and run the platform around it.


The mandate

Make the Brain something a creative strategist trusts enough to work from every day. Not to replace them: to make one strategist as productive as five, by answering the questions they currently answer on instinct.

  • Why is this angle working in one market and dying in another?
  • Why does this hook hold, and that one lose most of the audience in three seconds?
  • What should the next iteration of this brief change, and what must it not touch?
  • Is this brief actually failing, or is the sample too thin to say?

And make the platform around the Brain fast and reliable enough that the whole team works in it, not around it.


What you would actually do

  1. Own the platform end to end. The Brain is only useful if strategists, media buyers and editors can act on what it says. You will build and run the app they use every day, the Meta, Google, TikTok and CRM syncs that feed it, and the nightly jobs behind it.
  2. Prove it. Build the evaluation layer until we can state the Brain’s accuracy with a straight face: diagnosis quality against expert judgment, verdict accuracy against realized cost per signed case, and calibration of its stated confidence. We already have a prediction ledger that settles bets nightly and throws out any bet a human contaminated. Today it reports a plain hit rate per prediction type, and nothing checks whether its confidence is calibrated. Turn it into a number leadership can read.
  3. Open the genome. Subjective attributes such as hook pattern, claim type and emotional arc are extracted but held back from the performance rollups by an agreement gate: at least 85 percent agreement between AI and human tags across 30 or more compared ads. Make extraction consistent enough to clear it.
  4. Put real statistics under it. Signed cases are rare events. At genome depth our slices have single-digit counts, and today any row below a spend, signed-case or creative-count floor is simply dropped. Replace that with partial pooling, so a thin slice borrows strength from its parent instead of disappearing and the Brain can say “probably” without lying.
  5. Decompose the diagnosis. Today it is one large prompt. Make it a system: retrieval over past briefs and outcomes, grounding verification per claim, and graceful behavior when the evidence is genuinely thin.
  6. Make it causal. Nothing runs holdouts yet; the Brain only proposes test candidates. Design and run real creative holdouts, so we can separate “this changed and things improved” from “this caused things to improve.”


What you are walking into

The codebase is large, young and heavily documented.

  • Size. About 136,000 lines of TypeScript, roughly 25,000 of them in the Brain.
  • History. One engineer built it in under three months, working closely with AI coding tools. It has a long README, dated design plans, comments that explain why, and strong conventions of its own.
  • Video. A multimodal model, currently Gemini, returns the transcript and scene breakdown in one structured call. There is no separate speech-to-text or segmentation pipeline.
  • Retention. Meta reports viewers at 25, 50, 75 and 100 percent, plus plays, ThruPlays and average watch time. There is no second-by-second curve, so the Brain estimates the drop-off line between checkpoints. Making that estimate honest is part of the diagnosis work.
  • Schema. Drizzle generates one schema file from the code and applies the difference to the database. There is no migration history, and you may want to change that.


Must have

All three, shown in real work you shipped. Everything else on this page can be learned here.

  1. Full-stack TypeScript. You ship features end to end, from schema to screen. You are strong in Next.js App Router and React (we are on Next 16 and React 19): server components, server actions, streaming responses. You care how the result feels to the person using it. This is not a Python role.
  2. PostgreSQL and backend engineering. Analytical SQL (window functions, CTEs, aggregation across a deep dimensional model) and schema design, not only queries. Background and scheduled jobs, long-running pipelines, and a containerized app in production. The Brain is a data modeling problem wearing an AI costume.
  3. Production AI, measured. You have shipped LLM features real people use: structured JSON output, schema validation with Zod or similar, tool calling, retries, cost control, and retrieval where every claim traces to a citation. And you measured them, with golden sets, LLM-as-judge or regression checks. You can tell us the accuracy of something you built and exactly how you measured it.

One baseline on statistics: you know why a 40 percent conversion rate on 5 observations means nothing. The deeper methods are a strong plus, below.


Strong plus

  • Statistics for sparse data and experiments. Hierarchical or empirical Bayes, shrinkage and partial pooling; holdouts, power and significance. There is no PyMC or Stan in TypeScript, so pooling here is built by hand: beta-binomial for rates, gamma-Poisson for counts. If you have done this in Python or R, the math carries over, and Python is fine for offline analysis. The Brain itself stays in TypeScript.
  • Vector search. pgvector with HNSW indexes and cosine similarity. Our embeddings, transcript lines, creative segments and insight clusters all live in Postgres.
  • Model-agnostic habits. We route through OpenRouter across Gemini, GPT, Claude and Grok, and switch when one gets better or cheaper.
  • Ad platform APIs. Meta Marketing API, Google Ads API, TikTok Marketing API. Conversions API and offline conversion uploads help too; we do not send them yet.
  • Background that maps directly. Creative intelligence or ad-analytics products, growth engineering at a heavy-spending brand or agency, experimentation platforms, mobile UA tooling, or lead-gen, insurance or legal marketplaces.
  • Creative teams. You have worked alongside them and understand exactly why they distrust dashboards.
  • Startup ownership. You were trusted with significant technical decisions, not only tickets.

Not a fit

  • Python-only or data-science-only engineers
  • Frontend-only or backend-only engineers
  • Research or academic AI without production software
  • AI work limited to chat wrappers or simple demos

Our stack

TypeScript end to end on Bun. Next.js 16 and React 19, with Tailwind, shadcn, and TanStack Query and Table. PostgreSQL with pgvector, Drizzle ORM, and Supabase auth with role-based access. OpenRouter for models. Docker on Railway, with nightly jobs run by a scheduler.

Beyond the three must-haves, experience with any specific piece is a plus, not a requirement.

Not required

  • A marketing background. We will teach you the domain.
  • A PhD. We care what you have shipped and whether you measured it.


How we work

Small team. Direct access to the person spending the money. Everything is judged against cost per signed case.

Every change ships end to end: schema, query, server action, access rules and the screen, checked in a browser against real data. You will work across the whole stack, often in a single change.

There is a rule in this codebase we do not break: the Brain proposes, a person decides, and it never states a number or a quote it cannot cite.

You will read a lot of someone else’s documented work before you write your own. Bring that kind of temperament.


To apply

Send a short note with:

  1. One product or feature you built end to end, from database to UI
  2. One AI feature you shipped, with its accuracy and exactly how you measured it
  3. One time your own evaluation caught your system being confidently wrong, and what you changed
  4. A link to a TypeScript code sample or repository you wrote

Skip the cover letter. We read every note and reply to all of them.

Similar jobs

Apply on LinkedIn