PI

Staff Applied AI Engineer

People In AIApplies on LinkedInData & Analytics
Salary
$240K–$275K
USD
Moves you to
United States
Support
Visa sponsorship
Posted
Sep 26, 2026
Is this job info correct?

Staff Applied AI Engineer, Agents & Model Adaptation


Compensation: $240,000-$275,000 + very strong equity

Location: New York, NY


Company: An early-stage AI startup building intelligent systems to identify complex fraud and risk across the insurance ecosystem.


We’re partnering with a venture-backed AI company using advanced data and machine learning to tackle sophisticated fraud across property and casualty insurance. By connecting signals across multiple parts of the ecosystem, the team is helping uncover complex networks, improve decision-making, and automate investigative workflows that have traditionally required significant manual effort.


The Role

This is a senior individual-contributor role for an engineer who wants to operate at the intersection of LLM research, model adaptation, agentic systems, and production AI infrastructure.


You’ll own key architectural decisions around which frontier and open-weight models should power different workflows, how those models are adapted to domain-specific tasks, and how agent performance is evaluated in production.


A major focus will be solving the trade-off between the reasoning quality of frontier models and the latency, cost, privacy, and scalability benefits of smaller open-weight models.


You’ll remain deeply hands-on while setting the technical direction for the wider AI engineering function.


What You’ll Do

  • Own the base-model strategy across AI workflows, evaluating frontier and open-weight models and driving decisions around when to adopt, replace, or adapt them.
  • Fine-tune open-weight LLMs using techniques such as LoRA, QLoRA, supervised fine-tuning, and preference optimization / DPO.
  • Deploy adapted models as production-grade components within agentic workflows.
  • Build a continuous data flywheel, transforming human corrections and production agent traces into training data, synthetic datasets, evaluation cases, and future fine-tuning runs.
  • Develop evaluation frameworks that benchmark models and agents against realistic, domain-specific scenarios.
  • Measure and optimize performance across accuracy, hallucination rates, false positives, cost, and latency.
  • Design and productionize agent infrastructure covering tool use, context construction, orchestration, structured outputs, guardrails, and observability.
  • Implement eval-gated deployment processes and full traceability across production agent runs.
  • Own serving infrastructure for self-hosted models, including batching, quantization, latency optimization, and inference cost management.
  • Red-team models and agent systems for issues including prompt injection, data leakage, and adversarial manipulation.
  • Set a high technical bar through architecture reviews, mentorship, and clear written technical direction.


What You’ll Bring

  • Experience across software engineering, applied machine learning, or related disciplines.
  • Experience acting as a technical leader and making architectural decisions for production ML or LLM systems.
  • Deep hands-on experience fine-tuning large language models, with production results you can discuss in detail.
  • Experience building production AI agents that use external tools and generate structured outputs.
  • Strong experience developing evaluation systems that meaningfully influence model or product decisions.
  • Advanced proficiency with Python and PyTorch.
  • Familiarity with modern LLM training and distributed-compute tooling.
  • Comfort working in ambiguous environments where ground truth is imperfect and existing research does not provide an off-the-shelf solution.


Bonus experience includes:

  • Shipping quantized or distilled models into latency- or cost-constrained environments.
  • Diagnosing and addressing accuracy regressions caused by compression or serving optimizations.
  • Entity resolution, graph ML, or classical ML for highly imbalanced detection problems.
  • Agent benchmarking datasets and designing domain-specific evaluation suites.


Tech Stack

  • Python
  • PyTorch
  • Hugging Face ecosystem
  • LoRA / QLoRA
  • Supervised fine-tuning
  • DPO / preference optimization
  • vLLM, TGI, Triton, or equivalent serving infrastructure
  • DeepSpeed, Ray, or equivalent distributed training tooling
  • Open-weight and frontier LLMs
  • Agent orchestration and tool-use frameworks
  • Evaluation and observability infrastructure


Why Join?

You’ll have unusually broad ownership across the AI stack from model selection and fine-tuning through agent architecture, evaluation, serving, and safety.


Rather than simply wrapping third-party APIs, you’ll be solving technically difficult questions around how reasoning-capable models can be adapted and deployed efficiently at significant production scale.

You’ll join an early-stage environment where technical decisions have immediate product impact, work directly on a challenging real-world problem, and help establish the foundations and engineering standards for the company’s AI platform.

The package includes company equity, 100% medical coverage plus dental and vision, and potential visa-transfer sponsorship for eligible candidates currently based in the U.S.

About People In AI

People In AI is a specialist recruitment partner connecting exceptional AI, machine learning, and data talent with some of the most ambitious technology companies in the market.

We work closely with candidates and businesses across the AI ecosystem, from early-stage startups to established technology organizations, helping build teams working on genuinely impactful problems.

Similar jobs

Apply on LinkedIn