Senior Data & LLM Engineer (AI‑Ready Data & Agentic Systems) Company overview: Blue Orange Digital is a boutique data & AI consultancy that delivers enterprise-grade results. We design and build modern data platforms, analytics, and ML/AI Agent solutions for mid‑market and enterprise clients across Private Equity, Financial Services, Healthcare, and Retail. Our teams work with technologies like Databricks, Snowflake, dbt, and the broader Microsoft ecosystem to turn messy, real-world data into trustworthy, actionable insight. We’re a builder‑led, client‑first culture that prizes ownership, clear communication, and shipping high‑impact work. Note: Please submit your resume in English, as all application materials must be in English for review and consideration. Position overview Blue Orange Digital is seeking a Senior Data Engineer with genuine agentic systems depth for a full-time, 6-month embedded engagement in a private-capital data environment. This is a deliberately merged role: roughly half senior data engineering — Snowflake, dbt, Airflow, AWS, Python ingestion — and half senior LLM/agentic engineering, spanning production agents, MCP servers, evaluation practice, and observability. You will work on two core initiatives: an AI-Ready Data Catalog & Context Layer built alongside the Data Analytics team, and Unstructured Data Extraction at scale, with agentic implementation, evaluation, and orchestration as the throughline across both. This is not a pipeline-maintenance seat; we are looking for someone who has already shipped agents into production and can tell us candidly what broke. Responsibilities Agentic development (primary focus) Build production agents and multi-step agentic workflows, and design and implement MCP (Model Context Protocol) servers and tool interfaces as the primary integration surface between agents and the data platform — including custom internal MCP servers over proprietary datasets. Make deliberate architectural calls on agent vs. deterministic pipeline, single-agent vs. multi-agent decomposition, and where human-in-the-loop checkpoints belong. Agent evaluation and observability Stand up an evaluation practice from scratch — eval datasets, LLM-as-judge, regression testing — and define and track the metrics that matter: task success rate, faithfulness, extraction accuracy, tool-call correctness, latency, and cost per run. Instrument tracing and observability across agent runs. Agentic orchestration Own agentic workloads in Airflow 3.x on Astronomer — modeling LLM and agent calls as named, independently retriable tasks, using dynamic task mapping for fan-out/fan-in, assets and data-aware scheduling for event-driven triggers, deferrable operators for long-running calls, and human-in-the-loop operators for approvals. Design the workloads that do not belong in a DAG (long-running or interactive agents on AWS ECS/Fargate, queue- and event-driven execution, API-triggered services) and articulate clearly where the orchestration boundary sits and why. AI-ready catalog and semantic context layer Model and implement semantic definitions over a Snowflake + dbt estate — metric definitions, entity relationships, business glossary, and ownership — so agents receive governed meaning rather than raw table names. Build and maintain Snowflake Semantic Views and Cortex Analyst semantic models, keeping them in sync with dbt as the source of truth, and establish how context is versioned, tested, and promoted. Metadata and governance instrumentation Capture lineage, freshness, column-level descriptions, classification and sensitivity tags, and usage signals; automate description and tag generation with LLMs where it is safe to do so, behind human review gates. Evaluate catalog and metadata tooling and deliver a clear build-vs-buy recommendation. Unstructured data extraction Design extraction pipelines for PDFs, decks, spreadsheets, emails, and web content, including layout-heavy financial and operational documents where tables and nested structure matter. Build schema-constrained extraction using structured outputs and tool calling with frontier LLM APIs and Cortex LLM functions, with confidence scoring, citation back to source page, and explicit handling of low-confidence fields. Implement chunking, embedding, and retrieval strategies serving both RAG and agentic retrieval, and design the human review loop — what auto-accepts, what escalates, and how corrections flow back into eval datasets. Senior data engineering and architecture Define and implement architectural solutions that ensure the delivery speed and scalability of data pipelines, including tiered/medallion structures; drive best practices for advanced ELT, data governance, and data quality; and define and enforce data security and access-control policies across pipelines and platforms. Build and maintain Python ingestion pipelines (dlt and similar) across REST APIs, SaaS connectors, file drops, and incremental loads into Snowflake, and contribute to dbt models, tests, and documentation — treating dbt metadata as a first-class input to the context layer. Production operations for non-deterministic systems Build in the operational basics agents usually lack — idempotency, retry semantics for non-deterministic steps, budget and token caps, circuit breakers, structured logging, and audit trails — under Git-based workflow discipline with code review, CI/CD, and environment promotion. Technical leadership and client-facing consulting Serve as a technical leader and key client-facing resource: mentor peers, present complex architectural decisions, write designs down before building them, work independently against ambiguous problem statements, and operate effectively as a contractor embedded in an existing engineering team — asking for context early and communicating status without being asked. Requirements 7+ years of hands-on industry experience in data engineering or software engineering, with at least 5 years on major cloud platforms. Bachelor's degree or higher in Computer Science, Engineering, or a related technical field (or equivalent experience). Agentic systems (must-have, weighted heaviest): demonstrated experience shipping LLM-powered or agentic systems to production, with concrete examples of what was built, how it was evaluated, and how it failed. Practical command of agent design patterns: tool and function calling, structured outputs, retrieval/RAG, planning loops, multi-agent decomposition, and human-in-the-loop. Experience with agent evaluation and observability as a discipline — eval datasets, LLM-as-judge, regression testing, tracing, and cost/latency monitoring. Experience with the Anthropic Claude API or comparable frontier model APIs, including prompt and context engineering for reliability rather than novelty; plus experience with an agent framework such as the Claude Agent SDK and/or OpenAI Agents SDK / Assistants API, and a point of view on when to use a vendor framework versus building your own loop. Experience connecting models to enterprise data through MCP and/or comparable tool-integration patterns, and daily working fluency with Claude and ChatGPT as development tools (Claude Code, Projects, Custom GPTs, or equivalent). Mastery of Python — typing, testing, packaging, async — and advanced proficiency in SQL with strong analytical and dimensional data modeling skills. Deep, hands-on Apache Airflow experience: not just authoring DAGs, but debugging schedulers, managing dependencies, and designing retries, idempotency, and task granularity. Airflow 2→3 migration experience and familiarity with Astronomer are a strong plus. Hands-on proficiency with Snowflake (warehouses, RBAC, Snowpark, cost management), dbt, and AWS (ECS/Fargate, IAM, S3, Secrets Manager) with Docker and infrastructure-as-code — the core stack on this engagement. Broader exposure across the major cloud and data platforms (Azure, GCP, Databricks) is valued. Experience designing and building robust CI/CD pipelines, managing Infrastructure as Code, and working under Git-based review and environment-promotion discipline. Production experience extracting structured data from documents at scale, including at least one layout-heavy domain (financial documents, contracts, forms, or scientific papers), with familiarity with modern document parsing tooling and honest views on their tradeoffs. Experience with embeddings, chunking strategies, and vector or hybrid retrieval. Self-driven and autonomous, with excellent problem-solving and critical thinking, the ability to work against ambiguous problem statements, and the discipline to write a design down before building it. Excellent verbal and written professional English, sufficient for technical design discussions, written proposals, and collaboration with technical and non-technical stakeholders. Full availability during US Eastern Time business hours (09:00–17:00 ET). Preferred qualifications Hands-on experience with Snowflake Cortex specifically — Cortex Analyst, Cortex Search, Cortex Agents, or Cortex AISQL functions. Experience with a semantic layer product (dbt Semantic Layer, Cube, AtScale) or a data catalog platform (DataHub, OpenMetadata, Atlan, Collibra, or Snowflake Horizon). Knowledge graph or ontology modeling experience. Experience with dlt specifically, or having authored custom connectors and sources. Streaming or event-driven experience (Kafka, Kinesis, SNS/SQS). MLOps depth beyond agents: end-to-end ML lifecycle experience, model deployment and monitoring (e.g., MLflow), and advanced ML techniques such as regression, classification, or Bayesian methods. Experience fine-tuning LLMs and deploying them; publications or active contribution in relevant AI/ML communities; advanced degree in a relevant field. Prior work in private equity, venture capital, or financial services data environments. Prior experience as an embedded contractor or consultant influencing technical direction on client engagements. Benefits: Fully remote Flexible Schedule Unlimited Paid Time Off (PTO) Paid parental/bereavement leave Worldwide recognized clients to build skills for an excellent resume Top-notch team to learn and grow with Background checks may be required for certain positions/projects. Blue Orange Digital is an equal-opportunity employer. Department Engineering Role Machine Learning Engineer Locations Multiple locations Remote status Fully Remote Monthly salary $6,100 - $7,000 About Blue Orange Digital Blue Orange Digital is a data and AI consulting firm that helps companies turn complex data into real business outcomes. We partner with organizations across industries to design and deploy scalable data infrastructure, advanced analytics, and AI-powered solutions. Our team is fully remote, globally distributed, and driven by curiosity, impact, and innovation.
LLM Fine-Tuning Engineer (Open-Weight Models / Secure Environments)
Trenchant
Research Engineer Intern (Multimodal LLM)
Tether Operations Limited
Senior NLP/LLM Engineer
Social Discovery Ventures
Mobile Games QA
YallaPlay
Engineering Manager [DWH Data Tools]
Plata Card
Site Reliability Engineer (SRE)
Social Discovery Ventures