Applied AI Engineer
OBytesAbout OBX
OBX is an Obytes initiative helping engineering teams adopt AI-enabled software-delivery workflows.
We combine platform capabilities with technical expertise to assess opportunities, implement agent-enabled workflows inside customer environments, and operate them reliably.
We are looking for an Applied AI Engineer who can turn rapidly evolving AI capabilities into secure, measurable, production-grade systems.
The role
You will design, build, evaluate, and improve the AI and agent capabilities behind OBX.
This is not a research-only, data-science, or prompt-engineering role. You will be expected to write production software, integrate models with tools and customer environments, create evaluation systems, investigate failures, and improve reliability, latency, and cost.
You will work closely with our platform engineers, principal architect, product lead, engineering and delivery lead, and Forward Deployed Engineering team.
What you will own
- AI-system behaviour and agent-workflow quality.
- Model and provider integration.
- Tool calling, structured outputs, workflow orchestration, and state management.
- Evaluation datasets, test scenarios, and quality metrics.
- Prompt, model, and configuration versioning.
- Failure classification and root-cause analysis.
- AI observability and production diagnostics.
- Reliability controls, fallbacks, retries, and human approval points.
- Model-cost and latency optimization.
- Reusable AI components that can support multiple customer workflows.
- Technical documentation and operational runbooks for AI capabilities.
What you will do
- Build agent-enabled workflows for software engineering and delivery use cases.
- Integrate large language models with internal services, developer tools, APIs, and customer systems.
- Design evaluation frameworks covering accuracy, task completion, safety, reliability, latency, and cost.
- Create representative test datasets and regression suites.
- Identify and reduce hallucinations, tool-selection errors, invalid outputs, and non-deterministic failures.
- Develop safeguards, approval mechanisms, and clear boundaries for autonomous actions.
- Implement tracing, monitoring, alerting, and diagnostic capabilities.
- Compare models and providers using measured performance rather than assumptions.
- Work with product and commercial teams to assess whether proposed customer use cases are feasible and valuable.
- Support Forward Deployed Engineers during technical discovery, integration, deployment, and incident investigation.
- Turn customer-specific learning into reusable platform capabilities.
- Contribute to architecture reviews, engineering standards, code reviews, and technical documentation.
- Stay current with relevant AI technologies while avoiding unnecessary framework churn.
The role owns the quality and operational behaviour of AI components, while working within the broader product, architecture, and delivery model.
What your first 90 days will look like
Days 1–30: Understand, baseline, and expose weaknesses
- Understand the OBX architecture, agent workflows, demonstrations, and intended customer use cases.
- Review current model integrations, prompts, tools, data flows, and failure-handling mechanisms.
- Create an initial failure taxonomy covering the most important agent failure modes.
- Establish baseline measurements for quality, latency, reliability, and model cost.
- Define the first evaluation dataset and regression suite.
- Recommend the highest-priority technical improvements.
Days 31–60: Improve one priority workflow
- Take ownership of the AI behaviour of one strategically important OBX workflow.
- Implement automated evaluations and regression testing for that workflow.
- Improve tool use, output validation, error handling, observability, and recovery behaviour.
- Introduce appropriate human approval and safety controls.
- Document model-selection decisions and trade-offs.
- Demonstrate measurable improvement against the original baseline.
Days 61–90: Prove production readiness
- Validate the workflow in a realistic or customer-like environment.
- Define acceptable quality, latency, reliability, and cost thresholds.
- Establish monitoring, incident-diagnosis procedures, and operational runbooks.
- Support integration with the Forward Deployed Engineering team.
- Extract reusable components from implementation work.
- Present a 90-day technical assessment and the next two quarters of AI priorities.
How success will be measured
Success will be evaluated using evidence such as:
- Coverage and quality of automated AI evaluations.
- Reduction in critical agent failures.
- Reproducibility of previously observed failures.
- Task-completion and output-quality improvements.
- Reliability of tool calls and structured outputs.
- Model latency and cost per successful workflow.
- Production observability and diagnostic quality.
- Number of reusable capabilities created from customer implementations.
- Quality of technical documentation and operational handoffs.
- Collaboration with the platform and Forward Deployed Engineering teams.
The role will not be evaluated by the number of prompts written, frameworks tested, or demos produced without measurable reliability.
What we’re looking for
- Strong software-engineering fundamentals and experience building production systems.
- Hands-on experience building applications using large language models or agent-based workflows.
- Experience with model APIs, tool calling, structured outputs, orchestration, retrieval, and context management.
- Ability to design evaluations and reason about probabilistic system behaviour.
- Experience debugging complex workflows involving models, tools, services, and external data.
- Strong backend development skills using Python, TypeScript, or comparable technologies.
- Experience with APIs, databases, cloud infrastructure, automated testing, and CI/CD.
- Understanding of production concerns including security, privacy, monitoring, latency, availability, and cost.
- Ability to explain technical trade-offs clearly to engineering, product, and business stakeholders.
- Comfort operating in an early-stage environment where requirements and technologies evolve quickly.
Valuable additional experience
- Building developer tools or internal engineering platforms.
- Deploying AI systems into customer-controlled or private environments.
- AI observability, tracing, and evaluation platforms.
- Retrieval-augmented generation and enterprise search.
- Multi-agent or long-running workflow orchestration.
- Model routing, caching, and cost optimization.
- AI security, prompt-injection mitigation, and data-governance controls.
- Supporting customer-facing technical deployments.
- Open-source contributions related to AI infrastructure or developer tooling.
What good looks like
A strong candidate does not merely make a demonstration look impressive. They can explain:
- How the system is evaluated.
- Where and why it fails.
- What actions require human approval.
- How failures are detected and diagnosed.
- How customer data is protected.
- How model choices affect quality, latency, and cost.
- What should become reusable platform capability.
- What OBX should refuse to automate.
We value engineers who are excited about AI but sceptical of unmeasured claims.