Description Salla is looking for a Senior SRE Engineer (MLOps) to join our Salla AI team. This role focuses on running our AI and ML systems as real production systems, not side experiments — owning the operational layer around models, prompts, agents, inference services, and retrieval systems. You will be responsible for enabling Agentic AI and Generative AI features to operate reliably, securely, and cost-effectively at scale within the Salla ecosystem. This role is SRE- and platform-engineering-first, with a strong emphasis on reliability, observability, safe releases, cost, and governance, while collaborating closely with engineering, data, and AI teams to give every pod a fast, safe path to production. It exists because AI systems fail differently from normal services — a prompt change can behave like a code change, an agent calling tools needs auditability, and latency, quality, and cost can move together in uncomfortable ways. Key Responsibilities Own reliability for ML and agentic AI services in production — SLOs, dashboards, alerts, runbooks, and incident follow-ups Build observability across the AI stack — latency, errors, traces, tool calls, cost, and user impact Design safe-release patterns for models, prompts, agents, tools, and configuration, including canary, rollback, feature-flag, and evaluation-gate strategies Provide operational support for inference APIs, queues, retrieval layers, and AI workflows running on Kubernetes/EKS Establish ownership, traceability, and guardrails around what agentic systems are allowed to do, including how they call internal tools Defend agent tool-calling against prompt injection and untrusted-data risks — establish and enforce data-trust boundaries so that untrusted store/merchant content cannot manipulate agent decisions, tool calls, or actions Drive AI cost governance — per-model and per-pod spend visibility, token-cost tracking, and anomaly alerting Build automation and self-service paths so product teams have a known safe path to production instead of rebuilding it each time Turn recurring operational pain into simple, reusable platform standards that other teams adopt Participate in architecture discussions, code reviews, and technical decision-making
Junior AI Engineer - Computer Vision
Innovationteam
Senior AI Engineer - Professional Services - KSA
Datarobot
Generative AI Engineer
tamm
AI Engineer
Codeninjapk
Senior AI Engineer
Sulava MEA
AI Platform Engineer - Hybrid - KSA - (10 Months) - RTG
robusta