AI Data Engineer
- Hiring from
- South Africa
- Work type
- Hybrid
- Posted
- Oct 2, 2026
Is this job info correct?
We are looking for an AI Data Engineer to build and optimize high-throughput, AI-ready data pipelines that power enterprise GenAI, RAG applications, and predictive machine learning models on a 12-month hybrid/remote contract (with option to renew).
As an AI Data Engineer, you will serve as the architect of our data foundation, designing ingestion engine capabilities for both structured and unstructured data. You will implement vector databases, enforce rigorous data quality frameworks, and optimize inference datasets for production AI workloads.
The Tech Stack Deep-Dive
Must-Haves:
- Data Engineering & Pipelines: Scalable data pipeline construction for structured and unstructured data ingestion.
- Vector Databases & GenAI: Vector database implementation and configuration supporting RAG and Generative AI applications.
- ML Data Preparation: Dataset curation, transformation, and optimization for ML model training, validation, and inference.
- Data Governance & Quality: Data quality frameworks, lineage tracking, metadata catalogues, and data privacy controls.
Nice-to-Haves:
- Cloud AI service configurations and automated pipeline orchestration.
- Experience with POPIA compliance and enterprise model risk standards.
Key Responsibilities
- Build AI Pipelines: Construct scalable ingestion pipelines capable of handling massive structured and unstructured datasets.
- Operationalize Vector Stores: Implement and manage vector databases to drive fast, accurate context retrieval for enterprise RAG applications.
- Optimize ML Datasets: Prepare, clean, and structure high-quality datasets explicitly designed for model training, validation, and real-time inference.
- Embed Governance & Lineage: Establish metadata catalogues, lineage tracking, and automated data quality checks directly within pipeline architectures.
Why Join Us?
- High Technical Impact: Build the core data foundation that directly dictates the performance of enterprise-scale AI models.
- Modern AI Tech Stack: Work hands-on with cutting-edge vector databases, unstructured data parsers, and automated governance tools.
- Hybrid / Remote Autonomy: Work with a high-performing engineering team under flexible hybrid or remote arrangements