Senior Data Engineer
- Salary
- €35–€41/hr
- Hiring from
- Czech Republic
- Work type
- Remote
- Posted
Show job descriptionHide job description
We are looking for a Senior Data / Ingestion Engineer for our client, a consultancy that helps banks and insurers transform their operations through AI, automation, and advanced analytics, with a strong focus on Anti-Financial Crime (fraud prevention, AML, KYC) and enterprise AI platforms. You will join their delivery for a Dutch insurance client.
We're looking for an engineer with deep experience building robust ingestion pipelines for unstructured documents and integrating OCR and document extraction technologies. You will turn high volumes of insurance documents (PDFs, scans, emails, Office files) into structured, high-quality, AI-ready data that feeds downstream RAG and AI models.
This is a long-term remote-first contract position for candidates based in Europe.
Responsibilities
Design and implement scalable pipelines for ingesting high-volume unstructured insurance documents (PDFs, scans, emails, Word, Excel, PowerPoint).
Build connectors to document sources such as SharePoint and email.
Integrate, configure, and optimise OCR and document parsing technologies to extract high-accuracy text and layouts.
Build automated workflows for text cleaning, normalisation, semantic chunking, and metadata tagging.
Design vector storage schemas and robust retrieval mechanisms (RAG) to feed downstream AI models.
Ensure document processing pipelines meet enterprise security and low-latency SLA requirements.
Build automated error monitoring and extraction validation loops that flag low-confidence OCR outputs.
Apply engineering best practices across the pipeline lifecycle: version control, CI/CD, and testing.
Work Conditions
Start Date: ASAP
Location: Remote within Europe (CEE preferred)
Long-term contract-based role: until July 2027 with possible extension
Contract with EU LCC
5–10 years' experience in data engineering.
Proven experience building data processing and document ingestion pipelines on public cloud platforms.
Hands-on experience processing unstructured documents (PDF, Word, Excel, PowerPoint, scans, emails).
Experience building connectors to enterprise sources such as SharePoint and email.
Practical experience with document extraction / OCR tools, e.g. AWS Textract or equivalent.
Strong engineering practices: Git, CI/CD, automated testing.
Experience in banking or insurance is an advantage.
Technical Requirements
Mandatory
Python, SQL
AWS: S3, Step Functions, CloudWatch
OCR / document extraction: AWS Textract or equivalent
Unstructured document processing and ingestion pipelines
Git, CI/CD, testing
Nice to have
Vector databases and RAG architectures
Azure, Databricks
Financial services / insurance domain experience