Senior Data Engineer
Posted 15hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Data Engineer building OCR and ingestion pipelines for an AI consultancy serving insurers. Converting unstructured insurance documents into structured data for RAG and anti-financial-crime models.
Responsibilities:
- Design and implement scalable pipelines for ingesting high-volume unstructured insurance documents
- Build connectors to document sources such as SharePoint and email
- Integrate, configure, and optimise OCR and document parsing technologies to extract high-accuracy text and layouts
- Build automated workflows for text cleaning, normalisation, semantic chunking, and metadata tagging
- Design vector storage schemas and robust retrieval mechanisms (RAG) to feed downstream AI models
- Ensure document processing pipelines meet enterprise security and low-latency SLA requirements
- Build automated error monitoring and extraction validation loops that flag low-confidence OCR outputs
- Apply engineering best practices across the pipeline lifecycle, including version control, CI/CD, and testing
- Turn high volumes of insurance documents into structured, high-quality, AI-ready data for downstream RAG and AI models
- Support a Dutch insurance client through a consultancy focused on AI, automation, advanced analytics, and anti-financial crime
Requirements:
- 5–10 years' experience in data engineering
- Proven experience building data processing and document ingestion pipelines on public cloud platforms
- Hands-on experience processing unstructured documents (PDF, Word, Excel, PowerPoint, scans, emails)
- Experience building connectors to enterprise sources such as SharePoint and email
- Practical experience with document extraction / OCR tools, e.g. AWS Textract or equivalent
- Strong engineering practices: Git, CI/CD, automated testing
- Mandatory: Python and SQL
- Mandatory AWS: S3, Step Functions, CloudWatch
- Mandatory OCR / document extraction: AWS Textract or equivalent
- Mandatory unstructured document processing and ingestion pipelines
- Mandatory Git, CI/CD, testing
- Experience in banking or insurance is an advantage
- Nice to have: Vector databases and RAG architectures
- Nice to have: Azure and Databricks
- Nice to have: Financial services / insurance domain experience
- Based in Europe; candidates in CEE preferred
- Ability to start within 30 days
Benefits:
- Long-term remote-first contract position
- Contract through July 2027 with possible extension
- Remote within Europe (CEE preferred)
















