Data Engineer
Summary
Senior Data Engineer building scalable data and AI platforms on AWS using Python, SQL, Docker, and IaC for mission-critical HR tech solutions used by governments worldwide. Hybrid role in Utrecht.
Are you a senior Data Engineer who wants to leverage data engineering, MLOps, and architecture skills to build production-grade data and AI platforms with real-world impact? At WCC, we build high-impact, mission critical HR tech solutions used by governments and public institutions across the globe to help millions of people find suitable jobs today and prepare for the labor market of tomorrow. We are expanding our team with a senior Data Engineer to help us responsibly build scalable data and AI platforms that turn real-world client data into reliable product capabilities.
This is a hands-on role where you will work with cross-functional teams to design, build, deploy, and operate data and AI platform capabilities for solutions that impact and improve the lives of millions of people around the world. You will combine strong engineering skills with pragmatic judgment and ownership over production outcomes, ensuring that we do the right things and do them right.
What you’ll do
Data platforms, data pipelines, databases, and data products
- Design, build, and operate scalable data platforms, databases, and pipelines that ingest, transform, validate, and connect data from multiple internal and external sources, including messy, incomplete, and heterogeneous real-world client data.
- Design and optimize relational, analytical, operational, and vector-oriented data stores for large-scale workloads, semantic search, matching, recommendations, and other AI-enabled capabilities.
- Create reusable data products, canonical data models, curated datasets, and feature-ready assets for Data Scientists, product teams, implementation teams, and customer-facing applications.
- Build in data quality, metadata, observability, and lineage so teams can trust the data and understand where it came from, how it changed, and how it is being used.
Cloud infrastructure and operational reliability
- Design, deploy, and maintain cloud-native data and AI platform components on AWS, using automation and infrastructure-as-code wherever possible.
- Work hands-on with storage, compute, networking, access control, secrets management, monitoring, logging, CI/CD, orchestration, and deployment automation.
- Ensure data and AI workloads are secure, scalable, cost-effective, maintainable, and supported by pragmatic architectural trade-offs across environments.
MLOps, AI governance, and responsible operation
- Enable Data Scientists and AI engineers to turn experiments, models, prompts, embeddings, and retrieval pipelines into reliable production services.
- Support automated workflows for model training, validation, deployment, versioning, monitoring, retraining, lifecycle management, and lineage across datasets, features, models, prompts, embeddings, evaluations, deployments, and production outcomes.
- Implement monitoring and alerting for data quality issues, model performance, data drift, concept drift, embedding drift, and unexpected changes in production behavior.
- Translate AI governance requirements into workable platform capabilities such as audit trails, approval gates, access controls, evaluation workflows, release controls, and operational dashboards.
Collaboration and technical ownership
- Work closely with Data Scientists, Architects, DevOps engineers, Product Owners, and project teams to translate data and AI needs into maintainable technical solutions.
- Act as a senior technical sparring partner on data architecture, platform design, MLOps, AI governance, operational readiness, and production support, with a strong focus on practical decisions that keep solutions understandable, supportable, and reliable over time.
- Document data flows, model flows, platform designs, operational procedures, and architectural decisions so solutions can be understood, supported, audited, improved, and kept reliable in production.
- Take ownership of production outcomes: not just building pipelines and services, but ensuring they keep working, remain understandable, and support the people and products that depend on them.
Work conditions
- Hybrid work setup when not traveling (60% in our Utrecht office, 40% from home).
- Travel internationally as required (incidentally, depending on project needs).
What you bring
Core requirements
- A completed degree in Computer Science, Data Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics, Econometrics, or another quantitative or technical field; a Master’s degree is a plus.
- At least 5 years of professional experience in data engineering, platform engineering, MLOps, cloud engineering, or a closely related role.
- Strong experience designing, building, and operating scalable data platforms, pipelines, data lakes, or data products in production environments, including ETL development, workflow orchestration, data quality controls, metadata, lineage, and operational monitoring.
- Strong architectural judgment and a pragmatic, hands-on mindset, with the ability to balance speed, scalability, cost, security, governance, reliability, and maintainability while taking ownership in complex technical environments.
- Experience supporting production AI or machine learning services with monitoring, versioning, lifecycle management, reproducibility, deployment traceability, lineage, drift detection, audit trails, access control, and operational controls across datasets, features, models, and deployments.
- Strong hands-on experience with AWS-based infrastructure and services, Python, Docker, infrastructure-as-code, CI/CD, and production-grade deployment practices for data or platform workloads.
- Ability to turn messy, incomplete, inconsistent, or fast-changing real-world data into reliable, usable, and well-documented assets.
- Advanced SQL skills and substantial experience with database design, query optimization, canonical data models, scalable schema design, and modeling for complex business domains.
- Strong collaboration and communication skills, with the ability to work effectively across Data Science, DevOps, Architecture, Product, implementation, Security, and Privacy stakeholders.
- Professional-level English, both written and spoken.
Preferred qualifications
- Experience with AWS data, compute, networking, security, monitoring, and deployment services, such as S3, Lambda, Glue, Step Functions, EventBridge, Athena, EMR, RDS, Redshift, DynamoDB, ECS, ECR, EC2, VPC, IAM, CloudWatch, CloudTrail, Systems Manager, Secrets Manager, KMS, API Gateway, or AWS CDK.
- Experience with Kubernetes, Terraform, Bitbucket Pipelines, AWS Lake Formation, AWS DataZone, AWS Glue Data Catalog, or lakehouse architectures.
- Experience with MLflow, Amazon SageMaker, AWS Bedrock, feature stores, model registries, experiment tracking, evaluation stores, model monitoring platforms, drift detection, or AI observability tools.
- Familiarity with lakehouse architectures, Data Vault Modeling, data mesh concepts, data contracts, master data management, schema registries, data lineage frameworks, metadata management, or enterprise data cataloging.
- Experience with vector databases, embedding storage, vector indexing techniques, retrieval-augmented generation architectures, or semantic search platforms.
- Full-stack software development experience or experience building APIs and services that expose data or AI capabilities to products.
- Familiarity with frameworks such as GDPR, the EU AI Act, ISO 27001, and ISO 42001.
- Ability to translate data governance, AI governance, compliance, and risk management needs into pragmatic platform capabilities such as approval gates, audit trails, evaluation workflows, release controls, and post-deployment monitoring.
- AWS certifications are a plus.
- Experience working with public-sector, HR tech, labor market, staffing, or other socially impactful data domains is a plus.