AWS Data Engineer
Summary
Design and build an AWS-based data lake and lakehouse, including ingestion pipelines, transformations, and governance, using Glue, Redshift, Lambda, and Lake Formation.
Key Responsibilities
Architecture & Design
- Design and architect the end-to-end AWS Data Lake and Lakehouse solution, including Landing Zone, Transformed Zone, and Curated/Consumption Zone layers
- Define and govern data architecture standards, patterns, and best practices across the platform
- Architect reusable data ingestion pipelines supporting REST APIs, JDBC databases, S3 file uploads, and SaaS connectors (e.g. Salesforce via AWS AppFlow)
- Design data storage strategies including hot, warm, and cold storage tiers, encryption, and data lifecycle policies
Development & Deployment
- Develop and deploy data ingestion pipelines using AWS Glue, Lambda, Step Functions, EventBridge, and API Gateway
- Build and maintain data transformation workflows (batch and stream processing) using AWS Glue and Amazon Redshift
- Implement orchestration, monitoring, logging, and notification frameworks for pipeline operations
- Develop and maintain the AWS Glue Data Catalogue, including schema evolution tracking and metadata tagging
Security & Governance
- Configure and enforce data security policies using AWS Lake Formation, IAM, and Secrets Manager
- Implement granular access controls at database, table, and column levels
- Ensure compliance with data classification, retention, and audit requirements
- Support data quality frameworks and observability monitoring
Maintenance & Operations
- Monitor platform health, performance, and pipeline reliability
- Troubleshoot and resolve data pipeline failures and data quality issues
- Maintain documentation for architecture decisions, pipeline configurations, and operational runbooks
- Continuously optimise platform performance and cost efficiency on AWS
Requirements
Essential
- Minimum 3 to 5 years of experience in data engineering, data architecture, or cloud infrastructure roles
- Hands-on expertise with core AWS data services: Amazon S3, AWS Glue, Amazon Redshift, AWS Lambda, Amazon Kinesis, AWS Step Functions, Amazon EventBridge, AWS AppFlow, AWS Lake Formation
- Strong proficiency in SQL and at least one scripting language (Python or Scala)
- Experience designing and implementing Data Lake or Lakehouse architectures
- Solid understanding of data governance, data cataloguing, and metadata management
- Experience with batch and streaming data processing patterns
- AWS Certified Data Engineer – Associate or AWS Certified Solutions Architect certification (or equivalent)
Preferred
- Experience integrating with Tableau or similar BI visualisation tools via Amazon Redshift or S3
- Familiarity with MLOps frameworks and AI/ML model deployment on AWS SageMaker
- Experience with Salesforce data integration using AWS AppFlow
- Knowledge of Change Data Capture (CDC) and incremental data load patterns
- Prior experience in a government or public sector data environment