Data Engineer
Summary
A hybrid Data Engineer role in Glasgow (2-3 days per week onsite) designing, building, and maintaining scalable data pipelines, data lakes, and data warehouses on AWS. Core stack: PySpark, Spark, Python, SQL, AWS services (S3, Glue, Lambda, Step Functions), CloudFormation, and GitLab CI/CD.
We are looking for Data Engineer at Glasgow, Scotland – 2-3 days per week Onsite
Purpose of the Role
To design, build, and maintain scalable data pipelines, data lakes, and data warehouse solutions on AWS. The role focuses on developing high-performance data engineering solutions using PySpark, Spark, Python, and AWS services, enabling secure, reliable, and efficient data processing and analytics across enterprise platforms.
Key Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines using PySpark, Spark, Python, and AWS services.
- Build and optimize data lakes and data warehouse solutions ensuring data quality, security, and accessibility.
- Develop reusable, production-grade ETL/ELT frameworks and data processing solutions.
- Implement orchestration workflows using AWS Step Functions, Airflow, and other automation tools.
- Develop and maintain cloud infrastructure using AWS CloudFormation.
- Collaborate with business stakeholders to understand requirements and translate them into scalable technical solutions.
- Optimize data processing performance, monitoring, and operational support.
- Implement unit testing, code reviews, and CI/CD best practices using GitLab.
- Support platform modernization and migration initiatives leveraging Spark-based architectures.
- Work closely with Data Scientists and Analytics teams to enable AI/ML use cases.
Required Skills & Experience
- Strong hands-on experience in Data Engineering with delivery of production-grade solutions.
- Expertise in PySpark, Apache Spark, Python, and SQL.
- Strong experience designing and optimizing complex data pipelines and ETL/ELT frameworks.
- Hands-on experience with AWS services including:
- S3
- Glue
- Lambda
- Step Functions
- ECS
- IAM
- KMS
- VPC
- SageMaker (preferred)
- Experience with AWS CloudFormation for Infrastructure as Code.
- Strong understanding of data lakes, data warehouses, and distributed data processing.
- Experience with GitLab, CI/CD, Unit Testing, and DevOps practices.
- Excellent problem-solving skills and ability to work independently.
- Strong stakeholder management and communication skills.
Nice to Have
- Experience with Databricks, Delta Lake, Unity Catalog, and migration projects.
- Knowledge of AI/ML and MLOps frameworks.
- Experience with streaming technologies such as Kafka or Kinesis.
Ideal Candidate:
A hands-on Data Engineer with strong expertise in Spark, PySpark, AWS, and CloudFormation, capable of building scalable enterprise data solutions while driving modernization and cloud transformation initiatives.
Skills
- Accessibility
- AI
- Airflow
- Analytics
- Automation
- AWS
- CI/CD
- Cloud
- CloudFormation
- Data Engineering
- Data Pipelines
- Data Quality
- Data Warehousing
- Databricks
- Delta Lake
- DevOps
- ECS
- ELT
- ETL
- GitLab
- IAM
- Infrastructure as Code
- Kafka
- Kinesis
- Lambda
- Machine Learning
- MLOps
- PySpark
- Python
- S3
- SageMaker
- Spark
- SQL
- Stakeholder Management
- Unit Testing
- Unity
- VPC