Senior Data Engineer
Summary
The Senior Data Engineer will design, build, and operate cloud-native data platforms using Python, Spark, and modern orchestration tools. The role focuses on applying software engineering best practices to data pipelines and cloud infrastructure.
Role Summary
We are seeking engineers who combine strong software engineering fundamentals with modern data engineering expertise. Candidates should be capable of designing, building, deploying and operating cloud-native data platforms while applying software engineering best practices throughout the delivery lifecycle.
Strong SQL and data warehousing experience remain important but should complement broader engineering capability rather than define the candidate's profile.
Minimum Technical Requirements (Non-Negotiable)
Candidates must demonstrate practical project experience with:
Software Engineering
- Python as a primary programming language
- Software engineering principles and clean coding practices
- Object-oriented programming
- Testing and code quality practices
- Software Development Life Cycle (SDLC)
- Git and collaborative development workflows
- API development and integration
Candidates should demonstrate the ability to:
- Solve business problems through code
- Design scalable solutions
- Work within engineering teams
- Contribute to production systems
- Follow engineering standards and best practices
Core Data Engineering Requirements
Candidates should have hands-on experience with several of the following:
Data Processing
- Spark
- PySpark
- Databricks
- Data Lake architectures
- Batch processing
- Streaming architectures
- Data transformation frameworks
Data Platform Technologies
- Kafka
- Flink
- Airflow
- Modern orchestration platforms
Data Storage & Analytics
- SQL
- Data warehousing concepts
- Relational databases
- Analytical data platforms
Candidates should have practical experience delivering solutions on at least one major cloud platform:
Preferred Order
- Amazon Web Services (AWS)
Typical Technologies
GCP
- BigQuery
- Dataflow
- Pub/Sub
- GKE
AWS
- Glue
- EMR
- Redshift
- Kinesis
- EKS
- S3
- Data Factory
- Synapse
- Databricks
- Event Hubs
- AKS
Cloud experience should reflect real project delivery rather than certifications alone.
Strong candidates should also demonstrate exposure to:
Containerisation & Deployment
- Docker
Infrastructure
- Terraform
- Infrastructure as Code
Operations
- Monitoring
- Observability
- Production support
Candidates should possess working knowledge of:
Data Warehousing
- Data warehouse design
- Data governance concepts
Data modelling should support modern platform development, not exist in isolation.
Ideal Candidate Profile
The strongest candidates will demonstrate: