Mid-Level Data Engineer
Summary
Mid-level (3-5 yrs) data engineer who designs, builds and troubleshoots production ETL/ELT pipelines, data warehouses and models day to day. Core stack: Python (Pandas/PySpark), advanced SQL, Airflow/dbt, Snowflake/BigQuery/Redshift/Databricks and AWS/Azure/GCP, with support for AI/ML data workloads.
Data Engineer role
Core experience
- 3–5 years' professional experience in Data Engineering, data-focused backend engineering or data architecture.
- Proven commercial experience actually building, maintaining and troubleshooting production ETL/ELT pipelines rather than purely academic/project exposure.
- Experience working with data warehouses, data architecture, data modelling and production data environments.
- Bachelor's degree in Computer Science, Software Engineering, Data Engineering, Information Systems or related field, or equivalent practical experience.
Python & SQL — essential
- Strong/advanced Python for data engineering, ideally including Pandas and PySpark; FastAPI exposure is also specified at mid-level.
- Strong to expert SQL, including complex joins, window functions, debugging, query optimisation and query-plan/performance optimisation.
- Must be capable of troubleshooting and optimising both SQL queries and Python data jobs.
- This is a significant upgrade from the original requirement, where only good SQL and basic programming/scripting were required.
ETL/ELT & pipelines — essential
- Design, build, test, maintain and optimise scalable ETL/ELT pipelines.
- Batch processing experience, with real-time/streaming exposure highly valuable.
-
Data ingestion from multiple sources including:
- REST APIs
- PostgreSQL/MySQL or other relational databases
- Third-party/SaaS platforms
- Structured, semi-structured and unstructured data
- Ideally Kafka/RabbitMQ or other message queues.
- Production pipeline monitoring, error handling, debugging, incident/root-cause resolution and preventative improvements.
Data warehousing & modelling — essential
- Hands-on experience with modern data warehouse/lakehouse platforms such as Snowflake, BigQuery, Amazon Redshift or Databricks.
- Strong understanding of data models, schemas and dimensional modelling.
- Practical exposure to Kimball/star schema; Data Vault is advantageous.
- Understanding of storage/query optimisation including indexing, partitioning and compression.
Cloud — essential for the upgraded role
-
Solid working knowledge of at least one major cloud platform:
AWS, Azure or GCP. - Hands-on exposure to the platform's data services rather than simply having a cloud certification.
- The junior specification only required familiarity with cloud platforms such as Fabric, AWS or BigQuery; the upgraded role requires solid working knowledge.
Orchestration & modern data stack — essential
- Commercial experience with Apache Airflow, dbt, Prefect or a comparable orchestration/workflow framework.
- Experience transforming data using dbt, Spark or cloud-native warehouse tools.
- Understanding of modern scalable data architecture.
Git & CI/CD
- Strong Git experience.
- GitHub/GitLab workflows.
- Basic CI/CD pipeline automation for deploying data code.
- Should be accustomed to collaborative development, code reviews and version-controlled production code.
Data quality, governance & security
- Automated data validation/testing, preferably dbt tests, Great Expectations or equivalent.
- Data quality monitoring and alerting.
- Data lineage, metadata and data dictionaries.
- Understanding of access controls, masking and encryption.
- Awareness of data governance/compliance requirements.
AI/ML data engineering
- Must understand how Data Engineering supports AI/ML workloads.
- Preparing clean, reproducible datasets for model training, evaluation and inference.
- Data preprocessing and validation for ML pipelines.
- Exposure to feature stores, MLOps and generative-AI integrations is highly advantageous.
- The upgraded requirements build directly on the newer Junior specification, which already introduced structured/unstructured data preparation, vector databases and RAG pipelines.
- Bonus technologies include MLflow, Feast, Vertex AI and SageMaker.
Highly advantageous / differentiators
- Kafka, AWS Kinesis or Apache Flink
- Docker
- Kubernetes
- Terraform or CloudFormation
- NoSQL — MongoDB, DynamoDB or Redis
- Spark/PySpark
- Real-time streaming pipelines
- MLOps
- Cost optimisation within cloud data environments.
non-negotiables requirements
3–5 years relevant commercial experience + strong Python + advanced SQL + hands-on ETL/ELT pipeline development + data warehousing / modelling + AWS / Azure / GCP + Airflow / dbt / Prefect or equivalent + Git + genuine production troubleshooting / optimisation experience, production Data Engineering experience.
additional advantage Snowflake / Databricks / BigQuery / Redshift, PySpark/ Spark, CI/CD, automated data testing, AI/ML data pipelines and ideally Kafka/streaming.
Skills
- AI
- Airflow
- API
- Automation
- AWS
- Azure
- BigQuery
- CI/CD
- Cloud
- Cloud Native
- CloudFormation
- Data Engineering
- Data Governance
- Data Ingestion
- Data Lineage
- Data Modeling
- Data Pipelines
- Data Quality
- Data Warehousing
- Databricks
- dbt
- Dimensional Modeling
- Docker
- DynamoDB
- ELT
- ETL
- FastAPI
- Flink
- GCP
- Generative AI
- Git
- GitHub
- GitLab
- Kafka
- Kinesis
- Kubernetes
- Lakehouse
- Machine Learning
- MLflow
- MLOps
- MongoDB
- MySQL
- NoSQL
- pandas
- PostgreSQL
- Prefect
- PySpark
- Python
- RabbitMQ
- Redis
- Redshift
- REST
- SaaS
- SageMaker
- Snowflake
- Spark
- SQL
- Terraform
- Vault
- Vector Databases
- Vertex AI