Senior Data Engineer – Python / Spark / AWS & GCP @ Square One Resources
Summary
Senior Data Engineer / Data Platform Engineer joining an international data team in the e-commerce and digital marketplace space, building cloud-based batch and streaming pipelines that ingest petabytes across data lakes and warehouses on AWS and GCP. Core stack: Python/Java/Scala, SQL, Spark, Hadoop/Hive, Airflow, and BigQuery.
The team is responsible for building next-generation, cloud-based data solutions capable of ingesting and processing petabytes of data across data lakes and data warehouses.
Our requirements
- 10+ years of experience in Data Engineering, Software Engineering, Distributed Systems, or a related field.
- Strong programming skills in Python, Java, or Scala.
- Strong knowledge of SQL and experience with SQL and NoSQL databases such as BigQuery, Teradata, MySQL, PostgreSQL, or Cassandra.
- Experience with Big Data technologies, including Hadoop, Hive, and Apache Spark.
- Hands-on experience with Airflow or another orchestration tool.
- Strong understanding of ETL, data lineage, data quality, backfills, and data pipeline troubleshooting.
- Experience developing both batch and streaming data pipelines.
- Experience across both software development and data engineering.
- Experience with AWS or GCP, particularly data processing at scale.
- Experience designing and supporting production services with strict SLAs.
- Experience with distributed applications and monitoring/logging tools such as Elasticsearch and Wavefront.
- Good understanding of testing methodologies, CI/CD, and software engineering best practices.
- Excellent written and verbal communication skills.
- Scrum Master experience.Strong expertise in Python.
- Experience with Google Data Streams and Google Dataproc.
- Experience with Change Data Capture (CDC) technologies.
- Experience with modern data warehouse technologies such as Delta Lake.
The team is responsible for building next-generation, cloud-based data solutions capable of ingesting and processing petabytes of data across data lakes and data warehouses. ,(Design and develop high-volume batch and streaming data ingestion pipelines across AWS and GCP., Build and launch next-generation data ingestion and data curation platforms., Participate in system and data architecture discussions., Design scalable and high-performance distributed data solutions., Troubleshoot and resolve issues related to ETL, data lineage, data quality, backfills, and data pipelines., Design, implement, and support production services with strict SLAs., Collaborate with Software Engineers, Data Engineers, ML Engineers, Data Analysts, and other stakeholders., Provide technical guidance and mentor junior engineers., Contribute to best practices across software development, data engineering, testing, and CI/CD., Work with logging, monitoring, metrics, and alerting solutions to ensure platform reliability.) Requirements: Data engineering, Software Engineering, Distributed systems, Python, Java, Scala, SQL, NoSQL, BigQuery, Teradata, MySQL, PostgreSQL, Cassandra, Hadoop, Hive, Apache Spark, Airflow, ETL, AWS, GCP, SLA, Testing methodologies, CI/CD, Communication skills, Scrum Master, Google Data Streams, Google Dataproc, Change Data Capture, CDC, Delta Lake Additionally: Sport subscription, Private healthcare.