Senior Data Engineer (GCP, PySpark, Dataproc)
Summary
Senior Data Engineer builds and maintains PySpark ETL pipelines on GCP Dataproc, integrating PostgreSQL, SAP HANA, and S3-compatible storage for a public-sector data platform in the Middle East.
We are looking for a Senior Data Engineer with strong hands-on experience in Apache Spark, PySpark and GCP Dataproc to join Implex and work on a long-term data platform project for a public-sector in the Middle East.
The engineer will contribute to the development and support of data-processing pipelines that integrate multiple enterprise data sources, including PostgreSQL, SAP HANA and S3-compatible object storage. The role involves not only writing PySpark jobs but also troubleshooting distributed workloads, configuring Dataproc environments, securing external connections and ensuring reliable delivery of data into the target data warehouse.
Responsibilities
Develop, maintain, and optimize ETL pipelines using Apache Spark and PySpark.
Configure, run, and troubleshoot Spark workloads on GCP Dataproc.
Package and submit Spark jobs, manage dependencies, and analyze driver and executor logs.
Identify and resolve performance issues related to memory usage, shuffling, partitioning, data skew, and distributed data processing.
Integrate Spark workloads with S3-compatible object storage using the Hadoop S3A connector.
Configure and troubleshoot Spark connectivity with MinIO, including bucket policies, custom endpoints, path-style access, TLS, signature compatibility, and redirect handling.
Configure custom CA certificates and Java truststores for Spark drivers and executors.
Implement and optimize scalable Spark JDBC reads and writes for PostgreSQL and SAP HANA, including partitioning, batching, pushdown, and data type mapping.
Build reliable incremental data loads, retries, backfills, idempotent processing, and data reconciliation mechanisms.
Design and support staging-to-publish data flows, bulk data loads, and data warehouse integration patterns.
Contribute to data modeling, schema evolution, Slowly Changing Dimension patterns, data lineage, technical documentation, and data quality practices.
Apply secure secrets management and least-privilege access principles across data integrations.
Must-Have Technical Skills
Apache Spark / PySpark development (Dataproc): driver/executor behavior, job packaging/submission, performance tuning
GCP Dataproc operations: cluster configuration, init actions, dependency management, troubleshooting via logs/metrics
Hadoop S3A connector: `fs.s3a.*` configuration, endpoint/path-style access, credential providers, S3 semantics
MinIO (S3-compatible) integration: bucket policies, TLS endpoints, signature/redirect troubleshooting
TLS/SSL & PKI with custom CA: certificate chains, SAN/hostname validation, diagnosing handshake/PKIX errors
Java truststores (JKS/PKCS12) & JVM SSL config: `keytool`, distributing truststores, setting driver/executor JVM options
PostgreSQL integration: Spark JDBC reads/writes at scale, indexing/performance basics, data type mapping
SAP HANA integration: JDBC/ODBC connectivity, driver management, calculation views vs tables, pushdown/performance tuning
ETL engineering: incremental loads/CDC concepts, idempotency, retries, backfills, data quality/reconciliation
Data Warehousing integration: strong SQL, staging-to-publish patterns, SCD concepts, bulk load strategies
Data modeling & governance basics: dimensional modeling, schema evolution, lineage/documentation practices
Nice-to-Have Skills (but not must)
Linux + networking fundamentals: DNS, routing, firewall/LB/proxy basics; tools like `curl`/`openssl s_client` for validation
Secure secrets handling: GCP Secret Manager (or equivalent), least-privilege access, avoiding hardcoded credentials
KAFKA knowledge if we ever bring KAFKA into the architecture again
Hiring process
HR interview
PM/Technical interview with Implex
Dev Lead interview on the client side
What we offer
A long-term international project
Opportunity to work on a national-scale digital platform used by thousands of users
Remote full-time collaboration
Professional and supportive team environment
Challenging technical tasks and a meaningful product with real-world impact