Senior Data Engineer (PySpark, Kafka, GCP)
Summary
Build and optimize batch and real-time data pipelines using PySpark, Kafka, and GCP tools like BigQuery and Dataproc for large-scale data processing.
Who we are
Apex Systems is a leading Data and Digital Transformation professional services organisation focused on delivering solutions that deliver real business value. We build authentic partnerships with our clients, offering objective counsel from concept to deployment to ensure a consistent voice in a dynamic IT environment
Required Experience & Qualifications
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent experience).
- 5+ years of experience in Data Engineering or Software Engineering.
- Strong hands-on experience with Python.
- Proven experience building and optimizing batch and real-time data pipelines using PySpark, Spark SQL, and Spark Streaming.
- Strong expertise writing and optimizing complex SQL queries for large-scale datasets.
- Experience working with Apache Kafka and event-driven architectures.
- Experience with Google Cloud Platform (GCP), including BigQuery and Dataproc.
- Experience working with RESTful APIs for data ingestion and integration.
- Hands-on experience implementing CI/CD pipelines using GitLab CI.
- Strong understanding of data modeling, distributed systems, and scalable data architectures.
- Experience working with relational and NoSQL databases.
- Experience supporting high-volume, real-time data processing environments.
Preferred Qualifications
- Experience with Scala.
- Exposure to Java and Spring Boot development.
- Familiarity with pair programming, code reviews, and collaborative engineering practices.
- Experience designing and supporting event-driven data platforms.
- Strong communication skills and ability to explain technical decisions and trade-offs.
- Experience partnering with cross-functional teams in agile environments.
Technical Stack Core
- Python
- Apache Spark (PySpark, Spark SQL, Spark Streaming)
- Advanced SQL
- BigQuery
Additional Technologies
- REST APIs
- Relational Databases
- NoSQL Databases
- Scala (Preferred)
Concepts
- Batch Data Processing
- Real-Time Data ProcessingData Modeling
- Distributed Systems
- Data Integration & ETL/ELT Workflows
Tools
- BigQuery
- Spark Ecosystem (PySpark, Spark SQL, Spark Streaming)
- RESTful APIs
- 100% payroll
- 30-day Christmas bonus
- 15-day vacations + 5 floating days (70% of vacation bonus)