DE&A - Data Engineer
Summary
Data Engineer building large-scale ETL/ELT pipelines on GCP (BigQuery, Dataflow, Composer/Airflow, Pub/Sub), ensuring data quality and governance, and mentoring junior engineers.
Data Management & Governance
- Ensure data quality, integrity, and reliability through robust validation and monitoring frameworks.
- Implement best practices for data lineage, metadata management, and data cataloging using tools like Data Catalog and Looker.
- Collaborate with data governance, security, and compliance teams to enforce data access controls and encryption.
Collaboration & Leadership
- Work closely with data scientists, analysts, and business stakeholders to understand data requirements and deliver high-quality solutions.
- Provide technical mentorship and code reviews for junior data engineers.
- Contribute to architecture reviews and technology evaluations for continuous improvement of the data platform.
Preferred Qualifications
- GCP Professional Data Engineer Certification or equivalent.
- Experience with Looker, dbt (data build tool), or similar data modeling tools.
- Prior work with streaming data architectures and technologies such as Kafka, Spark Streaming, or Flink.
- Experience in data privacy, GDPR/CCPA compliance, and implementing data security at rest and in transit.
- Exposure to DevOps and container orchestration tools (e.g., Kubernetes, Cloud Run).
- Strong analytical and problem-solving skills with attention to performance and scalability.
Minimum Qualifications
- Bachelor's or Master’s degree in Computer Science, Engineering, Information Systems, or a related field.
- 5+ years of experience in data engineering, with at least 3+ years working extensively on GCP.
- Strong proficiency with SQL, Python, and/or Java/Scala for data processing and scripting.
- Proven experience with GCP services such as:
- BigQuery
- Dataflow / Apache Beam
- Pub/Sub
- Cloud Storage
- Cloud Composer (Airflow)
- Cloud Functions
- Experience in building large-scale ETL/ELT pipelines and optimizing query performance on cloud platforms.
- Solid understanding of data modeling, partitioning, and schema design for analytical workloads.
- Familiarity with CI/CD practices, Terraform, and version control (e.g., Git).