Data Engineer III
Summary
Designs and builds scalable cloud data pipelines and lakehouse architectures on GCP, using BigQuery, dbt, Airflow, Python, and SQL to power analytics and AI initiatives for a global edtech leader.
Data Engineer – Global Data Team
Role Overview
We are looking for an experienced Data Engineer with 6–10 years of hands-on experience to join our Global Data team and help build reliable, scalable, and self-service data platforms that power enterprise-wide analytics and data products.
In this role, you will design, develop, and optimize modern cloud-based data platforms, lakehouse architectures, and distributed data pipelines. You will work extensively with technologies such as Google Cloud Platform (GCP), Google BigQuery, GCP Data engineering services, Data flow, Data Fusion, Data Streams, Cloud Storage, Cloud Composer/Airflow, dbt, Python, and SQL, while contributing to data governance, observability, automation, and AI-powered data engineering capabilities.
The ideal candidate is a hands-on engineer who enjoys solving complex data challenges, building reusable engineering frameworks, improving platform reliability, and enabling data teams through scalable self-service capabilities.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines supporting enterprise data products and analytics.
- Build and optimize data solutions using Google BigQuery, GCS, Cloud Composer/Airflow, dbt, Python, and SQL.
- Develop robust data ingestion and transformation pipelines from databases, SaaS applications, APIs, files, and other enterprise data sources.
- Implement and maintain API integrations and data ingestion frameworks for cloud and enterprise applications.
- Develop reusable dbt models, transformation frameworks, and data quality processes.
- Build and manage workflow orchestration using Apache Airflow / Google Cloud Composer.
- Work with CDC and data ingestion technologies such as Informatica CDC, Airflow, Composer, dbt, or similar platforms.
- Design and implement modern data lakehouse architectures using BigQuery and related cloud technologies.
- Optimize data pipelines and BigQuery workloads for performance, scalability, reliability, and cost efficiency.
- Implement engineering best practices including CI/CD, version control, automated testing, deployment automation, and DevOps practices.
- Contribute to data observability and monitoring frameworks, including pipeline health, data quality, SLA monitoring, and operational metrics.
- Partner with data architects, analysts, data scientists, product teams, and business stakeholders to deliver high-quality data solutions.
- Contribute to data governance, metadata management, lineage, and data discovery capabilities.
- Explore and implement AI/ML and GenAI capabilities to improve data engineering workflows, automation, data discovery, and developer productivity.
- Troubleshoot complex data pipeline and platform issues and drive root-cause analysis and long-term improvements.
- Help establish engineering standards, reusable frameworks, and best practices across the Global Data organization.
- Champion self-service data capabilities that improve the experience of data consumers and engineering teams.
Required Skills and Experience
- 6–10 years of professional experience in Data Engineering, Data Platform Engineering, or a closely related field.
- Strong hands-on experience developing ETL/ELT pipelines and data transformation workflows.
- Strong programming experience with Python.
- Strong SQL skills, including experience with complex queries, optimization, and large-scale data processing.
- Hands-on experience with dbt and modern data transformation practices.
- Hands-on experience with Apache Airflow and/or Cloud Composer for workflow orchestration.
- Strong experience working with Google BigQuery or a comparable cloud data warehouse.
- Experience with Google Cloud Storage (GCS) and cloud-based data platforms.
- Experience building and integrating data pipelines using REST APIs and other data integration mechanisms.
- Strong understanding of data modeling, data warehousing, and modern lakehouse architectures.
- Experience working with large-scale distributed data pipelines and production data environments.
- Experience with Git, CI/CD, automated testing, and DevOps practices.
- Strong problem-solving and analytical skills with the ability to troubleshoot complex data engineering issues.
- Ability to work effectively in a global, collaborative, and cross-functional environment.
- Excellent written and verbal communication skills.
Preferred Skills
- Strong experience with GCP data services, particularly BigQuery, GCS, Cloud Composer, and related services.
- Experience with Informatica CDC / Mass Ingestion or other change-data-capture technologies.
- Experience with additional cloud platforms such as AWS or Azure.
- Experience with data governance and metadata management platforms, such as Collibra or similar tools.
- Experience implementing data observability frameworks and tools.
- Experience with data quality, lineage, cataloging, and metadata management.
- Experience with Terraform or Infrastructure as Code.
- Experience with containerization and modern DevOps technologies.
- Exposure to AI/ML and GenAI technologies, including using LLMs to automate or enhance data engineering workflows.
- Experience building self-service data platforms, reusable engineering frameworks, or data products.
- Experience with data security, privacy, access controls, and enterprise data governance.
- Experience optimizing cloud data platforms for performance and cost.
What You Will Bring
- A hands-on engineering mindset with a passion for building production-grade data solutions.
- Strong ownership and accountability for the reliability and quality of data pipelines and platforms.
- A continuous-improvement mindset and enthusiasm for automation, standardization, and reusable frameworks.
- Ability to simplify complex technical problems and develop scalable solutions.
- Curiosity and willingness to learn and adopt emerging technologies, particularly AI/ML and GenAI.
- A strong focus on platform reliability, data quality, observability, and developer/user experience.
- Ability to collaborate effectively with globally distributed engineering, architecture, product, and business teams.
- Strong communication skills and the ability to explain complex technical concepts to both technical and non-technical audiences.
- Passion for building self-service capabilities that enable teams to discover, access, understand, and use data effectively.
Why Join Us
- Build and scale next-generation cloud data platforms powering enterprise-wide analytics, data products, and AI initiatives.
- Join Pearson, a global leader transforming lives through learning and innovation.
- Work hands-on with GenAI, AI-powered data engineering, and intelligent automation.
- Shape the future of self-service data, data products, governance, and observability at scale.
- Collaborate with high-performing global data and technology teams solving real-world, high-impact problems.
- Work with modern technologies across GCP, BigQuery, dbt, Airflow/Composer, Python, and AI/ML.
- Have the opportunity to influence data engineering standards, architecture, and platform strategy.
- Accelerate your career in a fast-evolving, innovation-driven global data ecosystem.