Data Engineer
The Data Engineer designs, builds, and maintains scalable data architecture, pipelines, and platforms that support analytics, reporting, and AI/ML solutions across the organization. This role is responsible for preparing, integrating, governing, and optimizing data from multiple sources to ensure reliable, accessible, and high-quality data assets for business and technical use. The Data Engineer works closely with data scientists, software engineers, analysts, and business stakeholders to deliver cloud-based data solutions that support operational goals, innovation, and emerging AI capabilities.
What You Will Do:
Design, develop, and maintain scalable data pipelines, data models, and lakehouse or data warehouse structures to support analytics, reporting, and AI/ML workloads.
· Build, test, and optimize ETL/ELT workflows that ingest, transform, and deliver data from multiple structured, semi-structured, and unstructured sources.
· Support AI/ML initiatives by preparing and managing datasets for feature engineering, model training, inference, and other AI-enabled applications.
· Implement and maintain data quality controls, validation processes, monitoring, and observability practices to improve reliability, accuracy, and performance of data pipelines.
· Manage data integration processes across internal and external systems to ensure timely, secure, and efficient data availability.
· Support data governance practices, including lineage, cataloging, access management, and documentation, to promote trusted and well-controlled data assets.
· Develop and deploy cloud-based data engineering solutions using modern platforms, distributed processing tools, and containerized environments where applicable.
· Build and maintain batch and near real-time data pipelines to meet business, reporting, and application requirements.
· Collaborate with cross-functional teams to gather requirements, translate business needs into technical solutions, and support enterprise data initiatives.
· Troubleshoot, enhance, and continuously improve data infrastructure, workflows, and related tools to increase scalability, efficiency, and business value.
· Remain current with developments in data engineering, cloud platforms, and AI-enabling technologies and apply practical improvements where beneficial.
· Comply with applicable ABS Health, Safety, Quality, and Environmental Management System requirements and other internal policies and procedures.
What You Will Need:
Education and Experience
- Degree in a technical field or equivalent combination of education and experience
- 8+ years of relevant experience in data engineering or a closely related discipline, including significant experience building and maintaining cloud-based data pipelines and architectures.
- Databricks certification (preferred)
Knowledge, Skills, and Abilities
Strong knowledge of data engineering concepts, including data modeling, data integration, ETL/ELT design, and pipeline orchestration.
· Strong programming skills in Python and SQL, with the ability to build, test, and maintain production-grade data workflows.
· Experience working with cloud-based data platforms such as AWS, Azure, or Google Cloud.
· Experience designing, implementing, and supporting data warehouse or lakehouse architectures such as Databricks, Snowflake, Redshift, or similar technologies.
· Knowledge of distributed data processing frameworks and tools, such as Spark or equivalent technologies.
· Ability to work with structured, semi-structured, and unstructured data from multiple data sources.
· Experience implementing data quality checks, validation frameworks, monitoring, and performance optimization practices.
· Knowledge of data governance, lineage, metadata management, and cataloging concepts and tools.
· Familiarity with data infrastructure requirements that support AI/ML workflows, including feature engineering, data preparation, and embedding pipelines.
· Experience with real-time or streaming data technologies such as Kafka, Kinesis, or similar tools is preferred.
· Familiarity with containerization and deployment tools such as Docker, Kubernetes, or similar technologies is preferred.
· Ability to analyze technical requirements, solve complex data problems, and deliver practical solutions in a fast-paced environment.
· Strong collaboration and communication skills, with the ability to work effectively across technical and business teams.
· Ability to manage multiple priorities, meet deadlines, and maintain a high standard of quality and accuracy.
· Working knowledge of the ABS Health, Safety, Quality, and Environmental Management System.
Reporting Relationships:
Reports directly to the Chief Data Scientist in Global Technology or other management or executive level position.
Working Conditions
Work is primarily sedentary; exerting up to 10 pounds of force occasionally and/or a negligible amount of force frequently or constantly to lift, carry, push, pull or otherwise move object.
Skills
- AI
- Analytics
- AWS
- Azure
- Cloud
- Containerization
- Data Engineering
- Data Governance
- Data Modeling
- Data Pipelines
- Data Quality
- Data Warehousing
- Databricks
- Docker
- ELT
- ETL
- Feature Engineering
- GCP
- Kafka
- Kinesis
- Kubernetes
- Lakehouse
- Machine Learning
- Metadata Management
- Observability
- Python
- Redshift
- Snowflake
- Spark
- SQL