Data Engineer
Summary
Designs and maintains scalable data pipelines and cloud-native platforms to ensure reliable, high-quality data for analytics, BI, and AI use cases using SQL, Python, and Azure.
Job Title: Data Engineer
Reporting Line: Sr. Data Engineer
Division: Support Services - Business Excellence and IT
Department: Information Technology
The Data Engineer designs, builds, and optimizes scalable data pipelines and platforms that enable advanced analytics, BI, and AI use cases across the organization. This role ensures data availability, quality, lineage, and reliability, while leveraging modern cloud‑native architectures and engineering best practices. The Data Engineer collaborates closely with data scientists, BI analysts, and business teams to deliver secure, governed, and high‑performance data solutions.
Key Responsibilities and Performance Standards
- Design, develop, and maintain scalable batch and real‑time pipelines (ELT/ETL).
- Build and maintain data lakes, lakehouse layers, and curated datasets.
- Develop high‑performance SQL and Python code optimized for large datasets.
Data Reliability, Quality & Observability
- Set up monitoring, alerting, and logging using observability tools.
- Improve pipeline performance, optimize compute and storage costs, and ensure SLA adherence.
- Contribute to data quality rules, governance, and metadata management.
- Develop and deploy data solutions on cloud platforms (preferably Azure).
- Build infrastructure‑as‑code (Terraform, ARM, Bicep) for repeatable deployments.
- Implement CI/CD pipelines for data workflows (GitHub Actions, Azure DevOps).
- Partner with data scientists to prepare training datasets and MLOps pipelines.
- Work with BI teams to provide optimized semantic layers and analytical datasets.
- Participate in design reviews, architecture discussions, and sprint planning.
Documentation & Standards
- Maintain technical documentation, data dictionaries, lineage, and mappings.
- Contribute to engineering standards, best practices, and reusable templates.
Key Technical Skills and Proficiency Levels
- SQL and Relational Databases – Advanced
- Big Data Technologies (Spark, Hadoop) – Advanced
- Data Warehousing (Fabric preferred) – Intermediate
- Programming (Python, Scala) – Advanced
- Data Orchestration – Advanced
- Data Quality and Governance – Intermediate
Communications and Working Relationships
- Internal: Collaborates with Data Scientists, BI Analysts, and business stakeholders.
- External: May interact with technology vendors and service providers.
Selection Criteria – Essential
- Bachelor’s degree in Computer Science, Engineering, or related field
- 5+ years of experience in data engineering
- Proficiency in SQL and Python
- Experience with cloud data platforms
Desirable
- Master’s degree
- Experience with real‑time data processing
- Knowledge of data governance practices