Data Engineering & Analytics Engineer (Public Sector)
- 1-year contract, renewable
- Government project
- Hybrid work arrangement
We are looking for a Data Engineering & Analytics Engineer to design, build, and operate the data pipelines and data products supporting the future SSOE platform.
You will transform data from enterprise systems, operational platforms, applications, and infrastructure into trusted and usable data products for applications, operational reporting, analytics, capacity planning, and decision-making.
The role spans on-premise, GCC and hybrid environments, with an emphasis on production-grade data engineering using cloud-native and managed data capabilities.
What You Will Be Working On
As a Data Engineering & Analytics Engineer, you will own the data lifecycle from source systems through ingestion, transformation, modelling, quality, and serving.
You will build pipelines that extract and ingest data from enterprise and operational systems, transform it into consistent and trusted datasets, and make that data available to applications, dashboards, reporting, analytics, and machine-learning use cases.
You will work closely with the Logging & Data Platform Engineer on shared platform capabilities and with Software Engineers and other consumers to define reliable data interfaces and products.
Key Responsibilities
Data Pipeline Engineering
- Design, build, and operate production-grade data pipelines for data extraction, ingestion, transformation, and loading (ETL/ELT)
- Integrate data from on-premises systems, enterprise applications, APIs, databases, SaaS platforms, files, streams, cloud services, and other operational data sources
- Develop batch, incremental, change-data-capture (CDC), streaming, and event-driven ingestion patterns based on source-system and business requirements
- Build transformation pipelines that clean, enrich, standardise, join, aggregate, and structure raw data into trusted datasets
- Design secure and resilient mechanisms for transferring and synchronising data between on-premises, GCC, AWS, Azure, and other approved environments
- Design pipelines for failure handling, retry, recovery, idempotency, scalability, and changing data volumes
- Automate pipeline deployment, configuration, testing, and operation
Data Architecture & Modelling
- Design and maintain cloud-native and hybrid data stores, data lakes, and analytical datasets
- Develop data models that provide consistent representations of enterprise, operational, and asset information
- Define schemas and data contracts between data producers and downstream consumers
- Design data structures appropriate for operational applications, reporting, analytics, and machine-learning workloads
- Apply backwards-compatible schema changes and coordinate changes that may affect downstream consumers
- Maintain data lineage and metadata so datasets are traceable and discoverable
- Work with platform and application teams to define appropriate data-serving and integration patterns
Data Quality & Reliability
- Implement automated data validation, reconciliation, completeness, consistency, and quality controls throughout the pipeline lifecycle
- Monitor data freshness, pipeline health, processing latency, and data-quality indicators
- Detect and investigate ingestion failures, source-system changes, data-quality anomalies, and reconciliation differences
- Prevent invalid or incomplete data from silently propagating to downstream consumers
- Define appropriate SLOs for data freshness, availability, and pipeline reliability
- Build monitoring, alerting, error handling, and recovery into data pipelines from the outset
Analytics & Data Products
- Build trusted datasets and reusable data products for applications, dashboards, operational reporting, and analytics
- Develop datasets supporting asset intelligence, operational visibility, capacity planning, trend analysis, and decision-making
- Enable advanced analytics and machine-learning use cases using cloud-native data, analytics, and AI/ML capabilities
- Work with users and stakeholders to translate operational questions into appropriate datasets, metrics, and analytical products
- Support exploratory analysis and prototyping where required before operationalising successful approaches
- Ensure analytical outputs are based on governed, traceable, and reproducible data
Data Integration
- Design data architectures spanning on-premise infrastructure and cloud platforms
- Integrate traditional enterprise systems with modern cloud-native data capabilities
- Design for connectivity constraints, network boundaries, security zones, and data-residency requirements
- Implement appropriate buffering, checkpointing, retry, and reconciliation where data crosses environment boundaries
- Select appropriate integration patterns based on data volume, latency, source-system capability, and operational requirements
- Work with infrastructure, network, security, and platform teams to establish secure data flows
Security & Governance
- Ensure data is collected, transmitted, stored, processed, and accessed according to applicable security requirements
- Enforce appropriate access controls and least-privilege principles for data platforms and pipelines
- Ensure sensitive information is appropriately classified and protected throughout the data lifecycle
- Maintain auditability and traceability of data-processing activities
- Apply retention, archival, lifecycle, and deletion requirements to data products
- Participate in security, architecture, data-governance, and operational-readiness reviews
Reliability & Operations
- Operate and support production data pipelines and data products
- Participate in operational support and on-call responsibilities for owned services
- Investigate production incidents and contribute to root-cause analysis and preventative improvements
- Monitor pipeline performance, capacity, reliability, and cost
- Maintain architecture documentation, data definitions, operational procedures, and runbooks
- Continuously improve pipeline automation, reliability, performance, and maintainability
What We Are Looking For
Experience
- Minimum 3–5 years of experience in data engineering, cloud data engineering, analytics engineering, software engineering, or a related discipline
- At least 2 years of hands-on experience designing, building, and operating production-grade data pipelines
- Demonstrated experience with data extraction, ingestion, ETL/ELT, transformation, data modelling, and data quality
- Experience using AWS and/or Azure native data capabilities
- Experience integrating data from APIs, databases, enterprise systems, files, or streaming sources
- Experience implementing batch, incremental, CDC, and/or event-driven data pipelines
- Experience working with on-premises and/or cloud environments, with an understanding of hybrid integration patterns
- Experience applying software-engineering practices such as version control, automated testing, CI/CD, monitoring, and Infrastructure as Code to data solutions
Technical Skills
- Cloud: AWS/Azure-native logging, streaming, storage, search and data services
- On-Premise: Enterprise servers, networks, applications, databases, virtualised infrastructure, and log sources
- Data Engineering: Python, SQL, ETL/ELT, batch, incremental, CDC, streaming, and event-driven patterns
- Data Modelling: Relational, dimensional, analytical, and domain-oriented data modelling
- Data Quality: Validation, reconciliation, quality monitoring, lineage, and anomaly detection
- Infrastructure as Code: Terraform / OpenTofu
- CI/CD: GitLab CI/CD, SHIP-HATS or equivalent automated deployment practices
- Analytics: Data preparation, analytical datasets, reporting, statistical analysis, and ML enablement
Engineering Practices
- Treats data pipelines and data products as production software, not one-off scripts
- Keeps pipeline code, schemas, infrastructure, and configuration under version control
- Uses automated testing, CI/CD, Infrastructure as Code, and monitoring
- Designs pipelines for failure, retry, idempotency, scalability, and changing workloads
- Validates data at ingestion and transformation boundaries
- Establishes explicit data contracts between producers and consumers
- Understands when to use managed cloud-native capabilities rather than unnecessarily building and operating infrastructure
- Considers downstream consumers before making schema or behavioural changes
- Automates repeatable data-processing and operational activities
- Balances technical excellence with pragmatic delivery and operational sustainability
Nice to Have
- Experience with Singapore Government platforms such as TechPass, SHIP-HATS, SEED, and GCC
- Familiarity with OC/SN data-classification requirements
- AWS or Azure cloud certifications
- Experience designing data architectures spanning on-premises and cloud environments
- Experience building data products consumed by applications, dashboards, operational teams, or leadership
- Experience with managed cloud analytics and AI/ML capabilities
- Experience with data lineage, metadata management, and data cataloguing
- Experience working with enterprise asset management, MDM, network, procurement, or operational systems