Senior Data Engineer

Summary

Designs and builds scalable data pipelines and cloud-native architectures for government and enterprise clients, focusing on high-performance data systems and AI integration.

Overview

At Xtremax, we help government agencies and enterprises build robust, scalable, and future-ready digital systems. As a Senior Data Engineer, you will play a key role in designing scalable data architectures, translating business requirements into robust technical specifications, and building high-performance data pipelines. This role combines technical leadership, data systems solutioning, and stakeholder collaboration to deliver impactful, data-driven digital solutions. Candidates with public sector experience are preferred, as this role supports IT projects for government agencies.

Responsibility:

  • Data Pipeline Infrastructure & Architecture
  • Design and implement scalable data architectures on cloud data platforms with high availability, security, and performance
  • Lead development of Data Lakehouse solutions
  • Collaborate with stakeholders to understand requirements and translate them into technical specifications
  • Pipeline Development & Optimisation
  • Build and maintain robust ETL/ELT pipelines using modern data engineering tools and frameworks
  • Optimise data processing workflows for performance, cost-effectiveness, and reliability
  • Implement automated data quality checks and monitoring systems to ensure data integrity
  • Data Systems Architecting & Solutioning
  • Design and architect comprehensive cloud-native Data & AI solutions aligned with business objectives and technical requirements
  • Lead cloud migration strategies and oversee implementation of complex multi-cloud environments
  • Drive innovation through integration of Data & AI capabilities into enterprise platform architectures Data & AI platform product architectures
  • Conduct technical assessments and recommend modernised approaches using cloud native technologies
  • Maintain architectural documentation
  • Cloud Platform Operations
  • Leverage Cloud Native Services to build and manage data infrastructure
  • Implement infrastructure as code practices using Terraform
  • Ensure compliance with security standards and data governance policies
  • Technical Leadership & Collaboration
  • Mentor junior data engineers and provide technical guidance on complex challenges
  • Participate in architectural reviews and contribute to data strategy evolution

Requirements:

  • Bachelor’s degree in computer science, Information Technology, Computer Engineering, or related field
  • Minimum 3 years of relevant experience in data systems architecture, data systems integration, and data pipeline setup at production scale
  • Good understanding of cloud computing principles including infrastructure as code, containerisation, microservices architecture, cloud security frameworks, identity and access management, network architecture, and distributed systems
  • Proven ability to translate business requirements into technical solutions
  • Excellent communication skills for presenting complex concepts to diverse audiences
  • Experience with cloud security frameworks, compliance requirements, and risk management
  • Experience in data domains (e.g. DataOps, Data Lakehouse) and AI/ML Domains (e.g. MLOps, LLMOps)
  • Strong Knowledge and Hands-on experience with SQL, Python and Apache Spark
  • Hands-on experience with Apache Kafka, Airflow,or similar technologies

Good to Have (Optional):

  • Proficiency in Amazon Web Services (AWS) ecosystem and relevant cloud certifications (e.g., AWS Solutions Architect Professional, AWS Data Engineer Associate).
  • Hands-on experience with Data & AI cloud-native services (e.g., Amazon SageMaker, AWS S3, AWS Glue, AWS Lake Formation, AWS Bedrock).
  • Familiarity with MLOps, ML model deployment pipelines, serverless computing, edge computing, or IoT architectures.
  • Knowledge of metadata management tools, data governance frameworks, or business intelligence / data visualization tools (e.g., Power BI, Tableau).