Cloud Engineer
Summary
Designs and maintains scalable cloud data pipelines on Azure, using Python, PySpark, Databricks, and Fabric to ingest, transform, and deliver data for analytics and reporting.
Essential Requirements:
- Undergraduate Degree
- General Experience: Experience enables the job holder to deal with the majority of situations and to advise others (2 to 5 years in an IT or BI environment)
- Managerial Experience: Basic experience of coordinating the work of others (4 to 6 months)
- Data Collection and Analysis: Works with full competence to determine and analyze trends from collected data to assist in compiling reports that support business decisions
- Typically works without supervision and may provide technical guidance
- Engineering: Deep expertise in the major cloud Azure platform, Azure Data Factory, Microsoft Fabric, Databricks, Python, PySpark. Infrastructure as Code
- Works with full competence to create and run queries and interact with various database interfaces and query languages
- Typically works without supervision and may provide technical guidance. Experience with low-latency data Ingestion and processing using Message queues, Kafka, Azure Event Hub, Real time Intelligence
- Data Conversion: Works with full competence to use data conversion tools and techniques to encode data in various formats
- Typically works without supervision and may provide technical guidance
- Database Reporting: Works with full competence to use database reporting tools and techniques. Typically works without supervision and may provide technical guidance
- Proven experience with workflow orchestration and Implementing CI/Cd pipelines for data solutions
- Application Development: Works with full competence to develop software through use of programming languages
- Typically works without supervision and may provide technical guidance
- Works with full competence to utilize systems and tools to support categorizing and classifying data and information
- Typically works without supervision and may provide technical guidance
- Architect, engineer and maintain robust, scalable and fault-tolerant data pipelines using technologies like Ab Initio, Python/Scala, Apache Spark, Microsoft Fabric, and Databricks for ingestion, transformation, and delivery of structured, semi-structured, and unstructured data. Ensure efficient performance-tuned extraction and loading of data into the enterprise Data Warehouse and Data Lake, focusing on high availability, cost efficiency, and reliability
- Architect and manage the robust, hybrid cloud data platform utilising Azure Services (Microsoft Fabric, Ab Initio, Synapse, Databricks, SAS) and potentially on-premises technologies (DB2 Warehouse, Netezza, Denodo, Ab Initio). Demonstrate proficiency in infrastructure as code
- Establish and manage comprehensive observability for all data infrastructure components such as databases, data lakes, and data warehouses to meet strict service level objectives and aligned with architectural standards
- Embed data quality as code by implementing automated, unit and end-to-end data validation, reconciliation, and auditing frameworks within the CI/CD data pipelines
- Design and maintain the automated capture and maintenance of technical and operational metadata, ensuring complete automated data lineage within the enterprise metadata hub (Ab Initio)
- Serve as trusted technical partner, collaborating with business stakeholders, Data Scientists, Analysts, and Architects to translate requirements into scalable data solutions
- Drive the full data lifecycle of data products and end-to-end ownership of data engineering initiatives, ensuring timely delivery and alignment with business objectives
- Drive continuous optimisation of data engineering practices, standards, and processes to improve efficiency and performance
- Mentor and guide junior engineers, providing coaching and technical support to build capability
- Stay abreast of emerging technologies and industry trends to enhance data engineering maturity
- Contribute to innovation initiatives that align with the data-driven strategy