Cloud Engineer
Summary
Design and maintain petabyte-scale data pipelines and cloud platforms for a bank, using Azure, Spark, and Python to deliver governed, low-latency data products.
We are seeking a highly skilled Cloud Engineer to architect, engineer, and optimize our enterprise-level, petabyte-scale data infrastructure. In this role, you will be pivotal in translating raw, heterogeneous data into governed, low-latency, and actionable data products. You will own the end-to-end data lifecycle from ingestion and streaming to transformation and delivery ensuring data quality, semantic consistency, and metadata integrity. You will partner with cross-functional teams to advance the bank’s data-driven strategy by building scalable, fault-tolerant solutions that empower advanced analytics and machine learning.
Key Responsibilities:
- Data Pipeline Development: Architect and maintain robust data pipelines using Ab Initio, Python, Scala, Apache Spark, and Microsoft Fabric/Databricks for ingestion and transformation.
- Platform Management: Manage hybrid cloud platforms (Azure/Microsoft Fabric/SAS) and on-premises technologies (DB2, Netezza, Denodo), utilizing Infrastructure as Code (IaC) principles.
- Data Governance & Quality: Embed "data quality as code" by implementing automated validation, reconciliation, and auditing frameworks; manage technical metadata and lineage via the enterprise hub.
- Security & Compliance: Partner with CISO and Data Governance teams to enforce security policies, including data masking and anonymization, ensuring strict adherence to privacy regulations (e.g., POPIA).
- Cross-Functional Collaboration: Serve as a technical partner to Data Scientists and Business Analysts to translate business requirements into scalable, secure data solutions.
- Operational Excellence: Provide L2/L3 support for complex data incidents, ensuring minimal mean time to resolution and continuous optimization of data workflows through DataOps practices.
Required Skills and Qualifications:
- Education: Undergraduate Degree in a relevant field (Computer Science, Engineering, or IT).
- Experience: 2–5 years of professional experience in an IT or BI environment, with basic experience coordinating the work of others.
- Technical Expertise:
- Cloud Platforms: Deep expertise in Azure (Data Factory, Microsoft Fabric, Databricks).
- Programming: Proficiency in Python, PySpark, and SQL.
- DevOps/DataOps: Strong experience with CI/CD pipelines, orchestration tools, and Infrastructure as Code.
- Data Lifecycle: Competence in data conversion, profiling, and metadata management.
- Soft Skills:
- Tech Savvy: Ability to adopt and experiment with emerging technologies.
- Complex Problem Solving: Ability to distill complex, high-volume information into actionable insights.
- Communication: Highly effective at collaborating with diverse stakeholders and creating clear, compelling technical documentation.
Preferred Qualifications:
- Experience with legacy or hybrid data environments such as Netezza, DB2, or Denodo.
- Professional certification in Azure Data Engineering or Cloud Architecture.
- Experience in the financial services sector, specifically dealing with regulated data environments.