Senior DevOps / SRE Cloud Engineer
Summary
A senior engineer who owns an Azure-based Kubernetes (AKS) platform end-to-end for a utility-scale power generation and energy storage company: writing Terraform IaC, building GitHub Actions CI/CD pipelines, running observability/SRE (Prometheus, Grafana, SLOs), and securing the platform with Entra ID and Key Vault. Requires 5+ years in DevOps/SRE, incl. production K8s and Azure.
A company that develops, owns, and operates utility-scale power generation and energy storage plants across the African continent is seeking a Senior DevOps / SRE Cloud Engineer who will own the company's Azure-based Kubernetes platform end-to-end - remote or hybrid based on location .
Responsibilities:
Infrastructure as Code: Provision and operate production Kubernetes (AKS) and Azure infrastructure using modular Terraform.
CI/CD Automation: Maintain GitHub Actions pipelines featuring quality gates, security checks, and automated deployments.
Observability & SRE: Own monitoring, alerting, SLOs, capacity planning, and incident response across environments.
Security & Access: Manage platform secrets, network controls (VNets, private endpoints), and Entra ID identity governance.
Developer Enablement: Partner with engineering teams to resolve operational bottlenecks and drive reliability standards.
Minimum Requirements:
Experience: 5+ years in DevOps/SRE/Platform Engineering (including 2-3 years operating K8s in production and 2+ years on Azure).
Education: Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
Certifications: CKA, CKAD, or Azure Solutions Architect / DevOps Engineer certifications are advantageous.
Soft Skills: Strong problem-solving ability, clear technical communication, and the capacity to operate autonomously within a hybrid/remote team.
Azure Platform: Production experience with AKS, ADLS Gen2, Key Vault, Entra ID (Workload/Managed Identities), VNets/Private Endpoints, and Service Bus (or equivalent broker).
Production Kubernetes: Advanced operational depth in Helm chart authoring, operator deployment, node-pool sizing, pod troubleshooting, and cluster upgrades.
Terraform: Proven ability to author modular IaC, manage state safely, and maintain strict plan/apply disciplines.
Containerization & CI/CD: Hands-on Docker (multi-arch builds, optimization) and GitHub Actions pipeline development with required status checks.
Observability: Practical experience configuring Prometheus + Grafana, structured logging, and SLO-based alerting.
Linux & Automation: Strong Bash scripting with clean, code-maintained operational tooling.
Preferred Experience
Data Platforms: Apache Spark on K8s (Spark Operator/Connect), JupyterHub, Delta Lake, Trino, Hive Metastore, or Apache Ranger.
Advanced Telemetry: OpenTelemetry (SDKs/Collector topology) and OpenLineage/Marquez integration.
Languages & Databases: Intermediate Python (infrastructure testing/pytest) and basic DBA management for SQL Server and PostgreSQL.
Compliance & Isolation: Multi-tenant architecture design, dependency auditing, image provenance, and POPIA/GDPR/ISO 27001 controls.
Benefits:
Competitive salary based on experience (salary can potentially be more based on experience/skills)
What they ask for
Required
- 5+ years in DevOps/SRE/Platform Engineering, including 2-3 years operating Kubernetes in production and 2+ years on Azure
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
- Strong problem-solving, clear technical communication, and ability to operate autonomously in a hybrid/remote team
- Production experience with AKS, ADLS Gen2, Key Vault, Entra ID (Workload/Managed Identities), VNets/Private Endpoints, and Service Bus (or equivalent broker)
- Advanced production Kubernetes: Helm chart authoring, operator deployment, node-pool sizing, pod troubleshooting, cluster upgrades
- Terraform: authoring modular IaC, safe state management, strict plan/apply discipline
- Hands-on Docker (multi-arch builds, optimization) and GitHub Actions pipeline development with required status checks
- Observability: configuring Prometheus + Grafana, structured logging, SLO-based alerting
- Strong Bash scripting with clean, code-maintained operational tooling
Preferred
- CKA, CKAD, or Azure Solutions Architect / DevOps Engineer certifications
- Data platforms: Apache Spark on K8s (Spark Operator/Connect), JupyterHub, Delta Lake, Trino, Hive Metastore, or Apache Ranger
- Advanced telemetry: OpenTelemetry (SDKs/Collector topology) and OpenLineage/Marquez integration
- Intermediate Python (infrastructure testing/pytest) and basic DBA management for SQL Server and PostgreSQL
- Multi-tenant architecture design, dependency auditing, image provenance, and POPIA/GDPR/ISO 27001 controls