Kafka Platform / Site Reliability Engineer
Summary
Maintains and scales Kafka clusters, Kubernetes operators, and GitOps workflows for a streaming platform, troubleshooting distributed systems in a Kubernetes-native environment.
• Operate and
maintain production Kafka clusters and surrounding ecosystem (Schema Registry,
Kafka Connect)
• Develop and manage
Kubernetes CRDs and Operators for Kafka resource lifecycle management
• Build and maintain
Infrastructure-as-Code (Terraform/Terragrunt/Ansible) for the streaming
platform
• Implement and
maintain GitOps workflows for declarative infra and Kafka resource management
• Troubleshoot and
resolve incidents in distributed streaming systems running on Kubernetes
• Collaborate with
an international engineering team
Requirements
• Hands-on
experience operating Apache Kafka in production, including Schema Registry and
Kafka Connect
• Strong working
knowledge of Kubernetes, including writing/managing CRDs and Operators
• Experience with
Infrastructure as Code (Terraform; ideally Terragrunt and Ansible)
• Familiarity with
GitOps-style workflows (declarative infra/Kafka resource management via code +
K8s CRDs)
• Experience
troubleshooting distributed streaming systems in a Kubernetes-native
environment
• Strong English
communication skills, international team
Nice to Have
• Go knowledge
(Operators/internal tooling written in Go)
• Kafka on Aiven
experience
• German language