IT Operations Engineer (Kafka Platform Support)

Open 39d posting dated 3 weeks ago

In Cyclad we work with top international IT companies in order to boost their potential in delivering outstanding, cutting-edge technologies that shape the world of the future. We are seeking an experienced IT Operations Engineer (Kafka Platform Support) to join a global platform team. In this role, you will be responsible for the operation and support of a Kafka-based messaging platform in production, ensuring its availability, stability, and performance.

Project information:

  • Office location: Poland

  • Work mode: 100% Remote

  • Office location: Remote from Poland

  • Budget: 110- 130 PLN net/ h- B2B

  • Project length: Long-term

  • Only candidates with citizenship in the European Union and residence in Poland

  • Start date: ASAP (depending on candidate’s availability)

Project scope:

  • Operate, monitor, and maintain a Kafka-based messaging platform in a production environment

  • Ensure platform availability, stability, and performance in line with operational SLAs

  • Monitor system health using logs, metrics, and alerting tools

  • Perform routine operational checks and maintenance activities

  • Handle incidents and service requests via ticketing systems and support channels

  • Troubleshoot issues across Kafka components (brokers, producers, consumers, integrations)

  • Analyze logs, metrics, and system behavior to identify root causes of incidents

  • Execute operational procedures based on runbooks and standard operating procedures (SOPs)

  • Perform configuration changes (topics, access controls, settings) following established processes

  • Maintain and continuously improve operational documentation and runbooks

  • Act as a primary support contact for internal users of the Kafka platform

  • Provide technical support via collaboration tools (e.g., Slack, Teams)

  • Assist users with troubleshooting and best practices

  • Translate user-reported issues into actionable insights for technical teams

  • Collaborate closely with engineering and platform teams to resolve incidents

  • Participate in incident reviews and post-mortems

  • Identify recurring operational issues and suggest improvements or automation opportunities

  • Contribute to improving platform usability and support processes

Competence demands:

  • Minimum 3–6 years of experience in IT operations, production support, or platform support roles

  • Hands-on experience with Apache Kafka or similar event streaming platforms

  • Strong understanding of distributed systems (partitioning, replication, scaling)

  • Strong troubleshooting skills in production IT environments

  • Experience with monitoring, logging, and alerting tools (e.g., Grafana, Prometheus)

  • Knowledge of Git and version control practices

  • Familiarity with GitLab CI/CD and working with existing pipelines

  • Experience working with incident management processes and support tools

  • Experience working with runbooks, SOPs, and structured operational environments

  • Strong communication skills and ability to explain technical issues clearly

  • Experience working with internal customers and cross-functional teams

  • Customer-focused mindset with a proactive approach to support

  • Fluency in English (written and spoken)

Nice to have:

  • Experience with AWS or other cloud platforms

  • Familiarity with Kubernetes and containerized environments

  • Experience with monitoring tools such as Grafana and Prometheus

  • Understanding of automation in IT operations and platform support environments

We offer:

  • Remote working model

  • Full-time job agreement based on b2b

  • Private medical care with dental care (covering 70% of costs)

  • Multisport card (also for an accompanying person)

  • Life insurance

Recruitment process:

  • Introductory call

  • Technical interview

  • Final decision