Site Reliability Engineer (SRE) / DevOps Engineer -
Summary
Build and operate Kafka-based messaging platforms, applying SRE principles to improve reliability and performance while automating operations with Ansible, scripts, and CI/CD.
Who are we: Fulcrum Digital is an agile and next‑generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.
We’re looking for a Kafka Messaging / SRE Engineer to join our growing platform engineering team and help build, operate, and scale mission‑critical messaging services.
What You’ll Do
- Own and operate Kafka‑based messaging platforms in production environments.
- Apply SRE principles to improve reliability, availability, and performance.
- Drive DevOps & automation initiatives to reduce toil and manual operations.
- Build and enhance automation using Ansible, scripts, and CI/CD pipelines.
- Perform incident management, RCA, capacity planning, and operational readiness.
- Collaborate closely with application and platform engineering teams.
- Contribute to Java‑based tooling and platform enhancements.
What We’re Looking For
- 3‑6 years of experience working with Kafka / messaging systems.
- Strong understanding of Kafka architecture (brokers, topics, partitions, replication).
- Hands‑on experience with SRE / DevOps practices.
- Proven skills in automation (Ansible, scripting, CI/CD).
- Java development background (ability to debug, enhance, or build platform tools).
- Experience with Linux, distributed systems, monitoring & alerting.
- Exposure to incident response, production support, and operational excellence.