Kafka + Java + Data Engineer

Summary

Designs and builds high-volume data pipelines and streaming systems using Java, Kafka, and Hadoop/Spark, with a focus on real-time processing and cloud integration.

Kafka + Java + Data Engineers\:

- Extensive experience designing and developing streaming pipelines, data platform to handle high volume data processing, distributed system for both batch and realtime

- Strong Expertise in– Java, Kafka and Hadoop ecosystem (Spark, Hive, YARN, HDFS), Scala, Cassandra, Linux (on prem), CI-CD (Jenkins, Github), Obs (Dynatrace).

- Design and implement real-time streaming pipelines using Kafka

- Develop producers and consumers for high-throughput, fault-tolerant systems

- Responsible for both development and strong expertise to debug and fix any post production issues

- Implement Kafka-based event-driven architectures

- Experience with Databricks, Azure Cloud PaaS Services, Azure Data Factory, ADLS, Kubernetes, GitHub, Docker, and CI/CD workflows

- Experience in scripting, ETL (Pandas, Pola.rs, PySpark, Ibis), SQL, and NoSQL databases (Cassandra, MongoDB)

- Understanding of system architecture concepts\: distributed systems, caching, replication, data modeling

- Leads \: Role is 70% of time in Design, Architect, Coding and 30% of time in team mentoring, guidance, coordination

- Responsible for both development and strong expertise to debug and fix any post production issues for their developed code

- Interview ask \: To write optimal code (in Java/Kafka/Python/Pyspark) in timeboxed period for scenario given (covering data quality, transformation) without AI tool and code to be accurate. To resolve code issue for a bug introduced in sample code

- Excellent written and verbal communication skills

- Experience working in Agile environments; familiarity with cloud platforms

- Experience with orchestration tools (e.g. NiFi, Griffin, Hamilton, Airflow)

- Functional experience of Retail Banking, capital markets, securities processing will be advantage