Kafka + Java + Data Engineer
Summary
Designs and builds high-volume data pipelines and streaming systems using Java, Kafka, and Hadoop/Spark, with a focus on real-time processing and cloud integration.
Kafka + Java + Data Engineers\:
- Extensive experience designing and developing streaming pipelines, data platform to handle high volume data processing, distributed system for both batch and realtime
- Strong Expertise in– Java, Kafka and Hadoop ecosystem (Spark, Hive, YARN, HDFS), Scala, Cassandra, Linux (on prem), CI-CD (Jenkins, Github), Obs (Dynatrace).
- Design and implement real-time streaming pipelines using Kafka
- Develop producers and consumers for high-throughput, fault-tolerant systems
- Responsible for both development and strong expertise to debug and fix any post production issues
- Implement Kafka-based event-driven architectures
- Experience with Databricks, Azure Cloud PaaS Services, Azure Data Factory, ADLS, Kubernetes, GitHub, Docker, and CI/CD workflows
- Experience in scripting, ETL (Pandas, Pola.rs, PySpark, Ibis), SQL, and NoSQL databases (Cassandra, MongoDB)
- Understanding of system architecture concepts\: distributed systems, caching, replication, data modeling
- Leads \: Role is 70% of time in Design, Architect, Coding and 30% of time in team mentoring, guidance, coordination
- Responsible for both development and strong expertise to debug and fix any post production issues for their developed code
- Interview ask \: To write optimal code (in Java/Kafka/Python/Pyspark) in timeboxed period for scenario given (covering data quality, transformation) without AI tool and code to be accurate. To resolve code issue for a bug introduced in sample code
- Excellent written and verbal communication skills
- Experience working in Agile environments; familiarity with cloud platforms
- Experience with orchestration tools (e.g. NiFi, Griffin, Hamilton, Airflow)
- Functional experience of Retail Banking, capital markets, securities processing will be advantage