Kafka + No SQL DB

Summary

Designs and builds high-volume streaming pipelines and distributed data platforms using Kafka, NoSQL databases, and big-data tools, while mentoring a team.

KAFKA, No SQL DB\:

- Extensive experience designing and developing streaming pipelines, data platform to handle high volume data processing, distributed system for both batch and realtime

- Strong Expertise in–KAFKA and Hadoop ecosystem (Spark, Hive, YARN, HDFS), Scala, Cassandra, Linux (on prem), CI-CD (Jenkins, Github), Obs (Dynatrace).

- Design and implement real-time streaming pipelines using Kafka for onprem systems

- Develop producers and consumers for high-throughput, fault-tolerant systems

- Implement Kafka-based event-driven architectures

- Leads \: Role is 70% of time in Design, Architect, Coding and 30% of time in team mentoring, guidance, coordination

- Responsible for both development and strong expertise to debug and fix any post production issues for their developed code

- Interview ask \: To write optimal code (in Kafka/SQL) in timeboxed period for scenario given (covering data quality, transformation) without AI tool and code to be accurate. To resolve code issue for a bug introduced in sample code

- Experience in scripting, ETL (Pandas, Pola.rs, PySpark, Ibis), SQL, and NoSQL databases (Cassandra, MongoDB)

- Understanding of system architecture concepts\: distributed systems, caching, replication, data modeling

- Excellent written and verbal communication skills

- Experience working in Agile environments; familiarity with cloud platforms (Azure) is a plus

- Experience with orchestration tools (e.g. NiFi, Griffin, Hamilton, Airflow)

- Functional experience of Retail Banking, capital markets, securities processing will be advantage