Kafka + No SQL DB
Summary
Designs and builds high-volume streaming pipelines and distributed data platforms using Kafka, NoSQL databases, and big-data tools, while mentoring a team.
KAFKA, No SQL DB\:
- Extensive experience designing and developing streaming pipelines, data platform to handle high volume data processing, distributed system for both batch and realtime
- Strong Expertise in–KAFKA and Hadoop ecosystem (Spark, Hive, YARN, HDFS), Scala, Cassandra, Linux (on prem), CI-CD (Jenkins, Github), Obs (Dynatrace).
- Design and implement real-time streaming pipelines using Kafka for onprem systems
- Develop producers and consumers for high-throughput, fault-tolerant systems
- Implement Kafka-based event-driven architectures
- Leads \: Role is 70% of time in Design, Architect, Coding and 30% of time in team mentoring, guidance, coordination
- Responsible for both development and strong expertise to debug and fix any post production issues for their developed code
- Interview ask \: To write optimal code (in Kafka/SQL) in timeboxed period for scenario given (covering data quality, transformation) without AI tool and code to be accurate. To resolve code issue for a bug introduced in sample code
- Experience in scripting, ETL (Pandas, Pola.rs, PySpark, Ibis), SQL, and NoSQL databases (Cassandra, MongoDB)
- Understanding of system architecture concepts\: distributed systems, caching, replication, data modeling
- Excellent written and verbal communication skills
- Experience working in Agile environments; familiarity with cloud platforms (Azure) is a plus
- Experience with orchestration tools (e.g. NiFi, Griffin, Hamilton, Airflow)
- Functional experience of Retail Banking, capital markets, securities processing will be advantage