Senior Data Engineer / Scala, Spark
Summary
A senior data engineer who designs, builds, and maintains large-scale distributed data systems and pipelines for knowledge graph quality and integrity. Day-to-day work centers on Scala/Java, Spark, streaming (Kafka), and distributed storage (Cassandra) in a multi-datacenter environment.
- Design, build, and maintain large-scale distributed data systems
- Develop data processing pipelines using technologies such as Spark and streaming solutions
- Work on systems that evaluate and ensure the quality and integrity of large-scale knowledge graphs
- Contribute to data-intensive projects involving analytics, data discovery, and algorithm development
- Collaborate with cross-functional teams across multiple locations and product areas
- Participate in architecture and design decisions for scalable, multi-datacenter solutions
- Continuously improve system performance, scalability, and reliability
- Challenge existing assumptions and propose innovative data engineering solutions
- Hands-on professional software development experience of 5+ years
- Strong programming skills in Scala and/or Java
- Experience working with Big Data technologies such as Spark and distributed systems
- Hands-on experience with data streaming (e.g. Kafka) and distributed data storage (e.g. Cassandra)
- Strong problem-solving mindset and ability to work on complex, unsolved challenges
- Excellent written and verbal communication skills
- Ability to work effectively in distributed, cross-functional teams
- Proactive and self-driven approach to engineering challenges
- Experience with Knowledge Graphs or graph-based data systems
- Background in machine learning or data science
- Experience in music information retrieval
- Exposure to advanced data modeling and algorithm development