Senior Software Engineer (Java/C++) — Query Engine / Data Platform

Summary

Senior backend engineer building and optimizing a distributed SQL query engine in Java/C++ for a large-scale data lakehouse platform used by Fortune 500 companies across finance, energy, manufacturing, and logistics.

About the Client:
Our client is a leading enterprise data platform company building an open, high-performance data lakehouse for AI and analytical workloads. The platform combines an intelligent SQL query engine, an AI-ready semantic layer, and an open catalog built on Apache Iceberg — enabling Fortune 500 companies across finance, energy, manufacturing, and logistics to unify, query, and govern data at massive scale across cloud and on-premise sources.

About the Role:
We are looking for a Senior Software Engineer to design, build, and own significant components of the core query engine of a large-scale distributed data platform. You will lead work across query planning, optimization, and execution — driving architectural decisions, mentoring middle engineers, and pushing the platform's performance and scalability boundaries across all major clouds.
This is a systems-level, backend engineering role focused on distributed data processing internals — not application development or CRUD services.

Responsibilities:
Architect, design, and implement significant components of the query engine — planner, optimizer, execution operators, memory management, and data access layers.
Write and review high-quality, performance-critical code in Java and/or C++.
Own end-to-end delivery of features from technical design through production rollout.
Drive performance optimization work — profiling, benchmarking, and eliminating bottlenecks in latency- and throughput-sensitive paths.
Integrate deeply with columnar formats, open table formats, and connectivity drivers.
Mentor middle engineers, lead code reviews, and set technical standards for the team.
Contribute to CI/CD architecture and quality processes across Jenkins, containerization, and multi-cloud deployment (GCP, AWS, Azure via Kubernetes / Docker).
Debug complex, cross-layer issues spanning query planning, distributed execution, memory management, and I/O.
Partner with US-based engineering leads on design reviews, architectural decisions, and roadmap execution.

Required Qualifications:
Education: B.S. or M.S. in Computer Science, Computer Engineering, or a related technical field.
Programming: Strong proficiency in Java or C++, with deep understanding of OOP, systems design, memory management, and concurrency. Candidates strong in both are especially valued.
SQL & Data: Advanced SQL skills and strong understanding of query execution internals — planning, optimization, and operator implementation.
Data processing systems experience: Demonstrated experience building or extending data processing systems — query engines, distributed databases, ETL/ELT frameworks, or analytical platforms.
Engineering experience: 5+ years in backend/systems software engineering, with a track record of owning significant components.
CI/CD & DevOps: Solid experience with Jenkins and modern engineering workflows.
Containers & Orchestration: Strong working knowledge of Docker and Kubernetes.
Cloud: Hands-on experience with at least one major cloud (GCP, AWS, or Azure); exposure to more than one is a strong plus.
Version Control: Confident with Git / GitHub workflows and rigorous code review practices.
English: Upper-Intermediate or higher (B2+) — daily written and verbal communication with a US-based engineering team.
Availability: Able to work EU hours with a 2–3 hour shift toward US West Coast time to ensure daily overlap with the client team.
Leadership: Prior experience mentoring engineers, leading technical designs, or owning components end-to-end.

Desired Skills:
Deep experience with Apache Arrow (columnar in-memory format).
Hands-on experience with Apache Calcite (SQL parsing, planning, optimization framework).
Experience with Gandiva or LLVM-based expression compilation and code generation.
Experience with open table formats — Apache Iceberg, Delta Lake, or Hudi.
Deep familiarity with distributed computing frameworks (e.g., Apache Spark, Kafka) and MPP SQL query engines (e.g., Presto, Trino, or similar).
Deep understanding of query planning, optimization (cost-based, rule-based), and execution internals.
Experience with data connectivity drivers: JDBC, ODBC, Arrow Flight.
Kubernetes on managed services (GKE / EKS / AKS) and multi-cloud exposure.
IaC tools such as Terraform.
Track record of performance optimization at the systems level — profiling, vectorization, cache/memory-conscious design.
Prior experience contributing to open-source data infrastructure projects.