freehire launches on Product Hunt on 26 August.

Follow →

Middle SDET — Data Platform / Query Engine

Open 25d

Summary

Middle SDET role testing a high-performance, distributed data lakehouse platform with SQL query engine, semantic layer, and Apache Iceberg catalog, validating correctness, performance, and reliability across cloud and on-premise environments.

About the Client:
Our client is a leading enterprise data platform company building an open, high-performance data lakehouse for AI and analytical workloads. The platform combines an intelligent SQL query engine, an AI-ready semantic layer, and an open catalog built on Apache Iceberg — enabling Fortune 500 companies across finance, energy, manufacturing, and logistics to unify, query, and govern data at massive scale across cloud and on-premise sources.

About the Role:
We are looking for a Middle-level SDET to design and maintain automated tests for a large-scale distributed data platform. You will validate the correctness, performance, and reliability of the query engine, connectivity layer, and backend services across all major clouds.
This is a data-heavy, backend-focused role. We are not looking for web/UI QA engineers — the work centers on validating distributed query execution, data correctness at scale, connectivity drivers, and backend microservices.

Responsibilities:
Develop and maintain automated tests in Python / pytest for backend services, REST APIs, and distributed data components.
Write end-to-end, integration, and regression tests covering SQL query execution, data correctness, and platform APIs.
Support performance and load testing with JMeter, including workloads over JDBC / ODBC / Arrow Flight drivers.
Validate data-intensive scenarios — query plans, result correctness across data sources, metadata consistency, and behavior under concurrency.
Contribute to CI/CD test pipelines in Jenkins — configure, maintain, and troubleshoot.
Provision and manage test environments in Kubernetes (GKE / EKS / AKS) across GCP, AWS, and Azure using Docker.
Investigate failures across the stack — query engine, distributed services, drivers, infrastructure — and drive them to resolution with engineering.
Collaborate with US-based developers on testability and quality gates.

Required Qualifications:
Education: B.S. or M.S. in Computer Science, Computer Engineering, or a related technical field.
Programming: Strong proficiency in Python, including pytest, with a solid grasp of OOP and software design principles.
SQL & Data: Strong SQL skills and solid understanding of how relational and analytical data systems work.
Data-intensive testing experience: Hands-on experience testing data-intensive systems — query engines, ETL/ELT pipelines, streaming platforms, or analytical databases. Candidates with only web/UI QA backgrounds are not a fit.
Testing experience: 3+ years in backend/system test automation or SDET roles.
CI/CD & DevOps basics: Practical experience with Jenkins pipelines and deploying/maintaining test environments.
Containers: Working knowledge of Docker and basic Kubernetes (running workloads, debugging pods, kubectl fluency).
Cloud: Comfortable operating in at least one major cloud (GCP, AWS, or Azure).
Version Control: Confident with Git / GitHub workflows.
English: Upper-Intermediate or higher (B2+) — daily written and verbal communication with a US-based engineering team.
Availability: Able to work EU hours with a 2–3 hour shift toward US West Coast time to ensure daily overlap with the client team.

Desired Skills:
Experience testing REST APIs and backend microservices at scale.
Familiarity with distributed computing frameworks (e.g., Apache Spark, Kafka) and MPP SQL query engines (e.g., Presto, Trino, or similar).
Understanding of modern data lakehouse concepts, open table formats (Apache Iceberg), and data warehousing.
Experience with data connectivity drivers: JDBC, ODBC, Arrow Flight.
Performance testing with JMeter or comparable load-testing tools.
Kubernetes on managed services (GKE / EKS / AKS) and multi-cloud exposure.
IaC tools such as Terraform.
Understanding of query plan generation, query acceleration / materializations, and metadata integrity in distributed data systems.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available