Python/spark/ai developer
Our client is looking for a Python/Spark/AI Developer to replat form legacy T-SQL onto Spark/Delta Lake pipelines and build the APIs and AI features around them, writing type-safe, spec-first Python.
Key roles & responsibilities: Build Spark/Py Spark pipelines on Delta Lake, replat forming T-SQL into Spark SQL. Write type-hinted, tested Python with pytest suites and CI lint/type-check gates. Build REST APIs (Fast API) for pipeline jobs, with auth, idempotency, and status semantics. Validate migrated pipelines against the legacy system with parity evidence. Work spec-first against design docs/ADRs, documenting changes as you go. Must-have technical skills / experience: Python (3.12) as a primary language — modern idiomatic Python: type-hinted code that passes strict static analysis (pyright/mypy), Pydantic models, ABC-based provider patterns, packaging with modern tooling (uv or Poetry), Click or similar CLI frameworks. Apache Spark / Py Spark — production experience building data pipelines on OSS Spark (not only a managed vendor platform): Data Frame API and Spark SQL, partitioning and performance tuning, understanding of driver/executor architecture and Spark Connect. Delta Lake or an equivalent Lakehouse table format (Iceberg/Hudi) — MERGE INTO, schema evolution, time travel, idempotent write patterns. SQL — strong, dialect-portable — able to read legacy T-SQL (stored-procedure-era logic) and re-express its semantics faithfully in Spark SQL; comfortable reasoning about hashing, surrogate keys, and deduplication logic in set-based terms. Automated testing with pytest — fixtures, markers, tiered suites (unit/integration/e2e); test-first habits and comfort being held to parity/regression evidence. REST API development — Fast API or equivalent: request validation, auth (HMAC or similar signed-request schemes), idempotency, job-status semantics. Docker-based development — working daily against a Compose stack (Spark cluster, object storage, metastore, databases). Git + CI discipline — Git Hub flow, PR-driven work with lint (ruff), type-check, and test gates on every change. Preferred / nice-to-have technical skills: LLM integration engineering — building provider-neutral AI features behind an abstraction: prompt construction, structured output validation, local/self-hosted inference (Ollama, v LLM, llama.cpp, LM Studio) as well as hosted APIs; evaluation and guard-railing of model output used in data workflows. Data governance / privacy engineering — PII tokenization and hashing schemes, k-anonymity concepts, re-identification risk, POPIA/GDPR-adjacent data handling; multi-tenant isolation awareness. Lakehouse platform components — Hive Metastore, Trino, Apache Ranger; how catalogs, views, and row/column policies compose into a governed query surface. Azure data estate familiarity — ADLS Gen2, Synapse/ADF concepts (the legacy being replaced), Azure Key Vault, Service Bus. Observability instrumentation — structlog/structured logging, Open Telemetry metrics and traces, Open Lineage. Migration/parity experience — replat forming pipelines with byte/cell-level output comparison against a legacy system. Certifications — Databricks Certified Developer for Apache Spark or equivalent Spark credential. Seniority and experience: – Intermediate-to-senior. 4+ years professional Python development, with 2+ years building Spark (or comparable distributed) data pipelines in production. Must work autonomously against written design/requirements docs and ADRs — the codebase is heavily spec-driven, and changes are expected to arrive with tests and documentation, not just code. Required qualifications: Bachelor’s degree in Computer Science, Engineering, or equivalent experience. Spark credential advantageous; proven record shipping production Python/Spark pipelines and working independently in a hybrid/remote team. Why you’ll love working the companyThey believe in taking care of their team and creating an environment where you can thrive. As part of the company, you’ll enjoy:
Flexible Working Arrangements : Whether you are a night owl or an early bird, they offer hybrid and remote options to suit your lifestyle Comprehensive Benefits : From a wellness program to home office reimbursements and continuous learning opportunities, they have got you covered. Team Culture: Fun team-building activities, regular socials, and a supportive, inclusive culture that values transparency, accountability, and work-life balance. Performance Incentives : Competitive salaries, ESOP, and recognition for your hard work.