Senior Data Engineer(Databricks)
Summary
Senior data engineer owning production Databricks pipelines end-to-end: Spark SQL/PySpark, Delta Lake medallion architecture, DLT, and streaming ingestion from sources like Kafka or Kinesis. Also integrates third-party REST APIs, does entity resolution on messy text data, and builds metrics/semantic-layer models, working async with a US team from an on-site Lahore office.
Location: Near FC College Lahore (On-site)
Must-have experience
- 5+ years in data engineering with substantial production Databricks experience: Spark SQL, PySpark, Delta Lake, medallion architecture, and DLT.
- Structured Streaming or equivalent streaming ingestion in production (Event Hubs, Kafka, or Kinesis sources), including checkpoint recovery, watermarking, and dedup strategies.
- Demonstrated experience debugging source-vs-warehouse discrepancies — you can walk through a real incident where counts didn't match and explain how you found the cause.
- Production integration with third-party REST APIs, including pagination edge cases, rate/row limits, retries, and schema-drift handling.
- Entity resolution or data-matching work on messy real-world text data.
- Experience with a metrics/semantic layer (Holistics AML/AQL, dbt metrics, or LookML) and a working understanding of why non-additive measures can't be computed from pre-aggregated rollups.
- Strong SQL and Python; comfortable owning pipelines end-to-end without heavy oversight.
- Written and spoken English strong enough for async collaboration with a US team.
Nice to have
- Healthcare data experience: referrals, payer taxonomy, claims/eligibility, or other PHI-adjacent datasets; familiarity with HIPAA handling expectations.
- Voice-agent, call-center, or telephony/conversation data (call transcripts, containment/outcome metrics).
- Holistics specifically (AML/AQL modeling).