Lead Software Engineer - DataBricks, Spark, Terraform
As a Lead Software Engineer at JPMorganChase within the Corporate - Employee Platforms, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.
We use AI-assisted development as part of our day-to-day workflow, including GitHub Copilot for coding and native Databricks tools such as Databricks SQL Assistant and Genie to accelerate development, troubleshooting, and self-service analytics—while maintaining strong engineering controls and review practices.
Job responsibilities
- Platform leadership & architecture — Define and drive the technical roadmap for our Databricks lakehouse/database platform (ingestion, storage, modeling, serving) with clear standards and reference patterns.
- Data modeling & database engineering — Design curated datasets (e.g., medallion architecture) using Delta Lake, dimensional/semantic modeling where appropriate, and enforce consistent naming, partitioning, and performance practices.
- Reliability & operations — Build for availability and predictable performance; establish SLOs, runbooks, alerting, incident response, and operational hygiene for pipelines and SQL workloads.
- Security, governance & access controls — Implement and maintain strong governance (e.g., Unity Catalog), least-privilege access, auditing, data classification, and lifecycle management.
- Performance & cost management — Tune Spark/SQL workloads, optimize clusters/warehouses, manage caching and storage patterns, and implement cost observability/chargeback as needed.
- Engineering excellence — Set standards for code quality, testing, CI/CD, branching strategy, documentation, and review. Establish reusable libraries/templates and enforce consistency across teams.
- Mentorship & collaboration — Coach engineers, lead design reviews, and partner with stakeholders to translate business needs into scalable data platform capabilities.
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
- 7+ years in data engineering/platform engineering, with hands-on Databricks experience in production environments.
- Strong proficiency in Spark (PySpark/Scala) and SQL, including performance tuning and troubleshooting.
- Proven experience building and operating data platforms: batch/stream ingestion, transformation frameworks, orchestration, and curated data layers.
- Experience with Data Lake, (DataBricks, SnowFlake, or AWS) table design, and optimization (partitioning, Z-ORDER, file sizing).
- Familiarity with Unity Catalog (or equivalent governance tooling): permissions, catalogs/schemas, lineage/auditing concepts.
- Solid software engineering fundamentals: Git, CI/CD, automated testing, code reviews, modular design, and documentation.
- Experience implementing observability (logs/metrics/traces), data quality checks, and monitoring for pipelines and SQL workloads.
- Strong communication skills and demonstrated ability to lead technical decisions across teams.
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
- Databricks features: Workflows, Delta Live Tables (DLT), Structured Streaming, Databricks SQL Warehouses.
- Transformation frameworks (e.g., dbt) and semantic layer patterns.
- Infrastructure-as-code (e.g., Terraform) and automated environment provisioning.
- Experience with regulated-data environments, privacy controls, and enterprise data governance programs.