Lead Software Engineer - Data Engg. - Databricks / Snowflake
Summary
Lead a team building scalable data pipelines and a control plane for enterprise identity data using Databricks, Python, PySpark, and AWS, ensuring quality, security, and observability.
We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.
As a Lead Software Engineer at JPMorgan Chase within the (IAM) Identity and Access Management Data team, you will play a crucial role in designing, developing, and maintaining scalable data processing solutions using Databricks, Python, and AWS. You will collaborate with cross-functional teams to deliver high-quality data solutions that support our business objectives.
Job responsibilities
- Execute creative, data-driven software solutions end-to-end (design, development, troubleshooting), thinking beyond routine approaches to solve complex technical problems.
- Design and build a control plane for enterprise data pipelines, standardizing pipeline definition, scheduling, deployment, governance, and run-time management (Databricks today; extensible for future engines).
- Develop self-service APIs/SDKs, templates, and configuration-driven onboarding with consistent guardrails (standards, validation, environment promotion, approvals) and centralized pipeline metadata (ownership, SLAs/SLOs, dependencies, schema/parameter/version tracking).
- Design, develop, and maintain scalable data pipelines and processing workflows using Python, PySpark, SQL, Databricks on AWS; develop fact/dimension models for analytics and reporting.
- Ensure data quality, security, lineage, and operational transparency via standardized observability (logs/metrics/traces), dashboards, alerting, runbooks, and automated remediation patterns (retries/backfills, common-failure automation).
- Lead and participate in the full SDLC (requirements, design, build, test, deploy, maintain), acting as SRE/production support for pipeline and platform services to improve stability and reliability.
- Collaborate with stakeholders to shape data management strategy and translate requirements into scalable, compliant solutions; document data flows, logic, and transformation rules for knowledge sharing.
- Mentor engineers and lead communities of practice to drive adoption of modern engineering practices and tools, fostering an inclusive, high-performing culture; utilize firm-approved AI-assisted development tools to accelerate delivery and testing
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Proven experience in data management and ETL/ELT for large-scale processing, including strong SQL, Python, and PySpark with performance tuning and query optimization.
- Hands-on experience with Databricks/Spark and cloud data lake patterns, integrating compute/workflows with AWS services (e.g., S3, ECS, SNS/SQS, Lambda).
- Proven experience building platform services/control planes (or similar orchestration/automation platforms), including API/service design, configuration-driven systems, and versioning/backward compatibility.
- Strong understanding of data quality, security-by-design, and lineage/auditability, including IAM/least privilege and secrets management principles.
- Strong production engineering mindset: observability (logs/metrics/traces), monitoring/alerting, incident response, and operational excellence for always-on services.
- Proficiency in CI/CD and release engineering (quality gates, automated testing, safe deployments/rollbacks) using firm-standard tooling (e.g., Jenkins/Jules, Spinnaker, Sonar).
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
Preferred qualifications, capabilities, and skills
- Experience with orchestration/execution frameworks (Databricks Workflows/Jobs, Airflow, Step Functions) and operational patterns such as dependency graphs (DAGs), replays, and backfills.
- Experience with data governance integrations (e.g., Unity Catalog concepts such as cataloging, permissions, and lineage hooks), where applicable.
- Infrastructure-as-Code experience (Terraform/CloudFormation) and developer-platform “golden path” enablement (internal CLIs, templates, paved roads, onboarding automation).
- Experience with FinOps/cost controls for Spark/Databricks workloads (telemetry, quotas, chargeback/showback) and data formats (Parquet, JSON, CSV, Avro, Delta Lake), Knowledge of regulatory reporting and financial data aggregation techniques; Databricks and/or AWS certifications.