Sr Data Engineer
This role sits at the intersection of engineering, analytics, and platform operations. You will write production-grade Python and SQL, enforce data contracts, and help Protective treat data as a product with named owners and real consumers. You will be embedded on a delivery pod that owns its data products end to end, rather than servicing tickets from a queue.
On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description.
Key Responsibilities
• Build Bronze-layer ingestion that reliably captures data from APIs, relational databases, flat files, cloud storage, and SaaS platforms — using dlt (dltHub) and Databricks-native ingestion where each fits — including incremental loading, pagination, watermarking, state management, and replay after failure.
• Develop Silver-layer transformations in dbt and Python over Delta Lake that cleanse, standardize, type, deduplicate, validate, conform, and enrich data so that it is reusable across domains. A meaningful share of this role is making messy source data trustworthy.
• Create Gold-layer data products: dimensional models, slowly changing dimensions, fact and bridge tables, aggregates, and serving tables aligned to how consumers actually query.
• Produce and maintain the curated datasets ML engineering trains and serves models from — feature and training tables that are versioned and reproducible, not one-off extracts.
• Author and maintain data contracts using the Open Data Contract Standard (ODCS) — schema with real semantics, named owner, known consumers, quality rules, and freshness expectations — and assess backward compatibility before every change.
• Implement data quality as code: uniqueness and not-null on keys at minimum, plus referential, accepted-value, freshness, and custom business-rule tests, surfaced to producers and consumers rather than buried in logs.
• Orchestrate ingestion and transformation as assets in Dagster, deployed to Dagster Cloud, and operate what you build across development, branch, and production deployments — schedules and sensors, asset dependencies, backfills, and run observability.
• Apply governance through Unity Catalog — catalogs, schemas, external locations, grants, row- and column-level security, and lineage — and handle credentials through Azure Key Vault rather than in code.
• Implement incremental and merge-based processing with Delta Lake (MERGE, schema evolution, time travel, OPTIMIZE) and tune Spark jobs, table layouts, and compute for performance and cost.
• Troubleshoot production failures, data-quality issues, source-system changes, and late-arriving or duplicate data — including backfills and recovery — and take part in the pod’s on-call rotation for the pipelines it owns, with root-cause analysis that closes the gap rather than reopening the ticket.
• Build and maintain CI/CD for data assets in Azure DevOps — automated tests and CI checks on dlt, dbt, and Dagster changes, promotion from development through branch deployments to production, and releases that are repeatable and auditable.
• Instrument what you own for observability: freshness, volume, quality, latency, and cost, with alerting tied to the SLAs and SLOs your contract commits to instead of depending on someone noticing.
• Work inside the platform’s control expectations — least-privilege access, secrets in Azure Key Vault, change management through pull request and pipeline, and audit evidence that falls out of the deployment path rather than being reconstructed later.
• Participate in code review and document architecture, runbooks, and data products so others can discover, trust, and reuse them.
• Work with data architects, analysts, product owners, and business stakeholders to translate requirements into maintainable data solutions.
Qualifications
• Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered.
• 3+ years building and supporting production data pipelines in a cloud data platform environment.
• Strong hands-on Python and SQL. Both are used daily and neither substitutes for the other.
• Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake tables, MERGE, and incremental load patterns.
• Practical understanding of medallion / multi-layer lakehouse design, and the judgment to say what belongs in Bronze versus Silver versus Gold.
• Experience ingesting data from APIs, relational databases, files, or SaaS applications, including the incremental and state-management problems that come with it.
• Working knowledge of dimensional modeling — grain, keys, facts and dimensions, slowly changing dimensions — and of ELT design patterns and data quality practice.
• Experience with orchestration and scheduling using Dagster, Databricks Workflows, Airflow, Azure Data Factory, or similar.
• Git-based source control, pull request review, automated testing, and CI/CD as normal practice — Azure DevOps or comparable.
• Experience troubleshooting production data failures, performance bottlenecks, and source-system changes.
• Experience with pipeline monitoring and alerting, and a working understanding of what a freshness or quality SLA means once real consumers depend on it.
• Ability to explain technical designs and trade-offs to both technical and non-technical partners.
• Databricks certification (Data Engineer Associate or Professional) or equivalent demonstrated depth.
• Unity Catalog experience: catalogs, schemas, volumes, external locations, storage credentials, permissions, and lineage.
• dbt on Databricks, or another transformation framework used alongside Spark.
• Python-based modeling frameworks over Delta Lake, and experience implementing Type 2 history, surrogate keys, and merge strategies in code.
• Experience with a declarative Python ingestion framework such as dlt (dltHub), Airbyte, Meltano, or Fivetran.
• Dagster experience specifically, including assets, asset checks, sensors, schedules, and branch deployments.
Skills
As published by lever · 9 questions · 3 written answers
Basics
Which location are you applying for?, Resume/CV, Full name, Email, Phone, Current location, Current company, LinkedIn URL, GitHub URL, Portfolio/video resume URL, Other website
Short answers (1)
- What is your target compensation (salary and bonus)?
Pick from a list (5)
- As part of any recruiting and hiring process, Protective collects and processes personal information relating to job applicants. The Company is committed to being transparent about how it collects and uses that information and to meeting its information protection obligations. In the recruitment and hiring process the “personal information” we collect includes: General Identifying and Contact Information (e.g., name, mailing address, phone number, email address), Employment History Information, Education History Information. Please be aware that the personal information you provide during the application process may be shared internally with employees of Protective as well as our designated third-party vendors for the purposes of the recruitment and hiring process, including contacting you regarding your application and assessing your qualifications for job openings with Protective. California applicants should review the state-specific notice provided in accordance with the California Privacy Rights Act. By clicking “Consent” below you understand and freely give your consent to Protective and our designated third-party vendors collecting and processing your personal information relating to your potential employment with Protective. Applicant Notice - California: As part of Protective’s commitment to transparency related to personal information, and in order to comply with the California Privacy Rights Act (“CPRA”), we are providing you with this disclosure regarding the categories of personal information we may collect from you and the purposes for which that information may be used. Information We Collect: In the recruitment and hiring process the categories of “personal information” we collect include: General Identifying and Contact Information (e.g., name, mailing address, phone number, email address), Employment History Information, Education History Information. Use of Personal Information: The personal information we collect during the application process is used to assess your qualifications for the purposes of the recruitment and hiring process, including contacting your regarding your application and assessing your qualifications for job openings with Protective. The personal information you provide will not be “sold to” or “shared with” a third-party as those terms are defined by the CPRA. The Company will not collect additional categories of personal information or use the personal information we collect for materially different, unrelated, or incompatible purposes without providing you notice. Retention of Information: Each of the categories of information you provide are subject to our internal policies on retaining employee and applicant information. These policies are based on applicable legal retention requirements for retaining personal information of applicants and/or employees. As part of our internal policies, we have a process in place to determine when this information is no longer needed and can be disposed of in a secure manner. Additional Information: For more information on your rights under the CPRA, please see the Company’s California Privacy Rights Act Policy. optional
- Are you legally authorized to work in the United States?
- Do you now, or will you in the future, require immigration sponsorship for work authorization (e.g., H-1B)?
- Have you ever been employed with Protective or one of its affiliates?
- How did you hear about this opportunity? optional
Written answers (3)
- Do you have any family members who are employed with Protective?
- How have you handled issues with schema drift in the past (specific example)?
- Explain the relationships between a data contract, the generated schema.yml, and the dbt SQL with the actual column casts.