Lead DataOps / MLOps
The DataOps/MLOps Lead owns the platform and operating model that lets data and ML engineers ship reliably. You build the paved paths — CI/CD, orchestration, observability, environments, and governance automation — on our Databricks Lakehouse on Microsoft Azure, so that pipelines and models move from development to governed production quickly, safely, and repeatably.
KEY RESPONSIBILITIES
• Lead CI/CD standards and pipelines in Azure DevOps (ADO) for data pipelines and ML models — build, test, and release automation, environment promotion, and repeatable, auditable deployments.
• Standardize orchestration on Dagster — reusable assets, scheduling, backfills, dependency management, and run observability across the pod's pipelines.
• Operationalize the ingestion and transformation stack — dlt (dltHub) and dbt — with automated testing, CI checks, and safe deployment of changes.
• Build MLOps foundations with the ML engineering team — MLflow model registry, Databricks Model Serving, automated deployment, monitoring, drift detection, and retraining triggers.
• Establish data and model observability — freshness, quality, lineage, latency, drift, and cost — with alerting and clear SLAs/SLOs.
• Administer and govern the Databricks Lakehouse on Azure — workspace configuration, Unity Catalog governance, access controls, and policy automation.
• Manage infrastructure as code and environments — reproducible dev/test/prod setups (e.g., Terraform), secrets management, and least-privilege access.
• Own reliability and incident practices — on-call, runbooks, root-cause analysis, and continuous improvement for data and ML services.
• Drive cost visibility and optimization (FinOps) across compute, storage, and model serving.
• Automate governance and compliance controls — audit logging, model and pipeline inventories, approval workflows, and evidence collection for a regulated environment.
• Provide technical leadership and mentoring — coach engineers on operational excellence and set the platform standards the pod builds on.
QUALIFICATIONS
• 8+ years in data, ML, or platform engineering, or in SRE/DevOps, including several years operating production data and/or ML systems.
• Demonstrated technical leadership — setting standards, building paved paths and automation, and mentoring engineers (formal people management not required, but valued).
• Strong CI/CD expertise with Azure DevOps (ADO) — build/release pipelines, environment promotion, automated testing — and Git-based workflows.
• Hands-on experience with orchestration (Dagster or equivalent) and the modern data stack — dlt (dltHub) ingestion and dbt modeling — on a Databricks lakehouse (Delta Lake).
• MLOps experience — MLflow model registry, model deployment/serving, monitoring, drift detection, and retraining automation.
• Infrastructure-as-code and cloud platform administration on Microsoft Azure (compute, storage, identity, networking basics); Terraform or equivalent.
• Strong Python and SQL for automation and tooling.
• Experience with observability/monitoring tooling and SRE practices — SLAs/SLOs, alerting, and incident management.
• Demonstrated rigor in security, access control, and secure, compliant handling of sensitive data.
• Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent practical experience.
PREFERRED QUALIFICATIONS
• Experience in financial services or insurance platform work, and familiarity with model risk and regulatory audit expectations.
• Databricks administration — Unity Catalog, cluster policies, and Mosaic AI — and familiarity with Azure Machine Learning.
• Containerization and orchestration (Docker, Kubernetes; Azure AKS or Container Apps).
• Experience with data-quality / observability tooling (e.g., Great Expectations, Monte Carlo, or similar).
• Experience automating responsible-AI and model-governance controls.
• Relevant certification such as Databricks Certified Data Engineer/ML Engineer, Microsoft Azure DevOps Engineer, or Azure Administrator.
Skills
- AI
- AKS
- Automation
- Azure
- Azure DevOps
- Backbone.js
- CI/CD
- Cloud
- Containerization
- Dagster
- Data Pipelines
- Data Quality
- Databricks
- dbt
- Delta Lake
- DevOps
- Docker
- FinOps
- Git
- Infrastructure as Code
- Kubernetes
- Lakehouse
- Machine Learning
- MLflow
- MLOps
- Model Deployment
- Networking
- Observability
- Python
- Root Cause Analysis
- Secrets Management
- SQL
- Terraform
- Test Automation
- Unity
As published by lever · 10 questions · 4 written answers
Basics
Which location are you applying for?, Resume/CV, Full name, Email, Phone, Current location, Current company, LinkedIn URL, GitHub URL, Portfolio/video resume URL, Other website
Short answers (1)
- What is your target compensation (salary and bonus)?
Pick from a list (5)
- As part of any recruiting and hiring process, Protective collects and processes personal information relating to job applicants. The Company is committed to being transparent about how it collects and uses that information and to meeting its information protection obligations. In the recruitment and hiring process the “personal information” we collect includes: General Identifying and Contact Information (e.g., name, mailing address, phone number, email address), Employment History Information, Education History Information. Please be aware that the personal information you provide during the application process may be shared internally with employees of Protective as well as our designated third-party vendors for the purposes of the recruitment and hiring process, including contacting you regarding your application and assessing your qualifications for job openings with Protective. California applicants should review the state-specific notice provided in accordance with the California Privacy Rights Act. By clicking “Consent” below you understand and freely give your consent to Protective and our designated third-party vendors collecting and processing your personal information relating to your potential employment with Protective. Applicant Notice - California: As part of Protective’s commitment to transparency related to personal information, and in order to comply with the California Privacy Rights Act (“CPRA”), we are providing you with this disclosure regarding the categories of personal information we may collect from you and the purposes for which that information may be used. Information We Collect: In the recruitment and hiring process the categories of “personal information” we collect include: General Identifying and Contact Information (e.g., name, mailing address, phone number, email address), Employment History Information, Education History Information. Use of Personal Information: The personal information we collect during the application process is used to assess your qualifications for the purposes of the recruitment and hiring process, including contacting your regarding your application and assessing your qualifications for job openings with Protective. The personal information you provide will not be “sold to” or “shared with” a third-party as those terms are defined by the CPRA. The Company will not collect additional categories of personal information or use the personal information we collect for materially different, unrelated, or incompatible purposes without providing you notice. Retention of Information: Each of the categories of information you provide are subject to our internal policies on retaining employee and applicant information. These policies are based on applicable legal retention requirements for retaining personal information of applicants and/or employees. As part of our internal policies, we have a process in place to determine when this information is no longer needed and can be disposed of in a secure manner. Additional Information: For more information on your rights under the CPRA, please see the Company’s California Privacy Rights Act Policy. optional
- Are you legally authorized to work in the United States?
- Do you now, or will you in the future, require immigration sponsorship for work authorization (e.g., H-1B)?
- Have you ever been employed with Protective or one of its affiliates?
- How did you hear about this opportunity? optional
Written answers (4)
- Do you have any family members who are employed with Protective?
- Describe a CI/CD framework, deployment standard, or engineering platform capability that you designed or significantly improved for data or machine learning teams. What problem were you solving, what standards did you establish, and how did you measure adoption and success?
- Describe the most significant production incident involving a data platform, ML platform, or shared engineering service that you helped lead. How was the issue detected, how was the response coordinated, what was the root cause, and what permanent improvements were implemented afterward?
- A delivery team wants to bypass an existing CI/CD, governance, or deployment standard because they believe it will allow them to move faster. How would you evaluate the request, balance innovation against operational risk, and decide whether to evolve the platform or require the team to follow the existing paved path?