Forward Deployed Data Engineer
Summary
Senior forward-deployed data engineer owning a healthcare organization's data platform end to end on Azure Databricks: building pipelines (Salesforce, SQL Server, Snowflake), acting as technical product owner with business stakeholders, and using Claude Code and agentic AI workflows to ship fast. US residency required.
Role Overview
The data engineer role has changed. Building a pipeline, handing it off, and waiting for someone else to confirm the data is right no longer creates much value. The engineers who matter now go deep into the business domain, understand the problem directly, and use AI tools to move from question to working solution in hours rather than sprints.
We are seeking a Senior Forward Deployed Data Engineer who operates as a hybrid technical product owner. You will build and operate the data platform on Azure Databricks, and you will also sit with business stakeholders, interrogate the problem, and own the outcome end to end. You will not wait for requirements to arrive. You will go find them, validate them, and ship, using Claude Code and agentic development patterns to collapse the distance between understanding a business problem and solving it in production.
This is what it means to shift the engineer left into the business. You are not a pipeline builder waiting on a ticket. You are the person who understands how physician network growth, member engagement, and operational performance actually work, and who builds the data infrastructure that makes the entire analytics function faster and sharper.
This position supports a healthcare organization, so US residency is required and experience handling regulated data is valued.
What You Will Do
Own the Data Platform
- Design, build, and operate the data platform on Azure Databricks, covering ingestion, transformation, storage, and serving layers that power analytics, AI models, and operational reporting.
- Build and maintain pipelines across the business ecosystem, including Salesforce, SQL Server, Snowflake, third-party sources, and a new cloud-native payments platform.
- Engineer for quality and trust through validation checks, anomaly detection, lineage tracking, and documentation that every downstream consumer can rely on.
- Write clean, version-controlled, production-grade code. Think like a software engineer building a product, not a script runner maintaining jobs.
- Own architecture decisions across storage, transformation, orchestration, and serving, balancing delivery speed against long-term maintainability.
Go Deep Into the Business Domain
- Partner directly with stakeholders across physician growth, member services, finance, and operations to understand how data drives decisions, then build for those decisions rather than for abstract requirements.
- Act as technical product owner for your domain areas. Own the backlog, prioritize by business impact, and ship iteratively without waiting for a PM to sequence your work.
- Translate ambiguous business questions into data models, feature tables, and curated datasets that analysts and data scientists can build on immediately.
- Push back on vague requirements until they are sharp enough to build against.
- Close the loop. Follow your data through to the dashboard, the model, or the operational workflow and validate that it is actually driving the outcome.
- Communicate tradeoffs, limitations, and timelines clearly to technical and non-technical audiences alike.
Drive AI-First Engineering Practices
- Use Claude Code and agentic development as your primary workflow, including AI-driven pipeline generation, automated testing, and rapid prototyping, to ship at a pace traditional approaches cannot match.
- Build data infrastructure that is AI-ready: well-documented, semantically clear, and structured so that AI tools and agents can reason over it effectively.
- Support data workflows behind AI and LLM applications, including curated context, retrieval-ready datasets, and the evaluation data needed to measure whether those systems are working.
- Scout, evaluate, and adopt emerging AI tools and platforms that make the data team faster, separating real value from hype through hands-on testing.
- Share what you learn. Document patterns, run demos, and help the broader team adopt AI-first workflows with confidence.
Handle Data Responsibly
- Apply appropriate handling, access control, and compliance practices for sensitive and regulated healthcare data.
Who You Are
- A data engineer who refuses to stay in the technical silo. You go find the business problem rather than waiting for it to arrive as a ticket.
- Someone who thinks like a product owner. You prioritize by impact, ship incrementally, and own the outcome, not just the pipeline.
- The kind of engineer who was already using AI tools to write, test, and deploy code before anyone asked, and who knows when the output needs a human eye.
- Equally comfortable writing a Spark transformation, debugging a Salesforce data sync, presenting findings to leadership, and pushing back on a vague requirement.
- Pragmatic over perfectionist. You optimize for business impact and speed to value rather than theoretical elegance.
- A strong collaborator who elevates the people around you through clear communication, reusable patterns, and generous knowledge sharing.
Required Qualifications
- BS in Computer Science, Data Science, or a related field, with 6 or more years in data engineering or a hybrid data engineering and analytics role.
- Deep hands-on experience with Azure Databricks, including notebooks, Delta Lake, Unity Catalog, and production-scale pipelines.
- Strong Python and SQL, with experience in PySpark and distributed data processing.
- Experience building and operating pipelines that serve analytics, ML models, and operational systems, not only batch ETL jobs.
- Direct experience working with business stakeholders to define requirements, shape data products, and deliver measurable outcomes.
- Active, daily use of AI coding tools as a force multiplier, with Claude Code as the primary platform and familiarity across other leading models.
- Data architecture experience, including designing systems rather than only implementing against someone else's design.
- Experience building data products for marketing or growth teams, such as customer segmentation, campaign attribution, or engagement and retention analytics.
- Strong communication skills and a track record of presenting technical work to non-technical audiences.
- Proven ability to take an ambiguous business problem and turn it into a working data solution with minimal direction.
Nice to Have
- Healthcare, PHI, or other regulated-industry data experience.
- Experience with Salesforce data models and integrations.
- Experience with payments platforms or financial data pipelines.
- Experience with MCP, agent architectures, RAG, embeddings, or vector databases.
- Experience with data governance, cataloging, access control, and compliance frameworks.
- Experience with LLM evaluation, observability, prompt optimization, and AI system cost management.
- Experience with BI and visualization tooling, or building analytics for non-technical users.
- Experience in a forward-deployed, consulting, client-facing, startup, or highly autonomous engineering environment.
Ideal Candidate
The ideal candidate is a senior data engineer who understands the business, not only the platform. They move comfortably between data, AI, product, and operations, and they can hold a productive conversation with someone who has no technical background at all.
You should be able to carry a problem through the full arc:
Business Question → Discovery → Data Architecture → Platform and Pipelines → Production → Adoption → Optimization
The ability to build reliably, understand why the data matters, and work directly with stakeholders matters as much as technical depth.
Your partner for AI, consulting, software development, and nearshore staffing.
Skills
- Agentic AI
- AI
- Analytics
- Anomaly Detection
- Azure
- Claude Code
- Cloud
- Cloud Native
- Data Engineering
- Data Governance
- Data Pipelines
- Data Science
- Databricks
- Delta Lake
- Embeddings
- ETL
- LLM
- Machine Learning
- MCP
- Observability
- Prototyping
- PySpark
- Python
- Salesforce
- Snowflake
- Spark
- SQL
- SQL Server
- Test Automation
- Unity
- Vector Databases
As published by greenhouse · 7 questions · 2 written answers
Basics
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
Short answers (3)
- Preferred First Name optional
- LinkedIn Profile optional
- (Video Required) Pick a data problem you solved that a non-technical stakeholder cared about. Explain it the way you would to that stakeholder, not to an engineer. We are listening for how you make technical work land with people who do not share your background.
Pick from a list (2)
- Do you use AI coding tools such as Claude Code as part of your daily engineering workflow?
- Have you worked with PHI or other regulated data in a production environment?
Written answers (2)
- Which version of Azure Databricks work have you owned end to end? Name the components you were responsible for, for example Delta Lake, Unity Catalog, orchestration, serving.
- Describe one production pipeline you shipped using Claude Code or another agentic tool. What did the AI generate, and where did you have to correct it?