Databricks Data Engineer
Summary
Builds and maintains a Databricks-based lakehouse platform for a multi-tenant SaaS environment, writing PySpark pipelines, modeling data, and optimizing Delta Lake tables.
- Quickly integrates with the engineering team and contributes meaningfully to the data platform build
- Takes ownership of assigned pipeline and infrastructure work end-to-end, from design through production
- Brings architectural recommendations and solutions proactively, rather than waiting for direction
- Demonstrates strong collaboration and communication across engineering and product teams
- You have 5+ years of deep, hands-on experience building production lakehouses on Databricks. You write clean PySpark and Python, model data thoughtfully, and know how to build for a multi-tenant SaaS environment.
- Deep production experience across the Databricks platform including Unity Catalog, Delta Live Tables, Databricks SQL, and Workflows
- Delta Lake as a production table format — ACID transactions, schema evolution, performance optimization, and multi-tenant governance via Unity Catalog Experience building and maintaining dbt transformation projects using the Databricks adapter in a production environment
- PySpark for large-scale data transformation and batch pipeline authoring
- Strong understanding of batch ingestion pipeline design — migrating from relational sources like MySQL and PostgreSQL into a lakehouse architecture
- Experience with a modern pipeline orchestrator such as Dagster, Prefect, or Databricks Workflows; Dagster experience is a strong positive
- Familiarity with vector databases, embedding pipelines, and RAG patterns for AI workloads — using tools such as Databricks Vector Search, pgvector, or Amazon OpenSearch
- Exposure to AI agent and LLM-serving infrastructure including Amazon Bedrock, AgentCore, and Strands
- Experience with data cataloging and governance tools such as Unity Catalog or OpenMetadata
- Data modeling for multi-tenant analytical workloads — partitioning strategy, schema design, and tenant isolation patterns
- Databricks on AWS — workspace configuration, S3 integration, IAM, and cost governance
- Infrastructure as code using Databricks Asset Bundles or Terraform
- Strong Python and SQL skills
- Major Medical Expense Insurance
- Life Insurance
- Dental and Vision Insurance
- Mental Health Support
- IMSS (Mexican Social Security)
- Seniority Bonus
- Savings Fund Program
- Career Development Plan
- Christmas Bonus (Aguinaldo)
- Vacation Bonus
- Corporate Retirement Plan
- Certifications and Training Programs
- Internal Events
- TotalPass Wellness Program
- Additional Paid Time Off
- Auto and Motorcycle Insurance
- Pet Insurance
- Personal Belongings Insurance