Sr Data Engineer -
Summary
Build and optimize scalable data pipelines, ML workflows, and AI systems on Databricks using Spark, Delta Lake, MLflow, and Mosaic AI.
Job Summary
We are seeking a highly skilled Databricks Engineer with AI/ML experience to design, build, and optimize scalable data and machine learning platforms on Databricks. The role involves end-to-end ownership of data pipelines, ML workflows, and production AI systems.
Key Responsibilities
- Design and implement scalable ETL pipelines using Databricks & Spark
- Build Lakehouse architecture using Delta Lake
- Develop and deploy ML models using MLflow
- Implement MLOps pipelines for training, testing, and serving models
- Optimize cluster performance and reduce compute cost
- Build RAG and LLM-based solutions using Mosaic AI
- Integrate analytics with BI tools (Power BI, Tableau)
- Implement data governance using Unity Catalog
- Collaborate with Data Scientists and Business teams
- Ensure data quality, security, and compliance
Required Skills
Mandatory
- 5+ years of Databricks & Apache Spark
- Strong Python & PySpark
- Experience with Delta Lake & Lakehouse
- MLflow & MLOps experience
- Cloud platform (AWS/Azure/GCP)
- Git & CI/CD
Preferred
- Experience with LLMs & Generative AI
- RAG pipelines & Vector Databases
- Deep Learning frameworks
- Databricks certifications
- Power BI integration
Requirements
Job Summary
We are seeking a highly skilled Databricks Engineer with AI/ML experience to design, build, and optimize scalable data and machine learning platforms on Databricks. The role involves end-to-end ownership of data pipelines, ML workflows, and production AI systems.
Key Responsibilities
- Design and implement scalable ETL pipelines using Databricks & Spark
- Build Lakehouse architecture using Delta Lake
- Develop and deploy ML models using MLflow
- Implement MLOps pipelines for training, testing, and serving models
- Optimize cluster performance and reduce compute cost
- Build RAG and LLM-based solutions using Mosaic AI
- Integrate analytics with BI tools (Power BI, Tableau)
- Implement data governance using Unity Catalog
- Collaborate with Data Scientists and Business teams
- Ensure data quality, security, and compliance
Required Skills
Mandatory
- 5+ years of Databricks & Apache Spark
- Strong Python & PySpark
- Experience with Delta Lake & Lakehouse
- MLflow & MLOps experience
- Cloud platform (AWS/Azure/GCP)
- Git & CI/CD
Preferred
- Experience with LLMs & Generative AI
- RAG pipelines & Vector Databases
- Deep Learning frameworks
- Databricks certifications
- Power BI integration