Design, develop, and maintain scalable batch and real-time data pipelines using Databricks, PySpark, and Spark SQL.
Build and implement Delta Lake-based solutions, including data ingestion, transformation, and Medallion Architecture (Bronze, Silver, Gold) layers.
Develop and orchestrate end-to-end data workflows using Databricks Workflows and other scheduling tools.
Optimize Spark workloads for performance, scalability, and cost efficiency through effective cluster management and tuning practices.
Ensure data quality, reliability, and governance through validation, monitoring, and testing frameworks.Integrate and process data from multiple enterprise sources, including databases, APIs, files, and streaming platforms.
Collaborate with cross-functional teams, including data scientists, analysts, architects, and business stakeholders, to deliver data-driven solutions.
Contribute to code management, CI/CD implementation, documentation, and continuous improvement initiatives.