freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineer

Summary

Build and optimize scalable data pipelines on Databricks and Azure for AI systems, using PySpark, SQL, and Delta Lake to enable advanced analytics and insights.

Senior Data Engineer

Location: Bengaluru, India

Department: Projects & Delivery

Data Engineer

Experience: 2–4 years

Location: Bengaluru (Hybrid)

About the Company

Akaike Technologies is a fast-growing AI-first organization focused on building real-world, high-impact AI systems across industries. We work at the intersection of Generative AI, Multimodal AI, and Large-Scale ML Engineering, enabling enterprises to operationalize cutting-edge AI solutions at scale. We foster a culture of ownership, deep technical rigor, and continuous learning.


Experience Pre-Requisite: 2–4 years of hands-on experience in Data Engineering, with strong exposure to Databricks (PySpark, Spark SQL) and Azure Data Services (ADF, ADLS, Azure SQL). Candidates must demonstrate real-world experience in designing, building, deploying, and optimizing scalable data pipelines end-to-end.


Job Description: We are seeking a highly skilled Data Engineer to design, develop, and deploy scalable data solutions on Databricks and Azure Data Services. This role requires deep technical expertise in PySpark, SQL, Delta Lake, and Medallion Architecture, combined with strong problem-solving ability and collaboration skills.

The ideal candidate will take end-to-end ownership of data engineering workflows—from data ingestion and transformation to data modeling and governance—while ensuring performance, reliability, and security. You will work in a fast-paced, cloud-native environment, partnering with cross-functional teams to deliver production-grade data pipelines that enable advanced analytics and business insights.


Key Responsibilities

1. Design and Development

Design, develop, and deploy scalable data pipelines using Databricks (PySpark, Spark SQL), Azure Data Factory, and other Azure data services.

Implement ETL/ELT processes to ingest, transform, and load data from diverse sources into data lakes and data warehouses.

Optimize and tune data pipelines for performance, scalability, and cost efficiency.

2. Data Processing

Write and optimize complex SQL queries for data extraction, transformation, and analysis.

Utilize PySpark for large-scale data processing and advanced analytics.

Implement data partitioning, bucketing, and indexing strategies for efficient data retrieval.

3. Data Integration

Integrate data from multiple sources, including structured, semi-structured, and unstructured formats.

Work with APIs, streaming data, and batch processing to ensure seamless data integration.

4. Data Governance and Quality

Apply data governance practices to maintain data quality, consistency, and security.

Monitor and troubleshoot data pipelines to ensure accuracy, availability, and reliability.

5. Collaboration

Partner with data scientists, analysts, and other stakeholders to understand requirements and deliver solutions.

Work closely with DevOps teams to deploy and monitor data pipelines in production environments.

6. Documentation

Document data pipelines, workflows, and processes for knowledge sharing and future reference.

Maintain up-to-date documentation on data architecture and data models.


Must Have Technical Skills: Databricks: Hands-on experience with Databricks for data processing, analytics, and Serverless SQL Warehouse, Unity Catalog, Lakehouse, and Medallion Architecture.

Azure Data Services: Proficiency in Azure Data Factory, Azure Data Lake Storage, Azure SQL Database, and Key Vault.

SQL: Strong expertise in writing and optimizing complex SQL queries.

PySpark: Experience in using PySpark for data processing and transformation.

ETL/ELT: Solid understanding of ETL/ELT processes and tools.

Data Modeling: Knowledge of data modeling techniques and best practices.

Data Governance: Familiarity with data governance, data quality, and data security practices.


Must Have Soft Skills: Communication: Ability to clearly articulate technical ideas to both technical and non-technical audiences, especially in client-facing discussions.

Problem Solving: Strong analytical mindset with the ability to structure ambiguous problems and deliver actionable AI solutions quickly.

Ownership & Bias for Action: Self-driven, accountable, and proactive in driving outcomes end to end.

Collaboration: Empathy, active listening, and effective conflict resolution in cross-functional teams.


Relevant to Have: Experience with Azure DevOps for CI/CD pipelines.

Knowledge of Delta Live Tables (DLT) in Databricks.


Benefits and Perks: Competitive compensation and ESOPs

Opportunity to work on cutting-edge Generative AI and multimodal systems

High visibility across teams and leadership

Support for continuous learning, certifications, and conference participation


See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available