Data Engineer

Summary

Hands-on Data Engineer in Chennai building and maintaining automated data pipelines for data governance and analytics initiatives — ingestion, transformation, data quality checks, lineage capture, cost/usage metrics, and metadata integration — primarily using Azure Databricks, Azure Data Lake, PySpark, SQL, Delta Lake, and Microsoft Purview, working a US Eastern-time shift.

Job title : Data Engineer
Experience: 6+ years
Location: Chennai


Shift: US Eastern Time ( 5:00 PM – 2:00 AM )







Requirements

We are seeking a hands-on Data Engineer to develop, optimize, and maintain automated
data pipelines supporting data governance and analytics initiatives. This role will focus on
building production-ready workflows for ingestion, transformation, quality checks, lineage
capture, access auditing, cost usage analysis, retention tracking, and metadata integration,
primarily using Azure Databricks, Azure Data Lake, and Microsoft Purview.
Experience: 6+ years in data engineering, with strong Azure and Databricks experience
Pipeline Development – Design, build, and deploy robust ETL/ELT pipelines in Databricks
(PySpark, SQL, Delta Lake) to ingest, transform, and curate governance and operational
metadata from multiple sources landed in Databricks.
Granular Data Quality Capture – Implement profiling logic to capture issue-level metadata
(source table, column, timestamp, severity, rule type) to support drill-down from dashboards
into specific records and enable targeted remediation.
Governance Metrics Automation – Develop data pipelines to generate metrics for
dashboards covering data quality, lineage, job monitoring, access & permissions, query cost,
usage & consumption, retention & lifecycle, policy enforcement, sensitive data mapping, and
governance KPIs.
Microsoft Purview Integration – Automate asset onboarding, metadata enrichment,
classification tagging, and lineage extraction for integration into governance reporting.
Data Retention & Policy Enforcement – Implement logic for retention tracking and policy
compliance monitoring (masking, RLS, exceptions).
Job & Query Monitoring – Build pipelines to track job performance, SLA adherence, and
query costs for cost and performance optimization.
Metadata Storage & Optimization – Maintain curated Delta tables for governance metrics,
structured for efficient dashboard consumption.
Testing & Troubleshooting – Monitor pipeline execution, optimize performance, and resolve
issues quickly.
Collaboration – Work closely with the lead engineer, QA, and reporting teams to validate
metrics and resolve data quality issues.
Security & Compliance – Ensure all pipelines meet organizational governance, privacy, and
security standards.



See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available