Data Engineer
Summary
Build and maintain Azure-based data pipelines, troubleshoot PySpark/Python jobs, and support Power BI dashboards in a 24×7 production environment.
Position:
Data EngineerJob Description:
Key Responsibilities
Data Pipeline Support & Monitoring
- Monitor and support DataOps pipelines across Azure Data Factory, Azure Databricks, and related services.
- Identify pipeline failures, performance degradation, and data quality issues.
- Ensure SLA adherence and operational stability in a 24×7 production environment.
Incident Management & Troubleshooting
- Perform troubleshooting of failed pipelines, Databricks jobs, and Python/PySpark scripts.
- Execute resolution (job restarts, pipeline re-runs, alert analysis) and escalate when needed.
- Perform deep-dive analysis, identify root causes, and implement permanent fixes.
- Conduct and document Root Cause Analysis (RCA) for recurring and high-severity incidents.
Development & Fix Implementation
- Analyze and fix issues in Python, PySpark, SQL, and pipeline configurations.
- Improve error handling, stability, and performance of data workflows.
- Follow change management and deployment processes for production fixes.
Power BI & Data Validation
- Support Power BI datasets and dashboards, including refresh failures and data inconsistencies.
- Validate data accuracy, completeness, and freshness across pipelines and reporting layers.
- Resolve advanced data/model issues and performance concerns.
Collaboration & Continuous Improvement
- Act as escalation support (L2) and guide L1 engineers during incidents.
- Maintain runbooks, incident records, and shift handover documentation.
- Identify automation and monitoring improvements to reduce operational overhead.
Required Skills & Qualifications
- 4-year bachelor’s degree or equivalent
- 1–2 years of work experience
- Hands-on experience with Azure Data Factory, Azure Databricks, or similar data pipelines technologies.
- Strong understanding of Python programming and OOP concepts.
- Working knowledge of PySpark and data processing frameworks.
- Familiarity with Power BI dataset refreshes and data troubleshooting.
- Basic understanding of statistics and data science concepts (descriptive statistics, regression, time series).
- Strong analytical, debugging, and communication skills.
- Willingness to work in a 24×7 rotational shift model.
