Data Engineer
Summary
RefRelay is hiring a Data Engineer (5+ years) to build and maintain Databricks DLT pipelines for real-time data, with strong hands-on work in PySpark, Python, and SQL. The role also covers unit/integration testing, Git-based code reviews, Bamboo CI/CD deployments, and JIRA-managed Dev→QA→Prod release flows.
We are looking for an experienced Data Engineer (5+ years) with strong hands-on expertise in Databricks, DLT, PySpark, Python, SQL, and CI/CD for our client.
Key Responsibilities & Requirements
- Databricks DLT: Hands-on experience in creating, executing, troubleshooting, and maintaining DLT pipelines, particularly for real-time data — Mandatory.
- Unit Testing: Hands-on experience with PyTest, mocking, unit testing, and integration testing — Mandatory.
- PySpark: Strong experience with DataFrames, joins, transformations, and aggregations.
- Python & SQL: Strong hands-on experience in both Python and SQL — Mandatory.
- Git: Experience with branching, commit, push, pull, and merge.
- Bamboo / CI-CD: Experience with build and deployment processes.
- Repository / Code Review: Experience with PR creation, code review, and code push.
- JIRA: Experience with ticket creation, bug tracking, and sprint workflows.
- Deployment & Debugging: Experience managing Dev → QA → Prod deployment flows and troubleshooting issues across environments.