Data Engineer - ETL/PySpark (Banking Domain)
Summary
The Data Engineer will design, build, and maintain ETL pipelines and data marts within a banking environment using PySpark and Python. The role involves managing the full SDLC, from development and UAT to production deployment and support, while working with diverse data types.
Role Summary
We are looking for a hands-on Data Engineer with strong ETL and PySpark expertise to design, build, and support data pipelines and data marts within a banking environment. The ideal candidate will own the full SDLC lifecycle — from build through UAT, production deployment, and post-production support — while working across structured, semi-structured, and unstructured data.
Key Responsibilities
- Design, develop, and maintain ETL pipelines and data marts using PySpark and Python
- Write clean, maintainable, and production-grade Python code following software engineering best practices
- Own end-to-end SDLC activities: build, UAT support, UAT bug fixes, production deployment, and post-production support
- Perform data analysis and debugging using Oracle SQL and PySpark
- Work across structured, semi-structured, and unstructured data sources
- Build and maintain data warehousing solutions supporting banking/financial reporting needs
- Debug and optimize PySpark jobs for performance and reliability
- Collaborate with cross-functional teams (QA, DBAs, business analysts) through the release cycle
- Participate in CI/CD pipeline processes, including testing and validation of data pipelines
- Ensure data pipeline reliability, scalability, and adherence to banking data governance/compliance standards
Required Skills & Experience
- 5+ years of commercial experience in a data-driven engineering role
- Hands-on experience building data marts and ETL pipelines
- Expert-level PySpark and Python for ETL scripting
- Strong command of Oracle SQL for data analysis and debugging
- Proven experience across the full SDLC — build, UAT, bug fixing, deployment, post-prod support
- Strong understanding of software engineering concepts and best practices for production pipelines
- Experience working with structured, semi-structured, and unstructured data
- Prior experience with banking clients or strong banking domain knowledge
- Strong data warehousing fundamentals
Tech Stack (Daily Use)
- Languages: Python
- Big Data: Spark / PySpark, Hadoop, MapReduce, Hive
- Data Libraries: Pandas
- Databases: SQL and NoSQL DBMS
- Tools: Jupyter
- Practices: CI/CD, data testing & validation
Nice to Have (optional — add if applicable)
- Cloud experience (AWS/Azure/GCP) — not mentioned in your input, confirm with client
- Airflow or other orchestration tools
- Experience with regulatory/compliance reporting in banking
As published by workable
First name, Last name, Email, Headline, Phone, Address, Photo, Education, Experience, Summary, Resume, Cover letter