Senior Data Engineer - ODS
Summary
A senior data engineer on Grupo Santander's Data Management team designs and optimizes large-scale batch and near-real-time data pipelines for analytics, ML and business decision-making. Day to day is Apache Spark (Scala/PySpark) development on AWS (S3, Glue, EMR, Redshift, etc.), plus data quality, testing and CI/CD practices supporting platforms like Openbank.
As a Senior Data Engineer in the Data Management team, you design and optimize scalable data pipelines for analytics, ML, and decision-making. You will build robust, production-grade data solutions in large, complex environments using Spark-based technologies and cloud platforms. You’ll collaborate with Data Engineering, Data Science, ML and business teams to deliver practical data products. Your work directly supports an advanced digital and omnichannel platform used by flagship partners like Openbank. This role offers the opportunity to shape data architecture and elevate data quality at scale.
Responsabilidades- Design, develop, and optimize large-scale data pipelines with Apache Spark (Scala/Spark).
- Build batch and near-real-time processing solutions for high-volume data environments.
- Create reliable datasets for Analytics, Data Science, ML and business teams.
- Work with cloud platforms (ideally AWS) to process, transform, store and expose data.
- Implement data quality, validation, monitoring and documentation across pipelines.
- Collaborate with Data Engineering, Data Science, ML and business teams to deliver practical solutions.
- Contribute to engineering best practices around code quality, testing, CI/CD and automation.
- 5+ years in Data Engineering, Big Data or similar roles.
- Hands-on Spark experience in real projects, preferably Scala.
- Experience with cloud data platforms.
- Experience in banking, fintech or regulated environments.
- Fluent Spanish; professional English.
- Strong PySpark/Spark SQL or Spark-based processing skills.
- Strong Python or Scala coding for data engineering.
- Solid SQL and relational/analytical databases experience.
- ETL/ELT design and optimization experience.
- AWS or similar cloud ecosystem experience (S3, Glue, Athena, EMR, Redshift, IAM or Lake Formation).
- Git and collaborative software development practices.
- Data quality, validation, monitoring and performance optimization understanding.
- CI/CD or DevOps tools such as Jenkins, Sonar, Nexus, Jira, Splunk or similar.
- Strong problem-solving and autonomous work style
- Ownership mindset for reliable, maintainable solutions
- Clear communication with technical and non-technical stakeholders
- PySpark, Spark SQL or Spark-based distributed processing
- Python or Scala coding for data engineering
- SQL and relational/analytical databases