freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineer

Summary

Design and maintain scalable data pipelines using Databricks, Spark, and Delta Lake, integrating diverse data sources for analytics and ML workloads.

At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same

Key Responsibilities

  • Design, develop, and support data replication and integration solutions using HVR
  • Design, build, and maintain scalable data pipelines using Databricks (Spark, Delta Lake)
  • Develop and optimize ETL/ELT processes for structured and unstructured data
  • Work with large datasets to ensure data quality, integrity, and performance optimization
  • Implement data models and transformations for analytics and reporting
  • Collaborate with data scientists and analysts to enable advanced analytics and ML workloads
  • Integrate data from multiple sources including databases, APIs, and streaming systems
  • Optimize Spark jobs for performance tuning and cost efficiency
  • Implement data governance, security, and access controls
  • Monitor and troubleshoot data pipelines and production issues
  • Support CI/CD pipelines and DevOps best practices for data engineering workflows
  • Support data migration and modernization initiatives.
  • Ensure data quality, governance, security, and compliance standards.
  • Create operational documentation, runbooks, and support procedures.
  • Participate in production support, issue resolution, and performance tuning activities.

Technical Skills

  • Hands-on experience with data engineering or big data development
  • Strong knowledge of:
  • SQL / Oracle / SQL Server
  • Web-based applications and APIs (REST/SOAP)
  • Strong experience with:
  • Databricks Platform
  • Apache Spark (PySpark/Scala)
  • SQL & Python
  • Experience with Delta Lake and data lake architecture
  • Hands-on experience with cloud platforms (Azure, AWS, or GCP)
  • Familiarity with data orchestration tools (Azure Data Factory, Airflow, etc.)
  • Knowledge of data warehousing concepts (Star schema, Snowflake schema)
  • Experience with version control (Git) and CI/CD pipelines
  • Strong understanding of data pipeline optimization and performance tuning

What We Offer

  • Stable Employment: Permanent contract offering long-term job security.
  • Learning & Development: Access to a wide range of online training platforms and professional development resources.
  • Language Training: Weekly virtual English classes and conversation sessions with certified instructors. Online Courses for different languages.
  • Health Coverage: Comprehensive prepaid medical and dental plans.
  • Insurance Protection: Life and accident insurance for peace of mind.
  • Wellness Perks: Discounts and benefits through fitness and technology partnerships.

About Capgemini

At Capgemini Colombia, we aim to attract the best talent and are committed to creating a diverse and inclusive work environment, so there is no discrimination based on race, sex, sexual orientation, gender identity or expression, or any other characteristic of a person. All applications welcome and will be considered based on merit against the job and/or experience for the position.

See also