freehire launches on Product Hunt on 26 August.

Follow →

Databricks Engineer

Summary

Maintain and optimize Databricks clusters on AWS for a government cloud environment, automate infrastructure with Terraform, and secure CI/CD pipelines using GitLab, SonarQube, and Fortify.

Job Role Name: Databricks Engineer

Qualifications

  • 3 or more years of experience in data engineering with scalable pipelines
  • Strong experience designing data solutions including data modelling and distributed computing architectures
  • Hands-on experience with data processing jobs using PySpark, Spark SQL, and Databricks notebooks/jobs
  • Experience orchestrating data pipelines with ADF, Airflow, or similar tools
  • Experience with both real-time and batch data processing
  • Experience building pipelines on Azure, with AWS experience beneficial
  • Proficiency in SQL including window functions and performance optimization
  • Understanding of DevOps tools, Git workflows, and CI/CD pipelines
  • Familiarity with Scrum methodology and experience working in Scrum teams
  • Ability to apply Scrum practices in a practical project context
  • Strong problem-solving and collaborative mindset
  • Experience with streaming technologies such as Apache Kafka, Apache Flink, or AWS Kinesis
  • Ability to design and implement real-time data processing pipelines

Certification:
Databricks Certified Data Engineer Associate and Databricks Certified Data Engineer Professional are preferred.

Job Description

The Data Engineer will be responsible for designing, developing, and maintaining scalable and reliable data pipelines on Databricks and cloud platforms. The role requires integrating diverse data sources, ensuring high-quality data processing, and supporting analytics, reporting, and machine learning workloads. The role involves collaborating closely with analytics, product, and infrastructure teams to enhance the company’s data platform while adhering to best practices for governance, monitoring, and reliability.

What will you do?

  • Develop and maintain ETL pipelines for centralized data storage systems (e.g. Delta Lake).
  • Integrate data from databases, APIs, log files, streaming platforms, and external providers
  • Develop data transformation routines to clean, normalize, and aggregate data
  • Apply data processing techniques to handle complex or inconsistent datasets
  • Contribute to frameworks and best practices for code development and deployment
  • Implement data governance in alignment with company standards
  • Partner with analytics and product leaders to design and operationalize pipelines
  • Collaborate with infrastructure leaders to advance cloud-based data platforms
  • Explore new tools and techniques leveraging Azure, Databricks, or related platforms
  • Monitor data pipelines to detect and resolve issues promptly
  • Develop monitoring tools, alerts, and automated error-handling mechanisms
  • Analyze business requirements and identify data extraction requirements
  • Attend and refinement sessions with users
  • Develop and maintain ETL pipelines for ingestion, transformation, validation, and loading
  • Optimize performance and batch scheduling
  • Develop dashboards, reports, scorecards, and data visualizations
  • Perform SIT, data profiling and confirm data accuracy
  • Validate completeness and consistency of ETL Loads
  • Support UAT and production implementation

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available