Staff Engineer, Big Data

Summary

Designs and builds scalable data pipelines using Python, SQL, Spark, and Databricks to process batch and real-time data on Azure, while collaborating with cross-functional teams to ensure reliable, high-performance data solutions.

REQUIREMENTS:

  • Total experience: 5.5+ years.
  • Strong hands-on experience in Python and SQL programming.
  • Must-have expertise in Databricks, Apache Spark, Apache Kafka, and Terraform.
  • Strong experience building scalable batch and real-time data pipelines on Azure Databricks or similar cloud platforms.
  • Hands-on experience with ETL/ELT development, data ingestion, transformation, and orchestration.
  • Experience working with streaming technologies such as Apache Kafka or Azure Event Hub.
  • Good understanding of Delta Lake, Unity Catalog, and Medallion Architecture (Raw, Trusted, Curated).
  • Experience with Databricks Workflows, Airflow, or similar workflow orchestration tools.
  • Strong knowledge of CI/CD, Git, GitHub Actions, and Infrastructure as Code using Terraform.
  • Experience developing high-quality, scalable, and maintainable Python, SQL, and Spark code.
  • Good understanding of cloud-based data platforms, preferably Azure Databricks (GCP BigQuery or equivalent is acceptable).
  • Experience with metadata-driven data ingestion frameworks and automated pipeline development.
  • Familiarity with dbt and modern data engineering best practices.
  • Working knowledge of Kubernetes is an added advantage.
  • Exposure to Cybersecurity data domains (SIEM, EDR, Cloud Security Logs, OCSF) is desirable.
  • Experience with Cribl or log routing/observability pipelines is a plus.
  • Exposure to Machine Learning, AI, or GenAI solutions on Databricks is preferred.
  • Strong understanding of Agile development methodologies and DevOps practices.
  • Excellent analytical, problem-solving, communication, and stakeholder management skills.

RESPONSIBILITIES:

  • Design, develop, and maintain scalable batch, near real-time, and streaming data pipelines using Databricks, Apache Spark, and Python.
  • Build and operationalize end-to-end data ingestion pipelines from multiple data sources into the enterprise data lake.
  • Develop data transformation frameworks supporting Raw, Trusted, and Curated data layers.
  • Contribute to the design and implementation of scalable cloud-based data architectures.
  • Develop high-quality Python, SQL, and Spark code while participating in peer code reviews.
  • Build and optimize ETL/ELT pipelines following metadata-driven engineering standards.
  • Implement and maintain streaming data integrations using Apache Kafka and related technologies.
  • Automate deployment, configuration, and infrastructure provisioning using Terraform, GitHub Actions, and CI/CD pipelines.
  • Collaborate closely with Platform Engineering, Operations, Security, and Business teams to deliver reliable data solutions.
  • Monitor, troubleshoot, and optimize data pipelines to ensure high availability and performance.
  • Develop operational runbooks and support production incident analysis and resolution.
  • Ensure data quality, governance, and compliance across cloud data platforms.
  • Contribute to engineering best practices, coding standards, documentation, and knowledge sharing.
  • Mentor junior engineers and support continuous improvement initiatives across the data engineering team.
  • Evaluate and adopt modern cloud technologies, Databricks capabilities, and AI/ML innovations to enhance the data platform.

Bachelor’s or master’s degree in computer science, Information Technology, or a related field