Staff Engineer, Big Data
Summary
Designs and builds scalable data pipelines using Python, SQL, Spark, and Databricks to process batch and real-time data on Azure, while collaborating with cross-functional teams to ensure reliable, high-performance data solutions.
REQUIREMENTS:
- Total experience: 5.5+ years.
- Strong hands-on experience in Python and SQL programming.
- Must-have expertise in Databricks, Apache Spark, Apache Kafka, and Terraform.
- Strong experience building scalable batch and real-time data pipelines on Azure Databricks or similar cloud platforms.
- Hands-on experience with ETL/ELT development, data ingestion, transformation, and orchestration.
- Experience working with streaming technologies such as Apache Kafka or Azure Event Hub.
- Good understanding of Delta Lake, Unity Catalog, and Medallion Architecture (Raw, Trusted, Curated).
- Experience with Databricks Workflows, Airflow, or similar workflow orchestration tools.
- Strong knowledge of CI/CD, Git, GitHub Actions, and Infrastructure as Code using Terraform.
- Experience developing high-quality, scalable, and maintainable Python, SQL, and Spark code.
- Good understanding of cloud-based data platforms, preferably Azure Databricks (GCP BigQuery or equivalent is acceptable).
- Experience with metadata-driven data ingestion frameworks and automated pipeline development.
- Familiarity with dbt and modern data engineering best practices.
- Working knowledge of Kubernetes is an added advantage.
- Exposure to Cybersecurity data domains (SIEM, EDR, Cloud Security Logs, OCSF) is desirable.
- Experience with Cribl or log routing/observability pipelines is a plus.
- Exposure to Machine Learning, AI, or GenAI solutions on Databricks is preferred.
- Strong understanding of Agile development methodologies and DevOps practices.
- Excellent analytical, problem-solving, communication, and stakeholder management skills.
RESPONSIBILITIES:
- Design, develop, and maintain scalable batch, near real-time, and streaming data pipelines using Databricks, Apache Spark, and Python.
- Build and operationalize end-to-end data ingestion pipelines from multiple data sources into the enterprise data lake.
- Develop data transformation frameworks supporting Raw, Trusted, and Curated data layers.
- Contribute to the design and implementation of scalable cloud-based data architectures.
- Develop high-quality Python, SQL, and Spark code while participating in peer code reviews.
- Build and optimize ETL/ELT pipelines following metadata-driven engineering standards.
- Implement and maintain streaming data integrations using Apache Kafka and related technologies.
- Automate deployment, configuration, and infrastructure provisioning using Terraform, GitHub Actions, and CI/CD pipelines.
- Collaborate closely with Platform Engineering, Operations, Security, and Business teams to deliver reliable data solutions.
- Monitor, troubleshoot, and optimize data pipelines to ensure high availability and performance.
- Develop operational runbooks and support production incident analysis and resolution.
- Ensure data quality, governance, and compliance across cloud data platforms.
- Contribute to engineering best practices, coding standards, documentation, and knowledge sharing.
- Mentor junior engineers and support continuous improvement initiatives across the data engineering team.
- Evaluate and adopt modern cloud technologies, Databricks capabilities, and AI/ML innovations to enhance the data platform.
Bachelor’s or master’s degree in computer science, Information Technology, or a related field