freehire launches on Product Hunt on 26 August.

Follow →

Staff Engineer, Big Data

REQUIREMENTS:

Total experience: 5.5+ years.
• Strong experience in Data Engineering, Cloud Engineering, or Big Data Engineering.
• Must-have expertise in Google BigQuery, Python, SQL, PySpark, GCP fundamentals, and Kubernetes.
• Strong hands-on experience with BigQuery and SQL, including writing and optimizing complex queries for large-scale data processing.
• Strong programming experience in Python, with hands-on experience in developing scalable applications and APIs.
• Strong experience with PySpark/Spark and familiarity with Big Data technologies such as Hive.
• Hands-on experience with Apache Airflow or similar orchestration services for building and managing data workflows.
• Experience with GCP serverless services, particularly Cloud Functions and Cloud Run.
• Good experience with Docker and Kubernetes for containerization, deployment, orchestration, scaling, and troubleshooting.
• Experience with Python API frameworks such as FastAPI, Flask, or Django; FastAPI is preferred.
• Experience with CI/CD pipelines, preferably using GitLab CI/CD and Octopus Deploy, along with Git-based version control.
• Good understanding of Terraform and Infrastructure as Code (IaC) for provisioning and managing cloud resources.
• Good understanding of GCP networking, VPC, load balancing, IAM, API security, monitoring, logging, and alerting.
• Strong troubleshooting, analytical, problem-solving, communication, and collaboration skills.

RESPONSIBILITIES:

• Design, implement, and maintain scalable and reliable Big Data and cloud solutions using GCP, BigQuery, PySpark, Python, and Airflow.
• Develop and maintain data workflows and pipelines using Apache Airflow and other orchestration services.
• Design, develop, and optimize BigQuery and SQL queries for high-volume data processing and analytics workloads.
• Develop and deploy cloud-native and serverless applications using GCP Cloud Functions and Cloud Run.
• Develop scalable APIs using Python frameworks such as FastAPI, Flask, or Django.
• Containerize applications using Docker and deploy, manage, and troubleshoot workloads on Kubernetes.
• Implement and maintain CI/CD pipelines for automated build, testing, and deployment using GitLab CI/CD, Octopus Deploy, or similar tools.
• Implement Infrastructure as Code (IaC) using Terraform to provision and manage GCP cloud resources.
• Configure and manage IAM policies, VPC networking, load balancing, API access controls, and security mechanisms.
• Implement monitoring, logging, alerting, and observability solutions to proactively identify and resolve system issues.
• Write comprehensive unit tests and implement quality practices to ensure application reliability and maintainability.
• Troubleshoot production issues, perform root cause analysis, and implement preventive measures to improve system stability.
• Collaborate with Data Engineering, Cloud, DevOps, Infrastructure, Security, and Application teams to deliver scalable and reliable solutions.
• Continuously improve data pipelines, cloud architecture, automation, application performance, security, and operational efficiency through industry best practices.

Bachelor’s or master’s degree in computer science, Information Technology, or a related field

    See also

    Tailor your CV for this role?

    We couldn't check your fit for this role — add a CV to your profile to see it next time.

    A new version of freehire is available