Point your AI agent at freehire and let it find you a job.

Get the CLI →

Galent

NewBe an early applicant

Big Data Developer

Posted 1 view
Discussion

Summary

Senior backend/big data developer who designs, builds, and maintains scalable Spark-Scala data pipelines on Hadoop/Cloudera CDP, extracts data from REST APIs with Python/curl, and automates jobs with bash, Ansible, and schedulers like Control-M — using AI assistants such as GitHub Copilot. Onsite in Toronto.

We are looking for a Senior Backend Developer with 5+ years of experience in big data engineering, API integration, and AI-assisted development. The ideal candidate will design, build, and maintain scalable data pipelines and backend systems in a enterprise environment.

Key Responsibilities

  • Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clusters
  • Build and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL
  • Tune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)
  • Migrate Spark 2 applications to Spark 3 on Cloudera CDP platforms
  • Work with Parquet, ORC, Avro file formats on HDFS
  • Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations
  • Design and maintain Hive external/managed tables and partitioned datasets
  • Optimize slow-running queries and resolve correlated subquery issues
  • Work with HDFS encryption zones and data governance requirements

Unix / Shell Scripting

  • Develop and maintain bash shell scripts for job orchestration and automation
  • Handle error management, return codes, logging and alerting in shell scripts
  • Manage HDFS operations (hdfs dfs commands), file transfers, and data validation

API Extraction & Integration

  • Build scripts and pipelines to extract data from REST APIs using curl and Python
  • Parse and process JSON API responses and load into HDFS/Hive
  • Manage pagination, error handling and retry logic for API calls
  • Work with enterprise API gateways and URL parameter construction
  • Leverage GitHub Copilot / AI coding assistants to accelerate development
  • Use AI tools for code review, SQL generation, script debugging and documentation
  • Contribute to AI-assisted data quality and anomaly detection pipelines
  • Explore and implement LLM-based automation for repetitive data engineering tasks

Scheduling & Orchestration

  • Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cron
  • Build and maintain Ansible playbooks for automated deployments
  • Manage deployment pipelines including artifact versioning, Vault secret injection and environment-specific configuration
  • Monitor job health, handle failures and implement alerting

Nice to Have

  • Experience with Cloudera CDP (7.x) and migration from HDP
  • Knowledge of Kerberos, Vault, HDFS encryption zones
  • Familiarity with CI/CD pipelines (Helios, GitHub Actions)
  • Experience with MSSQL / JDBC connectivity from Spark
  • Understanding of AML / Financial regulatory data domains.

We are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, citizenship status, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law. https://www.e-verify.gov/sites/default/files/everify/posters/IER_RighttoWorkPoster.pdf

Skills

Apply

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available