Point your AI agent at freehire and let it find you a job.

Get the CLI →

Apple

NewBe an early applicant

Data Engineer, Apple Ads

Posted Updated
Discussion

Summary

Data engineer on Apple's Ads platform designing, building, and operating large-scale batch and streaming data pipelines that turn high-volume advertising events into reliable datasets for reporting, measurement, and analytics. Core stack: Java/Scala, Apache Spark, Kafka, SQL, and modern data lake technologies like Iceberg, S3, and Kubernetes.

At Apple, we build products and services that enrich people’s lives. Apple Ads helps customers discover relevant products and content while enabling developers, publishers, and advertisers to grow their businesses. Privacy is fundamental to how we design and build our advertising platform. The Apple Ads engineering organization operates large-scale data systems that process and transform high-volume advertising events into reliable datasets and products used for reporting, measurement, analytics, and other critical business functions. We are looking for a Software Engineer with strong software engineering fundamentals and experience building large-scale distributed data processing systems. In this role, you will design, develop, and operate production data pipelines using technologies such as Apache Spark, Kafka, Java/Scala, cloud storage, and modern data lake technologies. You will work on challenging problems involving large-scale batch and streaming data processing, data correctness, privacy, reliability, scalability, and performance. You will have opportunities to own systems end-to-end—from architecture and implementation through deployment, observability, and production support.

As a Software Engineer on the Apple Ads data engineering team, you will help build the next generation of scalable data processing and reporting infrastructure. You will design and implement distributed data pipelines that process large volumes of advertising events across batch, near-real-time, and streaming execution environments. You will work on systems where correctness, data quality, performance, reliability, and privacy are critical. The ideal candidate is a strong software engineer who also has hands-on experience with Apache Spark and large-scale data processing. You should be comfortable reasoning about distributed systems, debugging complex data pipelines, optimizing Spark workloads, designing data models, and building production-quality software around big-data processing frameworks. You will collaborate with engineers, product teams, data scientists, SREs, and other cross-functional partners to translate business and technical requirements into scalable data solutions.

Minimum Qualifications

  • 3+ years of professional software engineering or data engineering experience building production systems.
  • Strong computer science fundamentals, including data structures, algorithms, concurrency, and distributed systems concepts.
  • Strong programming skills in Java and/or Scala, with experience writing production-quality software.
  • Hands-on experience building and operating large-scale data pipelines using Apache Spark.
  • Strong understanding of Spark concepts including partitioning, shuffles, joins, caching, execution plans, memory management, and performance tuning.
  • Experience designing distributed batch and/or streaming data processing systems.
  • Experience with technologies such as Kafka, Hadoop, S3/object storage, or equivalent large-scale data infrastructure.
  • Strong SQL skills and experience working with large analytical datasets.
  • Expertise in distributed systems and data processing technologies (e.g. Spark, Kafka, Flink)
  • Understanding of data modeling, partitioning strategies, schema evolution, and efficient storage formats such as Parquet.
  • Experience building reliable production systems with appropriate testing, monitoring, alerting, and operational support.
  • Strong debugging and problem-solving skills, particularly across complex distributed systems.
  • Ability to communicate effectively and collaborate with technical and non-technical cross-functional partners.
  • Bachelor’s degree in Computer Science, Software Engineering, Computer Engineering, or a related technical field, or equivalent practical experience.

Preferred Qualifications

  • Experience designing and operating petabyte-scale data processing systems.
  • Deep expertise in Apache Spark performance tuning and optimization.
  • Experience with Apache Iceberg or similar modern data lake/lakehouse technologies.
  • Experience with both batch and real-time/streaming architectures, including Kafka and/or Flink.
  • Experience building data platforms or processing frameworks that are reused by multiple teams or pipelines.
  • Experience with AWS technologies, including S3 and Kubernetes/EKS or equivalent cloud platforms.
  • Experience running distributed workloads using Kubernetes and containerized environments.
  • Familiarity with analytical data stores and query engines such as Druid, Trino, or similar technologies.
  • Experience designing systems that support replay, backfills, reprocessing, and late-arriving data.
  • Experience implementing data-quality, reconciliation, lineage, or data-contract frameworks.
  • Understanding of privacy-preserving data processing and secure handling of large-scale datasets.
  • Experience building systems for advertising, measurement, reporting, or analytics.
  • Demonstrated ability to take ownership of complex projects and drive them from design through production.

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available