freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer — Data Platforms, IS&T Ai & Data Platforms

Summary

Owns the reliability and performance of Apple’s large-scale data and ML platforms, ensuring high availability and cost efficiency while automating operations and incident response.

AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning — including generative AI — along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.

The AiDP Data Platforms team builds and operates data platforms at scale on the Cloud, helping Apple process, store, and access petabytes of data. We’re seeking an SRE to own the reliability, performance, and operability of our data and ML platforms — someone who thinks in SLOs, failure modes, and blast radius, and is passionate about keeping large-scale distributed systems fast, available, and cost-efficient.

You’ll operate and harden our big data platform — built on open source and other technologies — that powers critical applications like analytics, reporting, and AI/ML. This means driving down MTTR, automating operational toil, tuning performance and cost, running capacity planning, and root-causing production incidents before and after they happen. You’re an independent, self-directed problem-solver who communicates clearly with both engineers and non-technical partners, and you’ll work across many teams to keep the platform running at Apple’s standard.

Minimum Qualifications

  • 3+ years of experience operating and supporting critical, large-scale distributed systems in production, with scripting/programming ability in Python, Go, Java, Scala, or Bash for automation and tooling.
  • Deep understanding of reliability principles — fault tolerance, high availability, low latency, graceful degradation — and how to instrument and enforce them (SLIs/SLOs, alerting, on-call practices, incident response, postmortems).
  • Hands-on experience operating data processing ecosystems and distributed computing frameworks (Spark, Flink) and MPP query engines (Trino, StarRocks), including performance tuning and capacity management.
  • Proficiency operating Kubernetes/Helm at scale, building and maintaining CI/CD pipelines (GitHub Actions, Jenkins), managing infrastructure as code (Terraform, Pulumi), and running service-oriented architectures across multi-cloud environments.
  • Strong troubleshooting and performance analysis skills in complex production environments; fluency in Unix/Linux and command-line diagnostics, with excellent problem-solving and communication skills.

Preferred Qualifications

  • Experience contributing to open source projects or operating across multiple public cloud providers.
  • In-depth operational knowledge of specific distributed frameworks — Spark, Flink, or Kafka Streams, Trino, Iceberg — including debugging via component-specific logs, and experience managing multi-tenant Kubernetes clusters at scale.
  • Experience with workflow/pipeline orchestration tools (Airflow, dbt) and understanding of data modeling and warehousing concepts.
  • Experience debugging Kubernetes/Spark production issues via logs and metrics, with a continuous-improvement mindset for self, team, and org.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available