Point your AI agent at freehire and let it find you a job.

Get the CLI →

Electronic Arts

NewBe an early applicant

Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer

Posted Updated
Discussion

« Pour visualiser la description de poste en français, veuillez sélectionner le français dans le menu déroulant au haut de la page. »

At EA, we believe games are powerful because they bring together multiple ways people engage: play, watch, create, and connect. And increasingly, the biggest entertainment platforms aren't just places to consume content — they're places where communities build.

Creator-made content is already a proven part of EA's history — from community creation tools in Battlefield to The Gallery in The Sims 4. We believe new creative technologies and tools will expand how players engage with and contribute to our experiences, supported by thoughtful product design, safety systems, and global reach. Our focus is on enabling more players to participate in creative expression by making creation easier, safer, and more rewarding.

As Lead Infrastructure Engineer, you will own the GPU fleet our researchers train on including capacity, scheduling, diagnostics, and support. You will set technical direction for GPU operations and infrastructure architecture. You will additionally lead Infrastructure as Code setup, granting permissions, and debugging infrastructure problems.

This is a hybrid role, working three days per week in Redwood City, Montreal, or Vancouver.

You will report to the Head of Data and Infrastructure.

Responsibilities:

  • You will own GPU fleet operations across our AWS estate.

  • You will build the scheduling layer from zero.

  • You will diagnose GPU and node failures fast and completely and drive hardware evidence and replacement through AWS support and capacity-block channels.

  • You will run researcher support as a first-class product including holding office hours, owning the support channel, and driving the recurring causes out of existence with self-service tooling, preflight checks, and documentation

  • You will instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics.

  • You will partner with our external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation.

  • You will author runbooks, decision records, and onboarding docs.

Qualifications:

  • 8+ years of experience operating production infrastructure, with deep, current, hands-on AWS depth — EC2 GPU fleets, EKS, IAM and cross-account security, VPC and networking, S3 and FSx

  • Experience scheduling, diagnosing, and managing GPUs in AWS specifically

  • Experience operating GPU fleets at 1000+ GPU scale

  • Expertise in scripting and automation with Python, PowerShell, bash, or equivalent

  • Expertise in infrastructure as code (Terraform or equivalent)

  • Familiarity with a GPU scheduling or orchestration layer (like Slurm, Kubernetes with Kueue or Volcano, Ray, dStack or SkyPilot)

  • Observability practice including Grafana, Prometheus, or equivalent

Skills

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available