Point your AI agent at freehire and let it find you a job.

Get the CLI →

Artificial.Agency

NewBe an early applicant

Senior Infrastructure Engineer (ML Inference)

Posted Updated
Discussion

Senior Infrastructure Engineer (ML Inference)

Artificial Agency
Location:
Edmonton, Alberta (Hybrid; remote possible for exceptional candidates)

Job Description

Artificial Agency, the leading player-facing agentic AI company for games, is seeking a Senior Infrastructure Engineer to build, operate, and scale the inference services that make our AI agents responsive, reliable, and cost-effective.

Reporting to the Head of Platform Engineering, you will join the team running our multi-service, multi-cloud Kubernetes platform. Your primary focus will be GPU-backed inference across clouds and managed providers, alongside shared responsibility for the broader platform.

You will collaborate with Game Technology, Platform, Agents, ML, and evaluation engineers to turn models and prototypes into dependable production services. Model research, training pipelines, and shared evaluation infrastructure are outside your primary scope.

You will work with ex-DeepMind researchers, world-class game developers, and product builders on a new category of AI-powered entertainment.

Responsibilities

  • Build and operate GPU inference infrastructure, including scheduling, capacity planning, autoscaling, and warm-up strategies. Evaluate AWS, GCP, and managed providers against workload requirements, availability, reliability, and cost.

  • Share ownership of platform reliability, scaling, security, and spend, including Kubernetes, Terraform, Helm, ArgoCD, PostgreSQL, and S3. Partner with security and compliance colleagues on access controls and SOC 2 / ISO 27001 requirements.

  • Benchmark representative agent workloads with backend and ML engineers. Optimize latency, throughput, caching, and memory use under expected concurrency, guiding hardware selection, serving configurations, and capacity purchases.

  • Build repeatable deployment, validation, and rollback workflows for models, adapters, and serving configurations, partnering with ML and evaluation engineers on release readiness.

  • Implement health-aware routing and provider failover with backend engineers, including timeouts, retries, and overload handling. Validate fallback configurations with ML and Agents teams.

  • Help define and meet service commitments. Monitor request latency, errors, queueing, GPU health, and inference economics against agreed performance targets. Maintain actionable alerts and runbooks, participate in on-call, and address recurring incidents.

  • Write and review production code and infrastructure configuration, automate operational work, test deployment and recovery changes, and maintain technical documentation.

Requirements

  • An AI-first approach to engineering: hands-on use of AI tools, enthusiasm for making AI-driven workflows central to your work, and continued experimentation as tools improve: while owning the quality of what you deliver.

  • 5+ years building and operating production cloud infrastructure, with meaningful ownership of uptime, incidents, and cost.

  • Strong production Kubernetes experience, infrastructure-as-code with Terraform and Helm, and GitOps delivery using ArgoCD or similar tools.

  • Hands-on experience deploying, configuring, and troubleshooting GPU-backed inference workloads, including GPU scheduling, memory, and capacity constraints. Production LLM-serving experience is strongly preferred.

  • Experience load testing and diagnosing performance bottlenecks, with an understanding of request latency, time-to-first-token, throughput, batching, and KV-cache trade-offs.

  • Strong AWS and/or GCP networking and security fundamentals, including IAM and secrets management, plus production observability experience with Prometheus/Grafana, OpenTelemetry, or similar tools.

  • Proficiency in Python, sound testing and code-review practices, and the ability to collaborate across platform, backend, and ML teams in a fast-moving startup.

Experience with vLLM, TensorRT-LLM, Triton, quantized models, LoRA adapters, managed inference providers, or multi-provider routing is valuable; expertise in every tool is not required. Production caching, storage, PostgreSQL operations, compliance experience, and an interest in games are also welcome.

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available