Senior Infrastructure Engineer - AI Engineering
WHAT MAKES US EPIC?
At the core of Epic's success are talented, passionate people. Epic prides itself on creating a
collaborative, welcoming, and creative environment. Whether it's building award-winning games
or crafting engine technology that enables others to make visually stunning interactive
experiences, we're always innovating.
Being Epic means being a part of a team that continually strives to do right by our community
and users. We're constantly innovating to raise the bar of engine and game development.
AI ENGINEERING
What We Do
The AI Engineering team enables teams across Epic to harness AI seamlessly, empowering
Epic to lead the industry in our approach to building with AI. We build platforms that remove
friction from AI adoption and deployment for all of our developers. Our mission is to transform
Epic into an AI-native industry leader where we can build without friction.
What You'll Do
As a Senior Infrastructure Engineer on the AI Engineering team, you'll design, deploy, and
operate the critical infrastructure that powers AI across Epic. You'll own systems operating at
Fortnite scale, providing reliable AI model access to thousands of users, processing significant
request volumes, and enabling the next generation of AI-powered development tools and
autonomous agents.
This role requires both depth in infrastructure engineering and breadth in AI systems. You'll build
production-grade platforms that serve diverse use cases—from real-time code review to
business automation—while maintaining operational excellence through monitoring, alerting,
and self-service capabilities.
In this role, you will
● Build and operate critical AI infrastructure that enables Epic-wide access to AI
capabilities—examples include unified model gateways (Portkey, MCP Gateway), agent
platforms (like n8n), and observability systems (such as langfuse) for AI applications
● Develop foundational infrastructure components for AI systems—examples include
agent knowledgebases, code indexing systems, vector database deployments, and LLM
logs analytics pipelines
● Implement security and access control systems for AI services, including authentication,
authorization, and multi-tenancy across Epic's security model
● Build deployment automation, CI/CD pipelines, and infrastructure-as-code for AI
platforms
● Ensure operational excellence through SLIs, SLOs, monitoring, and incident response
● Partner with product engineers to translate requirements into scalable infrastructure
solutions
● Drive adoption across Epic by building self-service capabilities and documentation
What we're looking for
● B.S. in Computer Science or equivalent experience
● 7+ years designing and operating production infrastructure at scale
● Deep expertise in cloud platforms (AWS preferred), including ECS, Lambda, S3, and
networking
● Strong infrastructure-as-code experience with Terraform
● Proficiency with container orchestration and serverless architectures
● Expert-level understanding of observability tools (Grafana and OTEL stack)
● Experience with API gateways, load balancing, and service mesh architectures
● Strong knowledge of security best practices, authentication systems (SSO, OAuth), and
network security
● Ability to design for high availability, disaster recovery, and graceful degradation
● Excellent troubleshooting skills across distributed systems
● Comfortable working directly with AI platforms and understanding their operational
characteristics
Bonus qualifications
● Experience operating AI/ML infrastructure (model serving, inference optimization)
● Knowledge of vector databases (Chroma, Pinecone, Weaviate) and embedding systems
● Familiarity with agent orchestration platforms (LangChain, n8n, Temporal)
● Experience with cost optimization for high-volume API services
● Background in site reliability engineering or platform engineering
● Experience integrating with authentication providers (Okta, Active Directory)
● Understanding of Epic's internal systems (Perforce, Team City, URC)