Infrastructure Engineer — Performance Optimization
Summary
Optimize cloud infrastructure costs and performance by rightsizing compute, tuning JVM workloads, and automating changes across thousands of services on GKE.
Role Summary
We're looking for a strong infrastructure engineer to join the Performance Optimization Squad within the Platform mission. You'll work alongside senior engineers and a data scientist to execute on fleet-wide cost and performance optimization initiatives — building automation, implementing changes across production infrastructure, and helping the squad scale its impact across thousands of services.
This is a hands-on execution role. You'll take well-scoped optimization initiatives like rightsizing compute resources, JVM tuning, improving workload placement and drive them to completion across the fleet.
The Opportunity
The infrastructure teams operates at a massive scale; the Performance Optimization Squad is a newly formed team with a mandate to systematically reduce infrastructure cost and improve resource efficiency across the entire fleet.
You will join a small, focused team of 4 engineers, a data analyst, and an engineering manager. The work is technical, cross-cutting, and directly measured in dollars saved and efficiency gained.
Main Responsibilities
Implement cost and performance optimization initiatives across production infrastructure.
Build automation processes to improve resource utilization.
Drive execution of optimization initiatives to completion.
Collaborate with the team on technical challenges and solutions.
Monitor and report on cost savings and efficiency improvements.
Key Requirements
3–5 years of experience in infrastructure, platform, or backend engineering roles.
Solid experience with Kubernetes (ideally GKE).
Proficient in at least two of: Java, Go, Python, with a preference for strong scripting and automation skills.
Comfortable with Google Cloud Platform; compute, networking, IAM, and cost monitoring.
Familiarity with IaC tools (Terraform, Helm) and CI/CD pipelines.
Solid understanding of reliability engineering, performance tuning, and incident response.
Familiarity with cloud cost monitoring tools (GCP Billing, BigQuery cost exports) is a plus.
Background in reliability engineering or SRE is a plus; understanding of SLOs, error budgets, and safe rollout practices is preferred.
Experience with JVM-based services at scale is a plus.
Nice to Have
Go or Python proficiency.
GCP/cloud experience.
Engineering degree is advantageous.
Experience with automation, autoscaling, or workload placement.
Other Details
Workplace: Stockholm or remote within Sweden.
Start: asap - 2026-12-31 with possibility to extend.