Senior Infrastructure Engineer
Summary
Design and build the distributed systems behind EVE Online's single-shard MMO simulation in Reykjavík — partitioning load across nodes, keeping state consistent under partial failure, and tuning millisecond-latency messaging. Expect hands-on work with systems languages (Go, C++, Rust, Java, Python), AWS/Azure, Kubernetes, and observability tooling like OpenTelemetry, Grafana, and Prometheus.
- System Design: Design and build the distributed systems that partition and rebalance simulation load across nodes.
- Consistency and Failure: Keep state consistent under partial failure, and shape the failure semantics so a bad node degrades one region instead of taking down a fight.
- Performance: Design messaging and RPC paths where latency budgets are measured in milliseconds.
- Services and Data: Work on service decomposition, service interfaces, data replication, and caching strategies.
- Implementation and Delivery: Write code, take a design from design review through to production, and manage priorities, deadlines, and deliverables.
- Measurement and Iteration: Instrument your work, load test it, and use that evidence to shape the next iteration.
- Collaboration: Work with product owners and game designers to understand what a feature demands of the platform, then design systems that meet those demands within real constraints.
- 5+ years in professional software development, with real depth in distributed or concurrent systems.
- Strong in at least one systems or backend language (Go, C++, Rust, Java, or Python), and the judgment to reach for a different one when the problem calls for it.
- A track record designing distributed systems: service and API boundaries, consistency and coordination models, replication, partitioning, and graceful degradation when a node fails.
- Solid fundamentals: complexity analysis, concurrency and synchronization, queueing behaviour, and a real feel for the CAP and latency tradeoffs you're making.
- Comfortable with distributed communication patterns and the tools behind them: RPC, pub/sub, message queues, event streaming.
- Experience with data stores and how they scale, relational and nonrelational, plus the caching and replication strategies built on top.
- You run what you build: operating services in production on cloud (AWS or Azure) and on Kubernetes.
- You use telemetry to reason about live behaviour (tracing, metrics, structured logging with tools like OpenTelemetry, Grafana, or Prometheus) and you load and failure test a design rather than assume it holds.
- A bachelor's in computer science, computer engineering, or a related technical field. Equivalent practical experience is just as welcome.
- Infrastructure as code and automated delivery (Terraform, Packer, Ansible, or similar) as the way your systems ship and change safely.
- Networking and systems fundamentals (UNIX, TCP/IP, load balancing, CDNs) good enough to debug your own systems down to the wire.
- The challenge of working on ambitious projects with amazingly smart and creative co-workers.
- A multicultural work environment that encourages growth, creativity and innovation.
- Double work station setup and flexible work environment.
- An active fun division that hosts regular events.
- An excellent canteen that offers a weekly breakfast and lunch menu as well as drinks and snacks.
- Discretionary quarterly and annual performance sharing plan.
- Annual sports grant.
- A mobile phone as well as a mobile usage package.
- Home internet.
- A conditional monthly transportation grant.
- Work environment that focuses on employee well-being.
- On-site doctor, free of charge as well as other on-site services at a discounted price.
- Relocation Package.
