Tools and Infrastructure Reliability Specialist - TGQF
Posted Updated
Summary
Ensures the reliability, stability, and performance of the tools and infrastructure used across Ubisoft Montreal's studio activities — automating deployments, improving observability, managing incidents, and coaching dev teams on best practices. Core tech spans Docker, Git, Terraform, Grafana/Prometheus/Splunk/OpenTelemetry, CI/CD, and cloud/on-prem systems.
#### Company description
Ubisoft is a world leader in the video game industry, with teams around the world creating original and memorable experiences, from Assassin’s Creed and Rainbow Six to Just Dance and much more. We are convinced that diverse perspectives allow both players and teams to thrive. If you are passionate about innovation and pushing the boundaries of entertainment, join us and help create the unknown!
#### Job description
In this role, you will ensure the proper functioning of tools and infrastructure used for the studio's various activities. Specifically, you will be responsible for their viability, stability, longevity, and performance.
As a true chameleon, thanks to your deep technical expertise and observability tactics, you will maintain systems across various domains. Furthermore, you will act as a key resource person to resolve and prevent incidents that may arise within them.
With agility, you will move from technical work to teaching to facilitate the bridge between development and operations. You will coach these teams on best practices for testing, development, validation, and automation. This will result in the delivery of stable, high-quality products in a timely manner.
**What you will do**
- Support development teams in technological choices to improve system visibility, control, and robustness.
- Automate processes as much as possible to facilitate everyone's work.
- Implement tools and work methods to facilitate the secure and controlled deployment of services + implement or improve incident management processes.
- Participate in the diagnosis and permanent resolution of anomalies and outages related to tools and infrastructure.
- Coordinate various resources to restore and ensure service level objectives.
- Design, deploy, secure, and maintain various reliable environments.
- Provide ongoing technical support + address and resolve issues proactively.
- Create and maintain deployment guides + document infrastructure implementation and technical specifications, as well as encountered issues and their solutions for knowledge sharing.
#### Qualifications
**What you bring to the team**
- An undergraduate degree in Computer Science, Computer Engineering, or equivalent
- Extensive experience in software development, system administration, and database administration (or other relevant experience)
- Experience in infrastructure automation (Cloud and on-premise)
- In-depth knowledge of programming languages, observability technologies (Grafana, Splunk, Elasticsearch, Prometheus, Opentelemetry, etc.), development tools (Docker, Git, Terraform), CI/CD processes, cloud services, and network infrastructure
- Familiarity with configuration management software (Ansible, Chef, etc.) and system administration (Linux and Windows)
- An understanding of the design, analysis, and debugging of redundant and scalable architecture, as well as code optimization and routine task automation
- The ability to work and adapt in a fast-paced environment
- Excellent interpersonal and communication skills, combined with a strong sense of collaboration
- A solution-oriented mindset, capable of analysis and synthesis
- An insatiable thirst for learning that keeps you up to date with technological advancements