Site Reliability Engineer
Summary
Maintains and improves the reliability of IFS Cloud services by resolving incidents, managing cloud infrastructure (Azure/GCP/AWS, AKS), and collaborating with teams to meet SLAs and KPIs.
A Site Reliability Engineer works within Cloud Operations to drive efficiency, reliability, and scalability, while also technically leading event, incident, case, and problem management, as well as service‑request fulfilment, to ensure the availability, security, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning of IFS Cloud services through close collaboration with architects, engineering & product development teams, account management, platform vendors and operations, provide ongoing support, and continuously improve operational processes.
As a Site Reliability Engineer, you will be responsible for core activities including, but not limited to:
Solid knowledge of IFS products, architecture, and industry solutions, able to independently resolve medium-complexity issues.
Contributes to knowledge management (KBAs, SOPs) and utilizes IFS support tools effectively.
Solid understanding of IFS product versions, support policies, and scope, able to handle issues independently with occasional guidance.
Actively works toward achieving team and organizational goals, contributing to customer satisfaction by following SOPs, SLAs, and KPIs.
Efficiently triages and resolves issues, escalating as necessary, and owns issues to closure through collaboration with other teams.
Communicates proactively with resolver groups and customers, providing regular updates on issue resolution.
Ensures clear, concise, and timely communication tailored to the audience, following established guidelines.
Actively participates in mentoring and coaching, providing basic guidance to junior engineers.
Collaborates to address knowledge gaps and contributes to training efforts.
Solid understanding of how applications and solutions are used, with guidance from senior team members.
Good understanding of application installation, configuration, and administration, with some independence in handling technical tasks.
Solid working knowledge of cloud platforms (Azure/GCP/AWS), capable of handling standard tasks and configurations with some supervision.
Practical knowledge of Docker, containerization, and container orchestration tools.
Familiar with Azure Kubernetes Service (AKS) and able to manage basic containerized environments independently.
Proficient in Linux/Unix and Windows Server (2016 preferred but other versions will be considered) administration.
Competent in Azure VPN/Express Route and Cloud Service Routers.
Hands-on experience with Oracle DB and MS SQL Server.
Good understanding of ITIL, ServiceNow, Jira Service Desk, and knowledge management tools.
Fluent in English and Japanese
We believe that coming together as a community, in person, is important for innovation, connection and fostering a sense of belonging. Our roles have the right balance of remote and in-office working to enable flexibility for managing your life along with ensuring a real connection with your colleagues and the broader IFS community.