System Engineer
Summary
Maintains and optimizes IT systems by monitoring performance, automating deployments, and managing incidents for a global payments company using Linux, cloud tools, and monitoring stacks.
The opportunity
You will be focusing on the harmonization and automatization of the tooling landscape. To move from historical grown environment with many different solutions to a simplified and better optimized ecosystem with the usage of AI.
You will be part of a team of specialists of different flavors to collect current solutions and develop a common solution.
Day-to-day responsibilities
1) Monitoring of Availability and Reliability:
- Monitor the performance, availability, and reliability of our systems and applications, ensuring adherence to SLAs.
- Implement and maintain monitoring and alerting systems to proactively identify issues.
- Release/Deploy:
- Have a holistic end-to-end view on the service: application, underlying infrastructure and other dependencies.
- Guarantee that the services we are in charge are installed, configured, changed, and operated to guarantee requirements compliance with regulation, security and the expected service levels.
- Improve through automation, the pipeline from development to operations (CI/CD) and all operations on the service. Participate in and validate service designs and changes, with the mission to ensure that quality of operations will remain at the proper level.
2) Incident Management and Alerting:
- Lead incident response efforts during outages or major incidents, coordinating with cross-functional teams to ensure timely resolution.
- Conduct post-mortem analyses to identify root causes and implement corrective actions.
- Availability for on call activities out of business hours – as Stand-by
3) Reporting and SLA management:
- Ensure that the 2nd Line of support has all information available to manage incidents and problems.
- Identify gather and analyze metrics and events from both infrastructure and application to provide capacity planning, performance improvements and incident analysis.
4) Analyze communication and ticketing tools to find a common approach.
Who are we looking for
- Good understanding tooling landscape.
- Good experience in Linux; PostgreSQL; HA/DR or common technologies.
- Cloud oriented: GCP knowledge is recommended, other provider could be a plus (i.e. AWS) Cloud management tools: Terraform, Puppet, Gitlab
- Experience with Prometheus, Grafana, ELK (Elasticsearch, Logstash, Kibana), CI/CD, Network management (debugging network issues)
- Fluent in English
Perks & Benefits
- Hybrid Working Policy
- Gift vouchers on the occasion of Christmas/Easter Holidays
- Private medical services
- 21 vacation days/year
- Referral bonuses for new hires recommended by you
- WFH & Flexible Working Hours
- Full access to the “Learning” platform