DevOps
Summary
Build and maintain AWS-based infrastructure for a next-gen payment routing platform, focusing on high availability, zero-downtime deployments, and automated CI/CD pipelines using Terraform, Ansible, and containerization.
We’re building a next-generation payment infrastructure designed for scale, speed, and resilience.
Our system dynamically routes transactions across multiple providers in real time — optimizing for performance, reliability, and approval rates.
As we enter a scaling phase, we’re focused on building robust, high-performance systems that can handle complex traffic flows and continuous growth.
This is not about maintaining infrastructure — it’s about engineering a system that powers a new standard in modern payment ecosystems.
Primary Objective of the Role
We are looking for a DevOps Engineer who can maintain and develop our production infrastructure in AWS, ensure high system availability, enable safe zero-downtime releases, automate infrastructure management, and establish effective monitoring.
Mandatory Requirements
- Hands-on experience with AWS.
- Understanding of High Availability principles and fault-tolerant infrastructure design.
- Experience implementing zero-downtime deployments.
- Hands-on experience with Infrastructure as Code, preferably Terraform.
- Experience with configuration management tools, such as Ansible or Puppet.
- Experience building and maintaining CI/CD pipelines.
- Strong understanding of Docker and application containerization.
- Experience with load balancers, health checks, and horizontal scaling.
- Understanding of network infrastructure: VPCs, subnets, routing, security groups, and firewall rules.
- Good understanding of DNS, SSL/TLS, and domain and certificate management.
- Experience with managed and self-hosted databases.
- Understanding of backup, restore, disaster recovery, and backup restoration testing.
- Experience setting up monitoring, centralized logging, and alerting.
- Experience with tools such as Prometheus, Grafana, Loki, ELK, Datadog, or CloudWatch.
- Good knowledge of Linux and confidence working with the command line.
- Experience troubleshooting production incidents.
- Understanding of secrets management and secure access control.
Expectations for the Candidate
The candidate should be able to:
- design a production architecture appropriate for the actual workload without unnecessary overengineering;
- identify single points of failure;
- explain trade-offs between cost, complexity, and reliability;
- build infrastructure that can be reproduced automatically;
- automate server and environment configuration using Ansible, Puppet, or similar tools;
- minimize manual operations and deployments via SSH;
- organize a safe rollback after a failed release;
- define critical metrics, logs, and alerts;
- propose a system scaling plan;
- independently analyze and resolve production issues.