Senior SRE / DevOps Engineer (Kubernetes / A2P Messaging)
This environment will suit you if you enjoy building infrastructure from the ground up, solving hard reliability problems and staying close to the systems you operate in production.
You’ll build and operate the infrastructure behind a new high-volume messaging platform, taking ownership from the Kubernetes cluster and networking layer through observability, deployment and incident response.
The platform processes around 1 million messages every day, supports approximately 175 customers and 120 suppliers, and maintains more than 300 long-lived messaging connections.
This is not a typical stateless web environment. The platform runs long-lived TCP sessions that need stable network endpoints and predictable behaviour during deployments, failures and traffic spikes.
You’ll help build the platform across two infrastructure sites and stay with it through production readiness, migration, hypercare and steady-state operation.
- Experience with telecom, carrier, messaging or other environments built around persistent network connections.
- Knowledge of SMPP, SIP, SS7 or similar telecom protocols.
- Strong production experience operating Kubernetes on self-managed, on-premise or similarly infrastructure-heavy environments - not only managed cloud services.
- Experience with stateful, long-lived TCP workloads on Kubernetes, including connection draining, stable ingress/egress, Layer 4 load balancing and deployment behaviour.
- Strong Linux and networking fundamentals. You are comfortable diagnosing problems involving routing, NAT, firewalls, MTU, TLS or packet-level behaviour.
- Practical troubleshooting experience with tools such as tcpdump and production network diagnostics.
- Experience designing monitoring and alerting, not only maintaining dashboards somebody else created. You understand what deserves to wake a human up, and what does not.
- Strong infrastructure-as-code and CI/CD experience for containerised systems.
- Practical PostgreSQL operations experience, including replication, failover, backup and importantly - verified restore.
- Ability to work directly with engineers from external infrastructure and network providers.
- Professional English.
- Real production on-call and incident-response experience.
