Senior Software Engineer, Reliability
Summary
Senior engineer builds and maintains Klaviyo’s cloud infrastructure, automating ops, enforcing SLOs, and improving system resilience and observability.
Senior Software Engineer, Reliability at Klaviyo.
About the role
You will design and maintain the foundational systems that keep Klaviyo's global platform stable, performant, and scalable. This role focuses on applying software engineering to infrastructure challenges, ensuring our services remain resilient while supporting rapid product development.
Key facts
- Location: Dublin, IE
- Engagement: Full-time
What you'll do
- Develop and manage security-critical services with a focus on high availability and fault tolerance.
- Automate infrastructure tasks to minimize manual operational work.
- Establish and monitor SLIs, SLOs, and error budgets to influence engineering priorities.
- Enhance observability and alerting frameworks to improve incident response times.
- Participate in on-call rotations, prioritizing automated remediation and sustainable operations.
- Conduct capacity planning and performance analysis to identify system bottlenecks.
- Collaborate with security and product teams to guide architectural decisions.
- Mentor team members to improve overall operational maturity and engineering standards.
Requirements
- Proficiency in writing production-ready code using languages such as Python or Go.
- Experience building and operating distributed, cloud-native systems.
- Practical knowledge of managing containerized workloads in production, specifically Kubernetes.
- Ability to diagnose production issues and contribute to post-incident reviews.
- Hands-on experience with infrastructure as code and declarative configuration tools like Terraform.
- Familiarity with observability systems and building actionable, user-impact-focused alerts.
- Experience with capacity planning, load testing, and performance analysis.
- Interest in exploring and applying AI tools to improve engineering workflows.
Nice to have
- Background in supporting security-critical platforms or internal security tooling.
- Knowledge of identity, access management, secrets management, or policy enforcement.
- Experience operating systems at scale within AWS environments.
- Familiarity with chaos engineering, fault injection, or resilience testing.
- Strong understanding of algorithms and data structures.
Skills & tools
- Languages: Python, Django, FastAPI
- Infrastructure: AWS, Kubernetes, Terraform
- Data & Messaging: MySQL, Redis, Memcached, RabbitMQ, Celery, Apache Kafka, Apache Pulsar
Practical notes
- Salary range: 92,000 to 138,000 EUR.
- Total compensation includes potential annual cash bonuses, equity, sign-on payments, and health/wellbeing benefits.
- This position may require up to 10% travel.
- Recruiters can provide specific salary details for your location during the interview process.