Site Reliability Engineer (SRE)
Summary
Design and develop infrastructure and reliability solutions, automate deployments and provisioning, and build observability tools (metrics, logs, dashboards) for complex customer environments, collaborating with R&D, Deployment, and Support teams. Core technologies include Python, Bash, Linux, Kubernetes, Helm, Prometheus, Grafana, Elastic, Fluentd/Vector, and Ansible.
Paragon is on a mission to transform the world of cyber intelligence.
Based in Tel Aviv, our innovative team is made up of top-tier talent who are passionate about making an impact. At Paragon, you’ll find the freedom to think boldly, collaborate with purpose, and grow alongside a team united by a shared mission - striving for excellence, and always looking out for one another.
Responsibilities
- Design and develop infrastructure and reliability solutions for complex customer environments, from architecture and development through production deployment.
- Develop automation for deployment, configuration, provisioning, and lifecycle management of our product and its infrastructure.
- Design and build observability solutions that provide deep visibility into system health, performance, reliability, and customer experience.
- Develop metrics, logs, events, dashboards, and alerting mechanisms to identify failures, understand system behavior, and proactively detect reliability issues.
- Build internal tools for debugging, diagnostics, troubleshooting, and operational visibility across application, container, system, and network layers.
- Investigate complex system and networking issues end-to-end, perform deep root cause analysis, and drive issues to resolution.
- Work closely with R&D, Deployment, and Support teams to improve observability, reliability, scalability, and operational efficiency.
Requirements
- 3+ years of hands-on experience in SRE, DevOps, or infrastructure engineering, working with production environments.
- Strong software engineering skills in Python, including developing production-quality solutions and automation. Bash scripting experience.
- Strong Linux knowledge, including system troubleshooting, networking, processes, configuration, and kernel-level behavior.
- Strong understanding of networking and troubleshooting, including TCP/IP, DNS, routing, NAT, VPNs, firewalls, and proxies.
- Hands-on experience with Kubernetes, including cluster troubleshooting, networking, deployments, and Helm. Experience creating and maintaining Helm charts.
- Strong experience with observability, including metrics, logs, dashboards, alerting, and end-to-end troubleshooting. Experience with Prometheus, Grafana, Elastic, Fluentd/Vector, or similar technologies.
- Experience with infrastructure automation, using Ansible or similar technologies.
- Strong debugging, root cause analysis, and problem-solving skills, with the ability to troubleshoot complex issues across multiple infrastructure layers.
- Strong ownership, self-management, and technical curiosity, with the ability to independently investigate and drive complex technical problems to resolution.
- Strong communication and collaboration skills, working effectively across R&D, Deployment, Support, and other engineering teams.
Nice to have:
- Experience with Terraform / Terragrunt or similar Infrastructure-as-Code technologies.