Site Reliability Developer (python/java) / SRE
Summary
Site Reliability Engineer at WatchGuard Technologies in Valencia, Spain, driving reliability and security of the production cloud across AWS, Azure, and hybrid environments. Day to day: leading incident response, defining operational standards, guiding teams on SLOs, and automating cloud operations with Python/Java/Go, Terraform/CloudFormation, and New Relic.
In this role you drive reliability and security for WatchGuard’s production cloud, partnering with application teams to deliver a stable customer experience. You lead large-scale incident response, define operational standards, and help teams meet service levels. You’ll automate and optimize cloud operations across AWS, Azure, and hybrid environments while growing as a cloud and SRE expert. This is a hands-on role that emphasizes impact, collaboration, and continuous learning.
Compensaciones / Beneficios- flexible work options
- caregiver support benefits
- parental leave
- DEI commitment
- disability accommodations
- equal employment opportunities
- Ensure smooth production operations with development teams and lead large-scale incident response
- Define operational and security policies, standards, and processes for development teams
- Guide teams to establish, monitor, and achieve service level indicators and objectives
- Collaborate with application teams in production environments to ensure monitoring, security, reliability, and automation
- Drive operational excellence through simplification, automation, analysis, and process evolution
- Champion security and operational best practices to become a cloud expert across global teams
- Participate in on-call rotation and coordinate production troubleshooting efforts
- Develop automation or assist with debugging complex production issues
- Customer-focused and data-driven mindset
- Proficiency in Python, Java, or Go
- Experience with full software development lifecycle practices (coding standards, code reviews, security, CI/CD, automated testing)
- Proven ability to lead production incident response and postmortems
- Strong analytical and problem-solving abilities with excellent verbal and written communication
- Knowledge of cloud technologies and DevOps/SRE practices
- Familiarity with tools and technologies listed in the description (see external_tools)
- Analytical thinking
- Verbal and written communication
- Collaboration and teamwork
- Cloud platforms (AWS, Azure, hybrid)
- Automation and IaC (CloudFormation, Terraform)
- Monitoring and observability (New Relic)