Site Reliability Engineer
Summary
Maintain and scale Qlik’s cloud platforms, ensuring high availability and reliability for millions of transactions using Kubernetes, Terraform, and observability tools like Prometheus.
- Exciting Challenges: Take on the responsibility of maintaining the reliability and availability of our cloud platforms, tackling complex problems and driving improvements to enhance performance and scalability.
- Collaborative Environment: Work closely with our Engineering organization, collaborating with Architecture, Platforms, and Domains teams to design and develop new infrastructure features and optimize cloud-related practices.
- Innovative Solutions: Design and develop effective tooling, alerts, and responses to identify and address reliability risks, utilizing your expertise in cloud technology and backend systems.
- Professional Growth: Act as a resource for fellow engineers, sharing your knowledge and expertise in cloud engineering, production service operations, incident management, and troubleshooting.
- Continuous Learning: Stay updated on the latest industry trends and technologies, contributing to the adoption of best practices and driving continuous improvement within our cloud environment.
- Reliability and Scalability: Ensure high reliability and availability of our cloud platforms, collaborating with cross-functional teams to implement new infrastructure features and optimize performance.
- Cloud Optimization: Define and evangelize cloud-related optimizations and best practices, driving improvements in reliability, scalability, and performance.
- Problem Solving: Analyze complex issues at the infrastructure, systems, network, and application levels, making recommendations and decisions to resolve them effectively.
- Knowledge Sharing: Share your expertise with fellow engineers, providing guidance on cloud technologies, automation, security, and best practices.
- On-Call Support: Participate in on-call duties to maintain the availability and performance of our cloud infrastructure, providing regular updates on project status and activities.
- Bachelor's or Master’s degree in Computer Science or a relevant field.
- Self-motivated with the ability to work autonomously and multitask effectively.
- Strong analytical skills for solving complex problems and driving innovative solutions.
- 3+ years’ experience with Infrastructure as Code (IaC) tools such as Terraform, Crossplane, Ansible, or similar
- 3+ years’ experience working alongside a production system running on Kubernetes
- 3+ years of professional experience in cloud engineering, preferably with AWS and Azure.
- 3+ years of Professional experience with operating and/or building microservices.
- Proficiency in scripting and automation (e.g., Bash, Python, Go) and software engineering concepts.
- Experience with CI/CD, cloud and microservice autoscaling.
- Experience with networking security and secret-management tools (e.g. Vault, AWS SSM).
- Proficiency with observability tooling such as Prometheus, Open Telemetry, distributed tracing, and SIEM such as Splunk.
- Experience with Helm including but not limited to managing helm charts as well as creating custom charts from existing ones or building new.
- Excellent English communication skills, both oral and written.
- Curiosity and desire to learn.
- Knowledge of infrastructure security review and compliance frameworks.
- Experience working with database concepts and tooling such as MongoDB
- Demonstrated ability to collaborate with development teams and provide expert guidance on implementing reliability best practices, ensuring systems are robust, scalable, and highly available.
- Where applicable, additional experience with other tools such as Temporal, Clik House, Fire Hydrant, Grafana, Solace, Gloo, and other cloud native related tools.
- Ability to obtain sufficient clearance status to work on IL5 systems with Qlik support.
- Due to this requirement: Must be a USA Citizen or be in process to become one by January 2027
- Ability to take a rotating on-call shift (24/7).
- Certifications such as CKD, CKS, AWS Certified Solutions Architect Associate/Professional, AWS Certified Advanced Networking Specialty, AWS Certified Security Specialty.
- Experience supporting FedRAMP or DoD IL4 certification initiatives by implementing security controls, driving audit readiness, and operationalizing compliant cloud infrastructure.
- Experience with self-hosted Temporal workflow infrastructure, including deployment, upgrades, scaling, monitoring, troubleshooting, and performance optimization across Kubernetes environments.
- Named in Newsweek’s ‘Americas Greatest Workplaces 2025’: .
- Genuine career progression pathways and mentoring programs.
- Culture of innovation, technology, collaboration, and openness.
- Flexible, diverse, and international work environment.
