Site Reliability Engineer [SRE]
Summary
An SRE ensures high-availability banking systems by designing scalable cloud infrastructure on GCP, automating .NET-based ops tasks, and optimizing reliability through observability and incident response.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer [SRE] based in Brazil.
This is a fully remote opportunity for an SRE focused on building reliable, scalable, and high-performing production environments.
You will play a key role in monitoring applications, detecting failures, improving resilience, and ensuring operational continuity.
The position combines cloud engineering, automation, observability, incident response, and infrastructure optimization.
You will contribute to a cloud migration to GCP, applying strong practices around security, scalability, automation, and cost management.
The role sits within a banking environment, where reliability and performance are critical to business operations.
You’ll collaborate across technical and business teams, connecting system behavior and business requirements to practical engineering solutions.
This is an environment that values autonomy, continuous learning, collaboration, and proactive problem-solving.
Accountabilities
- Continuously monitor application environments to maintain high availability, performance, and rapid detection of production incidents.
- Analyze existing application architecture and infrastructure, identifying opportunities to improve reliability, scalability, performance, and operational efficiency.
- Support and contribute to the migration of workloads and environments to Google Cloud Platform (GCP), applying best practices for security, automation, scalability, and cost optimization.
- Apply Site Reliability Engineering principles to production systems through automation, effective monitoring, incident response, and root-cause analysis.
- Develop automation solutions using .NET to reduce manual operational tasks and improve the reliability of technology environments.
- Analyze technical environments and integrate business rules and requirements into systems, processes, and delivery pipelines.
- Identify recurring operational issues and recommend improvements that increase system resilience and reduce the likelihood and impact of incidents.
- Collaborate with development, operations, and other stakeholders to promote reliable and efficient technology delivery within a banking environment.
- Professional experience working with Google Cloud Platform (GCP), applying cloud best practices for scalability, security, reliability, and performance.
- Practical experience with .NET development, particularly for automating operational and infrastructure-related tasks.
- Hands-on knowledge of Site Reliability Engineering (SRE) principles, including monitoring, automation, incident response, reliability engineering, and root-cause analysis.
- Ability to understand business rules and translate them into technical solutions, systems, and pipelines.
- Strong analytical and problem-solving skills, with a proactive approach to identifying and resolving reliability and performance issues.
- Ability to collaborate effectively with professionals from different technical areas and organizational levels.
- Self-management skills, autonomy, and comfort working in dynamic environments and outside of established routines.
- Strong communication skills and a collaborative mindset, with willingness to learn, share knowledge, and contribute to team development.
- Experience with Agile methodologies, such as Scrum or Kanban, and alignment with DevOps practices is a plus.
- Familiarity with advanced monitoring and observability tools is a plus.
- Relevant cloud certifications, such as Google Cloud Professional certifications, are considered a differentiator.
- 100% remote / home-office work model.
- Health insurance.
- Dental assistance.
- Meal and food allowance through Flash.
- Home-office allowance.
- Gympass / Wellhub access.
- Life insurance.
- Extended maternity and paternity leave.
- Partnerships and discounts in education, healthcare, wellness, fitness, language schools, and leisure.
- Continuous feedback culture, including semiannual feedback cycles, 1:1 meetings, Individual Development Plans (PDI), and development initiatives.
- Inclusive and collaborative work environment focused on professional growth and continuous learning.