SRE | Analista DevOps Pleno
Summary
This is a mid-level Site Reliability Engineering (SRE) and DevOps role focused on building, maintaining, and evolving reliable cloud infrastructure. The role involves cloud administration, infrastructure as code, CI/CD automation, observability, troubleshooting, and operational support, with a strong emphasis on automation and reliability.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE | Analista DevOps Pleno based in Brazil.
This is a mid-level Site Reliability and DevOps role focused on building, maintaining, and evolving reliable cloud infrastructure.
You will help ensure high levels of availability, stability, scalability, security, and performance across technology environments.
The role combines cloud administration, infrastructure as code, CI/CD automation, observability, troubleshooting, and operational support.
You will work closely with development and security teams to enable faster, safer, and more reliable software delivery.
A strong automation mindset is essential, with opportunities to eliminate repetitive work and continuously improve infrastructure processes.
You will also contribute to incident response, production support, metrics analysis, and service reliability within a collaborative technical environment.
This position offers flexible working arrangements, professional development opportunities, and the chance to work on modern cloud and DevOps practices.
Accountabilities:
- Cloud infrastructure: Deploy, configure, administer, and maintain cloud resources including servers, networks, firewalls, databases, and load balancers using public cloud platforms.
- Architecture and reliability: Design and evolve infrastructure architectures with a focus on security, stability, scalability, availability, and reliability, supporting strong SLA/SLO performance.
- Development enablement: Partner with development and security teams to define infrastructure solutions and provide the technical foundation required for efficient application delivery.
- Monitoring and operations: Monitor production and non-production environments, troubleshoot operational issues, and provide support to maintain service continuity.
- On-call support: Participate in an 18x7 on-call rotation, contributing to incident response, production troubleshooting, and service recovery.
- Infrastructure automation: Automate infrastructure and support activities to reduce manual work, improve consistency, and increase operational efficiency.
- Infrastructure as Code: Develop and maintain infrastructure as code using appropriate IaC and GitOps practices, ensuring code remains secure, documented, and up to date.
- CI/CD: Build, maintain, and automate continuous integration and continuous delivery pipelines to improve the speed, quality, and reliability of software releases.
- Deployment support: Work alongside developers throughout application deployment cycles and participate in technical incident or crisis rooms when required.
- Observability and metrics: Generate and analyze performance, availability, and reliability metrics, providing actionable insights and feedback to development teams.
- Documentation: Keep infrastructure, operational procedures, and technical documentation accurate and current.
- Continuous improvement: Identify opportunities to improve infrastructure, reliability, automation, monitoring, and support processes through a strong “automate everything” mindset.
- Cloud: Practical experience with public cloud platforms such as AWS and/or GCP.
- Containers: Hands-on experience with container engines such as Docker, containerd, and/or Podman.
- Infrastructure as Code: Experience with IaC and GitOps tools such as Terraform, AWS CloudFormation, AWS CDK, Ansible, Chef, and/or Puppet.
- CI/CD: Experience creating or maintaining continuous integration and continuous delivery pipelines using tools such as GitLab CI, Jenkins, Azure DevOps, CircleCI, and/or GitHub Actions.
- Operating systems: Strong knowledge of Linux and/or Windows administration, troubleshooting, and operational support.
- Scripting: Solid scripting skills, particularly with Bash and PowerShell.
- Observability: Good knowledge of monitoring and observability platforms such as Prometheus, Grafana, Kibana, Fluentd, and/or Elasticsearch.
- Architecture: Familiarity with microservices architectures and distributed systems.
- Web infrastructure: Knowledge of administering and troubleshooting web servers such as NGINX, Apache, and/or IIS.
- Networking: Good understanding of networking protocols including TCP and HTTP.
- Automation mindset: Strong interest in automation, reliability engineering, operational efficiency, and continuous improvement.
- Container orchestration: Experience managing, monitoring, or administering container orchestration services such as GKE, AWS EKS, or AWS ECS is a plus.
- DevOps culture: Previous experience working in environments that adopt DevOps practices and principles is advantageous.
- Programming: Proficiency in at least one programming language such as Python, Ruby, or Go is a plus.
- Databases: Advanced knowledge of relational databases such as MySQL, PostgreSQL, and/or SQL Server is an advantage.
- Flexible working hours with multiple schedule options to align with team and personal routines.
- Home office allowance.
- Medical and dental insurance, subject to applicable eligibility rules.
- Life insurance, subject to applicable eligibility rules.
- Childcare assistance, subject to applicable eligibility rules.
- Psychological, legal, and financial support through an employee assistance program.
- Wellhub (Gympass) access.
- Transportation allowance.
- Flexible meal and/or food allowance.
- Birthday day off.
- Language-learning incentives.
- Education incentives.
- Reimbursement and incentives for professional certifications.
- Structured performance reviews and a continuous feedback culture.
- Profit Sharing and Results (PLR), subject to applicable eligibility rules.
- Collaborative environment with opportunities for technical growth and continuous learning.
- Exposure to modern cloud infrastructure, DevOps, automation, observability, and reliability engineering practices.
- Full-time opportunity in Brazil with a strong focus on professional development and technical evolution.