Engenheiro SRE Especialista
Summary
Specialist SRE engineer working primarily as a Cloud DBA for mission-critical AWS database platforms — Amazon RDS SQL Server, Amazon Aurora, and DynamoDB — owning availability, performance, security, and disaster recovery. Blends deep DBA expertise with SRE practices, observability (CloudWatch, Dynatrace), IaC, automation, and CI/CD.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engenheiro SRE Especialista based in Brazil.
This role is focused on ensuring the reliability, performance, security, and continuous evolution of critical cloud database environments on AWS. You will act primarily as a Cloud DBA, with strong ownership of Amazon RDS SQL Server, Amazon Aurora, and DynamoDB platforms. The position combines deep database expertise with SRE, DevOps, automation, and Infrastructure as Code practices to improve scalability and operational efficiency. You will work with mission-critical environments where availability, resilience, performance, and data protection are essential. The role also provides an opportunity to modernize operational practices through observability, automation, CI/CD, and cloud-native technologies. You will contribute to robust and scalable data platforms while collaborating with technology teams in a dynamic and innovation-oriented environment
Accountabilities:
- Take ownership of the administration, availability, performance, security, troubleshooting, and continuous improvement of AWS database environments, with a primary focus on Amazon RDS SQL Server, Amazon Aurora, and Amazon DynamoDB.
- Manage and optimize mission-critical relational and NoSQL databases, supporting operational stability, workload performance, capacity, and the evolution of data platforms in distributed cloud environments.
- Design and maintain high-availability and disaster-recovery strategies, including AlwaysOn, Multi-AZ, Read Replicas, failover mechanisms, backup and restore procedures, snapshots, replication, and data recovery strategies.
- Perform advanced database performance tuning by analyzing queries, indexes, partitioning strategies, workloads, and other factors affecting SQL and NoSQL performance.
- Develop and maintain database observability practices using AWS CloudWatch, Dynatrace, and related monitoring tools, ensuring visibility into availability, performance, capacity, and operational health.
- Manage AWS Backup configurations, retention policies, data protection practices, compliance requirements, and governance standards for cloud database environments.
- Administer DynamoDB environments, including Global Tables, DynamoDB Streams, partitioning strategies, and performance-oriented data modeling, while optimizing the use of AWS resources supporting the data layer.
- Apply SRE principles to database operations by improving observability, reliability, capacity management, incident response, service resilience, and root-cause analysis.
- Automate infrastructure and database administration activities using Infrastructure as Code and scripting technologies, helping reduce manual effort and increase operational consistency.
- Contribute to CI/CD pipelines and automated deployment processes related to databases, while applying Git and GitHub practices for version control and change management.
- Collaborate with engineering and application teams to support performance analysis, troubleshooting, and integrations involving .NET and Node.js applications and their database environments.
- Proven experience working as a DBA in AWS environments, with strong hands-on expertise in Amazon RDS SQL Server, Amazon Aurora, and Amazon DynamoDB.
- Solid experience administering, supporting, troubleshooting, and optimizing databases in mission-critical environments, with a strong understanding of operational reliability and performance requirements.
- Strong knowledge of relational and NoSQL data modeling, including the ability to design structures that support scalability, performance, and distributed workloads.
- Advanced expertise in high availability and disaster recovery, including AlwaysOn, Multi-AZ, Read Replicas, failover, contingency strategies, backup, restore, snapshots, replication, and distributed data recovery.
- Advanced database performance tuning skills, including query analysis, indexing, partitioning, workload optimization, and troubleshooting across both SQL and NoSQL environments.
- Experience with database monitoring and observability using CloudWatch, Dynatrace, or comparable tools, together with practical knowledge of AWS Backup, retention policies, compliance, and data governance.
- Deep knowledge of DynamoDB, including Global Tables, DynamoDB Streams, partitioning strategies, and performance-oriented data modeling, as well as strong knowledge of EC2 instances and AWS resources supporting database environments.
- Knowledge of SRE practices applied to the data layer, including observability, availability, capacity management, incident management, service resilience, and root-cause analysis.
- Hands-on experience with Infrastructure as Code using Terraform and/or CloudFormation to provision and manage AWS environments.
- Experience automating administrative and operational tasks using Python, PowerShell, Shell Script, or comparable scripting technologies.
- Knowledge of CI/CD pipelines and automated deployment practices for database environments, together with Git and GitHub for version control and change management.
- Familiarity with .NET and Node.js applications for performance analysis, troubleshooting, and database integration support.
- Strong analytical and problem-solving abilities, combined with a proactive approach to operational challenges, continuous improvement, and reliability engineering.
- Opportunity to work with mission-critical AWS cloud database environments and modern data technologies.
- Hands-on exposure to Amazon RDS SQL Server, Amazon Aurora, DynamoDB, AWS Backup, CloudWatch, Dynatrace, and other cloud-native services.
- Opportunity to combine deep DBA expertise with SRE, DevOps, Infrastructure as Code, automation, observability, and CI/CD practices.
- Professional development supported by technical training, knowledge-sharing initiatives, technology communities, leadership development, and partnerships with learning organizations.
- Collaborative and agile environment focused on technology, innovation, continuous learning, and knowledge sharing.
- Opportunity to contribute to scalable and resilient solutions while working on complex technical challenges with meaningful business impact.
- Inclusive environment that values diversity, respect, ethics, and equal opportunities for professionals from different backgrounds.
Requirements
Benefits
Skills
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company
