Senior DevOps Engineer
Summary
Remote (Brazil-based) Senior DevOps Engineer who designs, runs, and improves core infrastructure — Kubernetes, compute, networking, CI/CD, and observability — using Terraform and Python. Day to day involves managing clusters, automating infrastructure as code, leading production incident triage, and improving developer experience.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer based in Brazil.
This is an opportunity to shape the infrastructure that enables modern software teams to build, deploy, and operate products at scale.
You’ll own critical infrastructure across compute, networking, CI/CD, observability, and Kubernetes environments.
The role combines hands-on engineering with architectural decision-making and a strong focus on developer experience.
You’ll help reduce operational friction, establish reliable engineering practices, and make the easiest path the safest one.
You’ll also play a key role in production incident response, using data, automation, and sound engineering judgment to solve complex problems.
Working closely with cross-functional engineering teams, you’ll influence technical direction while enabling others to deliver more effectively.
This is a highly collaborative, remote role suited to an engineer who enjoys ownership, continuous improvement, and practical innovation.
Accountabilities
- Design, build, operate, and continuously improve core infrastructure covering compute, networking, CI/CD, Kubernetes, and observability.
- Manage Kubernetes environments at the infrastructure level, including cluster configuration, node pools, upgrades, scaling, and networking.
- Develop and maintain infrastructure-as-code using Terraform or comparable technologies, ensuring infrastructure remains consistent, scalable, and maintainable.
- Improve developer experience by reducing operational friction, eliminating repetitive work, and establishing intuitive, reliable engineering paths and tooling.
- Lead production incident triage by analyzing logs and operational data, developing diagnostic scripts when useful, identifying root causes, and driving issues through to resolution.
- Shape infrastructure architecture with an emphasis on simplicity, scalability, reliability, and ease of use for engineering teams.
- Partner closely with software engineers and other stakeholders to understand technical needs and translate them into effective infrastructure solutions.
- Use AI-powered tools thoughtfully to increase productivity, improve problem-solving, and help the broader engineering organization adopt effective AI-assisted workflows.
- Contribute to internal platform and developer-experience initiatives that improve consistency, automation, and engineering efficiency across teams.
- Demonstrated experience managing and operating Kubernetes or a comparable container orchestration platform, including hands-on responsibility for clusters rather than only deploying applications to them.
- Strong practical experience with infrastructure-as-code, particularly Terraform or equivalent tools such as Pulumi or CloudFormation.
- Solid DevOps and infrastructure engineering fundamentals, combined with strong judgment around simplicity, scalability, maintainability, and architectural trade-offs.
- Proven ability to own production troubleshooting, including interpreting logs, forming hypotheses, writing scripts to investigate issues, and working effectively under operational pressure.
- A combination of infrastructure/DevOps expertise and software engineering skills, particularly with Python.
- Strong understanding of modern cloud infrastructure and distributed systems.
- Clear and effective communication skills, with the ability to explain technical concepts and trade-offs to both engineering and non-technical stakeholders.
- A proactive, collaborative mindset focused on enabling other engineers and improving the broader development ecosystem.
- Genuine hands-on experience with AI tools and an interest in using them as practical productivity and engineering multipliers.
- Experience with Google Cloud Platform, particularly GKE Autopilot or a comparable managed Kubernetes environment, is desirable.
- Familiarity with gRPC, PostgreSQL at scale, service-oriented architectures, growth-stage infrastructure, regulated environments such as HIPAA, or internal platform/paved-path initiatives is a plus.
- 100% remote working model.
- Comprehensive company-paid medical insurance and mental health support programs.
- 5 undocumented sick-leave days per year.
- 20 working days of paid annual vacation, plus local public holidays.
- Regular internal learning events and opportunities to develop your technical expertise.
- Access to a global community of experienced technology professionals for knowledge sharing and collaboration.
- Internal mobility opportunities to support career growth and the possibility of moving between projects when appropriate.
- Exposure to large-scale, international projects with fast-growing clients.
- Friendly and collaborative working environment with an informal culture, open communication, and regular team-building activities.
- Long-term employment opportunities within a globally distributed technology organization.
- Opportunities to work with modern cloud, infrastructure, automation, and AI technologies.
Requirements
Benefits
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company