DevOps Engineer
Summary
Metal Toad is seeking a DevOps/Cloud Engineer to design, maintain, and monitor scalable infrastructure for enterprise-grade applications. The role requires extensive AWS experience and strong scripting skills to support AI-focused projects in a fully remote environment.
Metal Toad is a strategic AI partner. We help our customers to use AI to build the engines of their future growth, aligning AI’s capabilities with their most important business opportunities. We are an AWS, Anthropic, and OpenAI partner.
Our team includes consultants, software engineers, devops engineers, UX designers, project managers, marketers, and supporting staff. It should go without saying, but everyone working at Metal Toad should be interested in the impact of AI, and believe in a human-first, AI-enabled future.
Metal Toad is a fully remote company, offering all team members the ability to work from home.
Although Metal Toad is a remote company, we are currently seeking contractors residing in Brazil for this position. Compensation will be made in Brazilian Reais (R$), in accordance with the local currency of the country where the position is located. Due to legal limitations, we are unable to sponsor any type of visa.
To apply for this position, please submit your resume in English as a PDF.
Job Description
The Cloud Engineer position at Metal Toad requires experience in designing and maintaining infrastructure for high-availability, scalable, enterprise-grade applications. You will be part of a talented team that works on mission-critical applications.
Responsibilities
Planning
- Analyzing customer requirements for software components, system availability, security, and performance.
- Designing and documenting complete cloud hosting systems, including capacity planning software and instance type selection, allocation, and network design.
- Estimating the costs of the recommended system design.
- Building systems by executing installation, configuration, and testing of cloud resources.
- Using automation and configuration management to ensure repeatability and traceability of changes.
Managed Services
- Troubleshooting system hardware, software, networks, and operating systems.
- Protecting the integrity and security of systems through proper use of controls and monitoring tools, and providing written evaluations and recommendations for ongoing improvement.
- Maintaining system performance through system monitoring and analysis, performance tuning, and planning for future growth.
- Designing and running load and stress tests, documenting outcomes, debugging infrastructure issues, and escalating documented application problems to the development team.
- Maintaining internal systems and customer deployment documentation.
- Partnering with project managers, technical consultants, software architects, and developers to validate infrastructure deliverables against the requirements and document all technical hand-offs.
- Experience with Amazon Web Services (AWS).
- Responding to support tickets and incidents in a timely manner that corresponds to SLA commitments.
Expertise
- Contributing to the definition of best practices, operational policies, and procedures.
- Establishing, documenting, and testing disaster recovery procedures, documenting outcomes, and making recommendations for ongoing improvement.
- Updating job knowledge by participating in educational opportunities, reading professional publications, maintaining personal networks, and participating in professional organizations.
Qualifications
- Read and agree to our Corporate Values Statement.
- Believe in the company's mission: to help people.
- Advanced to fluent English communication skills are essential for this role.
- 4+ years of experience with Amazon Web Services (AWS).
- Additional experience with other cloud providers is a bonus.
- Experienced in Linux and/or Windows Systems administration (at least one required).
- Scripting (bash shell and Python preferred, PowerShell acceptable).
- Knowledge of TCP/IP networking and HTTP protocols.
- Experience with web accelerators, load balancers, reverse proxies, and CDNs.
- Problem solver and willing to work in an agile/fast-paced environment.
- Customer-oriented with good communication skills.
- Willing to participate in a 24/7 on-call rotation with approximately one shift per month compensated.
Nice to have
- AWS Certifications or be willing to get certified.
- Interest in Generative AI technologies.
Please note that Metal Toad will never ask for payment or financial information during the hiring process. If you are asked for such information, it is a scam. Report it immediately.
What they ask for
Required
- Advanced to fluent English communication skills
- 4+ years of experience with Amazon Web Services (AWS)
- Experience in Linux and/or Windows Systems administration
- Scripting (bash shell and Python preferred, PowerShell acceptable)
- Knowledge of TCP/IP networking and HTTP protocols
- Experience with web accelerators, load balancers, reverse proxies, and CDNs
- Willing to participate in a 24/7 on-call rotation
Preferred
- Additional experience with other cloud providers
- AWS Certifications
- Interest in Generative AI technologies
