Site Reliability Engineer
Summary
SRE III operating and scaling GoDaddy's OpenStack-based global hosting platform — automating toil with Python and Puppet, improving observability, driving cloud migrations, and participating in on-call incident response across thousands of servers.
Location Details:
At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.
Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings.
About The Team....
Global Compute runs Optimised Hosting, GoDaddy's global platform for all customer hosting products. Squad R is the engineering team responsible for operating, scaling, and continuously improving the OpenStack-based clouds that power that platform. We treat reliability as an engineering problem: we automate toil away, we plan capacity ahead of demand, and we instrument everything so that we understand our systems before they surprise us. As an SRE III on the team, you'll be a senior technical contributor who others lean on for the hard problems.
What you'll get to do...
-
Operate and scale GoDaddy's cloud infrastructure, including our OpenStack-based hosting platform. You'll troubleshoot and improve services spanning compute, networking, and storage in large-scale production environments.
- Drive the OpenStack migration. Help move customer hosting workloads onto the platform safely — designing and executing migration tooling, validation, and rollback strategies that protect customer experience.
-
Work within a large-scale global hosting environment supporting thousands of servers and customer workloads across multiple regions.
- Eliminate toil through automation. Build and maintain automation in Python and Puppet to replace manual operational work. Treat repeated manual effort as a bug to be fixed.
- Strengthen observability. Improve monitoring, alerting, and dashboards so that signal reaches the right engineer at the right time, and so that we can reason about system behavior from data.
- Participate in on-call and incident response. Take a fair share of the on-call rotation, lead incident response when you're the responder, and drive the blameless post-incident process that turns failures into permanent fixes.
- Raise the engineering bar. Review peers' code and designs, document systems and runbooks clearly, and mentor SRE I/II engineers.
- Contribute to the AI/MCP initiative. Help bring AI-assisted workflows and internal MCP tooling into our operations so internal customers can resolve problems and incidents faster.
-
Participate in a shared on-call rotation (after onboarding) and help lead incident response activities when needed.
Your experience should include...
- 5+ years in SRE, infrastructure, platform, or systems engineering roles operating production systems at scale.
- Strong Linux systems fundamentals — networking, storage, processes, performance troubleshooting.
- Proficiency in Python for automation and tooling (writing maintainable, tested code — not just scripts).
- Hands-on experience operating distributed systems and diagnosing issues across service boundaries.
- Experience with infrastructure-as-code and configuration management (Puppet, Ansible, or equivalent).
- Comfort owning production reliability: on-call experience, incident response, and a track record of reducing toil through automation.
- Clear written and verbal communication — able to document systems, write post-incident reviews, and collaborate across a distributed, remote team.
You might also have...
-
Experience operating OpenStack (Nova, Neutron, Ceph) or comparable cloud infrastructure platforms
-
Experience with Docker, Kolla, or containerised infrastructure
-
Experience supporting large distributed systems at scale
- Experience operating Ceph or other software-defined storage at scale.
- Exposure to OpenStack migration or cloud-migration programs.
- Interest in applying AI/LLM tooling to operational workflows.
We encourage you to apply even if your experience or skillset doesn’t align perfectly with every requirement. We value a wide range of backgrounds and transferable skills, and we are excited to support learning and growth.
We've got your back... We offer a range of total rewards that may include paid time off, retirement savings (e.g., 401k, pension schemes), bonus/incentive eligibility, equity grants, participation in our employee stock purchase plan, competitive health benefits, and other family-friendly benefits including parental leave. GoDaddy’s benefits vary based on individual role and location and can be reviewed in more detail during the interview process.
About us... GoDaddy is empowering everyday entrepreneurs around the world by providing the help and tools to succeed online, making opportunity more inclusive for all. GoDaddy is the place people come to name their idea, build a professional website, attract customers, sell their products and services, and manage their work. Our mission is to give our customers the tools, insights, and people to transform their ideas and personal initiative into success. To learn more about the company, visit About Us.
At GoDaddy, we know diverse teams build better products—period. Our people and culture reflect and celebrate that sense of diversity and inclusion in ideas, experiences and perspectives. But we also know that’s not enough to build true equity and belonging in our communities. That’s why we prioritize integrating diversity, equity, inclusion and belonging principles into the core of how we work every day—focusing not only on our employee experience, but also our customer experience and operations. It’s the best way to serve our mission of empowering entrepreneurs everywhere, and making opportunity more inclusive for all. To read more about these commitments, as well as our representation and pay equity data, check out our Diversity and Pay Parity annual report which can be found on our Diversity Careers page.
We also embrace our diverse culture and offer a range of Employee Resource Groups (Culture). Have a side hustle? No problem. We love entrepreneurs! Most importantly, come as you are and make your own way.
GoDaddy is proud to be an equal opportunity employer. GoDaddy will consider for employment qualified applicants with criminal histories in a manner consistent with local and federal requirements. Refer to our full EEO policy.
Our recruiting team is available to assist you in completing your application. If they could be helpful, please reach out to myrecruiter@godaddy.com.
GoDaddy doesn’t accept unsolicited resumes from recruiters or employment agencies.
As published by greenhouse
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
- GoDaddy Employment History: choose one
- If you are a former employee, please indicate your GoDaddy work email address (if you remember). optional
- Do you have the legal right to perform this role in the country in which you are applying? choose one
- Is your right to work based (either now or in the future) on obtaining a visa, work permit or equivalent? choose one
- If you answered "yes" to the above, please provide details. written answer · optional
- Are you related to, or have you had (now or previously) a close personal or business relationship with, any current GoDaddy employees? Relevant business relationships include joint business ownership, shared investments or other financial interests. choose one
- If yes, what is the name of the employee(s) and the nature of the relationship (e.g. spouse, sibling, business partnership, etc.) written answer · optional
- Have you ever previously interviewed with GoDaddy or any of its subsidiaries (123-Reg, Freedom Voice, HEG, Media Temple, Main Street Hub, Sellbrite, Sucuri, etc.)? choose one
- Have you been introduced to GoDaddy or any of its subsidiaries by a recruitment agency within the last 12 months? choose one
- Have you ever been offered employment with GoDaddy or any of its subsidiaries (123-Reg, Freedom Voice, HEG, Media Temple, Main Street Hub, Sellbrite, Sucuri, etc.) choose one
- What is the soonest you would be able to start? written answer
- How did you hear about us? choose one
- If you selected "Employee Referral", please list the referring employee's name. optional
- LinkedIn profile (if you do not have one or if you prefer not to provide one, enter N/A):
- GitHub profile (N/A if you do not have one) optional
- In addition to email communications, GoDaddy utilizes a SMS/WhatsApp texting tool to send quick updates and communicate with you throughout our candidate processes. GoDaddy Executives do not contact candidates or employees by SMS or WhatsApp. Per GoDaddy policy, we will never text you a link to click. Please indicate below if you wish to receive SMS/WhatsApp messages from GoDaddy. Standard messaging and data rates may apply. choose one
- By submitting your application: (i) you confirm that the information you provide is accurate and complete to the best of your knowledge and you will inform us immediately if you subsequently become aware of any errors or inaccuracies; (ii) you acknowledge you are applying for a role with a company within the GoDaddy group, details of which will be provided at a later stage of the recruitment process if applicable; and (iii) you confirm you have read and understood the terms of our privacy policy (https://www.godaddy.com/legal/agreements/privacy-policy?isc=gdbb3376d). If you have questions regarding the application process or experience technical issues, please email myrecruiter@godaddy.com. If you have questions about our privacy practices, please email privacy@godaddy.com. choose one
- Tell us about your experience working in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or similar roles. What types of systems have you supported? written answer
- What programming languages, scripting languages, or automation tools have you used regularly in your current or recent roles? written answer
- Describe the largest infrastructure environment you have operated. Include approximate server count, users/customer impact, and your role. written answer
- Which country are you currently located in?
- This role includes participation in a shared on-call rotation after approximately six months, allowing time to onboard and become familiar with our systems. The rotation provides 24/7 coverage for a seven-day period, typically starting on Fridays. SREs are responsible for responding to production incidents, leading investigation and resolution efforts when required, and driving improvements that enhance long-term reliability. Are you comfortable participating in this on-call rotation?