SRE/DevOps Specialist
Summary
Designs and maintains AWS and Kubernetes infrastructure, writes Go/Python tools, and sets SLOs while leading incident response and chaos engineering for a Latin American food-tech and AI platform.
About the Company
iFood is a leading Brazilian technology company in Latin America. We connect thousands of restaurants to millions of consumers daily through innovative solutions, including food delivery, grocery, pharmacy, and pet markets, as well as our fintech arm, iFood Pago.
Responsibilities
- Investigate and respond to critical incidents, providing technical leadership in AWS environments, Kubernetes clusters, and network components.
- Develop internal solutions such as APIs and workers using Go and Python/LangGraph for AI solutions.
- Participate in metric standardization projects, including the definition of SLOs/SLIs, error budget alerts, and burn-rate.
- Collaborate with engineering teams and SREs to evolve internal solutions.
- Prepare applications for high-impact events.
- Contribute to technical refinement to increase team maturity.
- Define and govern standards through RFCs/IRCs.
- Lead deep Post-Incident Reviews (PIRs) and GameDays (Chaos Engineering).
Requirements
- Experience investigating and responding to critical incidents in AWS and Kubernetes.
- Ability to develop internal solutions using Go and Python.
- Experience with metric standardization (SLOs/SLIs).
- Ability to lead technical refinements and define standards via RFCs.
Preferred Qualifications
- Proficiency in Go (Golang) and Python.
- Knowledge of Chaos Engineering tools (e.g., Litmus) and performance testing (e.g., K6).
- Familiarity with Service Mesh (Istio) and Datadog.