Principal Production Engineer

As a Principal Production Engineer you will build and operate a secure highly scalable and cost-effective AWS and Kubernetes based cloud platform. You will work across infrastructure automation, CI/CD pipelines, observability, and production reliability, helping ensure the platform stays reliable, scalable, and continuously improving for customers. You will participate in an on-call rotation and act as the subject matter expert for the production infrastructure stack and software build pipelines. You'll work independently on initiatives to improve security, stability, and responsiveness of applications, mentor teammates, and collaborate across teams to share knowledge. You'll leverage GenAI tools to accelerate infrastructure development and automation, build infrastructure-as-code with Terraform, develop tooling in Go or Python, and improve CI/CD pipelines. You will define monitoring and observability systems, respond to production incidents with root cause analysis, and automate operational runbooks, occasionally supporting deployments during off-hours.

Responsibilities

  • Serve as subject matter expert for the entire production infrastructure stack and software build pipelines
  • Work independently on initiatives to improve security, stability, and responsiveness of applications
  • Train and mentor other members of the team
  • Collaborate across teams to share knowledge and best practices
  • Support and operate the AWS-based cloud platform and Kubernetes (EKS) environments
  • Leverage GenAI tools to accelerate infrastructure development, automation, and auto-remediation of production issues
  • Build and maintain infrastructure-as-code using Terraform
  • Develop automation and internal tooling using Go or Python
  • Improve CI/CD pipelines to increase deployment safety and velocity
  • Define and improve monitoring, alerting, and observability systems
  • Respond to production incidents, conduct root cause analysis, and implement systemic improvements
  • Develop and automate operational runbooks and remediation workflows
  • Support production deployments, including during off-hours as needed

Requirements

  • 8+ years of experience in SRE, DevOps, or SaaS production operations
  • 5+ years of hands-on experience operating large scale production workloads in AWS
  • Strong experience with Terraform and infrastructure-as-code practices
  • 5+ years of experience with containerized environments using Docker and Kubernetes (EKS preferred); familiarity with Helm
  • Proficiency in Go or Python (or similar programming language)
  • Experience building and maintaining CI/CD systems (Git-based workflows, Argo, Jenkins or similar)
  • Strong Linux/Unix systems experience
  • Bachelor's degree in Computer Science or equivalent practical experience

Benefits

  • Health Benefits
  • Paid Time Off and Paid Holidays
  • Parental Leave
  • Equity
  • Monthly Wellness Reimbursement
  • Monthly Lunch on Legion