Infrastructure Engineer

Open 26d

You are responsible for supporting application teams by turning around infra requests such as permissions, roles, service setup, and project peering so engineers can focus on delivering features. You will own CI/CD and deployments, maintain Terraform modules, manage Kubernetes deployments with Helm, and improve observability across the stack. You will also strengthen security controls, enable SOC 2 readiness, and explore AI-assisted automation to streamline routine tasks.

Responsibilities

  • Support the application teams by turning around infra requests (permissions, roles, service setup, project peering) so product engineers stay focused on shipping.
  • Own CI/CD and deployments by maintaining and extending GitHub Actions workflows and migrating toward a dedicated CD tool with proper permissioning for fully automated, locked-down deployments via service accounts and no direct engineer access to production.
  • Build and maintain infrastructure as code by authoring and updating Terraform modules for new and existing services across GCP environments.
  • Run Kubernetes the right way by managing service deployments via Helm (Helm 4) and keeping async workloads healthy on Dagster.
  • Unify observability by consolidating today's per-team alerting into a single view with system-to-system dashboards and incident alerting that routes upstream failures to the right teams and on-call rotations.
  • Advance resilience by moving toward a fully region- and cloud-agnostic posture so services can move if something fails.
  • Strengthen security and access by applying IAM, secrets management, least privilege, and auditability; contribute to SOC 2 readiness.
  • Automate with AI by building agent skills/agents.md so routine provisioning tasks can be handled by an agent and using AI to reason through bigger problems.

Requirements

  • Strong software engineering fundamentals in at least one production language (Python, Go, TypeScript, or Rust).
  • Python is especially valued, plus comfort with scripting and working in the shell.
  • Hands-on experience with cloud infrastructure and core cloud services, especially GCP.
  • Experience operating large-scale Kubernetes production systems.
  • Experience with infrastructure as code, especially Terraform.
  • Familiarity with CI/CD systems, especially GitHub Actions or Octopus Deploy.
  • Ability to debug production issues using logs, metrics, traces, shell tools, and source code.
  • Security and access-control fundamentals: IAM, secrets management, least privilege, and auditability.
  • Clear written communication around incidents, design decisions, and operational procedures.