Senior Infrastructure Engineer, Government Systems
You'll be a technical leader on the Government Systems team, acting as the connective tissue between infrastructure operations and product delivery for security-sensitive government customers. You'll lead the design and evolution of Kubernetes clusters, own the CI/CD platform end-to-end, and drive GitOps-based delivery. You'll architect infrastructure as code, define container image lifecycle policies, and eliminate operational toil through automation. You'll also establish monitoring and incident response practices, serve as a technical point of contact for product teams, own the security posture of the deployment platform, and mentor other engineers.
Responsibilities
- Lead the design and evolution of RKE2 Kubernetes clusters across development, acceptance, and production environments, driving upgrades, capacity planning, networking strategy, and operational resilience
- Architect and maintain infrastructure as code (Terraform) to provision and manage AWS-based environments, establishing patterns and modules that the team can scale with confidence
- Own the CI/CD platform (GitHub Actions) end-to-end, designing pipeline architecture, improving build performance, and ensuring releases flow reliably through staged environments
- Drive GitOps-based delivery strategy using Flux CD, defining standards for Helm charts and Kustomize overlays and ensuring consistent reconciliation across clusters
- Define container image lifecycle policies, building, signing, storing, and distributing images across OCI registries for multiple deployment targets
- Identify and eliminate operational toil through automation, improving environment provisioning, configuration management, and deployment processes at a systemic level
- Establish and maintain monitoring, alerting, and incident response practices, owning dashboards, runbooks, and post-incident reviews that raise platform reliability over time
- Serve as a technical point of contact for product service teams integrating into the deployment pipeline, unblocking teams, troubleshooting complex environment issues, and advocating for infrastructure best practices
- Own the security posture of the deployment platform, managing secrets, certificates, RBAC policies, and security configurations to meet compliance and operational security requirements
- Mentor and grow other engineers on the team through code review, pairing, design discussions, and knowledge sharing
Requirements
- Deep experience operating Kubernetes in production at scale, including cluster lifecycle management, Helm, networking, persistent storage, performance tuning, and complex workload troubleshooting
- Expert-level proficiency with Terraform for provisioning and managing cloud infrastructure across multiple environments, including module design and state management strategies
- Strong Linux systems administration skills (RHEL or similar, including networking, storage, systemd, performance analysis, shell scripting)
- Extensive experience designing and maintaining CI/CD pipelines (GitHub Actions, Jenkins, or similar), with a focus on reliability, speed, and developer experience
- Deep familiarity with AWS services commonly used in infrastructure (EC2, S3, VPC, IAM, EBS, Lambda, and related networking and security services)
- A strong operational mindset with proactive thinking about failure modes and automation to prevent recurring issues
- Eligibility to access U.S. Government information and systems; only U.S. citizens are eligible for this role
Techstars