Senior/Staff DevOps Engineer
Summary
Design and run secure, scalable cloud infrastructure for a Series A startup serving regulated sectors like defense and life sciences, using IaC, GitOps, and AI-assisted tooling.
About the Company
Ethos is on a mission to bridge the human readiness gap by transforming how training is developed, consumed, and aligned with strategic business outcomes. As a well-funded Series A startup, we are a trusted partner to enterprise customers across the U.S. military, life sciences, manufacturing, and professional sports.
Responsibilities
- Architect, implement, and run secure, scalable, multi-tenant infrastructure using IaC, immutable artifacts, and GitOps
- Use AI coding and agentic tools for IaC authoring, pipeline development, and toil reduction
- Build and harden CI/CD pipelines for multi-environment delivery, including air-gapped workflows
- Establish observability through SLOs, metrics, logs, and traces
- Integrate supply-chain security, secrets management, and baseline hardening
- Optimize infrastructure spend and performance through capacity planning and autoscaling
- Lead design reviews, author RFCs, and mentor engineers
- Support IL-4/IL-5-aligned patterns and RMF documentation
Requirements
- 5+ years building and operating cloud platforms
- 3+ years deploying SaaS in production
- Strong expertise with Terraform, Helm/Kustomize, and containers (Docker, Kubernetes)
- Deep AWS experience (VPC, EKS, EC2, S3, RDS, ECR, IAM/KMS)
- CI/CD expertise (GitHub Actions, CircleCI, or Argo Workflows) and GitOps (Argo CD or Flux)
- Experience with observability tools (Prometheus/Grafana, OpenTelemetry, or ELK)
- Active or eligible for a Secret Clearance
- Proficiency using AI development/operations tools in daily workflows
Preferred Qualifications
- Experience with supply-chain security (SBOMs, SLSA, image signing)
- Background with DoD/regulated customers and familiarity with IL-4/IL-5 or Platform One
- Knowledge of STIG/CIS hardening and air-gapped architectures
- Experience operating AI/ML workloads in production (GPU scheduling, vector DBs, or inference serving)