Senior DevOps Engineer
Summary
Senior DevOps engineer owning the build, test, and deployment pipeline for a distributed computer vision edge platform, working with Bazel, Python, Ansible, Docker, Golang, and LLM-driven automated testing systems.
What You'll Do
- Own and evolve the end-to-end release pipeline — branching strategy, build orchestration, artifact promotion, and rollback — across our Bazel monorepo and Python deployable units
- Design and maintain Ansible-driven fleet automation for heterogeneous Linux edge nodes (Ubuntu LTS, NVIDIA driver stacks, Docker with NVIDIA runtime)
- Manage all update tooling, currently written in Golang
- Build LLM-powered automated testing systems: test generation from specs, flake triage, log/failure analysis, regression diffing, and release-note synthesis from commit and ticket history
- Harden CI/CD for offline and bandwidth-constrained deployment targets (airgap wheel distribution, signed artifacts, deterministic builds)
- Drive observability for releases — deployment telemetry, version drift detection, and post-deploy health validation across the fleet
- Mentor engineers on release hygiene, reproducible builds, and infrastructure-as-code practices
- 10+ years of professional experience in release engineering, DevOps, or SRE roles shipping production Linux systems
- Deep curiosity for software, infrastructure, and applied AI — particularly using LLMs as production engineering tools, not just chat assistants
- Expert-level Python (3.8+) with a strong grasp of packaging, dependency resolution, and PEP 440 versioning discipline
- Demonstrated ownership of Linux fleets at scale — kernel, systemd, networking, package management
- Excellence in technical communication, runbook authorship, and post-incident documentation
- Strong systems thinking — comfortable reasoning about failure modes across hardware, OS, container, and application layers
- Expert proficiency with Ansible (roles, dynamic inventory, idempotent design); working knowledge of Terraform
- Expert proficiency with Docker, including creation and lifecycle management of containers, image hardening, registry management and installing & configuring the NVIDIA container runtime
- Production experience with Linux administration: systemd, networking (VLANs, DHCP, DNS), kernel/driver management (especially NVIDIA/DKMS), package and APT internals
- Strong Python skills focused on tooling, automation, packaging (wheels, pip, private indexes), and subprocess/CI integration
- Proficiency with Git workflows, branching strategies, and modern CI/CD systems (GitHub Actions, GitLab CI, or equivalent)
- Experience designing and operating automated test infrastructure — unit, integration, hardware-in-the-loop, and end-to-end
- Practical experience using LLMs (Anthropic, OpenAI, or local) as part of engineering workflows — test generation, code review augmentation, log analysis, or agentic tooling
- Bazel or similar monorepo build systems
- Edge or embedded deployment experience
- Tailscale, WireGuard, or zero-trust networking in production
- gRPC/protobuf service ecosystems
- Vault, PKI, or secrets management at fleet scale
- Background in regulated or compliance-driven environments (CMMC, ISO 27001, SOC 2)
Skills
- Agentic AI
- AI
- Ansible
- Anthropic
- Automation
- CI/CD
- Cmmc
- Computer Vision
- DevOps
- DHCP
- DNS
- Docker
- Git
- GitHub
- GitHub Actions
- GitLab
- Go
- gRPC
- Infrastructure as Code
- ISO 27001
- Linux
- LLM
- Networking
- Observability
- OpenAI
- PKI
- Protobuf
- Python
- Secrets Management
- SOC 2
- Terraform
- Test Automation
- Ubuntu
- Vault
- VLAN
- Zero Trust