Platform DevOps Engineer (5 years of experience Noght Shift)
Title: Platform / DevOps Engineer (Reliability & BCDR)
- Location: DHA phase 1
- Shift: 6pm - 3am
- Type: Full-time
- 2 year Contract
We’re looking for a Platform/DevOps engineer to own production reliability end-to-end. You’ll automate deployments, prove backups and disaster recovery with real exercises, keep production uptime healthy, tighten observability, maintain runbooks, and automate secrets rotation and access hygiene.
This is a hands-on role: you build and operate the systems, not just write tickets about them.
What you’ll own
Automate and harden CI/CD so production releases are repeatable, reviewable, and low-risk (less manual deploy work)
Design and run backup verification and restore exercises on a set cadence; document RTO/RPO results and close gaps
Own BCDR: keep business continuity / DR plans current, run tabletop and technical recovery drills, track follow-ups to done
Protect production uptime: monitoring, health checks, capacity, incident response, and postmortems
Keep Grafana dashboards and alerts useful—accurate views of production, less noise, faster debugging
Automate secrets management and rotation; remove long-lived / hardcoded credentials from repos and hosts
Improve infra-as-code and environment consistency (Terraform/OpenTofu, Docker, cloud resources)
Support compliance operations (SOC2 / healthcare-adjacent): evidence, access reviews, audit trails, control remediations
Partner with engineering on PostgreSQL operational work across environments (backup/restore, migration coordination, availability)
Must-have experience
Terraform or OpenTofu in production
Docker (and Compose) day-to-day
CI/CD ownership with GitHub Actions (or similar)
Strong PostgreSQL operations: backup/restore, basic HA, migration coordination—not just SQL queries
Observability ownership with Grafana + Prometheus/Loki (or equivalent)
Real BCDR / DR drill experience (not docs-only)
Secrets managers and rotation automation (Vault, Infisical, cloud secret managers, etc.)
Incident response and clear runbook writing
Comfort scripting ops automation (Bash, Python, or JS/TS)
Nice to have
PaaS / container-host operations (any modern app platform like Fly.io)
AWS security & compliance services (IAM, CloudTrail, GuardDuty, Config, SES)
SOC2 or HIPAA-adjacent operational experience