Senior Systems Engineer, IT Operations and Infrastructure
Summary
Systems engineer on Firestorm's IT Operations team supporting business-critical infrastructure across cloud and on-premise environments, including Microsoft 365 GCC-High and AWS-Gov. Day to day: provisioning servers/VMs and cloud instances, monitoring, incident response, automation, networking, backups/DR, and security compliance with Linux/Windows Server, Bash/Python/PowerShell, AWS/Azure/GCP, an
Who We Are
About the Role
What You’ll Do
- Infrastructure Provisioning & Maintenance: Deploy, configure, patch, and manage physical servers, virtual machines, cloud instances (AWS, Azure, GCP), and enterprise storage.
- System Monitoring & Health Checks: Implement and maintain observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch) to track system performance, resource utilization, and uptime.
- Incident Response & Triage: Serve as an escalation point for system outages and critical alerts; lead root-cause analysis (RCA) and document post-mortems to prevent recurrence.
- Automation & Scripting: Write and maintain scripts (Bash, Python, PowerShell) and infrastructure-as-code configurations (Terraform, Ansible) to automate routine operational tasks and deployments.
- Network & Connectivity Management: Monitor core network services (DNS, DHCP, VPNs, routing, and firewalls) to maintain secure, low-latency communication across environments.
- Backup & Disaster Recovery: Design, schedule, test, and verify automated backups, data replication, and disaster recovery plans to ensure strict Recovery Point and Recovery Time Objectives (RPO/RTO).
- Security & Compliance: Apply system hardening standards, manage access control (IAM/Active Directory/SSO), audit system logs, and remediate CVE vulnerabilities in coordination with the security team.
- Capacity Planning & Performance Tuning: Analyze compute, storage, and network trends to optimize system performance and forecast hardware or cloud resource scaling needs.
- Documentation & Standard Operating Procedures: Author and update runbooks, system architecture diagrams, standard operating procedures (SOPs), and operational workflows for team use.
Required Qualifications
- U.S. Citizenship and ability to obtain and maintain a U.S. Government security clearance
- Operating Systems: Deep operational proficiency with Linux distributions (RHEL, Ubuntu, Rocky) and/or Windows Server environments, including OS hardening, kernel tuning, and troubleshooting.
- Scripting & Automation: Strong practical scripting skills in at least one language (Bash, Python, or PowerShell) to automate administrative tasks and repetitive workflows.
- Networking Fundamentals: Solid understanding of core networking protocols and services (TCP/IP, DNS, DHCP, VLANs, HTTP/S, SSH, VPNs, and firewall rules).
- Virtualization & Cloud: Hands-on experience administering virtualized infrastructure (VMware ESXi/vCenter, Hyper-V, KVM) or public cloud platforms (AWS, Azure, or GCP).
- Monitoring & Tooling: Experience configuring and operating enterprise monitoring, logging, and APM tools (e.g., Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch).
- Experience with Microsoft GCC-High environments, AWS-Gov and similar secure environments that have CUI and ITAR data
- Experience with CMMC v2 and NIST 800.171 compliance needs
Preferred Qualifications
- Infrastructure as Code (IaC): Working knowledge of configuration management and provisioning tools such as Terraform, Ansible, Puppet, or SaltStack.
- Containerization: Familiarity with Docker and container orchestration platforms (Kubernetes, ECS).
- Identity & Security: Experience managing enterprise directory services, IAM, and SSO protocols (Active Directory, Entra ID, Okta, SAML, LDAP).
- Disaster Recovery: Direct experience designing, executing, and auditing multi-site backup strategies and failover simulations (e.g., Veeam, Zerto, AWS Backup).
- Experience with High Performance Compute (HPE, Supermicro)
- Experiene with Enterprise Storage (NetApp)
Work Environment
- This role is based in San Diego, CA.
- We welcome candidates who are local or open to relocating; relocation assistance is available and may be included in the offer package where appropriate.
Compensation
Benefits & Perks
- We offer comprehensive medical, dental, and visions plans
- 401(k) Retirement Savings Plan to invest in your long-term retirement goals
- Equity grants for new hires
- Unlimited PTO
- Extremely generous company holiday calendar, including a holiday hiatus in July & December.
- Generous Parental Leave
- Lifestyle Spending Account
- FSA
- DCFSA
- HSA
- Hospital Indemnity insurance
- Critical Illness insurance
- Accident insurance
- Basic Life/AD&D, short-term and long-term disability insurance, 100% covered by Firestorm. Plus, the option to purchase additional life insurance for you and your family.
- Mental Health Resources: We provide free mental health resources 24/7 including therapy and more. Additional work-life services, such as free legal and financial support, are available to you as well.
Export Control Compliance
Equal Opportunity Statement
Skills
- Active Directory
- Ansible
- Automation
- AWS
- Azure
- Bash
- Cloud
- CloudWatch
- Cmmc
- Containerization
- Datadog
- DHCP
- DNS
- Docker
- ECS
- ELK
- Entra ID
- Firewall
- GCP
- Grafana
- Hyper-V
- IAM
- Infrastructure as Code
- Itar
- Kubernetes
- LDAP
- Linux
- Networking
- Nist
- Observability
- Okta
- PowerShell
- Prometheus
- Puppet
- Python
- RHEL
- SAML
- Splunk
- SSH
- SSO
- TCP/IP
- Terraform
- Ubuntu
- Virtualization
- VLAN
- VMware
- Windows Server
