Site Reliability Engineer
TXSE is building the next-generation exchange infrastructure to support transparent, efficient, and resilient capital markets. With SEC approval and $275MM in funding, we are currently hiring a Site Reliability Engineer to help with a greenfield infrastructure build out.
Requirements
- Proficiency in automation and scripting tools (Ansible and Terraform are required as they are a large part of our environment).
- Proficiency supporting Linux Systems, general Linux Sys Administration, and Linux command line expertise.
- Monitoring & Incident Response: Monitor platform and infrastructure health, respond to incidents, troubleshoot service issues, and assist with root cause analysis to reduce downtime and operational risk. Exposure to Grafana, Prometheus or similar monitoring systems.
- Cloud Operations: Support workloads in AWS and Azure, including troubleshooting issues related to compute, networking, storage, access, and security in hybrid and multi-cloud environments.
- Enterprise Hardware & Storage Support: Support infrastructure running on Dell PowerEdge servers and assist in maintaining enterprise storage platforms, including performance monitoring, replication support, and capacity tracking.
- Experience with both on-prem and cloud infrastructure
- Experience with CI/CD and developer platform tools such as Jenkins, ArgoCD, GitHub Enterprise, SonarQube, Artifactory, or GitHub Actions runners.
Nice to have
- Red Hat OpenShift and/or Kubernetes in production or non-production environments.
- Multi-tenant cloud experience
- Systems/Software Development experience in Python
- Authentication services, DNS, and time synchronization concepts in distributed environments.
- Experience with enterprise server hardware and storage platforms is a plus
- Experience working in low latency environments, start-ups, prior experience within exchanges/trading firms
- Ability to context switch, and work highly effective in autonomous environments.