Operational Resiliency Engineer
You will provide front-line operational support for account acquisition, compliance, AML, and treasury processes. You will triage incidents, apply and improve runbooks, implement observability, maintain process dependencies and documentation, manage event platforms, establish service targets, and provide operational feedback that improves resilience.
Responsibilities
- Manage incident response by acknowledging and triaging events
- Apply runbooks to restore service
- Design and implement metrics and monitors for distributed applications
- Improve runbooks for FCM processes
- Maintain the DAG for FCM process dependencies
- Implement monitoring and logging with observability tooling
- Maintain technical documentation for critical FCM processing paths
- Establish SLA, SLO, and SLI targets
- Manage Kafka and IBM MQ event platform systems
- Provide operational feedback to engineering teams
- Use AI tooling to identify gaps in monitoring, analysis, and documentation
Requirements
- Computer Science or Software Engineering education or equivalent practical experience
- 2+ years of distributed systems experience in a production environment
- AI tooling
- Security best practices for backend services
- Verbal communication
Benefits
- 401k with up to 3.5% company match
- 18 days of paid time off annually
- Seven paid holidays
- Hybrid schedule with remote work on Mondays and Fridays
- 20 additional flex remote days annually
- 5 company-wide office-optional weeks tied to major holidays
- 1 service day annually
- Paid parental bonding leave
- Health coverage
- Vision coverage
- Dental coverage
- Life insurance
- Disability insurance