Senior DevOps Engineer
Who We Are:
Outlook Amusements operates a leading B2C consumer marketplace through its flagship brand, California Psychics. For over 30 years, the company has been dedicated to connecting individuals with trusted, expert guidance. Through personalized psychic readings and life guidance delivered by top-rated professional advisors, we help people gain clarity, confidence, and direction at meaningful moments in their lives.
What We’re Looking For:
We're looking for a Sr. DevOps Engineer to join our team! Reporting to the Director of TechOps, this role is responsible for the design, implementation, administration, and continuous improvement of secure, scalable, and highly available AWS-based infrastructure and supporting operational services. This role works closely with engineering, security, and operations teams to enhance deployment automation, strengthen system reliability, improve observability, and support efficient delivery of business-critical applications and services.
This position is instrumental in advancing infrastructure automation, Azure DevOps pipeline maturity, Datadog monitoring and alerting, cloud security, disaster recovery, and business continuity readiness. The Senior DevOps Engineer also helps ensure that infrastructure solutions are resilient, operationally effective, and aligned with business objectives, growth, and cost optimization goals.
What You’ll Do:
- Design, implement, and continuously improve secure, scalable, and highly available AWS-based cloud infrastructure and supporting services.
- Build, manage, and optimize Infrastructure as Code (IaC) solutions for AWS environments, including networking, compute, storage, and related services.
- Develop, maintain, and enhance Azure DevOps pipelines to improve deployment speed, consistency, reliability, and quality across development, test, and production environments.
- Automate infrastructure provisioning, configuration management, system maintenance, and operational workflows to reduce manual effort and improve operational efficiency.
- Monitor infrastructure, application health, logging, alerting, and service availability using Datadog; proactively identify trends, risks, and issues to minimize outages and improve service reliability.
- Partner with engineering, security, and operations teams to support application deployments, infrastructure changes, production readiness, and incident response.
- Implement and maintain disaster recovery, backup, resiliency, and business continuity capabilities to support operational readiness and recovery objectives.
- Strengthen cloud security by supporting identity and access management, patching, vulnerability remediation, system hardening, and adherence to established policies and standards.
- Evaluate existing infrastructure and operational practices and recommend improvements in scalability, performance, reliability, security, and cost optimization.
- Monitor and optimize AWS resource utilization and cloud spend while maintaining performance, availability, and scalability requirements.
- Perform capacity planning and forecast infrastructure resource requirements to support current and future business needs.
- Provide technical leadership, mentorship, and guidance through collaboration, documentation, knowledge sharing, and operational best practices.
- Participate in troubleshooting, root cause analysis, and post-incident remediation efforts to improve overall system stability and resilience.
- Work with third-party vendors and service providers as needed to support infrastructure, tooling, and service delivery objectives.
- Contribute to shared team and organizational objectives through strong cross-functional partnership, communication, and execution.