Cloud Operations Engineer (IGT1)

Summary

Maintains and automates cloud infrastructure for a SaaS platform, ensuring uptime and performance while deploying and troubleshooting systems in AWS and Kubernetes.

  • Ensure platform availability, stability, and performance.
  • Investigate, troubleshoot, perform RCA for system/application issues; automate fixes for global rollout.
  • Orchestrate and deploy infrastructure and services in AWS using Infrastructure as Code.
  • Build automation tooling for deployments, OS updates, failure detection/remediation.
  • Implement and maintain deployment automation with Ansible; improve playbooks/roles and CI integration.
  • Participate in team projects; own technical details, implementation, and timelines.
  • Develop and maintain documentation and architecture diagrams.
  • Maintain operational/configuration procedures and documentation.
  • Repair and recover from HW/SW failures; coordinate and communicate with impacted teams.
  • Patch, upgrade, and harden systems regularly; assist with performance tuning and resource optimization.
  • Willingness to learn new technologies.
  • Participate in rotational on call; support Saturday deployments as required.

Mandatory Requirements:

  • Bachelor’s degree or equivalent experience in Computer Science, Information Systems, Engineering, or related field.
  • 2–5 years supporting Linux or Windows server operating systems, preferably for a SaaS organization.
  • Linux server administration experience with production workloads.
  • Kubernetes operations experience in production environments (cluster ops, deployments, observability, upgrades)
  • Microsoft Windows Server administration basics (install, configure, maintain, troubleshoot).
  • Hands-on experience with Amazon Web Services (AWS).
  • Experience with configuration management and deployment automation, strongly preferring Ansible.
  • Effective English communication, verbal and written.

Optional Requirements:

  • Scripting for automation (Python, shell, PowerShell).
  • Terraform and container tooling (Docker); familiarity with AWS-native services and CI/CD integration.
  • Experience diagnosing and troubleshooting large, complex environments.
  • Oracle RDBMS management — plus.
  • ITIL framework exposure or incident management experience — plus.

We champion flexibility and hybrid work options to support varying lifestyles and personal needs. At the same time, we value the power of in-person collaboration to build community, spark innovation, and strengthen connections. Our approach ensures you can work in ways that suit you best while still engaging with colleagues to share ideas and grow together. #LI-Hybrid #LI-DNP