DevOps Engineer
Summary
DevOps engineer on the Cloud Service Desk (CSD) team, acting as the bridge between engineering and cloud infrastructure: running shift-based operational support, incident response, and deployment coordination across a hybrid AWS and Azure environment supporting Microsoft/.NET and Java application stacks.
- Act as the primary liaison between Engineering and Cloud Infrastructure teams, ensuring smooth deployment, support, and operational processes.
- Manage and support workloads across AWS and Azure cloud environments, including:
- AWS ECS, EKS, S3, CloudFront, EC2
- Azure IaaS, PaaS, and associated cloud services
- Support Microsoft (.NET) and Java application environments hosted across Windows and Linux platforms.
- Configure, maintain, and troubleshoot API Gateways, web servers, and application servers.
- Perform IAM access management, security policy administration, and lifecycle management activities.
- Investigate and resolve infrastructure, platform, application, and deployment issues across multi-cloud environments.
- Recommend and implement improvements to platform reliability, performance, security, and cost optimisation.
- Support CI/CD pipelines and change deployments within production and non-production environments.
- Participate in incident, change, problem, and release management activities.
- Primarily work within the Cloud Service Desk (CSD) shift structure, providing operational support for cloud-hosted services and applications.
- Act as an escalation point for 1st and 2nd Line Support teams on cloud, infrastructure, and application-related incidents.
- Monitor cloud services, application health, infrastructure alerts, and operational dashboards.
- Manage support tickets, requests, incidents, service restoration activities, and change records in accordance with ITIL best practices.
- Perform initial triage, root cause analysis, troubleshooting, and service recovery activities.
- Coordinate with Engineering, Infrastructure, Security, and Vendor teams during major incidents and production outages.
- Ensure adherence to SLA and KPI targets through timely resolution and effective communication.
- Participate in shift handovers and maintain operational documentation, runbooks, and knowledge base articles.
- Support maintenance activities, cloud platform upgrades, patching, and scheduled operational tasks.
- Contribute to continuous service improvement initiatives and operational automation.
- 3-5 years of experience in DevOps, Cloud Infrastructure, or Site Reliability roles or Cloud Operations roles
- Hands-on experience with AWS services (ECS, EKS, S3, CloudFront) and Azure
- Experience managing Windows-based EC2 instances alongside Linux environments
- Familiarity with both Microsoft/.NET and Java application stacks
- Experience with API gateways, web servers (e.g., IIS, Nginx, Apache), and application servers
- Working knowledge of IAM policies, access management, and lifecycle/retention policies
- Strong communication skills, comfortable working cross-functionally between engineering and infrastructure teams
- Experience with CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation, ARM/Bicep)
- Relevant AWS or Azure certifications
What they ask for
Required
- 3-5 years of experience in DevOps, Cloud Infrastructure, Site Reliability, or Cloud Operations roles
- Hands-on experience with AWS services (ECS, EKS, S3, CloudFront) and Azure
- Experience managing Windows-based EC2 instances alongside Linux environments
- Familiarity with both Microsoft/.NET and Java application stacks
- Experience with API gateways, web servers (e.g., IIS, Nginx, Apache), and application servers
- Working knowledge of IAM policies, access management, and lifecycle/retention policies
- Strong communication skills, comfortable working cross-functionally between engineering and infrastructure teams
- Experience with CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation, ARM/Bicep)
Preferred
- Relevant AWS or Azure certifications