DevOps Engineer
Summary
DevOps Engineer responsible for designing CI/CD pipelines, automating infrastructure with IaC, and managing containerized workloads using Kubernetes across AWS, Azure, or GCP. The role focuses on improving platform reliability, implementing security best practices, and managing monitoring and observability tools.
Key Responsibilities
Design, implement, and manageCI/CD pipelines.
Automate infrastructure provisioning usingInfrastructure as Code (IaC).
Deploy and maintain applications across cloud and on-premises environments.
Manage containerized workloads usingDocker and Kubernetes.
Monitor system availability, performance, security, and capacity.
Identify and resolve infrastructure, application, and deployment issues.
Implement logging, alerting, backup, and disaster-recovery solutions.
Apply security best practices throughout the software development lifecycle.
Manage source-control strategies, release processes, and environment configurations.
Improve platform reliability, scalability, and deployment frequency.
Develop scripts and tools to reduce manual operational tasks.
Maintain technical documentation, runbooks, and architecture diagrams.
Participate in incident response, root-cause analysis, and on-call rotations.
Collaborate with cross-functional teams to promote DevOps practices and continuous improvement.
Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
Proven experience in DevOps, cloud engineering, site reliability engineering, or systems administration.
Experience with at least one major cloud platform:AWS, Microsoft Azure, or Google Cloud.
Strong knowledge of Linux administration and networking fundamentals.
Experience with CI/CD tools such asGitHub Actions, GitLab CI/CD, Jenkins, or Azure DevOps.
Proficiency with Infrastructure as Code tools such asTerraform, CloudFormation, or Pulumi.
Experience with configuration-management tools such asAnsible, Puppet, or Chef.
Hands-on knowledge of Docker and container-orchestration platforms.
Scripting experience usingPython, Bash, or PowerShell.
Familiarity with monitoring and observability tools such asPrometheus, Grafana, Datadog, Splunk, or the ELK Stack.
Understanding of secrets management, identity and access management, and cloud security principles.
Strong troubleshooting, communication, and collaboration skills.
Preferred Qualifications
Relevant cloud or Kubernetes certifications.
Experience managing production Kubernetes environments.
Familiarity with GitOps tools such as Argo CD or Flux.
Knowledge of service meshes, serverless platforms, and microservices architecture.
Experience with database operations, performance tuning, and disaster recovery.
Understanding of SRE concepts, including SLIs, SLOs, error budgets, and incident management.
Experience supporting high-availability and large-scale distributed systems.
Key Performance Indicators
Deployment frequency and lead time for changes
Application and infrastructure availability
Mean time to detect and recover from incidents
Change-failure rate
Percentage of infrastructure managed through automation
Security and compliance findings
Reduction in manual operational work
Skills
- Ansible
- Automation
- AWS
- Azure
- Azure DevOps
- Bash
- CI/CD
- Cloud
- Cloud Security
- CloudFormation
- Datadog
- DevOps
- Distributed Systems
- Docker
- ELK
- Flux
- GCP
- GitHub
- GitHub Actions
- GitLab
- GitOps
- Grafana
- Infrastructure as Code
- Jenkins
- Kubernetes
- Linux
- Microservices
- Networking
- Observability
- PowerShell
- Prometheus
- Pulumi
- Puppet
- Python
- Secrets Management
- Serverless
- Splunk
- Terraform