DevOps - Site Reliability Engineer
Summary
Builds and maintains CI/CD pipelines for web-based IoT applications and firmware, manages AWS infrastructure with infrastructure-as-code, and performs site reliability work such as monitoring, incident response, and security. Core stack: AWS, Terraform/CloudFormation, Docker/Kubernetes, Jenkins/GitHub Actions, Python/Bash, Ansible, Prometheus/Grafana/ELK.
This job is based in Des Moines, Iowa.
Essential Functions
1. Pipeline Engineering:
Design, develop, and maintain robust CI/CD pipelines to automate the build, test, and
deployment processes for both web-based IoT applications and firmware.
Collaborate with engineering teams to implement effective branching strategies, code
reviews, and automated testing frameworks.
Continuously optimize pipeline performance and reliability to accelerate delivery cycles.
2. Cloud Infrastructure Engineering:
Manage and enhance our AWS infrastructure, including provisioning, configuration, and
scaling of resources.
Automate routine tasks and implement infrastructure as code (IaC) practices to improve
efficiency and consistency.
Monitor system performance and proactively identify and resolve potential issues.
Implement robust security measures to protect our cloud environment.
3. Site Reliability Engineering
Ensure high availability, performance, and reliability of our applications and infrastructure.
Respond to incidents and outages promptly, implementing effective incident response
procedures.
Analyze system logs and metrics to identify potential issues and bottlenecks.
Collaborate with development teams to improve software quality and reduce failure rates.
4. Strong proficiency in scripting languages (Python, Bash, etc.) and configuration management
tools (Ansible, Puppet, Chef, etc.)
5. Experience with CI/CD tools (Jenkins, GitHub CI/CD, GitHub Workflows and Actions, etc.) and
version control systems (Git)
6. Deep understanding of cloud platforms, particularly AWS
7. Knowledge of containerization technologies (Docker, Kubernetes) and orchestration tools
8. Knowledge of containerization security best practices, tools and techniques.
9. Knowledge of storage administration – NFS/EFS.
10. Knowledge of disaster recovery best practices and tools.
11. Knowledge of network security best practices and tools.
12. Experience with infrastructure change management best practices.
13. Experience with infrastructure as code (IaC) practices (Terraform, CloudFormation)
14. Solid understanding of networking concepts (TCP/IP, DNS, load balancing)
15. Familiarity with monitoring and logging tools (Prometheus, Grafana, ELK Stack)
16. Strong problem-solving and troubleshooting skills
17. A passion for automation and continuous improvement
Requirements
•
Education: BS degree in Computer Science, Software Engineering or relevant field – or equivalent experience.
Experience:
•
Minimum 5 years total experience in progressively responsible positions in the field of specialty.
•
Excellent analytical and time management skills
•
Teamwork skills with a problem-solving attitude
•
Experience working with other engineers to define constraints and requirements.
•
Professional and organizational skills are essential.
•
Ability to effectively communicate at all levels of the organization and externally (ie. customers, suppliers, auditors, etc.) through good written and verbal communication skills.
•
Ability to multitask and manage a variety of assigned projects.
•
Team building skills and the ability to foster an environment of cooperation and teamwork.
•
Ability to effectively and appropriately delegate workload.
•
Ability to prioritize tasks, direct the tasks of subordinates, and achieve objectives with minimal supervision.
•
Propensity to listen and approach new ideas and challenges with an open mind and evaluate situations on the merit of fact and data.