Senior Linux & DevOps Engineer
Summary
Senior engineer who administers and troubleshoots Linux environments (on-prem servers and VMs), automates infrastructure with Ansible, and builds monitoring with Grafana and Prometheus (PromQL). Day to day also includes creating and maintaining Azure DevOps CI/CD pipelines, artifacts, and branch policies.
As a Senior Linux & DevOps Engineer, you will:
- Administer and support Linux environments, including troubleshooting CPU, memory, disk, and other system-level issues.
- Handle issues related to on-premises servers and virtual machines (VMs).
- Develop and maintain Ansible playbooks, roles, and modules for infrastructure automation.
- Create and maintain monitoring dashboards and alerts using Grafana.
- Work hands-on with Prometheus, including writing PromQL queries, configuring alerts, and understanding how metrics are collected and delivered to Prometheus.
- Create, maintain, and troubleshoot Azure DevOps pipelines.
- Work with Azure DevOps artifacts and branch policies.
- Support continuous integration and continuous delivery processes across development and infrastructure environments.
- Troubleshoot infrastructure and DevOps issues and drive them through to resolution.
- Collaborate with technical teams to maintain reliable, automated, and well-monitored environments.
- Proactively identify potential infrastructure, monitoring, and deployment issues and take corrective action.
What You Bring to the Table:
- 8–10 years of overall professional experience in Linux administration, DevOps, infrastructure engineering, or a closely related field.
- Strong hands-on experience with Linux administration.
- Proven experience troubleshooting CPU, memory, disk, and other Linux system issues.
- Experience handling on-premises server and VM-related issues.
- Strong hands-on experience with Ansible.
- Ability to write and maintain Ansible playbooks and roles.
- Good understanding of Ansible modules and module concepts.
- Hands-on experience with Grafana, including dashboard creation and alert configuration.
- Hands-on experience with Prometheus, including PromQL queries and alerting.
- Understanding of Prometheus metrics collection and data flow.
- Hands‑on knowledge of Azure DevOps.
- Experience creating and troubleshooting Azure DevOps pipelines.
- Knowledge of Azure DevOps Artifacts and branch policies.
- Familiarity with scripting using Bash and/or Python.
- Knowledge of RDBMS concepts, including indexes and keys.
- Strong critical‑thinking and problem‑solving skills.
- Proactive approach to infrastructure and operational responsibilities.
You should possess the ability to:
- Diagnose and resolve complex Linux system and infrastructure issues.
- Troubleshoot CPU, memory, disk, and other operating‑system‑level problems.
- Manage and troubleshoot on‑premises servers and virtual machines.
- Automate infrastructure and operational tasks using Ansible.
- Develop effective Ansible playbooks and reusable roles.
- Configure Grafana dashboards and alerts for infrastructure monitoring.
- Write effective PromQL queries and configure Prometheus alerts.
- Understand how application and infrastructure metrics are generated, collected, and exposed to Prometheus.
- Create, maintain, and troubleshoot Azure DevOps CI/CD pipelines.
- Work with Azure DevOps artifacts and enforce appropriate branch policies.
- Use Bash and/or Python scripting to automate repetitive operational tasks.
- Apply RDBMS fundamentals when troubleshooting or working with database‑dependent systems.
- Think critically when investigating technical problems and identify root causes.
- Work proactively to prevent recurring infrastructure and deployment issues.
- Communicate effectively with technical teams and take ownership of assigned issues and deliverables.
What we bring to the table:
- The opportunity to work across Linux infrastructure, DevOps, automation, and monitoring.
- Exposure to enterprise infrastructure spanning on‑premises servers, virtual machines, and cloud/DevOps environments.
- Hands‑on work with Ansible, Grafana, Prometheus, and Azure DevOps.
- Opportunities to contribute to infrastructure automation and CI/CD improvements.
- A collaborative technical environment focused on reliability, monitoring, automation, and continuous improvement.
- Opportunities to solve complex infrastructure and DevOps challenges while improving operational efficiency.
Let’s Connect
Want to discuss this opportunity in more detail? Feel free to reach out.