DevOps / Site Reliability Engineer
Summary
The DevOps/SRE will build and maintain infrastructure for an AIOps monitoring platform, manage multi-cloud environments, and support cloud security operations. The role requires proficiency in Python and experience with cloud monitoring tools like Prometheus, Grafana, and ELK.
What you'd do:
- Build and maintain core infrastructure for the AIOps monitoring platform
- Maintain and optimize internal R&D infrastructure including GitLab and Nexus
- Manage monitoring data collection and alert governance across multi-cloud environments
- Support cloud security operations including alert management and compliance audits
What they want:
- 3+ years of DevOps or SRE experience with observability platform work a plus
- Proficient in Python, familiar with at least one of Go or Java
- Hands-on experience with Alibaba Cloud or AWS and cloud monitoring tools
- Familiar with monitoring stacks such as Prometheus, Grafana, and ELK
Nice to have:
- Full-stack capability with React or Vue frontend plus backend API
- Experience with AI/LLM application development such as RAG or agents
- Experience optimizing CI/CD toolchains like GitLab CI and container registries 3+ years of DevOps or SRE experience with observability platform work a plus
3+ years of DevOps or SRE experience with observability platform work a plus
Proficient in Python, familiar with at least one of Go or Java
Hands-on experience with Alibaba Cloud or AWS and cloud monitoring tools
Familiar with monitoring stacks such as Prometheus, Grafana, and ELK