DevOps Engineer
Responsibilities
Design, deploy and maintain backend infrastructure to ensure high availability, scalability and reliability of production systems.
Manage cloud infrastructure on AWS and AliCloud, including performance optimization, cost management and operational excellence.
Deploy, monitor and maintain Kubernetes clusters, distributed backend systems and supporting infrastructure services.
Build and maintain CI/CD pipelines and infrastructure automation using GitHub Actions, Terraform and Ansible.
Collaborate with software engineers to support backend service deployment, troubleshooting and production operations.
Investigate production incidents, perform root cause analysis and implement preventive improvements.
Implement infrastructure security best practices to ensure secure and reliable production environments.
Develop internal DevOps platforms and automation tools to improve engineering productivity and operational efficiency.
Implement monitoring, observability and alerting solutions to enhance service reliability.
Research and integrate AI technologies into infrastructure operations, including intelligent alert analysis, ChatOps and operational automation.
Qualifications: 5+ years of hands-on experience in Kafka and Redis operations in large-scale production environments, be able to cooperate with developers to optimize code Proficient in Python / Go / Java (at least one language) and SQL programming languages Hands-on experience with containerization and orchestration (Docker, Kubernetes) Strong experience with CI/CD tools such as GitHub Actions, Ansible, Terraform etc At least 3 years of experience with AWS cloud platform. GCP, Azure, or Ali Cloud is a plus Excellent problem-solving and troubleshooting skills Strong team collaboration attitude and develop partnership with other teams and business Practical experience building or operating AIOps systems (anomaly detection, alert correlation, automated healing, or RCA) Familiarity with LLM-based DevOps automation (e.g., building chat-based ops assistants or AI-driven observability workflows) Experience using or integrating tools like Dify, Agno, or LangChain into operational workflows
Qualifications: 5+ years of hands-on experience in Kafka and Redis operations in large-scale production environments, be able to cooperate with developers to optimize code Proficient in Python / Go / Java (at least one language) and SQL programming languages Hands-on experience with containerization and orchestration (Docker, Kubernetes) Strong experience with CI/CD tools such as GitHub Actions, Ansible, Terraform etc At least 3 years of experience with AWS cloud platform. GCP, Azure, or Ali Cloud is a plus Excellent problem-solving and troubleshooting skills Strong team collaboration attitude and develop partnership with other teams and business Practical experience building or operating AIOps systems (anomaly detection, alert correlation, automated healing, or RCA) Familiarity with LLM-based DevOps automation (e.g., building chat-based ops assistants or AI-driven observability workflows) Experience using or integrating tools like Dify, Agno, or LangChain into operational workflows