Backend Software Engineer (SRE) - Cloud Infrastructure
Summary
Design and maintain scalable, high-availability cloud infrastructure and SRE practices for a large-scale platform using Kubernetes, Docker, and distributed databases.
Responsibilities
- Reliability: Ensuring the reliability and efficiency of our core infrastructure, focusing on system capacity and stability; setting up reliability standards and recovery SOP.
- Reliability: Troubleshooting and locating the technical issues, bottleneck analysis, managing system high availability architecture transformation and upgrading.
- Efficiency: Building automated operation solutions for large-scale systems; partnering with system development teams for system iteration.
- Efficiency: Designing and implementing software platforms and monitoring frameworks for efficient, automated, and intelligent service-oriented architecture (SOA) governance.
- Cost: We should build delivery standards, and monitor and budget systems to optimize the cost of the company.
- Compliance: Designing and setting up new IDC; designing and implementing data protection plan to meet the standard requirement.
Minimum Qualifications
- Bachelor's / Master's Degree in Computer Science or related major, with at least 5 years of relevant experience.
- Solid basic knowledge of computer software, understanding of Linux operating system, storage, network IO and other related principles.
- Familiar with one or more programming languages, such as Python, Go, and Java. Knowledge of design patterns and coding principles is necessary.
Preferred Qualifications
- Experience with storage, and relevant system experience with the following: KV, Table, Graph, Redis, MySQL, MongoDB, MQ, and Kafka.
- Experience with computing & big data, and system experience with the following: Kubernetes, Docker/Containers, AIops, Spark, Flink, Function as a service, RPC Framework, and Service Mesh.