Big Data Site Reliability Engineer(A237998A)
Summary
Site Reliability Engineer maintaining Xiaomi's large-scale big data infrastructure (HDFS, HBase, ElasticSearch, Doris, RocksDB, ClickHouse) in Singapore. Day to day: automate system management with tooling, resolve performance and production failures, collaborate with app developers, and join a 24x7 on-call rotation. Backend coding in Python, Golang, or Java is expected.
Responsibilities
- 1. Implement and maintain reliable big data infrastructure systems (e.g., HDFS, HBase, ElasticSearch, Doris, RocksDB, ClickHouse, etc.)
- 2. Design and develop efficiency tools to automate manage large-scale big data system.
- 3. Identify and resolve production issues related to performance and system failures
- 4. Collaborate with application developers teams to trouble shooting and improve system reliability.
- 5. Participate in 24x7 on-call rotation and incident response
Qualifications
- 1. Bachelor''s or higher degree in Computer Science or related discipline.
- 2. Familiar with TCP/IP, Data Structures, Algorithms and other protocols, and have good knowledge of operation system, network, computer architecture.
- 3. Experience in at least one backend language (Python, Golang, Java etc) .
- 4. Familiar with Unix/Linux operating systems and networking is preferred.
- 5. Strong communication to collaborate with cross-functional teams.
- 6. Fluent in Chinese and English is preferred.