Big Data Site Reliability Engineer(A158938)
Summary
Maintains and automates big-data infrastructure (HDFS, HBase, ClickHouse, etc.) to keep systems reliable and performant, with on-call support.
Implement and maintain reliable big data infrastructure systems (e.g., HDFS, HBase, ElasticSearch, Doris, RocksDB, ClickHouse, etc.)
Design and develop efficient tools to automate management of large-scale big data system.
Identify and resolve production issues related to performance and system failures.
Collaborate with application developers teams to troubleshoot and improve system reliability.
Participate in 24x7 on-call rotation and incident response.
Qualifications
Bachelor's or higher degree in Computer Science or related discipline.
Familiar with TCP/IP, data structures, algorithms and other protocols, and have good knowledge of operating systems, network, computer architecture.
Experience in at least one backend language (Python, Golang, Java, etc.).
Familiar with Unix/Linux operating systems and networking is preferred.
Strong communication skills to collaborate with cross‑functional teams.
Fluent in Chinese and English is preferred.