Site Reliability Engineer(Senior SRE)
Summary
Senior SRE responsible for maintaining and improving the reliability of Xiaomi’s overseas production systems through automation, observability, and incident response.
Job Responsibilities:
- Ensure the stability, reliability, and high availability of the company’s overseas production environment, continuously improving system availability and service quality.
- Manage resource provisioning, capacity planning, monitoring, change management, incident response, and daily operations to maintain business continuity and stability.
- Review system architecture and technical solutions, identify potential risks, and implement mitigations to optimize system stability, performance, and resource efficiency.
- Participate in on-call rotations, responding promptly to production incidents to safeguard business operations.
- Build and enhance observability systems, including monitoring, logging, and distributed tracing, to improve monitoring capabilities and fault detection efficiency.
- Develop and optimize automation platforms and engineering efficiency tools to advance operational automation and team delivery effectiveness.
- Explore and promote the application of AI technologies in operations scenarios, leveraging AI tools to improve automation, fault analysis, knowledge management, and R&D efficiency.
- Collaborate closely with R&D, product, security, and infrastructure teams to drive stability initiatives, implement best practices, and support ongoing business development.
Requirements:
- Bachelor’s degree or above in Computer Science, Software Engineering, Information Technology, or a related field, with 5+ years of experience in Site Reliability Engineering (SRE), Computer Systems Administrator, Platform Engineer, or Cloud Engineer.
- Proficiency in at least one programming language (e.g., Python, Go, Java, or C++), with strong software development and automation skills.
- Familiarity with cloud computing services; experience with multi-cloud or hybrid cloud platforms (e.g., Alibaba Cloud, Azure, AWS, GCP) is a plus.
- Solid understanding of Linux, computer networking, load balancing, distributed systems, and high-availability architectures.
- Ability to quickly diagnose issues, communicate across teams, and drive solutions—developing system optimization and stability plans aligned with business goals, including dependency management, traffic governance, and disaster recovery planning.
- Experience with system monitoring and observability tools (e.g., Prometheus, Grafana, ELK, or similar), along with scripting knowledge (Bash or Python) and familiarity with CI/CD concepts is preferred.
- Familiarity with AI tools and their applications in software development, automated operations, or R&D efficiency—understanding of AI Agents or AIOps technologies; practical experience is a plus.
- Strong communication, teamwork, and project management skills; ability to adapt to a fast-paced technical environment and continuously learn and apply new technologies.