freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer(Senior SRE)

Open 22d

Summary

Senior SRE responsible for maintaining and improving the reliability of Xiaomi’s overseas production systems through automation, observability, and incident response.

Job Responsibilities:

  • Ensure the stability, reliability, and high availability of the company’s overseas production environment, continuously improving system availability and service quality.
  • Manage resource provisioning, capacity planning, monitoring, change management, incident response, and daily operations to maintain business continuity and stability.
  • Review system architecture and technical solutions, identify potential risks, and implement mitigations to optimize system stability, performance, and resource efficiency.
  • Participate in on-call rotations, responding promptly to production incidents to safeguard business operations.
  • Build and enhance observability systems, including monitoring, logging, and distributed tracing, to improve monitoring capabilities and fault detection efficiency.
  • Develop and optimize automation platforms and engineering efficiency tools to advance operational automation and team delivery effectiveness.
  • Explore and promote the application of AI technologies in operations scenarios, leveraging AI tools to improve automation, fault analysis, knowledge management, and R&D efficiency.
  • Collaborate closely with R&D, product, security, and infrastructure teams to drive stability initiatives, implement best practices, and support ongoing business development.

Requirements:

  • Bachelor’s degree or above in Computer Science, Software Engineering, Information Technology, or a related field, with 5+ years of experience in Site Reliability Engineering (SRE), Computer Systems Administrator, Platform Engineer, or Cloud Engineer.
  • Proficiency in at least one programming language (e.g., Python, Go, Java, or C++), with strong software development and automation skills.
  • Familiarity with cloud computing services; experience with multi-cloud or hybrid cloud platforms (e.g., Alibaba Cloud, Azure, AWS, GCP) is a plus.
  • Solid understanding of Linux, computer networking, load balancing, distributed systems, and high-availability architectures.
  • Ability to quickly diagnose issues, communicate across teams, and drive solutions—developing system optimization and stability plans aligned with business goals, including dependency management, traffic governance, and disaster recovery planning.
  • Experience with system monitoring and observability tools (e.g., Prometheus, Grafana, ELK, or similar), along with scripting knowledge (Bash or Python) and familiarity with CI/CD concepts is preferred.
  • Familiarity with AI tools and their applications in software development, automated operations, or R&D efficiency—understanding of AI Agents or AIOps technologies; practical experience is a plus.
  • Strong communication, teamwork, and project management skills; ability to adapt to a fast-paced technical environment and continuously learn and apply new technologies.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available