[LPS] AI Platform Engineer (Technical)
Summary
Build and optimize Kubernetes-based cluster platforms and intelligent scheduling systems for AI/ML workloads across global data centers.
We are Lenovo. We do what we say. We own what we do. We WOW our customers. Lenovo is a US$83 billion revenue global technology powerhouse, ranked #196 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full‑stack portfolio of AI‑enabled, AI‑ready, and AI‑optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world‑changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY). To find out more visit www.lenovo.com and read about the latest news via our StoryHub.
About The Job Duty
- Engineer Hyper-Scale Cluster Platforms: Enhance Kubernetes‑based cluster management systems to deliver superior performance, scalability, and resilience—supporting resource orchestration across client’s infrastructure.
- Advance Unified Scheduling: Design and maintain a comprehensive scheduling framework that supports diverse workloads, including containers, VMs, online services, offline computing, AI/ML, and CPU/GPU‑intensive tasks within massive‑scale resource pools.
- Develop Intelligent Scheduling Systems: Optimize workload performance and resource utilization across heterogeneous resources—CPU, GPU, memory, network, and power—spanning global data centers.
- Deliver Excellence and Innovation: Produce high‑quality, maintainable code while staying ahead of advancements in open‑source technologies, AI/ML research, distributed systems, and serverless computing.
Qualifications
- B.S./M.S., degree in Computer Science, Computer Engineering or a related area with 2+ years of relevant industry experience.
- Proven experience designing, architecting and building cloud and infrastructure related but not limited to resource management, allocation, job scheduling and monitoring.
- Familiarity with container and orchestration technologies such as Docker and Kubernetes.
- Proficiency in at least one major programming language such as Python, Go, C++, Rust, and Java.
- Proficient in spoken and written Cantonese/ Mandarin and English.
- Experience in one large scale cluster management systems, e.g., Kubernetes, Ray, Yarn, or Mesos, is highly desirable.
- Experience in large scale resource efficiency management and job scheduling development is highly desirable.
- Project experience in application scaling, workload co‑location, and isolation enhancement is highly desirable.
We are an Equal Opportunity Employer and do not discriminate against any employee or applicant for employment because of race, color, sex, age, religion, sexual orientation, gender identity, national origin, status as a veteran, and basis of disability or any federal, state, or local protected class.
