Server Engineer
SUPERPOWER X AI (SINGAPORE) TECHNOLOGY PTE. LTD. Server Engineer
Key Responsibilities
1. Server Deployment & Management
Install, configure, deploy, and maintain physical servers to ensure stable operation.
Monitor and optimize server resources, including CPU, GPU, memory, storage, and network performance.
Identify and resolve hardware and performance issues in a timely manner.
2. System Maintenance
Perform daily Linux server administration, system updates, patching, and kernel upgrades.
Develop and maintain automation tools to improve deployment, monitoring, and operational efficiency.
3. Troubleshooting & Incident Response
Diagnose and resolve server hardware, operating system, and infrastructure issues.
Respond to server incidents and minimize service interruption.
Conduct root cause analysis and prepare incident reports to improve system reliability.
4. Security & Compliance
Perform server security hardening, access control, log management, and security configuration.
Support vulnerability scanning, security monitoring, and incident response activities.
5. Documentation
Maintain server configurations, operation manuals, troubleshooting guides, and other technical documentation.
Follow and improve standard operating procedures for server operations and maintenance.
Requirements
Bachelor's degree or above in Computer Science, Network Engineering, Information Security, or related fields.
3+ years of server operation and maintenance experience; experience in cloud, internet, or data center environments is preferred.
Hands-on experience with GPU servers and hardware components such as memory, hard drives, and other server parts.
Strong knowledge of Linux systems, including CentOS, Ubuntu, or Red Hat, with experience in system tuning and performance optimization.
Familiar with automation tools and scripting languages such as Shell, Python, Ansible, or similar.
Familiar with Docker, Kubernetes, VMware, OpenStack, or other virtualization/container technologies.
Good understanding of storage technologies such as RAID, LVM, NFS, and iSCSI.
Strong troubleshooting, analytical, and incident response skills.
Good communication skills and ability to work with network, security, and other technical teams.
Experience with Prometheus, Grafana, Zabbix, GPU server operations, or HPC environments is an advantage.
RHCE, LPIC, MCSE, or other relevant certifications are preferred.