Senior Data Centre Operations Engineer
Skills
About the role
Senior Data Centre Operations Engineer
Location: Senai, Johor, Malaysia
Work Mode: Onsite
Employment type: Permanent
Our client is a leading specialist in the repair and maintenance of high-end AI computing infrastructure, with a state-of-the-art facility located in Johor Bahru, Malaysia. They are dedicated to providing mission-critical support and have established a reputation for precision and reliability in the Southeast Asian market. With a strong focus on transparency and accountability, they ensure that every repair process is documented and approved by clients, maintaining a high standard of service excellence.
We are seeking an experienced Senior Data Centre Operations Engineer to manage and support server and GPU infrastructure in large-scale environments, ensuring reliable AI and data centre operations.
Responsibilities
Set up, configure, and troubleshoot server hardware, including CPUs, memory, storage, RAID, NICs, and power supplies.
Monitor server health and review IPMI, BMC, and operating system logs.
Manage BIOS, BMC, iDRAC, iLO, and firmware upgrades.
Manage RAID configurations and monitor SSD and NVMe health.
Troubleshoot GPU servers and replace faulty hardware components.
Support GPU cluster performance, network topology, and system stability.
Collaborate with networking, storage, and virtualisation teams to resolve technical issues.
Automate routine tasks such as firmware upgrades, inspections, and hardware alert management.
Prepare technical guides, troubleshooting documentation, and standard operating procedures.
Coordinate with hardware vendors and manage RMAs and spare parts.
Monitor rack power, temperature, and air-cooling conditions.
Requirements
Must-have:
Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
Willingness to travel and support overtime, night shifts, on-call duties or weekend work when required.
At least 5-7 years of experience in server operations.
Experience using NVIDIA diagnostic tools, including NVIDIA-SMI and DCGM.
Strong Linux administration and troubleshooting knowledge.
Good hardware troubleshooting and problem-solving skills.
Ability to work effectively with internal teams and external vendors.
Good technical communication skills in English.
Nice-to-have:
Experience operating large-scale GPU clusters.
Experience with Shell or Ansible automation.
Familiarity with InfiniBand, RoCE, RDMA networking and optical modules.
Exposure to air-cooled server environments.
Experience with liquid-cooled servers.
Knowledge of Ceph, KVM, VMware or hyper-converged infrastructure.
Relevant certifications such as RHCE, RHCA, CompTIA Server+ or server vendor certifications.
Education:
Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
Why Join Us
Be part of a dynamic team at the forefront of data centre technology, supporting mission-critical AI infrastructure for leading enterprises.
Opportunity to work with advanced hardware, collaborate with skilled professionals, and contribute to the reliability of high-performance computing environments across the region.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.