Senior HPC Systems Administrator
KAUST (King Abdullah University of Science and Technology) · La Mecque
Job description
About the role
The KAUST Supercomputing Laboratory is looking for a highly motivated Senior HPC Systems Administrator to manage a large-scale HPC cluster of roughly 600 CPU and GPU nodes, storage systems, and high‑speed networks. The role provides critical support to researchers working on computational science, engineering, big data, and AI/ML workloads.
Key responsibilities
- Provide timely user support via phone, walk‑in, email and ticketing system while maintaining high service standards.
- Install, configure and manage compute nodes, storage systems, InfiniBand/Ethernet networks and configuration‑management tools such as Ansible or Puppet.
- Deploy and maintain cluster management software, monitoring tools and the Slurm workload manager, including QOS policies, accounts and automation scripts.
- Develop Bash and Python automation scripts to streamline administration tasks.
- Manage container environments (Singularity/Apptainer, Docker) for HPC workloads.
- Benchmark CPU, memory, InfiniBand and storage components and tune performance.
- Enforce security best practices, node hardening and kernel patching.
- Administer parallel file systems (Lustre, GPFS, Weka, Vast) with capacity planning and performance tuning.
- Collaborate with faculty, researchers and vendors to support research projects and drive technology evaluations.
- Produce user documentation, SOPs and training material for internal wiki.
Required profile
- Minimum Bachelor’s or Master’s degree in Computer Science, Engineering or related field.
- At least five years of experience supporting large‑scale HPC platforms and related subsystems.
- Proven track record managing complex HPC clusters, parallel file systems and job schedulers.
- Strong analytical, problem‑solving and decision‑making abilities.
- Excellent written and verbal communication skills in English.
Required skills
- Linux system administration (RHEL, Rocky Linux, CentOS).
- Slurm workload manager (or LSF/PBS).
- Configuration‑management tools (Ansible, Puppet).
- Scripting languages: Bash, Python, C++.
- HPC programming models: MPI, OpenMP, CUDA, OpenACC.
- Parallel file systems: Lustre, GPFS, Weka, Vast.
- Container technologies: Singularity/Apptainer, Docker (Kubernetes desirable).
- High‑speed networking: InfiniBand, Ethernet.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Saudi Arabia.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 9 hours ago
Expires 1 month from now
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
KAUST (King Abdullah University of Science and Technology)
La Mecque