KAUST
Thuwal, Saudi ArabiaPosted yesterday
Senior HPC Systems Administrator Job Details | King Abdullah University of Science & Technology Skip to main content Home Professional Life at KAUST Living at KAUST Career Areas Academic and Research Jobs Innovation and Entrepreneurship Jobs Administration and Support Services Jobs Engineering and Facilities Management Jobs The KAUST School Jobs Elevate Program Join our Talent Community Search by Keyword Home Professional Life at KAUST Living at KAUST Career Areas Academic and Research Jobs Innovation and Entrepreneurship Jobs Administration and Support Services Jobs Engineering and Facilities Management Jobs The KAUST School Jobs Elevate Program Join our Talent Community View Profile Search by Keyword Select how often (in days) to receive an alert: Create Alert × Select how often (in days) to receive an alert: Senior HPC Systems Administrator Apply now » Date: Sep 30, 2026 Location: Saudi Arabia Company: King Abdullah University of Science & Technology Position Summary We are seeking a highly motivated and skilled Senior HPC Systems Administrator to join the KAUST Supercomputing Laboratory (KSL). The successful candidate will be responsible for managing an HPC cluster of approximately 600 CPU and GPU nodes, HPC storage systems, InfiniBand and Ethernet networks, and day-to-day operational issues. The role provides broad support to researchers and end-users across computational science, engineering, big data analysis, and artificial intelligence/machine learning workloads. Major Responsibilities – include but are not limited to - Provide timely and effective user support via telephone, walk-in, email, and ticketing system for all inquiry types while maintain high customer service standards. Install, configure, and manage HPC subsystems including compute nodes, high-performance storage systems, InfiniBand, Ethernet, and configuration management tools (e.g., Ansible, Puppet). Deploy and manage cluster management software, monitoring tools, and supporting services for operating HPC clusters. Install and administer the Slurm workload manager, manage QOS policies, accounts, accounting, and related automation scripts (Python and C++). Develop and maintain automation scripts in Bash and Python to streamline system administration tasks. Deploy and manage container environments (Singularity/Apptainer, Docker) for HPC workloads. Benchmark HPC system components like CPU, memory, InfiniBand, and storage periodically to ensure optimal performance and identify tuning opportunities across hardware, driver, and application layers. Enforce security best practices including node hardening, kernel patching, and compliance across all systems. Manage parallel file systems such as Lustre, GPFS, Weka, or Vast, including performance tuning and capacity planning. Directly support research activities in computational science, engineering, data analysis, and AI/ML by working closely with faculty, researchers, collaboration partners, and industrial partners in collaboration with application support teams. Develop software tools and utilities as needed to support research projects on cluster systems and subsystems. Drive proof-of-concept projects and technology evaluations end-to-end and research industry best practices and advocate system enhancements. Coordinate with vendors and third-party service providers to report and resolve issues in a timely manner. Develop and maintain user documentation, standard operating procedures, and training materials in the internal wiki. Stay at the forefront of HPC advancements through continuous learning, industry conferences, and professional collaboration, while driving benchmarking initiatives to inform future hardware procurement. Competencies Expertise in supporting users of computational science and engineering, data analysis, and artificial intelligence applications and libraries in different HPC environments. Strong expertise in Linux system administration (RHEL, Rocky Linux, or CentOS) in large-scale HPC environments. Proficiency with HPC applications and programming models (Fortran, C/C++, Python, MPI, OpenMP, CUDA, OpenACC). Demonstrated track record of managing complex HPC systems, including parallel file systems, job schedulers, InfiniBand/Ethernet networks, and monitoring systems. Experience with configuration management tools (Ansible, Puppet, or equivalent). Familiarity with computational science, data analysis, and AI/ML applications and libraries used in HPC environments. Knowledge of project management principles and practices. Demonstrated ability to support research activities in a highly collaborative HPC environment. Strong analytical, problem-solving, and decision-making skills. Proactively identifies and implements system improvements; takes initiative and sees tasks through to closure. Ability to manage multiple concurrent projects and deliver high-quality results within deadlines. Proven ability to collaborate cross-functionally with researchers, application teams, and ven
Bachelor's degree in Computer Science, Electrical Engineering, or a related field (or equivalent experience). Extensive experience administering large-scale HPC clusters (Linux, HPC storage, InfiniBand/Ethernet networks). Proficiency with Slurm workload manager, configuration management tools (Ansible, Puppet), container technologies (Singularity/Apptainer, Docker), and scripting (Python, Bash, C/C++). Experience with parallel file systems (Lustre, GPFS, Weka, Vast), job schedulers, and performance tuning. Strong analytical, problem-solving, and project management skills; ability to work collaboratively with researchers and external partners. Excellent communication and documentation skills.
Manage and operate an HPC cluster (~600 CPU/GPU nodes), HPC storage, and high-performance networks. Install, configure, and administer HPC subsystems; deploy cluster management and monitoring tools. Administer Slurm (including QOS, accounts, and automation). Develop automation scripts (Python, Bash). Manage containers for HPC workloads. Benchmark and optimize hardware, software, and middleware; apply security best practices and patch management. Manage parallel file systems; conduct performance tuning and capacity planning. Support researchers across computational science, engineering, AI/ML workloads; collaborate with application support teams. Develop software tools and utilities; drive POC projects and technology evaluations. Coordinate with vendors; maintain user documentation and SOPs. Stay updated on HPC advancements and participate in benchmarking for procurement decisions.
Not sure you fit this role?
Upload your CV and see how you score against KAUST and every other live job. It's free.
Get my free matchesAED 15k–22k a month· est.