📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

This job is no longer available

This job expired on 13/08/2026. It no longer accepts applications.

Senior Infrastructure Engineer (HPC)

CONNECT Professional Services · Riyad

Senior 🇬🇧 English
RHEL Ubuntu NVIDIA GPU Base Command Manager AI Enterprise Suite NeMo Triton RAPIDS NVIDIA NIM NVIDIA Blueprint Slurm Kubernetes GPU Operator Network Operator GitLab CI GitHub Actions InfiniBand RoCE Prometheus Grafana DCGM MPI

Job description

About the role

We are looking for a Senior Infrastructure Engineer to design, deploy and operate large‑scale High‑Performance Computing (HPC) clusters. The role combines deep Linux system administration, NVIDIA GPU infrastructure, workload scheduling and Kubernetes automation to enable cutting‑edge AI and scientific workloads.

Key responsibilities

  • Design, implement and maintain end‑to‑end HPC clusters, including compute nodes, storage, InfiniBand/RoCE networking and management layers.
  • Provision and manage NVIDIA Base Command Manager (BCM) for bare‑metal deployment, OS lifecycle and GPU fleet monitoring.
  • Deploy and integrate NVIDIA AI Enterprise Suite, including NeMo, Triton, RAPIDS and NVIDIA NIM inference services.
  • Operate NVIDIA GPU Operator and Network Operator in Kubernetes to automate driver, CUDA, DCGM exporter and MIG configuration.
  • Install, configure and optimise Slurm workloads, partitions, QoS, fair‑share scheduling and MPI integration, including hybrid Slurm‑on‑Kubernetes scheduling.
  • Build and maintain highly available Kubernetes clusters (kubeadm, etcd backup, zero‑downtime upgrades).
  • Develop CI/CD pipelines with GitLab CI and GitHub Actions for infrastructure provisioning and software delivery.
  • Implement monitoring with Prometheus, Grafana and DCGM, set up alerts and capacity planning.
  • Enforce security hardening, kernel patching, RBAC and compliance across the HPC environment.

Required profile

  • Bachelor's degree in Computer Science, IT, Computer Engineering or related field.
  • Minimum 10 years of hands‑on HPC and infrastructure engineering experience.
  • Active Red Hat Certified Engineer (RHCE) and Certified Kubernetes Administrator (CKA) certifications.
  • Proven experience designing, deploying and managing large‑scale HPC environments.

Required skills

  • Linux administration (RHEL, Ubuntu)
  • NVIDIA GPU infrastructure, Base Command Manager, AI Enterprise Suite
  • Slurm workload manager
  • Kubernetes cluster deployment and administration (kubeadm, HA, upgrades)
  • GPU Operator, Network Operator, NIM, Blueprint reference architectures
  • CI/CD tools: GitLab CI, GitHub Actions
  • High‑speed networking (InfiniBand, RoCE)
  • Monitoring: Prometheus, Grafana, DCGM
  • MPI integration, performance analysis and tuning

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec CONNECT Professional Services.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in Saudi Arabia.

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 3 months ago

24 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

CONNECT Professional Services

Riyad