📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Technical Lead - GPU Infrastructure

Jobgether

New Remote
Remote 🇬🇧 English
GPU infrastructure Slurm Kubernetes NVIDIA GPU Operator Network Operator NVIDIA drivers CUDA Fabric Manager NVSwitch DCGM MIG KubeVirt VFIO observability metrics logging alerting SLOs managed inference multi-node parallelism autoscaling confidential-compute

Job description

About the role

This hands‑on technical leadership position is responsible for architecting and delivering a large‑scale GPU infrastructure platform in Saudi Arabia. You will guide the transition from managed Kubernetes workloads to a bare‑metal GPU environment that supports research, model training, and managed inference.

Key responsibilities

  • Own end‑to‑end platform architecture, including high‑level and low‑level designs, technical reviews, and documentation.
  • Lead and line‑manage a distributed engineering team covering backend, frontend, DevOps, QA, and documentation.
  • Design, build, and operate a managed Slurm service for research and model‑training workloads.
  • Oversee GPU infrastructure operations on bare‑metal, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, and MIG.
  • Bootstrap and maintain Kubernetes clusters on bare metal using NVIDIA GPU Operator and Network Operator.
  • Define and implement managed inference architecture with multi‑GPU, multi‑node parallelism, autoscaling, and confidential‑compute capabilities.
  • Establish observability across control plane, GPU fleet, and applications through metrics, logging, alerting, and SLOs.
  • Lead incident response and post‑incident reviews.

Required profile

  • Proven experience in technical leadership and team management.
  • Strong background in large‑scale GPU compute platforms.
  • Ability to make deep technical decisions while overseeing delivery.
  • Excellent communication with infrastructure partners, hardware providers, and internal consumers.

Required skills

  • GPU infrastructure and bare‑metal provisioning.
  • Slurm workload manager.
  • Kubernetes, NVIDIA GPU Operator, Network Operator.
  • NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG.
  • KubeVirt, VFIO for GPU isolation.
  • Observability tools: metrics, logging, alerting, SLO definition.
  • Managed inference design, multi‑GPU and multi‑node scaling, autoscaling, request routing, confidential‑compute.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Jobgether.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in Saudi Arabia.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 1 day ago

Expires 1 month from now

8 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Jobgether