📢 جديد: تابع عروض اليوم على قناتنا في واتساب
Jobiglo

لا توجد نتائج.

Technical Lead - GPU Infrastructure

Jobgether

جديد Remote
Remote 🇬🇧 English
GPU infrastructure Slurm Kubernetes NVIDIA GPU Operator Network Operator NVIDIA drivers CUDA Fabric Manager NVSwitch DCGM MIG KubeVirt VFIO observability metrics logging alerting SLOs managed inference multi-node parallelism autoscaling confidential-compute

وصف الوظيفة

About the role

This hands‑on technical leadership position is responsible for architecting and delivering a large‑scale GPU infrastructure platform in Saudi Arabia. You will guide the transition from managed Kubernetes workloads to a bare‑metal GPU environment that supports research, model training, and managed inference.

Key responsibilities

  • Own end‑to‑end platform architecture, including high‑level and low‑level designs, technical reviews, and documentation.
  • Lead and line‑manage a distributed engineering team covering backend, frontend, DevOps, QA, and documentation.
  • Design, build, and operate a managed Slurm service for research and model‑training workloads.
  • Oversee GPU infrastructure operations on bare‑metal, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, and MIG.
  • Bootstrap and maintain Kubernetes clusters on bare metal using NVIDIA GPU Operator and Network Operator.
  • Define and implement managed inference architecture with multi‑GPU, multi‑node parallelism, autoscaling, and confidential‑compute capabilities.
  • Establish observability across control plane, GPU fleet, and applications through metrics, logging, alerting, and SLOs.
  • Lead incident response and post‑incident reviews.

Required profile

  • Proven experience in technical leadership and team management.
  • Strong background in large‑scale GPU compute platforms.
  • Ability to make deep technical decisions while overseeing delivery.
  • Excellent communication with infrastructure partners, hardware providers, and internal consumers.

Required skills

  • GPU infrastructure and bare‑metal provisioning.
  • Slurm workload manager.
  • Kubernetes, NVIDIA GPU Operator, Network Operator.
  • NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG.
  • KubeVirt, VFIO for GPU isolation.
  • Observability tools: metrics, logging, alerting, SLO definition.
  • Managed inference design, multi‑GPU and multi‑node scaling, autoscaling, request routing, confidential‑compute.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Jobgether.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

لماذا تبلغ عن هذا العرض؟

شكراً لإبلاغك. سنراجع هذا العرض.

اكتشف المزيد

الرواتب والأدلة وعمليات البحث في المملكة العربية السعودية.

قدم طلبك في 30 ثانية

أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.

بالمتابعة، أنت توافق على شروط الاستخدام.

لديك حساب بالفعل؟ تسجيل الدخول

💬 راسلنا على تيليجرام الدردشة عبر واتساب

منشور منذ 4 أيام

ينتهي شهر من الآن

15 مشاهدات · 0 مهتم

عزز فرصك

حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.

جاري تحليل سيرتك الذاتية...

Jobgether