هذه الوظيفة لم تعد متاحة
انتهت صلاحية هذه الوظيفة في 13/08/2026. لم تعد تقبل الطلبات.
Senior Infrastructure Engineer (HPC)
CONNECT Professional Services · Riyad
وصف الوظيفة
About the role
We are looking for a Senior Infrastructure Engineer to design, deploy and operate large‑scale High‑Performance Computing (HPC) clusters. The role combines deep Linux system administration, NVIDIA GPU infrastructure, workload scheduling and Kubernetes automation to enable cutting‑edge AI and scientific workloads.
Key responsibilities
- Design, implement and maintain end‑to‑end HPC clusters, including compute nodes, storage, InfiniBand/RoCE networking and management layers.
- Provision and manage NVIDIA Base Command Manager (BCM) for bare‑metal deployment, OS lifecycle and GPU fleet monitoring.
- Deploy and integrate NVIDIA AI Enterprise Suite, including NeMo, Triton, RAPIDS and NVIDIA NIM inference services.
- Operate NVIDIA GPU Operator and Network Operator in Kubernetes to automate driver, CUDA, DCGM exporter and MIG configuration.
- Install, configure and optimise Slurm workloads, partitions, QoS, fair‑share scheduling and MPI integration, including hybrid Slurm‑on‑Kubernetes scheduling.
- Build and maintain highly available Kubernetes clusters (kubeadm, etcd backup, zero‑downtime upgrades).
- Develop CI/CD pipelines with GitLab CI and GitHub Actions for infrastructure provisioning and software delivery.
- Implement monitoring with Prometheus, Grafana and DCGM, set up alerts and capacity planning.
- Enforce security hardening, kernel patching, RBAC and compliance across the HPC environment.
Required profile
- Bachelor's degree in Computer Science, IT, Computer Engineering or related field.
- Minimum 10 years of hands‑on HPC and infrastructure engineering experience.
- Active Red Hat Certified Engineer (RHCE) and Certified Kubernetes Administrator (CKA) certifications.
- Proven experience designing, deploying and managing large‑scale HPC environments.
Required skills
- Linux administration (RHEL, Ubuntu)
- NVIDIA GPU infrastructure, Base Command Manager, AI Enterprise Suite
- Slurm workload manager
- Kubernetes cluster deployment and administration (kubeadm, HA, upgrades)
- GPU Operator, Network Operator, NIM, Blueprint reference architectures
- CI/CD tools: GitLab CI, GitHub Actions
- High‑speed networking (InfiniBand, RoCE)
- Monitoring: Prometheus, Grafana, DCGM
- MPI integration, performance analysis and tuning
Questions fréquentes
لماذا تبلغ عن هذا العرض؟
اكتشف المزيد
الرواتب والأدلة وعمليات البحث في المملكة العربية السعودية.
الرواتب حسب المهنة
لديك سؤال حول هذا العرض؟
اطرحه هنا: ستصلك تفاصيل العرض كاملة عبر البريد الإلكتروني، فوراً.
عزز فرصك
حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.
جاري تحليل سيرتك الذاتية...
CONNECT Professional Services
Riyad
عروض عمل ذات صلة
-
Master & Reference Data Management Consultant
EJADA Riyad -
Data Privacy Consultant
Protiviti Middle East Member Firm Riyad -
Information Technology Internal Auditor
Trivers Riyad -
Expert Generative AI Engineer (n8n) – 12‑Month Contract
Nextwo Co. Gouvernorat d'Amman -
Game Developer & AR/VR Developer
AZCO CONSULTING & CAREER SERVICES Gouvernorat Tunis