Jobiglo

لا توجد نتائج.

هذه الوظيفة لم تعد متاحة

انتهت صلاحية هذه الوظيفة في 29/08/2026. لم تعد تقبل الطلبات.

Infrastructure & Site Reliability Engineer – Datacentre AI

Qualcomm · Riyad

Mid 🇬🇧 English
AI inference systems SRE fundamentals monitoring alerting incident response Prometheus Grafana CloudWatch automation MLOps tools

وصف الوظيفة

About the role

Qualcomm is expanding its data‑centre footprint in Riyadh and seeks an Infrastructure & Site Reliability Engineer to design, operate, and continuously improve large‑scale AI inference systems. You will ensure that AI workloads run reliably, at scale, and are production‑ready for advanced machine‑learning applications.

Key responsibilities

  • Design, deploy, and operate large‑scale AI inference systems supporting critical AI workloads.
  • Ensure reliability, availability, and scalability of Qualcomm AI clusters.
  • Develop and maintain software tools and supporting infrastructure for AI software stacks.
  • Collaborate with architecture and hardware teams to build components for LLM inference, agentic AI workflows, and AI services.
  • Identify and implement performance optimisations for multi‑SoC and multi‑card deployments.
  • Apply SRE fundamentals: monitoring, alerting, incident response, and performance tuning.
  • Support production ML systems using MLOps tools and best‑practice operational processes.
  • Build and maintain observability dashboards, alerts, and telemetry using Prometheus, Grafana, CloudWatch, and custom tools.
  • Create and update technical documentation, runbooks, and knowledge‑base articles.
  • Automate repetitive tasks and support CI/CD pipelines to improve system reliability.

Required profile

  • 2–8 years of experience in infrastructure, site reliability, or AI‑focused engineering.
  • Strong systems and software engineering fundamentals with hands‑on execution ability.
  • Capacity to work independently on complex problems while collaborating across hardware, software, and ML teams.

Required skills

  • AI inference systems and large‑scale AI cluster management.
  • SRE practices (monitoring, alerting, incident response, performance optimisation).
  • Observability tools: Prometheus, Grafana, CloudWatch.
  • Automation and CI/CD pipeline development.
  • MLOps tooling and production ML system support.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Qualcomm.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

لماذا تبلغ عن هذا العرض؟

شكراً لإبلاغك. سنراجع هذا العرض.

اكتشف المزيد

الرواتب والأدلة وعمليات البحث في المملكة العربية السعودية.

💬 راسلنا على تيليجرام الدردشة عبر واتساب

منشور منذ شهرين

19 مشاهدات · 0 مهتم

عزز فرصك

حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.

جاري تحليل سيرتك الذاتية...

Qualcomm

Riyad