📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

This job is no longer available

This job expired on 16/08/2026. It no longer accepts applications.

Senior SRE Engineer (MLOps) – AI

Salla · La Mecque

Senior 🇬🇧 English
Kubernetes Python GitOps Incident response Infrastructure-as-code Secrets management Networking

Job description

About the role

Salla is seeking a Senior SRE Engineer (MLOps) to join its AI team. You will own the operational layer for AI and ML services, ensuring they run reliably, securely, and cost‑effectively at scale. The role bridges reliability engineering, platform development and close collaboration with engineering, data and AI groups.

Key responsibilities

  • Define and maintain SLOs, dashboards, alerts, runbooks and incident follow‑ups for production ML and agentic AI services.
  • Build end‑to‑end observability across the AI stack (latency, errors, traces, tool calls, cost, user impact).
  • Design safe‑release patterns for models, prompts, agents and tools, including canary, rollback, feature‑flag and evaluation‑gate strategies.
  • Provide operational support for inference APIs, queues, retrieval layers and AI workflows running on Kubernetes/EKS.
  • Establish ownership, traceability and guardrails for agentic systems, defending against prompt injection and untrusted‑data risks.
  • Drive AI cost governance with per‑model spend visibility, token‑cost tracking and anomaly alerting.
  • Automate and create self‑service pathways so product teams can ship safely without rebuilding infrastructure each time.
  • Turn recurring operational pain into reusable platform standards adopted by other teams.
  • Participate in architecture discussions, code reviews and technical decision‑making.

Required profile

  • 4+ years of experience in SRE, platform engineering, DevOps or production infrastructure for distributed systems.
  • Hands‑on experience operating Kubernetes and cloud‑native systems in production.
  • Familiarity with deploying machine‑learning projects and AI workloads.
  • Strong command of CI/CD, GitOps, observability and incident response practices.
  • Solid experience with infrastructure‑as‑code, secrets management and networking.
  • Ability to write automation or platform tooling in Python or a similar language.
  • Demonstrated production judgment – making systems measurable, debuggable, repeatable and safe to change.
  • Excellent cross‑team communication skills and ability to explain trade‑offs clearly.

Required skills

  • Kubernetes
  • AWS/EKS
  • Python
  • CI/CD pipelines
  • GitOps
  • Observability tools
  • Incident response
  • Infrastructure‑as‑code (e.g., Terraform)
  • Secrets management
  • Networking

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Salla.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in Saudi Arabia.

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 3 months ago

28 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Salla

La Mecque