This job is no longer available
This job expired on 16/08/2026. It no longer accepts applications.
Senior SRE Engineer (MLOps) – AI
Salla · La Mecque
Job description
About the role
Salla is seeking a Senior SRE Engineer (MLOps) to join its AI team. You will own the operational layer for AI and ML services, ensuring they run reliably, securely, and cost‑effectively at scale. The role bridges reliability engineering, platform development and close collaboration with engineering, data and AI groups.
Key responsibilities
- Define and maintain SLOs, dashboards, alerts, runbooks and incident follow‑ups for production ML and agentic AI services.
- Build end‑to‑end observability across the AI stack (latency, errors, traces, tool calls, cost, user impact).
- Design safe‑release patterns for models, prompts, agents and tools, including canary, rollback, feature‑flag and evaluation‑gate strategies.
- Provide operational support for inference APIs, queues, retrieval layers and AI workflows running on Kubernetes/EKS.
- Establish ownership, traceability and guardrails for agentic systems, defending against prompt injection and untrusted‑data risks.
- Drive AI cost governance with per‑model spend visibility, token‑cost tracking and anomaly alerting.
- Automate and create self‑service pathways so product teams can ship safely without rebuilding infrastructure each time.
- Turn recurring operational pain into reusable platform standards adopted by other teams.
- Participate in architecture discussions, code reviews and technical decision‑making.
Required profile
- 4+ years of experience in SRE, platform engineering, DevOps or production infrastructure for distributed systems.
- Hands‑on experience operating Kubernetes and cloud‑native systems in production.
- Familiarity with deploying machine‑learning projects and AI workloads.
- Strong command of CI/CD, GitOps, observability and incident response practices.
- Solid experience with infrastructure‑as‑code, secrets management and networking.
- Ability to write automation or platform tooling in Python or a similar language.
- Demonstrated production judgment – making systems measurable, debuggable, repeatable and safe to change.
- Excellent cross‑team communication skills and ability to explain trade‑offs clearly.
Required skills
- Kubernetes
- AWS/EKS
- Python
- CI/CD pipelines
- GitOps
- Observability tools
- Incident response
- Infrastructure‑as‑code (e.g., Terraform)
- Secrets management
- Networking
Questions fréquentes
Why are you reporting this job?
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Salla
La Mecque
Related job offers
-
Senior HPC Systems Administrator
KAUST (King Abdullah University of Science and Technology) La Mecque -
Communications Technician
Four Points by Sheraton La Mecque -
IT Infrastructure & Support Specialist
Majestic Office Furniture La Mecque -
Junior IT Engineer – Hospitality & F&B Support
Leylaty Group Riyadh -
Senior Frontend Engineer
Hudhud Maps Riyadh