Jobiglo

No results.

Site Reliability Engineer

TestCrew | Quality Engineering & Software Testing · Al Ahsa

Mid 🇬🇧 English
Python Bash PowerShell Grafana Prometheus Datadog Instana Zabbix Ansible Terraform Jenkins GitLab CI Azure DevOps

Job description

About the role

TestCrew is seeking Site Reliability Engineers (SREs) to join an upcoming enterprise engagement in Al Ahsa. The role is hands‑on and focuses on ensuring availability, reliability, performance and security of critical IT systems through proactive monitoring, automation and disciplined incident response. The position requires onsite work at the client location and participation in on‑call or shift rotations.

Key responsibilities

  • Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
  • Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively.
  • Design and implement monitoring strategies, dashboards, alerts, and performance metrics.
  • Perform root cause analysis for incidents and implement preventive measures.
  • Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
  • Manage CI/CD pipelines and support release management processes.
  • Optimize system performance while ensuring compliance with SLAs, operational standards, and best practices.
  • Support business continuity, disaster recovery, and operational resilience initiatives.
  • Implement security best practices, including system hardening, patch management, and secure operational procedures.
  • Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
  • Maintain operational documentation, runbooks, knowledge articles, and incident reports.
  • Participate in major incident management, post‑incident reviews, and continuous improvement initiatives.
  • Provide on‑call support and participate in shift rotations as required.

Required profile

  • Bachelor’s degree in Computer Science, Information Technology, Engineering or a related field.
  • 2–5 years of hands‑on experience in Site Reliability Engineering, DevOps, Infrastructure Operations or Production Support.
  • Experience managing enterprise monitoring and observability platforms such as Grafana, Prometheus, Datadog, Instana, Zabbix or equivalent.
  • Strong troubleshooting and root cause analysis skills across infrastructure, applications and networking.
  • Hands‑on experience with scripting and automation using Python, Bash or PowerShell.
  • Experience with automation and Infrastructure as Code tools such as Ansible or Terraform.
  • Practical knowledge of CI/CD tools such as Jenkins, GitLab CI or Azure DevOps.
  • Good understanding of business continuity, disaster recovery and operational risk.

Required skills

  • Python
  • Bash
  • PowerShell
  • Grafana
  • Prometheus
  • Datadog
  • Instana
  • Zabbix
  • Ansible
  • Terraform
  • Jenkins
  • GitLab CI
  • Azure DevOps

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec TestCrew | Quality Engineering & Software Testing.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in Saudi Arabia.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 1 month ago

Expires 1 week from now

31 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

TestCrew | Quality Engineering & Software Testing

Al Ahsa