Site Reliability Engineer
TestCrew | Quality Engineering & Software Testing · Al Ahsa
Job description
About the role
TestCrew is seeking Site Reliability Engineers (SREs) to join an upcoming enterprise engagement in Al Ahsa. The role is hands‑on and focuses on ensuring availability, reliability, performance and security of critical IT systems through proactive monitoring, automation and disciplined incident response. The position requires onsite work at the client location and participation in on‑call or shift rotations.
Key responsibilities
- Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
- Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively.
- Design and implement monitoring strategies, dashboards, alerts, and performance metrics.
- Perform root cause analysis for incidents and implement preventive measures.
- Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
- Manage CI/CD pipelines and support release management processes.
- Optimize system performance while ensuring compliance with SLAs, operational standards, and best practices.
- Support business continuity, disaster recovery, and operational resilience initiatives.
- Implement security best practices, including system hardening, patch management, and secure operational procedures.
- Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
- Maintain operational documentation, runbooks, knowledge articles, and incident reports.
- Participate in major incident management, post‑incident reviews, and continuous improvement initiatives.
- Provide on‑call support and participate in shift rotations as required.
Required profile
- Bachelor’s degree in Computer Science, Information Technology, Engineering or a related field.
- 2–5 years of hands‑on experience in Site Reliability Engineering, DevOps, Infrastructure Operations or Production Support.
- Experience managing enterprise monitoring and observability platforms such as Grafana, Prometheus, Datadog, Instana, Zabbix or equivalent.
- Strong troubleshooting and root cause analysis skills across infrastructure, applications and networking.
- Hands‑on experience with scripting and automation using Python, Bash or PowerShell.
- Experience with automation and Infrastructure as Code tools such as Ansible or Terraform.
- Practical knowledge of CI/CD tools such as Jenkins, GitLab CI or Azure DevOps.
- Good understanding of business continuity, disaster recovery and operational risk.
Required skills
- Python
- Bash
- PowerShell
- Grafana
- Prometheus
- Datadog
- Instana
- Zabbix
- Ansible
- Terraform
- Jenkins
- GitLab CI
- Azure DevOps
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Saudi Arabia.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 month ago
Expires 1 week from now
31 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
TestCrew | Quality Engineering & Software Testing
Al Ahsa
Related job offers
-
Machine Learning Engineer
شركة رؤى المتطورة للتدريب الصحي -
Software Developer
Al Watania Poultry - دواجن الوطنية Qassim -
IT Senior Security In-charge
Al Watania Poultry - دواجن الوطنية Qassim -
IT Development Division Head
Al Watania Poultry - دواجن الوطنية Qassim -
مهندس تنفيذ وتشغيل نظم أول
Employeur non precise