Site Reliability Engineer
TestCrew | Quality Engineering & Software Testing · Al Ahsa
وصف الوظيفة
About the role
TestCrew is seeking Site Reliability Engineers (SREs) to join an upcoming enterprise engagement in Al Ahsa. The role is hands‑on and focuses on ensuring availability, reliability, performance and security of critical IT systems through proactive monitoring, automation and disciplined incident response. The position requires onsite work at the client location and participation in on‑call or shift rotations.
Key responsibilities
- Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
- Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively.
- Design and implement monitoring strategies, dashboards, alerts, and performance metrics.
- Perform root cause analysis for incidents and implement preventive measures.
- Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
- Manage CI/CD pipelines and support release management processes.
- Optimize system performance while ensuring compliance with SLAs, operational standards, and best practices.
- Support business continuity, disaster recovery, and operational resilience initiatives.
- Implement security best practices, including system hardening, patch management, and secure operational procedures.
- Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
- Maintain operational documentation, runbooks, knowledge articles, and incident reports.
- Participate in major incident management, post‑incident reviews, and continuous improvement initiatives.
- Provide on‑call support and participate in shift rotations as required.
Required profile
- Bachelor’s degree in Computer Science, Information Technology, Engineering or a related field.
- 2–5 years of hands‑on experience in Site Reliability Engineering, DevOps, Infrastructure Operations or Production Support.
- Experience managing enterprise monitoring and observability platforms such as Grafana, Prometheus, Datadog, Instana, Zabbix or equivalent.
- Strong troubleshooting and root cause analysis skills across infrastructure, applications and networking.
- Hands‑on experience with scripting and automation using Python, Bash or PowerShell.
- Experience with automation and Infrastructure as Code tools such as Ansible or Terraform.
- Practical knowledge of CI/CD tools such as Jenkins, GitLab CI or Azure DevOps.
- Good understanding of business continuity, disaster recovery and operational risk.
Required skills
- Python
- Bash
- PowerShell
- Grafana
- Prometheus
- Datadog
- Instana
- Zabbix
- Ansible
- Terraform
- Jenkins
- GitLab CI
- Azure DevOps
Questions fréquentes
لماذا تبلغ عن هذا العرض؟
اكتشف المزيد
الرواتب والأدلة وعمليات البحث في المملكة العربية السعودية.
الرواتب حسب المهنة
قدم طلبك في 30 ثانية
أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.
بالمتابعة، أنت توافق على شروط الاستخدام.
لديك حساب بالفعل؟ تسجيل الدخول
عزز فرصك
حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.
جاري تحليل سيرتك الذاتية...
TestCrew | Quality Engineering & Software Testing
Al Ahsa