Evaluation Scenario Writer - AI Agent Testing Specialist
Mindrift
وصف الوظيفة
About the role
Mindrift is looking for specialists to design and execute evaluation scenarios for AI agents. You will work on short‑term, paid projects that help improve the reliability and safety of large language models.
Key responsibilities
- Create structured test cases that simulate complex human workflows.
- Define gold‑standard behavior and scoring logic to assess agent actions.
- Analyze agent logs, failure modes, and decision paths.
- Work with code repositories and test frameworks to validate scenarios.
- Iterate on prompts, instructions, and test cases to improve clarity and difficulty.
- Ensure scenarios are production‑ready, easy to run, and reusable.
Required profile
- Minimum 3 years of software development experience with a strong focus on Python.
- Proficiency with Git and code repositories.
- Comfortable using structured data formats such as JSON and YAML.
- Understanding of core LLM limitations (hallucinations, bias, context limits).
- Familiarity with Docker.
- English proficiency at least B2 level.
Required skills
- Python
- Git
- JSON
- YAML
- Docker
What we offer
- Project‑based paid work, rates up to $40 per hour.
- Flexibility to choose when and how to work.
- Opportunity to contribute to cutting‑edge AI evaluation efforts.
Questions fréquentes
لماذا تبلغ عن هذا العرض؟
اكتشف المزيد
الرواتب والأدلة وعمليات البحث في المملكة العربية السعودية.
الرواتب حسب المهنة
قدم طلبك في 30 ثانية
أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.
بالمتابعة، أنت توافق على شروط الاستخدام.
لديك حساب بالفعل؟ تسجيل الدخول
لديك سؤال حول هذا العرض؟
اطرحه هنا: ستصلك تفاصيل العرض كاملة عبر البريد الإلكتروني، فوراً.
عزز فرصك
حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.
جاري تحليل سيرتك الذاتية...
Mindrift