Senior Quality Assurance Automation Engineer

🏢 GCS
📍 United Arab EmiratesFull-timeRemote
📅 Posted: 1w ago🔄 Updated: 1w ago
CV%
✨ AI Summary
The Senior QA Automation Engineer, AI & Agentic Systems role involves owning quality for AI and agentic systems in a UAE-based Sovereign AI company. Responsibilities include end-to-end testing of agentic systems (tool-use, grounding, guardrails), building evaluation for non-deterministic AI outputs, and managing prompt/model regression suites. The role also covers integration, API, and performance testing. This is a fully remote, contract-based position with an immediate start.
Required Skills
Other
DeepEvalLLM-as-judgeRAGASLangSmithcustom harnessesLangfuseTruLens
Information Technology
Test AutomationObservabilityPythonOpenTelemetry
Finance, Legal & Governance
Valuation
Requirements
Requires 5+ years in QA or quality engineering, with at least 2 years specifically testing ML, LLM, or AI agent systems. Proficiency in Python (Pytest, custom harnesses), Selenium/Playwright, API testing, and LLM/agent evaluation frameworks (DeepEval, TruLens, RAGAS, custom Python evaluators, LLM-as-judge) is essential. Experience with agent observability and tracing tools like LangSmith, Langfuse, or OpenTelemetry is also required.
Description
Title: Senior QA Automation Engineer, AI & Agentic Systems (Remote) | Sovereign AILocation: RemoteAbout the role:We're hiring a Senior QA Automation Engineer for one of the leading players in Sovereign AI, a UAE-based company with frontier-model access and agentic AI products live across critical sectors including large-scale construction, government, healthcare, and defense.This is not traditional QA. You'll own quality for AI and agentic systems, defining what good looks like when a system reasons, calls tools, and can be wrong in subtle ways. If you've figured out how to test something whose output is different every time, this is for you.What you'll do:Test agentic systems end to end: tool-use and trajectory evaluation, grounding and citation verification, honest-failure and refusal calibration, guardrail and adversarial testing (prompt injection, jailbreaks)Build evaluation for non-deterministic AI: scoring for hallucination, consistency, and accuracy against gold datasets, including LLM-as-judge calibrated against human judgementOwn prompt and model regression suites that catch quality drift before it shipsRun continuous evaluation on live production traffic, with quality monitoring, alerting, and incident responseCover the fundamentals too: integration and API testing, secure gateway validation (RBAC, PII), and performance/load testingWhat we're looking for:5+ years in QA or quality engineering, with 2+ years testing ML, LLM, or AI agent systemsHands-on with LLM/agent evaluation frameworks (DeepEval, TruLens, RAGAS, or custom Python evaluators) and LLM-as-judgeAgent observability and tracing (LangSmith, Langfuse, OpenTelemetry)Expert Python (Pytest, custom harnesses), plus Selenium/Playwright and API testingThe setup:Fully remote, contract-basedOnboarding ASAPWorld-class environment: elite talent from FAANG and frontier AI labs
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00