Web Research Specialist - LLM Evaluation

🏢 Quikr
📍 EgyptFull-timeOn-site
📅 Posted: 1w ago🔄 Updated: 1w ago
CV%
✨ AI Summary
This role involves designing research problems for frontier AI browsing agents to evaluate their capabilities. The primary responsibility is to construct natural-language research questions that are difficult for state-of-the-art AI to solve, even with full web access. This requires investigative research skills, starting from a verifiable fact and working backward to create a challenging question, then providing a complete, auditable evidence trail. The output will include the research question, clues spanning multiple fact types with specific constraints, and a validation record of searches performed. Minimum qualifications include demonstrated open-web research ability, precision in sourcing, native or near-native written English, and a high tolerance for structured documentation. Experience with LLM evaluation, red-teaming, or benchmark construction is a plus. Preferred experience areas include reference librarianship, archival research, investigative journalism, OSINT, due diligence, patent search, genealogy research, and puzzle-hunt construction. Familiarity with JSON and structured data delivery formats is also considered a plus.
Required Skills
Information Technology
HiveAI Red Teaming
Other
structured data delivery formatsbenchmark constructiongovernment and institutional databasesstructured documentationRegistries
Soft Skills & Professional Competencies
Research
Finance, Legal & Governance
Valuation
Business, Sales & Management
Recruitment
Requirements
Requires 3+ years of experience in open-web research, including locating primary records and navigating government and institutional databases. Must have precision with sourcing, citing exact pages, tables, and sections. Native or near-native written English and high tolerance for structured documentation are essential. Experience with LLM evaluation, red-teaming, or benchmark construction is preferred.
Description
Employment Type: ContractorStart Date: ImmediatelyEngagement Length: 8 weeksYears of Experience: 3+ years of experienceCommitment Required: Full-time. 40 hours per week with at least 4 hours PST overlapPay Rate: (Average Handling Time per Task: 3 hours) Pay per task, $30/task, effective $10/hrAbout the ClientOur client is one of the world's fastest-growing AI companies, accelerating the advancementand deployment of powerful AI systems.It helps customers in two ways: Working with the world's leading AI labs to advance frontiermodel capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality,STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solvemission-critical priorities for companies.About the RoleWe are building an evaluation benchmark for frontier AI browsing agents. Your job is to designresearch problems that a state-of-the-art AI cannot solve, even with full web access and multipleattempts. This is not a subject matter expert role, nor is it a content-writing role. It isinvestigative research.You will start from a verifiable fact, work backwards to construct a question that makes that factextremely hard to locate, and then prove your work with a complete, auditable evidence trail.What You Will ProduceA natural-language research question with a short, stable, objectively verifiable answerClues, each independently checkable, spanning multiple fact types - including dates,people, places, organisations, works, events, records, and quantities - with specificconstraintsA validation record showing the obvious searches you ran and the results theyreturnedMinimum QualificationsDemonstrated open-web research ability, including locating primary records andnavigating government and institutional databases, archives, registries, and PDFdocumentsPrecision with sourcing. You cite exact pages, tables, and sections—not justhomepagesComfort researching unfamiliar subjects from scratchNative or near-native written EnglishHigh tolerance for structured documentation. The evidence trail is the majority of theworkExperience with LLM evaluation, red-teaming, or benchmark constructionExperience in one or more of the following domains: 1. Reference librarianship, archival research, or special collections 2. Investigative journalism or professional fact-checking 3. OSINT, due diligence, KYC, or investigative research 4. Patent, prior-art, or legal-discovery search 5. Genealogy and records research 6. Competitive quizzing or puzzle-hunt constructionNice to HaveExperience with LLM evaluation, red-teaming, or benchmark constructionFamiliarity with JSON and structured data delivery formats
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00