Requirements
Requires 3+ years of experience in open-web research, including locating primary records and navigating government and institutional databases. Must have precision with sourcing, citing exact pages, tables, and sections. Native or near-native written English and high tolerance for structured documentation are essential. Experience with LLM evaluation, red-teaming, or benchmark construction is preferred.
Description
Employment Type: ContractorStart Date: ImmediatelyEngagement Length: 8 weeksYears of Experience: 3+ years of experienceCommitment Required: Full-time. 40 hours per week with at least 4 hours PST overlapPay Rate: (Average Handling Time per Task: 3 hours) Pay per task, $30/task, effective $10/hrAbout the ClientOur client is one of the world's fastest-growing AI companies, accelerating the advancementand deployment of powerful AI systems.It helps customers in two ways: Working with the world's leading AI labs to advance frontiermodel capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality,STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solvemission-critical priorities for companies.About the RoleWe are building an evaluation benchmark for frontier AI browsing agents. Your job is to designresearch problems that a state-of-the-art AI cannot solve, even with full web access and multipleattempts. This is not a subject matter expert role, nor is it a content-writing role. It isinvestigative research.You will start from a verifiable fact, work backwards to construct a question that makes that factextremely hard to locate, and then prove your work with a complete, auditable evidence trail.What You Will ProduceA natural-language research question with a short, stable, objectively verifiable answerClues, each independently checkable, spanning multiple fact types - including dates,people, places, organisations, works, events, records, and quantities - with specificconstraintsA validation record showing the obvious searches you ran and the results theyreturnedMinimum QualificationsDemonstrated open-web research ability, including locating primary records andnavigating government and institutional databases, archives, registries, and PDFdocumentsPrecision with sourcing. You cite exact pages, tables, and sections—not justhomepagesComfort researching unfamiliar subjects from scratchNative or near-native written EnglishHigh tolerance for structured documentation. The evidence trail is the majority of theworkExperience with LLM evaluation, red-teaming, or benchmark constructionExperience in one or more of the following domains: 1. Reference librarianship, archival research, or special collections 2. Investigative journalism or professional fact-checking 3. OSINT, due diligence, KYC, or investigative research 4. Patent, prior-art, or legal-discovery search 5. Genealogy and records research 6. Competitive quizzing or puzzle-hunt constructionNice to HaveExperience with LLM evaluation, red-teaming, or benchmark constructionFamiliarity with JSON and structured data delivery formats