HyperBrowseComp stress-tests whether web-browsing agents can find and verify obscure answers on the open web. Its deliberately difficult questions are authored in their source languages; many require multi-step searches that connect evidence across webpages, videos, maps, images, audio, and documents.
Examples
Loading examples…
How the questions are built
Native or highly proficient speakers write each question from a publicly verifiable answer and its supporting sources. A second annotator checks the answer, the evidence, and whether the question is unambiguous. Candidate questions are also tested without internet access, and easier ones are excluded.
Languages
Questions by source language
Primary domains
One domain per question
domains
Hover or focus on a slice to inspect it.
Evidence and task types
Questions can belong to several categories
Baseline accuracy
Native search leaves most questions unanswered. Exa improves the GPT-5.6 Sol result but lowers both Gemini results. The OWL run also scores below Gemini's native-search result. Exa uses substantially more tokens than native search for each matched model.
Accuracy by browsing setting
Accuracy versus token use
Select a point to see exact values.
Totals include research and grading. OWL combines workforce traces with separate grading calls.
Citation
@misc{aji2026hyperbrowsecomp,
title = {HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents},
author = {Aji, Alham Fikri and Ramadhan, Faiz Rizki and Zuhri, Zayd M. K. and Han, Seung Hun Eddie and Diandaru, Ryandito and Cui, Qinrong and Cruz, Jan Christian Blaise and Chandana, Badrinath and Chomphooyod, Peerawat and Attia, Ahmed and Mansurov, Jonibek and Villa-Cueva, Emilio and Nguyen, Canh Duong and Turganov, Imran and Wu, Minghao and Limkonchotiwat, Peerat and Nikishina, Irina},
year = {2026},
eprint = {2610.03574},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2610.03574}
}