HyperBrowseComp

A Multilingual and Multimodal Stress Test for Web-Browsing Agents

Alham Fikri Aji1, Faiz Rizki Ramadhan1, Zayd M. K. Zuhri2, Seung Hun Eddie Han3, Ryandito Diandaru1, Qinrong Cui1, Jan Christian Blaise Cruz1, Badrinath Chandana1, Peerawat Chomphooyod4, Ahmed Attia1, Jonibek Mansurov1, Emilio Villa-Cueva1, Canh Duong Nguyen1, Imran Turganov1, Minghao Wu5, Peerat Limkonchotiwat6, Irina Nikishina1

1Mohamed bin Zayed University of Artificial Intelligence 2Mila – Quebec Artificial Intelligence Institute 3Inception AI 4Chulalongkorn University 5Alibaba Group 6AI Singapore

HyperBrowseComp stress-tests whether web-browsing agents can find and verify obscure answers on the open web. Its deliberately difficult questions are authored in their source languages; many require multi-step searches that connect evidence across webpages, videos, maps, images, audio, and documents.

… Human-authored questions
… Source languages
Multimodal Video, maps, documents, and more

Examples

Example question

Loading examples…

Answer —
Initial German example evidence screenshot

Evidence

Loading evidence…

How the questions are built

Native or highly proficient speakers write each question from a publicly verifiable answer and its supporting sources. A second annotator checks the answer, the evidence, and whether the question is unambiguous. Candidate questions are also tested without internet access, and easier ones are excluded.

Languages

Questions by source language

Primary domains

One domain per question

Primary
domains

Hover or focus on a slice to inspect it.

Evidence and task types

Questions can belong to several categories

Baseline accuracy

Native search leaves most questions unanswered. Exa improves the GPT-5.6 Sol result but lowers both Gemini results. The OWL run also scores below Gemini's native-search result. Exa uses substantially more tokens than native search for each matched model.

Accuracy by browsing setting

Accuracy versus token use

Select a point to see exact values.

Totals include research and grading. OWL combines workforce traces with separate grading calls.

Citation

@misc{aji2026hyperbrowsecomp,
  title  = {HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents},
  author = {Aji, Alham Fikri and Ramadhan, Faiz Rizki and Zuhri, Zayd M. K. and Han, Seung Hun Eddie and Diandaru, Ryandito and Cui, Qinrong and Cruz, Jan Christian Blaise and Chandana, Badrinath and Chomphooyod, Peerawat and Attia, Ahmed and Mansurov, Jonibek and Villa-Cueva, Emilio and Nguyen, Canh Duong and Turganov, Imran and Wu, Minghao and Limkonchotiwat, Peerat and Nikishina, Irina},
  year   = {2026},
  eprint = {2610.03574},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url    = {https://arxiv.org/abs/2610.03574}
}