HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
Abstract
We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.
Community
We release a benchmark for stress-testing browsing capabilities of LLMs in multimodal, multilingual setting!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities (2026)
- Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents (2026)
- VisDocAgentBench: Benchmarking Agents for Visually Rich Document Retrieval (2026)
- Projecting BrowseComp-Plus onto ClimbMix: Toward More Realistic Corpora for Agentic Search (2026)
- DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents (2026)
- BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents (2026)
- SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.03574 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper