PriorBench Technical Report
Introducing PriorBench: A New Frontier in Search Evaluation
Nicholas HopeAvirath SundaresanLiam McDonald
Talarion ·
Abstract
Traditional search benchmarks measure the accuracy of responses to narrowly scoped, objective questions. Our approach is motivated by the difficulty of evaluating response quality on challenging, open-ended questions directly. PriorBench is a search benchmark consisting of 100 research questions (RQs) and 1,782 associated evaluation questions (EQs) spanning events from September 1 2025 to May 1 2026.1 Search engines were evaluated by assigning an RQ to a research LLM (gpt-5.4 or gpt-5.6), allowing it to perform agentic search, and then evaluating the quality of its understanding by eliciting a forecast on each EQ. Evaluation questionnaires averaging 17.8 ± 4.3 EQs were produced for each RQ to capture a necessary (although possibly insufficient2) set of priors upon which any credible response to the RQ should be predicated. Crucially, forecasts were produced while holding agentic search results in context, but no additional search was permitted after EQs were revealed. Thus PriorBench is designed explicitly to quantify an LLM's latent understanding of world-state after finishing a research task. We measure context sufficiency as a proxy for response quality conditioned on fixed inferential ability.
Background
Existing benchmarks for agentic web search generally evaluate an agent's ability to recover specified information from the web. Prominent examples include BrowseComp, which evaluates retrieval of a single factual answer using a set of entangled clues; DeepSearchQA, which requires agents to follow chains of dependent searches and to judge when the resulting answer set is complete; and WideSearch, where agents must exhaustively collect and structure large volumes of individually easy-to-find facts. Benchmarks of temporal freshness such as FreshQA and RealTimeQA likewise pose the question before search begins, measuring correctness on current world knowledge.
PriorBench instead evaluates the informational state induced by search. The research LLM and agent scaffold are held fixed across search systems, allowing us to compare them under a common research policy. Evaluation questions are withheld until the research process is complete, and no further search is permitted once they are revealed. The score therefore reflects the sufficiency of the context accumulated through search, rather than an agent's ability to retrieve information it was explicitly asked to retrieve.
Results
Talarion robustly improves world-state understanding relative to existing search solutions. Exa, Perplexity and Brave perform similarly, with comparable scaling behavior as the number of search turns is increased. For additional results including Parallel and Tavily, see Appendices A and B. We attribute Talarion's performance to a combination of proprietary retrieval technology and a surprisal-aware, deduplicated corpus of information designed specifically to mitigate LLM calibration errors. We view these results as a clear indication that the raw contents of the open web are ill-suited to grounding LLMs in current world-state, necessitating a novel approach.
Generating Evaluation Questions
Generating fair and distributionally representative evaluation questions is a challenging problem. We considered three possible approaches, of which we implement two.
- Take a student-teacher approach, in which agentic search is performed by a student model with a knowledge cutoff that predates the evaluation window, and EQs are enumerated by a teacher model with a knowledge cutoff that postdates the evaluation window.
- Have human experts with up-to-date RQ-specific domain knowledge enumerate EQs that they consider to be necessary priors.
- Allow a search agent to enumerate EQs.
The primary results (displayed above) were produced using approach (1). We tasked claude-opus-5 with producing EQs that are relevant, important, and genuinely uncertain as of the gpt-5.4 knowledge cutoff of August 31, 2025. This has the advantage of sampling EQs from a plausible construction of all possible knowledge (that is, the claude-opus-5 pre-training corpus). However, any such evaluation must be run retroactively, compromising measurements of live performance and limiting the selection of search providers to those with robust infrastructure for timestamping results. In order to perform live measurements, we additionally tested two distinct implementations of approach (3), with results reported in Appendices A and B.
True Negatives
The above results include both positively and negatively resolved EQs. However, true negatives (evaluation questions that resolve to No) can present several subtle problems. For a representative example, consider the April 28th 2026 announcement by the UAE of their imminent departure from OPEC. Consider the following three possible true negatives.
- Did the UAE announce their imminent departure from OPEC on April 29th?
- Did the UAE not announce their imminent departure from OPEC on April 28th?
- Did the UAE and Saudi Arabia issue a joint pledge of solidarity with fellow OPEC member states on April 12th?
Question (1) is essentially mutually exclusive with a sufficiently precise snapshot of true world state (if you're certain that the UAE announced their departure on the 28th specifically, you can be confident that they did not announce their departure on the 29th). However, it is positively correlated to a fuzzier understanding of true world state. An LLM that understands that the UAE is considering departing OPEC but isn't certain as to the date of their announcement will assign a higher forecast to (1) than an LLM with no knowledge whatsoever of the underlying situation. Thus we consider (1) undesirable, since it will in many instances penalize better-informed models.
Question (2) is the complement of the true statement. Prediction market traders will be familiar with the core problem here: nothing ever happens. It is relatively straightforward to compile high-surprisal events; it is far harder to compile high-surprisal non-events. This leads to a preponderance of true positives of the form “X happened” and true negatives of the form “X did not happen”. Any such pattern defeats the purpose of including true negatives, since the forecasting LLM is once again capable (in theory although not in practice) of gaining an unfair advantage based solely on question phrasing and a predisposition towards “things happen”.
Question (3) we understand to be negatively correlated to latent world state. We therefore believe that (3) represents by far the best true negative EQ. Defining latent world state objectively is exceedingly difficult; we nonetheless steer EQ generation towards oblique contradiction rather than trivial variation or explicit negation wherever possible.
Sample Research Questions and Evaluation Questions
Each evaluation question is posed as a binary forecasting question.
Would Putin support a Chinese invasion of Taiwan?
- TrueDid China suspend or restrict imports of Japanese seafood after November 2025 amid the dispute with Tokyo?0.210.220.180.240.55
- TrueDid Vladimir Putin travel to Beijing to attend China's September 3, 2025 military parade marking the 80th anniversary of the end of World War II?0.800.810.850.990.89
- TrueBy May 1, 2026, had China imposed formal restrictions on rare earth exports that were subsequently suspended or paused as part of a deal with the United States?0.240.520.520.630.61
- FalseBy May 1, 2026, had Ukraine agreed to a US-brokered framework that included recognizing Russian control of all of Donetsk oblast?0.130.110.180.060.28
- FalseBy May 1, 2026, had Russia and Ukraine signed a comprehensive peace treaty formally ending the war?0.070.070.080.030.03
- TrueWas Sanae Takaichi serving as Prime Minister of Japan on May 1, 2026?0.180.070.200.100.95
- FalseDid Russia and Ukraine complete a ceasefire covering the entire front line that was still holding on May 1, 2026?0.190.280.420.240.12
- TrueDid the United States conduct military strikes on targets inside Venezuela before May 1, 2026?0.050.040.050.040.80
- TrueBy May 1, 2026, had China announced a defense budget increase of about 7 percent or more for 2026?0.780.880.820.860.91
- FalseDid Russia formally withdraw from the New START treaty framework and announce resumption of nuclear weapons testing with an actual nuclear explosive test conducted before May 1, 2026?0.040.040.030.020.06
- TrueDid the United States announce a new National Defense Strategy or equivalent strategy document prioritizing homeland and Western Hemisphere defense over Europe before May 1, 2026?0.310.430.500.470.88
- FalseBy May 1, 2026, had Russia formally recognized any Chinese claim by publicly endorsing the use of military force to unify Taiwan with the mainland?0.120.330.380.180.18
- TrueDid Xi Jinping and Donald Trump meet face-to-face in South Korea in late October 2025?0.270.580.680.720.80
- TrueBy May 1, 2026, had the United States and China agreed to a trade truce that included China resuming large-scale purchases of American soybeans?0.290.630.580.680.62
- FalseBy May 1, 2026, had the People's Republic of China initiated a military invasion of Taiwan?0.080.050.080.020.08
- TrueBy May 1, 2026, had Japan's Takaichi government faced a major diplomatic rupture with China over remarks by the Japanese prime minister about a Taiwan contingency?0.090.100.120.090.76
- TrueDid Taiwan's President Lai Ching-te remain in office through May 1, 2026?0.940.960.950.970.97
- FalseBy May 1, 2026, had the United States and Russia restored full ambassadorial-level diplomatic relations with a new US ambassador confirmed and posted in Moscow?0.220.160.420.280.38
- TrueDid Nicolás Maduro cease to be the de facto leader of Venezuela before May 1, 2026?0.220.170.140.140.24
- FalseBy May 1, 2026, had Russia signed a binding final agreement with China to construct the Power of Siberia 2 pipeline, including agreed pricing terms?0.360.240.220.220.34
- FalseDid Vladimir Putin and Donald Trump hold an in-person summit meeting in Budapest before May 1, 2026?0.090.190.120.260.40
- TrueDid the United States impose new sanctions on the Russian oil companies Rosneft and Lukoil after August 31, 2025?0.310.360.350.270.66
A simple search-assisted generation implementation: 150 handwritten research questions with 4,154 EQs produced by a research agent, all phrased as true positives. The search traces used for EQ generation were produced by a search agent performing 20 rounds of search via Parallel and 10 rounds of search via Talarion.
This appendix regenerates EQs with tighter provenance controls. For each RQ, an EQ generation agent produced up to 150 searches, each of which was routed uniformly at random to one of six search systems (Exa, Parallel, Tavily, Brave, Perplexity, and Talarion). For the results below, EQs were filtered by provenance: only the 478 EQs (78% YES) for which a citation could be found under both Talarion and a standard web search engine were included, so every retrieval family is scored on facts it can, in principle, reach. Search traces used for evaluation were produced by gpt-5.6 (knowledge cutoff Feb 16 2026); EQs resolve as of July 27 2026.
Notes
- 1.This date range was chosen to span the intervening period between the gpt-5.4 and claude-opus-5 pretrains. For other date ranges, see Appendices A and B. ↩
- 2.That is, given the open-ended nature of the RQs, informational sufficiency is undefined. An LLM would almost certainly further benefit from additional priors not captured by the provided EQs. ↩