31 August 2026

You.com vs built-in web search

Key takeaways
Search added +34.2 to +76.4 points
You.com search increased question accuracy for all four models. Their accuracy range narrowed from 47.91 points without search to 5.62 points with the same retrieval layer.
You.com led by +3.66 points for GPT
You.com increased GPT-5.6 Terra's question accuracy by 3.66 points over built-in search. GPT also answered 3.40 seconds faster and cost $0.0051 less per question.
Wide search helped 2 of 4 models
Wide You.com search increased question accuracy by 3.2 points for Claude Sonnet 5 and 2.4 points for DeepSeek V4. Statistical tests could not distinguish the GLM-5.2 and GPT-5.6 Terra differences from zero.

LLMs have training cutoffs, so agents need web search to answer questions about current events. Search results vary by provider, retrieval depth, and model integration. Their effect on accuracy is harder to inspect because a model can issue several queries, select evidence, and combine facts before answering.

An agent can use a model provider's built-in search or a shared external search API. The search tool can also return a small ranked set or a wider set with more context. These choices affect the evidence available to the model, the number of tokens it processes, response time, and cost.

The eval tested 1,329 current events questions from LiveNewsBench across 14 conditions and four models. Every condition logged the search trajectory and final answer for paired comparisons of accuracy, latency, cost, and evidence quality.[1][2]

Methodology

SetupHow search worksWhat the model receives
No searchThe setup disables searchOnly the question
You.comShared You.com codeUp to 5 web results and 5 news results
You.com wideShared You.com codeUp to 20 web results and 20 news results
Built-in searchOpenAI or Anthropic toolThe provider's own search results and result format

The model matrix used GPT-5.6 Terra, Claude Sonnet 5, and two models served through Baseten: GLM-5.2 (zai-org/GLM-5.2) and DeepSeek V4 Flash 0731 (deepseek-ai/DeepSeek-V4-Flash-0731). The plots use “DeepSeek V4” as the shorter display label.

The You.com setups normalize every result into the same fields. Each model receives the same title, URL, and text excerpt, which supports direct model comparisons. Built-in search tests the model and its provider's search product as one system. Every setup answered the same questions under the same search-call budget. The analysis pairs comparisons by question.[3]

Question accuracy is the primary outcome throughout this analysis. The SimpleQA grader measures it with a model-based evaluation that accepts equivalent wording and more closely matches the accuracy measure that LiveNewsBench publishes than literal string matching does. A separate scorer-design note near the end examines the deterministic answer match and the cases where the two measures disagree.

Existing search leaderboards hold the model fixed and rank providers.[4] This eval varies the retrieval layer, result depth, and model together. The design measures when retrieval improves accuracy, how much evidence to return, and where relevant evidence still produces a wrong answer.

Results

You.com narrowed the model accuracy range from 47.91 to 5.62 points

Without search, question accuracy ranged from 2.80% for GLM-5.2 to 50.71% for GPT-5.6 Terra. With the same You.com retrieval layer, the four models scored between 79.20% and 84.82%.

Search narrowed the gap between models
Question accuracy on 1,329 current-events questions, measured with the SimpleQA grader and shown with 95% intervals. Paired improvements compare the two setups question by question.
0%20%40%60%80%100%GLM-5.22.8%79.2%+76.4 pairedDeepSeek V48.1%80.7%+72.6 pairedClaude Sonnet 524.7%79.2%+54.5 pairedGPT-5.6 Terra50.7%84.8%+34.2 paired
No searchYou.com
ModelNo-search question accuracyYou.com question accuracyPaired improvementMean total cost per You.com question
GLM-5.22.80%79.20%+76.4 points$0.0464
DeepSeek V4 Flash 07318.05%80.66%+72.6 points$0.0163
Claude Sonnet 524.68%79.23%+54.5 points$0.0747
GPT-5.6 Terra50.71%84.82%+34.2 points$0.0417

Total cost includes model inference and search charges recorded during each run. It uses the pinned list prices for the experiment, excluding promotional and negotiated rates.

You.com increased question accuracy by 76.4 points for GLM and 34.2 points for GPT.[5] The 2.2-fold difference shows that retrieval reduced the effect of the base model on this task. A no-search ranking did not predict the ranking after the eval added search, and a cheaper model with the same retrieval layer approached the accuracy of a more expensive model.

News category produced a 13.5-point accuracy range

Across news categories, question accuracy ranged from 88.1% for Sports to 74.6% for Disasters and accidents. The 13.5-point range exceeded the 5.7-point range between models and the 3.7-point difference between You.com and GPT's built-in search. Science and technology scored 88.2% on 68 questions, compared with 252 Sports questions, so the Sports estimate has a larger sample.[6]

Topic moves the score more than any setup does
Question accuracy with no search and with standard You.com search, by news category, pooled over the four models. Categories were not assigned at random, so these describe the questions that fall in each category rather than an effect of the category itself.
0%20%40%60%80%100%Science and technology88.2%n=68Science and technology: 20.6% with no search, 88.2% with You.com, 88.2% with wide (n=68 per arm)Sports88.1%n=252Sports: 21.8% with no search, 88.1% with You.com, 87.3% with wide (n=252 per arm)Law and crime85.8%n=1116Law and crime: 19.2% with no search, 85.8% with You.com, 87.7% with wide (n=1116 per arm)Business and economy83.9%n=260Business and economy: 17.7% with no search, 83.9% with You.com, 85.8% with wide (n=260 per arm)Health and environment82.7%n=104Health and environment: 14.4% with no search, 82.7% with You.com, 83.7% with wide (n=104 per arm)International relations82.3%n=616International relations: 26.5% with no search, 82.3% with You.com, 84.7% with wide (n=616 per arm)Armed conflicts and attacks80.6%n=1284Armed conflicts and attacks: 26.6% with no search, 80.6% with You.com, 82.4% with wide (n=1284 per arm)Arts and culture78.6%n=168Arts and culture: 30.4% with no search, 78.6% with You.com, 82.7% with wide (n=168 per arm)Politics and elections75.4%n=924Politics and elections: 20.6% with no search, 75.4% with You.com, 77.5% with wide (n=924 per arm)Disasters and accidents74.6%n=524Disasters and accidents: 10.9% with no search, 74.6% with You.com, 80.2% with wide (n=524 per arm)
No searchYou.comSpread across topics with search: 13.6 points

Sports questions often ask for a score, seed, or round number that appears verbatim in reporting. Disasters and accidents often ask for casualty figures and damage estimates that changed as a story developed. Sources can disagree, and the expected value may not appear in the retrieved text.

You.com search increased question accuracy by 48.2 to 68.3 points in every category. The remaining category differences determined the post-search accuracy range.

You.com raised GPT accuracy by 3.66 points and matched Claude accuracy

ModelSearchQuestion accuracyAverage tokensAverage total costCost per correct answerAverage time
Claude Sonnet 5Built-in78.78%42.6k$0.1148$0.145712.74s
Claude Sonnet 5You.com79.23%29.9k$0.0747$0.094211.48s
GPT-5.6 TerraBuilt-in81.16%17.3k$0.0467$0.057612.04s
GPT-5.6 TerraYou.com84.82%21.9k$0.0417$0.04918.64s
You.com answered faster in both head-to-head comparisons
Question accuracy against mean response time on the 1,329 matched questions. Hover or focus a point for tokens and cost.
76%79%82%85%8s10s12s14sYou.comGPT-5.6 Terra · You.com: 84.82%, 8.64s, $0.0417Built-inGPT-5.6 Terra · Built-in: 81.16%, 12.04s, $0.0467You.comClaude Sonnet 5 · You.com: 79.23%, 11.48s, $0.0747Built-inClaude Sonnet 5 · Built-in: 78.78%, 12.74s, $0.1148GPT-5.6 TerraClaude Sonnet 5Mean response time, lower is better
GPT-5.6 Terra · You.com
84.82%
Question accuracy
21.9k
Mean tokens
$0.0417
Mean total cost

GPT gained 3.66 points of question accuracy with You.com and answered 3.40 seconds faster. Total cost fell by $0.0051 per question, and cost per correct answer fell from $0.0576 to $0.0491. The statistical test could not distinguish Claude's 0.45-point accuracy difference from zero. With You.com, Claude answered 1.26 seconds faster, cost $0.0402 less per question, and reduced cost per correct answer from $0.1457 to $0.0942. Built-in search sent Claude 12,700 more tokens.[7]

The provider's tool changes how the model searches as well as what comes back. On its own toolchain GPT issues 0.36 fewer searches per question and takes about 36% longer to answer. Claude issues 0.15 more and takes about 9% longer.

Retrieval gains fell as event age increased

Search closes the freshness gap
Question accuracy by the month the event happened. Select a diamond below the lines to open a question the two search methods scored differently.
0%25%50%75%100%MayJunJulAugSepOctNovDecMay: No search 63.9%, n=108Jun: No search 69.3%, n=88Jul: No search 58.7%, n=143Aug: No search 53.2%, n=62Sep: No search 52.6%, n=211Oct: No search 52.0%, n=171Nov: No search 47.5%, n=303Dec: No search 34.2%, n=243May: You.com 79.6%, n=108Jun: You.com 85.2%, n=88Jul: You.com 81.8%, n=143Aug: You.com 87.1%, n=62Sep: You.com 85.3%, n=211Oct: You.com 86.0%, n=171Nov: You.com 86.1%, n=303Dec: You.com 85.6%, n=243May: You.com wide 80.6%, n=108Jun: You.com wide 83.0%, n=88Jul: You.com wide 86.0%, n=143Aug: You.com wide 93.5%, n=62Sep: You.com wide 86.7%, n=211Oct: You.com wide 84.2%, n=171Nov: You.com wide 87.5%, n=303Dec: You.com wide 86.8%, n=243May: Built-in 78.7%, n=108Jun: Built-in 77.3%, n=88Jul: Built-in 80.4%, n=143Aug: Built-in 85.5%, n=62Sep: Built-in 84.8%, n=211Oct: Built-in 80.1%, n=171Nov: Built-in 82.2%, n=303Dec: Built-in 79.8%, n=243INDIVIDUAL QUESTIONS, BY WHICH ROUTE PASSEDYou.com onlyBuilt-in onlyBoth failBoth passIn the 2025 MLB Draft, which university did the younger brother of A.J. Puk attend?Subtract the defendants in the Gisèle Pelicot trial from the military personnel who marched in the 2025 Bastille Day parade. What is the result?How many days elapsed between Ukraine striking Russia's Ryazan refinery and the reported expiry of Trump's later deadline for Putin to end the war?Add the service games Carlos Alcaraz lost during his 2022 US Open run to the sets he played during his 2025 title run.How many players from the 2025 men's club of the year appeared in the men's Ballon d'Or top 30?Compared with 2023, by how many seats did Forum for Democracy increase its representation in the 2025 Dutch election?By how many novels did Booker Prize submissions exceed the books the five judges considered in 2025?How many seats nationwide did the party that won the most seats in Nineveh secure in Iraq's 2025 parliamentary election?
No searchYou.comYou.com wideBuilt-inOne question, placed in its outcome lane
You.com onlyNov 12, 2025·Politics and elections
How many seats nationwide did the party that won the most seats in Nineveh secure in Iraq's 2025 parliamentary election?
Expected
26 seats
You.com
26 seats
Built-in
30 seats

GPT's no-search question accuracy falls from 64% on May events to 34% on December events, and Claude's from 29% to 12%. The search lines stay between 74% and 94% across the same months. No-search question accuracy rises 5.7 percentage points for every three months the event is older, and the improvement from adding search falls 8.2 points over the same span.[8]

What search is worth depends on how old the questions are
Gain in question accuracy from You.com search over no search, by event age at the time of the run. Bands are 95% intervals.
0%20%40%60%80%GLM-5.2, 8–9 mo: +80.6 points (75.3 to 85.5)GLM-5.2, 10–11 mo: +76.3 points (72.1 to 80.6)GLM-5.2, 12–13 mo: +78.7 points (73.8 to 83.3)GLM-5.2, 14–16 mo: +71.4 points (65.8 to 76.6)GLM-5.2DeepSeek V4, 8–9 mo: +76.5 points (70.5 to 81.9)DeepSeek V4, 10–11 mo: +73.2 points (68.3 to 77.9)DeepSeek V4, 12–13 mo: +73.2 points (68.0 to 78.2)DeepSeek V4, 14–16 mo: +67.5 points (61.7 to 73.0)DeepSeek V4Claude Sonnet 5, 8–9 mo: +72.2 points (66.1 to 78.1)Claude Sonnet 5, 10–11 mo: +51.1 points (45.7 to 56.4)Claude Sonnet 5, 12–13 mo: +52.6 points (45.9 to 59.0)Claude Sonnet 5, 14–16 mo: +44.1 points (38.1 to 50.2)Claude Sonnet 5GPT-5.6 Terra, 8–9 mo: +53.0 points (46.3 to 59.7)GPT-5.6 Terra, 10–11 mo: +34.8 points (29.4 to 40.4)GPT-5.6 Terra, 12–13 mo: +31.3 points (25.0 to 37.7)GPT-5.6 Terra, 14–16 mo: +17.9 points (12.4 to 23.4)GPT-5.6 Terra8–9 mo10–11 mo12–13 mo14–16 moEvent age when the eval ran

The no-search baseline improved as events aged, while search accuracy stayed between 74% and 94%. The change also differed by model. GPT recovered 11.6 points of no-search accuracy per three months of event age, compared with 0.7 points for GLM.

ModelNo-search question accuracy for events 3 months olderGain in question accuracy from search for events 3 months older
GPT-5.6 Terra+11.6 pts−13.6 pts
Claude Sonnet 5+8.5 pts−10.8 pts
DeepSeek V4+1.8 pts−5.3 pts
GLM-5.2+0.7 pts−3.3 pts

The events were 236 to 482 days old when the eval ran. The estimated gain from retrieval fell from about 45 points at the recent end of that range to about 24 points at the older end. A benchmark's mean retrieval gain therefore depends on the age of its questions, and its refresh schedule affects the measurement.

The search methods disagreed on 186 GPT and 198 Claude questions

Claude's aggregate scores differed by 0.45 points, but the two search methods produced different outcomes on 198 questions. GPT's methods disagreed on 186 questions.

Similar averages hide different failures
Question accuracy outcomes for every question answered under both You.com and built-in search. Select a segment to browse the questions in it.
The shared harness found or prioritized the right fact while built-in search selected a different source, number, or scope.
You.com onlySep 07, 2025·Sports
Add the service games Carlos Alcaraz lost during his 2022 US Open run to the sets he played during his 2025 title run.
Expected
46
You.com
46
Built-in
44

For GPT, You.com passed 117 questions that built-in search failed, while built-in passed 69 that You.com failed. The difference between those counts is the entire 3.66-point lead. For Claude, You.com alone passed 102 and built-in alone passed 96.

Most disagreements arise when one search method returns a different number, date, or location. The model then calculates correctly with whatever evidence the search method returned.

  • GPT with You.com returns 26 Iraqi parliamentary seats, while built-in search returns 30.
  • GPT's built-in search gets the 18-novel Booker Prize difference. You.com uses 153 for both values and returns zero.
  • Claude with You.com finds Aemet's Level 2 avalanche warning and keeps the Level 3 warning for higher elevations as a detail. Built-in search returns Level 3 as the main answer.
  • Claude's built-in search gets the Thai parliament calculation right at 211. You.com uses 251 instead of 247 for the majority threshold and returns 215.

The Alcaraz row provides a representative example. You.com returns one passage saying he lost 24 service games in 2022 and another saying he played 22 sets in 2025, and the model adds them to reach 46. The wide and built-in runs do not include a passage supporting the 24. Both answer 44 by counting 22 twice. The selected source causes the error. The arithmetic follows from the evidence provided.

Provider-native search does not support the same passage-level audit. The OpenAI and Anthropic APIs return the queries and cited pages without exposing the passages the model reads.

Wide retrieval improved two models and reduced evidence quality

Increasing the maximum result count from five to twenty produced a small gain in question accuracy for two of the four models.[9] It also changed the composition of the retrieved context.

Twenty results buy coverage and cost context quality
Wide minus standard You.com, in percentage points, one dot per model. Every difference is statistically significant except answer coverage for GPT.
-10-8-6-4-20+2+4Answer coverageAnswer coverage (Any retrieved text contains the expected answer as a string), GLM-5.2: +1.23 ptsAnswer coverage (Any retrieved text contains the expected answer as a string), DeepSeek V4: +1.21 ptsAnswer coverage (Any retrieved text contains the expected answer as a string), Claude Sonnet 5: +1.19 ptsAnswer coverage (Any retrieved text contains the expected answer as a string), GPT-5.6 Terra: +0.53 ptsToken-discounted gainToken-discounted gain (Coverage weighted against the tokens it took to get there), GLM-5.2: +1.75 ptsToken-discounted gain (Coverage weighted against the tokens it took to get there), DeepSeek V4: +1.39 ptsToken-discounted gain (Coverage weighted against the tokens it took to get there), Claude Sonnet 5: +1.18 ptsToken-discounted gain (Coverage weighted against the tokens it took to get there), GPT-5.6 Terra: +2.04 ptsEvidence precisionEvidence precision (Share of retrieved text that directly helps answer the question), GLM-5.2: -1.24 ptsEvidence precision (Share of retrieved text that directly helps answer the question), DeepSeek V4: -1.27 ptsEvidence precision (Share of retrieved text that directly helps answer the question), Claude Sonnet 5: -1.48 ptsEvidence precision (Share of retrieved text that directly helps answer the question), GPT-5.6 Terra: -1.77 ptsSource diversitySource diversity (Spread of distinct websites across the results), GLM-5.2: -6.05 ptsSource diversity (Spread of distinct websites across the results), DeepSeek V4: -5.53 ptsSource diversity (Spread of distinct websites across the results), Claude Sonnet 5: -6.24 ptsSource diversity (Spread of distinct websites across the results), GPT-5.6 Terra: -4.90 ptsDistinctnessDistinctness (Inverse of how much the returned results repeat each other), GLM-5.2: -6.90 ptsDistinctness (Inverse of how much the returned results repeat each other), DeepSeek V4: -6.69 ptsDistinctness (Inverse of how much the returned results repeat each other), Claude Sonnet 5: -6.93 ptsDistinctness (Inverse of how much the returned results repeat each other), GPT-5.6 Terra: -6.10 ptsTemporal groundingTemporal grounding (Retrieved text is dated close to the event), GLM-5.2: -8.00 ptsTemporal grounding (Retrieved text is dated close to the event), DeepSeek V4: -9.69 ptsTemporal grounding (Retrieved text is dated close to the event), Claude Sonnet 5: -7.98 ptsTemporal grounding (Retrieved text is dated close to the event), GPT-5.6 Terra: -7.54 pts
GLM-5.2DeepSeek V4Claude Sonnet 5GPT-5.6 Terra
The wide payload also stops the model searching
-0.11
GLM-5.2
-0.18
DeepSeek V4
-0.09
Claude Sonnet 5
-0.28
GPT-5.6 Terra
Change in searches issued per question, wide minus standard.
Retrieved-context measurementWide minus standard
Answer coverage+0.5 to +1.2 pts
Token-discounted gain+1.2 to +2.0 pts
Evidence precision−1.2 to −1.8 pts
Temporal grounding−7.5 to −9.7 pts
Source diversity−4.9 to −6.2 pts
Distinctness of the retrieved text−6.1 to −6.9 pts
Searches issued per question−0.09 to −0.28

Every difference above has adjusted p below 0.04 except GPT's answer coverage.

Source diversity and distinctness together identify syndicated reporting.[10] Twenty outlets can republish the same wire story, producing many distinct hostnames and repeated text. Both metrics fell in the wide setup, showing that the additional results contained fewer new sites and less new text. Across all four models, wide retrieval increased answer coverage, reduced evidence precision and temporal grounding, and led to fewer follow-up searches after the larger first response.

Retrieval gains varied more by question than by model

Search setup explains more variance than model identity
How much of the question accuracy is set by which question was asked, next to how much each system choice adds on top.
Explained by the question
21.0% of the score is set by which question was asked, before anything you chose.
38.2%
All choices together
60.3%
Choices plus question
Added by each choice you make
Which search setup0.0263
Model and setup together0.0248
Question type and wording0.0141
Which model0.0006
Which model you pick adds the least of the four to the explained variance in question accuracy.

Letting each question have its own retrieval effect fit decisively better than one shared effect. The spread of the per-question effects was wider than the spread of the question baselines and the full accuracy range between models.[11] A benchmark's mean retrieval gain summarizes a wide distribution of question-level effects.

The question also accounted for roughly 82% of the variance in answer coverage, evidence precision, and token-discounted gain. Model and result depth explained little additional variance. Retrieval cannot return an answer absent from the indexed corpus.

A per-question routing policy can use event age, question composition, and whether the first search returns a dated source. These variables distinguished question difficulty more clearly than the choice of search provider.

Runs with five searches had 19% to 49% accuracy

The fifth search is worth much less than the first
Question accuracy by how many searches the model issued. Select a point for the number of questions and the time taken.
25%50%75%100%You.com: 1 searches, 91.1% question accuracy, 6.2s, n=541You.com: 2 searches, 90.0% question accuracy, 9.1s, n=448You.com: 3 searches, 76.8% question accuracy, 11.2s, n=177You.com: 4 searches, 74.2% question accuracy, 15.9s, n=66You.com: 5+ searches, 48.5% question accuracy, 24.4s, n=97Built-in: 1 searches, 89.5% question accuracy, 9.1s, n=771Built-in: 2 searches, 78.2% question accuracy, 14.9s, n=349Built-in: 3 searches, 65.3% question accuracy, 21.8s, n=121Built-in: 4 searches, 55.1% question accuracy, 31.6s, n=49Built-in: 5+ searches, 26.3% question accuracy, 43.8s, n=3812345+Search calls
You.comBuilt-in
You.com, 5+ searches
48.5%
Semantic accuracy
24.4s
Mean latency
97
Rows

For GPT with You.com, question accuracy is 91% after one search and 90% after two, then falls to 49% among runs reaching five or more. GPT with built-in search falls from 90% to 26%, and its average response time rises from 9 to 44 seconds. Claude's runs with five or more searches score 19% with You.com and 27% with built-in, taking about 37 seconds either way.

Search count is observational because hard questions can cause both longer search trajectories and wrong answers.[12] During a run, a third or fourth call indicates that the initial search plan has failed. The agent can change its query strategy before issuing a fifth variation.

Search service failures were rare. Across 13,290 rows where search was available, the model answered without searching in 126 cases, or 0.95%, and Claude produced 119 of those. Another 38 searches encountered a technical problem and still returned a result. Every search attempt returned a response.

Deterministic matching rejected 6,548 grader-accepted answers

A deterministic match measures a different construct
Scorer-design note: how far each score moves when the measurement rises by one standard deviation. Hollow markers are not significant.
-4-20+2+4Answer coverageDeterministic answer match, Answer coverage: +3.75 ptsQuestion accuracy (SimpleQA), Answer coverage: +0.51 pts, not significantEvidence precisionDeterministic answer match, Evidence precision: +3.35 ptsQuestion accuracy (SimpleQA), Evidence precision: +1.55 ptsSearches issuedDeterministic answer match, Searches issued: -1.59 ptsQuestion accuracy (SimpleQA), Searches issued: -4.07 ptsTemporal groundingDeterministic answer match, Temporal grounding: -0.65 pts, not significantQuestion accuracy (SimpleQA), Temporal grounding: +0.39 pts, not significantAnswer lengthDeterministic answer match, Answer length: +0.56 pts, not significantQuestion accuracy (SimpleQA), Answer length: +0.44 pts, not significantSource diversityDeterministic answer match, Source diversity: +0.43 pts, not significantQuestion accuracy (SimpleQA), Source diversity: +0.13 pts, not significantDistinctnessDeterministic answer match, Distinctness: +0.11 pts, not significantQuestion accuracy (SimpleQA), Distinctness: -0.51 pts, not significantPercentage points per 1 standard deviation
Deterministic answer matchQuestion accuracy (SimpleQA)

The deterministic scorer checks whether the expected answer appears verbatim in the response, which shows whether retrieval returned the reference string. The SimpleQA grader measures the primary outcome, question accuracy, because it accepts semantically equivalent answers and more closely follows the LiveNewsBench evaluation.

Answer coverage and evidence precision most strongly predict deterministic answer match. Evidence-quality measurements had much smaller associations with question accuracy. Search count had the largest association and was negative.[13] Long search trajectories usually identify questions where the model is already struggling.

The disagreement runs almost entirely one way. The two differ on 35.5% of rows. Deterministic matching rejects an answer the question-accuracy grader accepted 6,548 times, while the question-accuracy grader rejects one that deterministic matching accepted 52 times. The deterministic scorer behaves like an extra literal-answer filter after the question-accuracy grader.

Substring matching causes those 52 deterministic false positives. One question asks by how many days China's visa-exempt stay for South Koreans exceeds South Korea's allowance for Chinese tour groups. Claude's search returned the 15-day South Korean figure from four separate sources and never established the current 30-day Chinese one. It answered that both are 15 days, so the difference is 0. The expected answer is 15 days, the response contains that string, and deterministic matching scored it correct. Answer coverage has the same blind spot because it also uses string matching.

GPT and Claude shared 132 to 180 failures across search methods

GPT failed 132 questions under both methods, Claude 180.

Questions that combine facts or require a calculation were 13.3 points harder after standard You.com search and gained 4.6 points less from search than single-fact questions.[14] Search finds each value. The model still has to pick values from the same time period and do the arithmetic.

Some shared failures indicate possible dataset errors. Every setup says FvD rose from 3 seats to 7 in the Dutch election, an increase of four, and the dataset expects three. When every setup produces the same well-supported answer, the expected answer warrants review.

What happensLikely problemWhat to do
No search fails and both search methods passThe model lacked a recent factUse search
You.com and built-in search disagreeThey found different sources or datesCheck an official source
Both find facts and calculate the wrong answerThe model chose or combined values incorrectlyExtract values and calculate in code
Both agree with each other and miss the expected answerThe dataset answer may be wrongAsk a person to review the example
The model searches 3 to 5 timesIts first search plan failedSearch for the missing item or stop
Type of questionStart withIf that fails
Recent eventYou.com searchCheck with built-in search
GPT answers an older, low-impact questionNo searchSearch when the model is unsure
GLM or DeepSeek answers an older questionYou.com searchKeep search enabled
Several facts or a calculationSearch and list each needed valueCheck dates and calculate in code
No clear answer after 3 searchesSearch for the missing itemTry the other method or return uncertainty
Search methods return different valuesPause before answeringResolve the source, date, or location

Limitations

  • The live web changes between runs. Changing web content prevents exact page-level reproduction.
  • The events were 8 to 16 months old when the eval ran, and retrieval gain varied with event age. The three-month age trend may differ for breaking news measured in hours.
  • Each condition answered each question once. A single answer per condition bundles run-to-run variance into the residual and makes the reported question-level shares a lower bound.

Run your own eval

The LiveNewsBench dataset and benchmark code are public. The search eval library records the model, search method, event date, queries, returned text, answer, response time, and both accuracy scores as one experiment per condition.

You can use this study's search matrix to compare retrieval methods on your own current events questions. Sign up for free to run the eval, or book a demo to review the design with your team.

  1. 1.The eval defines 14 conditions across four models. Every condition sees the same dataset snapshot and answer prompt, and every search-enabled condition gets the same five-call budget. The analysis pairs results by task, and each condition logs the whole search trajectory alongside the final answer. Checkpoints pin the dataset version and condition order, recover interrupted runs at the row level, and keep a resumed experiment from silently changing the study design. See the eval code.
  2. 2.Each condition logs as one Braintrust experiment. The wrapped OpenAI and Anthropic clients trace every LLM call, and each search call emits its own tool span carrying the query, the normalized results, the resolved You.com parameters, and per-call latency, tokens, retries, and cost. Every evidence report draws every evidence measurement and scorer disagreement from those spans.
  3. 3.Every question ran in all 14 conditions, producing 18,606 rows over 1,329 dataset questions and 1,129 distinct task keys. Averaging by condition throws that pairing away, so the analysis fits one mixed-effects model per outcome with the task key as a grouping factor. These are linear probability models on 0/1 outcomes, so every coefficient reads in percentage points and can in principle fall outside 0 to 100. A logistic version clustered on the task key agrees on the sign and significance of every model, search-setup, and interaction term. The analysis fits evidence models on the 10,632 rows where four models crossed the standard and wide setups. It applies the Benjamini-Hochberg adjustment to per-model contrasts within each scorer.
  4. 4.Artificial Analysis holds one answer model fixed and ranks search APIs on quality, cost, and latency, which is the design you want when selecting a provider. Vals AI compares native and independent search tools across finance and legal tasks. Provider performance depends on the task domain and model-tool combination.
  5. 5.Gain in question accuracy from You.com search, with 95% intervals: GLM-5.2 +76.4 (74.1 to 78.7), DeepSeek V4 +72.6 (70.3 to 74.9), Claude Sonnet 5 +54.5 (52.3 to 56.9), GPT-5.6 Terra +34.2 (31.9 to 36.5). Every interval excludes zero by a wide margin. GLM gains 42.2 accuracy points more than GPT (39.0 to 45.5), with adjusted p below 0.0001. Dropping model identity from the full model for question accuracy costs 0.0006 of explained variance. Dropping the model-by-search-setup interaction costs 0.0248, about forty times as much.
  6. 6.The dataset does not randomly assign news categories. These estimates describe each category slice and do not identify a causal effect. The question-level analysis provides stronger evidence for the same association.
  7. 7.Inside the paired model, after the analysis controls for question and task context, GPT's built-in search costs 3.6 points of question accuracy (1.3 to 5.9), with adjusted p below 0.006. Claude's 0.5-point difference is indistinguishable from zero, with adjusted p near 0.70. Resuming the experiment created two duplicate answers in one GPT You.com run. The paired comparison excludes them. Built-in search ran only for GPT and Claude, so it is a within-model contrast and does not enter the model comparison for DeepSeek or GLM. The two methods also return different fields and allow different controls, so the comparison tests each complete search method.
  8. 8.The age regression accounts for model, news category, question length, calculations, date wording, and named sources. Fitting event age separately inside each setup, no-search question accuracy climbs 1.9 points per 30 days of event age (1.3 to 2.5, p below 1×10⁻⁹), while all three search setups hold flat or drift down between 0.6 and 0.8 points per 30 days, none reaching p below 0.05.
  9. 9.After adjusting for testing four models, Claude gains 3.2 points of question accuracy from the wide tier (0.9 to 5.5, p 0.012) and DeepSeek 2.4 (0.1 to 4.7, p 0.050). GLM's 2.2 points is borderline at p 0.074 and GPT's 1.2 points is indistinguishable from zero at p 0.320. Every model-by-tier interaction on the evidence measurements has a unique R² at or below 0.0010. Pooled, the wide effect is +2.66 points and adjusting for all seven evidence and behavior measurements leaves +1.85 unexplained, so about 30% runs through the measured changes.
  10. 10.Answer coverage checks whether any retrieved text contains the expected answer as a string. Token-discounted gain weights that coverage against the retrieval token count. Evidence precision measures the share of retrieved text that directly helps answer the question. Temporal grounding checks how closely the retrieval date matches the event date. The scorer returned no temporal-grounding value for 10.0% of standard rows and 11.4% of wide rows, so interpret the wide effect only among rows with a score. Source diversity divides the Shannon entropy of the result hostnames by the log of the result count. It reads 1.0 when every result comes from a different site and 0.0 when they all come from one. Distinctness measures how little the returned passages repeat each other.
  11. 11.Question identity accounts for 21.0% of the variance in question accuracy before any fixed effect enters. The fixed effects explain 38.2%, and adding the question's own baseline reaches 60.3%. The standard deviation of the per-question retrieval effect is 0.276, larger than the question baseline at 0.181, and the per-question version fits decisively better, with likelihood-ratio χ² of 1,053 on 2 degrees of freedom. Roughly 82% of the variance in answer coverage, evidence precision, and token-discounted gain also sits between questions. Tier unique R² is at most 0.0015, and model identity does not significantly predict answer coverage at p 0.20.
  12. 12.Search-call counts are observational because the logs record them after the model answers. They show which questions were hard and cannot establish that searching more caused the error. Answers that skipped search break down as Claude 48/1,329 (3.6%) on standard You.com, 59 (4.4%) on wide, and 12 (0.9%) on built-in. DeepSeek skipped search 3 times (0.2%) on standard and once (0.1%) on wide. GPT skipped search twice (0.2%) on wide and once (0.1%) on built-in. Neither GLM setup nor standard GPT You.com skipped search. DeepSeek had 4 and 12 degraded calls, while GPT built-in search had 22.
  13. 13.Associations per one standard deviation, holding model, result tier, and task context fixed. Deterministic answer match: answer coverage +3.75, evidence precision +3.35, searches issued −1.59, temporal grounding −0.65 (p 0.06), source diversity +0.43, distinctness +0.11, and answer length +0.56. Statistical tests do not distinguish the last four from zero. Question accuracy: evidence precision +1.55 and searches issued −4.07 are significant. Statistical tests do not distinguish answer coverage +0.51, temporal grounding +0.39, source diversity +0.13, distinctness −0.51, and answer length +0.44 from zero. The logs record these values after search runs, so they describe which evidence properties accompany a correct answer. They do not establish that increasing one would raise accuracy.
  14. 14.For question accuracy, retrieval adds 4.9 fewer points for a quantitative or composed question than for a single-fact one (2.5 to 7.4 points fewer), with p below 0.0001. Search can retrieve the component facts, but the model still has to align their dates, select the right values, and compose the answer.
Share

Read more evals

Behavior scoring vs output scoring for coding agents
20 August 2026
Compare Kimi K3 and DeepSeek V4
12 August 2026
Testing whether language model harnesses transfer the wrong strategy
7 August 2026

Subscribe to the Department of Evals

A newsletter for unfiltered thoughts on eval methodology, analysis, and failures

Subscribe