Skip to content
Resources > Product Articles

The Library You Can't Crawl: Open Web Search vs. Licensed Data

By Daniel Campos, Distinguished Engineer and Charlie Schwartz, Staff Software EngineerAugust 20, 2026
A scatter plot shows 'The full AlphaSense retrieval stack' provides 1.9x citable evidence for $0.27, outperforming other web search options in both cost and evidence quality.

Figure 1: The cost and role of data provenance as it relates to source quality. Purely using web data can often be much more expensive and lower quality than focused content. Price reflects input and output tokens plus the additional cost for web search.

When people think about building AI systems for research and intelligence, there are usually a series of reactions: “We can develop this in-house” andall of that data is already on the web" are a couple that come to mind.

This assumption is foundational to many expensive decisions. It's a simple justification to the build-versus-buy debate, where the “buy” column is mostly vendor fees, and the “build” column is unfounded optimism. It’s the same reason why many AI research tools look incredible for 90-second demos but fail when faced with actual questions people ask for their job.

To show the difference, we decided to measure it rigorously using 241 questions that mirror the type of questions financial services professionals ask every day. Each question runs through AlphaSense’s agentic research stack, measuring the effect of purely using the open web against our proprietary content (while our customers actually get a trained and optimized version of the two). What we found is:

  1. Poor source quality: 80% of what web search returns sits at the bottom rung of the authority ladder, and less than 7% is a primary source.
  2. Dead footnotes: Nearly half of the pages agents tried to open from search results no longer exist or can be accessed.
  3. No dates: 32% of provider 1’s web documents and 47% of provider 2's carry no date at all, which means agents cannot reliably find fresh and relevant information.
  4. Flimsy evidence everywhere: The web is full of documents that could plausibly answer most questions. They are often not right, nor do they withstand scrutiny. This leads to less citable answers, more input and output tokens, and more spend per query. Frontier intelligence with web search is on average 2.2 times more expensive and is 3.6 times less citable.
  5. A pile of owned documents is not a library: Moving from web search to plain vectors over our own documents is a smaller gain than moving from plain vectors to full retrieval over those same documents.

The Test

This test is simple: We use the same research system and foundational model to do the searching and discovering in every arm, so the model is never the variable. What changes is the search stack it can reach: plain vector search over AlphaSense documents, the full AlphaSense search system, open-web search. For the web, we ran two of the most common providers. They fail in the same ways: bottom-rung sources, missing dates, pages that won't open, but not equally: the neural engine returns more undated evidence, satisfies fewer requirements, and costs 84% more per question.

The experiments use the same 241 queries and evaluation methodology from our prior piece, Frontier AI Models Need Frontier Context: Why the Smartest Model Alone Won't Win. Each question has its own rubric that specifies which pieces of information are needed to answer user needs. Each question is graded not only on writing a good answer but on the information found, where it is from, and how fresh it is. If the question is time-sensitive and the evidence is stale, the answer fails no matter how well it reads.

Show Me Your Sources

Just like your middle school history teacher, we ask the question that never stops mattering: is that a reliable source?

Figure 2: Breakdown of the sources and data that have been retrieved. Class is based on document type; for open-web results, on the publishing domain.

Read clearly: 80% of what the web finds is coming from the bottom rung of the authority ladder, and less than 7% is coming from a primary source. A very small percentage of the web-cited documents yield primary data because documents like transcripts and AlphaSense professional research are generally not accessible on the web.

Compare that to AlphaSense retrieval, where filings, transcripts, expert calls, and licensed research carry the evidence, and the bottom rung is a rounding error.

Here is what that looks like in practice. An analyst asked: “Can you give me the latest news on Boeing and Airbus since mid-May 2026? No report before May 15, 2026. Deliveries, backlog, news on Q1 etc.”

Among the 82 documents cited: an Airbus press release stamped 2022 with a URL that dates to 2019, an Instagram reel, Crypto Daily, and Wikipedia. 23 of the 82 fall outside the requested window or carry no date. AlphaSense kept 53 documents, each one dated. Three predate May 15, and all three are Q1 primary documents the analyst asked for: Boeing's Q1 call, Airbus's Q1 call, and Boeing’s 10-Q.

Figure 3: The unbounded nature of the web means there are many answer-shaped documents. More documents are not always better, as their source quality tends to be lower and the token cost to support more documents is higher.

Nothing went wrong with the web model intelligence, and scaling the model will not fix it. It searched exactly the way you would expect it to, finding documents that plausibly answer the question. The trustworthiness of a source was never questioned, and a social reel was a happily received catchall.

Suppose the best web index in the world were designed exactly for the questions our users ask. How much of the evidence could it reach? Nearly half (41%) cannot be bought on the open web at any price (broker research and expert call transcripts). The other 59% is public information: regulatory filings (22%), company press releases (15%), earnings call transcripts (8%), news and wires (8%), and investor presentations (6%). We index it so it can be found, dated, and cited.

The Limitations of Web Search

A web search does not hand an agent documents. It hands the agent snippets: a median of 1,400 characters for provider 1 and 3,683 for provider 2 (approximately 230 and 470 words, respectively), written when the page was crawled or generated at runtime. Throughout this experiment, agents accepted nearly 30,000 of these snippets as evidence. Provider 1 attempted to open the real document only 3.3% of the time, while provider 2 had a 4.8% open rate. Almost every claim in a web-researched answer rests on a cached summary that no human or model ever checked against the page.

Figure 4: Window of limited agent interaction with the actual source leads to difficult-to-uncover data provenance issues.

We issued these queries to web search during the week of August 10, and the questions themselves date from the week of July 20. That is how fast the web changes. When search engines tell you they hold 100 billion documents, they are not telling you how many are still active. The failures are also unevenly distributed: shallow articles that cite secondary or primary sources open about 60% of the time, while PDFs and direct documents fail 72% of the time.

Revisiting the analyst question on Boeing and Airbus news on Q1: The web run bounded its queries by date (May 15, 2026 through July 31, 2026), then went after the pages its snippets could not cover. Two of the four pages it tried never opened: Reuters on the FAA restoring Boeing's certification authority, and Airbus's own mid-term outlook release. The model wrote from search snippets for both. Nothing about either page marked it out beforehand as one that would fail. That per-page asymmetry is what makes the open web impossible for an agent to plan around.

The Web Has Few Dates. Vector Search Ignores Them.

Another problem with the open web is that about 40% of evidence has no date or temporal component, which makes it impossible to tell if the page cited was written today or in 2010. We explored this deficiency further by taking undated documents retrieved from a web search and attempting to extract (or update) the publication date in five different ways, and were still only able to find dates for 11% of those documents.

This has a measurable effect. In the query discussed in Figure 5, “For about 18 months now, RoRo companies are transiting the Panama Canal more often. What explains this increase?,” the agent that purely uses web search leans its entire answer on a document that is relevant but not fresh. The agents with provider 1 choose a 2023 update from Wallenius Wilhelmsen (with provider 2 instead providing a 2024 update), placing it a whopping 15 months before the beginning of the timeframe the analyst asked about.

Figure 5: Each dot is one document the agent kept as evidence, plotted by its publication date. The web run scattered across five years, while AlphaSense retrieval stayed inside the period the analyst asked about. Web documents with no date at all are not shown.

When we don't use the full intelligence of our search stack and undermine its depth, the opposite problem exists. Every document has an ingestion date, and the naive search system retrieves roughly 50% of content that is 18 months or older, with only 12% under 90 days old. Vector search can be incredible, but it does not do a good job of understanding the role of time. Ask about this quarter’s earnings, and it hands you a beautifully on-topic document from three years ago. With AlphaSense’s full intelligence, 65% of content is procured from within 90 days.

Like Peanut Butter and Jelly, Better Together

Everything above could be interpreted as a case against using the open web. It isn’t. It’s a case against using the open web alone and uninformed. The same grading rigor makes a case for both avoiding using it alone and its inherent value. Using our rubrics, we can go deeper than “who won?” and “by how much?” We can iterate criterion by criterion and ask: did the web get it, did AlphaSense get it, or did you need a combination of the two?

Figure 6: While it can be easier to rely on one source, the true answer carries nuance. To quantify this, we evaluate and quantify source coverage, with the goal of providing context for a more complete document pool.

Start with the requirements where web and AlphaSense are well matched to the explicit requests in the query. You'll find that 50% of the time both surfaces got there. Unsurprisingly, there are facts available from more than one place. 26% of the time, the information was only reachable by the open web with documents from local news reports, a trade blog, a private company that files nothing, and an expert’s podcast. The information may be something you cite, or it may be something that helps you find the thing you cite. Turn the web off, and you get holes in your answer. This is why our product uses web data when it is the right call, and why it supports answers on roughly 8% of real user queries.

Turn your focus to where the web is lacking: a source with citations is coming from AlphaSense 60% of the time vs 13% for open web. Similarly, breadth of coverage stacks 63% to 17%. Evidence of the right timeliness: 40% to 18%. The web has a unique set of documents that are real and broad; however, they are not commonly primary sources or current and fresh.

The open web is how to find the unfiled or unpublished document. The library is how to prove what the open web found. A system that cannot leverage both inherently has a narrow view of the world. Better together is not a tradeoff; the requirements each side satisfies barely overlap, which makes the union much stronger than by itself.

Owning the Documents is Not the Same as Having the Library

If getting the data was all you needed, our full search stack would score no better than plain vector search over the same files. Same documents, same model, same questions. The full stack returns 82% more citable answers.

The bar is not whether an answer appeared. Any agent can produce answer-shaped output tokens whether the underlying corpus is the open web or an embedded pile of licensed documents; so “did we get an answer?” indicates nothing. This is why demos are worthless evidence. The bar is whether the source behind the answer is one you would let an analyst cite.

Plain vectors and the full stack differ by one thing: the layer on top that knows what each document is, when it was published, who published it, and which questions it should answer. Vectors fail in this regard. They match topics and erroneously hand you a three-year-old document for a question about this quarter, and fail to distinguish an audited filing from a blog post that mentions the same number.

That layer is invisible in a demo, which is why it keeps getting poorly rebuilt by people who assumed it was a weekend of embeddings.

You would never let an analyst cite an Instagram reel for any kind of financial data. Why would you let your model?

Figure 7: Impact of provenance of the sources for answering by retrieval stack.

Key Takeaways

Three tools, three jobs.

The open web is a broad instrument. You can mostly find an answer to any question on the internet, and some questions need information that exists nowhere else. That is a real job, and AlphaSense uses web data to augment and support results when it is the right call. In fact, web documents support answers in about 8% of user queries. What this study measured is what happens when the web is the only instrument: 80% of what it returns sits on the bottom rung of the authority ladder, 7% is a primary source for Provider 1 and 10% for Provider 2, 40% carry no date at all, nearly half of web origin pages cannot be reached, with two-thirds of the unreachable being gone for good. The web answers. It does not substantiate.

Licensed content is the deep instrument. 41% of the evidence these questions needed cannot be scraped on the open web. You cannot vibe-code your way into institutional-grade financial data. It is slower and more expensive than crawling other people’s content. It is also the only version that passes compliance, and the only one whose documents an analyst would cite by name.

The retrieval layer is what turns documents into a library. Buying and embedding the content is not the finish line. Vector search matches topics; a topic is not the only thing a question asks for. It lacks date awareness: only 12% of vector search returns are less than 90 days old, against 65% for the full stack. It lacks document type awareness as well, because to a vector, a blog post about revenue and an audited filing are equally on-topic. That judgment does not live in the vectors. It lives in the layer above them.

Investment-grade research relies on all three in harmony. It needs the library for depth, the open web for breadth, and the reasoning and orchestration layer at the top deciding what to trust and when. Take either component away, and the other two cannot cover for it.

About the Authors
  • Daniel Campos

    Daniel Campos, Distinguished Engineer

    Daniel Campos is a Distinguished Engineer at AlphaSense, where he researches how to improve language model inference, search, and information discovery for investment professionals and other knowledge workers. Before AlphaSense, Daniel founded Zipf AI and held research roles at Snowflake, Neeva, and Microsoft. He holds a PhD in Computer Science, an MS in Computational Linguistics, and a BS in Computer Science.
  • charlie schwartz

    Charlie Schwartz, Staff Software Engineer

    Staff Software Engineer
    Charlie is a Staff Software Engineer at AlphaSense, where he builds the software and infrastructure that turns research ideas into production systems. Before AlphaSense, he was the Founding Engineer at Zipf AI, and previously held engineering roles at Amazon and Red Ventures.

Explore more

Frontier AI Models Need Frontier Context: Why the Smartest Model Alone Won't Win

Today’s bottleneck on answer quality is no longer raw model intelligence — it’s context. Learn what we’re building next at AlphaSense.
model quality vs. cost with AlphaSense Search

Building Trustworthy Agentic AI for Financial Services

New FCA research shows agentic AI demands grounded, traceable, auditable oversight. See how AlphaSense delivers decision-grade AI for finance.
decision-grade agentic ai for financial services

The Agentic Enterprise

The next phase of AI transformation isn't about what your models can do. It's about what your company learns from every decision it makes and who owns that learning.
agentic enterprise

Transform intelligence
into advantage

Develop bold strategies, seize opportunities,
and lead with clarity and confidence.