The Page Is Not the Unit of Retrieval

Watch the video demonstration:

When people think about search, they often think in terms of pages.

A page ranks. A page appears in search results. A page gets traffic.

AI search introduces a different way to think about content:

The page is not necessarily the unit of retrieval.

An AI system may work with smaller passages taken from a page. For a particular question, one paragraph or section may become important while most of the page contributes very little to the answer.

This changes how we can think about articles, documentation, product pages, and other web content.

A 2,000-word article may effectively contain dozens of potential pieces of evidence. Which pieces matter depends on the question being asked.

That is the central idea behind the Oppvera Content Showdown demonstration.

An Article Can Become a Collection of Candidate Passages

Consider an article with these sections:

  • About NASA
  • Earth science programs
  • Satellite missions
  • Artificial intelligence
  • Climate observations
  • Research partnerships
  • Data access

A person can open that article and read the whole thing.

An AI retrieval system can process the content into smaller units and search those units for information relevant to a particular query.

OpenAI documents this process directly in its Retrieval API. Files placed into an OpenAI vector store are automatically chunked, embedded, and indexed. OpenAI’s current documentation says its default chunking strategy divides a file into chunks of up to 800 tokens (about 600 words), with overlap between adjacent chunks.[1]

Both OpenAI’s documented retrieval systems and Oppvera then search those chunks using semantic similarity.

So imagine someone asks:

How does NASA use artificial intelligence for Earth observation?

The retrieval system may find that two passages are especially relevant:

NASA uses machine-learning systems to classify satellite imagery and identify changes in Earth’s surface.

and:

Artificial intelligence helps researchers analyze large Earth-observation datasets.

Those passages could have a much stronger role in generating the answer than the rest of the article.

The article still matters. Its source, context, authority, accessibility, and surrounding information can all matter.

But at the moment of retrieval, a much smaller section of the article may become the useful evidence.

The Question Changes Which Part of the Page Matters

Now change the question:

How can researchers download NASA Earth-observation data?

The passages about machine learning may become irrelevant.

The retrieval system may instead find a passage in the “Data access” section.

Ask:

Which NASA satellites measure changes in Earth’s climate?

Now passages from the satellite and climate sections may become important.

The same page therefore has different retrieval characteristics depending on the query.

This gives us a useful mental model:

Query → candidate passages → selected evidence → LLM answer

OpenAI confirms that ChatGPT Search can rewrite a user’s question into one or more targeted searches and use web sources to construct an answer.[2]

OpenAI’s Retrieval API provides a concrete example of how OpenAI systems can turn larger documents into smaller searchable chunks.[1]

That makes passage-level retrieval a useful model for thinking about what may happen between a user’s question and the final AI answer.

Why Break Pages Into Smaller Pieces?

There is a practical reason retrieval systems do this.

A page can contain several unrelated ideas.

Imagine embedding the meaning of an entire long page into a single numerical representation.

The page might discuss:

  • satellite hardware

  • climate science

  • machine learning

  • grant programs

  • data APIs

  • educational programs

  • NASA history

Representing all of those subjects as one vector can blur the distinction between them.

Microsoft describes this issue in its Azure AI Search documentation. It notes that even when an entire document fits within model limits, retrieval can improve when a document covering several topics is divided at a finer level.[3]

Smaller chunks can produce more focused representations.

The system can then compare the meaning of the user’s question with the meaning represented by individual passages.

That creates a much more granular search.

Embeddings Let the System Search by Meaning

Chunking is only part of the process.

The passages can also be converted into embeddings.

An embedding is a numerical representation of the meaning of text.

OpenAI describes semantic search as a method that can retrieve relevant results even when the result shares few or no exact words with the query.[1]

For example, a user might ask:

When did humans first land on the moon?

A relevant passage might say:

The first lunar landing occurred in July 1969.

Those sentences do not use exactly the same wording.

Their meanings are closely related.

OpenAI’s Retrieval documentation uses this same type of example to illustrate the difference between keyword matching and semantic similarity.[1]

This is important for AI search because content does not necessarily have to repeat the user’s exact question.

It needs to contain information that a retrieval system can recognize as relevant to the meaning of the question.

Retrieval Can Combine Several Techniques

Modern retrieval systems can also use more than embeddings.

Anthropic describes a common Retrieval-Augmented Generation, or RAG, architecture that:

  1. breaks documents into smaller chunks,

  2. creates embeddings for those chunks,

  3. uses keyword-based retrieval such as BM25,

  4. uses semantic retrieval,

  5. combines the results,

  6. selects top chunks,

  7. supplies those chunks to the language model.[4]

This is particularly interesting because it shows that AI retrieval can operate at the passage level while using several signals to decide which passages deserve attention.

A passage may perform well because its meaning closely matches the question.

Another may perform well because it contains an important exact term.

A reranking system can then compare the candidates again before the final passages are sent to the LLM.

The resulting pipeline starts looking less like:

Find the best article.

And more like:

Find the strongest pieces of evidence for this question.

A Page May Have Strong and Weak Retrieval Areas

This leads to an important idea for people creating content.

A page does not necessarily have one AI-search performance level.

Different sections can behave very differently.

Imagine an article containing an excellent explanation of an organization’s product, but that explanation is buried inside a long paragraph:

Our organization has spent years advancing an integrated approach to information access across rapidly evolving digital environments, bringing together multiple technologies and strategic capabilities that help organizations adapt as the information landscape changes, including tools that allow teams to analyze how content performs when processed by emerging artificial intelligence systems.

Now compare it with:

Oppvera compares two versions of website content and shows which passages an AI retrieval system selects for a specific question.

The second passage describes the idea much more explicitly.

A retrieval system trying to answer:

What tool can compare which website passages an AI retrieves?

has a very clear candidate passage.

This does not imply that writers should fill pages with repetitive keywords.

It suggests that clear, specific statements can create stronger pieces of retrievable evidence.

That is useful for people too.

Passage Visibility May Matter Alongside Page Visibility

Traditional search has encouraged people to ask:

Does my page rank?

For AI search, another question becomes useful:

Does my page contain a strong passage for this particular question?

That distinction gives us two different ideas.

Page visibility asks whether the source can be found.

Passage visibility asks whether the relevant information inside that source can be found and selected.

A page might be accessible to an AI system while a particular idea buried inside the page rarely becomes useful evidence.

Another page may contain a concise, clearly written passage that strongly answers the question.

The second passage may become more useful during retrieval.

Context Around a Passage Also Matters

There is a complication.

Breaking documents into chunks can remove context.

Consider this passage:

Revenue increased by 18% during the quarter.

On its own, that sentence leaves important questions unanswered.

Which company?

Which quarter?

Which year?

Anthropic discusses exactly this type of problem in its work on “Contextual Retrieval.” Its approach adds information from the surrounding document to individual chunks before indexing them.[4]

That research illustrates another important principle:

The passage may be the retrieval unit, while the surrounding document still helps establish what that passage means.

Microsoft’s documentation makes a similar point when discussing semantic chunking. It recommends preserving meaningful structures such as headings, paragraphs, and sentences because higher-quality, semantically coherent chunks can improve retrieval relevance.[5]

So article structure can matter.

A heading such as:

How NASA Uses AI to Analyze Satellite Imagery

provides useful context for the paragraphs beneath it.

Clear sections can help preserve relationships between ideas when content is divided into smaller pieces.

The Page Still Provides the Container

The page remains important in this model.

It provides:

  • the source of the passages
  • surrounding context
  • titles and headings
  • relationships between concepts
  • links
  • metadata
  • organizational identity
  • broader topical meaning

The useful shift is to stop thinking of the page as a completely indivisible object.

A page can be viewed as a container holding many potential pieces of evidence.

A retrieval system can select different evidence from that container depending on the user’s question.

This Changes How We Can Think About Writing for AI Search

Suppose your company wants to be associated with a particular concept.

You could ask:

Do we have a page about this?

The passage-level model suggests several additional questions:

Where on the page is the actual answer?

Can the relevant passage stand on its own?

Does the section clearly identify what it is talking about?

Would the passage make sense if it were retrieved separately from the rest of the article?

Does another source contain a clearer answer?

Which version of our content produces stronger evidence for the same question?

These questions move AI-search analysis closer to information retrieval.

They also create something that can be tested.

Testing Two Versions of the Same Content

That is what I demonstrate in the video.

Oppvera Studio includes two sample NASA-related campaigns:

Optimized for Human Experience

and

Optimized for AI Experience

The content is prepared for retrieval, which includes breaking documents into chunks and representing those chunks numerically.

Then both campaigns can be tested with the same questions.

Oppvera’s Content Showdown shows:

  • the question
  • the generated response
  • the source documents
  • the retrieved passages
  • which campaign supplied stronger evidence
  • why particular passages were selected
  • which chunks ranked more strongly

Instead of looking only at the final LLM response, you can inspect what happened earlier in the pipeline.

That earlier stage is where the concept of passage-level retrieval becomes visible.

A Content Experiment

Consider two versions of a NASA-related article.

The first contains this information spread across several paragraphs:

NASA operates numerous programs involving advanced computational systems. These activities support researchers working with increasingly large datasets produced by Earth-observing missions.

The second says:

NASA uses artificial intelligence and machine learning to analyze satellite imagery and large Earth-observation datasets.

Now ask:

How does NASA use AI for Earth observation?

A retrieval experiment can determine which passage is more strongly associated with that question.

You can change the wording.

Run the question again.

Change the heading.

Run it again.

Add a clearer explanation.

Run it again.

The result becomes an experiment about information retrieval rather than a general guess about whether content is “AI optimized.”

From Articles to Evidence

The main concept to understand in AEO/GEO strategy is:

An article can be thought of as a collection of potential evidence.

A user asks a question.

The retrieval system searches for useful evidence.

Some passages become candidates.

Those candidates can be ranked.

A small subset can then become context for the language model.

The LLM uses that context to formulate an answer.

This provides a practical way to understand AI search without treating the entire system as a black box.

Despite an imprecise model of the proprietary retrieval pipeline used by ChatGPT, we know that chunking, embeddings, semantic search, keyword retrieval, reranking, and passage selection are established techniques used by major AI and search platforms.[1][3][4][5]

OpenAI explicitly uses chunking, embeddings, and semantic retrieval in its own documented retrieval systems.[1]

That gives us a useful framework for experimentation:

The page is not necessarily the unit of retrieval. The passage may be the piece of content competing to become evidence for the answer.

That is the idea explored in the Oppvera Content Showdown.

Watch the full video:

Learn AI Search Techniques with Oppvera for free
https://oppvera.com/


Sources

[1] OpenAI — Retrieval API
OpenAI documents semantic search, vector stores, embeddings, and automatic document chunking. Files added to a vector store are automatically chunked, embedded, and indexed.

[2] OpenAI — Searching the Web with ChatGPT
OpenAI explains that ChatGPT Search can rewrite a user’s question into targeted searches, obtain web results, and provide responses with source citations.
https://help.openai.com/en/articles/9237897-searching-the-web-with-chatgpt

[3] Microsoft — Chunk Documents for RAG and Vector Search in Azure AI Search
Microsoft explains why dividing documents into smaller chunks can improve retrieval and discusses fixed-size, structural, and semantic chunking strategies.

[4] Anthropic — Contextual Retrieval
Anthropic describes a RAG pipeline using document chunks, embeddings, BM25 keyword retrieval, rank fusion, reranking, and selection of top chunks for an LLM. It also discusses the loss of context that can occur when passages are separated from their parent documents.

[5] Microsoft — Chunk and Vectorize with the Document Layout Skill
Microsoft describes structure-aware chunking using headings, paragraphs, and sentences and explains how semantically coherent chunks can improve retrieval relevance.