LLM Memory vs Retrieval
2026-08-31 04:07:56

Exploring How LLMs Decide Between Memory and Retrieval in Responses

Investigating LLMs: Memory or Retrieval?



In the realm of artificial intelligence, large language models (LLMs) have garnered attention, particularly in how they respond to category-related queries. A recent study conducted by Todonada, Inc., headquartered in Taito, Tokyo, has shed light on the decision-making process of LLMs regarding whether to answer from their memory or to leverage retrieval-based approaches, specifically in the beauty domain across 84 subcategories using two models: GPT and Claude.

The research utilized the company's memory measurement tool, Digidigi, which focuses on the 'recall set' of knowledge that these models generate for specific categories. The results reveal a fascinating relationship: the stability of the recall set (the collection of retrieved memories) inversely correlates with the likelihood of activating retrieval tools. In simpler terms, the more stable the recall set for a category, the less the model relies on retrieval methods for answering. The study also explored how specific terms in the queries, such as 'recommended' and 'popular,' trigger retrieval independently of the recall state, with significant variations between the two models.

Key Findings



1. Inverse Correlation Between Recall Stability and RAG Activation Rate
When queries were standardized to the format 'What comes to mind when we say...?' and analyzed by the strength of the recall set, both models indicated that as the strength diminished, the activation rates for retrieval (RAG activation rate) increased significantly. The findings were as follows:
- Claude:
- Stable: ~6%
- Firm: ~7%
- Moderate: ~30%
- Unstable: ~60%
- GPT:
- Stable: ~2%
- Firm: ~1%
- Moderate: ~8%
- Unstable: ~25%

2. Two Layers of Activation
The research identified two distinct layers influencing retrieval activations:
- Recall Point Activation: This generally occurs when the model's recall set is unstable, leading the model to retrieve information if it cannot confidently access the necessary category knowledge.
- Question Point Activation: This involves activating retrieval based on query terms that prompt the model to seek current insights, regardless of its recall state. For instance, changing the question from 'What comes to mind when we say nose and ear hair trimmer?' to 'What are the recommended nose and ear hair trimmers?' led to a stark increase in retrieval calls.

3. Varied Influences of Query Terms by Model
The impact of query phrasing significantly varies by model. In GPT, the term 'popular' triggered retrieval in all measured categories consistently. Conversely, in Claude, while the influence of query terms was less pronounced, variations remained based on recall stability. Claude displayed a tendency to select retrieval probabilistically during repeated trials, resulting in greater variability in the outcomes.

4. Retrieval Does Not Replace Recall
Importantly, when retrieving information, the LLM constructs its own search queries which often include brands or model numbers that existed in its recall set. This indicates that retrieval acts more as a verification or update process rather than a replacement for memory recall. For example, when using GPT to ask about recommended products, the model named current models based on established knowledge, such as Panasonic and Philips for nose and ear hair trimmers, thereby confirming its initial recall.

Investigative Methodology


The study was structured around defining a set of beauty-related subcategories and deploying models such as gpt-5.6-luna and claude-sonnet-5. The recall sets were derived from structured query trials while using Digidigi’s measurement engine to calculate the RAG activation rate based on attempted queries. The entire analysis was performed under controlled conditions with repeat trials across selected beauty categories.

Conclusions and Limitations


While this study offers significant insights, it’s essential to recognize that RAG activation patterns observed are specific to the models tested and do not directly correlate to product implementations like ChatGPT. The inherent variability in trial outcomes should be factored into interpretations of results, suggesting that these fluctuations are better viewed as tendencies within strength bands rather than absolute values. Further research is warranted to solidify our understanding of the observed behaviors and to explore the implications for future model training and application in real-world scenarios.

For extended insights and data, readers can explore the publicly accessible 'Chappyru Log' at Chappyru Log and the functionalities of the Digidigi tool here.

About Todonada, Inc.


Located at Ueno, Taito, Tokyo, Todonada, Inc. specializes in PR Tech solutions, particularly through its flagship products, Qlipper and Digidigi. The methodologies developed aim to enhance the effectiveness of communication strategies for PR professionals.


画像1

画像2

画像3

画像4

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.