English

NewsRaphaël Pierre

6 Configuration Proposals to Start RAG Simply and Gradually Increase Complexity

This article is a translation. Read the Japanese original

The mechanism that allows Large Language Models (LLMs) such as ChatGPT to search external documents and generate answers based on that information is called "Retrieval-Augmented Generation (RAG)."

Raphael Pierre, an AI engineer with experience at Hugging Face and other organizations, proposes a method of advancing the configuration according to usage rather than building a complex system from the start.

The first method is "Full-Text Search." This is a common method of searching for input words within a document and is suitable when the search terms are clear, such as product names or error codes.

The second is "Query Rewriting." By using an LLM to rewrite search terms, it ensures that words with similar meanings, such as "car" and "automobile," are effectively linked.

The third is "Hybrid Search." This method combines full-text search with "semantic search," which examines the closeness of meaning between sentences.

The fourth is "On-Demand Embedding." Only the candidates found through search are converted into "embeddings (vectorization)" on the fly. This reduces the cost of handling data with high update frequencies.

The fifth is "Hot/Cold Tiering." This is an intermediate configuration where only frequently searched documents are processed in advance, while others are processed only when necessary.

The sixth is "Pre-Embedding." All documents are vectorized in advance and stored in a database. This is suitable for cases where updates are infrequent and high-speed searching across a massive volume of documents is required.

Pierre indicated as a guideline that 60% of RAG challenges can be solved with full-text search and query rewriting.

He stated that complex mechanisms should not be prepared from the beginning.


Source: AIの検索を強化する「RAG」をシンプルに始めて高度化する6つの方法 (GIGAZINE, 2026-09-04)