Skip to content
HN On Hacker News ↗

RAG for AI Agents: Why Agentic Retrieval Is Replacing Fixed Vector Pipelines

▲ 3 points • 9 comments • by gael_dev • 2h ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

100 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,110
PEAK AI % 100% · §1
Analyzed
Oct 7
backend: pangram/v3.3
Segments scanned
1 windows
avg 1110 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,110 words · 1 segments analyzed

Human AI-generated
§1 AI · 100%

Retrieval-Augmented Generation has become one of the standard ways to connect language models to external information. The traditional approach is familiar: split documents into chunks, generate embeddings, store those vectors in a database and retrieve the nearest passages before asking the model to answer.That architecture solved an important early problem. But AI agents are changing what retrieval needs to mean. Agents do not only answer questions. They choose tools, compare sources, inspect intermediate results, perform multi-step workflows and decide when they have enough information to act. The next generation of RAG is moving away from a fixed retrieve-then-generate pipeline and toward agentic RAG: retrieval controlled by the agent as part of a broader decision process. Microsoft describes this shift as treating retrieval as a tool that an agent can invoke, evaluate and repeat when necessary.This does not mean vector databases, embeddings or chunking disappear overnight. It means they are no longer the only architecture worth considering, and for many document-agent workflows they should not be the first thing developers build.What is RAG for AI agents?RAG for AI agents is a system that gives an agent access to external information while it reasons and acts. Traditional RAG usually follows one fixed sequence. Agentic RAG changes the sequence so the agent decides when to retrieve.Traditional RAGUser question → Embed the question → Search a vector database → Retrieve chunks → Send chunks to the model → Generate an answerAgentic RAGUser task → Agent analyses the task → Agent decides what information it needs → Agent calls a retrieval or document tool → Agent evaluates the result → Agent calls another tool if necessary → Agent takes an action or respondsThe important difference is that retrieval is no longer an invisible preprocessing step. It becomes a capability available to the agent. LangChain’s documentation describes this principle directly: an agent can use one or more tools to fetch external knowledge, including document loaders, APIs and database queries, instead of relying exclusively on a pre-built retrieval chain.Traditional RAG versus agentic RAGTraditional RAGAgentic RAGRetrieval happens before generationThe agent decides when to retrieveUsually uses one retrieval pipelineCan use multiple tools and sourcesRetrieves similar chunksRetrieves information required for a taskOften returns a single answerCan continue through multiple stepsFixed query-to-context flowDynamic planning and iterationDesigned primarily for question answeringDesigned for reasoning and actionContext is selected by similarityContext is selected by task relevanceRetrieval is an infrastructure layerRetrieval is an agent capabilityAgentic RAG is therefore not simply RAG with an agent added. It is a different way of designing the relationship between agents, documents and external tools.The limitations of fixed vector RAGVector-based RAG remains useful, especially for large, persistent collections of text. But the standard pipeline has several weaknesses when it becomes the default solution for every agent.Chunking can destroy document structureChunking divides a document into smaller passages so they can be indexed and retrieved. Documents are not naturally collections of arbitrary text fragments. Meaning can depend on a section title, a table header, a footnote, a previous paragraph, a page reference, the relationship between rows, a clause and its amendment, or the position of a value inside a form.A table row without its column headers is not equivalent to a complete table. A contract clause without the definitions section may be misleading. A sentence extracted from a policy may depend on exceptions stated several paragraphs later.Embeddings measure similarity, not business relevanceEmbeddings represent semantic relationships. They can identify text that resembles a query, but similarity is not the same as relevance to a business task. An agent asked whether an invoice matches a purchase order needs supplier, product, quantity, unit price, tax, total and currency. A vector search may retrieve text about both documents. It does not automatically perform the comparison.One retrieval call may not be enoughA fixed RAG chain assumes that one retrieval operation can provide the necessary context. Agents often need several steps: find the contract, find the latest amendment, check the renewal clause, compare the date, determine whether notice is required, then create a reminder. The agent decides what to inspect next based on what it discovered. A fixed retrieve-then-answer pipeline is not designed for that process.Vector databases add infrastructureDocument ingestion, parsing, chunking and embedding generation.Vector storage, metadata filtering, index updates and document deletion.Version management, access control and retrieval evaluation.Citation handling and reprocessing when parsing changes.Frameworks such as LlamaIndex and LangChain simplify parts of this process, but they do not eliminate the underlying architectural decisions. LlamaIndex remains particularly useful for data connectors, indexing and query engines, while other frameworks focus on orchestration and tool use. The question is not whether these tools are valuable. The question is whether every agent should begin with a full vector retrieval stack.Chunking, embeddings and vector databases are not the whole futureIt is too simplistic to say that vector databases, embeddings and chunking are obsolete. They remain valuable for large documentation libraries, persistent enterprise knowledge bases, semantic search, similarity discovery, recommendation systems, long-term collections and high-volume repeated queries.The real shift is architectural. Vector retrieval is becoming one tool inside an agentic information system, not the universal foundation of every AI agent. Agents will use the most appropriate tool for each task: structured extraction for invoices, SQL for financial records, APIs for live business data, full-document processing for short files, document comparison for case workflows, knowledge graphs for relationships, search indexes for lexical queries, vector retrieval for semantic discovery, and human review for ambiguity.What is agentic RAG?Agentic RAG is a retrieval architecture in which an AI agent controls when and how external information is accessed. The agent may decide whether retrieval is necessary, choose the source, formulate a more precise sub-question, retrieve from more than one source, evaluate the result, ask a follow-up, compare outputs, detect insufficient evidence and stop before taking an unsupported action.Microsoft’s agentic RAG guidance describes this pattern as dynamic query planning, multi-step reasoning and autonomous information gathering rather than a single fixed retrieval call.Retrieval becomes a toolsearch_documents() and query_document()process_multiple_documents() and compare_documents()extract_fields() and query_database()search_web(), get_contract_amendment() and request_human_review()Asked whether an invoice matches a purchase order, the agent can process both files together, compare supplier and totals, return an exception if they differ, and request review if the result is unclear. It does not need a generic semantic search that hopes the retrieved chunks contain the comparison.The evolution of RAG for agentsStage one: retrieve and generate Query → Vector search → Chunks → LLM answer Stage two: retrieve and rerank Query → Vector search → Reranking → Better chunks → LLM answer Stage three: hybrid retrieval Query → Vector + keyword + metadata → Combined context → LLM answer Stage four: agentic retrieval Task → Agent plans → Selects tools → Retrieves or processes