← All articles

RAG vs OKF: two ways to give AI agents knowledge

How retrieval-augmented generation differs from Google's Open Knowledge Format, why agents are adopting curated knowledge, and when to combine both.

  • RAG
  • OKF
  • AI agents
  • Knowledge management

Every AI agent eventually needs to know things that are not in the model: your company's policies, how a metric is defined, which system owns which data. For the last few years the default answer has been retrieval-augmented generation (RAG). In mid-2026 a different approach started getting attention: the Open Knowledge Format (OKF), an open specification published by Google Cloud for giving agents curated knowledge as plain files.

The two are often presented as rivals. They are not. They solve different halves of the same problem, and understanding which half each one solves is the key to designing an agent that is both broad and precise.

RAG in one paragraph

RAG keeps knowledge in its original documents. A pipeline splits those documents into chunks, converts each chunk into an embedding, and stores the embeddings in a vector index. At question time, the system searches for the chunks most similar to the question, places them in the prompt, and asks the model to answer from them. RAG scales to millions of documents and needs no manual curation. Its weakness is that it reconstructs meaning on the fly from fragments: the answer is only as good as which chunks happened to rank highest. I cover the production side in detail in building a production-grade RAG system.

OKF in one paragraph

OKF takes the opposite approach. Instead of searching raw documents, a team writes down important knowledge once, as a small set of curated Markdown files. Each file describes a single concept, such as a policy, a metric or a system, with YAML frontmatter for metadata and ordinary Markdown links to related concepts. An index.md file lists what exists. An agent reads the index, opens the files it needs, and follows links when it needs more context. Because it is just text in folders, the knowledge is versioned in Git and changed through pull requests, like code.

A concept file looks like this:

---
type: metric
title: Weekly active users
description: Distinct users with at least one qualifying session in a rolling 7-day window.
tags: [product, engagement]
---
 
# Weekly active users (WAU)
 
Count of distinct `user_id` values with one or more **qualifying sessions** in the last
7 days, in UTC.
 
A qualifying session lasts at least 10 seconds and is not from an internal account.
See [Internal accounts](../policies/internal-accounts.md) for the exclusion list.
 
Source of truth: the `analytics.sessions` table, refreshed hourly.
Owner: Product Analytics.

The only required frontmatter field is type; fields such as title, description and tags are recommended. Newer drafts of the specification add optional fields for provenance and trust, such as where a fact came from, whether a human verified it, and when it should be considered stale.

The real difference: finding versus defining

The simplest way to separate them:

  • RAG is a retrieval pattern. It decides how an agent finds relevant text in a large body of content.
  • OKF is a knowledge format. It decides how important knowledge is written down so an agent can rely on it.

That leads to different strengths:

RAGOKF
Best forLarge, unstructured, changing contentSmall, stable, high-stakes facts
How knowledge is preparedAutomatically: chunk, embed, indexDeliberately: written and reviewed by people
How the agent gets itSimilarity search at question timeReads named files and follows links
Relationships between factsInferred from whatever chunks rank highlyExplicit links between concepts
PrecisionGood, varies with retrieval qualityVery high for what has been curated
ScaleMillions of documentsHundreds of concepts, not millions
GovernanceDepends on your pipelineGit history, reviews and ownership built in
Main costInfrastructure and tuningPeople's time to curate and maintain

Why agents are adopting OKF now

Three shifts made a format like OKF attractive in 2026.

Agents act, not just answer. A chatbot that paraphrases a policy slightly wrong is annoying. An agent that applies the wrong refund rule, or computes the wrong revenue figure in an automated report, causes real damage. For facts that drive actions, "probably the right chunk" is not good enough.

Agents are good at reading files. Coding agents showed that a model with file-reading tools can navigate a well-organised folder, read an index, open the right file and follow references. OKF reuses that ability instead of building retrieval infrastructure. Files can be served through ordinary file tools or through the Model Context Protocol.

Governance became a requirement. Teams need to answer "why did the agent say that, and who approved that rule?" A Git history of curated files answers it directly. With embeddings in a vector store, the same question is much harder.

Where OKF falls short

OKF is not free. Someone has to write and maintain every concept, so it does not scale to large document collections like support tickets, contracts or research archives. Loading files also consumes context, so a bundle with thousands of concepts becomes slow and expensive to navigate. And curated knowledge can go stale like any documentation unless owners and review dates are taken seriously.

The hybrid most production agents need

In practice the strongest design uses both, each for what it does best:

QuestionRouteKnowledgeAnswerUser or taskquestionAgent routerwhat kind of fact?OKF bundlecurated conceptsRAG indexdocuments at scaleLLMgrounded + citeddefinitionslong tail
A router sends questions about defined business knowledge to the OKF bundle and everything else to RAG; both feed one grounded answer.
  • Put definitions, rules and canonical facts in OKF: metric definitions, pricing and refund rules, data ownership, compliance requirements, system descriptions.
  • Leave the long tail in RAG: documentation, tickets, meeting notes, product manuals.
  • Let the agent check OKF first for any question about a defined concept, and fall back to RAG for everything else. When both return something, the curated concept wins.
  • Use OKF to improve RAG as well. A clean concept file is excellent input for an index, and links between concepts tell retrieval which documents belong together.

How to start

  1. List the twenty questions where a wrong answer from your agent would be expensive. Those are your first OKF concepts.
  2. Give each concept an owner and a review date, and store the bundle in Git next to the agent's code.
  3. Keep RAG for everything else, and log which questions fall through to it. Questions that appear often and need precision are candidates for new concepts.
  4. Measure both paths with the same evaluation set, so you can see where precision actually improved.

OKF is young and the specification is still evolving, so check the current version before you commit to field names. The underlying idea, though, is worth adopting with or without the format: write down the knowledge your agent must never get wrong, review it like code, and let retrieval handle the rest.

For the bigger picture on structuring agents, see agentic AI frameworks and best practices.