AI & Technology

Google Just Measured the Gap Between What AI Knows and What It Will Say. Your Dealership Lives in That Gap.

On August 12, Google Research published findings from 4 million responses across 13 models: frontier AI encodes 95 to 98% of the facts it is tested on, then fails to recall 26 to 34% of them on demand. The gap is worst on rare entities and on questions asked in reverse, which is exactly how every shopper asks about a dealership.

Adam Gillrie - Founder & CEO, Savvy Dealer
August 17, 2026
8 min read

Adam founded Savvy Dealer and has spent 30 years at the intersection of automotive retail and digital strategy.

AI
AI Search
GEO
SEO
Dealer Websites
Google Just Measured the Gap Between What AI Knows and What It Will Say. Your Dealership Lives in That Gap.

Want to Learn More?

Book a quick demo to see these strategies in action.

On August 12, two research scientists at Google published a post with a title that sounds like it came off a parts counter: "Empty shelves or lost keys?"

The question behind it is the one every dealer should have been asking for two years. When an AI gets a fact about your store wrong, or leaves you out of an answer entirely, is that because it never learned the fact, or because it learned the fact and cannot get to it?

Google's answer, measured across 4 million responses from 13 different models, is that it is almost always the second one.

Here are the numbers. On Google's benchmark, frontier models including GPT-5 and Gemini-3 encode 95 to 98% of the facts tested. Those same models then fail to directly recall 26 to 34% of those facts. Turn on extended reasoning and they still miss 11 to 12%.

Read that again, because it reframes the entire conversation dealers have been having about AI visibility. The information is in there. A quarter to a third of the time, the model cannot produce it on demand.

The shelves are stocked. The keys are lost.

Why This Is a New Answer to an Old Problem

In 2023, a team of researchers documented a strange blind spot they named the "reversal curse." Train a model on "A is B" and it often fails at "B is A."

Their demonstration became famous. GPT-4 could name Mary Lee Pfeiffer as Tom Cruise's mother 79% of the time. Asked who Mary Lee Pfeiffer's son is, the same model got it right 33% of the time. Same fact, same model, one direction working and the other collapsing.

For two years the standard reading of that result was pessimistic and simple. The model does not really know things. It has learned a one-way word association, and the appearance of knowledge is a parlor trick.

Google's paper attacks that reading with a method worth understanding, because it is the same method you can use on your own dealership in about five minutes.

The team built a benchmark called WikiProfile, 2,150 facts pulled from Wikipedia, and tested each fact ten different ways. Two tasks probe whether the fact was ever encoded. Four ask for it open-ended. Four offer it as multiple choice. That last split is the whole experiment. If a model cannot generate an answer but can pick it correctly out of a lineup, the fact is clearly in there.

That is exactly what happened. In the authors' words, "in open-ended generation (i.e., recall), reverse questions are consistently harder than direct questions." But in multiple choice, reverse questions showed no meaningful disadvantage at all. The models recognize what they cannot retrieve.

Google's write-up puts it in one blunt sentence: "The reversal curse is a recall problem."

One more finding matters more to a franchised dealer than any other line in the paper. Rare facts, the long tail, showed only modest encoding gaps compared to famous ones. The recall gap on rare facts was substantially larger. Models learn the obscure stuff roughly as well as the popular stuff. They just cannot find it again.

There is no entity on the open web more long-tail than a single-rooftop dealership in a mid-size market.

The Honest Caveats, Before Anyone Sells You Something

The SEO trade press picked this up under the headline that entity order affects AI answers, which is accurate and is also exactly the framing that gets flattened into a vendor pitch within a month. So let me flag the two limits now, because they change what you should actually do.

First, this study is about parametric memory only. That is the knowledge baked into the model during training. It does not test retrieval, search grounding, or RAG. When a shopper asks ChatGPT or Gemini a question and the model goes and looks something up on the live web, this paper says nothing about that path. And for a lot of dealership queries (hours, current inventory, this week's price) that is the path being used.

So if a vendor tells you this study means you need to rewrite your VDPs in a particular word order to win AI search, they are selling you something the research does not support.

Second, the benchmark is 2,150 Wikipedia facts. Google published no breakdown of entity types. There is no evidence in this paper that anyone measured recall on businesses, on local entities, or on a car dealership. Nobody has tested your vertical. Treat what follows as a structural argument about how these systems behave, not as a measured result about dealerships.

With those two stakes in the ground, the part that does apply is still significant.

What This Means for Your Dealership

Start with how a fact about your store enters a model in the first place. It gets written somewhere on the web, in a sentence. Google's paper defines a fact as an ordered pair of entities, a subject and an object, where the subject is whichever one appears first in the source document.

Nearly every sentence ever written about your dealership puts you first. "Smith Ford is a Ford dealer serving Toledo." "Smith Ford has been family owned since 1994." Your name is the subject. The city, the brand, the specialty, all objects.

Now consider how a shopper actually asks. Almost nobody types your name. They ask the question from the other end: "what Ford dealers are in Toledo?" "who sells certified used trucks near me?" "which dealership near me has been around the longest?"

That is a reverse question. The shopper is handing the model the object and asking it to produce the subject, which is precisely the direction Google measured as consistently harder in open-ended generation.

This is the mechanism behind something dealers have been reporting anecdotally for a year. Ask an AI about your dealership by name and it knows you, describes you accurately, sometimes gets the details impressively right. Ask it the question a real shopper would ask and you are simply absent, while three competitors get named.

Both of those things can be true at once. The model has the facts. It cannot retrieve them in the direction the shopper is asking.

That is a very different problem from the one most of the AI visibility industry is selling against. The pitch you keep getting assumes you are missing from the training data and need to be inserted. Google's numbers say the shelves are already stocked for 95 to 98% of facts. The work is on the retrieval side.

What To Actually Do About It

Five things, in order of how much they will move the needle relative to effort.

1. Run the reverse-question test on yourself this week. Open ChatGPT or Gemini, turn web search off so you are testing memory rather than live lookup, and ask the object-first questions: "list Ford dealers in Toledo, Ohio." "which dealerships in [your county] sell certified pre-owned Hondas?" Do not ask about yourself by name first, because that primes the answer. If your competitors appear and you do not, you have just reproduced the paper's finding on your own store, for free.

2. Write the sentence in both directions somewhere on your site. You almost certainly have a hundred pages that say "Smith Ford serves Toledo." You probably have none that say "the Ford dealers serving Toledo include Smith Ford." One clean, natural paragraph on your about or location page that states the relationship from the category side costs you nothing. Do it once, in real prose. Repeating it forty times is the kind of thing that gets pages classified as spam.

3. Get your core facts stated consistently in more places. The long-tail finding is the actionable one. Recall degrades when a fact is rare in the training corpus. Your name, brand, city, and category appearing in the same consistent form across your Google Business Profile, your OEM locator listing, Yelp, your site, and the local press is the cheapest available way to make a fact less rare. This is the least glamorous item on this list and the highest-leverage.

4. Make your pages legible as connected facts, not just readable prose. We wrote about this in July when Google published its Open Knowledge Format, and this paper is the same argument arriving from the research side. Structured, linked, machine-parseable facts about your inventory, service, and location are the durable version of this work.

5. Ignore anyone selling "entity order optimization." It will exist by Q4. The paper's own recommendation is aimed at model builders, not at content creators. The authors argue that gains will come from "better utilization of knowledge already encoded in the model," which is a statement about inference-time reasoning, something Google and OpenAI control and you do not.

The Part That Fixes Itself

There is a genuinely optimistic finding buried in this paper, and it deserves the last word.

Extended reasoning, the "thinking" modes now shipping in every frontier model, recovered 40 to 65% of the facts that were encoded but not directly recallable. For facts the model never learned, thinking helped only 5 to 15%.

That asymmetry is the important part. Thinking cannot invent knowledge the model does not have. It is very good at finding knowledge the model does have. As reasoning modes become the default rather than a premium tier, a large share of this recall gap closes on its own, with no action from you.

Which tells you where to spend. Being present and consistent in the training data is the durable investment, because that is the half thinking cannot fix. Fussing over the word order of your sentences is optimizing against a bottleneck that the model builders are actively engineering away.

The dealers who win the next two years of AI search will be the ones whose facts are everywhere and identical, not the ones who reverse-engineered a phrasing trick.

If you want to know what an AI actually says about your store when a shopper asks the question sideways, we will run the test with you and show you the transcript.

Get Our Answers in Your Google Results

Add Savvy Dealer as a preferred source and Google highlights our articles with a Preferred badge in AI Overviews, AI Mode, and Top Stories. One click, then check the box next to savvydealer.com.

Ready to Transform Your Dealership's Marketing?

Schedule a free demo to see how Savvy Dealer can help you sell more cars.