New

Shopify launched Agentic Storefronts. We make AI agents recommend you - not just list you.

See the Shopify integration
Tru Commerce
PricingStart free →

← Insights

The AI Re-Crawl Lag — Your PDP Fix Takes Weeks to Surface

You fix a product page and check the assistant the next morning. It still shows the old answer. In one case we tracked, a PDP corrected to the Beauty category kept getting cited under Fashion for weeks after the edit went live. AI answers run on a re-crawl clock, not an edit clock, and the two are 2–4 weeks apart. This piece is about measuring the propagation, not the change — and planning campaigns around the lag instead of being surprised by it.

Viren Inaniyan · September 17, 2026 · Citation Rank & Share of Voice

Timeline chart showing amber edit markers and purple first-observed-in-citation markers for three example product pages, with the 2 to 4 week propagation gap shaded.

You fix a product page, check the assistant the next morning, and it still gives the old answer. That is not a bug in your edit — it is the propagation lag. AI answers run on a re-crawl clock, not an edit clock, and in our measurement the two sit 2–4 weeks apart. In one case a PDP corrected to the Beauty category kept getting cited under its old Fashion attribute for weeks after the change was live. The fix here is not a faster edit. It is measuring the propagation instead of the change, and planning campaigns around the lag.

A brand manager on a call last month described a small crisis: they had corrected a mis-categorised product page, confirmed it live, and then watched ChatGPT keep recommending the product with the old, wrong attribute for the rest of the week. "Did the change not take?" they asked. It took. What had not happened yet was the re-crawl. The edit was done; the answer had not caught up. This piece is about that gap — how long it runs, why it exists, and how to stop being surprised by it.

This is the fifteenth piece in our Winning in AI Visibility with Amazon series, built on the same spine as the rest: a locked panel of 425 real buyer prompts (mixer grinders) re-run monthly, plus the beauty-category expansion, all stored in our geo_vis schema. This topic is more qualitative than most in the series — it is about time, not share — so the honesty bar is higher, and I have flagged exactly what we measured versus what we are estimating.

The case: a fix that was live and still wrong

The clearest example we have sits in the beauty data. A product page was corrected from the wrong top-level category — it had been tagged Fashion — to the right one, Beauty. The change went live on the page. And for weeks afterward, ChatGPT continued to surface and cite the product under its old Fashion attribute. The model's picture of the product trailed the live page by a matter of weeks, not hours.

That is the whole insight in one sentence: the attributes the model cites are a lagged copy of your page, and the lag is measured in weeks. We can bound it, roughly, at 2–4 weeks from this and similar observations. I want to be precise about the precision: we do not have a clean, logged edit-timestamp-to-first-citation-timestamp for every case, so treat "2–4 weeks" as an observed band in one category, not a guaranteed constant. It is enough to plan around; it is not a number to quote to three decimal places.

The mechanism behind it is not mysterious once you separate the clocks:

Stage What happens Who controls the cadence Latency it adds
Edit lands Change saved in Seller Central / your CMS / the PDP You Instant
Re-crawl The page is re-fetched by a crawler or pulled via a feed Crawler / feed cadence Days
Re-index The new attributes enter the retrieval index Index refresh Days–weeks
Surfacing The model selects and cites the new version in an answer Model + retrieval The rest of the 2–4 weeks

Only the last row changes what a buyer sees. The first row is the one most teams celebrate.

Why the lag exists — two clocks, not one

In the click era there was effectively one clock. You changed a page, the search crawler came back on its own cadence, and your ranking moved when the index refreshed. Painful, but a single well-understood pipeline.

AI shopping answers add a second clock on top of the first. The assistant does not read your live page when a shopper asks a question — it reads a retrieval index and hands candidates to a model. So a change has to clear three gates before it reaches an answer: it has to be re-crawled or re-pulled through a feed, it has to be re-indexed, and then the model has to actually start selecting the new version over whatever it had cached in its working picture of the product. Each gate has its own cadence, and they compound. A fast crawl behind a slow index still surfaces slowly; a fast index behind a model that is still favouring the old attribute still surfaces slowly.

This is why the failure feels so counter-intuitive. Your edit is genuinely, verifiably live. The buyer-facing answer is genuinely, verifiably stale. Both are true at once, because they are reading from two different points on two different clocks. It is the same split we described in the two-layer model — the thing you control and the thing the model reads are not the same object — except here the gap is measured in time rather than in layers.

There is a news hook that turns this from a lament into a lever. On June 2 2026, OpenAI launched product feeds for . A feed has a refresh cadence, and a cadence is a number. The interval between your update and the next feed pull is the front half of the propagation lag — and unlike a mystery crawler, a feed-refresh schedule is something you can find out, monitor, and design around. Feed hygiene and feed-refresh timing are quietly becoming the visibility SLA: the contract for how fast your truth reaches the shelf.

What this means if you run a marketplace

The operational shift is to stop reporting edits as outcomes. "We corrected 400 category tags this sprint" is a task-completion metric. It tells you the work happened. It tells you nothing about whether a single buyer saw a corrected answer, because the surfacing is still two to four weeks downstream.

Split the reporting in two. One number for edits shipped — useful for tracking throughput. A separate number for propagation: of the attributes you corrected, how many have actually surfaced in a citation on the locked prompt set, and how long each one took. The second number is the one that maps to visibility. It is also the one that catches a fix that silently failed — a change that never propagates at all, as opposed to one that is merely slow, looks identical for the first fortnight and only the propagation clock tells them apart.

What this means if you run a D2C brand

For a brand the lag has a sharper edge, because brands change pages around moments: a launch, a reformulation, a claim update, a seasonal repositioning, a price move. If the AI answer surfaces the new page 2–4 weeks after you flip the switch, then the edit has to lead the campaign, not accompany it.

Concretely: if a launch is on the 1st, the PDP and feed changes that support it should be live around two to four weeks earlier, so that the assistant is citing the new story when the campaign actually spends. Brands routinely do the opposite — update the page the morning of the launch — and then wonder why the assistant spends the first fortnight of the campaign describing the old product. Availability is the same story from a different angle: as we showed in out-of-stock is an invisibility switch, a stock change also takes propagation time to clear the answer, so a restock timed to a promo needs the same head start.

The rule is simple to state and easy to forget: on AI surfaces, you are editing the answer your buyer will see three weeks from now. Plan the calendar accordingly.

What not to do

Do not re-edit in a panic. The single most common own-goal here is watching a stale answer for a few days, concluding the fix did not take, and changing the page again — often changing it back, or to a third variant, to "force" an update. That does not accelerate the crawl. It resets the clock, adds a second lagged version into the pipeline, and makes the propagation impossible to measure because you can no longer tell which edit any given answer is reflecting. Ship the fix once, then measure. Patience is a measurement discipline, not a personality trait.

Two qualifications keep this honest. First, the 2–4 week band is an observed range in our data, not a physical constant — a specific category, feed, or model update can be faster or slower, and you should measure your own lag rather than inherit ours. Second, propagation lag is a story about time to surface, not about whether you are eligible to surface at all; a page that is not citation-eligible in the first place will not appear no matter how long you wait. The lag governs when a change lands, not whether the underlying eligibility gate is cleared.

The measurement habit

Same discipline as the rest of this series, pointed at the time axis. Take your locked prompt set, re-run it on a schedule, and for every attribute you change, record two dates: when the edit went live, and when the corrected attribute first appears in a citation. The distance between them is your propagation lag, per category and per change type. Watch it the way you would watch a deployment pipeline — because that is what it is.

Do that for a quarter and the lag stops being a surprise and becomes a planning input. You will know that category corrections in beauty take, say, closer to the far end of the band while a price change clears faster; you will know which changes propagate reliably and which quietly die in the pipeline. That is the difference between hoping a fix worked and knowing when it will. It is the same argument we make in measurement is the moat: the number you do not track is the number that lies to you, and here the untracked number is a clock. If you want a system that watches the propagation clock for you across categories, book a demo or start with our .

The edit is not the event. The surfacing is the event, and it happens two to four weeks after you think you are done. Measure the propagation, and you get to plan around the lag instead of being ambushed by it.


Next in the series: The Breadcrumb Myth — why "wrong category" is hygiene, not the lever everyone thinks it is.

FAQ

Sources

  1. 1.Tru Commerce geo_vis panel — mixer-grinder locked set (425 prompts) plus beauty-category expansion, Jan–Jul 2026 pulls
  2. 2.OpenAI — ChatGPT product feeds for shopping (Ads), announced June 2 2026

Continue reading

September 19, 2026

From Recommended to Transacted — AEO Is the On-Ramp, Agentic Commerce Is the Destination

Getting recommended is the on-ramp; being transactable is the destination. ChatGPT-referred ecommerce converts at ~15.9% against ~1.76% for Google (Adobe, 2025), yet most brands still make an agent discover them and then dead-end at a checkout the agent cannot complete. This is the finale of the series: why discovery without transactability leaks the value, what the agentic-commerce protocols already shipping (ACP, UCP, AP2) actually do, and how to turn a recommendation into a purchase inside the chat.

September 15, 2026

The Breadcrumb Myth — Category Path Is Hygiene, Not Leverage

One of the most repeated theories in AI shopping is that a 'wrong' category breadcrumb quietly suppresses your citations. We tested it on a locked panel and the theory collapsed: the odds ratio between breadcrumb-correctness and being cited was 1.0, with a Fisher's exact p of 1.0 — no association at all. Correct categorization is basic hygiene. It is not a visibility lever. Here is the test, why the intuition is wrong, and the three levers that actually move the needle.

September 15, 2026

The Variant-Fragmentation Myth — Consolidating Listings Can Hurt Your AI Visibility

One of the most repeated PDP recipes says merge your color, size, and pack variants onto a single listing to concentrate authority. In our locked-panel data, consolidation ran the wrong way: the more variants a brand collapsed onto one listing, the lower its AI citation rate (Spearman ρ = −0.44 to −0.48). The tidy hypothesis that fewer listings win was disproven across 158 ASINs. Here is the data, the honest correlational caveat, and the reframed play — promote one canonical variant instead of collapsing everything.