Methodology
One site, one treatment, 91 days of observation.
The site
remoteteamer.com, a remote work job board with a blog. It is not an impressive site and that is part of the point. The content is 100% AI generated. It is no better and no worse than most of what is published on the web. Domain authority is unremarkable and it has no notable backlink profile.
If the effect only showed up on a strong domain, it would not be much use to anybody. This one has nothing going for it.
The window
| Experiment start | 9 May 2026 |
| Treatment applied | ~11 May 2026 |
| Experiment end | 7 August 2026 |
| Duration | 91 days |
| Pages treated | 98 existing blog posts |
| Pages published | 0 |
The treatment was applied once, over roughly two days. After that there was no ongoing work on the site at all. No new posts, no link building, no technical changes. The 89 days that follow are a clean observation window on a single intervention.
What was measured
AI citations come from Bing Webmaster Tools' AI visibility reporting, which covers both ChatGPT and Microsoft Copilot. Organic clicks, impressions, positions and index counts come from Google Search Console and Bing Webmaster Tools.
A citation here means a page was used as a source in an AI answer. It is not a click and it is not a visit. It is the model reading your page and attributing it.
What this is not
There is no control group. This is a before and after on one site, not a randomized trial. The correlation analysis inside the window is stronger evidence than the lift itself, because it compares pages against each other under identical conditions. Read the limitations at the end before you treat any single number as settled.
Finding 1: One refresh pass, roughly 10x the citations
Before the refresh the site averaged 72.5 AI citations a day. That was the baseline for as far back as the data goes.
Ten weeks after the treatment it was running around 700 a day, with single days above 1,000. The best week averaged 841. Against the pre-treatment baseline that is 9.6x sustained and 11.6x at peak.
The inputs were deliberately boring:
- 98 blog posts refreshed
- 0 new pages created
- 0 new domains or links bought
- 0 ongoing work after the initial pass
Same content, same URLs, same site. The only thing that changed was how the text on those 98 pages was structured.
That last point is the one worth sitting with. Every competitor on this topic is spending their budget publishing more. This site published nothing and multiplied its AI visibility by ten.
Finding 2: It took ten weeks, and week one looked like a failure
The lift did not arrive. It accumulated.
| Week | Multiplier vs baseline |
|---|---|
| 1 | 1.3x |
| 2 | 2.5x |
| 3 | 1.6x |
| 4 | 3.0x |
| 6 | 3.7x |
| 8 | 6.1x |
| 10 | 10.8x |
| 11 | 11.6x |
| 13 | 9.7x |
Week one was 1.3x, which is inside the range of ordinary fluctuation. Week two jumped to 2.5x and then week three fell back to 1.6x. At that point the honest read of the data is that nothing is happening.
Week four hit 3.0x, one month after the treatment, and that is the first point where the trend is clearly not noise. The 10x range did not arrive until week ten. Week eleven peaked at 11.6x, and week thirteen came back down to 9.7x as the model rollout hit.
This has a practical consequence. If you run a content refresh and check your AI visibility after three weeks, the data will not tell you whether it worked. It will look like noise, because at that stage it is indistinguishable from noise. You need to give retrieval systems several weeks to recrawl, reprocess and start selecting the new text.
Most people will quit somewhere around week three. The curve says that is exactly the wrong moment to stop.
Finding 3: Half the corpus went from invisible to cited
Total citations going up could just mean a handful of already-strong pages got stronger. That is not what happened.
Before the refresh, 12.8% of the corpus was being cited by AI. After the refresh, 52% was. Four times as many distinct pages were being pulled into answers.
Search Console's indexed page count was flat across the same period. No new pages entered the index. The pages that started getting cited were pages that already existed and were already indexed, sitting there uncited.
This is the difference between a visibility problem and a content problem. The content was already there. It was already crawlable. It just was not written in a way that made it selectable.
Finding 4: Google rank does not predict AI citation
This was the result that surprised us most.
Plotting all 41 pages by average Google position against ChatGPT citation count produces a cloud with no structure in it. Spearman rho is 0.017 with a p value of 0.914. That is not a weak relationship. It is the absence of one.
The two pages at the extremes make it concrete.
| /remote-job-search-websites | /netflix-reviewer-job | |
|---|---|---|
| Google position | 40 | 10 |
| Google clicks | 1 | 294 |
| ChatGPT citations | 5,453 | 1,056 |
The Netflix page is the site's best performer in Google by a distance. It accounts for 64% of all Google traffic to the domain.
The remote job search websites page ranks 40th. It has received exactly one Google click across its lifetime. It is, by every conventional measure, a buried page.
It outcites the site's best ranking page five to one.
This is not just our site
The pattern shows up in third-party data too. seoClarity's analysis of the top-cited pages from ChatGPT looked at the top 1,000 cited URLs and found 25% of them have zero organic visibility in Google. Among the top 3 cited URLs it is 50%.
Their rank correlations came in at 0.034 with browsing on and 0.022 with browsing off. Ours was 0.017. Three separate measurements, all sitting on zero.
They also found the relationship runs slightly backwards: organic visibility declines as a URL is cited more often.
Different dataset, different scale, same conclusion. Ranking and citation are separate systems selecting on different criteria.
Finding 5: 69 AI citations per Google click
Across the 91 day window, over the same 82 pages, the site earned:
- 31,687 AI citations
- 456 Google clicks
That is roughly 69 AI citations for every Google click.
One content cluster makes the point harder still. The /hiring-remotely/ cluster is 39 pages. Over the window it took 13,131 citations and 8 Google clicks. A ratio of 1,641 to 1.
In Search Console that cluster looks like a write-off. Eight clicks across 39 pages is the kind of number that gets a project cancelled. In ChatGPT it is the most productive thing on the site.
Beyond that, 41 cited pages had zero Google impressions. Not zero clicks, zero impressions. Google was not showing them to anybody. Those pages carried 35.7% of all citations the site earned.
If your only instrument is Search Console, more than a third of your AI visibility is invisible to you, and the parts that are working look like the parts that failed.
Finding 6: Bing indexing is the gate. Bing ranking is not.
Three numbers that should be the same size, and are not:
| Pages | |
|---|---|
| Indexed in Google | 41 |
| Indexed in Bing | 118 |
| Cited by ChatGPT | 75 |
ChatGPT cites more pages than Google has indexed. That is only possible because it is not reading Google's index.
The 25 pages that are live in Bing but absent from Google carried 6,248 citations between them. Those citations do not exist in any Google-based view of the site.
Google, in other words, sees 42% of the refreshed corpus. Bing is visibly more permissive about what it will index: the same content got 118 pages into Bing and 41 into Google, without any special effort on the Bing side.
ChatGPT does not run purely on Bing. OpenAI started on a lot of Bing infrastructure and has been moving toward its own crawling and retrieval since. But right now, how many of your pages are in Bing's index is a better predictor of what gets cited than anything Google reports.
Ranking in Bing does not help either
The obvious next question is whether Bing rankings predict citations where Google rankings do not. They do not. Across 55 pages Bing rho came in at 0.067 with p at 0.627, just as flat as Google's 0.017.
So the rule is narrower than "optimize for Bing". It is:
- Being indexed in Bing matters. Pages Bing has not indexed do not get cited.
- Where you rank in Bing does not. Position within that index carries no signal.
Get in the index. Stop caring about the position.
What the refresh actually changed
If rank does not decide it and links did not change, something about the text itself is doing the work.
The scoring model
We built a tool called Chunk Score to answer that question directly. It measures how retrievable a passage is, using the method described in the research paper What Evidence Do Language Models Find Convincing?.
It works on passages, not pages. Give it a question and three competing passages from three sites, and it scores each one on relevance and n-gram overlap against the model from that paper, then tells you which is most likely to be selected.
Take the question "is intermittent fasting effective for weight loss?". The winning passage opened with "Yes, intermittent fasting is effective for weight loss," then added context. It scored 110 relevance with an n-gram overlap of 4. The passage that opened with "whether intermittent fasting works depends on the individual" scored lower on both. The one that opened by defining time-restricted eating scored lowest.
The pattern is consistent: answer the question in the first sentence, in the questioner's own words, then elaborate.
What the scores look like on real pages
We ran it against the site's own pages.
An uncited page ranking 17th in Google, with 3,000 impressions and 19 clicks, scored 99 relevance with an n-gram overlap of 5. Borderline. Not terrible, not good. That page has zero AI citations.
A cited page scored 106 relevance and 3 n-gram overlap. The competing page for the same query, from a site with far more Google traffic, scored 103 and 2.
That is the entire gap. Three points of relevance and one of overlap, on two pages covering the same topic with the same search intent. The competitor is bigger, better trafficked, and better ranked. Our page gets cited far more often.
You do not need to be dramatically better. You need a small structural edge over the other candidates, consistently, across every passage on the page.
The per-post playbook
This is what was actually done to each of the 98 posts:
- Read the page and the live search results to identify what the page does not cover that it should.
- Add roughly three new sections, each an H2 with about three paragraphs. More sections means more passages eligible to be selected.
- Rewrite the weakest existing sections with the specific goal of lifting their relevance and n-gram overlap scores.
- Add an FAQ block, typically three to five questions. Question-shaped headings with direct answers underneath are high-value retrieval targets.
- Update the title tag and meta description.
- Add three to five new internal links pointing at the page from elsewhere on the site.
- Cut the word count. Streamline, remove filler, replace padding with data. One post went from 7,000 words to 3,000.
That last step is worth calling out because it runs against the instinct. The pages got shorter, not longer. Filler dilutes the passages around it. Removing it raised the scores.
The retrieval readiness breakdown on a treated page shows where the gains came from: the intro was already decent and was left alone, one section improved its n-gram score, two weak sections jumped substantially, and the newly added sections came in with solid scores of their own.
None of this was done by hand
98 posts at this depth is not a manual job. An AI agent ran the entire cycle. Once the process was defined, it was handed to the agent on a cron that fired every 15 minutes against a different page, and it ran for several hours until all 98 were done.
That is the difference between a technique and a system. The technique works on one page. The system is what makes it work on the whole corpus while you are asleep.
The dip was not the treatment
There is a large dip in AI visibility across July and early August where the site nearly falls off the map. Nothing was changed on the site during that period, and it recovered on its own.
What happened was a model transition. OpenAI was rolling out GPT-5.6 during that window. Sol, Terra and Luna began rolling out on 9 July in tier waves, and Luna became the default for Free and Go users on 6 August. The retrieval stack shifted underneath the measurement.
This is normal and it is documented elsewhere:
- seoClarity tracked five markets and found citation volumes down 86% to 94% between February and April, then rebounding in May. Their conclusion was that these are platform-level shifts in OpenAI citation behaviour, not a reflection of individual site performance.
- SISTRIX sampled 3.8 million German-language ChatGPT responses across one model transition and found 47% of citations redistributed within 48 hours, against a normal day-to-day variation of 1% to 2%.
The tell that this was not a penalty on the site: citations per page went up during the dip. Fewer total citations were being handed out, but the site's share of them improved. That is a supply-side change in the model, not a demand-side change in the content.
Treat major model releases the way you treat a Google core update. Expect volatility, do not react to the first week of it, and check whether your per-page rate moved before you conclude anything about your own site.
What to do with this
The findings collapse into a short list.
Refresh before you publish. Every page you already have that is indexed and uncited is a cheaper win than a new page. This site's entire result came from pages that already existed.
Check your Bing index count, not your Google one. If Bing has not indexed a page, it does not get cited. If Google has not, that tells you very little about AI visibility either way.
Stop using rank as the scorecard. Rank and citation are independent here at rho 0.017. A page can be buried at position 40 and be your single most cited asset.
Optimize passages, not pages. Answer the question in the first sentence. Use the questioner's language. Add more distinct, well-scoped sections so there are more passages to select from. Cut the filler between them.
Give it ten weeks. The curve is slow, non-monotonic, and looks like failure for the first month.
Do it across the whole corpus. A three point scoring edge on one page is worth very little. The same edge on 98 pages is what produced this result.
Limitations
Read these before you take any number here as settled.
One site, no control group. This is a before and after on a single domain. The lift is consistent with the treatment causing it, but a 91 day window with no control cannot rule out other causes. The within-window correlation results are stronger, because they compare pages against each other under identical conditions.
The corpus is AI generated. Every page on this site was written by AI before the refresh. It is possible that AI-generated content responds differently to structural optimization than human-written content. We do not know.
Bing AI visibility is a proxy. The citation data covers ChatGPT and Copilot as reported by Bing Webmaster Tools. It is the best available instrument and it is still an instrument, not ground truth from OpenAI.
A model transition sits inside the window. GPT-5.6 rolled out during the observation period and measurably moved citation volume. The 10x figure compares a post-treatment period against a pre-treatment baseline across that transition. The per-page rate held up through it, which is why we believe the effect is real, but the specific multiplier would land differently on a different window.
The treatment is a bundle. Seven changes were applied to each post at once. We cannot say from this data which one carries the weight, or how much comes from the internal links versus the passage rewrites versus the added FAQ sections.
Small absolute numbers on the Google side. 456 clicks over 91 days is not much traffic. Ratios built on small denominators are unstable, and the 69:1 figure would move considerably if the site had a few hundred more clicks.
This study covers one site over 91 days, measured from 9 May to 7 August 2026. Citation data is from Bing Webmaster Tools AI visibility reporting; organic data is from Google Search Console and Bing Webmaster Tools.