Why pasting your URL into ChatGPT isn't a website audit
ChatGPT does not run your JavaScript and cannot measure anything. What it actually sees when you paste a URL, what the research says about invented specifics, and what to do instead.
Pasting your URL into ChatGPT gets you a confident, well-written answer about a version of your page it could mostly not see. That is a different thing from an audit.
Almost everyone reading this has already tried it. You paste the link, you ask what is wrong, and back comes a tidy list. Strengthen the headline. Add social proof. The call to action needs work. It reads well. It is also roughly what you would get for any website, which is the first clue.
Can ChatGPT audit your website?
It can review whatever text sits in your raw HTML, and it does that well. It cannot render the page, time the load, open a phone screen or see your analytics. So you get plausible advice about a partial copy of your site, with nothing to say which finding is costing you customers.
That second gap matters more than the missing pixels. An audit is not a list of observations, it is a ranked list, and ranking needs numbers the model does not have. We priced the alternatives in what a CRO audit actually costs.
What does ChatGPT actually see when you paste a URL?
A fetcher called ChatGPT-User requests the page and hands the raw HTML back to the model. It does not run your JavaScript, wait for your API calls to resolve, or finish hydration. Anything that arrives after the page loads, which on a modern site is often the content itself, never reaches it.
Vercel and MERJ tested this across 569 million GPTBot fetches and found no evidence of JavaScript execution by any major AI crawler, OpenAI's included. The crawlers do download JavaScript files, ChatGPT in 11.50 percent of requests, but they read them as text instead of running them.
Cloudflare, measuring separately on its own network, splits AI bot traffic by purpose. The category covering your pasted link is user action, and ChatGPT-User is nearly three quarters of it. Both companies sell something, hosting and bot control, though neither sells conversion work.
569M GPTBot fetches analysed, zero JavaScript executed
11.5% of ChatGPT requests pull JS files it never runs
34.8% of ChatGPT fetches land on a page that is not there
That last figure is the quiet one. A third of the time, the thing it went to read did not exist.
What does the research show?
That models invent specifics, and invent more of them the less the subject appears in training data. Researchers at Deakin University checked all 176 references GPT-4o produced across six literature reviews. Thirty-five were entirely made up, and nearly two thirds were either invented or carried errors.
The headline rate is not the useful part. The pattern underneath is. The study, published in JMIR Mental Health, ran the same task across three topics with very different amounts of coverage behind them.
Heavily researched ██ 6%
Moderately known ███████████ 28%
Barely covered ████████████ 29%
Fabrication climbed as coverage thinned. Now ask what your website is. One domain, no literature, nothing in training data about it. The conditions that produced 29 percent are the generous end of the scale compared with asking a model about you.
The model is not lying to you. It is finishing a sentence about a page it never fully read.
How do you tell which answers to trust?
Ask it to quote. Pick something that only exists on your page after JavaScript runs, a price, a review count, a stock figure, and ask ChatGPT to repeat it back word for word. If it paraphrases, hedges or returns a number you cannot find, it did not read that part.
Then ask for your mobile conversion rate. It will either decline, which is the honest answer, or estimate one, which tells you how much of the rest was estimated too. The same sorting applies to any audit tool including ours, and the three-run method is in do AI website audits hallucinate findings.
Crawl your site properly and see what the paste missed, free.
Where ChatGPT genuinely helps
Rewriting a headline once you know the headline is the problem. Explaining why a paragraph reads badly. Drafting five versions of a button label. It is a strong editor and a poor instrument. The trouble starts when you ask a language tool to take a measurement.
How do you get an audit you can act on?
Separate the two jobs. Something has to look at the rendered page and produce numbers. Then something can interpret them. Doing it in that order is the entire difference, and three of these four steps cost nothing.
- Render it before you read it. Load your own page with JavaScript disabled. Whatever vanishes is what ChatGPT was working from.
- Bring your own numbers. Paste real analytics next to the URL, mobile against desktop over 90 days. Now the model has something to reason about rather than guess at.
- Price each finding before you fix it. Run
traffic × CVR gap × AOV × mobile weightover the candidates. A finding with no euro figure attached is an opinion. - Fix one thing at a time. Ship the largest number on its own and measure it. Bundled fixes teach you nothing about which one worked.
The method for step three is written out in the conversion leak index. If the answer you keep arriving at is to rebuild everything, read redesign or fix before you sign anything.
Paste your URL and get findings priced, not guessed.
When should you ignore this?
If your site is server rendered, plain and small, ChatGPT sees nearly all of it, and its notes on your copy are worth having. The gap described here is widest on the sites that need help most: heavy front ends, ecommerce, anything where the content arrives after the page does.
Now the obvious objection. We sell an audit, so of course we say the free thing is not one. So here is our own evidence, stated plainly. Revslip has audited 134 sites against roughly 200 checks, and that group is not a random sample of the web. Those owners already suspected a problem. Our frequency numbers describe them, not you by default, and every sourced claim above stands without us.
Questions people ask about AI website audits
Three come up repeatedly. Whether a better prompt closes the gap, whether the newer agent modes change the answer, and whether any of this is about to be fixed.
Does a better prompt fix it?
It improves the writing, not the input. No prompt makes a model see content that was never in the HTML it received. Hand it your rendered text and your analytics and it gets much better, because you have removed the guessing rather than dressed it up.
What about agent modes that drive a real browser?
Different thing, and better. An agent opening a real browser does render the page, so it sees what you see. It is slower, it works one screen at a time, and it still measures nothing. Worth using. Not the same as pasting a link into a chat window.
Will this change?
Partly, already. The same Vercel and MERJ work found Google's Gemini and Apple's crawler both render JavaScript, because both sit on browser infrastructure that already existed. The limit is a choice about cost, not a law, so expect it to narrow unevenly and test rather than assume. If traffic looks healthy and nothing sells, start with the traffic but no sales diagnostic.
Can it tell me whether my traffic or my site is the problem?
No. That answer lives in your analytics, split by source, and the model cannot see your analytics at all. The four checks that settle it are in is it your traffic or your website.
Are the free audit tools any better?
Better at what they cover, and narrower. Lighthouse and WAVE actually load the page instead of guessing at it, but they only test machine-checkable rules. Here is what a free website audit checks and the gap all of them leave.
Is the tool you sell any different?
It loads and renders the pages rather than guessing at them, which is the whole difference. We also published the unflattering version: the Revslip review covers what it does, what it costs and who should not buy it.