We Audited 50 Sites for AI Visibility — What We Found
Nobody is blocking the AI crawlers. The real gap is that a third of these homepages give an answer engine nothing structured to read.
Photo by Chris Ried on Unsplash
On August 27, 2026 we ran a machine-readable audit of fifty well-known B2B websites to see how ready they are for AI search. The headline is not what we expected. The fight the industry keeps bracing for — whether to shut AI crawlers out — is effectively over, and almost nobody turned up to fight it. Not one site in the sample blocked GPTBot, ClaudeBot or PerplexityBot. What we found instead was quieter and more damaging: a third of these homepages hand an AI system nothing structured to read at all.
What We Measured, and How
We picked fifty companies across four groups: twenty marketing and sales software firms, ten operations, HR and fintech platforms, ten developer-infrastructure and payments companies, and ten professional-services firms. All are established brands with real marketing budgets. If anyone has had the time and money to prepare for AI search, this cohort has.
For each domain we checked four things a machine can verify without judgment calls.
First, the robots.txt file, and specifically what it says about thirteen named AI crawlers — among them GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and Bytespider. Second, whether the site serves a real llms.txt, the plain-text summary file some teams now publish for language models. Third, the JSON-LD structured data on the homepage: does it declare an Organization, a WebSite, a breadcrumb trail, an FAQ. Fourth, two pieces of ancient on-page hygiene that still matter — a single h1 heading and a meta description.
Everything below is a direct observation from that run. Where a site behaved inconsistently we re-tested until the result held.
Finding One: The Crawler Blocking Debate Is Theatre
Forty-five of the fifty served us a readable robots.txt. Of those forty-five, zero blocked GPTBot. Zero blocked ClaudeBot. Zero blocked PerplexityBot, OAI-SearchBot, ChatGPT-User or Google-Extended. One site turned away CCBot, the Common Crawl bot. One blocked Bytespider. One blocked Meta's agent. That is the entire resistance.
More telling: only ten of the forty-five mention any AI crawler by name. The other thirty-five say nothing whatsoever about them, which means those bots inherit whatever the wildcard rule allows — and in every single case, the wildcard let them through.
So the default is wide open, and most companies got there by accident rather than by decision. That is worth sitting with. Two years of conference panels about whether to let the models train on your content, and the practical answer at these fifty firms is: nobody set a policy, so everything is permitted.
Five sites never gave us their robots.txt at all. Four returned a 403 to automated requests and one returned a 404. A site that refuses to serve robots.txt to a non-browser client is making a bet about which visitors matter, and that bet now includes the crawlers behind AI answers.
Finding Two: Two-Thirds Publish llms.txt, and the Evidence Says It Does Little
Thirty-four of the fifty sites — 68 percent — serve a genuine llms.txt. The median file runs about 15KB and lists roughly 85 links. Thirty-two of the thirty-four are structured link indexes; two are prose with no links at all, which defeats the purpose.
Set that against the wider web. SE Ranking crawled roughly 300,000 domains and found llms.txt on 10.13 percent of them, work covered by Search Engine Journal in November 2025. Our cohort adopts the file at nearly seven times the rate of the general web.
Here is the awkward part. That same SE Ranking analysis found no relationship between publishing the file and being cited by AI models. Google has said its Search systems ignore llms.txt. No major model provider has committed to reading it in production answers. A server-log study published in April 2026, covering around 900 monitored domains, recorded 1,227 requests for llms.txt files — and not one came from a frontier AI lab's crawler. Most requests came from a data aggregator and from humans in Chrome.
One site in our sample returns an HTML page at /llms.txt. It looks like a file to a casual check and is worthless to a parser. That is the whole trend in miniature: the artifact got adopted, the substance did not.
Finding Three: A Third of Homepages Carry No Structured Data
We fetched forty-nine homepages. One firm refuses automated requests to its homepage while still serving its llms.txt happily.
Sixteen of those forty-nine — nearly a third — contain no JSON-LD structured data of any kind. Thirty-one declare an Organization, so 63 percent tell a machine, in a format built for machines, who they actually are. Sixteen declare a WebSite. Five publish a breadcrumb trail. Four use FAQ markup.
That last number deserves a second look. Four sites out of forty-nine, in a sample stuffed with companies whose own blogs preach about AI search, use the one markup format explicitly designed to express a question and its answer.
Finding Four: The Fundamentals Are Worse Than the Fashionable Stuff
Nine of the forty-nine homepages have no h1 at all. Not a wrong h1 — none. Six have more than one, and a well-known email platform ships six of them on a single page.
Eleven have no meta description. Among the ones that do, lengths run from 7 characters to 203.
Think about what an extraction system does with that. It arrives, looks for the clearest statement of what this company is, and finds a hero animation with no heading and no summary. It can still read the rendered text, but you have removed every signal you controlled and left the machine to guess.
Finding Five: Only Fifteen of Fifty Clear a Low Bar
We set a deliberately modest test. Does the homepage declare an Organization, carry exactly one h1, include a meta description, and serve a working llms.txt? Four things, all of them a morning's work.
Fifteen of fifty passed. Drop llms.txt from the test and twenty-seven pass. Either way, roughly half this cohort of well-funded brands fails a check that any competent developer could clear before lunch.
What We Think This Actually Means
The obvious read is "add schema, add llms.txt, win AI search." We do not believe that, and the evidence does not support it.
Ahrefs tracked 1,885 pages that added JSON-LD markup and reported in May 2026 that AI citations barely moved: a 2.4 percent change in AI Mode and 2.2 percent in ChatGPT, both indistinguishable from noise, alongside a small decline in AI Overviews. The llms.txt research points the same way. If you are hoping a markup sprint will make ChatGPT recommend you, the data says no.
So why does any of this matter? Because these signals are not levers, they are hygiene — and hygiene is how you find out whether the underlying facts exist at all. A team that cannot produce an Organization block usually cannot answer the question underneath it: what do we do, who exactly for, and what makes us the right call. The markup is downstream of clarity. Its absence is a symptom.
The stakes are real even if the fix is unglamorous. Pew Research Center tracked the browsing of over 900 US adults and found that when a Google AI summary appeared, people clicked a search result 8 percent of the time, against 15 percent when no summary was present. They clicked a link inside the summary in 1 percent of visits. The traffic does not come back. Being described accurately inside the answer is what is left.
That is why our own audit work starts with the claims on the page, not the markup around them. If a buyer cannot extract your pricing model, your specialty and your proof from plain text, no structured data will rescue you. Structure a page so the answer comes first and the machine has less room to guess.
How to Run This on Your Own Site This Afternoon
Pull your robots.txt and read it properly. Find out whether you have a deliberate position on AI crawlers or an inherited one. Both are defensible; not knowing which you have is not.
Fetch your own /llms.txt and confirm it returns plain text rather than an HTML error page. If you do not have one, that is a reasonable choice on current evidence, and lower priority than everything else here.
View source on your homepage and search for application/ld+json. Confirm there is an Organization block and that the details in it match your footer, your Google Business Profile and your LinkedIn page. Contradictions between those sources are a genuine problem.
Count your h1 tags. You want one, and it should say what you do rather than what you feel.
Read your meta description out loud. If it does not name the buyer and the outcome, rewrite it.
Then do the part that matters more than all of it: open your five highest-intent pages and check whether a stranger could extract your specialty, your pricing shape and one piece of verifiable proof without scrolling. Our guide to getting cited by ChatGPT, Perplexity and AI Overviews covers that process in detail, and our primer on generative engine optimization explains where it sits inside search strategy.
What This Study Does Not Show
We measured machine-readable eligibility, not outcomes. We did not ask any assistant whether it recommends these brands, and a site can score badly here and still get named constantly because its reputation is built elsewhere. Reputation beats markup, every time.
This is also one point in time, one page per domain, and one requesting client. Sites behave differently for different visitors, as the 403 responses showed. A fuller audit would test many pages, repeat over weeks, and compare against actual assistant answers.
We publish the limits because a study that only flatters its own method is marketing, not research.
Frequently Asked Questions
Should we block AI crawlers?
For most B2B companies, no. Blocking removes you from the systems buyers increasingly use to build shortlists, and our sample shows your competitors are not blocking either. The stronger move is to decide deliberately and write it down.
Do we need an llms.txt file?
On current evidence it is optional. Adoption is high among large brands, but no major provider has committed to reading it, and the crawl logs show almost no frontier bot requesting it. Build one if it is cheap. Do not treat it as a strategy.
Does schema markup get us cited?
Not on its own. Ahrefs found little citation movement from adding markup. Schema helps machines parse facts you have already made clear, which is worth doing for its own sake.
Which fix gives the most return?
Answer-first writing on the pages that carry buying intent. Lead each page with a direct statement of what you do and who it serves, then support it with specifics a system can quote.
How often should we re-run this check?
Quarterly for the technical signals, and after every significant site release. The AI surfaces change faster than that, but your own signals do not need weekly attention.
Where does this fit in a wider marketing audit?
It belongs alongside demand, positioning and channel economics rather than in a silo. Our marketing audit covers the technical layer and the commercial one together, because a site that reads well to a model and badly to a buyer has solved the wrong problem.
Want This Run on Your Own Site?
Our marketing audit covers the technical AI-visibility layer and the commercial one together, so you find out what buyers and answer engines can actually extract about you.
Schedule a Strategy Call