// Technical AI Search Optimisation
A Robots.txt Line Gets You In. A Blog Post About Schema Doesn’t.
What actually controls whether AI crawlers can reach and read your site, and why most of what’s sold as “technical AI SEO” is either already solved by your CMS or isn’t a real lever at all.
Find out what your robots.txt is blockingThe Problem
Four Different Bots, Four Different Rules, One Robots.txt File
OpenAI, Anthropic, and Perplexity each run separate crawlers for model training versus search retrieval, and Google runs one crawler for everything. Getting a single token wrong in robots.txt doesn’t just risk your content being used for AI training, it can quietly remove you from being cited at all, and most sites have never actually audited which of these tokens they’re blocking.
On top of that, the crawlers that actually retrieve content for AI answers, OAI-SearchBot, ClaudeBot, PerplexityBot, don’t execute JavaScript. If your key content only exists after client-side rendering fires, these crawlers see an empty page, no matter how clean your robots.txt is.
The Folklore
Why Schema Markup and llms.txt Files Aren’t the AI-Citation Levers They’re Sold As
Two of the most heavily marketed “technical AI SEO” fixes have no documented mechanism behind them at all.
Schema.org Markup
Google’s own documentation states plainly that no special schema.org structured data is needed to appear in AI Overviews or AI Mode. The strongest controlled test available, Ahrefs tracking 1,885 pages that added JSON-LD against 4,000 matched control pages, found no meaningful citation lift: a small negative movement on AI Overviews, gains within statistical noise elsewhere.
llms.txt Files
No AI platform, OpenAI, Anthropic, Perplexity, or Google, documents llms.txt as required or beneficial. Google’s own search staff have said outright they don’t use it and aren’t planning to. Server-log analysis across hundreds of thousands of domains found the file carries no measurable predictive signal for AI visibility, and most published llms.txt files receive zero crawler requests at all.
What Actually Works
Two Things Actually Gate Your AI Visibility: Access and Rendering
A live server-log study of more than 500 million real crawler requests, tracking Googlebot, GPTBot, ClaudeBot, AppleBot, and PerplexityBot, found that none of the major AI crawlers except Google’s render JavaScript. Content that only appears after client-side rendering fires is invisible to GPTBot, ClaudeBot, and PerplexityBot, full stop.
Real-time indexing splits by ecosystem too. IndexNow pushes updates directly to Bing, which powers Microsoft Copilot and partially feeds ChatGPT and Perplexity retrieval. Google doesn’t participate in IndexNow at all, and its Indexing API is restricted to job listings and livestream events, using it for ordinary pages does nothing for AI Overviews.
Where We Fit
Where Technical AI Search Optimisation Fits
We audit which AI crawlers can actually reach and read your site, not just whether robots.txt looks tidy: which bots are blocked, whether your highest-value pages are readable before JavaScript runs, and whether your indexing pipeline reaches the platforms your buyers actually use.
What we don’t do is bill you for a schema rewrite or an llms.txt file and call it AI optimisation. The evidence doesn’t support that being the fix, and we’re not going to sell you one anyway.
What Changes
One Fix, Three Audiences
01 / The Buyer
They never see your robots.txt. They do see whether your product shows up when they ask an AI assistant instead of searching for it.
02 / The Search Engine
Googlebot already renders JavaScript and follows your existing SEO investment. Nothing here changes that relationship, this is about the crawlers that don’t.
03 / The AI Agent
A crawler that can’t execute your JavaScript and isn’t explicitly allowed in your robots.txt can’t recommend you, no matter how good your content is.
Questions
Frequently Asked
What does a Technical AI Search Optimisation audit actually check?
Which AI crawlers, GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, and Googlebot, can reach your site, whether your key content is visible before JavaScript executes, and whether your indexing setup reaches Bing as well as Google.
Will this get schema markup added to our site?
Where it earns something real, classic rich results and shopping feeds, yes. We won’t sell it to you as a fix for AI citation, because the best available controlled test found no meaningful citation lift from adding it.
Do we need an llms.txt file?
No platform documents it as required or beneficial, and Google’s own staff have said they don’t use it. We won’t build one and bill you for a visibility improvement that isn’t proven to exist.
Is this the same as Answerability, Entity and Brand Optimisation, Content Architecture, or Measurement?
No. This is infrastructure-level: can AI crawlers physically access and read your site at all. Answerability, Entity and Brand Optimisation, Content Architecture for AI Citation, and AI Search Performance Measurement all assume that access already works and focus on what happens once a crawler can read your content, or on measuring the results.
When is this not the right fit?
If your site is fully server-rendered already with no history of blocking crawlers, there may be very little to fix here, and confirming that takes an afternoon, not a full engagement. This is most valuable for sites built on JavaScript-heavy frameworks, or anyone who has never actually checked which AI bots their robots.txt blocks.
Find Out What Your Robots.txt Is Actually Blocking
We’ll show you exactly which AI crawlers can reach your site today, and what’s invisible to them, before you spend anything on schema or content.
Register interest