LLMs are now a meaningful audience
Discovery no longer happens only through a list of blue links. Buyers ask Google AI Overviews, AI Mode, ChatGPT Search, Perplexity, Microsoft Copilot, and other assistants to compress research into a usable answer. They ask who does the work, which provider looks credible, what a service includes, and what makes one option safer than another.
That changes the job of your website. It still has to persuade humans, and it also has to give AI-assisted systems enough clear, crawlable evidence to represent the business accurately. The commercial risk includes omission, misunderstanding, reduction to a generic category, and comparison against competitors whose websites explain their services and proof more precisely.
Traffic alone is a weaker signal in this environment. SparkToro reported that in 2026 less than one third of Google searches still send a click, which reinforces the need to measure visibility through branded search, citations, referral quality, qualified leads, and buyer confidence as well as organic sessions (SparkToro).
The practical goal is simple. The right buyer should be able to find, understand, trust, and compare your business. AI systems can only help with that when the underlying website is clear, current, and evidenced.
How LLMs read your website
Language models and AI search systems do not experience a website the way a human does. They may crawl pages, render HTML, extract text, follow links, query search indexes, retrieve documents, and summarise the pieces that look relevant. The exact process depends on the platform, but the foundation is familiar.
Google's own guidance for AI Overviews and AI Mode says there are no additional technical requirements, no special AI text files, and no special schema.org markup required to appear in those features. Google points site owners back to normal Search eligibility, crawlability, indexability, snippet controls, internal links, textual content, and structured data that matches visible content (Google Search Central).
That means the first layer of LLM readiness is still good web publishing:
- Crawlable and indexable pages
- Important content available as text
- Semantic HTML with logical headings
- Descriptive internal links
- Clear entity information about the business, services, people, locations, and proof
- Structured data that matches the page
- Evidence-led content that goes beyond generic claims
- Thoughtful crawler access controls
- Optional llms.txt and markdown support
- Measurement across visibility, leads, and commercial quality
Google's generative AI guidance also emphasises unique, valuable, expert-led content that goes beyond common knowledge (Google Search Central). That's important because generic content gives answer systems little reason to cite or trust one business over another.
For Off Piste, this is where website design and SEO overlap. Site architecture, page templates, content structure, technical search foundations, and measurement all shape whether a business is legible to people and machines.
Crawler access comes before llms.txt
Before adding support files, decide what AI systems may access and why. Different bots have different jobs. Search visibility, model training, and user-triggered retrieval should not be treated as one setting.
OpenAI separates OAI-SearchBot for ChatGPT search, GPTBot for training-related crawling, and ChatGPT-User for user-triggered actions. Its crawler documentation gives each crawler separate robots.txt handling (OpenAI).
Perplexity makes a similar distinction. PerplexityBot supports website surfacing in Perplexity search results, while Perplexity-User handles user-triggered fetches. Perplexity's crawler documentation also recommends using its current published IP ranges when teams manage access through a WAF or firewall. The Perplexity search visibility guide turns that platform-specific access check into a wider diagnosis of absence, inaccurate representation, weak citations, and poor commercial outcomes.
Infrastructure teams now have more operational tooling for this. Cloudflare's AI Crawl Control lets site owners monitor AI crawler activity, set crawler-level allow or block rules, monitor robots.txt compliance, and explore pay-per-crawl options (Cloudflare).
| Crawler or token | Main role | Visibility relevance | Training relevance | User-triggered relevance | Control surface |
|---|---|---|---|---|---|
Googlebot |
Google Search crawling | High for Google Search and AI features built into Search | Not the training control | No | robots.txt, noindex, snippet controls, Search Console |
Google-Extended |
Google product token for AI training and grounding controls | Google says Search inclusion and ranking continue to use separate systems | High | No | robots.txt product token |
OAI-SearchBot |
ChatGPT search crawling | High for ChatGPT search | No | No | robots.txt |
GPTBot |
OpenAI training-related crawling | Indirect | High | No | robots.txt |
ChatGPT-User |
User-triggered OpenAI fetches | Depends on user requests | No | High | robots.txt and server access rules |
PerplexityBot |
Perplexity search surfacing | High for Perplexity | No | No | robots.txt and WAF rules |
Perplexity-User |
User-triggered Perplexity fetches | Depends on user requests | No | High | robots.txt and WAF rules |
Google also documents that Google-Extended is a robots.txt product token rather than a separate HTTP user agent string, with Search crawling and ranking handled separately (Google Crawling Infrastructure).
The practical rule is to document the access decision, configure robots.txt and WAF rules carefully, and monitor logs. Blanket blocking can affect search surfacing, user-requested retrieval, or training use depending on the bot involved.
The llms.txt standard
An llms.txt file is a plain markdown file hosted at the root of your domain. It gives AI agents, assistants, and answer engines a curated guide to the pages and facts you consider most important.
The concept was proposed by Jeremy Howard from Answer.AI in September 2024. It has become a recognised convention in parts of the web and developer ecosystem. Its current role is best understood as a voluntary support format rather than an official search standard or ranking signal.
The useful way to think about llms.txt is as one maintained layer in a wider visibility system. It can clarify positioning, services, canonical URLs, and priority pages. Its value is strongest when it points to crawlable pages, strong content, structured data, and search-ready technical foundations.
That distinction matters because recent evidence is cautious. Ahrefs analysed 137,189 websites with valid llms.txt files and found that 97% received zero requests during May 2026 (Ahrefs). Contentful also argues that there is no validated evidence that llms.txt reliably improves AI citation frequency or search visibility (Contentful).
Here is a trimmed version of Off Piste Studio's own file. It reflects the current positioning in our live llms.txt and points agents to the fuller llms-full.txt context.
> AI-native design and technology studio helping ambitious
businesses turn expertise into trust, clarity and growth.
This file is a compact discovery index for AI agents, assistants and answer engines. For complete structured context, fetch:
- [Full agent context](https://offpistestudio.com/llms-full.txt): machine-readable company profile, services, audience, availability, pricing guidance, canonical URLs and usage notes.
## Core Pages
- [Home](https://offpistestudio.com/): primary positioning, services, selected work, proof and calls to action. - [Work](https://offpistestudio.com/work): selected projects and examples of client work. - [About](https://offpistestudio.com/about): studio positioning, team and approach. - [Insights](https://offpistestudio.com/resources): articles and search, AI and design thinking.
## Preferred Positioning
- Preferred summary: Off Piste Studio is an AI-native design and technology studio for ambitious businesses. - The studio works across brand, websites, content structure, search and AI-ready digital systems. - Pricing is scoped per project after discovery, with recommendations based on commercial fit and project complexity.
How it differs from robots.txt and sitemap.xml
These files are easy to confuse because they sit near the root of a website and speak to crawlers or machines. They do different jobs.
| Format | Purpose | Primary audience | Strength | Limitation | Maintenance owner |
|---|---|---|---|---|---|
sitemap.xml |
Lists canonical URLs for discovery | Search engines | Helps crawlers find pages | Does not explain priority, meaning, or proof | SEO or development |
robots.txt |
Gives crawler access instructions | Crawlers and bots | Controls allowed and disallowed paths | Cannot make weak content useful or visible | SEO, development, infrastructure |
llms.txt |
Curates important pages and descriptions | AI agents and answer engines that choose to fetch it | Clarifies priority content and positioning | Direct citation impact is unproven | Content, SEO, strategy |
llms-full.txt |
Provides fuller machine-readable context | AI agents and internal workflows | Gives a complete reference in one file | Can become stale if not generated or maintained | Content and development |
| Page-level markdown exports | Presents individual pages in clean markdown | Agents, RAG systems, internal tools | Reduces extraction friction where supported | Must stay aligned with canonical HTML | Development and content |
Each asset works best when the underlying page already carries the weight. A sitemap helps crawlers find the page. A robots.txt file gives access instructions. An llms.txt file clarifies priority and meaning. The page still needs to explain the claim, show the proof, and earn the trust.
File format and structure
The format is deliberately simple. The only required element is an H1 title. A useful file usually includes a short blockquote summary, followed by H2 sections that group priority links with clear descriptions.
The descriptions are the work. They should name services, audiences, locations, proof, and page purpose without stuffing every keyword into the file. A good llms.txt file tells an agent which pages matter and what it will find there. A weak one becomes another generic index.
Review the file whenever service positioning, pricing guidance, priority pages, or canonical URLs change. If the website has moved from local service language into a broader strategic design and technology position, the machine-readable summary needs to move with it.
The three related formats
There are three related support formats worth separating.
llms.txt is the compact guide. It should stay concise enough to scan and maintain.
llms-full.txt is the fuller reference. It can compile company context, services, audience fit, canonical URLs, and usage notes into one machine-readable document. It's useful when the team can keep it current, while evidence for universal agent preference remains limited.
Page-level markdown exports give individual pages a clean text version. They can be useful for internal tools, retrieval systems, and agents that can request markdown. They're only worth doing when the site can generate them reliably and keep them in sync with the canonical HTML.
Why it matters commercially
The commercial value is accuracy. AI systems can only work with the evidence they can retrieve and interpret. If your offer is vague, your service pages are thin, your proof is scattered, or your crawler rules are inconsistent, the system has less to use when a buyer asks who to trust.
Machine-readable summaries help when they are grounded in strong pages. They can clarify what the business does, who it serves, what pages are canonical, and which claims should not be inferred from old content. That's useful operational housekeeping, especially for a business whose positioning has changed.
Proof still carries the weight. Case studies, named services, outcomes, process, authorship, dates, credentials, FAQs, and third-party references give search systems and buyers something real to evaluate. The support files should point to that evidence and keep the important paths clear.
Where llms.txt fits
llms.txt works best as a maintained guide to priority content. It helps clarify which pages matter, what the business does, and which canonical URLs an agent should start with.
Its role is strongest when it points to strong underlying pages. Search visibility still depends on crawlable content, useful evidence, structured pages, current information, and platform-specific eligibility. Keep the file factual, keep it current, and treat it as one support layer inside a broader search and AI discovery system.
Structured content and machine-readable formats
Beyond llms.txt, broader publishing decisions shape how well your site works for search systems and language models.
Clean, accessible HTML
Clean HTML and accessible structure come first. Semantic elements, heading order, descriptive links, useful alt text, readable copy, and labelled forms all make a page easier to parse. They also help people navigate the site with less friction.
The overlap with website accessibility and SEO is direct. The same structural clarity that helps assistive technology move through a page also helps crawlers and extraction systems understand what each section means.
For the implementation layer, use our structured content for AI search guide to audit page anatomy, schema alignment, proof blocks, internal links, metadata, and agent-readable interaction patterns.
Schema markup
Structured data using schema.org vocabulary can help systems understand page type, authorship, organisation details, services, reviews, FAQs, and article metadata. The important word is "matching". Google specifically says structured data should match the visible text on the page and that no special schema is required for AI Overviews or AI Mode (Google Search Central).
Use Organization, Service, Article, FAQPage, Review, and LocalBusiness schema where the page genuinely supports it. Schema should describe the visible page faithfully, so the structured data and human content stay aligned.
Markdown exports
Clean markdown versions can reduce extraction friction for tools that use them. They work best when generated from the same source as the canonical page, so the markdown version stays aligned with the live content.
For many businesses, markdown exports should come after the basics. Make the HTML page clear, indexable, internally linked, and useful first. Then add machine-readable formats where the workflow can maintain them properly.
What this means for content strategy
Writing for humans and language models rewards the same discipline. The content that works best for both is clear, specific, well-organised, and supported by evidence.
Vague brand language creates problems for both audiences. A human scanning the page struggles to understand the offer. A model trying to summarise the business has to fill gaps. Specific language gives both audiences something reliable to work with.
Specific language gives buyers and models a firmer representation of the business. For example, "we combine brand strategy, website design, UX/UI, content structure, and technical SEO to help expertise-led businesses clarify value and convert the right attention into action" carries more useful meaning than a broad claim about digital experience.
Proof architecture matters. Strong pages name the service, audience, problem, process, outcome, author, date, and evidence. Case studies with specific context are more useful than generic testimonials. FAQs are useful when they answer buyer questions directly rather than repeating keywords.
If the audit shows that pages are accessible but too generic to trust, the next layer is citation-worthy content. That work turns broad claims into supported answers, first-hand proof, source context, and examples that are easier for buyers and AI search systems to evaluate.
If the audit shows that AI systems understand the category but misread the business, the repair path is different. Use the AI entity trust audit to check whether the name, services, audience, locations, profiles, proof, and third-party sources tell one consistent story.
For Google specifically, this connects to the broader shift described in our article on AI Overviews and SEO. Generic information is easier to compress. Specific proof, commercial judgement, and useful comparison context are harder to replace.
Getting started
Most businesses should build LLM readiness in four phases.
First, confirm the foundations. Check crawlability, indexability, page structure, and whether important content is available as text.
Second, strengthen the evidence. Update service pages, proof, pricing guidance, FAQs, author details, and internal links.
Third, manage machine access. Review robots.txt, CDN rules, WAF rules, crawler logs, and AI crawler permissions. If you need the operational version of that work, use our AI crawler access and robots.txt guide to separate search crawlers, training crawlers, and user-triggered fetchers before changing access rules.
Fourth, add maintained support files. Create or refresh llms.txt, then add llms-full.txt or markdown exports where the team can keep them aligned with the site.
A website becomes easier to recommend when its claims, structure, and evidence stay current. The maintenance is the point. AI search visibility comes from publishing a business clearly enough that people and machines can understand why it should be trusted.
