Glossary
The AI visibility glossary.
Every AI-search term that matters: AEO, GEO, llms.txt, GPTBot, structured data, defined in one plain, correct sentence you can actually use. No hype; where a term is oversold (looking at you, llms.txt), we say so.
28 terms · Reviewed July 14, 2026
Core concepts
- AI visibility#
- AI visibility is whether AI assistants, ChatGPT, Perplexity, Claude and Google’s AI Overviews, can read your website, understand it, and cite your business when someone asks a relevant question. It combines classic SEO with answer-first content, structured data, AI-crawler access and off-site trust signals.
- Answer Engine OptimizationAEO#
- Answer Engine Optimization (AEO) is the practice of structuring content so a search or answer engine can extract it as the direct answer: a featured snippet, a voice response, or a Google AI Overview, rather than just ranking it as a link. It relies on answer-first writing, question-shaped headings and structured data.
- Generative Engine OptimizationGEO#
- Generative Engine Optimization (GEO) is the practice of making a website one of the sources that generative AI assistants quote and cite. The term comes from a 2023 Princeton-led study which found that adding statistics, quotations and cited sources measurably increased a page’s visibility in AI answers.
- Search Engine OptimizationSEO#
- Search Engine Optimization (SEO) is the practice of improving a website so it ranks higher in search engine results. In the AI era it is the shared foundation for AEO and GEO, Google itself states that optimizing for AI search is still, fundamentally, SEO.
- AI citation#
- An AI citation is when an AI assistant names or links to your website as a source in its answer. Unlike a search ranking, a citation makes your business the quoted authority, and citations are chosen largely on readability, extractable structure and off-site corroboration rather than on Google position.
- Answer-first content#
- Answer-first content opens a page or section with a direct, self-contained answer to the exact question a reader asked, before any preamble. It is the single most important on-page pattern for AEO and GEO, because answer engines extract clean, standalone answers and ignore buried ones.
- E-E-A-T#
- E-E-A-T stands for Experience, Expertise, Authoritativeness and Trustworthiness: the qualities Google’s raters (and, increasingly, AI systems) use to judge whether a source is reliable. Named authorship, cited claims, transparent methodology and a consistent entity across the web all raise it.
Crawlers & access
- AI crawler#
- An AI crawler is an automated bot that fetches web pages on behalf of an AI company, either to train a model (a training crawler) or to answer a user’s question in real time (a retrieval crawler). If an AI crawler cannot access your site, the assistant it feeds cannot read or cite you.
- GPTBot#
- GPTBot is OpenAI’s training crawler, it downloads public web content that may be used to train OpenAI’s models. It is separate from OpenAI’s retrieval crawlers (OAI-SearchBot and ChatGPT-User), so a site can allow one and block the other via robots.txt.
- OAI-SearchBot#
- OAI-SearchBot is OpenAI’s crawler for ChatGPT Search, it indexes pages so they can appear as cited results in ChatGPT’s search feature. Because it drives live citations with clickable links, it is one of the highest-priority crawlers to allow for AI visibility.
- ChatGPT-User#
- ChatGPT-User is the user-agent OpenAI uses when ChatGPT fetches a specific page live to answer a user’s question. It represents real-time demand: a person is asking something your page could answer right now, so blocking it removes you from those live answers.
- ClaudeBot#
- ClaudeBot is Anthropic’s crawler for Claude. Anthropic also operates Claude-User and Claude-SearchBot for live retrieval. Anthropic states it respects robots.txt, so a site controls access to Claude through its robots directives.
- PerplexityBot#
- PerplexityBot is the crawler for the Perplexity answer engine, which indexes pages so they can be cited in Perplexity’s answers. Perplexity is notably citation-heavy and favors recently-updated content, so freshness matters more here than on most engines.
- Google-Extended#
- Google-Extended is a robots.txt token that controls whether your content may be used to train Google’s Gemini models and improve its generative features. It does not affect classic Googlebot indexing, blocking it keeps you in Search while opting out of AI training.
- CCBot#
- CCBot is the crawler for Common Crawl, an open repository of web data that many AI models are trained on. Because so many models draw on Common Crawl, allowing CCBot increases the chance your content is present in their training data.
- robots.txt#
- robots.txt is a plain-text file at the root of a website that tells crawlers which parts they may access. The major AI operators publicly honor it, so it is the primary lever for allowing or blocking AI crawlers, but it cannot override a block applied at the CDN or firewall layer.
- llms.txt#
- llms.txt is a proposed plain-text file that lists a site’s key pages in an LLM-friendly form, like a curated sitemap for AI. As of 2026 no major AI platform has committed to reading it, and an Ahrefs study of 137,000 domains (May 2026) found 97% of them received zero AI requests, so it is best treated as low-cost hygiene, not a proven citation lever.
Technical & structured data
- Structured data#
- Structured data is machine-readable markup that labels the meaning of content on a page. This is the author, this is a question and its answer, this is the price. It helps search engines and AI parse and attribute content, but it supports good content rather than replacing it.
- Schema.org#
- Schema.org is a shared vocabulary of structured-data types (Article, FAQPage, Organization, Person and hundreds more) maintained by Google, Microsoft, Yahoo and Yandex. It is the standard vocabulary used to describe web content to machines.
- JSON-LD#
- JSON-LD (JavaScript Object Notation for Linked Data) is the recommended format for adding schema.org structured data to a page: a script block of JSON that describes the page’s entities. Using stable @id references, JSON-LD lets separate blocks link into one coherent entity graph.
- Server-side renderingSSR#
- Server-side rendering (SSR) means the server returns fully-formed HTML on the first request, rather than sending an empty shell that assembles content with JavaScript in the browser. It matters for AI visibility because many AI crawlers do not run JavaScript and see a client-rendered page as blank.
- FAQPage schema#
- FAQPage is a schema.org type that marks up a list of questions and their answers so machines can extract them cleanly. It maps directly onto how AI assistants pull question-and-answer content, which is why answer-first pages pair it with visible FAQ sections.
- Speakable#
- Speakable is a schema.org specification that marks the parts of a page best suited to be read aloud by a voice assistant, using CSS selectors. It signals which sentences are the clean, self-contained answer, useful for voice search and answer extraction.
- IndexNow#
- IndexNow is a protocol that lets a website instantly notify participating search engines (including Bing and Yandex) that a URL has been added or changed, so they recrawl it quickly instead of waiting to rediscover it. It speeds up how fast fresh content is picked up.
AI engines & surfaces
- Google AI Overviews#
- Google AI Overviews are AI-generated summaries shown at the top of Google results for many queries, assembled from and citing multiple web sources. They have measurably reduced clicks to traditional links, which is why being cited inside the overview now matters as much as ranking below it.
- ChatGPT Search#
- ChatGPT Search is OpenAI’s feature that lets ChatGPT retrieve and cite live web results within a conversation, fed by OAI-SearchBot and ChatGPT-User. It turns ChatGPT into a discovery channel where being a cited source can send real referral traffic.
- Zero-click search#
- A zero-click search is one where the user gets their answer directly on the results page, from an AI Overview, featured snippet or knowledge panel, without clicking through to a website. Zero-click behavior has risen sharply with AI Overviews, making citation within the answer the goal, not just the click.
- Retrieval-augmented generationRAG#
- Retrieval-augmented generation (RAG) is the technique behind most AI search: the model retrieves relevant documents from the live web (or an index) and generates its answer grounded in them, citing sources. It is why being crawlable and extractable determines whether you can be part of the answer at all.
Cite or link this glossary
Found this useful? Link to it.
This glossary is free to quote and reference. A link helps more people find accurate, evidence-cited answers, and you're welcome to reuse it with attribution under CC BY 4.0. Grab a ready-made link or citation:
<a href="https://footnoteworks.com/glossary">AI Visibility Glossary, Footnote Works</a>
[AI Visibility Glossary, Footnote Works](https://footnoteworks.com/glossary)
Footnote Works. “AI Visibility Glossary.” Footnote Works, 2026, https://footnoteworks.com/glossary.
Enough theory
See how your own site scores.
The glossary explains the terms; the free scan shows you which of them are costing you. It checks crawler access, structured data and rendering the way an AI crawler would.