GLOSSARY
The language of AI search.
Plain-English definitions of the terms behind AI search visibility and Answer Engine Optimization.
A
- AEO (Answer Engine Optimization)
- The practice of improving how AI answer engines find, understand, trust, and cite a brand when users ask them questions. Where SEO optimizes for ranking on a results page, AEO optimizes for being named or cited in a generated answer.
- AI Answer Engine
- An AI system that responds to a question with a direct, generated answer rather than a list of links. ChatGPT, Google Gemini, Anthropic's Claude and Perplexity are the four in widest use.
- AI Citation
- When an AI answer references a specific source, page or brand as the basis for part of its answer, often with a link. Being cited is stronger than being mentioned, because it identifies you as the trusted source rather than just a name in a list.
- AI Crawler
- An automated agent operated by an AI company to fetch web pages for training, indexing or answering. Common ones include GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity) and Google-Extended (Google). Each can be allowed or blocked independently in robots.txt.
- AI Discoverability
- How findable and identifiable a brand is to AI engines: whether they can reach your content, understand what you are, and surface you in relevant answers.
- AI Mention
- When an AI answer names a brand within its response, whether or not it cites a specific source. A brand can be mentioned without being cited, and cited without being recommended.
- AI Referral Traffic
- Visitors who arrive at a site from a link inside an AI answer. Some engines pass identifiable referrer information and some send unmarked links, so this traffic is only partially attributable.
- AI Visibility
- How present and prominent a brand is across AI answer engines: how often it is named, cited or recommended when users ask questions in its category.
- Answerability
- Whether a page contains content an engine can lift a clean, direct answer from. Content written as continuous prose without clear question-and-answer structure is harder for an engine to extract.
- Attribution
- Determining which engine or source sent a visitor. In AI search this is incomplete by design: engines differ in whether they tag outbound links, so some AI-referred traffic cannot be traced to its origin.
B
- Brand Entity
- How an AI system represents a company as a distinct thing in the world, separate from its website. An engine can know a brand as an entity without having crawled its site, and can hold outdated or incorrect beliefs about it.
C
- Canonical URL
- The address a site declares as the authoritative version of a page when the same content is reachable at several URLs. Ambiguous canonicals split signals and make it harder for engines to attribute content correctly.
- Client-Side Rendering
- Building page content in the browser with JavaScript after the initial HTML loads. If an AI crawler does not execute that JavaScript, it sees an almost empty page regardless of what a human sees.
- Crawl Budget
- The practical limit on how much of a site an automated agent will fetch in a given period. Large sites with slow responses or redundant URLs can have important pages left unfetched.
E
- Entity Recognition
- An AI system identifying that a name in text refers to a specific known thing — a company, product or person — rather than a generic word. Brands with ambiguous or common names are frequently misrecognised.
G
- GEO (Generative Engine Optimization)
- Optimizing for visibility inside generative AI answers. Used largely interchangeably with AEO; GEO emphasises the generative nature of the output, AEO the answer format.
- Grounding
- Connecting a model's answer to retrieved source material rather than relying only on what it learned in training. Grounded answers are more current and usually carry citations.
H
- Hallucination
- When an AI system states something false with apparent confidence. In a brand context this often means inventing products, pricing or claims a company never made.
J
- JSON-LD
- The most widely supported format for embedding structured data in a page, written as a script block rather than woven through the HTML. The format search and AI systems generally expect.
K
- Knowledge Cutoff
- The date beyond which a model's training data ends. Anything after it is unknown to the model unless it retrieves live information at the time of answering.
- Knowledge Graph
- A structured store of entities and the relationships between them. Used to resolve what a name refers to and what is known about it, independently of any single web page.
L
- LLM (Large Language Model)
- The type of model behind current AI assistants, trained on large volumes of text to predict and generate language. It produces answers rather than retrieving documents.
- llms.txt
- A proposed plain-text file at a site's root that gives AI systems a curated map of its most important content. Adoption is not universal, but it is inexpensive to publish and explicitly aimed at machine readers.
N
- Non-Determinism
- The property that the same question can produce different answers on different runs. It means a single AI answer is a sample, not a measurement, and position should be read as a trend across repeated runs.
P
- Prompt
- The input given to an AI system. In an AEO context, the set of prompts used to test a brand should mirror what real buyers actually ask, not what the brand wishes they asked.
- Prompt Set
- A defined, repeated collection of buyer questions used to measure visibility consistently over time. Changing the prompt set changes the results, so it must be held stable to compare periods.
R
- RAG (Retrieval-Augmented Generation)
- An architecture where the system retrieves relevant documents first and generates an answer from them. It is why current, well-structured web content can influence an answer even when the underlying model was trained earlier.
- Robots.txt
- A file at a site's root that tells automated agents which paths they may fetch. AI crawlers can be permitted or blocked separately from search engine crawlers, and blocking them removes a site from consideration entirely.
S
- Schema.org
- The shared vocabulary used to describe what a page is about in a machine-readable way — an organisation, a product, a FAQ, an article. Structured data is written using its types.
- SEO (Search Engine Optimization)
- The practice of improving how a site ranks in traditional search results. It optimizes for a position in a list of links, which is a different objective from being named inside an answer.
- Server-Side Rendering
- Producing complete HTML on the server before it reaches the client. It guarantees an agent sees the actual content whether or not it runs JavaScript.
- Sitemap
- A file listing a site's URLs to help automated agents discover pages. Useful where internal linking is sparse or content is deeply nested.
- Structured Data
- Machine-readable markup that states explicitly what a page contains, rather than leaving it to be inferred from prose. It reduces the chance an engine misunderstands what a business does.
T
- Token
- The unit of text a language model processes, roughly a word fragment. Context limits are measured in tokens, which is why very long pages may be truncated before an engine reaches the end.
U
- User Agent
- The identifier an automated agent sends when requesting a page. It is how a server can tell GPTBot from a browser, and how AI crawler access is allowed or denied selectively.
Z
- Zero-Click Search
- A query answered directly on the results page or inside an assistant, with no visit to any source site. It is the mechanism by which a brand can be influential in a category while receiving no measurable traffic.
These terms are covered in depth in the AI Search Academy.
