The Founders ClubLifetime access for 25 agencies - pay once, no subscription, white-label reporting.Claim your seat

Category: AI Crawler Tracking

robots.txt for AI Crawlers

robots.txt for AI crawlers uses User-agent and Disallow lines to allow or block bots like GPTBot and ClaudeBot, but the file is voluntary, so some crawlers ignore it entirely. robots.txt for AI crawlers is a plain-text file at your site root that names each AI bot in a User-agent line and uses Disallow to control […]

llms.txt Explained and How to Create One

llms.txt is a proposed plain-text file that lists your most important pages in Markdown so AI models can find clean content fast, but adoption stays limited and no major AI company has confirmed its crawlers read it. llms.txt is a plain-text file you place at the root of your website that points AI language models […]

AI Visibility Audit Using Crawler Data to Identify Gaps in 2026

TLDR: An AI visibility audit finds the gap between what Google ranks and what AI cites by reading real crawler data, so you fix the exact pages AI bots ignore instead of guessing. An AI visibility audit maps where your brand actually stands in the gap between ranking on Google and getting cited by AI. […]

AI Citation Patterns 2026: Which Bots Reference Your Content and Why

TLDR: Every AI platform has its own citation fingerprint, ChatGPT leans on Wikipedia, Perplexity on Reddit, Claude on authoritative sources, so tailor content by platform and track which bots actually crawl your pages. AI citation patterns differ by platform: ChatGPT leans on Wikipedia almost half the time, Perplexity leans on Reddit, and Claude wants authoritative, […]

PerplexityBot Explained: User Agents, Verification, and the robots.txt Controversy

PerplexityBot is Perplexity's search crawler that indexes public pages so they can be surfaced and linked in Perplexity answers, and it identifies itself as PerplexityBot/1.0 in your server logs. PerplexityBot is Perplexity's crawler for search, not for model training. It sends the user agent PerplexityBot/1.0, reads server-delivered HTML, and indexes pages so Perplexity can cite […]

ClaudeBot Explained: Anthropic's Crawlers, User Agents, and robots.txt Control

ClaudeBot is Anthropic's web crawler that collects public content to train the models behind Claude, and it identifies itself as ClaudeBot/1.0 in your server logs. ClaudeBot is Anthropic's training crawler. It sends the user agent ClaudeBot/1.0, reads the HTML your server returns, and feeds that text into the data used to train Claude. Anthropic also […]

GPTBot Explained: OpenAI's Crawler, User Agent, and How to Block It

GPTBot is OpenAI's web crawler that collects public page content to train its AI models, and it identifies itself as GPTBot/1.4 in your raw server logs. GPTBot is OpenAI's training crawler. It sends the user agent GPTBot/1.4, fetches raw HTML from public pages, and adds that text to the data pool used to train the […]

Prompt-Based vs Log-Based AI Visibility: What Actually Gets Measured

Prompt-based AI visibility samples what AI models say; log-based AI visibility counts what AI crawlers actually did on your site. Prompt-based vs log-based AI visibility is the core split in AEO tools. Prompt-based tools like Profound, Peec AI, Otterly, and Scrunch ask AI models sample questions and count brand mentions, which is sampling. Log-based tools […]