Updated for GA4's AI Assistant channel, May 2026.
Say your team pulled the logs. GPTBot hit 40,000 pages last month. ChatGPT-User showed up 200 times. PerplexityBot crawled your pricing page nine times this week. Now what?
The scale behind those numbers isn't hypothetical. Fastly's traffic research puts automated requests at 37% of all web activity, and AI crawlers now account for roughly 80% of that share. Radware's telemetry shows AI crawler volume has grown to nearly match Googlebot and Bingbot combined. This is a line item in your traffic, not a rounding error.
Most AI-visibility advice stops at finding this data. GA4 can't see any of it, crawlers skip the browser entirely, so no tag ever fires, which means someone has to pull it from your server or CDN logs instead. We've written a full step-by-step walkthrough for that in the AI Visibility Playbook, or you can skip the manual pull entirely and let Hardal track it for you, continuously, with no log-diving required. This post is about what comes after: once the report lands in your inbox, what do you do with it?
Start With What Kind of Bot It Is
Not every hit means the same thing, and the decision you make depends on which bucket it falls into:
| Type | What it means | Example bots |
|---|---|---|
| Trainers | Your content is becoming training data for a future model | GPTBot, ClaudeBot, Bytespider |
| Searchers | Your pages are being indexed for live AI citations | OAI-SearchBot, PerplexityBot |
| Fetchers | Someone just asked, and the AI is reading your page right now | ChatGPT-User, Claude-User |
A thousand Trainer hits and a thousand Fetcher hits should trigger completely different responses. Treating them as one number called "AI traffic" is where most of this data goes to waste.
Build a Content Priority List
Sort your crawled pages by hit count and you get a ranked list of what AI models currently treat as your most important content, whether you intended that or not. Pages Trainers return to often are already shaping how a model describes your product, which makes them the highest-leverage place to spend time on accuracy and clarity. If your pricing or comparison pages top that list, keep them current. If a three-year-old blog post outranks them, that's worth knowing too.
Find Your Invisible Pages
Cross-reference the crawl list against the pages you actually want AI to cite: product pages, comparison pages, documentation. Anything with zero Searcher hits isn't in contention to show up in an AI answer, no matter how good the page reads to a human. That's a fix-it list, not a vague SEO concern. AI-referred sessions grew 527% year over year, per Search Engine Land, so a page that's invisible to Searchers today is missing out on a channel that's compounding, not a nice-to-have. Tighten the copy, add internal links pointing to those pages, and check again next month.
Treat Fetcher Spikes as Live Intent
A Trainer hit means a model read your page at some point during a crawl. A Fetcher hit means someone typed a question into ChatGPT or Perplexity a few seconds ago, and the AI went and read your page to answer it. That's closer to a hot lead than a pageview. Fastly clocked ChatGPT's fetcher bots alone generating 98% of all real-time AI retrieval requests, peaking at 39,000 requests a minute, so this traffic is bursty and it's already the dominant form of live AI activity hitting most sites. If you already track brand mentions or alert on branded search terms, add Fetcher spikes to the same watchlist. A jump in ChatGPT-User traffic to your pricing page deserves same-day attention, not a mention in next month's report.
Set robots.txt Rules on Purpose
Training inclusion and citation visibility are separate trade-offs, and robots.txt doesn't have to make the same call for both. You can block GPTBot if you'd rather your content not train the next model, while still allowing OAI-SearchBot and PerplexityBot so you stay eligible to be cited in live answers. Blocking every AI bot by default trades away the second thing to guard against the first, and they were never bundled to begin with.
Put It on a Recurring Calendar
New bots ship often enough that a one-time check goes stale within a quarter. Fold this into whatever reporting cadence you already run: a monthly pull, reviewed next to GA4 and Search Console, tracking crawl volume by bot, top crawled pages, and any Fetcher spikes. It doesn't need its own meeting. It needs a recurring line in a report that already exists.
Mistakes That Waste the Data
- Reading "AI traffic" as one number instead of three signals that each call for a different response.
- Blocking every AI bot out of caution, which trades away citation visibility to guard against training inclusion, when you can set that per bot instead.
- Pulling the report once and never checking again. The bot list and the crawl volumes both move fast.
- Reacting to Fetcher spikes on a monthly cadence instead of treating them as the near-real-time signal they are.
- Copying robots.txt rules from a generic list instead of reading your own crawl data. The mix of bots hitting your site is specific to you.
The Takeaway
AI crawler data is only as useful as the decisions it changes. Sort it by bot type, use it to prioritize content and find gaps, treat Fetcher spikes like the intent signal they are, and put the whole thing on a calendar so it doesn't go stale.
Ready to see this for yourself?
- Check your site free at isvisible.ai, no signup, results in seconds
- See the Hardal AI Visibility product page for what continuous tracking looks like
- Read the AI Visibility Playbook for the full GA4 and server-log walkthrough
- Join the Hardal AI Visibility waitlist to track both automatically