AhrefsBot is the web crawler Ahrefs runs to build its backlink index and power Yep, its search engine. It fetches public pages, reads the links on them, and adds what it finds to the database behind Site Explorer, Content Explorer and Ahrefs' backlink reports.
It is an SEO tool crawler. It is not an LLM training bot, not a search engine indexing your site for Google, and not a human visitor. That distinction decides what you do with the traffic.
AhrefsBot at a glance
| Detail | Answer |
|---|---|
| Bot name | AhrefsBot |
| Run by | Ahrefs Pte Ltd |
| Current user agent | Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/) |
| Powers | Ahrefs' link index and the Yep search engine |
| Volume | Up to 8 billion pages a day |
| Rank | The most active crawler in the SEO category |
| Respects robots.txt | Yes, including Allow, Disallow and Crawl-Delay |
| Renders JavaScript | Yes, so it also requests scripts, images and stylesheets |
| Category in Hardal | SEO tools, never a conversion source |
Ahrefs publishes the live bot details, IP lists and crawl controls on its official crawler page. Use that page as the source of truth when a user-agent version changes.
The AhrefsBot user agent string
Ahrefs runs two crawlers, and people mix them up constantly. Each has variants.
Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)
There is a news variant that carries an extra token:
Mozilla/5.0 (compatible; AhrefsBot/7.0; News; +http://ahrefs.com/robot/)
And the Site Audit crawler, which is a different bot doing a different job:
Mozilla/5.0 (compatible; AhrefsSiteAudit/6.1; +http://ahrefs.com/robot/site-audit)
Mozilla/5.0 (Linux; Android 13) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/151.0.7922.173 Mobile Safari/537.36 (compatible; AhrefsSiteAudit/6.1; +http://ahrefs.com/robot/site-audit)
AhrefsBot crawls the open web for the link index. AhrefsSiteAudit only crawls sites that someone has pointed the Site Audit tool at, which is usually you. Blocking one does not block the other.
Write your log filters against the AhrefsBot token, not the full string. Ahrefs has shipped 3.1, 4.0, 5.0, 5.2, 6.1 and 7.0 over the years. A filter pinned to one version quietly stops matching when the next one ships.
AhrefsBot/5.2 is a legacy user agent, not a separate bot. If it appears in current traffic, verify the source IP before you allow or block it. The old version number is a reason to check, not proof that the request is fake.
How to verify a real AhrefsBot
The user agent proves nothing. Anyone can send that header, and scrapers borrow well-known bot names precisely because sites wave them through.
Ahrefs publishes its crawler addresses in two machine-readable endpoints, so there is no reason to maintain a copied IP list that goes stale:
- CIDR ranges for firewall and WAF rules
- Individual IP addresses for exact matching
Match the source IP against those before you give the request any special handling. A reverse DNS lookup works as a second check: genuine requests resolve to hostnames ending in ahrefs.com or ahrefs.net.
This matters in both directions. An impostor claiming to be AhrefsBot is one you probably want to rate-limit. A genuine AhrefsBot you have accidentally counted as a visitor is a number in a report that somebody will act on.
How to test AhrefsBot access
Ahrefs' crawler page includes a status checker that tests whether AhrefsBot and AhrefsSiteAudit can crawl a domain. Use that first when you are debugging a block.
You can also test how your server responds to the published user agent:
curl -I -A 'Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)' https://example.com/
A 200 response shows that your user-agent rule allows the request. It does not prove that a real Ahrefs crawler sent it, and it does not test an IP allowlist. Use the published IP ranges to verify incoming traffic.
How to limit AhrefsBot without blocking it
Most people searching for a way to limit AhrefsBot do not actually want it gone. They want it quieter. Crawl-Delay does that, and Ahrefs honours it:
User-agent: AhrefsBot
Crawl-Delay: 10
The value is the minimum number of seconds between HTML requests. Ten seconds caps page requests at roughly six a minute. Ahrefs may still fetch a page's scripts, images and stylesheets in parallel while rendering, so your access log can show more than six total requests.
Two other things slow it down without any configuration. Persistent 4xx and 5xx responses make Ahrefs reduce crawl speed automatically. And AhrefsSiteAudit is already capped at 30 URLs a minute on sites nobody has verified, so the bot people blame for a traffic spike is usually AhrefsBot rather than the audit crawler.
How to block AhrefsBot
Full block:
User-agent: AhrefsBot
Disallow: /
Partial block, which is the better default if your load problem is one recursive section:
User-agent: AhrefsBot
Disallow: /search
Disallow: /filter
Faceted navigation and search-result URLs are where crawl budget goes to die, for every crawler, not just this one. Blocking those paths usually fixes the load without cutting yourself out of the link index.
Before you reach for the full block, be clear about the trade. Ahrefs' index is how you see your own backlinks, and it is how competitors see theirs. Blocking AhrefsBot degrades your own reporting more reliably than it protects anything. It also has no effect on Googlebot, your rankings, or your Search Console data, which are separate systems entirely.
Does AhrefsBot crawl your ads?
Not in the way the question usually means. Ahrefs' paid search and ads reports come from its own SERP and ads datasets, not from AhrefsBot clicking display placements on your site.
Ahrefs explicitly states that its crawler does not trigger ads. Its paid search reports use separate data sources, so seeing an Ahrefs report about an ad does not mean AhrefsBot clicked that ad.
GA4 also automatically excludes known bot and spider traffic. Your server, CDN and WAF logs still record the requests because those systems sit in front of analytics filtering.
How to measure AhrefsBot traffic
AhrefsBot should not appear in GA4 or Ads Manager, and it should never produce an order_created event. So the measurement question is not "how many sessions", it is "how much of my infrastructure is this consuming, and is it in the right bucket".
Three numbers are worth reporting:
- Crawl volume against Googlebot. Ahrefs should not be out-crawling Google on a site you want to rank. If it is, your crawl budget is going somewhere it should not.
- CDN and WAF cost. Recursive crawling of parameterised URLs is the usual culprit when bot traffic turns into a bill.
- Explicit exclusion from conversion destinations. This is the one that actually causes damage if you get it wrong.
That last point is the reason this page exists. AI crawler roundups tend to dump every user agent into one list, which is how an SEO tool crawler ends up in a report about AI visibility. AhrefsBot hitting /blog means Ahrefs' link graph wants the URL. It does not mean ChatGPT cited you. Citations come from OAI-SearchBot and ChatGPT-User.
How Hardal classifies it
Hardal files AhrefsBot under SEO tools in AI Visibility, separate from OpenAI, Anthropic and the rest of the trainer and fetcher categories, on the same server-side pipeline as Hardal Analytics. Same treatment as SemrushBot.
Because the classification happens server-side, the bot is identified before anything is forwarded, so it never reaches a conversion destination in the first place. For ChatGPT Ads conversions, see the hub and Hardal's OpenAI Ads destination. For the decision layer on trainers against fetchers, read AI crawler traffic for marketers.
What blocked traffic costs your attribution
Crawler traffic belongs in infrastructure reporting, not conversion reporting. Human visits have the opposite failure mode: an ad blocker can stop the analytics request while the person still buys. That leaves a real order with an incomplete acquisition path.
Read the marketing attribution guide for credit rules and the incrementality testing guide for causal lift. Hardal's measurement platform keeps bot filtering and first-party conversion collection in the same server-side pipeline.
Series: ChatGPT Ads · GPTBot · ClaudeBot · Google-Extended · PerplexityBot · OAI-SearchBot · ChatGPT-User · Meta-ExternalAgent · Googlebot · Bingbot · SemrushBot · AhrefsBot