Computational Marketing

How Bots Interact With Ecommerce Category Pages

Bots exploit category pages for pricing data, inventory signals, and crawl budget waste.

Reporter · · 10 min read
Cover illustration for “How Bots Interact With Ecommerce Category Pages”
Bots as a New Buyer Class · September 6, 2026 · 10 min read · 2,240 words

Category pages carry a disproportionate share of ecommerce bot traffic, and the reason comes down to what sits on them: prices, stock levels, SKUs, filters, and full product listings, all crawlable from one URL. During the 2024 holiday season, bots accounted for 57% of online shopping traffic, according to Radware's 2025 E-Commerce Bot Threat Report, meaning actual human shoppers were, for a stretch, the minority in their own marketplace. That single number is worth sitting with, and it supports a position worth stating up front: category pages get treated like SEO leftovers when they should be treated as the most fought-over real estate on the entire site.

Homepages sell a brand, and product detail pages sell one item at a time. Category pages carry dense grids of exactly the structured data a bot was built to grab.

The spectrum of bots that land on category pages and what each is after

Diagram: Bot Traffic on Category Pages: Who's There and Why. Visualizes: Visualize a ranked spectrum of bot types that land on category pages, ordered from 'beneficial' to 'harmful,' showing each bot's name and its primary target on a category page.

Non-human visitors range widely in intent, and lumping all of it under "bot traffic" is the first mistake most site owners make.

Googlebot and Bingbot crawl category pages to find and index content, governed by the crawl budget logic covered further down. AI training crawlers, GPTBot from OpenAI, ClaudeBot from Anthropic, CCBot from Common Crawl, sweep public pages to build or refine language models. A separate group, AI retrieval and real-time fetchers, crawls on behalf of a live user question, which behaves differently from a training crawler even though both get filed under the same "AI" label. Retailers also run their own price-monitoring tools, using the same scraping tricks as their competitors, just pointed outward instead of inward.

Then there's the other side of the ledger. Competitive price scrapers harvest pricing and stock data to feed intelligence platforms, while scalper bots watch for restock signals and sprint to checkout. Denial-of-inventory bots load up carts without ever paying, so real shoppers see "out of stock" on items that are, technically, sitting untouched in someone's abandoned cart. Content scrapers lift listings wholesale to build imposter storefronts or to train models without asking first.

Telling these apart by behavior alone keeps getting harder. Imperva's 2025 Bad Bot Report found that nearly 60% of bad bots now mimic human patterns: varying their timing, faking mouse movement, imitating a real person browsing closely enough to fool a casual glance at the logs. The sections that follow walk through what each bot class actually does on a category page, because behavior is what makes analytics readable and mitigation possible.

How price scraping bots use category pages as their primary data source

Hand a category page to a scraping bot and it functions like a spreadsheet: product names, prices, discount flags, stock status, SKUs, the whole filter and pagination structure, pulled in one request. A product detail page, by comparison, gives up one data point at a time. A bot chasing competitive intelligence on 400 SKUs doesn't bother visiting 400 PDPs; it recurses through a category page's filter combinations once and rebuilds the whole catalog map from the responses.

Competitor retailers do this, and so do third-party price intelligence vendors, marketplace aggregators, and scalpers timing purchases off scraped pricing signals. That's the part worth naming plainly: this is a standard input into how retail pricing gets set, and pretending otherwise is what leaves category pages unguarded.

The cost lands on infrastructure. Repeated, high-frequency crawling slows the page down for actual shoppers and shows up on the server bill as unexplained compute spend rather than a line item called "fraud." In analytics, the signature is unmistakable once someone knows where to look: session counts spike, engagement flatlines near zero, and the timing lines up a little too well with a competitor's pricing changes. Server logs will often show bot user agents or headless browser fingerprints sitting right there in plain sight, for anyone who bothers to check them.

How scalper and inventory bots use category pages as their early-warning system

For scalper bots, the category page is a tripwire. In-stock badges, "limited" labels, restock indicators, all of it shows up here before a shopper ever clicks through to a product page. The sequence runs the same way each time: watch the category page for a restock signal, jump to the PDP, add to cart, check out, then relist at a markup on a secondary market.

The highest-value targets are the ones where scarcity itself is the product: limited sneaker drops, GPU launches, concert tickets, collectibles. NVIDIA's RTX 5090 and 5080 launch in early 2025 is a clean, documented case: inventory disappeared within minutes across Best Buy, Newegg, and NVIDIA's own storefront. That's category-page monitoring at scale, working exactly as designed, and it's worth sitting with what "working" means here: a script beat every human who had a browser tab open and a credit card ready.

Denial-of-inventory bots run a quieter, meaner version of the same game: add to cart, never buy, repeat. That artificially suppresses availability, pushes genuine buyers toward resellers, and wrecks two metrics at once, cart abandonment rate and conversion rate, in ways that mislead whoever is making merchandising calls off that data.

This is hard to catch for a simple reason: a scalper bot's whole design goal is to look exactly like a human closing out a legitimate purchase. Of every bot type covered here, these are the best mimics, built to disappear into normal traffic rather than stand out from it, which is exactly what makes them the harder problem to solve.

How faceted navigation turns category pages into a crawl budget trap

Every filter combination on a faceted category page, size, color, brand, price range, star rating, generates its own URL, and search bots find these URLs and try to crawl every single one, so the math gets out of hand fast.

A catalog with a few thousand real products and half a dozen filterable attributes can expose a URL space that dwarfs the actual product count, with filter permutations multiplying into the millions. Faceted navigation is widely recognized as one of the single largest structural problems category pages create for search visibility, with filter-generated URL sprawl consistently cited as a leading source of crawl inefficiency. Every hour Googlebot spends crawling the filter combination for "size 9, blue, under $50, 4 stars and up" is an hour it isn't spending on the canonical category page, or on a product that just got added.

The knock-on effects compound. Duplicate content piles up as near-identical product sets sit behind dozens of different URLs, while internal link equity spreads thin across thousands of low-value facet pages instead of concentrating on the ones that should rank. New products and promotions get indexed slower, sometimes too slow to matter for a seasonal item that's half sold through by the time it shows up in search results.

AI crawlers add to this. Their crawl-to-referral ratio is rough: AI training crawlers generally consume crawl budget at a high rate relative to the referral traffic they return. Faceted URL bloat burns through that budget with even less return than it burns through a search bot's.

The standard fixes are well known: robots.txt rules that disallow facet parameter URLs, the Post-Redirect-Get pattern that hides filter links from bots while leaving the human browsing experience untouched, and canonical tags pointing every facet variant back to the main category page. Good bots respect these rules, but malicious ones largely don't, and that's the real gap: the tools that fix the SEO problem and the tools that stop the bad bots are separate tools built for separate jobs. Mistake one for the other and a site ends up protected from Googlebot while standing wide open to everyone else.

How bot activity on category pages corrupts the analytics that drive merchandising decisions

Bots don't just cost bandwidth; they distort the data itself. Session counts, pageview depth, bounce rate, click-through rate by SKU, all of it skews the moment automated traffic outweighs real visits, and category pages feed straight into decisions about inventory forecasting, promotional placement, ad budget allocation, and conversion optimization.

The effect on conversion metrics can be severe: when bot sessions dominate the denominator, a conversion rate that once meant something can become statistically meaningless. Nothing broke on the site; the denominator just got swamped by bot sessions, and a conversion rate that used to mean something stopped meaning anything.

Category pages sit upstream of product popularity signals, so bot inflation here doesn't just distort one page's numbers; it corrupts whatever recommendation or promotion logic downstream gets built on those numbers. Denial-of-inventory bots add another wrinkle: they generate add-to-cart events that never convert, which makes cart abandonment rate meaningless for whatever category they're targeting. Paid traffic takes the hit too. If a bot lands on a category page after a paid search click, the ad platform logs it as a visit, the campaign's conversion rate drops, and someone in marketing pulls budget off a keyword that might have been working fine for real customers all along.

Fixing this starts with recognizing what bot traffic looks like against human behavior: session length, click depth, timing that's a little too regular to be a person browsing on a Tuesday afternoon.

The AI shopping agent as a new kind of category page visitor

Diagram: AI-Referred Traffic: From Worst to Best Converter. Visualizes: Show a before-and-after or directional shift for AI-referred retail traffic conversion: in March 2025 it converted 38% worse than other sources; by recent accounts it converted…

Here's where the picture flips. An AI shopping agent acts on behalf of an actual person with an actual thing they want to buy, gathering data to inform that one purchase rather than harvesting it for resale. Same category page, different intent, and treating it like a scraper is the second big mistake site owners are about to make, right after they finally learn to catch the first kind.

What the agent evaluates looks less like browsing and more like due diligence: does the product match the stated need, what's the price, is it in stock, what are the delivery and return terms, what do the reviews actually say. It needs structured data that supports a decision, gathered from a page built to inform rather than one built only to nudge someone toward a click.

The conversion numbers moved fast. In March 2025, AI-referred traffic to US retail sites converted 38% worse than traffic from other sources. By recent accounts, that same AI-referred traffic was converting meaningfully better. The direction of the shift is what matters: agents that looked like noise not long ago are now sending customers who complete purchases at a meaningful rate.

That shift changes what a category page needs to be. Metadata completeness stops being a nice-to-have and becomes the entire discovery surface, since an agent that can't parse availability or specs in structured form simply can't recommend the product. Filter and sort logic needs to be readable by machines, not just clickable by humans. Emerging interoperability standards aim to let AI systems query a retailer's live data directly instead of scraping the rendered page, and newer protocols seek to let agents and retail systems exchange structured information across discovery, checkout, and post-purchase support.

The same category page that needs walls up against scrapers also needs a clear, well-lit door for the agents actually bringing paying customers. Those are two different design problems, and solving one does nothing for the other; confusing them is the most expensive mistake in this whole piece.

What site owners can actually do differently once they understand which bots are doing what

Bot mitigation was never going to be one tool, because a category page is fighting four different battles at once, and they don't share a single fix. Buying one vendor's "bot protection" and calling the problem solved is the third mistake, and it's the one that undoes the first two.

Crawl budget waste from legitimate bots hitting facet URLs gets solved with structural fixes: robots.txt rules, the PRG pattern, canonical tags. Competitive price scraping calls for rate limiting, session-level bot detection, and token-based access controls on high-frequency category requests. Scalping and inventory manipulation need queue systems for high-demand drops, anomaly detection at the add-to-cart step, and limits on how long inventory can sit held in a cart before it's released. AI agent readiness is a content problem before it's a security problem: structured data, complete metadata, clean taxonomy, and a plan for protocols that are still taking shape.

Analytics hygiene has to come first, honestly, ahead of everything on that list. Filtering known bot signatures out of reporting before drawing conclusions about product performance or ad spend matters more than most teams give it credit for; bad data drives bad decisions faster than bad bots drain a server budget.

Worth repeating: robots.txt is a request that well-behaved crawlers honor and malicious ones ignore outright, so any strategy that leans on robots.txt alone only ever solves the polite half of the problem. Blocking indiscriminately cuts both ways, too: shutting out all non-human traffic also shuts out the AI shopping agents now sending genuinely purchase-ready customers, which is revenue worth protecting, not traffic worth throwing out with the bathwater.

None of this sits in one department's lap. SEO, engineering, fraud and security, and analytics are all looking at the same traffic mix, just from different windows, and each one sees a different symptom. Content and merchandising teams have a stake here too: if the product rankings and click-through numbers feeding their decisions are quietly full of bot noise, the promotional calls built on top of that noise need to account for it, or someone ends up optimizing for an audience that was never actually there.

Sources

  1. radware.com

More in Bots as a New Buyer Class