Shopify SEO. Crawling
How to Crawl a Shopify Store Without 429 Errors
Written by Ram Kr Shukla, SEO and growth consultant
Quick answer
Shopify rate-limits crawlers that don’t identify themselves, so Screaming Frog at its default speed gets 429 Too Many Requests on a big share of pages, and those rows look like broken, non-indexable URLs. Fix it before you read any data. Create a crawler access signature in Shopify (Online Store > Preferences > Crawler access) and add its three headers in Screaming Frog, then set Configuration > Speed to 1 thread and 1 URL per second. Save that as a config file, pass it with --config when you crawl from the command line, and check the Status Code column for 429 before you report anything.
The first crawl I ran on an Indian cosmetics store finished in 11 seconds. That should’ve been the giveaway. Screaming Frog reported 83 URLs as client errors, every one of them non-indexable, and almost every collection page was on the list. Read at face value, the store’s category pages were broken.
They weren’t. All 83 rows said the same thing in the Status column: Too Many Requests, status code 429. Shopify had throttled the crawler, and the crawler had written the throttling down as if it were site data.
This is the method I use to crawl a Shopify store for an audit without that happening. You’ll size the crawl from the sitemaps, set the exact Screaming Frog options, save them and run them headless, spot 429s in an export, re-crawl only the failures, and check what Googlebot sees in the meantime. It’s the step that comes before every finding in my Shopify SEO work, because a finding built on a throttled crawl is wrong before anyone reads it.
Why a Crawl Full of 429s Is Worse Than No Crawl
A 429 is an HTTP status code that means the server got too many requests from you in a short window. It’s a traffic signal about your crawler. Nothing on the page is wrong.
Screaming Frog files it under Client Error (4xx), the same filter that holds 404s and 410s (Screaming Frog tabs guide). In the Internal tab the row reads Non-Indexable, with Client Error as the reason. Put that export in front of a client or a developer and it looks like dead pages. That’s the first problem. Three more sit behind it.
Discovery stops at the throttled page. A crawler finds new URLs by reading links on pages it fetched. A page that came back 429 has no links to read, so anything linked only from it never enters the crawl. On the cosmetics store, the default crawl reached 86 HTML pages. The sitemaps declared 417.
Every other report is built on the survivors. Missing H1s, duplicate titles, thin pages: those counts come from whichever pages returned 200. On that crawl it was 3 pages. A report saying “1 page is missing an H1” would’ve been technically true and completely useless.
The errors cluster on one template. Collections took the worst of it: 54 of the 55 collection URLs requested came back 429, along with all 19 product URLs. A pattern like that reads like a template bug, which sends a developer hunting for a fault that doesn’t exist.
So treat a 429 crawl as misleading, not partial. My rule is blunt. If the Status Code column contains a single 429, nothing from that crawl goes into a report.
Why Shopify Throttles Your Crawler
Screaming Frog’s own FAQ says the rate limiting on Shopify sites comes from Cloudflare, which Shopify uses in front of stores (Screaming Frog FAQ). Shopify has tightened it since. Its developer changelog says it applies stricter rate limits to bots and agents that access Shopify-hosted storefront pages, and that bots which don’t sign their requests are subject to the strictest limits (Shopify developer changelog).
Signing is done with Web Bot Auth. Merchants create signatures in the admin under Crawler access, and Shopify lists SEO and accessibility audits among the intended uses (Shopify Help Center). An unsigned Screaming Frog crawl at default speed is exactly the traffic those limits exist to slow down.
Two things follow. Slowing down helps, but an unsigned crawler gets judged more harshly at any speed. And signing isn’t a guarantee either. In one Shopify developer forum thread, a user reported signatures working at first and 429s coming back within two weeks (Shopify developer community). Do both: sign the crawler and throttle it.
Size the Crawl From the Sitemaps First
Before you touch a setting, find out how big the store is. Shopify generates the sitemap automatically at the root of each domain, and it links to separate sitemaps for products, collections, blogs and pages (Shopify Help Center). Open https://yourstore.com/sitemap.xml and you’ll see an index of child sitemaps. The names follow a predictable pattern:
https://yourstore.com/sitemap_products_1.xml?from=...&to=...
https://yourstore.com/sitemap_pages_1.xml
https://yourstore.com/sitemap_collections_1.xml
https://yourstore.com/sitemap_blogs_1.xml
Big catalogues get more than one products file. The two live stores I looked at while writing this also listed a sitemap_agentic_discovery.xml, which held a single URL on each. Count the URLs in every child. From a terminal, this does it per file:
curl -s https://yourstore.com/sitemap_collections_1.xml | grep -c "<loc>"
Image entries in Shopify sitemaps use a different tag, so they won’t inflate that count.
For the cosmetics store, the four children declared 417 URLs: 62 products, 82 pages, 127 collections and 145 blog posts. That’s your yardstick. A finished crawl should land near that number of HTML pages, plus whatever variant and filter URLs the theme links to. If it lands at 86, something stopped it.
How long a throttled crawl will take
The count also tells you how long you’ll wait. At 1 URL per second, 417 pages is about 7 minutes of page requests. Screaming Frog also fetches images, CSS and scripts by default, so call it about 10 minutes all in. Here’s the rough math for bigger stores, counting page URLs only:
| Sitemap URLs | At 1 URL per second | At 2 URLs per second |
|---|---|---|
| 417 | About 7 minutes | About 3.5 minutes |
| 5,000 | About 83 minutes | About 42 minutes |
| 20,000 | About 5.5 hours | About 2.8 hours |
| 100,000 | About 28 hours | About 14 hours |
At the far end, the fashion store in my Shopify case study had 200,000+ URLs indexed when the work started. At 1 URL per second that’s more than two days of nonstop crawling. On a store that size, crawl by section from the child sitemaps and run each one overnight. It also means you should look at index bloat and crawl budget before anything else, since most of those URLs shouldn’t exist.
You can feed the sitemaps straight into the crawl. In the app, Configuration > Spider > Crawl has Crawl Linked XML Sitemaps, with an option to discover them from robots.txt (Screaming Frog configuration guide). Run Crawl Analysis when it finishes to fill the sitemap filters. From the command line, --crawl-sitemap takes a sitemap URL and crawls what’s in it.
The Exact Screaming Frog Settings for Shopify
Five settings matter. The first two do most of the work.
| Setting | Where in Screaming Frog | Shopify value |
|---|---|---|
| Crawler access headers | Configuration > HTTP Header | Signature-Input, Signature, Signature-Agent from Shopify |
| Max Threads | Configuration > Speed | 1 (default is 5) |
| Max URI/s | Configuration > Speed | 1, or 0.5 if 429s persist |
| User agent | Configuration > User-Agent | Screaming Frog default, or Chrome |
| Storage mode | Storage settings | Database |
1. Add a Shopify crawler access signature
In the Shopify admin, go to Online Store > Preferences, find the Crawler access section and click Create signature. Name it, pick the connected domain you’ll crawl, choose an expiry and create it. Shopify gives you the values for three headers: Signature-Input, Signature, and Signature-Agent, which must be the string “https://shopify.com” with the quotes included (Shopify Help Center).
In Screaming Frog, open Configuration > HTTP Header and add the three as custom headers. That setting exists for custom request headers like these (Screaming Frog configuration guide).
Screaming Frog’s FAQ has three details that’ll save you an hour (Screaming Frog FAQ). Paste the values straight from Shopify instead of typing them. Expect 30 minutes to an hour before it takes effect if you’ve already been collecting 429s. And signatures expire, 90 days by default. Shopify caps the expiry at 3 months and won’t renew one, so when it lapses you create a new signature and update the headers.
Each signature covers one connected domain. If the store runs a separate domain per market, you need one for each domain you crawl. No admin access? Ask the store owner to create it and send you the values.
2. Throttle Configuration > Speed
Screaming Frog crawls at 5 threads by default (Screaming Frog configuration guide). On an unsigned Shopify crawl that’s enough to trip the limits almost at once. It’s the setting that produced 83 throttled pages in 11 seconds.
Open Configuration > Speed. Set Max Threads to 1. Turn on the Max URI/s limit and set it to 1. Screaming Frog’s guide says that when you’re slowing a crawl down, Max URI/s is the easier control, since it caps requests per second directly.
If 429s still appear with a signature in place and 1 URL per second, halve it to 0.5 and run again. Slow is cheap on a store with a few hundred pages.
3. Leave the user agent alone
Screaming Frog can switch its user agent, including a Googlebot preset (Screaming Frog configuration guide). On a throttled site it’s tempting. Don’t. Google publishes how site owners verify real Googlebot traffic, using reverse DNS and Google’s published IP ranges (Google Search Central). A crawler on your own connection claiming to be Googlebot fails that check.
Keep the Screaming Frog default. If you’re still blocked on the very first URL, try a Chrome user agent, which is one of the steps in Screaming Frog’s own 429 troubleshooting (Screaming Frog status codes tutorial).
4. Use database storage
Switch storage to database mode. It saves the crawl to disk as it runs and reopens saved crawls faster (Screaming Frog user guide). That matters when a throttled crawl runs for hours and you don’t want to lose it to a laptop restart.
5. Don’t count on retries
Configuration > Spider > Advanced has a 5XX Response Retries option. A 429 is a 4XX code, so don’t expect that setting to pick them up. Plan to re-crawl failures yourself, covered below.
Save the Config and Run It From the Command Line
None of those settings can be passed as command line flags. I checked the full --help output of Screaming Frog 24.1: there’s --config for a saved configuration file and nothing for threads, URLs per second or request headers. The command line guide describes --config as the way to supply a saved configuration file.
So set everything up once in the app, then save it with File > Configuration > Save As. That writes a .seospiderconfig file. Give it an obvious name like shopify-throttled and keep it with your audit templates. A headless crawl on a Mac then looks like this:
"/Applications/Screaming Frog SEO Spider.app/Contents/MacOS/ScreamingFrogSEOSpiderLauncher" \
--crawl https://yourstore.com/ \
--headless \
--config ~/sf-configs/shopify-throttled.seospiderconfig \
--save-crawl \
--output-folder ~/crawls/yourstore \
--overwrite \
--export-format csv \
--export-tabs "Internal:All,Response Codes:Internal Client Error (4xx)"
On Windows the executable is ScreamingFrogSEOSpiderCli.exe and the flags are the same. A few things the help text won’t warn you about:
Export names must match exactly. A typo in --export-tabs only fails at export time, after the crawl has run. Run the launcher with --help export-tabs first and copy the names from its list.
Startup isn’t instant. Each launch starts a Java process. On my Mac that’s about 25 seconds before the crawl begins, so batch your help lookups.
Reload instead of re-crawling. In database mode, --save-crawl keeps the crawl and --load-crawl reopens it later for new exports. On a rate-limited store, that’s the polite way to slice the data again.
If you’ve built other automation on top of Screaming Frog exports, the same config file keeps those runs consistent too. I use a similar setup to turn crawls into a Markdown knowledge base of a website.
How to Spot 429s in a Screaming Frog Export
Check this before you open any other tab. It takes a minute. In the Internal:All export, a throttled row looks like this:
| Column | Value on a throttled row |
|---|---|
| Status Code | 429 |
| Status | Too Many Requests |
| Indexability | Non-Indexable |
| Indexability Status | Client Error |
In Google Sheets or Excel, =COUNTIF(C:C,429) counts them. Status Code is column C in a default Internal:All export, but check your header row. Inside the app, the same URLs sit under Response Codes > Client Error (4xx), and Screaming Frog’s docs list 429 among the codes in that filter (Screaming Frog tabs guide).
Three tells give a throttled crawl away before you count anything:
It finished too fast. 11 seconds for a store with 417 URLs in its sitemaps isn’t a crawl. Check the start and end times in the Crawl Overview report.
Assets came back fine and pages didn’t. On the cosmetics store, 160 images, stylesheets and scripts from Shopify’s CDN returned 200 while 83 of 86 HTML pages got a 429.
One page type took almost all the errors. 54 of 55 collection URLs. When a single template carries nearly every error, check the status code before you blame the template.
Re-crawl Only the Failed URLs in List Mode
When the 429 count is small, you don’t need to start over. List mode crawls exactly the URLs you give it, pasted or uploaded (Screaming Frog user guide).
Pull the failures. Filter the Internal export to Status Code 429 and copy the Address column into a plain text file, one URL per line.
Load the throttled config. In the app, load your saved .seospiderconfig, switch Mode to List, then paste or upload the file.
Or run it headless. Swap --crawl for --crawl-list failed-429.txt in the command above and keep --config pointing at the same file.
Check again. Count 429s in the new export. If any remain, slow down to 0.5 URLs per second and repeat.
There’s a catch. List mode fixes the rows you already have, but it can’t find the pages you never discovered. Pages linked only from a throttled page never made it into the first crawl, so they won’t be in your list either. That’s why the sitemap count matters.
A sensible line: if the 429s are a handful of URLs and your crawled HTML count is close to the sitemap total, patch them in list mode. Otherwise, crawl the whole store again with the throttled config. On the cosmetics store, with 83 of 86 pages throttled, patching was never an option.
What Googlebot Sees While Your Crawler Gets 429s
Your crawler’s 429s say nothing about Google’s access. Shopify’s help page states that stores can be indexed by search engines and large language models without signatures (Shopify Help Center).
On the cosmetics store, I took 8 collection URLs that had returned 429 to Screaming Frog and ran them through URL Inspection in Search Console. All 8 were Submitted and indexed, and Google had crawled each of them within two days of the check.
What Google does when it does get a 429 is documented. Its crawlers treat 429 as a sign the server is overloaded, count it as a server error and slow down in proportion to how many URLs return errors, then speed back up once responses are healthy (Google Search Central). The crawl budget guide says the same signal lowers the crawl capacity limit for the site. And Google warns that if Googlebot sees 429, 500 or 503 on the same URL for multiple days, that URL may be dropped from the index (Google Search Central).
Google also ignores crawl-delay in robots.txt (Google robots.txt spec). Shopify’s default robots.txt sets Crawl-delay: 10 for AhrefsBot, AhrefsSiteAudit and MJ12bot, and 1 for Pinterest. That slows those tools down. It has no effect on Googlebot, and Screaming Frog isn’t named at all.
So where do you see Google’s side? In Search Console, under Settings > Crawl stats. The By response breakdown has an Other client error (4XX) row, and Host status shows whether Google had trouble reaching the server (Search Console Help). I’ve written up how to read it in the Search Console reports most people skip. Shopify doesn’t publish what limits, if any, apply to Googlebot, so read that report instead of guessing. If the 4XX row is climbing while your crawler is throttled, dig in. If it’s flat, the 429s belong to your crawler alone.
What the Store Looked Like Once the 429s Were Gone
With the throttling out of the way, the real issue on the cosmetics store turned out to be a template defect. 35 of 35 collection pages checked had no H1, and 31 of 35 had no meta description. Product, policy and content templates were clean.
That’s one edit to the collection template in Liquid. Written up as 35 page-level tasks, it would likely have been quoted as a week of developer time. The default crawl could never have shown it, because it returned 200 for exactly one collection page.
If your collections have the same gap, the collection page guide covers what that template should output. The influencer collection pages post deals with the thin, campaign-only collections that often make up most of a store’s collection count. Title tags are the other template check worth running on the same crawl, and the truncated product title tag fix explains the most common one. Duplicate URLs from variants and filters are their own subject, covered in Shopify technical SEO traps.
The Whole Process, Step by Step
If you’re running a full audit, this sits at the top of the 40-point SEO audit checklist, before any on-page or content check. A clean crawl also feeds the crawl depth analysis, which is meaningless when a third of the links were never followed.
More from the Shopify series: the Shopify SEO hub, Shopify duplicate content traps, index bloat and crawl budget, collection page SEO, and truncated product title tags. This is core to my technical SEO and SEO audit work.
Shopify Crawling FAQ
Why does Screaming Frog get 429 errors on Shopify?
Shopify rate-limits bots and agents that request its storefront pages, and unsigned requests get the strictest limits. Screaming Frog’s default of 5 threads is fast enough to trip them, so the store answers with 429 Too Many Requests instead of the page.
Are 429 errors in a crawl real problems on my store?
Usually not. A 429 means your crawler asked too quickly, and the page itself is fine. Screaming Frog still labels those rows as non-indexable client errors, which is why a throttled crawl reads like breakage. Re-crawl slower before you report anything.
What crawl speed should I use for Shopify in Screaming Frog?
Start at 1 thread and 1 URL per second under Configuration > Speed, with a Shopify crawler access signature in your HTTP headers. If 429s still show up, drop to 0.5 URLs per second and run it again.
Can I set crawl speed from the Screaming Frog command line?
Not with a flag. Set speed, headers and user agent in the app, save them with File > Configuration > Save As, then pass the .seospiderconfig file with –config when you run a headless crawl.
How do I find 429 errors in a Screaming Frog export?
Open the Internal:All export and filter the Status Code column for 429. Those rows show Too Many Requests as the status and Client Error as the indexability status. In the app, they sit under Response Codes > Client Error (4xx).
Does Shopify rate-limit Googlebot too?
Shopify says stores can be indexed by search engines without signatures, and it doesn’t publish what limits apply to Googlebot. Check the Crawl stats report in Search Console for Other client error (4XX) responses. If that row is flat, the 429s are your crawler’s problem only.
How long does a Shopify crawl take at 1 URL per second?
About one second per URL, so count the sitemap URLs first. A store with 417 sitemap URLs needs about 7 minutes of page requests, or around 10 once images and scripts are included. A store with 5,000 URLs needs well over an hour.
Should I switch Screaming Frog’s user agent to Googlebot to avoid 429s?
No. Google publishes how site owners verify real Googlebot traffic through reverse DNS and its IP ranges, so a spoofed user agent from your own connection is easy to spot. Keep the default user agent and add a crawler access signature instead.
Did your last Shopify crawl come back full of client errors?
Check the status code before anyone fixes anything. I run signed, throttled crawls of Shopify stores and turn what they find into template-level fixes your developer can ship in one pass.
Shopify SEOBook an SEO AuditTags: Shopify SEOTechnical SEOScreaming FrogSEO Audits
Written by Ram Kr Shukla
SEO and growth consultant and Fractional Chief Digital Officer with 20+ years in SEO and growth across 50+ brands. Google Analytics and Google Ads certified. Works daily in Screaming Frog, Search Console and GA4 on Shopify and enterprise e-commerce sites, and audits D2C stores for the crawl, template and tracking faults that defaults hide.




