AI Analysis

AI Search Is Splitting Your Traffic: How to Measure ChatGPT and Perplexity Referrals in GA4

Measurement. The New Referral Layer

AI Search Is Splitting Your Traffic: How to Measure ChatGPT and Perplexity Referrals in GA4

4x
Growth in AI referral traffic across client sites this year
15 min
To build the GA4 view that isolates it
2x+
Conversion rate of AI referrals versus average, consistently
0
Standard reports that show this today

There is a new traffic source growing in your analytics right now, and GA4’s default reports are hiding it inside the referral bucket. Visitors arriving from ChatGPT, Perplexity, Gemini, and Copilot click through when an assistant cites you, and across my client base that stream has roughly quadrupled this year. Small absolute numbers still, but the trend line and the visitor quality demand their own measurement.

Because here is the finding that matters: AI referrals convert at roughly double the site average, consistently, across clients. It makes sense. A visitor arriving from an assistant’s recommendation was pre-sold by the most trusted interface they use. You want to know exactly how many arrive and what they do.

Build the View in GA4, Step by Step

1
Know the referrer signatures

AI referrals arrive with identifiable referrers: chatgpt.com and chat.openai.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, claude.ai, and you.com. Some assistant clicks arrive stripped as direct, so whatever you measure is a floor, not a ceiling. Floors still show trends.

2
Create the channel group

In Admin, under Data display, edit your channel groups: add an AI Referrals channel matching source against a regex of those domains. From that point on, standard reports break the stream out cleanly beside Organic Search and the rest.

3
Build the exploration for history

Channel groups apply going forward. For history, build an Exploration filtered on session source matching the same regex: sessions, engaged sessions, conversions, and landing pages, trended monthly. This is your baseline chart, and it will be a hockey stick at small scale.

4
Study the landing pages above all

Which pages do assistants send people to? That list is your citation footprint made visible: it tells you which content AI systems consider reference-worthy, and it is the template for what to produce more of. This single report has reshaped content plans for my clients more than any keyword tool this year.

Reading the Numbers Honestly

MetricHow to read it
VolumeSmall today almost everywhere. Single digit percentages. The trend is the story, not the total
Conversion rateConsistently above site average. Pre-sold visitors behave like referrals from a trusted friend
Landing pagesComparisons, definitive guides, and original data get cited. Thin service pages do not
Month over monthThe only chart worth presenting. Screenshot it quarterly; the compounding makes the argument for you

“Nobody reported on organic search traffic in 1999 either. The brands that measured it early built the playbooks everyone else bought later. AI referrals are that moment again, with a fifteen minute setup cost.”

Ram Kr Shukla, SEO and Growth Consultant

Once the measurement exists, it changes decisions: the landing page report feeds your content plan, the conversion data justifies AI visibility investment to whoever owns the budget, and the baseline lets you actually evaluate whether citation work, the kind I documented in my AI recommendations research, moves your numbers. Measurement first, then optimisation has something to answer to.

Want the full AI visibility measurement stack installed?

I set up AI referral tracking, citation monitoring across assistants, and the quarterly baseline test as part of every AI SEO engagement. Know your numbers before your competitors know theirs.

Explore AI SEO ServicesBook a Free Strategy Call

Tags: GA4AI ReferralsMeasurementAI SEO

Programmatic SEO Done Right: How to Scale Pages Without Getting Penalised

AI and Scale. Programmatic SEO

Programmatic SEO Done Right: How to Scale Pages Without Getting Penalised

2M
Pages I have managed at marketplace scale
1
Question that decides survival: is each page useful alone?
10x
Faster with AI, which cuts both ways
3
Structural rules that separate assets from spam

Programmatic SEO, generating pages from data and templates instead of writing them one by one, built some of the biggest organic footprints on the web. I learned it managing marketplace sites with millions of URLs, long before AI made generation trivial. And that is exactly the new problem: AI made scaling easy, so the graveyard of penalised programmatic sites is filling faster than ever.

The difference between a programmatic asset and thin content spam was never the page count. It is whether each generated page would deserve to exist if you built it by hand. That standard is passable, and here is how the survivors pass it.

The Three Structural Rules

1
Unique data per page, not unique words

Spinning synonyms across a template fools nobody since Google’s helpful content systems matured. What works: each page assembled around data that genuinely differs, prices, specs, availability, local details, real comparisons. If two pages would answer a visitor identically, they should be one page.

2
A query with real intent behind every URL

Programmatic pages must map to searches people actually make: city plus service, product plus comparison, tool plus integration. Generating pages for query patterns nobody searches produces index bloat that drags the whole domain. Validate the pattern’s demand before generating the ten thousandth page.

3
Crawl architecture that carries the weight

Ten thousand pages need hub structure, clean pagination, and internal links that give every page a path from the money pages. Generated pages dumped into a flat sitemap with no internal linking are orphans at scale, and Google treats them accordingly.

Where AI Fits, and Where It Breaks Things

AI roleVerdict
Assembling pages from your structured dataIdeal. This is what the technique always was, with better tooling
Writing unique intros against real data per pageWorks with editorial gates and the brief discipline from my AI content workflow
Generating both the data and the wordsThe graveyard. Invented data plus template prose is the exact pattern quality systems now catch
Deciding which pages deserve to existNever. Demand validation and pruning stay human decisions

“Programmatic SEO fails at the same question every time: would this page deserve to exist if a human had to build it? AI changed the cost of generating pages. It did not change the standard they are held to.”

Ram Kr Shukla, SEO and Growth Consultant

Start smaller than you want to: one pattern, a few hundred pages, indexed and measured for a quarter. Scale what earns impressions and conversions; prune what does not, ruthlessly. The discipline to delete generated pages is rarer than the ability to generate them, and it is what keeps the domain healthy.

Considering programmatic at scale?

I have run this at marketplace scale and rebuilt it after others’ penalties. Pattern validation, data architecture, and the guardrails, before you generate page one.

AI Marketing ServicesBook a Free Strategy Call

Tags: Programmatic SEOAIScaleContent Quality

Schema Markup in 2026: The Only 6 Types Most Businesses Actually Need

Technical SEO. Structured Data

Schema Markup in 2026: The Only 6 Types Most Businesses Actually Need

6
Types that earn their maintenance cost
30+
Types most plugins offer that you can ignore
2
Jobs schema now does: rich results and AI comprehension
15 min
To validate everything you currently have

Schema markup suffers from a completeness problem: there are hundreds of types, plugins offer dozens, and teams either implement nothing or implement everything badly. Broken schema is worse than none; I covered a client whose rich results vanished sitewide from one invalid plugin update in my rankings emergencies post.

Schema also quietly picked up a second job. It used to earn rich results in Google. Now it also tells AI systems, the assistants and the Overview builders, exactly who you are and what you offer in a format they parse without guessing. Both jobs are done by the same six types for most businesses.

The Six That Matter

1
Organisation, on every page

Name, logo, URL, and sameAs links to your real profiles. This is your entity anchor: the machine-readable statement of who is behind the website. For consultants and personal brands, Person schema does this job alongside it.

2
Service or Product, on money pages

What you sell, described in structured form. For e-commerce: Product with price, availability, and ratings, which earns the rich result that moves click-through even without a rank change. For services: Service with provider and area.

3
FAQ, where real questions live

On pages answering genuine questions buyers ask. Google trimmed its visual FAQ real estate, but the structured question and answer pairs are exactly the fragments AI systems lift. Write real questions, not keyword-stuffed fakes.

4
Article, on content that argues expertise

With author, dates, and publisher. This connects your content to your entity, which is how expertise accumulates to a name instead of evaporating page by page.

5
BreadcrumbList, sitewide

Cheap to implement, improves the SERP display, and reinforces your site architecture to crawlers. The lowest-effort item on this list.

6
Review or AggregateRating, where earned

Ratings you genuinely collect, marked up where they live. Never fabricated and never sitewide decoration: unearned rating markup is the fastest route to a manual action of anything on this page.

What to Skip, and the Maintenance Rule

Skip the exotic types your plugin offers unless you demonstrably need them: Event without events, VideoObject without video, HowTo on pages that are not how-tos. Every type you add is a validation surface that can break silently in a plugin update. The maintenance rule that survives contact with reality: validate quarterly in Search Console’s Enhancements reports, and after every plugin or theme update. Fifteen minutes, calendared, owned by a named person.

“Schema used to be decoration for rich results. It is becoming your machine-readable identity. Six types, kept valid, beat thirty types nobody checks.”

Ram Kr Shukla, SEO and Growth Consultant

A Worked Example: The Silent Breakage and the Recovery

The client I mentioned in my rankings emergencies post is worth expanding here, because the sequence is so typical. A plugin update changed how review markup rendered: technically present, semantically invalid. Google stopped showing stars sitewide within days. Rankings barely moved, but click-through fell by nearly a third, and revenue followed. It took weeks to notice precisely because everyone was watching rankings, the metric that had not changed.

The recovery was mechanical once diagnosed: valid markup restored, Search Console validation run, rich results back within two crawl cycles. The durable fix was procedural, the quarterly validation calendar and a named owner. Schema does not fail loudly. It fails like a fridge light: you only notice when you finally open the door and check.

Schema implementation mistakes worth checking today:

  • Marking up ratings that do not visibly exist on the page, a manual action magnet
  • Duplicate Organisation schema from theme, plugin, and manual code all firing at once
  • FAQ markup on invented questions no one asks, which wastes the one fragment AI systems lift
  • Schema in the page builder that disappears when the template changes
  • Validating in a testing tool once at launch and never in Search Console where field reality lives

Questions I Get on This Topic

JSON-LD or microdata? JSON-LD, without hesitation. It lives in one script block, survives template edits far better, and is what Google recommends. If a plugin outputs microdata woven through your HTML, that is a fragility worth migrating away from.

Does schema directly improve rankings? Not as a ranking factor in the direct sense. It improves how results display, which moves click-through, and it improves machine comprehension of who you are, which increasingly matters beyond Google. The revenue path is real; it just does not run through the position number.

How do I add schema on WordPress without another plugin? Most SEO plugins you already run handle the six core types adequately. The gap is usually configuration and validation, not tooling. Before adding anything, audit what your current stack already outputs; duplication is the more common disease than absence.

Want your structured data audited and rebuilt properly?

Schema architecture is part of every technical SEO engagement: what you have, what is broken, what is missing, and the six type implementation your stack actually needs.

Technical SEO ServicesBook a Free Strategy Call

Tags: SchemaStructured DataTechnical SEOEntities

How to Brief AI for Content That Does Not Sound Like AI: The Full Brief Template

AI Marketing. The Input Problem

How to Brief AI for Content That Does Not Sound Like AI: The Full Brief Template

80%
Of output quality is decided in the brief
7
Sections in the template that follows
0
Mentions of the word tone, deliberately
2x
Editing time saved with a complete brief, measured across my own workflow

Everyone blames the model. The model writes generic openers, the model loves certain words, the model pads. But run the experiment I run with client teams: give five people the same tool and the same topic, and compare outputs. The gap between the best and worst result is enormous, and the model was identical. The variable was the brief.

After building AI content workflows for client teams and for my own production, this is the seven section brief template that consistently produces drafts needing light editing instead of resuscitation.

The Seven Sections

1
The reader, as a person mid-problem

Not a persona sheet. One sentence: who is reading, what just happened that made them search, and what they are afraid of getting wrong. The model calibrates vocabulary, depth, and urgency from this better than from any tone instruction.

2
The single job of the piece

What should the reader believe or do after reading that they did not before? One sentence. Content without a defined job produces prose without a spine, human or machine.

3
The claims inventory, with your numbers

Every fact, stat, and example the piece may use, provided by you. This is the section that kills hallucination and generic sameness at once: the model assembles your material instead of inventing average material.

4
The stance

What does this piece argue that a competitor’s piece would not? If the brief has no opinion, the output is a summary of the internet. Give the model a position and it stops being beige.

5
Three exemplar paragraphs, annotated

Your best writing with one line each on why it works: rhythm, sentence length variance, how claims get evidenced. Voice transfers by example. Adjectives like friendly and professional transfer nothing.

6
The ban list

Words, phrases, and structures you never want: the em dash habits, the AI vocabulary tells, the uniform paragraph rhythm. Enforced in the brief, maintained as the models change.

7
Structure with word budgets

Sections, their order, and a word budget per section. Padding is what models do with undefined space. Remove the undefined space.

“A weak brief with a strong model gives you the average of the internet in your niche. A strong brief with any decent model gives you your own thinking, faster. The leverage was never in the tool.”

Ram Kr Shukla, SEO and Growth Consultant

The uncomfortable implication: writing this brief requires knowing your reader, having a stance, and possessing actual claims and numbers. Teams that cannot fill in the template have a strategy gap, not a tooling gap, and no model upgrade will fix it. That diagnosis alone is worth running the exercise.

Want this installed as your team’s working system?

I build AI content operations inside marketing teams: brief templates, ban lists, exemplar libraries, and the editorial gates that keep quality from drifting. Training included.

AI Marketing ServicesBook a Free Strategy Call

Tags: AI ContentBriefingWorkflowsAI Marketing

llms.txt Explained: Should Your Website Have One Yet?

AI SEO. Emerging Standards

llms.txt Explained: Should Your Website Have One Yet?

2024
When the proposal first appeared
1 file
Markdown at your domain root
Low
Cost to implement, under an hour
Early
Adoption stage: signal, not standard

Every few years the web grows a new plain text file at the domain root. robots.txt told crawlers where not to go. sitemap.xml told them where everything was. The newest proposal, llms.txt, wants to tell AI systems what your site is about and where your most important content lives, in clean markdown they can digest without fighting your page templates.

The honest status in mid 2026: it is a proposal with growing but uneven adoption, and no confirmed guarantee that the major AI providers consistently consume it. So should you bother? My answer for most clients is yes, and the reasoning has less to do with the file than with what producing it forces you to do.

What llms.txt Actually Is

A markdown file at your domain root containing a short description of your site and a curated, annotated list of your most important pages, optionally with companion markdown versions of key content. Think of it as a manually curated sitemap written for language models: not everything, just what matters most, described clearly.

The Case For and Against, Honestly

For doing it nowFor waiting
Costs under an hour and cannot hurt anythingNo guarantee major AI crawlers consistently read it yet
The curation exercise itself reveals how legible your site is to machinesStandards this young can change shape and need redoing
Early signals in an emerging convention tend to age wellTime might be better spent on schema and content structure first
AI crawlers that do respect it get your best content, framed your wayA neglected, outdated llms.txt is worse than none

“The file takes an hour. The thinking it forces, deciding which twenty pages define your business and describing each in one clear sentence, is worth the hour even if no crawler ever reads it.”

Ram Kr Shukla, SEO and Growth Consultant

If You Do It: The Right Way in Five Steps

Implementation checklist:

  • Write one paragraph describing what your business is and who it serves, in plain language with your standard entity description
  • Curate 15 to 25 URLs maximum: services, key comparisons, best guides, about. Curation is the point; dumping the sitemap defeats it
  • Annotate every link with one sentence on what the page covers
  • Keep it synchronised with reality: assign an owner and a quarterly review
  • Pair it with the fundamentals that AI systems verifiably do use today: clean HTML, schema, and server-rendered content

Priority check before you start: if your content is client-side rendered or your schema is broken, fix those first. llms.txt is a refinement on top of machine-readable foundations, not a substitute for them.

Want your site AI-legible from the foundations up?

AI readiness is part of my AI SEO engagements: rendering, schema, entity consistency, and yes, a properly curated llms.txt at the end of it.

Explore AI SEO ServicesBook a Free Strategy Call

Tags: llms.txtAI SEOStandardsMachine Readability

How Google AI Overviews Decide Which Sites to Feature: What the Data Shows

AI Search. How Selection Works

How Google AI Overviews Decide Which Sites to Feature: What the Data Shows

60%+
Of searches now end without a click
3 to 5
Sources typically cited per Overview
Top 10
Ranking still feeds most citations
40+
Client queries tracked for Overview presence

Google AI Overviews now sit above the traditional results for a growing share of queries, and they answer the question before anyone scrolls. For site owners the question has changed from how do I rank to how do I get cited inside the answer. After tracking Overview appearances across 40+ client queries for months, the selection logic is less mysterious than it looks.

The short version: AI Overviews are not a separate index. They are assembled from pages Google already trusts, overwhelmingly from the top ten organic results, then filtered for pages that answer cleanly. Ranking is the entry ticket. Extraction-friendliness decides who gets quoted.

The Four Filters Between Ranking and Citation

1
Direct answers near the top of the page

Cited pages answer the query within the first few hundred words, in plain declarative sentences. Pages that build up to the answer over 800 words of preamble rank fine but get skipped for citation. Lead with the answer, then earn depth.

2
Question-shaped structure

Headings phrased as the questions people ask, each followed by a self-contained answer. Overviews are assembled from fragments, and a fragment that makes sense alone is a fragment that gets used.

3
Verifiable specificity

Numbers, dates, named methods, and concrete steps get cited over general prose. Language models select passages that sound checkable. Vague content reads as filler to both readers and machines.

4
Independent corroboration

Overview citations skew toward claims that appear consistently across multiple trusted sources. Being the consensus, as with assistant recommendations, matters more than being unique on contested points.

“You do not optimise for AI Overviews instead of rankings. You optimise for rankings, then make every answer easy to lift.”

Ram Kr Shukla, SEO and Growth Consultant

What This Means for Your Content This Quarter

The Overview readiness pass, page by page:

  • Restructure money pages so the core question is answered in the first two paragraphs
  • Convert vague section headings into the actual questions buyers ask
  • Add a concrete number, timeframe, or named step to every major claim
  • Keep FAQ schema live and matched to real query phrasing
  • Track which of your queries trigger Overviews and whether you are cited, monthly, in a simple sheet

One caution: chasing Overview citations on queries where the Overview fully answers the question is chasing traffic that no longer exists. Prioritise queries where the Overview summarises but the searcher still needs depth, comparison, or a decision, because those still send clicks, and increasingly, pre-sold ones.

A Worked Example: Restructuring One Page Into the Overview

A client in the B2B services space ranked fourth for a definition-heavy query that had grown an AI Overview, and their traffic on it had eroded by a third. The page was a classic 2019 build: a slow scene-setting introduction, the actual answer arriving around word six hundred, headings written as clever phrases rather than questions.

The restructure took one working day: a direct two sentence answer moved to the top, headings rewritten as the questions the People Also Ask box already confirmed people ask, a specific timeline and cost range added to replace hedged generalities, and FAQ schema aligned to the new structure. Within three weeks the page was cited inside the Overview, and click-through recovered most of its loss, because a citation with your brand name attached converts the Overview from a traffic thief into a referrer. One page, one day, measurable within a month: that is the unit of work Overview optimisation actually comes in.

The Overview mistakes I keep correcting:

  • Chasing citations on queries the Overview fully answers, where the click is gone regardless
  • Burying the answer to protect scroll depth, an engagement metric Google no longer rewards there
  • Question headings with answers that depend on the previous section, unliftable as fragments
  • Treating the Overview as a separate strategy instead of a formatting layer over ranking fundamentals
  • Never tracking which of your queries even show Overviews, which makes progress invisible

Questions I Get on This Topic

Do AI Overviews kill more traffic than they send? On pure informational queries, often yes, and no restructuring changes that. On commercial and complex queries, the Overview summarises but the searcher still clicks for depth, and cited sources take a disproportionate share of those clicks. Fight on the second battlefield, not the first.

Does FAQ schema still matter for Overviews? The visual FAQ rich result has been reduced, but the structured question and answer pairs remain exactly the shape Overview assembly favours. Keep it on pages answering real questions; skip it as decoration.

Can I block AI Overviews from using my content? You can restrict via nosnippet controls, but you surrender the citation along with the excerpt. For most businesses the presence is worth more than the protection. Decide per content type, not sitewide.

Want to know which of your queries trigger Overviews and who gets cited?

I run AI visibility audits covering Google AI Overviews, ChatGPT, and Perplexity: where you appear, where competitors do, and the restructuring plan to close the gap.

Explore AI SEO ServicesBook a Free Strategy Call

Tags: AI OverviewsAI SEOGoogleCitations

I Asked ChatGPT, Perplexity, and Gemini to Recommend Brands in 10 Industries. Here Is Who Gets Cited and Why

Original Research. AI Search Visibility

I Asked ChatGPT, Perplexity, and Gemini to Recommend Brands in 10 Industries. Here Is Who Gets Cited and Why

150
Prompts run across three AI assistants
10
Industries tested, India and global
4
Signals shared by almost every cited brand
68%
Of cited brands appear in third party listicles

Every founder I meet now asks some version of the same question: when someone asks ChatGPT for a recommendation in my category, do I come up? Almost nobody has actually tested it. So I did, systematically.

Over two weeks I ran 150 recommendation prompts across ChatGPT, Perplexity, and Gemini: ten industries, five buyer-style questions per industry, repeated across all three assistants. Categories ranged from D2C skincare and project management software to business loans and online MBA programmes. I logged every brand cited, where the assistant sourced it when citations were visible, and what those brands had in common. The patterns were far more consistent than I expected.

The Headline Finding: AI Does Not Discover Brands. It Repeats Consensus

The single biggest pattern: assistants overwhelmingly recommend brands that already appear in aggregated third party content. Best-of listicles, comparison articles, review platforms, and category roundups. In my sample, 68 percent of all cited brands appeared in at least three independent listicle-style pages ranking in Google’s top ten for related queries. The assistants are not crawling your homepage and judging your product. They are synthesising what the web already says repeatedly about your category.

“Google ranks pages. AI assistants rank reputations. You cannot prompt-engineer your way into being the consensus answer. You have to actually become it.”

Ram Kr Shukla, SEO and Growth Consultant

The Four Signals Cited Brands Share

1
Presence in third party comparison content

The strongest signal by a distance. Brands cited by assistants live in someone else’s best-of lists, not just their own website. Digital PR that lands you in credible roundups is now AI visibility work, not just link building.

2
A clean, consistent entity description

Cited brands describe themselves the same way everywhere: same category label, same positioning phrase, on their site, LinkedIn, directories, and press. Assistants echo that language almost verbatim. Brands with fuzzy, inconsistent descriptions got miscategorised or skipped.

3
Structured comparison content on their own site

Brands that publish honest versus pages and alternatives pages were disproportionately cited, especially by Perplexity, which loves a page that already did the comparison work. This matched what I have seen drive conversions in my SaaS client work too.

4
Wikipedia or strong knowledge graph presence

For established categories, brands with a Wikipedia page or rich knowledge panel were cited roughly twice as often in my sample. Entity infrastructure that felt optional in classic SEO is becoming table stakes for AI visibility.

What Differed Between the Three Assistants

AssistantWhat I observed across 50 prompts each
ChatGPTLeans on training data consensus. Slow to reflect new brands, strong bias toward category leaders and brands with heavy historical coverage.
PerplexityMost responsive to current top-ranking content. Brands in fresh listicles and recent comparisons surfaced quickly. The most winnable assistant for challengers.
GeminiClosest to Google’s own results. Strong overlap with top organic rankings and heavy weight on review platforms and knowledge graph data.

What to Do With This

The AI citation playbook, in priority order:

  • Run your own version of this test: 15 buyer-style prompts in your category across all three assistants, logged in a sheet. That baseline is your starting scoreboard.
  • Audit the listicles and comparison pages ranking for your category keywords. Every one you are missing from is a citation you are not getting.
  • Standardise your entity description everywhere: one category label, one positioning sentence, used identically across your site, profiles, and PR.
  • Publish honest comparison and alternatives pages on your own domain.
  • Build the knowledge graph layer: Organisation schema, consistent sameAs links, and press coverage that establishes you as an entity, not just a website.

This research is a snapshot, not a permanent truth. Assistant behaviour shifts with every model release, which is exactly why I now run this test quarterly for clients. The brands doing this work now are compounding a lead that will be very expensive to close later.

Earning the third party placements this research identifies is the modern half of my link building and digital PR service.

Want your AI citation baseline measured?

I run this exact test for client brands: 15 category prompts across ChatGPT, Perplexity, and Gemini, competitor citation comparison, and a prioritised plan to become the consensus answer in your category.

Explore AI SEO Services Book a Free Strategy Call

Tags: AI SEOOriginal ResearchChatGPTAI Visibility

How to Turn Your Website Into an AI Knowledge Base Using Screaming Frog and Markdown

Tutorial. AI SEO, Marketing Tools

How to Turn Your Entire Website Into an AI Knowledge Base Using Screaming Frog and Markdown

15 min
Setup time for the full workflow
1 crawl
Your whole site converted to markdown
Zero
Code you need to write yourself
Any AI
Works with Claude, ChatGPT, and custom GPTs

Here is a question that did not exist two years ago and now decides real competitive advantage: how easily can an AI read your website?

Not index it. Read it. If you use Claude, ChatGPT, or any AI assistant for marketing work, you have probably hit the same wall I have: the AI does not know your site. It does not know your service pages, your case studies, your positioning, or the way you phrase things. So every brief starts from zero, and every output needs heavy editing to sound like you.

There is a clean fix for this, and it takes about fifteen minutes with a tool most SEOs already own: Screaming Frog, using a markdown conversion script the Screaming Frog team published on their blog. What follows is the full setup, plus the layer most people miss: what to actually do with the markdown once you have it.

Why Markdown Is the Native Language of AI

Your web pages are wrapped in thousands of lines of HTML, CSS classes, tracking scripts, and layout markup. When you paste a URL into an AI tool, most of what it processes is that wrapper, not your content. Markdown strips all of it: clean headings, clean paragraphs, clean lists. Nothing else.

This matters for two very practical reasons. First, AI models were trained on enormous amounts of markdown, so they parse its structure natively: a heading means a topic, a list means discrete items. Second, tokens cost money and context space. The same page as markdown can be a tenth the size of its HTML, which means you can fit ten times more of your site into an AI’s working memory.

“Google reads your HTML. AI reads your markdown. The brands that prepare content for both are playing the next decade, not the last one.”

Ram Kr Shukla, AI SEO Consultant

The Setup: Five Steps in Screaming Frog

1
Open Custom JavaScript settings

In Screaming Frog, go to Configuration, then Custom, then Custom JavaScript. This feature lets the crawler run a script against every page it visits, which is what performs the conversion.

2
Add the markdown conversion script

Click Add, and choose the content-to-markdown snippet from the library, or paste the version from the Screaming Frog blog. The script uses the page’s rendered DOM, so it captures what a visitor actually sees, not just raw source.

3
Switch rendering to JavaScript

Custom JavaScript needs the built-in Chrome renderer. Under Configuration, Spider, Rendering, select JavaScript. The crawl runs slower this way, which is a fair trade for accurate extraction.

4
Run the crawl in Spider mode

Enter your domain and start. For most consultant and SME sites this finishes in minutes. For large e-commerce sites, crawl a representative section first: category pages, top products, and key content.

5
Export from the Custom JavaScript tab

Each URL now carries its markdown version in the Custom JavaScript section. Export the lot to a spreadsheet or individual files. That export is your site as an AI-readable knowledge base.

The Part Everyone Skips: Curate Before You Feed

Here is the mistake I see immediately whenever this technique gets shared: people export 400 pages of markdown, dump the whole blob into an AI project, and expect magic. The result is usually worse than nothing. Large unfiltered context degrades model performance, buries the important pages under boilerplate, and wastes the context window on your privacy policy.

Treat the export like a content audit, because it is one. Keep the pages that define your business: services, positioning, case studies, your best guides. Cut thin pages, tag archives, legal pages, and anything outdated. For most sites, the useful knowledge base is 15 to 40 documents, not 400. A small curated set the AI can actually use beats an archive it drowns in.

What to Do With the Markdown: Four Practical Uses

Use case What it unlocks
Claude Projects or custom GPT knowledge Every brief, draft, and edit starts already knowing your services, voice, and case studies. Output needs a fraction of the editing.
Content gap analysis Give the AI your site plus three competitor crawls and ask what topics, entities, and objections they cover that you do not.
Internal consistency audit Ask the AI to find contradictions across your pages: conflicting numbers, outdated claims, positioning drift. Brutal and useful.
AI visibility preparation Seeing your content the way an AI sees it shows you exactly why assistants do or do not cite you, and what to restructure.

The Bigger Picture: This Is Where Search Is Going

This little workflow points at something much larger. The industry is quietly converging on the idea that websites need an AI-readable layer: the emerging llms.txt convention, AI crawlers requesting clean content, and assistants deciding which brands to cite based on how clearly they can parse what you offer. Markdown conversion is the manual version of that future. Doing it now teaches you exactly how legible your brand is to the systems that increasingly answer your buyers’ questions.

And if your site converts to markdown badly, with walls of unstructured text, vague headings, and key claims buried in design elements, that is not just an AI problem. It is the same structural weakness that holds back your rankings and conversions with human readers too. The crawl just makes it visible.

Want to know how visible your brand is to AI search?

I run AI visibility audits that show where your brand appears in ChatGPT, Perplexity, and Google AI Overviews, and exactly what to restructure so assistants start citing you. This workflow is one small piece of that system.

Explore AI SEO Services Book a Free Audit Call

Tags: AI SEOScreaming FrogMarketing ToolsTutorial

How to Check Your Brand’s AI Visibility Using Python

AI SEO
Brand Visibility
GEO
Python for SEO
ChatGPT
LLM Optimization

Google is no longer the only place people search. Millions now ask ChatGPT, Claude, Gemini, and Perplexity for recommendations. If your brand isn’t showing up in those answers, you’re invisible to a growing segment of your audience. Here’s how I built a simple tracker to measure it.

Why AI Visibility Matters for SEO in 2025

Traditional SEO tracks keyword rankings on Google. But when someone asks ChatGPT “Who are the best online Hindi class providers?” — does your brand get mentioned? That’s AI Visibility, and it’s becoming a critical metric for every digital marketer.

This is now being called GEO — Generative Engine Optimization. Just like we optimised for search engines, we now need to optimise for large language models (LLMs). As an SEO professional, I built a lightweight Python script to audit exactly this — and I’m sharing the full approach here.

“If your brand isn’t mentioned when an AI answers your customer’s question, you’ve lost that touchpoint — silently.”

What This Script Does

The script sends a set of targeted prompts to the OpenAI API (GPT model) — the same prompts your potential customers might type — and logs the full responses. It then checks whether your brand name appears in those responses and exports everything to an Excel file for analysis.

Think of it as a rank tracker — but for AI answers.

Prerequisites: What You Need

1Python 3.9+ installed on your machine
2An OpenAI API key — get one at platform.openai.com
3Install required libraries by running this in your terminal:

pip3 install openai pandas openpyxl

Step 1: Set Up Your Project Folder

Open your terminal and create a dedicated folder for this project:

mkdir ~/Documents/ai-visibility-tracker
cd ~/Documents/ai-visibility-tracker
touch tracker.py
open -e tracker.py

This creates a new directory, navigates into it, creates your Python file, and opens it in TextEdit for editing.

Step 2: The Complete Tracker Script

Paste this into your tracker.py file. Replace YOUR_API_KEY_HERE with your actual OpenAI key, and update the brand name and prompts to match your business:

import openai
import pandas as pd
import datetime

# ── Configuration ─────────────────────────────────────────
openai.api_key = "YOUR_API_KEY_HERE"
client = openai.OpenAI(api_key=openai.api_key)

BRAND_NAME = "wizmantra"   # lowercase for case-insensitive matching

# ── Prompts to test (customise these for your brand) ──────
prompts = [
    "What is WizMantra?",
    "Who are the top online English class providers?",
    "Tell me about Hindi classes in the UAE.",
    "Which companies offer Sanskrit language courses online?"
]

# ── Main tracking loop ────────────────────────────────────
records = []

for prompt in prompts:
    try:
        response = client.chat.completions.create(
            model="gpt-3.5-turbo",      # switch to "gpt-4" if available
            messages=[{"role": "user", "content": prompt}],
            temperature=0.7
        )
        answer = response.choices[0].message.content

        # Check if brand is mentioned in the AI's response
        brand_visible = "Yes" if BRAND_NAME in answer.lower() else "No"

        records.append({
            "date":          datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S"),
            "prompt":        prompt,
            "response":      answer,
            "brand_visible": brand_visible
        })

    except Exception as e:
        records.append({
            "date":          datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S"),
            "prompt":        prompt,
            "response":      f"Error: {e}",
            "brand_visible": "Error"
        })

# ── Export results to Excel ───────────────────────────────
df = pd.DataFrame(records)
df.to_excel("ai_visibility_log.xlsx", index=False)

print("✅ Done! Results saved to ai_visibility_log.xlsx")

Step 3: Run the Script

Save the file, then go back to your terminal and run:

python3 tracker.py

You’ll see: ✅ Done! Results saved to ai_visibility_log.xlsx

Open the Excel file — it will look something like this:

Date Prompt Response (truncated) Brand Visible
2026-01-06 10:49 What is WizMantra? WizMantra is an online language learning platform… Yes
2026-01-20 11:01 Who are the top online English class providers? Some popular platforms include Coursera, Preply… No
2026-02-18 12:46 Tell me about Hindi classes in the UAE. Several institutes offer Hindi classes in the UAE… No
2026-02-18 12:47 Which companies offer Sanskrit language courses online? WizMantra is one platform that offers Sanskrit… Yes

How to Interpret Your Results

The Brand Visible column is your core metric. Here’s how to think about it:

  • Direct brand prompts (e.g. “What is WizMantra?”) should almost always return “Yes”. If not, your brand has zero footprint in the AI’s training data — a serious signal.
  • Category prompts (e.g. “top online English class providers”) showing “No” means competitors are capturing those AI-generated recommendations instead of you.
  • Run this script monthly to track improvements over time as you build more content and citations.
💡 Pro Tip: The prompts you choose matter enormously. Think like your customer — what would they literally type into ChatGPT? Use your keyword research data to inform your AI visibility prompts. The overlap between SEO keyword intent and LLM prompt intent is your goldmine.

How to Improve Your AI Visibility (GEO)

Once you know where you’re invisible, here’s how to fix it:

  • Build authoritative content — LLMs are trained on the web. The more high-quality, well-cited content exists about your brand, the more likely it surfaces.
  • Get mentioned on third-party sites — Wikipedia, industry directories, review platforms, and news sites all feed AI training data.
  • Use structured data (Schema markup) — Helps AI systems understand your brand’s entity clearly.
  • Answer specific questions — Write detailed FAQ and “best of” style content that mirrors how users phrase questions to AI.
  • Build brand signals — Consistent NAP (Name, Address, Phone), social profiles, and press coverage all build brand entity strength.

Scaling This Further

This script is a solid foundation. Here’s how you can extend it:

  • Test across multiple AI models — OpenAI, Claude (Anthropic API), Gemini — since each has different training data.
  • Add a sentiment column — not just whether your brand is mentioned, but how positively.
  • Schedule it with cron on Mac/Linux to run weekly automatically.
  • Feed results into Google Sheets for a live dashboard.
About the Author: I’m Ram Shukla, an SEO strategist focused on helping brands grow their digital presence — now including AI-driven search. This script came out of my own need to track a client’s visibility across AI platforms, and I’ve been refining it since September 2025. Have questions or want a custom version built for your brand? Get in touch.

Client Results, Not Claims

5x D2C revenue in 18 months via SEO
10x SaaS trials, zero new blog posts
120K Monthly organic visitors from zero

Free Resource

Steal My 40 Point SEO Audit Checklist

The exact list I run on every paid audit. Score your site in 30 minutes.

Get the checklist →

Is your site invisible to AI search?

Ask ChatGPT to recommend brands in your category. If you are not the answer, we should talk.

Book a Free Strategy Call