AI Search Is Splitting Your Traffic: How to Measure ChatGPT and Perplexity Referrals in GA4
4x
Growth in AI referral traffic across client sites this year
15 min
To build the GA4 view that isolates it
2x+
Conversion rate of AI referrals versus average, consistently
0
Standard reports that show this today
There is a new traffic source growing in your analytics right now, and GA4’s default reports are hiding it inside the referral bucket. Visitors arriving from ChatGPT, Perplexity, Gemini, and Copilot click through when an assistant cites you, and across my client base that stream has roughly quadrupled this year. Small absolute numbers still, but the trend line and the visitor quality demand their own measurement.
Because here is the finding that matters: AI referrals convert at roughly double the site average, consistently, across clients. It makes sense. A visitor arriving from an assistant’s recommendation was pre-sold by the most trusted interface they use. You want to know exactly how many arrive and what they do.
Build the View in GA4, Step by Step
1
Know the referrer signatures
AI referrals arrive with identifiable referrers: chatgpt.com and chat.openai.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, claude.ai, and you.com. Some assistant clicks arrive stripped as direct, so whatever you measure is a floor, not a ceiling. Floors still show trends.
2
Create the channel group
In Admin, under Data display, edit your channel groups: add an AI Referrals channel matching source against a regex of those domains. From that point on, standard reports break the stream out cleanly beside Organic Search and the rest.
3
Build the exploration for history
Channel groups apply going forward. For history, build an Exploration filtered on session source matching the same regex: sessions, engaged sessions, conversions, and landing pages, trended monthly. This is your baseline chart, and it will be a hockey stick at small scale.
4
Study the landing pages above all
Which pages do assistants send people to? That list is your citation footprint made visible: it tells you which content AI systems consider reference-worthy, and it is the template for what to produce more of. This single report has reshaped content plans for my clients more than any keyword tool this year.
Reading the Numbers Honestly
Metric
How to read it
Volume
Small today almost everywhere. Single digit percentages. The trend is the story, not the total
Conversion rate
Consistently above site average. Pre-sold visitors behave like referrals from a trusted friend
Landing pages
Comparisons, definitive guides, and original data get cited. Thin service pages do not
Month over month
The only chart worth presenting. Screenshot it quarterly; the compounding makes the argument for you
“Nobody reported on organic search traffic in 1999 either. The brands that measured it early built the playbooks everyone else bought later. AI referrals are that moment again, with a fifteen minute setup cost.”
Ram Kr Shukla, SEO and Growth Consultant
Once the measurement exists, it changes decisions: the landing page report feeds your content plan, the conversion data justifies AI visibility investment to whoever owns the budget, and the baseline lets you actually evaluate whether citation work, the kind I documented in my AI recommendations research, moves your numbers. Measurement first, then optimisation has something to answer to.
Want the full AI visibility measurement stack installed?
I set up AI referral tracking, citation monitoring across assistants, and the quarterly baseline test as part of every AI SEO engagement. Know your numbers before your competitors know theirs.
Programmatic SEO Done Right: How to Scale Pages Without Getting Penalised
2M
Pages I have managed at marketplace scale
1
Question that decides survival: is each page useful alone?
10x
Faster with AI, which cuts both ways
3
Structural rules that separate assets from spam
Programmatic SEO, generating pages from data and templates instead of writing them one by one, built some of the biggest organic footprints on the web. I learned it managing marketplace sites with millions of URLs, long before AI made generation trivial. And that is exactly the new problem: AI made scaling easy, so the graveyard of penalised programmatic sites is filling faster than ever.
The difference between a programmatic asset and thin content spam was never the page count. It is whether each generated page would deserve to exist if you built it by hand. That standard is passable, and here is how the survivors pass it.
The Three Structural Rules
1
Unique data per page, not unique words
Spinning synonyms across a template fools nobody since Google’s helpful content systems matured. What works: each page assembled around data that genuinely differs, prices, specs, availability, local details, real comparisons. If two pages would answer a visitor identically, they should be one page.
2
A query with real intent behind every URL
Programmatic pages must map to searches people actually make: city plus service, product plus comparison, tool plus integration. Generating pages for query patterns nobody searches produces index bloat that drags the whole domain. Validate the pattern’s demand before generating the ten thousandth page.
3
Crawl architecture that carries the weight
Ten thousand pages need hub structure, clean pagination, and internal links that give every page a path from the money pages. Generated pages dumped into a flat sitemap with no internal linking are orphans at scale, and Google treats them accordingly.
Where AI Fits, and Where It Breaks Things
AI role
Verdict
Assembling pages from your structured data
Ideal. This is what the technique always was, with better tooling
Writing unique intros against real data per page
Works with editorial gates and the brief discipline from my AI content workflow
Generating both the data and the words
The graveyard. Invented data plus template prose is the exact pattern quality systems now catch
Deciding which pages deserve to exist
Never. Demand validation and pruning stay human decisions
“Programmatic SEO fails at the same question every time: would this page deserve to exist if a human had to build it? AI changed the cost of generating pages. It did not change the standard they are held to.”
Ram Kr Shukla, SEO and Growth Consultant
Start smaller than you want to: one pattern, a few hundred pages, indexed and measured for a quarter. Scale what earns impressions and conversions; prune what does not, ruthlessly. The discipline to delete generated pages is rarer than the ability to generate them, and it is what keeps the domain healthy.
Considering programmatic at scale?
I have run this at marketplace scale and rebuilt it after others’ penalties. Pattern validation, data architecture, and the guardrails, before you generate page one.
Schema Markup in 2026: The Only 6 Types Most Businesses Actually Need
6
Types that earn their maintenance cost
30+
Types most plugins offer that you can ignore
2
Jobs schema now does: rich results and AI comprehension
15 min
To validate everything you currently have
Schema markup suffers from a completeness problem: there are hundreds of types, plugins offer dozens, and teams either implement nothing or implement everything badly. Broken schema is worse than none; I covered a client whose rich results vanished sitewide from one invalid plugin update in my rankings emergencies post.
Schema also quietly picked up a second job. It used to earn rich results in Google. Now it also tells AI systems, the assistants and the Overview builders, exactly who you are and what you offer in a format they parse without guessing. Both jobs are done by the same six types for most businesses.
The Six That Matter
1
Organisation, on every page
Name, logo, URL, and sameAs links to your real profiles. This is your entity anchor: the machine-readable statement of who is behind the website. For consultants and personal brands, Person schema does this job alongside it.
2
Service or Product, on money pages
What you sell, described in structured form. For e-commerce: Product with price, availability, and ratings, which earns the rich result that moves click-through even without a rank change. For services: Service with provider and area.
3
FAQ, where real questions live
On pages answering genuine questions buyers ask. Google trimmed its visual FAQ real estate, but the structured question and answer pairs are exactly the fragments AI systems lift. Write real questions, not keyword-stuffed fakes.
4
Article, on content that argues expertise
With author, dates, and publisher. This connects your content to your entity, which is how expertise accumulates to a name instead of evaporating page by page.
5
BreadcrumbList, sitewide
Cheap to implement, improves the SERP display, and reinforces your site architecture to crawlers. The lowest-effort item on this list.
6
Review or AggregateRating, where earned
Ratings you genuinely collect, marked up where they live. Never fabricated and never sitewide decoration: unearned rating markup is the fastest route to a manual action of anything on this page.
What to Skip, and the Maintenance Rule
Skip the exotic types your plugin offers unless you demonstrably need them: Event without events, VideoObject without video, HowTo on pages that are not how-tos. Every type you add is a validation surface that can break silently in a plugin update. The maintenance rule that survives contact with reality: validate quarterly in Search Console’s Enhancements reports, and after every plugin or theme update. Fifteen minutes, calendared, owned by a named person.
“Schema used to be decoration for rich results. It is becoming your machine-readable identity. Six types, kept valid, beat thirty types nobody checks.”
Ram Kr Shukla, SEO and Growth Consultant
A Worked Example: The Silent Breakage and the Recovery
The client I mentioned in my rankings emergencies post is worth expanding here, because the sequence is so typical. A plugin update changed how review markup rendered: technically present, semantically invalid. Google stopped showing stars sitewide within days. Rankings barely moved, but click-through fell by nearly a third, and revenue followed. It took weeks to notice precisely because everyone was watching rankings, the metric that had not changed.
The recovery was mechanical once diagnosed: valid markup restored, Search Console validation run, rich results back within two crawl cycles. The durable fix was procedural, the quarterly validation calendar and a named owner. Schema does not fail loudly. It fails like a fridge light: you only notice when you finally open the door and check.
Marking up ratings that do not visibly exist on the page, a manual action magnet
Duplicate Organisation schema from theme, plugin, and manual code all firing at once
FAQ markup on invented questions no one asks, which wastes the one fragment AI systems lift
Schema in the page builder that disappears when the template changes
Validating in a testing tool once at launch and never in Search Console where field reality lives
Questions I Get on This Topic
JSON-LD or microdata? JSON-LD, without hesitation. It lives in one script block, survives template edits far better, and is what Google recommends. If a plugin outputs microdata woven through your HTML, that is a fragility worth migrating away from.
Does schema directly improve rankings? Not as a ranking factor in the direct sense. It improves how results display, which moves click-through, and it improves machine comprehension of who you are, which increasingly matters beyond Google. The revenue path is real; it just does not run through the position number.
How do I add schema on WordPress without another plugin? Most SEO plugins you already run handle the six core types adequately. The gap is usually configuration and validation, not tooling. Before adding anything, audit what your current stack already outputs; duplication is the more common disease than absence.
Want your structured data audited and rebuilt properly?
Schema architecture is part of every technical SEO engagement: what you have, what is broken, what is missing, and the six type implementation your stack actually needs.
How to Brief AI for Content That Does Not Sound Like AI: The Full Brief Template
80%
Of output quality is decided in the brief
7
Sections in the template that follows
0
Mentions of the word tone, deliberately
2x
Editing time saved with a complete brief, measured across my own workflow
Everyone blames the model. The model writes generic openers, the model loves certain words, the model pads. But run the experiment I run with client teams: give five people the same tool and the same topic, and compare outputs. The gap between the best and worst result is enormous, and the model was identical. The variable was the brief.
After building AI content workflows for client teams and for my own production, this is the seven section brief template that consistently produces drafts needing light editing instead of resuscitation.
The Seven Sections
1
The reader, as a person mid-problem
Not a persona sheet. One sentence: who is reading, what just happened that made them search, and what they are afraid of getting wrong. The model calibrates vocabulary, depth, and urgency from this better than from any tone instruction.
2
The single job of the piece
What should the reader believe or do after reading that they did not before? One sentence. Content without a defined job produces prose without a spine, human or machine.
3
The claims inventory, with your numbers
Every fact, stat, and example the piece may use, provided by you. This is the section that kills hallucination and generic sameness at once: the model assembles your material instead of inventing average material.
4
The stance
What does this piece argue that a competitor’s piece would not? If the brief has no opinion, the output is a summary of the internet. Give the model a position and it stops being beige.
5
Three exemplar paragraphs, annotated
Your best writing with one line each on why it works: rhythm, sentence length variance, how claims get evidenced. Voice transfers by example. Adjectives like friendly and professional transfer nothing.
6
The ban list
Words, phrases, and structures you never want: the em dash habits, the AI vocabulary tells, the uniform paragraph rhythm. Enforced in the brief, maintained as the models change.
7
Structure with word budgets
Sections, their order, and a word budget per section. Padding is what models do with undefined space. Remove the undefined space.
“A weak brief with a strong model gives you the average of the internet in your niche. A strong brief with any decent model gives you your own thinking, faster. The leverage was never in the tool.”
Ram Kr Shukla, SEO and Growth Consultant
The uncomfortable implication: writing this brief requires knowing your reader, having a stance, and possessing actual claims and numbers. Teams that cannot fill in the template have a strategy gap, not a tooling gap, and no model upgrade will fix it. That diagnosis alone is worth running the exercise.
Want this installed as your team’s working system?
I build AI content operations inside marketing teams: brief templates, ban lists, exemplar libraries, and the editorial gates that keep quality from drifting. Training included.
llms.txt Explained: Should Your Website Have One Yet?
2024
When the proposal first appeared
1 file
Markdown at your domain root
Low
Cost to implement, under an hour
Early
Adoption stage: signal, not standard
Every few years the web grows a new plain text file at the domain root. robots.txt told crawlers where not to go. sitemap.xml told them where everything was. The newest proposal, llms.txt, wants to tell AI systems what your site is about and where your most important content lives, in clean markdown they can digest without fighting your page templates.
The honest status in mid 2026: it is a proposal with growing but uneven adoption, and no confirmed guarantee that the major AI providers consistently consume it. So should you bother? My answer for most clients is yes, and the reasoning has less to do with the file than with what producing it forces you to do.
What llms.txt Actually Is
A markdown file at your domain root containing a short description of your site and a curated, annotated list of your most important pages, optionally with companion markdown versions of key content. Think of it as a manually curated sitemap written for language models: not everything, just what matters most, described clearly.
The Case For and Against, Honestly
For doing it now
For waiting
Costs under an hour and cannot hurt anything
No guarantee major AI crawlers consistently read it yet
The curation exercise itself reveals how legible your site is to machines
Standards this young can change shape and need redoing
Early signals in an emerging convention tend to age well
Time might be better spent on schema and content structure first
AI crawlers that do respect it get your best content, framed your way
A neglected, outdated llms.txt is worse than none
“The file takes an hour. The thinking it forces, deciding which twenty pages define your business and describing each in one clear sentence, is worth the hour even if no crawler ever reads it.”
Ram Kr Shukla, SEO and Growth Consultant
If You Do It: The Right Way in Five Steps
Implementation checklist:
Write one paragraph describing what your business is and who it serves, in plain language with your standard entity description
Curate 15 to 25 URLs maximum: services, key comparisons, best guides, about. Curation is the point; dumping the sitemap defeats it
Annotate every link with one sentence on what the page covers
Keep it synchronised with reality: assign an owner and a quarterly review
Pair it with the fundamentals that AI systems verifiably do use today: clean HTML, schema, and server-rendered content
Priority check before you start: if your content is client-side rendered or your schema is broken, fix those first. llms.txt is a refinement on top of machine-readable foundations, not a substitute for them.
Want your site AI-legible from the foundations up?
AI readiness is part of my AI SEO engagements: rendering, schema, entity consistency, and yes, a properly curated llms.txt at the end of it.
How Google AI Overviews Decide Which Sites to Feature: What the Data Shows
60%+
Of searches now end without a click
3 to 5
Sources typically cited per Overview
Top 10
Ranking still feeds most citations
40+
Client queries tracked for Overview presence
Google AI Overviews now sit above the traditional results for a growing share of queries, and they answer the question before anyone scrolls. For site owners the question has changed from how do I rank to how do I get cited inside the answer. After tracking Overview appearances across 40+ client queries for months, the selection logic is less mysterious than it looks.
The short version: AI Overviews are not a separate index. They are assembled from pages Google already trusts, overwhelmingly from the top ten organic results, then filtered for pages that answer cleanly. Ranking is the entry ticket. Extraction-friendliness decides who gets quoted.
The Four Filters Between Ranking and Citation
1
Direct answers near the top of the page
Cited pages answer the query within the first few hundred words, in plain declarative sentences. Pages that build up to the answer over 800 words of preamble rank fine but get skipped for citation. Lead with the answer, then earn depth.
2
Question-shaped structure
Headings phrased as the questions people ask, each followed by a self-contained answer. Overviews are assembled from fragments, and a fragment that makes sense alone is a fragment that gets used.
3
Verifiable specificity
Numbers, dates, named methods, and concrete steps get cited over general prose. Language models select passages that sound checkable. Vague content reads as filler to both readers and machines.
4
Independent corroboration
Overview citations skew toward claims that appear consistently across multiple trusted sources. Being the consensus, as with assistant recommendations, matters more than being unique on contested points.
“You do not optimise for AI Overviews instead of rankings. You optimise for rankings, then make every answer easy to lift.”
Ram Kr Shukla, SEO and Growth Consultant
What This Means for Your Content This Quarter
The Overview readiness pass, page by page:
Restructure money pages so the core question is answered in the first two paragraphs
Convert vague section headings into the actual questions buyers ask
Add a concrete number, timeframe, or named step to every major claim
Keep FAQ schema live and matched to real query phrasing
Track which of your queries trigger Overviews and whether you are cited, monthly, in a simple sheet
One caution: chasing Overview citations on queries where the Overview fully answers the question is chasing traffic that no longer exists. Prioritise queries where the Overview summarises but the searcher still needs depth, comparison, or a decision, because those still send clicks, and increasingly, pre-sold ones.
A Worked Example: Restructuring One Page Into the Overview
A client in the B2B services space ranked fourth for a definition-heavy query that had grown an AI Overview, and their traffic on it had eroded by a third. The page was a classic 2019 build: a slow scene-setting introduction, the actual answer arriving around word six hundred, headings written as clever phrases rather than questions.
The restructure took one working day: a direct two sentence answer moved to the top, headings rewritten as the questions the People Also Ask box already confirmed people ask, a specific timeline and cost range added to replace hedged generalities, and FAQ schema aligned to the new structure. Within three weeks the page was cited inside the Overview, and click-through recovered most of its loss, because a citation with your brand name attached converts the Overview from a traffic thief into a referrer. One page, one day, measurable within a month: that is the unit of work Overview optimisation actually comes in.
The Overview mistakes I keep correcting:
Chasing citations on queries the Overview fully answers, where the click is gone regardless
Burying the answer to protect scroll depth, an engagement metric Google no longer rewards there
Question headings with answers that depend on the previous section, unliftable as fragments
Treating the Overview as a separate strategy instead of a formatting layer over ranking fundamentals
Never tracking which of your queries even show Overviews, which makes progress invisible
Questions I Get on This Topic
Do AI Overviews kill more traffic than they send? On pure informational queries, often yes, and no restructuring changes that. On commercial and complex queries, the Overview summarises but the searcher still clicks for depth, and cited sources take a disproportionate share of those clicks. Fight on the second battlefield, not the first.
Does FAQ schema still matter for Overviews? The visual FAQ rich result has been reduced, but the structured question and answer pairs remain exactly the shape Overview assembly favours. Keep it on pages answering real questions; skip it as decoration.
Can I block AI Overviews from using my content? You can restrict via nosnippet controls, but you surrender the citation along with the excerpt. For most businesses the presence is worth more than the protection. Decide per content type, not sitewide.
Want to know which of your queries trigger Overviews and who gets cited?
I run AI visibility audits covering Google AI Overviews, ChatGPT, and Perplexity: where you appear, where competitors do, and the restructuring plan to close the gap.
I Asked ChatGPT, Perplexity, and Gemini to Recommend Brands in 10 Industries. Here Is Who Gets Cited and Why
150
Prompts run across three AI assistants
10
Industries tested, India and global
4
Signals shared by almost every cited brand
68%
Of cited brands appear in third party listicles
Every founder I meet now asks some version of the same question: when someone asks ChatGPT for a recommendation in my category, do I come up? Almost nobody has actually tested it. So I did, systematically.
Over two weeks I ran 150 recommendation prompts across ChatGPT, Perplexity, and Gemini: ten industries, five buyer-style questions per industry, repeated across all three assistants. Categories ranged from D2C skincare and project management software to business loans and online MBA programmes. I logged every brand cited, where the assistant sourced it when citations were visible, and what those brands had in common. The patterns were far more consistent than I expected.
The Headline Finding: AI Does Not Discover Brands. It Repeats Consensus
The single biggest pattern: assistants overwhelmingly recommend brands that already appear in aggregated third party content. Best-of listicles, comparison articles, review platforms, and category roundups. In my sample, 68 percent of all cited brands appeared in at least three independent listicle-style pages ranking in Google’s top ten for related queries. The assistants are not crawling your homepage and judging your product. They are synthesising what the web already says repeatedly about your category.
“Google ranks pages. AI assistants rank reputations. You cannot prompt-engineer your way into being the consensus answer. You have to actually become it.”
Ram Kr Shukla, SEO and Growth Consultant
The Four Signals Cited Brands Share
1
Presence in third party comparison content
The strongest signal by a distance. Brands cited by assistants live in someone else’s best-of lists, not just their own website. Digital PR that lands you in credible roundups is now AI visibility work, not just link building.
2
A clean, consistent entity description
Cited brands describe themselves the same way everywhere: same category label, same positioning phrase, on their site, LinkedIn, directories, and press. Assistants echo that language almost verbatim. Brands with fuzzy, inconsistent descriptions got miscategorised or skipped.
3
Structured comparison content on their own site
Brands that publish honest versus pages and alternatives pages were disproportionately cited, especially by Perplexity, which loves a page that already did the comparison work. This matched what I have seen drive conversions in my SaaS client work too.
4
Wikipedia or strong knowledge graph presence
For established categories, brands with a Wikipedia page or rich knowledge panel were cited roughly twice as often in my sample. Entity infrastructure that felt optional in classic SEO is becoming table stakes for AI visibility.
What Differed Between the Three Assistants
Assistant
What I observed across 50 prompts each
ChatGPT
Leans on training data consensus. Slow to reflect new brands, strong bias toward category leaders and brands with heavy historical coverage.
Perplexity
Most responsive to current top-ranking content. Brands in fresh listicles and recent comparisons surfaced quickly. The most winnable assistant for challengers.
Gemini
Closest to Google’s own results. Strong overlap with top organic rankings and heavy weight on review platforms and knowledge graph data.
What to Do With This
The AI citation playbook, in priority order:
Run your own version of this test: 15 buyer-style prompts in your category across all three assistants, logged in a sheet. That baseline is your starting scoreboard.
Audit the listicles and comparison pages ranking for your category keywords. Every one you are missing from is a citation you are not getting.
Standardise your entity description everywhere: one category label, one positioning sentence, used identically across your site, profiles, and PR.
Publish honest comparison and alternatives pages on your own domain.
Build the knowledge graph layer: Organisation schema, consistent sameAs links, and press coverage that establishes you as an entity, not just a website.
This research is a snapshot, not a permanent truth. Assistant behaviour shifts with every model release, which is exactly why I now run this test quarterly for clients. The brands doing this work now are compounding a lead that will be very expensive to close later.
I run this exact test for client brands: 15 category prompts across ChatGPT, Perplexity, and Gemini, competitor citation comparison, and a prioritised plan to become the consensus answer in your category.
How to Turn Your Entire Website Into an AI Knowledge Base Using Screaming Frog and Markdown
15 min
Setup time for the full workflow
1 crawl
Your whole site converted to markdown
Zero
Code you need to write yourself
Any AI
Works with Claude, ChatGPT, and custom GPTs
Here is a question that did not exist two years ago and now decides real competitive advantage: how easily can an AI read your website?
Not index it. Read it. If you use Claude, ChatGPT, or any AI assistant for marketing work, you have probably hit the same wall I have: the AI does not know your site. It does not know your service pages, your case studies, your positioning, or the way you phrase things. So every brief starts from zero, and every output needs heavy editing to sound like you.
There is a clean fix for this, and it takes about fifteen minutes with a tool most SEOs already own: Screaming Frog, using a markdown conversion script the Screaming Frog team published on their blog. What follows is the full setup, plus the layer most people miss: what to actually do with the markdown once you have it.
Why Markdown Is the Native Language of AI
Your web pages are wrapped in thousands of lines of HTML, CSS classes, tracking scripts, and layout markup. When you paste a URL into an AI tool, most of what it processes is that wrapper, not your content. Markdown strips all of it: clean headings, clean paragraphs, clean lists. Nothing else.
This matters for two very practical reasons. First, AI models were trained on enormous amounts of markdown, so they parse its structure natively: a heading means a topic, a list means discrete items. Second, tokens cost money and context space. The same page as markdown can be a tenth the size of its HTML, which means you can fit ten times more of your site into an AI’s working memory.
“Google reads your HTML. AI reads your markdown. The brands that prepare content for both are playing the next decade, not the last one.”
Ram Kr Shukla, AI SEO Consultant
The Setup: Five Steps in Screaming Frog
1
Open Custom JavaScript settings
In Screaming Frog, go to Configuration, then Custom, then Custom JavaScript. This feature lets the crawler run a script against every page it visits, which is what performs the conversion.
2
Add the markdown conversion script
Click Add, and choose the content-to-markdown snippet from the library, or paste the version from the Screaming Frog blog. The script uses the page’s rendered DOM, so it captures what a visitor actually sees, not just raw source.
3
Switch rendering to JavaScript
Custom JavaScript needs the built-in Chrome renderer. Under Configuration, Spider, Rendering, select JavaScript. The crawl runs slower this way, which is a fair trade for accurate extraction.
4
Run the crawl in Spider mode
Enter your domain and start. For most consultant and SME sites this finishes in minutes. For large e-commerce sites, crawl a representative section first: category pages, top products, and key content.
5
Export from the Custom JavaScript tab
Each URL now carries its markdown version in the Custom JavaScript section. Export the lot to a spreadsheet or individual files. That export is your site as an AI-readable knowledge base.
The Part Everyone Skips: Curate Before You Feed
Here is the mistake I see immediately whenever this technique gets shared: people export 400 pages of markdown, dump the whole blob into an AI project, and expect magic. The result is usually worse than nothing. Large unfiltered context degrades model performance, buries the important pages under boilerplate, and wastes the context window on your privacy policy.
Treat the export like a content audit, because it is one. Keep the pages that define your business: services, positioning, case studies, your best guides. Cut thin pages, tag archives, legal pages, and anything outdated. For most sites, the useful knowledge base is 15 to 40 documents, not 400. A small curated set the AI can actually use beats an archive it drowns in.
What to Do With the Markdown: Four Practical Uses
Use case
What it unlocks
Claude Projects or custom GPT knowledge
Every brief, draft, and edit starts already knowing your services, voice, and case studies. Output needs a fraction of the editing.
Content gap analysis
Give the AI your site plus three competitor crawls and ask what topics, entities, and objections they cover that you do not.
Internal consistency audit
Ask the AI to find contradictions across your pages: conflicting numbers, outdated claims, positioning drift. Brutal and useful.
AI visibility preparation
Seeing your content the way an AI sees it shows you exactly why assistants do or do not cite you, and what to restructure.
The Bigger Picture: This Is Where Search Is Going
This little workflow points at something much larger. The industry is quietly converging on the idea that websites need an AI-readable layer: the emerging llms.txt convention, AI crawlers requesting clean content, and assistants deciding which brands to cite based on how clearly they can parse what you offer. Markdown conversion is the manual version of that future. Doing it now teaches you exactly how legible your brand is to the systems that increasingly answer your buyers’ questions.
And if your site converts to markdown badly, with walls of unstructured text, vague headings, and key claims buried in design elements, that is not just an AI problem. It is the same structural weakness that holds back your rankings and conversions with human readers too. The crawl just makes it visible.
Want to know how visible your brand is to AI search?
I run AI visibility audits that show where your brand appears in ChatGPT, Perplexity, and Google AI Overviews, and exactly what to restructure so assistants start citing you. This workflow is one small piece of that system.
AI SEOBrand VisibilityGEOPython for SEOChatGPTLLM Optimization
AI Visibility · GEO Field Guide
How to Check If ChatGPT Mentions Your Brand: A Free Python AI Visibility Tracker
82%of AI citations come from third-party earned media, not your own pages
44%of LLM citations are pulled from the first 30% of a page
5 minto run this tracker against your own brand
Freeno SaaS subscription, just a short Python script
Short answer: to check your brand’s AI visibility, run a fixed set of buyer-style questions through the ChatGPT (OpenAI) API on a schedule and log whether your brand gets named in the answers. It is the AI equivalent of a rank tracker. Below is the exact free Python script I use, how to read the results, and the GEO moves that actually improve them.
Google is no longer the only place people search. Millions now ask ChatGPT, Claude, Gemini, and Perplexity for recommendations before they ever open a results page. If your brand is missing from those answers, you are invisible to a fast-growing slice of your market, and unlike a Google ranking, nothing tells you by default. You have to measure it deliberately, which is exactly what this guide does.
What this tracker does
The script sends a set of targeted prompts to the OpenAI API, the same questions your potential customers might type, and logs the full responses. It checks whether your brand name appears, repeats each prompt a few times because AI answers vary, and exports everything to an Excel file you can review or chart. Think of it as a rank tracker for AI answers.
The tracker sends your buyer questions to the model, then logs whether your brand is named in the answer.
Why AI visibility matters in 2026
AI answers now shape buying decisions long before someone reaches your website. Traditional SEO measures keyword rankings; AI visibility measures something newer, whether a language model names you when it answers your customer’s question. That metric now sits alongside rankings in any serious SEO consulting engagement, because a customer who gets a confident recommendation from ChatGPT often never runs the Google search you optimised for.
The discipline has a name now: GEO, or generative engine optimization. It is the same instinct as SEO, applied to language models instead of the ten blue links. And the early research is blunt about where visibility comes from. Studies of generative engines suggest roughly 82 percent of AI citations come from third-party earned media rather than a brand’s own site, which means the question is not only what you publish, but whether the wider web treats you as an authority.
If your brand is not mentioned when an AI answers your customer’s question, you have lost that touchpoint silently, and no dashboard is going to warn you.
How AI engines actually answer, and why it changes the fix
The three engines your customers use do not source answers the same way, so the way you earn a mention differs by platform.
ChatGPT
Blends static training data with a live retrieval layer for current queries. It rewards durable authority and consistent mentions across the web, which build up over time.
Perplexity
Retrieves from the live web on almost every query, so a fresh, well-structured page can surface in its answers within hours. Speed and structure win here.
Gemini
Leans on Google’s index and query fan-out, expanding your question into sub-questions and pulling a source for each. Strong classic SEO carries over directly.
This is why a single tactic rarely fixes everything. Perplexity rewards fresh, structured content fast; ChatGPT rewards being talked about across the web; Gemini rewards the technical SEO foundations you should already own. Measuring first tells you which gap is actually yours.
Before you start: what you need
1
Python 3.9 or newer installed on your machine.
2
An OpenAI API key, created at platform.openai.com. A few dollars of credit is plenty for this.
3
Install the libraries the script uses. In your terminal, run: pip3 install openai pandas openpyxl
The complete tracker script
Create a file called tracker.py and paste the code below. Replace YOUR_API_KEY_HERE with your real key, then swap BRAND_NAME and the prompts for your own business. The example uses WizMantra, an online language-learning brand I run, so you can see real category questions in action.
from openai import OpenAI
import pandas as pd
import datetime
# Configuration
client = OpenAI(api_key="YOUR_API_KEY_HERE")
BRAND_NAME = "wizmantra" # lowercase, for case-insensitive matching
RUNS_PER_PROMPT = 3 # AI answers vary, so ask more than once
MODEL = "gpt-4o-mini" # any current chat model works
# The buyer-style questions your customers actually ask
prompts = [
"What is WizMantra?",
"Who are the best online spoken English class providers?",
"Recommend online Hindi classes for beginners.",
"Which companies offer online Sanskrit courses?",
]
# Main tracking loop
records = []
for prompt in prompts:
hits = 0
for run in range(RUNS_PER_PROMPT):
try:
resp = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
answer = resp.choices[0].message.content
mentioned = BRAND_NAME in answer.lower()
hits += 1 if mentioned else 0
records.append({
"date": datetime.datetime.now().strftime("%Y-%m-%d %H:%M"),
"prompt": prompt,
"run": run + 1,
"brand_mentioned": "Yes" if mentioned else "No",
"response": answer,
})
except Exception as e:
records.append({
"date": datetime.datetime.now().strftime("%Y-%m-%d %H:%M"),
"prompt": prompt,
"run": run + 1,
"brand_mentioned": "Error",
"response": str(e),
})
# Visibility rate for this prompt, 0 to 100 percent
print(prompt, "->", round(100 * hits / RUNS_PER_PROMPT), "percent visible")
# Export the full log to Excel
pd.DataFrame(records).to_excel("ai_visibility_log.xlsx", index=False)
print("Done. Full log saved to ai_visibility_log.xlsx")
Run it and read the results
Save the file, then run python3 tracker.py in the same folder. The terminal prints a visibility rate for each prompt, and the full log lands in ai_visibility_log.xlsx. It looks something like this:
Prompt
Response (truncated)
Brand named
What is WizMantra?
WizMantra is an online language-learning platform offering…
Yes
Best online spoken English providers?
Popular options include Preply, Cambly, italki…
No
Recommend online Hindi classes.
Several institutes offer beginner Hindi classes…
No
Companies offering online Sanskrit courses?
WizMantra is one platform that offers Sanskrit…
Yes
How to read your AI visibility score
The brand-named column is your core metric. Read it in two layers:
Direct brand prompts (“What is [brand]?”) should almost always come back Yes. If they do not, your brand has close to zero footprint in the model, a serious signal that the wider web barely references you.
Category prompts (“best providers of X”) coming back No means competitors are capturing the AI recommendation you want. That is the gap worth money.
Track the rate, not one answer. Because you run each prompt three times, you get a visibility percentage per prompt. Watch that number move month over month as you build authority.
Pro tip: the prompts you choose decide everything. Think like your customer and reuse your real keyword research. The overlap between search intent and the questions people ask an AI is your map. If you already have a keyword and intent audit, you are halfway to a good prompt set.
How to actually improve your AI visibility (the GEO playbook)
Once you know where you are invisible, the fixes are specific, and several are backed by early GEO research rather than guesswork.
44%
of AI citations come from the first 30% of a page, so front-load the clear answer
Source: GEO citation studies, 2026
+115%
visibility lift from adding source citations to a mid-ranking page
Source: Princeton GEO research
+40%
lift from adding statistics and quotations to content
Source: Princeton GEO research
82%
of AI citations come from third-party earned media
Source: GEO citation studies, 2026
Front-load the answer. Put the clear, quotable answer near the top of the page, not buried under a long intro. Nearly half of AI citations are pulled from the opening portion of a page.
Add evidence. Cite sources, add real statistics, and include short quotations. Princeton’s GEO research found these lift AI visibility by 30 to 40 percent, and adding citations alone produced a 115 percent relative lift for pages mid page one. Models prefer content they can verify.
Earn third-party mentions. If most AI citations come from earned media, then digital PR, expert commentary, directories, and reviews are not optional. A brand the web talks about is a brand the models repeat.
Fix your entity. Keep your name, category, and description consistent across your site and every profile, and mark it up with structured data so machines resolve you to one clear entity.
Build a cluster, not a lonely post. Models cite brands that cover a topic across several linked pages. This is the same topical authority logic that already wins in search, and it is why sequencing matters, technical foundation first, content second.
Keep it fresh. Language models favour the most recent version of an article that matches a query, so update and re-date your cornerstone content instead of letting it age.
Scaling this further
The script is a foundation. Once it is running, extend it:
Test across multiple models, OpenAI, Claude, and Gemini, since each has different training and retrieval sources.
Add a sentiment column, so you track not just whether you are mentioned, but how positively.
Schedule it with cron to run weekly and build a trend line automatically.
Pipe results into Google Sheets for a live team dashboard.
Frequently asked questions
What is AI visibility?
AI visibility is whether large language models like ChatGPT, Gemini, and Perplexity name your brand when they answer your customers’ questions. It is the AI-era equivalent of a keyword ranking, measured by mention rather than position.
What is GEO (generative engine optimization)?
GEO is the practice of optimising your content and brand so AI engines cite and recommend you. It reuses SEO foundations, authority, structure, and clear entities, but the target is the AI answer, not the search results page.
Does ChatGPT use the live web or only training data?
Both. ChatGPT blends static training data with a live retrieval layer for current queries. Perplexity retrieves from the live web on almost every query, and Gemini leans on Google’s index. That is why fresh, well-structured content can surface quickly in some engines and slowly in others.
How often should I run the tracker?
Monthly is a sensible baseline, plus a run after any major content or PR push. Always repeat each prompt at least three times, because AI answers vary between runs and a single response can mislead you.
Do I need to know how to code?
Only a little. If you can open a terminal, install a package, and paste a script, you can run this. If you would rather not, the same measurement can be set up for you as part of a consulting engagement.
Is AI visibility replacing SEO?
No, it complements it. The same signals that earn AI citations, authority, clean structure, consistent entities, and third-party mentions, are the signals that already drive strong search rankings. You are extending SEO, not abandoning it.
Want to know if AI is recommending you or your competitors?
I run AI visibility and GEO audits alongside classic SEO, measuring where you show up across ChatGPT, Perplexity, and Gemini, then building the authority and structure that gets you cited. Start with a conversation.
About the author: I am Ram Kr Shukla, an SEO and growth consultant helping brands grow their organic and AI-driven visibility. This script came out of my own need to track a brand’s presence across AI platforms, and I have been refining it since 2025. Want a version built for your brand? Get in touch.