Skip to content
Marketing·Aug 31, 2026·19 min·

Generative Engine Optimization for Shopify: Getting Your Store Cited in AI Search

Generative engine optimization is a new label on an old job, and Google now says so in writing. In July 2026, Google Search Central published a page that names GEO and AEO by their acronyms and calls the work SEO. If someone has quoted you a retainer for a brand new discipline, that’s the first thing to know.

The complication is that the controls moved even though the work didn’t. Three things genuinely are new: a Search Console setting that gates whether your store can appear in AI features at all, a reporting surface whose only metric is impressions, and CDN defaults that are still changing this month. Around those sit four separate switches that get blurred together constantly, and a store can get any of them wrong without touching a single page of copy.

What Is Generative Engine Optimization?

The term comes from an academic paper, not from an agency. Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande submitted “GEO: Generative Engine Optimization” on November 16, 2023, and it was accepted to KDD 2024. They defined generative engine optimization as “the first novel paradigm to aid content creators in improving their content visibility in generative engine responses through a flexible black-box optimization framework.”

That paper is also the origin of the number you’ve seen on every agency slide. The authors report that GEO can “boost visibility by up to 40% in generative engine responses,” and each qualifier in that sentence does work. It measures visibility inside GEO-bench, a research benchmark the authors built, not traffic or revenue on live commercial sites. It’s an “up to” figure, the authors state that efficacy “varies across domains,” and the evaluation ran against generative engines as they existed in 2023 and 2024, before AI Mode, before ChatGPT search in its current form.

“Answer engine optimization” has no comparable origin. There’s no paper, no vendor announcement, and no dated first use from an authoritative source, so nobody should be credited with coining it.

Is GEO SEO a Separate Discipline? Google Says No.

Google’s guide to optimizing for generative AI features, last updated July 10, 2026, addresses both acronyms directly.

'AEO' stands for 'answer engine optimization' and 'GEO' for 'generative engine optimization'. These are both terms you may see used to describe work specifically focused on improving visibility in AI search experiences. From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.

The same page states that “the best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems.” Google’s older AI features page, last updated December 10, 2025, puts the eligibility rule in mechanical terms: to appear as a supporting link in AI Overviews or AI Mode, “a page must be indexed and eligible to be shown in Google Search with a snippet,” with “no additional technical requirements.”

So the inputs are ordinary SEO. Crawlable pages, indexable URLs, content a person would choose over the alternatives, internal links that make things findable. That’s the same argument we make when a team confuses a plugin’s green lights with an actual SEO strategy: the checklist isn’t the work.

The overcorrection is to conclude that nothing changed. Something did, and none of it is content.

Answer Engine Optimization and GEO Describe the Same Work

Google documents two mechanisms behind its AI answers. The first is retrieval-augmented generation, which Google also calls grounding, and which relies “on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index.” The second is query fan-out, “a set of concurrent, related queries generated by the model to request more information.” Google’s own example: a query about fixing a weedy lawn fans out into “best herbicides for lawns,” “remove weeds without chemicals,” and “how to prevent weeds in lawn.”

Because fan-out is documented, a whole tactic grew up around it, and Google closed that door in the same guide. Creating separate content for every variation of how people might search, including fan-out queries, “primarily to manipulate rankings or generative AI responses in Google Search violates Google’s scaled content abuse spam policy.” That’s not a note about ineffectiveness. It names a spam policy. Anyone selling fan-out page generation for your collection pages is selling you penalty risk.

Google’s guide also tells you to ignore chunking your content into small blocks, rewriting content just for AI systems, and chasing mentions. On mentions specifically: “seeking inauthentic ‘mentions’ across the web isn’t as helpful as it might seem.” If your store is carrying a pile of thin variation pages from an earlier round of this advice, that’s a cleanup project, and the diagnostic order matters. We’ve written the sequence for working out why rankings dropped before you start deleting things.

AI Search Optimization: The Four Controls, and What Each One Actually Gates

Four controls decide what Google may do with your store. They live in different places, they do different jobs, and blurring them is how a store ends up excluded from AI Overviews while its owner is certain they only blocked training. Only the first of the four is new. The other three predate generative AI entirely, which is exactly why people reach for the wrong one.

Control Where it lives What it actually does
Search generative AI control Search Console, under Settings Gates whether your site can appear in AI Overviews, AI Mode, and generative AI features in Discover. Default is include. Inherits from parent properties.
Google-Extended robots.txt Limits training and grounding in some of Google’s other AI systems. Does not affect inclusion in Google Search and is not a ranking signal.
noindex Meta tag or HTTP header Removes the page from Google Search entirely, AI features included.
nosnippet, data-nosnippet, max-snippet Meta tag or inline attribute Limits what Google can show from the page. A page has to be eligible to appear with a snippet to appear in AI features at all.

The Search Console Control That Gates AI Eligibility, and It Is the New One

Google’s Search generative AI control is a property setting, and the help page carries a dated note: “As of August 31, 2026, we’ve rolled out this control to all websites worldwide.” You’ll find it under Settings, then Search generative AI. It governs inclusion in AI Overviews, AI Mode, and generative AI features in Google Discover.

There are three values. “Include my site’s links and content in Search generative AI features” is the default, except that “Inherit control from parent” is, in Google’s words, “the default if your property has a parent property.” “Exclude” prevents your content from being visible in those features, and Google is blunt about the consequence: “You won’t receive any traffic or impressions from these features.” The third value is “Inherit control from parent.”

The July 2026 optimization guide makes the dependency explicit: “In addition to the technical requirements for Search, a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search.” A store can do everything else right and still be invisible because of one radio button.

Does the inheritance rule create a Shopify problem? It can, and this is the specific trap. Google says a property “inherits its Search generative AI control from its closest parent that has changed its control to stop inheriting,” defaulting up to the top-level domain property. A storefront on shop.brand.com, or a Markets subfolder property, inherits whatever somebody set at the domain level, possibly a legal or IT decision made about the corporate site with no thought for the store. If you run multiple markets or domains through Shopify Markets, check the control on every property, not just the canonical one.

Two more scoping facts. The control doesn’t affect AI training, and Google points you to Google-Extended for that. Changes take one to two days to propagate, sometimes longer because of caching.

One inconsistency to hold in mind: Google’s December 2025 AI features page still reads as though robots.txt and the snippet directives are the only levers, and it makes no mention of this Search Console control. Treat that page as the older position, since supplemented.

Generative Engine Optimization Beyond Google: The Crawler That Cites You Isn't the Crawler That Trains On You

OpenAI and Perplexity both run the same two-crawler pattern, and once you see it the robots.txt decision stops being a single choice. Each vendor runs one crawler that indexes you for citation and obeys robots.txt, and one user-triggered fetcher that may not.

User agent Vendor Job What blocking it does
OAI-SearchBot OpenAI Surfaces sites in ChatGPT’s search features Removes you from ChatGPT search answers
GPTBot OpenAI Crawls content that may be used to train foundation models Signals no training use. Does not remove you from ChatGPT search
ChatGPT-User OpenAI Fetches a page when a user asks something in ChatGPT Not the Search control. Robots rules may not apply to user-initiated fetches
OAI-AdsBot OpenAI Validates landing pages submitted as ChatGPT ads and judges ad relevance Only visits pages submitted as ads. Not used for training
PerplexityBot Perplexity Surfaces and links sites in Perplexity results Removes you from Perplexity results
Perplexity-User Perplexity Visits a page to answer a live user question Generally ignores robots.txt, because fetches are user-triggered

Those roles come from OpenAI’s crawler documentation and Perplexity’s bots guide, both fetched August 31, 2026. OpenAI states the settings are independent: a site owner “can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot.” Perplexity describes PerplexityBot as “not used to crawl content for AI foundation models,” so blocking it costs you citations and protects nothing.

OpenAI is explicit about the cost of getting this backwards: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.” So before you paste a blocklist, read which user agents are actually in it, because “block the AI bots” describes at least two opposite outcomes depending on which names it names. OpenAI also notes it can take about 24 hours from a robots.txt update for its search systems to adjust, so nothing here is instant in either direction.

Does Shopify block AI crawlers for you? No. We pulled the default robots.txt from a live storefront on August 31, 2026 and grepped it for GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Google-Extended, CCBot, Applebot, Bytespider, Amazonbot and meta-external. Zero matches. Only seven user-agent groups appear by default, and none of them is an AI crawler, so out of the box a Shopify store allows every AI crawler that respects the file.

Your CDN Can Make the GEO Decision For You

Google lists “ensuring that crawling is allowed in robots.txt, and by any CDN or hosting infrastructure” among the fundamentals. That second clause is where stores get caught. Cloudflare ships a managed robots.txt feature that “generates and maintains a robots.txt file that instructs known AI crawlers to stay away from your content,” emitting a content signal and prepending disallow blocks for named AI user agents.

Cloudflare lets customers manage three AI-related use cases directly, and classifies further behaviors beyond them: Search, which “collects or indexes your content so it can answer questions about it later”; Agent, “automated activity acting in real time on a person’s behalf”; and Training, which crawls “to train or fine-tune a model.”

On September 15, 2026, less than two weeks from this post’s publication, Cloudflare changes its defaults. For new domains onboarding to Cloudflare, crawlers categorized as Training and Agent will be blocked by default on pages that display ads, while Search remains allowed by default. Owners can opt out in Security settings before that date. Cloudflare’s pay per crawl feature is still in closed beta and shouldn’t be planned around.

A second change lands the same day. Cloudflare says that from September 15 “multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors,” and because “the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training.” If you ticked a training-blocking box at some point, that setting is about to take Googlebot with it on any zone you control.

Cloudflare’s own caveat applies to all of it: “robots.txt compliance is voluntary. The file expresses your preferences, but it does not prevent crawlers from accessing your content at a technical level.”

If your storefront runs on Shopify’s Online Store, Shopify serves it and this isn’t your dashboard to check. If you run Hydrogen on your own infrastructure, proxy a blog subdomain, or keep the brand’s marketing site on a separate Cloudflare zone, it is. Managed robots.txt has to be switched on rather than arriving that way, so this is a setting to audit rather than a trap you fall into, and the September 15 defaults leave Search allowed. The exposure is the multi-purpose rule above, and it applies to whoever ticked the box.

noindex, nosnippet and Google-Extended Are Three Different Levers

noindex removes a page from Google Search entirely, and AI features go with it. The snippet directives, nosnippet, data-nosnippet and max-snippet, limit what Google can display from the page. Because eligibility requires a page that can be shown with a snippet, a blanket nosnippet takes you out of AI Overviews and AI Mode without taking you out of the index, which is a failure mode that looks like nothing at all in a rank tracker.

Google-Extended sits in robots.txt and limits training and grounding in some of Google’s other AI systems. It is not the AI Overviews control, and Google’s crawler documentation is explicit: “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” Four controls, four jobs, and none of them substitutes for another.

Generative Engine Optimization on Shopify: What the Platform Already Handles

Shopify's Default robots.txt Hides the Pages GEO Answers Need

The default file blocks /policies/. Shipping policy, return policy, warranty terms, the exact comparison facts an AI shopping answer wants to state, sit behind a disallow rule that ships with every store. Nothing about that is malicious, and it predates AI answers by years, but it’s live on stores right now.

You can change it. Shopify’s robots.txt customization guide covers adding a rule to a group, removing a default rule, adding custom rules for crawlers not in the default set, and adding sitemap URLs, and it publishes the removal snippet for the policies rule. Three constraints come with it. The template supports only six Liquid objects (robots, group, rule, user_agent, sitemap, request), so nothing dynamic about products or collections is available. Editing it needs theme code access. And it lives in the theme, which means a theme swap silently reverts it.

That last one belongs on your regression checklist alongside everything else that quietly breaks when a theme or platform changes. Our SEO migration checklist covers the same class of failure.

Be honest about the size of the claim here. No engine documents that unblocking /policies/ produces citations. The argument is narrower and it comes from Google’s own instruction to “ensure your content is crawlable” so that its models can use it: content a crawler can’t fetch can’t ground an answer. Shopify also notes that its default rules “are updated regularly to ensure that SEO best practices are always applied,” so a hand-written override is a thing you now own and maintain.

What Shopify's llms.txt and agents.md Actually Do

Shopify’s changelog entry of May 28, 2026 states that every store includes a default agents.md file at /agents.md, and that /llms.txt and /llms-full.txt point to the same content by default. Shopify’s template documentation confirms it manages the file “by default for every store,” serves it at the bare primary domain with no locale or Markets subfolder prefix, and resolves an optional agents.md.liquid template ahead of the managed default for all three URLs.

We checked this on four live storefronts on August 31, 2026. Allbirds, Brooklinen and Death Wish Coffee each returned HTTP 200 with Content-Type: text/markdown on all three paths, served directly with zero redirects. Gymshark returned 404 on all three. So coverage isn’t universal, and the right way to find out is to open yourdomain.com/llms.txt in a browser rather than to assume.

The file’s contents explain what it’s for. Allbirds’ live /llms.txt opens by pointing at its sibling: “Agent discovery: the canonical agent-facing description of the store is at /agents.md. You’re reading /llms.txt, which mirrors that content.” From there the live examples advertise a Universal Commerce Protocol discovery endpoint at /.well-known/ucp, a Model Context Protocol endpoint at /api/ucp/mcp, read-only browsing URLs for products and collections, the sitemap, store policies, and a rule stating that checkout requires human approval.

So should you add an llms.txt file for AI search? No, on two counts. Google’s July 2026 guide states that you don’t need AI text files “to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them,” and adds that maintaining one “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.” On Shopify the file already exists anyway, and it’s a transaction-protocol discovery document for shopping agents rather than a ranking input. Two different files, sharing a filename convention, doing unrelated jobs. The Files-upload-plus-redirect workaround for getting llms.txt onto a Shopify root is obsolete. Shopify superseded it in May 2026.

Product Schema and Feeds: The Part of GEO That Isn't About AI

Google’s position on structured data is narrower than the industry’s. From the July 2026 guide: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add. However, it’s a good idea to continue using it as part of your overall SEO strategy, as it helps with being eligible for rich results on Google Search.” The December 2025 page says the same thing more bluntly: there’s “no special schema.org structured data that you need to add.”

The one structured-data instruction Google gives in the AI context is a consistency requirement: make sure your structured data matches the visible text on the page. That’s a validation task, not a markup expansion project.

There’s a real gap in the Shopify default worth fixing regardless. Shopify’s structured_data Liquid filter works on product and article objects, outputting a schema.org Product for products with no variants and a ProductGroup for products with one or more. Its documented output covers @type, brand, category, description, image, name, offers and url. It includes no aggregateRating and no review. On a serious DTC store, reviews live in Judge.me, Okendo, Yotpo or Stamped, and those apps inject their own markup separately. A team that assumes the Liquid filter covers ratings never checks whether the review app’s markup is actually rendering, and finds out when the merchant listing loses its stars.

Google documents two product experiences: product snippets, “for product pages where people can’t directly purchase the product,” and merchant listings, “for pages where customers can purchase products from you.” A Shopify PDP is the second kind. Get the markup right for rich result eligibility and for feed accuracy, not because it buys you an AI citation, and pair it with the product page tests that actually move conversion.

Feeds are the one product input Google does tie to AI surfaces. Its guide states that “using products like Merchant Center (such as Merchant Center feeds) and Google Business Profiles can help your products and services to be visible in both AI responses and other Google Search results.” If you’re going to spend a day on structured product data, spend it on the feed.

Measuring Generative Engine Optimization Without Fooling Yourself

Google launched Generative AI performance reports in Search Console on June 3, 2026, and the post now carries a note that “as of August 31, 2026, we’ve rolled out these insights to all websites worldwide.” Impressions is the only metric. Pages, countries, devices and dates are dimensions you can slice it by. Google says it’s “continuing to work with website owners” on additional metrics over time.

Does Search Console show which prompts cited you? No. There are no clicks, no CTR, no queries, no citations and no conversions in the report. You get how often URLs from your site appeared in generative AI features, and which URLs those were. Google’s help documentation adds two constraints: the report draws from the Web search type in the Performance report, so AI-surface data is a slice of your organic data rather than an addition to it, and Search Labs experiments are excluded. That same help page still carries a line saying “not all properties have access to the report, as we’re rolling out over time,” which contradicts the dated rollout note above it. Google hasn’t reconciled the two, so check your own property rather than assuming either.

Everything downstream of the impression is dark. AI Overviews and AI Mode clicks arrive in GA4 as ordinary organic referrals and can’t be separated from classic search. Ahrefs, publishing its own analysis of AI traffic across 81,947 sites on June 26, 2025, made the same point against its own interest, noting that AI Overview and AI Mode traffic “is at least somewhat underestimated” because it “currently gets tracked as ‘Search’ traffic in site analytics tools.”

That Ahrefs study is also the best-sourced answer to how big this channel is. Across those 81,947 sites, AI traffic averaged 0.25% of a site’s total traffic, while growing roughly ten times year over year. Large multiples on a small base. Treat vendor sampling with the caution it deserves, and note that Ahrefs sells SEO software.

For click behavior, the most independent dataset is Pew Research Center’s browsing panel study, published July 22, 2025, from 900 US adults who agreed to share their browsing activity during March 2025, Google only. Users who saw an AI summary clicked a traditional result link in 8% of all visits, against 15% for those who didn’t. Clicks on a link inside the summary happened in 1% of visits to pages carrying such a summary. Pew also found 88% of summaries cited three or more sources. That data is roughly 17 months old at publication, it’s a share of visits rather than a per-impression CTR, and it isn’t ecommerce, so quote it with all four qualifiers or don’t quote it.

Crawling and citing and clicking are three separate things, and Cloudflare’s crawl-to-refer analysis, published August 29, 2025 from its own network data for January through July 2025, put numbers on the gap. In July 2025 the ratios ran 5.4 to 1 for Google, 194.8 to 1 for Perplexity, 1,091.4 to 1 for OpenAI and 38,065.7 to 1 for Anthropic. Those are vendor measurements of one network, they’re a year stale, and they moved sharply within the seven months measured, so read the shape rather than the levels.

Server log analysis is the one method fully under an operator’s control, and on Shopify it mostly isn’t available. Shopify doesn’t hand Online Store merchants raw storefront access logs, so counting OAI-SearchBot or PerplexityBot hits the way a self-hosted store would isn’t an option for most merchants. If you run Hydrogen on your own infrastructure or sit behind a proxy you control, OpenAI and Perplexity both publish IP ranges and user-agent strings precisely so you can verify crawler identity.

One last guardrail, from Google: “Be wary of third-party tools that promise ranking success or claim to use ‘internal’ Google metrics. No third-party tool has access to our internal ranking or AI systems.” Any product claiming to measure your ChatGPT visibility is sampling prompts and inferring, not reading OpenAI’s data.

What Bing Reports About Your GEO That Google Doesn't

Microsoft added AI Performance to Bing Webmaster Tools in public preview on February 10, 2026, covering Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations. It reports Total Citations, “the total number of citations that are displayed as sources in AI-generated answers,” Average Cited Pages, and grounding queries, “the key phrases the AI used when retrieving content.” A June 16, 2026 expansion added Intents, Topics, Citation Share and Compare, all still in preview.

Microsoft scopes its caveats per metric. Of grounding queries it says “The data shown represents a sample of overall citation activity.” Of Average Cited Pages it says the count “does not indicate ranking, authority, or the role of any page within an individual answer.” And of the page-level metric: “This reflects how often pages are cited, not page importance, ranking, or placement.”

Even so, that’s citation counts and query-level grounding data, which Google doesn’t provide at any price. Verify your storefront domain in Bing Webmaster Tools. It’s free, and it’s currently the only first-party window into which of your pages get cited and for what, even though Bing and Copilot are the smaller surface.

The Highest-Leverage Generative Engine Optimization Work Is Still Content

Google’s guide ends where it started, and it’s the least fashionable advice in the document: “Creating content that people find unique, compelling, and useful will likely influence your website’s presence in generative AI search in the long run more than any of the other suggestions in this guide.”

The distinction Google draws is between commodity and non-commodity content, and it gives an example rather than a definition. “7 Tips for First-Time Homebuyers” is commodity content, “based on common knowledge, which could originate from anyone.” Its counterexample is “Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line.” Google adds that AI systems “take a look at a variety of sources, so it can be helpful to have a unique viewpoint that stands out,” and singles out first-hand review over summary. It also notes that its generative features can pull in relevant images and video, “which means more opportunities for your website to appear beyond web page links.”

For a DTC brand, that translates into things a model can’t synthesize from other people’s pages. What the fabric does after 40 washes. Why you changed the closure in version three. Fit notes from real returns data. Your own photography of the product in the condition customers actually receive it. Every one of those is expensive to produce and impossible to copy, which is the entire point.

GEO Is Visibility. Agentic Commerce Is Checkout.

These two get conflated constantly, and they’re separate problems with separate work. Getting cited is upstream. Getting transacted is downstream, and it runs on protocols rather than content.

Google announced the Universal Commerce Protocol on January 11, 2026, describing it as “a new open standard for agentic commerce that works across the entire shopping journey” and naming Shopify, Etsy, Wayfair, Target and Walmart as co-developers. Google said UCP would “soon power a new checkout feature on eligible Google product listings in AI Mode in Search and the Gemini app.” Google’s July 2026 optimization guide references it too: “Protocols like Universal Commerce Protocol (UCP) are emerging that will allow Search agents to do more.”

That’s what the UCP and MCP endpoints in your store’s agents.md are pointing at, and it’s why the file is a commerce document rather than an SEO one. The versions are dated and they move, so don’t hardcode one into a runbook. If the agent-checkout side is where your questions actually are, we’ve covered agentic commerce on Shopify separately.

Where Generative Engine Optimization Pays Off

Hold two things at once. The content work is SEO, and Google’s own July 2026 documentation says so by name, which means the panic version of generative engine optimization is selling you a discipline that no search engine recognizes. And four controls now exist that no amount of good content substitutes for.

Here’s the short list for a Shopify store, in the order we’d run it:

  • Open Search Console, check Settings and confirm the Search generative AI control is set to include, on every property including subdomains and Markets properties
  • Read your live robots.txt and confirm no AI citation crawler is disallowed
  • Decide the <code>/policies/</code> rule deliberately instead of inheriting it
  • If anything of yours sits behind a CDN, check its AI crawler defaults before September 15
  • Verify the domain in Bing Webmaster Tools for citation data
  • Then go back to doing the SEO, because that's where the returns are

None of that is a new discipline. It’s four controls, one of them genuinely new, plus a platform default and a reporting surface, on top of work you already knew how to do.

If you want someone to run that pass with you, BLKDG offers generative engine optimization services built on the same principle Google documents: optimize for the search experience, and the AI surfaces follow. Or book a free Growth Audit and we’ll show you exactly where your store is invisible and why. No assumptions. No selling you work you don’t need.

You built something worth finding. We make sure it gets found.

EJ Ulery

About the author

EJ Ulery

Co-Founder & CTO

EJ Ulery on LinkedIn

Not sure where the gap is? That's exactly what the Digital Marketing Growth Audit is for.

A free, no-obligation look at where your site can win more traffic and conversions, with a clear digital marketing roadmap to get there. Just a straight read on where your digital presence stands and where it's headed.