GEO & AI Search
GEO and AIO: Get Recommended, Not Just Ranked
What Generative Engine Optimisation and AI Overview Optimisation actually involve for an e-commerce brand, and which work genuinely moves the needle.
· Danny Khow, CEO, Bridzia Sdn Bhd
Generative Engine Optimisation (GEO) is the practice of making a brand's information structured, quotable and independently corroborated enough that AI systems such as ChatGPT, Gemini, Perplexity and Claude include it when they generate an answer. AI Overview Optimisation (AIO) is the same discipline aimed specifically at the AI-generated summary that now sits above the traditional blue links on a search results page, most visibly Google AI Overviews.
Both are being sold hard at the moment, usually as a new channel with new tricks. The honest version is duller. Most of what earns a brand a mention inside an AI answer is work a competent search team would already recognise, done to a higher standard of precision, plus one genuinely new obligation: being verifiable somewhere other than your own website.
How a generative engine actually assembles an answer
Nothing is recalled from memory in the way people assume. When a buyer asks an assistant which brand of something to choose, the system typically does four things in sequence.
- Interprets the prompt and expands it into several narrower queries, including ones the buyer never typed.
- Retrieves candidate documents for each of those queries, usually from a conventional web index, sometimes supplemented by what the model absorbed during training.
- Extracts the passages within those documents that appear to answer each sub-query.
- Synthesises a single answer from the surviving passages and, depending on the interface, cites some of the sources.
Three consequences follow, and they set the whole agenda.
The unit of visibility is the passage, not the page. A page that ranks well but buries its answer inside a paragraph of positioning language gives the model nothing clean to lift.
You cannot be selected if you were never retrieved, which makes retrieval a search problem before it is anything else.
You are being compared. The model is weighing your passage against other candidates answering the same sub-query, so a vague superlative loses to a specific, checkable statement every time.
Traditional SEO is the foundation, not the previous era
The retrieval step above runs on a conventional web index. If a page is not crawlable, not indexable, not internally linked and not fast enough to be crawled deeply, it is not a candidate for anything downstream. Nothing about GEO replaces that.
Rendering is where e-commerce sites quietly fail. Many crawlers that feed generative engines fetch the raw HTML and do not execute JavaScript. If your product copy, specifications or FAQ answers only exist after client-side hydration, then as far as those crawlers are concerned the page is close to empty. Server-rendered or statically prerendered HTML is not a performance nicety in that context. It is the difference between having content and not having it.
On a large catalogue this is a platform problem as much as a content problem: faceted navigation spawning near-duplicate URLs, category pages too slow to crawl at depth, specifications trapped in tabs that load on click, thin variant pages competing with each other. That work belongs to whoever owns the platform. On Adobe Commerce (Magento) it usually means indexing strategy, cache configuration and template-level markup rather than anything a content team can reach.
Structured data tells a machine what the page means
Schema.org markup, expressed as JSON-LD, is an explicit statement of the facts a page contains, written for machines rather than readers. It does not make you rank. What it does is remove ambiguity, so an engine does not have to infer your page's meaning from prose and guess wrong.
| Schema type | What it settles | Where it belongs |
|---|---|---|
Organization | Who the brand is, where it operates, how to reach it | Sitewide |
Product and Offer | Product name, variant, availability, currency | Product detail pages |
BreadcrumbList | Where a page sits in the hierarchy | Every nested page |
FAQPage | Question and answer pairs, already in answer form | Service, product and FAQ pages |
Article with Person | Who wrote something, when, and who they are | Editorial content |
LocalBusiness | Physical presence, address, opening hours | Store and contact pages |
Two rules matter more than the list itself. Marked-up facts must match the visible content exactly, because a contradiction is worse than an omission. And the entity must stay consistent: one legal name, one address format, one spelling of each product line, everywhere. Every variation invites a model to treat you as two weak entities instead of one strong one.
llms.txt and AI crawler access
llms.txt is a proposed convention: a markdown file at the root of a domain listing the pages worth reading, each with a short description, so a model gets a curated map instead of inferring one from navigation. It is cheap to publish, it does no harm, and it is not the mechanism anyone should be selling you. Support across the major engines is unsettled, and a tidy map of a site that cannot be crawled properly changes nothing.
The genuinely load-bearing file is the older, duller one: robots.txt. Access for AI crawlers is a commercial decision, and it is often made by accident, by a security team blocking unfamiliar user agents.
| Crawler | Operator | What blocking it affects |
|---|---|---|
| GPTBot | OpenAI | Whether OpenAI may crawl your content at all |
| ClaudeBot | Anthropic | Whether your content is available to Claude |
| PerplexityBot | Perplexity | Inclusion in Perplexity's index and its citations |
| Google-Extended | Gemini and AI grounding use only. It does not affect Google Search or AI Overviews, which use the regular index | |
| CCBot | Common Crawl | Inclusion in an open crawl that many datasets and tools draw from |
Verify what your server actually returns to these agents, not what robots.txt claims. A firewall rule can refuse a crawler your policy welcomes, and nobody notices until someone reads the logs.
Content shaped as an answer
What makes a passage extractable
If the passage is the unit of selection, then the job is to write passages that survive being lifted out of context. In practice that means headings phrased as the questions buyers actually ask, an answer in the first sentence beneath the heading rather than after a run-up, and sentences that stand alone without depending on the paragraph above them. Name the subject explicitly instead of leaning on "it" and "this". Define the term on the page where you use it. Give the specific figure, material, region or timeframe rather than gesturing at one.
Extractability is also why a claim like "market-leading" is worse than useless. It cannot be corroborated, so a model reaching for a defensible sentence reaches past it.
The boring content wins
On an e-commerce catalogue the highest-value content is almost always the tedious content: sizing and fit, delivery timelines by region, returns conditions, materials and ingredients, compatibility, warranty and care. Those are precisely the questions people now put to an assistant, because finding them on a website is irritating. A well-maintained FAQ carrying real answers, marked up properly, will do more for AI visibility than a rewritten homepage.
Comparison content is worth the discomfort it causes internally. Buyers ask assistants to compare, and an engine assembling a comparison uses whichever source has laid the facts out most clearly. If you will not publish an honest account of where your product fits and where it does not, a competitor's account, or a forum thread, becomes the source instead.
Your own site cannot be the only source
A model that sees a claim on exactly one domain, the domain that benefits from it, has one unverified source. Corroboration is what turns a marketing claim into a fact a system will repeat: trade and business press, marketplace and retailer listings, review platforms, industry bodies and directories, supplier and partner sites, and the community forums where buyers actually talk to each other.
This is the part of GEO that most resembles public relations, and it is the reason the discipline cannot be completed in a sprint. It is earned slowly, it cannot be bought at volume without becoming obvious, and it depends on the same entity consistency described above: a brand listed under one legal name in one directory and a variant of it in another has split its own evidence in half.
How to measure any of this without fooling yourself
| Signal | What it tells you | What it cannot tell you |
|---|---|---|
| Prompt-set testing | Whether you appear, where, and how favourably across a fixed panel of buyer prompts | Anything reliable from one run, since answers vary by run, region and account |
| Referrals from AI interfaces in analytics | That an answer produced a click | Total exposure, because most answers never produce a click |
| AI crawler hits in server logs | That a named crawler fetched a specific URL | Whether the content was used in an answer |
| Search Console | The health of the retrieval foundation | Anything happening inside a generative engine |
| Branded search and direct traffic | Lagging evidence that demand is being created upstream | Attribution to any single channel |
Prompt-set testing is the closest thing to a real measure, and it works only if it is disciplined. Fix a panel of prompts matching how your buyers actually ask, run them on a schedule, and score each result for presence, position, sentiment and whether the mention carries a link. These systems are non-deterministic and personalised, so one favourable answer proves nothing. The trend across a fixed panel over months is the signal, and a single screenshot of a flattering ChatGPT response is an anecdote, whoever is showing it to you.
What nobody can promise you
No agency controls these models. Nobody can guarantee that a given assistant will recommend a given brand for a given prompt, because that selection happens inside systems that are opaque, updated without notice, and personalised to the person asking. An answer that names you this month can stop naming you after a model update that had nothing to do with you, your site or your competitors.
What can be committed to is the input side: content that is crawlable and actually rendered, structured data that matches reality, answers written to be extracted, corroboration built deliberately over time, and measurement honest enough to show when something is not working. Anyone offering a guaranteed AI recommendation is selling a result they do not control.
Where to start
- Check what AI crawlers receive from your server, including anything a firewall or bot-mitigation layer is rejecting on your behalf.
- Confirm your key commercial content exists in the HTML without JavaScript execution.
- Fix structured data so it is accurate, complete and consistent with what a visitor sees.
- Rewrite your highest-intent pages so each section answers one question in its first sentence.
- Build the FAQ and comparison content nobody enjoys writing.
- Establish a prompt panel and a baseline before you change anything else, so you can tell later whether it worked.
Bridzia has spent 14 years on Adobe Commerce (Magento), and more than 20 years in e-commerce and digital overall, across 100+ projects for retail and FMCG brands in Malaysia and across Asia, including Guardian Malaysia, MR.DIY, Kinokuniya and Sunway Education. GEO/AIO optimisation is one of 8 services we run end to end, and it sits close to the platform work, because most of what stops a generative engine reading a store properly lives in the platform rather than the copy.
If you want a clearer view of where your own site stands, the GEO/AIO optimisation page sets out how an engagement runs, and you can contact us to talk it through.
