GEO

GEO Readiness Audit

Generative Engine Optimization is about access and attribution: whether AI crawlers are allowed in, and whether there is a named, dated, sourced entity worth crediting. Paste a URL and we score the seven checks that decide it.

Published · Last updated

Quick answer

The GEO Readiness Audit fetches any URL and scores only the generative-engine checks: whether the content exists in server-rendered HTML, whether robots.txt permits AI crawlers, whether llms.txt exists, whether schema identifies the entity, and whether a named author, honest dates and cited sources give a model something concrete to credit. It returns a score out of 100 with evidence and ranked fixes.

We fetch the page once with no JavaScript and also read /robots.txt and /llms.txt on the same host, then score the crawler-access and attribution checks only. Nothing is stored.

GEO readiness

Enter a URL and run the audit. You will get a GEO score out of 100, pillar subscores, and a prioritised fix list with copy-paste markup for anything that fails.

New to the terms? Read what AEO is, what GEO is or how they differ from SEO.

What is the GEO Readiness Audit?

What it answersScore any URL out of 100 on generative-engine access and attribution — AI crawlers, llms.txt, authorship.
How the answer is producedGenerative Engine Optimization asks two questions an extraction audit cannot answer: is the crawler allowed in, and once it has your text, is there anything here worth naming as the source?
What you need to enterPaste the URL you want cited by ChatGPT, Perplexity or Gemini, and run the scan.
Where it stops being reliableIt audits one URL and cannot reach pages behind a login or a bot wall.
Cost and sign-upFree, runs in your browser, no account and no stored inputs.

How is the GEO score calculated?

Generative Engine Optimization asks two questions an extraction audit cannot answer: is the crawler allowed in, and once it has your text, is there anything here worth naming as the source? A page can be beautifully formatted for extraction and still be invisible because robots.txt quietly disallows GPTBot, or be paraphrased without credit because nothing on it identifies a publisher or a human being.

Seven checks carry the weight. Server rendering appears again because it gates everything. Crawler access is next and is the single most common hard failure — often added years ago by someone blocking scrapers, with no one aware AI agents were caught in the rule. Then llms.txt, entity schema, a named author, honest published and updated dates, and outbound citations to sources you actually used.

Each check is scored from what was found in the response: the raw HTML of the page, plus /robots.txt and /llms.txt fetched from the same host. Weights reflect impact — access failures cost far more than a missing llms.txt, which is cheap to add and carries a low weight precisely because adoption is still early.

The fix list is ranked by residual weight. For most sites the top item is either a robots rule or a missing human author, and both are afternoon-scale changes with durable benefit.

How do you use the GEO Readiness Audit?

  1. 1.Paste the URL you want cited by ChatGPT, Perplexity or Gemini, and run the scan.
  2. 2.Check the crawler-access verdict first. If robots.txt blocks AI agents, nothing else on the page matters until that rule changes.
  3. 3.Read the attribution checks — author, dates, schema, citations. These decide whose version of a contested claim gets repeated by name.
  4. 4.Fix from the top, re-scan, and run the AEO audit separately to confirm there is a quotable passage for the crawler to take.

What can this tool not tell you?

  • It audits one URL and cannot reach pages behind a login or a bot wall.
  • It cannot confirm whether a model has already ingested or cited you; referral logs and direct prompting are the only evidence for that.
  • llms.txt has no public commitment from any major engine yet, which is why it is weighted lightly rather than presented as essential.
  • Access rules and crawler names change; re-check after any infrastructure or CDN change.

What should you know about access failures are silent, and usually accidental?

The most expensive GEO failure is a single line in a file nobody has opened in three years. A disallow rule written to stop content scrapers now catches the agents that decide which sites get named in AI answers, and there is no error, no warning and no report anywhere that tells you it is happening.

Attribution failures are quieter still. Plenty of strong pages carry no byline, or a company name where a person should be, and models weigh resolvable human authorship heavily when choosing whose wording to repeat. The page is read, the claim is used, and the credit goes to whoever looked like a source.

Dates are the third silent leak. Pages that were genuinely updated show stale published dates, or worse, show automated updated dates that change weekly without the content changing at all. Both undermine the freshness signal, and the dishonest version is worse than the missing one.

None of this is difficult work. It is a robots review, an author page, real dates and a handful of outbound links — but it is invisible work, which is why it stays undone on sites that have otherwise invested heavily in content.

What do worked examples look like?

A publisher scoring 31

Excellent content, question-form headings, clean rendering — and a robots.txt disallowing GPTBot and CCBot added during a scraping incident two years earlier. Removing two lines took the same page to 79 with no other change.

An agency blog scoring 58

Crawlers were welcome, but every post was authored by the company with no person, no bio and no Article schema. Adding named authors with profile pages and matching schema recovered twenty-one points and changed how the posts were referenced in AI answers within a month.

What do people ask most about this tool?

How is this different from the AEO audit?

GEO covers access and attribution — can a generative crawler fetch this, and is there a nameable source to credit. AEO covers extraction — is there a clean passage to lift. A page can pass one and fail the other badly, so they are scored apart.

Should I allow AI crawlers at all?

It is a business decision, not a technical one. Blocking them protects content from being reused without payment and guarantees you are never cited. If citations and referral traffic are the goal, the block has to go.

Which crawlers does this check for?

The common named agents including GPTBot, PerplexityBot, ClaudeBot, Google-Extended and CCBot, plus any blanket rule that would catch them.

Does llms.txt actually do anything yet?

No engine has publicly committed to it. It is generated from the same data as your sitemap, costs nothing if ignored, and carries a low weight here for exactly that reason.

Why does authorship matter so much?

When two pages make the same claim, a resolvable human with a bio page and matching schema is a stronger thing to credit than a company name. Adding real named authors is one of the few advantages a competitor cannot copy in a week.

Do outbound citations hurt my rankings?

No. Linking to the sources you actually used is a trust signal for both search and generative systems, and it is one of the cheapest checks on this list to pass honestly.

Which related tools should you try next?

Written and reviewed by Jim Vernon, Editor, AI Intelligence International. Published by AI Answer Engine, a service of AI Intelligence International, and checked against our editorial standards.