
geo-score
Can AI engines cite your site, and do they? Free 0–100 readiness score on an open GEO rubric, plus citation tracking via the OpenAI, Perplexity, Gemini and C…
Install with your AI
Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.
Install and set up geo-score (claude-plugin project) into my current project. Found on https://claudeers.com/geo-score Repo: https://github.com/jianruntech/geo-score Homepage/docs: https://jianruntech.github.io/geo-score/ Detected install method: claude-plugin → /plugin install geo-score@jianruntech/geo-score Category: mcp-servers. Platforms: cli, api, web. Read the repo's README for exact setup and env vars, then install it and wire it into my project. Claudeers Health Verdict: unknown; community-verified: false. Confirm the source before running anything.
⚠ Unverified / not recently updated — review before pasting a run-this config.
/plugin marketplace add jianruntech/geo-score /plugin install geo-score@jianruntech/geo-score
git clone https://github.com/jianruntech/geo-score
// compatibility
| Platforms | cli, api, web |
|---|---|
| Operating systems | — |
| AI compatibility | claude |
| License | MIT |
| Pricing | open-source |
| Language | Python |
geo-score
Can AI engines cite your site? A free 0–100 score in 20 seconds. Do they? Check through their APIs with your own keys.
Score (free, no key) → Ask one question (your API keys) → Watch a question set every week (your API keys). Three levels, one tool · the citation half is never added to the score.
curl -sL https://raw.githubusercontent.com/jianruntech/geo-score/main/cli/geo_score.py \
| python3 - stripe.com --brief
One command. About twenty seconds. Every check, and what the next tier needs.
Full output — every check, its evidence, and what the next tier asks for
AIV READINESS https://stripe.com
──────────────────────────────────────────────────────────────────────────
71 / 100 Solid
12 points to Leading
Reachable 11/15
◐ Crawlers allowed in robots.txt ███████████░░░░░░░ 3/5
✓ Reachable to retrieval agents ██████████████████ 5/5
◐ Main content server-rendered ███████████░░░░░░░ 3/5
Understandable 15/22
◐ Sitemap discoverable and fresh █████████░░░░░░░░░ 2/4
✓ llms.txt present and structured ██████████████████ 5/5
✓ Organization + WebSite schema ██████████████████ 6/6
✗ BreadcrumbList on nested pages ░░░░░░░░░░░░░░░░░░ 0/3
◐ Page-type schema (Product, FAQ…) █████████░░░░░░░░░ 2/4
Content Citability 25/35
✓ Self-contained answer passages ██████████████████ 9/9
◐ Headings match how people ask ████████░░░░░░░░░░ 3/7
◐ Freshness signal present █████████░░░░░░░░░ 3/6
✓ Statistics carry a source ██████████████████ 7/7
◐ Named, verifiable authorship █████████░░░░░░░░░ 3/6
Brand Credibility 8/10
⊘ Third-party listings ·················· —
⊘ Independent mentions ·················· —
✓ Knowledge-graph entity ██████████████████ 4/4
◐ sameAs links resolve ████████████░░░░░░ 2/3
◐ Video and multimodal presence ████████████░░░░░░ 2/3
Answer Fit 2/4
◐ Content shaped for extraction █████████░░░░░░░░░ 2/4
⊘ Covers the questions people ask ·················· —
⊘ Chinese engine readiness ·················· —
Biggest gaps
+4 Headings match how people ask about half do
+3 Named, verifiable authorship and the name links to a verifiable identity page
+3 Freshness signal present most pages do, and dateModified agrees with the visible date
Scored 61 / 86 observable · 4 checks left the denominator · rubric v1.1
Needs judgement: p3.listings, p3.mentions, p4.question-coverage, p4.cn-engines
Full rubric and what each tier means:
https://github.com/jianruntech/geo-score
Python 3.8+, standard library only, nothing to install. It reads public URLs and prints a score against a published, versioned rubric — not a black box.
See how 314 well-known sites score → · a quarter of them are unreadable to AI crawlers.
GEO means Generative Engine Optimization — getting cited by ChatGPT, Perplexity, Google AI Overviews, Gemini and Copilot. Nothing to do with geography or maps.
Three levels, one tool
| Level | What it answers | Needs | Command |
|---|---|---|---|
| 1 · Score | Can AI engines reach, parse, trust and cite the site? 0–100 against the open rubric | Nothing: no key, no install | geo_score.py stripe.com |
| 2 · Ask | Right now, for a question you care about, do they cite it? | An API key for any of OpenAI, Perplexity, Gemini, Anthropic, OpenRouter | geo_score.py stripe.com --ask "best payments API for marketplaces" |
| 3 · Watch | How often are you cited, against which competitors and sources, week over week? | Your API keys and a fixed question list | geo_score.py watch run |
Level 1 is the score. Levels 2 and 3 measure the outcome the rubric keeps out of the score on purpose (two scores, never one): they are reported next to the 100 and never summed into it. A site can score 90 and still lose every answer to a competitor with more third-party coverage, and the reverse also happens, which is why you want both.
Levels 2 and 3 need cli/geo_watch.py next to cli/geo_score.py: clone the repo or download both
files. The one-line curl | python3 above runs level 1.
Why this is a different question from SEO
Classic SEO asks where do I rank. Answer engines don't rank — they retrieve passages, decide whether a source is worth quoting, and cite it. Different question, different failure modes: a site can sit at position 3 on Google and never be quoted, while a page nobody links to gets cited daily because its passages are clean.
Most of what determines this is mechanical and cheap to fix — a robots.txt line, a
JSON-LD block, a date in a template, a paragraph rewritten so it stands on its own. The
hard part is knowing which of them you are missing, and what each one is worth.
What it checks
21 tiered checks totalling 100 points, plus 4 bonus checks worth up to +6 outside the denominator. Full specification: rubric/v1.1.md · 简体中文
| Pillar | Pts | Asks |
|---|---|---|
| Reachable — gates | 15 | Can a retrieval crawler get the page at all? robots.txt, live reachability across 10 AI user-agents, server-rendered content |
| Understandable | 22 | Can it tell what the page and the company are? Organization + WebSite, llms.txt, sitemap, breadcrumbs, page-type schema |
| Content Citability | 35 | Is there anything here worth quoting? Self-contained answer passages, headings that match how people ask, sourced figures, real bylines, freshness |
| Brand Credibility | 18 | Why should an engine trust it? Knowledge-graph entity, third-party listings, sameAs that resolves, video presence |
| Answer Fit | 10 | Is the content shaped to be lifted into an answer? |
Content Citability carries the most weight on purpose: answer engines retrieve passages, not domains. Passage shape beats domain authority more often than classic SEO intuition expects.
Every scored check is tiered — 2 to 4 tiers, each naming a count out of the 8 sampled pages, so two people scoring the same site agree on the arithmetic. Three checks are gates: score zero on crawler access, live reachability or server-rendered content and the result caps at 40, because until a crawler can reach the content nothing else you change has any effect.
Bands
| 0–30 | 31–50 | 51–65 | 66–82 | 83–100 |
|---|---|---|---|---|
| Not started | Early | Growing | Solid | Leading |
Band names describe a stage, not a verdict. External benchmarks put most business sites in the 30–55 range, so a score in the forties is ordinary, not alarming.
314 sites, scored in public
A quarter of them are unreadable to AI crawlers. 80 sites have a gate check at zero — an
AI retrieval crawler cannot get the content, so it has nothing of theirs to quote. 19 block AI crawlers by name in
robots.txt, which is an editorial choice and reported as such — amazon.com lands at 12
for exactly this reason. 45 serve a page whose body only exists after JavaScript runs.
Their content is there, a browser sees it, and a crawler gets an empty shell. That group
almost certainly did not choose it. A further 16 hand a crawler an outright error.
Median 56. Range 12 to 98.
| Site | Score | Band |
|---|---|---|
| pulumi.com | 98 | Leading |
| minimaxi.com | 95 | Leading |
| lumalabs.ai | 93 | Leading |
| resend.com | 93 | Leading |
| ironcladapp.com | 91 | Leading |
| … | ||
| amazon.com | 12 | Not started |
| keepa.com | 12 | Not started |
| mercadolibre.com | 12 | Not started |
The full table, by sector → · markdown · raw data · re-run it
Two more findings worth the click: sites built for the Chinese market score 19 points lower than everyone else (median 40 against 59 — a gap that has held between 16 and 23 points across five separate samples, from 38 sites up to 314), and the same three cheap things — a date in the page template, an opening paragraph that stands on its own, one JSON-LD block — are missing from more than half the field.
Every number here is reproducible with the command at the top of this page (the benchmark
was run with 1.1.0 on 2026-09-09; since 1.1.2 a site can read a few points higher where
sameAs or knowledge-graph lookups were unreachable) — and we measured how reproducible. Running the whole benchmark twice and comparing every site:
96% land within ±5, 45% land identically. Read one site's score as ±5 rather than as
exact; medians are stable. The unstable part is the gate checks, where five sites flipped
between runs because their bot protection answered a crawler differently.
The band, the control experiment and the per-site pairs are in
benchmark/REPRODUCIBILITY.md.
For five reference sites we also publish hand-scored audits covering all 21 checks, with the evidence behind each one: examples/audits/v1.1/.
Ways to run it
CLI — level 1 is one file with no dependencies, 20 seconds. Levels 2 and 3 add
cli/geo_watch.py from the same release.
python3 cli/geo_score.py example.com # human-readable
python3 cli/geo_score.py example.com --explain # with the evidence behind every check
python3 cli/geo_score.py example.com --json # conforms to schema/report.v2.json
python3 cli/geo_score.py example.com --compare competitor.com # side by side
python3 cli/geo_score.py example.com --badge aiv-badge.svg # embeddable SVG
python3 cli/geo_score.py example.com --share # one line to paste somewhere
GitHub Action (level 1) — score on every push, fail the build when it regresses. For level 3 on a schedule, see examples/ci/watch-weekly.yml.
- uses: jianruntech/geo-score@v1
with:
url: https://example.com
fail-under: 40
Claude Code skill — the CLI measures what a static fetch can see. Four checks need off-site search or human judgement, and the skill does those too.
git clone https://github.com/jianruntech/geo-score ~/.claude/skills/geo-score
# then: /geo-score audit https://example.com
The CLI leaves those four checks out of the denominator rather than guessing, so it reads a little lower than a full audit — typically by 5 to 15 points on an established brand, which has listings and mentions the CLI cannot see.
MCP server (all three levels) — python3 cli/geo_score.py mcp, for Claude Code, Cursor and
other agents; see below.
Levels 2 and 3 · does AI actually cite you?
--ask and watch put the questions your buyers ask to ChatGPT, Perplexity, Gemini and Claude,
through each provider's search-enabled API with your own keys, and record who the answers cite:
you, your competitors, or the third-party pages (forums, review sites) the engines lean on
instead. Keys are read from the environment; geo-score never writes them anywhere.
git clone https://github.com/jianruntech/geo-score && cd geo-score/cli
export OPENAI_API_KEY=… PERPLEXITY_API_KEY=… # any subset of engines works
python3 geo_score.py acme.com --ask "best invoicing app for freelancers" # level 2
python3 geo_score.py watch init --brand Acme --domain acme.com --competitor "Rival=rival.com"
# put the questions your buyers ask an AI assistant into queries.csv, then:
python3 geo_score.py watch run --dry-run # the plan and the caps; no calls, no cost
python3 geo_score.py watch run # level 3: every question on every engine, saved
python3 geo_score.py watch diff # this run against the last, with a significance test
What a watch run prints (illustrative: made-up brands and canned answers, not a real measurement)
geo-score watch · Lumo · run 20260926T083000Z · api channel
3 questions × 3 engines × 1 = 9 planned · 9 answered · 0 failed · 0 skipped
Cited in 67% of answers (6 of 9, 95% CI 35–88%) · mentioned in 67%
Your site was cited somewhere for 3 of 3 questions.
By engine
engine model cited 95% CI mentioned avg rank
chatgpt-api gpt-6-luna 3/3 100% 44–100% 100% 1.0
perplexity-api sonar 0/3 0% 0–56% 0% —
gemini-api gemini-3.8-flash 3/3 100% 44–100% 100% 1.0
claude-api claude-sonnet-5 not measured no_key: set ANTHROPIC_API_KEY
Share of voice
cited 95% CI mentioned
Lumo (you) 67% 35–88% 67%
Pixa 67% 35–88% 67%
Reelcraft 33% 12–65% 33%
Sources the engines cite most (not yours)
domain answers share owner
pixa.example 6 67% Pixa
reddit.com 6 67%
reelcraft.example 3 33% Reelcraft
g2.com 3 33%
By question type
type cited 95% CI mentioned
alternative 2/3 67% 21–94% 67%
list 2/3 67% 21–94% 67%
pricing 2/3 67% 21–94% 67%
What the engines searched for (from 6 answers that show it)
times search
3 免费 ai 视频工具
1 best free ai video generator 2026
1 哪个 ai 视频生成器有免费额度 2026
1 pixa alternatives 2026
Ledger
9 questions asked · 9 API requests · 3,180 in / 2,400 out tokens · 9 searches · $0.05 + 3 answers with no price (add prices to geo-score-watch.json)
Caps: at most 120 questions.
Read this before quoting the numbers
- API channel. Answers come from each provider's search-enabled API, which is not the consumer app. Compare runs with runs; never pool them with answers sampled by hand in the apps.
- The same question gets different answers from one ask to the next. Rates carry a 95% Wilson interval, and diff only calls a change a change when a two-proportion test says so (p < 0.05).
- Engines without a key, calls that failed and calls skipped by a cap count as not measured, never as zero.
- This measures citations. It does not predict traffic, rankings or revenue.
Level 2 asks each engine each question once and prints the result under the readiness
report (and into the report's citation object with --json). One ask is an anecdote: use it to
see what the engines say today, not to measure a rate.
Level 3 keeps a fixed question list, saves every answer under .geo-score/watch/runs/
(schema), and reports:
| Meaning | |
|---|---|
| cited | The answer links to a URL you own: one of your domains (subdomains included), or a url_prefixes entry such as your Amazon store or GitHub org |
| rank | Your position among the distinct domains the answer cites. Rank 1 means you were the first source |
| mentioned | The answer names you, one of your aliases or your domain. Chinese, Japanese and Korean names match anywhere; all other names match whole words only |
| share of voice | The same two rates for each competitor, over the same answers |
| sources | The third-party domains cited most often: the pages the engines trust in your category |
| gaps | Questions where a competitor is cited and you are not, in any answer |
| searches | The searches the engine actually ran before answering, where the API exposes them (OpenAI, Gemini, Claude). This is the query fan-out, observed rather than guessed |
| ledger | Questions asked, API requests, tokens, searches and cost. Every run keeps its own ledger |
What it will not tell you.
- It is not the ChatGPT app. Answers come from each provider's API with web search switched on. The consumer apps use other models, prompts and personalisation. Compare API runs with API runs; never pool them with answers sampled by hand in the apps.
- One answer is an anecdote. Every rate carries a 95% Wilson interval, and
diffcalls something a change only when a two-proportion test says so (p < 0.05). - Not measured is not zero. An engine without a key, a failed call and a capped call are all reported as not measured, and none of them lowers your rate.
- Citations are not traffic. Nothing here predicts visits, rankings or revenue.
Engines. Pin the model you mean in the config and keep it fixed between runs; diff flags a
run where the model changed. OpenRouter covers hundreds of models with one key.
| Engine id (default) | Provider | Key | Default model | What counts as cited | Searches shown | Cost reported |
|---|---|---|---|---|---|---|
chatgpt-api | OpenAI Responses API + web_search | OPENAI_API_KEY | gpt-6-luna | url_citation annotations | yes | no, set prices |
perplexity-api | Perplexity Sonar | PERPLEXITY_API_KEY | sonar | numbered sources the answer uses | no | yes |
gemini-api | Gemini Interactions API + google_search | GEMINI_API_KEY or GOOGLE_API_KEY | gemini-3.8-flash | url_citation annotations | yes | no, set prices |
claude-api | Anthropic Messages + web_search tool | ANTHROPIC_API_KEY | claude-sonnet-5 | citations on the answer text | yes | no, set prices |
| any id you choose | OpenRouter + web plugin | OPENROUTER_API_KEY | openai/gpt-6-luna | url_citation annotations | no | yes |
Caps and cost. max_calls is an exact cap on questions asked in a run, and a plan that
exceeds it refuses to start. budget_usd is checked before every call against the spend so far
plus the most expensive call seen on that engine, so it can be exceeded by at most one call per
engine; with a budget set, engines whose cost cannot be known are left out unless you pass
--allow-unpriced. Every run keeps a ledger of requests, tokens, searches
and cost. Configuration, prices and all commands: cli/README.md.
From an agent. geo_score.py mcp is a stdio MCP server with no dependencies. It serves
score_site (level 1, free) and ask, run, list_runs, report, diff (levels 2 and 3). An
agent can only tighten your caps, never loosen them.
claude mcp add geo-score -- python3 /abs/path/geo-score/cli/geo_score.py mcp -c /abs/path/geo-score-watch.json
Every week. Citation rates move slowly and noisily: run on the same weekday with the same
questions and models, and read diff, not single runs.
examples/ci/watch-weekly.yml does it on a schedule with keys from
repository secrets and commits each run, so the history lives in git. Run files contain your
questions and the full answers: use a private repository if they are confidential.
Run it yourself, or have it run for you. Everything here is MIT; the tool has no paid edition. You pay your model providers directly and the ledger shows what each run used. What needs people rather than an API, Jianrun does as a service:
| Run it yourself (free) | Run by Jianrun | |
|---|---|---|
| Channel | Provider APIs with search | APIs and the consumer apps, sampled by hand each week |
| Engines | OpenAI, Perplexity, Gemini, Anthropic, anything on OpenRouter | ChatGPT with search, Perplexity, Gemini, Google AI Overviews, Copilot; Chinese engines when you sell into China |
| Report | Text, Markdown, CSV, JSON | A weekly report with the AIV dashboard: citation trend, per-engine rates, facts AI gets wrong about you, readiness history |
| When citations drop | Out of scope | We do the fixing |
| Price | Your API bill | AEO delivery system, from US$5,780 per 3 months, AIV dashboard included. See pricing |
Why a rubric, not just a tool
A score you cannot audit is a number someone made up. So the specification is the product, and the tools are implementations of it:
- Versioned. Every score reports the rubric version.
71 (v1.1)is a claim;71is not. - Tiered, with counts. Each tier names a page count out of 8, not "most".
- Evidence-bound. Every check requires an observation someone else can reproduce.
- Calibrated against public benchmarks, with the record published — including the four external sources the thresholds were checked against, and the eight specification ambiguities that real audits surfaced and v1.1 settled.
- Machine-readable.
rubric/v1.1.jsonwith stable check ids, andschema/report.v2.jsonso results from different implementations are comparable.
Implement it in your own stack, disagree with a weight, open a rubric proposal. That is the main thing we want contributions on.
Scope — what this does not do
This is the part most tools leave out, so it's stated plainly.
AIV Score measures. It does not fix.
| Not included | Why |
|---|---|
Fix templates — robots.txt, JSON-LD blocks, llms.txt boilerplate | Remediation is where the actual work and judgement live. It is a separate, non-open project |
| Content rewriting — how to shape a passage so it gets quoted | Same |
| Per-engine tactics — what to do differently for Perplexity vs Gemini | Same |
| A remediation roadmap | Same |
Other honest limits:
- It measures input-side readiness, not outcomes. A high readiness score means engines can cite you. Whether they do depends on competition, query intent and factors no external audit can observe. Citation performance is reported as a separate, unscored block and never folded into the 100 — see Two scores. Measure it with levels 2 and 3 above.
- Brand Credibility and the named-author check need human judgement. "Is this a real identifiable person" and "is this mention independent" are not fully automatable. Treat those ~24 points as assisted, not automatic.
- Tiers reduce disagreement, they do not remove it. Every tier names a count out of the 8 sampled pages, so two auditors agree on the arithmetic. They can still disagree on whether a given paragraph is a self-contained answer. The settled ambiguities are the ones we found; there will be more.
- Heavily client-rendered sites score low, sometimes unfairly. If your content only appears after hydration, most checks will read the pre-hydration HTML — which is also roughly what a crawler sees, so the low score is usually right, but verify by hand.
- Engine behaviour moves. The rubric is versioned for exactly this reason. A score from an older rubric version is not comparable to a current one.
Research behind the weights
The weights are opinionated but not invented. The two findings that most shaped them:
- Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024 — citing sources, adding statistics and quoting experts raise visibility by up to 40% (measured as Position-Adjusted Word Count, not citation count). Notably, the paper found an authoritative tone produced no significant improvement — which is why this rubric scores structure and attribution, not voice.
- llms.txt proposal, Answer.AI — the convention this rubric
checks for in the Understandable pillar (
p1.llms-txt).
Where a check rests on our own field observation rather than published research, the rubric says so. If you have evidence that a weight is wrong, open a rubric proposal — that is the main thing we want contributions on.
Related tools
Deliberately naming what this is not, so you can pick correctly:
| Project | What it does | Relationship |
|---|---|---|
| llms-txt | The llms.txt specification itself | AIV checks for compliance with it |
| yao-geo-skills | 21 categorized GEO skills, execution-oriented | Complementary — they do production, this does measurement |
| GEOFlow | Full GEO operations system for company sites | Much larger scope; AGPL |
If you need remediation and not just a score, those projects overlap with the part this repo deliberately excludes.
Who maintains this
Built and maintained by Jianrun Tech (见润科技), Shenzhen — we run GEO and AI-adoption programs for cross-border commerce companies. The rubric came out of client work and out of optimizing our own products; publishing it is how we'd like AI visibility to be measured consistently, including by people who never become our clients.
Commercial use of this repository is unrestricted under MIT — including inside paid consulting work. You do not need our permission, and there is no separate commercial licence.
Contributing
The most valuable contribution is evidence about the weights. See CONTRIBUTING.md.
Citation
If you reference the rubric in research or a report, see CITATION.cff.
License
// faq
What is geo-score?
Can AI engines cite your site, and do they? Free 0–100 readiness score on an open GEO rubric, plus citation tracking via the OpenAI, Perplexity, Gemini and Claude APIs with your own keys. Zero dependencies, MCP server.. It is open-source on GitHub.
Is geo-score free to use?
geo-score is open-source under the MIT license, so it is free to use.
What category does geo-score belong to?
geo-score is listed under mcp-servers in the Claudeers registry of Claude-compatible tools.
// embed badge
[](https://claudeers.com/geo-score)
// retro hit counter
[](https://claudeers.com/geo-score)
// reviews
// guestbook
// related in MCP Servers
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete…
A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Gemini CLI & Hermes Agent. Only official website: ccswitch.io
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
An open-source AI agent that brings the power of Gemini directly into your terminal.