Back to home@ssheleg

seo-aeo-audit

Evidence-first website audit for search AND answer engines. Ten tracks from crawl access to AI citation mechanics; every finding carries an observation, every recommendation an evidence tier. Output is a prioritized change plan plus a link-building brief with keyword CSV — not a score. Loads in DeepSeek Harness (dsh).

Stars
1
Language
Python
Created
Jul 27, 2026
Updated
Aug 25, 2026

Introduction

seo-aeo-audit

validate npm license site

Docs, and every skill → · this skill's page · follow @sshlg93 on X

Loads in DeepSeek Harness (dsh) with no plugin to write: it reads the Agent Skills standard directly, scanning ~/.agents/skills — where npx skills add puts this pack — at rank 500.

An agent skill that audits a website for search and answer-engine visibility, proves every finding with evidence, and hands back a prioritized plan of changes. Part of the sshlg-skills family.

Most SEO audits hand over a checklist and a health score. This one produces a diagnosis: what is wrong, where you can see it, why it happens, what to change, what it is worth, when to expect movement, and how you will know it worked — across Google, Yandex and Bing, and across AI Overviews / AI Mode, ChatGPT, Perplexity, Claude, Copilot, Gemini and Yandex Alice AI.


Install

Claude Code plugin (recommended for Claude Code):

/plugin marketplace add ssheleg/seo-aeo-audit
/plugin install seo-aeo-audit@seo-aeo-audit

Any other agent (Cursor, Codex, OpenCode, Zed, Windsurf, Gemini CLI, 70+):

npx skills add ssheleg/seo-aeo-audit

npx installer (no clone; installs into ~/.claude):

npx github:ssheleg/seo-aeo-audit

Update

Update the whole family — one package, every agent. A bundle with one member current and the rest stale is a combination nobody tested:

npx sshlg-skills update               # installed but behind — updates everything
npx sshlg-skills install              # nothing installed yet
npx --yes sshlg-skills@latest list    # what the current release of each member is

Restart your agent afterwards: skills and hooks load at session start, so the session that updates is not the session that gets the new ones.

Updating this one member only
claude plugin marketplace update seo-aeo-audit && claude plugin update seo-aeo-audit@seo-aeo-audit
npx skills update seo-aeo-audit --global --yes

Keep one channel per agent — do not leave a plain ~/.claude/skills/ copy next to the Claude Code plugin, or the stale copy shadows the fresh one. Restart Claude Code after updating.

Use

/seo-aeo-audit example.com
/seo-aeo-audit traffic dropped in May, /blog only
/seo-aeo-audit why doesn't ChatGPT cite us for "invoice reconciliation"

The skill detects the mode (first audit, re-audit, single-question diagnosis), states which inputs it has and which it is missing, runs the tracks in scope, and ends with exactly one recommended next action. Deliverables land in docs/seo/audit-<date>.md and docs/seo/plan-<date>.md.


The audit flow

#TrackThe question it answers
AAccess & indexation economicsCan bots fetch, render and afford to index this? Where is crawl budget burned?
BCanonicalization & duplicationWhich URL is the one true URL — and does the engine agree?
CArchitecture & link equityDo the money pages get authority, depth and crawl frequency?
DIntent & SERP fitDoes each page match what the SERP rewards? Do pages fight each other?
EContent valueIs there a reason to rank this that an AI cannot replicate?
FExtractability & AEO/GEOCan an answer engine retrieve, read and quote the answer?
GEntity & brand consensusDo the models know what this brand is — and name it?
HExperience, conversion & attributionDo users finish the task, does it convert, and is the conversion measured?
IRisk & threatsPenalties, hijacks, prompt injection, takedown abuse, adversaries.
JMeasurementWill anyone be able to tell whether the plan worked?

Each track has two halves: the diagnosis (what is wrong and why) and a mechanical sweep for completeness. Findings are triaged on four axes and no scalar — impact, irreversibility, uncertainty (the evidence tier), coordination — where the first axis that separates two findings decides, and grouped into Blockers → Leaks → Gains → Experiments. effort is recorded and never ranks: a product of the inputs cannot be argued with on them, and this README already refuses a score out of 100.

Rules the skill will not break

  • Evidence or silence. Every finding names the observation, its location, the value and the date. No finding that was not verified on the site being audited.
  • Tiered claims. Every recommendation carries CONFIRMED · STUDY · FIELD · HYPOTHESIS, and a hypothesis never outranks a confirmed blocker.
  • Diagnose before prescribing. "Add schema" is not a diagnosis.
  • A myth guard that refuses 33 popular tactics with published counter-evidence, and offers the working alternative instead.
  • Defense, not offense. Adversarial techniques appear only as things to detect and withstand, never as recommendations.
  • Honest horizons. Every change ships with a verification method and a realistic timeframe — including "we cannot promise a date".

What knowledge is packed inside

Twenty-five reference contracts ship inside the skill, so they travel to every agent, not just Claude Code. This is the substance:

AreaWhat it holdsWhy it is worth having
Ranking modelSystems vs signals vs "factors", the three that actually carry weight, what E-E-A-T really is, query-dependent weighting, personalization and localityStops audits built on documentation reshuffles and listicle "factor" claims
Technical & indexationCrawl access, rendering traps, robots wildcard failures, index tiering, crawl-budget killers, tracking parameters versus facets, soft-404 collapse patterns, canonical traps, migration protocol, plus a full mechanical sweepThe failure modes that silently cost the most traffic, each with its exact observable
Architecture & equityEquity distribution, hub-and-cluster, orphans, depth, anchor practice, and the answer-engine read budgetExplains why money pages starve while the homepage hoards authority — and why navigation now costs twice
Intent & content valueThe four intents and the page types they reward, cannibalization mechanics, information-gain findings, the content types that survive zero-click, and the AI-content patterns that now hurtTurns "write better content" into a specific, testable page-level decision
AEO / GEO mechanicsHow an answer is actually built (fan-out → retrieval → grounding → arbitration), what correlates with citation, per-engine retrieval paths, extractability rules, and a ready prompt set for measuring brand presenceThe part most audits either skip or fill with vendor folklore
Agent surfaceThe .well-known discovery set with each spec's real status, the Markdown boundary the myth guard stops at, the OpenAPI properties that decide whether an LLM can call you, the agent auth-discovery chain, and how to read a third-party "agent-readiness" score without being led by itSeparates what is specified from what is scored, so an agent-readiness push does not become a checklist nobody measures
Entity & brand consensusCross-profile consistency, the entity graph, ghost citations (cited but not recommended), and how review sentiment drives AI verdictsExplains why a technically perfect site still gets no recommendation
Experience & conversionCWV triage order, satisfaction-signal mechanics, CRO × SEO evidence, conversion elements per template, lead capture, the attribution gaps (calls, offline, AI referrals, cross-device), paid × organic alignmentThe post-click half of the funnel that audits usually leave on the table
Risk & defensePenalty behavior, subdomain and registrar risk, indirect prompt injection, takedown abuse, canonical hijacking, behavioral poisoning, proportionate link-risk handlingTurns "we got hit" into a specific, checkable hypothesis
Google update timelineEvery core, spam and Discover update with start and completion dates, the platform changes that retired old tactics, and an update-response protocolLets the skill date-align a traffic curve instead of guessing
Growth plays61 plays, each with the trigger that justifies it, the mechanism, the observed effect and its evidence tierA plan built from things that measurably worked, not from best-practice lists
BenchmarksDated figures for surface reach, click economics, citation mechanics, read budget, content correlations, operational targets and industry contextLets the report size an opportunity with numbers instead of adjectives
Link buildingTarget selection, both collection modes (Search Console reachable or not), anchor discipline, the exclusions a brief must name, and the CSV column contractTurns the audit into a brief a contractor can execute without misreading an assumption as a measurement
Bulk market dataThe Prowl MCP as a rung-5 source — ~448 provider tools behind one pay-per-call endpoint, routed by track, with measured per-call costs and the operating gotchasGives a no-seat audit a second independent index, which is what makes a demand finding hold up
MethodEvidence tiers, experiment design, the myth guard, check → tool routing with DevTools recipes, and the deliverable templatesKeeps two different runs of the audit comparable

Data freshness

  • Sources last re-fetched 2026-08-16. That date has one home — references/algorithm-updates.md — and the gate holds this line equal to it, because this bullet said "Verified as of 2026-08-10" for six days after the corpus was refreshed, and a staleness claim that is itself stale is the number a reader uses to decide whether to trust the rest. ~6,000 non-blank lines of distilled reference material across the twenty-five contracts, rounded to the nearest hundred and recomputed by python3 test/validate.py on every run. The update timeline covers March 2025 → July 2026, and carries two dates of its own — when the sources were last re-fetched and what its newest row is.
  • Most benchmarks carry a date and a sample; the ones that do not say so. The 2026-08-10 audit found 36 of 140 rows in benchmarks.md undated at both row and section level — the whole "Operational benchmarks" block had a two-column shape with nowhere to put a source. Those rows now name their source and are marked undated where the corpus has no date, because a number an audit uses to set a verification window is a promise with an expiry. (An exact line count used to sit here. It was wrong four edits later, which is the same defect class this skill refuses everywhere else — a number about a thing, kept next to the thing, with nothing reconciling them.)
  • The skill ships a refresh routine: re-fetch the update sources, append the new rows, re-check whether a shipped change retires a myth or invalidates a benchmark, and downgrade any claim older than ~18 months that nothing has confirmed since.
  • Where credible evidence conflicts — and in AI-surface research it regularly does — the claim is demoted to a hypothesis and routed to the experiment path instead of being asserted.

The bundled scripts

They live inside the skill, not at the repository root — scripts/ here is the documentation gate. Every command below is written against $SKILL_DIR, so set it once for whichever context you are in:

SKILL_DIR=plugins/seo-aeo-audit/skills/seo-aeo-audit          # in a clone of this repo
SKILL_DIR="${CLAUDE_PLUGIN_ROOT}/skills/seo-aeo-audit"          # installed as a Claude Code plugin

An agent gets the same instruction from SKILL.md, and the validator rejects a path relative to the caller in either file: all eleven invocations were once written that way and none of them resolved from where anybody actually stands.

page_audit.py — per-page evidence

Stdlib only, nothing to install, works offline:

python3 "$SKILL_DIR/scripts/page_audit.py" --url https://example.com/pricing
python3 "$SKILL_DIR/scripts/page_audit.py" --file saved.html --base-url https://example.com/pricing
python3 "$SKILL_DIR/scripts/page_audit.py" --url-list urls.txt --format json > audit.json

It catches what eyeballing misses: content="none" (≡ noindex, nofollow), a meta refresh + noindex conflict, a canonical silently discarded because it carries media/type/hreflang, invalid JSON-LD, missing alt text, a price that exists in the source but not in extractable text (so engines cite an aggregator for your pricing), and the answer-engine read budget — how much of a ~5,700-character first read your navigation eats before the answer.

gsc_pull.py — the half a crawl cannot see

Search Console evidence: which queries a property actually surfaces for, at what position, and whether a drop is a cliff that held rather than a decline. Stdlib only; auth is local Application Default Credentials, so no key file goes near the repo.

python3 "$SKILL_DIR/scripts/gsc_pull.py" --list
python3 "$SKILL_DIR/scripts/gsc_pull.py" --site sc-domain:example.com --quota-project my-proj

It prints the position split first, deliberately. Ranking findings by impressions is the mistake this script exists to prevent: a large impression count beyond position 30 is usually the biggest number in the account and worth the least, while a small cluster inside the top 20 carries most of the clicks.

The cliff detector only reports a drop that stayed down — a single bad day is a reporting gap, and updates redistribute rather than zero a property for weeks. When it fires, it says so plainly and points at what the API cannot answer: Manual Actions and Index Coverage are web-UI only, at every scope.

Three auth gates fail independently with unhelpful errors — the OAuth scope, the API being enabled on a project the account can use, and the quota-project header that client libraries add and raw HTTP does not. The script names which one you hit.

scripts/url_inspection.py — asks the index rather than inferring from a fetch. The Google-selected canonical against the declared one, coverage state, robots verdict, last crawl: the engine's own answers, so findings built on them are CONFIRMED rather than inferences. Quota is 2000/day and 600/minute per property, so it samples — a representative URL per template, plus the pages a finding is about.

python3 "$SKILL_DIR/scripts/url_inspection.py" --site sc-domain:example.com --urls https://example.com/pricing

scripts/sitemap_audit.py — the published side of "indexed vs published per template": declared URLs clustered into the template families the site actually ships, derived from its own URLs rather than a guessed taxonomy. It does not detect orphans; a sitemap carries no link graph, and inferring them from path shape would be a guess wearing the clothes of a finding.

scripts/psi_pull.py — Core Web Vitals with field (CrUX) and lab (Lighthouse) reported separately, judged at the 75th percentile. Where CrUX has no data for a URL that is reported as absent, never as a pass, and the lab score never stands in for it.

python3 "$SKILL_DIR/scripts/sitemap_audit.py" --url https://example.com/sitemap.xml
python3 "$SKILL_DIR/scripts/psi_pull.py" --url https://example.com/pricing

Every one of them states its own blind spot in its output. That is the eighth non-negotiable: an instrument that cannot see something must not let its silence read as a measurement.

Link-building extraction

The audit also produces a deliverable for someone else to execute: a brief plus a keyword CSV a contractor can work from. Two modes — with Search Console, every row is measured; without it, candidates are derived from what the site sells.

The rule that makes it safe to hand over: measured and assumed never share a column. A source column separates them, and the volume cells of an unmeasured row stay blank, not zero0 reads as "measured, no demand", blank reads as "nobody has checked". An auditor who overstates a finding wastes their own time; a brief that overstates one spends a client's budget.

See references/linkbuilding.md for target selection, the exclusions a brief must name, anchor discipline, and the CSV column contract.

Security posture

Text plus seven standard-library Python scripts, and nothing else runs. Four of them are read-only network clients (page_audit.py, psi_pull.py, sitemap_audit.py, agent_surface.py); two call Google APIs with a token minted locally (gsc_pull.py, url_inspection.py); one probes access (preflight.py). None of them writes anything, and none submits, requests indexing or changes a property.

page_audit.py makes plain http(s) GETs to the URLs you hand it — any other scheme is refused before a request is made, redirects off http(s) are refused, non-HTML responses are refused, no cookies or credentials are sent, responses are bounded by --timeout/--max-bytes (and a truncated response says so rather than reporting a fragment as a measurement), and it writes nothing.

gsc_pull.py, url_inspection.py and preflight.py shell out to gcloud auth application-default print-access-token — that is the only subprocess any of them starts, no key file goes near the repo, and the token never leaves the process.

No dependencies, no npm lifecycle scripts, no telemetry. The installers (install.sh, bin/seo-aeo-audit.js) only copy files into ~/.claude/ and only when you run them. Full statement in SECURITY.md.

Marketplace scanners rate skills that ship executable code above documentation-only skills by default; that rating is about the category, not a finding — the audit trail above is what it takes to check the claim yourself.

Repo layout

.claude-plugin/marketplace.json      root manifest
plugins/seo-aeo-audit/
  ├── .claude-plugin/plugin.json
  ├── commands/seo-aeo-audit.md      slash command
  └── skills/seo-aeo-audit/
      ├── SKILL.md                   the procedure
      ├── references/*.md            25 contract files (shipped on every channel)
      └── scripts/                   7 stdlib-only scripts: page_audit, gsc_pull,
                                     url_inspection, psi_pull, sitemap_audit, preflight,
                                     agent_surface
cursor/rules/seo-aeo-audit.mdc       Cursor rule (contracts inlined)
templates/*.template.md              deliverable skeletons for non-agent use
bin/seo-aeo-audit.js                 npx installer (zero dependencies)
install.sh                           POSIX installer for a local clone
package.json                         npm manifest for the installer CLI
CONTRIBUTING.md                      evidence discipline, checks, where things go
SECURITY.md                          what runs, what it touches, how to verify
scripts/check-docs.sh                the gate — runs the test files below
test/validate.py                     structural validator and doctrine guards
test/plant_guard_test.py             the plant guard itself: did a negative self-test's
                                     damage actually land (content AND mode)
test/test_page_audit.py              page-auditor tests (offline fixtures)
test/test_url_inspection.py          index-verdict tests (documented API shape)
test/test_collectors.py              psi / sitemap / gsc / preflight tests
test/test_agent_surface.py           track-K collector: parsers, tiers, blind spots
test/test_output_contracts.py        exit status and markdown safety, all seven
test/residue.py                      what a run leaves on disk — one ledger, printed
                                     on every path, `nothing` included
test/residue_test.py                 fixtures for it, run last: the final case reads
                                     the TMPDIR every suite above shared
test/fixtures/*.html                 pages the auditor is tested against
.github/workflows/validate.yml       CI, including negative self-tests
docs/research/                       provenance behind every claim in the references

Development

These run in a clone of this repository. The published npm package ships no test/ directory, so from an install they are names, not commands.

bash scripts/check-docs.sh      # the gate: runs exactly the files below, nothing else
python3 test/validate.py        # structure, version sync, references, links, anchors, drift
python3 test/test_page_audit.py # page auditor against offline fixtures
python3 test/test_url_inspection.py  # index verdicts against the documented API shape
python3 test/test_collectors.py      # psi, sitemap, gsc and preflight behaviour
python3 test/test_agent_surface.py   # the track-K collector, offline
python3 test/test_output_contracts.py  # exit status and markdown safety across all seven
python3 test/residue_test.py         # what the run left on disk, and that it said so
node --check bin/seo-aeo-audit.js
bash -n install.sh

Version sync is a hard rule: marketplace.json, plugin.json, package.json and the top CHANGELOG.md entry carry the same semver. Every finding the auditor emits points at a reference section, and the validator resolves those anchors against the real headings — a renamed heading fails the build. CI proves each check can fail by corrupting a copy of the repo and requiring a non-zero exit.

Part of a family

seo-aeo-audit is one of the ssheleg skills — install or update them all with a single command:

npx sshlg-skills install
npx sshlg-skills update

What this gives you

Vibe-coded sites ship fast and land invisible: rendered entirely client-side, no canonical story, a sitemap nobody ever submitted, and nothing an answer engine can quote. Then the advice you get is a 200-row checklist that never says what to do first.

  • Blockers first. If the site cannot be crawled, rendered or indexed, nothing else matters — the audit says so and stops, instead of burying it on row 140.
  • Answer engines, not only Google. What ChatGPT, Perplexity, Copilot and AI Overviews can actually extract from your pages, and why they cite a competitor instead of you.
  • Evidence or silence. Every finding carries an observation, a location and a date; every recommendation carries an evidence tier, so you can tell proven from worth-testing at a glance.
  • A myth guard. 33 popular tactics with published counter-evidence are refused outright — the ones an agent will otherwise recommend with total confidence.
  • The output is a prioritized change plan with verification steps and honest horizons, not a score out of 100.

Contributing

Issues and pull requests are welcome — see CONTRIBUTING.md. The short version: every claim carries an evidence tier, benchmarks.md owns the numbers, dated facts stay dated, and bash scripts/check-docs.sh must pass — it runs python3 test/validate.py, python3 test/plant_guard_test.py, python3 test/test_page_audit.py, python3 test/test_url_inspection.py, python3 test/test_collectors.py, python3 test/test_agent_surface.py, python3 test/test_output_contracts.py and python3 test/residue_test.py. Everyone taking part is expected to follow the Code of Conduct.

Author

Built by ssheleg — sshlg.me

Part of the ssheleg skill family: super-ux, task-pipeline, agent-sync, make-skill, sheleg-design, seo-aeo-audit. The family installs and updates as one package, for every agent you use — a bundle with one member current and the rest stale is a combination nobody tested:

npx sshlg-skills install              # nothing installed yet — the whole family, any agent
npx sshlg-skills update               # installed but behind — updates everything
npx --yes sshlg-skills@latest list    # what the current release of each member is

Restart your agent afterwards: skills and hooks load at session start, so the session that updates is not the session that gets the new ones.

License

MIT © 2026 ssheleg.