// WE SPEAK FLUENT ROBOT

Can AI even see your site?

ChatGPT, Perplexity and Gemini are the new front page. We crawl your site the way their bots do and find exactly what blocks you — before your customers ask a question you don't get to answer.

Free. No signup. ~60 seconds. Spoiler: it's often a 403.
0/100
GPTBot ClaudeBot PerplexityBot Amazonbot
// HOW IT WORKS
01
Crawl as the bots

We fetch your pages with the same user agents AI assistants use — and compare against what humans get.

02
Find what blocks you

WAF rules, robots.txt bans, missing llms.txt, broken metadata — scored and prioritized by severity.

03
Fix it in plain English

Every finding comes with the exact fix. No jargon, no 40-page PDF report.

// REAL FINDING
Your firewall is silently blocking ChatGPT
HTTP 403 for GPTBot on /features while the plain crawl got 200 — an edge/WAF rule is blocking AI crawlers.
// THE FULL SWEEP

One audit, both worlds.

AI visibility is the headline, but we run the unglamorous classic checks too — the ones that quietly cost you clicks.

Classic SEO
  • titles: missing, long, duplicated
  • doubled brand suffix in titles
  • meta descriptions: missing, short, long, truncated
  • canonicals: missing or mismatched
  • og:image present
  • noindex vs sitemap conflicts
  • sitemap hygiene: errors, redirects
  • new & removed pages tracked
  • Article JSON-LD structured data
AEO / AI visibility
  • robots.txt AI policy, 20 crawler tokens
  • WAF / edge blocking detection
  • llms.txt + soft-404 detection
  • content visible without JavaScript
  • 71 AI crawlers recognised in live traffic
  • weekly citation checks: ChatGPT, Claude, Gemini, Perplexity

Where those two numbers come from, since a number without provenance is a decoration: the robots.txt check evaluates your file against 20 named crawler tokens — 13 answer-engine fetchers (two of which, Googlebot and Bingbot, are classic crawlers that also ground AI answers) and 7 training crawlers. Separately, on connected sites, the edge recognises 71 AI crawler tokens in real traffic. What we can see there →

// BEYOND THE AUDIT

Findings → fixes → proof.

Plain-English findings

Every finding names the problem in human words and ships with the exact fix. No jargon safari.

Applied, not just advised

AI-drafted metadata you approve — or edit first, then approve — merged into your HTML at the edge. Server-side, so a crawler that never runs JavaScript still sees it. No CMS deploy, no plugin. One DNS record →

Structured data, drafted from your page

An article with no Article JSON-LD gets one, written only from what is already on the page and the facts you confirmed — no invented author, no invented date, no invented numbers. You approve it and we serve it as a <script type="application/ld+json"> at the edge, with no code on your side. It is added, not substituted: a page that already ships its own JSON-LD keeps it. This one is never automatic, and you tell us which paths are articles.

Impact, measured honestly

Before/after CTR and position at 14 and 28 days — labeled correlation, not causation. We tell you if a change helped, hurt, or didn't matter. No causation theater. And those verdicts are what earn a fix type the keys →

// NOT JUST THE DIAGNOSIS

We can serve the fix, too.

Finding a broken description is the easy half. The hard half is that fixing it means a CMS you don't control, a deploy you have to schedule, or a plugin you don't want. So we took that step out: point one DNS record at us and the changes you approve are merged into your HTML at the edge, before it leaves Cloudflare — server-side, so a crawler that never runs JavaScript still sees them.

It works the same on WordPress, Shopify, Squarespace, a Vercel app or a VPS, because from where we stand they're all just an origin. No plugin. No script tag. No Cloudflare account of your own. No code change.

It also means we're in the path your visitors take, which is a real thing to ask of you. Here's exactly what that costs and how you get out →

Two records, total

One TXT to prove the domain is yours, one CNAME when you're ready for us to serve it. See the real screens — including the ones where it fails →

The exit is a DNS record

Traffic follows DNS. Point it back and we're gone as soon as it propagates. Short of that, one button turns us into a plain forwarder within sixty seconds.

Optional, and it stays optional

Audits, findings, drafts and Search Console all work without it. And it's on every plan, including Free.

// THE NUMBERS NOBODY ELSE HAS

Your analytics can't see a crawler.

Google Analytics is a script in a browser. GPTBot doesn't run scripts, so as far as your dashboard is concerned, the assistant that read your entire site last Tuesday never visited.

We're in the response path, so we count what actually happened: which AI crawlers fetched which pages, which assistants sent you a human afterwards, and which search engine that referral really came from — per engine, not lumped into "organic". Gemini and Copilot count as assistants, not as search, because filing them under Google would erase the one number this product exists to show you.

Underneath that we also count everything your site served through us, sorted into 9 categories — human page views split by where the person came from, AI crawlers, other bots, assets, and two catch-alls for what we could not place. A count is all it is: a site, a day, a category, an integer. Daily, on every plan including Free.

In your analytics we keep the referring hostname and nothing else — not the path, not the query. The one measurement we take across all traffic, and what it cannot contain →

// PRICING
Free
$0
  • 1 site
  • 100 pages per audit
  • weekly audits
  • full AEO checks
  • 30 days of history
  • 10,000 edge events / month, pooled across your sites
  • 250k requests / month, pooled across your sites
start free
Studio
$79/mo
  • 10 sites
  • 2,000 pages per audit
  • daily audits
  • 2 years of history
  • 2,500,000 edge events / month, pooled across your sites
  • 10M requests / month, pooled across your sites
  • everything in Pro
start studio
// TWO METERS, AND WHY

Plans carry two allowances because they measure two different things, and rolling them into one would make both wrong. Edge events are the AI-and-search traffic we keep detail for — which crawler fetched which page, which assistant sent you a person, which engine a search visit came from. Requests are everything our edge served for you, assets included: it's the honest total, it's what we actually pay Cloudflare for, and on a normal site it's twenty to forty times the number of page views, because one page pulls that many stylesheets, scripts and images. Both allowances are per account, pooled across every site on the plan — not per site. Pooling means a busy site borrows headroom from quiet ones instead of paying overage while another site's allowance sits unused.

Past the included requests, the rate is $2.00 per additional million. Here is the arithmetic, because we'd rather you checked it than trusted it: a counted request costs us about $0.81 per million — roughly $0.80 to serve and about a cent to count — so $2.00 is about 2.5× our marginal cost. Those input prices are Cloudflare's published ones and the model that uses them is the same one we run internally.

Past the included edge events, the rate is $20.00 per additional million, prorated — same arithmetic, shown the same way: an accepted event costs us about $9.00 per million (it's the expensive lane — every event is a request we serve, rate-limit and roll up), so the rate is about 2.2× our marginal cost.

Nothing is billed today. There's no checkout, so the rate above is a statement of what we'd charge, not an invoice. And whatever the number does, your counters never stop — going over would change what you pay one day; it would never make your totals wrong, blank a chart, or start sampling. A product whose pitch is "measured, not asserted" does not get to stop measuring at the moment it gets interesting.

// ON EVERY PLAN, INCLUDING FREE

The managed edge — one DNS record, any hosting platform, no plan gate. Fail-open serving. A sixty-second pass-through switch that is yours, not support's. A public status page. Per-engine edge analytics and Search Console — including the full traffic breakdown, all 9 categories, daily, on Free. The MIT engine. And a human at hello@citeworthy.io.

What paid buys is not access, it's memory and cadence: how far back your history goes, how much traffic we'll count, and how often we audit. Individual edge hits are kept 7 days on every plan — the daily rollups are the record.

Paid plans aren't open yet — there's no checkout, so there's nothing to buy today. Pro launches soon; the free tier is real and works now. Weighing us against something else? See how we compare — or run the whole engine yourself: it's MIT-licensed and public.

// OPEN SOURCE
We started as open source, and the engine still is.

Citeworthy is built on seo-agent, an MIT-licensed Cloudflare Worker we publish in full — the crawler, the rule set, the AEO checks, the proposal model and the edge injector. It is not a trimmed-down demo of the paid product and there is no second, better engine behind the login. The version running this account is v1.30.1, the same release you can clone today.

So self-hosting is a real path, not a gesture at one: deploy it to your own Cloudflare account, point it at your own site, keep every byte of the data, pay us nothing. What you take on is the operating — one instance per site, your own bindings, your own upgrades, and no support line. What you buy from us is the other half: the judgment, the review lane, the verdicts, and somebody whose job it is when the edge misbehaves at 3am.

What the engine does, and what self-hosting actually requires → · github.com/citeworthyio/seo-agent

// WHAT CITEWORTHY IS

Citeworthy, in plain terms.

Citeworthy is a website audit application from Alan Wizemann, LLC. It crawls websites you own, finds SEO and AI-visibility problems (broken metadata, crawler blocks, missing AI-agent files), suggests fixes you review and approve, and — on connected sites — serves those approved fixes from the edge without changing your site's code.

If you connect Google Search Console, Citeworthy uses read-only access to display your own search performance (clicks, impressions, positions) inside your dashboard and to measure the before/after impact of changes. That data is only ever shown to you, is never used for anything else, and disconnecting removes it. Details in our privacy policy.