Declarations are hypotheses
A file’s presence is only the first check. Referenced agent, MCP, OAuth and OpenAPI surfaces must resolve, parse and agree with the site identity.
3Dogs Pointer · evidence-first benchmark
A frozen Pointer Strict rubric reranked every site scoring 100 on the current LLM Scan leaderboard, then compared the same public-evidence standard with 3dogs.ai.
Result
The current four-site 100-point cohort reranks between 50.8 and 64.7. This is a dated, reproducible snapshot—not a claim that no website can ever score 100.
Reranking
3dogs.ai is included as the requested comparator. Its current Cloudflare Agent Ready reference is 100/100; it was not one of the four LLM Scan 100-point leaderboard rows used to define the cohort.
| Pointer rank | Domain / manifest | Prior reference | Pointer Strict | Hard gates | Primary deductions |
|---|---|---|---|---|---|
| 1 | deutschfox.come7fc6a7b0375084e… | 100rank 3 | 64.7 raw 64.7 · cap 100 |
none | Protocol truth 0/10 · API and task contracts 0/10 · Freshness and versioning 1.5/6 |
| 2 | applikant.com156fd019d448adcf… | 100rank 1 | 56.5 raw 56.5 · cap 100 |
none | Claim evidence hygiene 0/12 · Protocol truth 0/10 · API and task contracts 0/10 |
| 3 | raison.ist175513f53b0c8e6f… | 100rank 4 | 52.9 raw 52.9 · cap 100 |
none | Claim evidence hygiene 0/12 · Protocol truth 0/10 · API and task contracts 0/10 |
| 4 | dreamkid-website.engineering-bb9.workers.dev18031bc015d40211… | 100rank 2 | 50.8 raw 50.8 · cap 69 |
HG-IDENTITY | HG-IDENTITY: Canonical identity points to another host · Claim evidence hygiene 0/12 · Protocol truth 0/10 · API and task contracts 0/10 |
| 5 | 3dogs.ai89a4f70efdb4a1ce… | not in 100 cohortCloudflare reference: 100/100 | 49 raw 79.9 · cap 49 |
HG-AUTH HG-SOFT404 | HG-AUTH: State-changing API operations lack explicit security · HG-SOFT404: Honest negative controls failed · Claim evidence hygiene 0.6/12 · Freshness and versioning 3/6 |
Each domain’s JSON preserves the request manifest SHA-256 shown above. Source leaderboard data and scan timestamps are included in the evidence bundle.
Why it is stricter
Pointer Strict is stricter specifically on deterministic public protocol and structural readiness. It does not claim to replace AEOlens model-observation capabilities or credentialed security testing.
A file’s presence is only the first check. Referenced agent, MCP, OAuth and OpenAPI surfaces must resolve, parse and agree with the site identity.
Identity conflicts, JavaScript-only critical content, soft-404s and unsecured state-changing API declarations can cap the final score.
Two deterministic nonexistent paths must return real 404/410 responses with bodies distinct from the homepage.
Initial HTML, negotiated Markdown, canonical identity, structured data and discovery artifacts are scored for parity—not merely existence.
The public tier uses only bounded GET requests and names what remains untested instead of implying a security certification.
Status, selected headers, byte counts, redirects, response-body hashes and the manifest hash can be downloaded for review.
Four-tool comparison
“Stricter” is meaningful only when the tested domain is named. The matrix separates structural proof, model visibility and public protocol evidence.
| Dimension | Cloudflare Agent Ready | AEOlens | LLM Scan | Pointer Strict v0.1 |
|---|---|---|---|---|
| Primary question | Can agents discover and use the site’s declared capabilities? | Is the brand structurally optimized and actually surfaced by major answer models? | Can LLM crawlers reach and interpret eight public readiness signals? | Do public declarations survive retrieval, parsing, parity and honest-negative-control tests? |
| Published scope | Content, discoverability, agent/API/auth/MCP capabilities; commerce protocols detected separately | 48 weighted structural checks plus buyer simulations across five model families | Eight areas: crawlability, robots, llms.txt, sitemap, Markdown, semantic HTML, structured data and Content-Signal | 11 evidence bands totaling 100 points; 14–15 bounded read-only requests in this run |
| Score discipline | Readiness score across published checks | Weighted structural score plus visibility/citation measures | Additive public-readiness score | Weighted score plus hard caps: a critical failure can override a high average |
| Declaration vs proof | Detects standards and published capabilities | Measures structural signals and model outcomes | Checks public files and page signals | Treats declarations as hypotheses: fetches, parses and cross-checks the endpoints they name |
| Adversarial controls | Not documented on the reviewed public methodology page | Not documented on the reviewed public methodology page | Not documented in the eight-area public documentation | Two deterministic nonexistent-path probes expose soft-404s and catch-all discovery routes |
| Auth safety gate | Validates published auth/discovery standards | Not the product’s stated focus | Not one of the eight published scoring areas | State-changing OpenAPI operations without explicit security cap the final score at 49 |
| Model visibility | Not the core public score | Strongest coverage: citation rate, position, sentiment and buyer simulation | Not in the published eight-area score | Explicitly untested by the deterministic public tier; a separate model-observation tier is planned |
| Evidence artifact | Interactive result and check details | Dashboard, audit and observability outputs | Interactive scan result and leaderboard | Downloadable per-request status, headers, byte counts, body hashes and a manifest SHA-256 |
Method
The benchmark did not change weights after seeing results. That protects the “no 100” finding from becoming a manufactured outcome.
The rules were challenged by Nova Grounding, Nova Pro, Gemini Pro and Gemini Flash. Their output shaped hard-gate and safety review, but advisory model text is not treated as benchmark proof; official specifications and deterministic scan artifacts remain the evidence.
The local scanner uses a lightweight HTML parser and pre-request DNS validation without connection pinning. A production arbitrary-domain service should add durable rate limiting, DNS pinning/rebinding defenses, robust standards parsers, queue isolation and repeated-run telemetry.
Responsible promotion line
“Nobody in the current LLM Scan 100-point cohort cleared Pointer Strict v0.1.”
Use with the snapshot date and methodology link. Avoid “no company can score 100” or “strictest on the market” until a broader, scheduled and independently reproducible dataset supports those claims.
These changes target the raw 79.9 evidence profile and remove the two caps; they do not guarantee 100 because the frozen ruleset will still apply.
Sources