3Dogs Pointer · evidence-first benchmark

A 100 elsewhere is the starting line.

A frozen Pointer Strict rubric reranked every site scoring 100 on the current LLM Scan leaderboard, then compared the same public-evidence standard with 3dogs.ai.

Snapshot: 8 August 2026pointer-strict-public-v0.1.0Current LLM Scan v1.0.4 cohortRead-only public scan
LLM Scan perfect-score sites4current leaderboard snapshot
Pointer Strict perfect scores0same frozen cohort
Best reranked score64.7deutschfox.com
3dogs.ai49raw 79.9, capped at 49

Result

No current LLM Scan 100 survives the stricter rubric.

The current four-site 100-point cohort reranks between 50.8 and 64.7. This is a dated, reproducible snapshot—not a claim that no website can ever score 100.

deutschfox.com moves from LLM Scan rank 3 to Pointer rank 1 at 64.7. The defining difference is that discoverability signals earn credit, but missing protocol/API proof and weak evidence hygiene still cost points.

Reranking

The same “100” sites, a different order.

3dogs.ai is included as the requested comparator. Its current Cloudflare Agent Ready reference is 100/100; it was not one of the four LLM Scan 100-point leaderboard rows used to define the cohort.

Pointer rankDomain / manifestPrior referencePointer StrictHard gatesPrimary deductions
1 deutschfox.come7fc6a7b0375084e… 100rank 3
64.7
raw 64.7 · cap 100
none Protocol truth 0/10 · API and task contracts 0/10 · Freshness and versioning 1.5/6
2 applikant.com156fd019d448adcf… 100rank 1
56.5
raw 56.5 · cap 100
none Claim evidence hygiene 0/12 · Protocol truth 0/10 · API and task contracts 0/10
3 raison.ist175513f53b0c8e6f… 100rank 4
52.9
raw 52.9 · cap 100
none Claim evidence hygiene 0/12 · Protocol truth 0/10 · API and task contracts 0/10
4 dreamkid-website.engineering-bb9.workers.dev18031bc015d40211… 100rank 2
50.8
raw 50.8 · cap 69
HG-IDENTITY HG-IDENTITY: Canonical identity points to another host · Claim evidence hygiene 0/12 · Protocol truth 0/10 · API and task contracts 0/10
5 3dogs.ai89a4f70efdb4a1ce… not in 100 cohortCloudflare reference: 100/100
49
raw 79.9 · cap 49
HG-AUTH HG-SOFT404 HG-AUTH: State-changing API operations lack explicit security · HG-SOFT404: Honest negative controls failed · Claim evidence hygiene 0.6/12 · Freshness and versioning 3/6

Why it is stricter

Critical failures cannot hide inside an average.

Pointer Strict is stricter specifically on deterministic public protocol and structural readiness. It does not claim to replace AEOlens model-observation capabilities or credentialed security testing.

Declarations are hypotheses

A file’s presence is only the first check. Referenced agent, MCP, OAuth and OpenAPI surfaces must resolve, parse and agree with the site identity.

Hard gates override averages

Identity conflicts, JavaScript-only critical content, soft-404s and unsecured state-changing API declarations can cap the final score.

Negative controls test honesty

Two deterministic nonexistent paths must return real 404/410 responses with bodies distinct from the homepage.

Representations are compared

Initial HTML, negotiated Markdown, canonical identity, structured data and discovery artifacts are scored for parity—not merely existence.

Safety boundaries are visible

The public tier uses only bounded GET requests and names what remains untested instead of implying a security certification.

Evidence is portable

Status, selected headers, byte counts, redirects, response-body hashes and the manifest hash can be downloaded for review.

Four-tool comparison

Different products answer different questions.

“Stricter” is meaningful only when the tested domain is named. The matrix separates structural proof, model visibility and public protocol evidence.

DimensionCloudflare Agent ReadyAEOlensLLM ScanPointer Strict v0.1
Primary questionCan agents discover and use the site’s declared capabilities?Is the brand structurally optimized and actually surfaced by major answer models?Can LLM crawlers reach and interpret eight public readiness signals?Do public declarations survive retrieval, parsing, parity and honest-negative-control tests?
Published scopeContent, discoverability, agent/API/auth/MCP capabilities; commerce protocols detected separately48 weighted structural checks plus buyer simulations across five model familiesEight areas: crawlability, robots, llms.txt, sitemap, Markdown, semantic HTML, structured data and Content-Signal11 evidence bands totaling 100 points; 14–15 bounded read-only requests in this run
Score disciplineReadiness score across published checksWeighted structural score plus visibility/citation measuresAdditive public-readiness scoreWeighted score plus hard caps: a critical failure can override a high average
Declaration vs proofDetects standards and published capabilitiesMeasures structural signals and model outcomesChecks public files and page signalsTreats declarations as hypotheses: fetches, parses and cross-checks the endpoints they name
Adversarial controlsNot documented on the reviewed public methodology pageNot documented on the reviewed public methodology pageNot documented in the eight-area public documentationTwo deterministic nonexistent-path probes expose soft-404s and catch-all discovery routes
Auth safety gateValidates published auth/discovery standardsNot the product’s stated focusNot one of the eight published scoring areasState-changing OpenAPI operations without explicit security cap the final score at 49
Model visibilityNot the core public scoreStrongest coverage: citation rate, position, sentiment and buyer simulationNot in the published eight-area scoreExplicitly untested by the deterministic public tier; a separate model-observation tier is planned
Evidence artifactInteractive result and check detailsDashboard, audit and observability outputsInteractive scan result and leaderboardDownloadable per-request status, headers, byte counts, body hashes and a manifest SHA-256

Method

Frozen before the cohort was scored.

The benchmark did not change weights after seeing results. That protects the “no 100” finding from becoming a manufactured outcome.

Public ruleset

  • 11 evidence bands totaling 100 points.
  • Maximum 2 MB per response, three redirects, eight-second request timeout.
  • Public targets only; private/reserved hosts and nonstandard ports blocked.
  • GET-only implementation in this prototype.
  • Response hashes and a manifest SHA-256 retained per scan.

Explicitly outside this tier

  • State-changing operation execution
  • Authenticated retrieval and tenant isolation
  • Semantic truth of claims
  • Actual LLM ranking, citation or recommendation behavior
  • Historical reliability, retry semantics and telemetry
  • Contractual controls and human operating practice
  • DNS connection pinning in this local prototype

Four-seat advisory panel

The rules were challenged by Nova Grounding, Nova Pro, Gemini Pro and Gemini Flash. Their output shaped hard-gate and safety review, but advisory model text is not treated as benchmark proof; official specifications and deterministic scan artifacts remain the evidence.

Prototype limitations

The local scanner uses a lightweight HTML parser and pre-request DNS validation without connection pinning. A production arbitrary-domain service should add durable rate limiting, DNS pinning/rebinding defenses, robust standards parsers, queue isolation and repeated-run telemetry.

Responsible promotion line

“Nobody in the current LLM Scan 100-point cohort cleared Pointer Strict v0.1.”

Use with the snapshot date and methodology link. Avoid “no company can score 100” or “strictest on the market” until a broader, scheduled and independently reproducible dataset supports those claims.

What 3Dogs has to work on

  • Return a real 404/410 for arbitrary nonexistent public routes.
  • Declare security on every state-changing commerce OpenAPI operation.
  • Add nearby sources, dates and evidence paths to consequential claims.
  • Improve explicit freshness/version metadata and task/error schemas.

These changes target the raw 79.9 evidence profile and remove the two caps; they do not guarantee 100 because the frozen ruleset will still apply.

Sources

Current product and standards references.

  1. LLM Scan leaderboard — cohort snapshot and ranking.
  2. LLM Scan documentation — eight published scoring areas and limitations of llms.txt.
  3. Cloudflare Agent Readiness — published readiness scope and supported standards.
  4. AEOlens — published 48-check audit and multi-model visibility features.
  5. RFC 9727 and RFC 9728 — API catalog and OAuth protected-resource metadata.
  6. MCP authorization specification and A2A protocol specification.
  7. WCAG 2.2 — accessibility reference.