Skip to main content

F1 — Audit Engine: Implementation Doc

Choose one page language / 选择页面语言

Overview

Status: Draft · Owner: Josh · Date: 2026-08-07 · Parent: Marketing-Engine-PRD.md §5/F1


1. The question first: why is this different from a dentist asking ChatGPT?

Straight answer: a single chat query and the audit are different instruments, the way stepping on a bathroom scale once differs from a monitored clinical measurement. The dentist asking their own phone is our sales demo — one emotionally persuasive sample. It cannot be the product, for reasons that are also exactly the build spec:

  1. One query is one sample of a non-deterministic system. AI answers vary by run, phrasing, engine, language, and location. "I asked ChatGPT and I wasn't there" proves nothing — and neither would our report, if that's all it were. The audit runs a battery: ~40–50 patient-realistic queries × 4 engine surfaces × EN/ZH × 3 trials each, geo-controlled, clean sessions. It reports citation frequency — "you appeared in 2 of 120 relevant answers; Dr. Chen appeared in 74" — a measurement, not an anecdote. Our own honesty rule (never claim certainty about AI outputs) requires this design.
  2. The diagnosis, not the symptom. A chat answer tells you you're invisible. It cannot reliably tell you why — ask a frontier model to "audit my website" and you get plausible generic advice, partly hallucinated. Our audit pairs the query battery with deterministic checks (robots.txt actually blocking GPTBot?, schema actually present?, Bing Places actually listed?) run against the real site. Every "why" in the report is a verified fact, and every fact maps to a fix we sell.
  3. Source attribution. Perplexity and Google AI Overviews cite sources. The audit captures which sources drive the answers in each market (Yelp? Healthgrades? a local news piece?) — that's the actionable map of where authority lives locally, and it differs per city. Invisible to someone reading answer text casually.
  4. The competitor league table. Same battery, run for the 5–10 nearest implant competitors, normalized into a ranking. A dentist could do this by hand in principle; nobody will do 600+ controlled queries per month by hand.
  5. The longitudinal line. Baseline stored, methodology frozen, re-run monthly. Month-3's "here is your citation frequency moving" chart is the entire proof of the 2–4 month GEO promise. A chat session has no memory and no baseline.
  6. Evidence-grade capture. Raw transcripts, screenshots, timestamps, and a methodology note ship with every report. "No fabricated audit data" stops being a compliance rule and becomes the differentiator: our numbers are reproducible; a chat anecdote isn't.

So the moat is not access to the models — the dentist has that. It's methodology + diagnosis + tracking at a cost per audit (see §6) no human process matches. If a competitor wants to replicate it, they have to build this same instrument; at that point they're us, minus the warm network.

2. Scope

In: query battery generation, multi-engine runners, mention/citation extraction, deterministic site & listings checks, competitor comparison, scoring, snapshot storage, report rendering. Out (F1): any fixes themselves (F2), campaign data (F3/F4), client dashboard, Chinese-platform presence checks (manual note in report for now).

3. Architecture

practice-config ──► [A] Query Battery Generator ──► [B] Engine Runners ──► raw transcripts
      │                                                                        │
      ├──────────► [C] Deterministic Checkers ──► verified facts               ▼
      │                                                │            [D] Mention/Citation
      │                                                │                Extractor
      ▼                                                ▼                       │
   [F] Snapshot Store (append-only) ◄──────────────────┴───────────────────────┘
                    │
                    ▼
   [E] Scorer & Competitor Table ──► [G] Report Renderer (PDF/HTML + evidence appendix)

A — Query battery generator

Templates instantiated from practice-config (city/geo, services, languages): discovery ("best implant dentist near {city}", "who does all-on-4 in {city}"), cost ("dental implant cost {city}", "affordable full-arch {city}"), trust ("is {competitor} good for implants"), and Chinese-language equivalents ("{city} 种植牙 推荐" class). ~40–50 queries per practice. Battery is versioned and frozen per client so monthly deltas compare like with like.

B — Engine runners (the realistic part)

Per-engine reality check — build in this order:

SurfaceAccessNotes
PerplexitySonar API — returns citations nativelyBest effort/value; build first
OpenAI (ChatGPT-class)Responses API with web-search toolCaveat: API answers ≈ but ≠ consumer app answers; disclosed in methodology note
Google AI Overviews / AI ModeNo first-party API → SERP providers (DataForSEO / SerpAPI capture AI Overview blocks, geo-parameterized)Main per-query cost; best geo control
GeminiAPI with search groundingCheap to add once runner abstraction exists
Bing CopilotNo practical APISkip for pilot; note in methodology

Rules: 3 trials per query per engine; no logged-in/personalized sessions; geo via SERP params where supported, via explicit prompt framing ("I live in {city}") where not — stated in the methodology note. Do not scrape consumer chat UIs (ToS + fragility); the handful of hero screenshots for pitch decks are captured manually from a clean session — real captures, per the honesty rule.

C — Deterministic checkers

No LLM judgment anywhere in this layer:

  • robots.txt / meta directives vs. the AI-crawler list (GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended, etc.)
  • Schema.org JSON-LD presence + validity (Dentist/LocalBusiness, FAQPage), sitemap, llms.txt
  • Google Business Profile via Places API: exists, claimed-signals, rating, review count/recency, photos
  • Bing Places, Yelp, Healthgrades, Zocdoc presence; NAP (name/address/phone) consistency across them
  • Basic site vitals (mobile viewport, HTTPS, load time bucket)

D — Mention/citation extractor

The hard problem: did the practice actually appear? Practices surface under variant names ("Sunrise Dental", "Dr. Wei Zhang DDS", domain, phone).

  • Pass 1: deterministic matching (name variants from config, domain, phone) against answer text and citation URLs.
  • Pass 2: LLM extraction constrained to quoting exact spans from the transcript, then string-verified against the raw text before storage. An extraction that can't be located verbatim in the transcript is discarded. This keeps the no-fabrication rule mechanical, not aspirational.
  • Same pipeline runs for each configured competitor.

E — Scorer

Per engine: citation frequency (appearances / relevant answers), citation position where rankable, and source-driver tally (which domains the engines cited). Overall: simple weighted visibility score (weights frozen in methodology doc — resist tuning it to flatter anyone). Competitor league table from identical runs.

F — Snapshot store

Append-only. Every run: battery version, raw transcripts (JSON), screenshots, checker outputs, extracted mentions, scores, timestamps. Postgres + object storage for artifacts. Monthly delta = diff of two snapshots; never recomputed retroactively.

G — Report renderer

One-page summary (visibility score, league table, top 5 verified problems, top 5 fixes) + evidence appendix (screenshots, methodology note, caveats). HTML → PDF. The methodology note ships in every report — it's both honesty and positioning ("this is an instrument, not an opinion").

4. Build phases

Phase 0 — pitch-ready (days): battery generator + Perplexity & OpenAI runners + manual hero screenshots + checker subset (robots.txt, schema, GBP, Bing Places) + hand-assembled report template. This is enough to make every prospect audit real. Corners cut consciously: no extractor pass 2 (Josh eyeballs mentions), no scoring, no store — but raw outputs saved to disk so nothing is lost for baselines.

Phase 1 — instrument (2–3 weeks): extractor with verification, full checker set, scorer, snapshot store, automated PDF. Every signed client gets a stored baseline from here.

Phase 2 — longitudinal (by month 2): scheduled monthly re-runs, delta report section, AI Overview via SERP provider, ZH battery fully automated, Gemini runner.

5. Interfaces

  • Input: practice-config record (PRD §10.2) — this doc adds required fields: name variants list, competitor list (5–10), geo coordinates, ZH name if any. - Output: snapshot ID + report artifact; consumed by F5 (reporting) and by the F2 backlog (each verified problem auto-creates a fix candidate in the approval queue).

6. Cost & effort per audit

~50 queries × 4 surfaces × 3 trials ≈ 600 calls: LLM APIs low single-digit dollars; SERP provider ~$5–15; total ≈ $10–25 per full audit, minutes of wall-clock time, near-zero marginal human time at Phase 1+. Compare: a human doing a degraded version of this is a day-plus of work. That ratio is the business.

7. Risks & honest caveats (these go in the methodology note)

  • API ≠ consumer app. Our automated numbers approximate what patients see; hero screenshots from real consumer sessions bridge the gap. Never present API-derived frequency as literally "what ChatGPT shows every patient." - Non-determinism cuts both ways. Report frequencies with denominators; never "you never appear." Also true post-engagement: month-3 improvements are trends, not guarantees — consistent with the no-promises rule. - Personalization variance: patients' logged-in results differ from clean sessions; disclosed. - Extractor false negatives (practice under an unconfigured alias) — mitigated by the name-variants field and a manual spot-check on each client's first run. - Engine surface churn: runners will break as products change; runner abstraction + per-engine health checks, and the methodology note is versioned so reports always describe what was actually measured.

8. Open items

  1. SERP provider selection (DataForSEO vs SerpAPI — trial both on one market, compare AI Overview capture quality/geo fidelity). 2. Whether prospect audits (pre-sale, free) run the full battery or a cheaper half-battery. Lean: half-battery, full on signing — it also gives the "your full baseline" moment some weight. 3. ZH battery phrasing needs a native-speaker pass before it's trusted (Jim or a study-group volunteer).