Skip to content
martechllc
Pillar I · Tools / LLM TrackerLive
martechllcLLM Tracker

Answer-engine visibility, measured honestly

Visibility scores with confidence intervals. Receipts for every claim.

LLM Tracker asks the questions your buyers ask — across ChatGPT, Google AI Overviews, Perplexity, Gemini, and Claude — then scores how your brand is found, cited, quoted, and recommended. Every prompt runs repeatedly, every score ships its plausible range, every citation is verified against the live page. No bare numbers, ever.

Sample prompt · What's the best issue tracker for an engineering team that ships every day?

The reading grammarillustration
Selection score

When the AI picks its sources for an answer, does anything you own make the cut?

62%(55–69, n=5)
plausible range 55% to 69%, based on 5 samples
within noisereceiptevidence 3fa1c97e

This is how every score in the product renders: the number, its plausible range and sample count, whether a movement cleared the noise gate, and the receipt that opens the stored evidence. Illustration only — no probe ran to draw it.

Curious what ChatGPT says about your domain right now?

The free probe runs one live, search-grounded ChatGPT question about your domain — real answer, real citations, no invented numbers.

Run my free probe

01 · The shift

Your buyers ask AI first. The answer names someone — is it you?

When someone asks ChatGPT, Perplexity, Gemini, or Google’s AI Overview which product to choose, the engine composes one answer, cites a handful of sources, and usually recommends by name. There is no page two. If a rival owns that answer, you lose the deal before you ever knew the question was asked.

Most teams answer this with a tool that prints a single “AI visibility score” — one number, no error bars, no evidence. But answer engines are non-deterministic: ask the same question twice and the answer changes. A score without a range is a coin flip wearing a suit.

LLM Tracker was built on the opposite premise: measure like it matters. Repeat every prompt, publish every score with its plausible range, verify every citation against the live page, and keep a receipt for every claim.

02 · The pipeline

Seven stages, from question to recommendation.

One visibility number hides where you actually lose. The tracker scores each stage of the path separately, so the fix is obvious: an engine that never searches for your market is a different problem from an engine that searches, finds rivals, and quotes them instead of you.

  1. 01

    Grounding score

    Did the AI actually search the live web before answering, or answer from memory?

  2. 02

    Entity resolution score

    Does the AI know who you are — does your brand show up as itself, not a lookalike?

  3. 03

    Selection score

    When the AI picks its sources for an answer, does anything you own make the cut?

  4. 04

    Absorption score

    When your page is cited, how much of the answer's actual wording comes from it?

  5. 05

    Behavioral eligibility score

    Does the answer actually steer the reader toward you — a recommendation, top pick, or concrete next step?

  6. 06

    Share of voice

    Of the answer real estate that goes to you and your competitors, what share is yours?

  7. 07

    Sentiment

    When the AI mentions you, is it saying something good, bad, or hedged?

  8. Each stage is measured independently, per engine, with its own range and sample count — and a stage that has not been measured says “not yet measured”, never an invented zero.

03 · The engines

The five surfaces your buyers already use.

Prompts run on every engine several times per run — repeated sampling is what turns a lucky answer into a defensible score. API-served engines are captured directly; Google’s AI Overview is captured from the rendered results page, because that is the only place it exists.

ChatGPTanswers captured directly
Google AI Overviewsread off the rendered page
Perplexityanswers captured directly
Geminianswers captured directly
Claudeanswers captured directly

Engines cite very differently — one picks few sources, another many — so scores are compared within one engine over time, never across engines. Stated on every chart, not buried in a methods page.

04 · The honesty contract

A tracker that can say “we don’t know yet.”

Three rules hold everywhere in the product, and they are the reason the numbers can be taken into a board meeting:

No bare numbers

Every score renders with its plausible range and sample count. The band below is the actual device, shown here as an illustration:

62% (55–69, n=5)
plausible range 55% to 69%, based on 5 samples
illustration — not a live reading

Noise is named

A movement only counts as change when it clears a statistical significance gate. Anything smaller is labeled for what it is:

within noise+9%real changeillustration — both chips, as they render

Receipts for every claim

Citations are verified against the live page — fabricated URLs never count for anyone — and every published score is sealed into an append-only evidence ledger you can open:

evidenceevidence 3fa1c97e, verifiedillustration — the receipt chip on every score

05 · The loop

Measuring is half the loop. Citerra is the other half.

LLM Tracker tells you where the engines find, cite, and recommend you — and where they don’t. Citerra scores the page behind each gap and hands back a paste-ready fix; a re-run here proves whether the fix moved the number. Measure, fix, re-measure — a score you can act on, then verify.

06 · Access

5 beta seats open. Real count, honestly kept.

LLM Tracker runs as an invite beta with 5 seats — none taken yet. That number is the same one our homepage and beta page publish, kept honest the same way the scores are: we would rather show you a real zero than an invented waiting list.

A seat gets the full instrument for your brand: guided onboarding, scheduled runs across all five engines, the seven-stage pipeline with confidence intervals, the evidence ledger, competitor tracking, and read-only share links for your stakeholders.

Request a beta seatnext door opens when the next tool ships

Answer layer

Questions worth answering plainly.

01

What is LLM Tracker?

LLM Tracker is Martech LLC's answer-engine visibility instrument.

It runs buyer-shaped prompts against ChatGPT, Google AI Overviews, Perplexity, Gemini, and Claude on a schedule, verifies every citation against the live page, and scores your brand across seven pipeline stages — every score published with its confidence interval and sample count, never as a bare number.

02

Why does every score have a range next to it?

Because answer engines are non-deterministic: the same prompt can produce different answers minutes apart.

LLM Tracker asks each prompt several times per run and publishes the score as a distribution — for example 62 (55–69, n=5) — so you can tell a real reading from a lucky sample. A movement smaller than the measurement's own spread is labeled "within noise" instead of being sold as a win.

03

Which engines does it measure?

Five: ChatGPT, Google AI Overviews, Perplexity, Gemini, and Claude.

API-served engines are captured directly; Google AI Overviews is captured from the rendered results page, because that is the only place it exists. Engines cite very differently, so scores are compared within one engine over time rather than across engines.

04

How is this different from other AI visibility trackers?

Most trackers hand back a single visibility number with no error bars and no evidence.

LLM Tracker treats measurement honestly: repeated sampling with confidence intervals, significance-gated change detection, citation verification that catches fabricated URLs, and an append-only evidence ledger where every published score can be traced to the stored answer it came from.

05

What do the seven stages measure?

They follow the path from question to recommendation: grounding (did the engine search the live web), entity resolution (does it know who you are), selection (do your pages make its source set), absorption (how much of the answer's wording comes from your page), behavioral eligibility (does the answer steer the reader to you), share of voice (how much of the competitive answer real estate is yours), and sentiment (how you are framed when mentioned).

06

What does the free probe do?

It runs one live, search-grounded ChatGPT probe about your domain — the same engine machinery the tracked runs use — and hands back the real answer with your domain highlighted, the grounding verdict, and every source the engine cited.

It deliberately shows no scores: one probe is a glimpse, not a measurement; measured scores with plausible ranges come from repeated runs inside the product.

07

How does LLM Tracker relate to Citerra?

They are two halves of one loop.

LLM Tracker measures whether answer engines find, cite, and recommend your brand; Citerra scores your page and hands back the paste-ready fix; then a re-run proves whether the fix moved the number. Measure, fix, re-measure.

We built LLM Tracker because every “AI visibility” number we saw in the wild was a point estimate pretending to be a fact. If a metric moves within its own noise, this product says so — even when the honest answer is less exciting than the chart a competitor would have shown you. That restraint is the feature.

— Sundar Ramesh Kumar, founder

The answers are being composed right now. Start measuring them.

Run one free live probe and see what ChatGPT says about your domain in about a minute, or take one of the 5 open beta seats and run the full instrument — confidence intervals, verified citations, and receipts included.

← All instruments· DM the founder (opens in a new tab) or [email protected]
Beta · 5 seats · request a seatRequest