Methodology

Why your results can look different from one scan to the next

Branch Tracker AI runs your queries against live AI models with live web search. That's the honest way to measure AI visibility - and it means no two scans will ever be perfectly identical. Here's exactly why, and how to read the numbers.

How each scan works

For every query you've added, we ask ChatGPT, Gemini, Claude and Perplexity the same question a real Australian user would type. We use each provider's live web search, pinned to an Australian location, with no chat history and no personalisation. The answer text and the URLs the model cited are stored, then we run brand and competitor detection over the response.

Models currently used: GPT-4o (ChatGPT), Claude Sonnet 4.5 (Claude), Gemini Flash (Gemini) and Sonar Pro (Perplexity). If a provider changes their default consumer model, we update ours to match.

Why results vary between scans

  1. Models are non-deterministic. LLMs sample their answers probabilistically. The same prompt can return a differently worded answer, a different set of recommended brands, or a different ordering, even seconds apart. This is true of every serious AI tracking tool - it is not a bug.
  2. The web changes constantly. Each scan triggers a fresh web search. New articles, review updates, price changes and index refreshes all shift which sources the model reads and cites.
  3. Citation order shifts. Gemini and Perplexity return sources in the order their retriever ranks them at that moment. Which URLs land in your "top cited sources" can move even when the answer text is similar.
  4. Sentiment is also sampled. We classify positive / neutral / negative with a lightweight model. Clear cases stay stable; borderline mentions can flip between neutral and positive.
  5. Your own detection is deterministic. Brand and competitor matching runs the same logic every time. If a run says you weren't mentioned, it's because the model genuinely didn't mention you in that answer - not because we missed it.

How to read the numbers

  • Treat a single scan as a snapshot, not a verdict. One run can swing 10-20% either way on visibility score.
  • The trend across weekly scans is the real signal. Four to eight scans smooths out the noise and shows whether you're actually gaining or losing ground.
  • Share of voice vs. competitors is more stable than any single query result, because it averages across every query and platform in the scan.
  • The Opportunities list (queries where competitors show up and you don't) is the most actionable output - those gaps repeat scan-to-scan.

What we do to reduce variance

  • Same prompt, same system message, same Australian location metadata every run.
  • No chat history, no personalisation, no logged-in accounts.
  • Deterministic brand and citation matching over the returned text and cited URLs.

If you want tighter results at higher cost (multiple samples per query, averaged), get in touch - it's available on custom plans.

Still seeing something that looks wrong?

If a run says you weren't mentioned and you can reproduce a mention in the consumer app right after, that's usually the randomness above - try re-running the same query in incognito mode a few times and you'll see the answer change too. If it keeps happening across every scan, email hello@branchtracker.ai with the query and we'll look into it.