Ghostart Labs

Original research, published whichever way it lands

AI tools now answer questions people used to ask Google — including “who should I use?” questions about businesses like yours. This industry is full of confident claims about how that works. Almost none of them have been tested. Labs is where we test them — including the ones our own product depends on.

The rules are simple. Before we collect any data, we publish the question we're asking and exactly how we'll answer it — so the numbers can't rewrite the question. Every result is published, including the ones that prove us wrong. And the raw data comes with every study, so you can check our work.

No published results yet

The first studies are registered — what each one is testing is listed below, locked before any data was collected, along with where it has got to. Results appear here the moment they're done, whatever they say.

What we're testing

The question each of these asks was locked before any data was collected — that's the point — and the results will appear here whatever they say. Each one says below whether it is currently collecting.

Question locked, starting soonQuestion locked 10 September 2026

Does substance predict citations? An observational cohort study

Analysis — we study patterns in anonymised data; no business or person is ever identifiable

What we're testing: Customer content with higher Beige-ometer substance and citability scores is associated with higher observed AI-answer citation rates than lower-scoring content, controlling for engine and topic coverage.

Question locked, starting soon

Does the content we suggest actually get cited? An observational study of Ghostart-suggested publishing

Analysis — we study patterns in anonymised data; no business or person is ever identifiable

What we're testing: Target answers that receive a Ghostart-suggested, published piece transition from absent to cited at a higher rate, and sooner, than target answers tracked for the same business over the same period that received no piece. Secondary: the published artefact itself appears in the engine's cited sources, at a rate varying by channel and engine.

Collecting data now

The Ghostart SME Visibility Panel — Edition 1

Analysis — we study patterns in anonymised data; no business or person is ever identifiable

What we're testing: Preregistered descriptive cuts of AI answer-engine visibility across 240 UK SME businesses in 9 sectors, per the cuts bank signed off 9 Aug 2026 (#2090). No directional hypothesis: Edition 1 reports rates with N and confidence intervals, nulls included.

The AI Visibility Index

Alongside the one-off studies, Labs runs a standing measurement programme. We've assembled a panel of 240 real UK small businesses — accountants, builders, physios, restaurants, shops, agencies and more, across 9 sectors and 18 cities — every one found in a public place, with the source recorded. Each month, our systems put the questions those businesses' customers actually ask (“who are the best accountants in Leeds?”) to the AI engines people actually use, and record who gets mentioned. Every question is asked three times, because AI answers change between asks.

Nobody can currently say what share of UK small businesses appear in AI answers at all — not Google, not the AI companies, not the SEO industry. Everyone measuring AI visibility today measures big brands. The Index measures the accountant in Leeds and the physio in Cardiff — and because one reading is weather but the same stations read every month are climate, the record only gets more valuable as it grows. Collection began in August 2026 and runs in monthly cycles; a cycle takes about four weeks of continuous sampling to complete.

  • The analyses are preregistered. Before anyone looked at a single aggregate, we wrote down exactly which cuts will publish and which won't — the discipline medical trials use against cherry-picking, dated in our method note.
  • Rates, never ranks. We will never say a business “ranks #3 in ChatGPT” — that number changes between asks. We report how often businesses appear, with sample sizes shown.
  • We name our bias up front. Panel businesses came from public directories and lists — exactly what AI tools read — so our numbers likely read rosier than the true national picture. The Index describes publicly listed UK small businesses, never all UK businesses.
  • No panel business is ever named, or contacted. Published output is aggregate and sector-level only, and we observe nothing a customer couldn't: public AI answers to normal customer questions.

The questions the first edition answers

Nobody can answer any of these about UK small businesses today. That is the whole reason to collect.

  • Which sectors are least visible? How often a sector's businesses get named at all when a customer asks who to use — broken down by sector and by size of business.
  • What gets recommended instead? When no real local firm is named, what fills the answer: directories, national chains, other independents, or press coverage.
  • Do the AI tools even agree? ChatGPT, Gemini, Perplexity and Google's AI Overviews, answering the same question on the same day, side by side.
  • How much does the answer flicker? Ask the same question three times and watch how often the businesses named change — which is precisely why we report rates and not rankings.
  • Discovery or validation? Whether businesses surface more when someone asks “who's the best?” or when they ask “is this firm any good?”

Which of these publish, and which don't, was written down before anyone looked at an aggregate. The question set itself is locked and versioned — reworded questions become a new version and are never compared against the old one — and it publishes in full alongside the first edition, with the cycles it produced.

What we won't publish

  • Any panel business, named. Not in an edition, not in a chart, not in a reply to a comment. They never agreed to be studied, so they are never identifiable.
  • Any rank. Not “#3 in ChatGPT”, not “position 2”. We measure where a business appears in an answer, and we keep that internal, because publishing it would be a ranking by another name.
  • Any single-sample claim, or any edition built on fewer than two complete collection cycles.
  • City-level breakdowns, at this width. Around thirteen businesses per city is too thin to defend, so cities wait until the panel is wider rather than being published thin.
  • Anything early. No teasers, no “early data suggests”. One leaked figure would undo the point of writing the analyses down in advance.

The full method publishes at the Beyond the Beige summit on 30 September 2026 — before any results, so it can be challenged while it can still be improved.

Edition 1 publishes when two complete collection cycles are in, and not before. We are deliberately not giving it a date: a date would be a promise about our schedule, and the only promise worth making here is about the evidence. If that means it lands later than we would like, it lands later.

The three kinds of study we run

  • Surveys. A research company asks real business owners a set of questions on our behalf. We never contact anyone ourselves — respondents come from the research company's own opted-in pool.
  • Experiments. We publish carefully matched pieces of content — identical except for the one thing we're testing — then track which version the AI tools actually cite, week after week.
  • Analyses. We look for patterns in data we already hold, always anonymised and aggregated. No business or person is ever identifiable in anything we publish.

The rules every study follows

  • The question is locked first. We publish what we're testing and how, before collecting any data — so the numbers can't rewrite the question.
  • Rates, never ranks. Ask an AI the same question three times and you can get three different answers. So we never trust a single sample — every number is measured across repeats, and we show how many.
  • Every result publishes. A study that finds nothing — or finds against us — goes up with the same prominence as one that flatters us.
  • Check our work. Methods are stated in enough detail to re-run them, and the raw data is downloadable from every study.

How we measure the product side is documented separately on How we measure.