Search & AI · 7 min read

How to Measure Whether LLMs Recommend Your Brand

A favourable answer from one LLM is one observation, not a trend. Learn how to measure visibility, evidence, accuracy and consistency across providers.

Joe HandayaJoe HandayaChief Product Officer
Repeated LLM recommendation results measured across four test runs

A favourable answer from an LLM is not a result. It is one observation.

If your brand appears in a recommendation, it is tempting to save the screenshot and call the work finished. Ask again and it may become a passing mention, disappear or show different sources.

OpenAI's evaluation guidance describes model behaviour as nondeterministic and recommends repeated, logged evaluation. Marketers therefore need a fixed test, clear answer labels and enough repetition to see a pattern.

This is how to measure whether LLMs recommend your brand without turning one favourable response into a claim you cannot defend.

One good answer proves very little

LLM platforms do not answer under identical conditions. Responses may use model knowledge, web retrieval or both. OpenAI's ChatGPT Search guide, for example, says ChatGPT may rewrite a request into several searches and attach citations or a Sources panel.

Account context can also affect results. On ChatGPT, memories, history and custom instructions may shape an answer. OpenAI says a Temporary Chat does not use or create memories. Use equivalent clean-session controls on each platform.

A SparkToro recommendation experiment collected 2,961 runs of 12 prompts from 600 volunteers. The same recommendation list appeared in less than one percent of runs. The work was not peer reviewed and did not control every account setting, but it shows why one answer is a weak baseline.

Measure four separate signals

Do not force every result into one ladder. A brand can be recommended without being cited. A brand can also be cited as a source without being recommended. Record four separate signals for every answer.

Four separate signals for measuring LLM brand visibility: visibility, evidence, accuracy, and consistency

Visibility tells you whether the brand appears in the answer.

  • Absent means no recognised brand name or alias appears.

  • Mentioned only means the brand appears but is not presented as suitable for the stated need.

  • Recommended means the brand is shortlisted, favoured or described as a good fit for that need.

Evidence tells you what sources are visible.

  • No citation means the response shows no supporting link for the relevant claim.

  • Owned citation means a page on your domain appears as a cited source.

  • Third-party citation means an external page supports the answer or recommendation.

Accuracy tells you whether the description is correct, incomplete or wrong. Check the product category, audience, features, availability and other facts that would affect a buyer's decision.

Consistency tells you how often the same result returns across repeated runs and over time. Track repeated presence and recommendation, not just where the brand appears in a list.

Keep negative framing and factual errors as separate flags. A mention is not useful if it attaches the wrong category or an outdated claim to your name.

Start with real buyer questions

Your prompt set should reflect what a buyer asks before they know which brand to choose. Start with a defined audience, not a list of phrases created inside the marketing team. StoryMint's guide to audience personas shows why broad demographic labels rarely capture the decision, pressure or objection behind a question.

Build a fixed core set across four types of intent.

  • Category prompts ask which products or providers fit a recognised category.

  • Problem prompts describe the job the buyer needs to complete without naming a product.

  • Comparison prompts ask what to compare or which options suit different needs.

  • Constrained use-case prompts add real conditions such as team size, market, language, budget sensitivity or required integrations.

Add a separate set of branded questions such as "What does [brand] do?" and "Who is [brand] for?" Use these to test factual accuracy. Do not include them in your discovery or recommendation rate because the prompt has already supplied the brand name.

Map the questions across discovery, consideration and decision stages. The customer journey mapping guide can help you avoid testing only the final purchase question.

Run the same test properly

Repeatable LLM test protocol with fixed conditions and four recorded runs
  1. Treat each provider as a separate test series. Do not mix ChatGPT, Gemini, Perplexity and Google AI Mode responses into one result before reviewing each platform.

  2. Start each run in a fresh conversation. Use the platform's available controls to reduce memory, history or personal context.

  3. Turn on web search or the equivalent retrieval mode when you are testing current web-backed answers. Do not mix retrieval-enabled and model-only answers in the same comparison.

  4. Keep the core prompt wording unchanged across providers. Put experimental variations in a separate set so they do not distort the trend.

  5. Record the provider, date, model or mode, account condition, market or location, prompt text and retrieval setting.

  6. Repeat each prompt across separate conversations. Use the same number of runs in each reporting period. Treat a small sample as directional, not statistically conclusive.

  7. Save the full answer, not only the sentence containing your brand. Capture competitors, reasons, citations, source domains, errors and negative framing.

Your log should make it possible to answer a basic question: did the brand become more consistently visible for the same buyer need under comparable conditions?

Read the pattern, not the position

LLM platforms do not provide one stable search ranking for brands. Recording whether you appeared first can be useful context, but it should not become the headline metric.

Track a compact set of rates and review the underlying answers alongside them.

  • Presence rate measures how often the brand appears in eligible discovery answers.

  • Recommendation rate measures how often the brand is explicitly presented as a suitable choice.

  • Owned citation rate measures how often your own pages appear as visible sources.

  • Third-party citation rate measures how often independent pages support the response.

  • Accuracy records how often important brand facts are correct, incomplete or wrong.

  • Consistency compares these results across the fixed prompt set and reporting periods for each provider.

Do not combine providers or signals into one score too early. A rising citation rate with no recommendation tells a different story from a rising recommendation rate supported by inaccurate facts. The labels tell you what happened. Human review tells you whether it matters.

Fix the failure you can see

Your brand stays absent

First check whether the affected provider can retrieve the site. For ChatGPT Search, OpenAI's publisher guidance says sites should allow OAI-SearchBot. Other providers have their own access rules, so check the current documentation for each one. Access only makes inclusion possible. It does not ensure that a page will be selected or that a brand will be recommended.

Then inspect whether your site states the category, use cases, audience and differentiators in plain language. Look beyond your own domain as well. Independent reviews, relevant coverage and accurate business profiles help establish how other sources describe the brand.

You appear without a recommendation

Read the reasons attached to the brands that were recommended. The gap may be product fit, proof, positioning or a missing answer to the buyer's constraint. Turn recurring missing questions into a content decision, not a burst of near-duplicate articles. StoryMint's content gap analysis workflow provides a practical route from a real audience question to a useful brief.

You are cited but not recommended

Your content is visible enough to be used as a source, but the answer does not present the brand as the right choice. More crawl access is unlikely to solve that by itself. Review the offer, evidence, buyer fit and independent support behind the competitors that were chosen.

You are recommended without citations

Record the recommendation, but do not assume you know what caused it. Check whether the description is accurate and whether the answer changes when Search is enabled. The useful signal is repeated recommendation for the same need, not a theory about an invisible source.

The answer is wrong

Correct the source of truth first. Make product, pricing, audience, feature and availability information easy to find and consistent across important pages. Clear brand and product stories help readers and AI systems separate what the company stands for from what the product actually does.

Put the method into StoryMint

StoryMint's AI brand visibility tools bring the measurement pieces into one campaign. LLM Visibility tracks appearances across ChatGPT, Gemini, Perplexity and Google AI Mode. AI Citation Audit shows which prompts cite the brand and which do not. Query Fan-Out helps reveal variations of the questions buyers ask.

Keep the dimensions separate when you review the results. Use the audience persona to decide which prompts matter, the visibility record to find the recurring gap and the content workflow to decide what deserves to be written. Human review remains necessary because a citation can be irrelevant, a mention can be inaccurate and a recommendation can be unstable.

Write down the questions a real buyer asks, run them under comparable conditions on each provider and let the repeated pattern choose your next content task. That is the difference between checking one AI answer and measuring LLM visibility.

Ready to put this into practice?

StoryMint is the audience-first content marketing suite. Plan, write and publish content your readers actually come back for.

Start your free trial

Keep reading