Table of Contents
- Intro
- When You Need a Vendor-Mention Baseline
- How to Build a Vendor-Mention Baseline in 10 Steps
- Step 1 — Define the category and its comparison prompts
- Step 2 — Lock the competitor cohort
- Step 3 — Run the prompt set across engines
- Step 4 — Record shortlist inclusion, not just mentions
- Step 5 — Capture competitor co-appearance
- Step 6 — Separate mentions from citations
- Step 7 — Identify the sources the engines cite
- Step 8 — Score and store the baseline
- Step 9 — Segment the scores by prompt tier and constraint
- Step 10 — Set the refresh cadence and change log
- Common Mistakes That Distort a Vendor-Mention Baseline
- Vendor-Mention Baseline: Steps, Metrics, and Expected Outcomes
- Conclusion: From Mention Counts to Shortlist Inclusion
- Frequently Asked Questions
- What is AI visibility for B2B SaaS?
- How is shortlist inclusion different from a brand mention?
- Which AI engines should B2B software teams track?
- How often should a vendor-mention baseline be re-run?
- Do G2 and Capterra citations drive AI software recommendations?
- How long does it take to build a baseline?
- Build Your B2B Software Vendor-Mention Baseline with Alef
- Sources
Intro
Gartner's survey of 645 B2B buyers found that 45% used generative AI to gather information on vendors and products, while 69% still turned to sales reps to validate what those tools told them (Gartner, May 2026). That split defines the measurement problem behind ai visibility for b2b saas: buyers arrive already primed by an engine's shortlist, then ask a rep to confirm it. This guide delivers a repeatable vendor-mention baseline for one software category — a fixed prompt set, run across engines, scored on shortlist inclusion and competitor co-appearance. The metric definitions behind AI visibility measurement live in the pillar guide; this spoke supplies the B2B operating procedure. Alef tracks mentions, citations, rankings, and competitor share across ChatGPT, Perplexity, Gemini, and Copilot from one workspace, mirroring the steps that follow. The gap is widening: Search Engine Journal reported ChatGPT-User made roughly 3.6x more requests than Googlebot over a 55-day window.
When You Need a Vendor-Mention Baseline
The trigger is rarely a dashboard alert. It is a forwarded screenshot: a prospect pastes a ChatGPT answer that names three competitors and stops there. Or a sales team reports that discovery calls now open with "the AI recommended X." Or a board member asks a deceptively simple question — where does the brand stand inside AI answers?
The stakes are measurable. Forrester found that 95% of B2B buyers plan to use generative AI in at least one area of a future purchase, and more than half said it led them to consider more or different vendors (Forrester). Meanwhile, Pew Research Center recorded a click-through collapse: users clicked a traditional result in only 8% of visits when an AI Overview appeared, versus 15% without one (Pew Research Center). The answer itself is now the decision surface.
Prerequisites are modest:
- An indexable product and category page set
- One defined category name
- A locked competitor list of three to five named vendors
Building a first baseline takes roughly three to four hours at beginner skill level, with no specialized tooling. Before that, though, a diagnostic matters: if the brand is invisible in AI answers despite real traffic, the problem is crawlability and answer readiness, not measurement. Pre-launch products with no indexable content should fix those foundations first, then use prompt intelligence to see which buyer questions the category already triggers.
How to Build a Vendor-Mention Baseline in 10 Steps
A vendor-mention baseline is a dated, reproducible snapshot of how AI answer engines describe a software category and which vendors they name inside it. Building one for a B2B software category takes roughly six to ten hours of focused work the first time, drops to two to three hours on a monthly refresh, and requires no engineering support beyond a spreadsheet and a shared login to the engines being tested. The prerequisite is narrow but firm: the team must already agree on the category label buyers use, because every downstream number inherits that definition. What follows is the full sequence, from prompt construction to a scored baseline that can be compared month over month.
The ten steps below assume the reader has read the metric definitions in the pillar guide. Where a term such as shortlist inclusion rate or share of voice needs a formal definition, that reference holds it.
Step 1 — Define the category and its comparison prompts
Write the exact category label buyers use, then generate 20 to 40 prompts across discovery, comparison, and buying intent.
The category label is the phrase a buyer types when they have budget but no vendor in mind. It is rarely the phrase a product marketer uses on the homepage. "Best enterprise CRM for mid-market retail" is a buyer label. "Unified revenue intelligence platform" is a positioning statement, and AI engines will not match it to a buyer's question.
Three prompt tiers should be represented, weighted roughly 40/40/20:
- Discovery prompts surface the category itself: "What are the leading warehouse management systems for third-party logistics providers?" These test whether the brand exists in the engine's category map at all.
- Comparison prompts force a head-to-head: "How does Vendor A compare to Vendor B for compliance-heavy manufacturing?" These are where shortlists form and where co-appearance data becomes richest.
- Buying-intent prompts carry constraints: "Which applicant tracking system works best for a 500-person healthcare company with EU hiring?" These produce the tightest three-to-five vendor lists and the highest commercial signal.
Long-form, specific prompts outperform short ones because they mirror how buyers actually phrase questions to an assistant. A prompt with a company size, a vertical, and a regulatory constraint will pull a narrower and more decision-relevant vendor set than "best CRM." Twenty prompts is the floor for a single category; forty is the practical ceiling before marginal returns flatten and maintenance cost rises.
Step 2 — Lock the competitor cohort
Name three to five direct vendors that appear in the same buyer consideration set, and freeze the list so the baseline stays comparable over time.
The cohort is not every company in the market. It is the set of vendors that a real buyer would evaluate against the brand in a single deal. Sales teams usually know this list better than marketing does, and a short conversation with two or three account executives will produce it faster than any analyst exercise.
Freeze the cohort for a minimum of two quarters. Every metric in the baseline — share of voice, co-appearance rate, displacement — is computed relative to this set, so adding or removing a vendor mid-stream invalidates the trend line. If a new entrant genuinely enters the consideration set, note the date of the change and start a new cohort version rather than editing the old one. A cohort of three to five keeps the co-appearance matrix readable; beyond six vendors, the matrix becomes a wall of cells with too few observations per cell to interpret.
Step 3 — Run the prompt set across engines
Execute every prompt in ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews; log the date, engine, and model version for each run.
Five engines multiplied by thirty prompts yields 150 runs per cycle. Each run needs four fields logged before the answer is even read: date, engine, model version or tier, and whether the session was logged in or anonymous. Model version matters because answers shift between releases, and a baseline without version stamps cannot distinguish a genuine visibility change from a model upgrade.
Two operational details determine whether the data holds up:
- Use a clean session for each prompt. Prior conversation turns contaminate answers. Opening a fresh thread per prompt prevents the engine from carrying context forward and biasing the vendor list.
- Capture the raw answer, not a summary. Screenshots or full-text exports preserve the exact vendor ordering and phrasing, which matters in Step 4 when accuracy of description is scored.
Google AI Overviews behaves differently from the assistant engines: results vary by query phrasing, location, and personalization signals, so it is worth running each prompt twice from separate sessions and recording both outputs. Where the two runs disagree, the prompt is flagged as unstable and weighted accordingly in scoring.
Step 4 — Record shortlist inclusion, not just mentions
For each answer, note whether the brand appears in the named vendor list, at what position, and whether it is described accurately.
This is the step that separates a shortlist baseline from a generic mention tracker. A brand can be mentioned in an answer and still be absent from the shortlist — for example, named in a historical aside ("Vendor X was an early entrant, though most teams now consider…") while the actual recommendation list contains four other vendors. That is a mention without inclusion, and it converts poorly.
Three fields per answer:
- Inclusion — binary. Did the brand appear in the set of vendors the engine presented as options?
- Position — the ordinal slot. First-named vendors carry disproportionate weight in how buyers read a list.
- Description accuracy — does the engine's one-line characterization match the brand's actual positioning, pricing model, and target segment? An inaccurate description is a distinct problem from absence and requires a different fix.
Inclusion rate, not mention count, is the number that maps to pipeline. Forrester's analysis of B2B buying behavior notes that buyers increasingly complete evaluation without visiting vendor sites at all, which means the answer itself is often the entire first impression a vendor gets (Forrester: B2B Buyers Make Zero-Click Buying Number One). A brand that is mentioned but not shortlisted has effectively not been seen.
Step 5 — Capture competitor co-appearance
Log which competitors appear alongside the brand and which appear when the brand does not — the co-appearance matrix is the diagnostic core of the baseline.
The co-appearance matrix is a grid with prompts as rows and cohort vendors as columns, with a mark in each cell where that vendor appeared in that prompt's answer. Reading the matrix reveals patterns that no single-prompt review surfaces:
- Constant companions. Vendors that appear in nearly every answer the brand appears in. These are the true head-to-head competitors in the engine's model of the category, regardless of how sales teams rank them.
- Exclusive occupants. Vendors that appear in prompts where the brand is absent. These indicate prompt clusters the brand has not entered at all — often a vertical, a company-size band, or a compliance requirement the brand's public content does not address.
- Substitution pairs. Prompts where the brand and a specific competitor alternate across engines or across runs, suggesting the engine treats them as interchangeable options.
The matrix also exposes category fragmentation. If no vendor appears in more than 40 percent of prompts, the engine has no stable category map, and the brand's problem is category definition rather than competitive displacement. That distinction changes the entire remediation plan, which is why the matrix is built before any scoring.
Step 6 — Separate mentions from citations
Record whether the answer links to the brand's domain, because Semrush's ghost-citations study found 61.7 percent of AI brand appearances were citations that never named the brand, and only 13.2 percent included both.
The distinction between being named and being linked is not academic. A citation — a URL pulled into the answer's source list — can occur without the brand name ever appearing in the answer text. Conversely, a brand can be named in the prose with no link back to its domain. Only a minority of appearances do both.
For a B2B software vendor, the two outcomes have different consequences. A named mention influences the buyer's shortlist. A citation without a name builds domain authority signals and can drive AI-referred traffic, but it does nothing for consideration at the moment of the question. Both are worth tracking, and they should be tracked in separate columns rather than collapsed into a single "visibility" field.
The practical implication: a baseline that reports only mention counts will overstate visibility for brands that are frequently cited but rarely named, and understate it for brands that are named but never linked. The mechanics of isolating each signal per engine are covered in the guide to tracking brand mentions in ChatGPT and Perplexity, which walks through the logging fields that keep the two categories clean.
Step 7 — Identify the sources the engines cite
Extract every cited URL per prompt and tally which domains recur, since these third-party pages are what drive inclusion or exclusion.
Every answer that includes citations carries a source list, and that list is the causal layer beneath the shortlist. If a review site, a comparison page, or an analyst write-up appears in the source list of most prompts where a competitor is shortlisted and the brand is not, that page is doing the competitive work.
Extraction is mechanical: for each of the 150 runs, copy every cited URL into a flat table with columns for prompt, engine, and cited domain. Then pivot. The output is a ranked list of domains by citation frequency, split into three buckets:
- Domains that cite the brand. These are the assets already earning inclusion and should be protected and expanded.
- Domains that cite competitors but not the brand. These are the highest-priority targets, because the engine has already decided the domain is authoritative for the category.
- Domains that cite no cohort vendor. These are neutral category sources and represent lower-leverage but lower-competition opportunities.
Domain-level tallies matter more than individual URLs. An engine citing six different pages from the same review platform is signaling that the platform, not the page, is the trusted source. The remediation target is the platform relationship, not a single article.
Step 8 — Score and store the baseline
Compute shortlist inclusion rate (prompts where the brand appears divided by total prompts) and share of voice (brand appearances divided by all vendor appearances) per engine.
Two scores, calculated separately for each engine because engine behavior diverges sharply:
| Metric | Formula | Worked example | What it tells you |
|---|---|---|---|
| Shortlist inclusion rate | Brand appearances ÷ total prompts | 12 ÷ 30 = 40.0% | How often the brand makes the recommendation set |
| Share of voice | Brand appearances ÷ all cohort vendor appearances | 12 ÷ 84 = 14.3% | Relative prominence against the frozen cohort |
| Citation rate | Prompts citing brand domain ÷ total prompts | 5 ÷ 30 = 16.7% | Domain-level presence independent of naming |
| Named-and-cited rate | Prompts doing both ÷ total prompts | 2 ÷ 30 = 6.7% | The highest-value appearance type |
| Co-appearance concentration | Most frequent companion vendor's appearances ÷ brand appearances | 9 ÷ 12 = 75.0% | How tightly the brand is paired with one rival |
A single blended score across all five engines hides the most actionable finding in the dataset: a brand can post a 60 percent inclusion rate on Perplexity and 15 percent on Copilot, and the gap usually traces to which sources each engine weights. Storing per-engine scores also makes month-over-month comparison honest, since engine updates land at different times.
Store the baseline as a dated file with the cohort version, the prompt set version, and the model versions recorded in Step 3. The comparison methodology for reading those trend lines against competitors is laid out in the guide to benchmarking AI visibility against competitors, which covers how to normalize scores when a cohort changes or an engine shifts its source weighting.
Step 9 — Segment the scores by prompt tier and constraint
Break the aggregate scores apart by discovery, comparison, and buying-intent prompts, and by the constraints embedded in each prompt.
Aggregate inclusion rate is a headline, not a diagnosis. A brand at 40 percent overall might sit at 70 percent on discovery prompts and 10 percent on buying-intent prompts — a pattern that means the engine knows the brand exists but does not consider it a fit for constrained, late-stage questions. That is a content and proof problem, not an awareness problem.
Segmenting by constraint reveals which buyer attributes the engines associate with the brand. If inclusion collapses specifically on prompts mentioning regulated industries, EU data residency, or sub-200-employee companies, the engine's model of the brand has a boundary the brand may not have intended. Gartner's research on AI-generated buying insights found that 69 percent of B2B buyers turn to sales reps to validate what AI tools tell them (Gartner: 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights), which means an engine's exclusion on a specific constraint gets tested by a rep in a live deal — and the brand never learns it happened unless the baseline segmented for it.
Three segmentation cuts are worth maintaining: prompt tier, embedded constraint type, and company-size band. Each produces a different remediation list.
Step 10 — Set the refresh cadence and change log
Re-run the full prompt set monthly, log every change to the cohort or prompt set, and treat any movement under five percentage points as noise until confirmed across two cycles.
AI answers are non-deterministic. The same prompt run twice can produce different vendor lists, and a single-cycle swing of three or four percentage points on a thirty-prompt set is well within normal variance. Two consecutive cycles moving in the same direction is the threshold for treating a change as real.
The change log should record four event types, each with a date: cohort changes, prompt set additions or removals, engine model version updates, and any material shift in the cited-source list. Without this log, a 15-point inclusion jump looks like a content win when it may simply be a model update that changed how the engine weights review sites.
Monthly is the right default cadence for most B2B software categories. Quarterly is defensible for narrow categories with low competitive churn, and weekly is warranted only during an active product launch or a category redefinition push. The point of the cadence is not frequency for its own sake — it is having a comparable prior data point ready when someone asks whether the last quarter of content work moved the number.
Once the ten steps are complete, the baseline is a working instrument rather than a report. It answers three questions on demand: does the brand make the shortlist, who occupies the slots when it does not, and which third-party pages decide the outcome.
Common Mistakes That Distort a Vendor-Mention Baseline
A baseline is only as trustworthy as the method behind it. Seven errors account for most corrupted vendor-mention datasets in B2B software categories.
- Counting mentions instead of shortlist inclusion. A brand can be named in passing and still be absent from the three-to-five vendors an engine actually recommends, and only that placement influences a deal.
- Changing the prompt set mid-tracking. Altering even one prompt between runs breaks comparability, which makes every subsequent trend number meaningless.
- Treating a single run as the baseline. Answers vary between runs, so repeated sampling is required before any conclusion is drawn.
- Ignoring ghost citations. Assuming a citation equals visibility is unsafe: Semrush found most AI appearances never name the brand, a pattern covered in AI citation tracking for mentions without links.
- Assuming review aggregators dominate. B2B software citation studies conflict sharply — one 500-citation study attributed 42% of citations to review aggregators, while a 40-category ChatGPT study found them cited in under 1% of cases — so the cited-source mix must be measured, not assumed.
- Measuring only one engine. Gemini mentions brands far more often than it cites them, while ChatGPT cites far more often than it mentions, so single-engine tracking produces a distorted picture.
- Never linking the baseline to pipeline. Without connecting inclusion rate to AI-referred traffic and conversion, the number stays a vanity metric.
The common remedy is discipline: freeze prompts, sample repeatedly, log citations and mentions separately, and track every engine in one workspace.
Vendor-Mention Baseline: Steps, Metrics, and Expected Outcomes
The ten steps below produce a single artifact: a dated vendor-mention baseline that records not just whether a B2B software brand appears in AI answers, but whether it appears inside the shortlist. Each step has a defined metric and a verifiable outcome, so the baseline can be audited and repeated rather than treated as a one-off snapshot.
| Step | Action | Metric Recorded | Expected Outcome |
|---|---|---|---|
| 1 | Define category and comparison prompts | 20–40 prompts across three intent types (category, comparison, alternatives) | A frozen prompt set covering the full buying cycle |
| 2 | Assign prompt owners and intent tags | 3 intent tags mapped to 20–40 prompts | Every prompt traceable to a deal stage |
| 3 | Run prompts across engines | 5 engines × 20–40 prompts = 100–200 logged answers | Raw answer corpus with timestamps and engine labels |
| 4 | Record shortlist inclusion | Shortlist inclusion rate, expressed as a percentage of prompts | A single headline number for the category |
| 5 | Capture co-occurring competitors | Competitor co-occurrence count per prompt | A ranked list of vendors named alongside the brand |
| 6 | Extract cited sources per answer | Cited domain list per engine | The third-party pages driving inclusion |
| 7 | Classify sources by type | Source type distribution (review site, analyst, community, owned) | A prioritized target list of citable pages |
| 8 | Calculate share of voice per engine | Share of voice per engine, as a percentage | Engine-level visibility split |
| 9 | Score source influence on inclusion | Inclusion correlation per cited domain | The pages worth pursuing first |
| 10 | Set a diff cadence | Weekly or monthly baseline diff | A repeatable trend line, not a snapshot |
Engine behavior diverges sharply, which is why step 8 matters: Semrush data shows Gemini mentions brands in 83.7% of appearances but cites only 21.4%, while ChatGPT cites 87% but mentions only 20.7% (Semrush). A mention-heavy engine and a citation-heavy engine require different content strategies, and a baseline that averages them hides the difference. For teams deciding how to present these numbers to stakeholders, the guide to AI visibility reporting questions for a B2B dashboard covers which metrics belong in a recurring report.
Conclusion: From Mention Counts to Shortlist Inclusion
For B2B software, shortlist inclusion — not raw mention volume — is the AI visibility metric that maps to pipeline. The operating loop stays fixed: a locked prompt set, a stable competitor cohort, multi-engine runs, shortlist and citation logging, cited-source prioritization, and re-runs on a fixed cadence.
Two findings most change how a baseline reads. Engine splits mean a single aggregate score hides where inclusion is won or lost. Ghost citations mean a brand can be named without being cited — or cited without being named.
Metric definitions live in the pillar guide; continuous tracking across ChatGPT, Perplexity, Gemini, and Copilot runs through Alef's AI visibility solution.
Key takeaways - Shortlist inclusion predicts pipeline; mention count alone does not. - Keep prompts, competitors, and cadence fixed so runs stay comparable. - Read results per engine — aggregate scores mask real gaps. - Check ghost citations before trusting any single-source answer. - Prioritize the third-party pages engines actually cite.
Frequently Asked Questions
What is AI visibility for B2B SaaS?
AI visibility for B2B SaaS is the measurable share of answer-engine responses in which a vendor is named, cited, or shortlisted for a defined set of buyer prompts. It is not a sentiment score or a traffic estimate; it is a ratio. If a software company runs 40 category and comparison prompts across four engines and appears in 22 of the 160 resulting answers, its baseline AI visibility sits at roughly 14 percent. The metric only becomes actionable once the prompt set is fixed, because a shifting prompt inventory makes period-over-period comparison meaningless.
How is shortlist inclusion different from a brand mention?
A mention is the brand name appearing anywhere in an answer; shortlist inclusion is the brand appearing inside the recommended vendor list that the engine presents as the answer. The distinction matters because buyers act on the shortlist, not on incidental references. An engine may name a product while explaining why it was excluded, which registers as a mention and as a negative outcome simultaneously. Separating the two counts prevents inflated visibility reporting and keeps the focus on the responses that shape vendor selection.
Which AI engines should B2B software teams track?
ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews should all be tracked, because mention and citation behavior differs sharply between them. An engine that leans on live web retrieval will cite different sources than one drawing primarily on training data, and the same prompt can produce a shortlist in one system and a definitional answer in another. Running an identical prompt set across all five surfaces is the only way to see where a brand is structurally absent versus merely under-cited. Teams that want to improve citation rates per engine can start with how to get cited by ChatGPT, which covers the retrieval and formatting factors that drive inclusion.
How often should a vendor-mention baseline be re-run?
Weekly re-runs suit fast-moving categories, while monthly is sufficient for stable ones, and the prompt set must remain identical across every cycle. Answer engines update their retrieval indexes and model weights on different schedules, so a baseline captured in January may not reflect March behavior. Holding prompts constant isolates genuine visibility movement from measurement noise. When a category shifts — a major acquisition, a new entrant, a pricing change — an off-cycle run is worth the effort.
Do G2 and Capterra citations drive AI software recommendations?
The evidence conflicts, so the only reliable answer is to measure the cited-source mix in your own category. One study found review aggregators accounting for 42 percent of citations, while another found them below 1 percent, a gap wide enough to suggest the result depends heavily on prompt type and engine. Rather than adopting either figure, teams should log which domains each engine cites for their prompts and rank those domains by how often they co-occur with shortlist inclusion. Alef's citation tracking records the sources behind each answer so that this mix is visible per prompt rather than inferred from aggregate reports.
How long does it take to build a baseline?
Roughly three to four hours for the first pass at beginner skill level, with no specialized tooling required. Most of that time goes into prompt construction and manual logging, not analysis. A defensible first pass includes 30 to 50 prompts, four engines, and a spreadsheet capturing brand presence, competitor presence, and cited domains. Teams that want to shorten the cycle should review how to optimize content for AI search engines before drafting prompts, since prompt phrasing determines which sources an engine retrieves. Note that buyers still verify AI-generated shortlists with human sources; Gartner found that 69 percent of B2B buyers turn to sales reps to validate AI-generated insights, which makes shortlist presence a qualification step rather than a closed deal.
Build Your B2B Software Vendor-Mention Baseline with Alef
A manual baseline is a starting point, not a system. Alef's AI visibility engine converts it into continuous monitoring: mention tracking, citation capture, ranking movement, and competitor share across ChatGPT, Perplexity, Gemini, and Copilot from one workspace. For teams weighing how much weight those answers now carry, Gartner reports that 69% of B2B buyers turn to sales reps to validate AI-generated insights — which makes shortlist presence the metric worth defending. Running a first visibility check on the Alef platform reveals whether the category prompts that matter already name the brand.
Sources
- Gartner: 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights
- Forrester: From Keywords to Context — AI-Powered Search in B2B Marketing
- Forrester: B2B Buyers Make Zero-Click Buying Number One
- Pew Research Center: Google Users Are Less Likely to Click on Links When an AI Summary Appears
- Semrush: Why 62% of AI Citations Don't Lead to Brand Mentions (Ghost Citations Study)
- Search Engine Journal: ChatGPT-User vs Googlebot Crawl Data
- Timothy Prestianni: The State of AI Visibility — How 5 AI Engines Cite Sources (500-Citation Study)
- DerivateX: The B2B SaaS AI Citation Study — How ChatGPT Recommends Software