aeod.app

Notes on AI visibility

AI Comparison Prompts: How X vs Y Queries Pick a Winner

Ask an AI assistant “X or Y?” and the response looks different from “Recommend a plumber.” A comparison prompt hands the assistant two candidates and asks for a verdict. ChatGPT returns a structured answer: what each option does well, then a sentence naming a preferred option. Gemini produces a similar shape. Perplexity attaches citations from comparison articles and review pages next to its pick.

The buyer reads that verdict as a decision. The first name becomes the default. The second becomes the alternative. The assistant skips discovery and ranks two candidates.

That narrow job changes what the buyer takes away. In a recommendation prompt, a name in the top three has value even without winning. In a comparison prompt, the answer arrives as a binary. One option gets the endorsement. The other gets a conditional mention, which reads as second place.

The buyer wrote the prompt, so the field was narrow before the question was typed. The assistant confirms or reverses that pair. aeod.app measures which option receives the endorsement and how stable that endorsement stays across repeated prompts.

A comparison prompt closes the field before the assistant speaks. The answer carries a verdict on two candidates, and that structure is what separates it from a shortlist.

A comparison prompt is a closed test, not a shortlist

A category prompt asks the assistant to build a candidate set. The buyer asks, “What is the best project management tool for a 20-person agency?” The assistant decides how many names to include and what order to present them. Recall is part of the task. The answer becomes a shortlist, and every named business enters the buyer’s consideration set. The businesses left out are absent from the conversation.

A comparison prompt fixes the set before the assistant speaks. The buyer asks, “Asana vs Monday for a 20-person agency?” Two names are fixed. The assistant’s job is adjudication. It must decide which of the two fits the stated need. Recall drops out of the task. Only ranking remains. A third option appears only as a caveat. The response remains structured around the pair. The question has already narrowed the field. This is a closed test. The model has no recall burden. The output does not need a tenth name to satisfy breadth.

That structure changes the output. The answer gets shorter because the model covers two entities instead of ten. The tone gets more confident because a binary choice demands a verdict. The model writes “Asana is better for…” or “Monday wins if…” and then qualifies by segment or price. The buyer receives a decision. The menu was fixed before the prompt. The measurement gets cleaner. Each response maps to one head-to-head outcome: which name is mentioned first and where each sits in the answer. aeod.app records those pair-level outcomes when it audits AI answers. The assistant spends its tokens on comparison criteria. Candidate discovery is already done.

The structural difference matters for diagnosis. A category prompt measures inclusion. A comparison prompt measures preference conditional on inclusion. If a business never appears in category answers, it never reaches the comparison stage. If it appears in both, the comparison reveals rank. The same incumbent advantage shows up in category prompts. That pattern is the subject of why AI assistants recommend the same businesses. Pair prompts make that advantage legible. A pair prompt tests the ranking of known names. A category prompt tests whether the name is known at all. The incumbent starts ahead when specs match.

The incumbent starts ahead when specs match

A June 2026 arXiv study on brand bias in LLM recommendations tested GPT-4o-mini, Claude Sonnet, and Gemini 3 Flash on skincare products. When competing products carried identical specifications, well-known brands took 100% of recommendations. The authors called this pattern a conditional monopoly. The condition is parity. Once the retrievable facts match, the model has no factual basis to separate the options. It falls back on familiarity.

Head-to-head prompts expose that fallback. Ask an assistant to compare two moisturizers with identical active ingredients and concentration, at the same price and size. The model can retrieve both products. It can list the shared attributes. Then it picks one. That pick is preference conditional on inclusion. If the two options are near-equal on retrievable facts, the better-known name wins. The less-known name can be fully described and still lose. The comparison becomes a familiarity test wearing a product comparison costume.

This matters for diagnosis because comparison prompts are where rank becomes visible. A category prompt asks whether a business enters the shortlist. A comparison prompt asks which shortlisted business the assistant favors. The same incumbent advantage shows up in category prompts, as covered in why AI assistants recommend the same businesses. In head-to-head prompts, the advantage persists even when the factual case for both options is identical. aeod.app measures that preference stage by running paired prompts and recording which name appears and which citation supports it.

The study has a category limit. It tested skincare and three models in one set of identical-specification scenarios. The 100% figure does not transfer automatically to industrial equipment, legal services, or B2B software. The mechanism does transfer. When the model cannot distinguish options on retrievable facts, brand familiarity becomes the deciding signal. That signal comes from training data and web presence. Prior mentions in other answers reinforce it. A business with a thinner footprint starts behind before the comparison begins.

The practical implication for comparison prompts is that factual parity does not produce answer parity. Two products can match on every attribute a buyer cares about. The assistant still returns one winner. The next variable is prompt order. Which name you type first changes the answer.

Which name you type first changes the answer

When a buyer types “Brand A vs Brand B,” the assistant reads a sequence. The first name arrives before the second. That sequence shapes the output. Research on ChatGPT and other large language models shows order sensitivity: when the same options are presented in a different order, the model’s choice changes. One set of experiments on multiple-choice prompts found that reordering answer labels shifted ChatGPT’s selections, with a skew toward earlier positions. The model does not treat “A vs B” and “B vs A” as identical questions. It treats them as two prompts with different token histories.

The mechanism is ordinary next-token prediction. The model conditions on everything before the answer. Earlier labels sit earlier in the context window and influence the probability distribution over the next tokens. In pairwise comparisons, that produces a primacy effect: the first brand gets a statistical advantage that has nothing to do with product quality or price. The second brand starts from behind.

The practical consequence is direct. “Brand A vs Brand B” and “Brand B vs Brand A” are two different tests. A single run of one ordering is a sample, not a result. If a business sees itself lose a comparison prompt, the loss can be an artifact of position. Run the same pair in both orders before reading anything into the outcome. If the winner flips when the order flips, the prompt is measuring order sensitivity. If the same brand wins in both orders, the result has survived a basic control.

This matters for how AI visibility gets measured. A test that runs each pair once, in one order, produces noisy data. A test that runs both orders separates a stable recommendation from a positional accident. In aeod.app audits, the same pair run in reverse order sometimes flips the winner, which is why each ordering is treated as a separate sample. The same principle applies to the broader question of why AI assistants recommend the same businesses: repeated mentions across prompts are evidence, while a single mention in a single order is not.

Order is one variable. It is not the only one. After controlling for order, the next question is what the assistant draws on when it picks a winner. That depends on the comparable facts present in the prompt and the third-party comparisons the model has seen elsewhere. Those inputs set the bounds for the verdict. Follow-up questions can also shift the frame across turns, which changes the context the model uses for the final answer. Where the verdict comes from: comparable facts and third-party comparisons.

Where the verdict comes from: comparable facts and third-party comparisons

When a buyer asks “Tool A vs Tool B,” the assistant parses two entities and one requested dimension. If the prompt adds “for small teams,” the dimension is team size. Retrieval then searches for passages that mention both names and the same dimension. A passage about A’s pricing and a separate passage about B’s pricing is weaker than one passage that lists both prices in one table. The model needs aligned facts to produce a verdict.

Self-published comparison pages create a specific problem. The publisher sells one side. The page describes both names, yet the comparison is authored by an interested party. Assistants still retrieve it. The weight shifts when third-party sources discuss both names in the same passage. A spec table with both products on one row is usable. A price band covering both products is usable. A review roundup that scores both on the same criteria is usable. A community thread where owners compare renewal costs for both tools is usable. A vendor page that names a competitor only to list advantages is less usable because the dimension is framed by the seller.

The dimension also shifts across turns; retrieval handles that change differently, as covered in multi-turn AI visibility follow-up questions.

The pattern shows up in sourcing. If one side has no published price, the assistant finds a price for one entity and nothing for the other. There is no second number to hold against it. The same happens when specs sit in a gated PDF or reviews split across different years. The assistant shifts to qualified language: “depends on your needs,” “A is often recommended for teams like yours.” That hedging is a retrieval symptom. The model has enough to identify both names, yet not enough aligned evidence to choose.

This is why comparison queries behave differently from single-business queries. Single-business answers reward any credible passage. Comparison answers reward passages that put both names on the same row. aeod.app tracks which dimensions appear in those paired passages and which dimensions stay one-sided. The gap shows where the answer goes vague.

That asymmetry grows when the query changes from “A vs B” to “alternatives to A.” In the first form both names must be compared. In the second, A is the anchor and every other name is an option. The evidence rules bend around that anchor, and the burden of proof lands on different pages.

The ‘alternatives to’ variant is asymmetric

“Alternatives to X” is a comparison prompt with one fixed name. The assistant treats X as the incumbent. The second slot stays open. Retrieval follows the anchor term. Pages that mention X in the title and body carry more weight, and listicles titled “alternatives to X” match the query pattern. X gets framed by its category and by the attributes the listicles emphasize. Challengers arrive as entries in a list.

The answer structure follows from that frame. A direct “X vs Y” prompt forces the assistant to evaluate both names against the same criteria. It has to resolve conflicts and pick a side. An alternatives prompt lets the assistant treat X as the reference point and arrange the other names around it. A challenger appears in position two and still loses the head-to-head. In the alternatives list, the assistant writes “Y is a lower-cost option” or “Y works for teams that need Z.” In the direct comparison, the assistant weighs Y against X on the buyer’s criteria. If Y’s evidence covers category presence and lacks direct comparison coverage, it shows up as an alternative and disappears from the head-to-head.

That gap is why the two prompts need separate tracking. The two prompts produce different answer structures and carry different position meanings. Their evidence demands diverge. A brand’s share of alternatives mentions measures whether it enters the consideration set. Its share of direct head-to-head wins measures whether the assistant picks it when both names are on the table. Collapsing them into a single comparison metric hides the difference.

2026 advertising analyses reported a brand-protection problem around this split. Campaigns that kept a brand active in category content still failed to surface on direct comparison prompts. The brand appeared in alternatives lists while the direct “A vs B” answer named a different option. The cause sits in the evidence map: listicle inclusion and head-to-head proof come from different pages. This compounds the concentration pattern described in why AI assistants recommend the same businesses.

aeod.app tracks these variants separately because the prompt wording sets the comparison rules. The next layer appears when the assistant avoids a clean winner. That response has its own structure, which the next section takes up under hedges, use-case splits, and refusals to pick.

Hedges, use-case splits, and refusals to pick

A comparison prompt returns one of four verdict types. Clear winner: the assistant names one business and ties the choice to the dimension in the question, such as response time or warranty. Split by use case: the assistant assigns each business a condition, such as one for emergency calls and one for scheduled replacements. Explicit tie: the assistant states both options perform the same on the asked dimension. Refusal to choose: the assistant declines to rank, lists criteria, or asks for more detail.

A hedge is a measurement result rather than a dead end. The retrieval layer returns passages about both businesses. When those passages cover different dimensions, or neither contains a direct comparison on the buyer’s dimension, the assistant has no basis to separate the two options. The hedge reflects the evidence set the assistant retrieved. It is a finding about content coverage and comparison framing.

Consider a prompt that asks which CRM is better for a 12-person sales team. One page describes pricing. Another describes integrations. Neither page compares pricing at 12 seats. The assistant splits by use case or refuses to choose. The verdict is a signal that the retrieved passages do not separate the two options on the dimension the buyer asked about.

Order sensitivity produces a different pattern. Run the same pair twice with the order reversed. A refusal in one ordering and a clear winner in the other points to order sensitivity rather than an evidence gap. If the evidence were missing, both orderings would refuse. If the evidence is present and surfaces only when one name appears first, the verdict changes with prompt order. The assistant’s answer depends on token order and retrieval ranking. The evidence set is sufficient for a winner in one ordering.

Follow-up questions change the evidence set, as covered in multi-turn AI visibility follow-up questions. aeod.app records the verdict type alongside the pair and prompt order. A single run cannot separate a stable verdict from an order artifact. That is why the next step is running a head-to-head audit: pairs, both orders, five runs.

Running a head-to-head audit: pairs, both orders, five runs

Build the pair list from buyer wording. Eight to ten pairs fit one audit cycle. Each pair places the business against a named competitor the way buyers name them. Source the phrasing from sales call notes, chat transcripts, and search queries. A pair reads like “X vs Y for small business payroll” or “is X better than Y for a two-person law firm.” Ten pairs give enough coverage to see whether a competitor wins across the category or only in one niche.

For each pair, run two prompt orders: “X vs Y” and “Y vs X.” Use at least three assistants, such as ChatGPT, Claude, Gemini, or Perplexity. Repeat every prompt five times. That produces 2 orders x 3 assistants x 5 runs = 30 runs per pair. Ten pairs produce 300 runs. Five runs per prompt separate a stable verdict from a one-off answer. Both orders expose whether the first name in the prompt gets the recommendation.

Log five fields per run: verdict type, winner if any, domains cited, hedge language, and assistant name. Verdict types include no winner, single winner, tie, and both named with one favored. Domains cited show which pages the assistant treated as evidence. Hedge language includes phrases like “typically,” “often,” “depends,” and “for most small teams.” Those phrases mark the conditions under which the assistant recommends one business over the other.

aeod.app records mentions, citations, and position per question, which maps to the same fields needed here plus a verdict label per pair. That structure keeps the audit comparable across pairs and assistants. When the same competitor wins in both orders and in four of five runs, the verdict is stable. When the winner flips with the prompt order, the result is an order artifact. When no winner appears, the pair is a tie or the assistant lacks enough evidence.

Patterns across pairs matter more than one result. If one competitor wins most pairs, the shortlist is shaped by source consensus, as covered in why AI assistants recommend the same businesses. If a follow-up question changes which domains get cited, the evidence set shifts, as covered in multi-turn AI visibility follow-up questions. The audit ends with a per-pair label: win, loss, tie, or order artifact. That label determines what to fix when you lose the head-to-head.

What to fix when you lose the head-to-head

A loss label from a pairwise audit is a diagnosis with three common causes. Each cause has a different repair. The repair depends on the dimension that decided the pair: spec, price, reputation, or order.

If the prompt names a spec or price dimension and the competitor wins, the assistant found that fact in the competitor’s indexed content and did not find a comparable fact on your side. The repair is to state the fact plainly where the assistant can retrieve it. Put the number and unit in text. A page that says “same-day service” or “$99 per month” gives the model a value to compare. A page that says “flexible scheduling” or “affordable plans” gives the model nothing to match. Tables help because they keep the attribute next to the value.

If both businesses meet the stated spec and you still lose on reputation, the gap is third-party corroboration. Assistants weigh independent sources: review sites, directories, trade press, and customer case studies. When independent sources repeat the same claim, that claim has more support than one vendor page. The repair is to earn mentions in those independent sources. The content should match the fact you want repeated. A review that mentions your response time reinforces that dimension. A directory listing that omits it does not.

If you lose only when your business is named second, the evidence is sufficient. The problem is order sensitivity. Large language models show position bias in ranking and comparison tasks. The repair is more independent mentions across the web, because that increases the chance the model retrieves your name in the top set regardless of prompt order. Publishing more pages on your own domain does not change the order effect. It adds one source. Independent corroboration requires sources the business does not control. Follow-up prompts shift the evidence set, as covered in multi-turn AI visibility follow-up questions.

A self-published comparison page is a weak first move. It is one source controlled by the business. Assistants treat it as vendor content when independent sources disagree or when the page lacks third-party validation. The page also does not fix a missing spec fact if the fact sits on a separate product page. Fix the fact on the relevant page first. Then seek independent mentions. Patterns across competitors point to source consensus, as covered in why AI assistants recommend the same businesses.

An aeod.app audit labels each pair as win, loss, or tie. The label tells you which repair to run first. The next step is what to do with this list. No edit guarantees a win because the assistant decides at question time.

What to do with this

A comparison prompt is a closed test. The buyer supplies two names and a fixed order: X vs Y. The assistant retrieves evidence for both names and ranks the candidates before writing an answer. The starting line is set before the question is typed. Name order is part of the input. Incumbent familiarity adds retrievable mentions from directories and review sites. Press coverage adds more. Published comparables add pricing, service area, licensing, and case detail. Where one side has those facts and the other side has a homepage, the comparison is uneven before any content work begins.

The useful output of an audit is a verdict type per pair. A single win rate collapses the repair work. A loss and a tie point to different jobs. A loss means the competitor was retrieved and cited while the business did not surface, or surfaced below the competitor. That pair needs entity presence first: consistent name, category, location, and third-party corroboration so the assistant has a stable candidate to retrieve. A tie means both names appeared with similar treatment, or neither name had enough differentiating evidence. That pair needs sharper proof on the dimensions the buyer asked about: price, availability, credentials, response time, or service scope. The assistant has both names and chooses based on the evidence it finds.

Run the repair list by verdict. Losing pairs get foundational work. Tied pairs get differentiation work. Winning pairs get monitoring, because source coverage changes and a new competitor page shifts the evidence set. The audit’s per-pair label is the unit to act on. aeod.app reports those labels as part of the finding, so the queue starts with the pair type instead of an aggregate score.

No content change guarantees a head-to-head win. The assistant retrieves and decides at the moment the question is typed. A page edit alters one input among many. Retrieval depends on indexed sources and their facts. The query phrasing also shapes the result. The work is to make the business retrievable and comparable on the facts buyers ask about, then measure the next verdict. That is the limit of the method, and it is enough to prioritize.

Your own answers

See what AI says about your business.

One domain, one assessment. Mention and citation rates, competitors named instead of you, and a prioritized action list.

Get your report · $29 ↗

More notes