Multi-Turn AI Visibility: One Follow-Up Erases the List

A buyer asks an AI assistant for three plumbers. The assistant returns three names. That answer looks like a ranking. It is a snapshot of one turn.
Add a follow-up. The buyer asks which of those can come out today. The assistant answers again, and the original three names drop out. ChatGPT and Gemini generate each response from the full conversation plus their retrieval and ranking step. They do not lock a shortlist across turns. A constraint added in turn two changes the candidate set for turn two.
aeod.app runs this test by repeating the same buyer question across multiple turns and recording which businesses appear and where each name sits. The pattern holds across runs: the first answer is one turn, and the conversation continues.
The practical effect is that a business is named in turn one and absent by turn three. The buyer never sees a rejection. The buyer sees a new list. The earlier list is gone from the conversation unless the assistant repeats it.
That makes single-turn tracking incomplete. A dashboard that captures one prompt per buyer question measures the opening bid. It misses the follow-up that decides the purchase.
What changes between the first answer and the eighth? The candidate pool, the citation set, and the position of every name. The next section breaks down those changes turn by turn.
What changes between the first answer and the eighth
Multi-turn testing gives the first hard number. A buyer asks a broad question and gets a set of brand names. Add one buyer detail, “for a small team,” and 62% of those brands disappear from the reply. The remaining set is roughly 28% of the original candidate pool. One sentence of context erases most of the list.
The companion reading is starker. By turn eight, only 19% of the brands named in turn one are still named. Turn-one presence correlates weakly with turn-eight presence. A brand that appears in the opening answer has a small claim on the final answer. The opening answer is a snapshot of a wide query. The eighth answer is a snapshot of a narrow one.
The question that matters follows from that gap: which names survive a narrowing conversation, and what determines that? The answer sits in the mechanics of each turn. The candidate pool shrinks as constraints accumulate. The citation set changes because the assistant retrieves sources that address the narrowed detail. The position of every surviving name moves as the ranking criteria change. A business named for “CRM software” competes against a different set when the buyer adds “for a small team.” The assistant looks for evidence that the product fits a small team. Brands with that evidence stay in the reply. Brands without it drop out.
That process explains the weak correlation between turn one and turn eight. Turn one measures broad category association. Turn eight measures fit against the full conversation. The same businesses recur in broad prompts for reasons covered in Why AI assistants recommend the same businesses: training data frequency, citation density, and category association. Those forces weaken when the buyer adds a constraint. The assistant re-ranks against the constraint. The list shortens. The names that remain are the ones with matching evidence in the retrieved sources.
Recording mentions and citations at each turn makes the drop from turn one to turn eight visible step by step. Turn-one visibility is incomplete. A business appears in the first reply and vanishes by the eighth. A business absent from the first reply enters later when a later constraint matches its evidence. The full conversation decides the final shortlist.
The next section breaks down why a narrower question returns a shorter list.
Why a narrower question returns a shorter list
Turn one is a recall problem. A buyer asks which CRM a small agency should use, and the assistant pulls candidates that match the category at a coarse level. The query carries almost no filter, so there are few grounds to disqualify anyone. Recognizable names get named partly because they are recognizable.
Then a constraint arrives. A budget ceiling. A metro area. A team of eight. A required integration with a help desk. A live date six weeks out. The assistant reruns retrieval with the new constraint as a filter, checking the previous candidates against the added predicate. The confidence threshold for a mention effectively rises.
This is where the mechanism matters. The assistant runs a fresh pass over its sources, and the new predicate has to be satisfied by evidence it can retrieve: a pricing page, a case study naming an industry, documentation listing the integration. A vendor site that never states a price band gives the assistant nothing to match against “under $500 a month.” Absent evidence reads as a failed check at this stage. A candidate that fails leaves the set. There is no lower rank for it to occupy, because the assistant is tuned to avoid confident wrong answers. A bad recommendation costs user trust in a way an omission does not.
Five constraints across five turns behave as an intersection. Intersections shrink fast, and the surviving set is what the buyer sees.
The same logic explains the first shortlist. Those three names cleared a low bar on turn one: category match, public footprint, enough corroboration to state without hedging. The shortlist is a survivor set. Under a verification bar, a company is either inside it or invisible.
Multi-turn testing records this pattern often: a business is named on turn one and absent by turn four, with no change to its site. What changed was the question, and the business had no retrievable evidence written against the new constraint.
Recall on turn one, fit on turn five.
Recall on turn one, fit on turn five
A buyer asks ChatGPT for “commercial plumbers in Austin.” The first answer is a shortlist. Three names. The assistant is performing recall: which businesses appear often enough in directories and review sites to be known options? Inclusion here depends on surface area. A business with a complete Google Business Profile and a trade association listing enters the candidate set. The buyer has not stated a constraint yet. The assistant has not tested fit.
Turn five changes the task. The buyer adds “which of those can handle a restaurant grease trap permit?” or “which ones avoid weekend surcharges?” Now the assistant must confirm suitability. The evidence required shifts from “this is a known option” to “this option satisfies the condition.” Retrieval moves toward pages that contain the constraint language. A plumber’s homepage that says “residential service” provides no support. A permit filing or a review that names the exact job provides support. The assistant cites those sources or drops the name.
Rejection-shaped prompts sharpen the test. “Which of those should I avoid?” or “which ones have complaints about missed appointments?” asks the assistant to produce negative evidence. Models handle this with caution. If third-party sources are thin, the assistant has two moves: omit the brand from the avoid list, which protects it, or mention it as an unknown risk, which damages it. The outcome depends on what sources exist. Brands with thin third-party evidence fare worst because the assistant lacks material to defend inclusion and lacks material to resolve doubt. The same retrieval bias that decides turn one also makes turn five a test of sources.
Replaying buyer questions across turns shows this directly: compare inclusion on turn one with confirmation on turn five. A brand that appears on turn one and disappears on turn five lost the fit test. A brand absent on turn one never entered the test.
Longer conversations get looser, not more precise.
Longer conversations drift toward looser answers
When a buyer keeps asking, the assistant has more context. The context does not sharpen the answer. It expands it. Extended threads tend to produce longer responses as conversations become more complex. Hallucination rates rise alongside length. The mechanism is token pressure. Each turn adds prior messages and inferred intent. The model spends more capacity reconciling its own output and less on checking external facts. Continuity wins over verification.
Reviews of verified contradictions have found three assistants flatly disagreeing on brand facts within a single thread. The same assistant contradicts itself across turns. A brand gets named as a market leader in turn three. By turn nine it is a niche option. By turn fifteen it is attached to the wrong product line. Each statement sounds confident. The thread does not flag the conflict.
Late-turn mentions rest on looser verification. Early turns get a fresh retrieval pass or a direct answer. Later turns get compressed context. The model relies on its own previous claims instead of re-checking the source. Drift accumulates inside a single conversation. A follow-up about pricing pulls in a competitor’s feature set. A request for a shortlist drops the original constraint. The buyer sees a coherent conversation. The business sees a mention that no longer matches its facts.
A brand named early keeps getting named late because the assistant treats its own prior turn as evidence. Tracking mentions, citations, and position exposes whether a late-turn presence is stable or drifting from the cited source.
A single prompt captures one moment. It samples the assistant’s opening answer before context accumulates. The drift in turn twelve stays invisible to a tracker that never asks the follow-up.
Why single-prompt tracking cannot see this
Most AI visibility dashboards report a mention rate. They run one prompt per query, collect the assistant’s answer, and count whether the business appears. That number describes turn one. It measures the opening shortlist before the buyer adds a budget, a location, a constraint, or a follow-up.
Multi-turn data shows why that snapshot fails. Turn-one presence predicts later presence poorly. A business named in the first answer often drops out by turn three or turn twelve. Another business absent at the start enters after the buyer narrows the request. The opening answer and the final shortlist are different objects.
A single-prompt tracker cannot record that movement. It has no turn two. It has no mechanism to see the assistant revise its shortlist after the user says “only ones open on Saturday” or “who has experience with older homes.” The dashboard can show a stable mention rate for weeks while the conversational path collapses underneath it. The business remains visible in the metric and absent in the buying conversation.
The concentration pattern makes the gap sharper. Assistants tend to return a small set of names across many categories. Once a business loses its place in the opening shortlist, later turns rarely recover it, and follow-ups tend to reinforce the first answer rather than widen it.
Measuring a question as a sequence exposes this gap. Run the question as a conversation, record the assistant’s answer at each turn, and mark where names fall out. The output is a path with turns. That path shows whether a business holds position, drops after a constraint, or returns when the question changes.
This is the measurement layer that single-prompt reporting misses. The next problem is designing the sequence itself. Building a conversation-shaped prompt set starts with the buyer’s real follow-ups, because those follow-ups decide which names survive.
Building a conversation-shaped prompt set
Start with the broad category question in the buyer’s own wording. If the buyer types “Who should I hire to migrate our CRM?”, that exact phrase opens the thread. Replacing it with “best CRM migration vendors” changes the retrieval query. Assistants match against the tokens in the conversation, so the first turn sets the candidate pool.
Then add one constraint per turn in the order buyers add them: location, budget, team size, integration, timeline. Turn one is the category. Turn two adds “in Austin.” Turn three adds “under $15k.” Turn four adds “for a four-person ops team.” Turn five adds “that integrates with HubSpot.” Turn six adds “live by Q3.” Each turn keeps the same thread. The assistant carries prior constraints forward, reranks the shortlist, and drops names that no longer match.
Run that thread across several assistants, including ChatGPT, Claude, Gemini, and Perplexity. Keep each thread intact. A restarted thread resets retrieval because the assistant no longer carries the earlier constraints. Restarting before turn six makes the assistant answer as if location and budget were never stated. That reset erases the list being tracked.
Five to eight turns per thread reveals the survival curve. Fewer than five turns captures the broad category answer. Eight turns captures the late constraints that remove most names. The curve moves from broad to steep to flat. The steep section is where businesses disappear. Retrieval favors sources that match the accumulated constraint set, and most sources match only the first turn.
The method is mechanical. Write the category question. Add one constraint per turn. Keep the thread alive across assistants. Stop at eight turns. Then compare survival.
The three numbers to record per turn make that survival curve comparable across threads.
The three numbers to record per turn
Survival rate is the share of turn-one names still present at turn N. If turn one lists five names and turn four lists three of them, survival is three out of five. The number isolates coverage. It answers whether the assistant kept the same shortlist or rebuilt it. A high survival rate means the set is stable. A low rate means the conversation resets the candidate pool.
Position drift is the rank change for names that persist. Record the signed change from one turn to the next, using new rank minus previous rank. A business moving from second to fourth has a drift of +2, which means it lost two places. A business moving from fourth to second has a drift of -2. This isolates ordering. Coverage can stay flat while ordering moves. A name can survive every turn and still become less likely to be chosen because it slides down the list. Position drift shows whether the assistant’s preference among surviving names is strengthening or weakening.
Citation swap is the set of domains the assistant cites at each turn. At turn one, the assistant cites a local directory or a trade association page. At turn four, it cites a Reddit thread or a review site. Record the domains per turn and compare. This isolates evidence source. A shift from directory citations to review-thread citations at a later turn signals a different retrieval path. The assistant pulled from a different part of its index. That is a sourcing change. The assistant’s opinion can remain constant. The names can stay identical while the proof underneath them rotates.
aeod.app records these three fields per turn so the survival curve is comparable across threads. Survival rate tracks coverage. Position drift tracks ordering. Citation swap tracks evidence source. When a brand disappears at turn three, the first diagnostic is which number moved first. If survival falls to zero, coverage broke. If survival holds and drift rises, ordering broke. If citations swap before the name drops, the retrieval path changed before the shortlist did. That distinction sets the repair work for the next section.
What to fix when a brand disappears at turn three
The diagnostic order follows from that distinction. The turn where the name drops is a symptom. The cause sits in the evidence the assistant retrieved before that turn.
When the follow-up adds a budget constraint, the assistant filters the shortlist by plan structure. A brand with no published pricing or plan tiers has no retrievable answer. The assistant keeps names that expose price bands and starting costs. The fix is a pricing page with named plans and ranges. Third-party directories need the same structure. The absence is a data gap.
When the follow-up asks about team size or capacity, the assistant looks for evidence that the business can staff the work. A brand with no third-party coverage framing its team size has no anchor. The assistant drops it. The fix is coverage that states team size and service capacity. Profiles need employee count or crew size where fields exist.
When the follow-up asks about a specific neighborhood or service area, conflicting entity data across profiles causes the assistant to lose confidence. One profile says city center. Another says county. The assistant resolves the conflict by choosing a name with consistent location data. The fix is reconciling name, address, phone, service area, and category across Google Business Profile, Bing Places, Apple Business Connect, industry directories, and the site’s contact page.
The drop point determines the order. If survival falls to zero at budget, fix pricing before writing more blog posts. If survival holds and position drift rises at team size, fix third-party framing before adding case studies. If citations swap before the name drops, fix the retrieval path.
Recording the turn where the name drops and the field that moved first turns a vague disappearance into a testable cause. The same retrieval logic explains why the shortlist converges on a small set of names.
Diagnose the turn and the cause before producing any more content. Content volume does not repair a missing pricing page, a thin team-size frame, or conflicting location data.
Treat the conversation as the unit of measurement. A single prompt gives one reading. A chain of four turns shows whether a brand remains after location, budget, team size, and license constraints enter. That persistence rate separates brands with matching evidence from brands that only appear in broad category lists.
Start by replaying buyer questions as narrowing chains. Ask the broad category question, record the shortlist, then add one constraint per turn. Track mention status and citation sources at each turn, along with rank position. The drop turn is the diagnostic. A brand that disappears when the question asks for under $500 per month has a pricing evidence gap. A brand that disappears when the question asks for licensed in Texas has a location or compliance gap. The assistant is failing to verify it against the new constraint.
Fix the evidence at the constraint level. Publish the thing the follow-up asks for: price ranges, license numbers, service areas, team size, and roles. Keep those facts consistent across the site and third-party profiles. Conflicting addresses or outdated headcounts give the assistant no stable entity to cite.
Re-run the same chains after the changes. Measure survival across turns. No single edit guarantees survival in a later turn. No tool can promise it. What survives is whatever the assistant can still verify once the question gets specific.
Your own answers
See what AI says about your business.
One domain, one assessment. Mention and citation rates, competitors named instead of you, and a prioritized action list.
Get your report · $29 ↗More notes

31 Aug 2026 · 13 min read
Why AI Assistants Recommend the Same Three Businesses
Why AI assistants recommend the same businesses, explained by concentration data, retrieval sources, and entity consistency.

1 Aug 2026 · 9 min read
Blocking AI Crawlers: What It Changes in AI Answers
Blocking AI crawlers splits into training and answer bots, and 2026 data disagrees on the citation cost. What to check before you block.