Skip to content
Ghost DecisionsPaper · v0.2 · Working draft

Paper

Ghost Decisions

Chaining AI steps can multiply errors or catch them. What decides which way it goes is whether each step's decisions stay open to checking and change, or get locked in and passed downstream as fact.

David Paquet Pitts · Oct 10, 2026 · v0.2 · 14 min read

In brief · the one-minute read

  1. Chained AI steps multiply their errors when nothing checks them: six steps that are each right 95% of the time are all right together about 74% of the time. See it run.
  2. Chains with real checks do better than a single pass, but only when the check sees something the step couldn't. Why.
  3. What decides it is a decision's status, not the number of steps. A choice that shapes everything downstream, but that nobody can see, question or change, is a ghost decision.
  4. One reasonable grouping can produce a well-written, useless article, and rejecting it only reruns the same mistake. Watch it happen.
  5. The fix is correction boundaries: places where a decision can be challenged and everything built on it updated. Open the decisions with the widest reach and reuse first. The rules and where to start.

The promise: power without the manual

For decades, software made you pick. Simple tools did less. Powerful tools did more, but you had to learn to drive them: taxonomy editors, report builders, rule engines.

AI breaks that trade-off. You describe the result you want and the system runs the tools for you. Jakob Nielsen calls this intent-based outcome specification, the first new UI paradigm in 60 years: you say what you want, not how to get it.

Under the hood, that usually means a chain. One step classifies, the next summarizes, the next analyzes, the next writes. Each step is a specialist, and together they do things no single tool could.

Go deeperThree paradigms of interface, and the one a sealed chain brings back

Nielsen describes three paradigms of user interface [1]. In batch processing, users submitted a whole job at once, and a single error could make the output meaningless. Command-based interaction fixed that by letting users reassess after each command. Intent-based outcome specification has users state the outcome they want and leaves the method to the system. Nielsen notes the cost: when users don't know how a result was produced, it is harder for them to identify or correct problems, and he expects hybrid interfaces that keep graphical controls.

The observation this paper starts from: a chain whose intermediate decisions are hidden and fixed reintroduces batch processing's defining property. The user sees only the final output and can only resubmit, even though the interface looks conversational.

The catch: errors multiply

If every step in a chain builds on the one before, the whole thing is only right when every step is right. So the odds multiply. Run the same chain a hundred times and count:

Fig. 1The same chain, run a hundred times

78 / 100runs came out right. The arithmetic says 74%.

Each step is right95%
Steps in the chain6
Between steps
Run 1: rightRun 2: rightRun 3: rightRun 4: rightRun 5: rightRun 6: rightRun 7: rightRun 8: rightRun 9: rightRun 10: rightRun 11: step 6 was wrong and nothing caught it; 0 later steps built on itRun 12: rightRun 13: rightRun 14: step 3 was wrong and nothing caught it; 3 later steps built on itRun 15: step 4 was wrong and nothing caught it; 2 later steps built on itRun 16: rightRun 17: rightRun 18: rightRun 19: rightRun 20: rightRun 21: rightRun 22: step 4 was wrong and nothing caught it; 2 later steps built on itRun 23: rightRun 24: rightRun 25: rightRun 26: rightRun 27: rightRun 28: step 2 was wrong and nothing caught it; 4 later steps built on itRun 29: rightRun 30: rightRun 31: rightRun 32: rightRun 33: rightRun 34: step 4 was wrong and nothing caught it; 2 later steps built on itRun 35: step 1 was wrong and nothing caught it; 5 later steps built on itRun 36: step 4 was wrong and nothing caught it; 2 later steps built on itRun 37: rightRun 38: rightRun 39: rightRun 40: rightRun 41: rightRun 42: rightRun 43: rightRun 44: rightRun 45: rightRun 46: step 4 was wrong and nothing caught it; 2 later steps built on itRun 47: step 4 was wrong and nothing caught it; 2 later steps built on itRun 48: rightRun 49: rightRun 50: rightRun 51: step 3 was wrong and nothing caught it; 3 later steps built on itRun 52: rightRun 53: step 5 was wrong and nothing caught it; 1 later step built on itRun 54: step 5 was wrong and nothing caught it; 1 later step built on itRun 55: rightRun 56: rightRun 57: rightRun 58: rightRun 59: rightRun 60: rightRun 61: step 1 was wrong and nothing caught it; 5 later steps built on itRun 62: rightRun 63: step 4 was wrong and nothing caught it; 2 later steps built on itRun 64: rightRun 65: step 1 was wrong and nothing caught it; 5 later steps built on itRun 66: rightRun 67: rightRun 68: rightRun 69: rightRun 70: step 2 was wrong and nothing caught it; 4 later steps built on itRun 71: rightRun 72: rightRun 73: step 1 was wrong and nothing caught it; 5 later steps built on itRun 74: rightRun 75: rightRun 76: step 4 was wrong and nothing caught it; 2 later steps built on itRun 77: rightRun 78: rightRun 79: rightRun 80: rightRun 81: step 1 was wrong and nothing caught it; 5 later steps built on itRun 82: step 1 was wrong and nothing caught it; 5 later steps built on itRun 83: rightRun 84: rightRun 85: rightRun 86: rightRun 87: rightRun 88: rightRun 89: rightRun 90: rightRun 91: rightRun 92: rightRun 93: step 5 was wrong and nothing caught it; 1 later step built on itRun 94: rightRun 95: rightRun 96: rightRun 97: rightRun 98: rightRun 99: rightRun 100: right
  • right
  • wrong, and passed on
  • built on a wrong step
  • caught and sent back
A simulation, not a measurement: each step fails at random at the rate you set, the same random draws every time, and a failure is passed on unless a check catches it. A check here sees something the step didn't, so it catches its share of mistakes independently.

Six steps at 95% each come out right together only about 74% of the time. At 80% per step, it's about 26%. That assumes nothing catches mistakes along the way, which is exactly what a chain looks like when nothing is checking. Put a check between the steps and the picture changes; how much depends on what the check can see.

In the batch-processing era, you handed over a whole job at once, and if it held the slightest error, the output was meaningless. Interactive interfaces fixed that by letting you look after each step and adjust. A sealed AI pipeline quietly brings batch processing back.

Go deeperWhat the research measures

That chained and multi-step AI systems compound errors is well documented:

  • Compositional reasoning. Dziri et al. argue that autoregressive models' performance can decline rapidly as task complexity grows, and find that models often solve multi-step problems by pattern matching rather than systematic reasoning [2].
  • Agent reliability. τ-bench introduced pass^k, the chance an agent succeeds on all of k repeated trials. Leading models succeeded on under half of tasks, and pass^8 fell below 25% in the retail domain [3].
  • Task length. METR measures the length of task agents can complete with 50% reliability, a framing that exists because success falls as tasks get longer [4].
  • Multi-agent systems. An analysis of over 1,600 traces across seven multi-agent frameworks found 14 failure modes in three groups: specification and system design, inter-agent misalignment, and task verification. Most were design problems, not model limitations alone [5].

The simulation assumes each step fails independently. Real steps share data and models, so failures correlate, and later steps can sometimes repair earlier ones. It shows the direction of the effect, not a measurement.

The twist: checked chains do better

If that were the whole story, the answer would be easy: use fewer steps. It isn't.

Some of the best AI results come from chains with checking built in. OpenAI's Let's Verify Step by Step trained a model to check every reasoning step instead of only the final answer, and step-by-step checking clearly won on hard math problems [6]. METR found the length of tasks AI agents can finish has been doubling roughly every seven months, driven partly by better reliability and a better ability to adapt to mistakes [4].

So chaining isn't the villain. A well-designed chain can take on bigger work than any single model, because each step is smaller and something can check it. But the checking has to be real:

Taken together: chains compound errors when nothing independent can intervene, and catch them when something can.

What decides it: ghost decisions

A chain that compounds errors and a chain that catches them can look identical: same steps, same models. The difference is what happens to each step's decisions.

  • Open decisions can be questioned. A judge, a test or a person can challenge a step and send it back.
  • Locked decisions are passed downstream as fact. Every later step builds on them, and nothing can push back.

Ghost decision: a choice the system made that shapes everything after it, but that nobody can see, question or change.

Ghost decisions are what turn a self-correcting chain into a compounding one. A small error gets locked in. The next step makes a small decision on top of it, and that gets locked in too. The result drifts further from what the customer needed while staying perfectly consistent with itself: the compounding assumption effect.

A ghost decision doesn't have to come from a model. A default configuration, a hard-coded rule or the product team's own workflow design can make one; a deterministic rule can apply the wrong business definition perfectly. And a check only counts if it can see something the step couldn't. A judge with the same blind spots, or tests written from the same wrong assumption, isn't independent. Sometimes the only one with that outside view is the customer.

Go deeperWhere the idea comes from

The interaction-design side of the problem has a long history, and much of the vocabulary already exists:

SourceCore ideaRelevance here
Norman, 1990 [10]The problem with automation is inappropriate feedback and interaction, not automation itselfA good happy path doesn't make a good product; the exception path is part of the design
Woods, 1996 [11]Automation transforms work rather than removing it; tighter coupling spreads disturbances and complicates diagnosisChaining moves effort from operating tools into diagnosing and repairing outputs
Green & Blackwell, 1998 [12]Hidden dependencies, viscosity (resistance to change), progressive evaluationA ghost decision is a hidden dependency; a sealed chain is viscous
Horvitz, 1999 [13]Mixed initiative: design automation around uncertainty about the user's goals and chances to refine resultsAn inferred decision is a guess about intent, not permission to proceed
Shneiderman, 2020 [14]Automation and human control are separate dimensions; both can be highOne-click workflows can still support inspection and override
Amershi et al., 2019 [15]Guidelines for human-AI interaction, including efficient correction, global controls and conveying consequencesCorrection needs a clear scope and a preview of its effects
Wu, Terry & Cai, 2022 [16]Letting users edit LLM chains and intermediate results improved outcomes and perceived transparency and controlThe closest direct precedent for correction boundaries
Ehsan et al., 2024 [17]Seamful design: revealing useful mismatches can support understanding and agencyExposing a decision is a feature, not an imperfection

The closest precedent is AI Chains [16]. Its study combined decomposition with editable intermediate steps, so it does not isolate the value of the editable boundary itself, and it was not run on enterprise products at scale.

The proposal, stated plainly. Whether a chain compounds errors or catches them depends less on its number of steps than on the status of its intermediate decisions. Where decisions stay open to an independent check, errors can be caught. Where they are locked and passed downstream as fact, errors compound. This is a synthesis of the work above, not an empirical finding; see the limits.

A ghost decision in the wild

Picture a knowledge-base product that turns support tickets into help-center improvements. One request in, one finished recommendation out. Behind the scenes, it's a four-step chain.

In step one, the AI groups "refund requested," "refund delayed" and "chargeback disputed" into a single topic called Billing. It's a reasonable guess. But this company handles those three with different teams and different policies, and the surge this month is in payouts that arrive late. Try both ways out of it:

Fig. 2One reasonable grouping becomes the wrong article

Run 1 of the sealed product.

  1. 1Classify the ticketsGhost decision
    Billing · 25 tickets“Can I get a refund on order 4471?” (Billing)“Wrong size, how do I send it back for my money?” (Billing)“Refund please, the item arrived broken” (Billing)“Where do I request a refund?” (Billing)“I cancelled within 14 days, refund?” (Billing)“Charged twice, need one refunded” (Billing)“Refund for the annual plan I didn't use” (Billing)“How long does a refund request take to approve?” (Billing)“My payout still hasn't arrived after 5 days” (Billing)“Refund approved last week, nothing in my account” (Billing)“Payout says sent, my bank says no” (Billing)“Why is my refund delayed again?” (Billing)“It's been 10 business days, where's my money?” (Billing)“Refund delayed, I need it for rent” (Billing)“When exactly do payouts land?” (Billing)“Payout stuck in 'processing'” (Billing)“Approved refund but no transfer yet” (Billing)“Refund delayed over the weekend?” (Billing)“Still waiting on the refund you confirmed” (Billing)“Is there a holiday delay on payouts?” (Billing)“My bank opened a dispute, what now?” (Billing)“Chargeback filed by mistake, can I cancel it?” (Billing)“Customer disputed a charge I already refunded” (Billing)“What evidence do you need for a dispute?” (Billing)“Chargeback fee on my statement” (Billing)

    “Refund requested”, “refund delayed” and “chargeback disputed” look alike, so they become one topic.

  2. 2Report on each topicBuilt on it
    Billing18 → 25 (+39%)

    Billing is up 39%. True, and no help: the growth has nowhere to point.Bar: this month · tick: last month

  3. 3Find coverage gapsBuilt on it
    TopicTicketsArticlesGap
    Billing252Biggest

    “Not enough Billing content”: the two existing articles are spread over 25 tickets.

  4. 4Suggest an articleBuilt on it
    Draft article
    How to request a refund

    Requesting a refund is easy. Open your order, choose Request refund, pick a reason and submit. Most requests are reviewed within two business days…

    A help-center article with this title already exists.

Made-up tickets for a made-up help center, run through the four steps. Hover a ticket to read it. Reject the article and the next run repeats the same grouping; open the topic and split it, and each step downstream redoes its work.

The article the sealed chain writes is well made and beside the point: it already exists. Every step did its own job correctly. The report counted the topic it was given, the gap finder found the gap the report implied, the writer wrote. The only mistake was one reasonable grouping that nothing downstream was allowed to question.

Rejecting the article puts the customer in the reroll trap. The next run makes the same grouping and the same mistake. They're fixing the symptom over and over, because the cause is a ghost.

The fix: seamless, not sealed

The answer isn't to bring back every manual step. It's to keep the decisions that matter open.

Correction boundary: a place where a decision can be seen, challenged and changed, and everything that depends on it gets updated.

A correction boundary can be used by a machine or a person:

  • A judge or verifier checks the topic grouping against the company's existing help-center structure before reports get built.
  • A person sees the Billing topic with its example tickets, splits it in three, and gets a preview: "This recalculates 1 report, reassesses 3 gaps and flags 1 draft for review. Nothing is published."

The happy path stays one smooth flow. The boundaries are just there when someone needs them. Ben Shneiderman makes this point well: automation and human control aren't opposites, and you can have a lot of both [14].

That's the difference between seamless and sealed. A seamless product removes unnecessary work. A sealed one also removes the places where mistakes get caught and customers say how their business is different.

Seven rules for killing ghost decisions

  1. Make key decisions real records. A topic definition, a metric, a plan: give it a name, a version and a history. You can't check or change what only exists inside a model call.
  2. Give every check an outside view. A verifier needs something the step didn't have: different data, a different model, the customer's own definitions, or a person [7, 8].
  3. Show decisions, not machinery. Show the Billing topic and its tickets, not prompts and logs. Seeing every step isn't the same as being able to change anything.
  4. Let people trace from the result back to the cause. Every output should link to the decisions and evidence behind it. Real records, not a plausible story.
  5. Make the scope explicit. This ticket, this topic, or every future run? A local fix should never quietly become a global policy [15].
  6. Update what depends on the change, and protect what's approved. Recompute downstream work, keep edits people already signed off on, and never resend an email just because an analysis reran.
  7. Match control to consequence. A throwaway draft needs an optional edit. A shared taxonomy needs a preview and permissions. Anything that leaves the building needs approval first.

Where to invest first

Opening up a decision costs something. You have to pull it out of the model call, store it, check it and build a way to change it, which adds engineering time and latency. So start where it pays off: decisions with wide reach (lots of work depends on them) and long reuse (they stick around). Here are the decisions in the same help-center product:

Fig. 3Where to open decisions first, in the same product
Open these firstReuse: one run → every weekReach: outputs that depend on itTopic taxonomy: Every report, every gap and every draft, every week.Topic taxonomyHow a ticket counts: Each topic's volume and trend, and so which gaps rank first.How a ticket countsGap threshold: Which topics get a draft at all.Gap thresholdArticle tone: The wording of each draft; easy to fix in the draft itself.Article toneThis week's summary: One report for one week; forgotten by the next.This week's summaryOne headline's wording: Nothing downstream. Edit it and move on.One headline's wordingEmail to every customer: Nothing downstream, but it can't be unsent.Email to every customer

Every report, every gap and every draft, every week.

Illustrative estimates for the help-center product above: reach is how many outputs depend on a decision in one run, reuse how many runs a year use it. Both axes are logarithmic. Pick a decision to see what rests on it.

Consequence can jump the queue. A one-time action that's hard to undo, like emailing customers or deleting data, deserves a strong boundary even if it never repeats.

For each important decision, ask: when this is wrong, who or what notices, where does it get fixed, and what happens to everything built on it? If nobody can answer, you've found a ghost.

When not to bother

  • Low-stakes, short-lived steps. Nobody needs to edit how one sentence was phrased in a throwaway draft.
  • Decisions nobody can judge. If users lack the expertise to improve a decision, an edit field won't help. A better verifier or an expert reviewer will.
  • Honest narrow products. If your product deliberately does one thing, say so. Doing less isn't the problem; pretending to be flexible is.

The goal isn't a settings screen for everything or an approval step at every turn. It's the right check, in the right place, by something that can actually see the mistake.

Before you ship

Chaining AI steps isn't good or bad. Chains with open decisions catch mistakes and take on bigger work. Chains full of ghost decisions multiply mistakes and trap customers in the reroll loop. Before you ship an AI workflow, ask:

  • Which decisions shape everything downstream?
  • Can a judge, a test or a person challenge each one, with a view the step itself didn't have?
  • Can users get from a bad result back to the decision behind it?
  • When a decision changes, does the work that depends on it update, without wrecking approved work?
  • Have you tested how long it takes to fix a wrong result, not just how fast the first one appears?

Automate the work. Keep the decisions open.

Limits, and how to test this

This is an argument built from other people's evidence, not a study of its own. The parts that would change it are below.

Go deeperLimitations

The proposal is a synthesis, not an empirical finding. It has not been tested directly; the cited studies support its parts in different settings (math reasoning, agent benchmarks, lab studies of chain editing), not the combination in production products. The simulation assumes independent failures, which real chains violate. The worked example and its tickets are made up.

Go deeperOpen questions
  • Decomposition cost. Many current systems make intermediate decisions implicitly inside a single model call. How much latency, cost and quality does extracting them as records trade away?
  • Discovery. How can teams find their ghost decisions before customers do? Support logs, override frequency and churn reasons are candidate signals.
  • Verifier independence. How independent must a machine check be to help, given that model errors are becoming more correlated [8]?
  • Who can judge. When users lack the expertise to evaluate a decision, which boundary works best: an edit field, an expert reviewer, or stricter validation?
Go deeperA study that would test it

Compare three interfaces over the same underlying pipeline: a sealed end-to-end experience; one with visible provenance but no source-level edit; and one with provenance, scoped correction and dependent repair. Tasks should include a seeded upstream error, a reasonable but unsuitable business definition, a late-stage error that needs no upstream change, and a case where the default is already right, so the study does not reward editing for its own sake. Participants should be people who own those decisions, such as support leads and knowledge managers.

Measures: time to a validated result, successful source-level corrections, repeated symptom fixes, damage to approved work, and added cost on runs where nothing was wrong. For recurring workflows, check whether corrections persist in later runs.

References

  1. Nielsen, J. (2023). AI: First New UI Paradigm in 60 Years. Nielsen Norman Group.
  2. Dziri, N., et al. (2023). Faith and Fate: Limits of Transformers on Compositionality. NeurIPS 2023.
  3. Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.
  4. Kwa, T., West, B., et al. (2025). Measuring AI Ability to Complete Long Software Tasks. METR.
  5. Cemri, M., et al. (2025). Why Do Multi-Agent LLM Systems Fail? NeurIPS 2025.
  6. Lightman, H., et al. (2023). Let's Verify Step by Step. OpenAI.
  7. Huang, J., et al. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024.
  8. Goel, S., et al. (2025). Great Models Think Alike and this Undermines AI Oversight. ICML 2025.
  9. Chen, L., et al. (2024). Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI Systems. NeurIPS 2024.
  10. Norman, D. A. (1990). The "Problem" with Automation: Inappropriate Feedback and Interaction, Not "Over-Automation". Philosophical Transactions of the Royal Society B, 327, 585–593.
  11. Woods, D. D. (1996). Decomposing Automation: Apparent Simplicity, Real Complexity. In R. Parasuraman & M. Mouloua (Eds.), Automation and Human Performance, 3–17. Erlbaum.
  12. Green, T. R. G., & Blackwell, A. F. (1998). Cognitive Dimensions of Information Artefacts: A Tutorial.
  13. Horvitz, E. (1999). Principles of Mixed-Initiative User Interfaces. CHI '99.
  14. Shneiderman, B. (2020). Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy.
  15. Amershi, S., et al. (2019). Guidelines for Human-AI Interaction. CHI 2019.
  16. Wu, T., Terry, M., & Cai, C. J. (2022). AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. CHI 2022.
  17. Ehsan, U., Liao, Q. V., Passi, S., Riedl, M. O., & Daumé III, H. (2024). Seamful XAI: Operationalizing Seamful Design in Explainable AI. Proc. ACM HCI, CSCW1.

Terms this paper coins

Ghost decision
A choice the system made that shapes everything after it, but that nobody can see, question or change.
Correction boundary
A place where a decision can be seen, challenged and changed, and everything that depends on it gets updated.
Reroll trap
Rejecting and regenerating the visible output while the decision that caused it stays out of reach, so every run repeats the mistake.
Compounding assumption effect
A ghost decision becomes the premise for every later step; each step adds small decisions of its own, and the result drifts from what was needed while staying internally consistent.

Cite this

Paquet Pitts, D. (2026). Ghost Decisions (v0.2, working draft). https://davidpp.com/papers/ghost-decisions

David Paquet Pitts
Head of Delivery & Solutions at Botpress

Versions

  1. v0.2currentOct 10, 2026Renamed around the failure itself. Adds the prior art on verification, the Billing example, seven rules and the reach × reuse heuristic.
  2. v0.1Oct 10, 2026 · as “Seamless, Not Sealed”First take: sealed AI pipelines bring batch processing back, and the fix is a product that is seamless without being sealed.

A working draft: the argument is stated and sourced, and will keep changing. Each change gets a version.