An AI tax tool returns a VAT determination and a number beside it: 94%. The reviewer sees green, approves, moves on. Ask who set the threshold at which green appears, and the answer is almost never a person inside the tax function.

Confidence scores have arrived in enterprise tax tooling ahead of any governance vocabulary for them. Classification engines, HS code suggestion tools, document extraction, VAT treatment determination — each surfaces a probability alongside its answer. The number looks like a control. It behaves like one in the process design: above the threshold, straight through; below it, human review. But a confidence score is a statement the model makes about its own output. It is not a statement about whether the tax position is correct, and it is not evidence that anyone accepted responsibility for it.

This is Cognitive Responsibility Diffusion™ operating at its most subtle. The framework describes what happens when a decision passes through enough hands, systems, and defaults that no participant experiences themselves as the decision-maker. A confidence threshold is an unusually efficient diffusion mechanism, because it converts an unowned judgement into a piece of infrastructure.

Three Questions That Have No Owner

Take a threshold set at 90%. Three distinct decisions are embedded in it, and in most implementations none has a named owner.

Who chose 90? In practice the number usually arrives as a vendor default, a figure carried over from a proof of concept, or a round number proposed in an implementation workshop and never revisited. It is rarely derived from the enterprise's own error tolerance, its transaction mix, or the penalty exposure attached to the transactions it governs. The number is treated as a technical configuration parameter, so it is set by whoever configures the system — which is to say, by someone whose accountability does not extend to the tax position.

What does 90 mean? A confidence score is calibrated against the model's training distribution. It expresses how closely an input resembles cases the model has seen. It does not express the probability that a tax authority would agree with the output. Those two things diverge most sharply exactly where it matters: novel transaction types, restructured supply chains, newly designated goods, recently amended treatment. The transactions carrying the highest interpretive risk are the ones the model has seen least of — and a model can be confidently wrong about them, or unhelpfully uncertain about routine items whose treatment is obvious to any practitioner.

What happens at 89.4%? The item routes to human review. The reviewer receives it with a visible signal that the machine nearly approved it. That signal is not neutral. A queue of items each marked as marginal, reviewed under time pressure, produces a predictable pattern: the review becomes ratification. The human is present in the process and absent from the decision.

The Straight-Through Population Is the One to Worry About

Attention in AI tax governance discussions tends to concentrate on the exception queue — the items the model flagged. The exception queue is the governed population. Something happens to those items; someone looks at them; there is a record.

The population above the threshold receives no human attention by design. It is also, in most enterprises, the overwhelming majority of transactions. If the threshold is miscalibrated — set to a vendor default, tuned for throughput, or simply inherited — the error does not appear in the exception queue. It appears as a large volume of unreviewed determinations that everyone believes were reviewed by the model.

This is the Green Dashboard Paradox™ in a new setting. The dashboard reports a low exception rate and a high straight-through rate. Both are read as evidence that the system is working. Neither measures whether the determinations are correct. A tuned-down threshold improves every metric on the dashboard while increasing the population of unexamined positions.

What the Audit File Contains

Under a real-time regime the question stops being theoretical. When a tax authority queries a specific transaction, the enterprise must reconstruct why that treatment was applied. If the answer is that a model returned a treatment with 94% confidence and the threshold was 90%, the enterprise has described its process, not defended its position.

A defensible record needs the reasoning available for reconstruction: which rule was applied, what facts drove it, what alternatives were considered, and who accepted the outcome. A probability score is none of these. It is a summary statistic that has discarded the reasoning. Where the tooling retains only the score, the enterprise has automated the determination and deleted the evidence — precisely the gap that Tax Defensibility and Evidence Architecture™ exists to close.

Making the Threshold a Governed Object

The remedy is not to abandon confidence scores. They carry real information about where a model is operating outside familiar ground. The remedy is to stop treating the threshold as configuration and start treating it as a tax position.

Name an owner. The threshold should be owned by the person accountable for the tax positions it governs — not by the implementation team, not by the vendor, not by whoever holds admin rights. Ownership means the owner can state why the number is what it is.

Set it against exposure, not throughput. Different transaction populations warrant different thresholds. Domestic standard-rated supplies to established customers are not the same risk as newly designated goods, cross-border services, or transactions where a reverse charge determination turns on counterparty status. A single enterprise-wide threshold is a statement that all tax risk is equivalent.

Sample above the line. The straight-through population needs periodic human sampling for exactly the reason that it receives no human attention. Without it, threshold miscalibration is undetectable — the failure mode produces no signal.

Version it. A threshold change alters the tax treatment of every subsequent transaction in scope. It belongs under change control, with a record of the prior value, the new value, the rationale, and the effective date. When a period under review sits on one side of a threshold change, the enterprise needs to know that.

Require reasoning, not just scores. In tool selection and ASP evaluation, the question to ask is whether the system can produce the basis for a determination months later, or only the number it attached at the time. That capability is an architectural property. It cannot be retrofitted after an authority has asked.

The Underlying Point

Automation does not remove judgement from a tax function. It relocates it — usually upstream, into configuration decisions taken early, by people whose accountability does not extend to the outcome, and recorded nowhere that a tax reviewer would look.

The confidence threshold is one of the clearest examples available. It is a small number in a settings panel that determines how many of an enterprise's tax positions receive human attention. Whoever sets it is making a tax policy decision. The governance question is simply whether that person knows it.

Cognitive Responsibility Diffusion™, the Green Dashboard Paradox™, and Tax Defensibility and Evidence Architecture™ are frameworks developed in Real-Time Tax Transformation (forthcoming) and the 17-module practitioner course.