The AI Liability Chasm Is a Verifiability Problem — And the Log Is the Battleground

Wootoshi
Guide

Hook

Here is a data signal that has nothing to do with price and everything to do with positioning. Effective January 1, 2026, Verisk's generative AI exclusion attaches to commercial general liability renewals — not one insurer's appetite change, but a standard-form move inside the ISO ecosystem. Once an exclusion lives in the form library, individual underwriters cannot easily write against it. Cross-reference that with the claim that only 22% of enterprise AI contracts carry uncapped vendor indemnity, and a split of the remainder rendering to 41/33/26 — percentages that sum to 100 while sitting beside a separate bucket labeled as the rest. Precise numbers, no named research house, arithmetic that does not close. In a sideways market where everyone is waiting for a catalyst, that combination is itself the signal: the liability plumbing is being rebuilt faster than anyone can verify who is doing the rebuilding.

Context

The legal architecture underneath this is real, not speculative. EU Directive 2024/2853, the revised Product Liability Directive, extends strict liability to software and AI systems, and its transposition deadline of December 9, 2026 sits close enough to matter. More consequentially, it introduces a defect presumption and an evidence-disclosure duty: where technical complexity makes proof unreasonably difficult, the burden shifts toward the provider. Strict liability means no negligence to prove. Defect presumption means the provider defends from behind. Separately, the EU AI Liability Directive has been effectively withdrawn from the Commission's work programme — a quieter and more final outcome than the stalled language usually applied to it.

FTC Chair Andrew Ferguson has publicly rejected the autonomous actor defense. Garcia v. Character Technologies, Raine v. OpenAI, and the Pennsylvania action against Character proceed on product liability and failure-to-warn theories rather than copyright. And Amazon v. Perplexity tests whether a third-party agent may act on a platform without authorization. Read that list as an engineer rather than a lawyer and it stops looking like a regulatory trend. It looks like a missing capability. Open books, open ledgers, open hearts — the market is asking for a record.

Core

Agent failure is not one thing. I keep returning to a taxonomy I built during my 2017 contract audits, when I spent three months reading token distribution logic line by line and found three flaws in a storage project's allocation mechanism that nobody had flagged. The flaws were never in the math. They were in the assumptions the math rested on. Agent failures split the same way. Under-specified instructions: the agent did exactly what it was told, and the instruction was incomplete — a contract problem between deployer and user. Capability defects: hallucination and reasoning error — a design-defect and failure-to-warn problem. Contaminated input: injection through a webpage, a document, a poisoned tool response — where damage originates with a third party and the instruction-to-harm causal chain simply breaks.

That third category is where the FTC model fails. Ferguson's position — that an out-of-control agent's audit trail consistently shows it executed the specific instructions it was given — presupposes the trail is complete, faithful, and replayable. It is not. Tool-calling traces get truncated and summarized. Browser agents log actions but not the reasoning that selected them. And chain-of-thought faithfulness has been repeatedly shown to be unreliable: the stated reasoning is often not the operative one. The regulator's causal model depends on an observability property that current agent frameworks do not have. Tracing the code back to the conscience only works if the code kept a record of the conscience in the first place.

There is also a category error running through the debate: autonomy is treated as binary when it is a dial. Almost every enterprise agent shipping now is bounded autonomy — orchestration, tool calls, human approval gates. The real dividing line is not whether it is an agent, but whether it has an untrusted input channel. Any agent reading web pages, email, documents, or third-party APIs inherits injection risk that prompt engineering cannot remove. The legal language never makes that distinction, which conveniently expands jurisdiction for regulators and demand for anyone selling cover.

The gap nobody has closed is authorization. In 2025 I ran a workshop series for two hundred executives at a Japanese bank, explaining self-sovereign identity through the structure of a tea ceremony — the point being that consent is a sequence of deliberate, witnessed acts, not a checkbox. Fifteen clients signed on to pilot DID-based KYC. What that taught me is that the hardest question in agent liability is not whether the model erred, but who authorized this agent to act as whom, and whether that can be proven without relying on the counterparty's own database. Card networks and platform vendors began pushing agent payment authorization standards in 2025, and those documents define technical interaction and liability allocation in the same breath. Whoever writes that document writes the answer.

The AI Liability Chasm Is a Verifiability Problem — And the Log Is the Battleground

Here is where the blockchain conversation stops being decorative. The infrastructure that makes agent risk measurable is boring: write-once storage, Merkle-anchored append-only logs, versioned model snapshots, signed guardrail attestations, DID-bound delegation chains verifiable independently. None of it requires a token. It requires a settlement layer cheap enough to anchor proofs to, and a willingness to treat observability as a product rather than an afterthought. Based on my audit experience, the failure mode is predictable: teams log what is convenient, retain it for the shortest defensible window, and discover during litigation that the evidence they need rotated out a year ago. The audit is not the end, but the beginning — and most teams treat it as neither.

The AI Liability Chasm Is a Verifiability Problem — And the Log Is the Battleground

Now the harder arithmetic. Classic insurability requires five conditions: many independent homogeneous risk units; actuarially estimable frequency and severity; a bounded maximum loss; no severe moral hazard or adverse selection; and an affordable premium. Agent risk fails all five. Every enterprise shares a handful of foundation models, so losses correlate rather than diversify. There are no adjudicated large-loss samples, therefore no severity curve. Indirect damages have no natural ceiling. The deployer fully controls guardrail strength, which is textbook moral hazard. And you cannot price what you cannot measure.

The AI Liability Chasm Is a Verifiability Problem — And the Log Is the Battleground

On-chain verifiability repairs exactly two of those five. Signed guardrail attestations and immutable configuration snapshots address moral hazard, because an insurer can verify deployment posture instead of trusting a questionnaire. Persistent tamper-evident traces begin generating the loss data severity curves require — over years, not quarters. It does nothing about correlation and does not bound the tail. The blockchain fixes the measurement problem, not the correlation problem. Anyone selling on-chain agent insurance as a solved category is selling half a sentence.

Watch what happens next, because we have watched it before. When capital wants to price a risk it does not understand, it invents a curve. Aave and Compound's interest rate models are the clearest example: utilization-based slopes presented as market discovery, in practice arbitrary parameters tuned by governance vote rather than any observable equilibrium. The moment agent risk becomes a tradable line item, expect an agent risk premium curve with kink points, a governance vote on the slope, and no loss history underneath any of it. The difference is that a mispriced lending curve redistributes yield. A mispriced liability curve accumulates unrecognized tail exposure until one event reprices the whole book.

One technical note the on-chain crowd will get wrong. The instinct will be to reach for a dedicated data availability layer for agent traces. Do not. An enterprise agent session generates kilobytes to megabytes of structured trace data, orders of magnitude below what justifies purpose-built DA. You need an append-only anchor with verifiable inclusion proofs, not a bandwidth-optimized data layer with a token attached. Hash the trace, anchor the root, store the payload in tiered storage with retention guarantees. That is the entire architecture.

Contrarian

Here is what should worry anyone treating verifiability as a clean answer. An immutable log is an exhibit. On-chain evidence cannot be quietly deleted before discovery, which makes it the only architecture that can plausibly support underwriting — and simultaneously the only one that hands a plaintiff a timestamped, cryptographically attested record of every decision the agent made. The same property that makes an agent insurable makes it maximally discoverable. Some enterprises will run that arithmetic and go the other direction: lower observability, shorter retention, less evidence to produce. When liability pressure exceeds insurability, the rational move is not caution. It is amnesia. That reflexivity is the real systemic risk, and no liability regime currently models it.

Which raises the unglamorous point. Culture is the ultimate consensus mechanism, but the consensus that matters here is between legal, actuarial, and engineering teams who do not share a vocabulary. Building bridges where others build walls is not sentiment; it is the actual work of making a deployment posture legible to an underwriter without converting the engineering team into a compliance function.

Takeaway

Wrong question: when does the liability chasm close. It closes when measurement, correlation, and tail-bounding all resolve, and one of those has no known solution. Right question: under the constraint that agent risk stays partially uninsurable, how does risk get redistributed — into contract terms, into authorization protocols, into architectures that make evidence production optional, and onto the balance sheets of whoever can absorb the tail. The log is the battleground. Whoever defines the authorization standard defines liability, and that standard is being written right now, mostly by people who are not thinking about it.