Null Fields, Full Formatting: The Hallucination Pipeline Inside AI Crypto Research

0xPlanB
Industry

I received a research report last Tuesday. Nine sections. Forty-one tables. Roughly four thousand words. The header promised a "full-spectrum teardown" of a protocol I had not heard of.

Every field was null.

Not empty. Null. An empty field is a gap. A null field is a value that was never assigned. The report had been generated by a pipeline that received no input, and it responded by producing the shape of an analysis without the substance. Technical score: N/A. Token supply: N/A. Team assessment: N/A. Risk matrix: six rows, six unknowns. The document closed with a disclaimer about doing your own research and a recommendation to "pause the process and return to stage one."

That last line was the only true statement in the document. It was also the tell.

The pipeline's heart.

The validation layer fired. The system above it shipped anyway.

I have spent nine years auditing systems that lie to their users. Most of them lie by omission. This one lied by formatting. It is the cleanest example I have seen of a failure mode that is about to become the default in crypto research: the artifact that looks like an answer because it has the shape of one.

The Industrialization of Crypto Research

Between 2023 and 2026, the volume of published crypto "analysis" grew by roughly an order of magnitude. I track this crudely. The number of distinct token reports indexed by three aggregators I monitor moved from about forty thousand to something north of four hundred thousand. The number of humans capable of independently verifying a token's on-chain claims did not grow. If anything, it shrank. The 2022 drawdown pushed a generation of analysts out of the industry and they did not come back.

The gap between the volume of claims and the capacity to verify them is a market. Markets get filled. The filler, in this cycle, is automation.

The economics are simple. A human analyst produces two to four deep reports a month. A pipeline produces two to four hundred a day. The publisher is paid per report, or per pageview, or per "insight credit" sold to a fund. None of those units of account reward accuracy. They reward throughput.

I understand the demand side because I have been on it. In 2020, during the first DeFi summer, I wrote a Python simulation of Compound's interest rate model to test its behavior under stress. I found a theoretical liquidation cascade in the oracle pricing path. The model held in live conditions, which is why the report was dismissed. But the exercise taught me something about research economics. The value of a risk report is only realized when the risk materializes. Until then, it is a cost.

Automated research does not change that incentive. It amplifies it. A pipeline that produces confident output is more valuable to its operator than a pipeline that produces correct refusals, right up until the confident output causes a loss.

The Pipeline, Stage by Stage

I have audited three of these systems. They vary in implementation and converge in architecture. The canonical pipeline has six stages.

Ingestion. The system pulls from RPC endpoints, Dune queries, GitHub repositories, token lists, and price oracles. This layer is deterministic and cheap. It rarely fails loudly. If an endpoint is down, the stage returns empty.

Extraction. Raw chain data becomes "metrics." Total value locked. Holder count. Unlock schedule. Governance participation. This is the first stage where assumptions enter. Every Dune query has a WHERE clause. The WHERE clause encodes a worldview. A query that excludes contract addresses below a threshold silently deletes a cohort of holders. Nobody reads the WHERE clause. It is not in the report.

Scoring. The nine-dimension framework. Technical, tokenomics, market, ecosystem, regulatory, team, risk, narrative, and supply-chain. Each dimension receives a score. The scores are not measured. They are inferred from the extracted metrics by a model that was trained on other reports. The model has learned the distribution of scores, not the distribution of reality.

Formatting. The scores are rendered into tables, headers, star ratings, and disclaimers. This stage is where the artifact acquires authority. A table with nine rows reads as nine findings.

Validation. A rule engine checks for nulls, internal contradictions, and unsupported claims. This is the gate.

Publication.

The failure mode lives in the ordering. In most production pipelines, formatting is stage four and validation is stage five. Formatting runs first. By the time the gate fires, the document already exists. The gate can block publication. It cannot un-generate the artifact. And in the pipeline I received, the gate blocked nothing. It wrote its refusal into the output, and the output shipped.

The dependency graph is inverted. Facts should be upstream of format. In practice, format is upstream of facts, and facts are optional. If the input is null, the schema still fills. It fills with N/A. And N/A, rendered in a table, looks like a data point that was measured and found to be unavailable. It was never measured.

That is the dependency graph's heart.

The ecosystem dimension illustrates the point. Automated frameworks score an L2's ecosystem by counting deployed chains. This accidentally tracks reality better than the technical dimension does. The real difference between the OP Stack and the ZK Stack is not cryptographic. It is who convinces more projects to deploy first. The pipeline scores the outcome, not the architecture. It gets the right answer for the wrong reason, and it will keep getting the right answer until the deployment race ends.

The market dimension has a similar accidental accuracy. Automated frameworks flag "liquidity fragmentation" as a risk, which is what the venture narrative wants them to flag. Fragmentation is not the problem. It is the sales pitch for the next aggregator. The pipeline has learned the pitch. It reproduces the pitch. The report that flags fragmentation is also the report that recommends the product that solves it, because both sentences appeared in the training data.

The Completeness Heuristic

A fixed schema produces a completeness heuristic. If a report has nine sections, a reader assumes nine sections of information. The reader cannot distinguish "nine sections of data" from "nine sections of scaffolding." The schema launders absence into presence.

This is not a new trick. I saw it in 2021, in NFT metadata. I audited ten mid-tier ERC-721 contracts and found that seventy percent stored their critical assets on centralized servers, S3 buckets, Cloudflare-fronted endpoints, a handful of IPFS pins with no replication. The metadata JSON had every field. Name, description, attributes, image URI. The schema was complete. The decentralization was zero. Marketplaces rendered both identically, and buyers paid for the rendering.

I sent that report to eleven projects. Two responded. Neither disputed the server logs. One argued that "decentralization is a spectrum." I agreed and asked for the spectrum's coordinates. No reply.

The same pattern now repeats in AI-generated research. The schema is complete. The analysis is zero. The reader renders both identically. The difference is that the NFT case had a verifiable ground truth, the server logs. In the AI research case, the ground truth is often another AI summary, and the summary cites the first report. The loop closes. Nothing was ever measured.

I call this null-field laundering. A null is generated at the extraction stage, formatted at stage four, validated at stage five, and published at stage six. By the time it reaches a reader, it has acquired the texture of a finding. It has a row, a label, and a font.

What Null-Field Reports Do to Prices

In early 2022 I modeled UST's seigniorage flow as a feedback system. I wrote a geometric proof of the de-peg condition under high volatility. It was downvoted for being "too abstract." Three weeks later the mechanism executed the proof in real time. The abstraction was the point. The proof was about structure, not price.

The AI research pipeline has the same structural flaw in reverse. It is never too abstract. It is concrete about things that do not exist. It will report the unlock schedule of a token that has no contract, because the schema requires an unlock schedule and the model has seen a thousand unlock schedules.

I tracked three tokens that received automated "analysis" in a null-input state over a thirty-day window. I will not name them. The point is structural, not accusatory. In each case the report contained at least one fabricated metric, a TVL figure, a holder count, a partnership. The fabrications were plausible. They sat within an order of magnitude of the real values. That is worse than being wildly wrong, because plausible errors survive review. A number that is off by a factor of a hundred gets caught. A number that is off by fifteen percent gets cited.

Two of the three tokens showed abnormal volume within forty-eight hours of publication. I cannot establish causation from three samples, and I will not pretend otherwise. I can state that the correlation is not zero, and that the reports contained no verified data. In a market where the marginal buyer is a retail participant reading a summary, the summary is the fundamental.

The Regulatory Dimension

The regulatory scoring dimension deserves a separate note. Most automated frameworks score compliance by checking for a KYC gate, a jurisdiction disclaimer, or a legal entity. These are presence checks, not function checks.

I have written about this before. Most project KYC is theater. Buying a few wallet holdings bypasses it. The compliance cost is passed entirely to honest users, who complete the process while the sybil wallet does not. An automated framework that scores "regulatory risk" by detecting a KYC page is scoring the theater, not the compliance.

Null Fields, Full Formatting: The Hallucination Pipeline Inside AI Crypto Research

The pipeline does not know this. It has learned that KYC pages correlate with lower regulatory scores in its training data. So it rewards the page. The page is cheap. The incentive is to build the page, not the compliance.

I audited an AI-agent framework in 2026 that executed on-chain transactions through smart wallets. I found a race condition that allowed an agent to bypass multi-sig requirements under specific latency conditions. The report triggered regulatory interest because it provided a concrete technical basis for a rule. That is the correct sequence. Technical finding, then rule. The automated research pipeline inverts it. It generates the rule-shaped output first and back-fills the finding. Regulators reading those reports are reading a schema, not a system.

Why Nobody Fixes It

Three buyers sustain the market.

The fund analyst screening two hundred tokens a week uses the report as a filter, not a source. The filter's error rate is invisible to them because they never check the rejects. A false negative costs them a missed position. A false positive costs them money, but the false positive looks identical to a true positive until it does not.

The retail reader consumes the star rating. The star rating was inferred from nulls. It is a random number with a font.

The protocol itself. Coverage is coverage. A "neutral" report with a two-star technical score still puts the name in front of readers. The marketing team will screenshot the table and crop the stars.

The publisher is paid by volume. Nobody in the chain is paid to be right. That is the incentive misalignment. It is not a bug in the model. It is the model.

The mechanism's heart.

What the Bulls Got Right

Here is the counter-intuitive part, and I will state it plainly because the evidence supports it.

The pipeline I received refused to hallucinate. That is rarer than it should be. Most systems, handed an empty input, will invent. They will infer a TVL from a market cap and label it "estimated." They will guess a team from LinkedIn. They will generate a partnership from a co-marketing tweet. The validation layer in this pipeline held. It wrote N/A in every field and it told the operator to return to stage one.

The framework was also sound. Nine dimensions is a defensible decomposition of a protocol. The risk matrix had the correct rows. The disclaimers were accurate. The failure was not in the design of the analysis. It was in the design of the publication gate.

And the humans are not better. I have read analyst reports with the same null-field problem, confident prose wrapped around an unverified assumption, delivered over a week instead of four seconds. The difference is latency, not honesty. Scale is the variable.

The framework's heart.

There is a version of this technology that works. It requires one change. Move validation upstream of formatting. If the gate runs before the render, the pipeline produces fewer reports and better ones. If it runs after, it produces a disclaimer. A disclaimer is not a gate. It is a receipt.

The Question That Matters

The question is not whether AI writes crypto research. It will. The question is where the validation gate sits relative to the publication gate.

Watch for pipelines that publish their null rate. A system that reports how often it refuses to answer is a system you can trust with the answers it does give. A system that never refuses is not analyzing. It is formatting.

The report I received last Tuesday had a null rate of one hundred percent. It also had a section count of nine. Only one of those numbers was measured. The other was designed.