The Memory War: SanDisk's HBF and the Decentralization of AI Inference

Bentoshi
Video
In an investor day slide that sparked more heat than light, SanDisk presented a comparison between its proposed High Bandwidth Flash (HBF) and the industry-standard HBM memory. The slide claimed that with HBF, the same AI inference workload could be served by fewer GPUs, thanks to higher capacity per stack. But analyst Zephyr from Citrini quickly pointed out a discrepancy: SanDisk had used a conservative 12.8 TB/s bandwidth figure for HBM, corresponding to today's HBM3E, while the more realistic future HBM4E would deliver 32 TB/s — effectively tripling the performance. The debate was framed as a technical footnote, but it reveals a deeper tension: the conflict between centralized, high-performance solutions and the push for open, scalable alternatives. In blockchain, we face the same battle every day. We code the trust, but we must audit the soul. At its core, the HBF versus HBM controversy is not about memory specifications; it is about who controls the memory stack. HBM is a JEDEC standard, tightly controlled by a handful of DRAM manufacturers — SK Hynix, Samsung, Micron — and packaged using TSMC's CoWoS technology. The supply chain is a fortress, with high barriers to entry and pricing power concentrated in the hands of a few. HBF, on the other hand, is SanDisk's attempt to repurpose its NAND flash manufacturing capacity into a high-bandwidth, high-capacity memory solution. It is not a JEDEC standard; it is a proprietary proposal. The analogy to blockchain is immediate: HBM is like a centralized validator set with high throughput but locked governance, while HBF is a more permissionless alternative — slower, but with the potential for broader participation and lower cost. In my years auditing DeFi protocols, I've seen the same tension play out in smart contract design. Consider oracle feeds: the fastest, most reliable oracles are often centralized, like Chainlink's price feeds. But as I argued in my 2020 whitepaper "Liquidity as Liberty," speed without decentralization is a hollow promise. The 2022 bear market exposed that fragility when a single oracle failure could cascade into liquidations. SanDisk's HBF faces a similar critique: it offers higher capacity and lower cost per gigabyte, but at the cost of latency that is orders of magnitude higher. For AI inference, where the model is loaded once and then queried, latency is less critical. For training, where every nanosecond counts, HBF is unusable. The technical reality is that HBF is not a direct competitor to HBM; it is a complementary layer — a memory tier for capacity-sensitive workloads, much like how Layer 2 solutions complement Layer 1 in blockchain. Yet the controversy matters because it reveals the narrative battle. SanDisk chose to compare HBF against a conservative HBM3E configuration: 192 GB capacity, 12.8 TB/s bandwidth. Zephyr's counterexample used HBM4E with 512 GB and 32 TB/s. The difference is not just quantitative; it is strategic. By setting the baseline lower, SanDisk made HBF look more competitive. This is a classic framing trick — one I have seen in countless blockchain whitepapers where a new protocol compares itself against a straw-man version of Ethereum. The real question is: what is the appropriate benchmark? For AI inference in 2026, the cutting edge is HBM4E with FP4 quantization, which can squeeze a 480B-parameter MoE model like Qwen3-480B-A35B into 240-480 GB. That is within the range of 512 GB HBM4E. So SanDisk's claim that HBF saves GPUs is valid only if the model cannot fit in HBM. But with quantization, it can. This is where the blockchain parallel deepens. In decentralized storage, we see a similar tension between capacity and speed. Arweave offers permanent storage with high latency, while IPFS provides faster retrieval but with caching dependencies. Both serve different use cases. HBF, with its NAND-based architecture, is closer to a decentralized storage layer than to a high-frequency memory bus. It is designed for workloads where capacity trumps latency — exactly the profile of AI inference at scale, where models are loaded once and served to millions of users. The irony is that the same property that makes HBF attractive — high capacity — also makes it a candidate for on-chain AI, where models must be stored and executed in a trustless environment. Imagine a smart contract that runs a large language model; it cannot afford HBM's cost per gigabyte. HBF could be the middleware that makes decentralized AI economically viable. But the contrarian angle is that HBF, despite its promise, is not a decentralized solution. It is a proprietary design from a single company, SanDisk (now part of Western Digital). The controller logic, the high-bandwidth interface, and the 3D stacking are all closely guarded. If HBF were to become the de facto standard for AI inference memory, we would trade one oligopoly (DRAM) for another (NAND). The real victory would be a truly open memory standard, one that is auditable, permissionless, and resistant to censorship. In blockchain, we have learned that the protocol is neutral, but the user is human. A proprietary memory solution can be used to enforce compliance, as we saw with USDC's ability to freeze addresses. Circle's compliance-first strategy is its biggest risk: it can freeze any address within 24 hours — how is that decentralized? Similarly, if HBF is controlled by a single entity, the memory itself becomes a vector for control. My experience in 2021, when I curated a digital exhibition on Tezos to emphasize carbon-neutral minting, taught me that the medium matters. The choice of memory is not just a technical decision; it is an ethical one. HBM, for all its centralization, is a mature standard with multiple vendors. HBF, if it gains traction, could create a new dependency. The deeper truth is that the AI industry's memory bottleneck is not just about bandwidth; it is about trust. Whose memory are we using? Who can revoke access? These questions are as relevant to blockchain as they are to AI. During the 2022 bear market, I took a six-month sabbatical to reassess the industry's direction. What I realized was that every layer of the stack — from consensus to storage to memory — must be designed with governance in mind. The collapse of FTX was not a failure of technology; it was a failure of accountability. The same applies to memory. We need a memory architecture that is not only performant but also transparent. In a world of ledgers, who holds the memory? In 2026, as I lead a consortium to design a decentralized identity framework for AI agents, I see the HBF vs HBM debate as a microcosm of a larger struggle. The future of AI inference will be shaped by what memory we choose. If we choose capacity at the cost of openness, we risk repeating the mistakes of the past. But if we can build a memory layer that is both high-capacity and decentralized — perhaps using a network of NAND-based nodes with verifiable proofs — then we can realize the dream of sovereign AI. Proof is binary; meaning is fluid. The takeaway is not that HBF is better or worse than HBM. It is that we must look beyond the benchmark numbers and ask: who controls the memory? The answer will determine whether the AI revolution is truly open or just another garden of walls. We are not moving money; we are moving belief. And belief requires a memory that cannot be erased.

The Memory War: SanDisk's HBF and the Decentralization of AI Inference

The Memory War: SanDisk's HBF and the Decentralization of AI Inference