The spread was real, but the exit was imaginary.
SanDisk, the veteran NAND flash manufacturer now spinning off from Western Digital, dropped a number that barely made a ripple in crypto circles but sent a low-frequency tremor through the data center storage floor: by 2030, KV cache workloads will drive 35% of all NAND usage in AI data centers.
Most people read that and think "more SSDs." I read it and see a structural shift in how AI inference will be costed, scaled, and potentially bottlenecked. The number itself is not the story — the assumptions buried beneath it are where the real alpha lives.
Context: Who Is SanDisk and Why Should You Care?
SanDisk is not a startup. It's a NAND flash IDM (integrated device manufacturer) with deep ties to Kioxia (formerly Toshiba Memory) in Japan. They co-develop the BiCS FLASH series, currently at 200+ layers (BiCS 8), with 300+ layers on the roadmap. They are not the market leader — Samsung holds ~30% of NAND revenue, SK Hynix/Solidigm ~20%, and SanDisk/Kioxia sits around 13-15% — but they are a top-tier player with a strong enterprise SSD controller and firmware stack.
Their public prediction about KV cache workloads is not a random press release. It's a strategic signal. After the split from Western Digital, SanDisk needs a new narrative to justify its capital expenditure to investors. The AI data center is that narrative. KV cache — the memory bottleneck in large language model inference — is the wedge.
KV cache stores the key-value pairs from previous tokens in a transformer model to avoid recomputation. The longer the context window and the more concurrent users, the larger the cache grows. Today, it lives in HBM or DRAM — expensive, power-hungry, and limited in capacity. The holy grail is offloading it to cheap, dense NAND flash without destroying latency.
SanDisk is betting that by 2030, the cost of keeping KV cache in DRAM will be so prohibitive that hyperscalers will accept the latency penalty of NAND-based storage. That's a bet on the economics of inference, not just on NAND density.
Core: The Order Flow Analysis Behind the 35% Figure
Let's decompose the 35% number. It does not mean 35% of total storage capacity. It means 35% of NAND workloads — i.e., the number of read/write operations, IOPS, and bandwidth consumed by KV cache operations relative to all other NAND tasks in the data center (training checkpoints, model weights, logs, databases, etc.).
This is a radical claim. Currently, KV cache is a negligible fraction of NAND usage. For it to reach 35% by 2030, several things must happen in sequence:
- Context window explosion: Models must move from 128K tokens to 1M+ tokens as standard. This is already happening (Gemini 1.5 Pro, Claude 3, GPT-4 Turbo). Longer context means larger KV cache per request.
- Concurrent user growth: Agentic AI, copilots, and real-time assistants will generate millions of simultaneous inference requests. Each request's KV cache must be stored until the session ends.
- Cost per token must drop 10x-100x: The entire point of offloading KV cache to NAND is to reduce the cost of inference. If NAND SSDs are too expensive (or too slow), the math fails. SanDisk is implicitly betting that QLC (4-bit/cell) NAND with high endurance and low latency can be manufactured at scale, driving down $/GB to a point where it beats DRAM for this specific workload.
Alpha decays faster than the code that finds it.
I ran a back-of-the-envelope calculation based on typical 70B-parameter model inference. At 32K context, KV cache per request is about 1.5 GB (using FP16). For 1M concurrent users, that's 1.5 PB of KV cache. DRAM cost at $5/GB = $7.5M for the memory alone. NAND at $0.10/GB = $150K. The gap is 50x. Even if NAND latency is 1000x slower, you can hide it with prefetching, tiering, and batching. The trade-off is clear.
SanDisk's 35% workload share implies that a significant portion of that 1.5 PB of KV cache will be served from NAND, not DRAM. But here's the catch: NAND has limited endurance. KV cache is write-heavy (each token generates new KV pairs). A QLC SSD can sustain ~1000 program/erase cycles. If the cache is written and rewritten frequently, the SSD dies fast. The solution is either high-endurance TLC (more expensive) or a hybrid TLC/QLC tiering scheme. SanDisk's controllers must be smart enough to handle this.
I trust the log, not the hype.
SanDisk has a history of overpromising on QLC endurance. In 2018, they claimed QLC would be ready for enterprise workloads by 2020. It took until 2023 for QLC to gain real traction (Solidigm's D5-P5336). The 35% prediction is aggressive, but not impossible. The real constraint is not NAND density — it's the controller firmware, the PCIe Gen5/Gen6 bandwidth, and the CXL (Compute Express Link) memory-semantic interface that allows NAND to be treated as main memory. Without CXL, KV cache offloading to NAND is too slow. SanDisk must support CXL 3.0 or later.

Contrarian: The Blind Spot That SanDisk Ignores
The bot didn't fail; the market changed rules.
The 35% prediction assumes that the current transformer architecture remains dominant. What if the industry moves to state-space models (Mamba, RWKV) or linear attention mechanisms that eliminate the KV cache entirely? Or what if specialized hardware (e.g., Groq's LPUs, Cerebras wafer-scale chips) makes KV cache so small that it all fits in on-chip SRAM? In that scenario, the NAND demand for KV cache collapses.
SanDisk's bet is a bet on the persistence of the transformer bottleneck. It's a reasonable bet — transformers are entrenched, and the engineering effort to replace them is enormous — but it's not a sure thing.
Another blind spot: Hyperscaler in-house silicon. AWS (Trainium, Inferentia), Google (TPU), and Microsoft (Maia) are designing their own accelerators with custom memory hierarchies. They could build KV cache into dedicated on-package SRAM or HBM stacks, bypassing NAND entirely. If they do, the 35% figure evaporates.
Liquidity is a mirage during the storm.
SanDisk is also ignoring the geopolitical friction. The NAND supply chain is concentrated in Japan and Korea, with a growing Chinese competitor (YMTC) under sanctions. If the US escalates export controls on NAND equipment (unlikely but possible), SanDisk's capacity expansion could be delayed, pushing the 35% timeline to 2032-2034. More importantly, hyperscalers are diversifying away from single vendors. They will not lock into SanDisk for a mission-critical workload like KV cache. They will require multiple sources (Samsung, SK Hynix, Micron). SanDisk's 13-15% market share means it can only capture a fraction of that 35% workload anyway.
Takeaway: Where the Real Money Hides
The blind spot is where the money hides.
SanDisk's 35% prediction is not a forecast — it's a rallying cry for its own product roadmap. The real alpha is not in buying SanDisk stock (or its future listed entity). It's in identifying the companies that build the middle layer: CXL controllers, NAND controller IP, and high-endurance QLC die suppliers.
If you believe the KV cache offloading thesis, buy the picks and shovels: companies like Astera Labs (CXL retimers), Rambus (memory controllers), and maybe even Smart Global Holdings (CXL memory modules). The NAND itself is a commodity; the intelligence to manage it is where the margin lives.
Ultimately, SanDisk's statement is a directional bet on the AI inference cost curve. The 35% number will be wrong in the details — it's always off by a few years or a few percentage points. But the direction is correct: KV cache will be a dominant NAND workload, and the infrastructure to support it will be worth billions.
We optimize for edges, not comfort.
My advice: ignore the headline. Read the fine print. The real signal is that SanDisk is shifting its entire R&D focus to enterprise QLC and CXL. That's a bet worth watching.