English
Back
Open Account
链捕手ChainCatcher
wrote a column · Jul 28 04:02

From Hot Storage to Cold Memory: Decentralized Storage in the Midst of the AI-Driven Storage Boom

Author: Jacob Zhao @ IOSG
Today, the so-called 'first domestic memory stock'ChangXin Memory Technologiesofficially listed on the ChiNext board and stunned the market with an explosive 500% surge. Despite recent sector-wide volatility from ongoing corrections, AI storage continues to be aggressively repriced by capital amid the current wave of tech-driven narratives.Web3 Meanwhile, decentralized storage has fallen into prolonged silence and disappointment. Why do two sectors bearing the same 'storage' label exhibit such starkly divergent market performances? The fundamental reason lies in a complete divergence of their underlying value functions.
The AI-era revaluation of storage is essentially a frenzy centered on 'hot data efficiency,' aimed at maximizing compute utilization and commercial monetization. In contrast, decentralized storage adheres to a value proposition built on 'cold data trustworthiness,' safeguarding data equity, censorship resistance, and humanity's long-term collective memory. The former is an efficiency system for hot data; the latter is a trust system for cold data. Capital markets today undeniably stand firmly on the side of 'efficiency,' yet human civilization will ultimately still require an immutable foundation for memory. The long-term value of trustworthy cold storage has never vanished—it merely lies dormant in the shadow of the cycle, awaiting rediscovery and repricing by a new era.
Why storage has once again become a focal point in the AI supply chain
In the traditional IT era, storage was a 'capacity business.' Enterprises CIOs focused on cost per unit of capacity, hard drive reliability, disaster recovery solutions, archiving strategies, and equipment refresh cycles spanning three to five years. Storage was viewed as an accessory bundled with server procurement.
The current storage boom is not a traditional cyclical recovery, but rather a repricing of data mobility driven by AI. In the era of large models, the logic of storage has fundamentally shifted from 'capacity-first' to 'efficiency-above-all,' with intense focus on extreme metrics such as GPU feeding rates, checkpoint writes, and RAG ultra-low latency. This marks storage’s evolution from being merely the 'final parking spot for data' to becoming the 'high-speed conduit through which data enters computation.'
The shifting bottlenecks in AI infrastructure essentially reflect a race to fix the weakest link in the 'barrel effect.' Real computational utilization is not a linear sum of individual components, but governed by a stringent multiplicative relationship: real compute utilization = GPU × HBM × DRAM × SSD × networking × file system. A deficiency in any single component can cause the entire system’s compute utilization to collapse. For the first time in the AI era, storage has transformed from a 'cost center' into an 'efficiency engine.' This is the fundamental logic behind storage’s repricing.
Author: Jacob Zhao @ IOSG    Today, the so-called 'first domestic storage stock'ChangXin Memory Technologiesofficially debuted on the ChiNext board, surging an astonishing 500% and electrifying the market. Although the broader storage sector continues to face headwinds from recent pullbacks, AI-related storage remains at the center of today’s tech narrative, undergoing aggressive revaluation by capital markets.Web3 Meanwhile, decentralized storage has fallen into prolonged silence and disillusionment. Why do two sectors both labeled as 'storage' exhibit such starkly divergent market performances? The fundamental reason lies in a complete divergence of their underlying value functions. The revaluation of storage in the AI era is essentially a frenzy centered on 'hot data efficiency,' driven by the pursuit of maximizing computing utilization and commercial monetization. In contrast, decentralized storage adheres to a value proposition built on 'cold data trustworthiness,' safeguarding data fairness, censorship resistance, and humanity’s long-term memory. The former constitutes an efficiency system for hot data; the latter, a trust system for cold data. Capital markets today clearly favor 'efficiency,' yet human civilization ultimately still requires an immutable foundation for memory. The long-term value of trustworthy cold storage has never vanished—it merely lies dormant in the shadows of the cycle, awaiting rediscovery and repricing by history.    Why Storage Has Reemerged as a Focal Point in the AI Supply Chain    In the traditional IT era, storage was a 'capacity business.' Enterprises' CIO focuses on cost per unit of capacity, hardware...
AI Storage Architecture Overview: From HBM Bandwidth Organs to Data Lake Foundations
AI storage is far more than just stacking hardware—it is a tightly integrated, hierarchically scheduled, and highly complex system. Within this architecture, industrial value and investment focus are heavily concentrated on HBM, enterprise-grade SSDs, SSD controllers, NVMe/CXL protocols, and high-performance storage systems. To clearly map the flow of value, we divide the AI storage architecture into four core layers from top to bottom:
Author: Jacob Zhao @ IOSG    Today, the so-called 'first domestic storage stock'ChangXin Memory Technologiesofficially debuted on the ChiNext board, surging an astonishing 500% and electrifying the market. Although the broader storage sector continues to face headwinds from recent pullbacks, AI-related storage remains at the center of today’s tech narrative, undergoing aggressive revaluation by capital markets.Web3 Meanwhile, decentralized storage has fallen into prolonged silence and disillusionment. Why do two sectors both labeled as 'storage' exhibit such starkly divergent market performances? The fundamental reason lies in a complete divergence of their underlying value functions. The revaluation of storage in the AI era is essentially a frenzy centered on 'hot data efficiency,' driven by the pursuit of maximizing computing utilization and commercial monetization. In contrast, decentralized storage adheres to a value proposition built on 'cold data trustworthiness,' safeguarding data fairness, censorship resistance, and humanity’s long-term memory. The former constitutes an efficiency system for hot data; the latter, a trust system for cold data. Capital markets today clearly favor 'efficiency,' yet human civilization ultimately still requires an immutable foundation for memory. The long-term value of trustworthy cold storage has never vanished—it merely lies dormant in the shadows of the cycle, awaiting rediscovery and repricing by history.    Why Storage Has Reemerged as a Focal Point in the AI Supply Chain    In the traditional IT era, storage was a 'capacity business.' Enterprises' CIO focuses on cost per unit of capacity, hardware...
Compute-Proximal Memory Layer (Bandwidth Core): Dominated by HBM, complemented by DRAM and CXL-based memory pooling technologies. This layer sits directly adjacent to GPU/CPU packages or buses, aiming to break through the 'memory wall'—it is the first critical gate determining whether computational power can be fully unleashed.
High-Speed Persistent Storage Layer (I/O Hub): The core logic here is enterprise-grade SSD = NAND chips + SSD controller + NVMe/PCIe data pathway. This layer handles high-frequency checkpoint writes, massive training dataset loading, and caching of hot RAG data, representing the clearest incremental demand for persistent storage in AI data centers.
Low-cost, high-capacity storage tier (capacity foundation): Composed of HDDs, cold storage, and data lake archival systems. Despite the exponential growth of multimodal raw data, historical logs, and compliance backups, this layer continues to offer an irreplaceable TCO (Total Cost of Ownership) advantage.
AI storage systems and data software (orchestration brain): Includes high-performance parallel file systems, distributed object storage, vector databases, and RAG-based data governance layers. What AI truly consumes is not raw hardware but data that has been efficiently organized, indexed, and permissioned by the software stack to ensure usability.
As an ecosystem extension, decentralized storage does not directly compete in the millisecond-level race for AI hot data. Instead, it focuses on public dataset attestation, AI training data provenance, and long-term cold memory archiving, establishing its unique ecological niche as a 'trusted cold tier.'
HBM: The 'bandwidth organ' closest to compute in the AI storage hierarchy
High Bandwidth Memory (HBM)Not traditional storage, but a high-bandwidth memory layer adjacent to the GPU. Its core purpose is not data retention but continuously 'feeding' data to compute units at extremely high bandwidth. HBM is the segment of the AI storage chain nearest to compute and with the highest determinism—it directly determines whether GPUs can be fully utilized and represents the most critical current bottleneck in the supply chain.
The core architecture of HBM is '3D DRAM stacking + 2.5D advanced packaging': Through TSV vertical stacking and CoWoS heterogeneous integration, it minimizes the distance between memory and compute, achieving a generational leap in bandwidth. Its industrial barrier lies not only in DRAM design but in a comprehensive systems engineering challenge encompassing DRAM process technology, TSV, ultra-thin stacking, packaging, thermal management, testing, and customer certification. A yield defect in any single component can result in the failure of the entire HBM stack.
Currently, only globally SK hynix, Samsung, Micron The three industry leaders have achieved stable mass production, establishing a triple moat through leading-edge DRAM process technology, advanced packaging capabilities, and customer certifications from NVIDIA and AMD.
Author: Jacob Zhao @ IOSG    Today, the so-called 'first domestic storage stock'ChangXin Memory Technologiesofficially debuted on the ChiNext board, surging an astonishing 500% and electrifying the market. Although the broader storage sector continues to face headwinds from recent pullbacks, AI-related storage remains at the center of today’s tech narrative, undergoing aggressive revaluation by capital markets.Web3 Meanwhile, decentralized storage has fallen into prolonged silence and disillusionment. Why do two sectors both labeled as 'storage' exhibit such starkly divergent market performances? The fundamental reason lies in a complete divergence of their underlying value functions. The revaluation of storage in the AI era is essentially a frenzy centered on 'hot data efficiency,' driven by the pursuit of maximizing computing utilization and commercial monetization. In contrast, decentralized storage adheres to a value proposition built on 'cold data trustworthiness,' safeguarding data fairness, censorship resistance, and humanity’s long-term memory. The former constitutes an efficiency system for hot data; the latter, a trust system for cold data. Capital markets today clearly favor 'efficiency,' yet human civilization ultimately still requires an immutable foundation for memory. The long-term value of trustworthy cold storage has never vanished—it merely lies dormant in the shadows of the cycle, awaiting rediscovery and repricing by history.    Why Storage Has Reemerged as a Focal Point in the AI Supply Chain    In the traditional IT era, storage was a 'capacity business.' Enterprises' CIO focuses on cost per unit of capacity, hardware...
DRAM and CXL: The Foundation of System Memory and the Engine for Memory Pooling
HBM addresses extreme bandwidth demands near the GPU, DRAM underpins the server system memory foundation, and CXL aims to break physical boundaries and reshape how memory resources are organized in data centers.
DRAMDRAM primarily handles CPU-side caching, data preprocessing, intermediate state buffering, and system operations, forming the most fundamental system memory layer in servers. The global DRAM market is highly concentrated among the three giants—SK hynix, Samsung, and Micron—with ChangXin Memory Technologies (CXMT) representing the key variable in China’s domestic DRAM substitution efforts.
CXL (Compute Express Link)CXL is a next-generation cache-coherent interconnect protocol designed for data centers, aiming to overcome limitations imposed by traditional DIMM slots, local memory capacity constraints, and isolated server memory resources. It drives memory architecture toward scalability, pooling, and sharing. CXL is currently in an early phase, transitioning from platform support to large-scale deployment, with significant architectural value over the medium to long term. Key companies include Astera Labs and Montage Technology.
Author: Jacob Zhao @ IOSG    Today, the so-called 'first domestic storage stock'ChangXin Memory Technologiesofficially debuted on the ChiNext board, surging an astonishing 500% and electrifying the market. Although the broader storage sector continues to face headwinds from recent pullbacks, AI-related storage remains at the center of today’s tech narrative, undergoing aggressive revaluation by capital markets.Web3 Meanwhile, decentralized storage has fallen into prolonged silence and disillusionment. Why do two sectors both labeled as 'storage' exhibit such starkly divergent market performances? The fundamental reason lies in a complete divergence of their underlying value functions. The revaluation of storage in the AI era is essentially a frenzy centered on 'hot data efficiency,' driven by the pursuit of maximizing computing utilization and commercial monetization. In contrast, decentralized storage adheres to a value proposition built on 'cold data trustworthiness,' safeguarding data fairness, censorship resistance, and humanity’s long-term memory. The former constitutes an efficiency system for hot data; the latter, a trust system for cold data. Capital markets today clearly favor 'efficiency,' yet human civilization ultimately still requires an immutable foundation for memory. The long-term value of trustworthy cold storage has never vanished—it merely lies dormant in the shadows of the cycle, awaiting rediscovery and repricing by history.    Why Storage Has Reemerged as a Focal Point in the AI Supply Chain    In the traditional IT era, storage was a 'capacity business.' Enterprises' CIO focuses on cost per unit of capacity, hardware...
Enterprise SSDs: The Data Hub Built on NAND, Controllers, and NVMe
Enterprise SSDs represent the most critical high-throughput persistent storage increment in AI data centers. With extremely high throughput, ultra-low latency, and consistent QoS, they continuously feed data to GPUs throughout the entire AI lifecycle—including training data loading, checkpoint writes, RAG retrieval, inference caching, and log backfilling.
In AI storage architectures, SSDs are not standalone hardware components but highly integrated systems that can be distilled into an industry formula: Enterprise SSD = NAND dies + SSD controller + NVMe/PCIe data pathways. These three layers correspond to distinct segments of the industrial chain:
NAND die (raw material layer): Determines storage density and cost per unit; the controller governs performance delivery and lifespan management. Representative companies include Samsung, SK hynix (Solidigm), Micron, Kioxia, Western Digital, and Yangtze Memory Technologies (YMTC).
SSD controller (performance enablement layer): Governs performance delivery, error correction, QoS stability, and wear leveling. Representative companies include Phison, Silicon Motion, Marvell, and Maxio.
NVMe/PCIe (data pathway layer): Determines data transfer efficiency from storage to compute. Combined with GPUDirect Storage technology, it reduces CPU memory bounce buffers and CPU involvement, significantly alleviating I/O bottlenecks. Representative companies include Broadcom, Marvell, and Astera Labs.
HDD / cold storage / archiving: The low-cost foundation for AI data lakes
AI will not eliminate HDDs. As multimodal large models drive surging demand for video and image data, and enterprise compliance logs and historical datasets expand exponentially, the need for low-cost cold data storage is growing in tandem. In AI storage architectures, SSDs and HDDs work in a value-based tiered synergy: SSDs handle hot data and high throughput, while HDDs provide cost-effective, long-term retention. Representative companies include Seagate, Western Digital, and Toshiba.
AI storage software stack: The orchestration hub for data availability
What AI truly consumes is never raw drives, but rather 'data services' meticulously orchestrated by the software stack. This architecture transforms underlying hardware into knowledge assets directly callable by upper-layer AI systems, structured across four layers:
High-Performance Storage Systems (Feeding Systems): Centered on concurrent throughput and low latency, these systems address the 'data hunger' of GPU clusters through parallel file systems, ensuring ultra-fast data flow for training and inference. Representative companies: VAST Data, WEKA, Pure Storage.
Object Storage (Raw Data Lake): Focused on managing objects, keys, and metadata, it handles massive volumes of unstructured data. It prioritizes low cost and cloud-native characteristics over ultra-low latency, forming a scalable capacity foundation. Representative company: AWS S3
Vector Databases (Semantic Indexing Layer): Vector databases store, index, and retrieve vectors generated by embedding models, enabling AI to precisely locate relevant content from vast knowledge repositories. Representative companies: Pinecone, Milvus
RAG Data Layer (Knowledge Invocation Layer): Going beyond simple retrieval, this layer encompasses data chunking, cleansing, access control, and citation tracing to ensure enterprise data can be securely, accurately, and traceably accessed by large language models. Representative company: Databricks
From AI Hot Storage to Decentralized Cold Memory: Maximizing Efficiency vs. Maximizing Trustworthiness
AI storageIt is an efficiency-driven system optimized to maximize computational output.HBM bandwidthdetermines whether the GPU can be fully fed,SSD throughputdetermines the efficiency of dataset and checkpoint read/write operations; low latency is critical for real-time RAG and inference experiences. These metrics ultimately converge intoGPU utilizationand per-unittoken cost, which directly determines the profitability of AI applications. The ultimate goal of AI storage is not preservation, butAcceleration, servingProductivity
anddecentralized storagehas a fundamentally different value function. It asks whether data will still exist a decade from now, whether it has been tampered with, and whether it can withstand single-point censorship. Through cryptographic proofs and distributed networks, it builds an openly accessible and permanently preserved public data foundation. Its ultimate goal is to defend the absolute truthfulness and sovereign independence of data, servingfairness, censorship resistance, and civilizational memory
Author: Jacob Zhao @ IOSG    Today, the so-called 'first domestic storage stock'ChangXin Memory Technologiesofficially debuted on the ChiNext board, surging an astonishing 500% and electrifying the market. Although the broader storage sector continues to face headwinds from recent pullbacks, AI-related storage remains at the center of today’s tech narrative, undergoing aggressive revaluation by capital markets.Web3 Meanwhile, decentralized storage has fallen into prolonged silence and disillusionment. Why do two sectors both labeled as 'storage' exhibit such starkly divergent market performances? The fundamental reason lies in a complete divergence of their underlying value functions. The revaluation of storage in the AI era is essentially a frenzy centered on 'hot data efficiency,' driven by the pursuit of maximizing computing utilization and commercial monetization. In contrast, decentralized storage adheres to a value proposition built on 'cold data trustworthiness,' safeguarding data fairness, censorship resistance, and humanity’s long-term memory. The former constitutes an efficiency system for hot data; the latter, a trust system for cold data. Capital markets today clearly favor 'efficiency,' yet human civilization ultimately still requires an immutable foundation for memory. The long-term value of trustworthy cold storage has never vanished—it merely lies dormant in the shadows of the cycle, awaiting rediscovery and repricing by history.    Why Storage Has Reemerged as a Focal Point in the AI Supply Chain    In the traditional IT era, storage was a 'capacity business.' Enterprises' CIO focuses on cost per unit of capacity, hardware...
AI storage is the 'hot storage' that fuels future productivity, while decentralized storage is the 'cold memory' preserving humanity’s immutable historical record. The former prioritizes efficiency and极致 speed; the latter prioritizes trustworthiness and defends silent memory. The former determines how fast models run; the latter determines whether memories can be erased. Currently, market mechanisms reward productive efficiency—AI storage is at the forefront of attention, while decentralized storage appears to be entering a quiet phase marked by valuation collapse and narrative erosion.
The Vision and Reality of Decentralized Storage
There are numerous decentralized storage projects, but in terms of industry perception and ecosystem maturity, Filecoin and Arweave remain the core representatives. Although both fall under 'decentralized storage,' their underlying architectural philosophies follow nearly opposite paths—one uses market-driven contracts to approximate AWS’s elasticity, while the other relies on a one-time social contract to approach the permanence of a library.
Filecoin: By leveraging Proof-of-Replication (PoRep) and Proof-of-Spacetime (PoSt), it has built the most comprehensive verifiable economic system. Rather than continuing to compete head-on with AWS in consumer-grade cloud storage, it should pivot toward AI data provenance, public dataset hosting, and compliant archiving—providing a verifiable chain for model auditing and copyright verification. Its inevitable path forward is to wrap its services in S3-compatible APIs and support fiat payments, evolving from a 'cheap storage marketplace' into 'verifiable computing infrastructure.'
Arweave: With its 'pay once, store forever' narrative, Arweave uses Blockweave and the SPoRA mechanism to incentivize miners to preserve—and rapidly retrieve—as much historical data as possible, especially rare or scarce records. Its ideal role is as a foundational layer for humanity's public memory—preserving human rights documentation, evidence of war crimes, cultural classics, and archiving legal and financial history, thereby providing AI agents with permanent, long-term memory access. Arweave’s value lies not in speed, but in its capacity to carry civilizational memory across time cycles.
Author: Jacob Zhao @ IOSG    Today, the so-called 'first domestic storage stock'ChangXin Memory Technologiesofficially debuted on the ChiNext board, surging an astonishing 500% and electrifying the market. Although the broader storage sector continues to face headwinds from recent pullbacks, AI-related storage remains at the center of today’s tech narrative, undergoing aggressive revaluation by capital markets.Web3 Meanwhile, decentralized storage has fallen into prolonged silence and disillusionment. Why do two sectors both labeled as 'storage' exhibit such starkly divergent market performances? The fundamental reason lies in a complete divergence of their underlying value functions. The revaluation of storage in the AI era is essentially a frenzy centered on 'hot data efficiency,' driven by the pursuit of maximizing computing utilization and commercial monetization. In contrast, decentralized storage adheres to a value proposition built on 'cold data trustworthiness,' safeguarding data fairness, censorship resistance, and humanity’s long-term memory. The former constitutes an efficiency system for hot data; the latter, a trust system for cold data. Capital markets today clearly favor 'efficiency,' yet human civilization ultimately still requires an immutable foundation for memory. The long-term value of trustworthy cold storage has never vanished—it merely lies dormant in the shadows of the cycle, awaiting rediscovery and repricing by history.    Why Storage Has Reemerged as a Focal Point in the AI Supply Chain    In the traditional IT era, storage was a 'capacity business.' Enterprises' CIO focuses on cost per unit of capacity, hardware...
The challenges facing decentralized storage projects like Filecoin and Arweave do not stem from flawed value propositions, but rather from persistent misalignments between productization, retrieval experience, genuine demand, and long-term token incentives. This reveals the vast chasm between niche technical ideals and mainstream commercial adoption:
Misaligned Supply-Demand IncentivesEarly networks, exemplified by Filecoin, rapidly scaled capacity through token incentives but failed to build a sufficiently strong paying demand side, resulting in massive capacity with low utilization and poor monetization conversion. Rewards were based on 'I can store,' rather than 'someone needs me to store.'
Lack of enterprise-grade service capabilitiesAWS’s moat isn’t hard drives—it’s the 'data operating system' composed of APIs, SLAs, access control, compliance auditing, and technical support. Enterprises buy peace of mind, not experimental infrastructure that requires them to manage keys and select nodes themselves.
Weak retrieval experienceStoring data doesn't guarantee it can be reliably retrieved with low latency. Decentralized node distribution, complex topology, and lack of unified SLAs make these systems ill-suited for AI hot-data workflows, positioning them better for trusted cold archiving and data provenance.
Insufficient privacy and compliance safeguardsEnterprise private data cannot simply be written onto a public, permanent network; the right to delete fundamentally conflicts with immutable permanence. Decentralized storage is better suited for public data and long-term archives, not indiscriminate handling of core private data.
Token economics amplify market cyclesBull-market financialization masks underlying demand deficiencies, while bear markets expose commercialization weaknesses as miner ROI declines. Tokens can cold-start supply, but they don’t automatically create demand or sustainable revenue.
Other decentralized storage projects mostly focus on specific ecosystems or niche segments: Storj and Sia have weaker cross-cycle industry recognition and Web3 narrative influence compared to Filecoin and Arweave; BNB Greenfield and Walrus are tightly coupled to the BNB Chain or Sui ecosystems, respectively; Celestia and EigenDA belong to the data availability (DA) layer, serving rollup transaction confirmation rather than long-term archival; hybrid AI/DA narrative projects like 0G aim to integrate storage, data availability, computation, and AI agent settlement into an AI-native modular infrastructure, though their real-world demand, developer adoption, and commercial viability remain unproven.
Future Opportunities in Decentralized Storage: The Long-Term Pendulum Between Efficiency and Trustworthiness
During periods of explosive technological gains, capital fervently chases efficiency—assets like GPUs and HBM command steep premiums, while decentralized storage, which champions 'trustworthiness and fairness,' is naturally sidelined. Yet history’s pendulum never remains fixed at the extreme of efficiency forever. Events such as arbitrary bans and content removals by mega-platforms, surging AI copyright litigation forcing demands for data provenance, geopolitical conflicts igniting battles over data sovereignty, the disappearance of public archives due to data monopolies, and mounting regulatory pressure to audit the compliance of model training data could all catalyze a repricing of 'trustworthy storage.' Decentralized storage still holds distinctive future potential in the following areas:
AI Data Provenance: Leverage cryptographic proofs to establish 'data lineage verification,' addressing regulatory and audit pressures.
Public Datasets and Civilizational Archives: Anchor censored records and cultural heritage to build irreplaceable, immutable memory.
Trusted Archiving and Compliance Evidence: Enable trustworthy self-attestation through hash-based evidence, providing high-assurance digital notarization.
Integration of ZK/TEE/DID Technologies: Resolve privacy tensions and evolve from a standalone 'storage protocol' into 'trusted data infrastructure.'
Invisible product roadmap: Offering S3-compatible APIs and fiat-based billing, enabling users to directly purchase 'trusted archiving' services.
AI storage pursues extreme efficiency, fueling our journey into the future; decentralized storage safeguards silent memories, preserving our right to look back at the past. Today’s market rewards efficiency without reservation, leaving decentralized storage seemingly quiet—even collapsing. However, as the AI era further exacerbates data monopolies, copyright disputes, and the fragility of historical memory, decentralized storage may undergo a revaluation as a 'trusted cold layer.' Memories that cannot be easily erased by platforms, corporations, or any single authority might evolve from idealistic romance and fringe belief into essential infrastructure.
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Thumbs Up
2
214K Views
Report
Comments
Write a Comment...
2
3