IDC reported that 94.7% of surveyed companies stored more data in the past 12 months due to AI and generative AI adoption, a figure that reframes the storage conversation from raw compute to data retention. The white paper, sponsored by WD, argues that AI creates a continuous “data compound cycle” in which inference outputs, synthetic data, and model checkpoints become inputs for future workloads, making the storage tier a structural dependency rather than a back-end utility, according to TechNews.

61% saw 25%+ data growth in 12 months

The growth is not marginal: 61% of respondents saw data volume grow more than 25% in the past year, and 74% expect that pace to hold for the next three years. 59.4% identified AI-generated data - synthetic sets, inference logs, model records - as the main driver of data-lake expansion, meaning the storage stack must absorb a new, continuously generated class of content that did not exist in pre-AI workloads.

75.9% re-using cold data for AI workloads

The most consequential shift is the erosion of the active/archive boundary. 75.9% of companies are now re-utilizing previously archived cold-storage data to support AI workloads, and 96% anticipate needing to retrieve that historical data faster than before. For storage architects, this means cold tiers can no longer be treated as write-once, read-rarely silos; they must be reachable at speeds that support inference and RAG pipelines.

60%+ of pool still cold, 74.6% tiered

Architecturally, 74.6% of surveyed firms already store data across warm, cool, and cold tiers, yet over 60% of the total data pool remains cold or infrequently accessed. The tension is clear: the majority of bytes sit in the cheapest tier, but the fastest-growing demand is for faster access to those same bytes. This pushes the design problem from “where do we park data” to “how do we make parked data reachable without breaking the cost model.”

98.2% rank TCO per TB as critical

Cost remains the binding constraint: 98.2% of companies rank TCO per TB as an extremely important factor in storage decisions. The practical implication for the flash and HDD supply chain is that the next generation of enterprise storage must deliver higher density and lower power per terabyte while shortening access latency on cold tiers - effectively bridging the performance gap between archival media and active SSD tiers without inflating the per-TB price.

WD CEO: data is the true foundation of AI

WD CEO Irving Tan framed the shift as a correction to the compute-centric narrative: “the true foundation of AI is data,” and the need for large-scale storage, management, and access will only grow. Whether that demand translates into a structural uplift for high-density NAND, QLC/PLC enterprise SSDs, or high-capacity HDDs depends on how quickly vendors can close the latency-cost gap on cold tiers - a question the survey data raises but does not answer.

Focus shifts from compute to data

Global discussions surrounding Artificial Intelligence (AI) infrastructure are shifting focus from pure compute power to the foundational role of data itself.. This cycle dictates that AI not only generates massive volumes of new data but also assigns higher value to the data that is continuously retained and accumulated after the initial compute cycle concludes.