.avif)
For the past few years the conversation around AI infrastructure has centered on one thing: compute. More GPUs, faster GPUs, bigger clusters (i.e. groups of GPUs). That focus has made sense: GPUs are extraordinary engines, and access to them has been the defining constraint for AI development. But as deployments have scaled from hundreds to thousands to hundreds of thousands of units, another bottleneck has emerged: the data infrastructure layer that feeds and coordinates those systems in real time.
In modern AI systems, the challenge is not just storing data, but continuously delivering, processing, and coordinating it across thousands of GPUs in real time.
When a GPU cluster stalls because data simply can’t be delivered fast enough, every idle second is wasted capital. Engineers call it “compute starvation.” One GPU cloud operator put it to us plainly: “If GPUs are idle, we lose money.” This is the problem VAST Data was built to solve by rethinking how data infrastructure operates in AI environments, and it’s why we’re excited to announce our investment in VAST’s Series F.
Why has data infrastructure become the limiting factor?
The answer is architectural. Traditional data infrastructure systems, the kind that have powered enterprise data centers for decades, were built around a “shared-nothing” approach where storage, compute, and processing are tightly coupled and operate in isolation. Each node bundles compute and storage together, manages its own slice of the data universe, and coordinates with neighboring nodes when data needs to move. This works adequately when you have a few hundred servers doing conventional workloads. It breaks down under the demands of AI.
Modern GPU training runs are massively parallel: thousands of chips simultaneously demand high-throughput access to the same datasets, performing millions of small reads and writes, checkpoints, gradient updates, and metadata lookups. As leading AI cloud providers have discovered when scaling their GPU deployments, legacy architectures introduce bottlenecks that prevent GPUs from running at full utilization.
How does VAST solve this at the system level?

VAST’s answer is an architecture called DASE — Disaggregated Shared-Everything, which redefines how data infrastructure is built by separating data access, processing, and compute coordination into a unified, shared system. The core insight is elegant: separate (data) compute from storage entirely, so each can scale independently. All compute nodes have simultaneous, direct access to a global storage pool via NVMe-over-Fabric, a protocol purpose-built for low-latency SSD communication at scale.
The practical implications are significant. There are no per-node caches to synchronize, no data to pre-stage or duplicate, no “east-west” coordination traffic that grows exponentially with cluster size. Data is continuously accessible across the entire system, enabling GPUs, pipelines, and applications to operate on shared, real-time state without duplication or movement. This is a fundamentally different model from what most enterprises have been running, and it turns out to be exactly what AI workloads need.
Get G2 Venture Partners’s stories in your inbox
Join Medium for free to get updates from this writer.
Subscribe
Remember me for faster sign in
VAST also made an early bet on flash (SSDs) over traditional spinning hard drives. Solidigm’s analysis of a one-exabyte deployment comparing VAST on flash to conventional HDD storage found roughly 60% lower 10-year total cost of ownership, driven by 77% lower power costs and a 90% reduction in physical footprint. Flash provides the performance and efficiency foundation, but its real impact comes from how VAST’s software turns that hardware into a globally accessible, real-time data platform.
Where is VAST winning, and why?
- HPC and research labs were the first to embrace VAST fully, because it let them stop managing infrastructure complexity and instead focus on science. Customers like the Texas Advanced Computing Center, NERSC, WEHI, and the National Cancer Institute choose VAST because it delivers high-end performance in a unified enterprise-grade package, eliminating complex storage tiers, custom code, and global data transfers.
- GPU clouds and neoclouds have become the fastest-growing segment. For an AI-focused cloud provider, storage throughput can make or break GPU utilization, and GPU utilization determines revenue per rack. VAST has become the dominant data infrastructure platform for this segment, deployed across leading providers including CoreWeave, Crusoe, Nscale, Lambda, and many others. Our portfolio company, Crusoe, was even quoted in VAST’s Series F press release, saying “This [round] validates the essential role [VAST] plays in helping us give AI infrastructure engineers and developers the seamless, industrial-scale foundation they need to build the future of AI”. Across the neocloud segment, we consistently hear the same sentiment: VAST is a critical part of the foundation to run an efficient and successful neocloud business.
- Enterprises are a newer and increasingly exciting growth vector. AI-forward companies are discovering that as AI moves from a feature to the core of their product, the underlying data architecture has to change to support continuous, AI-driven applications. As we heard from one customer: “To become an AI-native company, you need VAST”.
- Physical AI is another emerging segment. Foundation models being trained for robotics, autonomous vehicles, and industrial automation are uniquely data-hungry. Unlike conventional LLMs trained primarily on text, Physical AI models require massive volumes of multimodal data: video, sensor feeds, LiDAR, and proprietary telemetry, captured continuously from the real world. Training datasets can dwarf anything in the conventional enterprise software world. VAST is one of a very short list of platforms built to manage continuous, multimodal data streams and make them usable in real time.
From Storage to AI Operating System
To this point, we have been discussing storage, but storage is just the entry point for VAST. What began as a breakthrough in data access has evolved into a broader rethinking of how data infrastructure should operate in AI systems. VAST has been systematically extending its platform into adjacent data infrastructure categories: a native database layer (supporting OLTP, OLAP, data lake, and vector queries in one system), an event-driven compute layer that executes directly on stored data, a global namespace for hybrid environments, and most recently, an AI agent runtime. Collectively, VAST calls this the VAST AI Operating System.

VAST’s AIOS replaces a fragmented stack of storage, databases, streaming systems, and orchestration tools with a single runtime where data, compute, and intelligence operate together.
The (underappreciated!) climate angle
Data infrastructure might not seem like a climate issue, but at the scale of AI, it becomes a defining factor in overall system efficiency.
First, when the data layer can’t keep pace with compute, those GPUs sit idle, consuming power without producing useful work. In modern AI environments, this is not just a storage problem. It is a coordination across data access, movement and processing. VAST addresses this by enabling continuous, high-throughput access to shared data, reducing idle time and shortening overall workload duration.
As a result, GPUs complete the same work in fewer hours, directly lowering energy consumption. VAST reports 25 to 75 percent improvements in wall-clock time when customers move from legacy architectures, with independent research from the University of Minnesota (Source: University of Minnesota, 2024) corroborating this range. At production-scale data volumes, these efficiency gains translate into meaningful reductions in energy usage across the entire system.
Second, the underlying infrastructure itself becomes more efficient. Flash provides a more compact and power-efficient foundation than traditional disk, but the larger impact comes from how that hardware is utilized. By eliminating redundant data movement, reducing system overhead, and enabling a unified data environment, VAST improves overall compute efficiency across workloads. Organizations such as NOAA and the NIH have reported significant gains in processing efficiency and reduced time to results after deploying VAST.
Taken together, these improvements compound. Based on internal modeling, VAST estimates that its deployments could enable up to 16 million metric tons of avoided CO₂ emissions over the next decade. This is not the result of a standalone sustainability initiative, but a direct outcome of building more efficient data infrastructure for AI at scale.
This impact is intrinsic to VAST’s value proposition. Customers choose VAST because it’s faster and cheaper — the emissions reduction is a consequence of that efficiency.
Why we’re excited now
We’ve been tracking VAST since 2023, when we identified them as a winner in HPC with an emerging opportunity in GPU clouds. What we’ve seen since has exceeded those early expectations. The neocloud segment has grown faster than we anticipated. The enterprise pipeline has developed real conviction. And the product has evolved from a high-performance storage system into something that looks increasingly like essential infrastructure for the AI era.
One more thing worth noting: the business itself is extraordinary. VAST publicly reports a Rule of X of 228%, a metric that combines growth rate and free cash flow margin, where anything above 40% is considered healthy and the best large-cap software companies average around 61%. At 228%, VAST is in a category of its own. That kind of profitability, at this pace of growth, is genuinely unique.
The AI build-out is still in its early innings. VAST is powering millions of GPUs already today, but more GPUs will be deployed over the next five years than have been deployed in all of history to date. Every one of those GPUs depends on a data infrastructure layer that can keep up with continuous, real-time workloads. That is where VAST is positioned. .
We’re proud to back Renen Hallak and the entire VAST team as they build toward that vision.




