GPU Memory: The Costly Resource in AI Production
GPU memory stands out as the most expensive resource in AI production and is also the fastest resource to deplete. Long context windows and multi-turn interactions compel AI models to recalculate previously processed information, consuming valuable GPU memory and computational power that could instead support more users or generate new responses.
Rather than viewing GPU memory as a limiting factor, why not explore affordable storage solutions?
Weka believes that economical flash storage can bridge this gap. The launch of the company’s NeuralMesh 6 software platform alongside the Wekapod 3, its first in-house designed hardware series, expands what Weka calls its Augmented Memory Grid. This innovative approach aggregates NAND flash to emulate GPU memory at a significantly lower cost.
This sector is rapidly growing and becoming increasingly competitive. Companies like Dell, NetApp, Pure Storage, and VAST have pivoted towards AI infrastructure in the last two years. Weka claims to be designed explicitly for this moment, as opposed to merely adapting to it.
Weka co-founder and CEO Liran Zvibel noted, “Customers are now seeking immediate compute availability. When they receive a new allocation, they want to start running it right away,” in an interview with VentureBeat.
The potential advantages are clear: leverage existing GPU investments, cut down on inference costs, and swiftly deploy new AI workloads without waiting months for additional GPU resources.
This technology is particularly pertinent for organizations already operating AI at scale or those expecting rapid growth, such as those developing internal co-pilots, customer service agents, software engineering assistants, or search systems with lengthy context windows. Smaller deployments might not see the same direct benefits compared to organizations that already find GPU utilization to be a limiting constraint.
Exploring Weka’s NeuralMesh 6
NeuralMesh 6 introduces four key features aiming to address competitive evaluation gaps, as highlighted by Zvibel.
Configurable Virtual Multi-Tenancy: Composable clusters offer anchor tenants complete hardware-level isolation, dedicated CPU, memory, and storage. Through Weka’s RDMA fabric, virtual multi-tenancy allows for network-level isolation for over 1,000 tenants per cluster, with provisioning in under 30 minutes. One cluster supporting 50 composable clusters can handle up to 50,000 tenants.
Unified Storage for Files and Objects: Traditional storage systems maintain separate file-based and object-based paths, leading to data redundancy. Weka asserts that the same physical data can be read from either path simultaneously, eliminating the need for secondary copies and conversion layers, particularly targeting non-AWS GPU clouds such as Lambda, Nebius, G42, and CoreWeave. This approach boasts nearly two orders of magnitude higher performance than conventional S3 with a capacity-based pricing model instead of per-API costs.
Metadata-First Replication: This approach allows for destination environment visibility before complete data transfer, hydrating data only upon access. Zvibel highlights that customers can now receive new GPU allocations and become operational within an hour, compared to previous wait times that could extend to weeks or even months.
AlloyFlash and Always-On Data Reduction: Weka employs a hybrid of TLC and QLC NAND flash memory. While TLC is faster and more durable, QLC is more economical. AlloyFlash seamlessly integrates both types, running high-volume workloads on QLC while routing latency-sensitive tasks to TLC. This strategy lowers the cost per terabyte without sacrificing performance for demanding applications. Additionally, Weka has made data reduction a default feature rather than an optional add-on.
Addressing AI Context Challenges
The combination of multi-tenancy and object storage enhances how enterprises and neo-clouds operate daily. However, more complex challenges persist. As context windows and multi-turn interactions expand, GPUs waste computation on redundant calculations previously handled by the model. The Augmented Memory Grid, a specific feature of NeuralMesh 6, provides Weka’s solution.
Each prompt triggers two stages: prefill computation for attention (a core mechanism in large language models) and decoding, which turns that computation into output. The prefill is resource-intensive, while decoding is relatively light.
This inefficiency becomes apparent in multi-turn sessions like chat or coding. For instance, new turns may retrigger prefill of all past content unless cached.
“If you have 10 turns, you might end up recalculating 100 times; with 20 turns, that could escalate to 400 times,” Zvibel explained. “We can integrate significantly more NAND in shared memory, letting us cache all precomputed tokens to prevent recomputation.”
Weka’s Strategic Positioning
Storage vendors have been shifting their focus towards AI over the last year and a half. Discerning genuine capability from mere marketing claims is now a pressing concern for buyers.
Steve McDowell, chief analyst at NAND Research, remarked, “The storage industry is transitioning from delivering bits for enterprise workloads to managing data at AI speeds. This shift has been particularly noticeable over the past 18 months with Dell, NetApp, and Pure.” He noted that companies like Weka and VAST are truly AI-native data companies that have been solving these challenges from the beginning.
McDowell highlighted expanded memory grids as Weka’s standout technological advantage. “With this expanded memory grid, Weka continues to offer the most advanced KV cache implementation on the market,” he stated. “They have been early adopters and are constantly innovating, which is crucial for efficient AI inference. This effectively enhances GPU efficiency, translating to substantial savings on GPU and memory costs in today’s constrained market.”
Additionally, he cautioned that Weka’s commitments to data reduction claims are often undervalued. “Weka is putting its money where its mouth is with contractual guarantees for data reduction promises,” McDowell added.
For prospective buyers assessing competing offerings from Weka, VAST, Pure, and NetApp, McDowell advises careful consideration of vendor promises versus actual deliverables. “Smart buyers will investigate how competing vendors tackle real-world problems today by consulting with organizations operating similar workloads at comparable scales,” he stressed. “If a vendor cannot provide such references, it’s a significant red flag.”
Source: venturebeat.com


