Primemas and Micron Unveil Abaco System to Transform Rack-Scale AI Memory Capacity

Primemas Inc. unveiled breakthrough CXL memory technology at the Future of Memory and Storage (FMS) 2026 conference, addressing a critical bottleneck in AI systems. PostRegister reported the company showcased innovations that deliver four times more memory capacity per CXL port than conventional alternatives, enabling massive pooled memory architectures.
The company also partnered with Micron on the Abaco Project, a Department of Energy system that packs over 100 terabytes of shared memory into a single rack. Las Vegas Sun noted this architecture allows multiple AI servers to access tens of terabytes of pooled memory without expensive external switches.
Primemas' CXL add-in cards solve a persistent AI infrastructure problem: not enough memory bandwidth and capacity close to processors. Yahoo Finance explained the company's chiplet-based design delivers four times the memory per CXL port versus monolithic competitors. CXL is a standard that connects CPUs and GPUs to memory modules through fast interconnects.
This matters because AI models like ChatGPT need enormous amounts of memory to store token sequences during inference—the process of running trained models. Traditional server memory runs out quickly. Primemas' approach treats memory as a dynamic, shared resource pooled across an entire rack rather than trapped inside individual servers.
The Abaco System Architecture, built with Micron, represents a new way to organize AI compute clusters. UK Yahoo Finance reported the rack-scale system gives hundreds of CPU and GPU cores access to over 100 terabytes of shared Micron memory. Traditional data centers would need external switches and complex networking to achieve this.
Abaco was designed for Pacific Northwest National Laboratory (PNNL), a Department of Energy research center. The system handles memory like a neighborhood utility rather than private resources. Multiple servers can grab what they need instantly, eliminating bottlenecks that slow down AI workloads.
Primemas introduced SLiM (Switchless Pooled Memory), a memory architecture designed specifically for KV-cache-intensive AI inference. KV-caches store key-value pairs that AI models use when generating responses token-by-token. PostRegister reported SLiM lets tens of thousands of inference requests share massive DRAM pools without the overhead of external switches.
This approach cuts infrastructure costs significantly. Companies running inference at scale—like large language model API providers—waste money on switching hardware and complex configurations. SLiM simplifies the design while delivering better performance for memory-hungry workloads that dominate modern AI services.
Memory bandwidth has become the hidden ceiling for AI scaling. Yahoo Finance UK noted that GPU clusters often sit idle waiting for data. Primemas' full-stack approach—from chiplets to rack-level pooling—removes this constraint by making memory abundant and instantly accessible.
The partnership with Micron signals that memory architecture is shifting from boutique problem to mainstream solution. As AI models grow larger and inference workloads multiply, companies need systems that treat memory as shared infrastructure, not individual server resources.
Publishers
5
Articles
4
Reach
5