Speak With An Expert

Platform

Memory and GPU Pooling

Get more tokens out of every GPU 

Own your Token Machine. Dynamically scale and share Memory and GPU capacity. Support any model or workload. Increase utilization and control your inference economics. More tokens per watt. More tokens per dollar. 

Request a Demo
image of liqid rack with memory and gpu pooling
Product

LIQID Unveils Scale-Up AI Platform with AMD

Pool up to 30x Instinct™ MI350P GPUs

"Customers can scale GPU resources on demand to achieve higher utilization, lower infrastructure costs, and industry-leading AI inference economics.” Suresh Andani, Corporate Vice President, Compute and Enterprise AI Group, AMD.

Learn More
amd and liqid logo horizontal layoutamd and liqid logo - portrait layout
Solutions

LIQID CXL Memory pooling at PNNL’s Large-Scale Computing Platform

HPC, AI inference, and post-training

"The Abaco system testbed design increases memory capacity by orders of magnitude, which is critical for scientific workloads” James A. Ang, Ph.D., Chief Scientist for Computing, IDSD, PNNL

Learn More
PNNL logo, DOE logo, micron logo and liqid logo
Platform
Software-Defined Memory and GPU Pooling Infrastructure
Product
LIQID UltraStack 30 - AMD MI350P
Solution
LIQID Launches Most Advanced CXL Memory Pooling...
Trusted by
Innovators
Trusted by
Innovators

The AI Inference Dilemma

Limit innovation or rethink how you manage Memory and GPUs?

demand pressure icon

DEMAND PRESSURE

Explosive Token Consumption is Outpacing IT Budgets

Always-on AI and agentic workloads make demand and costs increasingly difficult to predict.

New GPUs Every Year

Up to 8x GPU generations for every server lifecycle.

architecture pressure icon

Architecture PRESSURE

60% of GPU Cycles Are Wasted

Physically locked to individual servers, unable to be shared or reallocated.

The Memory Wall

AI inference models demand 3–5x more memory than servers support.

supply pressure icon

supply PRESSURE

Data Center Power Costs

Demand is growing faster than new generation capacity can be added.

GPU and Memory Costs

AI demand is competing for limited supplies of GPUs, HBM and DRAM.

Own Your Token Machine

Own the infrastructure that produces your AI tokens. Control cost, performance, capacity and data security.

LIQID combines purpose-built hardware and Matrix software to deliver dynamic memory and GPU capacity for more tokens per watt and dollar.

graphic illustration of cloud moving to secure on-prem
Bring cloud workloads on-prem without sacrificing performance or agility while gaining security and control.
server rack with lightning bolt above it
Reduce rack space and power by pooling memory and compute across fewer servers.
graphic representation of memory and gpu pools attached to liqid matrix
Pool and scale memory and GPUs on demand to maximize utilization across any workload.

What Can LIQID Help You Solve?

AI Inference workloads are creating unprecedented opportunities, but also new challenges. How can we help you get the most from your AI?

GPU Pooling

Stop assigning GPUs to servers. Assign them to jobs. Allocate from a shared pool per job, per tenant, per hour — through the scheduler you already run — and take them back when the job ends. Utilization goes up without buying anything.

PROOF
30% to 90%
shared GPU utilization, before and after
Up to 75%
fewer servers for the same work
~50%
the cost of the dedicated GPU-chassis alternative

LIQID pools up to 30 GPUs to one host or shared across hosts for maximum performance and agility.

Learn More
“We bought the GPUs. Half of them are on, powered, and doing nothing.”
Industry-average GPU utilization in 2026 sits in the single digits. The status quo isn't “pretty good,” you need to get the most out of your most expensive devices.

GPU Pooling

Size the prefill pool to the burst, then release it. The same GPUs go back to serving other work in seconds instead of waiting for the next spike.

PROOF
30% to 90%
shared GPU utilization, before and after
30 GPUs
pooled to a single host for maximum performance or shared across several for unmatched agility
Up to 75%
fewer servers for the same work

LIQID pools up to 30 GPUs to one host or shared across hosts for maximum performance and agility.

Learn More
“Prefill spikes for ten minutes, then those GPUs sit pinned all afternoon.”
Prefill is compute-heavy and lumpy. Pinning GPUs to it means paying for the peak every hour of the day.

GPU Pooling

Pool the GPUs the job needs, run it, then unbuild it. No purpose-built system sitting idle between jobs.

PROOF
30 GPUs
pooled to a single host for maximum performance or shared across several for unmatched agility
$52M
across three US DoD systems, 32 petaflops
30% to 90%
shared GPU utilization, before and after

LIQID pools up to 30 GPUs to one host or shared across hosts for maximum performance and agility.

Learn More
“This job needs more GPUs than one server has, but not a whole new class of machine.”

Memory Pooling

Put a real DRAM tier between the HBM you can't source and the NVMe that's too slow to backstop it. More concurrent sessions per GPU, longer context, and no extra HBM to go find.

PROOF
Up to 7x
tokens per second on KV-cache workloads
160TB
of DRAM to a single host, or shared across 16 nodes
1.8x
tokens per watt in a fixed power envelope

CXL 2.0 DRAM fabric — upto 100TB to a single host or shared across 16 nodes.

Learn More
“Context length and concurrency hit the wall long before the GPU does.”
KV cache outgrows HBM. Offloading it to NVMe buys capacity and pays for it in latency on every single token. You end up buying whole servers for memory you don't always need. NVMe-tier offload. HBM you can't get.

Memory Pooling

Give the memory to the node running the job, for as long as the job runs, then take it back. No new server, no new fabric for your team to learn.

PROOF
160TB
of DRAM to a single host, or shared across 16 nodes
Up to 30x
faster graph analytics
1.8x
tokens per watt in a fixed power envelope

CXL 2.0 DRAM fabric — up to 100TB to a single host or shared across 16 nodes.

Learn More
“The job needs more DRAM than the box can physically hold, so it stalls.”
The fix on offer today is buying another whole server for memory you'll use twice a quarter.

Memory Pooling

One pool, sized for the peak of the sum. In-memory databases, feature stores and analytics draw what they need from it and give it back — without a per-node hardware refresh.

PROOF
Up to 30x
faster graph analytics
160TB
of DRAM to a single host, or shared across 16 nodes
Up to 75%
fewer servers for the same work

CXL 2.0 DRAM fabric — up to 100TB to a single host or shared across 16 nodes.

Learn More
“The dataset is sized for memory. We're buying it by the node.”
Every node gets sized for its own peak, so the estate ends up sized for the sum of the peaks instead of the peak of the sum.

Dual Fabric — Memory and GPU

Both fabrics in one system under one control plane. Relieve each constraint from its own pool, so neither one forces you to over-buy the other.

PROOF
160TB
of DRAM to a single host, or shared across 16 nodes
30 GPUs
pooled to a single host for maximum performance or shared across several for unmatched agility
Only one
independent vendor ships both fabrics

Both fabrics, one control plane. Compose GPU-to-memory ratios per workload, then hand them back.

Learn More
“Two ceilings at once, not enough GPU, and not enough memory per GPU.”
Relieve one and you hit the other. Buy your way out of both and you've over-bought on both.

Dual Fabric — Memory and GPU

Pool different GPU-to-memory ratio for each phase, out of the same hardware, and change it when the traffic shape changes.

PROOF
30 GPUs
pooled to a single host for maximum performance or shared across several for unmatched agility
160TB
of DRAM to a single host, or shared across 16 nodes
1.8x
tokens per watt in a fixed power envelope

Both fabrics, one control plane. Compose GPU-to-memory ratios per workload, then hand them back.

Learn More
“Prefill and decode want different hardware. We only bought one shape.”
Decode is memory-bandwidth bound while prefill is compute bound. One fixed GPU-to-memory ratio can't serve both well, so one phase is always paying for the other.

Dual Fabric — Memory and GPU

Build the machine the job needs out of pooled GPU and pooled memory, then unbuild it and give the parts back to everything else.

PROOF
30 GPUs
pooled to a single host for maximum performance or shared across several for unmatched agility
160TB
of DRAM to a single host, or shared across 16 nodes
Only one
independent vendor ships both fabrics

Both fabrics, one control plane. Compose GPU-to-memory ratios per workload, then hand them back.

Learn More
“What this needs in GPU and memory together doesn't exist in a catalog.”
The alternative on the table is a purpose-built machine that's wrong for everything else you run.

Industries and Use Cases

Trusted partners
proven technology
Trusted Partners
proven technology

Explore The Latest Insights in
AI Technology

photo of sumit puri

What Can LIQID Help You Solve?

AI Inference workloads are creating unprecedented opportunities, but also new challenges. How can we help you get the most tokens and ROI from your AI?

Request a Demo