Speak With An Expert

Mail icon to contact Liqid about our composable infrastructure technology
0
AMD MI350P GPUs / Server
0
TB
Aggregate HBM3E Memory
0
PF
FP8 AI Performance
0
Target GPU Utilization
image representing how ai models have outgrown legacy infrastructure
The Problem

Large Models Have Outgrown Traditional Servers

Most servers top out at 8 GPUs — and roughly 80% of the market is limited to just 2–4 — forcing large models across many systems, with stranded capacity and a rising cost per token.

Legacy Multi-Server Approach

  • Capped at 2–8 GPUs per server
  • Large models split across 4+ servers
  • Multi-node overhead drags performance
  • Duplicated software licensing & networking
  • Stranded, underutilized GPU capacity

Liqid UltraStack 30

  • Up to 30 GPUs pooled into a single node
  • Deploy frontier models on one server
  • Predictable linear scaling, no multi-node tax
  • Consolidated licensing & networking
  • GPU utilization driven toward 100%
superior tokenomics

More Tokens Per Second, Dollar & Watt

Most servers top out at 8 GPUs — and roughly 80% of the market is limited to just 2–4 — forcing large models across many systems, with stranded capacity and a rising cost per token.

3.7x
Higher Tokens / Second
Legacy
1x
UltraStack 30
3.7x
~65% Lower
Deployment cost for large-model builds
2.1x
More Tokens / Dollar
Legacy
1x
UltraStack 30
3.7x
~50% Lower
Power consumption vs. Legacy designs
1.8x
Better Tokens / Watt
Legacy
1x
UltraStack 30
3.7x
Up to 4x
The GPU density of traditional servers
Why Liqid

Density, Economics & A Path To The Future

Delivers the highest AMD GPU density available — up to 30 AMD MI350P GPUs in a single server — for superior AI inference tokenomics: more tokens per dollar and more tokens per watt.

Breakthrough Scale-Up Density

Up to 30× AMD Instinct MI350P GPUs pooled into one server over a high-speed PCIe fabric — with native Kubernetes support for multi-model serving.

Superior Inference Economics

Up to 3.7× tokens/second, 2.1× tokens/dollar and 1.8× tokens/watt — with up to 65% lower deployment cost and 50% lower power.

CXL Memory-Pooling Ready

A clear path to terabyte-scale shared memory for KV cache and other memory-intensive workloads as CXL pooling matures.
Why Liqid

One Pooled Platform, Many AI Workloads

Large LLM Inference

Deploy frontier and large-context models that demand massive HBM — on a single server instead of four or more — at industry-leading cost per token.

69 PF
Long-Context AI

4.3 TB of aggregate HBM3E keeps long context windows resident in GPU memory, sustaining throughput where legacy nodes stall.

4.3 TB
Multi-Model Serving

Serve small, medium and large models from one shared GPU pool with native Kubernetes support — higher models-per-server density.

4.3 TB
Agentic AI

Run sales, HR, marketing, legal and coding assistants on-premises, where data stays private and costs stay predictable.

Agents
RAG & Vector Search

Retrieval-augmented generation and long-context inference benefit directly from more than terabytes of pooled HBM per system.

RAG
Scientific Computing (HPC)

Drug discovery, materials science and mechanical design get GPU capacity that flexes with demand — no re-cabling, no reboots.

HPC

Large LLM Inference

Deploy frontier and large-context models that demand massive HBM — on a single server instead of four or more — at industry-leading cost per token.

69 PF

Long-Context AI

4.3 TB of aggregate HBM3E keeps long context windows resident in GPU memory, sustaining throughput where legacy nodes stall.

4.3 TB

Multi-Model Serving

Serve small, medium and large models from one shared GPU pool with native Kubernetes support — higher models-per-server density.

K8s

Agentic AI

Run sales, HR, marketing, legal and coding assistants on-premises, where data stays private and costs stay predictable.

Agents

RAG & Vector Search

Retrieval-augmented generation and long-context inference benefit directly from more than terabytes of pooled HBM per system.

RAG

Scientific Computing (HPC)

Drug discovery, materials science and mechanical design get GPU capacity that flexes with demand — no re-cabling, no reboots.

HPC
At a Glance

Large Models Have Outgrown Traditional Servers

Most servers top out at 8 GPUs — and roughly 80% of the market is limited to just 2–4 — forcing large models across many systems, with stranded capacity and a rising cost per token.

Download the Full Datasheet

LIQID UltraStack™

30× AMD Instinct™ MI350P PCIe
AI Performance
GPU HBM Memory
Host Server
Orchestration
69 PFLOPS (FP8)
4.3 TB HBM3E
Liqid Matrix® · Kubernetes-native
Dual-socket AMD EPYC™ 9005 Series
GPUs
Memory Expansion
CXL memory-pooling ready
Total System Power
~22 kW

Pool GPUs. Proven. Production-Ready. Future-Ready.

Whether your needs are performance, agility, or efficiency, pool your enterprise GPU infrastructure with LIQID.

Request a demo