Eight Arc Pro B60 cards with combined 192GB VRAM tested

First hands-on testing of Intel’s Arc Pro B60 Battlematrix platform from Storage Review shows a dense local AI setup built around dual-GPU cards with 24GB of VRAM per GPU. Four of these Maxsun dual cards give a workstation up to eight B60 GPUs and 192 GB of total VRA M, aimed at setups that want to run large language models locally and avoid cloud costs or data-sharing concerns.
Intel has set Arc Pro B60 at around $600 per GPU, so a dual-GPU 48 GB card lands at about 1,200 USD. At those levels, 24–48 GB of VRAM per node comes in far cheaper than most professional GPUs with similar memory footprints, which often cost at least twice as much.
Source: Storage Review
So obviously this system, and this configuration in particular, is not for gaming. In fact, a single B60 dual-GPU card isn’t actually a normal “dual-GPU” card we know from the GTX 690 era. Maxsun already explained that there are two cards using a single PCIe slot through bifurcation . Basically, two GPUs share the same board and the same slot, but for the operating system they are two separate GPUs. Combined, the system sees eight Arc Pro B60 cards, each with 24 GB of VRAM. So obviously LLM testing is where this platform will shine.
For many models, the best way to use these GPUs is simple: use the minimum number of cards you need to fit the model. At low batch sizes, a single GPU can beat an eight-GPU spread, because shuffling data across PCIe adds overhead. The full eight-GPU setup starts to make sense once you push higher concurrency and bigger batches where raw throughput matters more. The software is still early though. Only MXFP4-based GPT-OSS models worked as intended with low-precision paths, while formats like standard INT4, FP8 and AWQ refused to start, so many dense models had to run in BF16.
Source: Storage Review
A consistent pattern emerges across all tested models: at low batch sizes with our 256 input/output token configuration, using the minimum number of GPUs required to fit the model delivers better per-user performance than distributing across all eight GPUs. The inter-GPU communication overhead via PCIe, even at PCIe 5.0 speeds, introduces latency that exceeds the parallelization benefits for single-user or low-concurrency scenarios.
— Storage Review
The dual B60 card is long, uses a dual-slot blower, and pulls up to 400W through a single 12V-2×6 connector. That raised some concerns in early coverage about how hot it might get on an open bench. The extra length of the card might also cause problems in some tower cases, though it fit fine in standard server enclosures.
The tests use early drivers, a pre-release LLM Scaler build, and an AMD EPYC system instead of the Xeon 6 platform Battlematrix is meant to ship with, so all numbers are marked as preliminary. Intel announced Battlematrix back in May, but the reviewer expects the hardware and software combo to feel mature only sometime in 2026.
Source: Storage Review