Laptop RTX 5090 Runs Circles Around M5 Max In Prompt Processing & Token Generation With Up To 133% Faster Performance, Until Bigger LLMs Or Context Length Come Into Play

NVIDIA’s laptop RTX 5090 boasts sufficient VRAM to be able to crank every visual setting in the most demanding AAA titles, but when running AI workloads, it’s a whole new ballgame for the GPU’s 24GB framebuffer. Still, a series of tests shows that it’s able to hold its own against Apple’s M5 Max by comprehensively beating it in a series of LLM workloads, that is, until the benchmarks fire up a heavier AI model coupled with increased Context Length.

The unified memory architecture of M5 Max comes to the rescue, significantly beating the RTX 5090 in token generation once the Context Length is increased

The updated Razer Blade 18 took on Apple’s latest and greatest 16-inch MacBook Pro, with Alex Ziskind comparing all sorts of benchmarks, including AI workloads. Starting things off with Qwen3 4B running at 4-bit quantization and a Context Length of 262,144, it’s clear that the raw graphics horsepower of the RTX 5090 isn’t going to be enough here, because nearly its entire 24GB VRAM has been used up, resulting in a tokens/second generation speed of 25.28.

In comparison, the M5 Max obtained 181.40 tokens/second. However, as soon as that Context Length is reduced to 68,014, a ton of VRAM is freed up for the RTX 5090, leading the GPU to achieve 160.36 tokens/second. As for why the graphics chip is still slower than the M5 Max, it likely has to do with the difference in programs used. In Windows, LM Studio was being run while MLX was utilized for Apple Silicon.

Fortunately, the RTX 5090 isn’t down and out just yet because the YouTuber ran a bunch of LLMs, such as Gemma 4 12B, Qwen3.5 27B, and Qwen3.6 35B A3B using llama.cpp. In both prompt processing and token generation, the RTX 5090 had the last laugh when compared to the M5 Max. Then again, you have to remember that NVIDIA’s fastest GPU for laptops will always be the king of the hill just as long as it doesn’t fill up its 24GB VRAM.

Unfortunately, if you’re going to run extremely dense models like Qwen3.5 122B, there’s absolutely zero chance for the RTX 5090 to come out of this test unscathed. As for the M5 Max and its 128GB of unified memory, it achieves prompt processing of 139 tokens/second and token generation of 47.4, while consuming 94GB of RAM. If you’re wondering why NVIDIA is eager to launch the RTX Spark , running heavier AI models will definitely be one reason.

For an in-depth look at the tests, you can check out Alex Ziskind’s video up top and let us know what you think in the comments.

News Source: Alex Ziskind

Follow Wccftech on Google to get more of our news coverage in your feeds.