Lexar Wants to Offload Local AI Models to SSD Amid the RAMpocalypse

Lexar has been experimenting with various technologies to help consumers achieve faster data throughput and more reliable storage. However, the company is now envisioning something entirely different as the PC evolves from a regular personal computer to a local AI-enhanced experience. We had the opportunity to interview Lexar's Chief Technical Officer (CTO), Daniel Guo, about the technology Lexar is developing to help offload some of the DRAM demand to much cheaper NAND Flash. According to Guo, DRAM is about six times more expensive to manufacture than NAND Flash, and there are opportunities for AI SSDs to reduce the DRAM requirements for running AI models on local hardware. This is where the Lexar AI Storage Core SSD comes into play, as the company is creating new storage solutions for consumers to support local AI deployments using much less DRAM by offloading large language models (LLMs) to SSDs. This approach allows larger and more powerful LLMs to fit into a PC build, reducing memory footprint by at least 40%

Based on internal testing, Lexar managed to run the Qwen 3.5 122B AI model on a local PC. Traditionally, users would need to spend about $4,500 on a PC with a decent CPU and 128 GB of DRAM to run this model. Through hardware and software optimization, the Lexar AI suite with the Lexar AI Storage Core SSD can reduce the DRAM requirement to 32 GB and run the model with 35 billion parameters at 15.6 tokens per second, compared to only 5.2 tokens per second using traditional frameworks. When attempting to load the 122B model on 32 GB of DRAM, the traditional Llama.cpp fails to load and crashes, while Lexar's SSD offloading provides about 4.4 tokens per second.
Read full story