Lexar has been exploring various technologies to help consumers achieve faster data throughput and more reliable storage. The company now envisions a shift as the PC evolves from a traditional personal computer to a local AI-enhanced experience. We interviewed Lexar's Chief Technical Officer, Daniel Guo, about the technology Lexar is developing to help reduce some DRAM demand and shift the balance towards integrated hardware-software AI storage solution with NAND Flash. According to Guo, DRAM is about six times more expensive to manufacture than NAND Flash, and AI SSDs present opportunities to reduce DRAM requirements for running AI models on local hardware. This is where the Lexar AI Storage Solution comes into play, as the company is creating new storage solutions to support local AI deployments in constrained DRAM capacity by offloading some parts of large language models to SSDs. This approach allows larger and more powerful LLMs to fit into a PC build, reducing DRAM capacity requirements by up to 40% in specific scenarios.
Based on internal testing, Lexar managed to run the Qwen 3.5 122B AI model on a local PC. Traditionally, users would need to invest about $4,500 in a PC with a decent CPU and 128 GB of DRAM to run this model. Through hardware and software optimization, the Lexar AI suite with the Lexar AI SSD can reduce the DRAM requirement to 32 GB and run the model with 35 billion parameters at 15.6 tokens per second, compared to only 5.2 tokens per second using traditional frameworks. When attempting to load the 122B model on 32 GB of DRAM, the traditional Llama.cpp fails to load and crashes, while Lexar's SSD offloading provides about 4.4 tokens per second.