A 150-billion-parameter Mixture-of-Experts model can generate text on a standard developer laptop. The experiment shows that RAM, not SSD speed, is the real performance limiter. This matters because it proves that frontier-scale models are testable locally without spending on cloud compute.
The Test
We ran the DeepSeek V4 Flash model—a REAP-pruned 150 B-parameter MoE—through the open-source Colibrì inference engine. The hardware was an AMD Ryzen AI 9 365 laptop with 61 GB of RAM and a 1 TB NVMe SSD. We timed a three-minute generation run and logged every disk access.
What the Numbers Reveal
- Disk activity was minimal. The SSD performed fewer than 11 seconds of reads during the three-minute generation. Even if the drive could read data instantly, total throughput would improve by only about 6 percent.
- RAM size directly impacted speed. Halving the RAM cache lowered throughput by roughly 15 percent. The CPU spent almost as much time handling cache misses—de-quantising, copying, and managing data—as it did waiting for the disk.
These figures overturn the common belief that storage bandwidth is the primary choke point for large-model inference on a laptop.
Why RAM Beats SSD
When a model parameter is not already resident in RAM, the system must:
- Pull the data from the SSD.
- Decode it and move it into CPU registers for computation.
Both steps consume cycles. The first step is limited by the SSD’s bandwidth; the second adds roughly the same delay because the CPU must de-quantise the data regardless of how fast it arrives. A faster SSD therefore yields diminishing returns, while more RAM lets more of the model stay resident and eliminates the costly round-trip.
Implications for Developers
- Invest in memory, not storage.
- Local testing becomes viable. Developers can run open-weight MoE models on existing machines, checking style, tool-calling behavior, and output quality without paying for cloud credits.
- Batch evaluation is realistic. Real-time chat may still feel sluggish, but queued jobs—generating dozens of prompts for research—run comfortably on a laptop.
Potential Counter-Argument
Some may argue that SSD latency still matters for workloads that repeatedly load new expert weights.
What to Watch Next
- Memory-efficient model variants.
- Hardware with larger on-chip caches.
- Software-level caching strategies.
The bottom line is clear: developers who want to experiment with massive MoE models on a laptop should buy more RAM. SSD upgrades shave off only a few percent, while extra memory can cut inference time by double-digit percentages. This shifts the cost calculus for AI prototyping and opens the door to local, cost-free testing of models that were previously thought to require dedicated cloud clusters.
