A 150-billion-parameter Mixture-of-Experts model can generate text on a standard developer laptop. The experiment shows that RAM, not SSD speed, is the real performance limiter. This matters because it proves that frontier-scale models are testable locally without spending on cloud compute.

The Test

We ran the DeepSeek V4 Flash model—a REAP-pruned 150 B-parameter MoE—through the open-source Colibrì inference engine. The hardware was an AMD Ryzen AI 9 365 laptop with 61 GB of RAM and a 1 TB NVMe SSD. We timed a three-minute generation run and logged every disk access.

What the Numbers Reveal

  • Disk activity was minimal. The SSD performed fewer than 11 seconds of reads during the three-minute generation. Even if the drive could read data instantly, total throughput would improve by only about 6 percent.
  • RAM size directly impacted speed. Halving the RAM cache lowered throughput by roughly 15 percent. The CPU spent almost as much time handling cache misses—de-quantising, copying, and managing data—as it did waiting for the disk.

These figures overturn the common belief that storage bandwidth is the primary choke point for large-model inference on a laptop.

Why RAM Beats SSD

When a model parameter is not already resident in RAM, the system must:

  1. Pull the data from the SSD.
  2. Decode it and move it into CPU registers for computation.

Both steps consume cycles. The first step is limited by the SSD’s bandwidth; the second adds roughly the same delay because the CPU must de-quantise the data regardless of how fast it arrives. A faster SSD therefore yields diminishing returns, while more RAM lets more of the model stay resident and eliminates the costly round-trip.

Implications for Developers

  • Invest in memory, not storage.
  • Local testing becomes viable. Developers can run open-weight MoE models on existing machines, checking style, tool-calling behavior, and output quality without paying for cloud credits.
  • Batch evaluation is realistic. Real-time chat may still feel sluggish, but queued jobs—generating dozens of prompts for research—run comfortably on a laptop.

Potential Counter-Argument

Some may argue that SSD latency still matters for workloads that repeatedly load new expert weights.

What to Watch Next

  • Memory-efficient model variants.
  • Hardware with larger on-chip caches.
  • Software-level caching strategies.

The bottom line is clear: developers who want to experiment with massive MoE models on a laptop should buy more RAM. SSD upgrades shave off only a few percent, while extra memory can cut inference time by double-digit percentages. This shifts the cost calculus for AI prototyping and opens the door to local, cost-free testing of models that were previously thought to require dedicated cloud clusters.