Most performance claims in the JavaScript tooling space have the shelf life of milk. Someone installs a few packages on a quiet Tuesday morning, captures the terminal output, and publishes a dramatic bar chart. By the next sprint, one of the tools has shipped a patch that invalidates the whole thing. The post remains indexed by search engines. The chart keeps getting shared. The numbers, however, are already lying to you.

This is the rot that infects nearly all package manager benchmarks. They are event photography when what we need is a live feed.

A project called depjs/canary treats this problem as a machine responsibility rather than a content calendar task. It is a living benchmark that watches npm, pnpm, Yarn, and dep, then re-runs its entire suite the moment any of them publishes a new version. The results are public, continuous, and unavoidable. When something breaks, the repository stays red until it heals. There is no cherry-picking, no hiding behind an old blog post, and no assumption that last month’s winner still holds the crown.

Why Speed Claims Need an Expiration Date

JavaScript package managers do not stand still. The gap between minor versions can include rewritten resolution algorithms, altered hoisting strategies, or changes to how the global cache is keyed. A benchmark that captures npm 10.2.1 and pnpm 8.11.0 tells you almost nothing about how those same tools behave two releases later. Yet the web is full of definitive statements like “Tool X is three times faster” based on exactly that kind of frozen snapshot.

Worse, many tests ignore the conditions that define real developer pain. A package manager might blaze through an install in a warm, well-rehearsed environment and then crawl on a CI runner that starts with an empty disk. Without testing both extremes, the benchmark becomes a press release rather than usable engineering data.

How the Canary Automates Comparison

Every two hours, a job polls the npm registry. If a new version of npm, pnpm, Yarn, or dep appears, the canary wakes up. It does not wait for a human to notice the changelog. It immediately executes a full test matrix that pits all four managers against five popular, real-world packages. The selection includes heavy hitters like React, Next.js, and Vite, codebases that actual developers install every day. These are not synthetic micro-projects designed to flatter one particular tool.

This release-triggered approach matters because it ties measurement directly to change. If the benchmark only ran on a nightly schedule, it might miss a mid-day hotfix or swallow a regression for hours. By running specifically on new versions, the canary asks a direct question each time: did this release make things better or worse?

The Four Scenarios That Stress Different Muscles

The test matrix is built around four distinct setups that map directly to workflows you will recognize.

  • Cold cache, no lockfile. This is the fresh clone on a new laptop, or the first install after deleting node_modules. Nothing is cached. Nothing is pinned. The package manager has to resolve, fetch, and write everything from scratch.
  • Warm cache, with lockfile. This is the happy path for continuous integration when things go right. The lockfile exists locally, and the cache still holds the tarballs from a previous run. The tool should move fast because most decisions are already made.
  • Cold cache, with lockfile. Here the lockfile is present, but the cache has been wiped. The manager can skip dependency resolution, yet it still has to download every byte across the network. This isolates network speed from resolution speed.
  • Warm cache, no lockfile. The cache is hot, but the lockfile is gone. The package manager must re-resolve the dependency tree before it can even begin extracting files. This tests the efficiency of the solver and the metadata parser under ideal network conditions.

Each scenario is executed five times, and the canary keeps the median result. That single choice kills a lot of noise. One transient network hiccup or a brief spike in registry latency cannot hijack the narrative. The outlier gets ignored; the typical experience gets recorded.

Smoke Tests Beat Empty Timers

Tốc độ thuần túy rất dễ làm giả nếu bạn không xác minh kết quả. Một trình quản lý gói có thể bỏ qua các bước sau khi cài đặt (postinstall), làm hỏng một vài liên kết tượng trưng (symlinks), hoặc cài đặt sai phiên bản mà vẫn đưa ra một mốc thời gian ấn tượng. Canary từ chối dừng lại ở bộ đếm thời gian. Sau khi quá trình cài đặt hoàn tất, nó thực sự thực thi mã đã được cài đặt.

Ví dụ, nó khởi chạy một ứng dụng Express và xác nhận rằng máy chủ bắt đầu lắng nghe trên cổng dự kiến. Nếu mã không chạy, bài benchmark sẽ thất bại ngay lập tức. Bài kiểm tra khói (smoke test) biến bộ công cụ này từ một cuộc đua thành một cuộc kiểm toán. Nó trả lời câu hỏi mà chỉ riêng tốc độ không thể làm được: liệu việc cài đặt có thực sự hoạt động hay không?

Sự trung thực tuyệt đối là một tính năng

Tác giả của canary đã đưa ba quy tắc vào quy trình mà hầu hết các tác giả benchmark khác chỉ coi là tùy chọn.

Sân chơi bình đẳng. Các cờ (flags) được sử dụng để chuẩn hóa hành vi giữa các công cụ. Nếu một trình quản lý gói che giấu khiếm khuyết về hiệu suất đằng sau một thiết lập mặc định, bài benchmark sẽ phơi bày nó thay vì để công cụ đó trông có vẻ tốt một cách ngẫu nhiên.

Khởi động lạnh thực sự. Trước mỗi lần lặp lại, chứ không chỉ lần đầu tiên, bộ nhớ đệm (cache) của npm và kho lưu trữ (store) của pnpm đều được xóa sạch. Từ "mỗi" đó mang một ý nghĩa rất lớn. Nhiều bài benchmark chỉ xóa cache một lần, sau đó chạy năm lần cài đặt liên tiếp. Từ lần thứ hai đến lần thứ năm không thực sự là "khởi động lạnh", và các con số sẽ bị thổi phồng tương ứng. Canary bắt đầu từ con số không trong mỗi lần chạy.

Trạng thái thất bại công khai. Khi một bản phát hành mới làm hỏng thứ gì đó, kho lưu trữ (repository) sẽ duy trì trạng thái thất bại màu đỏ. Nó nằm chình ình trên trang chủ, xấu xí và chưa được giải quyết, cho đến khi một bản sửa lỗi được phát hành. Không có sự che giấu âm thầm nào để giữ cho bảng điều khiển (dashboard) luôn hiển thị màu xanh. Chính sách này buộc mọi thứ phải minh bạch. Một người dùng đang đánh giá các công cụ có thể thấy không chỉ công cụ nào nhanh nhất, mà còn công cụ nào duy trì được sự tin cậy theo thời gian.

Khả năng kiểm tra là điều không thể thương lượng

Một bài benchmark mà bạn không thể tái lập thì chỉ là một khẩu hiệu chiến dịch. Canary giải quyết vấn đề này bằng một script bash duy nhất cho phép bất kỳ ai cũng có thể chạy bất kỳ phần nào của bộ công cụ tại máy cục bộ. Bạn không cần phải tin tưởng vào mạng lưới của nhà cung cấp đám mây hay môi trường được tinh chỉnh thủ công của người duy trì. Nếu bạn nghi ngờ các con số không chính xác, bạn có thể tự tạo ra kết quả của riêng mình.

Sự minh bạch đó cũng khiến dự án trở nên hữu ích cho những người duy trì. Khi một lỗi hồi quy (regression) xảy ra, một nhà phát triển ở hạ nguồn (downstream) có thể tải script về, thực hiện bisect các bản phát hành của công cụ, và gửi cho đội ngũ thượng nguồn (upstream) một