A config file that worked five seconds ago is now a 347-byte fragment of broken JSON. Your CLI tool will not start. The user, who simply pressed Ctrl-C because the update was taking longer than expected, is now staring at an error stack trace they did not ask for. The two kilobytes of valid configuration that existed before the write are gone, replaced by the digital equivalent of a half-printed receipt.
This happens because writeFile is not atomic. It opens the existing path, truncates it, streams data from Node into the kernel’s page cache, and eventually closes the file descriptor. The moment the truncation happens, the old content is already gone. Everything between that truncation and the final close is a window of vulnerability. A SIGINT, a power outage, or a laptop lid slamming shut during that window leaves the filesystem holding a truncated mess. Even if Node reports the Promise as resolved, the operating system may still be buffering writes in memory. Convenience methods hide that gap, but they do not remove it.
The fix is not to write in place. The fix is to separate the act of writing from the act of publishing.
Write to a sibling, then swap
The reliable pattern has five steps. None of them are complicated, but together they move the failure window from an entire streaming write down to a single filesystem metadata operation.
First, serialize the entire payload in memory. Do this before you create any temporary file. If JSON.stringify throws because someone passed a circular object, you want that exception to bubble up before you touch the disk.
Second, write the serialized data to a temporary file located in the same directory as the target. Use a randomized name so two concurrent runs do not collide. Keeping the temp file in the same directory matters because rename is only atomic within a single filesystem. If your temp file lives on a different partition, the operating system falls back to a copy-and-delete sequence, which introduces its own failure modes and is no longer atomic.
Third, ask the kernel to flush that temporary file to physical storage. Node’s fsync, exposed here as the sync method on a filehandle, blocks until the buffers are down on the metal. This is slow, but config writes happen rarely enough that the durability is worth the milliseconds.
Fourth, rename the temporary file over the original path. On both POSIX systems and Windows, this is the commit point. Readers opening the original path will see either the complete old file or the complete new file. There is no moment in time where a reader can open the path and observe a half-written buffer.
Fifth, sync the parent directory. This catches a subtle edge case. The rename updates the directory entry, but the directory metadata itself might sit in the kernel’s page cache. Sudden power loss after a successful rename can sometimes leave the filesystem in a state where the new inode reference was never recorded durably. Syncing the directory forces that metadata update to disk and seals the transaction.
A concrete Node.js implementation
Here is what that pattern looks like in practice using only the Node.js standard library:
import { open, rename, rm } from "node:fs/promises";
import { dirname, basename, join } from "node:path";
import { randomUUID } from "node:crypto";
export async function writeJsonAtomic(path, value) {
const directory = dirname(path);
const temporary = join(directory, `.${basename(path)}.${randomUUID()}.tmp`);
const body = `${JSON.stringify(value, null, 2)}\n`;
let handle;
try {
handle = await open(temporary, "wx", 0o600);
await handle.writeFile(body, "utf8");
await handle.sync();
await handle.close();
handle = undefined;
await rename(temporary, path);
const directoryHandle = await open(directory, "r");
try {
await directoryHandle.sync();
} finally {
await directoryHandle.close();
}
} catch (error) {
if (handle) await handle.close().catch(() => {});
await rm(temporary, { force: true }).catch(() => {});
throw error;
}
}
A few details in this code are worth attention.
The wx flag means “write, but fail if the file already exists.” This guards against a UUID collision or an abandoned temp file from a previous crashed process. If someone has dropped a malicious file where your temp file should be, you will hear about it immediately rather than overwriting whatever is there.
The 0o600 permission mask creates the temp file with owner-read and owner-write only. Config files frequently hold secrets, API tokens, or private repository URLs. There is no reason to let other users on the system peek at the temporary file while it is being prepared.
Notice the separate sync calls on the file and then on the directory. Many developers skip the directory sync because it feels redundant. It is not. Ext4, APFS, and NTFS all handle directory updates differently, but they share a common habit of batching metadata writes for performance. If you care about surviving power loss, the directory sync is the final seal.
Việc dọn dẹp trong khối catch được thực hiện một cách có chủ đích để phòng vệ. Nếu có bất kỳ lỗi nào xảy ra sau khi filehandle đã được mở, mã sẽ cố gắng đóng handle và xóa tệp tạm thời, đồng thời bỏ qua các lỗi phụ để ngoại lệ (exception) gốc có thể được truyền đi một cách trọn vẹn. Bạn sẽ không muốn một lỗi phân quyền trong quá trình dọn dẹp che lấp đi lỗi thực sự đã gây ra sự cố.
Khi mô hình này không còn hiệu quả
Việc thay thế tệp nguyên tử (atomic file replacement) giúp ngăn chặn lỗi ghi dở dang (torn writes). Tuy nhiên, nó không ngăn được việc mất cập nhật (lost updates). Nếu hai thực thể CLI của bạn cùng đọc một tệp cấu hình tại một thời điểm, cả hai đều thực hiện chỉnh sửa trong bộ nhớ, cả hai đều ghi ra các tệp tạm mới và cả hai đều thực hiện lệnh đổi tên, thì lệnh đổi tên thứ hai sẽ thắng. Tiến trình đầu tiên đã không nhận thấy những thay đổi của tiến trình thứ hai. Tùy thuộc vào ứng dụng của bạn, điều này có thể có nghĩa là một người dùng thêm một cài đặt trong một terminal và một người dùng khác lại xóa nó ở một terminal khác, và tệp cuối cùng chỉ phản ánh nội dung của người ghi cuối cùng.
Nếu công cụ của bạn cần hỗ trợ các tác nhân thay đổi đồng thời (concurrent mutators), bạn cần một cơ chế điều phối bổ sung trên nền tảng ghi nguyên tử. Một tệp khóa tư vấn (advisory lock file) sẽ hiệu quả cho các trường hợp đơn giản. Version vectors hoặc một số hiệu phiên bản tăng dần (monotonic revision number) nằm ngay trong chính tệp cấu hình có thể giúp phát hiện các xung đột để người ghi thứ hai có thể thử lại. Những điều này làm tăng thêm sự phức tạp, và sự phức tạp chính là nơi các lỗi ẩn náu.
Đó là lý do tại sao ranh giới lại quan trọng. Một khối JSON duy nhất mà một tiến trình thỉnh thoảng cập nhật là một ứng cử viên tốt cho việc ghi tệp nguyên tử. Một khi bạn thấy mình phải quản lý nhiều bản ghi, áp đặt các schema, hoặc lo lắng về các thay đổi đồng thời, nghĩa là bạn đã vượt quá khả năng của hệ thống tệp. SQLite tồn tại chính vì lý do này. Nó cung cấp cho bạn các giao dịch nguyên tử (atomic transactions), nhật ký hoàn tác (rollback journals) và khả năng xử lý hợp lý các trình đọc và trình ghi đồng thời, tất cả đều nằm trong một tệp cục bộ duy nhất. Một giao thức tệp thông minh không phải là một cơ sở dữ liệu, và bạn không nên lãng phí ngân sách bảo trì để giả vờ làm điều ngược lại.
Bài học thực sự rút ra
Lần tới khi bạn định sử dụng writeFile bên trong một công cụ CLI, hãy dừng lại một chút. Tuần tự hóa (serialization) không phải là phần khó nhất. Tính bền vững (durability) mới là vấn đề. Các tệp cấu hình quá nhỏ để truyền theo dạng stream và quá quan trọng để có thể bị cắt bỏ (truncate). Hãy ghi toàn bộ payload vào một tệp anh em ẩn, flush nó, commit bằng một lệnh đổi tên, và thông báo cho thư mục về điều đó. Người dùng của bạn có thể nhấn Ctrl-C, rút dây nguồn, hoặc gập màn hình laptop. Khi máy tính hoạt động trở lại, tệp sẽ chứa trạng thái cũ hoặc trạng thái mới. Không có trạng thái lấp lửng ở giữa.
