If your email tests run perfectly on your laptop and collapse the moment they hit CI, you are not alone. The usual response is to sprinkle sleep calls through the test code or bump the retry count until the build passes. That might quiet the noise for a day, but it does not fix the bug. It only hides it.
The real issue is how your test identifies which email to open.
The Shared Inbox Problem
On your local machine, you run one test at a time. One email arrives. You grab it. Simple.
CI is a different environment entirely. A single pull request might trigger four, eight, or sixteen parallel jobs. If they all share a test inbox—whether that is a Mailosaur server, a Mailtrap inbox, or a real account on a staging domain—they are all writing to the same bucket at the same time. Job A sends a password reset. Job B sends an invite. Job C retries a failed welcome flow. Meanwhile, background workers and delivery queues add jitter that you cannot control.
When every job reaches into that shared inbox and asks for the newest message with the subject "Reset your password," it becomes a race. The test that wins gets the right email. The test that loses clicks a link meant for another job, asserts against the wrong content, and fails with an error that looks like a timing problem. It is not a timing problem. It is an identity problem.
Why "Newest Message" Fails
The brittle pattern is easy to fall into because it feels intuitive:
- Trigger the user flow.
- Poll the inbox every few seconds.
- Open the most recent message that matches the subject line.
- Click the first link and run assertions.
This falls apart for several reasons beyond simple parallelism. A retry from a previous failed run can land late, suddenly becoming the newest message just as your current test polls. Background workers inside your application might queue two emails and deliver the second one before the first. Subject lines alone are weak identifiers; your staging application might send similar emails from different paths. Sorting by timestamp is worse than it looks because clock skew between the CI runner and the mail provider is real, and mail APIs often cache or batch their indexes.
Timestamps get fuzzy in busy environments. You need something direct.
What a Run Token Actually Is
A run token is nothing more than a unique string generated at the start of your test and injected into the email your application sends. It does not need to be user-facing, and it does not need to look elegant. It only needs to guarantee that you can prove this specific message belongs to this specific test execution.
Concrete examples work best. Before the test starts, generate a token such as:
- A UUID:
550e8400-e29b-41d4-a716-446655440001 - A build-scoped request ID:
req_ci_build_4821_a7f3 - An invite slug or metadata suffix:
signup-token-8k2m9n - A random hex string generated by the test runner:
test-run-a4f9c2d1
If you control the backend code, pass the token into the email context and render it somewhere in the body. If you are testing against a black-box application, see if the app already accepts a reference field you can hijack. If not, you can sometimes embed the token in the recipient local-part using plus addressing—testuser+a4f9c2d1@example.com—though that only works if your application preserves and echoes it back in the email.
The point is to stop matching on metadata the mail system already owns. Match on data your test owns.
The Reliable Pattern
Replace the "newest message" algorithm with a narrow, token-driven search:
- Generate the run token before you trigger any flow.
- Start the user action, ensuring the application will include the token in the outbound email.
- Poll the mail provider with filters constrained to that token. If the API supports body search, use it. If not, fetch candidate messages and grep their bodies client-side.
- Assert that the token exists in the message body before you touch any links, buttons, or verification codes.
- Only then extract the confirmation URL or code and continue.
This sequence matters. If you extract a link first and check the token second, you have already clicked the wrong email. The assertion is your gatekeeper.
Về mặt thực tế, hàm helper của bạn nên tìm kiếm theo kiểu Subject:"Welcome to AppName" AND Body:"a4f9c2d1" thay vì Subject:"Welcome to AppName" sort:-received. Nhiều dịch vụ kiểm thử email cung cấp các API tìm kiếm cho phép lọc theo nội dung body. Hãy tận dụng chúng. Nếu bạn đang làm việc với một nhà cung cấp đơn giản hơn, hãy tập trung logic polling vào một nơi duy nhất để có thể áp dụng bộ lọc phía client một cách nhất quán cho mọi bài kiểm thử.
Ba quy tắc để giữ cho hệ thống luôn đáng tin cậy
Một run token giúp cố định đối tượng được chọn, nhưng bạn vẫn cần sự kỷ luật trong cách thức polling và cách xử lý khi có lỗi xảy ra.
Ghi lại trạng thái inbox khi gặp lỗi. Khi một bài kiểm thử thất bại, hãy xuất ra định danh inbox, dòng tiêu đề bạn đã truy vấn, khoảng thời gian (timestamp) chính xác và số lượng tin nhắn khớp với tiêu chí của bạn. Điều này biến một lỗi "email not found" mơ hồ thành một câu chuyện cụ thể. Nếu job 7823 lấy nhầm tin nhắn thử lại từ job 7821 vì nó đến trễ ba giây, nhật ký của bạn phải làm rõ được điều đó. Nếu không có ngữ cảnh này, bạn sẽ đổ lỗi cho vấn đề thời gian và lại thêm một lệnh sleep nữa.
Giữ tất cả logic polling email trong một file helper duy nhất. Đừng rải rác các lệnh gọi setTimeout và cy.task khắp hai mươi file kiểm thử. Hãy tập trung hóa logic chờ tin nhắn, thử lại lệnh gọi API và áp dụng cơ chế backoff. Nếu mọi bài kiểm thử đều dùng chung một helper, các quy tắc lọc của bạn sẽ luôn nhất quán, và khi bạn cải thiện logic tìm kiếm, mọi bài kiểm thử đều được hưởng lợi. Điều này cũng giúp việc thực thi kiểm tra token dễ dàng hơn; nếu helper yêu cầu một đối số token, sẽ không ai có thể vô tình quay lại sử dụng "điểm tựa" là tin nhắn mới nhất.
Chú ý đến các lần thử lại (retries). Việc thử lại bài kiểm thử là phổ biến trong CI, nhưng mỗi lần thử lại sẽ tạo thêm một email khác trong inbox. Nếu bài kiểm thử của bạn vượt qua ở lần thử thứ ba, bạn có thể ăn mừng và bỏ qua. Nhưng điều bạn bỏ lỡ là lần thử thứ nhất và thứ hai đã bộc lộ một lỗi thực sự—một race condition, một lần gửi trùng lặp, hoặc một index bị thiếu—mà các tin nhắn dư thừa đã che lấp mất. Nếu bắt buộc phải dùng retry, hãy kiểm tra xem inbox có chứa các tin nhắn trùng lặp không mong muốn sau khi thất bại hay không. Tốt hơn hết, hãy cân nhắc việc dọn dẹp inbox hoặc sử dụng một địa chỉ duy nhất cho mỗi job nếu nhà cung cấp của bạn hỗ trợ dynamic inboxes. Retry không nên trở thành một chiến thuật để lấp liếm cho logic chọn đối tượng không đáng tin cậy.
Bài học cốt lõi
Sắp xếp inbox theo ngày và lấy kết quả đầu tiên không phải là kiểm thử. Đó chỉ là sự đoán mò được khoác lên mình lớp vỏ mã nguồn. Một run token gần như không tốn kém gì—chỉ cần một biến string, một tham số lọc bổ sung, hoặc có thể là một thay đổi nhỏ trong template—nhưng nó mang lại cho bài kiểm thử của bạn một định danh xác định (deterministic identity). Nó chứng minh rằng tin nhắn trước mặt bạn thuộc về lượt chạy mà bạn đang thực hiện ngay lúc này.
Đừng thêm các lệnh sleep và hy vọng mạng sẽ hoạt động ổn định nữa. Hãy tạo một token, đưa nó vào email và tìm kiếm trực tiếp. Các lượt chạy CI của bạn sẽ nhanh hơn, nhật ký sẽ dễ đọc hơn, và cuối cùng bạn sẽ thực sự tin tưởng vào những gì bộ công cụ kiểm thử email đang báo cáo cho bạn.
