เริ่มต้นด้วยข้อความใน Slack บิลด์เป็นสีแดง คุณเลื่อนดูรายการที่ล้มเหลว ขมวดคิ้ว แล้วลองรันเทสต์เดิมบนแล็ปท็อปของคุณ ผลคือสีเขียว คุณลองรัน CI job อีกครั้ง เผื่อว่ามันจะเป็นแค่ความผิดพลาดชั่วคราว แต่ความล้มเหลวก็ยังกลับมา มันยังคงดื้อรั้นและเกิดขึ้นซ้ำๆ บนเซิร์ฟเวอร์ แต่กลับมองไม่เห็นเมื่อรันในเครื่องคุณ

การทดสอบเบราว์เซอร์ที่ล้มเหลวใน CI แต่ผ่านในเครื่องตัวเองนั้นเป็นมากกว่าแค่ความน่ารำคาญ มันสร้างความไม่ไว้วางใจ ทีมงานเริ่มโทษเรื่องจังหวะเวลา (timing) พวกเขาเริ่มใส่การแก้ไขชั่วคราวที่ไม่เคยถูกลบออกไป เช่น ใส่ setTimeout ตรงนั้น หรือ .wait(5000) ตรงนี้ ชุดการทดสอบ (test suite) ก็ช้าลง ความล้มเหลวก็ยังคงวนเวียนกลับมา เทสต์ที่เอาแน่เอานอนไม่ได้ (flaky tests) เหล่านั้นจะกลายเป็นส่วนหนึ่งของระบบอย่างถาวร และในที่สุดทุกคนก็เริ่มมองว่า pipeline ที่เป็นสีแดงเป็นเพียงเสียงรบกวนที่คุ้นเคย

นั่นเป็นเรื่องอันตราย คุณคงไม่ต้องการชุดการทดสอบที่ "ร้องเตือนหลอกๆ" (cries wolf)

CI ไม่ได้เสีย แต่มันแค่แตกต่าง

สภาพแวดล้อมของ CI ไม่ใช่เรื่องสุ่ม แต่มันเป็นแบบ deterministic ปัญหาก็คือมันเป็น deterministic สำหรับระบบที่ไม่ใช่ MacBook หรือเวิร์กสเตชัน Linux ของคุณ การตั้งค่าในเครื่องของคุณซ่อนความแตกต่างที่ CI runner ที่สะอาดสะอ้านจะเปิดเผยออกมาทันที

ลองคิดดูว่ามีส่วนประกอบต่างๆ มากมายแค่ไหนที่แตกต่างกัน เครื่องของคุณอาจรัน development server ที่มี hot module reloading ในขณะที่ CI สร้าง production artifact ที่มีการทำ tree shaking และ minification ซึ่งลำพังแค่เรื่องนี้ก็สามารถตัดเส้นทางการทำงานของโค้ด (code paths) หรือเปลี่ยนลำดับการทำงานได้แล้ว ต้นไม้ของ dependency (dependency trees) ก็เปลี่ยนไป lockfile ที่ดูเหมือนกันเป๊ะอาจจะ resolve ต่างกันหากเวอร์ชันของ package manager ต่างกันเพียงแค่ minor release เดียว ลำดับของเครือข่ายก็เปลี่ยนไป Wi-Fi ในออฟฟิศของคุณอาจจะเชื่อมต่อกับ staging API ได้ในขั้นตอนเดียว แต่ CI runner อาจจะวิ่งไปหา cluster อื่นที่อยู่หลัง load balancer ซึ่งทำให้เกิด latency ที่คุณไม่เคยเจอ

ตัวเบราว์เซอร์เองก็ทำงานต่างกันในแต่ละสภาพแวดล้อม Chrome ในเครื่องของคุณอาจมี extension, ข้อมูลการล็อกอินที่แคชไว้, local storage ที่ค้างอยู่ และ GPU ที่มี hardware acceleration แต่ CI จะเริ่มจากโปรไฟล์ที่ว่างเปล่าในทุกๆ การรัน วงจรชีวิตของเบราว์เซอร์ (browser lifecycles) แตกต่างกัน เส้นทางการเรนเดอร์ (rendering paths) ก็ต่างกัน ฟอนต์ที่มีอยู่ในเครื่องของคุณอาจถูกแทนที่ใน CI ขนาดของ viewport และ device pixel ratios ก็ต่างกัน ซึ่งอาจทำให้จุดตัดการตอบสนอง (responsive breakpoints) เปลี่ยนไป หรือเปลี่ยนพฤติกรรมของ lazy-loading ได้

ช่องว่างเหล่านี้มีอยู่จริง มันเป็นเรื่องทางเทคนิค การแสร้งทำเป็นว่ามันเป็นเรื่องสุ่มไม่ได้ช่วยให้มันหายไป

Preview Environments หลอกคุณ

Preview environments ยิ่งทำให้ปัญหาซับซ้อนขึ้น แม้มันจะมีประโยชน์สำหรับการตรวจสอบโดยมนุษย์ แต่มันไม่ใช่ production บ่อยครั้งที่มันชี้ไปยัง api-staging แทนที่จะเป็น host ของ API จริง Feature flags อาจจะถูกตั้งค่าเป็น true สำหรับทุกการทดลอง ซึ่งซ่อน logic เงื่อนไขที่ production ต้องรัน การยืนยันตัวตน (authentication) อาจข้ามบางขั้นตอนหรือใช้ mock token คุกกี้อาจใช้ policy ที่ผ่อนปรนกว่า และชุดข้อมูลอาจจะมีแค่เพียงส่วนเล็กๆ เช่น มีแค่สิบแถวแทนที่จะเป็นหมื่นแถว ซึ่งหมายความว่า logic ของการทำ pagination, การจัดอันดับการค้นหา หรือ virtualization จะไม่ถูกทดสอบเลย

หากการทดสอบของคุณผ่านเมื่อรันกับ preview URL แต่ล้มเหลวใน production หรือในทางกลับกัน ตัวการไม่ใช่บั๊กในเทสต์ แต่เป็นที่สภาพแวดล้อม (environment) ต่างหาก

บันทึก Log ก่อนที่จะเดา

เมื่อความล้มเหลวปรากฏขึ้นครั้งแรก จงหักห้ามใจไม่ให้รีบแก้เทสต์แล้วภาวนาให้มันผ่าน หยุดเดาสุ่ม คุณต้องหยุดสถานะ (freeze) ของบริบทไว้ เพื่อที่คุณจะได้เปรียบเทียบการรันที่ผ่านกับครั้งที่ล้มเหลวได้

บันทึกข้อมูลที่น่าสงสัยให้ชัดเจน บันทึก URL ของหน้าเว็บในขณะที่ล้มเหลว, build ID และ commit SHA จดบันทึก feature flags ที่ใช้งานอยู่ บันทึก API host, เวอร์ชันของเบราว์เซอร์ที่แน่นอน และขนาดของ viewport รายละเอียดเหล่านี้จะเปลี่ยนความล้มเหลวที่ลึกลับให้กลายเป็นเงื่อนไขที่สามารถทำซ้ำได้ (reproducible condition)

อย่าพึ่งพาแค่ภาพหน้าจอ (screenshots) เพียงอย่างเดียว หน้าเว็บสองหน้าอาจดูเหมือนกันเป๊ะในระดับพิกเซล แต่รัน JavaScript ที่ต่างกันโดยสิ้นเชิง ภาพหน้าจอจะไม่บอกคุณว่า bundle ของ CI มี polyfill ส่วนเกินเข้ามา หรือ bundle ในเครื่องคุณข้ามบาง chunk ไปเพราะมันอยู่ในแคชของเบราว์เซอร์แล้ว

และอย่าลืมว่าการเปิด DevTools จะทำให้จังหวะเวลา (timing) เปลี่ยนไป DevTools สามารถทำให้การทำ garbage collection ล่าช้าลง, เปลี่ยนลำดับความสำคัญของเครือข่าย และปิดการเพิ่มประสิทธิภาพการเรนเดอร์บางอย่าง เทสต์ที่ผ่านในขณะที่คุณกำลังตรวจสอบ DOM อาจล้มเหลวทันทีที่คุณปิดแผงควบคุมและรันแบบ headless แม้ debugger จะเป็นเครื่องมือที่มีประโยชน์ แต่มันไม่ใช่ผู้สังเกตการณ์ที่เป็นกลาง

จำลองสถานที่เกิดเหตุ

หากคุณต้องการจำลองความล้มเหลวให้แม่นยำ คุณไม่สามารถแค่รัน development server ในเครื่องแล้วหวังว่าทุกอย่างจะดีขึ้น คุณจำเป็นต้องจำลองเงื่อนไขที่เหมือนกับ the_CI's_ ทุกประการ

Build the exact artifact that CI produced. Download it if you must. Serve that artifact locally with a simple static file server, not with Vite or Webpack dev middleware. Use the same environment variables that CI injected. Match the browser version precisely. Run it in the same mode, headed or headless, because focus events, media queries, and autoplay policies still diverge between the two in subtle ways. If your CI uses a Docker container, run the same image locally. Remove your personal browser profile entirely.

When the local reproduction finally fails, you have a real debugging session. Until then, you are chasing shadows.

Stop Sleeping, Start Waiting

The most common response to a flaky browser test is to add delay. Wait five seconds. Wait ten. This is not a fix. It is a surrender. Arbitrary delays slow your suite, create false confidence, and still fail under load when the network hiccups.

Instead, wait for evidence of state. If a notification should appear after a form submission, do not wait for time to pass. Wait for a specific notification ID to exist in the DOM. If a counter should increment, wait for the text to change values. If a loading state blocks interaction, wait for the loading marker to disappear. If you are working with a WebSocket or server-sent events, wait for the network stream to produce a specific event.

Explicit waits turn your test from a guessing game into a contract. The test says: "I will proceed once the application confirms it is ready." That is far stronger than saying, "I will proceed once enough seconds have passed."

Hydration and the Disappearing Button

In modern React applications, hydration causes a specific class of failures that local dev servers often mask. The server sends HTML. React boots up in the browser and attaches event listeners. During that window, your test might click a button. React then replaces or restructures that DOM node during hydration. The element handle your test framework was holding now points to a detached node, and you get an error about interacting with a removed element.

The fix is not to write a more complex selector that digs deeper into the component tree. The fix is to look for readiness signals. Wait until a root element gains a hydrated attribute or a known data property. Wait for a skeleton loader to disappear. Wait for a client-side event handler to become active. Let the application announce that it is stable before you fire clicks.

Hidden Culprits: Dependencies and Third-Party Scripts

Sometimes the environment changes even though your application code did not. A transitive update in a small utility library, three levels deep in your node_modules, can alter browser behavior. It might change how promises resolve, how styles get injected, or how mocks intercept requests. When tests begin failing after a routine dependency update, record your package manager version and the lockfile checksum. You need to know whether you are looking at the same tree you were last week.

Third-party scripts are another frequent saboteur. Analytics trackers, payment SDKs, and chat widgets load asynchronously. They inject iframes, shift layout, or steal focus at moments your test does not expect. In CI, these scripts might load more slowly, or they might fail to load entirely because of network restrictions, causing your application to follow a different error-handling path. Log which third-party resources loaded and their HTTP status. If a payment iframe takes three seconds to mount in CI but loads instantly on your fast local connection, your "element not clickable" error suddenly has a clear cause.

And "element not clickable" is never a diagnosis. It is a symptom. Treat the cause.

Build an Evidence Kit

Every CI failure should be actionable. A stack trace alone is not enough. You need an evidence kit that lets another engineer, or yourself next month, reconstruct what happened.

Keep screenshots and video recordings from the failing run. Capture the full browser console output, not just errors but warnings too. Log network failures, including 404s, CORS rejections, and dropped connections. Preserve the build IDs and feature flags that were active. Take a DOM snapshot at the exact moment the assertion failed. A snapshot lets you inspect the HTML structure after the fact, rather