The Founder Lab Shopify store was scanned six times with an AI-driven scoring tool, and only the “intent” component of the score moved—shifting between 4 and 6 while the overall result stayed effectively the same. The finding shows that a single-point change in a shop’s AI-generated rating can be pure statistical noise, not a real improvement.
Why the test mattered
Shop owners increasingly rely on automated tools that assign a single numeric rating to their sites. Those ratings promise a quick health check and a roadmap for optimization. The assumption is that a higher number equals a better store, so any uptick is taken as evidence that a change worked. Founder Lab wanted to verify that premise. By running the tool on an unchanged store, the team could see how much the score fluctuates on its own.
What the experiment looked like
- Date 1 (July 12): One scan, total score 78, intent 6.
- Date 2 (July 17): Five scans, each under a minute, producing the following results:
- Scan 1: 76 / 4
- Scan 2: 76 / 4
- Scan 3: 78 / 6
- Scan 4: 77 / 5
- Scan 5: 77 / 5
Every other sub-score—visual design, schema markup, trust signals, technical health, pricing, recommendation, and brand—remained identical across all runs. Only the intent sub-score changed, and the total score after subtracting intent (78 – 6, 76 – 4, 77 – 5) always equaled 72. In other words, the deterministic parts of the evaluation were perfectly stable; the only variable was the AI-driven intent assessment.
The hidden volatility of “intent”
Intent is the portion of the algorithm that tries to gauge how well a store aligns with a shopper’s purpose—whether the product mix, copy, or overall messaging matches what a buyer is looking for. When the total rating appears as a single number, that variance is easy to miss. A store that moves from 78 to 80 might look improved, yet the change could be nothing more than the same 2-point swing that the Founder Lab test observed.
Why developers should run multiple scans
The experiment suggests a practical rule of thumb: treat any single-scan change of one or two points as suspect until it can be reproduced. Running the tool several times on the same snapshot of a site gives a “noise floor” – the range of scores you can expect when nothing has changed. If a new optimization pushes the score outside that range consistently, you have stronger evidence that the change mattered.
We no longer trust a single rescan. Every real change now requires multiple scans to prove it is not noise. The focus also moved upstream: they first lock down the deterministic factors—technical health, structured data, and site architecture—because those elements are stable and directly under the developer’s control. Only after those foundations are solid do they experiment with product copy or other intent-related signals, and even then they verify results with multiple scans.
The stakes for merchants
For a shop owner, a false sense of progress can lead to wasted time and money. If you chase a perceived 2-point boost, you might invest in copy rewrites, image swaps, or pricing experiments that don’t actually move the needle. Conversely, ignoring a real improvement because it falls within the noise band could mean missing an opportunity to double down on a winning change.
Counter-argument: single scans still have value
Some practitioners argue that a single scan is sufficient for high-level monitoring, especially when resources are limited. They point out that the deterministic components (technical, schema, trust, etc.) are already reliable, and the intent score, while noisy, still offers directional insight. In fast-moving environments where waiting for multiple scans could delay action, a single reading may be the only practical option.
The trade-off is clear: a single scan gives speed but carries the risk of reacting to noise; multiple scans provide confidence at the cost of time. Merchants need to decide which balance fits their workflow and risk tolerance.
What to watch next
- Tool providers: Will scoring platforms publish their own noise estimates or suggest a minimum number of runs for reliable results? Transparency around model variance could become a differentiator.
- Shopify ecosystem: As more merchants adopt AI scoring, community best practices around repeat testing may emerge, possibly in the form of plugins that automate multiple scans and aggregate the data.
Takeaway
Penarafan kedai yang dijana oleh AI hanya boleh dipercayai setakat tahap konsistensi model asasnya. Eksperimen enam imbasan oleh Founder Lab menunjukkan bahawa bahagian "niat" boleh berubah sebanyak beberapa mata manakala elemen lain kekal tidak berubah. Oleh itu, pembangun harus menganggap perubahan skor yang kecil sebagai gangguan (noise), mengesahkan penambahbaikan melalui beberapa imbasan, dan mengutamakan pembaikan aspek kedai yang bersifat deterministik serta boleh dikawal sepenuhnya sebelum cuba mengubah bahagian yang sensitif terhadap AI. Usaha tambahan pada masa ini dapat mengelakkan pembangun daripada mengejar keuntungan palsu pada masa hadapan.