Eight months inside a GitHub Actions merge queue teaches you something that feature comparison matrices never will. A framework can ship fifty metrics, gorgeous dashboards, and citations from respected research labs. If it blocks your deploy because a "vibe check" score drifted from 0.72 to 0.68 against identical code, it is worse than useless. It becomes an active threat to your shipping velocity.
That is the filter most LLM evaluation roundups miss. They count capabilities. They rarely ask the only question that matters in a merge queue: does this check pass and fail the exact same way every single time it runs?
I learned this by doing the uncomfortable work. I wired six open-source LLM eval frameworks into a real CI pipeline. They ran against live production pull requests for eight months. Two earned the right to remain as gatekeepers. The rest were demoted to advisory dashboards, moved to nightly jobs, or removed entirely. The lesson was sharp and expensive: deterministic structure beats probabilistic quality when you are guarding the main branch.
The Real Job of a Merge Gate
A CI gate is not a research environment. It is a bouncer. Its entire purpose is to look at a specific change and answer yes or no. Yes, this PR can join the main branch. No, it cannot. That answer needs to arrive in seconds, cost pennies, and never flip retroactively. If you rerun the same pipeline against the same commit on a quiet Tuesday and a frantic Friday, the outcome must be identical.
This is where most LLM eval frameworks stumble. They are built by data scientists for data scientists. They optimize for insight, exploration, and nuanced scoring. A merge queue optimizes for binary decisions, speed, and zero flakiness. Those two goals only partially overlap.
Why LLM-as-Judge Breaks the Queue
The tools that failed in my test shared a single design sin: they relied too heavily on LLM-as-judge calls as the primary gate mechanism.
An LLM-as-judge prompt asks a model to score an output on a scale of one to ten, or to pick the better of two responses, or to rate factual correctness. The approach is powerful for understanding quality trends. It is poison for a blocking CI check. The same input can produce different scores on different days because temperature, model versioning, and prompt formatting all introduce noise. When that score is tied to a hard threshold and a hard exit code, your queue blocks on ghosts.
The failures cascade quickly. A nondeterministic check creates queue backups. Engineers learn to retry until the number lands favorably, which trains the team to ignore red builds. Token costs pile up because every retry burns more API credits. Worst of all, the signal becomes meaningless. A red build should mean "you introduced a bug." If it means "the judge model woke up picky today," trust erodes.
What the Survivors Do Differently
Promptfoo and DeepEval survived because they treat deterministic checks as first-class citizens and LLM judge scores as secondary, non-blocking signals. They understand that a gate needs an exit code, not a floating-point number with an opinion.
Promptfoo, released under the MIT license, is built for the command line. It runs assertions like regex matches, JSON schema validation, contains checks, and exact string comparisons. These are not fancy. They are glorified grep and jq commands. That is exactly why they work in CI. A regex either matches or it does not. A JSON schema either validates or it throws. Promptfoo returns standard Unix exit codes, so GitHub Actions natively understands when to stop a merge. It is language-agnostic because it operates as a CLI tool. You do not need to install a Python ecosystem inside a Node.js service repo just to validate outputs.
DeepEval, licensed under Apache 2.0, is the choice for Python teams. It integrates like pytest. You write tests in familiar syntax, and a failure blocks the suite naturally. DeepEval offers a huge catalog of metrics, but the critical detail is that you must use them carefully. Lean on deterministic or heuristic metrics for gates. If you pull in G-Eval or other judge-based scorers, wrap them in non-blocking report generators rather than hard asserts. When used this way, DeepEval gives you the ergonomics of a testing framework without the flakiness of a research notebook.
Where the Other Four Fit
The four frameworks that did not survive as gates still have value. They simply belong elsewhere in your toolchain.
Future AGI (Apache 2.0) પચાસથી વધુ મેટ્રિક્સ પૂરા પાડે છે અને કસ્ટમ SDKs બનાવતી ટીમોને લક્ષ્ય બનાવે છે. મેટ્રિક્સ સંપૂર્ણ છે. સમસ્યા એ છે કે આ ટૂલ એવી અપેક્ષા રાખે છે કે તમે CI ક્યુ (queue) માં તેને ચલાવવા માટે તમારું પોતાનું હાર્નેસ (harness) લખો. સંશોધનના સંદર્ભમાં, તે એક વ્યાજબી સમજૂતી છે. મર્જ ક્યુ (merge queue) માં, કસ્ટમ વાયરિંગનું દરેક સ્તર અસ્થિરતાનું નવું સ્ત્રોત છે. તે એક સક્ષમ ઇવેલ્યુએશન એન્જિન છે, પરંતુ તૈયાર ગેટકીપર (gatekeeper) નથી.
RAGAS (Apache 2.0) રિટ્રીવલ-ઓગમેન્ટેડ જનરેશન (retrieval-augmented generation) ની ગુણવત્તા માપવામાં શ્રેષ્ઠ છે. તેના faithfulness અને answer relevance મેટ્રિક્સ નોલેજ બેઝ સમય જતાં કેવી રીતે કામ કરે છે તે સમજવા માટે ખરેખર ઉપયોગી છે. કમનસીબે, તે મેટ્રિક્સ LLM જજ પર ખૂબ આધાર રાખે છે. Slack પર ટ્રેન્ડ્સ પોસ્ટ કરતા નાઈટલી ક્વોલિટી જોબ માટે તે ઉત્તમ છે. પરંતુ પુલ રિક્વેસ્ટ (pull request) માટે તે નબળા બાઉન્સર છે. RAGAS ને તમારા શેડ્યૂલ કરેલા એનાલિસિસ પાઇપલાઇનમાં મૂકો, તમારા મર્જ બ્લોકર્સમાં નહીં.
Arize Phoenix એ Elastic License 2.0 ધરાવે છે અને તે સંપૂર્ણપણે અલગ વળાંક પર છે. તે ડિસ્ટ્રિબ્યુટેડ ટ્રેસિંગને ઇવેલ્યુએશન સાથે જોડે છે, જે તમને મોડેલ શા માટે ચોક્કસ રીતે વર્ત્યું તેની અવલોકનક્ષમતા (observability) આપે છે. જ્યારે તમે પ્રોડક્શન ઇન્સિડન્ટનું ડીબગિંગ કરી રહ્યા હોવ અથવા હેલ્યુસિનેશનને ખરાબ રિટ્રીવલ ચંક સુધી ટ્રેસ કરી રહ્યા હોવ ત્યારે તમને આની જરૂર પડે છે. તમે એવું નથી ઈચ્છતા કે ટ્રેસિંગ ટૂલ નક્કી કરે કે જુનિયર ડેવલપરની ફીચર બ્રાન્ચ શિપ કરી શકાય કે નહીં. તેનું આર્કિટેક્ચર ઇનસાઇટ (insight) માટે બનાવવામાં આવ્યું છે, બાઈનરી ગેટ્સ (binary gates) માટે નહીં.
MLflow Evaluate (Apache 2.0) એ તેના પ્રયોગ ટ્રેકિંગ (experiment tracking) ના વારસામાંથી આવ્યું છે. તે ભારે (heavy) છે. તેને લીન (lean) CI ઇમેજમાં લાવવાથી સ્ટાર્ટઅપ સમય અને ડિપેન્ડન્સીઝ વધે છે જે દરેક જોબને ધીમી પાડે છે. જો તમારે પાઇપલાઇનની અંદર તેનો ઉપયોગ કરવો જ હોય, તો સ્ટ્રક્ચરલ ચેક્સ માટે તેના હ્યુરિસ્ટિક મેટ્રિક્સ (heuristic metrics) પૂરતા છે. તેમ છતાં, તમે ફ્રેમવર્કની મૂળભૂત ડિઝાઇન સામે લડી રહ્યા છો. MLflow રન લોગ કરવા અને અઠવાડિયા દરમિયાન પ્રયોગોની તુલના કરવા માંગે છે. મર્જ ક્યુ એક મિનિટથી ઓછા સમયમાં ચુકાદો ઈચ્છે છે.
ગેટિંગ માટેના વ્યવહારુ નિયમો
જો તમે આ પ્રયોગમાંથી બીજું કંઈ ન લો, તો આ ત્રણ નિયમો લો.
પ્રથમ, સ્ટ્રક્ચરને ગેટ કરો, વાઇબ (vibe) ને નહીં. તમે એ અમલમાં મૂકી શકો છો કે આઉટપુટ માન્ય JSON છે. તમે એ અમલમાં મૂકી શકો છો કે તેમાં જરૂરી કી (keys) છે. તમે એ અમલમાં મૂકી શકો છો કે ક્લાસિફિકેશન લેબલ માન્ય enum માંથી છે. આ ચેક્સ ઝડપી, સસ્તા અને નિશ્ચિત (deterministic) છે. તમે વિશ્વસનીય રીતે એ અમલમાં મૂકી શકતા નથી કે સારાંશ "ફ્રેન્ડલી" છે અથવા પુનઃલેખન "ક્રિએટિવ" છે. તે ગુણો માનવ સમીક્ષા અથવા સમયાંતરે બેચ ઇવેલ્યુએશનમાં હોવા જોઈએ, ઓટોમેટેડ ગેટ્સમાં નહીં.
બીજું, જો કોઈ સ્કોર બદલાયા વગરના ઇનપુટ પર બદલાય છે, તો તેને તરત જ નીચલા સ્તરે લઈ જાઓ. સમાન આર્ટિફેક્ટ (artifact) સામે તમારી ઇવેલ્યુએશન સૂટ બે વાર ચલાવો. જો કોઈ મેટ્રિક પાસમાંથી ફેલ થઈ જાય, તો તેણે મર્જને બ્લોક કરવાનો અધિકાર ગુમાવ્યો છે. તેને એડવાઇઝરી ડેશબોર્ડમાં મૂકો જ્યાં વિચલન (variance) અપેક્ષિત અને સહન કરી શકાય તેવું હોય.
ત્રીજું, એક્ઝિટ કોડ (exit code) નો આદર કરો. લાલ બેનર સાથેનો સુંદર HTML રિપોર્ટ મર્જને અટકાવતો નથી. નોન-ઝીરો (nonzero) એક્ઝિટ કોડ અટકાવે છે. તમારા ઇવેલ્યુએશન ટૂલે તમારા CI પ્લેટફોર્મની મૂળ ભાષામાં વાત કરવી જોઈએ. Standard out મનુષ્યો માટે છે. એક્ઝિટ કોડ મશીનો માટે છે.
મુખ્ય તારણ
LLM-સંચાલિત એપ્લિકેશન્સનું પરીક્ષણ કેવી રીતે કરવું તે સમજવામાં આપણે હજુ શરૂઆતના તબક્કામાં છીએ. ઇવેલ્યુએશનને માનવ ગ્રેડિંગ રૂબ્રિકની જેમ લેવાની લાલચ થાય છે: સૂક્ષ્મ, સંદર્ભિત અને થોડું વ્યક્તિલક્ષી. તે રિસર્ચ પેપરમાં કામ કરે છે. મર્જ ક્યુમાં તે નિષ્ફળ જાય છે.
આઠ મહિનાના પ્રોડક્શન ટ્રાફિક પછી, મારી પાઇપલાઇન હવે સર્વિસિસમાં સ્ટ્રક્ચરલ અને સ્કીમા એસરશન માટે Promptfoo અને પાઇથોન-સાઇડ બિહેવિયરલ ચેક્સ માટે DeepEval ચલાવે છે જે પાસ-ફેલ કન્ડિશન સાથે સ્પષ્ટ રીતે મેપ થાય છે. બાકી બધું નાઈટલી ડેશબોર્ડ્સમાં રિપોર્ટ થાય છે. ક્યુ સ્થિર છે. સિગ્નલ ચોખ્ખું છે. ટીમ ફરીથી રેડ બિલ્ડ પર વિશ્વાસ કરે છે.
તમારે તમારા ગેટ પર વધુ મેટ્રિક્સની જરૂર નથી. તમારે એવા ઓછા મેટ્રિક્સની જરૂર છે જે દર વખતે સત્ય કહે છે.
મૂળ પરીક્ષણ અને Dev.to પર શેર કરાયેલ લેખ પર આધારિત. વિશ્વસનીય AI સિસ્ટમ્સ બનાવવા પર વધુ ચર્ચા માટે, Telegram પર GyaanSetu કોમ્યુનિટીમાં જોડાઓ.
