Four independent judges read one run. None can overrule another, none is averaged into the others, and every one of them is allowed to say it does not know — which is what makes the times they do answer worth acting on.
A field has a value and a provenance — the DOM path the value was read from, plus a structural fingerprint of the region around it. Comparing both against the previous run produces four outcomes, and the third one is the entire reason this product exists.
Nothing happened. The overwhelming majority of runs, and the reason a healthy source is drawn silently.
an ordinary monitor sees thisThe page changed its price. Reporting this as drift is how an operator learns to close the alert without reading it.
an ordinary monitor sees thisThe warning. The value is still correct and the extraction has moved under it. Nothing else in the stack has a name for this.
nothing else sees thisThe break, arrived. By now the wrong number is already in somebody’s dashboard.
an ordinary monitor sees thisEach of these is drawn the same way everywhere it appears: reduced ink, dashed, left open. A reading with nothing behind it must never be comparable to a reading with something behind it — which is why none of them is a paler shade of a verdict.
Nothing checked this value. There was no plan spec to judge it against, or no history to judge it with. It is not a mild version of questionable.
The evidence does not carry a cause. The diagnosis engine will name a redesign, a removed element or a real price change — and where it cannot, it says so instead of choosing the most likely one.
Verification could not decide. It is a distinct outcome from failed, and it is treated as a failure for the purpose that matters: only a verification that passed can adopt a repair.
Too few comparisons to grade how settled a field is. Below the threshold the archive reports nothing rather than a stability grade computed from two data points.
Collection is the first stage and it is the only one this project did not build — Bright Data runs it. Everything from stage 02 onward is the judgement layer, and stage 07 is the only door between a proposal and production.
A run arrives with a value and a structural digest. Validation asks whether the number is believable against this field’s own history. Trust asks whether the place it came from has moved. Diagnosis reads both verdicts — never their underlying rows — and names a cause with a confidence attached.
Only then is a repair proposed — a row in a table nothing in the collection path reads. It becomes real only when a verification against recorded evidence passes and promotes it into a new immutable plan version.
Three judges content, one dissenting. The ring never resolves this into a single number.
A language model is used for exactly one thing: choosing between selectors that were already observed on the page. It cannot invent a selector, and it cannot publish one. Both limits are enforced by the shape of the code rather than by review.
The second guard has fired in production. On the first live call of the recovery tier, the model chose the correct selector over the struck-through-price trap at 0.95 confidence — and a separate fabricated selector offered at 0.99 was rejected, because nothing had observed it on the page.