Probabilistic testing means every check a model makes returns how likely it is to hold, and the verdict comes from comparing that probability with a threshold you set. The test still passes or fails, but you can see how sure it was.This page explains why Sedum works this way, how scores become verdicts and exit codes, and how to tune it.
Every AI test is already probabilistic
When a model reads a page and decides whether “a list of products with prices is shown”, it is making a judgment, and it can be more or less sure. That is true for every tool that uses a model to test a page. Most tools ask the model for a plain yes or no and treat the answer as a fact. The uncertainty does not go away. It shows up later as a test that passes on Monday and fails on Tuesday with nothing changed and no explanation. So the choice is not between probabilistic and deterministic tests. It is between showing the uncertainty and hiding it. Sedum shows it.What Sedum measures
Averify step asks the model two separate questions about the page:
- Holds: how likely is it that the claim is true?
- Contradicted: how likely is it that the page shows evidence against it?
From scores to verdicts
The defaults for averify step are:
A pass also gets the
contradiction flag when the contradicted score is 0.50
or higher. A failed step keeps its scores but has no flags. Comparisons use the
exact scores, not rounded ones.
A flagged pass is still a pass. The flags tell you where to look, without
breaking the build for a check that probably holds.
From verdicts to exit codes
CI needs a clear answer, so every run ends in one exit code:--strict changes only the exit code and the JUnit mapping. The verdicts and
flags in the result stay the same, so a report reads the same whichever mode
you ran in.
This gives you three outcomes instead of two. A failure means the page did
not show what the test expected. A flag means the model was not sure. Exit 3
means Sedum could not check at all. Your CI can block on the first, warn on the
second, and retry the third.
Why this is better
- “Not sure” is its own answer. A normal test reports “the product broke” and “the test could not tell” the same way. Sedum keeps them apart, so people look at real failures first.
- Flakiness shows up before it breaks the build. A yes or no test goes green, green, red, green, and you learn nothing until it fails. A score moves first. See catching flakiness early.
- Cheap is safe. A fast, inexpensive model is only trustworthy if it says when it is unsure. Scores let Sedum accept confident answers and flag the rest, which is what makes it safe to ask a model on every step of every run.
- You set the trade-off. Every test suite balances false alarms against missed bugs. Brittle selectors cause false alarms; loose assertions miss bugs. A threshold turns that balance into one setting you can see and change.
- Bad tests show up. A high contradicted score usually means the sentence is ambiguous, or the page is. The fix is often to rewrite the step, and the report tells you which step.
Catching flakiness early
A flaky test is one that passes and fails on the same code. With yes or no checks, flakiness is noise: you see a red build, rerun it, get green, and move on. Nothing tells you whether it will happen again, or which step is to blame. A score is a continuous signal. Consider one claim across five runs:
A yes or no test shows four passes and then a surprise. The scores show the
check getting weaker from run 3, and the flag in run 4 says so before anything
fails. Common causes are a page that renders slowly or differently, a change in
the product that the claim no longer quite describes, or a claim that was vague
from the start.
Scores also tell you what kind of flaky a step is:
- A score that jumps around a threshold means the claim is borderline on this page. Make the claim more specific, or check that the page really shows the same thing every time.
- A score that is steady but low means the claim never fitted the page well. Rewrite it.
- A high contradicted score means the page shows evidence for and against the claim. Either the claim is ambiguous or the page is.
.sedum/runs/<run-id>/result.json. To see how
each claim scored across your recent runs:
--retries reruns a failed test from a
fresh browser and keeps every attempt in the result. Instead of rerunning until
green, you can compare the attempts’ scores and see whether the failure was
borderline or clear.
Tuning thresholds
Set the thresholds insedum.config.yaml. They apply to every verify step in
the project:
- Raise
verifyto fail more often on weak evidence. You will catch more bugs and see more false alarms. - Narrow
lowConfidenceBandto turn borderline passes into failures instead of flags. - Run with
--strictin CI if any flag should block a merge.
Observing without a verdict
A step that starts withmeasure, note, or observe records both scores
without deciding the test:
verify, or to track something you want to watch but not gate on.
Reading the scores
For every failed or flagged step, the Markdown report shows the holds score against the fail and pass lines, the contradicted score against its cutoff, and the page text the model judged. When an element could not be found, it lists the ranked candidates, including “no match”. See CLI commands for the HTML, Markdown, JSON, and JUnit reports.Writing claims that score well
A claim scores clearly when it names something a person could point to on the page.
If a step keeps getting
low_confidence or contradiction, read the page
text in the report. Usually the claim is vague, or the page really does say
two things.
The judge reads the page’s text, not its pixels. Colors, icons without a
name, images and layout are not in that text, so a claim about them cannot be
confirmed. One visual case is covered: a bare count on an icon-only control,
such as the “1” on a cart icon, appears as cart icon badge: 1 when the
control’s class, id or test hook names a known icon (cart, notifications,
menu, and similar). For anything else visual, claim the text the change
produces instead, for example verify the Sauce Labs Backpack button now says Remove rather than a claim about a colored badge.
What the scores are not
The scores are the model’s confidence, not measured frequencies. A score of 0.9 does not yet mean the claim is right 9 times out of 10. The default thresholds are set to be safe, and the flags exist so that uncertain answers reach a person. Treat the numbers as a ranking of how sure each check was, and use the thresholds to decide what to do about it.Related
- Plain-English browser tests: the basics.
- Assertion engine: the exact verdict rules.
- Project configuration: all
sedum.config.yamlkeys.