tests[].goal with complete, secret-redacted
text and verify strings after data resolution, before browser execution.
HTML reports display these above the replay and steps. Authored-step tests omit
this field. Existing result files without it remain supported.
TypeScript ai.goal is represented on its aggregate action step by optional
step.goal metadata. Its completion is "planner", with the redacted goal
text, actions, requests, completion reason, and optional dataSeed.
This explicitly means accepted planner DONE, not independent verification.
It carries no Judge scores and must not be displayed as “verified.” A goal-only
TypeScript test can therefore pass with planner completion; any subsequent
ai("verify ...") or failed code assertion is a separate step and determines
the test outcome normally.
YAML goal tests publish an action step after each dispatched click or type, then an
aggregate step (kind: verify, operation: goal) when the bounded loop ends.
Action steps use data-binding names instead of typed values and include target
scores, incremental planner calls, and opt-in post-action replay frames subject
to normal privacy rules. The aggregate includes action/request counts and the
termination reason in detail, independent Judge scores when reached, and only
remaining calls with purpose: planner or judge, avoiding duplicate usage.
A flagged goal assertion fails the goal even without --strict. The aggregate
omits page metadata, screenshots and replay frames. Goal failures are not
automatically restarted by --retries.
That required, independent, unflagged verification policy applies to YAML goal
tests and the exported runGoal API. It is unchanged by TypeScript
ai.goal’s planner-completion contract. For TypeScript, entering a goal also
disables automatic whole-test retries for every later failure in that attempt,
including a successful goal followed by failed verification or cleanup.
@sedum-dev/core/run-result.schema.json is the Draft 2020-12 schema generated from the same Zod definition used to validate every published snapshot. The schema ID is https://sedum.dev/schemas/run-result/v1. The v1 schema may still change incompatibly before its first published release. After v1 is published, incompatible changes require a new major schema version. Consumers should reject unknown major versions.
A step from a TypeScript test carries group, the enclosing ai.group names
from the outermost, when it ran inside a group. An exception thrown by test
code is one failed step with kind: verify and operation: code, whose error
code is expect_failed for a failed expect() and code_error otherwise. A
TypeScript test’s id is <file>#<title> unless it sets an id, and its
description is the title.
Each attempt problem has a unique ID, consecutive attempt-local ordinal, phase, ordered source stack, outcome, typed error, and an origin of step or module_binding. A step problem links its executed step; a module binding problem has stepId: null because no sentence ran. The first encountered problem is primary. A failed setup or body stays primary even if later teardown reports an operational error. A primary operational error leaves the verdict null. Executed steps retain their own phase and source stack; there are no synthetic module-call steps.
For a syntactically valid sedum run, the CLI resolves project configuration,
then creates <outputDir>/<run-id>/progress.json before provider setup or test
loading and prints its path. The default outputDir is .sedum/runs. Config
must be read first because it chooses this location; an invalid config uses the
default location for its terminal operational result when that directory is
writable. Each update replaces the file atomically with a complete, validated
RunResult; pollers can parse it at any point. result.json is written from
the terminal snapshot. A setup or execution error has state: "error", a typed
error, and verdict: null; already completed tests and steps remain. Zero
executed tests are never a pass. A passed step with low_confidence or
contradiction stays passed; --strict may change the exit status and JUnit
mapping, not the canonical verdict.
SIGINT and SIGTERM received while a run is active request cancellation. The CLI waits for the current operation to unwind, then writes state: "interrupted", verdict: null, and a terminal result.json before exiting with the conventional signal exit code. Once the terminal outcome is committed, late signals are ignored while its final files are written so the process exit and persisted state cannot disagree. An interrupted provider request may have incurred unreported usage; its call is retained with unknown cost, so a receipt never presents that attempt as free.
The CLI derives ordinary run exits from this result: completed clean pass 0,
completed flagged pass 0 or 2 with --strict, failed 1, and a null or
untrustworthy verdict 3. Strict mode changes only that shell gate. Terminal
and redirected summaries show the same ordered test identities, verdict/state,
flags, and totals; only ANSI/live updates and the default visibility of
token/cost lines differ.
If the output sink itself fails, exit 3 is reported from the latest validated
in-memory snapshot. Sedum names the intended path and does not advertise an
earlier progress file as an authoritative final result. This is the sole case
where terminal artifact persistence cannot be promised through the failed
destination.
The model has ordered tests, whole-test attempts, steps, and observation attempts. Only the selected terminal attempt contributes to final test/step outcome counts; all attempts contribute to model usage and cost. Unknown actual cost is null, not zero. Each model call carries its model ID, tokens, rate provenance, and actual cost when known. Whole-test retries from --retries create separate attempts. Locator cache outcomes are recorded explicitly; no cache hit is inferred from an absence of model calls.
New provider calls also carry an optional provider identity such as clef.
It is optional so existing v1 results remain valid; an absent value is unknown,
not inferred from the model name. Readers that previously rejected every
unknown call field must be upgraded before reading new results. Failed Clef
attempts may have unreported usage, so retry cost remains unknown rather than
being reported as free.
execution, when present, records how the run was scheduled:
parallel.requestedis the--parallelvalue, a number or"auto".parallel.lanesis the lane count actually used.shardis{ index, count, globalSelectedTests }or null.providerConcurrencyis the provider request cap.
selectedTestCount counts only that shard’s tests. tests stays in selection order when lanes finish out of order, and each attempt records its zero-based lane. A model call that waited on a shared 429 pause has rateLimited: true and rateLimitWaitMs; queueWaitMs is the time it waited for a provider slot. A 429 response is not billed, so it does not make the call’s cost unknown. The locator cache reason conflict means another lane held the entry’s lock; the step kept its model result.
For directory runs, selectedTestCount and totals.selectedTests include tests selected before execution, including those not reached after an interruption. discoveryProblems names malformed or unreadable candidate files with safe relative paths and fixes; valid selected files may still execute, but a run with discovery problems is incomplete and exits 3. Whole-test retries create distinct attempts. Classification calls made for a retry belong to that attempt’s calls; step calls remain on their steps. Usage and cost totals include both kinds of call from every attempt.
Markdown, HTML and JUnit reporters consume this same result: verdicts and flags, source positions, safe details, exact scores and thresholds, ranked candidates including (no match), page URL/title tied to an accepted observation, evidence status/path, timings, and usage/cost are all present. --replay adds a referenced frame per executed step and a normalized target rectangle where available; the HTML reporter packages captured frames into a self-contained report.html. The markdown reporter writes report.md beside the result and links non-pass frames by the same relative path.
Report output is sensitive. Declared runtime secrets are replaced in text fields; URL credentials, query, and fragment are stripped by default. URL paths and screenshots may still contain private data. --sensitive-origin=URL suppresses metadata and capture for that origin; --no-evidence disables default non-pass/flagged screenshots. Replay is opt-in. Evidence lives in separate, private files under the run directory, not inline JSON: evidence/<attempt-key>/<frame>.jpg, where each attempt directory is created once and never reused, and path is relative to the run directory. Redaction cannot guarantee removal of undeclared secrets or text rendered in screenshot pixels; mark sensitive pages or disable evidence. Retain and delete run directories according to your own data policy.