Skip to main content
YAML goal tests populate optional tests[].goal with complete, secret-redacted text and verify strings after data resolution, before browser execution. HTML reports display these above the replay and steps. Authored-step tests omit this field. Existing result files without it remain supported. TypeScript ai.goal is represented on its aggregate action step by optional step.goal metadata. Its completion is "planner", with the redacted goal text, actions, requests, completion reason, and optional dataSeed. This explicitly means accepted planner DONE, not independent verification. It carries no Judge scores and must not be displayed as “verified.” A goal-only TypeScript test can therefore pass with planner completion; any subsequent ai("verify ...") or failed code assertion is a separate step and determines the test outcome normally. YAML goal tests publish an action step after each dispatched click or type, then an aggregate step (kind: verify, operation: goal) when the bounded loop ends. Action steps use data-binding names instead of typed values and include target scores, incremental planner calls, and opt-in post-action replay frames subject to normal privacy rules. The aggregate includes action/request counts and the termination reason in detail, independent Judge scores when reached, and only remaining calls with purpose: planner or judge, avoiding duplicate usage. A flagged goal assertion fails the goal even without --strict. The aggregate omits page metadata, screenshots and replay frames. Goal failures are not automatically restarted by --retries. That required, independent, unflagged verification policy applies to YAML goal tests and the exported runGoal API. It is unchanged by TypeScript ai.goal’s planner-completion contract. For TypeScript, entering a goal also disables automatic whole-test retries for every later failure in that attempt, including a successful goal followed by failed verification or cleanup. @sedum-dev/core/run-result.schema.json is the Draft 2020-12 schema generated from the same Zod definition used to validate every published snapshot. The schema ID is https://sedum.dev/schemas/run-result/v1. The v1 schema may still change incompatibly before its first published release. After v1 is published, incompatible changes require a new major schema version. Consumers should reject unknown major versions. A step from a TypeScript test carries group, the enclosing ai.group names from the outermost, when it ran inside a group. An exception thrown by test code is one failed step with kind: verify and operation: code, whose error code is expect_failed for a failed expect() and code_error otherwise. A TypeScript test’s id is <file>#<title> unless it sets an id, and its description is the title. Each attempt problem has a unique ID, consecutive attempt-local ordinal, phase, ordered source stack, outcome, typed error, and an origin of step or module_binding. A step problem links its executed step; a module binding problem has stepId: null because no sentence ran. The first encountered problem is primary. A failed setup or body stays primary even if later teardown reports an operational error. A primary operational error leaves the verdict null. Executed steps retain their own phase and source stack; there are no synthetic module-call steps. For a syntactically valid sedum run, the CLI resolves project configuration, then creates <outputDir>/<run-id>/progress.json before provider setup or test loading and prints its path. The default outputDir is .sedum/runs. Config must be read first because it chooses this location; an invalid config uses the default location for its terminal operational result when that directory is writable. Each update replaces the file atomically with a complete, validated RunResult; pollers can parse it at any point. result.json is written from the terminal snapshot. A setup or execution error has state: "error", a typed error, and verdict: null; already completed tests and steps remain. Zero executed tests are never a pass. A passed step with low_confidence or contradiction stays passed; --strict may change the exit status and JUnit mapping, not the canonical verdict. SIGINT and SIGTERM received while a run is active request cancellation. The CLI waits for the current operation to unwind, then writes state: "interrupted", verdict: null, and a terminal result.json before exiting with the conventional signal exit code. Once the terminal outcome is committed, late signals are ignored while its final files are written so the process exit and persisted state cannot disagree. An interrupted provider request may have incurred unreported usage; its call is retained with unknown cost, so a receipt never presents that attempt as free. The CLI derives ordinary run exits from this result: completed clean pass 0, completed flagged pass 0 or 2 with --strict, failed 1, and a null or untrustworthy verdict 3. Strict mode changes only that shell gate. Terminal and redirected summaries show the same ordered test identities, verdict/state, flags, and totals; only ANSI/live updates and the default visibility of token/cost lines differ. If the output sink itself fails, exit 3 is reported from the latest validated in-memory snapshot. Sedum names the intended path and does not advertise an earlier progress file as an authoritative final result. This is the sole case where terminal artifact persistence cannot be promised through the failed destination. The model has ordered tests, whole-test attempts, steps, and observation attempts. Only the selected terminal attempt contributes to final test/step outcome counts; all attempts contribute to model usage and cost. Unknown actual cost is null, not zero. Each model call carries its model ID, tokens, rate provenance, and actual cost when known. Whole-test retries from --retries create separate attempts. Locator cache outcomes are recorded explicitly; no cache hit is inferred from an absence of model calls. New provider calls also carry an optional provider identity such as clef. It is optional so existing v1 results remain valid; an absent value is unknown, not inferred from the model name. Readers that previously rejected every unknown call field must be upgraded before reading new results. Failed Clef attempts may have unreported usage, so retry cost remains unknown rather than being reported as free. execution, when present, records how the run was scheduled:
  • parallel.requested is the --parallel value, a number or "auto".
  • parallel.lanes is the lane count actually used.
  • shard is { index, count, globalSelectedTests } or null.
  • providerConcurrency is the provider request cap.
For a shard, selectedTestCount counts only that shard’s tests. tests stays in selection order when lanes finish out of order, and each attempt records its zero-based lane. A model call that waited on a shared 429 pause has rateLimited: true and rateLimitWaitMs; queueWaitMs is the time it waited for a provider slot. A 429 response is not billed, so it does not make the call’s cost unknown. The locator cache reason conflict means another lane held the entry’s lock; the step kept its model result. For directory runs, selectedTestCount and totals.selectedTests include tests selected before execution, including those not reached after an interruption. discoveryProblems names malformed or unreadable candidate files with safe relative paths and fixes; valid selected files may still execute, but a run with discovery problems is incomplete and exits 3. Whole-test retries create distinct attempts. Classification calls made for a retry belong to that attempt’s calls; step calls remain on their steps. Usage and cost totals include both kinds of call from every attempt. Markdown, HTML and JUnit reporters consume this same result: verdicts and flags, source positions, safe details, exact scores and thresholds, ranked candidates including (no match), page URL/title tied to an accepted observation, evidence status/path, timings, and usage/cost are all present. --replay adds a referenced frame per executed step and a normalized target rectangle where available; the HTML reporter packages captured frames into a self-contained report.html. The markdown reporter writes report.md beside the result and links non-pass frames by the same relative path. Report output is sensitive. Declared runtime secrets are replaced in text fields; URL credentials, query, and fragment are stripped by default. URL paths and screenshots may still contain private data. --sensitive-origin=URL suppresses metadata and capture for that origin; --no-evidence disables default non-pass/flagged screenshots. Replay is opt-in. Evidence lives in separate, private files under the run directory, not inline JSON: evidence/<attempt-key>/<frame>.jpg, where each attempt directory is created once and never reused, and path is relative to the run directory. Redaction cannot guarantee removal of undeclared secrets or text rendered in screenshot pixels; mark sensitive pages or disable evidence. Retain and delete run directories according to your own data policy.