nova replay, diff and diagnose commands
Part of the NovaFabric CLI reference. Both nova and novafabric run the same binary.
Replay commands (v0.3)
nova replay <capsule>
Replay a captured run. <capsule> is either the run id nova capture printed, or a
path to the capsule directory. A path that exists is always used as given; only a
reference that is not a usable path is looked up as a run id, in
$NOVAFABRIC_CAPSULE_DIR (default ~/.novafabric/capsules/).
# By run id — what `nova capture` hands you
nova replay 01HXAY7M5JZ8R7K4P9DPBYK2WX --mode forensic
# By path — including the ./.novafabric/runs/ the in-process SDK writes to
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode forensic
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode mocked
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode semantic
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode exact
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --dry-run
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --output-dir /mnt/replaysOptions:
--mode {mocked,forensic,semantic,exact,intervention}— replay mode (default:mocked). Tab-completion available vianova --install-completion.--dry-run— report what would execute without running; writes dry-run report and exits 0--allow-readonly— permit re-invocation ofread-onlytools--allow-mutating— permit re-invocation ofidempotent-writeandnon-idempotent-writetools--allow-external-side-effects— permit re-invocation ofexternal-side-effecttools--allow-unknown-mutation— permit re-invocation of tools withunknownmutation class--output-dir, -o PATH— base directory for replay output (default:.novafabric/replays/)--environment ENV— experimental (ADR-0126): only replay a capsule that recordedENVas itsdeployment_environment(exact match, case-sensitive). Otherwise exit 2 before anything runs; a capsule with no recorded environment is refused. Usable as a CI gate, e.g.nova replay --environment staging --dry-run <run-id>. Thereplay_mutatingpolicy input also carries the recorded value asinput.resource.deployment_environment(nullwhen absent).--intervention-file PATH— InterventionSpec YAML for--mode intervention(experimental, ADR-0086): one target selector (event_indexorspan_id) + exactly one substitution (replace_model_response/replace_tool_result/mutate_payload) + optional named check-functions (fatal: trueaborts). The output capsule is diffable against the baseline withnova diff.
Mode contracts:
| Mode | Re-executes command? | Re-executes models? | Re-executes tools? | Output |
|---|---|---|---|---|
forensic |
No | No | No | Inspection report |
mocked |
Yes | From cache | From cache | Replay result |
semantic |
No | No | No | Similarity score (0–1.0) across model call responses |
exact |
No | No | No | Eligibility check: deterministic env + seeded calls |
intervention |
Yes | Substituted + cached | From cache | Counterfactual capsule marked replay_mode: intervention (experimental, ADR-0086) |
Output is written to .novafabric/replays/<replay-ulid>/replay_result.yaml.
✓ Replay written: .novafabric/replays/01HXBM1Y3K2NGH9V0RD9P0ZDC4 (replay_id=01HXBM1Y3K2NGH9V0RD9P0ZDC4 mode=forensic)nova diff <capsule-a> <capsule-b>
Structurally compare two run capsules. Smart-routed: if both arguments contain @,
falls back to the asset registry diff (v0.1 behavior).
Each argument is either a run id or a capsule directory path, resolved the same way
as nova replay.
nova diff 01HXAY7M5JZ8R7K4P9DPBYK2WX 01HYBZ8N6KA9S8L5Q0EQCZL3XY
nova diff .novafabric/runs/01HX.../ .novafabric/runs/01HY.../
nova diff cap-a/ cap-b/ --output-format json
nova diff cap-a/ cap-b/ --output-format github-annotation
nova diff cap-a/ cap-b/ --assert-no-regressions
nova diff --group-by variant runs/arm-a/ runs/arm-b/
nova diff --group-by environment runs/prod-01/ runs/staging-01/
nova diff --environment production cap-a/ cap-b/
nova diff cap-a/ cap-b/ --graph-shape
nova diff cap-a/ cap-b/ --assert-same-shapeEnvironment gating, as run against a temporary NOVAFABRIC_HOME (capsule A captured with
--environment production, B with --environment staging):
$ nova replay <A> --environment staging
replay refused: --environment staging required, but .../<A> recorded 'production'
$ echo $?
2
$ nova diff --group-by environment <A> <B>
Environment groups (ADR-0126, recorded deployment_environment):
production: .../<A>
staging: .../<B>
Cross-environment diff: production → staging
$ nova diff --environment production <A> <B> # B recorded 'staging'
--environment production: .../<B> recorded 'staging'; both capsules must have ...
$ echo $?
2Options:
--output-format {text,json,github-annotation}— output format (default:text). Tab-completion available vianova --install-completion.--assert-no-regressions— exit 1 if any structural changes detected; useful as CI gate--group-by variant— experimental (ADR-0116). Group the two capsules by their recorded A/B-variant attribution — the(experiment_id, variant_id)of the optionalvariantblock — and label the diff as cross-arm (different groups) or within-arm (same group). A capsule without avariantblock groups under(no variant). Read-only over recorded facts: this never assigns variants and never mutates a capsule. Capsule paths only;text/jsonoutput only (jsonwraps the report in{variant_groups, cross_arm, diff}).--group-by environment— experimental (ADR-0126 P2). Group the two capsules by their recordeddeployment_environment(the typed top-level field set bynova capture --environment/NOVAFABRIC_ENVIRONMENT) and label the diff cross-environment or within-environment. A capsule with no value — or one violating the^[A-Za-z0-9._:-]{1,64}$rule — groups under(no environment); nothing is inferred.jsonwraps the report in{environment_groups, cross_environment, diff}. Same restrictions as--group-by variant.--environment ENV— experimental (ADR-0126 P2). Only compare capsules that both recordeddeployment_environment == ENV(verbatim, case-sensitive). Fails closed: exit 2 naming each capsule that recorded another value or none. AnENVoutside the value rule is a usage error. Capsule paths only; not combinable with--media/--significance.jsonoutput addsenvironment_filter. To list capsules by environment usenova query --where 'deployment_environment = production'—nova listlists registry assets, which carry no deployment environment.--graph-shape— experimental (ADR-0124 P3). Rebuild both capsules' agent execution graphs (asnova graph agent) and append an additivegraph_shapeblock:same shape,shape changed(node/edge deltas by structural path, e.g.span:agent.turn[0]/tool_call:git[0], with record ids; capped at 50 per list with true totals), orgraph unavailablewith a reason. Reports each side'sgraph_digest(equal only for byte-identical graphs — it binds capsule id, record ids and timings) and an id/timing- independentshape_digestthat decides "same shape". Injsonthe block is a top-levelgraph_shapekey; ingithub-annotationonenotice/error/warningline. Never fails the diff. Capsule diffs only; not combinable with--media/--significance.--assert-same-shape— experimental. Implies--graph-shape; CI gate on the shape.
Exit codes (capsule diff): 0 success; 1 a capsule ref did not resolve, --assert-no-regressions
found changes (checked first), or --assert-same-shape found a shape change; 2 usage error, or
--assert-same-shape could not build a graph for either capsule (fail closed). Without
--graph-shape/--assert-same-shape the output is byte-identical to earlier releases.
Diff sections: environment (Python, OS), model calls (aligned by span_id), tool calls (aligned by tool_name + arg hash), output files (by hash).
Diagnose commands (gap-006, ADR-0084)
nova diagnose <run-id>
Attribute a failed run to its most likely responsible step, and label the failure
with an AgentErrorTaxonomy category. Runs over the existing lineage / causal graph
plus the captured trace — it reads capsule.yaml, trace.jsonl, tool-calls.jsonl, and
model-calls.jsonl, and (when present) the lineage store's parent/child edges. It is
read-only and writes nothing.
# Diagnose a failed run in the default capsule directory
nova diagnose run-2026-06-11-abc123
# Diagnose a capsule stored elsewhere, as JSON for tooling
nova diagnose run-xyz --capsule-dir ./capsules --output jsonOptions:
--capsule-dir PATH— capsule storage directory (defaults to$NOVAFABRIC_CAPSULE_DIR).--output {text,json}— output format (default:text).--intervene— experimental (ADR-0101): verify the top hypothesis with a counterfactual intervention replay and record an evidence-based verdict (see below).--search-root-cause— experimental (ADR-0101 §NF-018): search for the earliest step whose correction flips the outcome (see below).--max-interventions N— bound on how many causal-root candidates--search-root-causewill test (default:8, hard ceiling:50). Only with--search-root-cause.--replay-dir PATH— base directory for the intervention replay output (default:.novafabric/replays). Only with--interveneor--search-root-cause.
Algorithm (ADR-0084): decompose the run into ordered steps (src-411 module
decomposition); a coarse pass scores each erroring step by explicit error signal,
earliest-root-cause bias (src-411 — an earlier error outranks a later downstream
symptom), and causal-depth bias (src-413 — a downstream delegated_to/spawned
step outranks its coordinator ancestor); a fine pass picks the single responsible
step and labels it.
AgentErrorTaxonomy categories: MEMORY, REFLECTION, PLANNING, ACTION, SYSTEM,
UNKNOWN. Scores are relative ranking weights, not calibrated probabilities. A run
with no error signal yields UNKNOWN and no fabricated culprit (exit 0); an unknown run
id exits 1.
Hypothesis verification — --intervene (experimental, ADR-0101). For the top
hypothesis only, nova diagnose auto-synthesizes an InterventionSpec (ADR-0086),
replays the capsule counterfactually under mocked semantics (zero-token), and appends a
verification block — hypothesis, intervention applied, original vs counterfactual
outcome, and an evidence-based verdict, never guessed:
nova diagnose run-xyz --capsule-dir ./capsules --intervene --output jsonCONFIRMED— the intervention replay re-executed the capsule's command and the original failure flipped to success (exit code 0).REFUTED— the re-execution still failed after the intervention.INCONCLUSIVE— the flip is not measurable; the reason is always recorded (no hypothesis, the original run did not fail, the hypothesis class is not auto-mappable, the capsule has no re-executable command, or the replay aborted).
The auto-mappable subset in this slice is model-call hypotheses only: the corrective
edit clears the error signal at the implicated model call via a mutate_payload
substitution. Tool/span/run hypotheses report an honest
cannot auto-intervene for this hypothesis class reason. The intervened output capsule
is hard-marked replay_mode: intervention and referenced from the verdict block.
Fidelity bound (ADR-0086): the verdict tests control-flow and downstream handling of the
recorded run, not fresh model behavior. Without --intervene, nova diagnose remains
read-only and its output is unchanged (an unverified hypothesis is never presented as a
proven root cause).
Counterfactual root-cause search — --search-root-cause (experimental, ADR-0101
§NF-018). Widens --intervene from testing only the top hypothesis into a search: it
sweeps the §NF-019 causal-root candidates — already ranked shallowest/earliest-first,
which is exactly the pruning the ADR calls for over a naive linear sweep of every step —
running a bounded number of zero-token intervention replays (default --max-interventions 8, hard ceiling 50) until one confirms an outcome flip. The first CONFIRMED candidate
is the decisive root cause; every attempt (confirmed, refuted, or honestly unmappable) is
recorded, so the search itself is auditable, not just its winner:
nova diagnose run-xyz --search-root-cause --max-interventions 5 --output jsonSame auto-mappable subset as --intervene (model-call hypotheses only). When no
candidate flips the outcome within the bound, the result says so explicitly
(bounded: true) rather than silently looking exhaustive. Composes with --intervene —
both flags can be passed together.