nova replay, diff and diagnose commands

Part of the NovaFabric CLI reference. Both nova and novafabric run the same binary.

Replay commands (v0.3)

nova replay <capsule>

Replay a captured run. <capsule> is either the run id nova capture printed, or a path to the capsule directory. A path that exists is always used as given; only a reference that is not a usable path is looked up as a run id, in $NOVAFABRIC_CAPSULE_DIR (default ~/.novafabric/capsules/).

# By run id — what `nova capture` hands you
nova replay 01HXAY7M5JZ8R7K4P9DPBYK2WX --mode forensic

# By path — including the ./.novafabric/runs/ the in-process SDK writes to
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode forensic
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode mocked
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode semantic
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --mode exact
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --dry-run
nova replay .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/ --output-dir /mnt/replays

Options:

Mode contracts:

Mode Re-executes command? Re-executes models? Re-executes tools? Output
forensic No No No Inspection report
mocked Yes From cache From cache Replay result
semantic No No No Similarity score (0–1.0) across model call responses
exact No No No Eligibility check: deterministic env + seeded calls
intervention Yes Substituted + cached From cache Counterfactual capsule marked replay_mode: intervention (experimental, ADR-0086)

Output is written to .novafabric/replays/<replay-ulid>/replay_result.yaml.

✓ Replay written: .novafabric/replays/01HXBM1Y3K2NGH9V0RD9P0ZDC4  (replay_id=01HXBM1Y3K2NGH9V0RD9P0ZDC4  mode=forensic)

nova diff <capsule-a> <capsule-b>

Structurally compare two run capsules. Smart-routed: if both arguments contain @, falls back to the asset registry diff (v0.1 behavior).

Each argument is either a run id or a capsule directory path, resolved the same way as nova replay.

nova diff 01HXAY7M5JZ8R7K4P9DPBYK2WX 01HYBZ8N6KA9S8L5Q0EQCZL3XY
nova diff .novafabric/runs/01HX.../ .novafabric/runs/01HY.../
nova diff cap-a/ cap-b/ --output-format json
nova diff cap-a/ cap-b/ --output-format github-annotation
nova diff cap-a/ cap-b/ --assert-no-regressions
nova diff --group-by variant runs/arm-a/ runs/arm-b/
nova diff --group-by environment runs/prod-01/ runs/staging-01/
nova diff --environment production cap-a/ cap-b/
nova diff cap-a/ cap-b/ --graph-shape
nova diff cap-a/ cap-b/ --assert-same-shape

Environment gating, as run against a temporary NOVAFABRIC_HOME (capsule A captured with --environment production, B with --environment staging):

$ nova replay <A> --environment staging
replay refused: --environment staging required, but .../<A> recorded 'production'
$ echo $?
2
$ nova diff --group-by environment <A> <B>
Environment groups (ADR-0126, recorded deployment_environment):
  production: .../<A>
  staging: .../<B>
Cross-environment diff: production → staging
$ nova diff --environment production <A> <B>      # B recorded 'staging'
--environment production: .../<B> recorded 'staging'; both capsules must have ...
$ echo $?
2

Options:

Exit codes (capsule diff): 0 success; 1 a capsule ref did not resolve, --assert-no-regressions found changes (checked first), or --assert-same-shape found a shape change; 2 usage error, or --assert-same-shape could not build a graph for either capsule (fail closed). Without --graph-shape/--assert-same-shape the output is byte-identical to earlier releases.

Diff sections: environment (Python, OS), model calls (aligned by span_id), tool calls (aligned by tool_name + arg hash), output files (by hash).


Diagnose commands (gap-006, ADR-0084)

nova diagnose <run-id>

Attribute a failed run to its most likely responsible step, and label the failure with an AgentErrorTaxonomy category. Runs over the existing lineage / causal graph plus the captured trace — it reads capsule.yaml, trace.jsonl, tool-calls.jsonl, and model-calls.jsonl, and (when present) the lineage store's parent/child edges. It is read-only and writes nothing.

# Diagnose a failed run in the default capsule directory
nova diagnose run-2026-06-11-abc123

# Diagnose a capsule stored elsewhere, as JSON for tooling
nova diagnose run-xyz --capsule-dir ./capsules --output json

Options:

Algorithm (ADR-0084): decompose the run into ordered steps (src-411 module decomposition); a coarse pass scores each erroring step by explicit error signal, earliest-root-cause bias (src-411 — an earlier error outranks a later downstream symptom), and causal-depth bias (src-413 — a downstream delegated_to/spawned step outranks its coordinator ancestor); a fine pass picks the single responsible step and labels it.

AgentErrorTaxonomy categories: MEMORY, REFLECTION, PLANNING, ACTION, SYSTEM, UNKNOWN. Scores are relative ranking weights, not calibrated probabilities. A run with no error signal yields UNKNOWN and no fabricated culprit (exit 0); an unknown run id exits 1.

Hypothesis verification — --intervene (experimental, ADR-0101). For the top hypothesis only, nova diagnose auto-synthesizes an InterventionSpec (ADR-0086), replays the capsule counterfactually under mocked semantics (zero-token), and appends a verification block — hypothesis, intervention applied, original vs counterfactual outcome, and an evidence-based verdict, never guessed:

nova diagnose run-xyz --capsule-dir ./capsules --intervene --output json

The auto-mappable subset in this slice is model-call hypotheses only: the corrective edit clears the error signal at the implicated model call via a mutate_payload substitution. Tool/span/run hypotheses report an honest cannot auto-intervene for this hypothesis class reason. The intervened output capsule is hard-marked replay_mode: intervention and referenced from the verdict block. Fidelity bound (ADR-0086): the verdict tests control-flow and downstream handling of the recorded run, not fresh model behavior. Without --intervene, nova diagnose remains read-only and its output is unchanged (an unverified hypothesis is never presented as a proven root cause).

Counterfactual root-cause search — --search-root-cause (experimental, ADR-0101 §NF-018). Widens --intervene from testing only the top hypothesis into a search: it sweeps the §NF-019 causal-root candidates — already ranked shallowest/earliest-first, which is exactly the pruning the ADR calls for over a naive linear sweep of every step — running a bounded number of zero-token intervention replays (default --max-interventions 8, hard ceiling 50) until one confirms an outcome flip. The first CONFIRMED candidate is the decisive root cause; every attempt (confirmed, refuted, or honestly unmappable) is recorded, so the search itself is auditable, not just its winner:

nova diagnose run-xyz --search-root-cause --max-interventions 5 --output json

Same auto-mappable subset as --intervene (model-call hypotheses only). When no candidate flips the outcome within the bound, the result says so explicitly (bounded: true) rather than silently looking exhaustive. Composes with --intervene — both flags can be passed together.