Offline query, search and drift commands
Part of the NovaFabric CLI reference. Both nova and novafabric run the same binary.
Offline metrics query (experimental, ADR-0129)
nova query (experimental, ADR-0129)
Aggregate metrics across many local Run Capsules with a small, bounded,
declarative query — offline, no server, no network, no raw SQL, and it never
writes to your capsules. (Since ADR-0225 it does maintain a derived index under
$NOVAFABRIC_HOME; the capsules themselves stay untouched, which is the property
that matters — see the persistent-index note below.) The
grammar is a closed allow-list (spec: capsule-query-dsl-v0): exactly four
clauses (select, where, group_by, time window); anything outside the
allow-list is a hard parse error. Metrics are the already-recorded facts —
cost from the capsule's recorded nova.cost (ADR-0066), tokens/latency from
model-calls.jsonl, eval scores from scores.jsonl — never recomputed.
# How many capsules are on disk?
nova query --select 'count()'
# Average cost and run count of one asset in production, per model, last 7 days
nova query \
--select 'avg(cost) AS avg_cost, count() AS runs' \
--where 'asset = summarizer AND deployment_environment = production' \
--group-by model --since 7d --json
# p95 latency by run status
nova query --select 'p95(latency) AS p95_ms' --group-by status
# The same query as a saveable JSON/YAML object (flags override file fields)
nova query --query-file q.yaml --capsule-dir ./capsulesOptions:
--select EXPR[, EXPR…]— required (here or in the query file). Aggregates:count(), andsum/avg/min/max/pXX(percentile, e.g.p95) overcost,total_tokens,prompt_tokens,completion_tokens,latency, orscore[<name>](eval scores are per named metric — barescoreis rejected). OptionalAS alias.--where 'FIELD OP VALUE [AND …]'— filters over the allow-listed dimensionsasset,deployment_environment,variant,log_level,model,model_id,status,tagwith operators=,!=,<,<=,>,>=,IN (...).ANDonly (noOR/nesting in v0).log_levelcomparisons use severity order (debug < info < warn < error) over the per-observation severity recorded on model-call records (ADR-0127); a record withoutlog_levelreads asinfo.--group-by DIM[,DIM…]— group by the same allow-listed dimensions.--since 7d|24h|P30D|<RFC 3339>/--until <RFC 3339>— time window over the capsulecreated_at(since inclusive, until exclusive; default: all history → now).--limit N— max result rows (default 100, hard ceiling 10 000; group cardinality is also capped at 10 000 — a pathologicalgroup_byis refused).--order-by '<alias> [asc|desc]'— sort before truncation (default: first select, descending), so a--limittop-N is stable.--query-file q.json|q.yaml— the equivalent query object; flags override its fields. This is the saveable form later features build on (ADR-0130, planned).--json— the canonical machine-readable result (the default table is a rendering of it):columns,rows,row_count,truncated,time_window, and anindex {engine, built_at, capsule_count}provenance block.--capsule-dir PATH— capsule storage directory (defaults to$NOVAFABRIC_CAPSULE_DIR).--rebuild-index— discard the cached rows and rebuild them from a full scan.--no-cache— ignore the persistent index entirely and scan every capsule. The authoritative answer, and slower.
The persistent index (ADR-0225). Since the cache landed, a capsule is re-parsed only when it has actually changed, which is worth roughly 5× on the scan (measured: 1841 ms → 375 ms at 10,000 capsules). Two things follow:
nova querywrites to$NOVAFABRIC_HOME/query-index.db. It is a command that previously wrote nothing at all, so this is called out rather than left to be discovered. Your capsules are still never written to — they are signed evidence, and the cache deliberately lives outside them.- A cache that is missing, stale or damaged costs you time, never correctness:
any problem falls back to the full scan. If you ever suspect it,
--no-cachegives the authoritative answer and--rebuild-indexreplaces the stored rows.
Semantics: count() counts distinct capsules; other aggregates run over the
matched model-call records, skipping capsules where the metric is absent (SQL
null semantics); a dimension a capsule never recorded matches no filter. Empty
result is rows: [], exit 0 — not an error. Parse errors exit 2; execution
errors exit 1.
Aggregation runs on an in-process index, built fresh on each query. The
default engine is stdlib SQLite, even when DuckDB is installed. Measured
(ADR-0222 OQ-3, bench/query/MEASURED_CEILING.md), DuckDB reaches only
parity — 0.86× at 1,000 capsules, 1.00× at 20,000 — because the capsule
directory scan is 86-89% of total query time and the index build is ~3%. SQLite
needs no extra dependency and ties or wins everywhere measured, so it is the
default. Results are identical either way, and the engine actually used is
reported in index.engine.
DuckDB's fast path needs
pyarrow, which[query]deliberately does not install (~154 MB, for no measurable gain here).[scale]and[serve]include it. With DuckDB selected but pyarrow absent, the index build falls back to a row-by-row insert that is ~20× slower than SQLite, and says so once in the log.
To use DuckDB anyway (it needs pip install 'novafabric[query]'):
NOVAFABRIC_QUERY_ENGINE=duckdb nova query --select 'count()' --group-by modelIf that variable is set but DuckDB is not installed, the query still runs on SQLite and logs a warning naming the extra — a read-only query is never failed over an engine preference.
Derived metrics — ratio() (ADR-0236, experimental)
ratio(<operand>, <operand>) [AS alias] divides one selected aggregate by another. Operands name
other select items in the same query, by explicit alias or by canonical expression text:
nova query --select "count(), sum(cost), ratio(sum(cost), count()) AS cost_per_run"⚠ Undefined is not zero. A zero denominator, or an absent numerator or denominator, yields
no value — not 0. "0 errors out of 0 runs" is not a 0% error rate; it is no information,
and charting it as 0% would invent a data point that never existed. A measured zero over a
non-zero denominator is still reported as 0.
Both operands must be selected in the same query, and a ratio cannot take another ratio as an operand.
ⓘ Not yet expressible: a per-aggregate filter, e.g. ratio(count where status:error, count)
for an error rate. The DSL's where is plan-level; a per-aggregate where is a separate decision
(recorded in ADR-0236's implementation status).
Filter scope over the capsule tree — --scope (ADR-0233, experimental)
When a filter matches something inside a distributed run's capsule tree, there are three
defensible answers, and --scope names them:
--scope |
Returns | Answers |
|---|---|---|
node (default) |
the matching capsules | "which capsules failed?" |
root |
root capsules whose tree contains a match | "which runs were affected?" |
tree |
every capsule in any tree containing a match | "what was happening around the failure?" |
nova query --select 'count()' --where 'status = error' --scope rootScope is part of the query plan, not a display option — so a dashboard view using it is
reproducible from the CLI, and it round-trips through --query-file.
⚠ An incomplete tree is reported as incomplete. The JSON result carries a tree_scope block:
{"scope": "tree", "capsules_selected": 7, "expansion_truncated": false,
"complete": false,
"incomplete_reasons": ["parent2: 2 of 3 children have arrived, so this tree is still filling"]}A tree that renders as whole while children are still in flight is a wrong answer presented
confidently, so children still arriving, orphan placeholders, and capsules falling outside the
query's time window are each named. Expansion is capped at 5,000 capsules; exceeding it sets
expansion_truncated rather than silently dropping capsules.
ⓘ Using --scope for the first time after upgrading rebuilds the query index — the indexer now
extracts the parent/child fields the scopes need.
nova view (experimental, ADR-0130)
Saved views: name a nova query once, persist it as a small human-readable
file, and re-run it anywhere. A view is data, not code — a verbatim
ADR-0129 query object plus optional advisory display preferences, stored one
file per view under .novafabric/views/<view_id>.yaml (JSON equally valid;
override the directory with --views-dir or NOVAFABRIC_VIEWS_DIR). View
files are meant to be committed to the project repo: sharing a view is
committing a file, reviewing a view is reading a diff. No server, no network.
Wire contract: schemas/saved-view.schema.json; spec: saved-views-v0.
# Save a query under a name (validated fail-closed at save time)
nova view save failed-runs-7d \
--select 'count() AS runs' --where 'status = error' --since 7d \
--description 'Failed runs, last week' --tags 'triage'
# Re-run it — exactly `nova query` over the stored query
nova view run failed-runs-7d --json
# Manage views
nova view list
nova view show failed-runs-7d
nova view rm failed-runs-7dSubcommands:
nova view save NAME— persist the given query underNAME. Accepts the same query flags asnova query(--select,--where,--group-by,--since,--until,--limit,--order-by,--query-file); the query is compiled through the ADR-0129 parser before anything is written — an invalid query exits 2 and no file is created. The stored query is the normalized query object (the same onenova query --jsonechoes back).--view-idsets an explicit slug id (otherwise derived fromNAME: lowercase, non-alphanumeric runs collapsed to-);--description,--columns,--sort 'FIELD [asc|desc], …',--format table|json|csv, and--tagsset optional metadata and advisory display preferences;--created-byrecords an author (never inferred);--jsonwrites a.jsonfile instead of YAML. An existingview_idis refused unless--force(which preservescreated_atand setsupdated_at). On success the command prints the file path and the view's content hash (view_hash,sha256:over the canonical definition — name, query, display, tags; timestamps and author excluded) so a report can record exactly which view version produced it.nova view run NAME— load the view (by id or name) and execute its stored query via the ADR-0129 engine; semantically identical tonova querywith the same clauses (invariant I2).--capsule-dir,--json, and--format table|json|csvbehave as onnova query; a command-line format always overrides the saveddisplay.format(display preferences are advisory, invariant I3 —display.columns/display.sortshape the table/CSV rendering only, never the canonical JSON, and never which capsules match). The JSON result is thenova queryresult plus aview {view_id, name, view_hash}provenance block. A stale view whose query no longer parses fails loudly (exit 2), never silently.nova view list— list saved views (view_id, name, description, tags);--jsonaddsview_hashper view. An empty or missing views directory prints(no views), exit 0. Corrupt view files are skipped with a warning — a broken view never blocks any other operation.nova view show NAME— print the stored view without executing it (YAML, withview_hashand path as trailing comments;--jsonfor machine output).nova view rm NAME— delete the view's file.
Exit codes: parse/validation errors (bad query, bad --view-id, bad
--format) exit 2; missing/corrupt view, existing view without --force, and
execution errors exit 1; an empty result set exits 0.
nova trend (experimental, ADR-0131)
Compute an offline trend report — one metric (cost, score:<name>, or
latency) bucketed by day / week (UTC calendar buckets) or asset
(categorical) over the local capsule directory. A read-only snapshot
artifact, never a live monitor: no server, no network, no thresholds, no
notifications (the "act on a threshold" concern is the ADR-0136 budget gate).
Runs on the ADR-0129 extraction/filter path; spec: trend-report-v0
(schemas/trend-report.schema.json).
# Weekly cost over the last 60 days (TrendReport JSON to stdout)
nova trend --metric cost --group-by week --since 60d
# Daily GAIA score, written to a file
nova trend --metric score:gaia --since 14d --json trend.json
# p99 latency per asset, plus a single self-contained static HTML artifact
nova trend --metric latency --stat p99 --group-by asset --html trend.html
# Use a saved view (ADR-0130) as the capsule selector
nova trend --metric cost --view prod-summarizer--metric cost|score:<name>|latency— exactly one metric per report (v0). Cost sums the recordednova.costamounts (USD; a capsule with an unresolvable currency is skipped with a warning, never silently converted);score:<name>averages the named suite fromscores.jsonl; latency takes a point statistic per bucket (--stat p50|p95|p99|mean, defaultp95, latency only).--group-by day|weekemits every calendar bucket in the--sincewindow — a bucket with no capsules appears as an explicit gap (value: null,n: 0), never dropped.--group-by assetis categorical (bucket_start: null, ordered by asset id; capsules without an asset fall under(none)).--since(default30d) /--until(default now) bound the input window.--view NAMEruns a saved view'swhereclause as the capsule selector;viewandfiltersare echoed into the report for provenance.- Output: canonical
TrendReportJSON to stdout by default;--json FILEwrites it to a file;--html FILEalso writes one self-contained static HTML file — inline CSS, a stdlib pre-rendered inline SVG chart (line for time, bars for asset; gaps render as visible breaks), the exact JSON embedded inline, no JavaScript and zero external requests (opens fromfile://). - Unreadable capsules and capsules missing the metric are tallied in
skipped_countwith a warning — the report is never aborted by one bad capsule. An empty directory or window succeeds with an empty series.
Exit codes: usage errors (unknown metric/group-by, --stat on a non-latency
metric, unparsable or pathological window) exit 2; runtime errors (missing
capsule directory, unresolvable saved view) exit 1; an empty result exits 0.
Capsule content search (experimental, ADR-0204)
nova search (experimental, ADR-0204)
Full-text search over the redacted text of local run capsules — "which run
mentioned invoice INV-2291?", "where did the agent call rm -rf?". Local-first:
reads the registry DB (SQLite FTS5) and capsule directory directly, no server
needed. The index is populated automatically when capsules are ingested by
nova serve / nova ingest-capsule; pre-existing capsules need one backfill.
nova search "invoice INV-2291" # all terms must match (AND)
nova search "rm -rf" # operators match literally
nova search "inv*" # trailing * = prefix search
nova search "429" --status failure --json # filter + machine output
nova search "timeout" --stream trace # restrict to one stream
nova search --reindex # backfill/repair the index
nova search --reindex --all # force full re-extractionOptions:
--limit N(default 20, max 200) — max runs returned.--json— emit results as JSON (same item shape as the APImatches).--since/--until/--status— filter on the run metadata index.--stream {model-call-messages|tool-call-arguments|trace|capsule-yaml}— restrict to one source stream.--reindex [--all]— idempotent backfill/repair from the capsule directory; also garbage-collects rows for deleted capsules.--allforces re-extraction of already-indexed capsules.--capsule-dir/--db-path— override the standard locations.
Guarantees (ADR-0204): only post-redaction capsule bytes are indexed, and
the corpus is a strict subset of the secret scanner's targets
(model-calls.jsonl, tool-calls.jsonl, trace.jsonl, capsule.yaml);
query input is quoted before it reaches the FTS5 MATCH grammar, so OR,
NEAR, -, : match literally. NOVA_CONTENT_INDEX=off disables indexing
at ingest.
Exit codes: 0 (success, even with no matches), 1 (FTS5 unavailable in
this Python's SQLite, or runtime error), 2 (usage error / empty query).
Offline drift and tool-schema analysis (experimental, ADR-0147/0148)
Read-only, offline detectors over sealed capsule evidence. Each command loads a
JSON document from a path, prints a terminal report (or a machine-readable one
with --json), and never mutates any capsule. A flagged drift, silent failure,
or breaking schema change is a detector observation to investigate, never a
verdict or gate — the commands exit 0 whether or not anything is flagged,
and 2 only on missing or malformed input. This first slice takes the
samples/runs directly in the document; a collector that reads them from sealed
capsules over a --baseline/--window range is a documented follow-on.
nova a2a-objects map (experimental, ADR-0149)
Map A2A Task / Message / Artifact objects onto capsule entities, recording a
field-by-field mapping_manifest and a roundtrip_digest.
nova a2a-objects map --objects objects.json --out facet.jsonOptions:
--objects(required) — JSON:{"tasks": [...], "messages": [...], "artifacts": [...]}--out— write the facet JSON here (default: print to stdout)
Message parts are bound, never stored. A
partsarray can carry user content, so the mapping keeps onlyparts_digest(ADR-0009). A re-export therefore carriesparts_digestin place ofparts— a bounded reconstruction, not a fabricated one.Nothing is dropped silently. Any field of a source object that this mapping does not carry is listed in that object's
unmapped[], and the command reports the count.
Exit codes: 0 (mapped, or nothing to map), 2 (missing/malformed input).
nova a2a-objects roundtrip (experimental, ADR-0149)
Re-export the mapped objects and assert they reproduce the roundtrip_digest.
nova a2a-objects roundtrip --facet facet.jsonOptions:
--facet(required) — a facet JSON, or a capsule carrying one infacets.a2a_objects
Exit codes: 0 (round-trip reproduces the digest), 1 (divergence — the
diverging object and its fields are named on stderr), 2 (bad input).
The digest is computed over the re-exported objects rather than over the stored facet, so the check exercises map → store → export. A digest taken over storage would be re-derived from the same bytes and match unconditionally.
nova a2a-card capture (experimental, ADR-0149)
Record an A2A 1.0 Agent Card as portable capsule evidence. The full card is stored, not a summary, so any A2A-aware tool can read the agent's identity out of the evidence without knowing anything about NovaFabric's schema.
nova a2a-card capture --card card.json \
--well-known-url https://acme.example/.well-known/agent.json \
--outputs ./outputs --out facet.jsonOptions:
--card(required) — the A2A Agent Card JSON document--well-known-url— where the card was served from--a2a-version— protocol version (default1.0)--outputs— write the standalone portable export into this directory--out— write the facet JSON here (default: print to stdout)
signature_okis not asserted. A2A 1.0 cards are JWS-signed and NovaFabric has no JWS verifier wired, so the facet recordssignature_status: "unverified: no JWS verifier configured"rather than a verdict nobody reached.signedis structural — it says a signature block is present, never that it is valid. Supply a verifier through the Python API (a2a.card.build_facet(..., verifier=...)) to get a realsignature_ok.
Exit codes: 0 (captured), 2 (missing or malformed card).
nova a2a-card verify (experimental, ADR-0149)
Recompute the card fingerprint from a stored facet and report whether the card has been altered since capture. Needs no key material and no network.
nova a2a-card verify --facet facet.jsonOptions:
--facet(required) — a facet JSON, or a capsule carrying one infacets.a2a_card
Exit codes: 0 (fingerprint matches), 1 (the card changed), 2 (bad input).
The recorded fingerprint is never rewritten to the observed one — that would redefine which card was presented and turn a detected alteration into a silently accepted new identity.
nova replay-equivalence regime (experimental, ADR-0144)
Report whether a run executed in a regime where a replay could be expected to match — the context an equivalence verdict needs to mean anything.
nova replay-equivalence regime --capsule ./capsuleOptions:
--capsule(required) — capsule directory holdingmodel-calls.jsonl
Exit codes: 0 (eligible), 1 (not eligible, or unknown), 2 (bad input).
Why this matters. Divergence from a run taken at
temperature=1.2tells you almost nothing; the same divergence from a fully pinned run is a finding.schemas/replay-attestation.schema.jsonalready names the classes —BIT_EXACT,BOUNDED_EQUIVALENT,NON_DETERMINISTIC, the last for "missing pins, unpinned seed/temperature" — and this supplies the input they need.
unknownis noteligible, and exits1. A run that recorded no temperature has not demonstrated a replay was expected to match. That is a statement about the evidence, not an accusation about the run.The facts travel with the verdict.
calls_without_temperature,calls_with_nonzero_temperature,calls_without_seedand the observedtemperaturesare all reported, so a caller who accepts temperature-0 without a pinned seed can reach that conclusion from the record rather than being bound by the stricter default.
nova replay-equivalence check (experimental, ADR-0144)
Emit the behavioral-equivalence verdict for two trajectories: were the tool calls the replay made equivalent to the baseline's, under a declared match mode and tolerance?
nova replay-equivalence check --baseline baseline.json --replay replay.json \
--mode ordered --tolerance 0.0Each trajectory is a JSON array of {"name": "...", "arguments": {...}} objects.
Options:
--baseline/--replay(required) — the two trajectories--mode— correspondence required:set|ordered(default) |edit--tolerance— maximum distance still counted as equivalent (default0.0, exact)--rule— canonicalization rule (repeatable)--commutable— a tool whose order does not matter (repeatable)--idempotent— a tool whose retries collapse (repeatable)
Exit codes: 0 (equivalent), 1 (not equivalent), 2 (bad input).
Why non-equivalence is a non-zero exit. Unlike
nova drift, whose detectors are observations and always exit0, non-equivalence is the condition a canary replay alarms on (ADR-0147 D3).The verdict records
rules_versionandrules_applied, because a verdict whose canonicalization is unknown cannot be re-derived later. Tolerated divergences are still listed — allowing slack does not hide what it allowed.This is not
nova replay --check-equivalence. ADR-0147 writes the surface that way;nova replayperforms a replay, and fusing the verdict into it needs integration with the replay engine's output. That remains future work. This is the verdict surface a scheduler calls.
nova assure-impact report (experimental, ADR-0147)
Aggregate per-run C3 equivalence verdicts for a corpus of pinned baselines
replayed through a substituted model, into one report for from_model → to_model.
nova assure-impact report --corpus corpus.json --worst 5 --out report.jsonThe corpus is {"from_model": "...", "to_model": "...", "runs": [...]}, where each
run carries baseline_id, the C3 equivalent verdict, its distance, and
optionally cost_before/cost_after (integer minor units + ISO-4217) and
tokens_before/tokens_after.
Options:
--corpus(required) — the corpus document--worst— how many regressions to list (default 5)--out— write the report JSON here
Exit codes: 0 (report produced — including when runs regressed), 2 (bad input).
It decides nothing. ADR-0147 states NF-154 must not decide whether to adopt the new model, so there is no recommendation field and a regression is not a non-zero exit — exiting non-zero would be the adoption decision, made by an exit code.
Three things the report refuses to fudge. A run with no cost data is not a run that cost zero, so each delta reports how many runs contributed and how many carried nothing. A corpus mixing currencies is refused rather than summed. And
equivalent + regressed + inconclusive == nis enforced — a baseline that could not be replayed is inconclusive, never silently counted as a pass.
nova assure-alarm check (experimental, ADR-0147)
The standing production regression alarm. Runs the shipped Wilson + Wald-SPRT
primitive over a window of run outcomes and emits no-regression, regression or
inconclusive — firing only on a statistically significant regression.
nova assure-alarm check --window window.jsonThe window document is {"metric": "...", "baseline": [1,0,...], "window": [1,0,...]},
where each outcome is 1 = healthy.
Options:
--window(required) — the window document--drift-flags— outcomes are1 = driftedrather than1 = healthy; both sides are inverted--p0/--p1— SPRT null / alternative pass-rates
Exit codes: 0 (no regression, or inconclusive), 1 (alarm fired), 2 (bad input).
Why an SPRT and not a threshold. Agentic pass@1 swings several points as noise even at temperature 0, so a bare delta gate fires on noise. A single-run dip does not fire here.
inconclusivemeans "not enough evidence yet" and is not a regression — treating undecided as regression is the false alarm the SPRT exists to prevent.⚠ Polarity matters. If your window is drift flags (
1 = drifted) and you omit--drift-flags, the alarm inverts silently: it fires on improvement and stays quiet on a real regression, while every number in the output still looks plausible.This is not the promote gate. ADR-0147 D4 leaves ADR-0080 unchanged, so this never uses the gate's exit-code contract (
3on regression). It exits1because an alarm exists to be noticed.
nova assure-canary record (experimental, ADR-0147)
Record one canary replay of a pinned baseline: which baseline, when, which stack, the C3 equivalence verdict, its drift score, and whether that alarms.
nova assure-canary record --run canary.json --out record.jsonThe run document is {"baseline_id", "ran_at", "stack": {component: version}, "equivalent": bool, "drift_score": float, "baseline_stack": {...}}.
Options:
--run(required) — the run document--out— write the record JSON here
Exit codes: 0 (equivalent), 1 (alarm — not equivalent), 2 (bad input).
⚠ This records a canary run; it does not schedule or perform one. NF-153 also requires re-running each pinned baseline against the current stack on a declared cadence. That orchestration needs live infrastructure and is not built — ADR-0147's standing production loop still does not exist.
The stack fingerprint is the point. It answers "was this canary judged against the same stack as the baseline?", because a verdict compared across two different stacks is not a comparison, and a "regression" that is really a stack change is a false alarm that looks exactly like a true one. The command says so when the stack changed — and says so distinctly when the baseline's stack is unknown, which is not the same as matched.
Equivalence is never scored here; the verdict comes from
nova replay-equivalence check.
nova assure-run record (experimental, ADR-0147)
Record that a scheduled continuous-assurance run executed — which baselines it checked, which detectors ran, how many alarms fired, and when the next run is due.
nova assure-run record --schedule nightly --ran-at 2026-07-12T00:00:00Z \
--cadence 86400 --baseline bl-support-agent-golden-2026Q2 \
--detector output-drift --alarms 1 --out attestation.jsonOptions:
--schedule(required) — identifier of the assurance schedule--ran-at(required) — RFC 3339 timestamp of the run (a UTC offset is required)--cadence(required) — expected seconds between runs; must be positive--baseline— abaseline_idthis run checked (repeatable)--detector— a detector that ran (repeatable)--alarms— how many alarms fired (default0)--out— write the attestation JSON here
next_dueis derived from--ran-atplus--cadenceand cannot be supplied. It is the mechanism by which a missed run is detected, so a caller-chosen value could promise a due date that never arrives.
Exit codes: 0 (recorded), 2 (bad timestamp, non-positive cadence, empty schedule id).
nova assure-run check (experimental, ADR-0147)
Report whether the next assurance run is overdue.
nova assure-run check --attestation attestation.json --now 2026-07-15T00:00:00ZOptions:
--attestation(required) — attestation JSON, or a capsule carrying one--now(required) — the RFC 3339 instant to judge against
Exit codes: 0 (on time), 1 (overdue), 2 (bad input).
How a run that never happened is detected. A missed run writes nothing, so there is no record of it to inspect. The verdict is computed against the previous attestation's
next_due— the artifact that proves the failure is the last success. Exactly-due is not yet overdue.
nova assure-baseline pin (experimental, ADR-0147)
Designate a sealed capsule as a golden baseline — the fixed reference every drift detector measures against. The pin binds the run to its capsule Merkle root, so it names specific bytes rather than a run id that could later point at different content.
A pin is immutable: it is never edited, and a new pin supersedes it. That is what keeps a comparison made last quarter reproducible after the baseline moves on.
nova assure-baseline pin --capsule ./capsule --run run-8f2a \
--id bl-support-agent-golden-2026Q2 --criterion goal \
--pinned-at 2026-07-01T00:00:00Z --out pin.jsonOptions:
--capsule(required) — directory of the sealed golden capsule--run(required) — run id of that capsule--id(required) — identifier for this baseline--criterion(required) — which axis this is a baseline for:goal|trajectory|output-dist|cost--pinned-at(required) — RFC 3339 timestamp--out— write the pin as JSON to this path (default: print to stdout)
Exit codes: 0 (pinned), 2 (missing capsule, unknown criterion, or malformed input).
nova assure-baseline verify (experimental, ADR-0147)
Recompute a pinned capsule's sealed root offline and report whether the pinned bytes are still the bytes on disk. The pin is never rewritten — not even on a mismatch. Updating it to match what was found would "repair" the record by redefining the object, turning detected corruption into a silently accepted new baseline.
nova assure-baseline verify --pin pin.json --capsule ./capsule --run run-8f2aOptions:
--pin(required) — a baseline pin JSON, or a capsule carrying one infacets.baseline--capsule(required) — directory of the capsule to re-hash--run(required) — which pinned run that capsule corresponds to
Exit codes: 0 (roots match), 1 (mismatch — the pinned bytes changed), 2 (bad input, or the run is not in this pin).
Not in this slice:
nova assure-baseline list. Listing pins across a fleet needs a global pin registry, which is a storage decision of its own. The pin lives infacets.baselineon the capsule, as ADR-0147 §4.1 specifies.
nova drift collect (experimental, ADR-0147)
Read sealed capsules into the document a drift detector consumes. Until this
existed, every nova drift subcommand took a hand-written document.
It reads through the same ADR-0129 scanner nova query uses — one definition of
what a capsule is, one meaning for a window (since inclusive, until
exclusive) — with the ADR-0225 row cache in front of it.
nova drift collect --capsules ./capsules --window 7d.. --json
nova drift collect --capsules ./capsules --window 7d.. --emit silent-failure \
--quality-metric pass-rate --threshold 0.8 --json > runs.json
nova drift collect --capsules ./capsules --emit detect --dimension cost \
--baseline 30d..2026-07-05T00:00:00Z --window 2026-07-05T00:00:00Z.. \
--statistic psi --threshold 0.2 --json > drift.json
nova drift collect --capsules ./capsules --emit fingerprint \
--run run-42 --baseline-run golden-1 --threshold 0.2 --json > fp.json
nova drift collect --emit root-cause --run run-42 --baseline-run golden-1 --json > rc.jsonOptions:
--capsules PATH(required) — directory of sealed Run Capsules--window SINCE..UNTIL(required) — the window being examined; either side may be empty.SINCEaccepts a duration (7d) or a timestamp--emit runs|detect|silent-failure|fingerprint|root-cause— defaultruns, the neutral per-run record list--baseline SINCE..UNTIL— the comparison window (--emit detect)--dimension D—cost,prompt-tokens,completion-tokens,total-tokens,latency,model-calls, orscore:<name>. The recordkindis derived from it:score:*is output-drift, the rest is behavioral-drift--statistic psi|ks— the two-sample statistic (--emit detect)--threshold FLOAT— your policy; never defaulted--quality-metric NAME— the score to read (--emit silent-failure; opt-in for--emit fingerprint, where folding scores in unasked would change the basis and the signature)--run ID/--baseline-run ID— the run to fingerprint and the one to compare it against (--emit fingerprint, which is keyed by run id rather than by the window)--commutable NAME/--idempotent NAME— repeatable; passed through to the C3 canonicalizer--kind KIND— provenance kind to compare; repeatable (--emit root-cause)--depth N— lineage walk depth, default5; recorded in the document because a walk too shallow to reach the change looks exactly like no change (--emit root-cause)--lineage-db PATH— lineage store to read; the default store when omitted (--emit root-cause)--baseline-id ID— optional pinned-baseline id to record--no-cache— authoritative full scan, ignoring the ADR-0225 row cache--json— emit the document (the form the detectors read)
It collects; it does not judge — no drifted flag is computed here. A run
that recorded no value is left out of a sample and counted as missing, never
entered as a zero; mixed currencies are refused rather than summed; and an empty
sample refuses to become a drift document, because a statistic over nothing is
no evidence rather than evidence of no drift.
--emit root-cause reads the lineage store rather than the capsule tree, so it
needs no --capsules. A run absent from the lineage graph is refused: an empty
ancestor list on both sides diffs to no_change, which would report "nothing
changed" from missing data.
Exit codes: 0 (document rendered), 2 (bad input, unknown dimension, missing policy flag, or an empty sample).
nova drift detect (experimental, ADR-0147)
Compute an offline two-sample drift record from supplied samples — no model re-invocation, zero token cost.
nova drift detect drift.json
nova drift detect drift.json --jsonOptions:
<document>(positional, required) — JSON:{kind: output|behavioral, ...samples..., threshold}.outputdocuments carry{metric, statistic, baseline[], window[], window_meta, baseline_id?};behavioraldocuments carry{dimension, distance, baseline, window}--json— emit the drift record as JSON
Exit codes: 0 (record rendered, drifted or not), 2 (missing or malformed input).
nova drift silent-failure (experimental, ADR-0147)
Flag runs that reported success but whose quality signal fell below a threshold. A silent failure is surfaced for review — it is a detector observation, not a determination that the run failed.
nova drift silent-failure runs.json
nova drift silent-failure runs.json --jsonOptions:
<document>(positional, required) — JSON:{runs: [{run_id, status, quality_signal}], threshold, success_statuses?}--json— emit the silent-failure report as JSON
Exit codes: 0 (report rendered, whether or not any run is flagged), 2 (bad input).
nova drift root-cause (experimental, ADR-0147)
Link an observed drift to the input(s) that changed between a baseline and a drifted run. Diffs the two runs' lineage provenance ancestors down to the model/prompt/tool/dataset that changed. The result is a correlation, not a cause — a hypothesis to investigate.
nova drift root-cause rc.json
nova drift root-cause rc.json --jsonOptions:
<document>(positional, required) — JSON:{baseline: [{kind, ref}], drifted: [{kind, ref}], kinds?}--json— emit the root-cause hypothesis as JSON
Exit codes: 0 (rendered, whether or not anything changed), 2 (bad input).
nova drift fingerprint (experimental, ADR-0147)
Compute a run's behavioral fingerprint — a deterministic, offline-reproducible signature over its C3-canonicalized trajectory, tool mix and score profile — and, when a baseline is supplied, the distance to it.
Because the trajectory is canonicalized by the same ADR-0144 canonicalizer the equivalence engine uses, a collapsed idempotent retry or a reordered pair of declared-commutable calls does not move the signature, while a different trajectory does. The distance is computed from the basis components, never from the signatures — a digest can only answer same/different.
nova drift fingerprint run.json
nova drift fingerprint run.json --jsonOptions:
<document>(positional, required) — JSON:{run: {run_id, calls: [{name, arguments}], scores?}, baseline?: {…same shape…}, threshold?, commutable?: [], idempotent?: []}. Scores are higher-is-better in[0, 1]; an out-of-range score is refused rather than clamped.thresholdis required wheneverbaselineis present.--json— emit the fingerprint (or the comparison) as JSON
A comparison whose two fingerprints share no basis component reports distance: null
and shifted: null with a stated reason — unknown is not unchanged. Comparing
fingerprints built under different canonicalization rule versions is refused rather
than scored.
Exit codes: 0 (rendered, shifted or not — it is an observation, not a gate), 2 (bad input).
nova capture-ui (experimental, ADR-0148 D3)
Read the GUI actions a computer-use agent took and the screens it saw (NF-166/167).
nova capture-ui show --capsule ./capsules/run-42
nova capture-ui show --capsule ./capsules/run-42 --json
nova capture-ui verify --capsule ./capsules/run-42
nova capture-ui verify --capsule ./capsules/run-42 --strictSubcommands:
show— the recordedui_actionsandui_observationsverify— check that stored observation bytes still hash to theircontent_hash
Options: --capsule PATH (required) · --json · --strict (verify only)
⚠ A digest of typed text is not a redaction — read this before enabling GUI capture.
Keystrokes are the highest-PII-risk content NovaFabric touches, and what people type into
GUIs is exactly the low-entropy material a dictionary defeats: passwords, PINs, coupon
codes, postcodes, names, dates of birth. A plain sha256 of that is a verifiable oracle —
anyone holding the capsule confirms a guess by hashing candidates. Three layers apply:
- Scanned before it is digested. Typed text goes through the ADR-0009 rule pack first.
On a match no digest is written at all — the action records
redacted: truewith aredaction_reasonnaming the rule. A digest of a detected secret is strictly worse than an absence. - Salted per capsule. What survives is digested under a random salt recorded once in the facet. "Was the same string typed twice in this run?" stays answerable — the only thing the field is for — while precomputed tables and cross-capsule correlation stop working.
- The residual is stated, not hidden. A salted digest does not stop someone holding
the capsule from brute-forcing a short input: the salt is right beside it.
showprints that caveat whenever it prints a digest.
Raw typed text is stored only when byte capture is opted in (NOVAFABRIC_CAPTURE_MEDIA=1,
ADR-0125) and the text survives the redaction pass. Opting in is not opting out of the
secret rules.
Observations are ADR-0125 MediaParts, not a second blob store. Screenshots and DOM
snapshots carry content_hash and byte_size always, blob_ref only under the same opt-in.
An observation with no blob_ref is not an integrity failure in verify —
reference-metadata-only is the default, and counting it as unresolved would make every
privacy-preserving capsule look corrupt.
Fail-open is a safety property (ADR-0148 I-3). NovaFabric sits beside an agent driving a
real browser. A capture hook that raises must not stop the agent's click, so nothing here
propagates an exception — a lost action increments dropped, which travels with the facet
so a short list never reads as a complete one.
⚠ NovaFabric never performs a GUI action. It records that one was declared. Driving or replaying a browser is acting, not recording, and ADR-0148 rejects it as scope.
Exit codes: 0 (shown, or verified — including a verify that finds a tampered
observation, because reporting one is the job succeeding), 1 (verify --strict with any
unresolved observation), 2 (bad input).
nova provenance (experimental, ADR-0148 D1)
Bind a C2PA / Content-Credentials manifest to the exact content_hash of a captured
media part, and read back the recorded watermark-presence claims.
nova provenance bind --capsule ./capsules/run-42
nova provenance bind --capsule ./capsules/run-42 --write
nova provenance bind --capsule ./capsules/run-42 \
--manifest ./aa11….c2pa.json --write
nova provenance bind --capsule ./capsules/run-42 --output-hash sha256:aa11… \
--producing-model img-gen-v3 --producing-run-id run_42 \
--art50-marking-claimed --nf094-receipt-digest sha256:ee55… --write
nova provenance show --capsule ./capsules/run-42 --json
nova provenance verify --capsule ./capsules/run-42
nova provenance verify --capsule ./capsules/run-42 --strict
nova provenance output --capsule ./capsules/run-42
nova provenance watermark show --capsule ./capsules/run-42 --bind --writeSubcommands:
bind— discover manifests and bind them to media hashes (NF-161/163). Writes nothing without--writeshow— print themedia_provenancefacet a capsule already carriesverify— per-entryactive_manifest_ok/hard_binding_ok/cert_chain_okoutput— the NF-163 per-artifact output receipts and their NF-094 cross-linkwatermark show— the NF-162 presence claims;--bindreads them from the manifests
Key options:
--capsule PATH(required) — one Run Capsule directory--manifest PATH— an explicit manifest named<sha256-hex>.c2pa.json, repeatable. Overrides sidecar discovery. The filename is the binding: a manifest named for other content will not attach to whatever media the capsule happens to hold--output-hash HASH—content_hashof media the agent produced, repeatable (NF-163)--producing-model/--producing-run-id/--art50-marking-claimed/--nf094-receipt-digest— the NF-163 fields carried on output entries only--capture-manifests— record that manifest bytes were retained (ADR-0125 opt-in)--write— persist the facet intocapsule.yaml--strict(verifyonly) — exit1when any binding is not established--json— emit as JSON, including on the empty path
cert_chain_ok is always unknown, and never inferred from a signature. The signer's
public identity is read out of the manifest, but verifying an X.509 chain needs a trust
store and a chain verifier, and NovaFabric ships neither offline and adds no dependency for
one. The field therefore stays null and carries the machine-readable reason
no_offline_cert_chain_verifier. A fabricated true would put a verification that never
happened into signed evidence.
A binding checked against stored bytes is stronger than one checked against a recorded
field. Byte capture is opt-in (ADR-0125 D2), so most capsules hold a content_hash and no
blob. Both comparisons are worth making, but they are not the same evidence, so every entry
records which one was used in bound_against (blob_bytes or recorded_hash). Only the
first can detect a blob whose bytes no longer match what was recorded.
Absent is not false. Media with no manifest produces no entry — never one with
hard_binding_ok: false. A manifest declaring no hard binding yields null, not false.
And a watermark present has three values: true, false (somebody looked and said no),
and unknown (no claim exists). unknown never renders or serialises as false.
Watermark presence is pattern-only. NovaFabric records a claim carried by C2PA soft-binding, a declared assertion, or a third-party detector's result handed to it as data. It never imports, ships, or depends at runtime on a proprietary detector (SynthID, Luna) — enforced by a test that imports the package in a fresh interpreter and asserts none is loaded.
⚠ Manifests are discovered as JSON sidecars at outputs/<sha256-hex>.c2pa.json, or
passed with --manifest. Extracting an embedded JUMBF manifest from image bytes is
not implemented and would need a C2PA library; a capsule whose media carries only
embedded manifests yields no entries. media_parts_scanned and manifests_found travel
with every answer so that reads as "nothing was found", not "nothing was there".
Fail-open (ADR-0148 I-3): a capsule with no manifests produces no facet and exits 0.
Exit codes: 0 (bound, shown, or verified — including a verify that found a broken
binding, because reporting one is the job succeeding), 1 (verify --strict with any
binding not established), 2 (bad input).
nova toolschema conformance (experimental, ADR-0148)
Seal a capsule's recorded ADR-0128 conformance verdicts so they are tamper-evident, or verify an existing seal.
nova toolschema conformance --capsule ./capsules/run-42 --json
nova toolschema conformance --capsule ./capsules/run-42 --predicate --json
nova toolschema conformance --capsule ./capsules/run-42 --verify sha256:49fd32…Options:
--capsule PATH(required) — one sealed Run Capsule directory--predicate— emit the in-toto predicate fragment (the digest and counts only) forcapsule_statement(extra_predicate=...)--verify DIGEST— recompute from the recorded verdicts and compare--json— emit as JSON
The verdicts and the digest live apart on purpose. The verdicts belong in the capsule
facet; the digest goes into the signed attestation. That is what leaves --verify two
independent sources to compare — a check that recomputed from the object carrying the
digest would re-hash its own input and could never fail.
A call that declared no schema is unchecked, never conforming. Otherwise a capsule
where nothing declared a schema would report perfect conformance having validated nothing.
⚠ An ADR-0128 verdict carries checked_at, so re-running validation yields a different
verdict and therefore a different digest. The seal covers this record, not the act of
validating; a re-derived digest differing from the sealed one is not a seal failure.
Exit codes: 0 (sealed, whatever the counts — it records, it does not gate; or verified),
1 (--verify mismatch — a broken seal is a finding), 2 (bad input, or a capsule with
no tool calls, since a digest over an empty list is a constant).
nova toolschema deprecations (experimental, ADR-0148)
Flag every sealed run still pinned to a retired tool version.
nova toolschema deprecations --capsules ./capsules --tool mcp://acme/search \
--version 1.0.0 --deprecated-at 2026-07-28 --successor mcp://acme/search@2Options:
--capsules PATH(required) — directory of sealed Run Capsules to scan--tool ID(required) — the tool whose version is retired--version V(required) — the retired version--deprecated-at DATE(required) — ISO-8601; an unparseable date is refused rather than stored--successor ID— replacement tool, when one was declared. Absent stays absent — never""--json— emit thetool_deprecationrecord as JSON
Three buckets, not two. tool_version is a required field of the ADR-0128 tool-call
schema, documented as "Semver if known; unknown otherwise". A run recorded as
unknown can be neither confirmed as pinned nor cleared, so it is reported in its own
unknown_version_run_ids list rather than folded into either answer. Deprecating the
literal version unknown is refused, since it would flag every run whose version the
capture path could not determine.
capsules_scanned travels with the result: without it, "no run is affected" and "no
capsule was searched" serialise identically.
⚠ Scope. The dependent set is read from each capsule's tool-calls.jsonl, not from
the lineage graph — lineage/_writer.py emits run, asset and artifact nodes only,
so there are no tool nodes to reference. Emitting them is a documented follow-on.
It reports; it does not gate — exit 0 whether or not any run is pinned, 2 only on bad input.
nova toolschema track (experimental, ADR-0148)
Classify a tool-schema change between two versions: additive, breaking,
deprecation, or unknown.
nova toolschema track --tool mcp://acme/search --from-schema v1.json --to-schema v2.json
nova toolschema track --tool mcp://acme/search --from-schema v1.json \
--to-schema v2.json --jsonOptions:
--tool ID(required) — stable tool identity, e.g.mcp://acme/search--from-schema PATH/--to-schema PATH(required) — the two JSON Schemas--max-diff N— bound on the recorded diff, default100. The class is computed over the full comparison, so truncating the evidence never softens the verdict;diff_truncatedanddiff_totalsay the list is partial--json— emit thetool_schema_changerecord as JSON
It follows the additive-safe rule: a new optional property is additive, while a
removal, a type change, an optional-to-required tightening, or a new required
property is breaking — that last one is an addition that breaks every payload ever
recorded. A schema built from allOf/anyOf/oneOf/not/$ref is not traversed and
is reported as unknown with a reason, because calling an unanalysed schema
additive would read as "safe to ship".
Digests are over canonical content, not file bytes, so reformatting a schema is not
a change. (nova toolschema impact hashes the file itself — it identifies which file
was checked, a different question.)
It classifies; it does not gate — breaking is a fact about two schemas, not a
decision about shipping one. Pair it with nova toolschema impact to see which past
runs the change would actually break.
Exit codes: 0 (record rendered, whatever the class), 2 (missing/malformed schema, or a schema that is not a JSON object).
nova toolschema impact (experimental, ADR-0148)
Report which historical runs break under a new tool schema: validates each recorded tool call's arguments against the proposed JSON Schema using the shipped ADR-0128 validator — it does not reimplement validation.
nova toolschema impact calls.json --new-schema new.json
nova toolschema impact calls.json --new-schema new.json --jsonOptions:
<document>(positional, required) — JSON:{tool_id, tool_calls: [{run_id, arguments}]}--new-schema PATH(required) — the new JSON Schema to test past runs against--json— emit theschema_impactreport as JSON
Exit codes: 0 (report rendered), 2 (missing or malformed input).