Registry, prompt and promotion commands

Part of the NovaFabric CLI reference. Both nova and novafabric run the same binary.

Registry commands (v0.1)

nova register <spec.yaml>

Register an asset from a YAML spec file. Exits 0 on success, 1 on validation or duplicate error.

nova suggest-register [<capsule-ref>] [OPTIONS]

Analyze captured run capsules and suggest assets to register. Inverts the onboarding workflow — capture first, then let NovaFabric propose what to register from observed evidence.

Three modes:

Options:

Flag Description
<capsule-ref> Run ID or path. Omit to scan 10 most recent capsules.
--output-dir / -o Directory for draft YAML files (default .novafabric/drafts).
--draft-only Write drafts without registering.
--auto Auto-register all suggestions above --min-confidence.
--min-confidence Confidence threshold for --auto mode (default 0.8).
--skip-types Comma-separated asset types to skip (e.g. prompt,agent).
--runs-dir Base directory containing capsule runs.

Confidence scoring: models = 1.0 (every observed call), tools = 0.8 (min 2 calls), agents = 0.7.

Post-capture hint: after every successful nova capture, a one-line hint is printed when unregistered assets are detected. Disable with NOVAFABRIC_SUGGEST=0. The first capsule written to a directory also gets a one-line pointer to GitHub Discussions (terminal only; disable with NOVAFABRIC_COMMUNITY_HINT=0).

Examples:

# Scan 10 most recent capsules interactively
nova suggest-register

# Analyze a specific run
nova suggest-register 01HXAY7M5JZ8R7K4P9DPBYK2WX

# Write draft YAML files for review
nova suggest-register --draft-only --output-dir ./drafts/

# Auto-register high-confidence suggestions, skipping agents
nova suggest-register --auto --min-confidence 0.8 --skip-types agent

Dashboard equivalent: Registry tab → Suggest Register panel (live suggestions table with one-click Register).

nova list [--type TYPE] [--status STATUS] [--stale] [--stale-days N]

List registered assets. Optional filters by asset type (model, agent, prompt, etc.) or lifecycle status (development, staging, production, archived).

nova inspect <name[@version]>

Show full metadata for an asset. If version is omitted, shows the latest registered version.

nova promote

v0.13.0: nova promote is now a sub-group with three sub-commands: direct, propose, and approve. Old scripts must add direct as the first positional argument.

nova promote direct <name@version> --to <status> [--force] [--significance-gate]

Promote an asset in a single step (no SoD enforcement). Valid transitions:

Current status Valid targets
development staging, archived
staging production, archived
production archived
archived staging

Agent assets require a passing eval result to promote to staging or production. Use --force to override the eval gate (interactive confirmation required; audit-logged).

--significance-gate (opt-in, ADR-0080) — replace the single-passing-eval check with a statistical regression gate. A Wald SPRT runs over the asset's recent pass/fail eval sequence and only blocks on a statistically significant regression (ACCEPT_H1). Noise (ACCEPT_H0) and inconclusive evidence (CONTINUE, too few runs) do not block, so a single-run dip cannot fire the gate. Default hypotheses are p0=0.9, p1=0.7, alpha=0.05, beta=0.05 (overridable via the promote_asset() API). The flag is opt-in: omitting it preserves the default single-passing-eval behavior. --force bypasses it.

--scores-file <path> [--metric task_pass] (NF-007, experimental) — source the gate's pass/fail sequence from an evidence-grade scores.jsonl (the boolean --metric) instead of the eval_results table. Omitting --scores-file preserves the exact eval_results behavior.

--slsa-provenance [--slsa-out <path>] [--identity <id>] (NF-031, experimental, ADR-0096) — on a successful promotion, also emit a DSSE-signed SLSA v1 provenance (slsa.dev/provenance/v1) recording the promotion decision (and the significance-gate/v1 gate when --significance-gate is set). The subject digest is the sha256 of the registered asset spec; the attestation is signed with the keyring Ed25519 key for --identity (default: OS username) and written to <name>-<version>.slsa.json (or --slsa-out). Verify it with nova verify-envelope. Opt-in: without the flag, promotion behavior and output are unchanged.

--slsa-ml-profile (NF-057, experimental, ADR-0105) — with --slsa-provenance, emit the SLSA-for-ML profile instead of the generic one: the build type becomes https://novafabric.dev/promote-ml/v1, a gate-rule byproduct records the gating rule, and an eval-verdict byproduct carries the sha256 digest of the gating eval verdict — binding the promoted model to the exact eval result that justified it. Use for model assets. Still a valid slsa.dev/provenance/v1 Statement, DSSE-signed and nova verify-envelope-verifiable.

nova promote direct my-agent@v1.1 --to staging                       # default gate
nova promote direct my-agent@v1.1 --to staging --significance-gate   # statistical gate (eval_results)
nova promote direct my-agent@v1.1 --to staging --significance-gate \
    --scores-file cand/scores.jsonl --metric task_pass               # gate on evidence-grade scores
nova promote direct my-model@1.0.0 --to staging --slsa-provenance \
    --slsa-out my-model.slsa.json                                    # emit SLSA v1 provenance
nova promote direct my-model@1.0.0 --to staging --slsa-provenance \
    --slsa-ml-profile --slsa-out my-model.slsa.json                  # SLSA-for-ML (NF-057)

The dashboard Promote dialog (Registry tab → PROMOTE →) maps to nova promote direct.

nova promote propose <name@version> --to <status> [--identity NAME]

Open a maker-checker proposal (maker step). Signs a canonical payload with the proposer's Ed25519 key (auto-generated at ~/.config/novafabric/keyring/<identity>.pem if absent). The proposal ID is printed and stored in the promotion_proposals table.

nova promote approve <name@version> [--identity NAME]

Approve the latest open proposal for name@version (checker step). Enforces Separation of Duties at the cryptographic level:

On success, promotes the asset and records a promote.approve audit event.

Opt-in Rego gate: Load src/novafabric/policies/novafabric/defaults/maker_checker_gate.rego to block nova promote direct to staging/production and require the two-step flow.


NovaSeal linked-envelope chain maker-checker (v0.14.0, ADR-0059)

Cryptographic two-person authorization at the capsule signing level. Two separate DSSE envelopes (Proposal + Approval) are linked via SHA-256 of RFC 8785 JCS-canonicalized bytes. Offline-verifiable without a database.

Distinct from nova promote propose/approve (ADR-0058), which enforces two-actor approval for asset registry lifecycle. Both can be used together.

nova policy sign

Sign and version a promotion policy document. The policy defines which certificate subjects are allowed as proposers and approvers. Stored in the NovaSeal Merkle log with a monotonically increasing version number.

nova policy sign \
  --key admin.pem \
  --cert admin_cert.pem \
  --proposer-subjects "alice,alice-ci" \
  --approver-subjects "bob,carol" \
  --bypass-valid-hours 24
# Policy signed and stored: version 1

Options:

Prints Policy signed and stored: version N on success.


nova seal propose <capsule-id>

Create a promotion proposal (maker step). Builds a promote/proposal/v1 DSSE predicate, validates it against the JSON Schema, signs with ECDSA P-256, and stores at {data_dir}/promote/{capsule_id}/proposal/{uuid}.json.

nova seal propose capsule-abc123 \
  --justification "All eval gates green; p99 latency < 200ms. Ready for staging." \
  --key proposer.pem \
  --cert proposer_cert.pem
# Proposal created: b7820d90-039c-4d1f-a586-5023c82e9dd4
#   capsule: capsule-abc123
#   proposer: alice
#   policy version: 1
#   Run nova seal approve b7820d90-... --capsule-id capsule-abc123 as a different identity to complete.

Options:

Exit codes: 0 (proposal created), 1 (validation error, signing error, or no policy found).

Prerequisites: a policy must exist (nova policy sign must have been run). The command reads the latest policy and embeds its version number in the predicate.


nova seal approve <proposal-uuid>

Counter-sign a proposal (checker step). Fetches the Proposal bundle, displays its contents, prompts for confirmation, then builds and stores an Approval bundle.

nova seal approve b7820d90-039c-4d1f-a586-5023c82e9dd4 \
  --capsule-id capsule-abc123 \
  --key approver.pem \
  --cert approver_cert.pem
# Proposal details:
#   capsule_id:    capsule-abc123
#   justification: All eval gates green; p99 latency < 200ms. ...
#   proposer:      alice
#   timestamp:     2026-05-15T10:30:00Z
# Approve this proposal? [y/N]: y
# Approval recorded: 8b778f50-ea38-49c2-bcdc-3a69a1f8e841

Options:

If the operator types anything other than y at the prompt → exits 0 with "Promotion approval cancelled." No bundle is written.

proposal_digest computation: SHA-256(JCS(proposal_envelope_bytes)) — RFC 8785 JSON Canonicalization Scheme, not json.dumps. This makes the approval cryptographically bound to the exact bytes the proposer signed.

Exit codes: 0 (approval recorded or cancelled), 1 (proposal not found or signing error).


nova seal verify <capsule-id> [--offline]

Run the five-check SoD verifier. Loads the Proposal and Approval bundles, loads the policy version recorded in the Proposal, and runs five checks in order (stops at first failure).

nova seal verify capsule-abc123
# SoD verification passed   (exit 0)

nova seal verify capsule-abc123 --offline
# SoD verification passed   (exit 0; --offline accepted, no Rekor calls in v0.1)

The five checks and their exit codes on failure:

Check Description Exit code
1 Proposer cert subject ∈ policy.proposer_key_ids 3
2 Approver cert subject ∈ policy.approver_key_ids 4
3 proposal_digest in Approval = SHA-256(JCS(Proposal envelope bytes)) 5
4 approver_subject ≠ proposer_subject (no self-approval) 6
5 Approval timestamp > Proposal timestamp (ordering) 7

Additional exit codes:

Options:

Policy-time loading: the verifier loads the policy version recorded in the Proposal's policy_version field, not the latest policy. A policy change after proposal creation does not retroactively invalidate in-flight proposals (ADR-0059 §Proposal-time policy).


nova seal bypass

Declare a supervised bypass — logs an approved deviation from the seal gate without blocking the run.

nova seal bypass <reason> [--valid-hours N] [--run-id RUN_ID]

Options:

The bypass is recorded in the NovaSeal bypass log. The dashboard SealTab shows all active and expired bypasses with approver identity.

ECDSA P-256 key is auto-generated at first use if not already present.


nova seal log verify

Verify the integrity of the NovaSeal Merkle log — checks that the log has not been tampered with since the last append. Supports both SQLite (default) and Postgres backends (Scale-S4).

nova seal log verify [--db URI] [--full] [--verbose]

Options:

Returns exit code 0 if the log is consistent, 1 if tampered or inconsistent, 2 if a --consistency proof fails.

Examples:

# SQLite (default)
nova seal log verify

# Postgres — fast sampled check at 1M+ entries
nova seal log verify --db postgresql://user:pass@db.example.com/nova

# Full audit (slower — re-hashes every entry_json)
nova seal log verify --db postgresql://... --full

Env var: NOVAFABRIC_SEAL_DB_PATH — sets the default --db value; accepts both file paths and postgresql:// DSNs.


nova seal ratchet (experimental, v0.50.0, ADR-0089)

Forward-secure per-node signing key ratchet. Opt-in — the static-key NovaSeal signing path remains the default. State lives under $NOVAFABRIC_HOME/seal/ratchet (override: NOVAFABRIC_RATCHET_DIR).

nova seal ratchet init --node-id node-a      # provision epoch-0 chain key
nova seal ratchet rotate --node-id node-a    # advance epoch; erase old chain key (best-effort)
nova seal ratchet status --node-id node-a    # current epoch + registry history

Properties:


nova eval <agent@version>

Run declared evaluation suites for an agent and store results in the registry. Suites are resolved via the novafabric.evals entry-point group.

nova eval agent

Evaluate an agent version using all registered eval suites for that asset type. Equivalent to nova eval run --all-suites.

nova eval agent <name@version> [--db-path PATH] [--timeout SECONDS]

Options:

Results are stored in eval_results table and checked against the Rego gate for promotion eligibility.

nova eval run

Run a specific eval suite against a named agent version.

nova eval run <name@version> --suite SUITE_NAME [--db-path PATH]

Options:

nova eval cost <document>

Render a self-reported eval-cost / compute disclosure (ADR-0154 D2, NF-229, experimental).

nova eval cost cost.json
nova eval cost cost.json --json

Options:

Scope, honestly: every figure is self-reported by the harness. NovaFabric discloses what it was given — it does not measure, verify, or certify these values, and it does not run the eval. Both output modes carry that honesty line, --json included, as an honesty_line field.

This slice reads the figures from the document you pass. Reading facets.eval_cost from a sealed capsule (--capsule <run_id>) is planned, not implemented.

nova eval compare

Compare eval results between two agent versions and generate a regression report.

nova eval compare <name@v1> <name@v2> [--suite SUITE_NAME] [--output FORMAT]

Options:

A regression is flagged when any metric drops by more than the threshold defined in regression_gate.rego.

nova eval list

List all registered eval suite adapters discovered via the novafabric.eval_suites entry-point group.

nova eval list

Output columns: Suite ID, Version, OCI Digest (host-env for suites that run locally), Entry Point (importable module path).

Built-in suites:

Suite ID Version Notes
novafabric-smoke-v1 0.1.0 Fast structural check; no OCI container
gaia-v1 0.1.0 GAIA Lvl-1/2/3 benchmark (OCI-pinned)
mmlu-v1 0.1.0 MMLU 57-subject knowledge benchmark (OCI-pinned)
truthful-qa-v1 0.1.0 TruthfulQA adversarial honesty benchmark (OCI-pinned)
swe-bench-verified-v1 0.1.0 SWE-bench Verified coding benchmark (OCI-pinned)
agentbench-v1 0.1.0 AgentBench multi-task agent benchmark (OCI-pinned)

Third-party suites register via [project.entry-points."novafabric.eval_suites"] in their pyproject.toml. Any load errors for registered adapters are shown inline without crashing the command.

nova eval card / nova eval score (experimental)

Experimental — NF-002 / NF-010, ADR-0099. Evidence-grade evaluation: a score is not a number — it is a signed record binding (value, evaluator-identity, subject-span-digest, verdict). Library + CLI are shipped behind the experimental label; the API may change and there is no dashboard UI yet.

An eval card is the reproducibility key for a score: it pins the exact evaluator (judge model identity + endpoint reference, prompt version, rubric, dataset version, human-agreement calibration), is content-addressed (eval_card_digest = sha256 over the card's canonical JSON excluding its signature), and is Ed25519-signed with the local keyring. A score is one line of an additive, optional scores.jsonl file; a capsule without that file remains valid.

# create → sign → register an evaluator
nova eval card new --source code --card-id exact-match --name "Exact Match" --out card.json
nova eval card sign card.json                       # local Ed25519 keyring; prints digest
nova eval card register card.json                   # into the eval-card registry (must be signed)
nova eval card show   exact-match@0.1.0             # card JSON + digest
nova eval card verify exact-match@0.1.0             # signature_ok / calibration → exit code

# record and list evidence-grade scores
nova eval score add  --card exact-match@0.1.0 --subject sha256:<hex> \
                     --value true --value-type boolean --source code \
                     --name exact_match --capsule ./my-capsule      # or --scores-file scores.jsonl
nova eval score list --capsule ./my-capsule [--source judge] [--json]

nova eval card verify exits non-zero on a broken signature, a missing local key (key_id mismatch → exit 2), or missing calibration on a judge card. nova eval score add refuses a --card ref that does not resolve to a registered eval card. Judge models are referenced by identity + a configurable endpoint (env:NOVA_JUDGE_ENDPOINT); no external URL is hardcoded. Writing a score into a --capsule directory seals it: the score log (scores.jsonl) is covered by the capsule Merkle root, so any Evidence Bundle built from that capsule (nova export-evidence) detects score tampering — no separate re-seal step is required.

nova eval score config (experimental)

Experimental — ADR-0117, spec score-config-v0.md. The score-configuration catalog: named, reusable, immutable, content-addressed definitions of what a valid score for a given metric name looks like, so scores under one name are coherent and comparable across capsules. Fully additive: the Score record and scores.jsonl are unchanged, and a metric without a config is a free score — exactly the previous behavior.

A ScoreConfig declares a metric's value_type (boolean | categorical | numeric — the same vocabulary as Score.value_type), plus the allowed categories (optionally ordinal-ranked, e.g. bad=0 < ok=1 < good=2) or the inclusive numeric range with an optional direction (higher-better | lower-better). Configs live in the local registry SQLite (no server, no internet); each is pinned by a sha256: content_digest over its canonical definition body (wire contract: schemas/score-config-v0.schema.json).

# declare metric shapes (immutable; a changed body bumps the version, never edits)
nova eval score config add --name helpfulness --value-type categorical \
    --description "How helpful the turn was." \
    --category bad:0 --category ok:1 --category good:2
nova eval score config add --name toxicity --value-type numeric \
    --description "Lower is better." --min 0 --max 1 --direction lower-better
nova eval score config add --name grounded --value-type boolean \
    --description "True iff every claim is supported."

nova eval score config list [--all] [--json]    # latest per name; --all = every version
nova eval score config get  toxicity@1          # canonical JSON (also: bare name, sha256:<hex>)
nova eval score config show toxicity            # human view + version history

# opt-in enforcement on the append path (default OFF — free scores stay legal)
nova eval score add --card tox-scan@0.1.0 --subject sha256:<hex> \
    --value 0.3 --value-type numeric --source code --name toxicity \
    --scores-file scores.jsonl --validate-scores

Re-registering an identical body is a no-op (same digest, same version). nova eval score add --validate-scores resolves the latest config for --name and refuses the append (exit 1, nothing written) when the score's value_type disagrees, a categorical value is not in the allowed set, or a numeric value falls outside the inclusive [min, max]. Without the flag — or when no config governs the name — the score is appended unchanged. Because a config is immutable and content-addressed, an aggregate ("avg helpfulness over 500 capsules") can pin the exact content_digest it was computed against, making cross-capsule comparability reproducible evidence. Today that pin is wired into dataset experiments: nova experiment run --score-config records the digest and nova experiment compare reports comparability keyed by it (experimental, ADR-0117 P4 — see the nova experiment section).

nova eval contamination-check (experimental)

Experimental — NF-028, ADR-0108. Flags a capsule that was run against a contaminated or superseded benchmark version (contamination silently inflates eval scores). Detection/flagging only — no remediation.

# check a capsule's dataset_provenance facets (recorded status only)
nova eval contamination-check ./my-capsule

# resolve against a configurable known-bad hash registry (no default URL)
nova eval contamination-check ./my-capsule --registry known-bad.json --json

Reads the additive dataset_provenance facets stored under the capsule's extensions/dev.novafabric.dataset-provenance/ namespace (schema schemas/dataset-provenance-v1.schema.json) — each carrying name, version, dataset_hash, split_hash, and status ∈ current|superseded|contaminated|unknown. The --registry JSON ({"contaminated": [...], "superseded": [...]} of sha256: hashes) upgrades a facet's status when a dataset/split hash matches; the registry never downgrades a facet's recorded severity. Exit codes: 4 when any dataset is contaminated or superseded (CI-gateable), 0 when all are current/unknown, 2 on a usage error.

nova eval import-inspect / export-inspect (experimental)

Experimental — NF-024, ADR-0108. Score-level bridge between Inspect AI (UK AISI) JSON eval logs and NovaFabric's evidence-grade scores.jsonl. Pure stdlib JSON parsing against the documented Inspect log structure — no inspect-ai dependency. The Solver-steps → span-tree import and byte-equal native round-trip from the NF-024 spec remain planned.

# import an Inspect JSON eval log's scorer results into a capsule's score log
nova eval import-inspect ./logs/hello.json --capsule ./my-capsule

# export a capsule's scores.jsonl as an Inspect-compatible JSON log
nova eval export-inspect ./my-capsule --output inspect-log.json   # or stdout

Import maps each sample scorer result (and each aggregate results metric) to a typed Score: Inspect "C"/"I" verdict strings stay categorical (no lossy coercion), numbers become numeric, booleans boolean; model_graded_* scorers map to source: judge, everything else to source: code. Every imported score is provenance-stamped — evaluator_id: inspect-ai:<scorer> plus a synthetic content-addressed eval_card_digest derived from the foreign scorer identity + mapping version (it is not a signed NovaFabric eval card). The mapping is versioned and only pinned Inspect log versions are accepted; an unsupported version errors naming it. Nothing is dropped silently: Inspect fields with no Score target are preserved in extensions/org.inspect/import.json (unmapped), and content-bearing fields (prompts, outputs, transcripts) are enumerated by name in omitted but never copied (ADR-0021 §4). The log's dataset name is also recorded as an NF-028 dataset_provenance facet (status unknown — Inspect logs carry no content hashes).

Export produces an Inspect-shaped JSON log from the capsule's scores.jsonl: one sample per span subject, per-scorer aggregates under results.scores (booleans → accuracy, numerics → mean), and NovaFabric identities carried in score metadata under dev.novafabric.* keys. A capsule imported from Inspect restores the preserved task/model/run-id header; a capsule without scores exports a valid empty log. Exit codes: 0 on success, 2 on a usage error (missing/invalid log, unsupported version, not a capsule directory).

nova eval offline (experimental)

Experimental — NF-009, ADR-0099. Trace-first structural checks over an already-stored capsule that run with zero model calls and emit a code score bound to the capsule Merkle root.

# did the run exercise every declared tool? (reads tool-calls.jsonl)
nova eval offline --capsule ./my-capsule --check coverage --declared-tools search,fetch,write [--emit-score]

# do recorded outputs satisfy a JSON-schema contract?
nova eval offline --capsule ./my-capsule --check contract --schema out.schema.json [--field output] [--emit-score]

# does a recorded input transform preserve a recorded invariant? (declarative check-spec)
nova eval offline --capsule ./my-capsule --check metamorphic --spec check-spec.yaml [--emit-score]

--emit-score appends the resulting score to <capsule>/scores.jsonl, so it is sealed by the capsule Merkle root (see nova eval score). Because the capsule is already on disk, these checks spend zero tokens — they are pure arithmetic/validation over recorded events.

The metamorphic check is driven by a declarative check-spec (schemas/features/metamorphic-check-v0.schema.json): records whose input collapses to the same value under transform form metamorphic pairs, and every pair's output must satisfy invariant. Example check-spec.yaml:

records_file: tool-calls.jsonl   # where the (input, output) records live (default)
input_field: input               # record field forming the pairing key (default)
output_field: output             # record field the invariant is asserted over (default)
transform: [lower, strip]        # identity | lower | strip | collapse_whitespace | remove_punctuation
invariant: equal                 # equal | equal_normalized | numeric_close | length_within
tolerance: 0                     # slack for numeric_close / length_within

A passing check means equivalent inputs produced consistent outputs (a zero-token consistency/robustness signal); it emits a boolean code score. A malformed spec or unknown transform/invariant exits 2.

nova diff <name@v1> <name@v2>

Show field-level differences between two registered versions of an asset. Both arguments must contain @ to trigger asset diff; otherwise capsule diff is used (see above).

nova diff --media (experimental, ADR-0148)

Experimental — NF-170. Compares two runs' media parts by exact sha256, and optionally by a perceptual hash, classifying each pair identical | near-duplicate | changed | added | removed.

nova diff --media runs/run-01/ runs/run-02/
nova diff --media --perceptual runs/run-01/ runs/run-02/ --json

Options:

⚠ changed needs an identity that content cannot supply. Content hashes alone give only identical / added / removed — a changed part hashes differently, so it looks like a removal plus an addition. Parts are therefore paired positionally, which is an assumption about the two runs rather than a fact about them, and the basis is reported as pairing rather than left implicit.

⚠ Perceptual comparison needs an image decoder, which is not a declared dependency. Where one is unavailable, --perceptual exits 2 rather than returning exact-only results — that would report "no near-duplicates" about a check that never ran. A pair that could not be compared stays changed and carries the reason in perceptual_unavailable; it is never promoted to near-duplicate on missing evidence.

⚠ An image with too little low-frequency structure is not compared perceptually. A flat fill or a purely high-frequency pattern hashes like every other one, so any answer would be a coincidence — such pairs are reported unavailable, not near-duplicate.

Exit codes: 0 (rendered, whatever the classifications — it reports, it does not gate), 2 (missing capsule, or --perceptual with no decoder available).

nova diff --significance (experimental)

Experimental — NF-007, ADR-0099; extends ADR-0080. Compares two run sets by statistical significance, not raw delta, so a single-run dip cannot fire a regression gate.

Reads a boolean pass/fail metric from stored scores.jsonl files (or capsule directories) and computes a Wilson interval per side plus a Wald SPRT over the candidate sequence, yielding a three-valued verdict (accept_h0 / accept_h1 / continue). It is offline and zero-token — pure arithmetic over already-recorded outcomes; the workload is never re-run.

nova diff --significance \
    --baseline base/scores.jsonl --candidate cand/scores.jsonl \
    --metric task_pass [--p0 0.9 --p1 0.7 --alpha 0.05 --beta 0.05] [--json]

Exit codes: 0 = accept_h0/continue (no block), 3 = accept_h1 (significant regression — gate CI on this), 2 = usage error (unknown metric, non-boolean metric, or invalid SPRT parameters). --baseline/--candidate accept either a scores.jsonl path or a capsule directory (reads <dir>/scores.jsonl). Known limitation: the SPRT assumes i.i.d. Bernoulli outcomes; correlated early-token cascades violate this (see ADR-0080). Related shipped surfaces: nova eval offline (zero-token structural checks) and nova promote direct --significance-gate (which can also read its sequence from a scores.jsonl).

nova experiment run | list | show | compare (experimental, ADR-0120)

Experimental — ADR-0120; spec dataset-experiment-v0. Run a target command across every item of a pinned local JSONL dataset (one Run Capsule per item — the capsule invariant is untouched) and record an immutable, content-addressed Experiment (schemas/experiment.schema.json). Compare two experiments per item and in aggregate; the verdict is produced verbatim by the shipped ADR-0080 significance gate, so CI fails only on a statistically significant regression. Fully local, offline, zero-token (scores come from the built-in exact-match code scorer — no model call, ever).

The dataset is a JSONL file, one item per line:

{"item_id": "q-001", "input": "What is 2+2?", "expected": "4"}

item_id is the A/B alignment key; expected (optional) drives the built-in boolean exact-match score over the item capsule's recorded stdout. Loading pins dataset_hash (sha256 of the file bytes) and split_hash (sha256 of the ordered item_id sequence) into the record and into each item capsule's ADR-0108 dataset-provenance facet, so per-item contamination checks keep working.

# run: one capsule per item; {input}/{item_id}/{expected} placeholders are substituted
nova experiment run --dataset items.jsonl --target my-agent@1.2.0 -- \
    python agent.py --question "{input}"

# read side
nova experiment list [--json]
nova experiment show <experiment_id> [--json]

# A/B diff: exit 3 on a statistically significant regression (ADR-0080, verbatim)
nova experiment compare <baseline_id> <candidate_id> --metric exact_match \
    [--p0 0.9 --p1 0.7 --alpha 0.05 --beta 0.05] [--json] [-o comparison.json]

# CI gate in one shot: run the candidate and compare against a stored baseline
nova experiment run --dataset items.jsonl --target my-agent@1.3.0 \
    --baseline <baseline_id> -- python agent.py "{input}"

# ADR-0117 D4 aggregation pin: record the score-config digest the aggregate was computed under
nova experiment run --dataset items.jsonl --target my-agent@1.3.0 \
    --score-config exact_match@1 -- python agent.py "{input}"
nova experiment compare <baseline_id> <candidate_id> --require-comparable

Score-config pin (experimental, ADR-0117 P4). run --score-config NAME[@VERSION]|sha256:<hex> resolves the ref read-only against the local score-config catalog (nova eval score config) before any item runs, and records the resolved score_config_ref (name@version) plus score_config_digest on the experiment. The config must govern the --metric name and be boolean (the built-in exact-match scorer); an unresolvable, mismatched, or tampered ref exits 2 with nothing captured — a requested pin is never silently dropped. Every comparison carries a score_config block keyed by digest: same digest ⇒ comparable: true; different digests ⇒ comparable: false (both digests + a message — the aggregates are not directly comparable); either side unpinned ⇒ comparable: null ("not pinned" — never assumed). The block is reported, not gated: exit codes stay the ADR-0080 contract unless --require-comparable (on compare, or on run --baseline) is set, which exits 2 for anything but comparable: true.

Records live under ./.novafabric/experiments/ (override: --experiments-dir / NOVAFABRIC_EXPERIMENTS_DIR); item capsules under ./.novafabric/runs/ (override: --runs-dir). A finalized experiment is never mutated — a re-run mints a new experiment_id, and the store refuses overwrites. Comparing experiments over different pinned datasets (any of name/version/dataset_hash/split_hash) is a hard error, never a silent skew; items present on one side only, errored items, and items without the metric are reported unmatched and excluded from the SPRT sequences. Exit codes: 0 = no significant regression, 3 = significant regression (accept_h1), 2 = usage error. The comparison record (schemas/experiment-comparison.schema.json) also renders a regression_report-shaped gate input for the existing Rego regression gate (ADR-0003/0019) — no new gate engine.


Asset lifecycle commands (v0.11.1)

nova asset diff <name@v1> <name@v2>

Produce a unified diff of the spec JSON between two registered versions of the same (or different) asset. Exits with code 1 if differences are found — useful as a CI gate.

nova asset diff fraud-model@1.0.0 fraud-model@2.0.0
nova asset diff my-agent@v1 my-agent@v2 --unified 5
nova asset diff fraud-model@1.0.0 fraud-model@2.0.0 --output-format json

Options:

JSON output shape:

{
  "ref_a": "fraud-model@1.0.0",
  "ref_b": "fraud-model@2.0.0",
  "identical": false,
  "added":   { "dotted.field": "new-value" },
  "removed": { "dotted.field": "old-value" },
  "changed": { "dotted.field": { "from": "old", "to": "new" } }
}

Both arguments must use the name@version format. If either version is not found, the command exits 1 with a clear error message.

Declared asset dependencies (C-1.3). An asset spec may declare its dependencies in the dependencies: field:

novafabric_spec_version: "1"
asset_type: model
name: my-model
version: 1.0.0
dependencies:
  - base-prompt@2.0.0
  - eval-dataset@3.1.0

When nova register processes the spec, a depends_on lineage edge (confidence: declared) is written for each dependency. These edges are immediately visible to nova lineage blast-radius <dep-ref> and nova lineage provenance <asset-ref>.

Status at consumption (C-1.1). When your SDK code calls record_asset_consumption(asset_ref, status, capsule_dir) before using an asset, the lifecycle status of the asset at that moment is stored in assets.jsonl and propagated into the facets.status_at_consumption field of the resulting consumed lineage edge. This makes blast-radius queries answer "was this asset already in production when run X consumed it?"

Prompt versioning commands (experimental, ADR-0112)

Prompts as first-class, immutable, content-addressed registry versions. Every edit registers a new version (never a mutation); each version's content_hash is a sha256 over the canonical {template, variables, config} body, so a capsule reference prompt:<id>@<version>+sha256:<hex> resolves to exactly one verifiable version. Stored in the existing local SQLite registry — no new store, no schema change, fully offline. NovaFabric records prompt evidence; it never renders templates or serves prompts to inference.

nova prompt register <prompt_id> [--template TEXT | --file PATH] [--var NAME]... [--config JSON] [-m MSG]

Register a new immutable prompt version. The version number auto-increments per prompt_id; re-registering identical content is idempotent (the existing version is returned, no new row). Prints the replay-verifiable frozen reference.

nova prompt register triage -t "You triage {ticket_body}." --var ticket_body -m "first cut"
nova prompt register triage -f messages.json --config '{"temperature": 0.2}'

Options:

nova prompt get <prompt_id>[@<version>]

Fetch one version (latest when @<version> is omitted), frozen to version + content hash. The printed ref: line is the exact prompt:<id>@<version>+sha256:<hex> form a Run Capsule records. --json prints the raw record.

nova prompt list [<prompt_id>] [--status STATUS]

Without an argument: one summary row per managed prompt (latest version, status, version count). With a prompt_id: all versions of that prompt. --status filters by lifecycle status.

nova prompt history <prompt_id>

Chronological version log with created-at, status, content hash, and commit message. --json prints the full version records.

nova prompt diff <prompt_id>@<a> <prompt_id>@<b>

Structural diff of the canonical content triple (template, variables, config) between two pinned versions. Exits 1 when the versions differ (CI-gateable), 0 when identical. Both refs must pin a version.

nova prompt diff triage@6 triage@7
nova prompt diff triage@6 triage@7 --output-format json

Options:

Promotion. Prompt versions move through the standard lifecycle with the existing eval-gated machinery — nova promote direct triage@2 --to staging works unchanged (ADR-0112 D3). A dedicated nova prompt promote <prompt_id>@<version> --to <status> [--force] alias is also available (works today, ADR-0112 P4): a thin, deliberate wrapper around nova promote direct so prompts go through the exact same eval/policy gates and audit trail as any other asset — it adds only a type check that refuses to "promote" a non-prompt asset through this surface.

nova prompt promote triage@1.2.0 --to staging
nova prompt promote triage@1.2.0 --to production

Prompt composition (experimental, ADR-0115)

A prompt body MAY reference other registered prompt assets inline:

{{@prompt:<asset-name>@<selector>}}

where <selector> is an explicit integer version (@3) or a deployment label (@production, @latest — ADR-0113). The reference splices in the referenced asset's resolved body (textual inclusion only; referenced children must be text-form templates). The composition graph is a bounded acyclic DAG: cycles, trees deeper than 8 levels, and unknown references are rejected at register time with named errors — a malformed composition never enters the registry. Each direct reference is snapshotted (the version + content hash it resolved to at register time) into the version's frozen composition block.

nova prompt compose <prompt_id>[@<version>|@<label>]

Resolve the full composition DAG for a prompt (latest version when bare) and print the flattened, content-addressed resolved_composition_manifest: every transitively-included prompt version + hash, the resolved DAG edges (what each label reference pointed at, at this instant), the deepest resolution level, and the sha256 of the final assembled prompt. Read-only — nothing is captured or written. Rebuilding from the manifest's pins reproduces the assembled prompt byte-identically even after children are edited or labels move.

nova prompt compose triage-agent
nova prompt compose triage-agent@production --json
nova prompt compose triage-agent@9 --assembled

Options:

nova prompt tree <prompt_id>[@<version>|@<label>]

Print the composition DAG as an indented tree — each node shows the reference as written, the version it resolved to, its content hash, and a [label] flag when pinned through a deployment label. Read-only.

nova prompt tree triage-agent
triage-agent@9  aaaa0000…
├── system-preamble  @production → v4  11aa22bb…  [label]
│   └── org-header   @2 → v2  55ee66ff…
└── safety-footer    @7 → v7  33cc44dd…

Capsule wiring (recording the manifest during nova capture) and replay verification are planned (ADR-0115 P4) — today compose/tree and the register-time gate are shipped.


Deployment label commands (experimental, ADR-0113)

A deployment label is a mutable named pointer (production, staging, or any custom lowercase name) from an asset name to exactly one immutable registry version. Labels are scoped per asset: production on prompt:triage is unrelated to production on prompt:router. Moving a label appends an audit row to the asset_label_history table (append-only — UPDATE/DELETE are blocked by SQLite triggers); the current pointer is always the newest row. The reserved label latest is auto-maintained (always the highest registered version) and is never user-settable. Stored in the existing local SQLite registry — additive table, fully offline, no new dependency.

Resolution freeze. A <asset_type>:<asset_name>@<label> reference (e.g. prompt:triage@production) resolves at capture time to the concrete version + content hash, recorded as a resolved-asset-ref record (schemas/resolved-asset-ref.schema.json), so a capsule stays deterministically replayable even after the label later moves. The library API is novafabric.registry.labels.resolve_asset_ref(); wiring into nova capture is planned (ADR-0113 P2).

Assets are named [<asset_type>:]<asset_name> — the type prefix is optional but recommended (e.g. prompt:triage).

nova label set <asset> <label> <version> [--reason TEXT] [--json]

Point a label at an existing immutable version, appending an audit row. Fail-closed: the target version must exist (no row is written otherwise); setting latest errors; mixed-case label names are rejected, never lowercased. Re-pointing a label at its current target is a no-op.

nova label set prompt:triage production 4
nova label set prompt:triage production 3 --reason "rollback: v4 regressed on eval GAIA"

Options:

nova label get <asset> <label> [--json]

Resolve a label to its current target version + content hash.

nova label get prompt:triage production      # → production → 4  (1a77…)
nova label get prompt:triage latest --json

nova label list <asset> [--json]

All labels on an asset and their current targets — the auto-maintained latest first, then explicit labels alphabetically.

nova label list prompt:triage

nova label history <asset> [<label>] [--json]

The append-only label-move audit log, newest first: when each move happened, previous → target, who moved it, and why. Answers "when did production point at the poisoned version, and who moved it there?".

nova label history prompt:triage
nova label history prompt:triage production --json

Protected labels — maker-checker moves (experimental, ADR-0114)

A protected label is a deployment label whose reassignment requires two distinct principals: a maker proposes the move, a checker approves it. Direct nova label set on a protected label is refused with guidance. Both steps are Ed25519-signed with the per-identity keyring (ADR-0058: ~/.config/novafabric/keyring/); self-approval is refused at the crypto level (matching key fingerprint or identity). The applied move lands in the same append-only asset_label_history audit table, reusing the pending move's ULID as its move_id. Free (unprotected) labels are completely unchanged. Fully offline; the optional --policy-ref Rego gate (ADR-0019) fails closed if the policy file is unreadable.

nova label protect <asset> <label> [--required-approvals N] [--policy-ref PATH] [--note TEXT] [--unprotect] [--json]

Mark a label protected (or --unprotect to revert to free ADR-0113 behaviour). Protecting an already-assigned label does not move it — it only governs future moves. Config changes are append-only events; the active setting is the newest one.

nova label protect prompt:triage production
nova label protect prompt:triage production --required-approvals 2 --note "gates live traffic"
nova label protect prompt:triage production --unprotect

nova label propose-move <asset> <label> --to <version> [--reason TEXT] [--identity NAME] [--json]

Maker step: create an Ed25519-signed pending move. The label does not move. Fail-closed: the target version must exist, the label must be protected (free labels use set), and at most one pending move per (asset, label) may exist at a time.

nova label propose-move prompt:triage production --to 8 --reason "v8 passed the eval gate"

nova label approve-move <asset> <label> <move_id> [--reject] [--note TEXT] [--identity NAME] [--json]

Checker step: approve (or --reject, terminal) a pending move. SoD is enforced before anything is recorded — the approver's keyring key fingerprint and identity must both differ from the proposer's. When distinct approvals reach required_approvals and the policy gate allows, the label is reassigned atomically with its audit row. A duplicate approver is recorded but counts once.

nova label approve-move prompt:triage production 01J2Q8ZK7M4YZ2K7N9DPBYK2WX --identity bob
nova label approve-move prompt:triage production 01J2Q8ZK7M4YZ2K7N9DPBYK2WX --reject --note "regressed"

nova label status <asset> [--label L] [--json]

Protection config, current label targets, and all recorded moves (pending and terminal) for an asset.

nova label status prompt:triage
nova label status prompt:triage --label production --json

nova report [--format {markdown,json,html,pdf}] [--output FILE]

Generate an asset inventory report. Defaults to Markdown on stdout. Use --output to write to a file. Valid --format values: markdown (default), json, html, pdf. Tab-completion available via nova --install-completion.

html (ADR-0201; works today) — one self-contained page (no JS, no external requests) with an assets-by-type inline-SVG chart, the inventory table, and the canonical JSON embedded for machine consumption.

pdf (works today, optional dependency) — the same page rendered through WeasyPrint. Requires the optional extra (pip install 'novafabric[compliance]') and --output; without the extra the command exits with the install hint instead of failing cryptically.

nova report --format html --output inventory.html
nova report --format pdf  --output inventory.pdf   # needs novafabric[compliance]

nova validate <path>

Smart routing:

An existing path always wins, so an asset spec, a capsule directory and a replay directory are all routed exactly as before.

nova validate .novafabric/runs/01HXAY7M5JZ8R7K4P9DPBYK2WX/
nova validate my-model.yaml

Tool-call schema conformance (ADR-0128; experimental) — capsule directories only:

Option Default Effect
--schemas off Also validate each tool call's arguments/result against its declared arguments_schema_ref/result_schema_ref JSON Schemas and print a conformance summary. Report-only: exit 0 even on violations. Records with no schema_ref are counted as null ("no schema declared" — not a failure); unresolvable refs are reported, never fatal. Schema resolution is local-only (relative refs inside the capsule directory or absolute local paths; http(s):// refs are never fetched).
--fail-on-schema-violation off With --schemas: exit non-zero if any checked payload violates its schema (CI gate).
--write off With --schemas: persist the computed schema_validation verdict blocks back into tool-calls.jsonl (backfill capsules captured before this feature; checked_at records the backfill time).
nova validate --schemas .novafabric/runs/<run-id>/                            # report-only
nova validate --schemas --fail-on-schema-violation .novafabric/runs/<run-id>/ # CI gate
nova validate --schemas --write .novafabric/runs/<run-id>/                    # backfill verdicts

At capture time the same verdict is attached automatically whenever a tool-call record declares a schema_ref — record-only, never raised into the workload. At replay time stored tool calls are re-validated against their current schemas: drift is recorded as schema_drift in replay_result.yaml in every mode, and blocks eligibility (exact_eligible: false) in exact mode only.


Asset lifecycle commands — v0.12 additions (C-1.4, C-1.5)

nova rollback <name> --actor <id>

Atomically roll back an asset to its most recent previous production version.

nova rollback my-agent --actor on-call-eng
nova rollback my-agent --to v1.8.0 --actor on-call-eng

Steps performed in one DB transaction:

  1. Find the current production version (error if none).
  2. Find the most recent prior production version (auto, or use --to).
  3. Archive the current production version.
  4. Promote the prior version back to production.
  5. Write an audit log entry with a rollback_reason field.

Options:

Errors clearly if the discovered prior version is archived (requires --to) or if no production history exists.

nova unregister <name@version>

Hard-delete an asset version from the registry. Removes the asset record, its eval results, and approvals. Blocked for staging, production, and pending_approval assets unless --force is given. Always writes an UNREGISTER audit entry.

nova unregister fraud-model@1.0.0
nova unregister fraud-model@1.0.0 --yes          # skip confirmation
nova unregister broken-agent@dev --force          # override status guard
nova unregister old-model@v2 --actor ops-eng --db ./registry.db

Options:

Status at deletion Default With --force
development, validated, archived ✅ allowed ✅ allowed
staging, production, pending_approval ❌ blocked ✅ allowed

Note: Consumed and depends_on lineage edges referencing the deleted asset are preserved — they record what actually happened in captured runs. Only the synthetic reg:<name>@<version> dependency edges (written at registration time) are removed.

nova capture ... --asset <ref> --require-asset-status <statuses>

Gate a capture run on the named asset's lifecycle status.

nova capture --asset my-agent@v1 --require-asset-status staging,production -- python agent.py
nova capture --asset my-agent@v1 --warn-if-asset-status development -- python agent.py
nova capture --asset my-agent@v1 --require-asset-status production --require-registered -- python agent.py

Options: