Concepts

This page is the conceptual reference for NovaFabric. It explains the nouns you will meet everywhere else in the docs and on the nova command line — what a Run Capsule is, how capture works without touching your code, the four replay modes, structural diff, lineage, the Asset Registry and its lifecycle, and how signed evidence is produced.

What you will learn

Maturity note. NovaFabric is local-first and in beta (v0.98.0). Nearly all shipped surfaces carry experimental maturity: they work today and are tested, but on-disk formats are not frozen until the v1.0 schema freeze. Anything labeled PLANNED or FUTURE DESIGN below is documented design intent (in the ADRs) and is not implemented — never treat it as shipped.


The five primitives at a glance

NovaFabric is framed around exactly five primitives, each with a public spec and JSON Schema. Cryptographic sealing is part of the Evidence Bundle / trust layer, not a sixth primitive.

# Primitive What it is Since
1 Asset Registry Local SQLite registry of versioned AI assets (name@version), pinned to a git SHA, with a six-state lifecycle v0.1
2 Run Capsule The fundamental unit of capture: a ULID-named directory holding every observable fact of one execution v0.2
3 Replay Re-execute or inspect a capsule with external calls controlled, in four honest modes v0.3
4 Lineage A directed provenance graph (SQLite cache) with mechanical edge types and OpenLineage emission v0.4
5 Evidence Bundle A signed, self-contained ZIP an auditor can verify offline with only sha256sum + an ed25519 verifier v0.4

The design invariant behind all five: the capsule is the source of truth; the registry, metadata DB, and lineage graph are derived, rebuildable indexes. Only two top-level formats exist — Run Capsule and Evidence Bundle — and introducing a third requires an accepted ADR.


Run Capsule

A Run Capsule is the fundamental unit of capture in NovaFabric. It is a directory containing all observable facts about a single command execution: the command, timing, environment, every LLM call made, every tool invoked, stdout/stderr, and a proof that no secrets escaped.

Capsules are identified by a ULID — a lexicographically sortable, time-prefixed unique identifier.

.novafabric/capsules/01HXAY7M5JZ8R7K4P9DPBYK2WX/
  capsule.yaml          ← run manifest
  trace.jsonl           ← execution spans (OTel)
  model-calls.jsonl     ← LLM API calls (OTel GenAI semconv)
  tool-calls.jsonl      ← tool invocations
  assets.jsonl          ← asset references
  env.lock              ← environment snapshot
  redaction-proof.json  ← secret scan proof
  replay.yaml           ← replay constraints
  lineage.jsonl         ← lineage edges
  inputs/
  outputs/
    stdout.txt
    stderr.txt

Capsules are written on success and failure. A failed run produces a complete capsule with status: failure and an error block — the evidence of a failure is just as durable as the evidence of a success.

Capsules support additive, optional extensions: blocks (for example slurm, kubernetes, ray, openlineage) so the schema can grow without introducing a new top-level format. Readers tolerate unknown fields; old capsules stay readable forever.

Sessions vs parent/child (two different groupings)

Two distinct, composable ways of relating capsules exist — do not conflate them:

A session member may itself be a distributed-run PARENT. A capsule that belongs to a session carries two additive optional back-reference fields (session_id, sequence); absence means a standalone run, byte-identical to before.

Related: Replay Modes (a replay is itself a new capsule you can diff), Lineage Graph, and Evidence Bundle.


Evidence facets (experimental)

A capsule can carry optional, additive evidence facets — a facets map (ADR-0196) whose keys are per-domain evidence objects, each owned by its own ADR. A capsule without the facets key is valid and byte-identical to before, so facets never change the base schema. facets holds first-party, NovaFabric-recorded evidence; it is deliberately distinct from extensions, which carries third-party vendor data — an auditor must be able to tell NovaFabric-recorded evidence from a vendor annotation. The registry is closed on purpose (ADR-0196 D2): only the registered names below are accepted, so a typo cannot silently produce a capsule missing the evidence its author believed it carried.

Not the same as OpenLineage facets. The word "facet" is also used, unrelatedly, for the OpenLineage run/dataset facets that nova lineage emit-openlineage --with-facets attaches to emitted lineage events. Those live in OpenLineage payloads, not on the capsule. This section is about the capsule facets container.

Each facet is populated by a record-only evidence module: it records what another system did and never orchestrates, enforces, adjudicates, moves funds, controls a device, or sits in a hot path. All are experimental Python APIs today; none of them registers a nova CLI command.

Module Capsule facet Records ADR
novafabric.a2a a2a_messages Multi-agent A2A messages & handoffs — the wire between agents. Distinct from the novafabric.adapters.a2a capture-time SDK adapter, which is a separate subsystem. ADR-0142
novafabric.settlement settlement Agentic-commerce settlement provenance — mandate reconciliation, finality, and non-repudiation binding. Never processes payments or holds/moves funds. ADR-0163
novafabric.embodied embodied Embodied / cyber-physical agent evidence — declared sensor streams and actuation records. Stores references, digests, and counts only (never frames, point clouds, audio, or control credentials); never in a control path. ADR-0162
novafabric.science science_provenance Scientific-reproducibility / research-integrity provenance as a verifiable DAG of research steps. Never runs experiments or adjudicates validity. ADR-0164
novafabric.memstore memstore_mutation Persistent-knowledge / organisational-memory governance — an append-only ledger of who changed which shared-KB entry and when (store-external). Distinct from the nova memory capture feature. ADR-0171
novafabric.retrieval fetch_provenance (+ source pinning) Retrieval-source authority & knowledge provenance — what an external retriever or the agent fetched, and whether a pinned source still matches. Never fetches, crawls, or ranks. ADR-0153
novafabric.context (standalone artifacts) Context provenance — an ordered manifest of what entered the model's context window (ContextManifest) and the span→chunk support map a producer claimed (GroundingMap). Records sha256: digests only; never scores groundedness. Ships as standalone artifacts, not a registered capsule facet (ADR-0196 D2). ADR-0143 (NF-112/113)

Capture Hook Mechanism

nova capture <cmd> works without modifying application code. It injects a sitecustomize.py loader into the subprocess via PYTHONPATH. When Python starts the subprocess, it imports this loader automatically, which installs monkey-patches for:

The layering is deliberate: per-SDK hooks capture rich, structured request and response objects; wire-level hooks form a safety net beneath them so a call that bypasses a known SDK (or uses a version whose surface changed) is still recorded once at the transport layer. Because the hooks reach down to urllib3, capture does not depend on your choice of HTTP client library.

Patches are removed after the run. If an SDK is not installed, its hook is silently skipped. Capture works even if none of the AI SDKs are present — the capsule is still written with environment, stdout/stderr, and timing.

For uninstrumented agents (Claude Desktop, Cursor, third-party SDKs that do not import the Python mcp package), nova mcp-proxy is a transparent stdio proxy that records the same tool-calls.jsonl schema by sitting between the client and an upstream MCP server. See docs/cli-reference.md and ADR-0015 §Secondary; experimental in v0.5.x.

For non-Python clients over HTTP, nova api-proxy is a transparent HTTP proxy that classifies traffic against the same vendored provider registry — so a client in any language can be captured without SDK hooks (v0.6).

Runners — where the captured workload executes

The orchestrator delegates subprocess execution to a runner (ADR-0025). Choose one with nova capture --runner <name>:

Runner Since What it does Key required options
local v0.5.x Runs the workload as a local subprocess. Default; no remote infrastructure needed.
docker v0.6 Runs the workload inside a Docker container. The image must already have NovaFabric installed. Capsule directory is mounted as a volume so artifacts land on the host. image
kubernetes v0.6.1 Runs the workload as a Kubernetes Job via kubectl shell-out. Artifacts are pulled back via kubectl cp. image, namespace
slurm v0.6.1 Runs the workload as a SLURM batch job via sbatch + sacct. Trusts a shared filesystem for artifacts (no rsync). partition

Runner details:

Note. The single-node runner is the smallest case of a distributed run. The cluster-scale tiers shipped in v0.10+ with honest maturity labels: parent/child capsules (nova run …) are a prototype — implemented and tested, not yet validated at target scale — and the collector tier, object capsule store, and Postgres metadata DB are experimental. See ROADMAP.md for per-component labels; the federation and at-scale graph tiers remain future design.


model-calls.jsonl

Every LLM API call intercepted by the capture hooks is recorded as one JSONL line following the OpenTelemetry GenAI semantic conventions:

{
  "schema_version": "0.1.0",
  "semconv_version": "1.30.0",
  "model_call_id": "01HXAY7M6FN9TQGE0V0M7PAY1Q",
  "parent_span_id": "657bff2c61ddad1c",
  "started_at": "2026-05-09T12:34:56.123456Z",
  "finished_at": "2026-05-09T12:34:57.353456Z",
  "duration_ms": 1230,
  "status": "success",
  "gen_ai.system": "openai",
  "gen_ai.operation.name": "chat",
  "gen_ai.request.model": "gpt-4o",
  "gen_ai.response.model": "gpt-4o-2024-08-06",
  "gen_ai.response.id": "chatcmpl-abc123",
  "gen_ai.request.temperature": 0.7,
  "gen_ai.request.max_tokens": 1024,
  "gen_ai.request.top_p": 0.95,
  "gen_ai.request.seed": 42,
  "gen_ai.request.stop_sequences": ["END"],
  "gen_ai.request.messages": [...],
  "gen_ai.response.choices": [...],
  "gen_ai.response.finish_reasons": ["stop"],
  "gen_ai.usage.input_tokens": 42,
  "gen_ai.usage.output_tokens": 18,
  "endpoint": "https://api.openai.com/v1/chat/completions"
}

The full set of OTel "Required when applicable" fields is extracted by both the per-SDK hooks (_openai, _anthropic) and the wire-level hooks (_httpx, _requests, _aiohttp, _urllib3) when present in the request body — so nova replay --mode exact has the determinism inputs it needs (temperature, top_p, seed) regardless of which transport the captured call took. See design/spec/model-call-v0.md for the full field reference.

This format is stable across providers and is the basis for two things:


Environment Lock (env.lock)

The environment lock records the full execution environment at capture time:

nova replay uses env.lock to warn about environment mismatches before re-executing the command. This is what makes the exact replay mode falsifiable: eligibility for byte-exact replay depends on a matching, deterministic environment plus a per-call seed, and env.lock is where that determination starts.


Secret Scanning and Redaction

Before a capsule is finalized, SecretScannerV0 scans all JSONL artifacts for 12 LLM provider key patterns (Anthropic, OpenAI, HuggingFace, Replicate, Langfuse, and others). Detected values are redacted in-place as [REDACTED:rule-id].

After scanning, redaction-proof.json is written. It records:

Capsules without redaction-proof.json are considered invalid by nova validate and cannot be exported into an Evidence Bundle. Verifiable redaction is therefore not optional — it is a precondition of every downstream audit artifact.


Replay Modes

A replay re-executes or inspects a capsule with all external calls controlled by NovaFabric. There are four honest, falsifiable modes, plus a fifth, experimental counterfactual mode. A replay is itself a new capsule, so you can diff a replay against the original run.

Mode Spawns subprocess? Network? Best for
forensic No No Audit / post-incident inspection
mocked Yes LLM served from cache; tools gated by safety ladder CI / regression
semantic Yes Yes (re-executes) Drifting remote LLMs — judges meaning, not tokens
exact Yes Controlled Local / on-prem / compliance byte-exact re-run
intervention (experimental, ADR-0086) Yes, under mocked semantics No Counterfactual root-cause: substitute one captured event per an InterventionSpec, re-execute downstream, and record whether the outcome flips

Honesty note. NovaFabric explicitly does not claim byte-exact replay of remote LLM calls. exact mode requires a deterministic environment and a per-call seed, which is realistic for local/on-prem models but not for a remote endpoint that can change under you. For remote LLMs that drift, use semantic mode, which scores similarity of meaning on a 0.0–1.0 scale.

forensic mode

Read-only inspection. NovaFabric does not spawn a subprocess. It reads and returns the manifest, traces, and model calls directly from the capsule. No network access, no mutation.

Use forensic mode to inspect what happened without any risk of side effects.

mocked mode

The original command is re-spawned as a subprocess. All LLM calls are intercepted and served from the capsule cache in order (MockModelDispatcher). Tool calls are mocked or denied according to the safety ladder below.

Safety ladder (mocked replay)

Mocked replay gates tool call re-execution behind explicit flags. By default, every tool call is denied — you opt in, rung by rung, to exactly the level of side effect you are willing to allow:

(none)               — deny all tool calls
  ↓  --allow-readonly
read-only            — permit read-only tool calls
  ↓  --allow-mutating
idempotent-write     — permit idempotent writes
non-idempotent-write
  ↓  --allow-external-side-effects
external-side-effect — permit calls with external effects
  ↓  --allow-unknown-mutation
unknown              — permit unclassified tool calls

Per-tool overrides can be specified in replay.yaml.

semantic mode

Re-executes the command against live models and judges meaning rather than tokens, returning a 0.0–1.0 similarity score. This is the honest answer to the reality that remote LLMs drift: two runs weeks apart may produce different tokens yet mean the same thing.

exact mode

Byte-exact eligibility requiring a deterministic environment and a per-call seed. This is the compliance-grade mode for local and on-prem models where determinism is achievable.

intervention mode (experimental)

Answers a counterfactual question: if this one recorded event had gone differently, would the run's outcome have changed? An InterventionSpec names one captured model or tool call and a substitute outcome for it; the engine re-executes everything downstream of that point under mocked semantics (zero live tokens) and writes a diffable capsule hard-marked replay_mode: intervention, never mistakable for a real run. This is the building block behind the no-LLM causal-graph diagnostic suite below — see Diagnose: causal-graph attribution and counterfactual root-cause search.


Diagnose: causal-graph attribution and counterfactual root-cause search (experimental)

nova diagnose <run-id> (ADR-0084) attributes a failed run to its most likely responsible step: it walks the captured trace plus the lineage graph and produces a ranked attribution — which agent/step is most likely responsible, plus an AgentErrorTaxonomy label (MEMORY / REFLECTION / PLANNING / ACTION / SYSTEM / UNKNOWN). Scores are relative ranking weights, not probabilities, and every candidate is honestly marked verification: unverified unless tested.

Two further, experimental layers (ADR-0101) turn a ranking into evidence:

None of this uses an LLM judge: attribution is structural (trace + lineage walk) and verification is a real, zero-token replay — a hallucination-risk or root-cause finding you can point at evidence for, not a model's opinion.


Accountability Spine (experimental)

Three complementary, append-only evidence surfaces (ADR-0093/0094/0095), collectively referred to as the Accountability Spine, sit alongside NovaSeal and the Evidence Bundle:

Surface CLI Records ADR
Energy-Anchored Action Receipts nova energy probe/attest/verify/report Per-action energy receipts + a conservation check, so an energy claim is falsifiable rather than asserted ADR-0093
Adversary-anchored accountability ledger nova ledger anchor/verify/status Sidecar hash chains + signed checkpoints for every capsule .jsonl stream, so tampering with a stream after the fact is detectable ADR-0094
Structured safety case nova safety-case build/verify/export A compiled, evidence-grounded Claims-Arguments-Evidence (CAE) safety case over a capsule, exportable as JSON, Markdown, an EU AI Act Annex IV section, or a NIST AI RMF report ADR-0095

The dashboard surfaces all three read-only under the Spine tab. Like the evidence facets above, these are record-only: they never enforce, adjudicate, or gate anything themselves — they make an accountability claim checkable.


Structural Diff

nova diff compares two capsules field by field:

DiffReport has three counters: changed, added, removed. The --assert-no-regressions flag exits 1 if any of these are non-zero, making it suitable as a CI gate:

# Fail the pipeline if today's run diverges from a known-good baseline
nova diff baseline-capsule/ candidate-capsule/ --assert-no-regressions

This is the mechanism behind the "worked yesterday, fails today" flaky-agent problem: wire nova diff --assert-no-regressions into CI and a structural regression stops the merge.

Offline analytics over capsules (experimental)

Because every metric a run produced — cost, tokens, latency, scores — is already recorded inside its capsule, analytics never needs a server. The experimental v0.59 surfaces read those recorded facts, offline and read-only: nova query is a bounded filter → group-by → aggregate DSL over the local capsule directory, nova view persists a named query as a small versionable file, and nova trend buckets one metric over time or by asset into a JSON/HTML report. Nothing is recomputed or fetched; the capsule directory is never written. See the CLI reference.


Lineage Graph

The lineage graph is a directed graph stored in SQLite — a rebuildable cache derived from each capsule's lineage.jsonl. It has two table types:

Because the graph is derived from the capsules, it can always be rebuilt from them — the capsules remain the source of truth. The SQLite backend is the local-mode default and is crash-safe (WAL); it is well suited below roughly one million edges. Four at-scale backends exist for larger graphs, all experimental and testcontainers-verified, selected via nova lineage-store migrate / profile: Kuzu (embedded, benchmark-cleared at 10M edges, p99 blast_radius 45.5ms), Postgres (recursive-CTE, no extension needed), Apache AGE (openCypher on Postgres), and JanusGraph (Gremlin, needs the GraphSON serializer). There are currently zero NotImplementedError stubs among them. Billion-edge cross-cluster federation remains FUTURE DESIGN, not implemented.

Edge types

Edge type Meaning
consumed The run read/used this asset or artifact
produced_by This artifact was produced by the run
replayed_from This run is a replay of the referenced run

Confidence levels

Confidence Meaning
observed Directly recorded at runtime (e.g., explicit API call)
inferred Derived from capsule structure (e.g., output file presence)

Graph queries

Query Direction Use case
provenance Backward (ancestors) "What did this run depend on?"
blast-radius Forward (descendants) "What runs consume this asset?"
replay-chain Backward via replayed_from "What is the original run for this replay?"
time-travel Backward, filtered by timestamp "What was the lineage state as of time T?"

Graph analytics (experimental)

Traversal answers "what connects to this node?". A read-only analytics layer (ADRs 0212–0215) answers "what does the whole graph mean?": centrality (nova lineage metrics — hubs and single points of failure), node-level root-cause ranking (nova lineage root-cause), interop export to GraphML/GEXF/Cypher (nova lineage export-graph), and a synthesized intelligence report (nova insights). All are descriptive rankings for attention, not calibrated importance, and degrade honestly when a data source (such as cost) is absent.

OpenLineage Integration

NovaFabric emits OpenLineage 2.0.2 events from capsules. Events are typed as START and COMPLETE (or FAIL).

Events can be emitted:

This enables integration with data-catalog tools such as Marquez, Atlan, and OpenMetadata — so NovaFabric's run-level provenance can join the broader data-pipeline lineage you already track.

Custom run facets (experimental). emit-openlineage --with-facets attaches NovaFabric-specific facets to the COMPLETE event — the capsule id/hash, the eval verdict, the promotion-policy decision, and reproducibility run params — and --otel-correlation adds trace_id/span_id so a lineage node links to its OTel GenAI spans. Facets are additive and schema-validated before emission, so a catalog that ignores them still receives valid core OpenLineage events. Run facets, dataset provenance cards, and benchmark-contamination checks are shown hands-on in the feature tour §17.

Note: these OpenLineage facets are unrelated to the capsule facets container (ADR-0196). Same word, two different places: OpenLineage facets live in emitted lineage events; evidence facets live on the capsule.


AI Asset

An AI asset is any versioned artifact in an ML or LLM system. NovaFabric tracks seven types:

Type What it represents
model A trained ML/LLM model — framework, artifact path
agent An AI agent — model ref, tools, prompts, policies, eval suites
prompt A prompt template used by one or more agents
tool A callable function exposed to an agent
dataset A dataset used for training, fine-tuning, or evaluation
evaluation An evaluation suite definition (test cases, scoring logic)
deployment A running endpoint — base URL, environment

Assets are described by YAML spec files and stored in the local Asset Registry. Every registered asset is addressed as name@version, pinned to a git commit SHA, and carries a lifecycle status that tracks where it is in your development-to-production pipeline.

Prompts as versioned assets (experimental, ADR-0112). Beyond the YAML spec route, nova prompt register turns a prompt into an immutable, content-addressed registry version: every edit registers a new version (never a mutation), and a run references it as prompt:<id>@<version>+sha256:<hex> — so a capsule can be tied to exactly the prompt bytes that ran. Mutable deployment labels (nova label, ADR-0113) point a name like production at one immutable version, and prompt composition (nova prompt compose, ADR-0115) snapshots pinned includes so an assembled prompt rebuilds byte-identically. NovaFabric never renders or serves prompts — it versions and proves them.


Asset Lifecycle

The state machine

development ──► validated ──► pending_approval ──► staging ──► production ──► archived
      │                                                  ▲
      └──────────────────────────────────────────────────┘
           (direct skip: development → staging or production)

All six statuses in plain language:

Status Meaning
development Work in progress. Default when registered. No constraints.
validated Passed automated checks (nova validate). Spec is well-formed, secrets are clean.
pending_approval Waiting for a human sign-off (nova approve). Useful in regulated environments.
staging Cleared for pre-production use. Agents require a passing eval to reach here.
production Live. Agents require a passing eval to reach here.
archived Retired. No further promotion possible. Capsules and lineage are still readable.

Direct skips (e.g. development → production) are allowed. Once archived, an asset cannot be re-promoted. The validated and pending_approval steps are optional — small teams often go straight from development to staging.

What actually changes when you promote an asset

Promoting an asset updates one database record. That is all NovaFabric does.

nova promote direct my-agent@v1.2.0 --to production --actor alice

Internally this:

  1. Checks the policy engine — allow or deny, written to the audit log.
  2. For agent type going to staging/production: verifies a passing eval result exists. Blocks if not (unless --force).
  3. Runs UPDATE assets SET status = 'production', promoted_at = ..., promoted_by = 'alice'.
  4. Returns the updated record.

Nothing else happens. Your agent does not restart. Your prompt does not reload. No traffic is rerouted. No container is restarted. The capture hooks and runners are unchanged.

The status is governance metadata — a reliable signal your tooling can read and act on. NovaFabric gives you the signal; your platform decides what to do with it. (Policy checks and maker-checker approval are backed by OPA/Rego and signed decision logs — see Eval-Gated Promotion.)

How external tooling consumes the status

The pattern is: your CI/CD pipeline (or any script) reads the status from NovaFabric and makes a deployment decision based on it.

# In a GitHub Actions deploy step or Argo CD sync hook:
STATUS=$(nova inspect my-agent@v1.2.0 | jq -r .status)

if [ "$STATUS" = "production" ]; then
  kubectl set image deployment/my-agent container=my-agent:v1.2.0
else
  echo "Asset not in production — deploy blocked (status=$STATUS)"
  exit 1
fi

NovaFabric doesn't own your deployment. It owns the evidence that the asset is ready for deployment.

Use cases

Use case 1 — CI/CD gate for an LLM-powered feature

You have a customer-facing summarisation agent. Your CI pipeline runs evals on every PR. Only if evals pass and a team lead approves does the agent reach production, which triggers your Helm chart to update the image tag.

# .github/workflows/deploy-agent.yml (simplified)
- run: nova eval summariser@${{ github.sha }} --suite regression
- run: nova promote direct summariser@${{ github.sha }} --to pending_approval --actor ci-bot
# ... human approves in dashboard or via CLI ...
- run: |
    nova promote direct summariser@${{ github.sha }} --to production --actor ${{ github.actor }}
    helm upgrade summariser ./chart --set image.tag=${{ github.sha }}

Use case 2 — Safe prompt rollout without a full redeploy

You update a system prompt template. You don't want to redeploy the entire agent service — just swap the prompt version your agent reads at startup. The registry tells your agent loader which version is production.

# agent startup code
from novafabric.registry.service import get_asset

asset = get_asset("system-prompt-v2", version="latest")
if asset["status"] != "production":
    raise RuntimeError(f"Prompt not promoted to production (status={asset['status']})")

prompt_text = json.loads(asset["spec_json"])["spec"]["template"]

If the prompt has a bug and you archive it, the next agent restart fails fast rather than silently loading a bad prompt.

Use case 3 — Audit trail for a regulated environment

Your company needs to demonstrate that no AI model reached production without a human sign-off. Every promotion through pending_approval creates an approval record (approver, timestamp, note) and an audit log entry. At audit time:

nova inspect risk-scorer@v3.1.0 --format json | jq '{
  status,
  promoted_by,
  promoted_at,
  forced_promotion
}'
# → { "status": "production", "promoted_by": "alice", "promoted_at": "2026-03-15T10:22:01Z", "forced_promotion": false }

No manual tracking spreadsheet. The registry is the record.

Use case 4 — Blast-radius control when an agent misbehaves

An agent in production starts producing bad outputs after an upstream model update. You archive it:

nova promote direct my-agent@v2.0.0 --to archived --actor on-call-eng

Your deployment hook detects the status change, rolls back to the previous production version (my-agent@v1.9.0), and pages the team. The capsules from the bad version are still fully intact for forensic replay — nothing is deleted.

Use case 5 — Multi-team handoff on an HPC cluster

A research team develops and validates an agent on their workstation. When they're ready for the ML engineering team to integrate it into the batch pipeline, they promote it to staging. The ML team's Slurm job template only launches agents whose status is staging or production:

# check_asset_status.sh — called from sbatch prolog
STATUS=$(NOVAFABRIC_DB_PATH=/shared/nova/registry.db nova inspect $AGENT_NAME@$AGENT_VERSION | jq -r .status)
if [[ "$STATUS" != "staging" && "$STATUS" != "production" ]]; then
  echo "PROLOG FAIL: $AGENT_NAME@$AGENT_VERSION is not ready (status=$STATUS)" >&2
  exit 1
fi

The registry on the shared filesystem (/shared/nova/registry.db) is the handoff protocol between the two teams — no Slack messages, no wiki pages.


Asset Registry

The registry is a local SQLite database at ~/.novafabric/registry.db (override: NOVAFABRIC_DB_PATH). It stores:

Table Contents
assets Every registered asset — name, version, type, status, full spec JSON, git SHA, promotion history
eval_results Pass/fail eval results per asset, per suite
approvals Human sign-off records: approver, timestamp, note

The registry is entirely local. There is no server, no credentials, and no network dependency for local mode.

In server mode (nova serve, v0.7, experimental), the same data lives in Postgres 16 and is accessible via the REST API at /api/v1/assets, guarded by OIDC and an RBAC model (reader < writer < admin, plus an orthogonal auditor). Server mode is strictly additive: local mode never requires it.


Eval-Gated Promotion

nova eval <agent@version> discovers evaluation suites declared in the agent spec (spec.evals) and resolves them via Python entry points in the novafabric.evals group. Results are stored in the eval_results table.

Promotion to staging or production queries eval_results for a passing result. If none exists, promotion is blocked unless --force is used. A forced promotion is recorded as such (forced_promotion = true) and appears in nova inspect output and the audit log — the escape hatch is visible, not silent.

Standard eval suites ship OCI-pinned for reproducibility: GAIA, SWE-bench, AgentBench, MMLU, and a Smoke suite (v0.9). Promotion can be Rego-gated to block on a regression against a prior result.


Evidence Bundle

An Evidence Bundle is the compliance primitive: a signed, self-contained ZIP built by nova export-evidence that embeds

Its defining property is offline verifiability: a recipient can verify it with only sha256sum plus an ed25519 verifier — no NovaFabric runtime required. That is what makes the capsule a durable, portable artifact you own rather than a row in someone else's database. A capsule missing its redaction-proof.json cannot be exported (see Secret Scanning and Redaction).

Shipped trust-layer surfaces (v0.4+): ed25519-signed Evidence Bundles, in-toto DSSE attestations, secret scanning with verifiable redaction proofs, OPA/Rego policy gates with maker-checker promotion, and WORM storage adapters (S3 Object Lock / Azure immutable blob / GCS Bucket Lock, with legal holds).

Compliance-export cohort (experimental, ADR-0107 + related). Beyond the general-purpose Evidence Bundle, nova export-compliance and a long tail of dedicated nova export-* commands render capsule/lineage evidence into regulator-shaped artifacts: EU AI Act (Annex IV, Art.12 record-keeping, Art.50 marking + C2PA/SynthID assertion, Art.53 GPAI hash-chained form, Art.73 incidents with a deadline clock chained to Art.72 post-market monitoring), ISO 42001/42005, NIST AI RMF, the NIST GenAI/CSA profile, GDPR (Art.30 RoPA, Art.17 erasure), CycloneDX AI-SBOM, and more. Every exporter is a pure projection over already-captured evidence — it renders facts, never adjudicates compliance — and each is marked with an evidence_source provenance tag (operator_asserted / capsule_verified / unverifiable, design/spec/evidence-source-provenance-marker.md) so a reader can tell what was actually observed from what an operator merely declared.

Experimental — NovaSeal signing core. The in-process NovaSeal core shipped in v0.10+ (experimental): DSSE signing (ECDSA P-256), best-effort RFC 3161 trusted timestamps, and an append-only SQLite Merkle log, verified offline by nova verify (signature_ok / timestamp_ok / log_integrity_ok), plus the maker-checker nova seal propose/approve/verify chain (ADR-0059). Still PLANNED / FUTURE DESIGN per ADR-0041: the dedicated network signing service, qualified timestamps, Sigstore-keyless signing as the default path, WORM-backed retention at the seal layer, and PQC/ML-DSA. The most stable portable proof today remains the ed25519 / in-toto DSSE Evidence Bundle described above.


SDK Decorator

@novafabric.agent is an in-process alternative to nova capture. It wraps an agent function with the same capture hooks used by the CLI, recording all LLM calls into a capsule. Without capsule_dir, it emits OTel spans only — this is the v0.1 observability mode.


ULID

Run IDs and replay IDs are Universally Unique Lexicographically Sortable Identifiers. They embed a millisecond timestamp as the high bits, so capsule directories naturally sort chronologically without a separate index.


Summary and next steps

You now have the vocabulary NovaFabric is built on:

Where to go next: