Collector binaries and HPC runners
Part of the NovaFabric CLI reference. Both nova and novafabric run the same binary.
Collector binaries (Phase 2 — cluster scale)
Phase 2 ships three Go binaries (collector/) that form the cluster-deployable
evidence ingestion tier. They are separate from the Python nova CLI. Install
them by building from source (Go 1.22+):
cd collector
go build -o bin/novafabric-collector ./cmd/novafabric-collector
go build -o bin/novafabric-verifier ./cmd/novafabric-verifier
go build -o bin/novafabric-hpc-hub ./cmd/novafabric-hpc-hubThese binaries complement the Python SDK: agents continue to emit events via the
existing nova capture / hook path; the collector tier signs and forwards those
events to Kafka at cluster scale.
novafabric-collector
Start the OTel Collector with the novaseal_batch_signer processor pre-bundled.
Suitable for both K8s (Deployment) and bare-metal gateway deployments.
novafabric-collector --config collector.yaml
novafabric-collector --config /etc/novafabric/collector.yamlThe collector is built using the OpenTelemetry Collector Builder (OCB) with the
novaseal_batch_signer custom processor registered. It accepts standard OTel
Collector configuration YAML.
Reference pipeline (see deploy/k8s/configmap-collector.yaml):
receivers:
otlp:
protocols:
grpc:
endpoint: "0.0.0.0:4317"
http:
endpoint: "0.0.0.0:4318"
processors:
batch:
send_batch_size: 1000
timeout: 5s
novaseal_batch_signer:
keystore_endpoint: "${NOVASEAL_KMS_ENDPOINT}"
fail_open: false # default: fail-closed — unsigned batches are refused
cache_ttl: 5m
key_rotation_interval: 1h
local_wal: false # true only in dev (sets NOVAFABRIC_KMS_LOCAL_WAL=1)
exporters:
kafka:
brokers: ["${KAFKA_BOOTSTRAP_SERVERS}"]
topic: "nova.evidence"
service:
pipelines:
logs:
receivers: [otlp]
processors: [batch, novaseal_batch_signer]
exporters: [kafka]novaseal_batch_signer processor options:
| Option | Default | Description |
|---|---|---|
keystore_endpoint |
— | NovaSeal KMS mTLS endpoint (required unless local_wal: true) |
fail_open |
false |
false = fail-closed (unsigned batches refused); true = continue without signature if KMS is down |
cache_ttl |
5m |
How long a fetched key is cached |
key_rotation_interval |
1h |
How often to rotate signing key |
local_wal |
false |
Dev only. Use a local Ed25519 key from ~/.novafabric/dev-keys/. Emits a WARN on every signing call. Refuses when NOVAFABRIC_ENV=production. |
The processor places each batch's signature and key ID in Resource.attributes:
nova.batch.signature— base64-encoded Ed25519 signature (88 chars)nova.batch.signing_key_id— UUID of the signing key
Prometheus metrics exposed at :8888/metrics:
nova_batch_sign_latency_seconds— histogram of per-batch signing timenova_batch_sign_errors_total{reason="..."}— counter by error categorynova_collector_forwarded_events_total— events forwarded to Kafka
K8s deployment: see deploy/k8s/ for a complete namespace, ConfigMap,
Deployment, DaemonSet (Fluent Bit node collector), and Service manifest set.
novafabric-verifier
Verify a NovaSeal Ed25519 batch signature offline. Returns exit 0 if the signature is valid, exit 1 if invalid.
# Verify with a PEM public key from the KMS
novafabric-verifier --batch batch.pb --pubkey kms-public.pem --key-id <key-uuid>
# Verify with the local dev key (matches what local_wal=true used to sign)
novafabric-verifier --batch batch.pb --local-walFlags:
| Flag | Required | Description |
|---|---|---|
--batch PATH |
yes | Path to a serialized OTLP ResourceLogs proto file |
--pubkey PATH |
one of | PEM-encoded Ed25519 public key file |
--key-id UUID |
with --pubkey |
Key ID to verify against |
--local-wal |
one of | Use the dev key from ~/.novafabric/dev-keys/ |
Output:
# Success
OK key_id=a1b2c3d4-...
# Failure
INVALID key_id=a1b2c3d4-... error=signature mismatchThe verifier re-runs the same canonical-encoding pre-pass that the signer used
(ADR-001: strip nova.batch.signature + nova.batch.signing_key_id, sort
Resource.attributes by key, sort each LogRecord.attributes by key, marshal
with proto.MarshalOptions{Deterministic: true}). If the bytes differ between
signer and verifier, the signature will fail.
novafabric-hpc-hub
Run the HPC cluster-hub NATS message handler. Subscribes to
nova.<cluster_id>.aggregate, signs each batch with NovaSeal, and publishes
signed batches to the Kafka regional topic.
# Production (uses NovaSeal KMS)
novafabric-hpc-hub \
--cluster-id hpc01 \
--hub nats://hub.cluster.local:4222 \
--kms https://kms.internal:8443
# Development (uses local dev key)
novafabric-hpc-hub \
--cluster-id dev \
--hub nats://localhost:4222 \
--local-walFlags:
| Flag | Env override | Default | Description |
|---|---|---|---|
--cluster-id |
NOVAFABRIC_CLUSTER_ID |
— | Required. Cluster identifier |
--hub |
NOVAFABRIC_HUB_ADDRESS |
— | NATS hub address (nats://...) |
--kms |
NOVASEAL_KMS_ENDPOINT |
— | NovaSeal KMS endpoint (required unless --local-wal) |
--local-wal |
NOVAFABRIC_KMS_LOCAL_WAL=1 |
false | Dev only. Use local Ed25519 key. |
--job-id |
SLURM_JOB_ID |
— | Slurm job ID (set automatically by Prolog) |
--spool |
— | — | Spool directory for this job |
The hub signs batches at the cluster boundary before forwarding to Kafka — compute nodes (leaf NATS instances) never touch the NovaSeal keystore. This preserves HPC air-gap security: only the hub node needs KMS network access.
Signals: SIGTERM / SIGINT trigger a clean shutdown.
HPC Slurm integration
For Slurm cluster deployments, the collector provides Prolog/Epilog scripts that initialize a per-job NATS leaf node and bounded-flush on job exit.
Installation (run on all Slurm compute nodes):
# 1. Copy scripts to compute nodes
cp deploy/hpc/prolog.sh /etc/novafabric/prolog.sh
cp deploy/hpc/epilog.sh /etc/novafabric/epilog.sh
chmod 755 /etc/novafabric/prolog.sh /etc/novafabric/epilog.sh
# 2. Add to slurm.conf (requires Slurm admin)
Prolog=/etc/novafabric/prolog.sh
Epilog=/etc/novafabric/epilog.sh
PrologFlags=AllocPer-job lifecycle:
- Prolog — creates
/tmp/novafabric/$SLURM_JOB_ID/, startsnovafabric-hpc-hubas a leaf NATS instance bound to that directory. Exits 0. - Job runs — agents write events to the spool via SDK hooks; the leaf node buffers and forwards to the cluster hub.
- Epilog — calls
nats stream flush --timeout=${NOVAFABRIC_EPILOG_FLUSH_TIMEOUT:-50}sto drain the spool, stops the leaf process. Always exits 0 to prevent Slurm from draining the node.
Lustre/NFS safety: The spool uses rename(2) as its only atomic commit primitive. No flock, mmap, or fcntl calls are used — these are unsafe on shared network filesystems. On nodes without local NVMe, set NOVAFABRIC_HPC_STORAGE=spool to use the JSONL spool as the NATS leaf store directory.
Reference: deploy/hpc/README.md, deploy/hpc/ansible/, ADR-0028, ADR-0043.
HPC runner commands (v0.16)
PBSRunner (Python API)
Run a capsule job on a PBS/Torque cluster:
from novafabric.runners import PBSRunner
from novafabric.runners.spec import RunnerSpec
runner = PBSRunner(RunnerSpec(image=None, command=["python", "agent.py"]))
result = runner.run(capsule_dir=Path(".novafabric/runs/01HXAY7M"))The runner submits via qsub, polls via qstat, and cancels via qdel on timeout or signal. Job script is injected at PBS_JOBSCRIPT.
LSFRunner (Python API)
Run a capsule job on an IBM Spectrum LSF cluster:
from novafabric.runners import LSFRunner
runner = LSFRunner(RunnerSpec(image=None, command=["python", "agent.py"]))
result = runner.run(capsule_dir=Path(".novafabric/runs/01HXAY7M"))The runner submits via bsub, polls via bjobs, and cancels via bkill. Job script is injected at LSF_JOBSCRIPT.