Collector binaries and HPC runners

Part of the NovaFabric CLI reference. Both nova and novafabric run the same binary.

Collector binaries (Phase 2 — cluster scale)

Phase 2 ships three Go binaries (collector/) that form the cluster-deployable evidence ingestion tier. They are separate from the Python nova CLI. Install them by building from source (Go 1.22+):

cd collector
go build -o bin/novafabric-collector ./cmd/novafabric-collector
go build -o bin/novafabric-verifier  ./cmd/novafabric-verifier
go build -o bin/novafabric-hpc-hub   ./cmd/novafabric-hpc-hub

These binaries complement the Python SDK: agents continue to emit events via the existing nova capture / hook path; the collector tier signs and forwards those events to Kafka at cluster scale.


novafabric-collector

Start the OTel Collector with the novaseal_batch_signer processor pre-bundled. Suitable for both K8s (Deployment) and bare-metal gateway deployments.

novafabric-collector --config collector.yaml
novafabric-collector --config /etc/novafabric/collector.yaml

The collector is built using the OpenTelemetry Collector Builder (OCB) with the novaseal_batch_signer custom processor registered. It accepts standard OTel Collector configuration YAML.

Reference pipeline (see deploy/k8s/configmap-collector.yaml):

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: "0.0.0.0:4317"
      http:
        endpoint: "0.0.0.0:4318"
processors:
  batch:
    send_batch_size: 1000
    timeout: 5s
  novaseal_batch_signer:
    keystore_endpoint: "${NOVASEAL_KMS_ENDPOINT}"
    fail_open: false          # default: fail-closed — unsigned batches are refused
    cache_ttl: 5m
    key_rotation_interval: 1h
    local_wal: false          # true only in dev (sets NOVAFABRIC_KMS_LOCAL_WAL=1)
exporters:
  kafka:
    brokers: ["${KAFKA_BOOTSTRAP_SERVERS}"]
    topic: "nova.evidence"
service:
  pipelines:
    logs:
      receivers: [otlp]
      processors: [batch, novaseal_batch_signer]
      exporters: [kafka]

novaseal_batch_signer processor options:

Option Default Description
keystore_endpoint — NovaSeal KMS mTLS endpoint (required unless local_wal: true)
fail_open false false = fail-closed (unsigned batches refused); true = continue without signature if KMS is down
cache_ttl 5m How long a fetched key is cached
key_rotation_interval 1h How often to rotate signing key
local_wal false Dev only. Use a local Ed25519 key from ~/.novafabric/dev-keys/. Emits a WARN on every signing call. Refuses when NOVAFABRIC_ENV=production.

The processor places each batch's signature and key ID in Resource.attributes:

Prometheus metrics exposed at :8888/metrics:

K8s deployment: see deploy/k8s/ for a complete namespace, ConfigMap, Deployment, DaemonSet (Fluent Bit node collector), and Service manifest set.


novafabric-verifier

Verify a NovaSeal Ed25519 batch signature offline. Returns exit 0 if the signature is valid, exit 1 if invalid.

# Verify with a PEM public key from the KMS
novafabric-verifier --batch batch.pb --pubkey kms-public.pem --key-id <key-uuid>

# Verify with the local dev key (matches what local_wal=true used to sign)
novafabric-verifier --batch batch.pb --local-wal

Flags:

Flag Required Description
--batch PATH yes Path to a serialized OTLP ResourceLogs proto file
--pubkey PATH one of PEM-encoded Ed25519 public key file
--key-id UUID with --pubkey Key ID to verify against
--local-wal one of Use the dev key from ~/.novafabric/dev-keys/

Output:

# Success
OK key_id=a1b2c3d4-...

# Failure
INVALID key_id=a1b2c3d4-... error=signature mismatch

The verifier re-runs the same canonical-encoding pre-pass that the signer used (ADR-001: strip nova.batch.signature + nova.batch.signing_key_id, sort Resource.attributes by key, sort each LogRecord.attributes by key, marshal with proto.MarshalOptions{Deterministic: true}). If the bytes differ between signer and verifier, the signature will fail.


novafabric-hpc-hub

Run the HPC cluster-hub NATS message handler. Subscribes to nova.<cluster_id>.aggregate, signs each batch with NovaSeal, and publishes signed batches to the Kafka regional topic.

# Production (uses NovaSeal KMS)
novafabric-hpc-hub \
  --cluster-id hpc01 \
  --hub nats://hub.cluster.local:4222 \
  --kms https://kms.internal:8443

# Development (uses local dev key)
novafabric-hpc-hub \
  --cluster-id dev \
  --hub nats://localhost:4222 \
  --local-wal

Flags:

Flag Env override Default Description
--cluster-id NOVAFABRIC_CLUSTER_ID — Required. Cluster identifier
--hub NOVAFABRIC_HUB_ADDRESS — NATS hub address (nats://...)
--kms NOVASEAL_KMS_ENDPOINT — NovaSeal KMS endpoint (required unless --local-wal)
--local-wal NOVAFABRIC_KMS_LOCAL_WAL=1 false Dev only. Use local Ed25519 key.
--job-id SLURM_JOB_ID — Slurm job ID (set automatically by Prolog)
--spool — — Spool directory for this job

The hub signs batches at the cluster boundary before forwarding to Kafka — compute nodes (leaf NATS instances) never touch the NovaSeal keystore. This preserves HPC air-gap security: only the hub node needs KMS network access.

Signals: SIGTERM / SIGINT trigger a clean shutdown.


HPC Slurm integration

For Slurm cluster deployments, the collector provides Prolog/Epilog scripts that initialize a per-job NATS leaf node and bounded-flush on job exit.

Installation (run on all Slurm compute nodes):

# 1. Copy scripts to compute nodes
cp deploy/hpc/prolog.sh /etc/novafabric/prolog.sh
cp deploy/hpc/epilog.sh /etc/novafabric/epilog.sh
chmod 755 /etc/novafabric/prolog.sh /etc/novafabric/epilog.sh

# 2. Add to slurm.conf (requires Slurm admin)
Prolog=/etc/novafabric/prolog.sh
Epilog=/etc/novafabric/epilog.sh
PrologFlags=Alloc

Per-job lifecycle:

  1. Prolog — creates /tmp/novafabric/$SLURM_JOB_ID/, starts novafabric-hpc-hub as a leaf NATS instance bound to that directory. Exits 0.
  2. Job runs — agents write events to the spool via SDK hooks; the leaf node buffers and forwards to the cluster hub.
  3. Epilog — calls nats stream flush --timeout=${NOVAFABRIC_EPILOG_FLUSH_TIMEOUT:-50}s to drain the spool, stops the leaf process. Always exits 0 to prevent Slurm from draining the node.

Lustre/NFS safety: The spool uses rename(2) as its only atomic commit primitive. No flock, mmap, or fcntl calls are used — these are unsafe on shared network filesystems. On nodes without local NVMe, set NOVAFABRIC_HPC_STORAGE=spool to use the JSONL spool as the NATS leaf store directory.

Reference: deploy/hpc/README.md, deploy/hpc/ansible/, ADR-0028, ADR-0043.


HPC runner commands (v0.16)

PBSRunner (Python API)

Run a capsule job on a PBS/Torque cluster:

from novafabric.runners import PBSRunner
from novafabric.runners.spec import RunnerSpec

runner = PBSRunner(RunnerSpec(image=None, command=["python", "agent.py"]))
result = runner.run(capsule_dir=Path(".novafabric/runs/01HXAY7M"))

The runner submits via qsub, polls via qstat, and cancels via qdel on timeout or signal. Job script is injected at PBS_JOBSCRIPT.

LSFRunner (Python API)

Run a capsule job on an IBM Spectrum LSF cluster:

from novafabric.runners import LSFRunner

runner = LSFRunner(RunnerSpec(image=None, command=["python", "agent.py"]))
result = runner.run(capsule_dir=Path(".novafabric/runs/01HXAY7M"))

The runner submits via bsub, polls via bjobs, and cancels via bkill. Job script is injected at LSF_JOBSCRIPT.