production toolkit plan - Capsize-Games/spikeforge GitHub Wiki
Umbrella design for the production-toolkit program. Companion documents: the user-facing use-case umbrella
production_use_cases.mdand the first fully scoped applicationuse_case_streaming_timeseries.md. Read this file first: it defines the shared interfaces every use case consumes.
Design/spec only. Every claim about current code is grounded in a file/line reference so a Code-mode agent can execute this file-by-file.
Objective. Close the gaps identified in the "Q1 / Q2" analysis so a user can
go from train to a running, observable service and a simulator-backed test
deploy without owning a neuromorphic chip. Today the repo trains well and
declares deployment honestly, but the path from a checkpoint to a served model
is missing the SNN-specific pieces (stateful inference, a portable artifact,
encoder-at-inference, a headless serving runtime) and the chip-less simulators
(Sinabs/Rockpool/SpiNNaker) are declared but not wired.
Two tracks, both hardware-free:
- Track A — runtime & serving. Stateful inference, the deployment bundle, the encode-at-inference contract, a headless inference service, client SDKs, compression/quantization extensions, observability exporters, a governed registry, I/O adapters, and serving-level benchmarks.
-
Track B — simulator-backed test deploy. Keep
reference, Norse, and the Lava Loihi 2 CPU emulator; wire the SynSense and SpiNNaker software simulators; unify them into onetest-deploymatrix with an honest report.
Explicitly out of scope (and why): measured power/timing and on-device capacity (needs silicon); a full temporal SNN in ONNX (ONNX has no temporal spiking semantics by design); reproducing vendor compilers' numerics. These are recorded in Implications and boundaries and stay there.
| Capability | Current reality | Anchor | Gap to close |
|---|---|---|---|
| Closed-loop inference |
run(module, spikes, ...) consumes a whole [T, …] tensor |
runner.py |
no per-step API |
| Temporal loop | one loop, two modes; Trajectory holds S/U/I
|
execution.py, trajectory.py
|
loop body not exposed as a reusable step |
| Inference entry point |
infer_spikes returns UI payloads (rasters) |
network/inference.py |
no plain scoring API |
| Server surface |
/ws, /health, static mount |
server/app.py, server/web.py
|
no REST/streaming predict, auth, batching |
| Session | per-connection dashboard state | server/session.py |
not a model-serving session |
| Checkpoint | weights + meta (topology, topology_params) |
checkpoint_mixin.py, model_store.py
|
no single deploy bundle; encode config not a serving contract |
| Manifest | config hash, seed, versions, metric history | manifest.py |
reproducibility, not a runtime contract |
| NIR serialization | version-stamped graph envelope | serialization.py |
not bundled with weights/encoder |
| ONNX | single forward step + metadata | onnx_bridge/export.py |
not a temporal runtime (by design) |
| Encoding |
SpikeEncoder owns rate/latency/delta/random |
spike_encoder.py, encode_config.schema.json
|
encoding not frozen into the artifact |
| Targets | 6 declared; reference/norse/lava_loihi2 executable |
catalog.py, backends/api.py
|
speck/xylo/spinnaker2 not wired |
| Lava path | defaults to Loihi 2 CPU emulator, device opt-in | lava_backend.py |
document it; broaden lowering |
| Energy | declared tables, measured: false; measurement hook exists |
accounting.py |
nothing (honest by design) |
| Quantization | weight-only schemes |
quantize.py, quantize_schemes.py
|
no activation/membrane quant |
| Compression | none | — | no pruning / sparsity tooling |
| Benchmarks | training-step timing/memory + store/compare | benchmark/ |
no serving p50/p99/throughput |
| Observability | opt-in logs; in-process registry | observability/ |
no Prometheus/OTel exporters |
| Registry | file store + search/diff | model_store.py |
no stages/approvals/signing/lineage |
-
referencestays available unconditionally (catalog.py); every report staysestimate: truewhere applicable. - The closed-loop
run()/run_production()surfaces, the WebSocket protocol (protocol/), and all existing payload keys stay additive. - Core stays headless and dependency-light: no
fastapi/uvicorn/pydanticinspikeforge(enforced bycheck_core_boundary.py); SDK imports stay confined to isolated probes (backends/api.py). - House style: one class per file, ≤250-line files, ≤20-line functions, 79 columns (
rules.md).
The keystone. SNN inference is stateful across timesteps; today that state is internal to the closed loop, so no streaming service can drive it.
-
New package
spikeforge/serving/(core; torch-only). Note:spikeforge/runtime/already exists (device, execution_mode, system_stats), so the serving API must not reuse that name. -
API:
-
InferenceSession.load(bundle, device=..., mode=PRODUCTION)— build the module from the bundle's spec and load weights. -
.reset()— zero every carried neuron state. -
.step(frame) -> Prediction(logits, spikes, class_totals)— advance one timestep. -
.run_stream(frames) -> Iterator[Prediction]— drivestepover an iterable. -
.state()/.load_state(state)— JSON/tensor-serializable carried state.
-
-
Mechanism: extract the per-step body now inside
execution.pyinto a sharedstep_stages(module, inputs, state) -> (outputs, state)that both the closed loop andInferenceSessioncall, so they cannot diverge.Trajectoryremains unchanged for closed runs. -
State object: a
StateTree(one name -> tensor per stage) withreset(),to_dict()/from_dict()(numpy-tagged likearray_codec.py), and a device move. -
Acceptance:
session.run_streamover a full sequence produces readouts equal (within tolerance) torun(module, spikes);.state()round-trips; a reset between sequences reproduces a fresh run.
One portable artifact a runtime can load, replacing today's scattered
checkpoint + manifest + NIR envelope + ONNX set.
-
Format: a zip
model.spkfwith:-
manifest.json—TopologySpec, resolved stage params, library versions, expected metrics, label map, protocol/app version. -
weights.pt— thestate_dict. -
encode_config.json— the frozenEncodeConfig(encode_config.py). -
preprocessing.json— normalization/windowing (W2). -
graph.nir.json— optional NIR envelope (serialization.py). -
SHA256SUMS+ optional detached signature.
-
-
Builder:
spikeforge.serving.bundle.build(checkpoint, encode_config, ...)and a CLIspikeforge-serve bundle --from-model <name> --out model.spkf. -
Reuse:
ReproducibilityManifest(manifest.py) for provenance;TopologySpecfor rebuild; schema version pinned likecompatibility.json. - Acceptance: a bundle rebuilds the exact module + weights on a fresh process; loading rejects a tampered bundle with a typed error.
Prevents train/serve skew, the most common production failure for coded inputs.
- Promote encoding to a frozen, executable preprocessing spec stored in the
bundle (W1). The training path and serving path must call the same pure
function
spikeforge.serving.preprocess.encode(sample, spec) -> spikesbacked byspike_encoder.py. - Add
encode_config+preprocessingto the checkpoint's_meta()(checkpoint_mixin.py) so the artifact is self-describing; existing keys stay. - Also ship the label map and input geometry
(
image_size.py). - Acceptance: encoding the same sample at train time and serve time yields byte-identical spike tensors; a mismatched spec version is refused.
A headless inference service, separate from the dashboard server (which is
session/WebSocket oriented, server/session.py).
-
Distribution:
spikeforge-serve, import rootspikeforge_serve; depends on core (+ optionalspikeforge-targetsfor lowered execution).fastapi/uvicornare this distribution's deps, never core's. -
Endpoints:
-
POST /v1/predict— batch of pre-encoded or raw frames -> logits/label. -
POST /v1/reset— reset a session id's temporal state. -
GET|WS /v1/stream— streamingstepfor continuous sensors. -
GET /health,GET /metrics. -
GET /v1/bundle— loaded bundle metadata (spec, versions, metrics).
-
- Runtime semantics: session store keyed by id; bounded batching; concurrency limits; request timeouts; backpressure on the stream; optional bearer auth.
-
CLI:
spikeforge-serve --bundle model.spkf --host 0.0.0.0 --port 8899. -
Container: a minimal inference image (distinct from the dashboard
Dockerfile) + a compose profile. -
Acceptance:
predictmatches the in-process reference;streammaintains state across frames and resets on demand;/metricsexposes W6 metrics; auth rejects an unauthenticated call.
-
spikeforge-clients: a small Python client + a TypeScript client + a CLI, generated/validated against an OpenAPI schema so the service and clients cannot drift. Mirrors the existing protocol-codegen discipline (protocol/codegen/generate_ts.mjs). -
Acceptance: a round-trip test drives
predictandstreamfrom the Python client and asserts parity with the server.
-
Compression (
spikeforge/compression/): magnitude and structured pruning, sparsity schedule/report (weight density), and lossless export to the bundle; a before/after drift check reusingdrift.py. -
Quantization (
spikeforge-targets): add activation/membrane schemes + a calibration-dataset hook alongside the existing weight-only path (quantize_schemes.py); keep weight-only the default; report unapplied schemes rather than ignoring them. - Acceptance: pruning a topology then exporting/validating reports the sparsity gained and the induced drift; a new quantization scheme reports per-layer ranges and is refused honestly when unsupported.
-
Exporters: add Prometheus and OpenTelemetry exporters over the existing
MetricsRegistry(observability/registry.py); add request-level tracing tospikeforge-serve. Persistence stays opt-in (SPIKEFORGE_METRICS_PERSIST). -
Serving benchmarks: extend
benchmark/with a serving mode reporting p50/p99 latency, throughput at concurrency N, cold-start, and peak memory; feed the existing regression gate (compare.py). - Acceptance: a CI job asserts p99 latency and throughput stay within a configured threshold.
-
Registry: extend the model store + hub import with explicit stages
(
dev -> staging -> prod), approvals, artifact signing/verification, and lineage (dataset -> model -> bundle -> deployment). Build onmodel_search.py,model_diff.py, andspikeforge_hub/import_model.py. -
I/O (
spikeforge-io): windowing/normalization plus adapters for event cameras, audio, MQTT, and Kafka, feedingspikeforge-serve's stream. -
Acceptance: a bundle can be promoted with a recorded approver and verified
signature; an I/O adapter replays a recorded stream into
/v1/stream.
The repo already executes against CPU simulators; this track completes the matrix so "test deploy" is a first-class, documented experience.
| Target | Simulator (no chip) | Install | Status now | Work |
|---|---|---|---|---|
reference |
in-process NIR interpreter | — | done | — |
norse |
Norse pure-PyTorch | spikeforge-targets[norse] |
done | — |
lava_loihi2 |
Lava Loihi2SimCfg CPU emulator |
spikeforge-targets[lava] |
done (default) | document; broaden lowering past linear chains |
speck |
Sinabs / Speck simulator | vendor SDK | declared only | isolated probe + backend |
xylo |
Rockpool Xylo simulator | vendor SDK | declared only | isolated probe + backend |
spinnaker2 |
sPyNNaker / py-spinnaker2 host sim | vendor SDK | declared only | isolated probe + backend |
-
Design: each backend follows
backends/api.py(isolated import,available(),compile(),run()) and returns aBackendResultwith apath+ notes that name the emulator (pattern:lava_backend.py). -
Unified matrix: a
spikeforge-targets test-deploycommand that runs every available simulator, compares each toreferencewithbackends/compare.py, and emits a JSON report; a CI job runs it with whichever extras install. -
Honesty: an emulator run is never reported as a device measurement; energy
stays
estimate: trueuntil a device reports its own timing (accounting.py). -
Acceptance: with
[lava,norse]installed, one command test-deploys a linear topology onreference,norse, andlava_loihi2(CPU emulator) and reports parity; absent SDKs reportavailable: falsewith a named reason.
| Distribution | Import root | Depends on | Change |
|---|---|---|---|
spikeforge |
spikeforge |
torch stack only | add serving/, compression/, encode-at-inference, exporters |
spikeforge-targets |
spikeforge_targets |
core | add simulator backends, activation quant, test-deploy
|
spikeforge-serve |
spikeforge_serve |
core (+ targets) | new |
spikeforge-io |
spikeforge_io |
core | new |
spikeforge-clients |
spikeforge_clients |
core (schema only) | new |
spikeforge-registry |
spikeforge_registry |
core + hub | new (or extend hub) |
Rules carried over from ARCH-0001: core
never depends on a satellite; satellites pin spikeforge~=0.3.0 and record it in
compatibility.json; SDK imports stay in isolated probes;
everything degrades with a typed error.
| ID | Workstream | Deliverable | Depends on |
|---|---|---|---|
| W1 | Stateful runtime + bundle |
spikeforge/serving/ step API + .spkf bundle |
— |
| W2 | Encode-at-inference | frozen preprocessing spec + shared encode fn | W1 |
| W3 | spikeforge-serve |
REST + streaming service, auth, container | W1, W2 |
| W4 | Client SDKs | Python + TS + CLI, OpenAPI-validated | W3 |
| W5 | Compression + quantization | pruning/sparsity + activation schemes | W1 |
| W6 | Observability + serving benchmarks | Prometheus/OTel exporters + p50/p99 | W3 |
| W7 | Registry governance + I/O | stages/approvals/signing + adapters | W1 |
A user with no neuromorphic hardware can: train a topology, export a
DeploymentBundle, serve it with spikeforge-serve, call it from a client SDK,
read Prometheus metrics, and test-deploy the same model on reference,
norse, and the Lava CPU emulator with a parity report — all with the existing
evidence rules intact (estimate: true, available: false when an SDK is
absent, additive protocol only).