112 Commits

Author SHA1 Message Date
Dobromir Popov
f0f9a0eed7 controller: record DGR-043 completion 2026-08-01 01:58:14 +03:00
Dobromir Popov
e6ad9fdca9 story: DGR-043 Expose GGUF compatibility and measured cost inputs to existing routing 2026-08-01 01:58:13 +03:00
Dobromir Popov
d53acb1145 controller: record DGR-042 completion 2026-08-01 01:52:43 +03:00
Dobromir Popov
fd10607033 story: DGR-042 Carry native frames through direct and existing relay seams 2026-08-01 01:52:42 +03:00
Dobromir Popov
eb986ddf10 controller: record DGR-041 completion 2026-08-01 01:47:29 +03:00
Dobromir Popov
f37c4352fe story: DGR-041 Register native Shard capabilities without redesigning Meshnet 2026-08-01 01:47:28 +03:00
Dobromir Popov
95f005f646 controller: record DGR-040 completion 2026-08-01 01:41:52 +03:00
Dobromir Popov
520ccb8266 story: DGR-040 Add node-side native worker supervision 2026-08-01 01:41:51 +03:00
Dobromir Popov
f4980491d2 controller: record DGR-039 completion 2026-08-01 01:35:16 +03:00
Dobromir Popov
3a67eea569 story: DGR-039 Pass local two-process dense acceptance 2026-08-01 01:35:14 +03:00
Dobromir Popov
4c6c78d837 controller: record DGR-038 completion 2026-08-01 01:32:48 +03:00
Dobromir Popov
49560b396f story: DGR-038 Implement isolated shard-local Hot KV State 2026-08-01 01:32:46 +03:00
Dobromir Popov
a1df87deb6 controller: record DGR-037 completion 2026-08-01 01:28:07 +03:00
Dobromir Popov
8217b4c4a2 story: DGR-037 Bind llama.cpp to the standalone worker 2026-08-01 01:28:06 +03:00
Dobromir Popov
dfa403adc6 controller: record DGR-036 completion 2026-08-01 01:20:17 +03:00
Dobromir Popov
6e8bf7a64d story: DGR-036 Prove dense fixture and real-model range parity 2026-08-01 01:20:16 +03:00
Dobromir Popov
6e88b3bd8f controller: record DGR-035 completion 2026-08-01 01:13:52 +03:00
Dobromir Popov
64c2046e5a story: DGR-035 Implement dense architecture boundary input/output 2026-08-01 01:13:50 +03:00
Dobromir Popov
79c9bbaf63 controller: record DGR-034 completion 2026-08-01 01:08:29 +03:00
Dobromir Popov
d339cfde25 story: DGR-034 Implement dense-Llama range-aware GGUF ownership 2026-08-01 01:08:28 +03:00
Dobromir Popov
27a0d89678 docs: align PRD source-of-truth status 2026-07-27 08:58:51 +03:00
Dobromir Popov
8c87fae1ac chore: restore canonical PRD metadata and projections 2026-07-27 08:53:16 +03:00
Dobromir Popov
4d530d702c chore: restore canonical DGR-033 metadata projection 2026-07-26 23:04:45 +03:00
Dobromir Popov
7473bb7e44 fix: DGR-033 repair native worker protocol per cross-review BLOCK
Address the Codex GPT-5.5 review of the standalone fake C++ gRPC Shard
worker. Four root protocol defects fixed:

- Fail closed before SessionOpen: a per-session `opened` flag gates
  chunk/decode so no activation bypasses lifecycle, cancellation, epoch
  or flow-control state (terminal ERROR_CODE_INTERNAL), even when an
  out-of-band Cancel created placeholder state.
- Strict flow-control negotiation: NegotiateFlow takes the strictest of
  peer-vs-worker bounds (mirrors codec.negotiate_flow_control) and the
  negotiated per-session max_chunk_bytes is enforced on every bundle
  instead of trusting the peer proposal.
- In-stream ReleaseSignal now erases session state immediately.
- SessionOpen rejects incompatible schema, fingerprint, and shard-range
  identity and reports the worker's own served fingerprint rather than
  echoing the caller.

Adds 9 regression tests (worker suite 18 -> 27). Real gates on the
rebuilt pinned-gRPC binary: cmake build exit 0; ctest 2/2; worker
pytest 27 passed; harness+protocol 63 passed; compileall 0; diff --check
clean; ldd/nm show 0 llama/ggml linkage. DGR-033 passes -> true.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-26 22:57:03 +03:00
Dobromir Popov
c073826374 chore: reblock DGR-033 after protocol review 2026-07-26 22:34:34 +03:00
Dobromir Popov
0c7d475335 chore: remove Ralph lane runtime artifacts 2026-07-26 22:33:37 +03:00
Dobromir Popov
84d75f4cd2 controller: record DGR-033 completion 2026-07-25 22:38:01 +03:00
Dobromir Popov
766e480ba5 story: DGR-033 Build a standalone fake C++ gRPC Shard worker 2026-07-25 22:38:00 +03:00
Dobromir Popov
25e53bfeab story: DGR-032 Implement deterministic fake ShardEngine 2026-07-23 11:09:16 +03:00
Dobromir Popov
c34ab059cc story: DGR-031 Introduce the project-owned ShardEngine interface 2026-07-23 11:00:33 +03:00
Dobromir Popov
fd742d35c0 story: DGR-030 Add accelerator build presets and native CI matrix 2026-07-23 10:51:08 +03:00
Dobromir Popov
254297660a add CLAUDE.md with full milestone map and task explanations 2026-07-23 10:17:55 +03:00
Dobromir Popov
966aa10854 distributed-gguf-runtime: add CMake skeleton, gRPC harness, split-GGUF provisioning, performance contracts
DGR-019  Lock alpha/beta performance contracts (evidence + contract framework)
DGR-020  Run controlled whole-model GGUF baseline (benchmark results & contracts)
DGR-024  Real generated-gRPC protocol harness (shard_runtime_server.py + tests)
DGR-026  split-GGUF provisioning outside /home (provision script + manifest + tests)
DGR-028  Numbered patch-stack apply & verify (llama_cpp_dependency.py + UPSTREAM_LOCK.json)
DGR-029  Native CMake skeleton + deterministic CPU lane (UPSTREAM_LOCK.json + cmake gating)

New modules:
  packages/node/meshnet_node/dgr_performance/  — performance contract framework
  packages/node/meshnet_node/split_gguf/        — split-GGUF manifest & provisioning
  scripts/provision_split_gguf.py               — artifact provisioning CLI
  tests/test_dgr_performance_contract.py        — contract validation tests
  tests/test_split_gguf_manifest.py             — manifest tests
  tests/test_split_gguf_provision.py            — provisioning tests
  tests/test_shard_runtime_harness.py           — gRPC harness tests
2026-07-23 09:55:00 +03:00
Dobromir Popov
47bad0b7e1 backlog updated 2026-07-21 21:56:08 +03:00
Dobromir Popov
aa148cc7aa fix(vscode): use dynamic interpreter path in launch.json for cross-machine debug
Configs hardcoded .venv-rocm/bin/python (this Linux box's ROCm venv), which
doesn't exist on the Windows dev machine. Switch to
${command:python.interpreterPath} so each machine resolves whatever
interpreter is selected in the VS Code Python extension locally.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-21 14:09:45 +03:00
Dobromir Popov
505f37dd8d logs 2026-07-21 14:00:30 +03:00
Dobromir Popov
cd6b4d9d48 Merge DGR-024 (real generated-gRPC protocol harness) from ralph-terra-loop lane
# Conflicts:
#	.scratch/distributed-gguf-runtime/prd.json
2026-07-21 13:46:53 +03:00
Dobromir Popov
5177db25b0 feat: implement real generated-gRPC protocol harness (DGR-024)
Real ShardRuntimeServicer process bound to a real localhost socket, driven
by a generated ShardRuntimeStub over grpc.insecure_channel from a separately
spawned subprocess. Proves direct-hop and opaque-relay (exact captured
request bytes re-sent, no reinterpretation) produce byte-identical server
responses, cross-checked against an independent server-side wire capture.

Fails closed on the required negative paths: stale route epoch, expired
deadline, malformed/non-tiling fragments, checksum failure, exhausted
flow-control credit (with in-band top-up), duplicate idempotency steps
(acked, not re-applied), and cancel — both in-band CancelSignal (single
work item vs whole session) and the out-of-band unary Cancel RPC, including
a Cancel that races ahead of SessionOpen.

Supersedes the earlier in-memory fake-seam approach for this ticket, which
a policy audit rejected under the no-fake-data rule; that code is not
reintroduced. Evidence README rewritten to describe the actual files.

11 passed in tests/test_shard_runtime_harness.py.
2026-07-21 13:39:58 +03:00
Dobromir Popov
159284b3b5 Merge DGR-025 (certified-artifact-byte recipe identity) from ralph-fable-loop lane 2026-07-21 13:23:17 +03:00
Dobromir Popov
732ee9f91a Merge DGR-028 (numbered patch-stack apply/verify) from ralph-kimi-loop lane 2026-07-21 13:23:12 +03:00
Dobromir Popov
7da90ef475 feat: implement numbered patch-stack apply/verify enforcement (DGR-028)
Split the range-loader patch into single-concern patches 0002-0005 (loader,
filtered state report, boundary I/O endpoint guard, worker range-report
hook), add UPSTREAM-ASSUMPTIONS.json describing each patch's assumptions,
and enforce control-plane/license boundary checks plus first-incompatible-
patch reporting in scripts/llama_cpp_dependency.py apply/reverse/verify.

7 passed in tests/test_llama_cpp_dependency.py; SHA256SUMS verified against
all five patches; focused native CTest (test-meshnet-range-ownership 1/1)
recorded in evidence README (build/ dir not present in this environment to
independently reverify).
2026-07-21 13:22:55 +03:00
Dobromir Popov
03e97ca31a fix: bind recipe identity to certified artifact bytes (DGR-025)
Append +artifact.<sha256> to the llama.cpp runtime axis, computed from the
exact bytes read by attest_loaded_runtime, so a differently-built shared
object with copied lock values can no longer forge a certified runtime
identity. Node/tracker parsers require the suffix; new test proves a
byte-identical-lock but different-binary artifact produces a different
recipe fingerprint. Regenerates conformance vectors accordingly.

105 passed in tests/test_native_identity_emission.py,
tests/test_runtime_pin_identity.py, tests/test_runtime_recipe_identity.py.
2026-07-21 13:22:02 +03:00
Dobromir Popov
54d19f9a29 chore: replace fake protocol story with real harness 2026-07-19 00:22:03 +03:00
Dobromir Popov
377bc3475c chore: reconcile DGR-023 completion projection 2026-07-18 15:57:10 +03:00
Dobromir Popov
673830eac8 chore: reconcile DGR-023 completion projection 2026-07-18 15:56:59 +03:00
Dobromir Popov
902ecde363 [verified] feat: pin native protobuf and gRPC generation 2026-07-17 23:43:03 +03:00
Dobromir Popov
db59caa8e9 [verified] fix: enforce canonical native runtime pin 2026-07-17 23:20:03 +03:00
Dobromir Popov
ad66f7a4d8 feat: pin runtime identity to the exact llama.cpp patch stack (DGR-025)
Derive the recipe's runtime_version axis from the DGR-027 lock manifest
(exact upstream commit + ordered patch-stack byte digest) in new
meshnet_node.runtime_pin, failing closed on any lock/series/SHA256SUMS/
patch disagreement, and enforce pin discipline on runtime_version in both
the node and tracker identity implementations.

Also repair pre-existing backlog consistency: add missing DGR-022/DGR-027
completionNotes, regenerate the DGR-022/025/027 issue projections, and
relocate three pre-DGR legacy GLM issue files to issues/legacy/. Mark
DGR-025 passes=true with evidence at evidence/DGR-025/README.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 22:59:22 +03:00
Dobromir Popov
f83cf331c3 [verified] feat: harden llama.cpp provenance workspace 2026-07-17 16:24:46 +03:00
Dobromir Popov
ae51526e85 Merge remote-tracking branch 'origin/ralph/distributed-gguf-runtime' into ralph/distributed-gguf-runtime 2026-07-17 15:32:17 +03:00
Dobromir Popov
521a7b108a merge: close alternate DGR-001 maintenance history 2026-07-17 15:05:20 +03:00
Dobromir Popov
b7d40c5bcf fix: finish master branch integration compatibility 2026-07-17 14:35:05 +03:00
Dobromir Popov
8563d218c9 Merge branch 'ralph/distributed-gguf-runtime' of https://git.d-popov.com/popov/neuron-tai into ralph/distributed-gguf-runtime 2026-07-17 13:33:08 +02:00
Dobromir Popov
f0197bfa83 fix: reconcile merged runtime behavior 2026-07-17 14:29:09 +03:00
Dobromir Popov
6aced6a005 fix: reconcile legacy branch runtime with current GGUF 2026-07-17 14:21:02 +03:00
Dobromir Popov
66d9888a11 memory 2026-07-17 13:19:14 +02:00
Dobromir Popov
c758106a42 Merge branch 'archived_ralph/proxy-stream-cancellation' into merge/all-branches-into-master 2026-07-17 13:45:04 +03:00
Dobromir Popov
f0ddb69d33 Merge branch 'archived_ralph/dgr-001-performance-contract' into merge/all-branches-into-master
# Conflicts:
#	.claude/memory/MEMORY.md
#	.scratch/distributed-gguf-runtime/PRD.md
#	.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md
#	.scratch/distributed-gguf-runtime/README.md
#	.scratch/distributed-gguf-runtime/architecture.md
#	.scratch/distributed-gguf-runtime/evidence/DGR-017/README.md
#	.scratch/distributed-gguf-runtime/implementation-strategy.md
#	.scratch/distributed-gguf-runtime/issues/07-add-isolated-concurrent-local-hot-kv-state.md
#	.scratch/distributed-gguf-runtime/issues/13-harden-failure-cancellation-and-restart-semantics.md
#	.scratch/distributed-gguf-runtime/milestones.md
#	.scratch/distributed-gguf-runtime/prd.json
#	docs/issues/distributed-gguf-runtime/01-lock-the-safetensors-versus-gguf-performance-contract.md
#	docs/issues/distributed-gguf-runtime/02-adopt-the-versioned-grpc-shard-protocol.md
#	docs/issues/distributed-gguf-runtime/03-define-exact-artifact-and-runtime-recipe-identity.md
#	docs/issues/distributed-gguf-runtime/05-implement-dense-llama-range-aware-gguf-ownership.md
#	docs/issues/distributed-gguf-runtime/06-implement-architecture-defined-boundary-input-output.md
2026-07-17 13:44:52 +03:00
Dobromir Popov
9cb334cced Merge branch 'archived_ralph/deepseek-v4-flash-epic' into merge/all-branches-into-master 2026-07-17 13:44:12 +03:00
Dobromir Popov
ffe937678a Merge branch 'ralph/distributed-gguf-runtime' into merge/all-branches-into-master 2026-07-17 13:44:07 +03:00
Dobromir Popov
a35d86f343 chore: archive historical task programs 2026-07-17 12:41:46 +03:00
Dobromir Popov
9b257d9a1b feat: define shard lifecycle status contract 2026-07-17 11:58:50 +03:00
Dobromir Popov
3611b2cf9e Revert "fix: support headless Gitea credentials"
This reverts commit efd1cf4ef6.
2026-07-17 11:31:30 +03:00
Dobromir Popov
efd1cf4ef6 fix: support headless Gitea credentials 2026-07-17 10:47:15 +03:00
Dobromir Popov
ab466ce6b6 feat: add activation stream envelope 2026-07-17 02:36:24 +03:00
Dobromir Popov
989b55970b fix: reconcile Gitea Ralph status labels 2026-07-17 00:32:08 +03:00
Dobromir Popov
9e70b94417 feat: sync Ralph stories with Gitea issues 2026-07-16 23:37:03 +03:00
Dobromir Popov
9db036f91a feat: DGR-018 canonical Ralph metadata schema 2026-07-16 23:01:55 +03:00
Dobromir Popov
369b2072cc chore: clean superseded GGUF scaffolding 2026-07-16 22:32:37 +03:00
Dobromir Popov
81b1fa6074 docs: define implementation-ready distributed GGUF roadmap 2026-07-16 22:19:50 +03:00
Dobromir Popov
994546f78e Merge remote-tracking branch 'origin/master' into ralph/proxy-stream-cancellation 2026-07-16 18:06:54 +03:00
Dobromir Popov
6fd9d93e4b fix: flush direct SSE proxy frames before completion 2026-07-16 18:06:50 +03:00
Dobromir Popov
02b3709311 feat: checkpoint batching and release-gate stories 2026-07-16 17:24:56 +03:00
Dobromir Popov
737bade989 COLIBRI RESEARCH 2026-07-16 16:22:58 +02:00
Dobromir Popov
254627629b Merge commit '47b243cd98fd94da7918cacf5725373b099208e5' into ralph/distributed-gguf-runtime 2026-07-15 23:04:52 +02:00
Dobromir Popov
1fe31ef38d feat: checkpoint distributed gguf runtime stories 2026-07-15 23:42:58 +03:00
Dobromir Popov
47b243cd98 model loading, dash 2026-07-15 13:55:38 +02:00
Dobromir Popov
2852b1f80b loading more 2026-07-15 12:54:51 +02:00
Dobromir Popov
eaf00f6add test: record public relay smoke benchmark 2026-07-15 13:42:22 +03:00
Dobromir Popov
22f28bd69a fix model load/unload 2026-07-15 12:35:32 +02:00
Dobromir Popov
97e2784b37 node registration fixes 2026-07-15 10:34:41 +02:00
Dobromir Popov
c035bad5b7 feat: wire live benchmark CLI endpoints 2026-07-15 10:34:20 +03:00
Dobromir Popov
a508768e8a feat: add live endpoint benchmark runner 2026-07-14 22:46:11 +03:00
Dobromir Popov
e6f6782995 feat: add deterministic CPU/GPU benchmark runner slice 2026-07-14 21:39:13 +03:00
Dobromir Popov
ba7c656364 node metrics 2026-07-14 20:33:02 +02:00
Dobromir Popov
b661590ac7 log window bigger 2026-07-14 17:47:20 +02:00
Dobromir Popov
5b33bf8b99 feat: compare safetensors and gguf on cpu and gpu 2026-07-14 18:45:12 +03:00
Dobromir Popov
c7554ef7d8 feat: add DGR-001 performance contract 2026-07-14 18:13:54 +03:00
Dobromir Popov
21e6c86147 fix: let admin placement recover joined nodes 2026-07-14 16:37:42 +02:00
Dobromir Popov
def47f1a42 Merge branch 'master' of https://git.d-popov.com/popov/neuron-tai 2026-07-14 16:11:26 +02:00
Dobromir Popov
8cb00e951f feat: show admin node pool capacity 2026-07-14 16:11:18 +02:00
Dobromir Popov
7b3399760e chore: wrap up completed story metadata 2026-07-14 17:09:04 +03:00
Dobromir Popov
22467f145c merge: distributed performance baseline benchmark 2026-07-14 17:01:08 +03:00
Dobromir Popov
35af1e21de fix: make model placement controls observable 2026-07-14 16:00:37 +02:00
Dobromir Popov
905ea16ce0 feat: complete route session baseline benchmark 2026-07-14 16:55:52 +03:00
Dobromir Popov
348b003d6e fix: restore responsive dashboard panel grid 2026-07-14 15:55:24 +02:00
Dobromir Popov
1e64a5b2b9 new dash update 2026-07-14 15:29:11 +02:00
Dobromir Popov
f102be1098 docs: retarget gguf epic to DeepSeek-V4-Flash 2026-07-14 16:24:39 +03:00
Dobromir Popov
e2f3ae32b8 feat: let admins manage model placement 2026-07-14 15:16:23 +02:00
Dobromir Popov
29351d6217 chore: ignore local model cache 2026-07-14 14:05:37 +02:00
Dobromir Popov
1749f9b4ad chore: triage maintenance review and close completed stories 2026-07-14 14:33:09 +03:00
Dobromir Popov
5c9a2f6c97 dash style fix 2026-07-14 13:29:51 +02:00
Dobromir Popov
6516a92c04 feat: MAINT-002 - Update evidence READMEs for all completed stories 2026-07-14 14:23:27 +03:00
Dobromir Popov
4eeec7fa7f feat: MAINT-001 - Fix Ruff violations across all Python source 2026-07-14 14:17:23 +03:00
Dobromir Popov
13d82f8032 dash, tests 2026-07-14 12:26:10 +02:00
Dobromir Popov
d1a1400db9 Move tracker hive to admin and expand nodes panel.
Give Nodes & coverage full width on overview with inference prices and live speed, and expose model pricing on /v1/models.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-14 12:19:25 +02:00
Dobromir Popov
5d87e81bc9 feat: harden node placement and partial model loading 2026-07-13 21:58:08 +02:00
Dobromir Popov
a6bcc69288 sol mainnet payouts tasks 2026-07-13 18:51:40 +02:00
Dobromir Popov
c938d38031 more docs review 2026-07-13 18:37:07 +02:00
Dobromir Popov
95245be512 documentation revision 2026-07-13 18:14:21 +02:00
Dobromir Popov
180a7674e6 configure ralph 2026-07-13 17:56:00 +02:00
Dobromir Popov
f420dc1092 matt's skills updatged with upstream 2026-07-13 14:23:13 +02:00
388 changed files with 39446 additions and 4582 deletions

View File

@@ -1,6 +1,6 @@
---
name: ask-matt
description: Ask which skill or flow fits your situation. A router over the user-invoked skills in this repo.
description: Ask which skill or flow fits your situation. A router over the skills in this repo.
disable-model-invocation: true
---
@@ -8,26 +8,28 @@ disable-model-invocation: true
You don't remember every skill, so ask.
A **flow** is a path through the skills. Most paths run along one **main flow**, and two **on-ramps** merge onto it. Everything else is standalone.
A **flow** is a path through the skills. Most paths run along one **main flow**, and two **on-ramps** merge onto it. Everything else is standalone, or a vocabulary layer that runs underneath.
## The main flow: idea → ship
The route most work travels. You have an idea and want it built.
1. **`/grill-with-docs`** — sharpen the idea by interview. Start here when you **have a codebase**: it's stateful, retaining what it learns in `CONTEXT.md` and ADRs. (No codebase? Use `/grill-me` — see Standalone.)
1. **`/grill-with-docs`** — sharpen the idea by interview. Start here when you **have a codebase**: it's stateful, retaining what it learns in `CONTEXT.md` and ADRs. (No codebase? Use `/grill-me` — see Standalone. Both run the same `/grilling` primitive; `grill-with-docs` is the one that leaves a paper trail.)
2. **Branch — can you settle every question in conversation?** If a question needs a runnable answer (state, business logic, a UI you have to see), detour through a prototype, bridged by **`/handoff`** in both directions (see Crossing sessions):
- **`/handoff`** out, then open a fresh session against that file,
- **`/prototype`** to answer the question with throwaway code,
- **`/handoff`** back what you learned, and reference it from the original idea thread.
3. **Branch — is this a multi-session build?**
- **Yes** → **`/to-prd`** (turn the thread into a PRD) → **`/to-issues`** (split the PRD into independently-grabbable issues). Because the issues are independent, **clear context between each one**: start a fresh session per issue and kick off **`/implement`** by passing it the PRD and the single issue to work on.
- **Yes** → **`/to-spec`** (turn the thread into a spec), then **`/to-tickets`** to split it into tracer-bullet tickets, each declaring its **blocking edges**. On a local tracker that's one file per ticket under `.scratch/<feature>/issues/`, worked blockers-first by hand; on a real tracker the edges become native blocking links, so any ticket whose blockers are done can be grabbed — kick off **`/implement`** per ticket, **clearing context between each one**.
- **No** → **`/implement`** right here, in the same context window.
Either way, **`/implement`** builds each issue by driving **`/tdd`** internally — one red-green slice at a time — then closes out by running **`/code-review`**, a two-axis review (Standards + Spec) of the diff, before committing. Reach for **`/tdd`** on its own when you just want to build a concrete behaviour test-first without a full spec, and **`/code-review`** on its own whenever you want to review a branch or PR against a fixed point.
### Context hygiene
Keep steps 13 in **one unbroken context window** — don't compact or clear until after `/to-issues` — so the grilling, PRD, and issues all build on the same thinking. Each `/implement` then starts fresh, working from the issue.
Keep steps 13 in **one unbroken context window** — don't compact or clear until after `/to-tickets` — so the grilling, spec, and tickets all build on the same thinking. Each `/implement` then starts fresh, working from the ticket.
The limit on this is the **[smart zone](https://www.aihero.dev/ai-coding-dictionary/smart-zone)**: the window (~120k tokens on state-of-the-art models) within which the model still reasons sharply. If a session approaches it before `/to-issues`, don't push on degraded — `/handoff` and continue in a fresh thread.
The limit on this is the **[smart zone](https://www.aihero.dev/ai-coding-dictionary/smart-zone)**: the window (~120k tokens on state-of-the-art models) within which the model still reasons sharply. If a session approaches it before `/to-tickets`, don't push on degraded — `/handoff` and continue in a fresh thread.
## On-ramps
@@ -35,13 +37,26 @@ A starting situation that generates work, then merges onto the main flow.
- **Bugs and requests piling up** → **`/triage`**. It moves issues through triage roles and produces agent-ready issues, which **`/implement`** later picks up.
Triage is only for issues **you didn't create** — bug reports, incoming feature requests, anything that arrives raw. Issues that `/to-issues` produced are already agent-ready, so **don't triage them**.
Triage is only for issues **you didn't create** — bug reports, incoming feature requests, anything that arrives raw. Tickets that `/to-tickets` produced are already agent-ready, so **don't triage them**.
- **Something's broken** → **`/diagnosing-bugs`**. For the hard ones: the bug that resists a first glance, the intermittent flake, the regression that crept in between two known-good states. It refuses to theorise until it has a **tight feedback loop** — one command that already goes red on *this* bug — then fixes with a regression test. Its post-mortem hands off to **`/improve-codebase-architecture`** when the real finding is that there's no good seam to lock the bug down.
- **A huge, foggy effort — a greenfield project or a huge feature build, too big for one session** → **`/wayfinder`**, the most cognitively demanding flow here. When the way from here to the destination isn't visible yet, it charts a **shared map** of **decision tickets** on the issue tracker and resolves them one at a time — producing **decisions, not deliverables** — until the fog is pushed back and the way is clear. Where **`/grill-with-docs`** sharpens an idea you can hold in one session, wayfinder is for the idea you can't — and it's slower and denser, so save it for exactly that, never a well-scoped feature.
When the map clears, **it hands off, it doesn't build**: merge onto the main flow at **`/to-spec`**, which collapses the map's linked decisions into a buildable plan, then `/to-tickets` and `/implement` as usual. Looping the map straight into `/implement` skips that collapse and throws the linked detail away — go straight to `/implement` only when the effort turned out genuinely small.
## Codebase health
Not feature work — upkeep.
- **`/improve-codebase-architecture`** — run whenever you have a spare moment to keep the codebase good for agents to operate in. It surfaces deepening opportunities; picking one _generates an idea_ you can take into the main flow at `/grill-with-docs`.
- **`/improve-codebase-architecture`** — run whenever you have a spare moment to keep the codebase good for agents to operate in. It surfaces **deepening opportunities**; picking one _generates an idea_ you can take into the main flow at `/grill-with-docs`. It's the survey that finds the candidates; **`/codebase-design`** (below) is the bench you design the chosen one on.
## Vocabulary underneath
Two model-invoked references that run *beneath* the other skills — each the single source of truth for its vocabulary. Reach for them directly when the **words**, not the process, are the problem; or let the skills above pull them in.
- **`/domain-modeling`** — sharpen the project's *domain* language: challenge a fuzzy term, resolve an overloaded word ("account" doing three jobs), record a hard-to-reverse decision as an ADR. It's the active discipline `/grill-with-docs` drives to keep `CONTEXT.md` a clean glossary.
- **`/codebase-design`** — the deep-module vocabulary (module, interface, depth, seam, adapter, leverage, locality) for designing a module's *shape*: a lot of behaviour behind a small interface at a clean seam. `/tdd` and `/improve-codebase-architecture` both speak it.
## Crossing sessions
@@ -53,6 +68,8 @@ Not feature work — upkeep.
Off the main flow entirely.
- **`/grill-me`** — the same relentless interview as `/grill-with-docs`, but for when you have **no codebase**. Stateless: it saves nothing locally, builds no `CONTEXT.md`. Reach for it to sharpen any plan or design that doesn't live in a repo.
- **`/prototype`** — a small, throwaway program that answers one design question: does this state model feel right, or what should this UI look like. Throwaway from day one — keep the answer, delete the code. It's the detour in step 2 of the main flow, but reach for it any time a design question is hard to settle on paper.
- **`/research`** — delegate reading legwork to a **background agent**: it investigates a question against **primary sources**, then leaves a cited Markdown file in the repo. Keep working while it reads. The file it produces is something to take *into* the main flow at `/grill-with-docs` — research feeds the thinking, it doesn't replace it.
- **`/teach`** — learn a concept over multiple sessions, using the current directory as a stateful workspace.
- **`/writing-great-skills`** — reference for writing and editing skills well.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Ask Matt"
short_description: "Find the right skill or workflow"
policy:
allow_implicit_invocation: false

View File

@@ -1,10 +1,12 @@
---
name: grilling
description: Interview the user relentlessly about a plan or design. Use when the user wants to stress-test a plan before building, or uses any 'grill' trigger phrases.
description: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.
---
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
Interview me relentlessly about every aspect of this until we reach a shared understanding. Walk down each branch of the decision tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question before continuing. Asking multiple questions at once is bewildering.
If a question can be answered by exploring the codebase, explore the codebase instead.
If a *fact* can be found by exploring the environment (filesystem, tools, etc.), look it up rather than asking me. The *decisions*, though, are mine — put each one to me and wait for my answer.
Do not act on it until I confirm we have reached a shared understanding.

View File

@@ -0,0 +1,3 @@
interface:
display_name: "Grilling"
short_description: "Stress-test thinking one question at a time"

View File

@@ -9,7 +9,7 @@ Write a handoff document summarising the current conversation so a fresh agent c
Include a "suggested skills" section in the document, which suggests skills that the agent should invoke.
Do not duplicate content already captured in other artifacts (PRDs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
Do not duplicate content already captured in other artifacts (specs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
Redact any sensitive information, such as API keys, passwords, or personally identifiable information.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Handoff"
short_description: "Compact a conversation into a handoff"
policy:
allow_implicit_invocation: false

View File

@@ -1,15 +1,15 @@
---
name: implement
description: "Implement a piece of work based on a PRD or set of issues."
description: "Implement a piece of work based on a spec or set of tickets."
disable-model-invocation: true
---
Implement the work described by the user in the PRD or issues.
Implement the work described by the user in the spec or tickets.
Use /tdd where possible, at pre-agreed seams.
Run typechecking regularly, single test files regularly, and the full test suite once at the end.
Once done, use /review to review the work.
Once done, use /code-review to review the work.
Commit your work to the current branch.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Implement"
short_description: "Build work from a spec or tickets"
policy:
allow_implicit_invocation: false

View File

@@ -17,6 +17,11 @@ This command is _informed_ by the project's domain model and built on a shared d
### 1. Explore
**Scope before you scan — YAGNI.** Deepening a module pays off by making future changes to it easier, so put extra weight on the parts of the codebase that have recently changed. Decide *where* to look before you look:
- If the user named a direction — a module, a subsystem, a pain point — take it, and skip the inference below.
- Otherwise, walk back a good stretch of the commit history (`git log --oneline`) to find the codebase's hot spots — the files and areas that keep coming up — and let those paths pull your attention first. If the changes are scattered with no clear hot spot, widen the net.
Read the project's domain glossary (`CONTEXT.md`) and any ADRs in the area you're touching first.
Then use the Agent tool with `subagent_type=Explore` to walk the codebase. Don't follow rigid heuristics — explore organically and note where you experience friction:
@@ -56,7 +61,7 @@ Do NOT propose interfaces yet. After the file is written, ask the user: "Which o
### 3. Grilling loop
Once the user picks a candidate, run the `/grilling` skill to walk the design tree with them — constraints, dependencies, the shape of the deepened module, what sits behind the seam, what tests survive.
Once the user picks a candidate, run the `/grilling` skill to walk the decision tree with them — constraints, dependencies, the shape of the deepened module, what sits behind the seam, what tests survive.
Side effects happen inline as decisions crystallize — run the `/domain-modeling` skill to keep the domain model current as you go:

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Improve Codebase Architecture"
short_description: "Find and grill architecture improvements"
policy:
allow_implicit_invocation: false

View File

@@ -36,7 +36,7 @@ The right shape depends on the question:
Pick whichever shape best fits the question being asked, *not* whichever is easiest to wire to a TUI. Keep it pure: no I/O, no terminal code, no `console.log` for control flow. The TUI imports it and calls into it; nothing flows the other direction.
This is what makes the prototype useful past its own lifetime. When the question's been answered, the validated reducer / machine / function set can be lifted into the real module — the TUI shell gets deleted.
This is what makes the prototype useful past its own lifetime: when the question's been answered, the validated reducer / machine / function set can be lifted into the real module on its own.
### 4. Build the smallest TUI that exposes the state
@@ -66,9 +66,9 @@ If the host project has no task runner, just put the command at the top of the p
Give the user the run command. They'll drive it themselves; the interesting moments are when they say "wait, that shouldn't be possible" or "huh, I assumed X would be different" — those are the bugs in the _idea_, which is the whole point. If they want new actions added, add them. Prototypes evolve.
### 7. Capture the answer
### 7. Capture the answer and the prototype
When the prototype has done its job, the answer to the question is the only thing worth keeping. If the user is around, ask what it taught them. If not, leave a `NOTES.md` next to the prototype so the answer can be filled in (or filled in by you, if you've watched the session) before the prototype gets deleted.
Once the prototype has answered its question, capture the answer, then capture the prototype the way the [SKILL](SKILL.md) describes. The logic-specific mapping: the validated reducer / machine / function set lifts into the real module (the decision, absorbed); the TUI shell rides along to the throwaway branch that keeps the prototype as a primary source.
## Anti-patterns

View File

@@ -1,7 +1,6 @@
---
name: prototype
description: Build a throwaway prototype to flesh out a design — a runnable terminal app for state/business-logic questions, or several radically different UI variations toggleable from one route.
disable-model-invocation: true
description: Build a throwaway prototype to answer a design question. Use when the user wants to sanity-check whether a state model or logic feels right, or explore what a UI should look like.
---
# Prototype
@@ -22,10 +21,6 @@ The two branches produce very different artifacts — getting this wrong wastes
1. **Throwaway from day one, and clearly marked as such.** Locate the prototype code close to where it will actually be used (next to the module or page it's prototyping for) so context is obvious — but name it so a casual reader can see it's a prototype, not production. For throwaway UI routes, obey whatever routing convention the project already uses; don't invent a new top-level structure.
2. **One command to run.** Whatever the project's existing task runner supports — `pnpm <name>`, `python <path>`, `bun <path>`, etc. The user must be able to start it without thinking.
3. **No persistence by default.** State lives in memory. Persistence is the thing the prototype is _checking_, not something it should depend on. If the question explicitly involves a database, hit a scratch DB or a local file with a clear "PROTOTYPE — wipe me" name.
4. **Skip the polish.** No tests, no error handling beyond what makes the prototype _runnable_, no abstractions. The point is to learn something fast and then delete it.
4. **Skip the polish.** No tests, no error handling beyond what makes the prototype _runnable_, no abstractions. The point is to learn something fast.
5. **Surface the state.** After every action (logic) or on every variant switch (UI), print or render the full relevant state so the user can see what changed.
6. **Delete or absorb when done.** When the prototype has answered its question, either delete it or fold the validated decision into the real code — don't leave it rotting in the repo.
## When done
The _answer_ is the only thing worth keeping from a prototype. Capture it somewhere durable (commit message, ADR, issue, or a `NOTES.md` next to the prototype) along with the question it was answering. If the user is around, that capture is a quick conversation; if not, leave the placeholder so they (or you, on the next pass) can fill in the verdict before deleting the prototype.
6. **Capture it when done.** Fold any validated decision into the real code, then capture the prototype itself as a **primary source**: commit it to a throwaway branch, out of main, and leave a context pointer to that branch on the implementation issue. Capture the answer too — the verdict and the question it settled — in the issue or a commit. The main branch keeps only the validated decision.

View File

@@ -97,12 +97,12 @@ Surface the URL (and the `?variant=` keys). The user will flip through whenever
### 6. Capture the answer and clean up
Once a variant has won, write down which one and why (commit message, ADR, issue, or a `NOTES.md` next to the prototype if running AFK and the user hasn't responded yet). Then:
Once a variant has won, capture the answer — which variant and why — then capture the prototype the way the [SKILL](SKILL.md) describes. Fold the winner into the real code and move the rest onto the throwaway branch, not into main:
- **Sub-shape A** — delete the losing variants and the switcher; fold the winner into the existing page.
- **Sub-shape B** — promote the winning variant to a real route, delete the throwaway route and the switcher.
- **Sub-shape A** — fold the winner into the existing page; drop the losing variants and the switcher from main.
- **Sub-shape B** — promote the winning variant to a real route; drop the throwaway route and the switcher from main.
Don't leave variant components or the switcher lying around. They rot fast and confuse the next reader.
The full set of variants is the primary source, so it lands on the throwaway branch, not the bin — variant components and the switcher left in the main branch rot fast and confuse the next reader.
## Anti-patterns

View File

@@ -0,0 +1,3 @@
interface:
display_name: "Prototype"
short_description: "Prototype to answer a design question"

View File

@@ -26,16 +26,18 @@ Look at the current repo to understand its starting state. Read whatever exists;
- `docs/adr/` and any `src/*/docs/adr/` directories
- `docs/agents/` — does this skill's prior output already exist?
- `.scratch/` — sign that a local-markdown issue tracker convention is already in use
- Is the `triage` skill installed? (a `triage` skill folder alongside this one, or `triage` in your available skills.) This decides whether Section B runs at all.
- Monorepo signals — a `pnpm-workspace.yaml`, a `workspaces` field in `package.json`, or a populated `packages/*` with its own `src/`. Present only in a genuinely large multi-package repo; their absence means single-context, which is almost every repo.
### 2. Present findings and ask
Summarise what's present and what's missing. Then walk the user through the three decisions **one at a time** — present a section, get the user's answer, then move to the next. Don't dump all three at once.
Summarise what's present and what's missing. Then take the sections in order — one section, one answer, then the next.
Assume the user does not know what these terms mean. Each section starts with a short explainer (what it is, why these skills need it, what changes if they pick differently). Then show the choices and the default.
Lead each section with the recommended answer so the user can accept it in a word. Give a one-line explainer only when the choice genuinely branches; skip the section entirely when exploration already settled it (Section B when `triage` isn't installed, Section C when there's no monorepo).
**Section A — Issue tracker.**
> Explainer: The "issue tracker" is where issues live for this repo. Skills like `to-issues`, `triage`, `to-prd`, and `qa` read from and write to it — they need to know whether to call `gh issue create`, write a markdown file under `.scratch/`, or follow some other workflow you describe. Pick the place you actually track work for this repo.
> Explainer: The "issue tracker" is where issues live for this repo. Skills like `to-tickets`, `triage`, `to-spec`, and `qa` read from and write to it — they need to know whether to call `gh issue create`, write a markdown file under `.scratch/`, or follow some other workflow you describe. Pick the place you actually track work for this repo.
Default posture: these skills were designed for GitHub. If a `git remote` points at GitHub, propose that. If a `git remote` points at GitLab (`gitlab.com` or a self-hosted host), propose GitLab. Otherwise (or if the user prefers), offer:
@@ -44,41 +46,26 @@ Default posture: these skills were designed for GitHub. If a `git remote` points
- **Local markdown** — issues live as files under `.scratch/<feature>/` in this repo (good for solo projects or repos without a remote)
- **Other** (Jira, Linear, etc.) — ask the user to describe the workflow in one paragraph; the skill will record it as freeform prose
If — and only if — the user picked **GitHub** or **GitLab**, ask one follow-up:
Record the choice in `docs/agents/issue-tracker.md`. The GitHub and GitLab templates carry a "PRs as a request surface" flag, defaulted **off** — leave it off and don't raise it; a user who wants external PRs in the triage queue can flip the flag in the file later.
> Explainer: Open-source repos often receive feature requests as pull requests, not just issues — a PR is an issue with attached code. If you turn this on, `/triage` pulls *external* PRs into the same queue and runs them through the same labels and states as issues (collaborators' in-flight PRs are left alone). Leave it off if PRs aren't a request surface for you.
**Section B — Triage label vocabulary.** Skip this section entirely if the `triage` skill isn't installed (exploration told you) — an uninstalled skill needs no labels.
- **PRs as a request surface** — yes / no (default: no). Record the answer in `docs/agents/issue-tracker.md`. For local-markdown and other trackers, skip this question — there are no PRs.
If it is installed, ask exactly one question:
**Section B — Triage label vocabulary.**
> Do you want to keep the default triage labels? (recommended: **yes**)
> Explainer: When the `triage` skill processes an incoming issue, it moves it through a state machine — needs evaluation, waiting on reporter, ready for an AFK agent to pick up, ready for a human, or won't fix. To do that, it needs to apply labels (or the equivalent in your issue tracker) that match strings *you've actually configured*. If your repo already uses different label names (e.g. `bug:triage` instead of `needs-triage`), map them here so the skill applies the right ones instead of creating duplicates.
The defaults are the five canonical roles, each label string equal to its name: `needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`. On **yes**, write them as-is. Only if the user says no — usually because their tracker already uses other names (e.g. `bug:triage` for `needs-triage`) — collect the overrides so `triage` applies existing labels instead of creating duplicates.
The five canonical roles:
**Section C — Domain docs.** Default to **single-context** — one `CONTEXT.md` + `docs/adr/` at the repo root. This fits almost every repo; write it without asking.
- `needs-triage` — maintainer needs to evaluate
- `needs-info` — waiting on reporter
- `ready-for-agent` — fully specified, AFK-ready (an agent can pick it up with no human context)
- `ready-for-human` — needs human implementation
- `wontfix` — will not be actioned
Default: each role's string equals its name. Ask the user if they want to override any. If their issue tracker has no existing labels, the defaults are fine.
**Section C — Domain docs.**
> Explainer: Some skills (`improve-codebase-architecture`, `diagnosing-bugs`, `tdd`) read a `CONTEXT.md` file to learn the project's domain language, and `docs/adr/` for past architectural decisions. They need to know whether the repo has one global context or multiple (e.g. a monorepo with separate frontend/backend contexts) so they look in the right place.
Confirm the layout:
- **Single-context** — one `CONTEXT.md` + `docs/adr/` at the repo root. Most repos are this.
- **Multi-context** — `CONTEXT-MAP.md` at the root pointing to per-context `CONTEXT.md` files (typically a monorepo).
Offer **multi-context** — a root `CONTEXT-MAP.md` pointing to per-context `CONTEXT.md` files — only when exploration found monorepo signals. Then confirm which layout they want.
### 3. Confirm and edit
Show the user a draft of:
- The `## Agent skills` block to add to whichever of `CLAUDE.md` / `AGENTS.md` is being edited (see step 4 for selection rules)
- The contents of `docs/agents/issue-tracker.md`, `docs/agents/triage-labels.md`, `docs/agents/domain.md`
- The contents of `docs/agents/issue-tracker.md`, `docs/agents/domain.md`, and `docs/agents/triage-labels.md` (the last only when `triage` is installed)
Let them edit before writing.
@@ -101,7 +88,7 @@ The block:
### Issue tracker
[one-line summary of where issues are tracked, plus whether external PRs are a triage surface]. See `docs/agents/issue-tracker.md`.
[one-line summary of where issues are tracked]. See `docs/agents/issue-tracker.md`.
### Triage labels
@@ -112,12 +99,14 @@ The block:
[one-line summary of layout — "single-context" or "multi-context"]. See `docs/agents/domain.md`.
```
Then write the three docs files using the seed templates in this skill folder as a starting point:
Include the `### Triage labels` sub-block, and write `docs/agents/triage-labels.md`, only when `triage` is installed and Section B ran. When it isn't, both are omitted.
Then write the docs files using the seed templates in this skill folder as a starting point:
- [issue-tracker-github.md](./issue-tracker-github.md) — GitHub issue tracker
- [issue-tracker-gitlab.md](./issue-tracker-gitlab.md) — GitLab issue tracker
- [issue-tracker-local.md](./issue-tracker-local.md) — local-markdown issue tracker
- [triage-labels.md](./triage-labels.md) — label mapping
- [triage-labels.md](./triage-labels.md) — label mapping (only if `triage` is installed)
- [domain.md](./domain.md) — domain doc consumer rules + layout
For "other" issue trackers, write `docs/agents/issue-tracker.md` from scratch using the user's description.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Setup Matt Pocock Skills"
short_description: "Configure a repo for the skills"
policy:
allow_implicit_invocation: false

View File

@@ -32,3 +32,14 @@ Create a GitHub issue.
## When a skill says "fetch the relevant ticket"
Run `gh issue view <number> --comments`.
## Wayfinding operations
Used by `/wayfinder`. The **map** is a single issue with **child** issues as tickets.
- **Map**: a single issue labelled `wayfinder:map`, holding the Notes / Decisions-so-far / Fog body. `gh issue create --label wayfinder:map`.
- **Child ticket**: an issue linked to the map as a GitHub sub-issue (`gh api` on the sub-issues endpoint). Where sub-issues aren't enabled, add the child to a task list in the map body and put `Part of #<map>` at the top of the child body. Labels: `wayfinder:<type>` (`research`/`prototype`/`grilling`/`task`). Once claimed, the ticket is assigned to the driving dev.
- **Blocking**: GitHub's **native issue dependencies** — the canonical, UI-visible representation. Add an edge with `gh api --method POST repos/<owner>/<repo>/issues/<child>/dependencies/blocked_by -F issue_id=<blocker-db-id>`, where `<blocker-db-id>` is the blocker's numeric **database id** (`gh api repos/<owner>/<repo>/issues/<n> --jq .id`, _not_ the `#number` or `node_id`). GitHub reports `issue_dependencies_summary.blocked_by` (open blockers only — the live gate). Where dependencies aren't available, fall back to a `Blocked by: #<n>, #<n>` line at the top of the child body. A ticket is unblocked when every blocker is closed.
- **Frontier query**: list the map's open children (`gh issue list --state open`, scoped to the map's sub-issues / task list), drop any with an open blocker (`issue_dependencies_summary.blocked_by > 0`, or an open issue in the `Blocked by` line) or an assignee; first in map order wins.
- **Claim**: `gh issue edit <n> --add-assignee @me` — the session's first write.
- **Resolve**: `gh issue comment <n> --body "<answer>"`, then `gh issue close <n>`, then append a context pointer (gist + link) to the map's Decisions-so-far.

View File

@@ -33,3 +33,14 @@ Create a GitLab issue.
## When a skill says "fetch the relevant ticket"
Run `glab issue view <number> --comments`.
## Wayfinding operations
Used by `/wayfinder`. The **map** is a single issue with **child** issues as tickets.
- **Map**: a single issue labelled `wayfinder:map`, holding the Notes / Decisions-so-far / Fog body. `glab issue create --label wayfinder:map`. (On GitLab tiers with native epics, an epic may hold the map instead; a labelled issue works everywhere.)
- **Child ticket**: an issue carrying `Part of #<map>` at the top of its description and labels `wayfinder:<type>` (`research`/`prototype`/`grilling`/`task`). Once claimed, the ticket is assigned to the driving dev.
- **Blocking**: GitLab's **native blocking link** — the canonical, UI-visible representation. Add it with the `/blocked_by #<n>` quick action, posted as a note (`glab issue note <child> --message "/blocked_by #<blocker>"`). Native blocking links are a Premium/Ultimate feature; on the free tier (or where unavailable) fall back to a `Blocked by: #<n>, #<n>` line at the top of the description. A ticket is unblocked when every blocker is closed.
- **Frontier query**: `glab issue list -F json` scoped to the map's children, drop any with an open blocker — a native `blocked_by` link to an open issue (`glab api projects/:id/issues/:iid/links`), or an open issue in the `Blocked by` line — or an assignee; first in map order wins.
- **Claim**: `glab issue update <n> --assignee @me` — the session's first write.
- **Resolve**: `glab issue note <n> --message "<answer>"`, then `glab issue close <n>`, then append a context pointer (gist + link) to the map's Decisions-so-far.

View File

@@ -1,12 +1,12 @@
# Issue tracker: Local Markdown
Issues and PRDs for this repo live as markdown files in `.scratch/`.
Issues and specs (you may know a spec as a PRD) for this repo live as markdown files in `.scratch/`.
## Conventions
- One feature per directory: `.scratch/<feature-slug>/`
- The PRD is `.scratch/<feature-slug>/PRD.md`
- Implementation issues are `.scratch/<feature-slug>/issues/<NN>-<slug>.md`, numbered from `01`
- The spec is `.scratch/<feature-slug>/spec.md`
- Implementation issues are one file per ticket at `.scratch/<feature-slug>/issues/<NN>-<slug>.md`, numbered from `01` — never a single combined tickets file
- Triage state is recorded as a `Status:` line near the top of each issue file (see `triage-labels.md` for the role strings)
- Comments and conversation history append to the bottom of the file under a `## Comments` heading
@@ -17,3 +17,14 @@ Create a new file under `.scratch/<feature-slug>/` (creating the directory if ne
## When a skill says "fetch the relevant ticket"
Read the file at the referenced path. The user will normally pass the path or the issue number directly.
## Wayfinding operations
Used by `/wayfinder`. The **map** is a file with one **child** file per ticket.
- **Map**: `.scratch/<effort>/map.md` — the Notes / Decisions-so-far / Fog body.
- **Child ticket**: `.scratch/<effort>/issues/NN-<slug>.md`, numbered from `01`, with the question in the body. A `Type:` line records the ticket type (`research`/`prototype`/`grilling`/`task`); a `Status:` line records `claimed`/`resolved`.
- **Blocking**: a `Blocked by: NN, NN` line near the top. A ticket is unblocked when every file it lists is `resolved`.
- **Frontier**: scan `.scratch/<effort>/issues/` for files that are open, unblocked, and unclaimed; first by number wins.
- **Claim**: set `Status: claimed` and save before any work.
- **Resolve**: append the answer under an `## Answer` heading, set `Status: resolved`, then append a context pointer (gist + link) to the map's Decisions-so-far in `map.md`.

View File

@@ -0,0 +1,102 @@
---
name: setup-ts-deep-modules
description: Wire dependency-cruiser into a TypeScript repo so each package is a deep module — implementation hidden in subfolders, reachable only through its entry-point files. User-invoked.
disable-model-invocation: true
---
# Setup TS Deep Modules
Make every package in this repo a **deep module**: a lot of behaviour behind a small interface. A package's public surface is its **entry points** — the files at the package root — and everything in its subfolders is hidden. This skill installs [dependency-cruiser](https://github.com/sverweij/dependency-cruiser) and the rules that make the entry points the only way in, then proves the rules bite.
For the vocabulary (deep module, interface, seam, depth), run the `/codebase-design` skill — use its language throughout.
## The shape this enforces
```
src/packages/
<name>/
index.ts ← an entry point (public). Import this from outside.
client.ts ← another entry point. Packages may expose SEVERAL.
lib/ ← implementation: hidden from outside, free to import each other.
tests/ ← co-located tests + fixtures (a subfolder, so private).
```
The public surface is the package's **root files** — not one designated `index.ts`. By convention implementation lives in `lib/` and tests in `tests/`, giving every package the same two-folder shape. The rule itself is general, though: *anything* in *any* subfolder is private, so you never extend the config to add a folder.
Four rules, all `error`:
1. **Entry-point boundary** — code outside a package (app code or another package) may import only that package's entry points (its root files), never anything in its subfolders.
2. **Intra-package freedom** — a package's own files import each other freely.
3. **Tests through the entry points** — files under `<pkg>/tests/` may import any package's entry points and their own `tests/` fixtures, but never any package's subfolder internals (not even their own). Integration tests across packages are fine; deep imports are not.
4. **No cycles** — no dependency cycles.
**Entry points, not a barrel.** Because the public surface is *every* root file, a package can expose several small entry points (`index.ts`, `client.ts`, `server.ts`) instead of funnelling everything through one giant `index.ts`. Barrel files that re-export a whole subtree are discouraged — keep entry points small and hide implementation in subfolders.
Layering (which packages may depend on which) is a *different* concern and is left as a commented stub in the config for this repo to fill in.
## Steps
### 1. Detect the environment
- **Package manager** — `pnpm-lock.yaml` → pnpm, `yarn.lock` → yarn, `bun.lockb` → bun, else npm. Use it for every command below (`pnpm`/`yarn`/`npm run`/`bunx`).
- **Packages root** — if `src/` exists use `src/packages`, else `packages`. Confirm the choice with the user if the repo already has a different obvious convention.
- **Existing config** — check for a `.dependency-cruiser.*` file. If one exists, do **not** overwrite it: merge the four rules and the options in, and tell the user what you added.
**Done when:** package manager, packages root, and existing-config status are all known.
### 2. Install dependency-cruiser
Install `dependency-cruiser` as a devDependency with the detected package manager.
**Done when:** `dependency-cruiser` is in `devDependencies`.
### 3. Write the config
Copy [`dependency-cruiser.config.cjs`](./dependency-cruiser.config.cjs) to the repo root as `.dependency-cruiser.cjs`. Set `PACKAGES_ROOT` to the root detected in step 1. The rules are path-depth based and extension-agnostic, so nothing else needs adapting.
**Done when:** `.dependency-cruiser.cjs` exists with the correct `PACKAGES_ROOT`, and the four forbidden rules are present.
### 4. Wire it into the checks
- Add a `lint:boundaries` script: `depcruise <packages-root>` (or `depcruise src`).
- Fold it into the repo's umbrella check command — the one that already runs typecheck (e.g. a `check` / `ci` / `validate` script). Do **not** touch `tsconfig` or add path aliases.
- If there is no umbrella script, add `lint:boundaries` and tell the user to include it in CI.
**Done when:** `lint:boundaries` exists and runs as part of the same command as typecheck.
### 5. Scaffold the example package
Create a committed `<packages-root>/example/` as a copy-me template:
- `index.ts` — an entry point. Export one function that delegates to an internal file (so the package is visibly *deep*, not a pass-through).
- `lib/impl.ts` — an internal file in a **subfolder**, imported by `index.ts`, not reachable from outside.
- `tests/example.test.ts` — imports **only** `../index` (an entry point), and asserts against the public function.
Tell the user this is a starter template to copy or delete.
**Done when:** the example package exists, exposes its behaviour through a root entry point, and hides `impl` in a subfolder.
### 6. Prove the rules bite
This is the completion criterion for the whole skill — a config that doesn't fail on a violation is worthless.
1. Run `lint:boundaries`. It must **pass** on the clean example.
2. Temporarily add a deep import to `tests/example.test.ts` (e.g. `import { thing } from "../lib/impl"`). Run `lint:boundaries` again — it must **fail** with `tests-through-entrypoints`.
3. Revert the deep import. Run once more — it must **pass**.
**Done when:** you have observed a pass, then a fail on the deep import, then a pass again. If step 2 does not fail, the rules are not wired correctly — fix before finishing.
### 7. Document the convention
Write a `README.md` **in the packages folder** (`<packages-root>/README.md`) — next to the packages it governs — covering: the `src/packages/<name>/` layout (entry points at the root, `lib/` for implementation, `tests/` for tests), "import only through a package's entry points (its root files)", and how to run `lint:boundaries`. **Discourage barrel files** explicitly — expose several small entry points instead of re-exporting a whole subtree through one index. Keep it to the copy-me snippet plus the four rules in one paragraph each.
Then add a **context pointer** to it from the repo's agent-instructions file — `CLAUDE.md` if present, else `AGENTS.md` (create `AGENTS.md` if neither exists). One line is enough, e.g. `Packages are deep modules — see [src/packages/README.md](./src/packages/README.md) before adding or importing one.` This is what makes an agent discover the boundary rule instead of tripping over it.
**Done when:** `<packages-root>/README.md` exists and discourages barrels, and the repo's `CLAUDE.md`/`AGENTS.md` links to it.
## Notes
- The config's `$1` back-references (dependency-cruiser's group matching) are what let a package reach its own internals while outsiders can't — don't flatten them into separate per-package rules.
- Public vs private is decided by **depth**: a package's root files are entry points; anything in a subfolder is private. The conventional subfolders are `lib/` (implementation) and `tests/`, but the rule doesn't hardcode them — any subfolder is private, so a new folder never needs a config change. Adding an entry point is just adding a root file — no barrel.
- Packages are **flat**: one tier of immediate children under the root. A package's internals may nest as deep as you like; a package may not contain another package.
- Use `.cjs` (not `.js`) so the config's `module.exports` works even in `"type": "module"` repos.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Setup TS Deep Modules"
short_description: "Enforce deep TypeScript modules"
policy:
allow_implicit_invocation: false

View File

@@ -0,0 +1,95 @@
// @ts-check
// Deep-module enforcement for dependency-cruiser.
//
// Each package under the packages root is a DEEP MODULE: a lot of behaviour
// behind a small interface. A package's PUBLIC SURFACE is its ENTRY POINTS —
// the files at the package root. Implementation lives in SUBFOLDERS and is
// private — by convention `lib/` for implementation and `tests/` for tests,
// though any subfolder is private. A package may expose several small entry
// points (index.ts, client.ts, server.ts, …) — prefer that over one giant
// barrel index.
//
// The only thing you should ever need to edit here is PACKAGES_ROOT.
/** Where packages live. One immediate child dir per package (flat, no nesting). */
const PACKAGES_ROOT = "src/packages";
// --- derived patterns (no need to edit) -------------------------------------
const R = PACKAGES_ROOT;
/**
* A package's private internals: anything nested inside a package subfolder.
* The package's root files are its entry points and are NOT matched here —
* they stay importable from outside.
*/
const PACKAGE_INTERNALS = `^${R}/[^/]+/[^/]+/`;
/** @type {import('dependency-cruiser').IConfiguration} */
module.exports = {
forbidden: [
{
name: "entrypoint-boundary-from-app",
comment:
"App/root code may import a package's entry points (its root files), but nothing inside its subfolders.",
severity: "error",
from: { pathNot: `^${R}/` }, // importer is NOT inside any package
to: { path: PACKAGE_INTERNALS },
},
{
name: "entrypoint-boundary-across-packages",
comment:
"A package's own files import each other freely, but may reach OTHER packages only through their entry points — never their internals.",
severity: "error",
// importer is inside a package ($1), but is not a test file
from: { path: `^${R}/([^/]+)/`, pathNot: `^${R}/[^/]+/tests/` },
to: {
path: PACKAGE_INTERNALS,
pathNot: `^${R}/$1/`, // same package → intra-package freedom
},
},
{
name: "tests-through-entrypoints",
comment:
"A package's tests exercise it through its entry points like everyone else: they may import any package's entry points and their own tests/ fixtures, but never any package's internals — not even their own.",
severity: "error",
from: { path: `^${R}/([^/]+)/tests/` }, // a test file, in package $1
to: {
path: PACKAGE_INTERNALS,
pathNot: `^${R}/$1/tests/`, // own tests/ fixtures → allowed
},
},
{
name: "tests-folder-is-private",
comment:
"A package's tests/ folder is reachable only from tests — nothing else may import fixtures.",
severity: "error",
from: { pathNot: `^${R}/[^/]+/tests/` }, // importer is not itself a test
to: { path: `^${R}/[^/]+/tests/` },
},
{
name: "no-circular",
comment: "No dependency cycles. Scope to `^${R}/` if you want to allow cycles outside packages.",
severity: "error",
from: {},
to: { circular: true },
},
// --- Layering (optional, off by default) ----------------------------------
// Interface-hiding controls HOW you import (through the entry points).
// Layering controls WHICH packages may depend on which. Add your own rules
// here, e.g.:
//
// {
// name: "ui-may-not-depend-on-billing",
// severity: "error",
// from: { path: `^${R}/ui/` },
// to: { path: `^${R}/billing/` },
// },
],
options: {
doNotFollow: { path: "node_modules" },
tsConfig: { fileName: "tsconfig.json" },
enhancedResolveOptions: {
extensions: [".ts", ".tsx", ".js", ".jsx", ".json"],
},
},
};

View File

@@ -5,104 +5,32 @@ description: Test-driven development. Use when the user wants to build features
# Test-Driven Development
## Philosophy
TDD is the red → green loop. This skill is the reference that makes that loop produce tests worth keeping: what a good test is, where tests go, the anti-patterns, and the rules of the loop. Every section applies on every cycle — consult them before and during the loop, not after.
**Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
When exploring the codebase, read `CONTEXT.md` (if it exists) so test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching.
**Good tests** are integration-style: they exercise real code paths through public APIs. They describe _what_ the system does, not _how_ it does it. A good test reads like a specification - "user can checkout with valid cart" tells you exactly what capability exists. These tests survive refactors because they don't care about internal structure.
## What a good test is
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
Tests verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't. A good test reads like a specification — "user can checkout with valid cart" tells you exactly what capability exists — and survives refactors because it doesn't care about internal structure.
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
## Anti-Pattern: Horizontal Slices
## Seams — where tests go
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" - treating RED as "write all tests" and GREEN as "write all code."
A **seam** is the public boundary you test at: the interface where you observe behavior without reaching inside. Tests live at seams, never against internals.
This produces **crap tests**:
**Test only at pre-agreed seams.** Before writing any test, write down the seams under test and confirm them with the user. No test is written at an unconfirmed seam. You can't test everything — agreeing the seams up front is how testing effort lands on the critical paths and complex logic instead of every edge case.
- Tests written in bulk test _imagined_ behavior, not _actual_ behavior
- You end up testing the _shape_ of things (data structures, function signatures) rather than user-facing behavior
- Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
- You outrun your headlights, committing to test structure before understanding the implementation
Ask: "What's the public interface, and which seams should we test?"
**Correct approach**: Vertical slices via tracer bullets. One test → one implementation → repeat. Each test responds to what you learned from the previous cycle. Because you just wrote the code, you know exactly what behavior matters and how to verify it.
## Anti-patterns
```
WRONG (horizontal):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
- **Implementation-coupled** — mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed.
- **Tautological** — the assertion recomputes the expected value the way the code does (`expect(add(a, b)).toBe(a + b)`, a snapshot derived by hand the same way, a constant asserted equal to itself), so it passes by construction and can never disagree with the code. Expected values must come from an independent source of truth — a known-good literal, a worked example, the spec.
- **Horizontal slicing** — writing all tests first, then all implementation. Bulk tests verify _imagined_ behavior: you test the _shape_ of things rather than user-facing behavior, the tests go insensitive to real changes, and you commit to test structure before understanding the implementation. Work in **vertical slices** instead — one test → one implementation → repeat, each test a **tracer bullet** that responds to what the last cycle taught you.
RIGHT (vertical):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
## Rules of the loop
## Workflow
### 1. Planning
When exploring the codebase, read `CONTEXT.md` (if it exists) so that test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching.
Before writing any code:
- [ ] Confirm with user what interface changes are needed
- [ ] Confirm with user which behaviors to test (prioritize)
- [ ] Identify opportunities for deep modules (small interface, deep implementation) — run the `/codebase-design` skill for the vocabulary and the testability checks
- [ ] List the behaviors to test (not implementation steps)
- [ ] Get user approval on the plan
Ask: "What should the public interface look like? Which behaviors are most important to test?"
**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case.
### 2. Tracer Bullet
Write ONE test that confirms ONE thing about the system:
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
This is your tracer bullet - proves the path works end-to-end.
### 3. Incremental Loop
For each remaining behavior:
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
Rules:
- One test at a time
- Only enough code to pass current test
- Don't anticipate future tests
- Keep tests focused on observable behavior
### 4. Refactor
After all tests pass, look for [refactor candidates](refactoring.md):
- [ ] Extract duplication
- [ ] Deepen modules (move complexity behind simple interfaces)
- [ ] Apply SOLID principles where natural
- [ ] Consider what new code reveals about existing code
- [ ] Run tests after each refactor step
**Never refactor while RED.** Get to GREEN first.
## Checklist Per Cycle
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
- **Red before green.** Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features.
- **One slice at a time.** One seam, one test, one minimal implementation per cycle.
- **Refactoring is not part of the loop.** It belongs to the review stage (see the `code-review` skill), not the red → green implementation cycle.

View File

@@ -0,0 +1,3 @@
interface:
display_name: "TDD"
short_description: "Test-driven red-green-refactor"

View File

@@ -59,3 +59,19 @@ test("createUser makes user retrievable", async () => {
expect(retrieved.name).toBe("Alice");
});
```
**Tautological tests**: Expected value restates the implementation, so the test passes by construction.
```typescript
// BAD: Expected value is recomputed the way the code computes it
test("calculateTotal sums line items", () => {
const items = [{ price: 10 }, { price: 5 }];
const expected = items.reduce((sum, i) => sum + i.price, 0);
expect(calculateTotal(items)).toBe(expected);
});
// GOOD: Expected value is an independent, known literal
test("calculateTotal sums line items", () => {
expect(calculateTotal([{ price: 10 }, { price: 5 }])).toBe(15);
});
```

View File

@@ -0,0 +1,75 @@
---
name: to-spec
description: Turn the current conversation into a spec and publish it to the project issue tracker — no interview, just synthesis of what you've already discussed.
disable-model-invocation: true
---
This skill takes the current conversation context and codebase understanding and produces a spec (you may know this document as a PRD). Do NOT interview the user — just synthesize what you already know.
The issue tracker and triage label vocabulary should have been provided to you — run `/setup-matt-pocock-skills` if not.
## Process
1. Explore the repo to understand the current state of the codebase, if you haven't already. Use the project's domain glossary vocabulary throughout the spec, and respect any ADRs in the area you're touching.
2. Sketch out the seams at which you're going to test the feature. Existing seams should be preferred to new ones. Use the highest seam possible. If new seams are needed, propose them at the highest point you can. The fewer seams across the codebase, the better - the ideal number is one.
Check with the user that these seams match their expectations.
3. Write the spec using the template below, then publish it to the project issue tracker. Apply the `ready-for-agent` triage label - no need for additional triage.
<spec-template>
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list of user stories. Each user story should be in the format of:
1. As an <actor>, I want a <feature>, so that <benefit>
<user-story-example>
1. As a mobile bank customer, I want to see balance on my accounts, so that I can make better informed decisions about my spending
</user-story-example>
This list of user stories should be extremely extensive and cover all aspects of the feature.
## Implementation Decisions
A list of implementation decisions that were made. This can include:
- The modules that will be built/modified
- The interfaces of those modules that will be modified
- Technical clarifications from the developer
- Architectural decisions
- Schema changes
- API contracts
- Specific interactions
Do NOT include specific file paths or code snippets. They may end up being outdated very quickly.
Exception: if a prototype produced a snippet that encodes a decision more precisely than prose can (state machine, reducer, schema, type shape), inline it within the relevant decision and note briefly that it came from a prototype. Trim to the decision-rich parts — not a working demo, just the important bits.
## Testing Decisions
A list of testing decisions that were made. Include:
- A description of what makes a good test (only test external behavior, not implementation details)
- Which modules will be tested
- Prior art for the tests (i.e. similar types of tests in the codebase)
## Out of Scope
A description of the things that are out of scope for this spec.
## Further Notes
Any further notes about the feature.
</spec-template>

View File

@@ -0,0 +1,5 @@
interface:
display_name: "To Spec"
short_description: "Turn a conversation into a spec"
policy:
allow_implicit_invocation: false

View File

@@ -0,0 +1,107 @@
---
name: to-tickets
description: Break a plan, spec, or the current conversation into a set of tracer-bullet tickets, each declaring its blocking edges, published to the configured tracker — edges as text in one file per ticket locally, or native blocking links on a real tracker.
disable-model-invocation: true
---
# To Tickets
Break a plan, spec, or conversation into a set of **tickets** — tracer-bullet vertical slices, each declaring the tickets that **block** it.
The issue tracker and triage label vocabulary should have been provided to you — run `/setup-matt-pocock-skills` if not.
## Process
### 1. Gather context
Work from whatever is already in the conversation context. If the user passes a reference (a spec path, an issue number or URL) as an argument, fetch it and read its full body and comments.
### 2. Explore the codebase (optional)
If you have not already explored the codebase, do so to understand the current state of the code. Ticket titles and descriptions should use the project's domain glossary vocabulary, and respect ADRs in the area you're touching.
Look for opportunities to prefactor the code to make the implementation easier. "Make the change easy, then make the easy change."
### 3. Draft vertical slices
Break the work into **tracer bullet** tickets.
<vertical-slice-rules>
- Each slice cuts a narrow but COMPLETE path through every layer (schema, API, UI, tests) — vertical, NOT a horizontal slice of one layer
- A completed slice is demoable or verifiable on its own
- Each slice is sized to fit in a single fresh context window
- Any prefactoring should be done first
</vertical-slice-rules>
Give each ticket its **blocking edges** — the other tickets that must complete before it can start. A ticket with no blockers can start immediately.
**Wide refactors are the exception to vertical slicing.** A **wide refactor** is one mechanical change — rename a column, retype a shared symbol — whose **blast radius** fans across the whole codebase, so a single edit breaks thousands of call sites at once and no vertical slice can land green. Don't force it into a tracer bullet; sequence it as **expandcontract**. First expand: add the new form beside the old so nothing breaks. Then migrate the call sites over in batches sized by blast radius (per package, per directory), each batch its own ticket blocked by the expand, keeping CI green batch to batch because the old form still exists. Finally contract: delete the old form once no caller remains, in a ticket blocked by every migrate batch. When even the batches can't stay green alone, keep the sequence but let them share an integration branch that all block a final integrate-and-verify ticket — green is promised only there.
### 4. Quiz the user
Present the proposed breakdown as a numbered list. For each ticket, show:
- **Title**: short descriptive name
- **Blocked by**: which other tickets (if any) must complete first
- **What it delivers**: the end-to-end behaviour this ticket makes work
Ask the user:
- Does the granularity feel right? (too coarse / too fine)
- Are the blocking edges correct — does each ticket only depend on tickets that genuinely gate it?
- Should any tickets be merged or split further?
Iterate until the user approves the breakdown.
### 5. Publish the tickets to the configured tracker
Publish the approved tickets. **How** depends on the tracker `/setup-matt-pocock-skills` configured — the tickets are the same either way, only the shape of the blocking edges changes:
- **Local files** → write one file per ticket under `.scratch/<feature-slug>/issues/<NN>-<slug>.md`, numbered from `01` in dependency order (blockers first). Each file's "Blocked by" lists the numbers/titles it depends on. Use the per-ticket file template below — one ticket per file, never a single combined file.
- **A real issue tracker (GitHub, Linear, …)** → publish one issue per ticket in dependency order (blockers first) so each ticket's blocking edges can reference real identifiers. Use the platform's native blocking / sub-issue relationship where it has one; otherwise set each ticket's "Blocked by" to the blocking issues. Apply the `ready-for-agent` triage label unless instructed otherwise — the tickets are agent-grabbable by construction.
Work the **frontier**: any ticket whose blockers are all done. For a purely linear chain that means top to bottom.
Do NOT close or modify any parent issue.
<local-ticket-template>
# <NN> — <Ticket title>
**What to build:** the end-to-end behaviour this ticket makes work, from the user's perspective — not a layer-by-layer implementation list.
**Blocked by:** the numbers/titles of the tickets that gate this one, or "None — can start immediately".
**Status:** ready-for-agent
- [ ] Acceptance criterion 1
- [ ] Acceptance criterion 2
</local-ticket-template>
<issue-template>
## Parent
A reference to the parent issue on the tracker (if the source was an existing issue, otherwise omit this section).
## What to build
The end-to-end behaviour this ticket makes work, from the user's perspective — not layer-by-layer implementation.
## Acceptance criteria
- [ ] Criterion 1
- [ ] Criterion 2
## Blocked by
- A reference to each blocking ticket, or "None — can start immediately".
</issue-template>
In either form, avoid specific file paths or code snippets — they go stale fast. Exception: if a prototype produced a snippet that encodes a decision more precisely than prose can (state machine, reducer, schema, type shape), inline it and note briefly that it came from a prototype. Trim to the decision-rich parts — not a working demo, just the important bits.
Work the frontier one ticket at a time with `/implement`, clearing context between tickets.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "To Tickets"
short_description: "Split a plan into tracer-bullet tickets"
policy:
allow_implicit_invocation: false

View File

@@ -1,9 +1,16 @@
---
name: wayfinder
description: Plan a huge chunk of work — more than one agent session can hold — as a shared map of investigation tickets on your issue tracker, and resolve them one at a time until the way to the goal is clear.
description: Plan a huge chunk of work — more than one agent session can hold — as a shared map of decision tickets on your issue tracker, and resolve them one at a time until the way to the destination is clear.
disable-model-invocation: true
---
A loose idea has arrived — too big for one agent session, and wrapped in fog: the route from here to a plan isn't visible yet. This skill charts it as a **shared map** on the repo's issue tracker, then works its tickets one at a time. The map is domain-agnostic — engineering work, course content, whatever fits the shape.
A loose idea has arrived — too big for one agent session, and wrapped in fog: the way from here to the **destination** isn't visible yet. Wayfinding is about finding that way, not charging at the destination. This skill charts the way as a **shared map** on the repo's issue tracker, then works its **decision tickets** — questions whose resolution is a decision, not slices of a build to execute — one at a time until the route is clear.
The destination varies per effort, and naming it is the first act of charting — it shapes every ticket. It might be a spec to hand off and iterate on, a decision to lock before planning starts, or a change made in place like a data-structure migration. The map is domain-agnostic — engineering work, course content, whatever fits the shape.
## Plan, don't do
Wayfinder is **planning** by default: each ticket resolves a decision, and the map is done when the way is clear — nothing left to decide before someone goes and does the thing. The pull to just do the work is usually the signal you've reached the edge of the map and it's time to hand off. An effort can override this in its **Notes** — carrying execution into the map itself — but absent that, produce decisions, not deliverables.
## Refer by name
@@ -15,13 +22,17 @@ The map is a single issue on this repo's issue tracker, labelled `wayfinder:map`
The map is an **index**, not a store. It lists the decisions made and points at the tickets that hold their detail; a decision lives in exactly one place — its ticket — so the map never restates it, only gists it and links.
**Where the map, its child tickets, blocking, and frontier queries physically live is tracker-specific.** Consult `docs/agents/issue-tracker.md` (the "Wayfinding operations" section) for how _this_ repo expresses them. If that doc is absent, default to the local-markdown tracker.
**Where the map, its child tickets, blocking, and frontier queries physically live is tracker-specific.** The issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if not. Consult the tracker doc's "Wayfinding operations" section for how _this_ repo expresses them. If no tracker has been provided, default to the local-markdown tracker.
### The map body
The whole map at low resolution, loaded once per session. Open tickets are **not** listed — they are open child issues, found by query.
```markdown
## Destination
<what reaching the end of this map looks like — the spec, decision, or change this effort is finding its way to. One or two lines; every session orients to it before choosing a ticket.>
## Notes
<domain; skills every session should consult; standing preferences for this effort>
@@ -32,9 +43,13 @@ The whole map at low resolution, loaded once per session. Open tickets are **not
- [<closed ticket title>](link) — <one-line gist of the answer>
## Fog
## Not yet specified
<!-- see "Fog of war" for what belongs here -->
<!-- see "Fog of war": in-scope fog you can't ticket yet; graduates as the frontier advances -->
## Out of scope
<!-- see "Out of scope": work ruled beyond the destination; closed, never graduates -->
```
### Tickets
@@ -57,36 +72,48 @@ The answer isn't part of the body — it's recorded on resolution (see [Work thr
## Ticket Types
- **Research**: Reading documentation, third-party APIs, or local resources like knowledge bases. Creates a markdown summary as a linked asset. Use when knowledge outside the current working directory is required.
- **Prototype**: Raise the fidelity of the discussion by making a cheap, rough, concrete artifact to react to — an outline, a rough take, a stub, or UI/logic code via the /prototype skill. Links the prototype as an asset. Use when "how should it look" or "how should it behave" is the key question.
- **Grilling**: Conversation with the agent. Uses the /grilling and /domain-modeling skills. Asks one question at a time. The default case.
- **Task**: Literal manual work that must be done before the discussion can move forward — nothing to decide, prototype, or research. Moving data, signing up for a service, provisioning access. The agent automates it where it can; otherwise it hands the human a precise checklist. Resolved when the work is done; the answer records what was done and any resulting facts (credentials location, new URLs, row counts) later tickets depend on.
Every ticket is either **HITL** — human in the loop, worked *with* a human who speaks for themselves — or **AFK**, driven by the agent alone. A HITL ticket only resolves through that live exchange; the agent never stands in for the human's side of it (a grilling agent that answers its own questions has broken this).
- **Research** (AFK): Reading documentation, third-party APIs, or local resources like knowledge bases to surface a fact a decision waits on. Resolved by a `/research` **subagent**. Use when knowledge outside the current working directory is required.
- **Prototype** (HITL): Raise the fidelity of the discussion by making a cheap, rough, concrete artifact to react to — an outline, a rough take, a stub, or UI/logic code via the /prototype skill. Links the prototype as an asset. Use when "how should it look" or "how should it behave" is the key question.
- **Grilling** (HITL): Conversation via the /grilling and /domain-modeling skills, one question at a time. The default case.
- **Task** (HITL or AFK): Manual work that must happen before a *decision* can be made — nothing to decide, prototype, or research, but the discussion is blocked until it's done. Signing up for a service so its API can be judged, provisioning access, moving data so its shape can be seen. This is the one type that *does* rather than decides — and it earns its place by unblocking a decision, not by delivering the destination. The agent drives it alone where it can (AFK); otherwise it hands the human a precise checklist (HITL). Resolved when the work is done; the answer records what was done and any resulting facts (credentials location, new URLs, row counts) later tickets depend on.
## Fog of war
The map is _deliberately_ incomplete: don't chart what you can't yet see. Beyond the tickets lies fog — the dim view of decisions and investigations you can tell are coming but can't yet pin down, because they hang on questions still open. Resolving a ticket clears the fog ahead of it, graduating whatever's now specifiable into fresh tickets — one at a time, until the way to the goal is clear and no tickets remain.
The map is _deliberately_ incomplete: don't chart what you can't yet see. Beyond the live tickets lies the **fog of war** — the dim view of decisions and investigations you can tell are coming but can't yet pin down, because they hang on questions still open. Resolving a ticket clears the fog ahead of it, graduating whatever's now specifiable into fresh tickets — one at a time, until the way to the destination is clear and no tickets remain.
The map's **Fog** section is where that dim view is written down: the suspected question, the area to revisit later, the risk you're deferring. Write as loosely or as fully as the view allows; it doubles as a signpost for collaborators reading where the effort is headed.
The map's **Not yet specified** section is where that dim view is written down: the suspected question, the area to revisit later. It's the undiscovered frontier _toward_ the destination — everything here is in scope, just not sharp enough to ticket. Write as loosely or as fully as the view allows; it doubles as a signpost for collaborators reading where the effort is headed.
**Fog or ticket?** The test is whether you can state the question precisely now — _not_ whether you can answer it now.
- **Ticket when** the question is already sharp — even if it's blocked and you can't act on it yet.
- **Fog when** you can't yet phrase it that sharply. Don't pre-slice fog into ticket-sized pieces: it's coarser than a ticket, and one patch may graduate into several tickets, or none, once the frontier reaches it.
- **Not yet specified when** you can't yet phrase it that sharply. Don't pre-slice the fog into ticket-sized pieces: it's coarser than a ticket, and one patch may graduate into several tickets, or none, once the frontier reaches it.
Fog excludes only what's already decided (that's Decisions so far) and what's already a ticket.
**Not yet specified** excludes what's already decided (Decisions so far), what's already a live ticket, and what's out of scope (the next section).
## Out of scope
Fog only ever gathers _toward_ the destination. The destination fixes the scope, so work beyond it is **out of scope** — it isn't fog, and it doesn't belong in **Not yet specified**. It gets its own **Out of scope** section on the map: work you've consciously ruled out of _this_ effort. Scope, not sharpness, lands it here.
Out-of-scope work never graduates — the frontier stops at the destination — so it returns only if the destination is redrawn, and then as a fresh effort, not a resumption.
Ruling something out of scope is a scoping act, not a step on the route. When a ticket that already exists turns out to sit past the destination — mis-scoped in while charting, or exposed by a resolution — **close it** (a closed ticket is unambiguously off the frontier) and leave one line in the **Out of scope** section: the gist plus why it's out of scope, linking the closed ticket. It stays out of **Decisions so far**, which records the route actually walked — a scope boundary isn't a step on it.
## Invocation
Two modes. Either way, **never resolve more than one ticket per session.**
Two modes. Either way, **never resolve more than one ticket per session** — with the exception of research tickets.
### Chart the map
User invokes with a loose idea.
1. Run a `/grilling` and `/domain-modeling` session to surface the open decisions.
2. **Create the map** (label `wayfinder:map`): Notes filled in, Decisions-so-far empty, Fog sketched.
3. **Create the tickets you can specify now** as child issues of the map — then wire blocking edges in a **second pass** (issues need ids before they can reference each other). Wiring sorts them into the frontier and the blocked; everything you can't yet specify stays in the Fog.
4. Stop — charting the map is one session's work; do not also resolve tickets.
1. **Name the destination.** Run a `/grilling` and `/domain-modeling` session to pin down what this map is finding its way to — the spec, decision, or change. The destination fixes the scope, so it's settled first.
2. **Map the frontier.** Grill again, **breadth-first** this time: fan out across the whole space rather than deep on any one thread, surfacing the open decisions and the first steps takeable now. **If this surfaces no fog** — the way to the destination is already clear, the whole journey small enough for one session — you don't need a map. Stop and ask the user how they'd like to proceed.
3. **Create the map** (label `wayfinder:map`): Destination and Notes filled in, Decisions-so-far empty, the fog sketched into **Not yet specified**.
4. **Create the tickets you can specify now** as child issues of the map — then wire blocking edges in a **second pass** (issues need ids before they can reference each other). Wiring sorts them into the frontier and the blocked; everything you can't yet specify stays in the fog — the **Not yet specified** section.
5. **Fire the research subagents.** For each `research` ticket you just created, spin up a `/research` subagent to resolve it in parallel, capturing its findings on a throwaway `research/<name>` branch with a context pointer from the ticket.
6. Stop — charting is one session's work; it hand-resolves nothing.
### Work through the map
@@ -96,6 +123,6 @@ User invokes with a map (URL or number). A ticket is **optional** — without on
2. Choose the ticket. If the user named one, use it. Otherwise take the first frontier ticket in order. **Claim it**: assign it to yourself before any work.
3. Resolve it — **zoom as needed**: fetch the full body of any related or closed ticket on demand; invoke the skills the `## Notes` block names. If in doubt, use `/grilling` and `/domain-modeling`.
4. Record the resolution: post the answer as a **resolution comment**, **close** the issue, and **append a context pointer** to the map's Decisions-so-far.
5. Add newly-surfaced tickets (create-then-wire); graduate any fog the answer has made specifiable, clearing each graduated patch from the Fog so it lives only as its new ticket. If the decision invalidates other parts of the map, update or delete those tickets.
5. Add newly-surfaced tickets (create-then-wire); graduate any fog the answer has made specifiable, clearing each graduated patch from **Not yet specified** so it lives only as its new ticket. If the answer reveals a ticket — this one or another — sits beyond the destination, **rule it out of scope** rather than resolving it on the route. If the decision invalidates other parts of the map, update or delete those tickets.
The user may run unblocked tickets in parallel, so expect other sessions to be editing the tracker concurrently.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Wayfinder"
short_description: "Map a large effort as decision tickets"
policy:
allow_implicit_invocation: false

View File

@@ -158,6 +158,12 @@ _Failure mode._ Ending the current step before it is genuinely done, because the
_Avoid_: premature closure, the rush, rushing, shortcutting
### Negation
_Failure mode._ Steering by prohibition — telling the agent what _not_ to do — which drags the forbidden behaviour into context and makes it _more_ available, not less. _Don't think of an elephant_, and the elephant is all there is; _never write verbose comments_, and verbosity is the pattern the agent has just read. The negation is a weak modifier the strongly-activated concept overruns, so the ban half-reads as an instruction to do the thing. Its **leading word** is the _elephant_: whatever a prohibition names into the frame. Cure: prompt the **positive** — describe the target behaviour ("write one-line comments") so the banned one is never spoken. A prohibition earns its place only as a hard guardrail on a behaviour you cannot phrase positively; even then, pair it with the positive target so attention lands on what to do.
_Avoid_: ironic rebound, don't-prompting, the pink elephant
## Pruning
Keeping a skill lean — each remedy paired with the failure it cures.

View File

@@ -80,3 +80,4 @@ Use these to diagnose issues the user may be having with the skill.
- **Sediment** — stale layers that settle because adding feels safe and removing feels risky. The default fate of any skill without a pruning discipline.
- **Sprawl** — a skill simply too long, even when every line is live and unique. Hurts readability and maintainability and wastes tokens. The cure is the ladder: disclose **reference** behind pointers, and split by **branch** or sequence so each path carries only what it needs.
- **No-op** — a line the model already obeys by default, so you pay load to say nothing. The test: does it change behaviour versus the default? A weak leading word (_be thorough_ when the agent is already thorough-ish) is a no-op; the fix is a stronger word (_relentless_), not a different technique.
- **Negation** — steering by prohibition backfires: _don't think of an elephant_ names the elephant and makes it more available, not less. Prompt the **positive** — state the target behaviour so the banned one is never spoken; keep a prohibition only as a hard guardrail you can't phrase positively, and even then pair it with what to do instead.

View File

@@ -0,0 +1,5 @@
interface:
display_name: "Writing Great Skills"
short_description: "Principles for predictable skills"
policy:
allow_implicit_invocation: false

View File

@@ -2,11 +2,9 @@
- [Product selling points](product-selling-points.md) — key differentiators and landing page angles for neuron-tai
- [User profile](user-profile.md) — who Dobromir is and how to work with him
- [Project status](project-status.md) — 35/35 stories done; alpha hardening next
- [Project status](project-status.md) — US-001…US-035 done; US-036…US-050 in docs/prd.json; alpha hardening + scratch features next
- **Alpha hardening** — `.scratch/alpha-hardening/` (22 issues, ADRs 00160019, [README](../../.scratch/alpha-hardening/README.md), [handoff](../../.scratch/alpha-hardening/handoff.md))
- [Alpha hardening navigation](alpha-hardening-navigation.md) — locked fraud/auth decisions, Bucket-1 order, handoff pointers
- **Node capability admission** — `.scratch/node-capability-admission/` (P0 plan: generic doctor/real-forward validation, fail-closed readiness, tracker admission gate; [PRD](../../.scratch/node-capability-admission/PRD.md), [README](../../.scratch/node-capability-admission/README.md), ADR-0023)
- **Node capability admission** — `.scratch/node-capability-admission/` (P0 plan; [ADR-0023](../../docs/adr/0023-model-agnostic-node-capability-admission.md), [ADR-0026](../../docs/adr/0026-node-assignment-ownership-and-managed-placement.md))
- **Distributed relay performance** — relay `/rpc` requester sockets are persistent per Route Session and Activation Seam as of 2026-07-10; `request_id` remains unique per activation while `X-Meshnet-Session` remains stable for KV state. Next low-risk priorities: persistent direct/loopback HTTP, seam byte/latency telemetry, then trace-driven zstd tuning.
- **Distributed GGUF direction** — benchmark-gated native runtime: compare controlled Transformers/safetensors and whole-model llama.cpp lanes before expensive work; ship only for measured speed or model-fit advantage. Public parallelism is contiguous Shards in an Inference Route; concurrency comes from per-node continuous batching across isolated Route Sessions, while tensor/expert collectives stay inside optional trusted composite providers. Native data plane uses versioned Protobuf over long-lived gRPC/HTTP2 seam streams, with existing relay carrying the same opaque frames when needed. llama.cpp/GGML remains the substrate behind a project-owned standalone worker and small pinned fork; vLLM is an optional complete managed provider and concept donor, not a fork. Nakshatra, `prima.cpp`, `llama-gguf`, LiGGUF and historical GPUStack are source/test donors only. Active plan: [README](../../.scratch/distributed-gguf-runtime/README.md), [architecture](../../.scratch/distributed-gguf-runtime/architecture.md), [PRD](../../.scratch/distributed-gguf-runtime/PRD.md), [Ralph backlog](../../.scratch/distributed-gguf-runtime/prd.json). Research: [landscape](../../docs/research/distributed-gguf-landscape.md), [GitHub follow-up](../../docs/research/distributed-gguf-github-followup.md), [vLLM](../../docs/research/vllm-distributed-gguf-assessment.md).
- [DGR ROCm setup](dgr-rocm-setup.md) — version-matched TheRock SDK layout, relocated devel payload, verified `gfx1151` HIP llama.cpp build, and GPU-diagnostic boundary.
- **DGR-004 llama.cpp boundary** — `packages/node/native/llama/` locks `e920c523e3b8a0163fe498af5bf90df35ff51d25`, with a one-patch CMake marker and fail-closed clean materialize/apply/build/smoke harness. This is infrastructure only; stock GLM dense fallback remains uncertified.
- **Distributed GGUF direction** — benchmark-gated native runtime: compare controlled Transformers/safetensors and whole-model llama.cpp lanes before expensive work; ship only for measured speed or model-fit advantage. Public parallelism is contiguous Shards in an Inference Route; concurrency comes from per-node continuous batching across isolated Route Sessions, while tensor/expert collectives stay inside optional trusted composite providers. Native data plane uses versioned Protobuf over long-lived gRPC/HTTP2 seam streams, with existing relay carrying the same opaque frames when needed. llama.cpp/GGML remains the substrate behind a project-owned standalone worker and small pinned fork; vLLM is an optional complete managed provider and concept donor, not a fork. Nakshatra, `prima.cpp`, `llama-gguf`, LiGGUF and historical GPUStack are source/test donors only. Active plan: [README](../../.scratch/distributed-gguf-runtime/README.md), [architecture](../../.scratch/distributed-gguf-runtime/architecture.md), [PRD](../../.scratch/distributed-gguf-runtime/PRD.md), [Ralph backlog](../../.scratch/distributed-gguf-runtime/prd.json). ADR: [0024](../../docs/adr/0024-distributed-gguf-runtime.md). Research: [landscape](../../docs/research/distributed-gguf-landscape.md), [GitHub follow-up](../../docs/research/distributed-gguf-github-followup.md), [vLLM](../../docs/research/vllm-distributed-gguf-assessment.md).

View File

@@ -20,13 +20,13 @@ Active workstream (started 2026-07-04): alpha hardening of the money/trust path.
**Launch-readiness grilling (2026-07-06):** Locked launch plan — devnet dev/test run now, then **real mainnet SOL/USDT** (not devnet, not a new public token) for the first cohort: friends (API clients) + hired VPS/VPC hosts (our own test infra, not third-party volunteers — stake-free, risk-free if something breaks, not a long-term topology). Pricing: clients are the only party spending real money; nodes only accumulate off-chain credit and get paid in batches (30min dev / 24h later) — a failed distribution leaves funds parked, not lost, so mainnet-vs-devnet mixups are lower-risk than initially assumed. TAI token: do NOT issue/list now — ADR-0002 already locks listing behind $50k volume + 25 nodes/15 wallets plus an unresolved securities-review gate; only a dormant mainnet mint (cheap, ~few $ SOL) for name/branding reservation is in scope, bundled with treasury-key work, not before it. Treasury custody: bare keypair file (current runbook 02) is not acceptable for real funds — plan is **free native SPL multisig** (`spl-token create-multisig`, no protocol fee unlike Squads' 0.5 SOL), 2-of-3 signers, at least one cold/offline, others one-per-hired-VPS-provider to avoid correlated compromise (not yet built — ops task, no issue filed). Stake/slash asymmetry (registry/slash is a local Python adapter per ADR-0007, not on-chain) accepted for now since hired hosts are our own infra and friends aren't node operators — revisit before opening to real third-party node operators. A mainnet-vs-devnet boot guardrail was proposed and explicitly declined by the owner given the safe-by-default money flow above.
**Two new issues from this session, both `ready-for-agent`:**
- **21 — Honest-noise calibration corpus** (`.scratch/alpha-hardening/issues/21-honest-noise-calibration-corpus.md`) rescoped from "prod gate" to a **hard alpha-release blocker**. Confirmed by code read: `verify_activation_proofs()` (`packages/validator/meshnet_validator/audit.py:94-127`) returns bool only, no raw divergence value; fleet-dispatch exists but wrong shape (`server.py:2998-3104`, pinned routes + latency, not full-fleet + TOPLOC divergence); storage wrong shape (`registry_events` has no divergence/hardware columns). Three-part build: (1) surface raw TOPLOC distance from audit.py, (2) extend dispatch to hit every registered node with fixed prompt/seed, (3) new SQLite table keyed by node+GPU+dtype. Small-fleet exception granted (N = actual hired-VPS fleet size). Hired VPS hosts stay stake-free until this closes.
- **23 — Dynamic HF-benchmarked pricing** (`.scratch/alpha-hardening/issues/23-dynamic-hf-pricing.md`), high priority but not a release blocker. Pricing today is 100% static (`DEFAULT_PRICE_PER_1K_TOKENS = 0.02`, `billing.py:21`; `model_presets.json` has no per-model price). Target: 80% of cheapest comparable provider on `https://huggingface.co/inference/models` (per-provider-per-model marketplace, `?search=` query param works, no confirmed JSON API — plain scrape attempted first, escalate to headless browser only if the table isn't in raw HTML). Human-verified `hf_aliases` + `hf_verified_match_note` (params/quantization) per model, not auto-discovered matching. Reuses the `_settlement_loop` daemon-thread pattern for a daily refresh; falls back silently to the static default on any failure.
**Two new issues from this session:**
- **21 — Honest-noise calibration corpus** `Status: ready-for-human` (engineering done 2026-07-06; blocked on human fleet calibration run before mainnet launch).
- **23 — Dynamic HF-benchmarked pricing** `Status: done` (see `23-dynamic-hf-pricing_completed.md`).
Both are already migrated into `.scratch/alpha-hardening/prd.json` (AH-021 updated, AH-023 added) and the README index — ready for Ralph to pick up unattended.
**Ralph note:** `scripts/ralph_progress.py` tracks `docs/prd.json` (35/35 done) and does NOT see `.scratch/alpha-hardening/issues/`. No ralph loop is running and no `.ralph-tui/` state exists. `.scratch/alpha-hardening/prd.json` now has 23 stories (AH-001…AH-023); point Ralph at that file for the alpha-hardening branch. Do NOT use `ralph auto --parallel` on server.py-touching issues — 21 and 23 both touch `server.py`/`billing.py`/`audit.py`; if run in the same Ralph pass, run them serially, not in parallel (merge-conflict risk, same lesson as 03/04 previously).
**Ralph note:** `scripts/ralph_progress.py` tracks `docs/prd.json` (US-001…US-047; base 35/35 done, friends-test arc 3647 open/in-progress). Alpha hardening uses `.scratch/alpha-hardening/prd.json` (AH-001…AH-023). Point Ralph at the prd.json for the branch you're running.
**Why:** three audits agreed the alpha blockers are unauthenticated gossip (anyone can inject billing events), the free-credit faucet, and ephemeral bans.
**How to apply:** work test-first per issue acceptance criteria; use `.venv`; `cryptography` belongs in node deps (wallet.py imports it — causes many of the 24 "failures" in a fresh env). See [[project-status]] and [[autonomous-work-style]].

View File

@@ -6,7 +6,18 @@ metadata:
type: project
---
# Project Status (2026-07-02)
# Project Status (2026-07-13)
## Selected-node model placement (2026-07-14)
- Admin Model placement now opens a node selector for load and release; the control-plane accepts optional `node_id` and targets only that registry assignment. Multi-model serving remains supported through `ADD_SHARD` and `max_loaded_shards`.
- Total node pool resource values are rendered from `/v1/network/map`'s `node.capacity` contract. Route selection remains assignment/capability/throughput/queue based; capacity is used for placement and falls back to tracker defaults only if a node truly omits it.
## Distributed inference performance (2026-07-14)
`DIP-001` is done in `.scratch/distributed-inference-performance/`: the deterministic two-node Route Session stub benchmark covers direct/relay plus cached/stateless prefill and decode. Its JSON and concise summary explicitly attribute model execution, activation encode/decode, compression, connection setup, relay queueing, local HTTP forwarding, and end-to-end seam latency. `PYTHONPATH=packages/node pytest -q tests/test_route_session_benchmark.py` passed (7); the fixture assertion checks output-token identity and connection attempts.
> Doc reconciliation 2026-07-13: `docs/prd.json` tracks US-001…US-050 (048 memory budget, 049 mainnet pilot, 050 Qwen demand placement). ADRs 00250026 added (TAI phase B/C, assignment ownership).
All 35 user stories in docs/prd.json are done (35/35), including the reward-system arc US-030…US-035 completed 2026-07-02:
@@ -33,6 +44,10 @@ Historical handoff note: `/mnt/c/Users/popov/Downloads/neuron-tai-alpha-handoff-
Planning is ready at `.scratch/node-capability-admission/` with five sequential Ralph stories and ADR-0023. The design is model-agnostic: a Node must validate its selected Model Artifact/shard with a bounded real forward before Tracker routing; Qwen3.6 is only an optional development fixture. P0 adds a versioned local recipe-manifest/report contract, `meshnet-node doctor`, fail-closed startup admission, and tracker route gating. It intentionally excludes dynamic recipe/dependency installation and the future signed Node updater.
## Gitea DGR sync (2026-07-17)
Gitea is ahead of the local Markdown backlog with open DGR-022..DGR-071. The first executable P0 dependency frontier is DGR-022 (Shard lifecycle and structured status RPCs), DGR-023 (reproducible protobuf generation), DGR-025 (artifact/runtime recipe identity), and DGR-027 (llama.cpp provenance manifest). DGR-021, the named-tensor stream envelope prerequisite for DGR-022/023/025, is closed. DGR-022 is the next dependency-ordered issue and blocks DGR-024, DGR-033, and DGR-037.
## Windows CUDA node (working as of 2026-07-01)
- miniforge3 base env, torch 2.7.1+cu118, torchvision 0.22.x+cu118
- RTX 4060 Laptop GPU, 8 GB VRAM, benchmark index ~11,200

1405
.fuse_hidden0002bd66000001f0 Normal file

File diff suppressed because it is too large Load Diff

1521
.fuse_hidden0002bd66000001f9 Normal file

File diff suppressed because it is too large Load Diff

2
.gitignore vendored
View File

@@ -12,6 +12,7 @@ dist/
# Ralph local runtime state
.ralph-tui/*
!.ralph-tui/config.toml
.ralph-lane/
.env
@@ -20,6 +21,7 @@ dist/
!.env.testnet
.rocm-local/*
.pytest-tmp/*
.cache/
# Local tracker/node sqlite databases (never commit runtime state)
*.sqlite

5
.ralph-supervisor.log Normal file
View File

@@ -0,0 +1,5 @@
[2026-07-23 10:24:53] supervisor started, tailer pid=1460238
[2026-07-23 10:24:53] cycle 1: running ralph-tui resume (log starts at line 978)
[2026-07-23 10:25:59] ralph-tui exited without a recognized stop reason; retrying resume in 5 min
[2026-07-23 10:33:51] supervisor started, tailer pid=1465293
[2026-07-23 10:33:51] cycle 1: running ralph-tui run (log starts at line 1150)

1304
.ralph-tui-run.log Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -1,12 +1,2 @@
# Ralph TUI Configuration
# Generated by setup wizard
# See: ralph-tui config help
configVersion = "2.1"
tracker = "json"
agent = "opencode"
maxIterations = 0
autoCommit = true
[trackerOptions]
[agentOptions]
configVersion = "2.1"

File diff suppressed because it is too large Load Diff

View File

@@ -1,319 +1,57 @@
# Ralph execution context: Performant Concurrent Distributed GGUF Runtime
# Ralph context: Distributed GGUF Runtime
Status: authoritative context for every fresh Ralph iteration
Last updated: 2026-07-13
> **Specification status:** planning artifacts only. No distributed GGUF runtime is implemented. DGR-017 cleanup is complete; no runtime implementation story has completion credit. `prd.json` is authoritative.
## Mandatory startup sequence
## Mandatory startup for every fresh story
Before changing code, every Ralph agent must:
1. Read this file and authoritative `prd.json` completely.
2. Read the generated source issue named in the selected story description.
3. Read every dependency evidence README; legacy DGR-001..016 evidence is provenance only.
4. Read `docs/adr/0024-distributed-gguf-runtime.md`, root `CONTEXT.md`, `.claude/memory/MEMORY.md`, and relevant live source/tests.
5. Inspect `git status`; preserve unrelated work. Never infer implementation from planning text or old pass states.
6. If blocked or oversized, keep `passes: false` and write an honest `BLOCKED.md`/`DECOMPOSITION.md`; never weaken criteria or fabricate evidence.
1. Read this file completely.
2. Read the selected issue under `.scratch/distributed-gguf-runtime/issues/`.
3. Read `.scratch/distributed-gguf-runtime/GLM-5.2-MAX-ALPHA-ROADMAP.md`, `.scratch/distributed-gguf-runtime/ADR-0020-distributed-gguf-runtime.md`, and the relevant part of `architecture.md`.
4. Read `.claude/memory/MEMORY.md` and root `CONTEXT.md` for current project vocabulary and constraints.
5. Inspect the current implementation and tests; do not assume historical scratch text describes live code.
6. Read the evidence/handoff directories for every declared dependency.
7. Inspect `git status` and preserve all pre-existing working-tree changes.
## Locked scope
A fresh Ralph iteration has no conversational memory. These files are the context contract.
- Existing Meshnet Tracker routing, load balancing, billing, telemetry, relay, and provider semantics are backend-agnostic and are **not redesigned**. GGUF contributes exact compatibility, range/capacity, queue/load, seam-cost, health/reliability, and certification inputs only.
- The data plane is a standalone project-owned C++ Shard worker with gRPC/Protobuf and a project-owned `ShardEngine` boundary.
- llama.cpp is fetched at one exact commit into an ignored workspace from an in-repo manifest, then a numbered minimal patch stack is applied. There is no submodule, vendored tree, or permanent-fork dependency.
- llama.cpp owns DeepSeek V4 graphs, mHC, MoE, attention, hash routing, and kernels. Meshnet adds only range-ownership hooks, typed boundary/local-state adapters, worker integration, and parity/certification.
- Quantization and placement are dynamic recipe inputs. The 24 and 10+ stage layouts are certification scenarios, never product constants.
- Per-shard Hot KV and V4 CSA/HCA/SWA/indexer/compressor state remain local and keyed by route session/epoch. The WAN seam carries the typed mHC 4×4096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations—not per-layer caches.
- Route changes use cache miss plus re-prefill/restart. There is no WAN KV or V4 auxiliary-cache migration.
- CPU/CUDA/ROCm/Vulkan/Metal compile lanes are planned; only exact real-hardware-certified backend/model/recipe lanes may be advertised.
- Alpha requires correctness and the pre-locked useful-speed gate. MTP is reserved and off for alpha; its ownership contract, implementation, and benchmark are required before beta.
## Story sizing and interruption rule
## Target identities
Each story is intended to fit one focused Ralph context. Before implementation, estimate whether every acceptance criterion can be completed and verified in the current iteration.
- DeepSeek V4 official target SHA: `60d8d70770c6776ff598c94bb586a859a38244f1`.
- llama.cpp V4 support lineage began at PR 24162 / merge `8c146a8366304c871efc26057cc90370ccf58dad`; DGR-027 later pins one exact validated current commit.
- V4 scope: 43 main layers plus MTP; mHC 4×4096 boundary; 256 routed + 1 shared experts with six routed active; token IDs required for the first three hash-routed layers.
- Exact split-GGUF artifacts are provisioned to mounted-drive storage with a complete hashed manifest and resumable verification; no model artifact may be placed under `/home`.
If the story is too large, an external dependency is unavailable, or the context/provider limit prevents completion:
## Control/data-plane contract
- Do not weaken criteria.
- Do not mark the issue done or set `passes: true`.
- Avoid leaving an unverified cross-cutting partial implementation when a smaller safe spike is possible.
- Write `evidence/<TASK-ID>/DECOMPOSITION.md` or `BLOCKED.md` with the exact blocker, current verified state, proposed child stories, dependency graph and rollback/continuation instructions.
- Stop for supervised review.
Meshnet continues to own registration, coverage, existing route selection/load balancing, route epochs/sessions, direct/relay behavior, capability admission, cancellation, telemetry, billing, validation, and attribution. The GGUF adapter exposes measured inputs to those existing mechanisms. Direct seams use long-lived gRPC streams; relay seams carry byte-identical protobuf frames opaquely.
If interrupted after code changes, record every changed file, command result and unresolved invariant so the next fresh loop can verify rather than guess.
The project-owned `ShardEngine` hides llama.cpp internals. A worker loads one exact artifact/recipe/range identity. Default tests use fake/tiny fixtures. Real runs are opt-in, preserve raw metrics, and never download models under `/home`.
## Product objective
## Gitea issue synchronization
Build performant, concurrent distributed inference that combines consumer machines to serve top open models that exceed one node's RAM/VRAM.
Gitea is a projection of `prd.json`, never a competing source of truth. Before and after every supervised Ralph run, invoke:
The alpha target is the exact pinned GLM-5.2 `UD-IQ1_S` artifact served with `reasoning_effort=max` across physical consumer machines. Dense Llama is a structural fixture. Synthetic workers, dense-attention compatibility fallback, a smaller model, or a single host cannot satisfy target alpha. The immutable target contract and resource envelope are in `GLM-5.2-MAX-ALPHA-ROADMAP.md`.
A distributed demo is not success. The product must provide:
- Useful measured prefill and decode speed.
- Multiple concurrent Route Sessions.
- No KV/token cross-talk.
- Bounded memory, queues, cancellation and failures.
- Real execution on every participating node.
- A model-fit or performance advantage over the current Transformers/safetensors route.
## Critical-path architecture
```text
Existing Meshnet control plane
|
Versioned Protobuf over gRPC/HTTP2
|
Project-owned standalone C++ Shard worker
|
Small exact-commit llama.cpp patch stack
```bash
python3 scripts/ralph_gitea_sync.py sync
```
Meshnet remains the only control plane and owns:
For a complete Ralph invocation with automatic state reconciliation, use:
- Tracker registration, Coverage Map, route selection and route epochs.
- Route Sessions and Activation Seams.
- Direct/relay routing.
- Capability admission.
- Cancellation, Generation Telemetry and backpressure.
- Billing, validation and per-node work attribution.
Do not introduce another scheduler/control plane from vLLM, Nakshatra, prima.cpp, llama-gguf, GPUStack or another project.
## Runtime decisions that are not open
1. Public-network Shards are contiguous transformer layer ranges.
2. llama.cpp/GGML is the native GGUF execution substrate.
3. The project owns a small standalone worker and a narrow pinned llama.cpp patch stack.
4. The native Shard protocol is Protocol Buffers over gRPC/HTTP2.
5. One long-lived bidirectional stream serves one Route Session Activation Seam.
6. The public activation boundary is a versioned named-tensor bundle.
7. Hot KV State remains local to the node serving the Shard.
8. `(Route Session ID, route epoch)` maps to an isolated llama sequence or bounded context.
9. Concurrency uses continuous batching of compatible active sessions inside each node.
10. Transformers/safetensors remains the correctness and performance baseline.
11. vLLM may be an optional complete managed provider and concept donor; it is not forked into public Shards.
12. Tensor/expert collectives are deferred to a trusted composite provider, not public WAN routes.
13. Unsupported architectures/backends remain registered-but-dark until real certification passes.
14. Alpha failure retries from token zero; unverified KV is never migrated silently.
15. Model artifacts must remain on mounted-drive storage and never under `/home`.
16. Unified system RAM and integrated-GPU memory are one physical pool and must never be double-counted for admission.
17. Alpha requires native GLM-5.2 MoE, DSA, and IndexShare semantics; MTP/speculative decoding and 1M-context certification are post-alpha.
18. DGR-006 amends the decode fast path to carry a versioned `TensorBundle` and defines a typed tail logits/token result; the current single-`NamedTensor` fast path is insufficient for GLM sidebands.
19. Alpha reserves at least `max(20% of physical usable memory, 8 GiB)` per node outside weight-plus-Q8-KV placement and uses a same-switch wired 2.5 GbE minimum route.
Changing one of these requires an explicit ADR update and human review, not an incidental story implementation.
## Performance discipline
GGUF performance is a hypothesis. Never write “GGUF is faster” without measurements.
DGR-001 locks controlled benchmark lanes and thresholds. DGR-014 enforces the final distributed comparison.
Always distinguish:
- Weight quantization from activation/compute/KV dtype.
- Runtime/kernel gains from quantization/model-fit gains.
- Single-request latency from aggregate concurrency throughput.
- Synthetic unit coverage from real distributed acceptance.
Required metrics where applicable:
```text
TTFT
prefill tokens/sec
decode tokens/sec
aggregate throughput
p50/p95 latency
seam bytes and latency
queue and batch occupancy
RSS and VRAM
KV pressure
output-quality drift
failures and cleanup
```bash
scripts/ralph-gitea-run.sh ralph-tui run --prd .scratch/distributed-gguf-runtime/prd.json --agent claude --model sonnet --iterations 1 --no-tui --no-setup --direct-merge --no-sandbox
```
Do not weaken or move performance thresholds after seeing implementation results.
The sync creates/reconciles one Gitea issue per `DGR-*` story, creates missing labels/milestones, closes issues whose `passes` is true, marks the selected next eligible story `status:in-progress`, and marks blocked stories `status:blocked`. `gitea-issues.json` is a derived mapping only.
## Transport discipline
## Evidence and completion
Do not invent a raw TCP protocol, new WebSocket protocol, QUIC layer or bespoke binary control format.
The `.proto` schema is the semantic contract. Direct transport uses gRPC. Existing relay infrastructure may carry the same serialized protobuf frames as opaque binary.
Protocol requirements:
- Schema/version negotiation.
- Request/work ID.
- Route Session ID and route epoch.
- Exact Model Artifact/runtime recipe fingerprint.
- Shard range and effective overlap-safe start.
- Prefill/decode/release/cancel phases.
- Position/token range and idempotency step.
- Named tensors with shape, dtype, byte order and bounded fragments.
- Compression/checksum.
- Cache expectation/result.
- Deadlines, cancellation, flow control and structured status.
Avoid per-token channel creation and unbounded unary payloads. Generated code and build tooling must be reproducible; do not require manual copying.
## Native runtime discipline
Reuse llama.cpp for GGUF, mmap, kernels, architecture graphs, tokenizer, KV, sequences and heterogeneous backends.
The project patch stack is limited to:
- Range-aware tensor registration/loading.
- Endpoint-specific embedding/final head ownership.
- Architecture-defined intermediate input/output.
- Intermediate output before final norm/head.
- Layer-filtered KV and session mapping.
Do not place Meshnet routing, transport, billing or authentication inside llama.cpp. Keep patches numbered, scoped, pinned and upstreamable.
Dense Llama is first only as a cheap range/boundary fixture. GLM-5.2 is the explicit product adapter and alpha target immediately afterward. Qwen3/Qwen3-MoE is post-alpha. Do not generalize through unchecked tensor-name substitutions.
## Existing code seams to inspect first
- `packages/node/meshnet_node/model_backend.py` — backend abstraction.
- `packages/node/meshnet_node/torch_server.py` — reference ranged execution and session behavior.
- `packages/node/meshnet_node/activation_compression.py` — current activation framing/compression.
- `packages/node/meshnet_node/route_session_benchmark.py` — existing benchmark infrastructure.
- `packages/tracker/meshnet_tracker/server.py` — registration, route and proxy behavior.
- `packages/tracker/meshnet_tracker/capability.py` — fail-closed capability admission.
- `tests/test_real_model_backend.py` — real backend coverage.
- `tests/test_tracker_routing.py` — route/session behavior.
- `tests/test_tracker_capability_admission.py` — recipe admission.
- `tests/test_route_session_benchmark.py` and `tests/test_manual_route_benchmark.py` — benchmark patterns.
- `docs/adr/0008-binary-activation-wire-format.md` — existing wire compatibility.
- `docs/adr/0012-start-layer-overlapping-shards.md` — effective start semantics.
- `docs/adr/0022-sharded-per-node-kv-cache.md` — Hot KV State contract.
- `docs/adr/0023-model-agnostic-node-capability-admission.md` — certification/admission.
Do not edit generated `build/`, `__pycache__`, egg-info, Ralph logs or unrelated scratch features.
## Planned source layout
Use these paths unless current code inspection proves a better project-consistent location. If changed, document the reason in task evidence.
```text
packages/node/native/
proto/shard_runtime.proto
cmake/
llama/
UPSTREAM_COMMIT
patches/
gguf_worker/
tests/
packages/node/meshnet_node/
native_protocol/
gguf_backend.py
runtime_recipe.py
.scratch/distributed-gguf-runtime/evidence/<TASK-ID>/
README.md
commands.txt
results.json or other machine-readable evidence
```
Generated protobuf/C++ build outputs belong in build directories unless packaging explicitly requires checked-in generated Python modules. The story must document the generation command and version.
## Story output map
| Story | Required durable outputs |
|---|---|
| DGR-001 | benchmark harness/tests; `evidence/DGR-001/performance-contract.json`; raw/summary benchmark evidence |
| DGR-002 | `packages/node/native/proto/shard_runtime.proto`; reproducible Python/C++ generation/build wiring; protocol round-trip/compatibility tests; `evidence/DGR-002/` |
| DGR-003 | exact runtime-recipe/fingerprint implementation and admission tests; `evidence/DGR-003/` |
| DGR-004 | exact upstream pin, numbered patch series, reproducible fetch/apply/build smoke; `evidence/DGR-004/` |
| DGR-005 | dense-Llama range ownership loader and memory evidence; `evidence/DGR-005/` |
| DGR-006 | decode `TensorBundle` protocol amendment, typed tail-result contract, architecture boundary adapter/parity tests and results; `evidence/DGR-006/` |
| DGR-007 | concurrent session/KV manager, isolation/cleanup tests; `evidence/DGR-007/` |
| DGR-008 | standalone C++ gRPC worker, fake-model integration tests, lifecycle evidence; `evidence/DGR-008/` |
| DGR-009 | Meshnet backend/registration/relay integration and tests; `evidence/DGR-009/` |
| DGR-010 | real local two-process commands, raw metrics and parity report; `evidence/DGR-010/` |
| DGR-011 | two-machine configuration, commands, hardware/network manifest and raw results; `evidence/DGR-011/` |
| DGR-012 | continuous scheduler/admission implementation and 1/2/4/8 concurrency report; `evidence/DGR-012/` |
| DGR-013 | failure/cancel/restart test matrix and resource-cleanup evidence; `evidence/DGR-013/` |
| DGR-014 | immutable final comparison against DGR-001 thresholds and ship/stop recommendation; `evidence/DGR-014/` |
| DGR-015 | Qwen3-family adapter, architecture-specific parity/admission/performance evidence; `evidence/DGR-015/` |
| DGR-016 | narrow upstream patches/tests, design note and human-ready outreach package; `evidence/DGR-016/` |
| DGR-017 | exact GLM-5.2/GGUF target manifest, resource planner, immutable alpha contract and upstream status; `evidence/DGR-017/` |
| DGR-018 | verified whole-model `UD-IQ1_S` oracle with native GLM semantic evidence; `evidence/DGR-018/` |
| DGR-019 | explicit range-owned GLM MoE/MLA/DSA/IndexShare adapter, fixtures and parity; `evidence/DGR-019/` |
| DGR-020 | real multi-node GLM-5.2 Max target evidence and immutable `alpha`/`stop` verdict; `evidence/DGR-020/` |
## Dependency handoff rule
For every dependency listed by Ralph:
1. Confirm its `passes` state in `prd.json`.
2. Read `.scratch/distributed-gguf-runtime/evidence/<DEPENDENCY-ID>/README.md`.
3. Verify referenced source paths and commands still exist.
4. Do not repeat completed work unless verification exposes a concrete defect.
5. If dependency evidence is missing or contradictory, stop and repair the dependency instead of guessing.
## Testing and hardware rules
Default tests must be deterministic, GPU-free, model-download-free and API-credit-free.
Real model tests require:
```text
MESHNET_ENABLE_REAL_INFERENCE_TESTS=1
```
On this machine:
- Use `.venv-rocm` for real Radeon 8060S ROCm execution.
- The default Python 3.14 `.venv` is unsuitable for real ROCm inference.
- Resolve model storage through the machine-specific `.env.<hostname>` configuration.
- Never download model artifacts under `/home`.
- Real acceptance must exercise actual Tracker-routed CPU/GPU computation; synthetic workers are only unit tests.
Record exact:
- Model/revision and Artifact hash.
- Quantization and runtime recipe.
- Host/hardware/backend/driver.
- Commands and environment names without secrets.
- Raw output and metrics.
- Whether the evidence is synthetic, local-real, or multi-machine-real.
## Worktree and commit discipline
This repository may contain pre-existing changes from research or another feature.
- Inspect `git status` before editing.
- Never reset, checkout over, stash, delete or reformat unrelated changes.
- Stage only files belonging to the selected story.
- Exclude `.ralph-tui`, iteration logs, caches, generated builds, FUSE artifacts and unrelated scratch work.
- Keep one scoped commit per completed story when the supervising loop requests commits.
- Do not modify `passes` for another story.
## Mandatory finish/handoff sequence
Before emitting `<promise>COMPLETE</promise>`:
1. Verify every acceptance criterion with real command output or file evidence.
2. Run story-specific gates and repository quality gates.
3. Write `.scratch/distributed-gguf-runtime/evidence/<TASK-ID>/README.md` containing:
- Summary of changes.
- Exact files changed.
- Commands run and their real results.
- Performance/correctness evidence.
- Known limitations and deferred work.
- Compatibility or migration notes.
- Clear handoff for dependent stories.
4. Save machine-readable evidence beside it when the story produces metrics or schemas.
5. Update the source issue status to `done` only after all gates pass.
6. Preserve failures honestly. Never fabricate model, benchmark, test or hardware output.
## Authoritative references
Active decisions:
- `.scratch/distributed-gguf-runtime/README.md`
- `.scratch/distributed-gguf-runtime/implementation-strategy.md`
- `.scratch/distributed-gguf-runtime/architecture.md`
- `.scratch/distributed-gguf-runtime/ADR-0020-distributed-gguf-runtime.md`
- `.scratch/distributed-gguf-runtime/PRD.md`
- `.scratch/distributed-gguf-runtime/prd.json`
Source research:
- `docs/research/distributed-gguf-landscape.md`
- `docs/research/distributed-gguf-github-followup.md`
- `docs/research/vllm-distributed-gguf-assessment.md`
If historical notes conflict with these files, the active decisions above win.
Each story writes `/run/media/popov/d/DEV/repos/d-popov.com/AI/.claude/worktrees/distributed-gguf-runtime/.scratch/distributed-gguf-runtime/evidence/<DGR-ID>/README.md` with exact files, commands/results, limitations, identities, and dependent-story handoff. Only `prd.json` may record `passes`; DGR-017 and DGR-018 are complete and DGR-019 onward remain false. Generated Markdown and Gitea issues cannot override it. One scoped commit per story is expected during future execution.

View File

@@ -1,48 +1,32 @@
# Performant concurrent distributed GGUF runtime
# Distributed GGUF Runtime planning workspace
Status: active benchmark-gated implementation program.
> **Implementation status:** DGR-017 through DGR-033 have verified lane evidence, including a fixture-only standalone C++ gRPC worker. These lane checkpoints still require serialized integration and remote publication; they do not claim real model inference. `prd.json` is authoritative.
## Objective
Serve the exact pinned GLM-5.2 `UD-IQ1_S` artifact in `reasoning_effort=max` mode across consumer machines with useful measured performance. Dense Llama is a structural fixture; the real multi-node GLM target is the alpha release gate.
## Locked scope
See **[GLM-5.2 Max distributed alpha roadmap](GLM-5.2-MAX-ALPHA-ROADMAP.md)** for the target identity, minimum hardware, immutable acceptance matrix, and revised execution order. The 224-GiB figure is an experimental hard-fit floor; recommended topology is 5×64 GiB or 3×96/128 GiB after the required per-node reserve.
- Existing Meshnet Tracker routing, load balancing, billing, telemetry, relay, and provider semantics are backend-agnostic and are **not redesigned**. GGUF contributes exact compatibility, range/capacity, queue/load, seam-cost, health/reliability, and certification inputs only.
- The data plane is a standalone project-owned C++ Shard worker with gRPC/Protobuf and a project-owned `ShardEngine` boundary.
- llama.cpp is fetched at one exact commit into an ignored workspace from an in-repo manifest, then a numbered minimal patch stack is applied. There is no submodule, vendored tree, or permanent-fork dependency.
- llama.cpp owns DeepSeek V4 graphs, mHC, MoE, attention, hash routing, and kernels. Meshnet adds only range-ownership hooks, typed boundary/local-state adapters, worker integration, and parity/certification.
- Quantization and placement are dynamic recipe inputs. The 24 and 10+ stage layouts are certification scenarios, never product constants.
- Per-shard Hot KV and V4 CSA/HCA/SWA/indexer/compressor state remain local and keyed by route session/epoch. The WAN seam carries the typed mHC 4×4096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations—not per-layer caches.
- Route changes use cache miss plus re-prefill/restart. There is no WAN KV or V4 auxiliary-cache migration.
- CPU/CUDA/ROCm/Vulkan/Metal compile lanes are planned; only exact real-hardware-certified backend/model/recipe lanes may be advertised.
- Alpha requires correctness and the pre-locked useful-speed gate. MTP is reserved and off for alpha; its ownership contract, implementation, and benchmark are required before beta.
## Critical path
## Target identities
```text
Meshnet control plane
-> versioned gRPC/Protobuf Shard protocol
-> project-owned standalone C++ worker
-> small pinned llama.cpp patch stack
```
- DeepSeek V4 official target SHA: `60d8d70770c6776ff598c94bb586a859a38244f1`.
- llama.cpp V4 support lineage began at PR 24162 / merge `8c146a8366304c871efc26057cc90370ccf58dad`; DGR-027 later pins one exact validated current commit.
- V4 scope: 43 main layers plus MTP; mHC 4×4096 boundary; 256 routed + 1 shared experts with six routed active; token IDs required for the first three hash-routed layers.
- Exact split-GGUF artifacts are provisioned to mounted-drive storage with a complete hashed manifest and resumable verification; no model artifact may be placed under `/home`.
Transformers/safetensors remains the correctness baseline. vLLM remains an optional complete managed provider and a design donor; it is not forked into the public mesh.
## Navigation
## Planning artifacts
- **[Mandatory Ralph context](RALPH-CONTEXT.md)**read first in every fresh iteration
- [Task evidence contract](evidence/README.md)
- [Implementation strategy](implementation-strategy.md)
- [Current architecture](architecture.md)
- [PRD](PRD.md)
- [Ralph backlog](prd.json)
- [ADR-0020](ADR-0020-distributed-gguf-runtime.md)
- [Milestones](milestones.md)
- [Issues](issues/)
- [Distributed GGUF research](../../docs/research/distributed-gguf-landscape.md)
- [GitHub follow-up](../../docs/research/distributed-gguf-github-followup.md)
- [vLLM assessment](../../docs/research/vllm-distributed-gguf-assessment.md)
## Ralph execution
Use supervised one-story iterations for this high-risk runtime:
```bash
ralph-tui run \
--prd .scratch/distributed-gguf-runtime/prd.json \
--agent claude --model opus \
--iterations 1 --no-tui --no-setup --verify
```
Inspect the diff, run the story gates, and commit one verified story before the next iteration. Real-model stories require the explicit environment gate and mounted-drive model storage.
- [`prd.json`](prd.json) — sole authoritative 55-story backlog, DGR-017..071.
- [`PRD.md`](PRD.md) — human-readable projection of goals, gates, and all stories.
- [`RALPH-CONTEXT.md`](RALPH-CONTEXT.md) — mandatory fresh-session context.
- [`architecture.md`](architecture.md), [`implementation-strategy.md`](implementation-strategy.md), [`milestones.md`](milestones.md) — design and execution sequence.
- [`issues/`](issues/) — generated story specs; files 01..16 are retained legacy artifacts pending DGR-017.
- [`evidence/`](evidence/) — provenance and future per-story handoffs.

View File

@@ -1,264 +1,45 @@
# Performant Concurrent Distributed GGUF Architecture
# Distributed GGUF Runtime architecture
Status: current target architecture
Last updated: 2026-07-13
> **Specification status:** planning artifacts only. No distributed GGUF runtime is implemented. DGR-017 cleanup is complete; no runtime implementation story has completion credit. `prd.json` is authoritative.
## Product invariant
The system exists to serve high-quality models that exceed one consumer node's memory while retaining useful interactive speed and aggregate concurrency. A feature that only produces a distributed demo but is slower, globally serialized, or impossible to operate on consumer hardware is not complete.
## Locked scope
The alpha target is the exact pinned GLM-5.2 `UD-IQ1_S` artifact in `reasoning_effort=max` mode. Its target-specific architecture/resource/acceptance contract is [GLM-5.2-MAX-ALPHA-ROADMAP.md](GLM-5.2-MAX-ALPHA-ROADMAP.md). Dense Llama is a structural fixture, not the product target.
- Existing Meshnet Tracker routing, load balancing, billing, telemetry, relay, and provider semantics are backend-agnostic and are **not redesigned**. GGUF contributes exact compatibility, range/capacity, queue/load, seam-cost, health/reliability, and certification inputs only.
- The data plane is a standalone project-owned C++ Shard worker with gRPC/Protobuf and a project-owned `ShardEngine` boundary.
- llama.cpp is fetched at one exact commit into an ignored workspace from an in-repo manifest, then a numbered minimal patch stack is applied. There is no submodule, vendored tree, or permanent-fork dependency.
- llama.cpp owns DeepSeek V4 graphs, mHC, MoE, attention, hash routing, and kernels. Meshnet adds only range-ownership hooks, typed boundary/local-state adapters, worker integration, and parity/certification.
- Quantization and placement are dynamic recipe inputs. The 24 and 10+ stage layouts are certification scenarios, never product constants.
- Per-shard Hot KV and V4 CSA/HCA/SWA/indexer/compressor state remain local and keyed by route session/epoch. The WAN seam carries the typed mHC 4×4096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations—not per-layer caches.
- Route changes use cache miss plus re-prefill/restart. There is no WAN KV or V4 auxiliary-cache migration.
- CPU/CUDA/ROCm/Vulkan/Metal compile lanes are planned; only exact real-hardware-certified backend/model/recipe lanes may be advertised.
- Alpha requires correctness and the pre-locked useful-speed gate. MTP is reserved and off for alpha; its ownership contract, implementation, and benchmark are required before beta.
## Existing control plane
## Target identities
Meshnet remains the only public control plane:
- DeepSeek V4 official target SHA: `60d8d70770c6776ff598c94bb586a859a38244f1`.
- llama.cpp V4 support lineage began at PR 24162 / merge `8c146a8366304c871efc26057cc90370ccf58dad`; DGR-027 later pins one exact validated current commit.
- V4 scope: 43 main layers plus MTP; mHC 4×4096 boundary; 256 routed + 1 shared experts with six routed active; token IDs required for the first three hash-routed layers.
- Exact split-GGUF artifacts are provisioned to mounted-drive storage with a complete hashed manifest and resumable verification; no model artifact may be placed under `/home`.
- Tracker registration, Coverage Map, route scoring and assignment.
- Contiguous Shards and overlap-safe effective starts.
- Stable Route Sessions and route epochs.
- Local per-Shard Hot KV State in the reference backend.
- Direct/relay transport, cancellation and backpressure.
- Generation Telemetry, billing, validation and per-node attribution.
- Model-agnostic capability admission.
No external engine replaces these responsibilities.
## Runtime topology
## Topology
```text
OpenAI-compatible client
|
Gateway / Tracker Node
|
ordered Inference Route
|
+-- head Shard: tokenizer/embedding + early layers
| local weights and Hot KV State
|
+-- middle Shard(s): architecture boundary + owned layers
| local weights and Hot KV State
|
+-- tail Shard: final layers + norm/head/sampling
local weights and Hot KV State
existing Meshnet Tracker/control plane
-> existing backend-agnostic route/load-balancing decision
-> direct gRPC or existing opaque relay
-> project-owned standalone C++ Shard worker
-> project-owned ShardEngine
-> pinned upstream llama.cpp + numbered range/boundary/state hook patches
-> GGUF mmap, upstream V4 graph/kernels, local per-shard state
```
Weights never move in the per-request hot path. Every node opens and verifies its local Model Artifact before becoming routable.
A route is ordered contiguous half-open ranges. Head owns token embedding; tail owns final norm/head/sampling. Compatibility fingerprints bind source/split hashes, tokenizer, architecture adapter, typed boundary, runtime pin/patches, backend, quant, activation/compute/KV layout, range, and certification.
## Primary execution substrate
## V4 boundary and state
```text
project-owned C++ Shard worker
|
small exact-commit llama.cpp patch stack
|
GGUF mmap, quantized kernels, architecture graphs,
KV/sequence operations, CPU/CUDA/HIP/Vulkan/Metal backends
```
The inter-stage boundary is semantic and versioned: mHC 4×4096 residual, positions, token IDs only where the first three hash-routed layers require them, and cache/schema expectations. CSA/HCA/SWA/indexer/compressor/KV state belongs to upstream layer execution on the owning worker and is isolated by `(route_session_id, route_epoch)`. On loss, return cache miss and re-prefill/restart. Never serialize those caches into the WAN bundle.
The patch stack adds only the missing local execution seam:
## Concurrency, failure, and admission
1. Range-aware tensor registration/loading.
2. Endpoint-specific embedding and final head ownership.
3. Architecture-defined intermediate input.
4. Architecture-defined pre-tail boundary output.
5. Layer-filtered KV and external session mapping.
The worker owns protocol translation and process lifecycle. llama.cpp never receives Tracker, relay, billing or volunteer-network code.
## Shard data plane
Use Protocol Buffers and gRPC over HTTP/2.
### Service shape
- Unary capability and health.
- Bidirectional Route Session stream.
- Explicit release and cancellation.
- Metrics suitable for capability admission and route scoring.
### Session stream
One long-lived stream represents one Route Session Activation Seam. It amortizes connection setup and inherits HTTP/2 flow control. Every message carries enough identity to reject stale or incompatible work.
```text
schema version
request/work id
Route Session id
route epoch
Model Artifact hash
runtime recipe fingerprint
Shard begin/end and effective start
prefill/decode/release/cancel phase
position and token range
idempotency step id
cache expectation/result
named tensor bundle
compression/checksum
```
Prefill tensors are split into bounded ordered frames. Decode messages carry one-step architecture boundary bundles and remain small. DGR-006 amends the current v1 decode fast path—which carries only one `NamedTensor`—to carry a versioned `TensorBundle`, while preserving compact one-tensor encoding and explicit compatibility behavior.
Tail completion is not inferred from an activation tensor name. The protocol exposes a typed logits and/or sampled-token result, and exact sampling parameters plus chat-template/reasoning mode are bound to request/runtime identity.
Direct nodes use gRPC. Nodes requiring the existing relay carry the same protobuf frames as opaque binary through the relay session. This preserves one semantic protocol instead of maintaining separate direct and relay payload contracts.
## Architecture boundary
The public boundary is a versioned named-tensor bundle:
```text
bundle schema/version
architecture adapter and boundary point
named tensors
per-tensor shape, dtype and byte order
payload fragments
compression/checksum
```
Dense Llama may use one residual tensor. Other adapters may require more. vLLM's Llama and Qwen3-MoE PP paths demonstrate a boundary with both `hidden_states` and `residual`; therefore the generic protocol must not assume one anonymous tensor.
GLM-5.2 normally exchanges a 6,144-element hidden state. If a memory-balanced Shard boundary splits an IndexShare Full producer from Shared consumers, the bundle also carries the typed top-k index sideband. The planner prefers boundaries that keep an IndexShare ownership group local, but the protocol validates the sideband rather than assuming it never crosses a seam.
Only the head owns token embedding. Only the tail owns final normalization, LM head and sampling. Middle Shards exchange the architecture-defined pre-tail boundary, not final normalized embeddings.
## Hot KV State and concurrency
```text
(Route Session id, route epoch)
-> local llama sequence or bounded context
-> KV for owned layers only
-> lease, memory accounting and lifecycle
```
Required operations:
- Prefill append.
- Decode append.
- Truncate after rejected speculative positions if later enabled.
- Explicit release.
- TTL/LRU eviction.
- Cache-miss response.
- Stale-epoch rejection.
A node must not clear global KV on a new stream or serialize all requests behind one logical serving sequence.
## Continuous batching
Autoregressive dependencies remain sequential inside one Route Session. Aggregate throughput comes from batching compatible decode steps across active sessions:
```text
time 0: session A token 1 + session B token 8 + session C token 3
-> one llama batch for this Shard
time 1: next ready positions from active sessions
-> next llama batch
```
The node scheduler:
- Admits work against weight, KV, scratch and queue budgets.
- Keeps per-session token positions and outputs separate.
- Prevents long prefill from starving decode.
- Applies bounded backpressure.
- Reports active sessions, queue depth, batch occupancy, KV pressure and throughput.
The initial deterministic gate is four concurrent sessions on a small model without cross-talk. Hardware-specific limits are measured and advertised through capability admission.
## Parallelism boundaries
| Mechanism | First-runtime use |
|---|---|
| Layer/pipeline parallelism | Public Inference Route across contiguous Shards |
| Continuous batching | Inside every node across active Route Sessions |
| Data parallelism | Multiple complete routes for independent requests |
| Tensor parallelism | Deferred to a trusted composite node/managed cluster |
| Expert parallelism | Deferred to a trusted composite node/managed cluster |
| Disaggregated prefill | Deferred until core route performance passes |
| Speculative decoding | Deferred optimization |
Public WAN tensor/expert collectives are rejected for the first runtime because their per-layer communication and static rank assumptions conflict with heterogeneous volunteer nodes.
## Optional providers
### Transformers/safetensors
Remains:
- Correctness/reference backend.
- Fallback for unsupported architectures.
- Baseline for performance and output quality.
### vLLM
May run unmodified as a complete model or managed TP/PP/EP cluster represented as one logical provider. Its internal ranks are not independently routed or rewarded.
Borrow only concepts such as named bundles, continuous batching, typed compatibility fingerprints, explicit transfer lifecycle and load telemetry.
### Whole-model llama.cpp
Provides a local proxy backend, correctness oracle and performance baseline. It is not the native distributed milestone.
## Artifact and recipe compatibility
A routable recipe identifies separately:
- Source Model Artifact hash and optional derivative/slice hash.
- Architecture and adapter version.
- Tokenizer revision and vocabulary.
- Weight quantization.
- Activation interchange dtype/schema.
- Backend compute dtype and backend implementation.
- KV dtype/layout.
- RoPE/context parameters.
- llama.cpp commit and project patch version.
- Shard range and endpoint ownership.
Compatibility fails closed. Similar quantization labels or model names are not enough.
## Admission and failure
A recipe becomes routable only after a real local and distributed forward passes. Synthetic tests remain unit coverage.
Alpha failure behavior:
- Deadline or node loss cancels the Route Session.
- Every node releases KV and queued buffers.
- Uncertain mutations are not replayed silently.
- Retry starts from token zero on a newly compatible route.
- No cross-node KV import is trusted until a later signed/compatible snapshot protocol exists.
## Performance release contract
Before native development proceeds, compare the current Transformers/safetensors backend with whole-model llama.cpp under controlled model/hardware/quality lanes.
Final release compares distributed GGUF with distributed safetensors using thresholds locked before seeing final results.
Required measurements:
- TTFT.
- Prefill and decode tokens/sec.
- Aggregate concurrency throughput.
- p50/p95 latency.
- Seam bytes and latency.
- Queue/batch occupancy.
- RSS, VRAM and KV pressure.
- Output-quality drift.
- Cancellation/failure cleanup.
The GGUF path ships only if it is faster at acceptable quality or enables a larger otherwise-unroutable model at useful measured speed.
## Implementation sequence
1. Preserve completed DGR-001 performance and DGR-002 protocol contracts.
2. DGR-017 locks exact GLM-5.2 Max artifact, resource, and alpha acceptance identity.
3. Define exact recipe identity and pin one reproducible llama.cpp boundary.
4. Run two lanes in parallel: DGR-018 establishes the whole-model `UD-IQ1_S` oracle on 224+ GiB usable memory, while DGR-005/DGR-006 implement range loading and named boundary parity with a cheap dense fixture.
5. DGR-019 adds explicit GLM-5.2 MoE/MLA/DSA/IndexShare semantics after both lanes pass.
6. Implement local KV; build and integrate the standalone worker.
7. Pass local two-process and real two-physical-machine execution.
8. Harden cancellation, node loss, restart, and cleanup required by alpha.
9. DGR-020 executes the exact multi-node target and emits immutable `alpha` or `stop`.
10. Post-alpha: continuous batching, final comparison, longer context, MTP, and package optimization.
11. Prepare narrow upstream patches/tests; add Qwen as later architecture expansion.
See [the Ralph backlog](prd.json) and [implementation strategy](implementation-strategy.md).
Compatible sessions may be continuously batched within a worker while retaining isolated positions/state. Admission bounds weights, local state/KV, scratch, fragments, and queues. Uncertain cross-route mutation is not replayed. Registration can show an uncertified lane, but existing admission keeps it unroutable until signed/versioned real-hardware evidence exists.

View File

@@ -1,270 +1,40 @@
# Distributed GGUF Decision Framework
> **Superseded for active implementation decisions.** The grill was resolved on 2026-07-13. Use [implementation-strategy.md](implementation-strategy.md), [architecture.md](architecture.md), [ADR-0020](ADR-0020-distributed-gguf-runtime.md), and [prd.json](prd.json). This file remains as historical decision rationale.
This framework is for grilling open decisions. It keeps decisions tied to project vocabulary and implementation gates instead of vague "distributed inference" language.
## Core Vocabulary
Use the existing domain terms this way:
- **Shard**: contiguous transformer layer range. This is the compute, routing, cache, and reward unit.
- **Shard Swarm**: storage/download group for artifacts needed by a shard.
- **Inference Route**: ordered node sequence that covers all layers for one request.
- **Route Session**: one active request bound to one inference route and stable session id.
- **Hot KV State**: live per-shard cache held by the route node during a route session.
- **Prefix Snapshot**: persisted route-session state used for reuse or failover, not the hot decode path.
- **Artifact Manifest**: canonical mapping from model artifacts to semantic model parts and runtime support.
- **Generation Telemetry**: realtime progress for a route session, including phase and tokens/sec, independent of whether token deltas are streamed.
## The Five Planes
### 1. Control Plane
Owner: Tracker.
Responsibilities:
- node registry
- coverage map
- route selection
- rebalance directives
- route-session creation
- health and telemetry
- client-visible Generation Telemetry
- billing/audit records
Must not do:
- serve hot KV during every token
- become the only place model artifacts can be fetched
### 2. Artifact Plane
Owner: Shard Swarms, local node storage, optional CDN/bootstrap mirrors.
Responsibilities:
- GGUF/safetensors/tokenizer download
- content-addressed verification
- local artifact inventory
- artifact-to-layer mapping
- cache eviction
Must not do:
- define execution order by file split alone
- imply that a downloaded file chunk equals a Shard
### 3. Execution Plane
Owner: active Inference Route.
Responsibilities:
- chunked prefill
- one-step decode
- hidden-state transfer across activation seams
- start-layer handling for overlapping shards
- backpressure
Must not do:
- resend full context activations during decode
- require cross-node tensor parallel all-reduce for public v1
### 4. Session State Plane
Owner: route nodes for hot KV; cache servers only for snapshots.
Responsibilities:
- per-shard local KV ownership
- cache allocation and eviction
- cache ABI compatibility
- session close/release
- optional prefix snapshots
Must not do:
- centralize hot KV in a remote service
- let a replacement node continue from incompatible state
### 5. Economics And Trust Plane
Owner: tracker plus settlement/validation components.
Responsibilities:
- distinguish storage/seeding work from inference work
- account for prefill and decode separately
- record route participation
- sample validation events
- slash proven fraud
Must not do:
- pay a node for merely holding files as if it generated tokens
- hide public-swarm privacy limits from clients
## Hard Invariants
These are the framework rules unless we deliberately write a new ADR:
1. Public-network Shards are contiguous layer ranges.
2. Hot KV State is local to the node serving that Shard in that Route Session.
3. Artifact distribution and route execution are separate systems.
4. Decode seam payload must be `O(hidden_size)`.
5. Prefill may be `O(sequence_length * hidden_size)`, but only in bounded chunks.
6. The tracker chooses routes; nodes do not negotiate route topology peer-to-peer.
7. Model/backend-specific cache internals stay behind backend capability reports.
8. PyTorch remains the correctness/reference backend while llama.cpp/GGUF becomes the performance backend.
9. Streaming responses are preferred when feasible; Generation Telemetry is always required.
## Resolved Gates
### Gate 1: Public Shard Semantics
Decision: public-network Shards are contiguous transformer layer ranges. Tensor-parallel or ring-style execution is allowed only inside one trusted node, one colocated pod, or a future composite node abstraction.
Rationale:
- Layer ranges match the existing `Shard`, `Coverage Map`, `Inference Route`, billing, and fraud vocabulary.
- Public volunteer nodes should not require cross-node all-reduce or tight per-layer synchronization in v1.
- Existing projects such as prima.cpp and Distributed Llama can still inform local-cluster/backend execution without becoming the public routing primitive.
Consequences:
- Artifact Manifests must map files/tensors to semantic layer ranges.
- Route selection remains ordered layer coverage.
- Rewards can be attributed to layer-range work.
- Hot KV State is naturally owned by the node serving that layer range for the Route Session.
### Gate 2: Hot KV Strategy
Decision: v1 rejects centralized hot KV. Hot KV State is local to the node serving the relevant Shard in the active Route Session. Cache servers may store Prefix Snapshots for reuse, retry, or failover, but they are not in the per-token decode path.
Rationale:
- Decode is the tight loop; adding remote cache I/O there makes latency and bandwidth worse at the worst point.
- Local KV naturally follows layer-range Shard ownership.
- Centralized hot KV increases privacy exposure and creates consistency problems.
- Prefix Snapshots preserve the useful part of central storage without making it mandatory for every generated token.
Consequences:
- Route Session must be sticky.
- Failover is limited in alpha unless a compatible Prefix Snapshot exists.
- Cache servers are optimization infrastructure, not required runtime infrastructure.
- Route repair requires compatible model revision, layer range, backend cache ABI, and snapshot position.
### Gate 3: First Runtime Proof
Decision: prove distributed Route Session and Hot KV State semantics in the existing PyTorch route before modifying llama.cpp/GGUF.
Rationale:
- PyTorch exposes model internals and cache objects more directly, so it is the fastest way to validate the distributed protocol.
- The current distributed PyTorch route already has the right high-level shape but disables cache and recomputes full prompts.
- Fixing that path gives us a reference implementation for correctness tests, telemetry, session lifecycle, and wire protocol behavior.
- llama.cpp/GGUF should receive a clear target ABI rather than becoming both the protocol experiment and the performance backend at once.
Consequences:
- Issue 02 precedes issue 05.
- llama.cpp collaboration has a concrete target ABI.
- The PyTorch route remains the architecture-coverage/reference backend even after GGUF becomes the preferred performance path.
- The first success metric is eliminating full-prompt recompute in distributed decode.
### Gate 3A: Client Feedback During Latency
Decision: streaming responses are preferred when feasible, and realtime Generation Telemetry is required regardless of streaming support.
Rationale:
- The product optimizes for access to large capable models, so some latency is acceptable.
- Users still need confidence that the route is alive and roughly how fast it is generating.
- Streaming token deltas give the best user experience when the backend exposes them cleanly.
- Tokens/sec remains useful during prefill, queueing, and any backend that cannot stream token deltas.
Consequences:
- The gateway should stream token deltas through an OpenAI-compatible response when possible.
- The gateway must expose progress through SSE, WebSocket, or polling.
- The final answer can be delivered after completion only as a fallback.
- Telemetry must include route phase, generated token count, and rolling tokens/sec.
- Non-streaming clients still need realtime telemetry.
### Gate 4: llama.cpp Collaboration Shape
Decision: target upstreamable `libllama`/ggml hooks instead of planning around a permanent fork.
Rationale:
- llama.cpp changes quickly across model support, quantization, kernels, and hardware backends.
- A permanent fork would become expensive to maintain and would lag upstream improvements.
- A short-lived prototype branch is acceptable if it proves the API and makes upstream collaboration concrete.
- Keeping tracker/routing logic outside llama.cpp makes the upstream ask smaller and cleaner.
Consequences:
- Need a minimal reproducible localhost demo before asking upstream to carry the design.
- Need to separate "what llama.cpp should expose" from "what our tracker does".
- Desired upstream surface is layer-range execution, hidden-state boundary I/O, partial loading/introspection, and per-session KV ownership.
- If upstream rejects the shape, we revisit whether to carry a narrow adapter fork or keep GGUF distributed execution as experimental.
### Gate 5: First Model Target
Decision: use a two-tier model target. Use a small, boring, llama.cpp-supported GGUF model for the first protocol smoke test. Use `deepseek-ai/DeepSeek-V4-Flash` as the first serious large-model target. Keep GLM-5.2 and Ornith as later support audits.
Rationale:
- The first protocol proof should isolate route/session/KV bugs from model-architecture bugs.
- DeepSeek-V4-Flash is a strong first serious target because it is much smaller than 1.6T-class models while still being large enough to validate the product thesis.
- DeepSeek-V4-Flash still has architecture-specific risks, so it should not be the first smoke test.
- GLM-5.2 and Ornith remain valuable targets, but they add DSA/MLA/hybrid attention uncertainty.
Consequences:
- 128K cache accounting can be modeled now.
- The first "real" target-model audit is DeepSeek-V4-Flash support in PyTorch, vLLM/SGLang, and any available GGUF/llama.cpp quantization path.
- Production support waits for backend capability reports and exact cache ABI support.
### Gate 6: Failure Semantics
Decision: alpha fails Route Sessions on route-node loss instead of attempting automatic route repair.
Rationale:
- Route repair requires compatible Prefix Snapshots, cache ABI checks, replacement-node selection, billing correction, and client stream/error recovery.
- Local Hot KV State means a replacement node cannot continue unless it has compatible state at the same position.
- Fail-fast keeps the first implementation correct while the session/KV protocol is still being proven.
Consequences:
- Better observability and explicit errors are required.
- Snapshotting becomes a later feature, not a blocker for first inference.
- Generation Telemetry must report the last known phase and failure reason.
- Client or gateway retry starts a new Route Session from scratch.
### Gate 7: Transport
Decision: keep binary HTTP for v1 activation transfer instead of jumping immediately to QUIC, WebRTC, or a custom transport.
Rationale:
- ADR-0008 already defines binary activation bodies with HTTP headers.
- HTTP keeps the first implementation debuggable with the existing server stack and tooling.
- The core risk is route/session/KV correctness, not transport optimization.
- QUIC/WebRTC can be introduced later behind the same activation protocol once semantics are proven.
Consequences:
- Focus benchmark work on payload shape, chunking, and cache behavior first.
- QUIC/WebRTC can be introduced as an optimization behind the same activation protocol.
- v1 implementation can reuse the current HTTP routing, relay, and observability infrastructure.
- Transport abstraction should be kept narrow enough that HTTP can be replaced later without changing backend cache semantics.
## Grilling Progress
Gates 1, 2, 3, 3A, 4, 5, 6, and 7 are resolved. The remaining work is to convert the resolved framework into implementation-ready issue briefs and prototype milestones.
# Distributed GGUF Runtime decision framework
> **Specification status:** planning artifacts only. No distributed GGUF runtime is implemented. DGR-017 cleanup is complete; no runtime implementation story has completion credit. `prd.json` is authoritative.
## Decision order
1. DGR-019 locks comparable lanes and thresholds before results.
2. DGR-020 runs safetensors and whole-model llama.cpp only, then returns `go`, `optimize baseline`, or `stop`.
3. Dense and V4 work must prove parity, independent per-stage execution, local-state isolation, bounded failure, and measured resources.
4. DGR-054 returns `alpha`, `optimize measured bottleneck`, or `stop`; MTP is explicitly off.
5. Post-alpha optimizations must be selected from profiles, not assumptions.
6. DGR-070 returns `beta`, `targeted optimization`, or `stop/rollback`, and requires MTP and the exact certified hardware/recipe matrix.
## Interpretation rules
- Quant/model-fit gains are separate from runtime/kernel/transport gains.
- Fixture, real-model, real-hardware, and release evidence are never interchangeable.
- 24 and 10+ stages are certification scenarios only.
- Existing routing policy is certified, not redesigned.
- Build success is not hardware certification; dark lanes remain unroutable.
- Route loss uses cache miss and re-prefill/restart, never WAN cache migration.
## Locked scope
- Existing Meshnet Tracker routing, load balancing, billing, telemetry, relay, and provider semantics are backend-agnostic and are **not redesigned**. GGUF contributes exact compatibility, range/capacity, queue/load, seam-cost, health/reliability, and certification inputs only.
- The data plane is a standalone project-owned C++ Shard worker with gRPC/Protobuf and a project-owned `ShardEngine` boundary.
- llama.cpp is fetched at one exact commit into an ignored workspace from an in-repo manifest, then a numbered minimal patch stack is applied. There is no submodule, vendored tree, or permanent-fork dependency.
- llama.cpp owns DeepSeek V4 graphs, mHC, MoE, attention, hash routing, and kernels. Meshnet adds only range-ownership hooks, typed boundary/local-state adapters, worker integration, and parity/certification.
- Quantization and placement are dynamic recipe inputs. The 24 and 10+ stage layouts are certification scenarios, never product constants.
- Per-shard Hot KV and V4 CSA/HCA/SWA/indexer/compressor state remain local and keyed by route session/epoch. The WAN seam carries the typed mHC 4×4096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations—not per-layer caches.
- Route changes use cache miss plus re-prefill/restart. There is no WAN KV or V4 auxiliary-cache migration.
- CPU/CUDA/ROCm/Vulkan/Metal compile lanes are planned; only exact real-hardware-certified backend/model/recipe lanes may be advertised.
- Alpha requires correctness and the pre-locked useful-speed gate. MTP is reserved and off for alpha; its ownership contract, implementation, and benchmark are required before beta.
## Target identities
- DeepSeek V4 official target SHA: `60d8d70770c6776ff598c94bb586a859a38244f1`.
- llama.cpp V4 support lineage began at PR 24162 / merge `8c146a8366304c871efc26057cc90370ccf58dad`; DGR-027 later pins one exact validated current commit.
- V4 scope: 43 main layers plus MTP; mHC 4×4096 boundary; 256 routed + 1 shared experts with six routed active; token IDs required for the first three hash-routed layers.
- Exact split-GGUF artifacts are provisioned to mounted-drive storage with a complete hashed manifest and resumable verification; no model artifact may be placed under `/home`.

View File

@@ -1,275 +1,114 @@
# DGR-017 — Lock the GLM-5.2 Max target and alpha contract
# DGR-017 evidence — superseded backlog cleanup
Status: **done**. Every acceptance criterion is met with real command output.
**Completed:** 2026-07-16
**Branch:** `ralph/distributed-gguf-runtime`
**Planning checkpoint before cleanup:** `81b1fa6`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
Evidence class: **real upstream metadata + deterministic arithmetic**. No weight
payload was downloaded, no model was loaded, no GPU was used, and no benchmark was
run — and none is claimed. This story makes the target *reviewable before* the
216.7 GB download, which is exactly its job.
## Outcome
## 1. Summary
The old DGR-001…016 completion claims and active artifacts were reconciled against the live branch. No old pass state transferred to the new implementation roadmap.
The alpha target is now pinned, planned, and sealed:
The active `packages/` and `tests/` trees were restored exactly to `origin/master`. The branch therefore no longer exposes a nominal GGUF startup path backed by unimplemented transport methods, a protobuf-only native scaffold, or isolated synthetic scheduler/cache/failure modules as if they were a working distributed GGUF runtime.
- **Identity.** `zai-org/GLM-5.2` @ `b4734de4facf877f85769a911abafc5283eab3d9` and
`unsloth/GLM-5.2-GGUF` @ `abc55e72527792c6e77069c99b4cb7de16fa9f23`, quantization
`UD-IQ1_S`, six shards, 216,715,360,960 bytes, every shard's LFS SHA-256 resolved.
- **Architecture.** The config/tokenizer/chat-template metadata the runtime cannot
shard without, hashed at the pinned revision.
- **Resources.** A deterministic planner that counts unified memory once, applies the
`max(20% , 8 GiB)` reserve, and reports the arithmetic minimum and the recommended
node count as two different numbers.
- **Contract.** The roadmap's section-5 acceptance matrix as a machine-readable,
digest-sealed document, locked before the target ever runs and cross-bound to the
exact manifest and architecture-snapshot digests.
- **Upstream.** A refreshed llama.cpp/donor status report.
## Classification and disposition
Three findings are worth a reader's attention.
### Retained
**Every number in the roadmap reproduced from primary sources.** The 216,715,360,960
byte total, the 201.832 GiB figure, the whole KV table (0.73 / 0.77 / 0.89 / 1.68 GiB
at 16K, through 46.62 / 49.41 / 56.98 / 107.25 GiB at 1M), and the whole tier table
(9 / 6 / 4 / 3 / 2 arithmetic minimum nodes) fall out of the exact config and the
exact shard bytes. The roadmap was not approximating. The planner is written as a
*reproduction* of those tables, so if the arithmetic ever stops matching, a test says
which numbers moved.
- Accepted ADRs and repository research, including `docs/research/colibri-implementation-audit.md`.
- The authoritative 55-story roadmap `DGR-017…071` and its generated issue specifications.
- The real public-relay smoke benchmark, moved with provenance to `legacy-public-relay-smoke-benchmark.json`.
- Git history containing the complete superseded implementation/reference work.
**The roadmap's "recommended" column is an imbalance factor of exactly 1.10.** Nodes
= `ceil(total x 1.10 / budget)` yields 10 / 6 / 5 / 3 / 3 for the 32 / 48 / 64 / 96 /
128 GiB tiers — precisely the roadmap's recommendations. That constant is now named
(`PLACEMENT_IMBALANCE_FACTOR`) and documented as a placeholder for measured
per-tensor placement, not a fudge factor to be tuned once results are in.
### Removed from the active tree
**224 GiB aggregate does not actually fit.** Two 112 GiB nodes hit the 224 GiB
"hard-fit floor" exactly, and still come up **23.5 GiB short** once each node honours
its reserve. That is what makes 224 GiB an *experimental floor* rather than an
envelope, and it is now a test, not a caveat in prose. Relatedly, the 2×128 and 4×64
"fit probe" topologies fit with only **2.08 GiB of headroom across the entire route**
which is why they require measured placement evidence and are not the recommendation.
- Legacy issue specifications DGR-001…016 and their stale/blocked/synthetic evidence directories.
- The nonfunctional `gguf_backend` startup path whose gRPC execution methods raised not-implemented errors.
- Synthetic/reference-only boundary, Hot KV, scheduler, failure, recipe, ownership, and native-protocol modules that were not a real llama.cpp Shard runtime.
- The protobuf round-trip-only native scaffold, placeholder llama.cpp patch, generated bindings/build workspace, and associated tests.
- Tracker/admission/source modifications coupled to that superseded scaffold.
## 2. Files changed
### Confirmed absent and still required
New — runtime-loadable package (single source of truth):
- Real standalone C++ gRPC Shard worker.
- Exact pinned llama.cpp manifest and verified patch stack.
- Range-aware GGUF tensor ownership and real ranged execution.
- Real Shard-local llama.cpp KV/V4 auxiliary state.
- DeepSeek V4 boundary adapter and ranged parity.
- Real multi-machine DeepSeek V4 alpha or beta acceptance.
| Path | What |
|---|---|
| `packages/node/meshnet_node/glm_alpha/__init__.py` | Public surface |
| `packages/node/meshnet_node/glm_alpha/manifest.py` | Target manifest + architecture snapshot; fail-closed identity |
| `packages/node/meshnet_node/glm_alpha/planner.py` | Memory / KV / seam planner; unified-memory de-duplication |
| `packages/node/meshnet_node/glm_alpha/contract.py` | Immutable, digest-sealed alpha acceptance contract |
| `packages/node/meshnet_node/glm_alpha/data/target-manifest.json` | The six pinned shards, sizes, SHA-256, URLs, licenses |
| `packages/node/meshnet_node/glm_alpha/data/architecture-snapshot.json` | Pinned architecture + config/template hashes |
| `packages/node/meshnet_node/glm_alpha/data/alpha-contract.json` | Sealed acceptance thresholds (`aab23220…`) |
| `scripts/refresh_glm_target_manifest.py` | Re-resolve/verify pins from upstream metadata (`--check` / `--write`) |
| `tests/test_glm_alpha_target.py` | 97 deterministic offline tests (99 after the late-review repair — see §4a) |
These remain `passes: false` in DGR-018…071.
New — evidence:
## Before-cleanup baseline
- `.scratch/distributed-gguf-runtime/evidence/DGR-017/README.md` (this file)
- `.../commands.txt` — exact commands and real results
- `.../resource-plan.json` — generated tier/route/seam/KV plan
- `.../upstream-status.json` — refreshed llama.cpp and donor status
Command:
Modified:
- `.scratch/distributed-gguf-runtime/issues/17-...md``Status: done`
- `.scratch/distributed-gguf-runtime/prd.json` — DGR-017 `passes: true` (this story only)
- `.ralph-tui/progress.md` — learnings
The data files live **in the package**, not in evidence, because the runtime must load
them (DGR-018 verifies downloads against these digests; DGR-003 folds the manifest
digest into the recipe fingerprint). Duplicating them into evidence would create two
sources of truth that could drift.
## 3. Acceptance criteria
| Criterion | Where it is proven |
|---|---|
| Pin both repos by exact observed revision; `UD-IQ1_S` is the alpha quant | `target-manifest.json`; `test_manifest_pins_both_repositories_by_exact_revision` |
| Six filenames, exact bytes, LFS SHA-256, aggregate GB/GiB, license, URLs, no payload download | HF `paths-info` API (LFS pointer metadata); `test_manifest_resolves_all_six_shards…`, `test_manifest_aggregate_bytes_are_exact_and_self_consistent` |
| Snapshot + hash architecture-critical config/tokenizer/chat-template metadata | `architecture-snapshot.json`; `test_snapshot_captures_the_architecture_critical_metadata`, `test_snapshot_hashes_the_config_and_chat_template_bytes` |
| Deterministic minimum-node calc from exact bytes, Q8_0 KV @16K/c1, imbalance, reserve | `planner.plan_topology`; `test_topology_planner_reproduces_the_published_tier_table` |
| 224 GiB is a hard-fit floor, not an envelope; recommend 5×64 or 3×96/128 | `test_224_gib_aggregate_is_a_hard_fit_floor_not_an_operational_envelope`, `test_the_recommended_topologies_are_five_by_64_or_three_by_96_or_128` |
| Unified memory counted once; additive RAM+VRAM rejected | `NodeMemory.from_host`; `test_adding_integrated_gpu_memory_to_system_ram_is_rejected`, `test_unified_memory_is_counted_once` |
| 2.5 GbE minimum / 10 GbE recommended; serial latency modelled apart from bandwidth | `planner.plan_seams`; `test_2_5_gbe_is_the_alpha_minimum_and_10_gbe_is_recommended`, `test_serial_seam_latency_is_modelled_separately_from_bandwidth` |
| Identity/semantic/target-run/performance/reliability/storage criteria locked before execution | `alpha-contract.json` (sealed, `locked_before_target_execution: true`); `test_the_contract_locks_every_roadmap_acceptance_section` |
| Refresh upstream llama.cpp + donor status; no broad fork/scheduler | `upstream-status.json`; `adoption_state: none adopted` |
| Tests reject changed revisions, missing shards, coordinated digest/config substitutions, inconsistent bytes, duplicate unified memory, malformed telemetry, and post-result threshold mutation | 97 tests; see §5 |
| Targeted pytest passes | `97 passed` |
| Installed wheel includes and loads all locked JSON resources | Real wheel build/install plus `load_locked_target()` outside the source tree |
| `compileall packages tests` | exit 0 |
| `git diff --check` | exit 0 |
| Default tests deterministic, download-free, credit-free, GPU-free | Pure JSON + arithmetic; the only network code is an opt-in script excluded from the suite |
| Full deterministic `pytest -q` | **852 passed, 13 skipped** on final rerun |
## 4. Real results
```
scripts/refresh_glm_target_manifest.py --check -> match upstream (exit 0)
pytest -q tests/test_glm_alpha_target.py -> 97 passed
build + install wheel; load_locked_target() -> INSTALLED_WHEEL_PASS
compileall -q packages tests -> exit 0
git diff --check -> exit 0
pytest -q -> 852 passed, 13 skipped (253.30s)
```bash
.venv-rocm/bin/python -m pytest -q \
tests/test_performance_contract.py tests/test_native_shard_protocol.py \
tests/test_gguf_ownership.py tests/test_boundary_adapter.py \
tests/test_hot_kv_state.py tests/test_gguf_backend.py \
tests/test_batch_scheduler.py tests/test_failure_semantics.py \
tests/test_llama_worker_build.py tests/test_node_admission.py \
tests/test_node_capability.py tests/test_tracker_capability_admission.py
```
Controller review found and fixed four gaps before the commit was accepted:
Result:
1. the initial wheel omitted `glm_alpha/data/*.json`, despite source-tree tests passing;
2. the self-sealed contract did not bind the manifest/snapshot digests, so coordinated
shard-hash or architecture substitutions could be accepted;
3. `--check` followed moving repository HEAD instead of validating immutable pins;
4. non-finite resource telemetry could bypass ordinary range checks.
```text
216 passed, 2 skipped, 1 failed, 1 warning
```
Each failure now has either an adversarial regression test or a real installed-wheel /
live-metadata gate. None of the provisional agent results were accepted on trust.
The failure was a synthetic capability-test helper `KeyError: 'compatibility_fingerprint'`. The warning was a pre-existing heartbeat-thread `SystemExit` warning.
The intermittent tracker cancellation race that DGR-001 and DGR-002 both recorded as
flaky on a clean tree failed once in the final validation (`404` after the request
completed), then passed **5/5** in isolation and passed in the integrated full-suite
rerun above. This story touches no tracker code; the failed run is retained in
`commands.txt` rather than hidden.
## Cleanup verification
### 4a. Late independent-review repair (2026-07-14)
### Source equality
During delayed DGR-003 review, two contract-continuity defects were found and
fixed here: v1 now has an independently trusted digest pinned in code
(`test_resealing_a_mutated_v1_contract_is_rejected`) and parsed nested contract
state is recursively immutable. This added two tests; the suite is now
**99 passed** (`commands.txt` §7 records the exact runs). All "97" figures
elsewhere in this README describe the suite at original completion.
Command:
Planner output (`resource-plan.json`):
```bash
git diff --quiet origin/master -- packages tests
```
| Route | Fits | Headroom |
|---|---|---:|
| 5×64 GiB unified (recommended) | yes | +53.28 GiB |
| 3×96 GiB unified (recommended) | yes | +27.68 GiB |
| 3×128 GiB unified (recommended) | yes | +104.48 GiB |
| 4×64 GiB (fit probe) | yes | **+2.08 GiB** |
| 2×128 GiB (fit probe) | yes | **+2.08 GiB** |
| 2×112 GiB (= 224 GiB floor) | **no** | 23.52 GiB |
| 3×64 GiB | no | 49.12 GiB |
Result:
## 5. How the "no silent swap" claim is earned
```text
packages_tests_match_origin_master=yes
```
The story's whole purpose is to stop a later agent from changing the target after
seeing a result. Each of those moves now has a test that names it:
The staged cleanup removes approximately 15.3k obsolete source/test/evidence lines from the active branch.
- swap the artifact → `test_a_changed_gguf_revision_is_rejected`
- pin a branch instead of a commit → `test_a_branch_name_is_not_an_acceptable_revision_pin`
- drop a shard → `test_a_missing_shard_is_rejected`
- shrink a shard so it "fits" → `test_a_shard_size_edited_to_make_the_model_look_smaller_is_rejected`
- change a valid-looking shard SHA without changing bytes → `test_coordinated_shard_hash_substitution_is_rejected_by_contract`
- change internally consistent architecture metadata → `test_internally_consistent_architecture_substitution_is_rejected_by_contract`
- take the bigger quant quietly → `test_swapping_in_a_different_quantization_is_rejected`
- add iGPU "VRAM" to system RAM → `test_adding_integrated_gpu_memory_to_system_ram_is_rejected`
- count one machine twice → `test_the_same_machine_counted_twice_in_a_route_is_rejected`
- lower the speed floor → `test_lowering_the_speed_floor_after_seeing_a_result_is_rejected`
- call a slow pass an alpha → `test_relabelling_a_speed_failure_as_a_pass_is_rejected`
- admit the dense fallback → `test_admitting_the_dense_attention_fallback_after_the_fact_is_rejected`
- shrink the reserve → `test_relaxing_the_per_node_reserve_after_the_fact_is_rejected`
### Cleanup-relevant regression suite
**Contract continuity is fail-closed.** The documents `contract_sha256` detects
accidental edits, and the approved v1 digest is pinned independently in code. A caller
that changes a threshold and re-seals it under `glm-5.2-max-alpha/v1` is rejected by
`test_resealing_a_mutated_v1_contract_is_rejected`; amendments require a new supported
contract identity under human review. Parsed nested state is recursively immutable,
so thresholds cannot change between validation and use; `to_dict()` returns an isolated
copy rather than exposing the validated object.
Command:
## 6. Upstream status — the gating risk for DGR-004/DGR-018
```bash
.venv-rocm/bin/python -m pytest -q \
tests/test_node_admission.py tests/test_node_capability.py \
tests/test_tracker_capability_admission.py \
tests/test_kv_cache_distributed.py tests/test_real_distributed_inference.py
```
Refreshed against live GitHub on 2026-07-13. One item **changed** since the roadmap:
Result:
- **#24231 is now MERGED** (2026-07-11) — a generic `GGML_OP_LIGHTNING_INDEXER` exists.
- #24770 MERGED (2026-06-20) — GLM-5.2 loads via a **dense-MLA compatibility path**.
- **#25407 still OPEN** (updated today) — the real GLM DSA/IndexShare wiring.
- #24730 still OPEN — the umbrella GLM-5.2 support request.
```text
119 passed, 2 skipped, 1 warning in 15.90s
```
**No released upstream llama.cpp performs native GLM-5.2 DSA + IndexShare today.** A
stock pin taken now would load the artifact and emit text through the dense fallback —
which the alpha contract explicitly refuses (`dense_attention_fallback_satisfies_alpha:
false`). DGR-018 must therefore prove those paths are *active*, not that output appeared.
The warning is the same pre-existing heartbeat-thread `SystemExit` warning.
Donor policy holds, and the evidence now supports it more strongly than before: PR
#25407 is **12 files, +414/7**. The semantics alpha needs are small enough to track
and reproduce upstream. That is the argument against adopting Mesh-LLM's 261-patch
fork — recorded as a donor (Apache-2.0, branch head `9bd18f15`, 2026-07-12), nothing
adopted here.
### Known `origin/master` limitations
## 7. Limitations and deferred work
The wider routing run produced `210 passed, 2 skipped, 4 failed, 1 warning`. Each failure reproduced individually while `packages/` and `tests/` matched `origin/master` exactly:
- **No artifact was downloaded or loaded.** Sizes and SHA-256 come from Hugging Face
LFS pointer metadata. DGR-018 must verify the digests against the real files on
mounted storage before route admission. A matching size with a wrong hash is exactly
the failure this manifest exists to catch, and only a local verify can catch it.
- **`PLACEMENT_IMBALANCE_FACTOR = 1.10` is a planning assumption, not a measurement.**
It reproduces the roadmap's recommendations, but the real per-node share depends on
exact tensor bytes (embeddings, output head, dense vs MoE layers, shared experts,
indexer tensors, quant block alignment). DGR-019 must replace it with measured
placement. Until then, arithmetic-minimum topologies (2×128, 4×64) stay fit probes.
- **KV numbers are planning estimates, not admission truth.** The planner deliberately
budgets the *conservative* indexer layout (keys across all 78 layers, not just the 21
Full ones) so a route admitted here cannot be surprised by the implementation it
actually gets. The runtime must still report measured allocated/resident MLA and
indexer cache per shard.
- **Peak scratch is unmodelled.** The reserve exists precisely because backend
workspaces and graph scratch are not predictable from the artifact; measured peak
must land inside the reserve, and the contract requires that evidence.
- **Upstream is moving fast.** #25407 was updated the same day it was observed. Refresh
`upstream-status.json` before DGR-004 picks a llama.cpp pin. The manifest script's
`--check` deliberately validates the immutable Hugging Face pins, not moving HEAD.
- `test_tracker_models_endpoint_lists_registered_hf_repo_and_short_name_alias`
- `test_torch_node_applies_tracker_load_shard_directive`
- `test_shard_heal_cycle_surviving_node_covers_dead_peers_gap`
- `test_a_node_with_an_unusable_precision_covers_no_layers`
## 8. Compatibility and migration notes
They are recorded as pre-existing baseline defects and were not repaired or hidden by this cleanup story.
- Purely additive. No existing module, wire format, or test changed. Nothing in this
story is on a live request path.
- `meshnet_node.glm_alpha` has **no heavy imports** — no torch, no transformers, no
network at import time — so a tracker or planner can read the target contract without
paying for a model runtime.
- Re-pinning is deliberately awkward: `--write` follows current HEAD and leaves the
existing contract binding invalid until a new contract is reviewed and sealed.
`--check` uses revision-specific APIs and exits non-zero rather than healing any
integrity drift in the already locked target.
## Dependency handoff
## 9. Handoff to dependent stories
**DGR-003 (recipe identity):** fold `TargetManifest.digest`
(`0b6aed04479d204902bb64c0203f1a46cab26a47b378ecccf85237b63f6c1962`) and
`ArchitectureSnapshot.digest` (`253fbd94…`) into the runtime recipe fingerprint. The
GLM fields the roadmap asks you to add (DSA/IndexShare metadata, context max, expert
counts) are already resolved in `architecture-snapshot.json` — read them, do not
re-derive them by hand. Populate the DGR-002 `Fingerprint` message; do not invent a
second identity struct.
**DGR-004 (llama.cpp pin):** read `upstream-status.json` first. Any pin taken before
#25407 merges gives you the dense-MLA fallback, which cannot satisfy alpha. Track
#25407 (12 files) as a numbered patch; do not adopt the Mesh-LLM fork.
**DGR-018 (whole-model oracle):** the six digests in `target-manifest.json` are what you
verify the download against. Your host needs ≥224 GiB runtime-accessible memory — and
note that 224 GiB is the *floor*, not a comfortable target (see §1). Prove DSA,
IndexShare, shared expert, and the Max template are **active**; the contract's
`require_rendered_reasoning_effort_marker` is `<|system|>Reasoning Effort: Max`. Assert
the rendered marker, not the request field: the template's only non-max level is
`'high'`, so *every other value — including an absent one — renders Max*, and "the
request said max" proves nothing.
**DGR-019 (GLM semantics):** the IndexShare split is **21 Full producer layers and 57
Shared consumers**, in a `[full, full, full] + repeating [shared, shared, shared, full]`
pattern. Prefer shard boundaries that keep an ownership group whole; the 8 KiB
(2048 × int32) top-k sideband is the cost when you cannot. Replace
`PLACEMENT_IMBALANCE_FACTOR` with measured per-tensor placement.
**DGR-020 (alpha verdict):** load the target with `load_locked_target()`. It verifies
the contract seal and cross-binds the manifest and architecture snapshot before
returning them. Judge against
`contract.threshold(section, key)` — an unlocked threshold raises rather than
defaulting, so a criterion cannot be invented at read time. The verdict is `alpha` or
`stop`; there is no third outcome, and a quality pass with a speed failure is `stop`.
**Everyone:** unified system RAM and integrated-GPU memory are one pool. Build nodes
with `NodeMemory.from_host(..., unified=True)` and it is impossible to write the
double-count; pass a GPU size alongside `unified=True` and it raises rather than
silently ignoring the argument.
DGR-018 and later stories must start from the cleaned upstream-equivalent runtime tree. Reuse concepts from superseded commits only by explicitly porting the smallest verified slice under the new storys contracts, tests, and evidence gates. Git history is provenance, not completion evidence.

View File

@@ -0,0 +1,83 @@
{
"schema_version": 1,
"executed_at_utc": "2026-07-15T10:41:14Z",
"test_kind": "public-relay-single-node-streaming-smoke-benchmark",
"target": {
"public_chat_endpoint": "https://meshnet.2.d-popov.com/v1/chat/completions",
"relay_url": "wss://meshnet.2.d-popov.com/ws",
"model": "qwen2.5-0.5b-instruct",
"quantization": "bfloat16"
},
"recovery": {
"problem": "The local node's capability proof had expired and its port-7000 HTTP server had wedged with CLOSE-WAIT sockets.",
"action": "Gracefully restarted the local public-tracker meshnet-node process on port 7000.",
"startup_validation": {
"device": "cuda",
"capability_proof_ms": 336,
"node_id": "7j77FsPY-b32476219492",
"relay_addr": "wss://meshnet.2.d-popov.com/rpc/7j77FsPY1evV8tuf-7000"
}
},
"tracker_admission_after_recovery": {
"node_id": "7j77FsPY-b32476219492",
"alive": true,
"status": "ready",
"capability_state": "admitted",
"routable": true,
"route_hops": 1
},
"client_measurements": {
"warmup": {
"http_status": 200,
"ttft_ms": 420.8,
"elapsed_ms": 610.23,
"response_text": "MeshNet Relay Benchmark Passed"
},
"runs": [
{
"run": 1,
"ttft_ms": 376.04,
"elapsed_ms": 458.65,
"response_text": "relay benchmark pass"
},
{
"run": 2,
"ttft_ms": 258.33,
"elapsed_ms": 336.71,
"response_text": "relay benchmark pass"
},
{
"run": 3,
"ttft_ms": 288.26,
"elapsed_ms": 363.2,
"response_text": "relay benchmark pass"
}
],
"p50_ttft_ms": 288.26,
"p50_elapsed_ms": 363.2
},
"tracker_relay_evidence": [
{
"status": 200,
"relay": true,
"node_id": "7j77FsPY-b32476219492",
"tokens": 11,
"elapsed_seconds": 0.1686,
"tokens_per_sec": 65.2541
},
{
"status": 200,
"relay": true,
"node_id": "7j77FsPY-b32476219492",
"tokens": 11,
"elapsed_seconds": 0.1891,
"tokens_per_sec": 58.1799
}
],
"scope_and_remaining_work": {
"validated": "Public HTTPS chat endpoint routed a streaming request through the tracker relay to the local CUDA node and completed with HTTP 200.",
"not_validated": "Two-node shard routing was not run because the remote node 5gMLrmyB-88f5cba044d0 still had an expired capability proof and was not routable.",
"next_gate": "Refresh the remote node capability proof, then load a multi-node-compatible assignment and repeat the benchmark through the public tracker relay."
},
"reproduction": "Use a valid bearer API key with the public /v1/chat/completions endpoint and stream a short qwen2.5-0.5b-instruct request. Do not connect directly to private node HTTP endpoints; the tracker relay is the required path."
}

View File

@@ -0,0 +1,204 @@
# DGR-018 evidence — canonical Ralph and Gitea metadata schema
**Completed:** 2026-07-16
**Branch:** `ralph/distributed-gguf-runtime`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
**Dependency:** DGR-017 (`evidence/DGR-017/README.md`) — cleaned backlog reconciled to `origin/master`; no old pass state transferred.
## Objective
Make `prd.json` the validated source from which Markdown (and, later, Gitea) issues
can be generated losslessly, per
`.scratch/distributed-gguf-runtime/issues/018-define-canonical-ralph-and-gitea-metadata-schema.md`.
## Pre-existing state found (not caused by this story)
Before any change in this session, `git status` showed `.scratch/distributed-gguf-runtime/prd.json`
already modified in the working tree relative to `HEAD` (commit `369b207`), with no corresponding
progress-log entry. Diffing against `HEAD` showed the working copy had **dropped** prd.json's
top-level `sourceOfTruth`, `qualityGates`, `metadataSchema`, `milestones`, and `supersededStories`
objects, while `userStories` itself was byte-identical to `HEAD`. This looked like an abandoned,
uncommitted partial edit from a prior session, not intentional current work — those fields are
exactly the schema/quality-gate/audit-provenance content this story depends on, and their loss
wasn't explained by any acceptance criterion. They were restored (see "Changes" below) rather than
silently accepted or discarded, per the instruction to investigate unexplained working-tree state
before building on top of it.
## Changes
### `scripts/ralph_prd_schema.py` (new)
Single module providing:
- **Parse:** `load_prd(path)` — JSON load with clear `PrdValidationError`s for missing file /
invalid JSON / non-object document.
- **Canonical schema registry:** `STORY_FIELDS` (name → required/type), `EXECUTION_MODES`,
`EVIDENCE_CLASSES`, `HARDWARE_FLAGS`, `UPSTREAM_FLAGS`, `TRIAGE_VALUES`. Covers every field named
in the acceptance criteria: stable `id`/`title`, `labels`, `milestone`, derived `type`
(`derive_type`), `dependsOn`, derived `blocks`, `triage`, `evidenceClass`, and
`hardware`/`model`/`upstream` flags.
- **Structural validation:** `validate_schema(data)` — required fields, types, enum membership,
ID convention, `type:`/`priority:` label cardinality, non-empty `acceptanceCriteria`.
- **Semantic validation:** `validate_semantics(data)` — unique IDs, unique titles, `dependsOn`
resolves to known stories (no self-dependency), dependency graph is acyclic (with a reported
cycle path on failure), `blocks` matches the dependency graph exactly (sorted set equality, not
superset), and `evidencePath` matches the per-story convention.
- **Fresh vs. in-progress backlog:** `validate_fresh_backlog(data)` additionally requires every
story to start `passes: false` (for a backlog that hasn't started execution yet);
`validate_backlog(data)` is the composed check for a real, in-flight backlog where some stories
have legitimately completed.
- **Self-consistency check:** `validate_metadata_schema_consistency(data)` — when prd.json declares
its own `metadataSchema`/`qualityGates` (as this one now does), verifies that self-documentation
hasn't drifted from what the validator actually enforces (enum sets, required/optional field
lists, presence of `qualityGates` and `generatedArtifactDisclaimer`). This is a no-op for minimal
fixture PRDs that don't carry that documentation.
- **Generation (one-directional, prd.json → artifact):** `render_issue_markdown(story, data)`
renders the exact Markdown convention already used by
`.scratch/distributed-gguf-runtime/issues/*.md`, sourcing the "Shared quality gates" bullets from
`data["qualityGates"]` and the leading disclaimer from
`data["metadataSchema"]["generatedArtifactDisclaimer"]` (falling back to a module default only
when `data` omits them) — not from a duplicated Python string literal.
`to_gitea_issue_payload(story, data)` wraps the same body into a Gitea create-issue-shaped payload
(`title`, `body`, `labels`, `milestone`).
- **Authority guard:** `check_generated_markdown_authority(text, disclaimer=...)` rejects generated
Markdown that's missing the disclaimer or that contains a conflicting authority claim (e.g. "this
file is authoritative"). There is deliberately no Markdown → prd.json parser, so a generated
artifact structurally cannot feed `passes` (or anything else) back into the authoritative source.
- CLI: `python scripts/ralph_prd_schema.py validate <prd.json> [--fresh]` and
`... render <prd.json> <STORY-ID>`.
### `.scratch/distributed-gguf-runtime/prd.json`
- Restored the top-level `sourceOfTruth`, `qualityGates`, `milestones`, and `supersededStories`
objects to their `HEAD` content (see "Pre-existing state" above); `userStories` was already
identical to `HEAD` and is unchanged in content.
- Extended `metadataSchema` (previously incomplete for this story's own acceptance criteria) with:
`requiredStoryFields` now also lists `notes` and `blocks` (present on all 55 stories); new
`optionalStoryFields: ["completionNotes"]`; new `hardwareValues`/`upstreamValues` enums (`model`
is documented as an open convention, not a closed enum, since quantization/model targets are
dynamic recipe inputs per `RALPH-CONTEXT.md`); new `typeDerivation` and `labelConventions`
(reserved prefixes, cardinality); new `generatedArtifactDisclaimer` (the exact string generated
artifacts must start with); extended `dependencyRules`/`authorityRule` prose to match what the
validator enforces.
- Reworded `sourceOfTruth`'s stale "All stories are unimplemented ... passes=false" clause, which
was no longer accurate once DGR-017 completed.
- Marked `DGR-018.passes = true` with `completionNotes` recording this story's outcome.
### `.scratch/distributed-gguf-runtime/issues/018-define-canonical-ralph-and-gitea-metadata-schema.md`
Regenerated via `render_issue_markdown` to reflect `passes: true` (checked acceptance criteria,
"completed" status line, "Verified evidence" handoff line) — matching the same convention DGR-017's
issue file already used for a completed story.
### `tests/test_ralph_prd_schema.py` (new)
108 deterministic, model-download-free, GPU-free tests:
- **Parse** (4 tests): real backlog parses to 55 stories; missing file, invalid JSON, and
non-object documents raise `PrdValidationError`.
- **Structural/semantic validation against the real backlog** (7 tests): passes `validate_schema`,
`validate_semantics`, and the composed `validate_backlog`; unique IDs/titles; all `dependsOn`
resolve; `blocks` matches the derived dependency graph for all 55 stories; no cycle; every
`passes: true` story carries `completionNotes` and an existing evidence README (a durable
invariant, not a hardcoded list of which stories have completed — that list will keep growing).
- **Structural/semantic failure-mode fixtures** (13 tests): missing required field, bad enum, wrong
type, empty `acceptanceCriteria`, multiple `type:` labels, duplicate ID, duplicate title, unknown
dependency, self-dependency, dependency cycle, mismatched `blocks`, bad `evidencePath`.
- **Fresh-backlog invariant** (3 tests): accepts all-`false`, rejects a premature `passes: true`,
and confirms `validate_backlog` (the in-progress variant) permits completed stories.
- **prd.json-as-source-of-truth for boilerplate** (9 tests): `qualityGates`/`metadataSchema`
self-consistency checks (no-op without them, catches a drifted enum, catches a missing
`qualityGates`), `quality_gate_bullets` flattening order, `authority_disclaimer` precedence and
fallback, and 3 tests asserting the real backlog's declared schema matches the code, its 7
quality-gate bullets are intact, and its disclaimer matches the module default.
- **`derive_type`** (4 tests): label-derived type, release-gate synthetic type for HITL gate
stories, `None` when absent, and confirmation that the real backlog's two release-gate stories
(`DGR-054`, `DGR-070`) derive `release-gate`.
- **Markdown generation round trips** (55 parametrized + 6 tests): `render_issue_markdown` for
every story `DGR-017`..`DGR-071` is byte-for-byte identical to the corresponding file already in
`.scratch/distributed-gguf-runtime/issues/`; determinism; leading disclaimer; `Blocks (derived)`
rendering (`None` vs. listed); checkbox reflects `passes`; filename convention.
- **Authority-claim rejection** (4 tests): accepts real generated text, rejects a missing
disclaimer, rejects an overriding claim, and confirms every committed issue file in
`.scratch/distributed-gguf-runtime/issues/` passes the check.
- **Gitea payload generation** (3 tests): payload shape, body carries no information beyond what's
in prd.json, and every real story's payload is well-formed and authority-clean.
## Commands and results
```bash
python3 -m pytest -q tests/test_ralph_prd_schema.py
```
```text
108 passed in 0.16s
```
```bash
python3 -m compileall -q packages tests
```
Exit code 0, no output (all files compile).
```bash
git diff --check
```
Exit code 0 (no whitespace errors).
```bash
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
```
```text
OK: 55 stories validated.
```
```bash
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json --fresh
```
```text
ERROR: DGR-017: fresh backlog requires passes=false, got True
ERROR: DGR-018: fresh backlog requires passes=false, got True
2 validation error(s).
```
Expected: `--fresh` is the invariant for a backlog that hasn't started execution; this backlog has
legitimately completed two stories, so it correctly fails that stricter check while passing the
plain (in-progress) `validate` command above.
### Baseline: full repository suite (ad hoc `python3`, not a project venv)
```bash
python3 -m pytest -q
```
```text
20 failed, 776 passed, 13 skipped, 2 warnings in 244.17s (0:04:04)
```
None of the failures touch `scripts/ralph_prd_schema.py` or `tests/test_ralph_prd_schema.py`
(neither file existed before this story; this story adds no changes to `packages/`). Four of the
20 failures reproduce exactly the pre-existing baseline defects DGR-017's evidence already recorded
(`test_tracker_models_endpoint_lists_registered_hf_repo_and_short_name_alias`,
`test_torch_node_applies_tracker_load_shard_directive`,
`test_shard_heal_cycle_surviving_node_covers_dead_peers_gap`,
`test_a_node_with_an_unusable_precision_covers_no_layers`). The remaining 16 (activation
compression, dynamic routing, gossip/relay, manual route benchmark, openai gateway, TOPLoC
calibration dispatch, tracker control plane) include a `ModuleNotFoundError: langchain` failure,
indicating this ad hoc `python3` lacks the project's `dev` extras (`langchain-openai`, etc.) rather
than a real regression; this environment has no project virtualenv (e.g. no `.venv-rocm`) to run
against instead. Not investigated further as out of scope for this story.
## Limitations
- No real Gitea instance or API integration exists; `to_gitea_issue_payload` defines the payload
shape (title/body/labels/milestone) only. Creating issues against a live Gitea server is future
work, not claimed here.
- `model` is intentionally validated as an open string, not a closed enum, per
`RALPH-CONTEXT.md`'s "Quantization and placement are dynamic recipe inputs" constraint; the schema
documents (`metadataSchema.modelConvention`) but does not restrict its value set.
- Validation and generation were exercised only against this feature's `prd.json`
(`.scratch/distributed-gguf-runtime/prd.json`); `docs/prd.json` and other `.scratch/*/prd.json`
files in this repo use a materially different (simpler) shape and are out of scope.
## Dependency handoff
DGR-021 and DGR-025 (this story's derived `blocks`) may treat `prd.json`'s `metadataSchema`,
`qualityGates`, and this validator/generator as stable. Any future field addition to a story shape
must extend `STORY_FIELDS` in `scripts/ralph_prd_schema.py` and the corresponding
`metadataSchema.requiredStoryFields`/`optionalStoryFields` in `prd.json` together —
`validate_metadata_schema_consistency` fails closed if they drift apart.

View File

@@ -0,0 +1,215 @@
# DGR-019 evidence — lock alpha and beta performance contracts
**Completed:** 2026-07-22
**Branch:** `ralph/distributed-gguf-runtime`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
**Dependency:** DGR-017 (`evidence/DGR-017/README.md`) — cleaned backlog reconciled to `origin/master`; no old pass state transferred.
## Objective
Freeze useful-speed, correctness, memory-fit, and stop/go thresholds for the DeepSeek V4 Flash
distributed GGUF track *before* any distributed implementation produces a benchmark result, per
`.scratch/distributed-gguf-runtime/issues/019-lock-alpha-and-beta-performance-contracts.md`.
## Pre-existing state found (not caused by this story)
Before any change in this session, `git status` showed `.scratch/distributed-gguf-runtime/prd.json`
already modified in the working tree relative to `HEAD` (commit `47bad0b`), with no corresponding
progress-log entry. Diffing against `HEAD` showed the working copy had **dropped** prd.json's
top-level `sourceOfTruth`, `qualityGates`, `metadataSchema`, `milestones`, and `supersededStories`
objects (replacing them with only a bare `metadata: {"updatedAt": ...}` stamp), while `userStories`
itself was byte-identical to `HEAD`. Running `tests/test_ralph_prd_schema.py` against the
as-found working tree confirmed the damage: 56 of 108 tests failed (every
`test_render_issue_markdown_matches_committed_file[...]` parametrization, since
`quality_gate_bullets`/`authority_disclaimer` fall back to module defaults once `qualityGates`/
`metadataSchema` are absent, which no longer match the committed issue files).
This is the same shape of problem DGR-018's evidence documented and fixed: an abandoned,
unexplained edit that silently dropped the schema/gates/milestone/provenance content this and
future stories depend on, while `scripts/ralph_prd_schema.py validate` did not catch it (those
top-level sections are optional-if-absent by design, so the CLI reported `OK: 55 stories
validated.` even with them missing). The most likely cause is `ralph-tui`'s own read/write of
`prd.json` as its task source, which only round-trips the fields it models
(`name`/`description`/`branchName`/`userStories`) and stamps its own `metadata.updatedAt`,
dropping any project-specific extension fields it doesn't know about.
Per `RALPH-CONTEXT.md`'s instruction to inspect `git status` and preserve unrelated work rather
than build on top of unexplained state, and following the DGR-018 precedent, the dropped fields
were restored verbatim from `HEAD` (`git show HEAD:.scratch/distributed-gguf-runtime/prd.json`)
while keeping the current `userStories` content (identical) and the current `metadata.updatedAt`
stamp. `tests/test_ralph_prd_schema.py` returned to `108 passed` immediately after the restore,
before any DGR-019-specific change was made.
## Changes
### `packages/node/meshnet_node/dgr_performance/` (new package)
- **`data/alpha-beta-contract-v1.json`** — the locked, versioned, machine-readable contract.
`schema_version`/`contract_version`/`contract_id` (`dgr-alpha-beta-performance/v1`), sealed with
a `contract_sha256` digest over its own canonical content (the repository's existing digest
convention, shared with `meshnet_node.glm_alpha.contract`). Contents:
- `prompt_set` — four fixed prompts (`short-instruction`, `code-completion`,
`multi-step-reasoning`, `long-context-fill`) referenced by ID from every lane, so no lane can
quietly drift onto a different workload.
- `sampling` — greedy (`temperature=0`, `top_p=1`, `top_k=1`, `seed=1234`), matching
`meshnet_node.recipe_benchmark.SamplingPolicy` defaults.
- `lanes` — all four lanes named in the acceptance criteria. `controlled-safetensors` and
`whole-model-gguf` are marked `locked_elsewhere: true` and point at the pre-existing immutable
DGR-001 lock (`meshnet_node.performance_contract`, `contract_version=1`,
`ContractThresholds`) rather than re-defining or risking a conflicting duplicate. Only
`dense-distributed-gguf` and `v4-flash-distributed` are newly locked here, each with fixed
`prompt_ids`, `context_tokens`/`output_tokens` (alpha- and beta-scale for the V4 lane),
`concurrency_levels`, `hardware` (named certification-scenario topology, network class, device
class, MTP-off note), and a `metrics` list drawn from the existing
`recipe_benchmark`/`performance_contract`/`route_session_benchmark` metric vocabulary
(`ttft_p50_ms`, `decode_tokens_per_sec`, `seam_bytes`, `seam_latency_ms`, ...).
- `gain_attribution` — two disjoint metric sets, `quantization_model_fit_metrics` and
`runtime_transport_batching_kernel_metrics`, plus the rule that a speed/fit claim must cite
which axis moved it.
- `certification_scenarios``quantization` (`Q4_K_M`, `Q8_0`, `bf16-reference`) and
`stage_count` (`2-4-stage`, `10-plus-stage`) as named labels only, with an explicit rule that
no product/runtime code path may hardcode them.
- `alpha` — correctness thresholds (greedy token agreement, mean state cosine similarity,
nonfinite-tensor/fail-closed checks, no dense-attention-fallback credit) plus a `useful_speed`
block whose ratios (`1.25`/`0.75`-class, matching the already-locked DGR-001 25% convention)
carry an explicit `human_approval` sub-block (`required: true`, `approved: false`,
`approved_by: null`, `approved_at: null`). The ratio alone cannot satisfy alpha; DGR-054 must
fill in the approval against real evidence. `mtp.reserved=true`/`enabled_for_alpha=false` per
`RALPH-CONTEXT.md`. `verdicts: ["alpha", "optimize", "stop"]`.
- `beta` — adds exactly `concurrency`, `long_context`, `failure`, `sustained_throughput` axes
(16k-token long-context threshold matching the V4 lane's `beta_context_tokens`, no-silent-KV-
migration and no-synthetic-workers failure rules, 30-minute sustained-throughput floor).
`verdicts: ["beta", "targeted-optimization", "stop-rollback"]`.
- `amendment_policy` — thresholds may not be weakened/moved/reinterpreted after results are
known; a change requires a new `contract_id`/`contract_version` under human review.
- **`contract.py`** — loader/validator mirroring the proven
`meshnet_node.glm_alpha.contract` pattern: `parse_contract` recomputes the canonical-JSON SHA-256
over the document (excluding the digest field) and requires it match both the document's own
declared `contract_sha256` *and* a digest pinned independently in code
(`CONTRACT_V1_SHA256`), so neither an in-place edit nor a resealed mutation can pass silently.
Structural checks enforce all four required lanes, that the two referenced lanes actually
declare `locked_elsewhere`, that the two newly-locked lanes carry full benchmark-plan fields,
that `alpha.verdicts`/`beta.verdicts` are exactly the three-outcome sets the release gates use,
and — the one property with no analogue in `glm_alpha` — that
`alpha.useful_speed.human_approval.required` is `true`. `seal_contract()` is the only supported
way to produce a new digest, kept separate from load-time verification for the same reason
`glm_alpha` keeps it separate.
- **`__init__.py`** — re-exports the public API, documented as the contract DGR-020, DGR-044,
DGR-054, and DGR-070 are judged against.
### `tests/test_dgr_performance_contract.py` (new, 28 tests)
Deterministic, offline, GPU-free, model-download-free. Covers: packaged load and identity; digest
recomputation; all four lanes present; the two referenced lanes point at the real DGR-001 module
and its actual immutable thresholds (`min_decode_speedup == 1.25`, `max_resident_memory_ratio ==
0.75`); the two newly-locked lanes carry complete benchmark plans, fixed context/output/
concurrency; the shared prompt set and every lane's `prompt_ids`/`beta_prompt_ids` are a subset of
it; sampling is greedy; `gain_attribution`'s two metric sets are non-empty and disjoint;
certification-scenario names and rule text; **a structural test that greps every `.py` file under
`packages/node/meshnet_node` (excluding this contract's own module and data file) for the literal
strings `2-4-stage`/`10-plus-stage` and fails if any product module hardcodes them** — the concrete
form of "no product logic may hardcode them"; alpha verdicts/correctness/`human_approval`/MTP-off;
beta verdicts/axes/long-context/failure semantics; digest-mutation rejection (in-place and
resealed); missing-digest rejection; `load_contract` from an explicit path matches the packaged
load; `seal_contract` reproduces the pinned digest; amendment policy text.
### `.scratch/distributed-gguf-runtime/prd.json`
- Restored the top-level `sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/
`supersededStories` objects dropped by the pre-existing unrelated edit (see above); kept the
current `metadata.updatedAt` tooling stamp.
- Marked `DGR-019.passes = true` with `completionNotes` summarizing this outcome.
### `.scratch/distributed-gguf-runtime/issues/019-lock-alpha-and-beta-performance-contracts.md`
Regenerated via `python scripts/ralph_prd_schema.py render` to reflect `passes: true` (checked
acceptance criteria, "completed" status line, "Verified evidence" handoff line), matching the
convention DGR-017/DGR-018's issue files already use.
## Commands and results
```bash
.venv-rocm/bin/python -m pytest -q tests/test_dgr_performance_contract.py
```
```text
28 passed in 0.14s
```
```bash
.venv-rocm/bin/python -m pytest -q tests/test_ralph_prd_schema.py tests/test_dgr_performance_contract.py \
tests/test_glm_alpha_target.py tests/test_recipe_benchmark.py tests/test_route_session_benchmark.py
```
```text
270 passed in 1.04s
```
```bash
.venv-rocm/bin/python -m compileall -q packages tests
```
Exit code 0, no output (all files compile).
```bash
git diff --check
```
Exit code 0 (no whitespace errors).
```bash
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
```
```text
OK: 55 stories validated.
```
```bash
.venv-rocm/bin/python -m pytest -q tests/ -k "not integration" --ignore=tests/test_shard_runtime_harness.py
```
```text
7 failed, 1146 passed, 11 skipped, 4 deselected, 3 warnings in 261.71s (0:04:21)
```
This full sweep was launched in the background while `prd.json`/the evidence README below were
still being written, so it raced its own inputs: one of its 7 failures
(`test_ralph_prd_schema.py::test_real_backlog_passed_stories_have_completion_evidence`) was this
story's own `passes=true`/evidence-README edit landing mid-run, not a real defect — re-running
`tests/test_ralph_prd_schema.py` alone afterward, against the finalized tree, gives
`108 passed`. The other 6 failures (`test_billing_ledger.py::
test_tracker_enables_billing_with_default_db`, `test_dynamic_routing.py::
test_admin_can_replace_a_served_model_and_release_it`, `test_dynamic_routing.py::
test_models_list_does_not_duplicate_a_preset_registered_by_hf_repo`, three cache tests in
`test_real_model_backend.py`) are in files this story's `git diff` never touches (`git diff --stat
HEAD -- tests/test_billing_ledger.py tests/test_dynamic_routing.py tests/test_real_model_backend.py`
is empty) and none of them import `dgr_performance`, `performance_contract`, or `glm_alpha`; they
are pre-existing baseline defects, not regressions from this story, in the same spirit as the
known `origin/master` limitations DGR-017's evidence recorded.
## Known limitations
- `tests/test_shard_runtime_harness.py` fails to *collect* in this environment
(`ModuleNotFoundError: No module named 'grpc'`). This is a pre-existing environment gap from
DGR-024's real generated-gRPC protocol harness, not something this story touched or caused; it is
excluded from the sweep above rather than silently masked.
- Alpha's `useful_speed` ratios (`1.25`/`0.75`-class) are proposed thresholds held at the same
margin already locked for the whole-model contract (DGR-001/v1). They are locked numbers, but
`human_approval.required=true` means DGR-054 may not treat them as self-certifying from the
ratio alone — a human must approve the observed ratio against real evidence. This session did
not, and could not, supply that approval: no distributed benchmark evidence exists yet.
- `v4-flash-distributed`'s `reference_baseline` documents that a safetensors DeepSeek V4 Flash
distributed baseline may not yet be pinned (that is DGR-044's job); until then, comparisons must
fall back to `dense-distributed-gguf` runtime/transport overhead as an explicit, stated
limitation rather than a silent substitution.
- This is a specification-materialization story; per the shared quality gates, it is intentionally
left uncommitted for manual review rather than given the "one scoped story commit" other stories
get.
## Dependency handoff
DGR-020 (run the controlled whole-model baseline) consumes the DGR-001 lock referenced — not
redefined — by this contract's `controlled-safetensors`/`whole-model-gguf` lanes.
DGR-044 (pin the DeepSeek V4 Flash target contract) and DGR-054/DGR-070 (enforce the alpha/beta
gates) must load `meshnet_node.dgr_performance.load_contract()` and judge results against its
`dense-distributed-gguf`/`v4-flash-distributed` lanes and `alpha`/`beta` sections without changing
any threshold. DGR-054 specifically must populate `alpha.useful_speed.human_approval`
(`approved`/`approved_by`/`approved_at`) as part of publishing its verdict — a satisfied ratio
without a filled-in approval is not alpha certification. Any amendment must open a new
`contract_id`/`contract_version` under human review per `amendment_policy`; this document and its
digest are not editable in place.

View File

@@ -0,0 +1,243 @@
# DGR-020 evidence — run the controlled whole-model GGUF baseline
**Completed:** 2026-07-22
**Branch:** `ralph/distributed-gguf-runtime`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
**Dependency:** DGR-019 (`evidence/DGR-019/README.md`) — locked the alpha/beta performance
contract, whose `controlled-safetensors` and `whole-model-gguf` lanes are `locked_elsewhere:
true` and point at the pre-existing immutable DGR-001 lock (`meshnet_node.performance_contract`,
`contract_id: dgr-001-controlled-whole-model-baseline-v1`) rather than redefining it.
## Objective
Per `.scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md`:
execute the exact locked safetensors and whole-model llama.cpp lanes — with locked prompts,
lengths, sampling, concurrency, hardware, and artifact/runtime identities — and publish a
threshold-based decision, before any distributed-implementation benchmark result can influence
it. Because DGR-019 references DGR-001's lock rather than defining a new one, "the exact DGR-019
safetensors and whole-model llama.cpp benchmark lanes" *is* the DGR-001
`dgr-001-controlled-whole-model-baseline-v1` plan. This story re-executes that exact plan live,
on the current real machine, rather than reusing DGR-001's prior numbers as inherited completion
credit.
## Pre-existing state found (not caused by this story)
Before any change, `git status` showed `.scratch/distributed-gguf-runtime/prd.json` already
modified relative to `HEAD` (`47bad0b`). Diffing against `HEAD` showed the same corruption
DGR-018 and DGR-019 documented: the working copy had dropped the top-level `sourceOfTruth`,
`qualityGates`, `metadataSchema`, `milestones`, and `supersededStories` objects (most likely from
`ralph-tui`'s own read/write of `prd.json`, which round-trips only the fields it models). The only
legitimate `userStories` difference from `HEAD` was DGR-019's own (uncommitted) `passes: true`
edit. Restored the five dropped top-level objects verbatim from `HEAD` while keeping the current
`userStories` (including DGR-019's edit) and `metadata.updatedAt`. `tests/test_ralph_prd_schema.py`
went from 56 failed / 108 passed to 108 passed immediately after the restore, before any
DGR-020-specific change.
## Reproducibility verification before running
Every identity DGR-001/DGR-019 pinned was independently re-checked against the current real
machine before the benchmark ran — nothing was assumed from prior evidence:
| Identity | Pinned (DGR-001) | Measured now | Match |
|---|---|---|---|
| llama.cpp commit | `e920c523e3b8a0163fe498af5bf90df35ff51d25` | `e920c523e3b8a0163fe498af5bf90df35ff51d25` | yes |
| `llama-server` SHA-256 | `fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd` | same | yes |
| BF16 GGUF artifact SHA-256 | `e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862` | same | yes |
| Q4_K_M GGUF artifact SHA-256 | `a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5` | same | yes |
| Torch / Transformers versions | `2.10.0+rocm7.13.0a20260513` / `5.13.0` | same | yes |
The safetensors snapshot, both GGUF artifacts, the pinned `llama-server` binary, and the pinned
Python runtime were all still present unmodified on `/run/media/popov/DATA/llm/`, so this session
reused them exactly rather than reconverting or requantizing (which would itself have been a
silent redefinition of an immutable artifact identity).
## Real results — fresh run on real hardware
`.scratch/distributed-gguf-runtime/evidence/DGR-020/benchmark-config.json` and
`performance-contract.json` are byte-identical copies of DGR-001's (same `plan_sha256`
`efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570` and `config_sha256`
`00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3`), so this is the same plan,
not a new one.
```bash
MESHNET_ENABLE_REAL_INFERENCE_TESTS=1 \
MESHNET_EVIDENCE_SIGNING_KEY=/home/popov/.config/neuron-tai/keys/dgr-001-evidence-ed25519.pem \
PYTHONPATH=packages/node .venv-rocm/bin/python -m meshnet_node.recipe_benchmark \
--config .scratch/distributed-gguf-runtime/evidence/DGR-020/benchmark-config.json \
--json-out .scratch/distributed-gguf-runtime/evidence/DGR-020/results.json \
--summary-out .scratch/distributed-gguf-runtime/evidence/DGR-020/results.txt
```
All three recipes completed every request with zero failures, on CPU, `fedora`
`7.0.14-101.fc43.x86_64`, 32 logical CPUs:
| Metric | Transformers BF16 (ref) | llama.cpp BF16 | llama.cpp Q4_K_M | DGR-001 (prior run, same plan) |
|---|---:|---:|---:|---|
| Decode tok/s, c=1 | 50.8 | 102.5 | 213.1 | 40.8 / 98.5 / 207.7 |
| Aggregate decode tok/s, c=4 | 48.8 | 218.1 | 235.7 | 46.5 / 222.8 / 195.7 |
| TTFT p50, c=1 | 32.9 ms | 15.1 ms | 17.3 ms | 40.0 / 15.1 / 21.6 ms |
| Peak resident memory, c=1 | 1.93 GB | 1.11 GB | 0.54 GB | 1.94 / 1.11 / 0.54 GB |
| Artifact size | 1.00 GB | 0.99 GB | 0.40 GB | (identical, same artifacts) |
| Failures | 0 | 0 | 0 | 0 / 0 / 0 |
| Exact match vs reference | — | 0.3333 | 0.00 (advisory) | 0.3333 |
| Mean similarity vs reference | — | 0.9471 | 0.456 (advisory) | 0.9471 |
Per-recipe measurements against the reference (`baseline.json`, `contract-evaluation.json`):
- `llama-cpp-near-lossless-quality` (BF16, quality lane): decode speedup **2.02x**, aggregate
throughput speedup (c=4) **4.47x**, resident-memory ratio **0.574x**, TTFT ratio **0.459x**
but `quality_pass: false` (exact match 0.33 < required 0.90).
- `llama-cpp-quantized-performance-fit` (Q4_K_M, performance-fit lane): decode speedup **4.19x**,
aggregate throughput speedup (c=4) **4.83x**, resident-memory ratio **0.280x**, artifact-size
ratio **0.398x**, TTFT ratio **0.525x**; drift is advisory only for this lane (never read as
quantization/bf16 numerical-equivalence evidence).
The absolute numbers move by ordinary machine-load variance (single-digit-percent) from DGR-001's
prior run of the identical plan; every pass/fail threshold crossing is identical, and the drift
figures (`exact_match_rate=0.3333`, `mean_similarity=0.9471`) are bit-for-bit the same greedy
divergence DGR-001 recorded, on the same three fixed prompts. This is a genuine independent
reproduction, not a copy: `results.json`'s `provenance.run_id`
(`59b12968-c5d0-4391-90f4-0cd2aff77b21`), `started_at`/`completed_at` timestamps, and Ed25519
`signature` are all freshly generated by this session's run, signed with the same DGR-001 evidence
key (`signer_public_key_sha256` `8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de`,
matching the sole active entry in `../../trusted-evidence-signers.json`).
## Gain attribution — quantization/model-fit versus runtime/transport/kernel
Per DGR-019's `dgr_performance` contract `gain_attribution` rule ("a speed or fit claim must cite
which axis moved it"):
- **Quantization/model-fit metrics** (`resident_memory_ratio`, `artifact_size_ratio`,
`exact_match_rate`, `mean_similarity`): the Q4_K_M recipe's memory win (0.280x) and size win
(0.398x) are attributable to the *weight-format/quantization* change (GGUF Q4_K_M vs Transformers
BF16 safetensors), not to any runtime/kernel change — the BF16 GGUF recipe, which changes runtime
but keeps the same near-lossless bit width, still shows a real (smaller) memory win of 0.574x
purely from the GGUF container/runtime being lighter-weight than the Transformers/PyTorch process,
which separates "quantization" memory savings (BF16→Q4_K_M: 0.574x→0.280x) from "runtime/format"
memory savings (safetensors→BF16 GGUF: 1.0x→0.574x). The quality-lane failure
(`exact_match_rate=0.3333`) is on the *quantization/model-fit* axis by the contract's own metric
list, even though the affected recipe (BF16 GGUF) is near-lossless — i.e. this is evidence of an
unexplained GGUF-runtime/conversion divergence at the same bit width, not a quantization
trade-off, and DGR-001's evidence already recorded that its root cause is undetermined.
- **Runtime/transport/batching/kernel metrics** (`decode_speedup`, `ttft_ratio`,
`aggregate_throughput_speedup`, `prefill_tokens_per_sec`): both GGUF recipes' decode-speed and
prefill-speed wins over the Transformers reference (2.02x/4.19x decode, 1740/1181 tok/s prefill
vs 700 tok/s) are attributable to the *llama.cpp GGML kernel and server runtime*, not to
quantization — the BF16 GGUF recipe reproduces almost the same speedup pattern as Q4_K_M despite
carrying the same bit width as the Transformers reference, so the dominant single-request speed
win here is a runtime/kernel effect, and only the *additional* Q4_K_M-over-BF16-GGUF delta
(102.5→213.1 tok/s decode, ~2.08x) is attributable to quantization on top of that runtime effect.
No distributed-lane (`dense-distributed-gguf`, `v4-flash-distributed`) result exists yet and none
was consulted; this story measures single-node recipe swap only.
## Failed / unavailable lanes
None. All three configured recipes (`transformers-safetensors-reference`,
`llama-cpp-near-lossless-quality`, `llama-cpp-quantized-performance-fit`) completed every request
at both concurrency levels with zero failures; nothing is reported as available-but-degraded or
silently skipped. There is no fourth lane to run here: DGR-019's contract explicitly does not
re-define `controlled-safetensors`/`whole-model-gguf` as separate artifacts from DGR-001's plan, so
running "the exact DGR-019 lanes" is exactly this one three-recipe experiment.
## Decision
`contract-evaluation.json` (evaluated with the unmodified, immutable
`meshnet_node.performance_contract` v1 thresholds — `min_decode_speedup=1.25`,
`max_ttft_ratio=1.25`, `min_aggregate_throughput_speedup=1.25`, `max_resident_memory_ratio=0.75`,
`min_quality_exact_match_rate=0.90`, `min_quality_mean_similarity=0.97`, `max_failure_rate=0.0`)
records:
```text
speed_benefit: true
fit_benefit: true
quality_lane_pass: false
stop_condition_met: true
verdict: stop
```
Mapped to this story's `go` / `optimize baseline` / `stop` vocabulary: **stop**. A meaningful speed
benefit and a meaningful fit benefit were both measured and would ordinarily be sufficient to
`go`/`optimize`, but the immutable v1 stop condition is explicit that a failed near-lossless
quality lane overrides speed/fit benefits ("indicates a broken runtime rather than a quantization
trade-off"). This decision uses only the locked v1 thresholds and this session's freshly measured
metrics; no threshold was changed, and no distributed-implementation result (DGR-024's gRPC
harness or any other distributed-lane evidence) was read or ingested to produce it.
This reproduces DGR-001's original `stop` verdict on the same plan on the same real machine,
confirming that verdict is stable over time and not an artifact of a single run.
## Limitations
- This is a **0.5B CPU baseline** (`Qwen/Qwen2.5-0.5B-Instruct`), the same generic model DGR-001
and DGR-019's `locked_elsewhere` reference use — not DeepSeek V4 Flash. DGR-019's evidence
already recorded that a DeepSeek V4 Flash `controlled-safetensors`/`whole-model-gguf` baseline is
not yet pinned; that is separate future work (see DGR-019's `v4-flash-distributed.reference_
baseline` note), not something this story's acceptance criteria ask it to create — it asks only
to run the exact already-locked lanes, which are this DGR-001 plan.
- The `whole-model-gguf` quality-lane exact-match divergence (0.33 vs 0.90 required) reproduces
identically and remains unexplained; this story does not diagnose it further beyond confirming
it reproduces (DGR-001's `quality-parity-diagnosis.md` documents the CPU-vs-ROCm split already
known).
- Absolute timings are single-developer-machine measurements with ordinary run-to-run variance;
the locked ratios/ratios-vs-threshold crossings are the durable evidence, not the raw absolute
tok/s figures.
- No new GPU (ROCm) diagnostic was re-run in this session — DGR-001's existing GPU diagnostic is
cited as prior evidence only; it uses a distinct signed `run_configured_gpu_diagnostic/v1`
producer that the v1 evaluator does not accept, so it cannot itself change the `stop` verdict
above.
## Files changed
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/benchmark-config.json` (new) — byte-identical
copy of DGR-001's locked plan.
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/performance-contract.json` (new) —
byte-identical copy of DGR-001's immutable v1 thresholds.
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/results.json` / `results.txt` (new) — raw
signed real evidence from this session's fresh run.
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/baseline.json` / `contract-evaluation.json`
(new) — distilled baseline and fail-closed v1 verdict for this session's run.
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/README.md` (new, this file).
- `.scratch/distributed-gguf-runtime/prd.json` — restored the dropped top-level
`sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/`supersededStories` objects (see
above); marked `DGR-020.passes = true` with `completionNotes`.
- `.scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md`
regenerated via `python scripts/ralph_prd_schema.py render` to reflect `passes: true`.
No source or test files under `packages/` or `tests/` were changed by this story.
## Commands and results
```bash
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
```
```text
OK: 55 stories validated.
```
```bash
.venv-rocm/bin/python -m pytest -q tests/test_recipe_benchmark.py tests/test_dgr_performance_contract.py tests/test_ralph_prd_schema.py
```
```text
164 passed in 0.69s
```
```bash
.venv-rocm/bin/python -m compileall -q packages tests
```
Exit code 0, no output (all files compile).
```bash
git diff --check
```
Exit code 0 (no whitespace errors).
## Dependency handoff
DGR-054 (enforce the alpha gate) may cite this evidence when it fills in
`alpha.useful_speed.human_approval` — this is fresh, independently-collected, signed real-hardware
evidence that the `controlled-safetensors`/`whole-model-gguf` v1 contract still holds `stop` on the
current machine, immediately before any distributed-lane result exists, but it is a 0.5B CPU
baseline, not the DeepSeek V4 Flash target; DGR-044 must still pin the V4 Flash reference baseline
separately before DGR-054/DGR-070 can judge `dense-distributed-gguf`/`v4-flash-distributed` against
it. No threshold in either `meshnet_node.performance_contract` or `meshnet_node.dgr_performance`
was changed by this story.

View File

@@ -0,0 +1,169 @@
{
"artifact_sha256": {
"llama-cpp-near-lossless-quality": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
"llama-cpp-quantized-performance-fit": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5",
"transformers-safetensors-reference": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6"
},
"backend_detail": {
"llama-cpp-near-lossless-quality": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
"llama-cpp-quantized-performance-fit": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
"transformers-safetensors-reference": "torch 2.10.0+rocm7.13.0a20260513; dtype bfloat16; device cpu; intra-op threads 16"
},
"evidence_class": "local-real",
"host": {
"accelerator_name": "Radeon 8060S Graphics",
"accelerator_runtime": "7.13.26183",
"benchmark_lane": "cpu-controlled-baseline",
"converter_sha256": "c819f18fb22927b49fabc3b35d1c9e21ee638b3817eccd1bd4efbcc7116eeb4d",
"cpu_count": 32,
"cuda_available": true,
"hostname": "fedora",
"llama_cpp_commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
"llama_cpp_version": "9991",
"llama_server_identities": {
"/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server": {
"sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
"version": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64"
}
},
"llama_server_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
"platform": "Linux-7.0.14-101.fc43.x86_64-x86_64-with-glibc2.42",
"python": "3.12.13",
"quantizer_sha256": "bd0cc8c7be6d48aad4755b31062e0e59a887cbadd43dbb8771853d5858bb198f",
"torch_version": "2.10.0+rocm7.13.0a20260513",
"transformers_version": "5.13.0"
},
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
"plan_sha256": "efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570",
"provenance": {
"completed_at": "2026-07-22T05:52:30.445799Z",
"config_sha256": "00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3",
"producer": "meshnet_node.recipe_drivers.run_configured_benchmark/v1",
"run_id": "59b12968-c5d0-4391-90f4-0cd2aff77b21",
"schema_version": 1,
"signature": "aExtG1Y0fWFaqlKEtUOOpXZrganVAxbLvpov2WVgm19eNJ50VheeI7CuRhlWx4SJX9OFto2WuLaVPhjwSA88Cw==",
"signature_algorithm": "ed25519",
"signer_public_key_sha256": "8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de",
"started_at": "2026-07-22T05:51:36.511891Z"
},
"recipe_runtime": {
"llama-cpp-near-lossless-quality": {
"device": "cpu",
"runtime": "llama.cpp-9991-e920c523",
"weight_format": "gguf",
"weight_quantization": "bfloat16"
},
"llama-cpp-quantized-performance-fit": {
"device": "cpu",
"runtime": "llama.cpp-9991-e920c523",
"weight_format": "gguf",
"weight_quantization": "Q4_K_M"
},
"transformers-safetensors-reference": {
"device": "cpu",
"runtime": "transformers-5.13.0",
"weight_format": "safetensors",
"weight_quantization": "bfloat16"
}
},
"recipes": {
"llama-cpp-near-lossless-quality": {
"artifact_bytes": 994156448,
"available": true,
"concurrency": {
"1": {
"aggregate_decode_tokens_per_sec": 89.2873,
"decode_tokens_per_sec": 102.5344,
"failures": 0,
"latency_p50_ms": 316.647,
"latency_p95_ms": 374.8515,
"peak_rss_bytes": 1110106112,
"peak_vram_bytes": 0,
"prefill_tokens_per_sec": 1740.0213,
"ttft_p50_ms": 15.067,
"ttft_p95_ms": 65.191
},
"4": {
"aggregate_decode_tokens_per_sec": 218.1128,
"decode_tokens_per_sec": 80.0623,
"failures": 0,
"latency_p50_ms": 403.9781,
"latency_p95_ms": 767.6557,
"peak_rss_bytes": 1139265536,
"peak_vram_bytes": 0,
"prefill_tokens_per_sec": 1064.6179,
"ttft_p50_ms": 36.611,
"ttft_p95_ms": 178.801
}
},
"device": "cpu",
"lane": "quality"
},
"llama-cpp-quantized-performance-fit": {
"artifact_bytes": 397807520,
"available": true,
"concurrency": {
"1": {
"aggregate_decode_tokens_per_sec": 149.8675,
"decode_tokens_per_sec": 213.1452,
"failures": 0,
"latency_p50_ms": 161.7164,
"latency_p95_ms": 282.5491,
"peak_rss_bytes": 541663232,
"peak_vram_bytes": 0,
"prefill_tokens_per_sec": 1181.0842,
"ttft_p50_ms": 17.252,
"ttft_p95_ms": 130.529
},
"4": {
"aggregate_decode_tokens_per_sec": 235.6963,
"decode_tokens_per_sec": 94.7604,
"failures": 0,
"latency_p50_ms": 373.7211,
"latency_p95_ms": 759.3151,
"peak_rss_bytes": 571027456,
"peak_vram_bytes": 0,
"prefill_tokens_per_sec": 567.7335,
"ttft_p50_ms": 42.086,
"ttft_p95_ms": 312.645
}
},
"device": "cpu",
"lane": "performance-fit"
},
"transformers-safetensors-reference": {
"artifact_bytes": 999586347,
"available": true,
"concurrency": {
"1": {
"aggregate_decode_tokens_per_sec": 44.4625,
"decode_tokens_per_sec": 50.8327,
"failures": 0,
"latency_p50_ms": 701.9146,
"latency_p95_ms": 776.2706,
"peak_rss_bytes": 1933221888,
"peak_vram_bytes": 0,
"prefill_tokens_per_sec": 699.7553,
"ttft_p50_ms": 32.8569,
"ttft_p95_ms": 173.7161
},
"4": {
"aggregate_decode_tokens_per_sec": 48.849,
"decode_tokens_per_sec": 13.4779,
"failures": 0,
"latency_p50_ms": 2503.1601,
"latency_p95_ms": 2600.6307,
"peak_rss_bytes": 2170908672,
"peak_vram_bytes": 0,
"prefill_tokens_per_sec": 264.5822,
"ttft_p50_ms": 95.7502,
"ttft_p95_ms": 425.4973
}
},
"device": "cpu",
"lane": "quality"
}
},
"reference_recipe_id": "transformers-safetensors-reference"
}

View File

@@ -0,0 +1,118 @@
{
"artifact_storage_root": "/run/media/popov/DATA/llm",
"evidence_class": "local-real",
"host": {
"benchmark_lane": "cpu-controlled-baseline",
"llama_cpp_commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
"llama_cpp_version": "9991",
"llama_server_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
"converter_sha256": "c819f18fb22927b49fabc3b35d1c9e21ee638b3817eccd1bd4efbcc7116eeb4d",
"quantizer_sha256": "bd0cc8c7be6d48aad4755b31062e0e59a887cbadd43dbb8771853d5858bb198f",
"transformers_version": "5.13.0"
},
"plan": {
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
"prompts": [
{
"id": "short-fact",
"text": "The capital of France is",
"context_class": "short"
},
{
"id": "medium-code",
"text": "Complete this Python function without commentary:\n\ndef fibonacci(n):\n \"\"\"Return the nth Fibonacci number for n >= 0.\"\"\"\n",
"context_class": "medium"
},
{
"id": "long-summary",
"text": "A distributed inference service divides a transformer across consumer machines. The tracker owns admission, routing, cancellation, accounting, and telemetry, while workers own only model execution. Every request carries an immutable model identity and revision. Workers must reject incompatible protocol versions and resource demands before allocating large buffers. Activation tensors are chunked, checksummed, bounded by negotiated limits, and propagated with explicit flow-control credits. A caller may disconnect at any time, so cancellation must release queued work, in-flight transfers, and cache reservations without double billing. Retries can occur after network failures, requiring idempotent request identifiers and deterministic completion accounting. The system keeps the existing safetensors path as a correctness reference while a native GGUF path is measured. Benchmarks compare the same prompts, output lengths, sampling policy, device, and concurrency, and they separate near-lossless quality checks from quantized speed and fit claims. Summarize the design priorities in three concise bullet points.",
"context_class": "long"
}
],
"sampling": {
"temperature": 0.0,
"top_p": 1.0,
"top_k": 1,
"seed": 1234,
"max_output_tokens": 32
},
"concurrency_levels": [1, 4],
"repeats": 3,
"warmup_requests": 2
},
"recipes": [
{
"id": "transformers-safetensors-reference",
"runtime": "transformers-5.13.0",
"weight_format": "safetensors",
"weight_quantization": "bfloat16",
"lane": "quality",
"device": "cpu",
"artifact_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
"artifact_sha256": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
"is_reference": true,
"notes": "artifact_sha256 is the deterministic digest of every snapshot path and file byte",
"driver": {
"type": "transformers",
"model_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
"device": "cpu",
"dtype": "bfloat16",
"threads": 16
}
},
{
"id": "llama-cpp-near-lossless-quality",
"runtime": "llama.cpp-9991-e920c523",
"weight_format": "gguf",
"weight_quantization": "bfloat16",
"lane": "quality",
"device": "cpu",
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
"artifact_sha256": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
"is_reference": false,
"notes": "Converted directly from the exact mounted safetensors revision while preserving BF16 weights with pinned llama.cpp",
"driver": {
"type": "llama-cpp-server",
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server",
"binary_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
"device": "cpu",
"threads": 16,
"n_parallel": 4,
"context_per_slot": 512,
"n_gpu_layers": 0
}
},
{
"id": "llama-cpp-quantized-performance-fit",
"runtime": "llama.cpp-9991-e920c523",
"weight_format": "gguf",
"weight_quantization": "Q4_K_M",
"lane": "performance-fit",
"device": "cpu",
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
"artifact_sha256": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5",
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
"is_reference": false,
"notes": "Quantized from the exact-revision F16 GGUF with pinned llama-quantize",
"driver": {
"type": "llama-cpp-server",
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server",
"binary_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
"device": "cpu",
"threads": 16,
"n_parallel": 4,
"context_per_slot": 512,
"n_gpu_layers": 0
}
}
]
}

View File

@@ -0,0 +1,71 @@
{
"contract_version": 1,
"fit_benefit": true,
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
"quality_lane_pass": false,
"rationale": [
"the near-lossless quality lane failed: the GGUF runtime disagrees with the safetensors reference beyond what near-lossless weights can explain",
"a meaningful speed benefit was measured",
"a meaningful fit benefit was measured"
],
"recipes": [
{
"comparable": true,
"failures": 0,
"fit_benefit": false,
"incomparable_reason": "",
"lane": "quality",
"measurements": {
"aggregate_concurrency": 4,
"aggregate_throughput_speedup": 4.465,
"artifact_size_ratio": 0.9946,
"artifact_size_win": false,
"compared_prompts": 3,
"decode_speedup": 2.0171,
"exact_match_rate": 0.3333,
"expected_prompts": 3,
"failure_rate": 0.0,
"mean_similarity": 0.9471,
"resident_memory_ratio": 0.5742,
"ttft_ratio": 0.4586
},
"quality_pass": false,
"reasons": [
"single-request decode 2.02x reference (>= 1.25x) at TTFT ratio 0.46",
"aggregate throughput at concurrency 4 is 4.46x reference (>= 1.25x)",
"peak resident memory is 0.57x reference (<= 0.75x)",
"quality lane exact-match 0.33 / similarity 0.947 versus the reference (fail)"
],
"recipe_id": "llama-cpp-near-lossless-quality",
"speed_benefit": false
},
{
"comparable": true,
"failures": 0,
"fit_benefit": true,
"incomparable_reason": "",
"lane": "performance-fit",
"measurements": {
"aggregate_concurrency": 4,
"aggregate_throughput_speedup": 4.825,
"artifact_size_ratio": 0.398,
"artifact_size_win": true,
"decode_speedup": 4.1931,
"failure_rate": 0.0,
"resident_memory_ratio": 0.2802,
"ttft_ratio": 0.5251
},
"quality_pass": null,
"reasons": [
"single-request decode 4.19x reference (>= 1.25x) at TTFT ratio 0.53",
"aggregate throughput at concurrency 4 is 4.83x reference (>= 1.25x)",
"peak resident memory is 0.28x reference (<= 0.75x)"
],
"recipe_id": "llama-cpp-quantized-performance-fit",
"speed_benefit": true
}
],
"speed_benefit": true,
"stop_condition_met": true,
"verdict": "stop"
}

View File

@@ -0,0 +1,87 @@
{
"schema_version": 1,
"contract_version": 1,
"locked_at": "2026-07-13T00:00:00Z",
"locked_by": "DGR-001",
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
"thresholds": {
"min_decode_speedup": 1.25,
"max_ttft_ratio": 1.25,
"min_aggregate_throughput_speedup": 1.25,
"max_resident_memory_ratio": 0.75,
"max_artifact_size_ratio": 0.6,
"min_quality_exact_match_rate": 0.9,
"min_quality_mean_similarity": 0.97,
"max_failure_rate": 0.0
},
"baseline": {
"status": "pending-real-evidence",
"required_evidence_class": "local-real",
"required_recipes": [
"transformers-safetensors-reference",
"llama-cpp-near-lossless-quality",
"llama-cpp-quantized-performance-fit"
],
"required_concurrency_levels": [
1,
4
],
"required_controlled_variables": [
"model architecture",
"model revision",
"machine and device",
"formatted prompts and context lengths",
"output length and greedy sampling policy"
],
"required_plan_sha256": "efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570",
"minimum_prompt_count": 3,
"minimum_repeats": 3,
"minimum_output_tokens": 32,
"required_device": "cpu",
"required_config_sha256": "00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3",
"required_signer_public_key": "zQ/qRMwF/ydazzaxEI24Xvnrl5bZxzw16JYpP0bfRuI=",
"required_artifact_sha256": {
"transformers-safetensors-reference": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
"llama-cpp-near-lossless-quality": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
"llama-cpp-quantized-performance-fit": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5"
},
"required_recipe_runtime": {
"transformers-safetensors-reference": {
"runtime": "transformers-5.13.0",
"weight_format": "safetensors",
"weight_quantization": "bfloat16",
"device": "cpu"
},
"llama-cpp-near-lossless-quality": {
"runtime": "llama.cpp-9991-e920c523",
"weight_format": "gguf",
"weight_quantization": "bfloat16",
"device": "cpu"
},
"llama-cpp-quantized-performance-fit": {
"runtime": "llama.cpp-9991-e920c523",
"weight_format": "gguf",
"weight_quantization": "Q4_K_M",
"device": "cpu"
}
},
"required_backend_detail": {
"transformers-safetensors-reference": "torch 2.10.0+rocm7.13.0a20260513; dtype bfloat16; device cpu; intra-op threads 16",
"llama-cpp-near-lossless-quality": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
"llama-cpp-quantized-performance-fit": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0"
},
"required_host_identity": {
"python": "3.12.13",
"torch_version": "2.10.0+rocm7.13.0a20260513",
"transformers_version": "5.13.0",
"llama_server_identities": {
"/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server": {
"sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
"version": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64"
}
}
}
},
"stop_condition": "Stop the native llama.cpp/GGUF track when, on the same machine and device as the Transformers/safetensors reference and under this plan, no performance-fit GGUF recipe delivers either a meaningful speed benefit (>=25% higher single-request decode tokens/sec without a >25% worse TTFT, or >=25% higher aggregate throughput under concurrency) or a meaningful fit benefit (>=25% lower peak resident memory), or when the near-lossless quality lane fails, which indicates a broken runtime rather than a quantization trade-off.",
"notes": "Quantized performance-fit output drift is reported as advisory only. It is not numerical-equivalence evidence. DGR-014 consumes this immutable v1 contract. Non-synthetic evidence must be Ed25519-signed by the pinned key and match the exact locked config, artifacts, runtimes, backends, and host runtime identity."
}

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,10 @@
Recipe benchmark dgr-001-controlled-whole-model-baseline-v1 (local-real)
model Qwen/Qwen2.5-0.5B-Instruct@7ae557604adf67be50417f59c2c2f167def9a775
transformers-safetensors-reference [quality ] c= 1 ttft p50/p95 32.9/ 173.7 ms; prefill 699.8 tok/s; decode 50.8 tok/s; aggregate 44.5 tok/s; rss 1.93 GB; vram 0.00 GB; artifact 1.00 GB; failures 0
transformers-safetensors-reference [quality ] c= 4 ttft p50/p95 95.8/ 425.5 ms; prefill 264.6 tok/s; decode 13.5 tok/s; aggregate 48.8 tok/s; rss 2.17 GB; vram 0.00 GB; artifact 1.00 GB; failures 0
llama-cpp-near-lossless-quality [quality ] c= 1 ttft p50/p95 15.1/ 65.2 ms; prefill 1740.0 tok/s; decode 102.5 tok/s; aggregate 89.3 tok/s; rss 1.11 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
llama-cpp-near-lossless-quality [quality ] c= 4 ttft p50/p95 36.6/ 178.8 ms; prefill 1064.6 tok/s; decode 80.1 tok/s; aggregate 218.1 tok/s; rss 1.14 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
llama-cpp-quantized-performance-fit [performance-fit ] c= 1 ttft p50/p95 17.3/ 130.5 ms; prefill 1181.1 tok/s; decode 213.1 tok/s; aggregate 149.9 tok/s; rss 0.54 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
llama-cpp-quantized-performance-fit [performance-fit ] c= 4 ttft p50/p95 42.1/ 312.6 ms; prefill 567.7 tok/s; decode 94.8 tok/s; aggregate 235.7 tok/s; rss 0.57 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
drift llama-cpp-near-lossless-quality vs transformers-safetensors-reference exact 0.33; similarity 0.947 (gated)
drift llama-cpp-quantized-performance-fit vs transformers-safetensors-reference exact 0.00; similarity 0.456 (advisory)

View File

@@ -0,0 +1,101 @@
# DGR-021 evidence — versioned named-tensor activation envelope
**Completed:** 2026-07-17
**Branch:** `distributed-gguf-runtime`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
**Dependency:** DGR-018 (`evidence/DGR-018/README.md`) — canonical backlog schema / issue projection contract
## Objective
Establish the backend-neutral activation envelope used by direct and relayed Shard traffic, with stable versioning, named tensors, bounded fragmentation, checksum validation, and reserved extensibility for future state.
## Changes
### `packages/node/meshnet_node/protocol.py` (new)
Added a self-contained activation-envelope module with:
- `SCHEMA_NAME = "meshnet.activation-stream"` and `SCHEMA_VERSION = 1`
- `TensorFragment`
- bounded byte fragments with offset, compression tag, checksum, and extension preservation
- deterministic `to_dict()` / `from_dict()` round-trip
- `NamedTensor`
- named tensor metadata: `name`, `shape`, `dtype`, `byte_order`, `compression`, `checksum`, `fragments`
- fragmentation via `from_bytes(..., max_fragment_bytes=...)`
- checksum validation over reconstructed tensor bytes
- unknown-field preservation via `extensions`
- `ActivationEnvelope`
- top-level fields for `request_id`, `work_id`, `route_session`, `route_epoch`, `shard_start`, `effective_start`, `phase`, `position`, and `idempotency_step`
- reserved extension fields for `token_id_sideband`, `architecture_state`, `recurrent_state`, and `mtp`
- deterministic canonical serialization (`to_bytes`) and round-trip parsing (`from_bytes`)
- size-limit enforcement (`to_bytes(max_bytes=...)`)
- conversion from a live `TensorPayload` into the envelope and back again
### `packages/node/meshnet_node/model_backend.py`
Extended `TensorPayload` with envelope conversion helpers:
- `TensorPayload.to_envelope(...)`
- `TensorPayload.from_envelope(...)`
These keep the existing activation payload interface intact while exposing the new versioned envelope as the shared protocol layer.
### `tests/test_activation_envelope.py` (new)
Added focused deterministic tests covering:
- deterministic envelope serialization and round-trip parsing
- tensor fragmentation and checksum validation
- unknown-field preservation at both envelope and tensor levels
- size-limit rejection
- `TensorPayload` ↔ envelope round-trip
### `.scratch/distributed-gguf-runtime/prd.json`
Marked `DGR-021.passes = true` and added completion notes recording the envelope implementation and verification commands.
## Commands and results
```bash
pytest -q tests/test_activation_envelope.py
```
```text
5 passed in 0.06s
```
```bash
pytest -q tests/test_activation_envelope.py tests/test_kv_cache_distributed.py -k 'session_is_stable_and_decode_payloads_are_single_token or large_prefill_activation_survives_zstd_compressed_hop'
```
```text
.. [100%]
2 passed, 21 deselected in 1.84s
```
```bash
python3 -m compileall packages/node/meshnet_node tests/test_activation_envelope.py
```
```text
Listing 'packages/node/meshnet_node'...
Listing 'packages/node/meshnet_node/native_protocol'...
Compiling 'tests/test_activation_envelope.py'...
```
```bash
git diff --check
```
```text
No whitespace errors
```
## Limitations
- The envelope is implemented as a canonical deterministic JSON contract with dataclasses and conversion hooks, not generated `.proto` classes. The environment had `protobuf` available but not the `grpc_tools` generation toolchain, so I did not materialize a compiled proto artifact here.
- The direct/relayed HTTP/WebSocket transports remain byte-oriented; the envelope is the shared structured contract layered above those transports.
## Dependency handoff
DGR-022 and later shard-control stories can reuse the envelope contract and its `TensorPayload` conversion hooks as the stable activation metadata layer. Future work that requires generated protobuf code can replace the JSON serialization with a generated wire codec without changing the top-level field contract defined here.

View File

@@ -0,0 +1,46 @@
# DGR-022 evidence — Shard lifecycle and structured status RPC contract
**Completed:** 2026-07-17
**Branch:** `ralph/distributed-gguf-runtime`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
## Outcome
Implemented the versioned, backend-neutral lifecycle/status contract consumed by a future generated gRPC binding. The contract keeps Meshnet routing, identity, authentication policy, billing, and llama.cpp ownership outside the worker contract.
## Implemented
- `packages/node/meshnet_node/shard_lifecycle.py`
- capability, health, session, cancellation, release, and metrics RPC names
- schema version negotiation and fail-closed unsupported-version handling
- structured status/error taxonomy with retryability and details
- lifecycle state machine for prefill/decode/cancel/release transitions
- monotonic idempotency-step enforcement and duplicate rejection
- bounded frame/byte flow control with cancellation-aware waits
- explicit cache expectation/result types
- deadline policy and TLS/auth transport hooks
- deterministic contract serialization round-trip
- `tests/test_shard_lifecycle.py`
- contract round-trip and RPC coverage
- unsupported-version rejection
- malformed transition and idempotency rejection
- cancellation/release behavior
- bounded flow-control behavior
- TLS hook and incomplete-contract fail-closed behavior
## Verification
```text
$ PYTHONPATH=packages/node pytest -q tests/test_shard_lifecycle.py tests/test_activation_envelope.py
17 passed in 0.10s
```
The existing DGR-021 activation-envelope tests remain green alongside DGR-022.
## Scope limitation
This story defines the lifecycle/status contract only. Generated Python/C++ protobuf bindings and the concrete `shard_runtime.proto` generation pipeline are DGR-023 and remain separate.
## Dependency handoff
DGR-023 may consume the RPC names, status taxonomy, version identity, deadlines, flow-control limits, and TLS/auth hooks when the canonical `.proto` schema and toolchain are provisioned.

View File

@@ -0,0 +1,126 @@
# DGR-023 evidence — reproducible Python and C++ protobuf/gRPC generation
**Status:** complete after controller verification and independent-review repairs on 2026-07-17.
**Authority:** live Gitea issue #7. The local PRD is a secondary projection.
## Implemented contract
- Python generation requires exactly `grpcio-tools==1.82.1`; the generator checks installed distribution metadata and rejects missing or different versions with an actionable exact install command.
- The C++ bootstrap builds one ignored toolchain prefix from exact inputs:
- Protobuf release `33.1` (`protobuf-config` version `33.1.0`);
- Abseil release `20250814.1`;
- gRPC C++ `1.82.1` at commit `acccf84c0df20487d64101f528e5d426541ca4e5`;
- gRPC's exact-commit submodules for c-ares, RE2, OpenSSL, and zlib.
- Protobuf is configured with local dependencies only after the exact Abseil build. gRPC uses the installed Protobuf/Abseil packages and commit-pinned module dependencies, avoiding unpinned system development packages and download fallbacks.
- CMake requires exact Protobuf `33.1.0` and gRPC `1.82.1`, requires the exported `gRPC::grpc_cpp_plugin` target, and always generates/builds both message and service stubs in the ignored build tree.
- Python bindings remain committed package output; `--check` regenerates into a temporary directory and compares output. C++ bindings are never committed.
- The C++ conformance test parses Python-produced vectors, validates fields/CRC32C, and emits `cpp_roundtrip.binpb`; Python compares that artifact byte-for-byte.
## Defects found and fixed
1. A relative bootstrap prefix was resolved after entering the temporary source directory, so successful output was deleted by cleanup. The script now canonicalizes the caller-relative destination first. The regression executes `--print-prefix` from a temporary working directory and validates the resulting path behavior.
2. The original native path omitted gRPC C++ and accepted any discoverable plugin. The bootstrap now builds exact gRPC/plugin sources, and CMake rejects absent/incompatible versions.
3. The Python script named the `grpcio-tools` pin but did not validate the installed distribution. It now refuses mismatched versions.
4. Protobuf ignored a stale provider option and attempted to download a different Abseil. The build was stopped; exact Abseil is now built first and Protobuf uses `LOCAL_DEPENDENCIES_ONLY`.
5. The host lacked OpenSSL development headers. Rather than add a floating system dependency, gRPC now uses the submodule pinned by its exact commit.
6. Documentation uses `bash scripts/bootstrap_native_toolchain.sh ...`, so a normal checkout does not depend on executable-mode preservation.
## Verified toolchain
```text
cmake version 4.4.0
c++ (GCC) 15.2.1 20260123 (Red Hat 15.2.1-7)
libprotoc 33.1
protobuf CMake package 33.1.0
grpcio-tools 1.82.1
grpcio 1.82.1
protobuf Python runtime 7.35.1
gRPC C++ 1.82.1
commit acccf84c0df20487d64101f528e5d426541ca4e5
grpc_cpp_plugin sha256 995ca8ac620fe83532b649a7c8c0a9341c7003da927fe0e4a8f821bfc579206d
```
The native toolchain and generated/build artifacts live under ignored mounted-drive `build/` paths; model/build artifacts were not stored under `/home`.
## Commands and results
```bash
bash scripts/bootstrap_native_toolchain.sh build/native-toolchain
```
```text
passed from a clean build directory
libprotoc 33.1
gRPC 1.82.1 commit acccf84c0df20487d64101f528e5d426541ca4e5
grpc_cpp_plugin sha256 995ca8ac620fe83532b649a7c8c0a9341c7003da927fe0e4a8f821bfc579206d
```
```bash
cmake -S packages/node/native -B build/native \
-DCMAKE_PREFIX_PATH="$PWD/build/native-toolchain"
cmake --build build/native -j"$(nproc)"
test -f build/native/shard_runtime.grpc.pb.cc
test -f build/native/shard_runtime.grpc.pb.h
test -f build/native/libshard_runtime_grpc.a
ctest --test-dir build/native --output-on-failure
```
```text
Pinned gRPC 1.82.1: building ShardRuntime service stubs
shard_runtime_proto built
shard_runtime_grpc built
1/1 shard_protocol_conformance passed
```
```bash
python3 -m pytest -q tests/test_native_shard_protocol.py
```
```text
50 passed, 2 optional-path skips
```
All DGR-023-required checks were selected explicitly:
```bash
python3 -m pytest -q -rs tests/test_native_shard_protocol.py \
-k 'cpp_and_python_agree_byte_for_byte or generated_python_stubs_match_the_proto or native_toolchain_bootstrap or wrong_grpcio'
```
```text
4 passed, 48 deselected
```
```bash
python3 scripts/generate_native_protocol.py --check
python3 scripts/generate_protocol_goldens.py --check
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
python3 -m compileall -q packages tests
git diff --check
```
```text
generated stubs are up to date
conformance vectors are up to date
OK: 55 stories validated
compileall passed
git diff --check passed
```
## Changed files
- `scripts/bootstrap_native_toolchain.sh`
- `scripts/generate_native_protocol.py`
- `packages/node/native/CMakeLists.txt`
- `packages/node/native/README.md`
- `tests/test_native_shard_protocol.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-023/README.md`
- `.scratch/distributed-gguf-runtime/prd.json` (secondary completion projection only)
## Limitations and dependency handoff
- This story proves exact schema/message/service generation and cross-language conformance. It does not implement or run the standalone worker service itself; DGR-033/DGR-037 own worker behavior.
- The plugin SHA is evidence for this verified build. Reproducibility authority is the exact gRPC commit plus its submodule graph, not an assumption that different compilers produce byte-identical executables.
- No model, GPU, API credits, or model download was used.
- DGR-024 and DGR-037 may consume this completed generation dependency but must provide their own transport/worker evidence.

View File

@@ -0,0 +1,180 @@
# DGR-024 evidence — real generated-gRPC protocol harness
**Status:** independently re-verified in a fresh worktree/environment (this session); `prd.json` `DGR-024.passes` is now `true`.
**Authority:** live Gitea #8 (revised); the local PRD is a secondary projection.
## Policy history
An earlier iteration of this lane implemented `FakeShardSeam` /
`InMemoryGrpcChannel`, an in-memory fake transport. A subsequent policy audit
rejected that approach outright under the no-fake-data/no-demo-implementation
rule (see `prd.json`, `DGR-024.notes`): "the former in-memory fake/stub seam
task was invalid... Existing fake-seam work is preserved as unaccepted
historical material and must not be integrated." That code
(`fake_shard_seam.py`, `test_fake_shard_seam.py`) is **not present** in this
worktree and must not be resurrected. This document supersedes any earlier
evidence describing it.
## Outcome
A real `ShardRuntimeServicer` (`packages/node/meshnet_node/shard_runtime_server.py`)
runs as an actual OS process, bound to a real localhost TCP socket, speaking
the generated `shard_runtime_pb2`/`shard_runtime_pb2_grpc` stubs over real
gRPC/HTTP2 — no in-memory channel, no synthetic model output. A test harness
(`tests/test_shard_runtime_harness.py`) spawns that process with
`subprocess.Popen`, waits for its real "listening on" readiness line, and
drives it with a generated `ShardRuntimeStub` over `grpc.insecure_channel`.
## Implemented
- `GetCapability` / `Health` unary RPCs over the real socket.
- `Session` bidirectional stream: `SessionOpen` handshake → `SessionAccepted`,
then `ActivationChunk` prefill and compact `DecodeStep` decode frames, each
echoed back after a real bounded forward (a CRC32C checksum derived from the
bytes actually deserialized off the socket — `derive_checksum`).
- **Wire fidelity proof**: the harness performs a DIRECT localhost hop and then
an OPAQUE RELAY that re-sends the exact captured request bytes verbatim
(`identity_send=True`, no reinterpretation), and asserts the server's
responses are byte-identical between the two paths. A server-side
`WireCapture` independently persists the same request bytes to a JSON-lines
file, cross-checked against what the client believes it sent.
- **Fail-closed negative paths** (`ShardRuntimeServicer.Session`, per-
`route_session_id` `SessionState`):
- Stale route epoch on an `ActivationChunk``ERROR_CODE_EPOCH_STALE`.
- Expired `deadline_unix_nanos` (chunk or decode) → `ERROR_CODE_DEADLINE_EXCEEDED`.
- Fragment tiling gap/overlap or CRC32C checksum mismatch on an uncompressed
tensor (`_validate_bundle`) → `ERROR_CODE_PAYLOAD_CORRUPT`.
- Exhausted flow-control credit → `ERROR_CODE_FLOW_CONTROL_VIOLATION`
(`retryable=True`); an in-band `FlowControl` top-up message tops the
session's remaining credit back up (capped at `max_inflight_chunks`).
- Duplicate `idempotency_step``Ack(duplicate=True)` instead of
re-executing the step.
- In-band `CancelSignal` with a `work_id` cancels only that item (session
continues, non-terminal `ShardStatus`); an empty `work_id` cancels the
whole session (terminal). The out-of-band unary `Cancel` RPC reaches the
same shared, lock-guarded `SessionState`, including a race where `Cancel`
arrives before the matching `SessionOpen` — the eventual session for that
id still fails closed.
- `Release` and `Cancel` unary RPCs operate on real per-session state rather
than a hardcoded response (`released` reflects whether the session existed;
`cancelled_work_items` reflects whether cancellation was newly recorded).
## Verification
The previous evidence for this story predated an environment with `grpc`
importable (`tests/test_shard_runtime_harness.py` could not even *collect* on
the ambient interpreter — see `.ralph-tui/progress.md`'s DGR-019 entry). This
session built a real, disposable `uv`-managed `.venv` at the repo root and
installed only the protocol-relevant floors already pinned in
`packages/node/pyproject.toml` (`grpcio==1.82.1`, `grpcio-tools==1.82.1`,
`protobuf==7.35.1`) plus `pytest==9.1.1`, then reran the full harness for
real — this is not a re-statement of the earlier claim, it is an independent
execution:
```bash
uv pip install grpcio grpcio-tools==1.82.1 protobuf pytest
PYTHONPATH=packages/node:packages/tracker .venv/bin/python -m pytest -q tests/test_shard_runtime_harness.py -v -s
```
```text
collected 11 items
tests/test_shard_runtime_harness.py .wire-frame sha256: requests=0eeae5943363a7cb7b74f6d4d819254d841397bb89fc36639b79595ec799765e responses=beeb3408d5401e362b2ebd3b2b0f20fd17be94cd7ba587ad7f9deaa7060d8ab1
..........
11 passed in 3.56s
```
Covers: `test_native_protocol_not_drifted` (generated stubs match
`shard_runtime.proto` exactly — reran `scripts/generate_native_protocol.py
--check`, which now succeeds with `grpc_tools` installed: `generated stubs
are up to date`), `test_shard_runtime_real_subprocess_harness` (the real
subprocess/socket/direct-vs-relay byte-identity proof, now extended with the
wire-frame-hash assertions below), and 9 negative-path tests — stale epoch,
expired deadline, malformed fragment tiling, checksum failure, duplicate
idempotency step, flow-control violation + top-up, in-band cancel of one work
item vs. the whole session, and an out-of-band `Cancel` RPC racing ahead of
`SessionOpen`.
### Wire-frame hashes (new this session)
The prior evidence proved wire fidelity only by raw byte-equality assertions;
it recorded no hash. `WireCapture.to_dict()`
(`packages/node/meshnet_node/shard_runtime_server.py`) now also persists
`requests_sha256`/`responses_sha256` — SHA-256 over the concatenation of the
exact serialized frame bytes the server captured, independent of the client's
own view. `tests/test_shard_runtime_harness.py::test_shard_runtime_real_subprocess_harness`
asserts these server-persisted hashes equal independently-computed SHA-256
hashes over the client-side captured bytes, and that the DIRECT and OPAQUE
RELAY hashes are identical:
```text
wire-frame sha256: requests=0eeae5943363a7cb7b74f6d4d819254d841397bb89fc36639b79595ec799765e responses=beeb3408d5401e362b2ebd3b2b0f20fd17be94cd7ba587ad7f9deaa7060d8ab1
```
### Generated artifact identities
SHA-256 of the committed generated stubs this harness runs against (produced
by `grpcio-tools==1.82.1` from `packages/node/native/proto/shard_runtime.proto`;
confirmed not-drifted by `test_native_protocol_not_drifted` above):
```text
759026b11bbd659f2caed713044a0584809c44bee733359e80a197635cd0c362 packages/node/meshnet_node/native_protocol/generated/shard_runtime_pb2.py
f16326da96991c2e9212c6ca7f113037a194d601533edfbff13a583dfafa1fc8 packages/node/meshnet_node/native_protocol/generated/shard_runtime_pb2.pyi
2f96f9ecac7f7358ce64a330a573f6da8d531b5a56b0e2b1c527c9ba759e5dbe packages/node/meshnet_node/native_protocol/generated/shard_runtime_pb2_grpc.py
```
```bash
.venv/bin/python -m compileall -q packages/node/meshnet_node/shard_runtime_server.py tests/test_shard_runtime_harness.py
.venv/bin/python -m compileall -q packages tests
git diff --check
```
```text
compileall (targeted): exit 0
compileall (packages tests, universal gate wording): exit 0
git diff --check: exit 0
```
Also re-ran `tests/test_ralph_prd_schema.py` (108 passed) after restoring
`prd.json`'s top-level `sourceOfTruth`/`qualityGates`/`metadataSchema`/
`milestones`/`supersededStories` fields — a recurrence of the known
prd.json-field-drop bug (see `.ralph-tui/progress.md` Codebase Patterns and
the DGR-019/DGR-020 evidence for two earlier occurrences); `userStories`
content (including the not-yet-committed DGR-019/DGR-020 completions already
present in this working tree) was untouched by the restore.
The full repository suite was not rerun from this worktree in isolation in
this session; the prior merge-time full sweep (after this lane was merged
into the integration branch alongside DGR-025 and DGR-028) produced 3
failures unrelated to this change (pre-existing billing-default-db and
dynamic-routing expectations) against 1116 passing — see the integration
branch merge commits.
## Limitations and handoff
- This is a model-free protocol/transport harness: `GetCapability` reports a
fixed test fingerprint, not a real validated model artifact, and the
"bounded real forward" is a checksum-and-echo, not real tensor compute.
- Checksum/tiling enforcement only covers `CHECKSUM_ALGORITHM_CRC32C` +
`COMPRESSION_NONE` tensors; a compressed tensor's fragment tiling is not
independently re-verified here (would require a real zstd decompressor).
- Flow control is a simple per-session credit counter, not a full HTTP/2-aware
admission model; it demonstrates the required violate/top-up/recover cycle
but does not enforce `max_chunk_bytes`/`max_prefill_chunk_tokens` size
limits yet — a real worker (DGR-029+) should add those checks.
- `CacheExpectation`/`CacheResult`/`CACHE_MISS` handling is not exercised: the
echo server has no real KV/session cache to miss against. A real worker
implementation owns that.
- Session state lives in process memory for the life of the server process;
there is no persistence or multi-process sharing story, which is fine for a
single-worker protocol harness but not for a production worker.
## Changed files
- `packages/node/meshnet_node/shard_runtime_server.py` (this session: added
`requests_sha256`/`responses_sha256` to `WireCapture.to_dict()`)
- `tests/test_shard_runtime_harness.py` (this session: added wire-frame-hash
assertions and a printed hash line to `test_shard_runtime_real_subprocess_harness`)
- `.scratch/distributed-gguf-runtime/evidence/DGR-024/README.md` (this session: independent
re-verification record, wire-frame hashes, generated-artifact identities)
- `.scratch/distributed-gguf-runtime/prd.json` (this session: restored
dropped top-level fields; `DGR-024.passes` flipped to `true`)

View File

@@ -0,0 +1,447 @@
# DGR-025 evidence — exact artifact and runtime recipe identity
**Status:** in progress — controller gates pass; final independent P0/P1 re-review is pending.
**Branch:** fixed detached Claude Fable provider lane
**Authority:** live Gitea #9; the local PRD is a secondary projection.
**Dependencies:** DGR-018 (`evidence/DGR-018/README.md` — canonical backlog schema and
issue projection), DGR-021 (`evidence/DGR-021/README.md` — versioned activation
envelope). Both read before changing code.
## Objective
Ensure the tracker and worker only combine numerically and operationally
compatible shards: fingerprint every axis that moves the numbers, bind shards to
exact half-open ranges, fail closed on any mismatch, and keep uncertified
recipes registered-but-dark.
## What was found live (verified, not inherited)
Per RALPH-CONTEXT, legacy pass states were not trusted. The DGR-003-lineage
identity core was inspected and exercised live before any change:
- `packages/node/meshnet_node/runtime_recipe.py` — node-side identity:
domain-separated digests (`meshnet.model-artifact.v1`,
`meshnet.runtime-recipe.v1`, `meshnet.shard-binding.v1`) over the source
artifact SHA (`source_digest`, with split artifacts bound to their exact
source via `DerivativeBinding`), tokenizer revision (pin-enforced),
architecture adapter + architecture/config digest, boundary and protocol
schema versions, backend, weight quantization, activation/compute dtypes, and
KV dtype/layout (`RECIPE_AXES`). Shard ranges are half-open
(`shard_start`/`shard_end`, end-exclusive, protocol convention) with no
topology or quant constants anywhere; `check_route` accepts any tiling of
`[0, layer_count)`. Route, handshake (`check_handshake`), and session-open
(`check_session_open`) checks fail closed with structured `RouteMismatch`
reasons mapped to specific protocol error codes (`handshake_error`).
- `packages/tracker/meshnet_tracker/recipe.py` — deliberately independent
tracker re-derivation (no `meshnet_node` import); declared fingerprints are
recomputed, never trusted (`parse_identity`, `FingerprintMismatch`). The
`CertificationLedger` keeps every registered recipe dark until a real
distributed forward — at least 2 distinct nodes, whole-model coverage,
non-synthetic, tokens actually generated — certifies it; dark recipes may
route only to certify.
- The two implementations are pinned by committed conformance vectors
(`tests/data/recipe_fingerprint_vectors.json`).
Live verification of that pre-existing core before changes:
`PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q tests/test_runtime_recipe_identity.py`
`45 passed`; plus `tests/test_native_identity_emission.py`,
`tests/test_tracker_capability_admission.py`, `tests/test_node_admission.py`
`59 passed`.
## Gap found and closed (this story's change)
**The `runtime_version` recipe axis was a label, not a pin.** It was an opaque
caller-supplied string: nothing derived it from the DGR-027 lock manifest, and
neither identity implementation rejected a moving reference (`"latest"` was
accepted), so two workers could run different llama.cpp pins or patch stacks
under one label and still agree on the recipe digest. The acceptance criterion
explicitly requires fingerprinting the "runtime pin/patch stack".
### Changed files
- `packages/node/meshnet_node/runtime_pin.py` (new) — derives the canonical
`runtime_version` axis value from the DGR-027 lock workspace
(`packages/node/native/llama`):
`<runtime>@<40-hex upstream commit>+patchstack.<sha256>` where the stack
digest commits, under the `meshnet.runtime-patch-stack.v1` domain, to the
*ordered* `(patch name, patch bytes sha256)` stack. Fails closed on: missing
or malformed `UPSTREAM_LOCK.json`, unknown schema version, non-40-hex/moving
commit, `UPSTREAM_COMMIT` disagreement, any disagreement among the lock's
`patch_series`, `patches/series`, and `patches/SHA256SUMS`, a missing patch
file, or a patch whose bytes don't match their recorded digest. Reads the
committed manifest only; fetching/patching stays with
`scripts/llama_cpp_dependency.py` (DGR-027).
- `packages/node/meshnet_node/runtime_recipe.py``runtime_version` is now
pin-enforced (`_require_pin`) exactly like `tokenizer_revision`; for the
llama.cpp backend it must also match the canonical
`llama.cpp@<40-hex>+patchstack.<64-hex>` grammar.
- `packages/node/meshnet_node/native_backend.py` — the production native
identity seam no longer accepts a caller-supplied runtime string. It derives
`runtime_version` directly through `load_runtime_pin()` from the committed
lock and rejects a non-llama backend at this llama.cpp-specific boundary.
- `packages/tracker/meshnet_tracker/recipe.py` — the independent tracker
implementation applies the same backend-specific grammar before re-deriving
the recipe digest, so forged operator labels cannot register or certify.
- `tests/test_runtime_pin_identity.py` and
`tests/test_native_identity_emission.py` — deterministic tests cover lock
derivation, production native emission, and node/tracker rejection of the
forged values from independent review. Conformance vectors were regenerated
through `scripts/gen_recipe_fingerprint_vectors.py` for the tightened wire
contract.
### Backlog-consistency repair (pre-existing damage, honestly recorded)
`tests/test_ralph_prd_schema.py` had 4 pre-existing failures before this story
touched anything, left by prior sessions and the alternate-history merge:
- DGR-022 and DGR-027 were marked `passes: true` without `completionNotes` and
without regenerated issue projections. Added their `completionNotes`
(explicitly labeled as added during this repair, content drawn from their own
evidence READMEs) and regenerated
`issues/022-…` / `issues/027-…` via `scripts/ralph_prd_schema.py render`.
- Three pre-DGR legacy GLM alpha issue files (`18-…`, `19-…`, `20-…`,
committed 2026-07-14, before DGR-018 established the generated-only
convention; they carry no authority disclaimer because they are *not*
generated from prd.json) were relocated via `git mv` to
`issues/legacy/` — preserved as provenance, out of the generated namespace.
### prd.json
Marked `DGR-025.passes = true` with `completionNotes`; regenerated
`issues/025-define-exact-artifact-and-runtime-recipe-identity.md`.
## Acceptance criteria → evidence
1. **Fingerprint all axes**`RECIPE_AXES` + `ArtifactIdentity` cover source
artifact SHA, tokenizer revision, architecture adapter/version (adapter axis
+ architecture/config digest), boundary schema (boundary + protocol schema
versions), backend, quant, activation/compute dtype, KV/state layout; the
runtime pin/patch stack is now committed via the derived `runtime_version`
axis (`runtime_pin.py`). Verified by `test_runtime_recipe_identity.py` and
`test_runtime_pin_identity.py`.
2. **Exact half-open range, no hardcoded topology/quant**`ShardIdentity`
end-exclusive ranges, `DerivativeBinding` coverage checks, `check_route`
tiling over arbitrary layouts; quant/dtype values are open strings
(dynamic recipe inputs). Verified by `test_runtime_recipe_identity.py`
(routes of 1, 2, and 5 shards; no product constants).
3. **Fail closed on any mismatch** — artifact, adapter, boundary/schema, cache
layout, backend, and runtime mismatches each produce structured
`RouteMismatch` reasons and protocol error codes; the tracker recomputes
digests and rejects inconsistent claims; moving runtime references are now
rejected on both sides.
4. **Registered-but-dark**`CertificationLedger`: unknown recipes cannot be
certified, registered recipes are dark, only a real ≥2-distinct-node
whole-model non-synthetic forward promotes; verified by
`test_runtime_recipe_identity.py` / `test_tracker_capability_admission.py`.
5. **Gates + this handoff** — below.
## Commands and results
```bash
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q tests/test_runtime_pin_identity.py
```
```text
23 passed in 0.15s
```
```bash
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q \
tests/test_runtime_pin_identity.py tests/test_runtime_recipe_identity.py \
tests/test_native_identity_emission.py tests/test_tracker_capability_admission.py \
tests/test_node_admission.py tests/test_node_capability.py tests/test_recipe_benchmark.py
```
```text
202 passed, 1 warning in 5.38s
```
```bash
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q tests/test_ralph_prd_schema.py
```
```text
108 passed
```
(4 failed before this story's backlog repair; 0 after.)
```bash
python3 -m compileall -q packages tests # exit 0
git diff --check # exit 0
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
# OK: 55 stories validated.
```
Default tests are model-download-free, API-credit-free, and GPU-free; no model
artifact was touched and nothing was written under `/home`.
## Limitations
- The production native identity seam now derives the manifest pin and cannot
accept an operator-supplied runtime label. It still cannot attest that the
running binary was built from those locked bytes. Embedding the patched-tree
hash at build time and echoing it through the DGR-022 status contract belongs
with DGR-028+/DGR-031; real distributed certification remains the final trust
boundary.
- The DGR-027-recorded blocker stands: `0002-dense-llama-owned-range-loader.patch`
does not apply cleanly against the pin (DGR-028). That does not affect this
story: the identity commits to the patch *bytes as committed*, which is
precisely what makes a later repaired patch a *different* runtime identity.
- No native/CMake change was made, so the native build/CTest gate is not
applicable; no llama.cpp patch content was changed, so apply/check/reverse
verification is not applicable (and is blocked by the DGR-028 defect anyway).
- Tracker routing, load balancing, billing, telemetry, and relay semantics are
untouched; the only behavior change outside the new module is the stricter
(fail-closed) rejection of moving `runtime_version` values.
## Dependency handoff
- **DGR-026** (split-GGUF provisioning): bind each provisioned split via
`DerivativeBinding` to the exact source digest recorded in its hashed
manifest; the per-split `shard_binding_digest` is what certification pins.
- **DGR-031** (`ShardEngine`): construct worker identity through
`shard_identity_from_native_report` and populate `runtime_version` from
`meshnet_node.runtime_pin.load_runtime_pin().runtime_version` — never from an
operator string. A build-time echo of the patched-tree hash through the
status contract would close the manifest-vs-binary gap noted above.
- **DGR-041** (capability registration): the tracker already re-derives and
fail-closes on presented identities (`parse_identity`); register recipes
through the `CertificationLedger` so they arrive dark.
- **DGR-044** (DeepSeek V4 Flash target): pin the target's artifact identity
the same way `glm_alpha_artifact` does — read locked manifests, never restate
digests — and note `layer_count` must count the routed transformer stack the
route tiles, excluding MTP (reserved for beta).
## Reopened P1 repair — 2026-07-18
The earlier evidence above is provenance only. Its stated limitation — that
the identity seam could not attest the executing runtime — was reproduced in
late review, along with the tokenizer-label weakness. This repair replaces
both claims at the production identity boundary.
### Changed files
- `packages/node/meshnet_node/runtime_recipe.py` — replaces the moving-ref
denylist with the sole valid `tokenizer.v1:<sha256>` form, derived from an
ordered map of named tokenizer/config byte digests. A label, tag, branch, or
symbolic ref cannot be a valid identity.
- `packages/tracker/meshnet_tracker/recipe.py` — independent tracker
derivation and validation of the same tokenizer byte identity; it does not
import node code.
- `packages/node/meshnet_node/runtime_pin.py` — adds patched source-tree and
numerically relevant build-recipe digest to the lock-derived runtime pin.
- `packages/node/meshnet_node/native_backend.py`
`NativeLoadedArtifactReport` now requires an executing-runtime attestation:
runtime/source-tree/patch-stack/build-recipe digests and boundary/protocol
ABI versions. `shard_identity_from_native_report` compares every field to
the lock/build-derived expectation before emitting an identity.
- `scripts/gen_recipe_fingerprint_vectors.py` and
`tests/data/recipe_fingerprint_vectors.json` — regenerate canonical vectors
for the strengthened wire contract.
- `tests/test_runtime_pin_identity.py`,
`tests/test_runtime_recipe_identity.py`, and
`tests/test_native_identity_emission.py` — cover mutable labels including
`origin/main`, `stable`, `release`, a tag, and `HEAD`; independent node and
tracker validation; distinct byte sets under one label; one-byte fingerprint
change; build-recipe change; and each executing-runtime attestation mismatch.
### Verification
```bash
PYTHONPATH=packages/node:packages/tracker python3 scripts/gen_recipe_fingerprint_vectors.py
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q \
tests/test_runtime_pin_identity.py tests/test_native_identity_emission.py \
tests/test_runtime_recipe_identity.py
```
```text
92 passed
```
```bash
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q \
tests/test_runtime_pin_identity.py tests/test_runtime_recipe_identity.py \
tests/test_native_identity_emission.py tests/test_node_admission.py \
tests/test_node_capability.py tests/test_recipe_benchmark.py
```
```text
188 passed, 1 pre-existing pytest thread warning
```
`python3 scripts/ralph_prd_schema.py validate
.scratch/distributed-gguf-runtime/prd.json`, `python3 -m compileall -q packages
tests`, and `git diff --check` each exit 0. The broader PRD pytest projection
suite has two unrelated existing DGR-023 failures: its `passes: true` entry has
no completion notes and its generated issue file is stale. The socket-backed
subset of `test_tracker_capability_admission.py` is additionally un-runnable in
this sandbox (`PermissionError: [Errno 1] Operation not permitted` creating an
AF_INET socket); its deterministic non-socket identity coverage is included in
the passing runs above.
### Remaining boundary (superseded 2026-07-18, same day — see below)
The attestation at this point was a native runtime report *contract*: a
Python dataclass the worker was trusted to populate. Late review reproduced
the obvious hole — `load_runtime_pin()` is world-readable, so any operator
could copy the lock's values into the dataclass and pass every comparison.
The section below closes that hole.
## Executing-artifact evidence binding — 2026-07-18 (this repair)
The executing native runtime's identity must not be forgeable by copying
repository lock values into a Python self-report. Attestation values are now
accepted only when *extracted from the native artifact itself*, through two
channels that must agree, and the seam fails closed until such native
evidence exists.
### The boundary
`meshnet_node.native_backend` now defines the attestation extraction
contract:
- **Static channel** — the artifact's bytes must embed exactly one
NUL-terminated `MESHNET-RUNTIME-ATTESTATION.v1:<canonical json>` marker.
The canonical payload (`attestation_payload` /
`expected_attestation_payload`) commits to runtime name, upstream commit,
patched tree, ordered patch-stack digest, build-recipe digest, and
boundary/protocol ABI versions; the DGR-027 CMake ABI-marker lane is where
a real native build bakes it in from the lock at configure time.
- **Dynamic channel** — the artifact must actually `dlopen`, and its exported
`llama_meshnet_runtime_attestation` symbol must return byte-identically the
embedded marker. A marker pasted into a plain file is not an executing
runtime.
- **Evidence capability** — `attest_loaded_runtime(artifact_path)` is the
only mint for `NativeArtifactEvidence` (module-private token). The evidence
records the artifact path, a sha256 over the artifact bytes
(`binary_digest`), and a sha256 over the extracted payload
(`payload_digest`). `NativeRuntimeAttestation` requires the evidence and
re-derives the canonical payload from its own field values on
construction: if the digest disagrees, construction fails — so
`dataclasses.replace`-style laundering of a mismatched runtime with copied
lock values also fails.
- `shard_identity_from_native_report` is unchanged downstream: it still
compares every attested field to the lock/build-derived expectation and
the `runtime_version` axis stays lock-derived, so the committed
conformance vectors are unchanged by this repair (regenerated and
byte-stable).
Fail-closed consequence: in a workspace with no built native artifact (this
one — the DGR-028 patch defect still blocks a native build), no attestation
and therefore no native identity can exist at all.
### Changed files
- `packages/node/meshnet_node/native_backend.py` — marker/symbol contract,
canonical payload encoding, `NativeArtifactEvidence` (token-guarded),
evidence-bound `NativeRuntimeAttestation`, `attest_loaded_runtime`
extractor with strict payload parsing (exact key set, types, canonical
re-encoding).
- `tests/test_native_identity_emission.py` — rewritten around real compiled
fixture artifacts: tests build tiny genuine/forged shared objects with
`cc -shared` at test time (skipped cleanly if no C compiler; one is
present here) and prove copied lock values alone cannot pass anywhere.
### Behavior tests proving copied lock values cannot pass
- Bare `NativeRuntimeAttestation(**lock_values)` (the pre-repair forgery) is
unconstructible; `evidence=None` and hand-authored/`object()`-token
`NativeArtifactEvidence` each raise.
- The true marker bytes written into a plain file fail (`not a loadable`).
- A loadable artifact with no marker, with conflicting markers, without the
exported symbol, whose symbol disagrees with its marker, or whose payload
is non-canonical (wrong keys, or right keys re-encoded with whitespace)
each fail closed.
- A self-consistent artifact built from the *wrong* values attests, then
fails identity emission per-field (runtime name, upstream commit, patched
tree, patch stack, build recipe, boundary/protocol ABI), and
`dataclasses.replace`-ing it with the lock's true values fails the
evidence binding (`edited after extraction`).
- The genuine path: an artifact embedding
`expected_attestation_payload(load_runtime_pin())` attests, emits the
lock-derived identity, and its evidence `binary_digest` equals the sha256
of the artifact bytes.
### Verification (all in this worktree, 2026-07-18)
```bash
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q \
tests/test_native_identity_emission.py tests/test_runtime_pin_identity.py \
tests/test_runtime_recipe_identity.py
```
```text
104 passed
```
```bash
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q \
tests/test_node_admission.py tests/test_node_capability.py \
tests/test_recipe_benchmark.py
```
```text
96 passed, 1 pre-existing pytest thread warning
```
```bash
PYTHONPATH=packages/node:packages/tracker python3 -m pytest -q \
tests/test_tracker_capability_admission.py
```
```text
34 passed # socket-backed subset ran in this session's sandbox
```
`PYTHONPATH=packages/node:packages/tracker python3
scripts/gen_recipe_fingerprint_vectors.py` reproduces the committed vectors
byte-for-byte; `python3 -m compileall -q packages tests` and
`git diff --check` each exit 0.
A controller full-suite run (`python3 -m pytest -q`) was also executed and is
not represented as green: `13 failed, 1104 passed, 22 skipped, 2 warnings`.
The failures are outside the DGR-025 changed paths: unavailable optional
`zstandard`/`langchain_openai` dependencies, unrelated billing/dynamic-routing/
tracker expectations, and the already recorded stale DGR-023 local projection.
The exact DGR-025 identity suites and broader admission coverage remain green as
recorded above.
### Remaining boundary
What is now proven: no identity can be constructed, registered, admitted, or
certified without evidence extracted from an actual loadable native artifact
that both embeds and reports the attestation, and the extracted values cannot
be edited afterward. What is deliberately not claimed: a cross-compiler
bit-reproducible binary SHA, defense against an adversary who *builds* a
native artifact that embeds lock-true values while lying about its source
(a categorically higher bar than authoring a Python dict), an OS-level swap
of the artifact file between the byte read and the `dlopen` (documented
residual race), or in-process tampering below Python semantics. Real
distributed certification (the registered-but-dark ledger) remains the final
backstop behind this boundary; the DGR-028+ native build lane must embed the
marker via the reserved CMake ABI-marker hook.
## Executing-byte identity repair — 2026-07-18 controller follow-up
A later controller review rejected the preceding remaining-boundary claim as
insufficient for DGR-025: a separately built loadable shared object could copy
all public lock values into both marker channels and receive the same
`runtime_version` as a certified artifact. The repair now appends
`+artifact.<sha256>` to the llama.cpp runtime axis, where the digest is computed
from the exact bytes read by `attest_loaded_runtime`. Node and tracker parsers
independently require this suffix. Consequently, copying lock values into a
different loadable artifact produces a different recipe fingerprint; only the
same artifact bytes can retain the same identity, and every new binary remains
dark until certified.
`test_copying_public_lock_values_cannot_forge_the_certified_runtime_identity`
builds a second loadable artifact with byte-identical lock attestation but
different executable bytes, and proves both its `runtime_version` and recipe
digest differ from the accepted artifact. Conformance vectors were regenerated
for the strengthened wire identity.
Controller verification:
```text
python3 scripts/gen_recipe_fingerprint_vectors.py
python3 -m pytest -q tests/test_native_identity_emission.py \
tests/test_runtime_pin_identity.py tests/test_runtime_recipe_identity.py
# 105 passed in 0.52s
python3 -m compileall -q packages/node/meshnet_node \
packages/tracker/meshnet_tracker tests scripts/gen_recipe_fingerprint_vectors.py
# exit 0
git diff --check
# exit 0
```

View File

@@ -0,0 +1,268 @@
# DGR-026 evidence — provision exact split-GGUF artifacts outside `/home`
**Status:** implemented and verified this session; live re-review, not inherited credit.
**Dependency:** DGR-025 (`evidence/DGR-025/README.md`) — read before changing code.
## Objective
Make exact split-GGUF inputs reproducibly available from mounted-drive
storage, bound by a hashed manifest that fingerprints the source artifact,
tokenizer/revision, and every split file, without embedding a quantization or
split-topology assumption anywhere in product code.
## What was found live (verified, not inherited)
Per RALPH-CONTEXT, legacy pass states were not trusted. No prior split-GGUF
manifest or provisioning module existed:
`grep -rln "provision\|mounted-drive" packages/ scripts/ tests/` found only
`packages/node/meshnet_node/recipe_drivers.py`'s existing
`artifact_storage_root` `/home` check (benchmark config validation, not
provisioning) and the RALPH-CONTEXT/prd.json prose itself. The pre-existing
`packages/node/meshnet_node/downloader.py` is a different mechanism entirely —
it fetches HuggingFace SafeTensors *layer* shards into `~/.cache/meshnet/shards`
(i.e. under `/home` by default) for the existing Tracker route/download flow,
with no manifest binding or split-GGUF concept; it was left untouched because
this story's provisioning target (mounted-drive-only, hash-manifest-bound
split-GGUF files) is a distinct concern from that peer/HF shard cache.
Two existing conventions were read and reused directly rather than
reinvented:
- `packages/node/meshnet_node/glm_alpha/manifest.py` (DGR-017) — the
per-shard identity manifest shape (name/size/sha256/revision, aggregate byte
cross-check) that this story's manifest schema follows for source/split
records.
- `packages/node/meshnet_node/runtime_recipe.py`'s `DerivativeBinding` (DGR-003)
— the half-open (`shard_start`, end-exclusive `shard_end`) range convention
a split is bound to its source under; this story's optional per-split range
fields use the same convention so a route already speaks the same layout
language.
- `packages/node/meshnet_node/recipe_drivers.py`'s `_validate_config` — the
exact `/home` rejection shape (`not root.is_absolute() or root ==
Path("/home") or Path("/home") in root.parents`) this story's
`reject_home_path` mirrors for provisioning destinations.
## What was built (this story's change)
### `packages/node/meshnet_node/split_gguf/` (new package)
- **`manifest.py`** — `SplitArtifactManifest`: binds a `SourceArtifact`
(artifact id, repo, 40-hex pinned revision, sha256, size), a `TokenizerRef`
(repo, 40-hex pinned revision, sha256), a free-form `quantization` string
(a recipe input, not a validated enum), and a tuple of `SplitFile` records —
each with `name`, `size_bytes`, `sha256`, `role`, optional `url`, and an
optional half-open (`shard_start`, `shard_end`) range. `total_bytes` is
cross-checked against the sum of split sizes (rejects a hand-edited "it fits
now" manifest, mirroring DGR-017's aggregate check); duplicate names and
duplicate content hashes are rejected; revisions must be full 40-hex commits
(a branch/tag/short-SHA is refused). Nothing in this module names a
quantization, shard count, or layout — `test_quantization_and_topology_are_manifest_data_not_constants`
parses a single-split, differently-quantized manifest to prove it.
- **`provision.py`** — `provision_split_artifact(manifest, dest_dir, fetch)`:
for each split, reuses an already-correct final file untouched (idempotent
re-run), discards and re-fetches a file with the wrong size/hash rather than
trusting it, stages fetches as `<name>.partial` so an interrupted run
resumes from the exact byte offset already on disk (a stale partial *larger*
than the manifest size is discarded and restarted, never trusted), and
promotes a partial to its final name only once its SHA-256 matches the
manifest exactly — a short, truncated, or hash-mismatched split is deleted
and raises `SplitProvisionError` rather than being silently accepted.
`verify_provisioned_split_artifact` is the standalone completeness/hash
check a downstream loader or a resumed run should call before trusting a
directory. `reject_home_path` is the fail-closed `/home` gate, called by
every entry point (provision, verify) before touching disk, and does not
require the destination to exist yet (provisioning creates it), unlike
`recipe_drivers.py`'s `strict=True` benchmark-root check. Two `SplitFetcher`
implementations are provided: `local_directory_fetcher` (byte-for-byte copy
with seek-based resume from a local directory — used by tests and for
splits already staged/mirrored on another local or mounted path) and
`http_split_fetcher` (Range-header resume over HTTP/HTTPS for real network
provisioning, with a fallback to a full restart if a server ignores
`Range`).
### `scripts/provision_split_gguf.py` (new)
A CLI wrapper: `--manifest`, `--dest`, optional `--source-dir` (uses
`local_directory_fetcher` instead of downloading each split's manifest `url`).
Manually smoke-tested end to end this session (see Commands below), including
a real `/home` destination rejection through the CLI, not just the library.
### Tests (new, deterministic, offline, GPU-free, download-free)
- `tests/test_split_gguf_manifest.py` (19 tests) — resolves source/tokenizer/
splits correctly; quantization/topology are manifest data, not constants
(single-split, differently-quantized manifest parses); digest stability;
rejects: split declaring only one of `shard_start`/`shard_end`, an empty
range, a missing required field, a duplicate split name, two splits sharing
one content hash, an inconsistent aggregate byte total, a shrunk split size,
a truncated SHA-256, a branch-name source/tokenizer revision, an unsupported
schema version, an empty `splits` array.
- `tests/test_split_gguf_provision.py` (12 tests) — covers exactly the four
scenarios the acceptance criteria name:
- **`/home` rejection** — a `/home/...` destination, `/home` itself, and a
nested `/home` subdirectory are refused by both `provision_split_artifact`
and `verify_provisioned_split_artifact`; a mounted-drive-style path is
accepted.
- **Interrupted download → resume** —
`test_an_interrupted_partial_download_resumes_from_its_exact_byte_offset`
plants a half-written `.partial` file, wraps the fetcher to record the
`resume_from_bytes` argument it's actually called with, and asserts
resume starts from the exact prior byte count (not 0) while an
unstarted split still starts from 0; a stale partial larger than the
manifest size is discarded and restarted from scratch.
- **Missing split** — a missing local source file raises
`SplitProvisionError` during provisioning; a split absent from an
already-provisioned destination is caught by
`verify_provisioned_split_artifact`.
- **Hash mismatch** — a same-size-but-wrong-content source file is rejected
(`SplitProvisionError`, and neither the corrupt final file nor its
`.partial` is left on disk); a destination file with the wrong hash (but
right size) is not trusted and is transparently replaced by a correct
re-fetch; a destination corrupted after a prior successful provisioning
run is caught by `verify_provisioned_split_artifact`.
- Also: idempotent no-op re-run over already-complete, correctly-hashed
splits (verified with the source files deleted, proving no re-fetch was
attempted).
## Acceptance criteria → evidence
1. **Exact manifest binding source artifact, tokenizer/revision, every split's
name/size/range-or-role/hash** — `SplitArtifactManifest`/`SourceArtifact`/
`TokenizerRef`/`SplitFile` in `manifest.py`; covered by
`test_split_gguf_manifest.py`.
2. **Resumable, hash-verifying provisioning targeting mounted-drive storage;
refuses `/home` and incomplete/mismatched splits** —
`provision_split_artifact`/`verify_provisioned_split_artifact`/
`reject_home_path` in `provision.py`; covered by
`test_split_gguf_provision.py` and the CLI smoke test below.
3. **Quantization/topology are manifest/recipe inputs, not hardcoded**
`quantization` is a free-form string; `SplitFile.shard_start`/`shard_end`
are optional per-split fields; no product module names a quant, node
count, or range constant. Verified by
`test_quantization_and_topology_are_manifest_data_not_constants` (a
single-split, differently-quantized manifest parses without any code
change).
4. **Deterministic model-download-free tests covering interrupted resume,
missing split, hash mismatch, `/home` rejection** — see the Tests section
above; all fixtures are in-memory or tiny `tmp_path` files, no network
access anywhere in the suite.
5. **Gates + this handoff** — below.
## Commands and results
```bash
python3 -m pytest -q tests/test_split_gguf_manifest.py tests/test_split_gguf_provision.py
```
```text
31 passed in 0.10s
```
```bash
python3 -m pytest -q tests/test_ralph_prd_schema.py
```
```text
108 passed
```
```bash
python3 -m compileall -q packages/node/meshnet_node/split_gguf tests scripts/provision_split_gguf.py
git diff --check
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
```
```text
(compileall exit 0; git diff --check exit 0)
OK: 55 stories validated.
```
CLI smoke test (manual, not part of the automated suite — exercises the real
network-capable code path against tiny local files instead of a real model):
```bash
python3 scripts/provision_split_gguf.py \
--manifest /tmp/dgr026-smoke/manifest.json --dest /tmp/dgr026-smoke/dest \
--source-dir /tmp/dgr026-smoke/source
# -> "provisioned 2 split(s) to /tmp/dgr026-smoke/dest"
python3 scripts/provision_split_gguf.py \
--manifest /tmp/dgr026-smoke/manifest.json --dest /home/popov/should-fail \
--source-dir /tmp/dgr026-smoke/source
# -> "error: refusing to provision split-GGUF artifacts under /home/popov/should-fail: ..."
# exit 1
```
The scratch directory (`/tmp/dgr026-smoke`) was removed after the smoke test;
nothing from it is committed or referenced by the test suite.
Default tests are model-download-free, API-credit-free, and GPU-free; no model
artifact was downloaded and nothing product-relevant was written under
`/home` (the CLI smoke test's `/home` path was rejected before any write).
## Changed files
- `packages/node/meshnet_node/split_gguf/__init__.py` (new)
- `packages/node/meshnet_node/split_gguf/manifest.py` (new)
- `packages/node/meshnet_node/split_gguf/provision.py` (new)
- `scripts/provision_split_gguf.py` (new)
- `tests/test_split_gguf_manifest.py` (new)
- `tests/test_split_gguf_provision.py` (new)
- `.scratch/distributed-gguf-runtime/prd.json` (`DGR-026.passes = true` +
`completionNotes`; also restored the top-level `sourceOfTruth`/
`qualityGates`/`metadataSchema`/`milestones`/`supersededStories`/
`branchName` fields — see Gotcha below)
- `.scratch/distributed-gguf-runtime/issues/026-provision-exact-split-gguf-artifacts-outside-home.md`
(regenerated via `scripts/ralph_prd_schema.py render`)
- `.scratch/distributed-gguf-runtime/evidence/DGR-026/README.md` (new, this file)
## Gotcha reproduced (pre-existing, documented pattern)
Before touching anything, `.scratch/distributed-gguf-runtime/prd.json`'s
top-level `sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/
`supersededStories`/`branchName` fields were already missing in the working
tree at session start (this is the fourth documented occurrence of the
round-trip-drop bug noted in DGR-018/019/020/025's evidence — `userStories`
itself was unaffected, only these top-level fields). Restored them from
`git show HEAD:.scratch/distributed-gguf-runtime/prd.json` before making any
DGR-026 edit; `scripts/ralph_prd_schema.py validate` reported `OK` both before
and after the restoration, confirming (again) that this validator does not
catch the drop on its own.
## Limitations
- `http_split_fetcher` (the real network-download path) is exercised only by
manual code review and the CLI's argument wiring, not by an automated test —
by design, since the default suite must stay network-free. Its Range-header
resume logic shares the same `provision_split_artifact` byte/hash
verification as the tested `local_directory_fetcher` path, so the
fetcher-specific risk surface is the HTTP interaction itself (server Range
support, redirects, auth), not the resume/verify contract.
- No real DeepSeek V4 Flash split-GGUF manifest exists yet — this story
defines the manifest schema and provisioning tooling; DGR-044/DGR-045
(below) are what will populate a real manifest against the pinned target.
- `python3 -m pytest -q` (unscoped full-repo sweep) was not run this session;
DGR-019/DGR-020/DGR-025's evidence already recorded several pre-existing,
unrelated failures in that sweep (missing optional `zstandard`/
`langchain_openai` dependencies, unrelated billing/dynamic-routing/cache
tests, and `tests/test_shard_runtime_harness.py`'s `grpc` import
requirement). This story's own targeted suites, `test_ralph_prd_schema.py`,
`compileall`, and `git diff --check` are all green as recorded above.
- Tracker routing, load balancing, billing, telemetry, and relay semantics are
untouched; this story adds a new, isolated package and does not modify any
existing runtime/identity module.
## Dependency handoff
- **DGR-044** (DeepSeek V4 Flash target contract): when pinning the real
target's split-GGUF artifact, express it as a
`meshnet_node.split_gguf.manifest.SplitArtifactManifest``source.sha256`
is the whole-model artifact digest DGR-003's `ArtifactIdentity.source_digest`
compares against, and each `SplitFile`'s `shard_start`/`shard_end` should
match the exact ranges the route's `ShardIdentity`s claim.
- **DGR-045** (V4 GGUF tensor/layer-ownership inventory): once layer ownership
per split is derived, populate each `SplitFile.role` and
`shard_start`/`shard_end` from that inventory rather than restating them —
this manifest is meant to bind, not redefine, DGR-045's ownership finding.
- Any future story that actually provisions a real split-GGUF artifact onto
mounted-drive storage should call `provision_split_artifact` with
`http_split_fetcher` (or `local_directory_fetcher` if mirroring from another
local/mounted path) and must call `verify_provisioned_split_artifact` before
trusting a directory a prior run may have left partially populated.

View File

@@ -0,0 +1,72 @@
# DGR-027 evidence — exact llama.cpp provenance manifest and fetch workspace
**Completed implementation:** 2026-07-17
**Branch:** `ralph/dgr-small-terra`
**Authority:** live Gitea issue #11. The controller fetched and claimed the issue
through the Gitea API before launch; the isolated agent received that exact body.
## Changed files
- `packages/node/native/llama/UPSTREAM_LOCK.json`
- `packages/node/native/llama/PATCH-STACK.md`
- `scripts/llama_cpp_dependency.py`
- `tests/test_llama_cpp_dependency.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-027/README.md`
## Provenance and retrieval contract
`UPSTREAM_LOCK.json` records the upstream Git URL, immutable 40-character
commit `e920c523e3b8a0163fe498af5bf90df35ff51d25`, expected Git tree
`6c91a11407a3a3fb160f5dac705f9c59718f54f1`, MIT license, and the sole
retrieval method: `git-clone-detached-commit` into `build/llama.cpp/source`.
`python3 scripts/llama_cpp_dependency.py fetch` has no branch, tag, ref, or
repository override. On a first fetch it clones the manifest URL, checks out
the detached commit, and verifies commit, tree, required upstream blobs,
license, and cleanliness. If the workspace already exists, it makes no network
request and accepts it only after the same verification. Dirty or mismatched
caches fail closed. The build directory is already ignored by `.gitignore`.
## Verification
| Command | Result |
| --- | --- |
| `python3 -m pytest -q tests/test_llama_cpp_dependency.py` | `7 passed in 0.22s` |
| `python3 -m compileall packages tests` | passed |
| `git diff --check` | passed (no output) |
| `python3 scripts/llama_cpp_dependency.py inspect` | passed; reports exact commit/tree, retrieval workspace, MIT license, and two-patch stack |
| `python3 scripts/llama_cpp_dependency.py fetch --workspace /tmp/not-llama-workspace` | failed closed with status 2: workspace outside the locked ignored build root |
| symlinked workspace regression | passed; both a `build/` ancestor symlink and a final `source` symlink escaping the repository are refused |
| attached-branch cache regression | passed; an exact commit on a local branch is refused until checked out as detached HEAD |
| ignored/excluded injection regression | passed; a file hidden by `.git/info/exclude` is detected and refused |
| tracked injection regression | passed; modified tracked content hidden by both `assume-unchanged` and `skip-worktree` is content-hashed and refused |
| executable-mode regression | passed on the POSIX fixture for both index flags; the mounted project workspace has `core.filemode=false`, so its exact index tree is the canonical mode record and physical mode bits are not treated as meaningful |
| `git check-ignore -v build/llama.cpp/source` | passed; `.gitignore:6:build/` |
| `git diff --summary` and `git ls-files build packages/node/native/llama` | no source checkout or new submodule introduced; only manifest/docs/patches/native wrapper are tracked |
| `python3 scripts/llama_cpp_dependency.py fetch` (controller network lane) | passed; fetched the exact detached commit and verified HEAD `e920c523e3b8a0163fe498af5bf90df35ff51d25` and tree `6c91a11407a3a3fb160f5dac705f9c59718f54f1` in the ignored workspace |
| `python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source` | failed on the pre-existing `0002-dense-llama-owned-range-loader.patch` as a corrupt patch at line 26; this is an explicit DGR-028 blocker and no native-build claim is made |
The targeted test suite creates a local Git fixture to prove offline cache reuse
after full identity verification, then proves a dirty cache is rejected. It
also proves the CLI rejects a repository/branch override and an arbitrary
workspace.
## Limitations
- The controller successfully materialized and verified the exact upstream
commit/tree, so the DGR-027 fetch and offline-cache boundary has real upstream
evidence rather than fixture-only evidence.
- The existing `0002-dense-llama-owned-range-loader.patch` is malformed and
cannot pass `git apply --check` against the exact pin. DGR-027 changes no patch
file; repairing and certifying the numbered patch stack belongs to DGR-028.
Until that story closes, the repository must not claim patched-tree, native
CMake/CTest, or reverse-apply certification.
- No model, API credits, GPU, or model artifact storage was used.
## Dependency handoff
DGR-028, DGR-029, and DGR-044 must invoke the manifest-owned `fetch` command
before touching llama.cpp source. They may use only the verified
`build/llama.cpp/source` checkout and must record any native build, CTest, and
patch apply/check/reverse evidence against the exact manifest pin. DGR-017's
cleanup remains provenance only and grants no inherited completion credit.

View File

@@ -0,0 +1,191 @@
# DGR-028 evidence — numbered llama.cpp patch-stack verification
**Status:** implementation complete; independently re-verified in a fresh Ralph session (2026-07-22) against live source and the real cached upstream checkout, per `RALPH-CONTEXT.md`'s "inspect live source/tests rather than trusting legacy pass states" mandate. `prd.json`'s `DGR-028.passes` is now `true`.
**Authority:** local `prd.json` is authoritative; live Gitea #12 is a projection.
**Upstream pin:** `e920c523e3b8a0163fe498af5bf90df35ff51d25` (`llama.cpp`)
## Implemented
- Replaced the stale non-applying range-loader patch with an ordered five-patch stack whose concerns are separated into build marker, dense-Llama owned-range loading, filtered state reporting, boundary I/O fail-closed guard, and worker range-report hook plus native fixture.
- Added `patches/UPSTREAM-ASSUMPTIONS.json`, binding each patch to the exact pre/post blob IDs and named upstream API assumptions for every touched file.
- Extended `scripts/llama_cpp_dependency.py` so `apply`, `reverse`, and `verify` validate patch digests, exact ordered coverage, assumptions, first-incompatible-patch behavior, pristine/patched Git trees, touched paths, license/attribution preservation, and exclusion of Meshnet control-plane concerns.
- `verify` performs the complete apply/check/reverse cycle and leaves the cached detached upstream checkout pristine.
- Updated the lock's exact patched tree and patch checksums. No model artifact was downloaded or created.
## Controller repairs during verification
The preserved Kimi output was not accepted from prose. Initial controller execution found and repaired:
1. a missing `_git` helper that made the dependency verifier raise `NameError`;
2. assumptions resolved relative to the repository root rather than the llama manifest directory;
3. the documented `verify`/`reverse` contract was not wired into the CLI or apply path;
4. assumptions and control-plane/license boundaries were defined but never enforced during apply;
5. a stale Python test hardcoded the old two-patch count;
6. the native fixture made an invalid strict resident-buffer-size comparison. Backend allocation granularity made a two-layer range and tail endpoint incomparable even though exact tensor ownership and mapped-byte behavior were correct. The assertion was narrowed to the deterministic mapped-byte invariant, and patch/blob/tree digests were regenerated.
## Verification
All commands below were re-executed in the continuation session on the exact
pin; results are from that run.
```text
cd packages/node/native/llama/patches && sha256sum -c SHA256SUMS
# all five patches OK
python scripts/llama_cpp_dependency.py inspect
# exact commit/tree, MIT license, five-patch series, no model downloads
python scripts/llama_cpp_dependency.py verify --workspace build/llama.cpp
# reused verified offline cache; apply/check/reverse succeeded; source returned to clean detached HEAD
git -C build/llama.cpp/source status --short --branch --untracked-files=all
# ## HEAD (no branch)
python -m pytest -q tests/test_llama_cpp_dependency.py
# 7 passed in 0.27s
python -m compileall -q scripts/llama_cpp_dependency.py tests/test_llama_cpp_dependency.py
# exit 0
python -m compileall -q packages tests
# exit 0
git diff --check
# exit 0
```
Focused native gate against the patched exact pin (apply first because `verify`
intentionally restores the source checkout to pristine state, then reverse after
the test):
```text
python scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
# patched index tree c0045714735ae5ee7b7334a480d8ac04e03e1b18 matches the lock
cmake -S build/llama.cpp/source -B build/llama.cpp/dgr028-build-verify \
-G 'Unix Makefiles' -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_TESTS=ON \
-DLLAMA_BUILD_EXAMPLES=OFF -DLLAMA_BUILD_SERVER=OFF \
-DLLAMA_BUILD_TOOLS=OFF -DLLAMA_BUILD_APP=OFF -DLLAMA_CURL=OFF
cmake --build build/llama.cpp/dgr028-build-verify --target test-meshnet-range-ownership -j2
# [100%] Built target test-meshnet-range-ownership
ctest --test-dir build/llama.cpp/dgr028-build-verify \
-R '^test-meshnet-range-ownership$' --output-on-failure
# 1/1 Test #27: test-meshnet-range-ownership ..... Passed 0.01 sec
python scripts/llama_cpp_dependency.py reverse --source-dir build/llama.cpp/source
git -C build/llama.cpp/source status --short --branch --untracked-files=all
# ## HEAD (no branch); HEAD e920c523e3b8a0163fe498af5bf90df35ff51d25, tree 6c91a11407a3a3fb160f5dac705f9c59718f54f1
```
Build-directory note: `build/llama.cpp/dgr028-build` is a stale configure from
before the fixture repair and does not know the
`test-meshnet-range-ownership` target (`No rule to make target`); the working
configure lives in `build/llama.cpp/dgr028-build-verify` with the flag set
recorded above (verified against its `CMakeCache.txt`). Both directories are
derived artifacts under the ignored `build/` tree; no tracked work depends on
them.
A broad `cmake --build ... --target test` was also attempted after building only the focused target. It reported 52 unrelated tests as `Not Run` because their executables had not been built, and exposed the original focused-fixture assertion failure. It is not presented as a full-suite gate. After the fixture repair, the exact focused target was rebuilt and its CTest passed as shown above.
A controller Python full-suite run (`python3 -m pytest -q`) was also executed
and is not represented as green: `12 failed, 1072 passed, 22 skipped, 2
warnings`. The failures are outside the DGR-028 changed paths: unavailable
optional `zstandard`/`langchain_openai` dependencies, unrelated billing/
dynamic-routing/tracker expectations, and the stale DGR-023 local projection.
The exact dependency verifier, patch apply/check/reverse cycle, Python tests,
and focused native CTest remain green as recorded above.
## Changed files
- `packages/node/native/llama/PATCH-STACK.md`
- `packages/node/native/llama/THIRD_PARTY_NOTICES.md`
- `packages/node/native/llama/UPSTREAM_LOCK.json`
- `packages/node/native/llama/patches/series`
- `packages/node/native/llama/patches/SHA256SUMS`
- `packages/node/native/llama/patches/0002-dense-llama-owned-range-loading.patch`
- `packages/node/native/llama/patches/0003-owned-range-filtered-state-report.patch`
- `packages/node/native/llama/patches/0004-dense-boundary-io-endpoint-guard.patch`
- `packages/node/native/llama/patches/0005-worker-range-report-hook.patch`
- `packages/node/native/llama/patches/UPSTREAM-ASSUMPTIONS.json`
- `scripts/llama_cpp_dependency.py`
- `tests/test_llama_cpp_dependency.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-028/README.md`
The superseded `0002-dense-llama-owned-range-loader.patch` is removed.
## Limitations and handoff
- This is patch-stack and model-free native fixture evidence, not real-model correctness, memory-fit, performance, or route certification.
- The range loader remains dense-Llama scoped and deliberately fails partial-range graph execution closed until the typed DGR-035 boundary adapters exist.
- DGR-029 may use the now-verifiable exact patch stack for the deterministic native CPU build lane. DGR-034 owns real dense-Llama range behavior and memory evidence.
## Independent re-verification (2026-07-22, fresh Ralph session)
The prior evidence above was carried over from an earlier session that recorded
a focused native CMake/CTest build (`test-meshnet-range-ownership`) it could
not independently reverify because `build/` was not present at commit time
(see the DGR-028 commit message, `7da90ef`). This session re-ran the
Python/Git-level contract live and end to end, and is explicit about what
could and could not be re-checked:
```text
cd packages/node/native/llama/patches && sha256sum -c SHA256SUMS
# all five patches: OK
python3 scripts/llama_cpp_dependency.py inspect
# exact commit e920c523e3b8a0163fe498af5bf90df35ff51d25, tree 6c91a114...,
# MIT license, five-patch series, no model downloads
python3 scripts/llama_cpp_dependency.py verify --workspace build/llama.cpp
# reused verified offline cache; apply -> assumption/boundary checks ->
# reverse succeeded; source left at pristine detached HEAD
python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
# git -C build/llama.cpp/source diff --cached --name-only ==
# CMakeLists.txt, cmake/meshnet-patch-stack.cmake, include/llama.h,
# src/llama-model.cpp, src/llama-model.h, src/models/llama.cpp,
# tests/CMakeLists.txt, tests/test-meshnet-range-ownership.cpp
# git -C build/llama.cpp/source write-tree ==
# c0045714735ae5ee7b7334a480d8ac04e03e1b18 (matches UPSTREAM_LOCK.json patched_tree)
python3 scripts/llama_cpp_dependency.py reverse --source-dir build/llama.cpp/source
# git -C build/llama.cpp/source status --short --branch --untracked-files=all
# -> ## HEAD (no branch)
# git -C build/llama.cpp/source rev-parse HEAD HEAD^{tree}
# -> e920c523e3b8a0163fe498af5bf90df35ff51d25 / 6c91a11407a3a3fb160f5dac705f9c59718f54f1
python3 -m pytest -q tests/test_llama_cpp_dependency.py tests/test_ralph_prd_schema.py
# 115 passed
python3 -m compileall -q packages tests
# exit 0
git diff --check
# exit 0
```
`cmake` is not installed in this environment (`which cmake` fails), so the
native CMake/CTest build claim from the prior session (`test-meshnet-range-ownership`
1/1 Passed) could **not** be independently re-executed here; it is neither
re-confirmed nor retracted, just carried forward from `7da90ef` without a new
build-verified claim in this session. Everything at the Python/Git contract
level — patch digests, assumption-blob enforcement, apply/reverse against the
real cached upstream checkout, patched-tree identity, and pristine-restore —
was independently re-verified against live source in this fresh session.
## prd.json repair (unrelated to DGR-028 itself)
Before editing `DGR-028.passes`, `prd.json` was found with its top-level
`sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/`supersededStories`
fields silently dropped again (`branchName` was also missing but had already
been restored by a prior in-flight edit) — the same ralph-tui round-trip bug
documented for DGR-019/DGR-020. Unlike those occurrences, `userStories` in the
working tree was *not* unchanged: it already carried legitimate uncommitted
`passes: true`/`completionNotes` updates for DGR-019, DGR-020, DGR-024, and
DGR-026 from other stories' sessions. The missing top-level sections were
restored from `git show HEAD:.scratch/distributed-gguf-runtime/prd.json`
while preserving the current `userStories` array verbatim, then
`DGR-028.passes` was set `true` with `completionNotes` added, and
`.scratch/distributed-gguf-runtime/issues/028-implement-numbered-patch-stack-apply-and-verification.md`
was regenerated via `scripts/ralph_prd_schema.py render` (which only prints;
the caller must redirect it into the issue file — it does not write in
place). `python3 scripts/ralph_prd_schema.py validate` and
`python3 -m pytest -q tests/test_ralph_prd_schema.py` (108 passed) both pass
against the repaired file.

View File

@@ -0,0 +1,198 @@
# DGR-029 evidence — native CMake skeleton and deterministic CPU lane
**Status:** implementation complete, live-verified in this session (2026-07-22).
**Authority:** local `prd.json` is authoritative; Gitea is a projection.
**Upstream pin:** `e920c523e3b8a0163fe498af5bf90df35ff51d25` (`llama.cpp`, from `UPSTREAM_LOCK.json`).
## What existed before this session
`scripts/llama_cpp_dependency.py` already had `build()`, `smoke()`, and `reproduce()`
functions and `UPSTREAM_LOCK.json` already had a `build` section (both landed as part
of DGR-028's commit `7da90ef`), but:
- No test in `tests/test_llama_cpp_dependency.py` ever exercised `build`/`smoke`/`reproduce`
— only `fetch`/`apply`/`reverse`/`inspect` had coverage.
- `cmake` was not installed in the DGR-028 session's environment ("`cmake` is not installed
in this environment," per its evidence), so this lane was never actually run end to end;
DGR-028's own live-verified CTest evidence used a one-off manual `cmake`/`ctest` invocation
with `-DLLAMA_BUILD_TESTS=ON` outside this driver, against a build directory that no longer
exists in this session.
- The locked `configure_flags` did not force CPU-only backend options (`GGML_CUDA`/`GGML_HIP`/
`GGML_VULKAN`/`GGML_METAL`/`GGML_BLAS`) — relying on upstream per-platform defaults (which
happen to default OFF on Linux, but are undocumented and platform-dependent), and
`LLAMA_BUILD_TESTS` was `OFF`, so no CTest lane existed at all — only a `--help` smoke check
against the unrelated stock `llama-gguf-hash` tool.
This session found and closed those three gaps rather than re-implementing from scratch.
## What changed in this session
- `packages/node/native/llama/UPSTREAM_LOCK.json`: the `build` section's `configure_flags` now
explicitly force `-DGGML_CPU=ON` and `-DGGML_CUDA=OFF -DGGML_HIP=OFF -DGGML_VULKAN=OFF
-DGGML_METAL=OFF -DGGML_BLAS=OFF`, so the CPU lane can never silently gain a GPU/BLAS backend
from a build machine's ambient toolchain. `-DLLAMA_BUILD_TESTS` flipped `OFF``ON` (required
so the `test-meshnet-range-ownership` CTest target exists at all — configuring
`LLAMA_BUILD_TESTS=ON` does not itself build every upstream test, only registers them; the
`native_targets` list still controls what actually gets compiled). Added `native_targets` entry
`test-meshnet-range-ownership` and a new `ctest_regex` field
(`"^test-meshnet-range-ownership$"`) naming the exact deterministic model-free fixture CTest
added by DGR-028's patch 0005.
- `scripts/llama_cpp_dependency.py`: factored `_cmake()`'s override/PATH/venv-sibling resolution
into a shared `_toolchain_binary(name, env_var)` and added `_ctest()` using the same resolution
(`CTEST` env override, PATH, or the sibling of the resolved `cmake` binary's environment). Added
`ctest_lane(build_dir)`, which loads the lock's `ctest_regex` and runs
`ctest --test-dir <build_dir> -R <regex> --output-on-failure`, printing output on success and
raising `DependencyError` (via the existing `_run` wrapper, which already attaches
stdout/stderr detail) on failure. Added a `ctest` CLI subcommand (`--build-dir`). Wired
`reproduce()` to run `fetch → apply → build → smoke → ctest_lane → reverse`, so a full
`reproduce` run leaves the cached upstream checkout pristine afterward (previously `reproduce()`
left the source permanently patched, which would have broken every *subsequent* `reproduce`/
`fetch` call's `require_clean=True` cleanliness check).
- `tests/test_llama_cpp_dependency.py`: added
`test_build_config_locks_an_explicit_cpu_only_deterministic_lane` (offline; asserts the lock's
`configure_flags` are CPU-only and that `ctest_regex`/`native_targets`/`smoke_binary` all agree
with each other and with `patched_paths`) and
`test_ctest_lane_raises_an_actionable_error_for_a_failing_named_test` (gated on `cmake`
availability via a `requires_cmake` marker mirroring `test_native_identity_emission.py`'s
`requires_cc` pattern; builds a tiny synthetic two-test CMake project — not the full llama.cpp
tree, so it runs in about a second — and proves `ctest_lane()` both passes silently on a passing
named test and raises `DependencyError` naming the failing test on a failing one).
## Toolchain note
Neither the ambient system Python nor `.venv-rocm` has `cmake`. This session installed `cmake`
(the PyPI wheel that bundles prebuilt binaries, version 4.4.0) into the pre-existing repo-root
`.venv` used by earlier DGR-024/DGR-026 sessions (`.venv/bin/cmake`, `.venv/bin/ctest`), which was
already on-disk from a prior session but had never had `cmake` installed into it. All commands
below were run with that `.venv/bin` prepended to `PATH`. This is the same "disposable venv for a
lightweight optional dependency" pattern DGR-024 used for `grpc`.
## Verification — full live `reproduce` run (fresh out-of-tree build)
```text
$ rm -rf build/llama.cpp/build
$ python3 scripts/llama_cpp_dependency.py reproduce
reused verified offline cache: .../build/llama.cpp/source
usage: .../build/llama.cpp/build/bin/llama-gguf-hash [options] GGUF_IN
Hash a GGUF file
options: ...
Test project .../build/llama.cpp/build
Start 27: test-meshnet-range-ownership
1/1 Test #27: test-meshnet-range-ownership ..... Passed 0.01 sec
100% tests passed out of 1
$ echo $?
0
```
Wall-clock: `real 2m16.227s` (fresh CPU compile of ggml/llama-common/llama plus the
`llama-gguf-hash` example and the `test-meshnet-range-ownership` fixture; no full `llama.cpp`
test suite or example set is built — only the two targets named in `native_targets`).
Post-run checks:
```text
$ ls build/llama.cpp/build/bin/*.so*
libggml-base.so libggml-base.so.0 libggml-base.so.0.16.0
libggml-cpu.so libggml-cpu.so.0 libggml-cpu.so.0.16.0
libggml.so libggml.so.0 libggml.so.0.16.0
libllama-common.so ... libllama.so ...
# no libggml-cuda*, libggml-hip*, libggml-vulkan*, or libggml-metal* — CPU-only backend built
$ grep -E '^GGML_(CPU|CUDA|HIP|VULKAN|METAL|BLAS):' build/llama.cpp/build/CMakeCache.txt
GGML_BLAS:BOOL=OFF
GGML_CPU:BOOL=ON
GGML_CUDA:BOOL=OFF
GGML_HIP:BOOL=OFF
GGML_METAL:BOOL=OFF
GGML_VULKAN:BOOL=OFF
$ cat build/llama.cpp/build/meshnet-build-metadata.json
{
"model_downloads": false,
"semantic_certification": false,
...
}
$ git -C build/llama.cpp/source status --short --branch --untracked-files=all
## HEAD (no branch)
$ git -C build/llama.cpp/source rev-parse HEAD HEAD^{tree}
e920c523e3b8a0163fe498af5bf90df35ff51d25
6c91a11407a3a3fb160f5dac705f9c59718f54f1
```
`reproduce()`'s final `reverse(source)` call restored the exact locked pin/tree — the cached
workspace is reusable for a subsequent `fetch`/`reproduce` without re-cloning.
## Verification — actionable toolchain failure (missing `cmake`)
```text
$ python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
$ env -i HOME="$HOME" PATH=/usr/bin:/bin python3 scripts/llama_cpp_dependency.py build \
--source-dir build/llama.cpp/source --build-dir /tmp/no-cmake-build
DGR-027 dependency error: cmake is unavailable; set CMAKE or activate the project toolchain
$ echo $?
2
$ python3 scripts/llama_cpp_dependency.py reverse --source-dir build/llama.cpp/source # restore pristine
```
## Verification — targeted test suites and shared gates
| Command | Result |
| --- | --- |
| `python3 -m pytest -q tests/test_llama_cpp_dependency.py` | `9 passed in 1.34s` (7 pre-existing + 2 new; the new gated CTest-wiring test ran for real, not skipped, since `cmake` is present in `.venv`) |
| `python3 -m pytest -q tests/test_llama_cpp_dependency.py tests/test_ralph_prd_schema.py` | `117 passed` |
| `python3 -m compileall -q packages tests` | exit 0 |
| `git diff --check` | exit 0 (no output) |
## Ensuring build success does not advertise capability
- The locked `configure_flags` disable every accelerator backend explicitly
(`GGML_CUDA/HIP/VULKAN/METAL/BLAS=OFF`) rather than relying on per-platform defaults, so a
successful configure/build can only ever mean "the CPU reference backend compiled" — never an
accelerator claim, and never dependent on whether the build host happens to have a GPU SDK
installed.
- `meshnet-build-metadata.json` (written by `build()`) records `model_downloads: false` and
`semantic_certification: false` alongside the exact commit/patch/flag identities — the artifact
itself, not just prose, states this build proves toolchain compilation only.
- The two targets actually compiled are `llama-gguf-hash` (a stock upstream file-hashing utility;
no inference) and `test-meshnet-range-ownership` (a model-free fixture that writes a tiny
synthetic GGUF and asserts range-ownership bookkeeping — no real model, no generation, no
numerical/backend correctness claim). Neither exercises inference, MoE, attention, or any
DeepSeek V4 semantic path.
## Changed files
- `packages/node/native/llama/UPSTREAM_LOCK.json`
- `scripts/llama_cpp_dependency.py`
- `tests/test_llama_cpp_dependency.py`
- `.scratch/distributed-gguf-runtime/prd.json`
- `.scratch/distributed-gguf-runtime/issues/029-create-the-native-cmake-skeleton-and-deterministic-cpu-lane.md`
- `.scratch/distributed-gguf-runtime/evidence/DGR-029/README.md` (new)
## Limitations
- This is a toolchain/compile-and-link gate plus one model-free ownership-bookkeeping fixture —
it proves the CPU lane builds and the DGR-027/DGR-028 patch stack functions structurally on
CPU. It proves nothing about real-model correctness, memory-fit, performance, or any
backend/model/recipe certification; `stock_glm_limitations` in `UPSTREAM_LOCK.json` and DGR-028's
own limitations continue to apply unchanged.
- `cmake`/`ctest` are not installed system-wide or in `.venv-rocm` in this environment; they were
installed only into the pre-existing repo-root `.venv` for this session's verification (and for
the new gated pytest test, which is skipped in any environment lacking `cmake`). A future session
without that `.venv` (or without re-installing `cmake` into it) will see the same "cmake is
unavailable" actionable failure demonstrated above, not a silent pass.
- Only the two named targets are compiled (`llama-gguf-hash`, `test-meshnet-range-ownership`); a
broad `cmake --build ... --target test` / full upstream test suite is out of scope here, exactly
as DGR-028 recorded ("not presented as a full-suite gate").
- CUDA/ROCm/Vulkan/Metal compile lanes remain unimplemented; this story only establishes the CPU
lane "before accelerator matrix work," per its objective. Those lanes are separate future work.
## Dependency handoff
DGR-030 and DGR-034 (this story's declared blockers) may rely on: an out-of-tree, CPU-only,
explicit-backend-flag native build (`scripts/llama_cpp_dependency.py build`/`reproduce`) that
compiles the exact DGR-027/DGR-028 patched pin and runs a real CTest lane
(`test-meshnet-range-ownership`) proving the patch stack's range-ownership bookkeeping compiles
and passes on CPU. Any accelerator (CUDA/ROCm/Vulkan/Metal) lane, any real-model load, and any
backend/model/recipe capability certification remain unimplemented and must not be assumed from
this story's green build alone.

View File

@@ -0,0 +1,275 @@
# DGR-030 evidence — accelerator build presets and native CI/build matrix
**Status:** implementation complete, live-verified in this session (2026-07-23).
**Authority:** local `prd.json` is authoritative; Gitea is a projection.
**Upstream pin:** `e920c523e3b8a0163fe498af5bf90df35ff51d25` (`llama.cpp`, unchanged from DGR-027..029).
## What existed before this session
DGR-029 locked exactly one build lane — the deterministic CPU-only lane — in
`UPSTREAM_LOCK.json`'s `build` section, plus `scripts/llama_cpp_dependency.py`'s
`build()`/`smoke()`/`ctest_lane()`/`reproduce()`. There was no accelerator
preset, no SDK-availability probing, and no matrix runner: only the one CPU
lane existed, and there was no mechanism that could ever advertise a GPU
backend as compiled or capable.
## What changed in this session
- `packages/node/native/llama/UPSTREAM_LOCK.json`: added a new top-level
`accelerator_presets` object with one entry each for `cuda` (`GGML_CUDA`),
`rocm` (`GGML_HIP`), `vulkan` (`GGML_VULKAN`), and `metal` (`GGML_METAL`).
Each entry names only the one backend flag it flips and an `sdk_probe`
(a binary to resolve on `PATH`, an optional env-var override, and — for
Metal — a `platform_only: "darwin"` gate). **The existing `build` section
— the deterministic CPU default DGR-029 locked — is untouched.**
- `scripts/llama_cpp_dependency.py`:
- `_load_lock()` now calls a new `_verify_accelerator_presets()`, which
fail-closed-rejects any preset whose named backend flag is not `OFF` in
the CPU default's `configure_flags` — structurally guaranteeing a preset
can only ever *add* one backend on top of the untouched CPU baseline,
never redefine it.
- `accelerator_configure_flags(lock, name)` returns a **new** flag list —
the CPU default's own `configure_flags` list is never mutated — with
exactly the named preset's backend flag flipped `ON` and every other flag
(including `GGML_CPU=ON`, the fallback ops backend GPU builds still need)
left exactly as the CPU default declares it.
- `_sdk_probe(probe)` / `accelerator_status(name, lock)` resolve a lane's
SDK without ever raising: an absent SDK is returned as
`{"available": false, "reason": "<binary> is unavailable on PATH"}` (or
a platform-mismatch reason for Metal), so "unavailable" is data a caller
reports, never an exception a caller has to remember to catch.
- `accelerator_build(source, name, build_dir)` compiles one lane into its
own out-of-tree `build_dir` (an isolated directory, never DGR-029's CPU
`build_dir`), using the same patched-source verification and
`native_targets` as the CPU lane, then writes a
`meshnet-build-metadata.json` recording the exact `commit`/`commit_tree`,
per-patch SHA-256 digests, the lane's overridden `configure_flags`, the
resolved `cmake`/`cxx`/SDK-binary versions/paths, and explicit
`model_downloads: false`, `hardware_execution: false`,
`hardware_certified: false`, `semantic_certification: false` fields plus
a `note` stating the lane is registered-dark until a real-hardware
certification record exists. It **never** calls `smoke()`/`ctest_lane()`
— running a binary linked against a real accelerator backend would touch
real hardware, which this story deliberately keeps out of scope.
- Added `accelerator-status --name <lane>` and
`accelerator-build --name <lane> --source-dir --build-dir` CLI
subcommands, mirroring the existing `ctest`/`build` subcommand pattern.
- `scripts/native_accelerator_matrix.py` (new): the native CI/build matrix.
`run_matrix(workspace)` fetches and applies the locked pin/patch stack once,
runs the unchanged CPU lane (build → smoke → ctest, exactly DGR-029's
contract), then for each `accelerator_presets` entry either reports
`{"status": "skipped", "reason": ...}` (SDK absent) or compiles it via
`accelerator_build` and reports `{"status": "built", ...}` — never silently
treating a skip as a pass. Any `DependencyError` from a lane (CPU or
accelerator) is caught per-lane and reported as `{"status": "failed", ...}`
without aborting the remaining lanes or skipping cleanup. `reverse()` always
runs in a `finally`, restoring the exact pristine pin/tree regardless of
lane outcomes. The CLI prints a JSON report and exits non-zero only if any
lane actually `failed` (a `skipped` lane never fails the run).
- `tests/test_llama_cpp_dependency.py`: added 7 new tests —
`test_accelerator_presets_isolate_one_backend_without_touching_the_cpu_default`
(every preset flips exactly its own flag and the CPU default list is never
mutated), `test_accelerator_configure_flags_rejects_an_unknown_lane`,
`test_accelerator_status_reports_unavailable_sdks_without_raising` (asserts
the exact reason string for cuda/rocm/vulkan/metal absence),
`test_accelerator_status_honors_an_explicit_sdk_override`,
`test_accelerator_status_rejects_an_unknown_lane`,
`test_accelerator_build_refuses_to_compile_an_unavailable_lane` (asserts no
build directory is created), and a `requires_cmake`-gated
`test_accelerator_build_compiles_the_available_lane_with_isolated_evidence`,
which builds a tiny synthetic CMake project (not the full llama.cpp tree) to
prove `accelerator_build`'s "SDK present" path really configures with the
overridden flag, compiles, and writes the registered-dark metadata — in
about a second, without a real GPU SDK.
- `tests/test_native_accelerator_matrix.py` (new): 3 offline tests exercising
`run_matrix`'s orchestration with `llama_cpp_dependency`'s
fetch/apply/reverse/build/smoke/ctest_lane/accelerator_status/
accelerator_build stubbed out — proving unavailable SDKs are reported
`skipped` (never a false pass), an available accelerator lane is compiled
without ever calling `smoke`/`ctest_lane`, and a lane failure is reported
per-lane without aborting sibling lanes or skipping the `reverse()` cleanup.
## Toolchain note
As in DGR-029, neither the ambient system Python nor `.venv-rocm` has `cmake`;
this session's `.venv` also had no `cmake` (a prior session's install did not
persist). This session ran `.venv/bin/python3 -m ensurepip --upgrade` (no
`pip` was present in `.venv` either) and then
`.venv/bin/python3 -m pip install cmake`, landing the same PyPI wheel
(`cmake==4.4.0`) DGR-029 used, at `.venv/bin/cmake` / `.venv/bin/ctest`. All
commands below were run with that `.venv/bin` prepended to `PATH`. No CUDA,
ROCm, or Vulkan SDK (`nvcc`, `hipcc`, `glslc`) is installed in this
environment, and the host platform is Linux, not `darwin` — so all four
accelerator lanes are genuinely `skipped` in this environment's own live run
below, which is real evidence for AC2 ("unavailable SDKs ... explicit
unavailable/skipped lanes"), not a simulated one.
## Verification — live native CI/build matrix run
```text
$ rm -rf build/llama.cpp/build build/llama.cpp/build-cuda build/llama.cpp/build-rocm build/llama.cpp/build-vulkan build/llama.cpp/build-metal
$ python3 scripts/native_accelerator_matrix.py
reused verified offline cache: .../build/llama.cpp/source
usage: .../build/llama.cpp/build/bin/llama-gguf-hash [options] GGUF_IN
...
Test project .../build/llama.cpp/build
Start 27: test-meshnet-range-ownership
1/1 Test #27: test-meshnet-range-ownership ..... Passed 0.01 sec
100% tests passed out of 1
{
"failed_lanes": [],
"hardware_certified": false,
"lanes": [
{
"build_dir": ".../build/llama.cpp/build",
"lane": "cpu",
"metadata": {
"cmake": "cmake version 4.4.0",
"commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
"commit_tree": "6c91a11407a3a3fb160f5dac705f9c59718f54f1",
"configure_flags": [
"-DCMAKE_BUILD_TYPE=Release", "-DLLAMA_BUILD_TESTS=ON",
"-DLLAMA_BUILD_EXAMPLES=ON", "-DLLAMA_BUILD_SERVER=OFF",
"-DLLAMA_BUILD_TOOLS=OFF", "-DLLAMA_BUILD_APP=OFF", "-DLLAMA_CURL=OFF",
"-DGGML_CPU=ON", "-DGGML_BLAS=OFF", "-DGGML_CUDA=OFF",
"-DGGML_HIP=OFF", "-DGGML_VULKAN=OFF", "-DGGML_METAL=OFF"
],
"cxx": "c++ (GCC) 15.2.1 20260123 (Red Hat 15.2.1-7)",
"model_downloads": false,
"patches": { "...": "... (5 entries, unchanged sha256 digests from DGR-029)" },
"semantic_certification": false
},
"status": "built"
},
{"lane": "cuda", "reason": "nvcc is unavailable on PATH", "status": "skipped"},
{"lane": "rocm", "reason": "hipcc is unavailable on PATH", "status": "skipped"},
{"lane": "vulkan", "reason": "glslc is unavailable on PATH", "status": "skipped"},
{"lane": "metal", "reason": "platform 'linux' is not 'darwin'", "status": "skipped"}
],
"note": "A `built` lane means it compiled with the exact recorded compiler/SDK/upstream-pin/patch-stack/build-option evidence — it never means an accelerator device was exercised. Every backend/model/recipe lane stays registered-dark until a separate real-hardware certification record exists."
}
$ echo $?
0
```
Wall-clock: `real 2m19.797s` — matches DGR-029's ~2m16s CPU-lane compile; no
accelerator lane actually compiled in this environment (all four SDKs are
genuinely absent), so this run's added cost over DGR-029's own CPU-only
`reproduce()` is just the four fast SDK probes.
Post-run checks (source checkout left pristine by the matrix's `reverse()`):
```text
$ git -C build/llama.cpp/source status --short --branch --untracked-files=all
## HEAD (no branch)
$ git -C build/llama.cpp/source rev-parse HEAD HEAD^{tree}
e920c523e3b8a0163fe498af5bf90df35ff51d25
6c91a11407a3a3fb160f5dac705f9c59718f54f1
$ ls build/llama.cpp/ | grep build
build
```
Only the CPU lane's `build/` directory was created — no `build-cuda`,
`build-rocm`, `build-vulkan`, or `build-metal` directory exists, because every
accelerator lane was genuinely skipped rather than attempted.
## Verification — targeted test suites and shared gates
| Command | Result |
| --- | --- |
| `python3 -m pytest -q tests/test_llama_cpp_dependency.py tests/test_native_accelerator_matrix.py` | `19 passed` (9 pre-existing + 7 new accelerator-lane tests in `test_llama_cpp_dependency.py`, 3 new in `test_native_accelerator_matrix.py`; the `requires_cmake`-gated compile test ran for real, not skipped) |
| `python3 -m compileall -q packages tests` | exit 0 |
| `git diff --check -- packages/node/native/llama/UPSTREAM_LOCK.json scripts/llama_cpp_dependency.py tests/test_llama_cpp_dependency.py scripts/native_accelerator_matrix.py tests/test_native_accelerator_matrix.py` | exit 0 |
| `python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json` | `OK: 55 stories validated.` |
`git diff --check` against the full working tree separately reports one
pre-existing trailing-whitespace line in `.ralph-tui-run.log`, which was
already modified before this session started (see the session's initial
`git status`) and is unrelated to this story's scope; it is excluded above by
naming this story's own changed files explicitly.
`python3 -m pytest -q tests/test_ralph_prd_schema.py` reports `55 failed, 53
passed` in this session (all `test_render_issue_markdown_matches_committed_file`
drift between `prd.json` and committed issue Markdown for other stories,
e.g. `DGR-053`..`DGR-071`). `git stash`-ing this session's changes and rerunning
reproduces `56 failed, 52 passed` identically — the same 56 failures minus the
one this session's own `DGR-030` regeneration fixed, confirming the remaining
55 predate this story and are out of scope to fix here. This session did
regenerate `.scratch/distributed-gguf-runtime/issues/030-add-accelerator-
build-presets-and-native-ci-matrix.md` via
`python3 scripts/ralph_prd_schema.py render ... DGR-030` so DGR-030's own
generated issue Markdown matches `prd.json` byte-for-byte (confirmed by the
`test_render_issue_markdown_matches_committed_file[DGR-030]` case no longer
appearing in the failure list).
## Ensuring build success does not advertise capability
- Every accelerator lane's `meshnet-build-metadata.json` explicitly records
`hardware_execution: false`, `hardware_certified: false`, and
`semantic_certification: false`, plus a `note` stating the lane is
registered-dark until a separate real-hardware certification record exists
— the same "artifact states this, not just prose" pattern DGR-029 used for
the CPU lane's `model_downloads`/`semantic_certification` fields.
- `accelerator_build` never runs `smoke()` or `ctest_lane()`: it only
configures and compiles the exact `native_targets` DGR-029 already locked
(`llama-gguf-hash`, `test-meshnet-range-ownership`) — no binary linked
against a real accelerator backend is ever executed by this story's code.
- `_verify_accelerator_presets()` structurally refuses any preset whose
backend flag is not `OFF` in the locked CPU default, so a preset can never
be defined in a way that redefines (rather than adds one backend on top of)
DGR-029's deterministic CPU lane.
- The matrix's top-level report always carries `"hardware_certified": false`
regardless of how many lanes built, and its `note` field states this
explicitly for any consumer reading only the report, not the per-lane
metadata.
## Limitations
- This story proves accelerator lanes *compile* with correct, isolated
flags and preserves exact evidence when a lane's SDK is present. It proves
nothing about numerical correctness, performance, or any backend/model/
recipe capability on real accelerator hardware — that is explicitly
deferred to DGR-041 (capability registration), DGR-053 (real 2-4 stage
certification), and DGR-067 (capability matrix certification), all of which
remain unimplemented.
- No CUDA, ROCm, or Vulkan SDK, and no macOS/Metal toolchain, is available in
this session's environment, so the "compile an available accelerator lane"
path is proven end-to-end only via the `requires_cmake`-gated synthetic-
project unit test and the offline matrix-orchestration tests, not via a
live compile of the real llama.cpp tree under `GGML_CUDA=ON` (etc.). A
future session with a real SDK installed will exercise
`accelerator_build`'s real-lane path against the genuine llama.cpp source
for the first time; nothing in this story's design assumes that hasn't
happened yet.
- The accelerator lanes reuse the CPU lane's exact `native_targets`
(`llama-gguf-hash`, `test-meshnet-range-ownership`), so a passing
accelerator compile also proves the DGR-027/DGR-028 patch stack's
range-ownership code compiles under that backend flag combination — but,
per the point above, only structurally; it says nothing about GPU
execution correctness.
- `cmake`/`ctest` remain absent system-wide in this environment; this session
reinstalled them into `.venv` exactly as DGR-029 did, and that install does
not appear to persist across sessions (this session found `.venv` without
`cmake` despite DGR-029's evidence recording its earlier install). A future
session without a `cmake`-equipped `.venv` will see the same actionable
"cmake is unavailable" failure DGR-029 demonstrated, not a silent pass, and
the new `requires_cmake`-gated tests will be skipped rather than failing.
- `git diff --check` and `tests/test_ralph_prd_schema.py` both carry
pre-existing, out-of-scope failures unrelated to this story (see the gates
table above); this story's own changed files pass both checks cleanly.
## Dependency handoff
DGR-053 (real 2-4 stage certification), DGR-067 (capability matrix
certification), and DGR-068 (packaged releases) may rely on: four isolated,
out-of-tree accelerator build presets (`cuda`/`rocm`/`vulkan`/`metal`) in
`UPSTREAM_LOCK.json`'s `accelerator_presets`, each toggling exactly one
backend flag on top of DGR-029's unchanged CPU default; a native CI/build
matrix (`scripts/native_accelerator_matrix.py`) that compiles every
SDK-available lane with full compiler/SDK/upstream-pin/patch-stack/build-
option evidence and reports SDK-unavailable lanes as explicit `skipped`
lanes, never a false pass; and a compile-only contract (no lane here ever
runs a binary against real accelerator hardware). Real-hardware execution,
numerical correctness, performance measurement, and backend/model/recipe
certification for any accelerator remain entirely unimplemented and must not
be assumed from any lane's green compile.

View File

@@ -0,0 +1,237 @@
# DGR-031 evidence — the project-owned `ShardEngine` interface
**Completed:** 2026-07-23
**Branch:** `ralph/distributed-gguf-runtime`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
**Dependencies:** DGR-021 (`evidence/DGR-021/README.md` — versioned activation
envelope, `NamedTensor`/`ActivationEnvelope` as the project-owned wire-envelope
layer), DGR-025 (`evidence/DGR-025/README.md` — exact artifact/runtime recipe
identity; both read before changing code).
## Objective
Isolate worker/protocol code from llama.cpp internals behind a stable
project-owned engine contract, so a fake fixture engine (DGR-032) and a real
llama.cpp-backed engine (DGR-037) are interchangeable subclasses of one
interface.
## What was found live before changing code
Per RALPH-CONTEXT, legacy pass states were not trusted; the live surrounding
contracts were read and exercised before designing this one:
- `packages/node/meshnet_node/shard_lifecycle.py` (DGR-022) already defines a
versioned RPC/session lifecycle contract — `StructuredStatus`, `StatusCode`,
`CacheExpectation`, `CacheResult`, `LifecycleState`, `SessionLifecycle` — but
it is explicitly the *wire RPC* contract "consumed by a future generated
gRPC binding," not an execution-engine boundary.
- `packages/node/meshnet_node/native_backend.py` (DGR-025) is the identity
boundary for the native GGUF artifact — it derives and attests a
`ShardIdentity`, but does not define an execution contract either.
- `packages/node/meshnet_node/protocol.py` (DGR-021) defines a project-owned
`NamedTensor`/`ActivationEnvelope` for activation traffic *between shard
hops over the network*, distinct from the generated-protobuf wire ABI in
`native_protocol`.
- `packages/node/meshnet_node/shard_runtime_server.py` (DGR-024) is today a
real gRPC servicer that proves wire fidelity by checksumming and echoing
bytes — it has no execution engine behind it yet; that seam is exactly
where `ShardEngine` plugs in for DGR-037.
- `packages/node/meshnet_node/architecture_boundary.py` established the
precedent this story follows for tail output: `TailOutput.sampled_token()`
never exposes raw logits, only a sampled token id.
- No `ShardEngine` (or `shard_engine`) symbol existed anywhere in the
repository prior to this story (confirmed by
`grep -rn -i "shardengine\|shard_engine"` across `.py`/`.md`, which returned
only planning-document prose naming it as future work).
Live verification of the pre-existing dependency contracts before adding new
code: `PYTHONPATH=packages/node:packages/tracker .venv/bin/python3 -m pytest -q
tests/test_shard_lifecycle.py tests/test_activation_envelope.py
tests/test_architecture_boundary.py tests/test_native_shard_protocol.py
tests/test_shard_runtime_harness.py``95 passed, 3 skipped`.
## What was added (this story's change)
### `packages/node/meshnet_node/shard_engine.py` (new)
The `ShardEngine` boundary: an `abc.ABC` with eight abstract operations —
`load`, `capabilities`, `prefill`, `decode`, `cancel`, `release`, `health`,
`metrics` — matching the acceptance criterion's list exactly (`prefill`/
`decode` share one operation family; their shared result type is what the
criterion calls the "boundary/logits result"). Every request/result type is a
frozen dataclass built from plain `str`/`int`/`bytes`/`Mapping` values:
- `EngineTensor` / `BoundaryBundle` — the project-owned named-tensor
activation crossing a shard boundary (head/middle/tail-in). Deliberately a
*new*, minimal type distinct from both `native_protocol.pb.TensorBundle`
(generated-protobuf ABI) and `protocol.NamedTensor`/`ActivationEnvelope`
(wire-framing/fragmentation concerns irrelevant to model execution) — a
fourth, execution-facing layer underneath the three that already existed.
- `TokenOutput` — a tail shard's sampled result: a token id (+ optional
decoded text), never a raw logits tensor.
- `MtpHook` — reserved multi-token-prediction hook; its own `__post_init__`
raises if constructed with `enabled=True`, so the type exists (fixing its
field shape for DGR-051/DGR-066) without any code path being able to turn it
on before DGR-066, matching RALPH-CONTEXT's "MTP is reserved and off for
alpha."
- `ArchitectureAuxStateHook` — reserved per-shard architecture auxiliary state
(V4 CSA/HCA/SWA/indexer/compressor and similar); has no wire encoding and is
never embedded in a `BoundaryBundle`, matching RALPH-CONTEXT's "remain local
... never carried over the WAN seam."
- `LoadRequest`/`LoadResult`, `EngineCapabilities`, `PrefillRequest`/
`DecodeRequest` (exactly one of `token_ids`/`token_id` (head) or `input`
(middle/tail) required — enforced in `__post_init__`), `StepResult` (a
successful result must carry an output; `cache_result` reuses
`shard_lifecycle.CacheResult`), `HealthResult`, `MetricsResult`.
- Status vocabulary is reused, not reinvented: `StructuredStatus`/
`StatusCode`/`CacheExpectation`/`CacheResult` are imported from
`shard_lifecycle` (already project-owned and version-stable) rather than a
parallel enum living alongside it.
- The module imports nothing from `native_protocol`, `grpc`, or `ctypes`
verified structurally, not just by convention (see tests below).
### `tests/shard_engine_contract.py` (new)
A reusable, non-`test_`-prefixed helper: `assert_shard_engine_contract(make_engine)`
takes a zero-arg engine factory and runs nine lifecycle checks — health before
load, load→capabilities range/MTP-off, prefill→decode determinism (byte-identical
output replayed on a fresh session), middle-shard boundary-bundle-in/out vs.
head/tail token-output, deterministic cache-miss on an unopened session,
stale-route-epoch rejection, cancel-then-decode rejection (+ cancel
idempotency), release-then-decode rejection (+ release idempotency), and
metrics reporting cancelled sessions. DGR-032's fixture and DGR-037's
llama.cpp binding are both expected to import this and pass it against their
own engine, proving identical lifecycle semantics without duplicating the
checks.
### `tests/test_shard_engine.py` (new)
- `_ReferenceEngine`: a minimal in-memory `ShardEngine` used only to prove the
shared contract is non-vacuous. It is explicitly *not* the DGR-032
deterministic fixture (no delay/memory-pressure/malformed/crash injection —
that is DGR-032's own, larger scope); the docstring says so to prevent this
story's evidence from being read as inherited completion credit for DGR-032.
- Dataclass validation tests: abstract-class instantiation refusal, tensor/
bundle/token-output field validation, MTP-hook enable refusal, exactly-one-
input-kind enforcement on `PrefillRequest`/`DecodeRequest`, `LoadRequest`
shard-range-vs-total-layers validation, `StepResult` output-required-on-OK.
- `test_shard_engine_module_imports_no_native_or_grpc_or_wire_abi_types`:
walks `vars(shard_engine_module)` and asserts no bound name's `__name__` is
`ctypes`, `grpc`, or `meshnet_node.native_protocol` — a structural check
(not a docstring-text grep, which produced a false positive on first draft
because the module's own docstring *names* `ggml_tensor` as an example of
what must never appear) that the ABI-isolation acceptance criterion holds.
### `.scratch/distributed-gguf-runtime/prd.json` / issue markdown
Marked `DGR-031.passes = true` with `completionNotes`; regenerated
`issues/031-introduce-the-project-owned-shardengine-interface.md` via
`scripts/ralph_prd_schema.py render` so it matches `prd.json` byte-for-byte.
## Acceptance criteria → evidence
1. **load/capabilities/prefill/decode/boundary-logits-result/cancel/release/
health/metrics** — `ShardEngine`'s eight abstract methods plus
`StepResult.output: BoundaryBundle | TokenOutput | None`. Verified by
`test_reference_engine_obeys_the_shared_shard_engine_contract` and the
middle-shard-vs-tail-shard assertion inside
`assert_shard_engine_contract`.
2. **No `ggml_tensor`/llama context/scheduler/ABI-owned structure** — every
type in `shard_engine.py` is a plain dataclass over `str`/`int`/`bytes`/
`Mapping`; no import of `native_protocol`, `grpc`, or `ctypes`. Verified by
`test_shard_engine_module_imports_no_native_or_grpc_or_wire_abi_types`.
3. **Reserved typed MTP/architecture-aux-state hooks, not enabled**
`MtpHook.__post_init__` raises on `enabled=True`; `ArchitectureAuxStateHook`
carries opaque shard-local state with no wire path. Verified by
`test_mtp_hook_is_reserved_and_refuses_to_enable` and
`test_architecture_aux_state_hook_carries_opaque_shard_local_state`, plus
`assert_shard_engine_contract`'s `caps.supports_mtp is False` check.
4. **Contract tests proving fake and future llama implementations obey
identical lifecycle semantics** — `tests/shard_engine_contract.py` is
written to be imported by DGR-032 and DGR-037 against their own engines;
`test_shard_engine.py` proves it is real by running it against
`_ReferenceEngine`.
5. **Gates + this handoff** — below.
## Commands and results
```bash
PYTHONPATH=packages/node:packages/tracker .venv/bin/python3 -m pytest -q tests/test_shard_engine.py
```
```text
12 passed in 0.13s
```
```bash
PYTHONPATH=packages/node:packages/tracker .venv/bin/python3 -m pytest -q \
tests/test_shard_engine.py tests/test_shard_lifecycle.py \
tests/test_architecture_boundary.py tests/test_activation_envelope.py \
tests/test_native_shard_protocol.py tests/test_shard_runtime_harness.py
```
```text
95 passed, 3 skipped in 3.65s
```
```bash
.venv/bin/python3 -m compileall packages/node/meshnet_node/shard_engine.py tests/shard_engine_contract.py tests/test_shard_engine.py
```
```text
Compiling 'packages/node/meshnet_node/shard_engine.py'...
Compiling 'tests/shard_engine_contract.py'...
Compiling 'tests/test_shard_engine.py'...
```
```bash
git diff --check
```
```text
(no output — clean)
```
## Limitations
- `tests/` as a whole does not collect cleanly in this environment: 27
pre-existing test modules fail to import for missing optional dependencies
(`cryptography`, etc.) unrelated to this story. Reproduced identically with
`git stash` before this session's change (`27 errors during collection`),
so this is pre-existing environment state, not a regression introduced
here. This story's own gates were run as the targeted, scoped test set
above per the shared quality gates' own wording ("Targeted deterministic
tests pass").
- The contract in `shard_engine_contract.py` proves *lifecycle* semantics
(gating, cache-miss/stale-epoch/cancel/release, boundary-vs-token output
shape) are identical across implementations. It does not — and cannot yet
— prove numerical parity between a fake and a real engine; that is
DGR-036's explicit job once DGR-032 and DGR-037 both exist.
- `_ReferenceEngine` in `test_shard_engine.py` is intentionally minimal
(no delay/memory-pressure/malformed-output/crash injection). DGR-032's
acceptance criteria require those independently; nothing here should be
read as satisfying them.
- No gRPC/CMake/native-build changes were needed or made — this story is
pure Python interface/type definition (`evidenceClass: model-free`,
`hardware: none`), so the native CMake/CTest and patch-stack gates in the
shared quality-gate list do not apply here (consistent with DGR-021/DGR-025,
which record the same non-applicability for non-native stories).
## Dependency handoff
- **DGR-032** (fake `ShardEngine`): subclass `ShardEngine`, add delay/memory-
pressure/malformed-output/crash injection, and pass the *same*
`assert_shard_engine_contract` from `tests/shard_engine_contract.py`
against it — no new contract vocabulary should be needed.
- **DGR-034/DGR-035** (range-aware GGUF ownership, boundary I/O): `LoadRequest`
already carries `shard_start`/`shard_end`/`total_layers`/`recipe`; `capabilities()`
reports the authoritative range via `EngineCapabilities.is_head`/`is_tail`.
`BoundaryBundle.token_id_sideband` is reserved for the first-three-hash-
routed-layers V4 requirement RALPH-CONTEXT documents.
- **DGR-037** (bind llama.cpp to the worker): implement `ShardEngine` as a
thin wrapper around the native artifact from `native_backend.py`/
`runtime_recipe.py`; `shard_runtime_server.py`'s `Session`/`GetCapability`/
`Health`/`Cancel`/`Release` handlers become the translation layer between
`pb.*` wire messages and this module's request/result types — this story
intentionally does not touch `shard_runtime_server.py` itself, since that
wiring is DGR-037's scope.
- **DGR-051** (V4 `ShardEngine` adapter): `MtpHook`/`ArchitectureAuxStateHook`
fix the field shape now so the V4 adapter does not need a breaking change
to enable MTP after DGR-066 or to carry CSA/HCA/SWA/indexer/compressor
state.

View File

@@ -0,0 +1,259 @@
# DGR-032 evidence — deterministic fake `ShardEngine`
**Completed:** 2026-07-23
**Branch:** `ralph/distributed-gguf-runtime`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
**Dependencies:** DGR-031 (`evidence/DGR-031/README.md` — the project-owned
`ShardEngine` abstract contract, `tests/shard_engine_contract.py`'s
`assert_shard_engine_contract`, and its own dependency-handoff note that
DGR-032 should "subclass `ShardEngine`, add delay/memory-pressure/malformed-
output/crash injection, and pass the *same* `assert_shard_engine_contract`
... — no new contract vocabulary should be needed").
## Objective
Provide an engine fixture that deterministically transforms typed boundary
bundles and session state: head/middle/tail, prefill/decode, cancellation,
release, isolated per-session epoch state, deterministic cache-miss/stale-
epoch failures, and configurable delay/memory-pressure/malformed-output/
crash-injection fault surfaces — all without llama.cpp, a GPU, or any I/O.
## What was found live before changing code
- `packages/node/meshnet_node/shard_engine.py` (DGR-031): the abstract
`ShardEngine` with eight operations (`load`, `capabilities`, `prefill`,
`decode`, `cancel`, `release`, `health`, `metrics`) and its project-owned
dataclasses (`LoadRequest`, `EngineCapabilities`, `PrefillRequest`/
`DecodeRequest`, `StepResult`, `BoundaryBundle`/`EngineTensor`,
`TokenOutput`, `HealthResult`, `MetricsResult`).
- `tests/shard_engine_contract.py` (DGR-031): the reusable
`assert_shard_engine_contract(make_engine)` helper — nine lifecycle checks
any implementation must pass, explicitly designed to be imported by
DGR-032 and DGR-037 against their own engines.
- `tests/test_shard_engine.py` (DGR-031): its `_ReferenceEngine` is
explicitly documented as *not* the DGR-032 fixture ("no delay/memory-
pressure/malformed/crash injection... that is a separate, larger story") —
confirming this story starts from nothing, not inherited credit.
- `grep -rn -i "fakeshardengine\|fake_shard_engine"` across `.py`/`.md`
returned no prior matches — no fake engine existed before this story.
- No file in `packages/node/meshnet_node/` wires a `ShardEngine` into
`shard_runtime_server.py` yet (confirmed by grep for `ShardEngine`/
`shard_engine` in that file — no matches); that wiring is DGR-037's scope,
so this fixture is a standalone, importable engine only.
Live verification of the pre-existing dependency contract before adding new
code:
```bash
PYTHONPATH=packages/node:packages/tracker .venv/bin/python3 -m pytest -q tests/test_shard_engine.py
```
```text
12 passed in 0.13s
```
## What was added (this story's change)
### `packages/node/meshnet_node/fake_shard_engine.py` (new)
`FakeShardEngine(ShardEngine)` — a pure-Python, deterministic fixture:
- **Determinism.** Every `prefill`/`decode` output is `SHA-256(seed_bytes +
idempotency_step)`, where `seed_bytes` is derived from `token_ids` (head)
or the input `BoundaryBundle`'s tensor bytes plus any `token_id_sideband`
(middle/tail-in). Replaying identical inputs on a brand-new session
produces byte-identical output — proven by
`assert_shard_engine_contract`'s own determinism check and reused directly.
- **Head/middle/tail.** Tail shards (`shard_end >= total_layers - 1`) return
a `TokenOutput` sampled into `[0, TOKEN_ID_VOCAB_SIZE)`; head/middle shards
return a `BoundaryBundle` tagged `boundary_point="post_head_residual"` or
`"post_middle_residual"` respectively, so the three cases are
distinguishable in fixture output, not just in the load request. A middle
shard's `token_id_sideband` passes through unchanged from its input bundle
to its output bundle (the V4 first-three-hash-routed-layers requirement
RALPH-CONTEXT documents), never invented or dropped.
- **Isolated session/epoch state.** `_sessions: dict[str, _SessionState]`
keyed by `session_id`; each session tracks its own `epoch`/`cancelled`
flag. A stale epoch, cancel, or release on one session never touches
another's state (`test_session_state_is_isolated_between_two_concurrent_sessions`
proves a stale-epoch rejection and a cancel on session `"a"` leave session
`"b"` fully serviceable). Decoding an unopened session is a deterministic
`NOT_FOUND`/`CacheResult.MISS`, not an exception.
- **Configurable delay.** `FakeShardEngineConfig.step_delay_seconds` +
injectable `sleep` hook (defaults to `time.sleep`, overridable in tests so
they don't block wall-clock time) — invoked once per `prefill`/`decode`
call before computing the deterministic output.
- **Configurable memory pressure.** `FakeShardEngineConfig.memory_budget_bytes`
— the engine accumulates `_bytes_used` across every step's seed bytes;
once a step would push cumulative usage past the budget, that step
deterministically returns `StatusCode.RESOURCE_EXHAUSTED` (`retryable=True`)
with no output, instead of computing one.
- **Configurable malformed output.** `FakeShardEngineConfig.malformed_output`
— when set, the engine still reports `StatusCode.OK` (the point is a
buggy-but-"successful"-looking response, not a status-coded failure) but
the payload is structurally valid, semantically wrong: a tail `TokenOutput`
is pushed past `MALFORMED_TOKEN_ID_FLOOR` (outside the fixture's own
advertised vocab), and a head/middle `BoundaryBundle` gets an
`architecture` field prefixed `"malformed:"` and its tensor `data`
truncated to one byte — both structurally valid per `EngineTensor`'s and
`BoundaryBundle`'s own `__post_init__` validation (which does not
cross-check `data` length against `shape`/`dtype`), so a consumer must
actually check shape/semantics, not just status codes, to catch it.
- **Configurable crash injection.** `FakeShardEngineConfig.crash_after_calls`
+ `crash_exception_factory` — after the configured number of
`prefill`/`decode` calls, the engine raises an arbitrary exception (default
`RuntimeError`, injectable) directly out of the call instead of returning a
`StepResult`. This is deliberately *not* wrapped in `EngineError`/
`StructuredStatus`: it simulates a whole-process failure (what a worker
supervisor — DGR-040 — must catch and restart around), which is a
different failure mode from a graceful status-coded rejection.
- **Fixture-vs-real marker.** `FakeShardEngine.EVIDENCE_CLASS = "fixture"` —
a structural constant (not just docstring prose) so DGR-036's fixture-vs-
real-model parity check can assert programmatically that it is comparing a
fixture engine against a real one, never two fixtures.
- Every fault-injection knob defaults to off (`0`/`None`/`False`), so a bare
`FakeShardEngine()` passes `assert_shard_engine_contract` unmodified —
fault injection is opt-in, never a baseline behavior change.
### `tests/test_fake_shard_engine.py` (new)
- `test_fake_shard_engine_obeys_the_shared_shard_engine_contract` — runs the
full DGR-031 contract against a bare `FakeShardEngine`.
- `test_fake_shard_engine_declares_fixture_evidence_class` — pins the
`EVIDENCE_CLASS` marker DGR-036 will rely on.
- Head/middle/tail output-shape tests (`boundary_point`, token-id-sideband
pass-through, tail vocab range).
- `test_session_state_is_isolated_between_two_concurrent_sessions` — a
stale-epoch rejection and a cancel on one session leave a second,
concurrently open session fully serviceable.
- One test per fault-injection knob (delay hook invocation, memory-budget
trip, malformed tail/boundary-bundle output, crash-after-N-calls,
configurable crash exception type) plus `FakeShardEngineConfig`'s own
`__post_init__` validation (negative delay, negative budget, non-positive
`crash_after_calls`).
- `test_load_result_and_capabilities_report_recipe_architecture` — the
fixture threads `LoadRequest.recipe["architecture"]` through to both
`LoadResult.architecture` and `EngineCapabilities.architecture` rather than
hardcoding `"dense"`/`"fake"` everywhere, so a future V4 recipe is visible
in fixture output too.
### `.scratch/distributed-gguf-runtime/prd.json` / issue markdown
Marked `DGR-032.passes = true` with `completionNotes`; regenerated
`issues/032-implement-deterministic-fake-shardengine.md` via
`scripts/ralph_prd_schema.py render` so it matches `prd.json` byte-for-byte.
## Acceptance criteria → evidence
1. **Head, middle, tail, prefill, decode, cancellation, release with
deterministic outputs** — `FakeShardEngine`'s `_transform`, boundary-point
tagging, and `assert_shard_engine_contract`'s own determinism/cancel/
release checks. Verified by
`test_fake_shard_engine_obeys_the_shared_shard_engine_contract`,
`test_head_shard_returns_boundary_bundle_with_post_head_residual_point`,
`test_middle_shard_returns_boundary_bundle_and_passes_through_token_sideband`,
`test_tail_shard_returns_token_output_within_advertised_vocab`.
2. **Isolated session/epoch state and deterministic cache-miss/stale-epoch
failures** — `_sessions` dict keyed per session;
`test_session_state_is_isolated_between_two_concurrent_sessions` plus the
shared contract's own cache-miss/stale-epoch checks.
3. **Configurable delay, memory pressure, malformed output, crash
injection** — `FakeShardEngineConfig`; verified by
`test_step_delay_seconds_invokes_the_configured_sleep_hook`,
`test_memory_budget_bytes_trips_deterministic_resource_exhausted`,
`test_malformed_output_is_structurally_valid_but_semantically_wrong_for_tail`,
`test_malformed_output_is_structurally_valid_but_semantically_wrong_for_boundary_bundle`,
`test_crash_after_calls_raises_instead_of_returning_a_structured_status`,
`test_crash_exception_factory_is_configurable`,
`test_config_rejects_invalid_knob_values`.
4. **Contract tests distinguish fixture evidence from real-model
certification** — module docstring and this README are explicit that
this is FIXTURE evidence only (numeric parity is DGR-036 onward); the
`EVIDENCE_CLASS = "fixture"` constant makes that distinction structurally
checkable, not just prose, pinned by
`test_fake_shard_engine_declares_fixture_evidence_class`.
5. **Gates + this handoff** — below.
## Commands and results
```bash
PYTHONPATH=packages/node:packages/tracker .venv/bin/python3 -m pytest -q tests/test_fake_shard_engine.py tests/test_shard_engine.py
```
```text
26 passed in 0.17s
```
```bash
PYTHONPATH=packages/node:packages/tracker .venv/bin/python3 -m pytest -q \
tests/test_fake_shard_engine.py tests/test_shard_engine.py tests/test_shard_lifecycle.py \
tests/test_architecture_boundary.py tests/test_activation_envelope.py \
tests/test_native_shard_protocol.py tests/test_shard_runtime_harness.py
```
```text
109 passed, 3 skipped in 3.78s
```
```bash
.venv/bin/python3 -m compileall -q packages tests
```
```text
(no output — clean; exit 0)
```
```bash
git diff --check
```
```text
(no output — clean)
```
## Limitations
- `tests/` as a whole does not collect cleanly in this environment: the same
pre-existing collection errors DGR-031's evidence recorded (missing
optional dependencies such as `cryptography`) are still present and are
unrelated to this story. This story's own gates were run as the targeted,
scoped test set above per the shared quality gates' wording ("Targeted
deterministic tests pass").
- This is FIXTURE evidence only. `FakeShardEngine` proves lifecycle,
session/epoch isolation, and fault-injection semantics; it proves nothing
about numerical parity with a real model. That is DGR-036's explicit job
once DGR-037's real engine exists, and DGR-053/054 for V4 alpha
certification.
- `FakeShardEngine` is not wired into `shard_runtime_server.py` or any gRPC
surface — it is a standalone, importable engine only. Wiring a
`ShardEngine` (fake or real) into the gRPC servicer is DGR-037's scope for
the real engine; DGR-033 covers a C++ worker surface, which is a separate
native executable, not a consumer of this Python module.
- No gRPC/CMake/native-build changes were needed or made — this story is
pure Python fixture code (`evidenceClass: fixture`, `hardware: none`), so
the native CMake/CTest and patch-stack gates in the shared quality-gate
list do not apply here, consistent with DGR-031's own README recording the
same non-applicability.
## Dependency handoff
- **DGR-033** (standalone fake C++ gRPC Shard worker): its own issue
describes a native C++ executable serving the lifecycle/stream RPC
contract "using the fake engine" — that is a native analogue, not a
consumer of this Python module; DGR-033 should still read this README for
the exact deterministic-output/session-isolation/fault-injection semantics
its C++ fake engine needs to reproduce so both fakes behave identically
from a client's point of view.
- **DGR-034/DGR-035** (range-aware GGUF ownership, boundary I/O):
`FakeShardEngine` already demonstrates range-driven head/middle/tail
behavior purely from `LoadRequest.shard_start`/`shard_end`/`total_layers`;
no new range vocabulary was needed.
- **DGR-036** (fixture vs real-model parity): compare a `FakeShardEngine`
instance's `EVIDENCE_CLASS` (`"fixture"`) against DGR-037's real engine's
equivalent marker (expected `"real"`) to assert the parity check is
actually comparing two different implementations; reuse
`assert_shard_engine_contract` against both to prove lifecycle parity
before attempting numeric parity.
- **DGR-037** (bind llama.cpp to the worker): `FakeShardEngine` is the
reference implementation to diff a real engine's lifecycle behavior
against — same request/result types, same session/epoch model, no new
contract vocabulary.
- **DGR-040** (worker supervision): the crash-injection knob
(`crash_after_calls`/`crash_exception_factory`) exists specifically so
supervision/restart logic has a deterministic way to trigger and test an
unhandled engine failure distinct from a graceful `StructuredStatus`
rejection.

View File

@@ -0,0 +1,281 @@
# DGR-033 evidence — standalone fake C++ gRPC Shard worker
**Completed:** 2026-07-25 (initial); **repaired:** 2026-07-26 after Codex
GPT-5.5 cross-review BLOCK (see "Cross-review repair" below).
**Branch:** `ralph/distributed-gguf-opus`
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
**Dependencies:** DGR-022 (lifecycle/status contract), DGR-024 (real generated
gRPC harness + `shard_runtime_server.py` reference semantics), DGR-032
(deterministic fake `ShardEngine` semantics).
## Objective
Prove the standalone worker process, stream, lifecycle, and supervision shape
before any llama.cpp integration: a real C++ executable that serves the whole
ShardRuntime lifecycle/stream contract over gRPC using a model-free fake engine,
driven end-to-end by Python integration tests over a real socket.
## What was found live before changing code
- `packages/node/native/proto/shard_runtime.proto` (DGR-021..023): the single
semantic contract. Its `ShardRuntime` service has exactly five RPCs —
`GetCapability`, `Health`, `Session` (bidi stream), `Release`, `Cancel`.
- `packages/node/meshnet_node/shard_runtime_server.py` (DGR-024): the reference
Python servicer. It performs a *bounded real forward* (a CRC over the received
bundle bytes) then echoes the chunk, and fails closed on stale epoch, expired
deadline, corrupt/mis-tiled fragments, exhausted flow-control credit, duplicate
idempotency step, and in-band/out-of-band cancellation, with per-`route_session_id`
state kept on the servicer so an out-of-band `Cancel` can reach a live session.
**Key finding:** despite the schema labelling the checksum `CRC32C`, this
runtime computes it with `zlib.crc32` (standard CRC-32, *not* Castagnoli). The
C++ worker mirrors `zlib.crc32` exactly so its checksum acceptance is
byte-identical to the existing Python surface (the committed C++ *conformance*
test, by contrast, uses true Castagnoli against separately-generated goldens —
the two are unrelated code paths).
- `packages/node/native/CMakeLists.txt` (DGR-029/030): configures against the
ignored `build/native-toolchain` prefix (pinned Protobuf 33.1 + gRPC 1.82.1),
always generates both message and service stubs, and registers a C++
conformance CTest. There was **no** worker executable and **no** Python
worker integration test before this story (confirmed by
`ls packages/node/native/worker` → absent, and grep for `shard_worker`).
- `packages/node/meshnet_node/fake_shard_engine.py` (DGR-032): the Python fake
engine, deliberately *not* wired into the gRPC surface. DGR-033's worker is
its native analogue — a separate executable, not a consumer of that module —
so both fakes present identical behaviour to a client (deterministic,
model-free bounded forward; per-session isolation; fail-closed lifecycle).
## What was added (this story's change)
### `packages/node/native/worker/fake_engine.h` (new)
`meshnet::worker::FakeShardEngine` — a header-only, model-free fixture engine.
Its only capability is to validate a `TensorBundle` (fragments tile exactly, the
uncompressed CRC-32 matches the declared checksum, the declared payload stays
within the negotiated `max_chunk_bytes`) and fold the fragment bytes through a
bounded forward. It links, loads, and dispatches to **nothing** — no llama.cpp,
no graph execution. Carries `kEvidenceClass = "fixture"` mirroring the Python
`FakeShardEngine.EVIDENCE_CLASS` for the later DGR-036 parity check.
### `packages/node/native/worker/shard_service.{h,cpp}` (new)
`ShardRuntimeServiceImpl : meshnet::shard::v1::ShardRuntime::Service` — a faithful
C++ port of the DGR-024 Python servicer: the same per-`route_session_id`
identity/credit/dedup state guarded by a mutex, the same fail-closed negative
paths, and the same lifecycle (open → prefill/decode → flow-control top-up →
release/cancel). Each per-request response is computed under the lock and written
*after* releasing it, so a blocking `Write` can never deadlock the out-of-band
`Cancel` RPC that needs the same lock. Bounded messages are enforced two ways: a
per-tensor `RESOURCE_EXHAUSTED` app check against `max_chunk_bytes`, plus a hard
transport receive ceiling.
### `packages/node/native/worker/shard_worker_main.cpp` (new)
The standalone `shard_worker` executable. Binds `MESHNET_SHARD_LISTEN_ADDR`
(or an `argv` address), prints one readiness line (`ShardRuntime worker listening
on <addr>`), and serves until `SIGTERM`/`SIGINT`. **Graceful shutdown** uses a
self-pipe: the async-signal-safe handler writes one byte, a drain thread reads it
and calls `server->Shutdown()`, so in-flight sessions finish and the process
exits `0` printing `ShardRuntime worker shut down cleanly`. A `--selftest` mode
binds an ephemeral port and self-drives capability/health/fragmented-prefill/
decode/release over a real loopback gRPC channel, giving a pure-C++ CTest that
needs no Python.
### `packages/node/native/CMakeLists.txt` (modified)
Adds the `shard_worker` executable (linking only `shard_runtime_grpc` +
`gRPC::grpc++` — no llama.cpp) and registers `shard_worker_selftest` as a CTest.
### `tests/test_native_shard_worker.py` (new)
18 integration tests that spawn the **real compiled binary** as a subprocess and
drive it with the committed generated stubs over a real localhost socket. When
the binary is not built they skip (the DGR-029/030 `requires_cmake` gating
pattern), locating it via `MESHNET_SHARD_WORKER_BIN` or `build/native/shard_worker`.
## Acceptance criteria → evidence
1. **Standalone C++ executable serves the complete lifecycle/stream contract
using the fake engine** — `shard_worker` builds and serves all five RPCs; the
`shard_worker_selftest` CTest drives open → fragmented prefill → decode →
release over real gRPC; the 18 Python tests cover the same against the
subprocess.
2. **Python integration tests cover startup, health, capability, fragmented
prefill, decode, release, cancellation, graceful shutdown** —
`test_worker_startup_and_health`, `test_worker_capability`,
`test_fragmented_prefill_echoes_reassembled_payload` (3-fragment tiling),
`test_decode_step_is_served`, `test_release_is_terminal`,
`test_in_band_cancel_of_single_work_item_does_not_end_stream`,
`test_in_band_cancel_of_whole_session_is_terminal`,
`test_out_of_band_cancel_rpc_races_ahead_of_open`,
`test_graceful_shutdown_on_sigterm` (SIGTERM → exit 0 + clean-shutdown line).
3. **Bounded messages, deadlines, flow control, independent session
cancellation enforced** — `test_bounded_message_is_rejected`
(`RESOURCE_EXHAUSTED` on an over-ceiling tensor),
`test_expired_deadline_is_rejected`, `test_flow_control_violation_and_topup`,
`test_independent_session_cancellation` (cancelling session A leaves session B
fully serviceable), plus `test_stale_route_epoch_is_rejected`,
`test_duplicate_idempotency_step_is_acked`,
`test_malformed_fragment_tiling_is_rejected`.
4. **Exposes neither llama.cpp RPC nor arbitrary graph execution**
`ldd build/native/shard_worker` shows no llama/ggml shared libs;
`nm -C build/native/shard_worker | grep -icE 'llama_|ggml_'``0`; the proto
exposes exactly one service with five lifecycle RPCs and no graph-exec entry.
5. **Gates + this handoff** — below.
## Commands and results
Toolchain (ignored `build/native-toolchain`, pinned Protobuf 33.1 + gRPC 1.82.1):
```bash
bash scripts/bootstrap_native_toolchain.sh "$PWD/build/native-toolchain"
# ... gRPC 1.82.1 commit acccf84c0df20487d64101f528e5d426541ca4e5
# grpc_cpp_plugin sha256 43705cf26ae9ce98bbcee76b3408f5e171eec746b50bf0dd42dd68d132c6a533
```
Focused out-of-tree CMake build + CTest:
```bash
cmake -S packages/node/native -B build/native -DCMAKE_PREFIX_PATH="$PWD/build/native-toolchain"
cmake --build build/native -j"$(nproc)"
ctest --test-dir build/native --output-on-failure
```
```text
1/2 Test #1: shard_worker_selftest ............ Passed 0.01 sec
2/2 Test #2: shard_protocol_conformance ....... Passed 0.00 sec
100% tests passed out of 2
```
Python integration tests against the real binary:
```bash
PYTHONPATH=packages/node:packages/tracker python -m pytest -q tests/test_native_shard_worker.py
```
```text
18 passed in 3.96s
```
AC4 (no llama.cpp / no graph exec):
```bash
ldd build/native/shard_worker | grep -iE 'llama|ggml' # -> (no matches)
nm build/native/shard_worker | grep -icE 'llama_|ggml_' # -> 0
```
Shared gates + regression:
```bash
python -m compileall -q packages tests # exit 0
git diff --check -- packages/node/native tests/test_native_shard_worker.py # exit 0
PYTHONPATH=packages/node:packages/tracker python -m pytest -q \
tests/test_shard_runtime_harness.py tests/test_native_shard_protocol.py
# -> 61 passed, 2 skipped (DGR-024 harness + native protocol untouched)
```
Toolchain used: `cmake`/`ctest` from the `distributed-gguf-runtime` worktree's
`.venv` (PyPI `cmake==4.4.0` wheel — no system cmake exists here, same as
DGR-029/030); the Python client uses that venv's `grpcio==1.82.1`,
`grpcio-tools==1.82.1`, `protobuf`, `pytest`. `g++ (GCC) 15.2.1`.
## Limitations
- This is FIXTURE evidence only. The worker's "forward" is a CRC-over-wire-bytes
echo, not real tensor compute; it proves process/stream/lifecycle/supervision
shape, nothing about numerical correctness. Real engine binding is DGR-037 and
numeric parity is DGR-036/052.
- The worker checksum path mirrors the DGR-024 runtime's `zlib.crc32` (standard
CRC-32 under a `CRC32C` label). Compressed-tensor tiling/checksum is not
independently verified (no zstd decompressor in the fixture) — identical to the
DGR-024 limitation.
- Default `pytest` runs skip `tests/test_native_shard_worker.py` unless the
worker binary is built (or `MESHNET_SHARD_WORKER_BIN` is set); this session
built it and ran all 18 for real (results above). Building requires the pinned
gRPC C++ toolchain, which is not present by default and must be bootstrapped.
- No CUDA/ROCm/GPU, no model download, no network at test time — all default
tests are fixture-only and offline.
## Dependency handoff
- **DGR-036** (fixture vs real-model parity): the worker's `FakeShardEngine`
carries `kEvidenceClass = "fixture"`; diff it against DGR-037's real engine's
equivalent marker, and reuse the same lifecycle/stream contract this worker
serves to prove behavioural parity before numeric parity.
- **DGR-037** (bind llama.cpp): replace `FakeShardEngine`'s bounded forward with
the real engine behind the *same* `ShardRuntimeServiceImpl` surface; the
service's session/epoch/credit/dedup/cancel machinery and the graceful-shutdown
supervision shape are reusable as-is.
- **DGR-040** (worker supervision): `shard_worker` already provides the
supervision primitives — a readiness line for start detection, `SIGTERM`
graceful drain with a clean-exit line, and a `--selftest` liveness probe.
A supervisor can start/monitor/restart the process around these.
## Cross-review repair (2026-07-26)
An independent Codex GPT-5.5 review BLOCKED the initial implementation. Four
root protocol defects in the native worker were fixed in this worktree
(`.claude/worktrees/distributed-gguf-opus`); the fake-engine echo semantics and
supervision shape are unchanged.
### Defects fixed
1. **Activation before SessionOpen bypassed all state.** A chunk/decode whose
`route_session_id` had no opened session fell through every `if (state && ...)`
guard and was echoed — bypassing lifecycle, cancellation, epoch and
flow-control. `SessionState` now carries an `opened` flag set only by a valid
`SessionOpen`; chunk and decode fail closed with a terminal
`ERROR_CODE_INTERNAL` and end the stream when it is false. A placeholder state
created by an out-of-band `Cancel` that races `Open` has `opened == false`, so
it can never admit work either.
2. **Flow control blindly trusted the peer proposal.** `SessionOpen` copied the
proposed `credits/max_inflight/max_chunk_bytes` verbatim into session state and
the accepted reply. New `ShardRuntimeServiceImpl::NegotiateFlow` takes the
strictest bound of peer-vs-worker for every field (mirroring
`negotiate_flow_control` in `native_protocol/codec.py`), stores the negotiated
ceilings on the session, and enforces the negotiated per-session
`max_chunk_bytes` on every bundle (`FakeShardEngine::Validate` now takes the
ceiling as an argument instead of a fixed construction-time value).
3. **In-stream `ReleaseSignal` leaked session state.** The stream `release` arm
wrote a terminal status but never dropped the session. It now erases the
session under the lock before responding, so KV/credits/dedup are freed
immediately (the out-of-band `Release` RPC already erased).
4. **`SessionOpen` echoed caller identity instead of validating it.** The handshake
now rejects an incompatible `schema_version` (`SCHEMA_UNSUPPORTED`), a
mismatched model/recipe `Fingerprint` (`FINGERPRINT_MISMATCH`), and a
`ShardRange` outside the worker's served range (`SHARD_RANGE_MISMATCH`), each
terminal; `SessionAccepted` now reports the worker's own served fingerprint
rather than a copy of the caller's.
### Changed files (repair)
- `packages/node/native/worker/shard_service.h``opened` +
`max_prefill_chunk_tokens` on `SessionState`; `NegotiateFlow` decl; engine now
default-constructed.
- `packages/node/native/worker/shard_service.cpp` — worker-identity constants +
fill helpers; `NegotiateFlow`; `SessionOpen` validation/negotiation; fail-closed
chunk/decode; per-session `max_chunk_bytes`; in-stream release erase.
- `packages/node/native/worker/fake_engine.h``Validate(bundle, max_chunk_bytes)`.
- `tests/test_native_shard_worker.py` — extended `_open` (schema/fingerprint/range/
flow overrides); fixed `test_release_rpc_is_idempotent` for the new erase
semantics; added 9 regression tests (chunk/decode before open, flow-control
clamp, negotiated-ceiling cap, in-stream release erase, schema/fingerprint/range
rejection, worker-fingerprint-not-caller).
### Re-run gates (real, rebuilt binary)
Build driven through the pinned `cmake` (Unix Makefiles + `gmake`, gRPC 1.82.1):
```text
cmake --build build/native --parallel 8 -> BUILD_EXIT 0
ctest --test-dir build/native --output-on-failure -> 100% (2/2) passed
shard_worker_selftest ....... Passed
shard_protocol_conformance .. Passed
python -m pytest -q tests/test_native_shard_worker.py -> 27 passed
python -m pytest -q tests/test_shard_runtime_harness.py \
tests/test_native_shard_protocol.py -> 63 passed
python -m compileall -q packages tests -> exit 0
git diff --check -> clean
ldd build/native/shard_worker | grep -iE 'llama|ggml' -> NONE
nm -C build/native/shard_worker | grep -cE 'llama_|ggml_' -> 0
```
The worker integration suite grew from 18 to 27 tests; all pass against the
freshly compiled binary. No `.ralph-lane` runtime artifacts were touched.

View File

@@ -0,0 +1,94 @@
# DGR-034 evidence — dense-Llama range-aware GGUF ownership
**Status:** implemented and live-verified on 2026-08-01. `prd.json` remains
the authority for story state.
## What changed
- The pinned llama.cpp patch stack adds `meshnet_owned_layer_start/end` and
filters dense-Llama GGUF registration to `blk.N.*` for the requested
half-open range. `token_embd.weight` belongs to the head; `output_norm` and
`output.weight` (or the tied embedding) belong to the tail.
- The load state exposes a C range report derived from the registered model
buffers, and a project-owned `meshnet-range-report` tool audits the live
registered tensor map. It rejects empty, inverted, out-of-model, missing,
outside-range, unexpected, and endpoint-inconsistent loads.
- `meshnet_node.range_report` accepts only audited tool output. It makes the
range and endpoint flags authoritative from loaded state rather than caller
assertions, and fails closed on malformed ownership or byte counts.
## Real-model memory evidence
Artifact: `Magistral-Small-2509-Q4_K_M.gguf`, 14,333,911,104 bytes, SHA-256
`a17a113480e7f55780ad1d100493c70ac158d1943e578bbdd75acef0872ab7dc`.
It stayed on the configured mounted drive; no artifact was downloaded or put
under `/home`.
The direct non-mmap lane proves resident storage tracks owned tensors:
| Range | Registered tensors | Resident bytes | Process peak RSS |
| --- | ---: | ---: | ---: |
| `[10, 20)` | 90 | 3,304,898,560 | 3,298,800 KiB |
| `[0, 40)` | 363 | 14,326,026,240 | 14,061,632 KiB |
Raw reports and timings are in `runs/default-mid-a.*` and
`runs/default-full-nommap.*`. The middle range is 23.1% of the full
resident allocation and owns 24.8% of the registered tensors.
## Commands and results
```text
python3 scripts/llama_cpp_dependency.py reverse --source-dir build/llama.cpp/source
python3 scripts/llama_cpp_dependency.py verify --workspace build/llama.cpp
python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
# apply/check/reverse succeeded against e920c523e3b8a0163fe498af5bf90df35ff51d25;
# the source was then applied for the focused native checks.
(cd packages/node/native/llama/patches && sha256sum -c SHA256SUMS)
# all six patches: OK
/home/popov/.hermes/hermes-agent/venv/bin/ctest \
--test-dir build/llama.cpp/dgr034-check \
-R '^test-meshnet-range-ownership$' --output-on-failure
# 1/1 passed
PYTHONPATH=packages/node MESHNET_RANGE_REPORT_BIN="$PWD/build/llama.cpp/dgr034-check/bin/meshnet-range-report" \
/home/popov/.hermes/hermes-agent/venv/bin/pytest -q \
tests/test_range_report.py tests/test_meshnet_range_report_tool.py \
tests/test_llama_cpp_dependency.py
# 56 passed in 0.87s
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python \
-m compileall -q packages tests
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
git diff --check && git diff --cached --check
# all exit 0; PRD validation: 55 stories validated
```
The model commands used the same `meshnet-range-report` binary with
`--no-mmap --no-extra-bufts`, first for `[10,20)` and then `[0,40)`; both
returned `ok: true` and their exact output is retained above.
## Changed files
- `packages/node/native/llama/PATCH-STACK.md`
- `packages/node/native/llama/UPSTREAM_LOCK.json`
- `packages/node/native/llama/patches/{series,SHA256SUMS,UPSTREAM-ASSUMPTIONS.json,0006-meshnet-range-report-tool.patch}`
- `packages/node/meshnet_node/range_report.py`
- `tests/test_range_report.py`
- `tests/test_meshnet_range_report_tool.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-034/*`
## Limitations and dependency handoff
- The mmap loader can retain broad contiguous file spans when GGUF tensor
order places a tail endpoint near the beginning of the artifact; the direct
non-mmap lane is the certified resident-memory result. The raw mmap report
is retained in `runs/default-head.json` and must not be presented as a
physical-RSS saving.
- This story proves loading/ownership only. Partial-range graph execution
remains fail-closed until DGR-035 provides typed dense boundary adapters.
- DGR-037 can bind the worker to `llama_model_meshnet_range_report` or the
strict Python consumer; it must use the reported range, not requested range,
for capability publication. DGR-051 must add its V4-specific ownership
rules separately.

View File

@@ -0,0 +1 @@
a17a113480e7f55780ad1d100493c70ac158d1943e578bbdd75acef0872ab7dc Magistral-Small-2509-Q4_K_M.gguf

View File

@@ -0,0 +1,24 @@
{
"ok": true,
"model": "/run/media/popov/DATA/llm/lmstudio-community/Magistral-Small-2509-GGUF/Magistral-Small-2509-Q4_K_M.gguf",
"architecture": "llama",
"n_layer": 40,
"file_bytes": 14333911104,
"requested_range": [0, 40],
"reported_range": [0, 40],
"mmap": false,
"touched": false,
"use_extra_bufts": false,
"has_token_embeddings": true,
"has_output_head": true,
"tied_output_head": false,
"mapped_bytes": 0,
"resident_bytes": 14326026240,
"registered_tensors": 363,
"registered_bytes": 14326026240,
"unexpected_registered_tensors": [],
"missing_owned_layers": [],
"vm_size_bytes": 14392061952,
"vm_rss_bytes": 14387003392,
"vm_hwm_bytes": 14399111168
}

View File

@@ -0,0 +1 @@
elapsed=0:02.48 maxrss_kib=14061632 exit=0

View File

@@ -0,0 +1,24 @@
{
"ok": true,
"model": "/run/media/popov/DATA/llm/lmstudio-community/Magistral-Small-2509-GGUF/Magistral-Small-2509-Q4_K_M.gguf",
"architecture": "llama",
"n_layer": 40,
"file_bytes": 14333911104,
"requested_range": [0, 10],
"reported_range": [0, 10],
"mmap": true,
"touched": false,
"use_extra_bufts": true,
"has_token_embeddings": true,
"has_output_head": false,
"tied_output_head": false,
"mapped_bytes": 6219366400,
"resident_bytes": 6219366400,
"registered_tensors": 91,
"registered_bytes": 3771596800,
"unexpected_registered_tensors": [],
"missing_owned_layers": [],
"vm_size_bytes": 16942260224,
"vm_rss_bytes": 16937005056,
"vm_hwm_bytes": 16947953664
}

View File

@@ -0,0 +1,24 @@
{
"ok": true,
"model": "/run/media/popov/DATA/llm/lmstudio-community/Magistral-Small-2509-GGUF/Magistral-Small-2509-Q4_K_M.gguf",
"architecture": "llama",
"n_layer": 40,
"file_bytes": 14333911104,
"requested_range": [10, 20],
"reported_range": [10, 20],
"mmap": false,
"touched": false,
"use_extra_bufts": false,
"has_token_embeddings": false,
"has_output_head": false,
"tied_output_head": false,
"mapped_bytes": 0,
"resident_bytes": 3304898560,
"registered_tensors": 90,
"registered_bytes": 3304898560,
"unexpected_registered_tensors": [],
"missing_owned_layers": [],
"vm_size_bytes": 3370934272,
"vm_rss_bytes": 3365814272,
"vm_hwm_bytes": 3377971200
}

View File

@@ -0,0 +1 @@
elapsed=0:00.82 maxrss_kib=3298800 exit=0

View File

@@ -0,0 +1,54 @@
# DGR-035 evidence — dense architecture boundary input/output
**Implemented:** 2026-08-01
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
## What changed
- `DenseRangeBoundaryExecutor` is a strict execution-facing adapter for the certified `dense-llama` architecture. A head range accepts non-empty token IDs and owns the embedding callback. Middle/tail ranges reject token IDs and require the named `dense.residual.v1` `BoundaryBundle`.
- Non-tail execution returns exactly the raw `hidden_states` residual from its local layer callback. Its constructor rejects a final-norm/output callback, preventing final normalization, logits projection, sampling, and tail-only row pruning before the tail.
- Tail execution is the only path allowed to own final output and returns an explicit `TailOutput`: either validated logits or a sampled token. The existing wire `TypedTailResult` now serializes and validates both choices.
- Unknown architectures, wrong boundary points, and tensor bundles other than one named `hidden_states` tensor fail closed.
## Changed files
- `packages/node/meshnet_node/architecture_boundary.py`
- `tests/test_dense_range_boundary.py`
- `tests/test_architecture_boundary.py`
- `.ralph-tui/progress.md`
- `.scratch/distributed-gguf-runtime/evidence/DGR-035/README.md`
## Commands and results
```bash
TESTPY=/home/popov/.hermes/hermes-agent/venv/bin/python
PYTHONPATH=packages/node:packages/tracker "$TESTPY" -m pytest -q tests/test_dense_range_boundary.py tests/test_architecture_boundary.py tests/test_shard_engine.py tests/test_fake_shard_engine.py
```
```text
37 passed in 0.22s
```
```bash
"$TESTPY" -m ruff check packages/node/meshnet_node/architecture_boundary.py tests/test_dense_range_boundary.py tests/test_architecture_boundary.py
PYTHONPATH=packages/node "$TESTPY" -m compileall -q packages tests
git diff --check
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
```
```text
All checks passed!
OK: 55 stories validated.
```
## Limitations
- This story adds and proves the project-owned boundary contract with deterministic, model-download-free tests. It does not claim real-model range parity; DGR-036 owns that numerical certification.
- The llama.cpp graph remains fail-closed for partial owned ranges until DGR-037 binds its worker to this execution contract. No native source or patch-stack file was changed here, so native CMake/CTest and patch-cycle gates are not applicable to this Python contract change.
- `.venv/bin/python3` has no `pytest` module in this worktree. The available project validation interpreter above ran the exact targeted tests.
## Dependency handoff
- DGR-036 should use `DenseRangeBoundaryExecutor` with its real-engine bridge to compare whole-model and split residual/logits outputs, including prefill and decode.
- DGR-037 must adapt the pinned llama.cpp dense graph to `embed_tokens`, `run_layers`, and tail-only `tail_output`; it must preserve `dense.residual.v1` unnormalized and avoid row pruning until the tail.
- DGR-069 can propose only a generic residual-in/residual-out llama.cpp hook; architecture names and Meshnet wire/session semantics remain outside upstream.

View File

@@ -0,0 +1,12 @@
# DGR-036 real-model lane blocker
`DGR-036` cannot receive completion credit yet. The live standalone worker is DGR-033's `FakeShardEngine`, a CRC/echo fixture; DGR-037's real llama.cpp `ShardEngine` binding has not been implemented. Consequently no code path can execute a whole or ranged dense GGUF and no real prefill/logit or greedy-token parity result exists.
The deterministic two-process fake-worker regression is implemented in `tests/test_native_shard_worker.py`, but this sandbox cannot open loopback sockets (`PermissionError: [Errno 1] Operation not permitted`), so that runtime test needs host-side execution as well.
Unblock in this order:
1. Complete DGR-037's real ranged llama.cpp worker binding without changing the DGR-035 dense boundary contract.
2. Provision a small exact dense GGUF on mounted-drive storage and record its artifact/split hashes plus runtime/backend/hardware/network identity.
3. Run whole-model and two-range prefill comparison, then record at least 32 greedy token IDs against the locked tolerance and retain raw metrics.
4. Run the deterministic two-process test on a host with loopback sockets, then update `README.md` and only then set `prd.json` completion truth.

View File

@@ -0,0 +1,63 @@
# DGR-036 evidence — dense fixture and real-model range parity
**Status:** incomplete; `prd.json` remains authoritative and keeps `DGR-036.passes` as `false`.
## Deterministic fixture proof implemented
`tests/test_native_shard_worker.py` now contains `test_two_disjoint_fake_worker_processes_preserve_prefill_and_decode_seam`. It starts two separate DGR-033 `shard_worker` OS processes, opens disjoint requested ranges `[0, 16)` and `[16, 32)`, forwards the first worker's actual protobuf output to the second, and checks one prefill plus 32 sequential decode positions. The test tops up the worker's 16-credit flow-control window before decode positions 16 and 32, so all 32 positions are exercised.
This is deliberately **fixture evidence only**. The worker's `FakeShardEngine` validates a bundle and echoes its bytes; it has no dense graph, logits, sampler, or GGUF load. The assertions prove the two-process protocol/lifecycle seam and that bytes survive a disjoint-range handoff. They do not claim numerical model or greedy-token parity.
## Real-model lane: blocked honestly
DGR-037, which is still `passes: false`, is the story that binds llama.cpp to the standalone worker. The live DGR-033 worker remains the fake CRC/echo fixture, and no `ShardEngine` implementation can load/run a GGUF range. DGR-034 proves tensor ownership and memory reporting, while DGR-035 proves the Python boundary contract; neither supplies a real ranged execution engine. Therefore there is no truthful way to run a small dense GGUF whole-model versus two-range prefill comparison or to compare 32 greedy generated tokens yet.
The real-model proof must be run after DGR-037 with an exact small dense GGUF, the pinned llama.cpp/runtime identity, two loaded worker ranges, and a raw report containing artifact and split hashes, backend/driver/hardware/network, prefill tolerance, all 32 token IDs, and raw metrics. It must remain opt-in, use mounted-drive artifact storage, and never download an artifact under `/home`.
## Commands and results
```bash
TESTPY=/home/popov/.hermes/hermes-agent/venv/bin/python
PYTHONPATH=packages/node:packages/tracker "$TESTPY" -m pytest -q tests/test_dense_range_boundary.py tests/test_architecture_boundary.py tests/test_shard_engine.py tests/test_fake_shard_engine.py
```
```text
37 passed in 0.18s
```
```bash
"$TESTPY" -m ruff check tests/test_native_shard_worker.py
PYTHONPATH=packages/node "$TESTPY" -m compileall -q packages tests
git diff --check
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
```
```text
All checks passed!
OK: 55 stories validated.
```
Attempted two-process fixture command:
```bash
PYTHONPATH=packages/node:packages/tracker "$TESTPY" -m pytest -q tests/test_native_shard_worker.py -k two_disjoint_fake_worker_processes_preserve_prefill_and_decode_seam
```
```text
FAILED: PermissionError: [Errno 1] Operation not permitted at socket.socket(AF_INET, SOCK_STREAM)
```
This is the workspace sandbox's known localhost-socket restriction, before any worker is spawned; it is not a test assertion failure. Run that exact command on a host that permits loopback sockets after building `build/native/shard_worker`.
## Changed files
- `tests/test_native_shard_worker.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-036/README.md`
- `.scratch/distributed-gguf-runtime/evidence/DGR-036/BLOCKED.md`
- `.ralph-tui/progress.md`
## Dependency handoff
- DGR-033 supplies the process, lifecycle, generated gRPC surface, fake engine, and bounded-flow-control behaviour used by the deterministic test.
- DGR-035 supplies the strict dense residual boundary and tail-only output contract. DGR-037 must preserve that contract when it replaces the echo fake with a real engine.
- Once DGR-037 is complete, return here to run the opt-in numerical lane. Do not turn this fixture test into a claim that a real GGUF can execute ranges.

View File

@@ -0,0 +1,77 @@
# DGR-037 evidence — bind llama.cpp to the standalone worker
**Date:** 2026-08-01
**Authority:** `.scratch/distributed-gguf-runtime/prd.json` (`passes` remains
`false` until the opt-in real-model worker lane and native CMake/CTest lane run).
## Implemented
- Replaced the native worker's `FakeShardEngine` member with a private C++
`ShardEngine` implementation backed by the pinned, patched llama.cpp API.
`LlamaShardEngine` owns `llama_model` and backend lifetime; neither type is
visible to the gRPC service interface.
- Startup now requires one node-provided artifact path/digest, recipe digest,
recipe/catalogue identity, and half-open layer range. It loads the artifact
with the pinned range-loader parameters and rejects startup unless
`llama_model_meshnet_range_report` attests the same range.
- `GetCapability`, `Health`, and `SessionOpen` derive identity/range and
resident memory from the loaded engine. An open must name the exact loaded
range and compatible artifact/recipe digests; stream values cannot select a
different artifact or range.
- Prefill/decode validation and admitted execution route through
`ShardEngine::Validate` / `ShardEngine::Execute`; session release reaches the
engine and process shutdown releases the model/backend handles.
- Added the opt-in `MESHNET_INJECT_PROCESS_DEATH_AFTER_EXECUTIONS` test hook.
The worker exits `70` after the configured admitted operation so DGR-040's
supervisor can observe bounded process death without an in-process recovery
path.
## Changed files
- `packages/node/native/CMakeLists.txt`
- `packages/node/native/README.md`
- `packages/node/native/worker/llama_shard_engine.{h,cpp}`
- `packages/node/native/worker/shard_service.{h,cpp}`
- `packages/node/native/worker/shard_worker_main.cpp`
- `tests/test_llama_shard_worker_binding.py`
## Commands and results
```text
python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
Applied the exact local DGR-027 patch stack; the resulting header exposed
meshnet_owned_layer_start/end and llama_model_meshnet_range_report.
c++ -std=c++17 -fsyntax-only [llama_shard_engine.cpp, shard_service.cpp, shard_worker_main.cpp]
All three translation units passed syntax checking. The gRPC toolchain emitted
only its existing deprecation warnings.
/home/popov/.hermes/hermes-agent/venv/bin/cmake -S packages/node/native -B build/native-dgr037 \
-DCMAKE_PREFIX_PATH="$PWD/build/native-toolchain" \
-DMESHNET_LLAMA_SOURCE_DIR="$PWD/build/llama.cpp/source" \
-DMESHNET_LLAMA_LIBRARY_DIR="$PWD/build/llama.cpp/build/bin"
/home/popov/.hermes/hermes-agent/venv/bin/cmake --build build/native-dgr037 -j2
/home/popov/.hermes/hermes-agent/venv/bin/ctest --test-dir build/native-dgr037 --output-on-failure
shard_worker built successfully; 1/1 shard_protocol_conformance passed.
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_llama_shard_worker_binding.py tests/test_native_shard_protocol.py
53 passed, 2 skipped
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python -m compileall -q packages tests
git diff --check
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
compileall passed; diff check passed; OK: 55 stories validated
```
## Limitations and dependency handoff
- No model artifact was selected for this session, so no opt-in real-model
process run, process-death observation, or raw hardware metrics are claimed.
- The pinned API currently attests range ownership/loading. Its typed
dense-boundary graph bridge remains intentionally separated from generated
wire bytes; DGR-038 owns per-session local KV/context state and DGR-039 owns
the real two-process range-parity exercise.
- DGR-040 can supervise this worker using its readiness line, health identity,
clean SIGTERM shutdown, and deterministic exit-70 injection hook. DGR-038
must make `ReleaseSession` dispose of local llama sequence/KV resources.

View File

@@ -0,0 +1,66 @@
# DGR-038 evidence — isolated shard-local Hot KV State
**Date:** 2026-08-01
**Authority:** `.scratch/distributed-gguf-runtime/prd.json` (`passes` remains
`false` until the opt-in real-model concurrency lane runs).
## Implemented
- The native `LlamaShardEngine` now creates one bounded llama.cpp context and
assigns a distinct `llama_seq_id` to each `(route_session_id, route_epoch)`.
It never accepts remote KV data; the loaded, range-attested llama model owns
the local cache layout and layers.
- Prefill/decode append state tracks local positions and expected past length.
A re-prefill at an earlier position truncates only that sequence with
`llama_memory_seq_rm`; a discontinuity or past-length mismatch returns a
retryable `CACHE_MISS`. Older route epochs return `EPOCH_STALE`.
- The token-reservation budget is bounded by per-session context, total Hot KV
budget, maximum sequence count, TTL, and LRU. Release, superseding epoch,
TTL, and LRU remove only the victim sequence and return its token reservation
and sequence id to the worker.
- The gRPC service converts native cache/stale/resource results to the typed
protocol errors and does not consume idempotency/flow-control credit on a
rejected append. Release is epoch-specific, so a stale release cannot erase
the active epoch's service state.
- Added opt-in configuration: `MESHNET_HOT_KV_MAX_SESSIONS`,
`MESHNET_HOT_KV_CONTEXT_TOKENS`, `MESHNET_HOT_KV_BUDGET_TOKENS`, and
`MESHNET_HOT_KV_TTL_SECONDS`.
## Changed files
- `packages/node/native/worker/llama_shard_engine.{h,cpp}`
- `packages/node/native/worker/shard_service.cpp`
- `packages/node/native/worker/shard_worker_main.cpp`
- `tests/test_llama_shard_worker_binding.py`
## Commands and results
```text
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_llama_shard_worker_binding.py tests/test_native_shard_protocol.py
54 passed, 2 skipped
/home/popov/.hermes/hermes-agent/venv/bin/cmake --build build/native-dgr037 -j2
shard_worker built successfully against the pinned, patched llama.cpp source.
/home/popov/.hermes/hermes-agent/venv/bin/ctest --test-dir build/native-dgr037 --output-on-failure
1/1 shard_protocol_conformance passed.
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python -m compileall -q packages tests
git diff --check
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
compileall and diff check passed; OK: 55 stories validated.
```
## Limitations and dependency handoff
- DGR-037 supplied the range-attested native model/engine boundary. DGR-038
adds local sequence ownership without changing its artifact or range
identity contract.
- No mounted GGUF artifact was selected. Therefore no opt-in real-model
four-session run, actual llama KV byte measurement, or hardware metrics are
claimed. The default tests intentionally remain model-download-free and the
source `prd.json` remains `passes: false`.
- DGR-039 should exercise the real two-process range-parity lane with four
sessions and the Hot-KV environment bounds, recording actual cache memory
and cancellation isolation evidence.

View File

@@ -0,0 +1,48 @@
# DGR-039 is blocked: no real dense ranged executor exists
**Date:** 2026-08-01
`DGR-039` remains `passes: false` in the authoritative `prd.json`.
## Verified blocker
The live native worker can load and range-attest a GGUF, and it maintains
per-session llama.cpp KV bookkeeping. It cannot execute a dense model range:
- `LlamaShardEngine::Execute` in
`packages/node/native/worker/llama_shard_engine.cpp` deliberately does not
convert the `TensorBundle` into a llama.cpp/ggml graph, call graph compute,
return a residual, or return tail logits/token IDs. Its only successful
effect is advancing `session.past_len` and the local token reservation.
- `ShardRuntimeServiceImpl::Session` in
`packages/node/native/worker/shard_service.cpp` returns the incoming prefill
bundle verbatim (`*response.mutable_chunk() = chunk`) and builds the decode
response from the same received bundle. It therefore cannot demonstrate
that either range performed prefill/decode, compare whole-model parity, or
greedily generate 32 tokens.
- `tests/test_architecture_boundary.py` proves a pure-Python fixture contract,
while `tests/test_native_shard_worker.py` proves an echo seam. Neither is a
real GGUF execution route. There is also no local coordinator/harness that
drives a whole-model baseline, two range workers, four route sessions,
cancellation/cleanup, process death, and the required metrics collection.
The prerequisite evidence READMEs describe this limitation, but their current
`prd.json` completion flags do not alter the live implementation above.
## Required follow-on before this acceptance can run
1. Bind the DGR-035 dense boundary adapter to a native llama.cpp graph bridge:
head accepts token IDs and emits its real pre-tail residual; tail consumes
that residual and emits real logits/sampled token IDs. Use the exact pinned
API and preserve the `ShardEngine` privacy boundary.
2. Add a real-model-only two-worker harness which opens disjoint ranges against
one exact mounted-drive artifact, records the whole-model baseline and all
raw identity/hardware/metric fields, and does not run by default.
3. Make the harness enforce bounded RPC deadlines and translate a killed
worker to an observed structured failure; test four concurrent sessions,
cancellation, and release without cross-talk.
4. Run it on a host with loopback sockets and an explicitly selected GGUF.
This managed sandbox denies `socket(AF_INET, SOCK_STREAM)` before a worker
starts, so it cannot supply even the fixture process evidence.
No criterion is weakened and no real-model evidence is claimed.

View File

@@ -0,0 +1,97 @@
# DGR-039 evidence — local two-process dense acceptance
**Date:** 2026-08-01
**Status:** blocked; `prd.json` remains authoritative and keeps
`DGR-039.passes` as `false`.
## Result
The requested acceptance run cannot truthfully be executed from the current
source. This is not a missing-model-artifact-only limitation: the live
`LlamaShardEngine::Execute` has no llama.cpp graph/boundary execution and the
gRPC service returns received boundary bytes unchanged. Consequently, two
workers could only prove protocol/KV bookkeeping, not real prefill/decode,
whole-model parity, greedy tokens, or tail output.
See [BLOCKED.md](BLOCKED.md) for the exact live-source blocker and the required
implementation seam.
## Dependency review
- **DGR-036:** its two-process proof is explicitly a `FakeShardEngine` echo
fixture; its real-model lane was blocked pending DGR-037.
- **DGR-037:** it loads and range-attests a GGUF, but its own handoff says the
typed dense-boundary graph bridge remains separate.
- **DGR-038:** it provides bounded per-session llama sequence/KV bookkeeping,
but its own handoff says DGR-039 must supply the real concurrency and metric
run.
The live source confirms those limits: `llama_shard_engine.cpp` increments
`past_len` without computing a graph, and `shard_service.cpp` echoes both
prefill/decode bundles.
## Commands and results
```bash
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_architecture_boundary.py tests/test_llama_shard_worker_binding.py \
tests/test_native_shard_protocol.py
```
```text
61 passed, 2 skipped in 0.51s
```
```bash
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python -m compileall -q packages tests
git diff --check
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
```
```text
OK: 55 stories validated.
```
```bash
/home/popov/.hermes/hermes-agent/venv/bin/python -m cmake --build build/native-dgr037 -j2
/home/popov/.hermes/hermes-agent/venv/bin/ctest --test-dir build/native-dgr037 --output-on-failure
```
```text
shard_worker built successfully.
1/1 shard_protocol_conformance passed.
```
Attempted existing two-worker fixture:
```bash
PYTHONPATH=packages/node:packages/tracker /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_native_shard_worker.py -k two_disjoint_fake_worker_processes_preserve_prefill_and_decode_seam
```
```text
FAILED before worker startup: PermissionError: [Errno 1] Operation not permitted
at socket.socket(AF_INET, SOCK_STREAM).
```
That is the managed sandbox's loopback restriction, not an assertion result.
Even on a socket-permitting host this test uses fake echo workers and does not
meet DGR-039's real-model acceptance criteria.
## Changed files
- `.scratch/distributed-gguf-runtime/evidence/DGR-039/README.md`
- `.scratch/distributed-gguf-runtime/evidence/DGR-039/BLOCKED.md`
- `.ralph-tui/progress.md`
## Limitations and dependency handoff
- No artifact was selected and no raw artifact/split hash, hardware/backend,
TTFT, prefill/decode rate, seam bytes/latency, RSS/VRAM, KV, queue, or
failure metric is claimed.
- No whole-model parity, 32-token greedy decode, four-session isolation,
cancellation/cleanup, or killed-worker structured-failure acceptance is
claimed.
- The next owner must first implement the native dense graph bridge and then
add/run the opt-in coordinator harness on a socket-permitting host. Keep
`DGR-039.passes` false until it has the required real run evidence.

View File

@@ -0,0 +1,89 @@
# DGR-040 evidence — node-side native worker supervision
**Date:** 2026-08-01
**Authority:** `.scratch/distributed-gguf-runtime/prd.json` (`passes` remains
`false`; this is fixture-only supervision evidence and does not claim a real
GGUF/gRPC process run in this sandbox).
## Implemented
- Added `NativeWorkerSupervisor`, the node-side owner of one standalone native
worker's process lifecycle. It verifies SHA-256-pinned executable and model
artifact bytes before `Popen`, passes the immutable artifact/recipe/range
identity through the worker's required environment, waits for the native
readiness line, and only then accepts a bounded capability/health probe whose
identity and half-open range exactly match the configured values.
- The default probe uses the generated gRPC `GetCapability` and `Health` RPCs.
The test seam accepts a model-free probe, so process supervision can be
proved without a mounted GGUF artifact or a listening socket.
- Both stdout and stderr are captured into a bounded in-memory log tail.
`stop()` sends SIGTERM to the owned process group, waits for graceful drain,
then sends SIGKILL only after the configured timeout. `restart()` withdraws
availability, stops the old child, and proves a new child before making it
available again.
- A monitor detects process exit and failed health probes, withdraws only the
native capability through an `on_unavailable` callback, and leaves existing
Transformers startup/server objects untouched. DGR-041 owns connecting those
callbacks to backend-agnostic tracker registration.
- Added deterministic fake-worker tests. The fake recognizes
`MESHNET_INJECT_PROCESS_DEATH_AFTER_EXECUTIONS` and exits 70 once, matching
DGR-037's production crash-injection exit code; the supervisor observes the
withdrawal and successfully restarts it.
## Changed files
- `packages/node/meshnet_node/native_worker_supervisor.py`
- `tests/test_native_worker_supervisor.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-040/README.md`
- `.ralph-tui/progress.md`
## Commands and results
```bash
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_native_worker_supervisor.py tests/test_llama_shard_worker_binding.py \
tests/test_native_shard_protocol.py
# 60 passed, 2 skipped in 0.97s
python3 -m compileall -q packages tests
# exit 0
git diff --check
# exit 0
/home/popov/.hermes/hermes-agent/venv/bin/python -m ruff check \
packages/node/meshnet_node/native_worker_supervisor.py \
tests/test_native_worker_supervisor.py
# All checks passed!
```
The system Python and repository `.venv` did not contain pytest; the existing
Hermes Python environment above supplied pytest 9.0.3 and grpc for the focused
checks. No model was downloaded, no GPU/API credits were used, and no native
source/patch changed, so an out-of-tree CMake/CTest or patch-apply gate was not
applicable to this story's Python-only change.
## Limitations
- The real worker requires a mounted GGUF artifact and a pinned native runtime;
this fixture run did not exercise the default socket-based gRPC probe. It
exercises the same identity and state transitions through an injected probe.
- Availability callbacks deliberately do not perform tracker registration or
deregistration yet. That integration is DGR-041; direct/relay stream handling
remains DGR-042.
- The supervisor exposes explicit restart rather than an automatic retry loop.
Retry policy/backoff and stream failure semantics belong to DGR-058, so this
story cannot accidentally re-advertise a repeatedly crashing capability.
## Dependency handoff
- DGR-033 supplied the readiness line and SIGTERM-clean-shutdown contract used
here. The supervisor captures both lines and bounds escalation if SIGTERM does
not complete.
- DGR-037 supplied startup identity environment names, range reporting via
capability/health, and deterministic exit-70 injection. The supervisor now
verifies all of those before availability and after failure.
- DGR-041 can use `on_available` only after `start()` returns a verified probe,
and must use `on_unavailable` to withdraw the native backend without changing
Transformers registration. DGR-042 can receive the verified native listen
address after DGR-041 publishes the capability.

View File

@@ -0,0 +1,99 @@
# DGR-041 evidence — backend-agnostic native Shard registration
**Date:** 2026-08-01
**Authority:** `.scratch/distributed-gguf-runtime/prd.json` (`passes` remains
`false`; this is model-free integration evidence, not a real hardware
certification).
## Implemented
- Added the optional, backend-neutral `ExecutionCapacity` capability-report
block: memory capacity in bytes, Hot-KV capacity in tokens, and maximum
concurrent Route Sessions. Existing Transformers reports omit it and keep
their previous serialized shape.
- Added `NativeShardRegistration`, which accepts only an exact `ShardIdentity`,
DGR-040 startup spec, and verified worker probe that all agree on artifact
digest, recipe fingerprint, recipe labels, and half-open range. It emits the
existing tracker registration payload and uses the capability report for
backend, capacity, and exact identity facts.
- Added `NativeCapabilityRegistrar.bind()` and additive supervisor callbacks:
publish happens only after DGR-040 has verified availability; a worker health
loss invokes caller-owned withdrawal. The adapter owns neither tracker HTTP
nor routing, billing, telemetry, relay, or provider policy.
- Tracker capability parsing/network state now preserves the three optional
capacity facts. Its existing `CertificationLedger` still registers the exact
native recipe as `dark` / `uncertified`, making it visible but unroutable.
No backend-name allowlist or routing special case was added.
## Changed files
- `packages/node/meshnet_node/capability.py`
- `packages/node/meshnet_node/native_registration.py`
- `packages/node/meshnet_node/native_worker_supervisor.py`
- `packages/tracker/meshnet_tracker/capability.py`
- `tests/test_native_registration.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-041/README.md`
- `.ralph-tui/progress.md`
## Commands and results
```bash
PYTHONPATH=packages/node:packages/tracker /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_native_registration.py tests/test_native_worker_supervisor.py \
tests/test_node_capability.py tests/test_runtime_recipe_identity.py
```
```text
101 passed in 0.71s
```
```bash
/home/popov/.hermes/hermes-agent/venv/bin/python -m ruff check \
packages/node/meshnet_node/capability.py \
packages/node/meshnet_node/native_registration.py \
packages/node/meshnet_node/native_worker_supervisor.py \
packages/tracker/meshnet_tracker/capability.py \
tests/test_native_registration.py
```
```text
All checks passed!
```
```bash
python3 -m compileall -q packages tests
git diff --check
```
```text
Both exit 0.
```
The default focused tests are model-download-free, API-credit-free, and
GPU-free. No model artifact was touched and nothing was written under `/home`.
## Limitations
- The full HTTP tracker-registration route suite could not run in this sandbox:
`PYTHONPATH=packages/node:packages/tracker /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q tests/test_tracker_capability_admission.py`
produced `25 passed, 9 failed`; every failure is the known sandbox
`PermissionError: [Errno 1] Operation not permitted` while creating an AF_INET
listening socket. The model-free direct tracker admission path is exercised
by `test_native_registration.py` and the existing identity suite.
- No native source/protobuf/patch changed, so an out-of-tree CMake/CTest build
and pin patch apply/check/reverse gates are not applicable.
- The registrar deliberately takes caller-owned register/withdraw callbacks.
DGR-042 owns the native direct/relay activation endpoint; deployment wiring
must provide its existing tracker transport rather than invent another one.
- No real backend/model/recipe combination is certified by this change.
`prd.json` remains false until the authoritative execution process grants
completion credit.
## Dependency handoff
- **DGR-025:** `ShardIdentity` and the tracker-owned `CertificationLedger` are
used directly; do not substitute labels for the fingerprint or promote a
recipe in node code.
- **DGR-040:** construct this registration from the post-`start()` verified
probe and call `NativeCapabilityRegistrar.bind(supervisor)` before startup.
Its unavailable callback must withdraw only the native capability.
- **DGR-042:** consume the registration's verified native endpoint through the
existing direct/relay route mechanism; keep its protobuf transport opaque to
tracker admission.

View File

@@ -0,0 +1,73 @@
# DGR-042 evidence — native frames through direct and relay seams
**Date:** 2026-08-01
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`.
## Implemented
- Added `NativeActivationSeam`, a Route-Session-scoped adapter with exactly two
selectable transports. Direct traffic calls the generated
`ShardRuntimeStub.Session()` once and keeps its bidirectional gRPC stream
open for the session. Its request and response hand-off queues are bounded.
- Relay traffic calls the existing persistent relay request shape with
`POST /native/session`, `application/x-protobuf`, and the exact
`SessionRequest.SerializeToString()` body. It parses only the returned
`SessionResponse`; neither the adapter nor the relay contract rewrites a
protobuf frame. Relay failure is explicitly uncertain and is never retried.
- `NativeFrameContext` validates Route Session, epoch, work, and deadline
fields against the versioned protobuf request before either path sends it.
The unchanged existing relay header contract receives request/billing ID,
node attribution, route, work, and deadline copies for control-plane
telemetry/billing correlation. `NativeSeamTelemetry` reports per-node,
per-request seam byte/latency observations without interpreting frames.
- Deterministic fake-worker tests cover a single direct stream, byte-identical
relay request frames, relay disconnect/no replay, cancellation, correlation
headers, telemetry, and bounded direct buffering.
## Changed files
- `packages/node/meshnet_node/native_activation_seam.py`
- `tests/test_native_activation_seam.py`
- `.scratch/distributed-gguf-runtime/prd.json`
- `.scratch/distributed-gguf-runtime/issues/042-carry-native-frames-through-direct-and-existing-relay-seams.md`
- `.scratch/distributed-gguf-runtime/evidence/DGR-042/README.md`
- `.ralph-tui/progress.md`
## Commands and results
```bash
PYTHONPATH=packages/node:packages/tracker /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_native_activation_seam.py tests/test_native_shard_protocol.py \
tests/test_native_worker_supervisor.py tests/test_native_registration.py \
tests/test_ralph_prd_schema.py
# 172 passed, 2 skipped in 2.01s
/home/popov/.hermes/hermes-agent/venv/bin/python -m ruff check \
packages/node/meshnet_node/native_activation_seam.py tests/test_native_activation_seam.py
# All checks passed!
python3 -m compileall -q packages tests
# exit 0
git diff --check
# exit 0
```
No model download, GPU, API credit, native worker build, or upstream patch was
required. Native CMake/CTest and patch-stack gates do not apply to this
Python-only transport adapter.
## Limitations and dependency handoff
- Relay is deliberately a sequence of opaque existing relay RPC bodies, not a
gRPC tunnel. The direct path alone is a long-lived gRPC stream; this avoids
changing relay behavior while preserving native frame bytes.
- This fixture lane uses an injected generated-stub-shaped fake worker and an
injected existing-relay-client-shaped callable. DGR-054/DGR-058 must use the
adapter with certified workers and add real route-loss/restart policy; they
must retain the no-replay rule after an uncertain relay send.
- DGR-024 supplied the versioned generated `Session` protocol and the prior
raw-frame identity proof. DGR-040 supplied the verified worker lifecycle;
its published native listen address is the direct endpoint for this seam.
- Existing Transformer HTTP routes and relay routing, load balancing, billing,
and peer behavior were not changed.

View File

@@ -0,0 +1,69 @@
# DGR-043 evidence — GGUF inputs through existing tracker routing
**Date:** 2026-08-01
**Authority:** `.scratch/distributed-gguf-runtime/prd.json` (`passes` remains
`false`; this is model-free integration evidence, not a hardware certification).
## Implemented
- Added optional backend-neutral `RoutingMeasurements` to the existing capability report. It carries measured tokens/second, queue depth, seam latency, health, and reliability; reports that omit it retain their exact previous serialized shape.
- Extended the trackers existing sanitized `CapabilityState` and network-map capability view to retain the routing measurements with exact recipe, artifact/runtime fingerprint, half-open-range-derived coverage, capacity, backend, and certification facts.
- `NativeShardRegistration` now accepts this generic measurement block and adapts throughput and queue depth to the existing registration/heartbeat scoring inputs. The tracker continues to apply its established queue-adjusted throughput selection; no GGUF routing, balancing, billing, relay, provider, quantization, topology, or architecture branch was added.
- Added deterministic coverage tests showing that existing route formation excludes a dark candidate, forms a complete route only from matching exact fingerprints, and rejects a range otherwise covered only by a mismatched recipe.
## Changed files
- `packages/node/meshnet_node/capability.py`
- `packages/node/meshnet_node/native_registration.py`
- `packages/tracker/meshnet_tracker/capability.py`
- `packages/tracker/meshnet_tracker/server.py`
- `tests/test_native_registration.py`
- `.scratch/distributed-gguf-runtime/evidence/DGR-043/README.md`
- `.ralph-tui/progress.md`
## Commands and results
```bash
PYTHONPATH=packages/node:packages/tracker /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_native_registration.py tests/test_node_capability.py \
tests/test_runtime_recipe_identity.py
```
```text
96 passed in 0.23s
```
```bash
PYTHONPATH=packages/node:packages/tracker /home/popov/.hermes/hermes-agent/venv/bin/python -m pytest -q \
tests/test_dgr_performance_contract.py tests/test_native_activation_seam.py \
tests/test_native_worker_supervisor.py tests/test_native_registration.py \
tests/test_ralph_prd_schema.py
```
```text
151 passed in 1.78s
```
```bash
/home/popov/.hermes/hermes-agent/venv/bin/python -m ruff check \
packages/node/meshnet_node/capability.py \
packages/node/meshnet_node/native_registration.py \
packages/tracker/meshnet_tracker/capability.py \
packages/tracker/meshnet_tracker/server.py tests/test_native_registration.py
python3 -m compileall -q packages tests
git diff --check
```
```text
All checks passed; both remaining commands exited 0.
```
Default tests were model-download-free, API-credit-free, and GPU-free. No native source, protobuf, patch, model artifact, or mounted-drive content was changed; therefore native CMake/CTest, patch-stack, and real-hardware gates do not apply to this Python-only adapter.
## Limitations
- The full HTTP tracker/admission and tracker-routing suites cannot bind an AF_INET listener in this sandbox. The attempted focused suite had 132 passes and 14 failures, all `PermissionError: [Errno 1] Operation not permitted` during socket creation. Model-free direct tracker parsing and route-formation tests cover this change; HTTP/billing/relay regression suites must be rerun in an environment that permits localhost sockets.
- Measurements are inputs, not self-certification. An exact native recipe remains `dark` until the existing tracker-owned certification ledger admits it, and worker health loss continues to withdraw the native capability.
- Seam latency is retained as a measured tracker capability input. Existing route latency learning remains the tracker-owned mechanism for end-to-end seam cost; this story intentionally does not alter its scoring algorithm.
## Dependency handoff
- **DGR-041:** `NativeShardRegistration`, `ExecutionCapacity`, exact `ShardIdentity`, and the tracker certification ledger remain the only registration/admission path. Supply `RoutingMeasurements` from verified worker/telemetry observations; do not infer values from backend names, quantization labels, architecture, or stage topology.
- **DGR-053/DGR-061:** use the exposed opaque measurements and existing tracker routing mechanisms for real certified routes. Any real-run evidence must add artifact/split hashes, worker/upstream pins, backend/driver, hardware/network details, commands, and raw metrics.

View File

@@ -1,15 +1,15 @@
# Ralph task evidence
# Distributed GGUF Runtime evidence
Each completed story creates `evidence/<TASK-ID>/README.md`. Fresh dependent iterations must read it before coding.
> **Specification status:** planning artifacts only. No distributed GGUF runtime is implemented. DGR-017 cleanup is complete; no runtime implementation story has completion credit. `prd.json` is authoritative.
Required README sections:
## Authority and classes
1. Summary and acceptance decision.
2. Exact files changed.
3. Commands run and real exit/results.
4. Correctness, performance and hardware evidence classification.
5. Known limitations and deferred work.
6. Compatibility/migration notes.
7. Explicit handoff for each dependent story.
Evidence supports but never overrides `prd.json`. Valid classes are `model-free`, `fixture`, `real-model`, `real-hardware`, and `release`; lower classes cannot satisfy higher-class acceptance. Legacy DGR-001..016 directories remain unchanged for DGR-017 provenance audit and confer no completion credit.
Store raw machine-readable metrics, manifests and protocol artifacts beside the README. Never store secrets, model weights, build outputs or Ralph iteration logs here.
## Future story layout
Each DGR-017..071 story writes `evidence/<ID>/README.md` with summary, exact changed files, exact commands and real outputs, limitations, compatibility/migration notes, and dependent-story handoff. Machine-readable contracts, manifests, metrics, and raw logs live beside it. Never fabricate output.
Real runs record exact model SHA and all split hashes, tokenizer, quant/recipe, llama.cpp pin+patch identity, backend/driver/toolchain, host/hardware/network, commands/environment (without secrets), raw parity/performance/resource results, and evidence class. Models live on configured mounted-drive storage, never `/home`.
Routing certification records prove only the exact exercised backend/model/recipe lane. Compile-only, fixture, failed, or unavailable lanes remain registered-dark. V4 cache/state evidence must show KV and CSA/HCA/SWA/indexer/compressor data remain shard-local/session-keyed; route recovery evidence must show cache miss plus re-prefill/restart, not migration.

View File

@@ -0,0 +1,335 @@
{
"repository": "https://git.d-popov.com/popov/neuron-tai",
"stories": {
"DGR-017": {
"number": 1,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/1",
"state": "closed",
"status": "completed"
},
"DGR-018": {
"number": 2,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/2",
"state": "closed",
"status": "completed"
},
"DGR-019": {
"number": 3,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/3",
"state": "closed",
"status": "completed"
},
"DGR-020": {
"number": 4,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/4",
"state": "closed",
"status": "completed"
},
"DGR-021": {
"number": 5,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/5",
"state": "closed",
"status": "completed"
},
"DGR-022": {
"number": 6,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/6",
"state": "closed",
"status": "completed"
},
"DGR-023": {
"number": 7,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/7",
"state": "closed",
"status": "completed"
},
"DGR-024": {
"number": 8,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/8",
"state": "closed",
"status": "completed"
},
"DGR-025": {
"number": 9,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/9",
"state": "closed",
"status": "completed"
},
"DGR-026": {
"number": 10,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/10",
"state": "closed",
"status": "completed"
},
"DGR-027": {
"number": 11,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/11",
"state": "closed",
"status": "completed"
},
"DGR-028": {
"number": 12,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/12",
"state": "closed",
"status": "completed"
},
"DGR-029": {
"number": 13,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/13",
"state": "closed",
"status": "completed"
},
"DGR-030": {
"number": 14,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/14",
"state": "open",
"status": "in-progress"
},
"DGR-031": {
"number": 15,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/15",
"state": "open",
"status": "ready"
},
"DGR-032": {
"number": 16,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/16",
"state": "open",
"status": "blocked"
},
"DGR-033": {
"number": 17,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/17",
"state": "open",
"status": "blocked"
},
"DGR-034": {
"number": 18,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/18",
"state": "open",
"status": "blocked"
},
"DGR-035": {
"number": 19,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/19",
"state": "open",
"status": "blocked"
},
"DGR-036": {
"number": 20,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/20",
"state": "open",
"status": "blocked"
},
"DGR-037": {
"number": 21,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/21",
"state": "open",
"status": "blocked"
},
"DGR-038": {
"number": 22,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/22",
"state": "open",
"status": "blocked"
},
"DGR-039": {
"number": 23,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/23",
"state": "open",
"status": "blocked"
},
"DGR-040": {
"number": 24,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/24",
"state": "open",
"status": "blocked"
},
"DGR-041": {
"number": 25,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/25",
"state": "open",
"status": "blocked"
},
"DGR-042": {
"number": 26,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/26",
"state": "open",
"status": "blocked"
},
"DGR-043": {
"number": 27,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/27",
"state": "open",
"status": "blocked"
},
"DGR-044": {
"number": 28,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/28",
"state": "open",
"status": "ready"
},
"DGR-045": {
"number": 29,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/29",
"state": "open",
"status": "blocked"
},
"DGR-046": {
"number": 30,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/30",
"state": "open",
"status": "blocked"
},
"DGR-047": {
"number": 31,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/31",
"state": "open",
"status": "blocked"
},
"DGR-048": {
"number": 32,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/32",
"state": "open",
"status": "blocked"
},
"DGR-049": {
"number": 33,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/33",
"state": "open",
"status": "blocked"
},
"DGR-050": {
"number": 34,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/34",
"state": "open",
"status": "blocked"
},
"DGR-051": {
"number": 35,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/35",
"state": "open",
"status": "blocked"
},
"DGR-052": {
"number": 36,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/36",
"state": "open",
"status": "blocked"
},
"DGR-053": {
"number": 37,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/37",
"state": "open",
"status": "blocked"
},
"DGR-054": {
"number": 38,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/38",
"state": "open",
"status": "blocked"
},
"DGR-055": {
"number": 39,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/39",
"state": "open",
"status": "blocked"
},
"DGR-056": {
"number": 40,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/40",
"state": "open",
"status": "blocked"
},
"DGR-057": {
"number": 41,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/41",
"state": "open",
"status": "blocked"
},
"DGR-058": {
"number": 42,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/42",
"state": "open",
"status": "blocked"
},
"DGR-059": {
"number": 43,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/43",
"state": "open",
"status": "blocked"
},
"DGR-060": {
"number": 44,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/44",
"state": "open",
"status": "blocked"
},
"DGR-061": {
"number": 45,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/45",
"state": "open",
"status": "blocked"
},
"DGR-062": {
"number": 46,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/46",
"state": "open",
"status": "blocked"
},
"DGR-063": {
"number": 47,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/47",
"state": "open",
"status": "blocked"
},
"DGR-064": {
"number": 48,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/48",
"state": "open",
"status": "blocked"
},
"DGR-065": {
"number": 49,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/49",
"state": "open",
"status": "blocked"
},
"DGR-066": {
"number": 50,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/50",
"state": "open",
"status": "blocked"
},
"DGR-067": {
"number": 51,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/51",
"state": "open",
"status": "blocked"
},
"DGR-068": {
"number": 52,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/52",
"state": "open",
"status": "blocked"
},
"DGR-069": {
"number": 53,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/53",
"state": "open",
"status": "blocked"
},
"DGR-070": {
"number": 54,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/54",
"state": "open",
"status": "blocked"
},
"DGR-071": {
"number": 55,
"url": "https://git.d-popov.com/popov/neuron-tai/issues/55",
"state": "open",
"status": "blocked"
}
}
}

View File

@@ -1,247 +1,40 @@
# Focused implementation strategy: performant concurrent distributed inference
# Distributed GGUF Runtime implementation strategy
Status: Accepted planning direction
Last updated: 2026-07-13
> **Specification status:** planning artifacts only. No distributed GGUF runtime is implemented. DGR-017 cleanup is complete; no runtime implementation story has completion credit. `prd.json` is authoritative.
## Product objective
## Execution model
Enable clients to run top open models that do not fit on one consumer machine by combining independently owned model Shards into performant, concurrent Inference Routes.
Execute one numerically ordered, dependency-ready story per fresh Ralph context. Read `RALPH-CONTEXT.md`, source issue, and dependency evidence first; use TDD/fixture-first verification; finish with exact evidence. `prd.json` is the only state authority.
The alpha-release target is the exact `zai-org/GLM-5.2` model, pinned by revision and served with `reasoning_effort=max`, using the smallest published Unsloth `UD-IQ1_S` GGUF across physical consumer machines. See [GLM-5.2-MAX-ALPHA-ROADMAP.md](GLM-5.2-MAX-ALPHA-ROADMAP.md). Dense Llama remains a cheap structural fixture; Qwen expansion is post-alpha.
## Sequence
The project is not trying to reproduce every vLLM feature or support every inference engine. It is optimizing for:
1. **M0 DGR-017..020:** reconcile legacy reality, lock metadata/performance contracts, and run the independent whole-model baseline.
2. **M1 DGR-021..033:** protocol/lifecycle/codegen, exact identities and split artifacts, pinned upstream/patches, CPU then accelerator builds, `ShardEngine`, fixtures, fake worker.
3. **M2 DGR-034..043:** dense ranged ownership/boundary/parity/local state, worker integration, supervision/direct-relay, and measured GGUF inputs to unchanged routing.
4. **M3 DGR-044..054:** pin/inventory V4, adapt upstream boundary/local state/MoE/hash execution, pass parity and real 24 scenario, then enforce alpha with MTP off.
5. **M4 DGR-055..067:** batching/backpressure/failure/recovery/long-context, existing-routing 10+ certification, real scale, measured optimization/compression, MTP contract+implementation, hardware certification.
6. **M5 DGR-068..071:** packages, human upstream collaboration, beta gate (including MTP), and pin/patch/certification maintenance.
1. Models larger than one node's RAM/VRAM.
2. Useful interactive decode speed on consumer CPU, AMD, NVIDIA, Vulkan, and mixed routes where certified.
3. Multiple concurrent Route Sessions without cache corruption or global serialization.
4. A lean runtime with one control plane and one primary GGUF engine.
5. Measured improvement over the existing Transformers/safetensors implementation.
## Guardrails
## Current reality
The existing project already owns the differentiating distributed control plane:
## Locked scope
- Tracker-selected contiguous Shards.
- Stable Route Sessions.
- Local per-Shard Hot KV State in the Transformers reference backend.
- Binary Activation Seams.
- Relay/direct routing, cancellation, telemetry, billing, and capability admission.
- Persistent relay and direct transport optimizations.
- Existing Meshnet Tracker routing, load balancing, billing, telemetry, relay, and provider semantics are backend-agnostic and are **not redesigned**. GGUF contributes exact compatibility, range/capacity, queue/load, seam-cost, health/reliability, and certification inputs only.
- The data plane is a standalone project-owned C++ Shard worker with gRPC/Protobuf and a project-owned `ShardEngine` boundary.
- llama.cpp is fetched at one exact commit into an ignored workspace from an in-repo manifest, then a numbered minimal patch stack is applied. There is no submodule, vendored tree, or permanent-fork dependency.
- llama.cpp owns DeepSeek V4 graphs, mHC, MoE, attention, hash routing, and kernels. Meshnet adds only range-ownership hooks, typed boundary/local-state adapters, worker integration, and parity/certification.
- Quantization and placement are dynamic recipe inputs. The 24 and 10+ stage layouts are certification scenarios, never product constants.
- Per-shard Hot KV and V4 CSA/HCA/SWA/indexer/compressor state remain local and keyed by route session/epoch. The WAN seam carries the typed mHC 4×4096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations—not per-layer caches.
- Route changes use cache miss plus re-prefill/restart. There is no WAN KV or V4 auxiliary-cache migration.
- CPU/CUDA/ROCm/Vulkan/Metal compile lanes are planned; only exact real-hardware-certified backend/model/recipe lanes may be advertised.
- Alpha requires correctness and the pre-locked useful-speed gate. MTP is reserved and off for alpha; its ownership contract, implementation, and benchmark are required before beta.
The missing production path is a native GGUF execution worker that can load and execute only an assigned layer range while retaining local Hot KV State for concurrent Route Sessions.
## Target identities
Whole-model llama.cpp, vLLM, and existing Transformers serving remain baselines or optional route kinds. They are not substitutes for native distributed Shards.
- DeepSeek V4 official target SHA: `60d8d70770c6776ff598c94bb586a859a38244f1`.
- llama.cpp V4 support lineage began at PR 24162 / merge `8c146a8366304c871efc26057cc90370ccf58dad`; DGR-027 later pins one exact validated current commit.
- V4 scope: 43 main layers plus MTP; mHC 4×4096 boundary; 256 routed + 1 shared experts with six routed active; token IDs required for the first three hash-routed layers.
- Exact split-GGUF artifacts are provisioned to mounted-drive storage with a complete hashed manifest and resumable verification; no model artifact may be placed under `/home`.
## Performance hypothesis—not an assumption
GGUF itself is a format. Performance comes from llama.cpp/GGML's quantized kernels, memory layout, mmap, backend scheduling, and reduced working set.
Quantized GGUF may be faster or may merely fit a larger model. Comparisons against safetensors must report both speed and quality because BF16 safetensors and Q4/Q8 GGUF are not numerically equivalent.
Before expensive native work, establish controlled lanes. DGR-001 remains immutable; DGR-017 adds a target-specific fit and semantics contract without rewriting DGR-001 evidence:
- Same model architecture and upstream revision.
- Same machine, prompt set, context, output length, sampling policy, and concurrency.
- Transformers/safetensors BF16 or the current production recipe.
- llama.cpp GGUF F16/BF16 or Q8 correctness lane where available.
- Q4_K_M or selected production quantization performance/fit lane.
- TTFT, prefill tok/s, decode tok/s, p50/p95 latency, RSS, VRAM, artifact size, energy where available, and output-quality drift.
The program proceeds only if llama.cpp/GGUF provides at least one meaningful advantage recorded in a machine-readable performance contract:
- Better decode or aggregate throughput at acceptable quality; or
- Materially lower memory that makes the target model routable while preserving useful throughput.
## Parallelism we will use
### Public Inference Route: layer/pipeline parallelism
Each node independently executes one contiguous Shard. Activations cross seams; weights and Hot KV State remain local.
This is the only public cross-machine model-parallel primitive in the first runtime.
### Per-node continuous batching
Autoregressive tokens remain sequential within one generation. Throughput comes from batching decode steps from multiple active Route Sessions inside each node using llama.cpp batches and sequence IDs or bounded context pools.
This is essential. A worker that globally serializes sessions is not production-ready.
### Multiple complete routes: data parallelism
The Tracker may select multiple complete routes for independent requests. This increases network throughput and availability without requiring collectives between routes.
### Trusted composite node: optional tensor/expert parallelism
Tensor parallelism and expert parallelism require frequent collectives and tight compatibility. They may be used later inside one operator-controlled composite node or managed cluster exposed as one logical provider. They are not public WAN routing primitives.
### Deferred mechanisms
- Disaggregated prefill and KV transfer.
- Speculative decoding.
- Cross-route prefix snapshots.
- Route repair with KV migration.
- Public tensor/expert parallel collectives.
They remain out of the critical path until the native layer route passes performance and concurrency gates.
## Reuse decisions
### llama.cpp/GGML: primary runtime substrate
Reuse:
- GGUF parsing and mmap.
- Quantized kernels.
- CPU, CUDA, HIP/ROCm, Vulkan, Metal, and other supported backends.
- Tokenizer and model architecture implementations.
- KV and sequence operations.
- Backend scheduler and graph execution.
Maintain a small exact-commit fork only for the missing local seam:
- Range-aware tensor ownership/loading.
- Architecture-defined boundary input/output.
- Intermediate boundary output without tail normalization.
- Layer-filtered KV and sequence mapping.
Keep networking, Tracker logic, billing, and public protocol outside llama.cpp. Upstream generic hooks where possible.
### vLLM: concepts and optional managed backend
Use unmodified vLLM only as:
- A whole-model node backend.
- A managed TP/PP/EP cluster represented as one logical provider.
- A performance/correctness baseline.
Adapt concepts, not runtime code:
- Named intermediate tensor bundles.
- Continuous batching and request-owner maps.
- Versioned KV-transfer compatibility fingerprints.
- Explicit send/receive/abort/failure lifecycle.
- Load telemetry and unbiased route selection.
Do not fork vLLM for public Shards and do not transplant PagedAttention, Torch process groups, or GGUF-plugin kernels into the llama.cpp worker.
### Nakshatra, prima.cpp, llama-gguf, LiGGUF, GPUStack
Use as source and test donors only:
- Nakshatra: partial-GGUF patches, daemon concepts, replay cases.
- prima.cpp: selected tensor ownership and local-layer KV evidence.
- llama-gguf: small protocol and integration-test patterns.
- LiGGUF: Q8 activation transport and tensor-reduction reference.
- historical GPUStack: resource preflight and role-oriented placement.
Do not adopt or fork their repositories wholesale.
### Mesh-LLM GLM branch: focused test/patch donor only
Use its GLM-5.2 branch to study DSA, IndexShare, stage-local KV, and sideband tests. Do not import its scheduler, discovery/control plane, package manager, or broad llama.cpp patch stack. Every adopted idea must be independently understood, minimized, attributed, and tested against our exact pin.
## Battle-proven transport decision
Use gRPC over HTTP/2 with Protocol Buffers for the native C++ Shard worker protocol.
Why:
- Mature Python and C++ implementations.
- Bidirectional streaming.
- HTTP/2 flow control and connection reuse.
- Deadlines, cancellation, status codes, TLS, authentication interceptors, and generated schemas.
- Avoids inventing a socket protocol.
Scope boundary:
- OpenAI-compatible client/Gateway APIs remain HTTP/SSE.
- Tracker/control APIs remain existing project interfaces.
- One long-lived bidirectional gRPC stream serves one Route Session Activation Seam.
- Existing relay/WebSocket infrastructure may carry the same versioned protobuf frames as opaque binary when direct gRPC reachability is unavailable.
- Large prefill tensors are chunked into bounded frames; decode bundles stay small.
- No QUIC/WebRTC/custom transport in this milestone.
The public boundary uses a versioned named-tensor bundle rather than one anonymous tensor because architecture boundaries can require more than `hidden_states`. DGR-006 updates the current single-`NamedTensor` decode fast path to carry the same bundle semantics and adds an explicit typed tail logits/token result with sampling/template identity.
Minimum identity:
```text
schema version
request/work id
Route Session id and route epoch
Model Artifact and runtime recipe fingerprint
Shard range and effective start
phase: prefill/decode/release/cancel
position/token range
named tensors with shape/dtype/byte order
compression and checksum
idempotency step id
cache expectation/result
```
## Concurrency model
A native worker must not use one global serving sequence or one lock around all model execution.
Required ownership:
```text
(Route Session id, route epoch)
-> local sequence/context
-> Shard-local Hot KV State
-> bounded lease and memory accounting
```
The node scheduler:
- Admits sessions against model memory and KV budget.
- Forms compatible decode batches from active sessions.
- Preserves per-session position and route order.
- Applies bounded queues and backpressure.
- Cancels/releases independently.
- Reports queue, batch, KV, prefill, decode, and seam telemetry.
Initial deterministic gate: at least four concurrent sessions on a small certified model with no token/KV cross-talk. Final concurrency targets are hardware/recipe-specific and recorded by capability admission rather than hardcoded globally.
## Stage gates
### Gate A: performance hypothesis
Controlled safetensors-versus-GGUF benchmark produces a signed/reproducible report and locks thresholds. Stop native work if there is no meaningful speed or fit benefit.
### Gate B: local range parity
Two local processes own disjoint GGUF ranges and match whole-model llama.cpp within the certified numerical tolerance for prefill and greedy decode.
### Gate C: concurrent KV
Multiple Route Sessions prefill/decode concurrently with isolated local KV, bounded memory, cancellation, and release.
### Gate D: real distributed route
Two physical machines execute one model that uses both Shards. Synthetic activation tests do not satisfy this gate.
### Gate E: consumer-hardware performance
On certified consumer hardware, the GGUF route beats the current distributed safetensors route under the locked performance contract or enables a larger otherwise-unroutable model at useful measured speed.
### Gate F: exact GLM-5.2 alpha target
After the generic dense fixture proves range and boundary mechanics, certify explicit GLM-5.2 MoE, MLA KV, DSA, IndexShare, and NextN policy. Alpha requires the exact `UD-IQ1_S` target across physical consumer nodes, native Max-mode semantics, locked parity/usefulness/performance thresholds, and bounded failure cleanup. Qwen3/Qwen3-MoE is later architecture expansion.
## Scope discipline
The following do not block the first production candidate:
- New cryptocurrency/economics work.
- New artifact P2P protocol.
- QUIC or WebRTC.
- vLLM fork.
- Whole-repository Nakshatra/prima adoption.
- Every GGUF architecture.
- Automatic route repair.
- Prefix snapshot migration.
- Speculative decoding.
- A large-model marketing demo before small-model parity and concurrency pass.
Every optimization must preserve output contract, session isolation, cancellation, resource cleanup, capability admission, and per-node attribution.
DGR-020 cannot use distributed results. DGR-054 does not depend on MTP. DGR-070 depends on DGR-066. Compile support and scenario success never imply general routability.

View File

@@ -0,0 +1,39 @@
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
# DGR-017: Reconcile and clean the superseded DGR backlog
- **Status / triage:** completed; `passes: true`
- **Execution mode:** `AFK`
- **Milestone:** `M0`
- **Dependencies:** None
- **Blocks (derived):** `DGR-018`, `DGR-019`, `DGR-027`, `DGR-054`
- **Labels:** `area:provenance`, `area:cleanup`, `type:audit`, `priority:p0`, `ready-for-agent`
- **Evidence class:** `model-free`
- **Hardware:** `none`
- **Model:** `generic`
- **Upstream:** `no`
## Objective / description
Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/017-reconcile-and-clean-the-superseded-dgr-backlog.md`, and evidence READMEs for dependencies (none) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Audit implementation reality, void inherited completion credit, and clean misleading backlog/stub baggage while preserving attributable evidence and accepted research.
## Acceptance criteria
- [x] Compare the branch, old DGR-001..016 issue/pass states, evidence, and actual runtime sources; classify each output as reusable, reference-only, blocked, obsolete, or absent.
- [x] Record an authoritative old-to-new disposition and provenance; explicitly give no completion credit to any new story and note absent implementation/evidence.
- [x] Remove or archive only artifacts the audit proves obsolete while preserving accepted ADRs, useful research, raw benchmark evidence, and attributable reusable work.
- [x] Protect ignored build workspaces, generated protobuf outputs, Ralph logs, and model artifacts from accidental commits, and document every retained legacy artifact.
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
## Shared quality gates
- Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`.
- `git diff --check` passes.
- Default tests are model-download-free, API-credit-free, and GPU-free.
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
## Evidence handoff
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-017/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.

View File

@@ -0,0 +1,39 @@
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
# DGR-018: Define canonical Ralph and Gitea metadata schema
- **Status / triage:** completed; `passes: true`
- **Execution mode:** `AFK`
- **Milestone:** `M0`
- **Dependencies:** `DGR-017`
- **Blocks (derived):** `DGR-021`, `DGR-025`
- **Labels:** `area:planning`, `area:gitea`, `type:infrastructure`, `priority:p0`, `ready-for-agent`
- **Evidence class:** `model-free`
- **Hardware:** `none`
- **Model:** `generic`
- **Upstream:** `no`
## Objective / description
Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/018-define-canonical-ralph-and-gitea-metadata-schema.md`, and evidence READMEs for dependencies (DGR-017) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Make `prd.json` the validated source from which Markdown and Gitea issues can later be generated losslessly.
## Acceptance criteria
- [x] Define fields for stable ID/title, labels, milestone, type, `dependsOn`, derived `blocks`, triage, evidence class, and hardware/model/upstream flags.
- [x] Validate that all stories start `passes: false`, use known dependencies, and have unique stable IDs.
- [x] Reject cycles, missing dependencies, mismatched generated `blocks`, duplicate titles/IDs, and generated artifacts claiming authority over `prd.json`.
- [x] Add deterministic model-free tests for parse, validation, and generation round trips.
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
## Shared quality gates
- Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`.
- `git diff --check` passes.
- Default tests are model-download-free, API-credit-free, and GPU-free.
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
## Evidence handoff
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-018/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.

View File

@@ -0,0 +1,40 @@
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
# DGR-019: Lock alpha and beta performance contracts
- **Status / triage:** completed; `passes: true`
- **Execution mode:** `HITL`
- **Milestone:** `M0`
- **Dependencies:** `DGR-017`
- **Blocks (derived):** `DGR-020`, `DGR-044`, `DGR-054`
- **Labels:** `area:performance`, `type:contract`, `priority:p0`, `gate:hitl`, `ready-for-human`
- **Evidence class:** `release`
- **Hardware:** `required`
- **Model:** `generic+deepseek-v4`
- **Upstream:** `no`
## Objective / description
Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/019-lock-alpha-and-beta-performance-contracts.md`, and evidence READMEs for dependencies (DGR-017) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Freeze useful speed, correctness, memory-fit, and stop/go thresholds before implementation results are visible.
## Acceptance criteria
- [x] Define controlled safetensors, whole-model GGUF, dense distributed GGUF, and V4 Flash distributed lanes with fixed prompts, context/output lengths, sampling, concurrency, hardware, and metrics.
- [x] Alpha requires correctness plus a human-approved useful-speed threshold; beta adds concurrency, long-context, failure, and sustained-throughput thresholds.
- [x] Separate quantization/model-fit gains from runtime, transport, batching, and kernel gains.
- [x] Treat quants and 24/10+ stage counts only as named certification scenarios; no product logic may hardcode them.
- [x] Lock thresholds and stop conditions in versioned machine-readable data before benchmark result ingestion.
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
## Shared quality gates
- Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`.
- `git diff --check` passes.
- Default tests are model-download-free, API-credit-free, and GPU-free.
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
## Evidence handoff
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-019/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.

View File

@@ -0,0 +1,39 @@
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
# DGR-020: Run the controlled whole-model GGUF baseline
- **Status / triage:** completed; `passes: true`
- **Execution mode:** `HITL`
- **Milestone:** `M0`
- **Dependencies:** `DGR-019`
- **Blocks (derived):** `DGR-054`
- **Labels:** `area:performance`, `type:benchmark`, `priority:p0`, `gate:hitl`, `ready-for-human`
- **Evidence class:** `real-hardware`
- **Hardware:** `required`
- **Model:** `generic`
- **Upstream:** `no`
## Objective / description
Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md`, and evidence READMEs for dependencies (DGR-019) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Execute the locked safetensors and whole-model llama.cpp lanes before distributed implementation results can influence the decision.
## Acceptance criteria
- [x] Run the exact DGR-019 safetensors and whole-model llama.cpp benchmark lanes with locked prompts, lengths, sampling, concurrency, hardware, and artifact/runtime identities.
- [x] Record raw machine-readable correctness, TTFT, prefill/decode, throughput, latency, memory, artifact-size, failure, and quality-drift metrics without ingesting distributed implementation results.
- [x] Separate quantization/model-fit effects from runtime/kernel effects and preserve failed or unavailable lanes honestly.
- [x] Publish a threshold-based `go`, `optimize baseline`, or `stop` decision without changing the locked contract.
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
## Shared quality gates
- Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`.
- `git diff --check` passes.
- Default tests are model-download-free, API-credit-free, and GPU-free.
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
## Evidence handoff
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-020/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.

View File

@@ -0,0 +1,39 @@
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
# DGR-021: Define the versioned named-tensor stream envelope
- **Status / triage:** completed; `passes: true`
- **Execution mode:** `AFK`
- **Milestone:** `M1`
- **Dependencies:** `DGR-018`
- **Blocks (derived):** `DGR-022`, `DGR-023`, `DGR-025`, `DGR-031`, `DGR-035`, `DGR-046`
- **Labels:** `area:protocol`, `type:infrastructure`, `priority:p0`, `ready-for-agent`
- **Evidence class:** `model-free`
- **Hardware:** `none`
- **Model:** `generic`
- **Upstream:** `no`
## Objective / description
Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/021-define-the-versioned-named-tensor-stream-envelope.md`, and evidence READMEs for dependencies (DGR-018) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Establish the backend-neutral protobuf envelope used by direct and relayed Shard activation traffic.
## Acceptance criteria
- [x] Define schema version, request/work ID, route session/epoch, shard range/effective start, phase, position, and idempotency step.
- [x] Define named tensors with shape, dtype, byte order, bounded fragments, compression identity, and checksum.
- [x] Reserve extensible fields for token-ID sidebands, architecture state, recurrent state, and MTP without claiming implementations.
- [x] Add deterministic serialization, fragmentation, checksum, unknown-field, and size-limit tests.
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
## Shared quality gates
- Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`.
- `git diff --check` passes.
- Default tests are model-download-free, API-credit-free, and GPU-free.
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
## Evidence handoff
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-021/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.

View File

@@ -0,0 +1,39 @@
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
# DGR-022: Define Shard lifecycle and structured status RPCs
- **Status / triage:** completed; `passes: true`
- **Execution mode:** `AFK`
- **Milestone:** `M1`
- **Dependencies:** `DGR-021`
- **Blocks (derived):** `DGR-024`, `DGR-033`, `DGR-037`
- **Labels:** `area:protocol`, `area:lifecycle`, `type:infrastructure`, `priority:p0`, `ready-for-agent`
- **Evidence class:** `model-free`
- **Hardware:** `none`
- **Model:** `generic`
- **Upstream:** `no`
## Objective / description
Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/022-define-shard-lifecycle-and-structured-status-rpcs.md`, and evidence READMEs for dependencies (DGR-021) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Complete the gRPC contract for worker capability, health, sessions, cancellation, release, and metrics.
## Acceptance criteria
- [x] Define capability, health, bidirectional session stream, cancellation, release, and metrics RPCs.
- [x] Specify deadlines, cancellation propagation, bounded flow control, cache expectations/results, and structured error taxonomy.
- [x] Specify TLS/auth hooks without moving Meshnet authentication or billing into the worker.
- [x] Add compatibility tests for supported versions and fail-closed tests for unsupported versions and malformed lifecycle transitions.
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
## Shared quality gates
- Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`.
- `git diff --check` passes.
- Default tests are model-download-free, API-credit-free, and GPU-free.
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
## Evidence handoff
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-022/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.

Some files were not shown because too many files have changed in this diff Show More