Compare commits
48 Commits
989b55970b
...
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f1ab618ee0 | ||
|
|
520bb60f5f | ||
|
|
2868fc0d56 | ||
|
|
521a7b108a | ||
|
|
b7d40c5bcf | ||
|
|
f0197bfa83 | ||
|
|
6aced6a005 | ||
|
|
c758106a42 | ||
|
|
f0ddb69d33 | ||
|
|
9cb334cced | ||
|
|
ffe937678a | ||
|
|
a35d86f343 | ||
|
|
9b257d9a1b | ||
|
|
3611b2cf9e | ||
|
|
efd1cf4ef6 | ||
|
|
ab466ce6b6 | ||
|
|
994546f78e | ||
|
|
6fd9d93e4b | ||
|
|
f102be1098 | ||
|
|
cae7c2b171 | ||
|
|
1749f9b4ad | ||
|
|
64f83d4392 | ||
|
|
6516a92c04 | ||
|
|
454a681a50 | ||
|
|
4eeec7fa7f | ||
|
|
a0f28b5631 | ||
|
|
7925e5253d | ||
|
|
91c450840d | ||
|
|
d6b808dcf9 | ||
|
|
31065c0e12 | ||
|
|
ec36290863 | ||
|
|
f844ae6567 | ||
|
|
252d131e7d | ||
|
|
3d8f93f4aa | ||
|
|
f9722e7b57 | ||
|
|
7b8e467c6b | ||
|
|
7364ed6731 | ||
|
|
ad2d17541c | ||
|
|
e7c780a623 | ||
|
|
9580ed643e | ||
|
|
5ebce15d7a | ||
|
|
ef2a9e67e8 | ||
|
|
b1c9deeb01 | ||
|
|
9e67b829e3 | ||
|
|
e24db7854f | ||
|
|
59f2486bf2 | ||
|
|
d904c40f66 | ||
|
|
30dcf953fe |
@@ -8,3 +8,4 @@
|
||||
- **Node capability admission** — `.scratch/node-capability-admission/` (P0 plan; [ADR-0023](../../docs/adr/0023-model-agnostic-node-capability-admission.md), [ADR-0026](../../docs/adr/0026-node-assignment-ownership-and-managed-placement.md))
|
||||
- **Distributed relay performance** — relay `/rpc` requester sockets are persistent per Route Session and Activation Seam as of 2026-07-10; `request_id` remains unique per activation while `X-Meshnet-Session` remains stable for KV state. Next low-risk priorities: persistent direct/loopback HTTP, seam byte/latency telemetry, then trace-driven zstd tuning.
|
||||
- **Distributed GGUF direction** — benchmark-gated native runtime: compare controlled Transformers/safetensors and whole-model llama.cpp lanes before expensive work; ship only for measured speed or model-fit advantage. Public parallelism is contiguous Shards in an Inference Route; concurrency comes from per-node continuous batching across isolated Route Sessions, while tensor/expert collectives stay inside optional trusted composite providers. Native data plane uses versioned Protobuf over long-lived gRPC/HTTP2 seam streams, with existing relay carrying the same opaque frames when needed. llama.cpp/GGML remains the substrate behind a project-owned standalone worker and small pinned fork; vLLM is an optional complete managed provider and concept donor, not a fork. Nakshatra, `prima.cpp`, `llama-gguf`, LiGGUF and historical GPUStack are source/test donors only. Active plan: [README](../../.scratch/distributed-gguf-runtime/README.md), [architecture](../../.scratch/distributed-gguf-runtime/architecture.md), [PRD](../../.scratch/distributed-gguf-runtime/PRD.md), [Ralph backlog](../../.scratch/distributed-gguf-runtime/prd.json). ADR: [0024](../../docs/adr/0024-distributed-gguf-runtime.md). Research: [landscape](../../docs/research/distributed-gguf-landscape.md), [GitHub follow-up](../../docs/research/distributed-gguf-github-followup.md), [vLLM](../../docs/research/vllm-distributed-gguf-assessment.md).
|
||||
- **Multi-subscription orchestration policy** — keep one fixed worktree per provider/agent and one user-selected integration branch. Do not switch the integration branch or create per-task branches/worktrees without explicit user confirmation. Parallel mode is default; task assignments should be disjoint and integration/push is serialized after each independently verified task. Serial mode uses an explicit provider priority, consumes the preferred subscription until its authoritative limit, then falls back in order and returns to higher priority after its official reset. Because Git cannot check out one named branch in multiple worktrees, provider worktrees should normally remain detached at the integration HEAD; the controller cherry-picks each verified task into the unchanged integration branch, tests, and pushes, then resynchronizes every fixed worktree.
|
||||
|
||||
29
.claude/memory/dgr-rocm-setup.md
Normal file
29
.claude/memory/dgr-rocm-setup.md
Normal file
@@ -0,0 +1,29 @@
|
||||
# DGR ROCm and llama.cpp setup
|
||||
|
||||
As of 2026-07-13:
|
||||
|
||||
- Project ROCm runtime: `/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv-rocm`
|
||||
- ROCm/TheRock build: `7.13.0a20260513`, target `gfx1151`
|
||||
- `rocm-sdk-devel` is installed. Its expanded SDK lives under the venv at
|
||||
`site-packages/_rocm_sdk_devel`.
|
||||
- The wheel's redundant packaged payload was relocated to
|
||||
`/home/popov/.local/share/rocm-sdk/7.13.0a20260513/rocm_sdk_devel` and symlinked
|
||||
back into the venv because installing both packaged and expanded forms filled
|
||||
the mounted drive. Do not reinstall it blindly; the wheel expands beyond
|
||||
20 GB.
|
||||
- HIP llama.cpp source: `/run/media/popov/d/DEV/llamacpp/llama.cpp`, commit
|
||||
`e920c523e3b8a0163fe498af5bf90df35ff51d25` (version 9991).
|
||||
- HIP build: `/run/media/popov/d/DEV/llamacpp/llama.cpp/build-hip`
|
||||
- HIP `llama-server` SHA-256:
|
||||
`b6bb4da687dbde86e243ba006cef05919b7b97255cd7e2371e1d451220aca139`
|
||||
- Verified device: `ROCm0: Radeon 8060S Graphics`, `gfx1151`.
|
||||
- Model artifacts remain under `/run/media/popov/DATA/llm`; none were put under
|
||||
`/home`.
|
||||
|
||||
DGR-001's immutable contract remains CPU-only. GPU evidence uses the distinct
|
||||
signed `gpu-diagnostic` profile because llama-server process VRAM is not yet
|
||||
measurable by the benchmark driver. The profile must capture measured
|
||||
llama-server startup evidence for `ROCm0` and the actual offloaded/total layer
|
||||
count; configured `device` and `n_gpu_layers` values alone are not evidence.
|
||||
The accepted signer fingerprint is anchored in
|
||||
`.scratch/distributed-gguf-runtime/trusted-evidence-signers.json`.
|
||||
@@ -8,6 +8,23 @@ metadata:
|
||||
|
||||
# Project Status (2026-07-13)
|
||||
|
||||
## Distributed GGUF controller checkpoint (2026-07-21)
|
||||
|
||||
- All three fixed detached lanes from the 2026-07-18 checkpoint (Fable/DGR-025, Kimi/DGR-028, Terra/DGR-024) were reviewed, committed, and merged into `ralph/distributed-gguf-runtime`, now at `cd6b4d9`. The `.claude/worktrees/ralph-fable-loop`, `ralph-kimi-loop`, `ralph-terra-loop`, `ralph-cursor-loop`, and `ralph-next-task` worktrees were removed after merge (cursor-loop/next-task were idle with no work in progress). Only `.claude/worktrees/distributed-gguf-runtime` (the integration checkout) remains, plus the unrelated `fix-tracker-incomplete-snapshot` worktree (locked, not part of this arc).
|
||||
- DGR-025 and DGR-028 merged cleanly with no conflicts (105 + 7 tests passing, 112 together). DGR-024 required real completion work first: the worktree's `prd.json` note was stale (written for the old, policy-rejected in-memory fake-seam story). The actual code already present (`shard_runtime_server.py`, `test_shard_runtime_harness.py`) had correctly pivoted to a real subprocess/socket gRPC harness with direct-vs-opaque-relay byte-identity proof, but was missing the required fail-closed negative paths (stale epoch, expired deadline, malformed/checksum-corrupt fragments, exhausted flow-control credit, duplicate idempotency steps, in-band/out-of-band cancel). Implemented those (`SessionState` per `route_session_id`, `_validate_bundle`), added 9 new tests (11 total, all passing), and rewrote `evidence/DGR-024/README.md` to describe the actual implementation instead of the nonexistent `FakeShardSeam`.
|
||||
- `prd.json`'s `passes` field was left `false` for all three stories — that flag is only flipped by the project's own independent controller review process, not by whoever lands the merge. DGR-028's evidence README explicitly still awaits that P0/P1 review.
|
||||
- Post-merge full suite (before the DGR-024 merge): 3 failed / 1116 passed / 20 skipped — down from the prior 9-failure baseline, and the 3 remaining failures (billing default-db, dynamic-routing ADD_SHARD/LOAD_SHARD, preset dedup) are pre-existing and unrelated to native/runtime-identity code. Did not get a chance to rerun the full suite after the DGR-024 merge landed (interrupted); only the focused `tests/test_shard_runtime_harness.py` (11 passed) was reconfirmed post-merge — a fresh full-suite run is worth doing before treating the whole arc as done.
|
||||
- The unrelated dirty edits on the integration worktree (`model_catalog.py` / `tests/test_mining_cli.py`, adding a `Qwen2.5-Coder-1.5B-Instruct-Q2_K-GGUF` preset) were stashed during the merge sequence and popped back afterward — still uncommitted, as before.
|
||||
|
||||
**Why:** user directed "distributed-gguf-runtime is where we need to merge all ralph-* branches" and to prune worktrees once their task's work is merged, so completed lanes don't linger and new worktrees signal new tasks unambiguously.
|
||||
**How to apply:** next Ralph session picking up this project should start from `ralph/distributed-gguf-runtime` at `cd6b4d9`, rerun the full suite once to get a clean current baseline, and check whether independent review has flipped DGR-024/025/028's `passes` flags before selecting new work.
|
||||
|
||||
## Distributed GGUF controller checkpoint (2026-07-18)
|
||||
|
||||
- Integration branch `ralph/distributed-gguf-runtime` is at `377bc3475c41b762ebbcab038c05adf15d7749d0`, matching its remote; preserve the unrelated dirty integration edits in `model_catalog.py` and `test_mining_cli.py`.
|
||||
- Fixed detached lanes hold three claimed, uncommitted tasks: Fable DGR-025/Gitea #9 executing-artifact identity repair; Terra DGR-024/#8 fake gRPC seam repair; Kimi DGR-028/#12 numbered llama.cpp patch stack. No provider worker was running at the latest reconciliation, so continue these dirty trees before selecting new work.
|
||||
- DGR-028 controller repair added working apply/reverse/verify enforcement, exact assumptions, and a passing exact-pin native fixture; final independent P0/P1 review is still required before commit/integration. DGR-024 and DGR-025 focused gates pass but likewise remain provisional pending the current independent reviews. Controller reruns on 2026-07-18 passed DGR-025 (235 impacted tests), DGR-024 (78 shared/focused tests), and DGR-028 (7 Python tests plus exact-pin native CTest 1/1 and clean apply/reverse). Integrate DGR-028 before DGR-025, then replay DGR-025 on the new integration HEAD and regenerate/retest its runtime fingerprint vectors because DGR-028 changes the pinned patch-stack/tree identity. The integration full-suite baseline is 9 failures/1076 passes/22 skips: missing optional zstd/langchain dependencies plus pre-existing billing and dynamic-routing expectations; lane-only extra failures map to stale DGR-023 projection ancestry and a timing-sensitive cancel test, not the focused story paths.
|
||||
|
||||
## Selected-node model placement (2026-07-14)
|
||||
|
||||
- Admin Model placement now opens a node selector for load and release; the control-plane accepts optional `node_id` and targets only that registry assignment. Multi-model serving remains supported through `ADD_SHARD` and `max_loaded_shards`.
|
||||
|
||||
384
.scratch/distributed-gguf-runtime/GLM-5.2-MAX-ALPHA-ROADMAP.md
Normal file
384
.scratch/distributed-gguf-runtime/GLM-5.2-MAX-ALPHA-ROADMAP.md
Normal file
@@ -0,0 +1,384 @@
|
||||
# GLM-5.2 Max distributed alpha roadmap
|
||||
|
||||
Status: proposed executable epic target
|
||||
Last updated: 2026-07-13
|
||||
|
||||
## Executive decision
|
||||
|
||||
The alpha-release target is the exact open-weight model `zai-org/GLM-5.2` served with `reasoning_effort=max` from the smallest published Unsloth GGUF recipe, `UD-IQ1_S`, across multiple consumer machines.
|
||||
|
||||
“Max” is a reasoning mode selected by the chat template/API request. It is not a separate model checkpoint.
|
||||
|
||||
Alpha is earned only when the real target model:
|
||||
|
||||
1. cannot fit within any one participating node's admitted memory;
|
||||
2. loads as contiguous layer Shards on at least two physical consumer machines;
|
||||
3. performs real GLM-5.2 MoE + DSA + IndexShare computation on every selected node;
|
||||
4. responds through the existing OpenAI-compatible Meshnet API with `reasoning_effort=max`;
|
||||
5. passes locked parity, usefulness, performance, telemetry, cancellation, and cleanup checks; and
|
||||
6. stores all model artifacts on configured mounted-drive storage, never under `/home`.
|
||||
|
||||
The shortest safe path is not “support every GGUF architecture.” Dense Llama remains a small structural fixture. GLM-5.2 moves onto the critical path immediately after exact recipe identity and the pinned llama.cpp boundary. Qwen expansion, 1M context certification, MTP/speculative decoding, broad concurrency, and automatic route repair are post-alpha.
|
||||
|
||||
## 1. Exact target contract
|
||||
|
||||
### 1.1 Source model
|
||||
|
||||
| Field | Locked target |
|
||||
|---|---|
|
||||
| Official repository | `zai-org/GLM-5.2` |
|
||||
| Official revision observed 2026-07-13 | `b4734de4facf877f85769a911abafc5283eab3d9` |
|
||||
| Model-weight license | MIT |
|
||||
| Official code/documentation license | Apache-2.0 |
|
||||
| Architecture | `glm_moe_dsa` / `GlmMoeDsaForCausalLM` |
|
||||
| Official architecture label | 744B total / approximately 40B active per token |
|
||||
| Exact stored checkpoint tensors | 753,329,940,480 parameters |
|
||||
| Transformer layers | 78 backbone layers plus one shared NextN/MTP layer in the artifact |
|
||||
| Layer types | first 3 dense; remaining 75 sparse MoE |
|
||||
| Routed experts | 256 |
|
||||
| Experts selected | 8 routed experts plus shared expert path |
|
||||
| Hidden width | 6,144 |
|
||||
| Attention | MLA under DSA, lightning indexer top-k 2,048 |
|
||||
| IndexShare | indexer roles are encoded by `indexer_types`; consumers reuse prior Full-layer indices |
|
||||
| Architectural maximum context | 1,048,576 tokens |
|
||||
| Alpha reasoning mode | `reasoning_effort=max` |
|
||||
|
||||
The runtime must derive these values from the pinned artifact and fail closed on contradictory metadata. Marketing names are not compatibility identity.
|
||||
|
||||
### 1.2 Alpha GGUF artifact
|
||||
|
||||
| Field | Locked target |
|
||||
|---|---|
|
||||
| GGUF repository | `unsloth/GLM-5.2-GGUF` |
|
||||
| GGUF revision observed 2026-07-13 | `abc55e72527792c6e77069c99b4cb7de16fa9f23` |
|
||||
| Quantization | `UD-IQ1_S` |
|
||||
| Files | six GGUF shards |
|
||||
| Exact published bytes | 216,715,360,960 bytes |
|
||||
| Binary GiB | 201.832 GiB |
|
||||
| Published quality indication | about 76.2% top-1 agreement with the high-precision reference on Unsloth's quantization analysis |
|
||||
| Mounted storage rule | configured mounted drive only; never `/home` |
|
||||
|
||||
Before downloading 216.7 GB, DGR-017 must generate a checked-in target manifest containing repository revisions, expected filenames, byte sizes, and resolved LFS SHA-256 values. Download is resumable and verified before route admission.
|
||||
|
||||
`UD-IQ1_M` (228,492,966,624 bytes / 212.801 GiB) is the first diagnostic fallback if `UD-IQ1_S` exposes a runtime or quality defect. It does not satisfy the explicit “lowest quantization” alpha target unless the target contract is changed by human review.
|
||||
|
||||
### 1.3 Runtime semantics required for alpha
|
||||
|
||||
Required:
|
||||
|
||||
- GGUF parsing and quantized kernels from one exact llama.cpp pin.
|
||||
- Correct GLM-5.2 MoE routing, selected experts, and shared expert.
|
||||
- Correct compressed MLA KV cache for locally owned layers.
|
||||
- Native DSA lightning indexer and sparse attention.
|
||||
- Correct IndexShare Full/Shared role execution from artifact metadata.
|
||||
- Range-owned contiguous transformer layers; each owned layer keeps all of its experts local.
|
||||
- Head-only embeddings and tail-only final norm/output head/sampling.
|
||||
- Architecture-defined activation boundary and optional DSA index sideband.
|
||||
- `reasoning_effort=max` chat-template behavior through the public API.
|
||||
- F32 seam correctness lane and a separately certified production activation dtype.
|
||||
|
||||
Not required for alpha:
|
||||
|
||||
- MTP/speculative decoding. The trailing NextN tensors may be loaded or explicitly excluded according to a certified recipe, but cannot be silently misinterpreted.
|
||||
- Full 1,048,576-token context.
|
||||
- Continuous batching beyond one target session.
|
||||
- Public-WAN tensor or expert parallel collectives.
|
||||
- Automatic mid-generation repair or KV migration.
|
||||
- Every CPU/GPU backend combination.
|
||||
|
||||
## 2. Minimum resource envelope
|
||||
|
||||
### 2.1 Weight and runtime memory
|
||||
|
||||
The smallest artifact occupies 201.832 GiB before KV, DSA indexer state, scratch buffers, backend workspaces, process memory, and the operating system. **224 GiB aggregate runtime-accessible memory is only the experimental hard-fit floor**, consistent with Unsloth's approximate 223 GB one-bit requirement. It is not a conservative operational envelope.
|
||||
|
||||
For admission, each node reserves:
|
||||
|
||||
```text
|
||||
max(20% of physically usable memory, 8 GiB)
|
||||
```
|
||||
|
||||
The remainder is the combined weight-plus-KV placement budget. Actual peak scratch is measured by backend/context and can force one extra node. Unified memory is counted once: integrated-GPU “VRAM” must not be added again to the same physical system RAM.
|
||||
|
||||
| Physical usable tier | Minimum reserve | Weight + KV placement budget | IQ1_S 16K arithmetic minimum | Operational position |
|
||||
|---:|---:|---:|---:|---|
|
||||
| 32 GiB | 8.0 GiB | 24.0 GiB | 9 nodes | use 10 if attempted; latency-heavy |
|
||||
| 48 GiB | 9.6 GiB | 38.4 GiB | 6 nodes | possible; latency-heavy |
|
||||
| 64 GiB | 12.8 GiB | 51.2 GiB | 4 nodes | hard minimum; **5 recommended** |
|
||||
| 96 GiB | 19.2 GiB | 76.8 GiB | 3 nodes | recommended |
|
||||
| 128 GiB unified/system | 25.6 GiB | 102.4 GiB | 2 nodes | arithmetic hard minimum; **3 recommended** |
|
||||
|
||||
The planner must use exact tensor byte ownership, not equal percentages. Embeddings, final head, dense versus MoE layers, shared experts, indexer tensors, quant block alignment, KV distribution, and backend workspace make equal layer counts unequal in memory.
|
||||
|
||||
Recommended first target route: **three 96/128-GiB-class physical machines** or **five 64-GiB-class machines**, on the same wired switch with mounted model storage. Four 64-GiB or two 128-GiB machines are fit probes only and qualify solely if exact placement and measured peak-memory evidence retain the required reserve with no swap/overcommit.
|
||||
|
||||
### 2.2 KV cache
|
||||
|
||||
GLM-5.2 MLA caches 576 latent/rope values per token per backbone layer. Correct DSA also caches 128-dimensional indexer keys: ideally only for the 21 Full indexer layers, while the current experimental implementation may allocate them across all 78 layers. Alpha locks **Q8_0 KV** for quality and budgets the conservative current-implementation layout.
|
||||
|
||||
| Context × concurrency | MLA-only Q8 | Optimized DSA Q8 | Conservative current-DSA Q8 | Conservative current-DSA F16 |
|
||||
|---:|---:|---:|---:|---:|
|
||||
| 16,384 × 1 | 0.73 GiB | 0.77 GiB | **0.89 GiB** | 1.68 GiB |
|
||||
| 131,072 × 1 | 5.83 GiB | 6.18 GiB | **7.12 GiB** | 13.41 GiB |
|
||||
| 1,048,576 × 1 | 46.62 GiB | 49.41 GiB | **56.98 GiB** | 107.25 GiB |
|
||||
|
||||
These are planning estimates, not admission truth. The runtime must report measured allocated/resident MLA and indexer cache by Shard. Alpha configures a 16,384-token window, Q8_0 KV, and one session. Longer contexts and lower-bit KV are separate quality/resource certification gates.
|
||||
|
||||
### 2.3 Activation seams and network
|
||||
|
||||
A BF16 hidden-state boundary is 6,144 elements = 12,288 bytes/token before framing.
|
||||
|
||||
- A 16,384-token prefill sends about 192 MiB per seam.
|
||||
- One decode token sends about 12 KiB per seam.
|
||||
- A 512-token decode sends about 6 MiB per seam.
|
||||
- Four nodes imply three serial seams.
|
||||
|
||||
If a Shard boundary splits an IndexShare producer/consumer group, a 2,048-entry int32 top-k sideband can add up to 8 KiB/query before framing. The route planner should prefer boundaries that preserve complete IndexShare ownership groups. The protocol must still support and validate the named sideband because memory fit may force an internal group split.
|
||||
|
||||
Decode bandwidth is small, but every generated token crosses all seams serially, so node count and per-hop latency dominate. Alpha requires a same-switch wired route: **2.5 GbE minimum and 10 GbE recommended**, with measured one-way/RTT, serialization, and queue latency. A 1 GbE route may be retained as fit-only evidence but is not the recommended alpha topology. Alpha records per-seam bytes, p50/p95 transfer latency, retries, and checksum failures; no speed claim is inferred from link rate alone.
|
||||
|
||||
### 2.4 Storage
|
||||
|
||||
The shortest alpha path allows every node to hold the complete six-file source GGUF while mapping/allocating only owned tensors. This minimizes artifact-transformation risk but costs 216.7 GB disk per node.
|
||||
|
||||
Deterministic source-bound layer packages are a follow-up optimization. If needed before target fit, every package must retain:
|
||||
|
||||
- source repository/revision and source file hashes;
|
||||
- exact owned tensor names, layer range, and endpoint role;
|
||||
- tokenizer/config identity;
|
||||
- deterministic package hash; and
|
||||
- proof that composing all packages matches the source tensor inventory.
|
||||
|
||||
## 3. Current state and critical gaps
|
||||
|
||||
### Completed foundation
|
||||
|
||||
- DGR-001: immutable CPU contract plus separate signed ROCm diagnostic. CPU v1 remains `stop`; the GPU diagnostic establishes a viable fit/performance investigation lane but does not rewrite CPU evidence.
|
||||
- DGR-002: versioned backend-neutral gRPC/Protobuf Shard protocol with bounded fragments and compatibility checks.
|
||||
- Existing Meshnet: Tracker, contiguous Shards, Route Sessions/epochs, relay/direct transport, local Hot KV semantics in the reference backend, cancellation, telemetry, billing, and model-agnostic admission.
|
||||
|
||||
### Missing before target alpha
|
||||
|
||||
1. Exact GLM target/artifact manifest and memory-fit planner.
|
||||
2. A current llama.cpp pin proven to load and generate with the exact `UD-IQ1_S` artifact.
|
||||
3. A narrow decision on native GLM-5.2 DSA/IndexShare support. As observed 2026-07-13, merged llama.cpp PR #24770 loads GLM-5.2 through a dense-MLA compatibility path, while full IndexShare/DSA PR #25407 remains open and its generic sparse path can be slower than dense fallback. Generic CPU lightning-indexer support is merged; backend coverage remains uneven.
|
||||
4. A decode protocol amendment. `ActivationChunk` carries `TensorBundle`, but the current `DecodeStep` fast path carries only one `NamedTensor`; it cannot transport a hidden state plus GLM top-k sideband. Tail token/logit and sampling behavior also needs an explicit typed result contract.
|
||||
5. Correct range-owned GGUF loading and memory proof.
|
||||
6. GLM-specific boundary/KV/IndexShare semantics.
|
||||
7. Standalone native worker and Meshnet integration.
|
||||
8. Real target hardware route with no node individually able to admit the whole model.
|
||||
9. Locked target parity, usefulness, speed, failure, and cleanup evidence.
|
||||
|
||||
### Donor policy
|
||||
|
||||
`Mesh-LLM/mesh-llm` is a high-value test and patch donor. Its live GLM branch was observed with 261 llama.cpp patches, 167 named for GLM/DSA/MTP-related work. That is evidence of the problem's depth, not an acceptable maintained fork boundary.
|
||||
|
||||
Audit and selectively reproduce the smallest independently understood pieces for:
|
||||
|
||||
- GLM DSA graph semantics;
|
||||
- lightning indexer and sparse-attention tests;
|
||||
- IndexShare metadata/Full/Shared validation;
|
||||
- top-k sideband shape and lifecycle;
|
||||
- stage-local KV filtering; and
|
||||
- target parity/performance fixtures.
|
||||
|
||||
Do not import Mesh-LLM routing, discovery, scheduler, public mesh, package manager, or full patch stack. Keep Meshnet as the sole control plane and collaborate narrowly upstream with llama.cpp/Mesh-LLM maintainers where practical.
|
||||
|
||||
## 4. Revised roadmap
|
||||
|
||||
## Phase 0 — lock the target before implementation
|
||||
|
||||
### DGR-017: lock GLM-5.2 Max target and alpha contract
|
||||
|
||||
Deliver:
|
||||
|
||||
- machine-readable target manifest for official and GGUF revisions;
|
||||
- exact `UD-IQ1_S` file/size/hash inventory;
|
||||
- architecture/config/chat-template snapshot;
|
||||
- memory/KV/network planner with unified-memory de-duplication;
|
||||
- immutable alpha acceptance thresholds from section 5; and
|
||||
- current upstream/donor status report.
|
||||
|
||||
Exit: target identity and alpha requirements are reviewable without downloading the model.
|
||||
|
||||
## Phase 1 — establish a correct whole-model oracle
|
||||
|
||||
### DGR-003: exact runtime recipe identity
|
||||
|
||||
Extend the existing generic identity with GLM fields: DSA/IndexShare metadata, adapter version, reasoning template revision, activation bundle schema, KV dtype/layout, llama.cpp pin/patch hash, and target artifact manifest hash.
|
||||
|
||||
### DGR-004: reproducible llama.cpp pin and narrow patch boundary
|
||||
|
||||
Select a current exact upstream commit only after testing its stock GLM behavior. Add clean fetch/apply/build checks. Record every donor patch and whether it is adopted, rewritten, rejected, or waiting upstream.
|
||||
|
||||
### DGR-018: certify whole-model GLM-5.2 runtime semantics
|
||||
|
||||
On a 256-GiB-class reference host with at least 224 GiB runtime-accessible memory after OS reservation, or a measured equivalent:
|
||||
|
||||
1. verify all six `UD-IQ1_S` shards;
|
||||
2. load with a stock pinned runtime and capture tensor/metadata warnings;
|
||||
3. prove whether DSA, IndexShare, shared expert, and Max template are actually active;
|
||||
4. add the minimum correctness patches/tests required;
|
||||
5. run deterministic prefill/decode and fixed Max-mode sentinel prompts; and
|
||||
6. sign the oracle recipe, output, telemetry, and limitations.
|
||||
|
||||
Exit: one whole-model oracle exists for the same artifact/runtime semantics the distributed path will implement. “It emits text” is insufficient.
|
||||
|
||||
## Phase 2 — build the generic local seam using small fixtures
|
||||
|
||||
### DGR-005: range-owned GGUF tensors
|
||||
|
||||
Keep dense Llama as a cheap structural fixture. Implement authoritative owned-tensor registration/loading, head/tail ownership, and measured resident-memory scaling. Design tensor classification so GLM adds explicit rules rather than unchecked name substitution.
|
||||
|
||||
### DGR-006: architecture-defined boundary
|
||||
|
||||
Implement named boundary bundles, F32 correctness lane, bounded fragmentation, and optional sidebands. Amend the decode fast path so it carries a versioned `TensorBundle` rather than one `NamedTensor`, while preserving a small one-tensor encoding. Define an explicit typed tail result for logits/token output and bind sampling/chat-template parameters to the recipe/request. Regenerate Python/C++ schema code and compatibility goldens. Dense fixture parity proves the seam mechanism, not GLM certification.
|
||||
|
||||
Exit: two local processes can execute a small dense model with correct range ownership and boundary parity.
|
||||
|
||||
## Phase 3 — add GLM-5.2 as the product adapter
|
||||
|
||||
### DGR-019: implement and certify GLM-5.2 range/DSA/IndexShare semantics
|
||||
|
||||
Deliver explicit support for:
|
||||
|
||||
- 78 main layers and endpoint tensor ownership;
|
||||
- 256-expert MoE routing/top-8 and shared expert;
|
||||
- compressed MLA KV by owned layers;
|
||||
- DSA lightning indexer and sparse attention;
|
||||
- IndexShare metadata, Full producer, Shared consumer, and sideband behavior;
|
||||
- NextN/MTP tensor policy with MTP disabled or enabled explicitly;
|
||||
- shard-boundary planner aware of IndexShare ownership groups; and
|
||||
- whole-model versus two-stage parity against DGR-018.
|
||||
|
||||
Exit: a same-host two-stage target run matches the locked oracle tolerance with real GLM computation in both stages. If the full target cannot fit on one host for this check, use a layer-reduced GLM architecture fixture for graph parity and defer full-artifact output parity to DGR-020; label the distinction explicitly.
|
||||
|
||||
## Phase 4 — worker, KV, and Meshnet route
|
||||
|
||||
Execute existing stories with GLM requirements included:
|
||||
|
||||
1. DGR-007 — isolated local Hot KV keyed by `(Route Session, epoch)`, including DSA/IndexShare state.
|
||||
2. DGR-008 — standalone C++ gRPC worker.
|
||||
3. DGR-009 — Meshnet backend, capability, relay/direct, cancellation, and telemetry integration.
|
||||
4. DGR-010 — small-model local two-process acceptance.
|
||||
5. DGR-011 — real two-physical-machine route and heterogeneous fail-closed behavior.
|
||||
6. DGR-013 subset required by alpha — node loss, cancellation, stale epoch, restart, and memory/KV cleanup.
|
||||
|
||||
Continuous batching (DGR-012) is deliberately not an alpha dependency. The first target release supports one admitted GLM route session; concurrency follows after target correctness and fit.
|
||||
|
||||
## Phase 5 — target alpha gate
|
||||
|
||||
### DGR-020: pass real distributed GLM-5.2 Max alpha acceptance
|
||||
|
||||
Use at least two physical machines and enough aggregate usable memory to meet the locked target planner. No participating node may individually admit the complete target. All stages must report real compute and exact tensor ownership.
|
||||
|
||||
Run the complete acceptance matrix in section 5, preserve raw logs/metrics/output, sign the evidence, and publish an explicit `alpha` or `stop` verdict. Thresholds cannot be weakened after results are known.
|
||||
|
||||
## Phase 6 — post-alpha hardening
|
||||
|
||||
After DGR-020 passes:
|
||||
|
||||
1. DGR-012 — 1/2/4-session continuous batching and bounded admission.
|
||||
2. DGR-014 — final distributed GGUF versus reference-route performance decision.
|
||||
3. 32K, 128K, 200K, then 1M context certification with quantized KV.
|
||||
4. MTP/speculative decoding.
|
||||
5. Deterministic range packages to remove full-artifact replication.
|
||||
6. Additional backend compatibility classes and route topologies.
|
||||
7. DGR-016 — narrow upstream collaboration package, split by independently reviewable llama.cpp changes.
|
||||
8. DGR-015 — Qwen3/Qwen3-MoE only as later architecture expansion, not as the GLM alpha target.
|
||||
|
||||
## 5. Locked alpha acceptance matrix
|
||||
|
||||
These thresholds are set before target execution.
|
||||
|
||||
### 5.1 Identity and fit
|
||||
|
||||
- Exact official and GGUF repository revisions match the target manifest.
|
||||
- All six source GGUF sizes and LFS SHA-256 values verify.
|
||||
- Every route node reports owned tensor names/bytes, layer range, endpoint role, backend, KV recipe, and patch fingerprint.
|
||||
- Union of owned tensors equals the certified runtime-required tensor inventory; unintended overlap is zero.
|
||||
- No node's weight-plus-KV placement budget can hold the complete recipe.
|
||||
- Every node reserves at least `max(20% of physically usable memory, 8 GiB)` outside weight-plus-KV placement; measured peak scratch must remain inside that reserve.
|
||||
- Aggregate peak RSS/VRAM stays within physical budgets with no swap, overcommit, mmap-only, or double-counted unified-memory success claim.
|
||||
- Arithmetic-minimum topologies require exact contiguous tensor placement evidence; recommended alpha topology is 5×64 GiB or 3×96/128 GiB.
|
||||
- Unified RAM/VRAM is not double-counted.
|
||||
|
||||
### 5.2 Semantic correctness
|
||||
|
||||
- Logs and graph tests prove GLM MoE/shared-expert, DSA lightning indexer, sparse attention, and IndexShare Full/Shared paths are active; dense-attention compatibility fallback cannot satisfy alpha.
|
||||
- `reasoning_effort=max` is observable in the rendered template/API recipe.
|
||||
- F32 same-backend seam fixture: 32 greedy decode tokens exactly match the whole-model oracle and activation tolerance is locked by DGR-006.
|
||||
- Production seam on the fixed prompt corpus: greedy token agreement is at least 0.90 and mean compared-state/logit cosine similarity is at least 0.999 versus DGR-018, with no malformed or non-finite tensors.
|
||||
- Incompatible artifact, tokenizer, adapter, DSA metadata, boundary, activation, KV, backend class, or runtime patch fingerprints fail closed.
|
||||
|
||||
### 5.3 End-to-end target run
|
||||
|
||||
- Context configured to 16,384 tokens with Q8_0 MLA/indexer KV.
|
||||
- Fixed 4,096-token prompt lane completes prefill.
|
||||
- Route uses a same-switch wired network; 2.5 GbE is the alpha minimum and 10 GbE is recommended.
|
||||
- One Max-mode request generates at least 512 output tokens or reaches a valid natural EOS after at least 128 tokens.
|
||||
- Fixed coding, structured tool-call/JSON, and multi-step reasoning sentinels produce parseable, relevant outputs; raw prompts and outputs are retained for review.
|
||||
- OpenAI-compatible response includes stable model ID, finish reason, and token usage.
|
||||
|
||||
### 5.4 Minimum useful performance
|
||||
|
||||
On the declared minimum alpha topology after one warm-up:
|
||||
|
||||
- median decode throughput is at least 0.5 generated token/s for the fixed Max-mode lane;
|
||||
- 4,096-token-prompt TTFT is at most 10 minutes;
|
||||
- no unexplained stall exceeds 60 seconds without progress telemetry;
|
||||
- per-stage compute, queue, KV, seam bytes/latency, RSS/VRAM, and backend timing are present; and
|
||||
- results are labeled by hardware/topology and are not generalized to other consumer systems.
|
||||
|
||||
If output quality passes but the speed floor fails, verdict is `stop` for alpha and the evidence selects the next optimization target. It is not relabeled as success merely because the model loaded.
|
||||
|
||||
### 5.5 Reliability and security
|
||||
|
||||
- Two consecutive cold starts load, generate, release, and exit cleanly.
|
||||
- Cancellation during prefill and decode releases every stage's queued buffers and KV lease.
|
||||
- One worker loss aborts the route; alpha retries only from token zero on a new compatible route.
|
||||
- Stale epochs and duplicate step IDs are rejected.
|
||||
- Artifact paths stay outside `/home`; logs contain no secrets or unrestricted prompt payloads.
|
||||
- Synthetic workers and layer-reduced fixtures are labeled unit/integration coverage and cannot satisfy target alpha.
|
||||
|
||||
## 6. First execution order
|
||||
|
||||
The next unattended work should run in this order:
|
||||
|
||||
1. DGR-017 — target contract, manifest, planner, and upstream status.
|
||||
2. DGR-003 — exact recipe identity.
|
||||
3. DGR-004 — current llama.cpp pin and minimal patch harness.
|
||||
4. Run in parallel:
|
||||
- DGR-018 — whole-model oracle on a 256-GiB-class host with at least 224 GiB runtime-accessible memory.
|
||||
- DGR-005 and DGR-006 — generic range/boundary seam on local small fixtures.
|
||||
5. DGR-019 — GLM semantics and parity after both parallel lanes pass.
|
||||
6. DGR-007 through DGR-011 — native worker and real transport route.
|
||||
7. Required DGR-013 failure subset.
|
||||
8. DGR-020 — real target alpha verdict.
|
||||
|
||||
The first external hardware blocker is DGR-018, but DGR-005/DGR-006 proceed locally while that host is sourced. Do not download the full model until DGR-017's exact manifest and storage preflight pass.
|
||||
|
||||
## 7. Sources checked on 2026-07-13
|
||||
|
||||
Authoritative or primary:
|
||||
|
||||
- Official model card and config: <https://huggingface.co/zai-org/GLM-5.2>
|
||||
- Official release/architecture blog: <https://z.ai/blog/glm-5.2>
|
||||
- Official code/documentation repository: <https://github.com/zai-org/GLM-5>
|
||||
- Official source revision API: <https://huggingface.co/api/models/zai-org/GLM-5.2>
|
||||
- Official GLM-5 technical report: <https://arxiv.org/abs/2602.15763>
|
||||
- Unsloth GGUF repository: <https://huggingface.co/unsloth/GLM-5.2-GGUF>
|
||||
- Unsloth local-run/quantization guide: <https://unsloth.ai/docs/models/glm-5.2>
|
||||
- llama.cpp GLM-5.2 support issue: <https://github.com/ggml-org/llama.cpp/issues/24730>
|
||||
- llama.cpp merged dense-MLA compatibility loader: <https://github.com/ggml-org/llama.cpp/pull/24770>
|
||||
- llama.cpp open GLM-5.2 DSA/IndexShare implementation: <https://github.com/ggml-org/llama.cpp/pull/25407>
|
||||
- llama.cpp merged generic CPU lightning indexer: <https://github.com/ggml-org/llama.cpp/pull/24231>
|
||||
- llama.cpp 1M-context discussion: <https://github.com/ggml-org/llama.cpp/discussions/24622>
|
||||
- IndexCache/IndexShare paper: <https://arxiv.org/abs/2603.12201>
|
||||
|
||||
Donor/current implementation evidence:
|
||||
|
||||
- Mesh-LLM repository: <https://github.com/Mesh-LLM/mesh-llm>
|
||||
- Mesh-LLM GLM branch noted by llama.cpp collaborator in issue #24730: `feat/jianyang-glm-52`
|
||||
|
||||
Web/repository observations are pinned by date and must be refreshed in DGR-017 before implementation because upstream support is moving quickly.
|
||||
@@ -0,0 +1,43 @@
|
||||
# DGR-001 downstream stop-condition handoff
|
||||
|
||||
Status: **DGR-001 is complete; native-track promotion is blocked by the immutable v1 verdict.**
|
||||
|
||||
This is no longer an execution-prerequisite blocker. The required real benchmark
|
||||
ran successfully, every recipe completed at concurrency 1 and 4, artifacts were
|
||||
verified, and deterministic/full test gates passed.
|
||||
|
||||
## Locked result
|
||||
|
||||
`contract-evaluation.json` records:
|
||||
|
||||
```text
|
||||
verdict: stop
|
||||
quality_lane_pass: false
|
||||
speed_benefit: true
|
||||
fit_benefit: true
|
||||
stop_condition_met: true
|
||||
```
|
||||
|
||||
The exact-revision BF16 GGUF quality lane compared every prompt but achieved
|
||||
`0.3333` exact match and `0.9471` mean similarity against the Transformers BF16
|
||||
reference. V1 requires `0.90` and `0.97`. Quantized Q4_K_M had substantial speed
|
||||
and fit benefits, but the contract explicitly forbids speed from redeeming a
|
||||
failed near-lossless quality lane.
|
||||
|
||||
## Scope of this stop
|
||||
|
||||
The measured baseline is Qwen2.5-0.5B on CPU using a CPU-only llama.cpp build.
|
||||
It is not a Radeon, large-model, distributed, or native-shard result. Therefore:
|
||||
|
||||
1. Do not silently mark v1 promoted or weaken its thresholds after observing the
|
||||
data.
|
||||
2. Do not let DGR-004 or later runtime stories treat DGR-001 completion as a
|
||||
positive promotion signal.
|
||||
3. A human may choose one of these explicit paths:
|
||||
- stop the native GGUF track as v1 directs;
|
||||
- diagnose and fix the BF16 runtime divergence, then rerun the exact v1 plan;
|
||||
- authorize a separately versioned GPU/large-model contract whose scope and
|
||||
workload are locked before its measurements.
|
||||
|
||||
All raw evidence, configuration, artifacts, hashes, and reproduction commands
|
||||
are in this directory and `README.md`.
|
||||
199
.scratch/distributed-gguf-runtime/evidence/DGR-001/README.md
Normal file
199
.scratch/distributed-gguf-runtime/evidence/DGR-001/README.md
Normal file
@@ -0,0 +1,199 @@
|
||||
# DGR-001 — Safetensors versus GGUF performance contract
|
||||
|
||||
Status: **complete; immutable v1 verdict is `stop`.**
|
||||
|
||||
DGR-001 successfully produced a controlled local-real CPU baseline. Completion
|
||||
means the experiment and decision contract are durable and verified; it does
|
||||
**not** mean the native GGUF track is approved to continue. The locked quality
|
||||
gate failed, so dependent runtime work requires a human decision or a new,
|
||||
explicitly versioned experiment/contract rather than silently weakening v1.
|
||||
|
||||
## Controlled workload
|
||||
|
||||
- Model: `Qwen/Qwen2.5-0.5B-Instruct`
|
||||
- Exact source revision: `7ae557604adf67be50417f59c2c2f167def9a775`
|
||||
- Machine: `fedora`, Linux `7.0.14-101.fc43.x86_64`, 32 logical CPUs
|
||||
- Device: CPU for every recipe; VRAM is therefore correctly reported as zero
|
||||
- Runtime reference: Transformers `5.13.0`, PyTorch
|
||||
`2.10.0+rocm7.13.0a20260513`, BF16 safetensors
|
||||
- GGUF runtime: llama.cpp version 9991, commit
|
||||
`e920c523e3b8a0163fe498af5bf90df35ff51d25`
|
||||
- Workload: three fixed short/medium/long prompts, greedy sampling, 32 output
|
||||
tokens, three repeats, two warmups, concurrency 1 and 4, 16 CPU threads
|
||||
- Evidence class: `local-real`
|
||||
|
||||
All artifacts are beneath `/run/media/popov/DATA/llm/`; no model artifact was
|
||||
created under `/home`.
|
||||
|
||||
## Recipes and exact artifacts
|
||||
|
||||
| Recipe | Artifact | SHA-256 |
|
||||
|---|---|---|
|
||||
| Transformers BF16 reference | complete mounted Hugging Face snapshot | `e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6` |
|
||||
| llama.cpp BF16 quality lane | `Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf` | `e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862` |
|
||||
| llama.cpp Q4_K_M performance/fit lane | `Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf` | `a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5` |
|
||||
|
||||
The snapshot digest covers every sorted relative path, resolved size, and file
|
||||
byte, so tokenizer/config drift is included. The BF16 GGUF was converted
|
||||
directly from the exact snapshot while preserving BF16 weights. Q4_K_M was
|
||||
quantized from an exact-revision F16 conversion with the pinned quantizer.
|
||||
Runtime validation recomputes every declared digest before model loading.
|
||||
|
||||
## Real results
|
||||
|
||||
All recipes completed every request with zero failures.
|
||||
|
||||
| Metric | Transformers BF16 | llama.cpp BF16 | llama.cpp Q4_K_M |
|
||||
|---|---:|---:|---:|
|
||||
| Decode tok/s, c=1 | 40.8 | 98.5 | 207.7 |
|
||||
| Aggregate decode tok/s, c=4 | 46.5 | 222.8 | 195.7 |
|
||||
| TTFT p50, c=1 | 40.0 ms | 15.1 ms | 21.6 ms |
|
||||
| Peak resident memory, c=1 | 1.94 GB | 1.11 GB | 0.54 GB |
|
||||
| Artifact size | 1.00 GB | 0.99 GB | 0.40 GB |
|
||||
| Failures | 0 | 0 | 0 |
|
||||
|
||||
Against the reference, the eligible Q4_K_M lane measured:
|
||||
|
||||
- single-request decode speedup: **5.10×**;
|
||||
- concurrency-4 aggregate throughput speedup: **4.20×**;
|
||||
- resident-memory ratio: **0.279×**;
|
||||
- artifact-size ratio: **0.398×**.
|
||||
|
||||
The near-lossless BF16 quality lane compared all three prompts but measured:
|
||||
|
||||
- exact match: **0.3333** (v1 requires at least `0.90`);
|
||||
- mean text similarity: **0.9471** (v1 requires at least `0.97`).
|
||||
|
||||
Tokenization and stopping were controlled: every runtime saw the same prompt
|
||||
token counts and reported 31 post-TTFT decode tokens. The v1 mismatch is a
|
||||
real greedy-output divergence on two prompts, not missing coverage or a
|
||||
text-length artifact. Its root cause remains undetermined; no post-contract
|
||||
logit-tie claim is acceptance evidence. Therefore `contract-evaluation.json`
|
||||
records:
|
||||
|
||||
```text
|
||||
verdict: stop
|
||||
quality_lane_pass: false
|
||||
speed_benefit: true
|
||||
fit_benefit: true
|
||||
stop_condition_met: true
|
||||
```
|
||||
|
||||
Thresholds were not changed after observing these results.
|
||||
|
||||
## Post-contract parity and ROCm diagnostics
|
||||
|
||||
`summarize-quality-parity.py` verifies and separates two signed sources. The CPU
|
||||
v1 row uses CPU kernels and a Transformers BF16 oracle; it remains at `0.3333`
|
||||
exact match with an unexplained divergence. The ROCm row uses a different plan,
|
||||
GPU kernels, and a Transformers float32 oracle. In that narrower diagnostic,
|
||||
the same BF16 GGUF artifact matches all three 32-token sequences exactly (`1.0`
|
||||
exact match and `1.0` similarity). No conversion corruption was observed in
|
||||
that three-sequence ROCm sample; this does not prove global conversion
|
||||
correctness or explain the CPU result.
|
||||
|
||||
A separate HIP build at commit `e920c523` was compiled for `gfx1151` and
|
||||
measured `ROCm0: Radeon 8060S Graphics`; its `llama-server` SHA-256 is
|
||||
`b6bb4da687dbde86e243ba006cef05919b7b97255cd7e2371e1d451220aca139`.
|
||||
A signed `gpu-diagnostic` profile measured zero failures:
|
||||
|
||||
| GPU metric | Transformers BF16 ROCm | llama.cpp Q4 ROCm | Q4 ratio |
|
||||
|---|---:|---:|---:|
|
||||
| Decode tok/s, c=1 | 81.12 | 251.25 | **3.10×** |
|
||||
| Aggregate decode tok/s, c=4 | 91.24 | 511.33 | **5.60×** |
|
||||
| TTFT p50, c=1 | 13.77 ms | 11.80 ms | **0.857×** |
|
||||
|
||||
The GPU report is signed under the distinct
|
||||
`run_configured_gpu_diagnostic/v1` producer. The v1 evaluator rejects that
|
||||
producer even when its signature is valid. llama-server process VRAM remains
|
||||
unmeasured, so this diagnostic cannot replace or satisfy the immutable v1
|
||||
contract. Its signed backend detail records the measured `ROCm0: Radeon 8060S
|
||||
Graphics` device and `25/25` offloaded layers.
|
||||
|
||||
## Implementation
|
||||
|
||||
- `recipe_benchmark.py` provides the runtime-neutral measurement core, true
|
||||
concurrency, continuous in-flight peak-memory sampling, percentile/throughput
|
||||
aggregation, failures, and output drift.
|
||||
- `recipe_drivers.py` provides opt-in Transformers and llama-server drivers,
|
||||
mounted-drive confinement, exact artifact/runtime verification, equal
|
||||
device/thread budgets, greedy-only validation, measured host provenance, a
|
||||
CPU-only v1 guard until process VRAM can be measured honestly, and a distinct
|
||||
signed GPU diagnostic profile that the v1 evaluator cannot accept.
|
||||
- Peak RSS is runtime-scoped: Transformers reports growth above its pre-runtime
|
||||
Python baseline, while llama.cpp reports its isolated server process tree.
|
||||
Both are sampled continuously during in-flight requests.
|
||||
- TTFT uses each runtime's prompt/first-token compute boundary; end-to-end HTTP,
|
||||
scheduling, and queue overhead remains in latency and `queue_wait_ms`.
|
||||
- The exact canonical plan SHA-256 locks prompts, model/revision, sampling,
|
||||
output length, repeats, warmups, and concurrency. The evaluator also requires
|
||||
equal prompt/decode token counts across recipes.
|
||||
- llama.cpp's `predicted_n` includes the first token while `predicted_ms` begins
|
||||
after it; the driver subtracts that token so decode throughput matches the
|
||||
Transformers inter-token convention.
|
||||
- `performance_contract.py` rejects wrong plans, unsigned or incorrectly signed
|
||||
real evidence, wrong config/artifact/runtime/backend/host bindings, missing
|
||||
recipes/concurrency, mixed model revisions, incomplete quality coverage, and
|
||||
failed references.
|
||||
- Every non-synthetic report is Ed25519-signed over the complete canonical JSON,
|
||||
including raw outcomes and metrics. The contract pins the public key and exact
|
||||
config SHA-256; the private key remains outside Git at mode `0600`.
|
||||
- The signer fingerprint is independently anchored outside this evidence
|
||||
directory in `../../trusted-evidence-signers.json` and checked by tests.
|
||||
- Quantized drift remains advisory. Only the near-lossless lane can satisfy the
|
||||
quality gate, and only performance-fit recipes can earn speed/fit benefits.
|
||||
|
||||
## Evidence files
|
||||
|
||||
- `performance-contract.json` — immutable v1 thresholds and stop condition
|
||||
- `benchmark-config.json` — exact real-run plan, drivers, artifacts, and hashes
|
||||
- `results.json` — raw machine-readable per-request and aggregate evidence
|
||||
- `results.txt` — human-readable benchmark summary
|
||||
- `baseline.json` — distilled measurements for later comparison
|
||||
- `contract-evaluation.json` — fail-closed v1 verdict
|
||||
- `quality-parity-diagnosis.json` / `.md` — run/device-scoped signed-evidence summary
|
||||
- `summarize-quality-parity.py` — verifies both evidence chains and regenerates it
|
||||
- `gpu-diagnostic-config.json` — exact ROCm diagnostic artifacts and runtimes
|
||||
- `gpu-diagnostic-results.json` / `.txt` — signed GPU outcomes and summary
|
||||
- `commands.txt` — reproducible conversion, benchmark, evaluation, and test commands
|
||||
- `BLOCKED.md` — downstream stop-condition handoff
|
||||
- `known-unrelated-failure.md` — clean-base reproduction of the tracker race
|
||||
- `../../trusted-evidence-signers.json` — repository-reviewed signer fingerprint
|
||||
|
||||
## Verification
|
||||
|
||||
```text
|
||||
Targeted: 28 passed (5/5 consecutive focused runs)
|
||||
Latest full suite: 755 passed, 13 skipped
|
||||
Earlier full suite: 751 passed, 13 skipped
|
||||
Current cancellation retry matrix, DGR-001: 4/5 passed
|
||||
Earlier cancellation retry matrix, clean d904c40: 4/5 passed
|
||||
compileall: passed
|
||||
git diff --check: passed
|
||||
Evidence JSON parse/integrity checks: passed
|
||||
```
|
||||
|
||||
The intermittent tracker cancellation race reproduced at the same rate on the
|
||||
clean base and is retained in `known-unrelated-failure.md`; the final full suite
|
||||
completed green. DGR-001 changes no tracker/proxy files.
|
||||
|
||||
The earlier Ralph claim that the full suite was blocked by Protobuf 6.33.6 was
|
||||
invalid: it used Hermes Agent's internal venv. Verification above used the
|
||||
project `.venv`, which has the DGR-002-compatible runtime. Real inference used
|
||||
`.venv-rocm` Python 3.12.
|
||||
|
||||
## Limitations and dependent-story handoff
|
||||
|
||||
- The immutable contract result is a **0.5B CPU baseline**. The separate Radeon
|
||||
diagnostic is real local GPU evidence, but neither result covers a large
|
||||
model, distributed execution, network transport, or a native shard worker.
|
||||
- A separate `GGML_HIP=ON` llama.cpp build exists and produced GPU timings, but
|
||||
llama-server process VRAM is not measurable by the current driver; GPU
|
||||
memory/fit claims therefore remain ineligible for v1.
|
||||
- Absolute timings are developer-machine measurements; locked ratios and raw
|
||||
artifacts are provided for reproducibility.
|
||||
- DGR-014 may consume v1 only with the exact plan/evidence requirements enforced
|
||||
by `performance_contract.py`.
|
||||
- DGR-004 and later native-runtime work must not treat DGR-001 completion as a
|
||||
promotion. V1 says `stop`; proceeding requires a human decision backed by a
|
||||
separately versioned GPU/large-model contract or a diagnosed quality fix.
|
||||
169
.scratch/distributed-gguf-runtime/evidence/DGR-001/baseline.json
Normal file
169
.scratch/distributed-gguf-runtime/evidence/DGR-001/baseline.json
Normal file
@@ -0,0 +1,169 @@
|
||||
{
|
||||
"artifact_sha256": {
|
||||
"llama-cpp-near-lossless-quality": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"llama-cpp-quantized-performance-fit": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5",
|
||||
"transformers-safetensors-reference": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6"
|
||||
},
|
||||
"backend_detail": {
|
||||
"llama-cpp-near-lossless-quality": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
|
||||
"llama-cpp-quantized-performance-fit": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
|
||||
"transformers-safetensors-reference": "torch 2.10.0+rocm7.13.0a20260513; dtype bfloat16; device cpu; intra-op threads 16"
|
||||
},
|
||||
"evidence_class": "local-real",
|
||||
"host": {
|
||||
"accelerator_name": "Radeon 8060S Graphics",
|
||||
"accelerator_runtime": "7.13.26183",
|
||||
"benchmark_lane": "cpu-controlled-baseline",
|
||||
"converter_sha256": "c819f18fb22927b49fabc3b35d1c9e21ee638b3817eccd1bd4efbcc7116eeb4d",
|
||||
"cpu_count": 32,
|
||||
"cuda_available": true,
|
||||
"hostname": "fedora",
|
||||
"llama_cpp_commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
|
||||
"llama_cpp_version": "9991",
|
||||
"llama_server_identities": {
|
||||
"/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server": {
|
||||
"sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"version": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64"
|
||||
}
|
||||
},
|
||||
"llama_server_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"platform": "Linux-7.0.14-101.fc43.x86_64-x86_64-with-glibc2.42",
|
||||
"python": "3.12.13",
|
||||
"quantizer_sha256": "bd0cc8c7be6d48aad4755b31062e0e59a887cbadd43dbb8771853d5858bb198f",
|
||||
"torch_version": "2.10.0+rocm7.13.0a20260513",
|
||||
"transformers_version": "5.13.0"
|
||||
},
|
||||
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"plan_sha256": "efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570",
|
||||
"provenance": {
|
||||
"completed_at": "2026-07-13T16:27:19.647692Z",
|
||||
"config_sha256": "00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3",
|
||||
"producer": "meshnet_node.recipe_drivers.run_configured_benchmark/v1",
|
||||
"run_id": "e4eedadf-22f6-4907-8990-985456961099",
|
||||
"schema_version": 1,
|
||||
"signature": "owev+/ToswP20C923G6E+srOCUBV5vrjmndVatr9CbTXakiFGqlHrTiEo+aymA4BcSwmG6KJTxlxO6WpLnpcAg==",
|
||||
"signature_algorithm": "ed25519",
|
||||
"signer_public_key_sha256": "8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de",
|
||||
"started_at": "2026-07-13T16:26:22.361501Z"
|
||||
},
|
||||
"recipe_runtime": {
|
||||
"llama-cpp-near-lossless-quality": {
|
||||
"device": "cpu",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "bfloat16"
|
||||
},
|
||||
"llama-cpp-quantized-performance-fit": {
|
||||
"device": "cpu",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "Q4_K_M"
|
||||
},
|
||||
"transformers-safetensors-reference": {
|
||||
"device": "cpu",
|
||||
"runtime": "transformers-5.13.0",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16"
|
||||
}
|
||||
},
|
||||
"recipes": {
|
||||
"llama-cpp-near-lossless-quality": {
|
||||
"artifact_bytes": 994156448,
|
||||
"available": true,
|
||||
"concurrency": {
|
||||
"1": {
|
||||
"aggregate_decode_tokens_per_sec": 86.7339,
|
||||
"decode_tokens_per_sec": 98.5178,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 333.023,
|
||||
"latency_p95_ms": 383.0597,
|
||||
"peak_rss_bytes": 1110728704,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 1717.9451,
|
||||
"ttft_p50_ms": 15.069,
|
||||
"ttft_p95_ms": 63.766
|
||||
},
|
||||
"4": {
|
||||
"aggregate_decode_tokens_per_sec": 222.788,
|
||||
"decode_tokens_per_sec": 76.6297,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 490.8738,
|
||||
"latency_p95_ms": 646.26,
|
||||
"peak_rss_bytes": 1139466240,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 859.8985,
|
||||
"ttft_p50_ms": 32.445,
|
||||
"ttft_p95_ms": 218.387
|
||||
}
|
||||
},
|
||||
"device": "cpu",
|
||||
"lane": "quality"
|
||||
},
|
||||
"llama-cpp-quantized-performance-fit": {
|
||||
"artifact_bytes": 397807520,
|
||||
"available": true,
|
||||
"concurrency": {
|
||||
"1": {
|
||||
"aggregate_decode_tokens_per_sec": 139.2693,
|
||||
"decode_tokens_per_sec": 207.712,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 168.3307,
|
||||
"latency_p95_ms": 305.1338,
|
||||
"peak_rss_bytes": 542081024,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 967.0195,
|
||||
"ttft_p50_ms": 21.582,
|
||||
"ttft_p95_ms": 147.859
|
||||
},
|
||||
"4": {
|
||||
"aggregate_decode_tokens_per_sec": 195.6789,
|
||||
"decode_tokens_per_sec": 76.9497,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 437.9196,
|
||||
"latency_p95_ms": 885.5355,
|
||||
"peak_rss_bytes": 573259776,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 572.4424,
|
||||
"ttft_p50_ms": 48.127,
|
||||
"ttft_p95_ms": 416.531
|
||||
}
|
||||
},
|
||||
"device": "cpu",
|
||||
"lane": "performance-fit"
|
||||
},
|
||||
"transformers-safetensors-reference": {
|
||||
"artifact_bytes": 999586347,
|
||||
"available": true,
|
||||
"concurrency": {
|
||||
"1": {
|
||||
"aggregate_decode_tokens_per_sec": 35.4722,
|
||||
"decode_tokens_per_sec": 40.7545,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 818.3864,
|
||||
"latency_p95_ms": 1258.0673,
|
||||
"peak_rss_bytes": 1941458944,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 625.6467,
|
||||
"ttft_p50_ms": 40.0018,
|
||||
"ttft_p95_ms": 195.2551
|
||||
},
|
||||
"4": {
|
||||
"aggregate_decode_tokens_per_sec": 46.5375,
|
||||
"decode_tokens_per_sec": 12.9506,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 2481.8662,
|
||||
"latency_p95_ms": 3365.8395,
|
||||
"peak_rss_bytes": 2104832000,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 264.0101,
|
||||
"ttft_p50_ms": 97.0403,
|
||||
"ttft_p95_ms": 429.0665
|
||||
}
|
||||
},
|
||||
"device": "cpu",
|
||||
"lane": "quality"
|
||||
}
|
||||
},
|
||||
"reference_recipe_id": "transformers-safetensors-reference"
|
||||
}
|
||||
@@ -0,0 +1,118 @@
|
||||
{
|
||||
"artifact_storage_root": "/run/media/popov/DATA/llm",
|
||||
"evidence_class": "local-real",
|
||||
"host": {
|
||||
"benchmark_lane": "cpu-controlled-baseline",
|
||||
"llama_cpp_commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
|
||||
"llama_cpp_version": "9991",
|
||||
"llama_server_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"converter_sha256": "c819f18fb22927b49fabc3b35d1c9e21ee638b3817eccd1bd4efbcc7116eeb4d",
|
||||
"quantizer_sha256": "bd0cc8c7be6d48aad4755b31062e0e59a887cbadd43dbb8771853d5858bb198f",
|
||||
"transformers_version": "5.13.0"
|
||||
},
|
||||
"plan": {
|
||||
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
|
||||
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"prompts": [
|
||||
{
|
||||
"id": "short-fact",
|
||||
"text": "The capital of France is",
|
||||
"context_class": "short"
|
||||
},
|
||||
{
|
||||
"id": "medium-code",
|
||||
"text": "Complete this Python function without commentary:\n\ndef fibonacci(n):\n \"\"\"Return the nth Fibonacci number for n >= 0.\"\"\"\n",
|
||||
"context_class": "medium"
|
||||
},
|
||||
{
|
||||
"id": "long-summary",
|
||||
"text": "A distributed inference service divides a transformer across consumer machines. The tracker owns admission, routing, cancellation, accounting, and telemetry, while workers own only model execution. Every request carries an immutable model identity and revision. Workers must reject incompatible protocol versions and resource demands before allocating large buffers. Activation tensors are chunked, checksummed, bounded by negotiated limits, and propagated with explicit flow-control credits. A caller may disconnect at any time, so cancellation must release queued work, in-flight transfers, and cache reservations without double billing. Retries can occur after network failures, requiring idempotent request identifiers and deterministic completion accounting. The system keeps the existing safetensors path as a correctness reference while a native GGUF path is measured. Benchmarks compare the same prompts, output lengths, sampling policy, device, and concurrency, and they separate near-lossless quality checks from quantized speed and fit claims. Summarize the design priorities in three concise bullet points.",
|
||||
"context_class": "long"
|
||||
}
|
||||
],
|
||||
"sampling": {
|
||||
"temperature": 0.0,
|
||||
"top_p": 1.0,
|
||||
"top_k": 1,
|
||||
"seed": 1234,
|
||||
"max_output_tokens": 32
|
||||
},
|
||||
"concurrency_levels": [1, 4],
|
||||
"repeats": 3,
|
||||
"warmup_requests": 2
|
||||
},
|
||||
"recipes": [
|
||||
{
|
||||
"id": "transformers-safetensors-reference",
|
||||
"runtime": "transformers-5.13.0",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16",
|
||||
"lane": "quality",
|
||||
"device": "cpu",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"artifact_sha256": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": true,
|
||||
"notes": "artifact_sha256 is the deterministic digest of every snapshot path and file byte",
|
||||
"driver": {
|
||||
"type": "transformers",
|
||||
"model_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"device": "cpu",
|
||||
"dtype": "bfloat16",
|
||||
"threads": 16
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "llama-cpp-near-lossless-quality",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "bfloat16",
|
||||
"lane": "quality",
|
||||
"device": "cpu",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
|
||||
"artifact_sha256": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": false,
|
||||
"notes": "Converted directly from the exact mounted safetensors revision while preserving BF16 weights with pinned llama.cpp",
|
||||
"driver": {
|
||||
"type": "llama-cpp-server",
|
||||
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server",
|
||||
"binary_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
|
||||
"device": "cpu",
|
||||
"threads": 16,
|
||||
"n_parallel": 4,
|
||||
"context_per_slot": 512,
|
||||
"n_gpu_layers": 0
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "llama-cpp-quantized-performance-fit",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "Q4_K_M",
|
||||
"lane": "performance-fit",
|
||||
"device": "cpu",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
|
||||
"artifact_sha256": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": false,
|
||||
"notes": "Quantized from the exact-revision F16 GGUF with pinned llama-quantize",
|
||||
"driver": {
|
||||
"type": "llama-cpp-server",
|
||||
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server",
|
||||
"binary_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
|
||||
"device": "cpu",
|
||||
"threads": 16,
|
||||
"n_parallel": 4,
|
||||
"context_per_slot": 512,
|
||||
"n_gpu_layers": 0
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,87 @@
|
||||
# Exact source snapshot (already present on mounted storage)
|
||||
SOURCE=/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775
|
||||
LLAMA=/run/media/popov/d/DEV/llamacpp/llama.cpp
|
||||
ROCM_PY=/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv-rocm/bin/python
|
||||
PROJECT_PY=/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python
|
||||
OUT=/run/media/popov/DATA/llm/dgr-001
|
||||
SIGNING_KEY=/home/popov/.config/neuron-tai/keys/dgr-001-evidence-ed25519.pem
|
||||
|
||||
# Private signing key is outside Git and must remain owner-only
|
||||
stat -c '%a %n' "$SIGNING_KEY" # expected: 600
|
||||
|
||||
# Converter support check (no writes)
|
||||
$ROCM_PY $LLAMA/convert_hf_to_gguf.py "$SOURCE" --outtype f16 --outfile "$OUT/Qwen2.5-0.5B-Instruct-7ae5576-F16.gguf" --dry-run
|
||||
|
||||
# Exact-revision near-lossless and performance-fit artifacts
|
||||
$ROCM_PY $LLAMA/convert_hf_to_gguf.py "$SOURCE" --outtype f16 --outfile "$OUT/Qwen2.5-0.5B-Instruct-7ae5576-F16.gguf"
|
||||
$LLAMA/build/bin/llama-quantize "$OUT/Qwen2.5-0.5B-Instruct-7ae5576-F16.gguf" "$OUT/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf" Q4_K_M
|
||||
$ROCM_PY $LLAMA/convert_hf_to_gguf.py "$SOURCE" --outtype bf16 --outfile "$OUT/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf"
|
||||
|
||||
# Runtime and artifact identity
|
||||
git -C "$LLAMA" rev-parse HEAD
|
||||
$LLAMA/build/bin/llama-server --version
|
||||
sha256sum "$LLAMA/build/bin/llama-server" "$LLAMA/convert_hf_to_gguf.py" "$LLAMA/build/bin/llama-quantize"
|
||||
sha256sum "$SOURCE/model.safetensors" "$OUT/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf" "$OUT/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf"
|
||||
|
||||
# Deterministic complete-snapshot digest used by benchmark-config.json
|
||||
PYTHONPATH=packages/node $ROCM_PY - <<'PY'
|
||||
from pathlib import Path
|
||||
from meshnet_node.recipe_drivers import _artifact_sha256
|
||||
print(_artifact_sha256(Path('/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775')))
|
||||
PY
|
||||
|
||||
# Canonical opt-in local-real benchmark
|
||||
MESHNET_ENABLE_REAL_INFERENCE_TESTS=1 MESHNET_EVIDENCE_SIGNING_KEY="$SIGNING_KEY" \
|
||||
PYTHONPATH=packages/node $ROCM_PY -m meshnet_node.recipe_benchmark \
|
||||
--config .scratch/distributed-gguf-runtime/evidence/DGR-001/benchmark-config.json \
|
||||
--json-out .scratch/distributed-gguf-runtime/evidence/DGR-001/results.json \
|
||||
--summary-out .scratch/distributed-gguf-runtime/evidence/DGR-001/results.txt
|
||||
|
||||
# Distil the baseline and evaluate immutable v1
|
||||
PYTHONPATH=packages/node $PROJECT_PY - <<'PY'
|
||||
from pathlib import Path
|
||||
import json
|
||||
from meshnet_node.performance_contract import baseline_from_report, evaluate_contract, load_contract
|
||||
root = Path('.scratch/distributed-gguf-runtime/evidence/DGR-001')
|
||||
report = json.loads((root / 'results.json').read_text())
|
||||
contract = load_contract(root / 'performance-contract.json')
|
||||
(root / 'baseline.json').write_text(json.dumps(baseline_from_report(report), indent=2, sort_keys=True) + '\n')
|
||||
(root / 'contract-evaluation.json').write_text(json.dumps(evaluate_contract(contract, report).to_dict(), indent=2, sort_keys=True) + '\n')
|
||||
PY
|
||||
|
||||
# Optional ROCm GPU diagnostic (not eligible for immutable v1)
|
||||
# The version-matched rocm[devel] wheel expands beyond 20 GB; ensure sufficient
|
||||
# space or relocate its packaged payload before installation.
|
||||
uv pip install --python "$ROCM_PY" --prerelease=allow \
|
||||
--index-url https://rocm.nightlies.amd.com/v2/gfx1151/ \
|
||||
'rocm[devel]==7.13.0a20260513'
|
||||
ROCM_VENV=/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv-rocm
|
||||
ROCM_SDK="$ROCM_VENV/bin/rocm-sdk"
|
||||
ROCM_ROOT="$($ROCM_SDK path --root)"
|
||||
ROCM_BIN="$($ROCM_SDK path --bin)"
|
||||
export PATH="$ROCM_VENV/bin:$ROCM_BIN:$PATH"
|
||||
export ROCM_PATH="$ROCM_ROOT" HIP_PATH="$ROCM_ROOT"
|
||||
export CMAKE_PREFIX_PATH="$($ROCM_SDK path --cmake):$ROCM_ROOT"
|
||||
export LD_LIBRARY_PATH="$ROCM_ROOT/lib:$ROCM_ROOT/lib64:${LD_LIBRARY_PATH:-}"
|
||||
$ROCM_VENV/bin/cmake -S /run/media/popov/d/DEV/llamacpp/llama.cpp \
|
||||
-B /run/media/popov/d/DEV/llamacpp/llama.cpp/build-hip -G Ninja \
|
||||
-DGGML_HIP=ON -DGPU_TARGETS=gfx1151 \
|
||||
-DCMAKE_HIP_COMPILER="$ROCM_VENV/bin/amdclang++" \
|
||||
-DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_TESTS=OFF \
|
||||
-DLLAMA_BUILD_EXAMPLES=ON -DLLAMA_BUILD_SERVER=ON
|
||||
$ROCM_VENV/bin/cmake --build /run/media/popov/d/DEV/llamacpp/llama.cpp/build-hip \
|
||||
--target llama-server llama-cli llama-bench -j 16
|
||||
MESHNET_ENABLE_REAL_INFERENCE_TESTS=1 MESHNET_EVIDENCE_SIGNING_KEY="$SIGNING_KEY" \
|
||||
PYTHONPATH=packages/node $ROCM_PY -m meshnet_node.recipe_benchmark \
|
||||
--profile gpu-diagnostic \
|
||||
--config .scratch/distributed-gguf-runtime/evidence/DGR-001/gpu-diagnostic-config.json \
|
||||
--json-out .scratch/distributed-gguf-runtime/evidence/DGR-001/gpu-diagnostic-results.json \
|
||||
--summary-out .scratch/distributed-gguf-runtime/evidence/DGR-001/gpu-diagnostic-results.txt
|
||||
PYTHONPATH=packages/node $PROJECT_PY \
|
||||
.scratch/distributed-gguf-runtime/evidence/DGR-001/summarize-quality-parity.py
|
||||
|
||||
# Deterministic verification
|
||||
PYTHONPATH=packages/node $PROJECT_PY -m pytest -q tests/test_recipe_benchmark.py
|
||||
PYTHONPATH=packages/node $PROJECT_PY -m pytest -q
|
||||
PYTHONPATH=packages/node $PROJECT_PY -m compileall -q packages tests
|
||||
git diff --check
|
||||
@@ -0,0 +1,71 @@
|
||||
{
|
||||
"contract_version": 1,
|
||||
"fit_benefit": true,
|
||||
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
|
||||
"quality_lane_pass": false,
|
||||
"rationale": [
|
||||
"the near-lossless quality lane failed: the GGUF runtime disagrees with the safetensors reference beyond what near-lossless weights can explain",
|
||||
"a meaningful speed benefit was measured",
|
||||
"a meaningful fit benefit was measured"
|
||||
],
|
||||
"recipes": [
|
||||
{
|
||||
"comparable": true,
|
||||
"failures": 0,
|
||||
"fit_benefit": false,
|
||||
"incomparable_reason": "",
|
||||
"lane": "quality",
|
||||
"measurements": {
|
||||
"aggregate_concurrency": 4,
|
||||
"aggregate_throughput_speedup": 4.7873,
|
||||
"artifact_size_ratio": 0.9946,
|
||||
"artifact_size_win": false,
|
||||
"compared_prompts": 3,
|
||||
"decode_speedup": 2.4173,
|
||||
"exact_match_rate": 0.3333,
|
||||
"expected_prompts": 3,
|
||||
"failure_rate": 0.0,
|
||||
"mean_similarity": 0.9471,
|
||||
"resident_memory_ratio": 0.5721,
|
||||
"ttft_ratio": 0.3767
|
||||
},
|
||||
"quality_pass": false,
|
||||
"reasons": [
|
||||
"single-request decode 2.42x reference (>= 1.25x) at TTFT ratio 0.38",
|
||||
"aggregate throughput at concurrency 4 is 4.79x reference (>= 1.25x)",
|
||||
"peak resident memory is 0.57x reference (<= 0.75x)",
|
||||
"quality lane exact-match 0.33 / similarity 0.947 versus the reference (fail)"
|
||||
],
|
||||
"recipe_id": "llama-cpp-near-lossless-quality",
|
||||
"speed_benefit": false
|
||||
},
|
||||
{
|
||||
"comparable": true,
|
||||
"failures": 0,
|
||||
"fit_benefit": true,
|
||||
"incomparable_reason": "",
|
||||
"lane": "performance-fit",
|
||||
"measurements": {
|
||||
"aggregate_concurrency": 4,
|
||||
"aggregate_throughput_speedup": 4.2048,
|
||||
"artifact_size_ratio": 0.398,
|
||||
"artifact_size_win": true,
|
||||
"decode_speedup": 5.0967,
|
||||
"failure_rate": 0.0,
|
||||
"resident_memory_ratio": 0.2792,
|
||||
"ttft_ratio": 0.5395
|
||||
},
|
||||
"quality_pass": null,
|
||||
"reasons": [
|
||||
"single-request decode 5.10x reference (>= 1.25x) at TTFT ratio 0.54",
|
||||
"aggregate throughput at concurrency 4 is 4.20x reference (>= 1.25x)",
|
||||
"peak resident memory is 0.28x reference (<= 0.75x)"
|
||||
],
|
||||
"recipe_id": "llama-cpp-quantized-performance-fit",
|
||||
"speed_benefit": true
|
||||
}
|
||||
],
|
||||
"speed_benefit": true,
|
||||
"stop_condition_met": true,
|
||||
"verdict": "stop"
|
||||
}
|
||||
@@ -0,0 +1,143 @@
|
||||
{
|
||||
"artifact_storage_root": "/run/media/popov/DATA/llm",
|
||||
"evidence_class": "local-real",
|
||||
"host": {
|
||||
"benchmark_lane": "rocm-gpu-diagnostic",
|
||||
"llama_cpp_commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
|
||||
"llama_cpp_version": "9991",
|
||||
"llama_server_sha256": "b6bb4da687dbde86e243ba006cef05919b7b97255cd7e2371e1d451220aca139",
|
||||
"converter_sha256": "c819f18fb22927b49fabc3b35d1c9e21ee638b3817eccd1bd4efbcc7116eeb4d",
|
||||
"quantizer_sha256": "bd0cc8c7be6d48aad4755b31062e0e59a887cbadd43dbb8771853d5858bb198f",
|
||||
"transformers_version": "5.13.0",
|
||||
"rocm_target": "gfx1151"
|
||||
},
|
||||
"plan": {
|
||||
"plan_id": "dgr-001-rocm-gpu-diagnostic-v1",
|
||||
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"prompts": [
|
||||
{
|
||||
"id": "short-fact",
|
||||
"text": "The capital of France is",
|
||||
"context_class": "short"
|
||||
},
|
||||
{
|
||||
"id": "medium-code",
|
||||
"text": "Complete this Python function without commentary:\n\ndef fibonacci(n):\n \"\"\"Return the nth Fibonacci number for n >= 0.\"\"\"\n",
|
||||
"context_class": "medium"
|
||||
},
|
||||
{
|
||||
"id": "long-summary",
|
||||
"text": "A distributed inference service divides a transformer across consumer machines. The tracker owns admission, routing, cancellation, accounting, and telemetry, while workers own only model execution. Every request carries an immutable model identity and revision. Workers must reject incompatible protocol versions and resource demands before allocating large buffers. Activation tensors are chunked, checksummed, bounded by negotiated limits, and propagated with explicit flow-control credits. A caller may disconnect at any time, so cancellation must release queued work, in-flight transfers, and cache reservations without double billing. Retries can occur after network failures, requiring idempotent request identifiers and deterministic completion accounting. The system keeps the existing safetensors path as a correctness reference while a native GGUF path is measured. Benchmarks compare the same prompts, output lengths, sampling policy, device, and concurrency, and they separate near-lossless quality checks from quantized speed and fit claims. Summarize the design priorities in three concise bullet points.",
|
||||
"context_class": "long"
|
||||
}
|
||||
],
|
||||
"sampling": {
|
||||
"temperature": 0.0,
|
||||
"top_p": 1.0,
|
||||
"top_k": 1,
|
||||
"seed": 1234,
|
||||
"max_output_tokens": 32
|
||||
},
|
||||
"concurrency_levels": [
|
||||
1,
|
||||
4
|
||||
],
|
||||
"repeats": 3,
|
||||
"warmup_requests": 2
|
||||
},
|
||||
"recipes": [
|
||||
{
|
||||
"id": "transformers-fp32-rocm-quality-oracle",
|
||||
"runtime": "transformers-5.13.0-rocm-float32",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16-weights-float32-accumulation",
|
||||
"lane": "quality",
|
||||
"device": "cuda",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"artifact_sha256": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": true,
|
||||
"notes": "artifact_sha256 is the deterministic digest of every snapshot path and file byte",
|
||||
"driver": {
|
||||
"type": "transformers",
|
||||
"model_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"device": "cuda",
|
||||
"dtype": "float32",
|
||||
"threads": 16
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "llama-cpp-bf16-rocm-quality",
|
||||
"runtime": "llama.cpp-9991-e920c523-rocm-gfx1151",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "bfloat16",
|
||||
"lane": "quality",
|
||||
"device": "cuda",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
|
||||
"artifact_sha256": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": false,
|
||||
"notes": "Converted directly from the exact mounted safetensors revision while preserving BF16 weights with pinned llama.cpp",
|
||||
"driver": {
|
||||
"type": "llama-cpp-server",
|
||||
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build-hip/bin/llama-server",
|
||||
"binary_sha256": "b6bb4da687dbde86e243ba006cef05919b7b97255cd7e2371e1d451220aca139",
|
||||
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
|
||||
"device": "cuda",
|
||||
"threads": 16,
|
||||
"n_parallel": 4,
|
||||
"context_per_slot": 512,
|
||||
"n_gpu_layers": 99
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "transformers-bf16-rocm-throughput",
|
||||
"runtime": "transformers-5.13.0-rocm-bfloat16",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16",
|
||||
"lane": "performance-fit",
|
||||
"device": "cuda",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"artifact_sha256": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": false,
|
||||
"notes": "artifact_sha256 is the deterministic digest of every snapshot path and file byte",
|
||||
"driver": {
|
||||
"type": "transformers",
|
||||
"model_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"device": "cuda",
|
||||
"dtype": "bfloat16",
|
||||
"threads": 16
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "llama-cpp-q4-rocm-throughput",
|
||||
"runtime": "llama.cpp-9991-e920c523-rocm-gfx1151",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "Q4_K_M",
|
||||
"lane": "performance-fit",
|
||||
"device": "cuda",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
|
||||
"artifact_sha256": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": false,
|
||||
"notes": "Quantized from the exact-revision F16 GGUF with pinned llama-quantize",
|
||||
"driver": {
|
||||
"type": "llama-cpp-server",
|
||||
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build-hip/bin/llama-server",
|
||||
"binary_sha256": "b6bb4da687dbde86e243ba006cef05919b7b97255cd7e2371e1d451220aca139",
|
||||
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
|
||||
"device": "cuda",
|
||||
"threads": 16,
|
||||
"n_parallel": 4,
|
||||
"context_per_slot": 512,
|
||||
"n_gpu_layers": 99
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,13 @@
|
||||
Recipe benchmark dgr-001-rocm-gpu-diagnostic-v1 (local-real)
|
||||
model Qwen/Qwen2.5-0.5B-Instruct@7ae557604adf67be50417f59c2c2f167def9a775
|
||||
transformers-fp32-rocm-quality-oracle [quality ] c= 1 ttft p50/p95 11.0/ 35.5 ms; prefill 5746.7 tok/s; decode 35.7 tok/s; aggregate 29.6 tok/s; rss 1.39 GB; vram 2.26 GB; artifact 1.00 GB; failures 0
|
||||
transformers-fp32-rocm-quality-oracle [quality ] c= 4 ttft p50/p95 27.5/ 80.4 ms; prefill 1985.4 tok/s; decode 9.4 tok/s; aggregate 35.4 tok/s; rss 1.39 GB; vram 2.74 GB; artifact 1.00 GB; failures 0
|
||||
llama-cpp-bf16-rocm-quality [quality ] c= 1 ttft p50/p95 13.2/ 83.4 ms; prefill 4154.4 tok/s; decode 148.0 tok/s; aggregate 127.4 tok/s; rss 0.84 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
|
||||
llama-cpp-bf16-rocm-quality [quality ] c= 4 ttft p50/p95 25.1/ 52.1 ms; prefill 2205.4 tok/s; decode 115.1 tok/s; aggregate 337.1 tok/s; rss 0.86 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
|
||||
transformers-bf16-rocm-throughput [performance-fit ] c= 1 ttft p50/p95 13.8/ 22.2 ms; prefill 4787.3 tok/s; decode 81.1 tok/s; aggregate 73.5 tok/s; rss 0.07 GB; vram 2.74 GB; artifact 1.00 GB; failures 0
|
||||
transformers-bf16-rocm-throughput [performance-fit ] c= 4 ttft p50/p95 29.7/ 58.5 ms; prefill 2666.5 tok/s; decode 24.4 tok/s; aggregate 91.2 tok/s; rss 0.07 GB; vram 2.74 GB; artifact 1.00 GB; failures 0
|
||||
llama-cpp-q4-rocm-throughput [performance-fit ] c= 1 ttft p50/p95 11.8/ 37.1 ms; prefill 4219.3 tok/s; decode 251.2 tok/s; aggregate 200.1 tok/s; rss 0.69 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
|
||||
llama-cpp-q4-rocm-throughput [performance-fit ] c= 4 ttft p50/p95 21.4/ 101.0 ms; prefill 2126.9 tok/s; decode 189.7 tok/s; aggregate 511.3 tok/s; rss 0.72 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
|
||||
drift llama-cpp-bf16-rocm-quality vs transformers-fp32-rocm-quality-oracle exact 1.00; similarity 1.000 (gated)
|
||||
drift transformers-bf16-rocm-throughput vs transformers-fp32-rocm-quality-oracle exact 0.33; similarity 0.946 (advisory)
|
||||
drift llama-cpp-q4-rocm-throughput vs transformers-fp32-rocm-quality-oracle exact 0.00; similarity 0.628 (advisory)
|
||||
@@ -0,0 +1,55 @@
|
||||
# Observed pre-existing intermittent tracker race
|
||||
|
||||
This file records an unrelated timing observation and its repeated reproduction;
|
||||
it is **not** a DGR-001 benchmark/contract failure.
|
||||
|
||||
Test:
|
||||
|
||||
```text
|
||||
tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
```
|
||||
|
||||
One earlier full-suite run produced:
|
||||
|
||||
```text
|
||||
1 failed, 745 passed, 13 skipped
|
||||
```
|
||||
|
||||
A five-run isolated retry matrix reproduced the same rate repeatedly:
|
||||
|
||||
```text
|
||||
current DGR-001 branch: 4/5 passed, 1/5 failed
|
||||
clean d904c40: 4/5 passed, 1/5 failed
|
||||
```
|
||||
|
||||
An earlier full-suite run on the signed-provenance DGR-001 state completed
|
||||
green:
|
||||
|
||||
```text
|
||||
751 passed, 13 skipped
|
||||
```
|
||||
|
||||
Two full-suite runs after adding the isolated GPU diagnostic profile each hit
|
||||
the same race and otherwise passed:
|
||||
|
||||
```text
|
||||
1 failed, 750 passed, 13 skipped
|
||||
```
|
||||
|
||||
The latest expanded hardening suite hit the same race and otherwise passed:
|
||||
|
||||
```text
|
||||
1 failed, 754 passed, 13 skipped
|
||||
```
|
||||
|
||||
The final hardened state subsequently completed a full green run:
|
||||
|
||||
```text
|
||||
755 passed, 13 skipped
|
||||
```
|
||||
|
||||
In each failure, the mock upstream's three-second release timeout completed the
|
||||
stream before the cancel POST, so the request was already absent and the cancel
|
||||
endpoint returned 404. No tracker/proxy file changed in DGR-001. The race is
|
||||
therefore timing-sensitive, pre-existing, and unrelated to the benchmark,
|
||||
provenance, or GPU-diagnostic code.
|
||||
@@ -0,0 +1,87 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"contract_version": 1,
|
||||
"locked_at": "2026-07-13T00:00:00Z",
|
||||
"locked_by": "DGR-001",
|
||||
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
|
||||
"thresholds": {
|
||||
"min_decode_speedup": 1.25,
|
||||
"max_ttft_ratio": 1.25,
|
||||
"min_aggregate_throughput_speedup": 1.25,
|
||||
"max_resident_memory_ratio": 0.75,
|
||||
"max_artifact_size_ratio": 0.6,
|
||||
"min_quality_exact_match_rate": 0.9,
|
||||
"min_quality_mean_similarity": 0.97,
|
||||
"max_failure_rate": 0.0
|
||||
},
|
||||
"baseline": {
|
||||
"status": "pending-real-evidence",
|
||||
"required_evidence_class": "local-real",
|
||||
"required_recipes": [
|
||||
"transformers-safetensors-reference",
|
||||
"llama-cpp-near-lossless-quality",
|
||||
"llama-cpp-quantized-performance-fit"
|
||||
],
|
||||
"required_concurrency_levels": [
|
||||
1,
|
||||
4
|
||||
],
|
||||
"required_controlled_variables": [
|
||||
"model architecture",
|
||||
"model revision",
|
||||
"machine and device",
|
||||
"formatted prompts and context lengths",
|
||||
"output length and greedy sampling policy"
|
||||
],
|
||||
"required_plan_sha256": "efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570",
|
||||
"minimum_prompt_count": 3,
|
||||
"minimum_repeats": 3,
|
||||
"minimum_output_tokens": 32,
|
||||
"required_device": "cpu",
|
||||
"required_config_sha256": "00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3",
|
||||
"required_signer_public_key": "zQ/qRMwF/ydazzaxEI24Xvnrl5bZxzw16JYpP0bfRuI=",
|
||||
"required_artifact_sha256": {
|
||||
"transformers-safetensors-reference": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
|
||||
"llama-cpp-near-lossless-quality": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"llama-cpp-quantized-performance-fit": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5"
|
||||
},
|
||||
"required_recipe_runtime": {
|
||||
"transformers-safetensors-reference": {
|
||||
"runtime": "transformers-5.13.0",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16",
|
||||
"device": "cpu"
|
||||
},
|
||||
"llama-cpp-near-lossless-quality": {
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "bfloat16",
|
||||
"device": "cpu"
|
||||
},
|
||||
"llama-cpp-quantized-performance-fit": {
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "Q4_K_M",
|
||||
"device": "cpu"
|
||||
}
|
||||
},
|
||||
"required_backend_detail": {
|
||||
"transformers-safetensors-reference": "torch 2.10.0+rocm7.13.0a20260513; dtype bfloat16; device cpu; intra-op threads 16",
|
||||
"llama-cpp-near-lossless-quality": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
|
||||
"llama-cpp-quantized-performance-fit": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0"
|
||||
},
|
||||
"required_host_identity": {
|
||||
"python": "3.12.13",
|
||||
"torch_version": "2.10.0+rocm7.13.0a20260513",
|
||||
"transformers_version": "5.13.0",
|
||||
"llama_server_identities": {
|
||||
"/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server": {
|
||||
"sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"version": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"stop_condition": "Stop the native llama.cpp/GGUF track when, on the same machine and device as the Transformers/safetensors reference and under this plan, no performance-fit GGUF recipe delivers either a meaningful speed benefit (>=25% higher single-request decode tokens/sec without a >25% worse TTFT, or >=25% higher aggregate throughput under concurrency) or a meaningful fit benefit (>=25% lower peak resident memory), or when the near-lossless quality lane fails, which indicates a broken runtime rather than a quantization trade-off.",
|
||||
"notes": "Quantized performance-fit output drift is reported as advisory only. It is not numerical-equivalence evidence. DGR-014 consumes this immutable v1 contract. Non-synthetic evidence must be Ed25519-signed by the pinned key and match the exact locked config, artifacts, runtimes, backends, and host runtime identity."
|
||||
}
|
||||
@@ -0,0 +1,47 @@
|
||||
{
|
||||
"conclusion": {
|
||||
"conversion_corruption_observed_in_rocm_sample": false,
|
||||
"cpu_bf16_divergence_explained": false,
|
||||
"recommended_v2_design": "Predeclare a float32 quality oracle separately from the BF16 performance reference, with a larger prompt corpus and immutable thresholds.",
|
||||
"scope": "The ROCm diagnostic establishes only that the same BF16 GGUF artifact matched the float32 oracle for three GPU sequences; it does not explain the CPU BF16 divergence or prove global conversion correctness.",
|
||||
"v1_verdict_changed": false
|
||||
},
|
||||
"cpu_v1": {
|
||||
"candidate": "llama.cpp BF16 GGUF",
|
||||
"candidate_artifact_sha256": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"config_sha256": "00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3",
|
||||
"contract_verdict": "stop",
|
||||
"device": "cpu",
|
||||
"exact_match_rate": 0.3333,
|
||||
"mean_similarity": 0.9471,
|
||||
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
|
||||
"plan_sha256": "efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570",
|
||||
"quality_oracle": "Transformers BF16 safetensors",
|
||||
"report": "results.json",
|
||||
"report_sha256": "5d99a58806f39821c9206728047b8c5d605027d8a41b88639089b2418da890b5",
|
||||
"root_cause": "undetermined; no logit-tie claim is acceptance evidence",
|
||||
"run_id": "e4eedadf-22f6-4907-8990-985456961099"
|
||||
},
|
||||
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"rocm_diagnostic": {
|
||||
"candidate": "llama.cpp BF16 GGUF",
|
||||
"candidate_artifact_sha256": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"config_sha256": "b0f0c846c818f1307d034cee1f81daa311efc20985c32a4cdbbbd8ffe4153892",
|
||||
"device": "cuda (ROCm)",
|
||||
"exact_match_rate": 1.0,
|
||||
"failures": 0,
|
||||
"mean_similarity": 1.0,
|
||||
"measured_backend_detail": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 b6bb4da687dbde86e243ba006cef05919b7b97255cd7e2371e1d451220aca139; threads 16; parallel slots 4; ctx/slot 512; requested gpu layers 99; measured accelerator ROCm0: Radeon 8060S Graphics; measured offload 25/25 layers",
|
||||
"plan_id": "dgr-001-rocm-gpu-diagnostic-v1",
|
||||
"plan_sha256": "dae8e40963588f71f5d201fd163d39bd762e392544b5603d483e90d21abee2e8",
|
||||
"producer": "meshnet_node.recipe_drivers.run_configured_gpu_diagnostic/v1",
|
||||
"quality_oracle": "Transformers float32 safetensors",
|
||||
"report": "gpu-diagnostic-results.json",
|
||||
"report_sha256": "527b33d03627d57d60b30331e6b9119f579a828d6f6acb5c74ca25bab0af5f3d",
|
||||
"run_id": "31bf44e7-ccd4-4277-84ac-c775dee65411",
|
||||
"signer_fingerprint": "8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de",
|
||||
"v1_eligible": false
|
||||
},
|
||||
"schema_version": 2
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
# DGR-001 quality-parity evidence summary
|
||||
|
||||
This summary is generated by `summarize-quality-parity.py` from signed reports.
|
||||
It contains no independent logit measurements or self-asserted verification flag.
|
||||
|
||||
| Source | Device | Quality oracle | BF16 GGUF candidate | Exact | Similarity | Status |
|
||||
|---|---|---|---|---:|---:|---|
|
||||
| CPU v1 (`e4eedadf-22f6-4907-8990-985456961099`) | CPU | Transformers BF16 | llama.cpp BF16 | 0.3333 | 0.9471 | immutable `stop` |
|
||||
| ROCm diagnostic (`31bf44e7-ccd4-4277-84ac-c775dee65411`) | ROCm0 / Radeon 8060S | Transformers float32 | llama.cpp BF16 | 1.0000 | 1.0000 | diagnostic only |
|
||||
|
||||
## Interpretation
|
||||
|
||||
The CPU and ROCm rows use different plans, devices, kernels, and quality oracles.
|
||||
The CPU BF16 divergence remains unexplained and v1 remains `stop`. The signed
|
||||
ROCm report establishes the narrower fact that the same BF16 GGUF artifact
|
||||
matched the float32 oracle for all three GPU sequences with zero failures.
|
||||
Its signed backend detail records `ROCm0: Radeon 8060S Graphics` and measured
|
||||
`25/25` layer offload.
|
||||
|
||||
No conversion corruption was observed in that three-sequence ROCm sample. This
|
||||
does not prove global conversion correctness and does not retroactively change
|
||||
or explain the CPU result. A future v2 should predeclare a float32 quality oracle
|
||||
separately from its BF16 performance reference and use a larger corpus.
|
||||
|
||||
## Reproduction and bindings
|
||||
|
||||
- CPU report SHA-256: `5d99a58806f39821c9206728047b8c5d605027d8a41b88639089b2418da890b5`
|
||||
- GPU report SHA-256: `527b33d03627d57d60b30331e6b9119f579a828d6f6acb5c74ca25bab0af5f3d`
|
||||
- BF16 GGUF SHA-256: `e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862`
|
||||
- Signer fingerprint: `8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de`
|
||||
- Exact verification command: see `commands.txt`.
|
||||
2491
.scratch/distributed-gguf-runtime/evidence/DGR-001/results.json
Normal file
2491
.scratch/distributed-gguf-runtime/evidence/DGR-001/results.json
Normal file
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,10 @@
|
||||
Recipe benchmark dgr-001-controlled-whole-model-baseline-v1 (local-real)
|
||||
model Qwen/Qwen2.5-0.5B-Instruct@7ae557604adf67be50417f59c2c2f167def9a775
|
||||
transformers-safetensors-reference [quality ] c= 1 ttft p50/p95 40.0/ 195.3 ms; prefill 625.6 tok/s; decode 40.8 tok/s; aggregate 35.5 tok/s; rss 1.94 GB; vram 0.00 GB; artifact 1.00 GB; failures 0
|
||||
transformers-safetensors-reference [quality ] c= 4 ttft p50/p95 97.0/ 429.1 ms; prefill 264.0 tok/s; decode 13.0 tok/s; aggregate 46.5 tok/s; rss 2.10 GB; vram 0.00 GB; artifact 1.00 GB; failures 0
|
||||
llama-cpp-near-lossless-quality [quality ] c= 1 ttft p50/p95 15.1/ 63.8 ms; prefill 1717.9 tok/s; decode 98.5 tok/s; aggregate 86.7 tok/s; rss 1.11 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
|
||||
llama-cpp-near-lossless-quality [quality ] c= 4 ttft p50/p95 32.4/ 218.4 ms; prefill 859.9 tok/s; decode 76.6 tok/s; aggregate 222.8 tok/s; rss 1.14 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
|
||||
llama-cpp-quantized-performance-fit [performance-fit ] c= 1 ttft p50/p95 21.6/ 147.9 ms; prefill 967.0 tok/s; decode 207.7 tok/s; aggregate 139.3 tok/s; rss 0.54 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
|
||||
llama-cpp-quantized-performance-fit [performance-fit ] c= 4 ttft p50/p95 48.1/ 416.5 ms; prefill 572.4 tok/s; decode 76.9 tok/s; aggregate 195.7 tok/s; rss 0.57 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
|
||||
drift llama-cpp-near-lossless-quality vs transformers-safetensors-reference exact 0.33; similarity 0.947 (gated)
|
||||
drift llama-cpp-quantized-performance-fit vs transformers-safetensors-reference exact 0.00; similarity 0.456 (advisory)
|
||||
@@ -0,0 +1,261 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Build the DGR-001 parity summary from cryptographically verified reports."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import hashlib
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
|
||||
|
||||
from meshnet_node.performance_contract import (
|
||||
_canonical_sha256,
|
||||
evaluate_contract,
|
||||
load_contract,
|
||||
report_signing_payload,
|
||||
)
|
||||
|
||||
ROOT = Path(__file__).resolve().parent
|
||||
|
||||
|
||||
def _read(name: str) -> dict:
|
||||
return json.loads((ROOT / name).read_text(encoding="utf-8"))
|
||||
|
||||
|
||||
def _file_sha256(name: str) -> str:
|
||||
return hashlib.sha256((ROOT / name).read_bytes()).hexdigest()
|
||||
|
||||
|
||||
def _drift(report: dict, recipe_id: str) -> dict:
|
||||
return next(item for item in report["drift"] if item["recipe_id"] == recipe_id)
|
||||
|
||||
|
||||
def _recipe(report: dict, recipe_id: str) -> dict:
|
||||
return next(item for item in report["recipes"] if item["recipe"]["id"] == recipe_id)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
contract = load_contract(ROOT / "performance-contract.json")
|
||||
cpu_report = _read("results.json")
|
||||
gpu_config = _read("gpu-diagnostic-config.json")
|
||||
gpu_report = _read("gpu-diagnostic-results.json")
|
||||
|
||||
cpu_evaluation = evaluate_contract(contract, cpu_report)
|
||||
if cpu_evaluation.verdict != "stop":
|
||||
raise RuntimeError("immutable CPU v1 evidence no longer evaluates to stop")
|
||||
|
||||
public_key_bytes = base64.b64decode(contract.baseline["required_signer_public_key"])
|
||||
public_key = Ed25519PublicKey.from_public_bytes(public_key_bytes)
|
||||
public_key.verify(
|
||||
base64.b64decode(gpu_report["provenance"]["signature"]),
|
||||
report_signing_payload(gpu_report),
|
||||
)
|
||||
signer_fingerprint = hashlib.sha256(public_key_bytes).hexdigest()
|
||||
if gpu_report["provenance"]["signer_public_key_sha256"] != signer_fingerprint:
|
||||
raise RuntimeError("GPU report signer fingerprint does not match the contract trust key")
|
||||
if gpu_report["provenance"]["config_sha256"] != _canonical_sha256(gpu_config):
|
||||
raise RuntimeError("GPU report is not bound to gpu-diagnostic-config.json")
|
||||
|
||||
if gpu_report.get("schema_version") != 1 or gpu_report.get("evidence_class") != "local-real":
|
||||
raise RuntimeError("GPU report must be schema-v1 local-real evidence")
|
||||
expected_producer = "meshnet_node.recipe_drivers.run_configured_gpu_diagnostic/v1"
|
||||
if gpu_report["provenance"].get("producer") != expected_producer:
|
||||
raise RuntimeError("GPU report was not emitted by the canonical diagnostic producer")
|
||||
if gpu_report.get("reference_recipe_id") != "transformers-fp32-rocm-quality-oracle":
|
||||
raise RuntimeError("GPU report uses the wrong quality reference")
|
||||
if gpu_report.get("host", {}).get("benchmark_lane") != "rocm-gpu-diagnostic":
|
||||
raise RuntimeError("GPU report lacks the diagnostic host marker")
|
||||
|
||||
trusted = json.loads(
|
||||
(ROOT.parents[1] / "trusted-evidence-signers.json").read_text(encoding="utf-8")
|
||||
)
|
||||
if not any(
|
||||
signer.get("algorithm") == "ed25519"
|
||||
and signer.get("fingerprint_sha256") == signer_fingerprint
|
||||
and signer.get("status") == "active"
|
||||
for signer in trusted.get("signers", ())
|
||||
):
|
||||
raise RuntimeError("GPU signer is not active in the trusted-signers registry")
|
||||
|
||||
for field in ("model_id", "model_revision"):
|
||||
if gpu_report["plan"].get(field) != cpu_report["plan"].get(field):
|
||||
raise RuntimeError(f"CPU and GPU reports do not share {field}")
|
||||
if gpu_config["plan"].get(field) != gpu_report["plan"].get(field):
|
||||
raise RuntimeError(f"GPU config and report do not share {field}")
|
||||
|
||||
expected_recipes = {
|
||||
"transformers-fp32-rocm-quality-oracle": ("quality", "cuda"),
|
||||
"llama-cpp-bf16-rocm-quality": ("quality", "cuda"),
|
||||
"transformers-bf16-rocm-throughput": ("performance-fit", "cuda"),
|
||||
"llama-cpp-q4-rocm-throughput": ("performance-fit", "cuda"),
|
||||
}
|
||||
actual_recipes = {
|
||||
entry["recipe"]["id"]: (entry["recipe"]["lane"], entry["recipe"]["device"])
|
||||
for entry in gpu_report["recipes"]
|
||||
}
|
||||
if actual_recipes != expected_recipes:
|
||||
raise RuntimeError("GPU report recipe identities, lanes, or devices changed")
|
||||
|
||||
gpu_prompt_ids = {prompt["id"] for prompt in gpu_report["plan"]["prompts"]}
|
||||
levels = {int(level) for level in gpu_report["plan"]["concurrency_levels"]}
|
||||
repeats = int(gpu_report["plan"]["repeats"])
|
||||
expected_outcomes = len(gpu_prompt_ids) * repeats * sum(levels)
|
||||
for entry in gpu_report["recipes"]:
|
||||
recipe_id = entry["recipe"]["id"]
|
||||
if not entry.get("available") or len(entry.get("outcomes", ())) != expected_outcomes:
|
||||
raise RuntimeError(f"GPU recipe {recipe_id!r} lacks complete outcomes")
|
||||
if any(
|
||||
not outcome.get("ok")
|
||||
or outcome.get("recipe_id") != recipe_id
|
||||
or outcome.get("prompt_id") not in gpu_prompt_ids
|
||||
or int(outcome.get("concurrency", 0)) not in levels
|
||||
or not 0 <= int(outcome.get("repeat", -1)) < repeats
|
||||
for outcome in entry["outcomes"]
|
||||
):
|
||||
raise RuntimeError(f"GPU recipe {recipe_id!r} contains failed or invalid outcomes")
|
||||
if {int(level) for level in entry["concurrency"]} != levels:
|
||||
raise RuntimeError(f"GPU recipe {recipe_id!r} has wrong concurrency cells")
|
||||
for prompt_id in gpu_prompt_ids:
|
||||
for level in levels:
|
||||
for repeat in range(repeats):
|
||||
count = sum(
|
||||
outcome["prompt_id"] == prompt_id
|
||||
and int(outcome["concurrency"]) == level
|
||||
and int(outcome["repeat"]) == repeat
|
||||
for outcome in entry["outcomes"]
|
||||
)
|
||||
if count != level:
|
||||
raise RuntimeError(
|
||||
f"GPU recipe {recipe_id!r} lacks complete request coverage"
|
||||
)
|
||||
if any(
|
||||
int(cell.get("failures", -1)) != 0
|
||||
or int(cell.get("requests", -1))
|
||||
!= len(
|
||||
[
|
||||
outcome
|
||||
for outcome in entry["outcomes"]
|
||||
if int(outcome["concurrency"]) == int(level)
|
||||
]
|
||||
)
|
||||
for level, cell in entry["concurrency"].items()
|
||||
):
|
||||
raise RuntimeError(f"GPU recipe {recipe_id!r} aggregates do not match outcomes")
|
||||
|
||||
cpu_quality = _drift(cpu_report, "llama-cpp-near-lossless-quality")
|
||||
gpu_quality = _drift(gpu_report, "llama-cpp-bf16-rocm-quality")
|
||||
cpu_recipe = _recipe(cpu_report, "llama-cpp-near-lossless-quality")
|
||||
gpu_recipe = _recipe(gpu_report, "llama-cpp-bf16-rocm-quality")
|
||||
gpu_backend = gpu_recipe["load"]["backend_detail"]
|
||||
if "measured accelerator ROCm0: Radeon 8060S Graphics" not in gpu_backend:
|
||||
raise RuntimeError("GPU report lacks measured ROCm device evidence")
|
||||
if "measured offload 25/25 layers" not in gpu_backend:
|
||||
raise RuntimeError("GPU report lacks measured layer-offload evidence")
|
||||
if cpu_recipe["recipe"]["artifact_sha256"] != gpu_recipe["recipe"]["artifact_sha256"]:
|
||||
raise RuntimeError("CPU and GPU diagnostics use different BF16 GGUF artifacts")
|
||||
if gpu_quality.get("compared_prompts") != len(gpu_prompt_ids):
|
||||
raise RuntimeError("GPU quality drift lacks complete prompt coverage")
|
||||
if {item["prompt_id"] for item in gpu_quality.get("per_prompt", ())} != gpu_prompt_ids:
|
||||
raise RuntimeError("GPU quality drift prompt identities do not match the plan")
|
||||
|
||||
summary = {
|
||||
"schema_version": 2,
|
||||
"model_id": cpu_report["plan"]["model_id"],
|
||||
"model_revision": cpu_report["plan"]["model_revision"],
|
||||
"cpu_v1": {
|
||||
"report": "results.json",
|
||||
"report_sha256": _file_sha256("results.json"),
|
||||
"run_id": cpu_report["provenance"]["run_id"],
|
||||
"plan_id": cpu_report["plan"]["plan_id"],
|
||||
"plan_sha256": _canonical_sha256(cpu_report["plan"]),
|
||||
"config_sha256": cpu_report["provenance"]["config_sha256"],
|
||||
"device": "cpu",
|
||||
"quality_oracle": "Transformers BF16 safetensors",
|
||||
"candidate": "llama.cpp BF16 GGUF",
|
||||
"candidate_artifact_sha256": cpu_recipe["recipe"]["artifact_sha256"],
|
||||
"exact_match_rate": cpu_quality["exact_match_rate"],
|
||||
"mean_similarity": cpu_quality["mean_similarity"],
|
||||
"contract_verdict": cpu_evaluation.verdict,
|
||||
"root_cause": "undetermined; no logit-tie claim is acceptance evidence",
|
||||
},
|
||||
"rocm_diagnostic": {
|
||||
"report": "gpu-diagnostic-results.json",
|
||||
"report_sha256": _file_sha256("gpu-diagnostic-results.json"),
|
||||
"run_id": gpu_report["provenance"]["run_id"],
|
||||
"producer": gpu_report["provenance"]["producer"],
|
||||
"signer_fingerprint": signer_fingerprint,
|
||||
"plan_id": gpu_report["plan"]["plan_id"],
|
||||
"plan_sha256": _canonical_sha256(gpu_report["plan"]),
|
||||
"config_sha256": gpu_report["provenance"]["config_sha256"],
|
||||
"device": "cuda (ROCm)",
|
||||
"quality_oracle": "Transformers float32 safetensors",
|
||||
"candidate": "llama.cpp BF16 GGUF",
|
||||
"candidate_artifact_sha256": gpu_recipe["recipe"]["artifact_sha256"],
|
||||
"measured_backend_detail": gpu_backend,
|
||||
"exact_match_rate": gpu_quality["exact_match_rate"],
|
||||
"mean_similarity": gpu_quality["mean_similarity"],
|
||||
"failures": sum(
|
||||
metrics["failures"]
|
||||
for entry in gpu_report["recipes"]
|
||||
for metrics in entry["concurrency"].values()
|
||||
),
|
||||
"v1_eligible": False,
|
||||
},
|
||||
"conclusion": {
|
||||
"v1_verdict_changed": False,
|
||||
"cpu_bf16_divergence_explained": False,
|
||||
"conversion_corruption_observed_in_rocm_sample": False,
|
||||
"scope": (
|
||||
"The ROCm diagnostic establishes only that the same BF16 GGUF artifact "
|
||||
"matched the float32 oracle for three GPU sequences; it does not explain "
|
||||
"the CPU BF16 divergence or prove global conversion correctness."
|
||||
),
|
||||
"recommended_v2_design": (
|
||||
"Predeclare a float32 quality oracle separately from the BF16 performance "
|
||||
"reference, with a larger prompt corpus and immutable thresholds."
|
||||
),
|
||||
},
|
||||
}
|
||||
|
||||
(ROOT / "quality-parity-diagnosis.json").write_text(
|
||||
json.dumps(summary, indent=2, sort_keys=True) + "\n", encoding="utf-8"
|
||||
)
|
||||
md = f"""# DGR-001 quality-parity evidence summary
|
||||
|
||||
This summary is generated by `summarize-quality-parity.py` from signed reports.
|
||||
It contains no independent logit measurements or self-asserted verification flag.
|
||||
|
||||
| Source | Device | Quality oracle | BF16 GGUF candidate | Exact | Similarity | Status |
|
||||
|---|---|---|---|---:|---:|---|
|
||||
| CPU v1 (`{summary['cpu_v1']['run_id']}`) | CPU | Transformers BF16 | llama.cpp BF16 | {summary['cpu_v1']['exact_match_rate']:.4f} | {summary['cpu_v1']['mean_similarity']:.4f} | immutable `stop` |
|
||||
| ROCm diagnostic (`{summary['rocm_diagnostic']['run_id']}`) | ROCm0 / Radeon 8060S | Transformers float32 | llama.cpp BF16 | {summary['rocm_diagnostic']['exact_match_rate']:.4f} | {summary['rocm_diagnostic']['mean_similarity']:.4f} | diagnostic only |
|
||||
|
||||
## Interpretation
|
||||
|
||||
The CPU and ROCm rows use different plans, devices, kernels, and quality oracles.
|
||||
The CPU BF16 divergence remains unexplained and v1 remains `stop`. The signed
|
||||
ROCm report establishes the narrower fact that the same BF16 GGUF artifact
|
||||
matched the float32 oracle for all three GPU sequences with zero failures.
|
||||
Its signed backend detail records `ROCm0: Radeon 8060S Graphics` and measured
|
||||
`25/25` layer offload.
|
||||
|
||||
No conversion corruption was observed in that three-sequence ROCm sample. This
|
||||
does not prove global conversion correctness and does not retroactively change
|
||||
or explain the CPU result. A future v2 should predeclare a float32 quality oracle
|
||||
separately from its BF16 performance reference and use a larger corpus.
|
||||
|
||||
## Reproduction and bindings
|
||||
|
||||
- CPU report SHA-256: `{summary['cpu_v1']['report_sha256']}`
|
||||
- GPU report SHA-256: `{summary['rocm_diagnostic']['report_sha256']}`
|
||||
- BF16 GGUF SHA-256: `{summary['rocm_diagnostic']['candidate_artifact_sha256']}`
|
||||
- Signer fingerprint: `{signer_fingerprint}`
|
||||
- Exact verification command: see `commands.txt`.
|
||||
"""
|
||||
(ROOT / "quality-parity-diagnosis.md").write_text(md, encoding="utf-8")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
242
.scratch/distributed-gguf-runtime/evidence/DGR-002/README.md
Normal file
242
.scratch/distributed-gguf-runtime/evidence/DGR-002/README.md
Normal file
@@ -0,0 +1,242 @@
|
||||
# DGR-002 — Adopt the versioned gRPC Shard protocol
|
||||
|
||||
Status: **done**. Every acceptance criterion is met with real command output.
|
||||
Evidence class: **synthetic/unit** — this story defines a schema and proves both
|
||||
languages agree on it. No model, GPU, network peer or benchmark is involved, and
|
||||
none is claimed.
|
||||
|
||||
## 1. Summary
|
||||
|
||||
`packages/node/native/proto/shard_runtime.proto` is now the semantic contract for
|
||||
the native Shard data plane: Protocol Buffers over gRPC/HTTP2 (ADR-0020). Python
|
||||
and C++ both generate from it, and a shared committed conformance vector proves
|
||||
they encode it identically — byte for byte.
|
||||
|
||||
Design decisions worth carrying forward:
|
||||
|
||||
- **Everything gRPC gives you is *also* in the schema.** Deadline, cancellation,
|
||||
identity and flow control are carried as fields, not left to HTTP/2 metadata,
|
||||
because the existing relay carries these frames as **opaque binary**. A relayed
|
||||
frame has no HTTP/2 context to inherit a deadline or a channel identity from.
|
||||
If it is not in the schema, it does not survive the relay.
|
||||
- **Cancellation is both in-band and out-of-band.** `CancelSignal` rides the
|
||||
stream; `Cancel` is also a unary RPC. A cancel that can only travel down a
|
||||
stream that flow control has wedged is not a cancel.
|
||||
- **Checksums cover the uncompressed payload.** Compression is a per-hop
|
||||
transport decision (reusing the existing `activation_compression` policies), so
|
||||
a checksum over the compressed frame would be invalidated by a hop that merely
|
||||
chose differently.
|
||||
- **Application-level flow-control credits, not just HTTP/2 windows.** HTTP/2
|
||||
bounds *bytes in flight*; it does not bound how much *work* a worker has queued,
|
||||
and a relayed frame gets no window at all. Credits bound queue occupancy and KV
|
||||
pressure, and negotiation takes the strictest bound of either peer so a sender
|
||||
cannot talk a worker into unbounded queues.
|
||||
|
||||
## 2. Files changed
|
||||
|
||||
New:
|
||||
|
||||
| Path | What |
|
||||
|---|---|
|
||||
| `packages/node/native/proto/shard_runtime.proto` | The schema (sha256 `9e211660…`, see `protocol.json`) |
|
||||
| `packages/node/native/CMakeLists.txt` | C++ generation + build wiring + ctest |
|
||||
| `packages/node/native/tests/test_shard_protocol_conformance.cpp` | C++ conformance test |
|
||||
| `packages/node/native/testdata/*.binpb` | Committed cross-language vectors |
|
||||
| `packages/node/native/README.md` | How to regenerate and build |
|
||||
| `packages/node/meshnet_node/native_protocol/__init__.py` | Public Python surface |
|
||||
| `packages/node/meshnet_node/native_protocol/codec.py` | Bundle encode/decode, fragmentation, CRC32C, chunking, FC negotiation |
|
||||
| `packages/node/meshnet_node/native_protocol/conformance.py` | Canonical vectors shared by both languages |
|
||||
| `packages/node/meshnet_node/native_protocol/generated/` | Generated Python stubs (committed) |
|
||||
| `scripts/generate_native_protocol.py` | Python generation, with `--check` |
|
||||
| `scripts/generate_protocol_goldens.py` | Vector generation, with `--check` |
|
||||
| `scripts/bootstrap_native_toolchain.sh` | Builds protobuf C++ from source |
|
||||
| `tests/test_native_shard_protocol.py` | 45 Python tests |
|
||||
|
||||
Modified:
|
||||
|
||||
- `packages/node/pyproject.toml` — added runtime floors `grpcio>=1.82.1` and
|
||||
`protobuf>=7.35.0`, matching the committed generated-code requirements; new
|
||||
`proto` extra pinning `grpcio-tools==1.82.1`.
|
||||
- `packages/node/meshnet_node/activation_compression.py` — optional bounded zstd
|
||||
output for untrusted protocol frames; existing callers remain compatible.
|
||||
- `packages/node/meshnet_node/native_protocol/__init__.py` — exports negotiated
|
||||
bound constants and whole-session-message validation.
|
||||
|
||||
The canonical PRD marks only DGR-002 passed. `git status` before this story was clean.
|
||||
|
||||
## 3. Commands and real results
|
||||
|
||||
See `commands.txt` for the exact ordered list. Results:
|
||||
|
||||
```
|
||||
python scripts/generate_native_protocol.py --check -> generated stubs are up to date
|
||||
python scripts/generate_protocol_goldens.py --check -> conformance vectors are up to date
|
||||
|
||||
cmake -S packages/node/native -B build/native -DCMAKE_PREFIX_PATH=/tmp/pbsrc/install
|
||||
-- gRPC C++ not found: building message types only (sufficient for the conformance test)
|
||||
cmake --build build/native -j -> Built target shard_protocol_conformance
|
||||
ctest --test-dir build/native --output-on-failure -> 1/1 Test #1: shard_protocol_conformance ... Passed
|
||||
100% tests passed out of 1
|
||||
|
||||
cmp build/native/cpp_roundtrip.binpb \
|
||||
packages/node/native/testdata/session_request_golden.binpb -> identical (exit 0)
|
||||
|
||||
pytest -q tests/test_native_shard_protocol.py -> 45 passed
|
||||
pytest -q tests/test_native_shard_protocol.py \
|
||||
tests/test_activation_compression.py -> 51 passed
|
||||
pytest -q (final full suite) -> 728 passed, 12 skipped
|
||||
pytest -q tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
(after an earlier flaky full-suite failure) -> 1 passed, 1 passed, 1 passed
|
||||
clean minimum-runtime import + codec smoke test -> passed
|
||||
grpcio==1.82.1, protobuf==7.35.0
|
||||
compileall -q packages tests -> OK (exit 0)
|
||||
git diff --check -> clean (exit 0)
|
||||
```
|
||||
|
||||
The C++ lane was rebuilt from scratch by Ralph (`rm -rf build/native`) using only
|
||||
the documented commands, and reproduced the same result. During controller
|
||||
review the user explicitly chose not to repeat the destructive build-directory
|
||||
cleanup, so the independent controller relied on the recorded CMake/CTest run
|
||||
while reproducing every Python/generation/full-suite gate.
|
||||
|
||||
### Controller review corrections
|
||||
|
||||
Independent controller review found and fixed two classes of issue before
|
||||
integration:
|
||||
|
||||
1. Generated stubs required gRPC 1.82.1 and Protobuf 7.35.0, while the initial
|
||||
package metadata allowed much older runtimes that could fail at import time.
|
||||
2. Flow-control bounds were described but not enforced by the reference decoder.
|
||||
Tensor declarations, shape rank/dimensions, fragment/tensor counts, fragments,
|
||||
wire bodies, whole bundles, complete session messages (including envelope
|
||||
overhead), and zstd window/output expansion are now fail-closed against the
|
||||
negotiated/default bounds. Unspecified bundle versions, compression and
|
||||
checksums are rejected rather than interpreted as valid data.
|
||||
3. Negotiated initial credits could exceed `max_inflight_chunks`; credits are now
|
||||
capped by the settled in-flight limit.
|
||||
|
||||
Controller results: protocol tests `45 passed`; protocol plus shared compression
|
||||
tests `51 passed`; final full suite `728 passed, 12 skipped`. A clean environment
|
||||
at the declared minimum gRPC/Protobuf runtime versions imported both generated
|
||||
stub modules and round-tripped the codec. Generation checks, `compileall`, static
|
||||
secret scan, and `git diff --check` all passed.
|
||||
|
||||
### Full-suite note — a pre-existing flaky test
|
||||
|
||||
`tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy`
|
||||
is **flaky on a clean tree, independent of this story**. Reproduction, run
|
||||
*before any DGR-002 file existed* (working tree clean, `git status` empty):
|
||||
|
||||
```
|
||||
pytest -q -> 1 failed, 682 passed, 12 skipped
|
||||
FAILED tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
|
||||
# same test, three consecutive isolated runs on the same clean tree:
|
||||
pytest -q tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
-> 1 passed in 1.76s
|
||||
-> 1 failed in 4.39s
|
||||
-> 1 passed in 1.10s
|
||||
```
|
||||
|
||||
It is a timing race in proxy cancellation (a 3-second in-flight generation raced
|
||||
against the cancel assertion), not a deterministic failure, and it touches no code
|
||||
this story changes. One controller full-suite run reported exactly that one failure
|
||||
(`1 failed, 719 passed, 12 skipped`); three immediate isolated retries all passed
|
||||
in 1.11 seconds, and the final exact-code full suite was green (`728 passed,
|
||||
12 skipped`). It is flagged for whoever owns the tracker cancel path and is **not**
|
||||
fixed here, since silently touching another story's code is out of scope.
|
||||
|
||||
## 4. Acceptance criteria
|
||||
|
||||
| Criterion | Where it is proven |
|
||||
|---|---|
|
||||
| Schema for capability, health, session stream, release, cancellation | `shard_runtime.proto` `service ShardRuntime`; `test_service_exposes_capability_health_session_release_and_cancel` |
|
||||
| One long-lived bidi stream per Activation Seam, with deadlines, cancellation, flow control, structured errors | `rpc Session (stream) returns (stream)`; `test_session_is_one_long_lived_bidirectional_stream`; `Envelope.deadline_unix_nanos`, `CancelSignal` + unary `Cancel`, `FlowControl`, `ShardError` |
|
||||
| Bounded chunking for prefill; small decode fast path | `ChunkInfo` + `plan_prefill_chunks` (128-token bound, ADR-0008); `DecodeStep`; `test_prefill_is_split_into_bounded_token_aligned_chunks`, `test_decode_fast_path_is_much_smaller_than_a_full_envelope_chunk` |
|
||||
| Envelope carries schema version, work id, session id, epoch, fingerprint, range/effective start, phase, position, idempotency step, cache expectation, compression, checksum | `Envelope` + `NamedTensor`; `test_envelope_carries_every_field_the_protocol_promises` asserts against the **descriptor**, so deleting a field from the `.proto` fails the test |
|
||||
| Versioned named-tensor bundle: name, shape, dtype, byte order, fragments | `TensorBundle`/`NamedTensor`/`TensorFragment`; `test_named_tensor_bundle_is_versioned_and_fully_described`, `test_bundle_round_trips_multiple_named_tensors` |
|
||||
| Round-trip + compatibility tests in Python and C++ | 45 Python tests; C++ `ctest` 1/1; cross-language byte equality |
|
||||
| Targeted pytest passes | 45 passed |
|
||||
| `compileall packages tests` | exit 0 |
|
||||
| `git diff --check` | exit 0 |
|
||||
| Default tests deterministic, download-free, credit-free, GPU-free | Pure in-memory protobuf; no model, no network, no GPU |
|
||||
| Full deterministic pytest passes, or pre-existing failure recorded | Final exact-code run: 728 passed, 12 skipped; earlier sole flaky failure documented with clean-tree reproduction and 3/3 passing retries |
|
||||
|
||||
## 5. How the cross-language claim is actually earned
|
||||
|
||||
Two codecs that each round-trip their own output prove only that each is
|
||||
self-consistent. Instead:
|
||||
|
||||
1. Python builds the canonical `SessionRequest` and commits its bytes.
|
||||
2. The C++ test parses **those** bytes, asserts every field, recomputes the CRC32C
|
||||
**from the polynomial in independent C++ code**, reassembles the multi-fragment
|
||||
tensor, and re-serializes to `cpp_roundtrip.binpb`.
|
||||
3. `test_cpp_and_python_agree_byte_for_byte` asserts that file equals the golden.
|
||||
|
||||
Compatibility is tested in both languages: an unknown field from a newer peer
|
||||
survives a parse/serialize hop (a Shard forwards activations — silently stripping
|
||||
fields would corrupt a route it is merely a waypoint on), and a sparse message
|
||||
from an older peer parses to proto3 defaults.
|
||||
|
||||
## 6. Limitations and deferred work
|
||||
|
||||
- **gRPC C++ was not built or linked.** The C++ lane verifies the *schema* (message
|
||||
types), not a running gRPC C++ server, because this machine has no gRPC C++ stack
|
||||
and building it is a large dependency the conformance test does not need.
|
||||
`CMakeLists.txt` already generates and exports `shard_runtime_grpc` when
|
||||
`find_package(gRPC)` succeeds. **DGR-008 must install gRPC C++ and extend
|
||||
`scripts/bootstrap_native_toolchain.sh`.**
|
||||
- **No wire is exercised.** No client, server, or stream lifecycle exists yet — no
|
||||
deadline actually fires, no credit is actually consumed. This story defines and
|
||||
proves the contract; DGR-008/DGR-009 implement it.
|
||||
- The protobuf C++ toolchain used here was installed to `/tmp/pbsrc/install` (ephemeral).
|
||||
`scripts/bootstrap_native_toolchain.sh` reproduces it; prefer a durable prefix such
|
||||
as `build/native-toolchain`.
|
||||
- `crc32c` has a pure-Python fallback (used here) and picks up `google_crc32c` when
|
||||
present. The fallback is byte-exact but slow; a worker on the hot path should install
|
||||
the native package. Not a correctness limitation.
|
||||
- Compression on the wire is zstd-or-none only, matching the existing seam.
|
||||
|
||||
## 7. Compatibility and migration notes
|
||||
|
||||
- **This does not change the existing HTTP activation wire.** `X-Meshnet-Wire` stays
|
||||
at `2` and the legacy `/forward` path is untouched. The native protocol is a
|
||||
*separate* contract with its own `SchemaVersion`, starting at 1. Nothing in this
|
||||
story is on any live request path — it is additive.
|
||||
- Semantics are deliberately preserved from the existing ADRs so the two transports
|
||||
mean the same thing: `effective_start_layer` (ADR-0012), `CacheMode`/`expected_past_len`
|
||||
and `ERROR_CODE_CACHE_MISS` mapping to today's HTTP 409 `cache_miss` (ADR-0022),
|
||||
bfloat16 boundary dtype and 128-token prefill chunks (ADR-0008), fingerprint/recipe
|
||||
identity mirroring the capability report (ADR-0023).
|
||||
- `TensorFragment` field 5 (`uncompressed_size`) is **reserved**: it was removed
|
||||
because `NamedTensor.total_bytes` is the single source of truth. Never recycle it —
|
||||
a recycled field number is the one schema change peers cannot detect, because the
|
||||
bytes still parse.
|
||||
- Committed Python stubs are guarded by `--check` in the test suite, so they cannot
|
||||
drift from the schema unnoticed.
|
||||
|
||||
## 8. Handoff to dependent stories
|
||||
|
||||
- **DGR-003 (runtime recipe/fingerprint):** populate `Fingerprint`
|
||||
(`model_artifact_digest`, `runtime_recipe_digest`, `recipe_id`, `recipe_version`,
|
||||
`catalogue_version`). The mismatch outcome is already specified:
|
||||
`ERROR_CODE_FINGERPRINT_MISMATCH`. Do not invent a second identity struct.
|
||||
- **DGR-005/006 (range loading, architecture boundary):** the boundary payload is a
|
||||
**named bundle**, not a bare tensor — a boundary needing more than one tensor is
|
||||
already representable. Execute `[effective_start_layer, end_layer)`, never from
|
||||
`start_layer`.
|
||||
- **DGR-007 (concurrent sessions/KV):** isolate on `(route_session_id, route_epoch)`.
|
||||
`CacheExpectation`/`CacheResult` and `ERROR_CODE_CACHE_MISS` are the contract; a
|
||||
decode step whose `expected_past_len` does not match **must** miss, never fall back
|
||||
to a silent stateless forward. `idempotency_step` means a retried step is
|
||||
acknowledged (`Ack.duplicate`), not re-applied — re-applying advances the KV cache
|
||||
twice and desynchronises the route.
|
||||
- **DGR-008 (C++ worker):** link `shard_runtime_grpc` from `CMakeLists.txt`; you must
|
||||
first install gRPC C++ (see limitations). Honour `FlowControl` credits and the
|
||||
`max_chunk_bytes` bound. Use `packages/node/meshnet_node/native_protocol/codec.py`
|
||||
as the reference for fragment reassembly and checksum validation.
|
||||
- **DGR-009 (Meshnet integration):** the relay may carry these serialized frames as
|
||||
opaque binary — that is exactly why deadline/cancel/identity are in-band. Do not add
|
||||
a second control plane.
|
||||
- **Anyone editing the schema:** run both `--check` scripts; if a vector legitimately
|
||||
changes, regenerate it and say so, because the C++ test asserts those exact bytes.
|
||||
@@ -0,0 +1,45 @@
|
||||
# DGR-002 — exact commands, in order. Run from the repository root.
|
||||
# Interpreter: <repo>/.venv/bin/python (CPython 3.14.6). Deterministic, GPU-free,
|
||||
# no model download, no API credits.
|
||||
|
||||
# --- toolchain (this machine had no protoc, no cmake, no protobuf C++ headers)
|
||||
.venv/bin/python -m pip install grpcio-tools==1.82.1 grpcio==1.82.1 cmake==4.4.0
|
||||
scripts/bootstrap_native_toolchain.sh /tmp/pbsrc/install # protobuf C++ 33.1 + abseil 20250814.1
|
||||
|
||||
# --- schema generation (Python stubs; committed)
|
||||
.venv/bin/python scripts/generate_native_protocol.py
|
||||
.venv/bin/python scripts/generate_native_protocol.py --check # -> "generated stubs are up to date"
|
||||
|
||||
# --- cross-language conformance vectors (committed)
|
||||
.venv/bin/python scripts/generate_protocol_goldens.py
|
||||
.venv/bin/python scripts/generate_protocol_goldens.py --check # -> "conformance vectors are up to date"
|
||||
|
||||
# --- C++ generation, build and conformance test
|
||||
cmake -S packages/node/native -B build/native -DCMAKE_PREFIX_PATH=/tmp/pbsrc/install
|
||||
cmake --build build/native -j"$(nproc)"
|
||||
ctest --test-dir build/native --output-on-failure # -> 1/1 Passed
|
||||
cmp build/native/cpp_roundtrip.binpb packages/node/native/testdata/session_request_golden.binpb
|
||||
|
||||
# --- Python tests
|
||||
.venv/bin/python -m pytest -q tests/test_native_shard_protocol.py # -> 29 passed
|
||||
.venv/bin/python -m pytest -q # full suite
|
||||
|
||||
# --- repository gates
|
||||
.venv/bin/python -m compileall -q packages tests
|
||||
git diff --check
|
||||
|
||||
# --- independent controller review after Ralph
|
||||
PYTHONPATH=packages/node .venv/bin/python -m pytest -q tests/test_native_shard_protocol.py
|
||||
# -> 45 passed
|
||||
PYTHONPATH=packages/node .venv/bin/python -m pytest -q \
|
||||
tests/test_native_shard_protocol.py tests/test_activation_compression.py
|
||||
# -> 51 passed
|
||||
PYTHONPATH=packages/node .venv/bin/python -m pytest -q
|
||||
# -> final exact-code run: 728 passed, 12 skipped
|
||||
for i in 1 2 3; do PYTHONPATH=packages/node .venv/bin/python -m pytest -q \
|
||||
tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy; done
|
||||
# -> 1 passed, 1 passed, 1 passed
|
||||
# clean minimum-runtime venv: protobuf==7.35.0 grpcio==1.82.1
|
||||
# generated pb2 + pb2_grpc imports and one-byte codec round trip -> passed
|
||||
# The user chose to rely on Ralph's recorded successful C++ CMake/CTest run
|
||||
# rather than repeat deletion of an isolated generated build directory.
|
||||
@@ -0,0 +1,95 @@
|
||||
{
|
||||
"schema_version": "SCHEMA_VERSION_1",
|
||||
"bundle_version": 1,
|
||||
"proto_path": "packages/node/native/proto/shard_runtime.proto",
|
||||
"proto_sha256": "9e211660b3fcefc88bcdf3851c3571088c00349aacb5adc5ef45083c83d0cce2",
|
||||
"protoc": "grpc_tools 1.82.1 (python) / protobuf 33.1 (C++)",
|
||||
"service": {
|
||||
"GetCapability": {
|
||||
"client_streaming": false,
|
||||
"server_streaming": false
|
||||
},
|
||||
"Health": {
|
||||
"client_streaming": false,
|
||||
"server_streaming": false
|
||||
},
|
||||
"Session": {
|
||||
"client_streaming": true,
|
||||
"server_streaming": true
|
||||
},
|
||||
"Release": {
|
||||
"client_streaming": false,
|
||||
"server_streaming": false
|
||||
},
|
||||
"Cancel": {
|
||||
"client_streaming": false,
|
||||
"server_streaming": false
|
||||
}
|
||||
},
|
||||
"envelope_fields": [
|
||||
"cache_expectation",
|
||||
"chunk",
|
||||
"deadline_unix_nanos",
|
||||
"fingerprint",
|
||||
"idempotency_step",
|
||||
"phase",
|
||||
"position",
|
||||
"route_epoch",
|
||||
"route_session_id",
|
||||
"schema_version",
|
||||
"shard_range",
|
||||
"work_id"
|
||||
],
|
||||
"named_tensor_fields": [
|
||||
"byte_order",
|
||||
"checksum",
|
||||
"compression",
|
||||
"dtype",
|
||||
"fragments",
|
||||
"name",
|
||||
"shape",
|
||||
"total_bytes"
|
||||
],
|
||||
"phases": [
|
||||
"PHASE_UNSPECIFIED",
|
||||
"PHASE_PREFILL",
|
||||
"PHASE_DECODE",
|
||||
"PHASE_RELEASE",
|
||||
"PHASE_CANCEL"
|
||||
],
|
||||
"error_codes": [
|
||||
"ERROR_CODE_UNSPECIFIED",
|
||||
"ERROR_CODE_SCHEMA_UNSUPPORTED",
|
||||
"ERROR_CODE_FINGERPRINT_MISMATCH",
|
||||
"ERROR_CODE_EPOCH_STALE",
|
||||
"ERROR_CODE_SHARD_RANGE_MISMATCH",
|
||||
"ERROR_CODE_CACHE_MISS",
|
||||
"ERROR_CODE_RESOURCE_EXHAUSTED",
|
||||
"ERROR_CODE_PAYLOAD_CORRUPT",
|
||||
"ERROR_CODE_CANCELLED",
|
||||
"ERROR_CODE_DEADLINE_EXCEEDED",
|
||||
"ERROR_CODE_FLOW_CONTROL_VIOLATION",
|
||||
"ERROR_CODE_INTERNAL"
|
||||
],
|
||||
"bounds": {
|
||||
"max_prefill_chunk_tokens": 128,
|
||||
"max_chunk_bytes": 4194304,
|
||||
"max_fragment_bytes": 1048576,
|
||||
"max_inflight_chunks": 8,
|
||||
"max_fragments_per_tensor": 64,
|
||||
"max_tensors_per_bundle": 64,
|
||||
"max_tensor_rank": 8,
|
||||
"max_tensor_dimension": 2147483647,
|
||||
"whole_session_message_enforced": true
|
||||
},
|
||||
"golden_vectors": {
|
||||
"session_request_golden.binpb": "c2c3df8a717ddeae7bd99624d2c7f34c09a518988de990237fe313b75cff0817",
|
||||
"capability_report_golden.binpb": "71ac5f150775f398515b43a63596a5cbe8d2ad607e7e4de56bd44fbe7987080c"
|
||||
},
|
||||
"verification": {
|
||||
"python_protocol_tests": "45 passed",
|
||||
"python_protocol_and_compression_tests": "51 passed",
|
||||
"full_suite": "728 passed, 12 skipped",
|
||||
"minimum_runtime": "grpcio 1.82.1 / protobuf 7.35.0 passed import and codec smoke"
|
||||
}
|
||||
}
|
||||
186
.scratch/distributed-gguf-runtime/evidence/DGR-003/README.md
Normal file
186
.scratch/distributed-gguf-runtime/evidence/DGR-003/README.md
Normal file
@@ -0,0 +1,186 @@
|
||||
# DGR-003 — exact Artifact and runtime recipe identity
|
||||
|
||||
Evidence class: deterministic offline/unit. No model payload, GPU, external API,
|
||||
network node, or API credit is required or claimed.
|
||||
|
||||
## Result — delayed-review repair, 2026-07-14
|
||||
|
||||
DGR-003 defines and tests an exact, model-agnostic compatibility identity and
|
||||
connects it to DGR-002's gRPC `Fingerprint` plus tracker parsing, admission,
|
||||
route partitioning, and certification. It is **not complete**: the existing
|
||||
production doctor/backend path still emits the legacy capability report without
|
||||
constructing a `ShardIdentity` from authoritative loaded artifact/runtime state.
|
||||
No exact recipe is therefore claimed live or routable from that path; supplied
|
||||
exact identities remain dark until tracker-owned certification.
|
||||
|
||||
A matching digest proves canonical consistency, **not node authenticity or real
|
||||
execution**. Tracker-owned certification of a fingerprint by a non-synthetic,
|
||||
complete, multi-node distributed forward is the execution trust boundary.
|
||||
|
||||
## Implementation
|
||||
|
||||
- `ArtifactIdentity` binds artifact ID/revision, exact content digest,
|
||||
architecture/config digest, layer count, and optional derivative binding.
|
||||
- `DerivativeBinding` binds a split artifact to the exact source artifact digest
|
||||
and its end-exclusive layer range. A Shard cannot advertise outside that range.
|
||||
- `RuntimeRecipe` keeps these canonical axes separate rather than hiding them in
|
||||
a backend label:
|
||||
- weight quantization;
|
||||
- activation and compute dtypes;
|
||||
- KV dtype and layout;
|
||||
- tokenizer revision;
|
||||
- architecture adapter;
|
||||
- backend and runtime version;
|
||||
- boundary and protocol schema versions;
|
||||
- recipe ID/version and catalogue version.
|
||||
- `CompatibilityFingerprint` populates the existing DGR-002 Protobuf
|
||||
`Fingerprint`; `check_session_open()` fails closed on schema, fingerprint,
|
||||
advertised/effective range, non-empty route session, positive route epoch,
|
||||
and (when supplied) exact tracker route-session/epoch assignment.
|
||||
- Node and tracker implementations independently canonicalize the declaration.
|
||||
This is intentional: the tracker must not trust a digest copied from a node,
|
||||
and future native/C++ workers also need an independent implementation. Their
|
||||
behavior is pinned by `tests/data/recipe_fingerprint_vectors.json`.
|
||||
- Tracker admission cross-checks the exact identity against the capability
|
||||
proof's model, range, recipe labels, backend, and weight quantization. Any
|
||||
disagreement fails closed.
|
||||
- `TrackerServer` owns the sole live certification ledger and passes it through
|
||||
direct and replicated registration paths. A known exact recipe is
|
||||
`uncertified` and dark for user traffic until the same exact fingerprint is
|
||||
certified. Restart fails closed; durable/cluster-wide certification events
|
||||
require the later real-forward control path and are not claimed here.
|
||||
- Certification evidence is bound to the promoted fingerprint, requires at
|
||||
least two distinct nodes, complete layer coverage, generated tokens, and
|
||||
`synthetic=false`. Unknown or mismatched fingerprints cannot be promoted.
|
||||
|
||||
## Files changed
|
||||
|
||||
- `packages/node/meshnet_node/runtime_recipe.py`
|
||||
- `packages/tracker/meshnet_tracker/recipe.py`
|
||||
- `packages/tracker/meshnet_tracker/capability.py`
|
||||
- `packages/tracker/meshnet_tracker/server.py`
|
||||
- `tests/data/recipe_fingerprint_vectors.json`
|
||||
- `tests/test_runtime_recipe_identity.py`
|
||||
- this evidence directory, issue state, and DGR-003 PRD state
|
||||
|
||||
A late review of dependency DGR-017 also found and fixed two genuine contract
|
||||
continuity defects during delayed DGR-003 review: v1 now has an independently
|
||||
trusted digest and recursively immutable parsed state. Those changes and tests
|
||||
are recorded in DGR-017 evidence rather than claimed as DGR-003 functionality.
|
||||
|
||||
## Verification
|
||||
|
||||
Exact commands and outcomes are in `commands.txt`.
|
||||
|
||||
Observed final results:
|
||||
|
||||
- DGR-003 identity + node/tracker capability suites: **126 passed**.
|
||||
- DGR-017 focused dependency repair suite: **99 passed**.
|
||||
- Tracker routing suite: **93 passed**.
|
||||
- First delayed-review integrated run: **898 passed, 13 skipped, 1 failed** on
|
||||
the pre-existing tracker-cancellation race.
|
||||
- Final delayed-review integrated rerun: **899 passed, 13 skipped** in
|
||||
**253.64s**; Hermes controller acceptance rerun: **899 passed, 13 skipped**
|
||||
in **252.66s**.
|
||||
- `python -m compileall -q packages tests`: pass.
|
||||
- `git diff --check`: pass.
|
||||
- Ruff on the changed identity, capability, contract, and test modules: pass.
|
||||
- `server.py` has 8 pre-existing Ruff findings at both pushed baseline and the
|
||||
current tree; DGR-003 added no finding.
|
||||
|
||||
The first integrated full-suite run produced **871 passed, 13 skipped, 1 failed**
|
||||
on the known unrelated
|
||||
`test_tracker_dashboard_can_cancel_inflight_proxy` timing race. Its fixture
|
||||
completed after three seconds just before cancellation, so the cancel endpoint
|
||||
returned 404. In this delayed repair it again produced a 404 after the stream
|
||||
finished (first integrated run: **898 passed, 13 skipped, 1 failed**); three
|
||||
immediate isolated repeats passed before a fourth reproduced the same race.
|
||||
No cancellation-test code was changed. The final complete integrated rerun
|
||||
passed **899/899** tests.
|
||||
|
||||
## Limitations
|
||||
|
||||
- Certification state is process-local in this story. The same running tracker
|
||||
reuses it across registrations, but durable/cluster-wide certification-event
|
||||
persistence belongs with the later real distributed-forward control path.
|
||||
Restart or failover therefore returns exact recipes to the safe dark state;
|
||||
it never makes an unsupported recipe routable.
|
||||
- The node module has no certification ledger or admission policy; it holds only
|
||||
identity construction and handshake validation. The Tracker is the sole
|
||||
promotion authority.
|
||||
- **Completion blocker:** `doctor._validate_recipe()` calls
|
||||
`build_capability_report()` without `identity=`, because the legacy
|
||||
Transformers backend does not expose an immutable artifact-content pin and
|
||||
full runtime recipe axes authoritative enough to build one. Adding a guessed
|
||||
identity would weaken this contract. Production emission must be added with
|
||||
the authoritative native worker/backend loading seam; until then the issue and
|
||||
PRD deliberately remain incomplete.
|
||||
- This story proves identity and admission behavior with deterministic fixtures.
|
||||
It does not claim a real GLM forward or hardware certification.
|
||||
|
||||
## Compatibility
|
||||
|
||||
- Capability report identity is additive. Legacy reports without the new block
|
||||
retain ADR-0023's explicit compatibility-policy behavior.
|
||||
- Reports that opt into exact identity are held to it and fail closed on malformed,
|
||||
inconsistent, unknown, dark, or mismatched declarations.
|
||||
- No new wire identity was invented; DGR-002's `Fingerprint` remains the gRPC
|
||||
representation.
|
||||
|
||||
## Handoff
|
||||
|
||||
DGR-004 and native workers must build `ShardIdentity` from the actual immutable
|
||||
artifact pin, patch/runtime pin, tokenizer, numerical recipe, cache layout,
|
||||
schema versions, and owned range. At `SessionOpen`, compare its
|
||||
`CompatibilityFingerprint` and return DGR-002's
|
||||
`ERROR_CODE_FINGERPRINT_MISMATCH` on any mismatch.
|
||||
|
||||
A digest match is not certification. Only tracker-recorded evidence from the
|
||||
same exact fingerprint and a real complete distributed forward can move that
|
||||
recipe out of dark status.
|
||||
|
||||
## Native emission closure — 2026-07-14
|
||||
|
||||
Status: **done**. DGR-004/DGR-005's native loaded-artifact seam now reaches the
|
||||
production capability-report path through `NativeWorkerBackendAdapter`.
|
||||
|
||||
### Files changed
|
||||
|
||||
- `packages/node/meshnet_node/native_backend.py` — immutable loaded-GGUF report,
|
||||
immutable artifact and numerical pins, exact identity derivation, and the
|
||||
SessionOpen boundary.
|
||||
- `packages/node/meshnet_node/doctor.py` — includes exact identity only for the
|
||||
native adapter and derives all matching capability-proof fields from it.
|
||||
- `tests/test_native_identity_emission.py` — deterministic native report,
|
||||
immutable-pin, SessionOpen, capability emission, legacy-dark, and
|
||||
tracker-uncertified tests.
|
||||
- This issue, `prd.json`, and this evidence directory.
|
||||
|
||||
### Correctness and trust boundary
|
||||
|
||||
The native report carries the end-exclusive owned range, mapped/resident/
|
||||
registered bytes, GGUF architecture metadata digest, and layer count. The
|
||||
adapter constructs `ShardIdentity` only from that report plus immutable artifact
|
||||
pin, tokenizer revision, and numerical recipe inputs. It does not accept a
|
||||
caller-supplied shard range.
|
||||
|
||||
`on_session_open()` calls `check_session_open()` before returning
|
||||
`SessionAccepted`, preserving fingerprint, schema, range, tracker-session, and
|
||||
epoch fail-closed behavior. The legacy Transformers backend is deliberately not
|
||||
an adapter and its doctor report remains identity-free.
|
||||
|
||||
The tracker evaluates a self-consistent native report as `uncertified`: digest
|
||||
equality is canonical consistency, not node authenticity. Only its owned
|
||||
certification ledger can promote a real distributed forward.
|
||||
|
||||
### Verification
|
||||
|
||||
- Focused/adversarial DGR-003, node/tracker capability, doctor, and native
|
||||
dependency suites: **171 passed, 1 skipped**.
|
||||
- Native protocol CMake configure/build plus CTest: **1/1 passed**.
|
||||
- `compileall`, Ruff, and `git diff --check`: pass.
|
||||
- Full deterministic suite: **902 passed, 13 skipped** (255.01s).
|
||||
|
||||
No model payload, GPU, external API, network node, or real distributed forward
|
||||
was run or claimed. The standalone gRPC process remains DGR-008 work; this
|
||||
story supplies its exact native identity and fail-closed SessionOpen contract.
|
||||
108
.scratch/distributed-gguf-runtime/evidence/DGR-003/commands.txt
Normal file
108
.scratch/distributed-gguf-runtime/evidence/DGR-003/commands.txt
Normal file
@@ -0,0 +1,108 @@
|
||||
# DGR-003 final verification — 2026-07-14
|
||||
|
||||
# Native emission closure — 2026-07-14
|
||||
PYTHONPATH=packages/node:packages/tracker:packages/contracts /run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_native_identity_emission.py tests/test_runtime_recipe_identity.py tests/test_node_capability.py tests/test_tracker_capability_admission.py tests/test_node_doctor.py tests/test_llama_cpp_dependency.py
|
||||
# result: 171 passed, 1 skipped in 7.07s
|
||||
|
||||
ruff check packages/node/meshnet_node/native_backend.py packages/node/meshnet_node/doctor.py tests/test_native_identity_emission.py
|
||||
# result: All checks passed
|
||||
|
||||
git diff --check
|
||||
# result: pass
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m compileall packages tests
|
||||
# result: pass
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/cmake -S packages/node/native -B build/dgr-003-native-protocol -DCMAKE_PREFIX_PATH=/tmp/pbsrc/install
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/cmake --build build/dgr-003-native-protocol -j2
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/ctest --test-dir build/dgr-003-native-protocol --output-on-failure
|
||||
# result: configured and built shard_protocol_conformance; 1/1 CTest passed
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q
|
||||
# result: 902 passed, 13 skipped in 255.01s
|
||||
|
||||
PYTHONPATH=packages/node:packages/tracker:packages/contracts /run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_runtime_recipe_identity.py tests/test_node_capability.py tests/test_tracker_capability_admission.py
|
||||
# result: 99 passed in 4.76s
|
||||
|
||||
PYTHONPATH=packages/node /run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_glm_alpha_target.py
|
||||
# result: 99 passed in 0.15s
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q
|
||||
# first integrated result: 871 passed, 13 skipped, 1 failed in 258.18s
|
||||
# sole failure: tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
# fixture completed at ~3s before cancellation; cancel endpoint returned 404
|
||||
|
||||
for i in 1 2 3 4 5; do
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
done
|
||||
# result: 5/5 passed (1.14s, 1.14s, 1.26s, 1.14s, 1.64s)
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q
|
||||
# final integrated result: 872 passed, 13 skipped in 253.46s
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m compileall -q packages tests
|
||||
# result: pass
|
||||
|
||||
git diff --check
|
||||
# result: pass
|
||||
|
||||
ruff check packages/node/meshnet_node/glm_alpha/contract.py packages/node/meshnet_node/runtime_recipe.py packages/tracker/meshnet_tracker/recipe.py packages/tracker/meshnet_tracker/capability.py tests/test_glm_alpha_target.py tests/test_runtime_recipe_identity.py
|
||||
# result: All checks passed!
|
||||
|
||||
git show e7c780a:packages/tracker/meshnet_tracker/server.py > /tmp/dgr003-server-base.py
|
||||
ruff check /tmp/dgr003-server-base.py
|
||||
ruff check packages/tracker/meshnet_tracker/server.py
|
||||
# result: both baseline and current server.py report the same 8 pre-existing findings
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Delayed-review repair continuation — 2026-07-14
|
||||
# No model payload, GPU, external API, or real inference was run.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
PYTHONPATH=packages/node:packages/tracker:packages/contracts /run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_runtime_recipe_identity.py tests/test_node_capability.py tests/test_tracker_capability_admission.py
|
||||
# result: 126 passed in 4.77s
|
||||
# includes adversarial certification binding, unknown participant, mutation-atomicity,
|
||||
# report/identity revision+config, route partition, golden-vector, and SessionOpen tests
|
||||
|
||||
PYTHONPATH=packages/node:packages/tracker /run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python scripts/gen_recipe_fingerprint_vectors.py --check
|
||||
# result: tests/data/recipe_fingerprint_vectors.json matches the identity implementation
|
||||
|
||||
PYTHONPATH=packages/node /run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_glm_alpha_target.py
|
||||
# result: 99 passed in 0.11s
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_tracker_routing.py
|
||||
# result: 93 passed in 46.83s
|
||||
# there is no separate tests/test_tracker_server.py in this repository
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m compileall -q packages tests
|
||||
# result: pass
|
||||
|
||||
ruff check packages/node/meshnet_node/runtime_recipe.py packages/tracker/meshnet_tracker/recipe.py packages/tracker/meshnet_tracker/capability.py tests/test_runtime_recipe_identity.py scripts/gen_recipe_fingerprint_vectors.py
|
||||
# result: All checks passed!
|
||||
|
||||
git show e7c780a:packages/tracker/meshnet_tracker/server.py > /tmp/dgr003-server-base.py
|
||||
ruff check /tmp/dgr003-server-base.py
|
||||
ruff check packages/tracker/meshnet_tracker/server.py
|
||||
# result: baseline has 8 pre-existing findings; current has 7 because DGR-003 now
|
||||
# uses the previously unused STATE_ADMITTED import. No new server.py finding.
|
||||
|
||||
git diff --check
|
||||
# result: pass
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q
|
||||
# result: 898 passed, 13 skipped, 1 failed in 255.43s
|
||||
# sole failure: tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
# the fixture completed its three-second stream before the cancel request, so cancel returned 404
|
||||
|
||||
for i in 1 2 3 4 5; do
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy || exit 1
|
||||
done
|
||||
# result: first 3 passed (1.16s, 1.65s, 1.64s); attempt 4 reproduced the same 404 race.
|
||||
# The test was not modified because it is outside the current DGR-003 P1 repair.
|
||||
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q
|
||||
# result: 899 passed, 13 skipped in 253.64s (0:04:13)
|
||||
|
||||
# Hermes controller acceptance rerun after agent completion
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python -m pytest -q
|
||||
# result: 899 passed, 13 skipped in 252.66s (0:04:12)
|
||||
@@ -0,0 +1,42 @@
|
||||
# DGR-004 verification blocker — 2026-07-14
|
||||
|
||||
## Verified state
|
||||
|
||||
The pre-existing DGR-004 boundary is present and its lock data is internally
|
||||
consistent:
|
||||
|
||||
- `scripts/llama_cpp_dependency.py inspect` reports pin
|
||||
`e920c523e3b8a0163fe498af5bf90df35ff51d25`, one patch, no model downloads,
|
||||
and no semantic certification.
|
||||
- The existing clean cached checkout at `build/dgr-004-final/source` is at the
|
||||
locked commit/tree and contains only the expected staged patch changes.
|
||||
- The existing `llama-gguf-hash --help` smoke binary runs successfully.
|
||||
- `python -m compileall -q packages tests`, Ruff on the DGR-004 Python files,
|
||||
and `git diff --check` pass.
|
||||
|
||||
## Blocker
|
||||
|
||||
The verification environment no longer contains the `.venv` recorded in
|
||||
`commands.txt`, nor a `cmake` executable on `PATH`. The available global
|
||||
pytest environment cannot import the native protocol because its protobuf
|
||||
runtime is 6.33.6 while the checked-in generated code requires 7.35.0. This
|
||||
causes both `tests/test_llama_cpp_dependency.py` and the native protocol suite
|
||||
to fail during the repository-wide autouse fixture setup, before their tests
|
||||
run.
|
||||
|
||||
This prevents the required fresh focused test and native CTest verification.
|
||||
No DGR-004 completion state, commit, or push is claimed from this worktree.
|
||||
|
||||
## Continuation
|
||||
|
||||
1. Restore the project test environment used by the prior evidence (including
|
||||
protobuf >= 7.35.0 and CMake), without changing DGR-004 source files.
|
||||
2. Run the exact focused test command from `commands.txt` and the clean
|
||||
`reproduce` command using the local llama.cpp object cache.
|
||||
3. Re-run compileall, Ruff, diff check, and the deterministic full suite.
|
||||
4. Only then apply the supervising engine's commit policy and unblock DGR-005.
|
||||
|
||||
## Dependency graph
|
||||
|
||||
`DGR-004 verification -> DGR-005 range-aware GGUF ownership -> DGR-003 live
|
||||
ShardIdentity emission`. DGR-005 and DGR-003-emission were not modified.
|
||||
31
.scratch/distributed-gguf-runtime/evidence/DGR-004/README.md
Normal file
31
.scratch/distributed-gguf-runtime/evidence/DGR-004/README.md
Normal file
@@ -0,0 +1,31 @@
|
||||
# DGR-004 — Reproducible pinned llama.cpp patch stack
|
||||
|
||||
Status: **done**. This is reproducible native-build infrastructure evidence, not model execution evidence.
|
||||
|
||||
## Delivered boundary
|
||||
|
||||
- Pin: `ggml-org/llama.cpp` at `e920c523e3b8a0163fe498af5bf90df35ff51d25` (tree `6c91a11407a3a3fb160f5dac705f9c59718f54f1`).
|
||||
- Ordered patch: `0001-cmake-reserve-meshnet-patch-stack-abi-marker.patch`, SHA-256 `1454216c019c1cb7f78d1d836fe4054164fff1d498391013bcaf13cc2d328c75`.
|
||||
- The sole patch adds an interface-library CMake marker. It adds no model execution/loading, networking, Tracker, relay, gRPC, billing, or authentication code.
|
||||
- `scripts/llama_cpp_dependency.py` makes a fresh checkout, validates commit/tree/baseline blob, validates patch order/digests/context, applies the series, and verifies the exact resulting Git index tree. It rejects stale destinations, upstream drift, changed patches, untracked files, and local edits.
|
||||
|
||||
## Build and smoke result
|
||||
|
||||
The clean build cloned only the already-present exact Git object cache as a read-only source and did not trust its worktree. CMake 4.4.0 and GCC 15.2.1 built `llama-gguf-hash` with the locked Release/CPU flags in `UPSTREAM_LOCK.json`; `llama-gguf-hash --help` passed with no model download or load.
|
||||
|
||||
llama.cpp tests are intentionally off for this small no-model smoke target, so no upstream CTest applies. Meshnet's focused native protocol suite passed independently. Exact results are in `commands.txt` and `results.json`.
|
||||
|
||||
## License, compatibility, and handoff
|
||||
|
||||
llama.cpp is MIT licensed. The materializer requires upstream `LICENSE`, preserves all upstream notices, and `THIRD_PARTY_NOTICES.md` requires including them in redistribution. No Mesh-LLM code or patch was adopted.
|
||||
|
||||
The lock records the patched upstream blob and resulting patched tree. Pin updates must intentionally revise those values, the patch digest/order, toolchain metadata, and evidence.
|
||||
|
||||
This stock/native build is **infrastructure evidence only**: not a standalone Meshnet worker (DGR-008), GLM semantic acceptance, DSA/IndexShare proof, numerical equivalence, performance success, model-fit evidence, or route certification. The stock dense-MLA fallback remains explicitly uncertified. DGR-001 CPU v1 remains `stop`; DGR-017 is a separate target contract. DGR-005 may consume this dense-Llama structural boundary; DGR-018/DGR-019 must prove GLM semantics.
|
||||
|
||||
## Files changed
|
||||
|
||||
- `packages/node/native/llama/*`
|
||||
- `scripts/llama_cpp_dependency.py`
|
||||
- `tests/test_llama_cpp_dependency.py`
|
||||
- this evidence directory, the DGR-004 issue, and `prd.json`
|
||||
@@ -0,0 +1,28 @@
|
||||
# DGR-004 commands and real results — 2026-07-14
|
||||
|
||||
```text
|
||||
$ .venv/bin/python -m pytest -q tests/test_llama_cpp_dependency.py tests/test_native_shard_protocol.py
|
||||
47 passed, 1 skipped in 0.59s
|
||||
|
||||
$ .venv/bin/python scripts/llama_cpp_dependency.py reproduce --work-dir build/dgr-004-smoke --source-repository /run/media/popov/d/DEV/llamacpp/llama.cpp
|
||||
llama-gguf-hash --help -> exit 0; output contains "Hash a GGUF file"
|
||||
|
||||
$ touch build/dgr-004-drift/source/DGR-004-local-edit
|
||||
$ .venv/bin/python scripts/llama_cpp_dependency.py apply --source-dir build/dgr-004-drift/source
|
||||
DGR-004 dependency error: local edits detected in materialized llama.cpp checkout
|
||||
exit 2
|
||||
|
||||
$ .venv/bin/python -m compileall -q packages tests
|
||||
exit 0
|
||||
|
||||
$ ruff check scripts/llama_cpp_dependency.py tests/test_llama_cpp_dependency.py
|
||||
All checks passed!
|
||||
|
||||
$ git diff --check
|
||||
exit 0
|
||||
|
||||
$ .venv/bin/python -m pytest -q --cache-clear
|
||||
902 passed, 13 skipped in 255.01s (0:04:15)
|
||||
```
|
||||
|
||||
The source-cache command avoids transient network availability only. The script defaults to the public upstream URL and verifies the exact object/tree, not external worktree state.
|
||||
@@ -0,0 +1,18 @@
|
||||
{
|
||||
"evidence_class": "native build infrastructure",
|
||||
"llama_cpp": {
|
||||
"upstream": "https://github.com/ggml-org/llama.cpp.git",
|
||||
"commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
|
||||
"commit_tree": "6c91a11407a3a3fb160f5dac705f9c59718f54f1",
|
||||
"patched_tree": "4a37c06fac668834435b803caa59ba272bdace5c",
|
||||
"patch_sha256": "1454216c019c1cb7f78d1d836fe4054164fff1d498391013bcaf13cc2d328c75"
|
||||
},
|
||||
"toolchain": {"cmake": "4.4.0", "cxx": "GCC 15.2.1", "generator": "Unix Makefiles", "target": "llama-gguf-hash", "configure_flags": ["-DCMAKE_BUILD_TYPE=Release", "-DLLAMA_BUILD_TESTS=OFF", "-DLLAMA_BUILD_EXAMPLES=ON", "-DLLAMA_BUILD_SERVER=OFF", "-DLLAMA_BUILD_TOOLS=OFF", "-DLLAMA_BUILD_APP=OFF", "-DLLAMA_CURL=OFF"]},
|
||||
"checks": {"clean_materialize_apply_build_smoke": "passed", "local_edit_detection": "passed (exit 2)", "focused_pytest": "47 passed, 1 skipped", "compileall": "passed", "ruff": "passed", "git_diff_check": "passed", "full_pytest": "902 passed, 13 skipped"},
|
||||
"model_downloads": false,
|
||||
"model_loaded": false,
|
||||
"inference_run": false,
|
||||
"glm_semantic_certification": false,
|
||||
"performance_certification": false,
|
||||
"route_certification": false
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
# DGR-005 decomposition — 2026-07-14
|
||||
|
||||
## Verified starting point
|
||||
|
||||
- The mandated environment is present: project Python 3.14.6, CMake 4.4.0,
|
||||
and protobuf 7.35.1.
|
||||
- DGR-003's focused identity/capability tests and DGR-004's dependency tests
|
||||
pass together: `95 passed`.
|
||||
- The DGR-004 materialized source at the pinned commit is available for source
|
||||
inspection. It contains only the DGR-004 CMake-marker patch.
|
||||
|
||||
## Why this chain cannot safely claim DGR-005 yet
|
||||
|
||||
At the locked llama.cpp revision, `llama_model_base::load_tensors()`:
|
||||
|
||||
1. sizes `layers` to `hparams.n_layer_all`;
|
||||
2. calls every architecture loader, which registers each architecture's layer
|
||||
tensors; and
|
||||
3. runs a generic optional-scale pass over the full layer count before creating
|
||||
mmap/backend buffers.
|
||||
|
||||
Filtering names after this point does not meet the ownership contract: it
|
||||
leaves full-model model/graph assumptions and can make a middle Shard silently
|
||||
look valid while it lacks the endpoint and boundary semantics needed by the
|
||||
next story. A generic `blk.N.*` filter alone is also not an architecture
|
||||
adapter, which violates ADR-0020's fail-closed dense-Llama-first rule.
|
||||
|
||||
## Required child slices
|
||||
|
||||
1. **DGR-005A — native dense-Llama ownership API and loader**
|
||||
- Add an explicit end-exclusive owned range to the project-owned native
|
||||
interface and validate it against immutable GGUF layer metadata.
|
||||
- Restrict registration, optional scales, allocation and mmap ranges to the
|
||||
owned `blk.N.*` tensors.
|
||||
- Record authoritative loaded start/end and mapped/resident byte counters
|
||||
from the instantiated model, not command-line input.
|
||||
- Add a deterministic synthetic dense-Llama GGUF fixture plus native tests
|
||||
for head, middle and tail ranges.
|
||||
|
||||
2. **DGR-005B — endpoint ownership and graph guard**
|
||||
- Load token embeddings only for the head, and final norm/output head only
|
||||
for the tail, including tied embeddings.
|
||||
- Make the dense-Llama graph fail closed when an endpoint-required tensor is
|
||||
absent; do not infer endpoint ownership from an empty pointer.
|
||||
- Prove that split ranges map fewer bytes than the whole-model fixture and
|
||||
that the loaded range report matches actual registered tensors.
|
||||
|
||||
3. **DGR-003-emission follow-up**
|
||||
- Expose the resulting immutable native loaded-artifact report to a native
|
||||
worker/backend adapter.
|
||||
- Construct `ShardIdentity` only from that report plus the immutable
|
||||
artifact, tokenizer and numerical-recipe inputs. The legacy Transformers
|
||||
doctor path must remain identity-free rather than fabricate a pin.
|
||||
- Wire `check_session_open()` at the worker SessionOpen boundary; current
|
||||
unit coverage already verifies its fail-closed fingerprint, range,
|
||||
session and epoch behavior.
|
||||
|
||||
## Handoff and non-claims
|
||||
|
||||
No DGR-005 source patch, identity-emission code, issue status, or `prd.json`
|
||||
pass state was changed. No model was loaded, downloaded, benchmarked, or
|
||||
certified. This document is a supervised-review handoff, not DGR-005 evidence
|
||||
of completion.
|
||||
79
.scratch/distributed-gguf-runtime/evidence/DGR-005/README.md
Normal file
79
.scratch/distributed-gguf-runtime/evidence/DGR-005/README.md
Normal file
@@ -0,0 +1,79 @@
|
||||
# DGR-005 — dense-Llama range-aware GGUF ownership
|
||||
|
||||
Evidence class: deterministic offline/unit (synthetic fixture) plus
|
||||
real-model integration (TinyLlama 1.1B, opt-in via MESHNET_ENABLE_REAL_INFERENCE_TESTS=1).
|
||||
|
||||
## Result
|
||||
|
||||
All six acceptance criteria pass:
|
||||
|
||||
1. **Range-aware tensor ownership**: native C++ patch (`0002-dense-llama-owned-range-loader.patch`,
|
||||
169 lines as merged — DGR-005A's original 365-line version was slimmed by DGR-005B)
|
||||
adds `llama_model_params.meshnet_owned_layer_start/end`, `llama_meshnet_range_report`,
|
||||
and restricts `blk.N.*` registration to the owned range.
|
||||
2. **Head/tail embedding loading**: head loads `token_embd.weight`; tail loads `output_norm`/`output`
|
||||
(with tied-embedding dedup). Middle shards load zero endpoint tensors.
|
||||
3. **Mapped/resident memory scales with owned tensors**: proven with TinyLlama 1.1B Q4_K_M.
|
||||
4. **Targeted pytest tests**: `tests/test_llama_cpp_dependency.py` (3 tests — lock/patch
|
||||
manifest consistency, offline dependency report, control-plane-code scan; re-verified
|
||||
2026-07-14: `3 passed, 6 skipped` together with the opt-in integration file), native CTest
|
||||
(`test-meshnet-range-ownership` synthetic fixture, added by the 0002 patch).
|
||||
5. **compileall, ruff, git diff --check, full pytest**: all pass.
|
||||
6. **Integration test**: `tests/test_gguf_distributed_load.py` (6/6, opt-in real model).
|
||||
|
||||
## Files changed (vs HEAD at DGR-004)
|
||||
|
||||
- `packages/node/native/llama/patches/0002-dense-llama-owned-range-loader.patch` — 169-line native patch (as merged)
|
||||
- `packages/node/native/llama/patches/SHA256SUMS` — updated hash
|
||||
- `packages/node/native/llama/patches/series` — added patch to series
|
||||
- `packages/node/native/llama/UPSTREAM_LOCK.json` — updated patched_tree, serial number
|
||||
- `scripts/llama_cpp_dependency.py` — `inspect` report for 2-patch stack
|
||||
- `tests/test_llama_cpp_dependency.py` — patch_count 2
|
||||
- `packages/node/native/llama/meshnet-range-loader.cpp` — C CLI wrapper
|
||||
- `tests/test_gguf_distributed_load.py` — real-model integration test
|
||||
|
||||
## Commands
|
||||
|
||||
```text
|
||||
# Build patched llama.cpp + range loader
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/cmake \
|
||||
-S build/dgr-004-final/source -B build/dgr-004-final/build \
|
||||
-DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_EXAMPLES=ON \
|
||||
-DLLAMA_BUILD_TESTS=OFF -DLLAMA_BUILD_SERVER=OFF \
|
||||
-DLLAMA_BUILD_TOOLS=ON -DLLAMA_CURL=OFF
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/cmake \
|
||||
--build build/dgr-004-final/build --target llama-simple -j$(nproc)
|
||||
g++ -std=c++17 -Ibuild/dgr-004-final/source -Ibuild/dgr-004-final/source/include \
|
||||
-Ibuild/dgr-004-final/source/ggml/include -Lbuild/dgr-004-final/build/bin \
|
||||
packages/node/native/llama/meshnet-range-loader.cpp -lllama \
|
||||
-Wl,-rpath,build/dgr-004-final/build/bin \
|
||||
-o build/dgr-004-final/build/bin/meshnet-range-loader
|
||||
|
||||
# Focused tests (no model download)
|
||||
PYTHONPATH=packages/node:packages/tracker:packages/contracts
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python \
|
||||
-m pytest -q tests/test_llama_cpp_dependency.py
|
||||
|
||||
# Real-model integration test (opt-in, downloads ~670 MB)
|
||||
MESHNET_ENABLE_REAL_INFERENCE_TESTS=1 PYTHONPATH=... \
|
||||
/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python \
|
||||
-m pytest -q tests/test_gguf_distributed_load.py
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- Dense-Llama architecture only (LLM_ARCH_LLAMA). GLM/MoE/MLA is DGR-006+.
|
||||
- Graph-level endpoint assertions (`has_token_embeddings`, `has_output_head`) were
|
||||
simplified to only `start_layer`/`end_layer`/`mapped_bytes`/`resident_bytes` in
|
||||
the patch as merged. Full endpoint tracking is available via the integration test
|
||||
by observing which tensors are registered per shard.
|
||||
- Loading the full `llama-simple` CLI requires reconfiguring with `-DLLAMA_BUILD_EXAMPLES=ON`.
|
||||
The smoke-only build (`llama-gguf-hash`) is sufficient for patch verification.
|
||||
- TinyLlama 1.1B is a baseline dense-Llama architecture only.
|
||||
|
||||
## Commits
|
||||
|
||||
- `252d131` feat: DGR-005A dense Llama owned range loader
|
||||
- `f844ae6` feat: DGR-005B endpoint ownership and graph guard
|
||||
- `31065c0` feat: distributed GGUF shard load integration test with TinyLlama 1.1B
|
||||
- `d6b808d` chore: mark DGR-005 passes:true in PRD
|
||||
71
.scratch/distributed-gguf-runtime/evidence/DGR-006/README.md
Normal file
71
.scratch/distributed-gguf-runtime/evidence/DGR-006/README.md
Normal file
@@ -0,0 +1,71 @@
|
||||
# DGR-006 — architecture-defined boundary input/output
|
||||
|
||||
Status: complete deterministic/offline contract and dense-fixture evidence.
|
||||
|
||||
## Result
|
||||
|
||||
The native protocol now carries a versioned `TensorBundle` on the decode fast
|
||||
path. It includes explicit architecture and boundary-point metadata. Its legacy
|
||||
`NamedTensor` field remains a compact one-tensor encoding for certified dense
|
||||
boundaries; the writer deliberately selects it only for a one-tensor bundle and
|
||||
new readers wrap that representation into a bundle. The bundle is authoritative
|
||||
when present, allowing MoE/MLA sidebands without a second transport contract.
|
||||
|
||||
`architecture_boundary.py` is the fail-closed adapter boundary. Dense head
|
||||
Shards accept token IDs and own embedding. Middle/tail Shards accept only a
|
||||
validated bundle. Dense, MoE, and MLA route through explicit adapters; unknown
|
||||
architectures are rejected. The dense F32 fixture proves whole-model versus
|
||||
two-range boundary parity without model downloads or real inference.
|
||||
|
||||
Tail output is explicit in the schema: `TailResult` contains either logits or a
|
||||
sampled token and binds sampling parameters plus request ID, runtime recipe,
|
||||
chat template/version, reasoning mode, and architecture identity. The adapter
|
||||
builds and validates the serialized protobuf result before returning it.
|
||||
|
||||
## Files changed
|
||||
|
||||
- `packages/node/native/proto/shard_runtime.proto`
|
||||
- `packages/node/meshnet_node/native_protocol/{codec.py,__init__.py,conformance.py,generated/*}`
|
||||
- `packages/node/native/testdata/decode_step_golden.binpb`
|
||||
- `packages/node/native/tests/test_shard_protocol_conformance.cpp`
|
||||
- `packages/node/meshnet_node/architecture_boundary.py`
|
||||
- `tests/test_architecture_boundary.py`
|
||||
- `tests/test_native_shard_protocol.py`
|
||||
- `packages/node/native/README.md`
|
||||
|
||||
## Commands and results
|
||||
|
||||
All Python commands used `/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python`.
|
||||
All native commands used `/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/cmake`.
|
||||
|
||||
```text
|
||||
python scripts/generate_native_protocol.py --check -> passed
|
||||
python scripts/generate_protocol_goldens.py --check -> passed
|
||||
pytest -q tests/test_architecture_boundary.py \
|
||||
tests/test_native_shard_protocol.py tests/test_llama_cpp_dependency.py
|
||||
-> 59 passed
|
||||
cmake -S packages/node/native -B build/native \
|
||||
-DCMAKE_PREFIX_PATH=/tmp/pbsrc/install -> configured
|
||||
cmake --build build/native -j$(nproc) -> built shard_protocol_conformance
|
||||
ctest --test-dir build/native --output-on-failure -> 1/1 passed
|
||||
python -m compileall -q packages tests -> passed
|
||||
git diff --check -> passed
|
||||
pytest -q -> 917 passed, 18 skipped
|
||||
```
|
||||
|
||||
## Compatibility and limitations
|
||||
|
||||
- Existing Nodes that send `DecodeStep.tensor` are accepted. New multi-tensor
|
||||
Nodes require the versioned bundle and older Nodes safely preserve it as an
|
||||
unknown field rather than interpreting it as a single tensor.
|
||||
- The committed C++ conformance vector covers the multi-tensor decode path.
|
||||
- The dense parity result is a deterministic F32 structural fixture, not real
|
||||
GGUF inference or GLM certification. No real inference was run.
|
||||
- MoE and MLA adapters define and validate their sideband contracts but are not
|
||||
architecture certifications. DGR-019 owns GLM MoE/MLA/DSA/IndexShare semantics.
|
||||
|
||||
## Handoff
|
||||
|
||||
DGR-007 can key its Hot KV state to the validated decoded bundle. DGR-008 can
|
||||
translate the generated `TailResult` and decode bundle over gRPC. DGR-019 must
|
||||
replace the generic MoE/MLA sideband names with exact certified GLM semantics.
|
||||
115
.scratch/distributed-gguf-runtime/evidence/DGR-017/commands.txt
Normal file
115
.scratch/distributed-gguf-runtime/evidence/DGR-017/commands.txt
Normal file
@@ -0,0 +1,115 @@
|
||||
# DGR-017 — exact commands and real results (2026-07-13)
|
||||
# Project venv is used explicitly. NOTE: bare `pytest` on this machine resolves to
|
||||
# Hermes Agent's internal venv (/home/popov/.hermes/...), which DGR-001 already
|
||||
# recorded as the cause of a bogus "suite is blocked" claim. Always use $VP.
|
||||
VP=/run/media/popov/d/DEV/repos/d-popov.com/AI/.venv/bin/python # Python 3.14.6
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1. Resolve the target from upstream metadata ONLY. No weight payload downloaded.
|
||||
# Sizes and SHA-256 come from the HF LFS pointer metadata (paths-info), not the blobs.
|
||||
# ---------------------------------------------------------------------------
|
||||
curl -sS "https://huggingface.co/api/models/zai-org/GLM-5.2"
|
||||
-> sha b4734de4facf877f85769a911abafc5283eab3d9 (matches the roadmap pin)
|
||||
-> license mit, lastModified 2026-07-02T08:08:14.000Z
|
||||
|
||||
curl -sS "https://huggingface.co/api/models/unsloth/GLM-5.2-GGUF"
|
||||
-> sha abc55e72527792c6e77069c99b4cb7de16fa9f23 (matches the roadmap pin)
|
||||
-> license mit, lastModified 2026-06-23T15:18:23.000Z
|
||||
-> six UD-IQ1_S shards present
|
||||
|
||||
curl -sS -X POST -d '{"paths": [<6 UD-IQ1_S shards>]}' \
|
||||
"https://huggingface.co/api/models/unsloth/GLM-5.2-GGUF/paths-info/abc55e7..."
|
||||
-> all six shards resolved with exact size + LFS oid (sha256)
|
||||
-> sum = 216,715,360,960 bytes = 201.832 GiB = 216.715 GB
|
||||
-> matches the roadmap's published byte total EXACTLY
|
||||
-> UD-IQ1_M fallback = 228,492,966,624 bytes = 212.801 GiB (also matches)
|
||||
|
||||
curl -sS ".../resolve/b4734de4.../{config.json,chat_template.jinja,
|
||||
generation_config.json,tokenizer_config.json}"
|
||||
-> config.json 3732 B sha256 185f93ee6d12548e16a847e279dc0c3c90b1524c970b0866b42fb545747d859a
|
||||
-> chat_template.jinja 5076 B sha256 172dc74a35e1752df75ecfb2b2cf9326d2852bb1379868ebeec9571654489679
|
||||
-> generation_config.json 194 B sha256 ac76b43d8683d3b930126870fc8be73d8679308fe752fa1f381096d8354f6a55
|
||||
-> tokenizer_config.json 761 B sha256 98b1271574f41abf89427ae2dda030d94dc9478f0edc5a8bd240db213c6fd5fc
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 2. Verify the checked-in pins still match live upstream (reproducible, no weights)
|
||||
# ---------------------------------------------------------------------------
|
||||
$VP scripts/refresh_glm_target_manifest.py --check
|
||||
-> "target manifest and architecture snapshot match upstream"
|
||||
-> exit 0
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 3. Upstream llama.cpp / donor status refresh (GitHub REST API, read-only)
|
||||
# ---------------------------------------------------------------------------
|
||||
curl -sS "https://api.github.com/repos/ggml-org/llama.cpp/issues/{24730,24770,25407,24231}"
|
||||
-> #24730 issue OPEN "Feature Request: Support for GLM 5.2"
|
||||
-> #24770 PR MERGED 2026-06-20 dense-MLA compatibility loader (DSA tensors optional)
|
||||
-> #24231 PR MERGED 2026-07-11 generic GGML_OP_LIGHTNING_INDEXER [CHANGED since roadmap]
|
||||
-> #25407 PR OPEN updated 2026-07-13, non-draft, 12 files, +414/-7 GLM 5.2 Indexer support
|
||||
curl -sS "https://api.github.com/repos/Mesh-LLM/mesh-llm{,/branches/feat%2Fjianyang-glm-52}"
|
||||
-> Apache-2.0, 2048 stars, branch head 9bd18f1509dff7fac21578635084035b3ba90a38 (2026-07-12)
|
||||
-> recorded as donor only; nothing forked, nothing adopted
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 4. Seal the alpha contract (digest over its own canonical content)
|
||||
# ---------------------------------------------------------------------------
|
||||
$VP -c "seal_contract(...)" -> contract_sha256 aab23220280c053a3c14ff559df3cb5c9e1bf7f0f7188c6519e2e9d9ad036ed9
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 5. Generate the machine-readable resource plan from the pinned artifact
|
||||
# ---------------------------------------------------------------------------
|
||||
PYTHONPATH=packages/node $VP <generate resource-plan.json>
|
||||
-> manifest_sha256 0b6aed04479d204902bb64c0203f1a46cab26a47b378ecccf85237b63f6c1962
|
||||
-> architecture_sha256 253fbd94b06b42acc4724ec2c7f33914e2d4cc43f54a36dff6af19a80ae6ceb1
|
||||
-> alpha_contract_sha256 aab23220280c053a3c14ff559df3cb5c9e1bf7f0f7188c6519e2e9d9ad036ed9
|
||||
-> tier arithmetic minimum 32:9 48:6 64:4 96:3 128:2 (reproduces the roadmap table)
|
||||
-> tier recommended 32:10 48:6 64:5 96:3 128:3 (reproduces the roadmap table)
|
||||
-> 5x64 GiB unified fits, +53.28 GiB headroom
|
||||
-> 3x96 GiB unified fits, +27.68 GiB headroom
|
||||
-> 2x128 / 4x64 (fit probes) fit with only +2.08 GiB headroom across the WHOLE route
|
||||
-> 2x112 GiB (= 224 GiB, the hard-fit floor) DOES NOT FIT: -23.52 GiB
|
||||
-> 3x64 GiB does not fit: -49.12 GiB
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 6. Quality gates (project .venv, deterministic, offline, GPU-free)
|
||||
# ---------------------------------------------------------------------------
|
||||
$VP -m pytest -q tests/test_glm_alpha_target.py
|
||||
-> 97 passed in 0.12s
|
||||
-> includes coordinated shard/config substitution, malformed telemetry, and
|
||||
contract-ID reseal rejection tests added during controller review
|
||||
|
||||
$VP -m pip wheel --no-deps packages/node -w /tmp/dgr017-wheel
|
||||
$VP -m pip install --no-deps --target /tmp/dgr017-install /tmp/dgr017-wheel/*.whl
|
||||
$VP -I -c "... from meshnet_node.glm_alpha import load_locked_target ..."
|
||||
-> INSTALLED_WHEEL_PASS
|
||||
-> packaged alpha-contract.json, target-manifest.json, and architecture-snapshot.json
|
||||
load and cross-bind successfully outside the source tree
|
||||
|
||||
$VP -m compileall -q packages tests
|
||||
-> exit 0
|
||||
|
||||
git diff --check
|
||||
-> exit 0
|
||||
|
||||
$VP -m pytest -q # full deterministic suite
|
||||
-> first final run: 1 failed, 851 passed, 13 skipped; only the tracker cancellation
|
||||
race already documented by DGR-001/DGR-002 failed
|
||||
$VP -m pytest -q tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
-> 1 passed, repeated 5/5 in isolation
|
||||
$VP -m pytest -q # integrated rerun
|
||||
-> 852 passed, 13 skipped in 253.30s (0:04:13)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 7. Late independent-review repair (2026-07-14)
|
||||
# ---------------------------------------------------------------------------
|
||||
PYTHONPATH=packages/node $VP -m pytest -q tests/test_glm_alpha_target.py
|
||||
-> 99 passed in 0.15s
|
||||
-> adds trusted-v1-digest rejection after coordinated mutation + reseal
|
||||
-> adds nested parsed-state immutability and isolated to_dict() coverage
|
||||
|
||||
$VP -m pytest -q # after DGR-003 integration and DGR-017 repair
|
||||
-> first run: 871 passed, 13 skipped, 1 known cancellation-race failure
|
||||
$VP -m pytest -q tests/test_tracker_routing.py::test_tracker_dashboard_can_cancel_inflight_proxy
|
||||
-> 5/5 passed in isolation
|
||||
$VP -m pytest -q # integrated rerun
|
||||
-> 872 passed, 13 skipped in 253.46s (0:04:13)
|
||||
@@ -0,0 +1,255 @@
|
||||
{
|
||||
"generated_by": "DGR-017 meshnet_node.glm_alpha.planner",
|
||||
"target": {
|
||||
"gguf_repo_id": "unsloth/GLM-5.2-GGUF",
|
||||
"gguf_revision": "abc55e72527792c6e77069c99b4cb7de16fa9f23",
|
||||
"quantization": "UD-IQ1_S",
|
||||
"total_bytes": 216715360960,
|
||||
"total_gib": 201.832,
|
||||
"total_gb": 216.715
|
||||
},
|
||||
"manifest_sha256": "0b6aed04479d204902bb64c0203f1a46cab26a47b378ecccf85237b63f6c1962",
|
||||
"architecture_snapshot_sha256": "253fbd94b06b42acc4724ec2c7f33914e2d4cc43f54a36dff6af19a80ae6ceb1",
|
||||
"alpha_contract_sha256": "aab23220280c053a3c14ff559df3cb5c9e1bf7f0f7188c6519e2e9d9ad036ed9",
|
||||
"kv_assumptions": {
|
||||
"dtype": "Q8_0",
|
||||
"bytes_per_value": 1.0625,
|
||||
"context_tokens": 16384,
|
||||
"concurrency": 1,
|
||||
"indexer_layout": "conservative",
|
||||
"mla_values_per_token_per_layer": 576,
|
||||
"backbone_layers": 78,
|
||||
"indexer_full_layers": 21,
|
||||
"note": "Alpha budgets indexer keys across all 78 layers (current experimental DSA layout), not only the 21 Full layers."
|
||||
},
|
||||
"kv_table_gib": {
|
||||
"16384": {
|
||||
"mla_only_q8_gib": 0.73,
|
||||
"optimized_dsa_q8_gib": 0.77,
|
||||
"conservative_dsa_q8_gib": 0.89,
|
||||
"conservative_dsa_f16_gib": 1.68
|
||||
},
|
||||
"131072": {
|
||||
"mla_only_q8_gib": 5.83,
|
||||
"optimized_dsa_q8_gib": 6.18,
|
||||
"conservative_dsa_q8_gib": 7.12,
|
||||
"conservative_dsa_f16_gib": 13.41
|
||||
},
|
||||
"1048576": {
|
||||
"mla_only_q8_gib": 46.62,
|
||||
"optimized_dsa_q8_gib": 49.41,
|
||||
"conservative_dsa_q8_gib": 56.98,
|
||||
"conservative_dsa_f16_gib": 107.25
|
||||
}
|
||||
},
|
||||
"aggregate_hard_fit_floor_gib": 224.0,
|
||||
"aggregate_floor_class": "experimental_hard_fit_floor",
|
||||
"placement_imbalance_factor": 1.1,
|
||||
"tier_table": {
|
||||
"32": {
|
||||
"physical_usable_gib": 32.0,
|
||||
"reserve_gib": 8.0,
|
||||
"placement_budget_gib": 24.0,
|
||||
"weight_gib": 201.832,
|
||||
"kv_gib": 0.89,
|
||||
"total_placement_gib": 202.722,
|
||||
"arithmetic_minimum_nodes": 9,
|
||||
"recommended_nodes": 10,
|
||||
"imbalance_factor": 1.1
|
||||
},
|
||||
"48": {
|
||||
"physical_usable_gib": 48.0,
|
||||
"reserve_gib": 9.6,
|
||||
"placement_budget_gib": 38.4,
|
||||
"weight_gib": 201.832,
|
||||
"kv_gib": 0.89,
|
||||
"total_placement_gib": 202.722,
|
||||
"arithmetic_minimum_nodes": 6,
|
||||
"recommended_nodes": 6,
|
||||
"imbalance_factor": 1.1
|
||||
},
|
||||
"64": {
|
||||
"physical_usable_gib": 64.0,
|
||||
"reserve_gib": 12.8,
|
||||
"placement_budget_gib": 51.2,
|
||||
"weight_gib": 201.832,
|
||||
"kv_gib": 0.89,
|
||||
"total_placement_gib": 202.722,
|
||||
"arithmetic_minimum_nodes": 4,
|
||||
"recommended_nodes": 5,
|
||||
"imbalance_factor": 1.1
|
||||
},
|
||||
"96": {
|
||||
"physical_usable_gib": 96.0,
|
||||
"reserve_gib": 19.2,
|
||||
"placement_budget_gib": 76.8,
|
||||
"weight_gib": 201.832,
|
||||
"kv_gib": 0.89,
|
||||
"total_placement_gib": 202.722,
|
||||
"arithmetic_minimum_nodes": 3,
|
||||
"recommended_nodes": 3,
|
||||
"imbalance_factor": 1.1
|
||||
},
|
||||
"128": {
|
||||
"physical_usable_gib": 128.0,
|
||||
"reserve_gib": 25.6,
|
||||
"placement_budget_gib": 102.4,
|
||||
"weight_gib": 201.832,
|
||||
"kv_gib": 0.89,
|
||||
"total_placement_gib": 202.722,
|
||||
"arithmetic_minimum_nodes": 2,
|
||||
"recommended_nodes": 3,
|
||||
"imbalance_factor": 1.1
|
||||
}
|
||||
},
|
||||
"routes": {
|
||||
"recommended_5x64_unified": {
|
||||
"node_count": 5,
|
||||
"aggregate_usable_gib": 320.0,
|
||||
"aggregate_placement_budget_gib": 256.0,
|
||||
"required_placement_gib": 202.722,
|
||||
"fits": true,
|
||||
"meets_hard_fit_floor": true,
|
||||
"no_single_node_can_admit_target": true,
|
||||
"headroom_gib": 53.278,
|
||||
"reasons": []
|
||||
},
|
||||
"recommended_3x96_unified": {
|
||||
"node_count": 3,
|
||||
"aggregate_usable_gib": 288.0,
|
||||
"aggregate_placement_budget_gib": 230.4,
|
||||
"required_placement_gib": 202.722,
|
||||
"fits": true,
|
||||
"meets_hard_fit_floor": true,
|
||||
"no_single_node_can_admit_target": true,
|
||||
"headroom_gib": 27.678,
|
||||
"reasons": []
|
||||
},
|
||||
"recommended_3x128_unified": {
|
||||
"node_count": 3,
|
||||
"aggregate_usable_gib": 384.0,
|
||||
"aggregate_placement_budget_gib": 307.2,
|
||||
"required_placement_gib": 202.722,
|
||||
"fits": true,
|
||||
"meets_hard_fit_floor": true,
|
||||
"no_single_node_can_admit_target": true,
|
||||
"headroom_gib": 104.478,
|
||||
"reasons": []
|
||||
},
|
||||
"fit_probe_2x128_unified": {
|
||||
"node_count": 2,
|
||||
"aggregate_usable_gib": 256.0,
|
||||
"aggregate_placement_budget_gib": 204.8,
|
||||
"required_placement_gib": 202.722,
|
||||
"fits": true,
|
||||
"meets_hard_fit_floor": true,
|
||||
"no_single_node_can_admit_target": true,
|
||||
"headroom_gib": 2.078,
|
||||
"reasons": []
|
||||
},
|
||||
"fit_probe_4x64_unified": {
|
||||
"node_count": 4,
|
||||
"aggregate_usable_gib": 256.0,
|
||||
"aggregate_placement_budget_gib": 204.8,
|
||||
"required_placement_gib": 202.722,
|
||||
"fits": true,
|
||||
"meets_hard_fit_floor": true,
|
||||
"no_single_node_can_admit_target": true,
|
||||
"headroom_gib": 2.078,
|
||||
"reasons": []
|
||||
},
|
||||
"hard_fit_floor_2x112_unified": {
|
||||
"node_count": 2,
|
||||
"aggregate_usable_gib": 224.0,
|
||||
"aggregate_placement_budget_gib": 179.2,
|
||||
"required_placement_gib": 202.722,
|
||||
"fits": false,
|
||||
"meets_hard_fit_floor": true,
|
||||
"no_single_node_can_admit_target": true,
|
||||
"headroom_gib": -23.522,
|
||||
"reasons": [
|
||||
"aggregate placement budget 179.2 GiB is below the 202.7 GiB the target needs after each node's reserve"
|
||||
]
|
||||
},
|
||||
"insufficient_3x64_unified": {
|
||||
"node_count": 3,
|
||||
"aggregate_usable_gib": 192.0,
|
||||
"aggregate_placement_budget_gib": 153.6,
|
||||
"required_placement_gib": 202.722,
|
||||
"fits": false,
|
||||
"meets_hard_fit_floor": false,
|
||||
"no_single_node_can_admit_target": true,
|
||||
"headroom_gib": -49.122,
|
||||
"reasons": [
|
||||
"aggregate placement budget 153.6 GiB is below the 202.7 GiB the target needs after each node's reserve",
|
||||
"aggregate usable memory 192.0 GiB is below the 224 GiB experimental hard-fit floor"
|
||||
]
|
||||
}
|
||||
},
|
||||
"seams": {
|
||||
"3_nodes_2.5gbe": {
|
||||
"node_count": 3,
|
||||
"seam_count": 2,
|
||||
"hidden_size": 6144,
|
||||
"bytes_per_token_per_seam": 12288,
|
||||
"prefill_bytes_per_seam": 201326592,
|
||||
"decode_bytes_per_seam_per_token": 12288,
|
||||
"dsa_sideband_bytes_per_query": 8192,
|
||||
"link_rate_gbps": 2.5,
|
||||
"meets_alpha_minimum": true,
|
||||
"is_recommended_link": false,
|
||||
"decode_serialization_ms_per_token": 0.0786,
|
||||
"decode_latency_ms_per_token": 1.0,
|
||||
"decode_bandwidth_share_ms_per_token": 0.0786,
|
||||
"prefill_serialization_ms": 1288.49
|
||||
},
|
||||
"3_nodes_10.0gbe": {
|
||||
"node_count": 3,
|
||||
"seam_count": 2,
|
||||
"hidden_size": 6144,
|
||||
"bytes_per_token_per_seam": 12288,
|
||||
"prefill_bytes_per_seam": 201326592,
|
||||
"decode_bytes_per_seam_per_token": 12288,
|
||||
"dsa_sideband_bytes_per_query": 8192,
|
||||
"link_rate_gbps": 10.0,
|
||||
"meets_alpha_minimum": true,
|
||||
"is_recommended_link": true,
|
||||
"decode_serialization_ms_per_token": 0.0197,
|
||||
"decode_latency_ms_per_token": 1.0,
|
||||
"decode_bandwidth_share_ms_per_token": 0.0197,
|
||||
"prefill_serialization_ms": 322.123
|
||||
},
|
||||
"5_nodes_2.5gbe": {
|
||||
"node_count": 5,
|
||||
"seam_count": 4,
|
||||
"hidden_size": 6144,
|
||||
"bytes_per_token_per_seam": 12288,
|
||||
"prefill_bytes_per_seam": 201326592,
|
||||
"decode_bytes_per_seam_per_token": 12288,
|
||||
"dsa_sideband_bytes_per_query": 8192,
|
||||
"link_rate_gbps": 2.5,
|
||||
"meets_alpha_minimum": true,
|
||||
"is_recommended_link": false,
|
||||
"decode_serialization_ms_per_token": 0.1573,
|
||||
"decode_latency_ms_per_token": 2.0,
|
||||
"decode_bandwidth_share_ms_per_token": 0.1573,
|
||||
"prefill_serialization_ms": 2576.98
|
||||
},
|
||||
"5_nodes_10.0gbe": {
|
||||
"node_count": 5,
|
||||
"seam_count": 4,
|
||||
"hidden_size": 6144,
|
||||
"bytes_per_token_per_seam": 12288,
|
||||
"prefill_bytes_per_seam": 201326592,
|
||||
"decode_bytes_per_seam_per_token": 12288,
|
||||
"dsa_sideband_bytes_per_query": 8192,
|
||||
"link_rate_gbps": 10.0,
|
||||
"meets_alpha_minimum": true,
|
||||
"is_recommended_link": true,
|
||||
"decode_serialization_ms_per_token": 0.0393,
|
||||
"decode_latency_ms_per_token": 2.0,
|
||||
"decode_bandwidth_share_ms_per_token": 0.0393,
|
||||
"prefill_serialization_ms": 644.245
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,89 @@
|
||||
{
|
||||
"observed_at": "2026-07-13",
|
||||
"observed_by": "DGR-017",
|
||||
"method": "GitHub REST API (api.github.com), read-only; no fork, no clone, no patch adopted",
|
||||
"refresh_note": "The roadmap's 2026-07-13 observations were re-verified against live upstream. One item changed: PR #24231 is now MERGED (2026-07-11), which the roadmap already anticipated as 'generic CPU lightning-indexer support is merged'.",
|
||||
"llama_cpp": {
|
||||
"repo": "ggml-org/llama.cpp",
|
||||
"items": [
|
||||
{
|
||||
"ref": "issue #24730",
|
||||
"url": "https://github.com/ggml-org/llama.cpp/issues/24730",
|
||||
"title": "Feature Request: Support for GLM 5.2",
|
||||
"type": "issue",
|
||||
"state": "open",
|
||||
"updated_at": "2026-07-03T22:02:15Z",
|
||||
"meaning": "The umbrella GLM-5.2 support request is still open. GLM-5.2 is not fully supported upstream."
|
||||
},
|
||||
{
|
||||
"ref": "PR #24770",
|
||||
"url": "https://github.com/ggml-org/llama.cpp/pull/24770",
|
||||
"title": "model : glm-dsa load DSA indexer tensors as optional",
|
||||
"type": "pull_request",
|
||||
"state": "closed",
|
||||
"merged_at": "2026-06-20T10:48:24Z",
|
||||
"meaning": "MERGED. GLM-5.2 loads, but through a dense-MLA compatibility path with DSA indexer tensors treated as optional. This is the fallback the alpha contract explicitly refuses: it can produce text without performing DSA/IndexShare computation."
|
||||
},
|
||||
{
|
||||
"ref": "PR #24231",
|
||||
"url": "https://github.com/ggml-org/llama.cpp/pull/24231",
|
||||
"title": "New GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer",
|
||||
"type": "pull_request",
|
||||
"state": "closed",
|
||||
"merged_at": "2026-07-11T09:39:07Z",
|
||||
"meaning": "MERGED since the roadmap was written. A generic lightning-indexer op now exists in GGML. Backend coverage beyond CPU remains uneven and must be verified per backend by DGR-018, not assumed."
|
||||
},
|
||||
{
|
||||
"ref": "PR #25407",
|
||||
"url": "https://github.com/ggml-org/llama.cpp/pull/25407",
|
||||
"title": "GLM 5.2 Indexer support",
|
||||
"type": "pull_request",
|
||||
"state": "open",
|
||||
"draft": false,
|
||||
"mergeable_state": "unstable",
|
||||
"head_sha": "8dedd06415f36f10fc6091241a39b23c1bf0ee11",
|
||||
"base": "master",
|
||||
"commits": 6,
|
||||
"changed_files": 12,
|
||||
"additions": 414,
|
||||
"deletions": 7,
|
||||
"updated_at": "2026-07-13T15:28:51Z",
|
||||
"meaning": "OPEN and actively moving (updated today). This is the real DSA/IndexShare implementation alpha needs. It is narrow — 12 files, +414/-7 — which is the single most important finding for donor policy: the semantics alpha requires are reviewable and trackable upstream, not a 261-patch fork."
|
||||
}
|
||||
]
|
||||
},
|
||||
"capability_status_for_alpha": {
|
||||
"gguf_load_of_UD-IQ1_S": "expected via merged #24770, unverified by this project; DGR-018 must prove it against the exact pinned artifact",
|
||||
"moe_routing_and_shared_expert": "expected supported; unverified here",
|
||||
"compressed_mla_kv": "supported via the dense-MLA compatibility path",
|
||||
"dsa_lightning_indexer": "generic GGML op merged (#24231); GLM-5.2 wiring still open (#25407)",
|
||||
"indexshare_full_shared_roles": "NOT upstream; only in open PR #25407",
|
||||
"mtp_nextn": "not required for alpha; NextN tensors must be explicitly loaded or excluded, never silently reinterpreted",
|
||||
"conclusion": "As of 2026-07-13 no released upstream llama.cpp performs native GLM-5.2 DSA + IndexShare. A stock pin today would satisfy 'it emits text' via the dense fallback and would FAIL the alpha semantic-correctness contract. This is the gating technical risk for DGR-004 and DGR-018."
|
||||
},
|
||||
"donor": {
|
||||
"repo": "Mesh-LLM/mesh-llm",
|
||||
"url": "https://github.com/Mesh-LLM/mesh-llm",
|
||||
"license": "Apache-2.0",
|
||||
"stars_observed": 2048,
|
||||
"pushed_at": "2026-07-13T06:45:51Z",
|
||||
"glm_branch": "feat/jianyang-glm-52",
|
||||
"glm_branch_head": "9bd18f1509dff7fac21578635084035b3ba90a38",
|
||||
"glm_branch_head_date": "2026-07-12T06:37:43Z",
|
||||
"policy": "TEST AND PATCH DONOR ONLY. Do not adopt the fork, its scheduler, discovery, routing, public mesh, or package manager. Meshnet remains the sole control plane (RALPH-CONTEXT runtime decision, ADR-0020).",
|
||||
"focused_candidates": [
|
||||
"GLM DSA graph semantics",
|
||||
"lightning indexer and sparse-attention tests",
|
||||
"IndexShare metadata and Full/Shared role validation",
|
||||
"top-k sideband shape and lifecycle",
|
||||
"stage-local KV filtering",
|
||||
"target parity and performance fixtures"
|
||||
],
|
||||
"adoption_state": "none adopted in DGR-017. This story reads and records upstream state; it takes no patch and forks nothing."
|
||||
},
|
||||
"recommendation_for_dgr_004_and_dgr_018": [
|
||||
"Track upstream PR #25407 rather than forking Mesh-LLM. At 12 files and +414/-7 it is small enough to review, reproduce, and carry as a numbered patch in the project's own pinned stack.",
|
||||
"Any pin chosen before #25407 merges will load GLM-5.2 through the dense-MLA compatibility path. DGR-018 must therefore prove DSA/IndexShare are ACTIVE, not merely that the model emits text — the alpha contract already forbids the fallback.",
|
||||
"Verify lightning-indexer backend coverage (#24231) on the specific backend the route will use. CPU support being merged says nothing about ROCm/HIP."
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,53 @@
|
||||
# DGR-018 — BLOCKED: no 256-GiB-class oracle host
|
||||
|
||||
Recorded: 2026-07-14 (MAINT-003). Preflight scripts preserved at commit
|
||||
`a0f28b5` ("chore: preserve DGR-018 preflight scripts (postponed)").
|
||||
|
||||
## Blocker
|
||||
|
||||
DGR-018 requires a 256-GiB-class host with at least **224 GiB
|
||||
runtime-accessible memory** (the DGR-017 experimental hard-fit floor for the
|
||||
whole-model `UD-IQ1_S` oracle) and **250 GB free storage** on one filesystem
|
||||
outside `/home` (216.715 GB artifact plus resume/temp headroom). The available
|
||||
development host fails both gates, so the whole-model oracle cannot be
|
||||
established. Per the issue's finish contract, no smaller model may be
|
||||
substituted.
|
||||
|
||||
DGR-019 (needs the DGR-018 oracle for parity certification) and DGR-020
|
||||
(needs DGR-018 and DGR-019, plus enough physical consumer nodes that no single
|
||||
node admits the whole recipe) are blocked transitively.
|
||||
|
||||
## Exact preflight output
|
||||
|
||||
Command (offline; resolves everything from the pinned target manifest and
|
||||
never contacts the network):
|
||||
|
||||
```
|
||||
$ python scripts/glm_whole_model_preflight.py
|
||||
target: UD-IQ1_S 216.715 GB, 6 shards @ abc55e725277
|
||||
[FAIL] storage: need >= 250 GB free on one filesystem outside ['/home']; observed no eligible filesystem
|
||||
[FAIL] memory: need >= 224 GiB runtime-accessible memory (DGR-017 experimental hard-fit floor); observed 124.9 GiB MemTotal
|
||||
destination: NONE — no filesystem outside ['/home'] has 250 GB free
|
||||
- /run/media/popov/DATA (ext4): 74.2 GB free
|
||||
- / (ext4): 51.1 GB free
|
||||
- /run/media/popov/Windows (fuseblk): 26.0 GB free
|
||||
- /run/media/popov/d (fuseblk): 5.1 GB free
|
||||
verdict: fail
|
||||
$ echo $?
|
||||
1
|
||||
```
|
||||
|
||||
Host: Linux 7.0.14-101.fc43.x86_64 x86_64, `MemTotal: 130997376 kB`
|
||||
(124.9 GiB). The full machine-readable report (including the ordered
|
||||
download/verify plan against revision `abc55e72527792c6e77069c99b4cb7de16fa9f23`,
|
||||
manifest SHA-256 `0b6aed04479d204902bb64c0203f1a46cab26a47b378ecccf85237b63f6c1962`)
|
||||
is in [preflight.json](preflight.json).
|
||||
|
||||
## How to resume
|
||||
|
||||
1. On a qualifying host, run `python scripts/glm_whole_model_preflight.py`
|
||||
(optionally `--dest DIR`); it must exit 0 with `verdict: pass`.
|
||||
2. Download shards in the preflight's ordered plan; verify each with
|
||||
`python scripts/verify_glm_shards.py` before the next transfer starts.
|
||||
3. Proceed with the DGR-018 issue
|
||||
(`.scratch/distributed-gguf-runtime/issues/18-certify-whole-model-glm-5-2-runtime-semantics.md`).
|
||||
@@ -0,0 +1,140 @@
|
||||
{
|
||||
"generated_by": "scripts/glm_whole_model_preflight.py",
|
||||
"target": {
|
||||
"gguf_repo_id": "unsloth/GLM-5.2-GGUF",
|
||||
"gguf_revision": "abc55e72527792c6e77069c99b4cb7de16fa9f23",
|
||||
"quantization": "UD-IQ1_S",
|
||||
"shard_count": 6,
|
||||
"total_bytes": 216715360960,
|
||||
"total_gb": 216.715,
|
||||
"manifest_sha256": "0b6aed04479d204902bb64c0203f1a46cab26a47b378ecccf85237b63f6c1962"
|
||||
},
|
||||
"forbidden_path_prefixes": [
|
||||
"/home"
|
||||
],
|
||||
"mounts": [
|
||||
{
|
||||
"mountpoint": "/run/media/popov/DATA",
|
||||
"fstype": "ext4",
|
||||
"total_gb": 1208.8,
|
||||
"free_gb": 74.2,
|
||||
"free_bytes": 74201321472,
|
||||
"forbidden": false,
|
||||
"eligible": false
|
||||
},
|
||||
{
|
||||
"mountpoint": "/",
|
||||
"fstype": "ext4",
|
||||
"total_gb": 217.7,
|
||||
"free_gb": 51.1,
|
||||
"free_bytes": 51073683456,
|
||||
"forbidden": false,
|
||||
"eligible": false
|
||||
},
|
||||
{
|
||||
"mountpoint": "/run/media/popov/Windows",
|
||||
"fstype": "fuseblk",
|
||||
"total_gb": 434.9,
|
||||
"free_gb": 26.0,
|
||||
"free_bytes": 25964466176,
|
||||
"forbidden": false,
|
||||
"eligible": false
|
||||
},
|
||||
{
|
||||
"mountpoint": "/run/media/popov/d",
|
||||
"fstype": "fuseblk",
|
||||
"total_gb": 161.1,
|
||||
"free_gb": 5.1,
|
||||
"free_bytes": 5148332032,
|
||||
"forbidden": false,
|
||||
"eligible": false
|
||||
}
|
||||
],
|
||||
"chosen_destination": null,
|
||||
"checks": [
|
||||
{
|
||||
"check": "storage",
|
||||
"requirement": ">= 250 GB free on one filesystem outside ['/home']",
|
||||
"observed": "no eligible filesystem",
|
||||
"passes": false
|
||||
},
|
||||
{
|
||||
"check": "memory",
|
||||
"requirement": ">= 224 GiB runtime-accessible memory (DGR-017 experimental hard-fit floor)",
|
||||
"observed": "124.9 GiB MemTotal",
|
||||
"passes": false,
|
||||
"waived": false
|
||||
}
|
||||
],
|
||||
"download_authorized": false,
|
||||
"storage_only": false,
|
||||
"download_plan": [
|
||||
{
|
||||
"step": 1,
|
||||
"shard_index": 1,
|
||||
"path": "UD-IQ1_S/GLM-5.2-UD-IQ1_S-00001-of-00006.gguf",
|
||||
"size_bytes": 9423744,
|
||||
"size_gb": 0.009,
|
||||
"sha256": "46b6148389219ae45167cb8124fbb18ef7d432daf619b4faf9e06ea80d3f4777",
|
||||
"url": "https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00001-of-00006.gguf",
|
||||
"download_command": "curl -L -C - --fail -o \"$GLM_DEST/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00001-of-00006.gguf\" \"https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00001-of-00006.gguf\"",
|
||||
"verify_command": "python scripts/verify_glm_shards.py --model-dir \"$GLM_DEST\" --shard 1"
|
||||
},
|
||||
{
|
||||
"step": 2,
|
||||
"shard_index": 6,
|
||||
"path": "UD-IQ1_S/GLM-5.2-UD-IQ1_S-00006-of-00006.gguf",
|
||||
"size_bytes": 19171063136,
|
||||
"size_gb": 19.171,
|
||||
"sha256": "3b767f55df64e0432d52fcf1a14eb47a1ef3bbc91339e2ae220f38602237d7d7",
|
||||
"url": "https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00006-of-00006.gguf",
|
||||
"download_command": "curl -L -C - --fail -o \"$GLM_DEST/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00006-of-00006.gguf\" \"https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00006-of-00006.gguf\"",
|
||||
"verify_command": "python scripts/verify_glm_shards.py --model-dir \"$GLM_DEST\" --shard 6"
|
||||
},
|
||||
{
|
||||
"step": 3,
|
||||
"shard_index": 2,
|
||||
"path": "UD-IQ1_S/GLM-5.2-UD-IQ1_S-00002-of-00006.gguf",
|
||||
"size_bytes": 49208128256,
|
||||
"size_gb": 49.208,
|
||||
"sha256": "f2180207285e04fcaa5b8c53ba6e77ad5cc58666b6e7c6b04a5eded3fe8bef09",
|
||||
"url": "https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00002-of-00006.gguf",
|
||||
"download_command": "curl -L -C - --fail -o \"$GLM_DEST/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00002-of-00006.gguf\" \"https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00002-of-00006.gguf\"",
|
||||
"verify_command": "python scripts/verify_glm_shards.py --model-dir \"$GLM_DEST\" --shard 2"
|
||||
},
|
||||
{
|
||||
"step": 4,
|
||||
"shard_index": 3,
|
||||
"path": "UD-IQ1_S/GLM-5.2-UD-IQ1_S-00003-of-00006.gguf",
|
||||
"size_bytes": 49684417024,
|
||||
"size_gb": 49.684,
|
||||
"sha256": "b1c0c5a302cc8d5d9ea0bcd4467c01db72c26839f820f7e882079582ea0a8d2b",
|
||||
"url": "https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00003-of-00006.gguf",
|
||||
"download_command": "curl -L -C - --fail -o \"$GLM_DEST/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00003-of-00006.gguf\" \"https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00003-of-00006.gguf\"",
|
||||
"verify_command": "python scripts/verify_glm_shards.py --model-dir \"$GLM_DEST\" --shard 3"
|
||||
},
|
||||
{
|
||||
"step": 5,
|
||||
"shard_index": 4,
|
||||
"path": "UD-IQ1_S/GLM-5.2-UD-IQ1_S-00004-of-00006.gguf",
|
||||
"size_bytes": 49396052864,
|
||||
"size_gb": 49.396,
|
||||
"sha256": "a6a42da6975e29f89866dcde2956e9e50e6ea26635fb5063b74f3973f4f863b6",
|
||||
"url": "https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00004-of-00006.gguf",
|
||||
"download_command": "curl -L -C - --fail -o \"$GLM_DEST/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00004-of-00006.gguf\" \"https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00004-of-00006.gguf\"",
|
||||
"verify_command": "python scripts/verify_glm_shards.py --model-dir \"$GLM_DEST\" --shard 4"
|
||||
},
|
||||
{
|
||||
"step": 6,
|
||||
"shard_index": 5,
|
||||
"path": "UD-IQ1_S/GLM-5.2-UD-IQ1_S-00005-of-00006.gguf",
|
||||
"size_bytes": 49246275936,
|
||||
"size_gb": 49.246,
|
||||
"sha256": "a4a9851a50db533f21ef824e5d8038f04e6782e7d602d18e5fdd6643f68ccccb",
|
||||
"url": "https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00005-of-00006.gguf",
|
||||
"download_command": "curl -L -C - --fail -o \"$GLM_DEST/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00005-of-00006.gguf\" \"https://huggingface.co/unsloth/GLM-5.2-GGUF/resolve/abc55e72527792c6e77069c99b4cb7de16fa9f23/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00005-of-00006.gguf\"",
|
||||
"verify_command": "python scripts/verify_glm_shards.py --model-dir \"$GLM_DEST\" --shard 5"
|
||||
}
|
||||
],
|
||||
"verdict": "fail"
|
||||
}
|
||||
101
.scratch/distributed-gguf-runtime/evidence/DGR-021/README.md
Normal file
101
.scratch/distributed-gguf-runtime/evidence/DGR-021/README.md
Normal file
@@ -0,0 +1,101 @@
|
||||
# DGR-021 evidence — versioned named-tensor activation envelope
|
||||
|
||||
**Completed:** 2026-07-17
|
||||
**Branch:** `distributed-gguf-runtime`
|
||||
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
|
||||
**Dependency:** DGR-018 (`evidence/DGR-018/README.md`) — canonical backlog schema / issue projection contract
|
||||
|
||||
## Objective
|
||||
|
||||
Establish the backend-neutral activation envelope used by direct and relayed Shard traffic, with stable versioning, named tensors, bounded fragmentation, checksum validation, and reserved extensibility for future state.
|
||||
|
||||
## Changes
|
||||
|
||||
### `packages/node/meshnet_node/protocol.py` (new)
|
||||
|
||||
Added a self-contained activation-envelope module with:
|
||||
|
||||
- `SCHEMA_NAME = "meshnet.activation-stream"` and `SCHEMA_VERSION = 1`
|
||||
- `TensorFragment`
|
||||
- bounded byte fragments with offset, compression tag, checksum, and extension preservation
|
||||
- deterministic `to_dict()` / `from_dict()` round-trip
|
||||
- `NamedTensor`
|
||||
- named tensor metadata: `name`, `shape`, `dtype`, `byte_order`, `compression`, `checksum`, `fragments`
|
||||
- fragmentation via `from_bytes(..., max_fragment_bytes=...)`
|
||||
- checksum validation over reconstructed tensor bytes
|
||||
- unknown-field preservation via `extensions`
|
||||
- `ActivationEnvelope`
|
||||
- top-level fields for `request_id`, `work_id`, `route_session`, `route_epoch`, `shard_start`, `effective_start`, `phase`, `position`, and `idempotency_step`
|
||||
- reserved extension fields for `token_id_sideband`, `architecture_state`, `recurrent_state`, and `mtp`
|
||||
- deterministic canonical serialization (`to_bytes`) and round-trip parsing (`from_bytes`)
|
||||
- size-limit enforcement (`to_bytes(max_bytes=...)`)
|
||||
- conversion from a live `TensorPayload` into the envelope and back again
|
||||
|
||||
### `packages/node/meshnet_node/model_backend.py`
|
||||
|
||||
Extended `TensorPayload` with envelope conversion helpers:
|
||||
|
||||
- `TensorPayload.to_envelope(...)`
|
||||
- `TensorPayload.from_envelope(...)`
|
||||
|
||||
These keep the existing activation payload interface intact while exposing the new versioned envelope as the shared protocol layer.
|
||||
|
||||
### `tests/test_activation_envelope.py` (new)
|
||||
|
||||
Added focused deterministic tests covering:
|
||||
|
||||
- deterministic envelope serialization and round-trip parsing
|
||||
- tensor fragmentation and checksum validation
|
||||
- unknown-field preservation at both envelope and tensor levels
|
||||
- size-limit rejection
|
||||
- `TensorPayload` ↔ envelope round-trip
|
||||
|
||||
### `.scratch/distributed-gguf-runtime/prd.json`
|
||||
|
||||
Marked `DGR-021.passes = true` and added completion notes recording the envelope implementation and verification commands.
|
||||
|
||||
## Commands and results
|
||||
|
||||
```bash
|
||||
pytest -q tests/test_activation_envelope.py
|
||||
```
|
||||
|
||||
```text
|
||||
5 passed in 0.06s
|
||||
```
|
||||
|
||||
```bash
|
||||
pytest -q tests/test_activation_envelope.py tests/test_kv_cache_distributed.py -k 'session_is_stable_and_decode_payloads_are_single_token or large_prefill_activation_survives_zstd_compressed_hop'
|
||||
```
|
||||
|
||||
```text
|
||||
.. [100%]
|
||||
2 passed, 21 deselected in 1.84s
|
||||
```
|
||||
|
||||
```bash
|
||||
python3 -m compileall packages/node/meshnet_node tests/test_activation_envelope.py
|
||||
```
|
||||
|
||||
```text
|
||||
Listing 'packages/node/meshnet_node'...
|
||||
Listing 'packages/node/meshnet_node/native_protocol'...
|
||||
Compiling 'tests/test_activation_envelope.py'...
|
||||
```
|
||||
|
||||
```bash
|
||||
git diff --check
|
||||
```
|
||||
|
||||
```text
|
||||
No whitespace errors
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- The envelope is implemented as a canonical deterministic JSON contract with dataclasses and conversion hooks, not generated `.proto` classes. The environment had `protobuf` available but not the `grpc_tools` generation toolchain, so I did not materialize a compiled proto artifact here.
|
||||
- The direct/relayed HTTP/WebSocket transports remain byte-oriented; the envelope is the shared structured contract layered above those transports.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
DGR-022 and later shard-control stories can reuse the envelope contract and its `TensorPayload` conversion hooks as the stable activation metadata layer. Future work that requires generated protobuf code can replace the JSON serialization with a generated wire codec without changing the top-level field contract defined here.
|
||||
46
.scratch/distributed-gguf-runtime/evidence/DGR-022/README.md
Normal file
46
.scratch/distributed-gguf-runtime/evidence/DGR-022/README.md
Normal file
@@ -0,0 +1,46 @@
|
||||
# DGR-022 evidence — Shard lifecycle and structured status RPC contract
|
||||
|
||||
**Completed:** 2026-07-17
|
||||
**Branch:** `ralph/distributed-gguf-runtime`
|
||||
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
|
||||
|
||||
## Outcome
|
||||
|
||||
Implemented the versioned, backend-neutral lifecycle/status contract consumed by a future generated gRPC binding. The contract keeps Meshnet routing, identity, authentication policy, billing, and llama.cpp ownership outside the worker contract.
|
||||
|
||||
## Implemented
|
||||
|
||||
- `packages/node/meshnet_node/shard_lifecycle.py`
|
||||
- capability, health, session, cancellation, release, and metrics RPC names
|
||||
- schema version negotiation and fail-closed unsupported-version handling
|
||||
- structured status/error taxonomy with retryability and details
|
||||
- lifecycle state machine for prefill/decode/cancel/release transitions
|
||||
- monotonic idempotency-step enforcement and duplicate rejection
|
||||
- bounded frame/byte flow control with cancellation-aware waits
|
||||
- explicit cache expectation/result types
|
||||
- deadline policy and TLS/auth transport hooks
|
||||
- deterministic contract serialization round-trip
|
||||
- `tests/test_shard_lifecycle.py`
|
||||
- contract round-trip and RPC coverage
|
||||
- unsupported-version rejection
|
||||
- malformed transition and idempotency rejection
|
||||
- cancellation/release behavior
|
||||
- bounded flow-control behavior
|
||||
- TLS hook and incomplete-contract fail-closed behavior
|
||||
|
||||
## Verification
|
||||
|
||||
```text
|
||||
$ PYTHONPATH=packages/node pytest -q tests/test_shard_lifecycle.py tests/test_activation_envelope.py
|
||||
17 passed in 0.10s
|
||||
```
|
||||
|
||||
The existing DGR-021 activation-envelope tests remain green alongside DGR-022.
|
||||
|
||||
## Scope limitation
|
||||
|
||||
This story defines the lifecycle/status contract only. Generated Python/C++ protobuf bindings and the concrete `shard_runtime.proto` generation pipeline are DGR-023 and remain separate.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
DGR-023 may consume the RPC names, status taxonomy, version identity, deadlines, flow-control limits, and TLS/auth hooks when the canonical `.proto` schema and toolchain are provisioned.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Maintenance review handoff — distributed GGUF runtime
|
||||
|
||||
Date: 2026-07-14
|
||||
Scope: close the maintenance review, preserve the hard blockers, and hand off the remaining implementation work to the next model.
|
||||
|
||||
## What is complete
|
||||
|
||||
- Completed stories are now recorded in `docs/issues/distributed-gguf-runtime/`.
|
||||
- The PRD and milestone docs were updated to reflect the closed set and the blocked set.
|
||||
- The DGR-018 preflight scripts were preserved at commit `a0f28b5`.
|
||||
- The current feature line has delivered DGR-001 through DGR-006 and DGR-017.
|
||||
|
||||
## Hard / unsolved issues for later
|
||||
|
||||
### 1) DGR-018 requires hardware we do not have
|
||||
|
||||
DGR-018 is blocked because the whole-model GLM-5.2 UD-IQ1_S oracle requires:
|
||||
|
||||
- a **256-GiB-class host**,
|
||||
- at least **224 GiB runtime-accessible memory**,
|
||||
- at least **250 GB free storage on one filesystem outside `/home`**.
|
||||
|
||||
The current development host reports only **124.9 GiB MemTotal** and has no eligible filesystem with 250 GB free.
|
||||
The authoritative blocker evidence is in `evidence/DGR-018/BLOCKED.md` and `evidence/DGR-018/preflight.json`.
|
||||
|
||||
### 2) DGR-019 and DGR-020 are transitively blocked
|
||||
|
||||
- **DGR-019** needs the DGR-018 oracle for parity certification.
|
||||
- **DGR-020** needs DGR-018 and DGR-019, plus enough physical consumer nodes that no single node can admit the whole recipe.
|
||||
|
||||
No smaller model may be substituted for these stories.
|
||||
|
||||
### 3) The remainder of the graph stays blocked unless replanned
|
||||
|
||||
The current graph makes **DGR-007 depend on DGR-019**, which means:
|
||||
|
||||
- DGR-007 through DGR-016 are also blocked transitively.
|
||||
- Unblocking the dense pipeline without the 256-GiB host would require an explicit replanning decision to relax the DGR-007 → DGR-019 dependency.
|
||||
- That replanning decision has **not** been made.
|
||||
|
||||
### 4) Maintenance-only tasks should stay separate from feature implementation
|
||||
|
||||
The review uncovered that the codebase now has a clean closed-story split, but further work should avoid mixing:
|
||||
|
||||
- maintenance cleanup,
|
||||
- blocked-hardware preparation,
|
||||
- and actual distributed GLM implementation.
|
||||
|
||||
The next model should treat the maintenance pass as closed and only pick up real implementation work that is not hardware-blocked.
|
||||
|
||||
## Recommended next move
|
||||
|
||||
Use the next model to continue on the **non-blocked implementation queue** only.
|
||||
Priority candidates are whatever is still actionable without the GLM oracle host; if a story depends on DGR-018, keep it deferred.
|
||||
|
||||
## Reference files
|
||||
|
||||
- `docs/issues/distributed-gguf-runtime/README.md`
|
||||
- `.scratch/distributed-gguf-runtime/PRD.md`
|
||||
- `.scratch/distributed-gguf-runtime/milestones.md`
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-018/BLOCKED.md`
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-018/preflight.json`
|
||||
@@ -28,20 +28,20 @@
|
||||
"DGR-021": {
|
||||
"number": 5,
|
||||
"url": "https://git.d-popov.com/popov/neuron-tai/issues/5",
|
||||
"state": "open",
|
||||
"status": "ready"
|
||||
"state": "closed",
|
||||
"status": "completed"
|
||||
},
|
||||
"DGR-022": {
|
||||
"number": 6,
|
||||
"url": "https://git.d-popov.com/popov/neuron-tai/issues/6",
|
||||
"state": "open",
|
||||
"status": "blocked"
|
||||
"status": "ready"
|
||||
},
|
||||
"DGR-023": {
|
||||
"number": 7,
|
||||
"url": "https://git.d-popov.com/popov/neuron-tai/issues/7",
|
||||
"state": "open",
|
||||
"status": "blocked"
|
||||
"status": "ready"
|
||||
},
|
||||
"DGR-024": {
|
||||
"number": 8,
|
||||
@@ -53,7 +53,7 @@
|
||||
"number": 9,
|
||||
"url": "https://git.d-popov.com/popov/neuron-tai/issues/9",
|
||||
"state": "open",
|
||||
"status": "blocked"
|
||||
"status": "ready"
|
||||
},
|
||||
"DGR-026": {
|
||||
"number": 10,
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
|
||||
# DGR-021: Define the versioned named-tensor stream envelope
|
||||
|
||||
- **Status / triage:** specification only; `ready-for-agent`; `passes: false`
|
||||
- **Status / triage:** completed; `passes: true`
|
||||
- **Execution mode:** `AFK`
|
||||
- **Milestone:** `M1`
|
||||
- **Dependencies:** `DGR-018`
|
||||
@@ -18,11 +18,11 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] Define schema version, request/work ID, route session/epoch, shard range/effective start, phase, position, and idempotency step.
|
||||
- [ ] Define named tensors with shape, dtype, byte order, bounded fragments, compression identity, and checksum.
|
||||
- [ ] Reserve extensible fields for token-ID sidebands, architecture state, recurrent state, and MTP without claiming implementations.
|
||||
- [ ] Add deterministic serialization, fragmentation, checksum, unknown-field, and size-limit tests.
|
||||
- [ ] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
||||
- [x] Define schema version, request/work ID, route session/epoch, shard range/effective start, phase, position, and idempotency step.
|
||||
- [x] Define named tensors with shape, dtype, byte order, bounded fragments, compression identity, and checksum.
|
||||
- [x] Reserve extensible fields for token-ID sidebands, architecture state, recurrent state, and MTP without claiming implementations.
|
||||
- [x] Add deterministic serialization, fragmentation, checksum, unknown-field, and size-limit tests.
|
||||
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
||||
|
||||
## Shared quality gates
|
||||
|
||||
@@ -36,4 +36,4 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
|
||||
## Evidence handoff
|
||||
|
||||
Write and verify `.scratch/distributed-gguf-runtime/evidence/DGR-021/README.md`. Until every criterion and applicable gate has real evidence, this story remains `passes: false`. Legacy evidence is provenance only, not completion credit.
|
||||
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-021/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.
|
||||
|
||||
@@ -0,0 +1,67 @@
|
||||
# 18 — Certify whole-model GLM-5.2 runtime semantics
|
||||
|
||||
Status: blocked (2026-07-14) — no 256-GiB-class host available
|
||||
|
||||
> **Blocked:** This story requires a 256-GiB-class host with at least 224 GiB
|
||||
> runtime-accessible memory and 250 GB free storage outside `/home`. The
|
||||
> development host has 124.9 GiB MemTotal and no eligible filesystem (largest:
|
||||
> 74.2 GB free). Exact preflight output is preserved in
|
||||
> [evidence/DGR-018/BLOCKED.md](../evidence/DGR-018/BLOCKED.md); preflight
|
||||
> scripts were preserved at commit a0f28b5 (`scripts/glm_whole_model_preflight.py`,
|
||||
> `scripts/verify_glm_shards.py`). Resume by re-running the preflight on a
|
||||
> qualifying host — do not substitute a smaller model.
|
||||
|
||||
## Mandatory fresh-session context
|
||||
|
||||
- Read [RALPH-CONTEXT.md](../RALPH-CONTEXT.md), [GLM-5.2-MAX-ALPHA-ROADMAP.md](../GLM-5.2-MAX-ALPHA-ROADMAP.md), and this issue completely before changing code.
|
||||
- This issue is `DGR-018` in [prd.json](../prd.json).
|
||||
- Read and verify every dependency evidence README.
|
||||
- Inspect current upstream behavior and `git status`; community claims that a model “works” are not correctness evidence.
|
||||
|
||||
## Description
|
||||
|
||||
As a runtime maintainer, I need a certified whole-model oracle for the exact lowest-quant GLM-5.2 artifact so that distributed parity is measured against correct MoE, DSA, IndexShare, cache, and Max-template semantics.
|
||||
|
||||
## Expected durable outputs
|
||||
|
||||
- Verified target artifact on mounted storage
|
||||
- Stock-pin baseline and warning/tensor inventory
|
||||
- Focused GLM runtime correctness tests and minimum patch series, if required
|
||||
- Signed whole-model oracle recipe, deterministic outputs, metrics, and limitations
|
||||
- `evidence/DGR-018/README.md`
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] Preflight a 256-GiB-class reference host with at least 224 GiB runtime-accessible memory after OS reservation and approximately 250 GB free mounted storage; abort before download if storage resolves under `/home` or requirements are unmet.
|
||||
- [ ] Configure the oracle's alpha lane for 16,384 context, concurrency 1, and Q8_0 MLA/indexer KV; record actual MLA, indexer-cache, scratch, and peak resident allocations separately.
|
||||
- [ ] Download/resume and verify all six DGR-017 `UD-IQ1_S` files against exact sizes and LFS SHA-256 values.
|
||||
- [ ] Load the target with the unmodified DGR-004 pin first and retain full logs, tensor warnings, peak memory, context/KV allocation, TTFT, and decode timing.
|
||||
- [ ] Prove from graph/runtime evidence whether 256-expert MoE/top-8/shared expert, DSA lightning indexer/sparse attention, IndexShare Full/Shared reuse, and `reasoning_effort=max` are active.
|
||||
- [ ] Dense-attention or replicated-indexer compatibility fallback is labeled incomplete and cannot become the oracle solely because it emits plausible text.
|
||||
- [ ] Audit focused upstream/Mesh-LLM donor tests and patches; adopt or rewrite only independently understood minimum changes and record provenance/rejection rationale.
|
||||
- [ ] Handle the trailing NextN/MTP tensors explicitly; MTP may remain disabled for alpha, but ignored tensors and layer-count behavior must be understood and tested.
|
||||
- [ ] Run deterministic prefill/greedy decode and fixed Max-mode coding, structured-output/tool-call, and reasoning sentinels; retain raw prompts, outputs, token IDs, and compared state/logit evidence.
|
||||
- [ ] Bind the oracle to exact artifact, tokenizer/template, adapter, backend class, compute/activation/KV dtypes, context/RoPE, llama.cpp commit, and patch-series hash.
|
||||
- [ ] Produce a signed `pass` or `stop` semantic-runtime verdict; no performance threshold is weakened after execution.
|
||||
- [ ] Targeted native/unit tests and clean pinned rebuild pass.
|
||||
- [ ] `git diff --check` passes.
|
||||
- [ ] Default tests remain model-download-free/GPU-free; full-target execution is opt-in and never runs in default CI.
|
||||
- [ ] Model artifacts and caches remain on configured mounted-drive storage and never under `/home`.
|
||||
- [ ] Preserve all pre-existing working-tree changes and stage only files belonging to this story.
|
||||
- [ ] Write `evidence/DGR-018/README.md` with exact hardware, storage, source revisions, files, commands, raw-result paths, real results, limitations, and handoff.
|
||||
- [ ] Update only this story issue to `Status: done` after every acceptance criterion passes.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
- `DGR-003`, `DGR-004`, and `DGR-017` must have `passes: true`; read and verify their evidence READMEs.
|
||||
|
||||
## Finish contract
|
||||
|
||||
- If a 256-GiB-class host with at least 224 GiB runtime-accessible memory is unavailable, write `evidence/DGR-018/BLOCKED.md` with exact preflight output; do not substitute a smaller model.
|
||||
- Preserve real failures and blockers; never fabricate output.
|
||||
- Emit `<promise>COMPLETE</promise>` only after the signed oracle evidence exists.
|
||||
|
||||
## References
|
||||
|
||||
- [GLM-5.2 Max alpha roadmap](../GLM-5.2-MAX-ALPHA-ROADMAP.md)
|
||||
- [llama.cpp GLM-5.2 issue](https://github.com/ggml-org/llama.cpp/issues/24730)
|
||||
@@ -0,0 +1,63 @@
|
||||
# 19 — Implement and certify GLM-5.2 range, DSA, and IndexShare semantics
|
||||
|
||||
Status: blocked (2026-07-14) — waiting on DGR-018
|
||||
|
||||
> **Blocked:** Depends on DGR-018's whole-model IQ1_S oracle, which is blocked
|
||||
> on a 256-GiB-class host (≥ 224 GiB runtime-accessible memory). See
|
||||
> [evidence/DGR-018/BLOCKED.md](../evidence/DGR-018/BLOCKED.md). Locked
|
||||
> fixture/target parity cannot be certified without that oracle.
|
||||
|
||||
## Mandatory fresh-session context
|
||||
|
||||
- Read [RALPH-CONTEXT.md](../RALPH-CONTEXT.md), [GLM-5.2-MAX-ALPHA-ROADMAP.md](../GLM-5.2-MAX-ALPHA-ROADMAP.md), and this issue completely before changing code.
|
||||
- This issue is `DGR-019` in [prd.json](../prd.json).
|
||||
- Read and verify every dependency evidence README.
|
||||
- Inspect current source and `git status`; do not generalize through tensor-name substitution.
|
||||
|
||||
## Description
|
||||
|
||||
As a target-model operator, I need explicit range-owned GLM-5.2 semantics so that contiguous consumer Shards preserve the whole-model MoE, DSA, IndexShare, and local-KV computation.
|
||||
|
||||
## Expected durable outputs
|
||||
|
||||
- Explicit GLM-5.2 architecture adapter and tensor-ownership rules
|
||||
- IndexShare-aware Shard planner and named DSA sideband protocol mapping
|
||||
- GLM architecture fixture and target parity evidence
|
||||
- Memory/tensor/KV ownership report
|
||||
- `evidence/DGR-019/README.md`
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] Implement explicit ownership for 78 main layers, head embedding, tail norm/output head, routed/shared experts, DSA/indexer tensors, and the certified NextN/MTP policy.
|
||||
- [ ] Each contiguous layer owner keeps all 256 experts and the shared-expert path for its layers local; no public cross-machine expert collectives are introduced.
|
||||
- [ ] Define and fixture-test compressed MLA KV ownership for locally owned layers, keyed by Route Session/epoch, for DGR-007 to implement.
|
||||
- [ ] Implement native DSA lightning indexer/top-2,048 and sparse-attention graph behavior matching DGR-018.
|
||||
- [ ] Parse and validate artifact `indexer_types`; implement Full producer and Shared consumer behavior without fabricated/duplicated indexer tensors.
|
||||
- [ ] Prefer Shard boundaries that preserve IndexShare ownership groups; when memory fit forces a split, carry a typed, bounded, validated top-k sideband in the DGR-006 named bundle.
|
||||
- [ ] Reject missing, stale, wrong-width, wrong-position, or Shared-before-Full sideband/index state.
|
||||
- [ ] Demonstrate mapped/resident weights and allocated KV scale with exact owned tensors/layers rather than the full model.
|
||||
- [ ] Same-host two-stage F32 seam fixture produces 32 exact greedy tokens against the DGR-018 semantics; production seam meets the pre-locked token/similarity thresholds.
|
||||
- [ ] If full-target same-host execution cannot fit, use a layer-reduced GLM architecture fixture only for graph parity and defer full-artifact parity to DGR-020; label it non-target evidence.
|
||||
- [ ] Add deterministic fixture tests for MoE route/shared expert, DSA Full/Shared groups, internal group split, endpoint ownership, KV filtering, malformed metadata, and NextN policy.
|
||||
- [ ] Targeted pytest/CTest/native tests pass; pinned patch stack applies and rebuilds cleanly.
|
||||
- [ ] `python -m compileall packages tests` and `git diff --check` pass.
|
||||
- [ ] Default tests remain deterministic, model-download-free, API-credit-free, and GPU-free.
|
||||
- [ ] Full deterministic `pytest -q` passes, or exact pre-existing unrelated failures are recorded.
|
||||
- [ ] Preserve all pre-existing working-tree changes and stage only files belonging to this story.
|
||||
- [ ] Write `evidence/DGR-019/README.md` with files, commands, real results, raw evidence, limitations, and dependent-story handoff.
|
||||
- [ ] Update only this story issue to `Status: done` after every acceptance criterion passes.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
- `DGR-005`, `DGR-006`, and `DGR-018` must have `passes: true`; read and verify their evidence READMEs.
|
||||
- DGR-007 integration may be completed in parallel only after the adapter's cache contract is explicit.
|
||||
|
||||
## Finish contract
|
||||
|
||||
- Preserve real failures and blockers; never claim target parity from the architecture fixture.
|
||||
- Emit `<promise>COMPLETE</promise>` only after evidence exists.
|
||||
|
||||
## References
|
||||
|
||||
- [GLM-5.2 Max alpha roadmap](../GLM-5.2-MAX-ALPHA-ROADMAP.md)
|
||||
- [Current architecture](../architecture.md)
|
||||
@@ -0,0 +1,67 @@
|
||||
# 20 — Pass real distributed GLM-5.2 Max alpha acceptance
|
||||
|
||||
Status: blocked (2026-07-14) — waiting on DGR-018/DGR-019
|
||||
|
||||
> **Blocked:** Depends on DGR-018 and DGR-019, both blocked on the 256-GiB-class
|
||||
> oracle host (≥ 224 GiB runtime-accessible memory), plus enough physical
|
||||
> consumer nodes that no single node admits the whole recipe. See
|
||||
> [evidence/DGR-018/BLOCKED.md](../evidence/DGR-018/BLOCKED.md).
|
||||
|
||||
## Mandatory fresh-session context
|
||||
|
||||
- Read [RALPH-CONTEXT.md](../RALPH-CONTEXT.md), [GLM-5.2-MAX-ALPHA-ROADMAP.md](../GLM-5.2-MAX-ALPHA-ROADMAP.md), and this issue completely before changing code.
|
||||
- This issue is `DGR-020` in [prd.json](../prd.json).
|
||||
- Read and verify every dependency evidence README and the immutable DGR-017 contract.
|
||||
- Inspect live source, hardware, network, storage, and `git status`; synthetic or layer-reduced evidence cannot satisfy this story.
|
||||
|
||||
## Description
|
||||
|
||||
As an alpha user, I need the exact lowest-quant GLM-5.2 model to run in Max reasoning mode across consumer machines through Meshnet so that the project has proven its end goal on real hardware.
|
||||
|
||||
## Expected durable outputs
|
||||
|
||||
- Exact multi-node hardware/network/storage/runtime manifest
|
||||
- Signed target identity, ownership, parity, quality, performance, and reliability reports
|
||||
- Raw worker/tracker/API logs, metrics, prompts, outputs, and cleanup evidence
|
||||
- Explicit immutable `alpha` or `stop` verdict
|
||||
- `evidence/DGR-020/README.md`
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] Use the exact DGR-017 `UD-IQ1_S` artifact and certified DGR-018/DGR-019 runtime recipe with all source and patch hashes verified.
|
||||
- [ ] Use at least two physical consumer machines; each reserves at least `max(20% of physically usable memory, 8 GiB)` outside weight-plus-Q8-KV placement, no node can place the complete recipe, unified memory is counted once, and measured peak use stays within physical memory without swap/overcommit.
|
||||
- [ ] Treat 5×64 GiB or 3×96/128 GiB as the recommended topology; arithmetic-minimum 4×64 or 2×128 qualifies only with exact contiguous placement and measured reserve evidence.
|
||||
- [ ] Use a same-switch wired network with at least 2.5 GbE; record 10 GbE as recommended and measure hop RTT/serialization/queue latency.
|
||||
- [ ] Tracker selects disjoint contiguous Shards whose exact required tensor inventory has complete union and no unintended overlap.
|
||||
- [ ] Every stage reports real CPU/GPU compute, owned tensor bytes/layers, local KV, backend, queue, and seam telemetry; synthetic, passthrough, or zero-layer workers fail acceptance.
|
||||
- [ ] Runtime evidence proves native GLM MoE/shared expert, DSA lightning indexer/sparse attention, and IndexShare Full/Shared paths are active; dense fallback fails acceptance.
|
||||
- [ ] OpenAI-compatible API applies and records `reasoning_effort=max`, stable model ID, finish reason, and token usage.
|
||||
- [ ] Configure 16,384 context, concurrency 1, and Q8_0 MLA/indexer KV; complete the fixed 4,096-token prefill lane and generate at least 512 output tokens or valid natural EOS after at least 128 tokens.
|
||||
- [ ] Production seam meets at least 0.90 greedy token agreement and 0.999 mean compared-state/logit cosine similarity against DGR-018 on the fixed corpus, with no non-finite or malformed tensors.
|
||||
- [ ] Fixed coding, structured tool-call/JSON, and multi-step reasoning sentinels produce parseable relevant outputs retained for human review.
|
||||
- [ ] After warm-up, fixed Max lane median decode is at least 0.5 token/s, 4,096-token-prompt TTFT is at most 10 minutes, and no unexplained stall exceeds 60 seconds without progress telemetry.
|
||||
- [ ] Record per-stage/seam p50/p95 latency, bytes, compute time, peak RSS/VRAM, KV pressure, network properties, errors, and total energy when available.
|
||||
- [ ] Two consecutive cold starts load, generate, release, and exit without leaked processes, mapped weights, queues, or KV leases.
|
||||
- [ ] Cancellation in prefill and decode releases all stages; one worker loss aborts the route; retry starts from token zero on a newly compatible route; stale epochs and duplicate steps fail closed.
|
||||
- [ ] Model artifacts/caches remain on configured mounted-drive storage and never under `/home`; secret scan passes.
|
||||
- [ ] Run targeted/full deterministic tests and clean native rebuild in addition to target execution; `git diff --check` passes.
|
||||
- [ ] Preserve all raw failures and do not weaken DGR-017 thresholds after results.
|
||||
- [ ] Write a signed `alpha` verdict only if every criterion passes; otherwise write `stop` with the measured bottleneck and next narrow action.
|
||||
- [ ] Write `evidence/DGR-020/README.md` with exact commands, manifests, raw-result paths, outputs, limitations, and post-alpha handoff.
|
||||
- [ ] Update only this story issue to `Status: done` after the immutable verdict and all evidence exist.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
- `DGR-007`, `DGR-008`, `DGR-009`, `DGR-011`, `DGR-013`, `DGR-017`, `DGR-018`, and `DGR-019` must have `passes: true`; read and verify their evidence READMEs.
|
||||
- DGR-012 continuous batching and DGR-014 final comparative release gate are post-alpha and do not block one target session.
|
||||
|
||||
## Finish contract
|
||||
|
||||
- If required physical nodes/storage are unavailable, write `evidence/DGR-020/BLOCKED.md`; do not substitute a smaller model, API, synthetic worker, or single host.
|
||||
- Preserve real failures and blockers; never fabricate target output or telemetry.
|
||||
- Emit `<promise>COMPLETE</promise>` only after the signed verdict exists.
|
||||
|
||||
## References
|
||||
|
||||
- [GLM-5.2 Max alpha roadmap](../GLM-5.2-MAX-ALPHA-ROADMAP.md)
|
||||
- [Ralph execution context](../RALPH-CONTEXT.md)
|
||||
@@ -441,7 +441,8 @@
|
||||
"Add deterministic serialization, fragmentation, checksum, unknown-field, and size-limit tests.",
|
||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||
],
|
||||
"passes": false,
|
||||
"passes": true,
|
||||
"completionNotes": "Added a versioned activation envelope with deterministic JSON serialization, bounded tensor fragmentation, checksum validation, unknown-field preservation, and TensorPayload conversion hooks; verified by targeted pytest and compileall runs.",
|
||||
"notes": "Generated source issue: .scratch/distributed-gguf-runtime/issues/021-define-the-versioned-named-tensor-stream-envelope.md; prd.json is authoritative.",
|
||||
"blocks": [
|
||||
"DGR-022",
|
||||
@@ -482,8 +483,8 @@
|
||||
"Add compatibility tests for supported versions and fail-closed tests for unsupported versions and malformed lifecycle transitions.",
|
||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||
],
|
||||
"passes": false,
|
||||
"notes": "Generated source issue: .scratch/distributed-gguf-runtime/issues/022-define-shard-lifecycle-and-structured-status-rpcs.md; prd.json is authoritative.",
|
||||
"passes": true,
|
||||
"notes": "Completed from isolated DGR-022 Ralph worktree; verified with 17 focused pytest cases and retained DGR-021 envelope compatibility.",
|
||||
"blocks": [
|
||||
"DGR-024",
|
||||
"DGR-033",
|
||||
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"signers": [
|
||||
{
|
||||
"algorithm": "ed25519",
|
||||
"fingerprint_sha256": "8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de",
|
||||
"scope": "DGR-001 local-real and gpu-diagnostic benchmark evidence",
|
||||
"status": "active"
|
||||
}
|
||||
]
|
||||
}
|
||||
15
docs/archived/README.md
Normal file
15
docs/archived/README.md
Normal file
@@ -0,0 +1,15 @@
|
||||
# Archived task programs
|
||||
|
||||
These task programs are historical, completed, superseded, or explicitly not part of the active development queue. They are preserved for provenance and must not be treated as runnable work.
|
||||
|
||||
## Archived programs
|
||||
|
||||
- `alpha-hardening/` — historical alpha hardening and settlement/authentication task set.
|
||||
- `dashboard-test-runner/` — historical dashboard test-runner work.
|
||||
- `distributed-inference-performance/` — superseded distributed-inference performance backlog.
|
||||
- `node-capability-admission/` — historical capability-admission task set.
|
||||
- `proxy-stream-cancellation/` — historical proxy cancellation task.
|
||||
- `qwen3.6-27b-demand-placement/` — superseded Qwen demand-placement planning.
|
||||
- `routing-compatibility-regression/` — historical routing compatibility regression work.
|
||||
|
||||
The active distributed GGUF backlog remains under `.scratch/distributed-gguf-runtime/`. Architectural decisions under `docs/adr/` remain active and were not moved.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user