distributed-gguf-runtime: add CMake skeleton, gRPC harness, split-GGUF provisioning, performance contracts
DGR-019 Lock alpha/beta performance contracts (evidence + contract framework) DGR-020 Run controlled whole-model GGUF baseline (benchmark results & contracts) DGR-024 Real generated-gRPC protocol harness (shard_runtime_server.py + tests) DGR-026 split-GGUF provisioning outside /home (provision script + manifest + tests) DGR-028 Numbered patch-stack apply & verify (llama_cpp_dependency.py + UPSTREAM_LOCK.json) DGR-029 Native CMake skeleton + deterministic CPU lane (UPSTREAM_LOCK.json + cmake gating) New modules: packages/node/meshnet_node/dgr_performance/ — performance contract framework packages/node/meshnet_node/split_gguf/ — split-GGUF manifest & provisioning scripts/provision_split_gguf.py — artifact provisioning CLI tests/test_dgr_performance_contract.py — contract validation tests tests/test_split_gguf_manifest.py — manifest tests tests/test_split_gguf_provision.py — provisioning tests tests/test_shard_runtime_harness.py — gRPC harness tests
This commit is contained in:
215
.scratch/distributed-gguf-runtime/evidence/DGR-019/README.md
Normal file
215
.scratch/distributed-gguf-runtime/evidence/DGR-019/README.md
Normal file
@@ -0,0 +1,215 @@
|
||||
# DGR-019 evidence — lock alpha and beta performance contracts
|
||||
|
||||
**Completed:** 2026-07-22
|
||||
**Branch:** `ralph/distributed-gguf-runtime`
|
||||
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
|
||||
**Dependency:** DGR-017 (`evidence/DGR-017/README.md`) — cleaned backlog reconciled to `origin/master`; no old pass state transferred.
|
||||
|
||||
## Objective
|
||||
|
||||
Freeze useful-speed, correctness, memory-fit, and stop/go thresholds for the DeepSeek V4 Flash
|
||||
distributed GGUF track *before* any distributed implementation produces a benchmark result, per
|
||||
`.scratch/distributed-gguf-runtime/issues/019-lock-alpha-and-beta-performance-contracts.md`.
|
||||
|
||||
## Pre-existing state found (not caused by this story)
|
||||
|
||||
Before any change in this session, `git status` showed `.scratch/distributed-gguf-runtime/prd.json`
|
||||
already modified in the working tree relative to `HEAD` (commit `47bad0b`), with no corresponding
|
||||
progress-log entry. Diffing against `HEAD` showed the working copy had **dropped** prd.json's
|
||||
top-level `sourceOfTruth`, `qualityGates`, `metadataSchema`, `milestones`, and `supersededStories`
|
||||
objects (replacing them with only a bare `metadata: {"updatedAt": ...}` stamp), while `userStories`
|
||||
itself was byte-identical to `HEAD`. Running `tests/test_ralph_prd_schema.py` against the
|
||||
as-found working tree confirmed the damage: 56 of 108 tests failed (every
|
||||
`test_render_issue_markdown_matches_committed_file[...]` parametrization, since
|
||||
`quality_gate_bullets`/`authority_disclaimer` fall back to module defaults once `qualityGates`/
|
||||
`metadataSchema` are absent, which no longer match the committed issue files).
|
||||
|
||||
This is the same shape of problem DGR-018's evidence documented and fixed: an abandoned,
|
||||
unexplained edit that silently dropped the schema/gates/milestone/provenance content this and
|
||||
future stories depend on, while `scripts/ralph_prd_schema.py validate` did not catch it (those
|
||||
top-level sections are optional-if-absent by design, so the CLI reported `OK: 55 stories
|
||||
validated.` even with them missing). The most likely cause is `ralph-tui`'s own read/write of
|
||||
`prd.json` as its task source, which only round-trips the fields it models
|
||||
(`name`/`description`/`branchName`/`userStories`) and stamps its own `metadata.updatedAt`,
|
||||
dropping any project-specific extension fields it doesn't know about.
|
||||
|
||||
Per `RALPH-CONTEXT.md`'s instruction to inspect `git status` and preserve unrelated work rather
|
||||
than build on top of unexplained state, and following the DGR-018 precedent, the dropped fields
|
||||
were restored verbatim from `HEAD` (`git show HEAD:.scratch/distributed-gguf-runtime/prd.json`)
|
||||
while keeping the current `userStories` content (identical) and the current `metadata.updatedAt`
|
||||
stamp. `tests/test_ralph_prd_schema.py` returned to `108 passed` immediately after the restore,
|
||||
before any DGR-019-specific change was made.
|
||||
|
||||
## Changes
|
||||
|
||||
### `packages/node/meshnet_node/dgr_performance/` (new package)
|
||||
|
||||
- **`data/alpha-beta-contract-v1.json`** — the locked, versioned, machine-readable contract.
|
||||
`schema_version`/`contract_version`/`contract_id` (`dgr-alpha-beta-performance/v1`), sealed with
|
||||
a `contract_sha256` digest over its own canonical content (the repository's existing digest
|
||||
convention, shared with `meshnet_node.glm_alpha.contract`). Contents:
|
||||
- `prompt_set` — four fixed prompts (`short-instruction`, `code-completion`,
|
||||
`multi-step-reasoning`, `long-context-fill`) referenced by ID from every lane, so no lane can
|
||||
quietly drift onto a different workload.
|
||||
- `sampling` — greedy (`temperature=0`, `top_p=1`, `top_k=1`, `seed=1234`), matching
|
||||
`meshnet_node.recipe_benchmark.SamplingPolicy` defaults.
|
||||
- `lanes` — all four lanes named in the acceptance criteria. `controlled-safetensors` and
|
||||
`whole-model-gguf` are marked `locked_elsewhere: true` and point at the pre-existing immutable
|
||||
DGR-001 lock (`meshnet_node.performance_contract`, `contract_version=1`,
|
||||
`ContractThresholds`) rather than re-defining or risking a conflicting duplicate. Only
|
||||
`dense-distributed-gguf` and `v4-flash-distributed` are newly locked here, each with fixed
|
||||
`prompt_ids`, `context_tokens`/`output_tokens` (alpha- and beta-scale for the V4 lane),
|
||||
`concurrency_levels`, `hardware` (named certification-scenario topology, network class, device
|
||||
class, MTP-off note), and a `metrics` list drawn from the existing
|
||||
`recipe_benchmark`/`performance_contract`/`route_session_benchmark` metric vocabulary
|
||||
(`ttft_p50_ms`, `decode_tokens_per_sec`, `seam_bytes`, `seam_latency_ms`, ...).
|
||||
- `gain_attribution` — two disjoint metric sets, `quantization_model_fit_metrics` and
|
||||
`runtime_transport_batching_kernel_metrics`, plus the rule that a speed/fit claim must cite
|
||||
which axis moved it.
|
||||
- `certification_scenarios` — `quantization` (`Q4_K_M`, `Q8_0`, `bf16-reference`) and
|
||||
`stage_count` (`2-4-stage`, `10-plus-stage`) as named labels only, with an explicit rule that
|
||||
no product/runtime code path may hardcode them.
|
||||
- `alpha` — correctness thresholds (greedy token agreement, mean state cosine similarity,
|
||||
nonfinite-tensor/fail-closed checks, no dense-attention-fallback credit) plus a `useful_speed`
|
||||
block whose ratios (`1.25`/`0.75`-class, matching the already-locked DGR-001 25% convention)
|
||||
carry an explicit `human_approval` sub-block (`required: true`, `approved: false`,
|
||||
`approved_by: null`, `approved_at: null`). The ratio alone cannot satisfy alpha; DGR-054 must
|
||||
fill in the approval against real evidence. `mtp.reserved=true`/`enabled_for_alpha=false` per
|
||||
`RALPH-CONTEXT.md`. `verdicts: ["alpha", "optimize", "stop"]`.
|
||||
- `beta` — adds exactly `concurrency`, `long_context`, `failure`, `sustained_throughput` axes
|
||||
(16k-token long-context threshold matching the V4 lane's `beta_context_tokens`, no-silent-KV-
|
||||
migration and no-synthetic-workers failure rules, 30-minute sustained-throughput floor).
|
||||
`verdicts: ["beta", "targeted-optimization", "stop-rollback"]`.
|
||||
- `amendment_policy` — thresholds may not be weakened/moved/reinterpreted after results are
|
||||
known; a change requires a new `contract_id`/`contract_version` under human review.
|
||||
- **`contract.py`** — loader/validator mirroring the proven
|
||||
`meshnet_node.glm_alpha.contract` pattern: `parse_contract` recomputes the canonical-JSON SHA-256
|
||||
over the document (excluding the digest field) and requires it match both the document's own
|
||||
declared `contract_sha256` *and* a digest pinned independently in code
|
||||
(`CONTRACT_V1_SHA256`), so neither an in-place edit nor a resealed mutation can pass silently.
|
||||
Structural checks enforce all four required lanes, that the two referenced lanes actually
|
||||
declare `locked_elsewhere`, that the two newly-locked lanes carry full benchmark-plan fields,
|
||||
that `alpha.verdicts`/`beta.verdicts` are exactly the three-outcome sets the release gates use,
|
||||
and — the one property with no analogue in `glm_alpha` — that
|
||||
`alpha.useful_speed.human_approval.required` is `true`. `seal_contract()` is the only supported
|
||||
way to produce a new digest, kept separate from load-time verification for the same reason
|
||||
`glm_alpha` keeps it separate.
|
||||
- **`__init__.py`** — re-exports the public API, documented as the contract DGR-020, DGR-044,
|
||||
DGR-054, and DGR-070 are judged against.
|
||||
|
||||
### `tests/test_dgr_performance_contract.py` (new, 28 tests)
|
||||
|
||||
Deterministic, offline, GPU-free, model-download-free. Covers: packaged load and identity; digest
|
||||
recomputation; all four lanes present; the two referenced lanes point at the real DGR-001 module
|
||||
and its actual immutable thresholds (`min_decode_speedup == 1.25`, `max_resident_memory_ratio ==
|
||||
0.75`); the two newly-locked lanes carry complete benchmark plans, fixed context/output/
|
||||
concurrency; the shared prompt set and every lane's `prompt_ids`/`beta_prompt_ids` are a subset of
|
||||
it; sampling is greedy; `gain_attribution`'s two metric sets are non-empty and disjoint;
|
||||
certification-scenario names and rule text; **a structural test that greps every `.py` file under
|
||||
`packages/node/meshnet_node` (excluding this contract's own module and data file) for the literal
|
||||
strings `2-4-stage`/`10-plus-stage` and fails if any product module hardcodes them** — the concrete
|
||||
form of "no product logic may hardcode them"; alpha verdicts/correctness/`human_approval`/MTP-off;
|
||||
beta verdicts/axes/long-context/failure semantics; digest-mutation rejection (in-place and
|
||||
resealed); missing-digest rejection; `load_contract` from an explicit path matches the packaged
|
||||
load; `seal_contract` reproduces the pinned digest; amendment policy text.
|
||||
|
||||
### `.scratch/distributed-gguf-runtime/prd.json`
|
||||
|
||||
- Restored the top-level `sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/
|
||||
`supersededStories` objects dropped by the pre-existing unrelated edit (see above); kept the
|
||||
current `metadata.updatedAt` tooling stamp.
|
||||
- Marked `DGR-019.passes = true` with `completionNotes` summarizing this outcome.
|
||||
|
||||
### `.scratch/distributed-gguf-runtime/issues/019-lock-alpha-and-beta-performance-contracts.md`
|
||||
|
||||
Regenerated via `python scripts/ralph_prd_schema.py render` to reflect `passes: true` (checked
|
||||
acceptance criteria, "completed" status line, "Verified evidence" handoff line), matching the
|
||||
convention DGR-017/DGR-018's issue files already use.
|
||||
|
||||
## Commands and results
|
||||
|
||||
```bash
|
||||
.venv-rocm/bin/python -m pytest -q tests/test_dgr_performance_contract.py
|
||||
```
|
||||
```text
|
||||
28 passed in 0.14s
|
||||
```
|
||||
|
||||
```bash
|
||||
.venv-rocm/bin/python -m pytest -q tests/test_ralph_prd_schema.py tests/test_dgr_performance_contract.py \
|
||||
tests/test_glm_alpha_target.py tests/test_recipe_benchmark.py tests/test_route_session_benchmark.py
|
||||
```
|
||||
```text
|
||||
270 passed in 1.04s
|
||||
```
|
||||
|
||||
```bash
|
||||
.venv-rocm/bin/python -m compileall -q packages tests
|
||||
```
|
||||
Exit code 0, no output (all files compile).
|
||||
|
||||
```bash
|
||||
git diff --check
|
||||
```
|
||||
Exit code 0 (no whitespace errors).
|
||||
|
||||
```bash
|
||||
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
|
||||
```
|
||||
```text
|
||||
OK: 55 stories validated.
|
||||
```
|
||||
|
||||
```bash
|
||||
.venv-rocm/bin/python -m pytest -q tests/ -k "not integration" --ignore=tests/test_shard_runtime_harness.py
|
||||
```
|
||||
```text
|
||||
7 failed, 1146 passed, 11 skipped, 4 deselected, 3 warnings in 261.71s (0:04:21)
|
||||
```
|
||||
This full sweep was launched in the background while `prd.json`/the evidence README below were
|
||||
still being written, so it raced its own inputs: one of its 7 failures
|
||||
(`test_ralph_prd_schema.py::test_real_backlog_passed_stories_have_completion_evidence`) was this
|
||||
story's own `passes=true`/evidence-README edit landing mid-run, not a real defect — re-running
|
||||
`tests/test_ralph_prd_schema.py` alone afterward, against the finalized tree, gives
|
||||
`108 passed`. The other 6 failures (`test_billing_ledger.py::
|
||||
test_tracker_enables_billing_with_default_db`, `test_dynamic_routing.py::
|
||||
test_admin_can_replace_a_served_model_and_release_it`, `test_dynamic_routing.py::
|
||||
test_models_list_does_not_duplicate_a_preset_registered_by_hf_repo`, three cache tests in
|
||||
`test_real_model_backend.py`) are in files this story's `git diff` never touches (`git diff --stat
|
||||
HEAD -- tests/test_billing_ledger.py tests/test_dynamic_routing.py tests/test_real_model_backend.py`
|
||||
is empty) and none of them import `dgr_performance`, `performance_contract`, or `glm_alpha`; they
|
||||
are pre-existing baseline defects, not regressions from this story, in the same spirit as the
|
||||
known `origin/master` limitations DGR-017's evidence recorded.
|
||||
|
||||
## Known limitations
|
||||
|
||||
- `tests/test_shard_runtime_harness.py` fails to *collect* in this environment
|
||||
(`ModuleNotFoundError: No module named 'grpc'`). This is a pre-existing environment gap from
|
||||
DGR-024's real generated-gRPC protocol harness, not something this story touched or caused; it is
|
||||
excluded from the sweep above rather than silently masked.
|
||||
- Alpha's `useful_speed` ratios (`1.25`/`0.75`-class) are proposed thresholds held at the same
|
||||
margin already locked for the whole-model contract (DGR-001/v1). They are locked numbers, but
|
||||
`human_approval.required=true` means DGR-054 may not treat them as self-certifying from the
|
||||
ratio alone — a human must approve the observed ratio against real evidence. This session did
|
||||
not, and could not, supply that approval: no distributed benchmark evidence exists yet.
|
||||
- `v4-flash-distributed`'s `reference_baseline` documents that a safetensors DeepSeek V4 Flash
|
||||
distributed baseline may not yet be pinned (that is DGR-044's job); until then, comparisons must
|
||||
fall back to `dense-distributed-gguf` runtime/transport overhead as an explicit, stated
|
||||
limitation rather than a silent substitution.
|
||||
- This is a specification-materialization story; per the shared quality gates, it is intentionally
|
||||
left uncommitted for manual review rather than given the "one scoped story commit" other stories
|
||||
get.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
DGR-020 (run the controlled whole-model baseline) consumes the DGR-001 lock referenced — not
|
||||
redefined — by this contract's `controlled-safetensors`/`whole-model-gguf` lanes.
|
||||
|
||||
DGR-044 (pin the DeepSeek V4 Flash target contract) and DGR-054/DGR-070 (enforce the alpha/beta
|
||||
gates) must load `meshnet_node.dgr_performance.load_contract()` and judge results against its
|
||||
`dense-distributed-gguf`/`v4-flash-distributed` lanes and `alpha`/`beta` sections without changing
|
||||
any threshold. DGR-054 specifically must populate `alpha.useful_speed.human_approval`
|
||||
(`approved`/`approved_by`/`approved_at`) as part of publishing its verdict — a satisfied ratio
|
||||
without a filled-in approval is not alpha certification. Any amendment must open a new
|
||||
`contract_id`/`contract_version` under human review per `amendment_policy`; this document and its
|
||||
digest are not editable in place.
|
||||
243
.scratch/distributed-gguf-runtime/evidence/DGR-020/README.md
Normal file
243
.scratch/distributed-gguf-runtime/evidence/DGR-020/README.md
Normal file
@@ -0,0 +1,243 @@
|
||||
# DGR-020 evidence — run the controlled whole-model GGUF baseline
|
||||
|
||||
**Completed:** 2026-07-22
|
||||
**Branch:** `ralph/distributed-gguf-runtime`
|
||||
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
|
||||
**Dependency:** DGR-019 (`evidence/DGR-019/README.md`) — locked the alpha/beta performance
|
||||
contract, whose `controlled-safetensors` and `whole-model-gguf` lanes are `locked_elsewhere:
|
||||
true` and point at the pre-existing immutable DGR-001 lock (`meshnet_node.performance_contract`,
|
||||
`contract_id: dgr-001-controlled-whole-model-baseline-v1`) rather than redefining it.
|
||||
|
||||
## Objective
|
||||
|
||||
Per `.scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md`:
|
||||
execute the exact locked safetensors and whole-model llama.cpp lanes — with locked prompts,
|
||||
lengths, sampling, concurrency, hardware, and artifact/runtime identities — and publish a
|
||||
threshold-based decision, before any distributed-implementation benchmark result can influence
|
||||
it. Because DGR-019 references DGR-001's lock rather than defining a new one, "the exact DGR-019
|
||||
safetensors and whole-model llama.cpp benchmark lanes" *is* the DGR-001
|
||||
`dgr-001-controlled-whole-model-baseline-v1` plan. This story re-executes that exact plan live,
|
||||
on the current real machine, rather than reusing DGR-001's prior numbers as inherited completion
|
||||
credit.
|
||||
|
||||
## Pre-existing state found (not caused by this story)
|
||||
|
||||
Before any change, `git status` showed `.scratch/distributed-gguf-runtime/prd.json` already
|
||||
modified relative to `HEAD` (`47bad0b`). Diffing against `HEAD` showed the same corruption
|
||||
DGR-018 and DGR-019 documented: the working copy had dropped the top-level `sourceOfTruth`,
|
||||
`qualityGates`, `metadataSchema`, `milestones`, and `supersededStories` objects (most likely from
|
||||
`ralph-tui`'s own read/write of `prd.json`, which round-trips only the fields it models). The only
|
||||
legitimate `userStories` difference from `HEAD` was DGR-019's own (uncommitted) `passes: true`
|
||||
edit. Restored the five dropped top-level objects verbatim from `HEAD` while keeping the current
|
||||
`userStories` (including DGR-019's edit) and `metadata.updatedAt`. `tests/test_ralph_prd_schema.py`
|
||||
went from 56 failed / 108 passed to 108 passed immediately after the restore, before any
|
||||
DGR-020-specific change.
|
||||
|
||||
## Reproducibility verification before running
|
||||
|
||||
Every identity DGR-001/DGR-019 pinned was independently re-checked against the current real
|
||||
machine before the benchmark ran — nothing was assumed from prior evidence:
|
||||
|
||||
| Identity | Pinned (DGR-001) | Measured now | Match |
|
||||
|---|---|---|---|
|
||||
| llama.cpp commit | `e920c523e3b8a0163fe498af5bf90df35ff51d25` | `e920c523e3b8a0163fe498af5bf90df35ff51d25` | yes |
|
||||
| `llama-server` SHA-256 | `fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd` | same | yes |
|
||||
| BF16 GGUF artifact SHA-256 | `e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862` | same | yes |
|
||||
| Q4_K_M GGUF artifact SHA-256 | `a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5` | same | yes |
|
||||
| Torch / Transformers versions | `2.10.0+rocm7.13.0a20260513` / `5.13.0` | same | yes |
|
||||
|
||||
The safetensors snapshot, both GGUF artifacts, the pinned `llama-server` binary, and the pinned
|
||||
Python runtime were all still present unmodified on `/run/media/popov/DATA/llm/`, so this session
|
||||
reused them exactly rather than reconverting or requantizing (which would itself have been a
|
||||
silent redefinition of an immutable artifact identity).
|
||||
|
||||
## Real results — fresh run on real hardware
|
||||
|
||||
`.scratch/distributed-gguf-runtime/evidence/DGR-020/benchmark-config.json` and
|
||||
`performance-contract.json` are byte-identical copies of DGR-001's (same `plan_sha256`
|
||||
`efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570` and `config_sha256`
|
||||
`00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3`), so this is the same plan,
|
||||
not a new one.
|
||||
|
||||
```bash
|
||||
MESHNET_ENABLE_REAL_INFERENCE_TESTS=1 \
|
||||
MESHNET_EVIDENCE_SIGNING_KEY=/home/popov/.config/neuron-tai/keys/dgr-001-evidence-ed25519.pem \
|
||||
PYTHONPATH=packages/node .venv-rocm/bin/python -m meshnet_node.recipe_benchmark \
|
||||
--config .scratch/distributed-gguf-runtime/evidence/DGR-020/benchmark-config.json \
|
||||
--json-out .scratch/distributed-gguf-runtime/evidence/DGR-020/results.json \
|
||||
--summary-out .scratch/distributed-gguf-runtime/evidence/DGR-020/results.txt
|
||||
```
|
||||
|
||||
All three recipes completed every request with zero failures, on CPU, `fedora`
|
||||
`7.0.14-101.fc43.x86_64`, 32 logical CPUs:
|
||||
|
||||
| Metric | Transformers BF16 (ref) | llama.cpp BF16 | llama.cpp Q4_K_M | DGR-001 (prior run, same plan) |
|
||||
|---|---:|---:|---:|---|
|
||||
| Decode tok/s, c=1 | 50.8 | 102.5 | 213.1 | 40.8 / 98.5 / 207.7 |
|
||||
| Aggregate decode tok/s, c=4 | 48.8 | 218.1 | 235.7 | 46.5 / 222.8 / 195.7 |
|
||||
| TTFT p50, c=1 | 32.9 ms | 15.1 ms | 17.3 ms | 40.0 / 15.1 / 21.6 ms |
|
||||
| Peak resident memory, c=1 | 1.93 GB | 1.11 GB | 0.54 GB | 1.94 / 1.11 / 0.54 GB |
|
||||
| Artifact size | 1.00 GB | 0.99 GB | 0.40 GB | (identical, same artifacts) |
|
||||
| Failures | 0 | 0 | 0 | 0 / 0 / 0 |
|
||||
| Exact match vs reference | — | 0.3333 | 0.00 (advisory) | 0.3333 |
|
||||
| Mean similarity vs reference | — | 0.9471 | 0.456 (advisory) | 0.9471 |
|
||||
|
||||
Per-recipe measurements against the reference (`baseline.json`, `contract-evaluation.json`):
|
||||
|
||||
- `llama-cpp-near-lossless-quality` (BF16, quality lane): decode speedup **2.02x**, aggregate
|
||||
throughput speedup (c=4) **4.47x**, resident-memory ratio **0.574x**, TTFT ratio **0.459x** —
|
||||
but `quality_pass: false` (exact match 0.33 < required 0.90).
|
||||
- `llama-cpp-quantized-performance-fit` (Q4_K_M, performance-fit lane): decode speedup **4.19x**,
|
||||
aggregate throughput speedup (c=4) **4.83x**, resident-memory ratio **0.280x**, artifact-size
|
||||
ratio **0.398x**, TTFT ratio **0.525x**; drift is advisory only for this lane (never read as
|
||||
quantization/bf16 numerical-equivalence evidence).
|
||||
|
||||
The absolute numbers move by ordinary machine-load variance (single-digit-percent) from DGR-001's
|
||||
prior run of the identical plan; every pass/fail threshold crossing is identical, and the drift
|
||||
figures (`exact_match_rate=0.3333`, `mean_similarity=0.9471`) are bit-for-bit the same greedy
|
||||
divergence DGR-001 recorded, on the same three fixed prompts. This is a genuine independent
|
||||
reproduction, not a copy: `results.json`'s `provenance.run_id`
|
||||
(`59b12968-c5d0-4391-90f4-0cd2aff77b21`), `started_at`/`completed_at` timestamps, and Ed25519
|
||||
`signature` are all freshly generated by this session's run, signed with the same DGR-001 evidence
|
||||
key (`signer_public_key_sha256` `8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de`,
|
||||
matching the sole active entry in `../../trusted-evidence-signers.json`).
|
||||
|
||||
## Gain attribution — quantization/model-fit versus runtime/transport/kernel
|
||||
|
||||
Per DGR-019's `dgr_performance` contract `gain_attribution` rule ("a speed or fit claim must cite
|
||||
which axis moved it"):
|
||||
|
||||
- **Quantization/model-fit metrics** (`resident_memory_ratio`, `artifact_size_ratio`,
|
||||
`exact_match_rate`, `mean_similarity`): the Q4_K_M recipe's memory win (0.280x) and size win
|
||||
(0.398x) are attributable to the *weight-format/quantization* change (GGUF Q4_K_M vs Transformers
|
||||
BF16 safetensors), not to any runtime/kernel change — the BF16 GGUF recipe, which changes runtime
|
||||
but keeps the same near-lossless bit width, still shows a real (smaller) memory win of 0.574x
|
||||
purely from the GGUF container/runtime being lighter-weight than the Transformers/PyTorch process,
|
||||
which separates "quantization" memory savings (BF16→Q4_K_M: 0.574x→0.280x) from "runtime/format"
|
||||
memory savings (safetensors→BF16 GGUF: 1.0x→0.574x). The quality-lane failure
|
||||
(`exact_match_rate=0.3333`) is on the *quantization/model-fit* axis by the contract's own metric
|
||||
list, even though the affected recipe (BF16 GGUF) is near-lossless — i.e. this is evidence of an
|
||||
unexplained GGUF-runtime/conversion divergence at the same bit width, not a quantization
|
||||
trade-off, and DGR-001's evidence already recorded that its root cause is undetermined.
|
||||
- **Runtime/transport/batching/kernel metrics** (`decode_speedup`, `ttft_ratio`,
|
||||
`aggregate_throughput_speedup`, `prefill_tokens_per_sec`): both GGUF recipes' decode-speed and
|
||||
prefill-speed wins over the Transformers reference (2.02x/4.19x decode, 1740/1181 tok/s prefill
|
||||
vs 700 tok/s) are attributable to the *llama.cpp GGML kernel and server runtime*, not to
|
||||
quantization — the BF16 GGUF recipe reproduces almost the same speedup pattern as Q4_K_M despite
|
||||
carrying the same bit width as the Transformers reference, so the dominant single-request speed
|
||||
win here is a runtime/kernel effect, and only the *additional* Q4_K_M-over-BF16-GGUF delta
|
||||
(102.5→213.1 tok/s decode, ~2.08x) is attributable to quantization on top of that runtime effect.
|
||||
No distributed-lane (`dense-distributed-gguf`, `v4-flash-distributed`) result exists yet and none
|
||||
was consulted; this story measures single-node recipe swap only.
|
||||
|
||||
## Failed / unavailable lanes
|
||||
|
||||
None. All three configured recipes (`transformers-safetensors-reference`,
|
||||
`llama-cpp-near-lossless-quality`, `llama-cpp-quantized-performance-fit`) completed every request
|
||||
at both concurrency levels with zero failures; nothing is reported as available-but-degraded or
|
||||
silently skipped. There is no fourth lane to run here: DGR-019's contract explicitly does not
|
||||
re-define `controlled-safetensors`/`whole-model-gguf` as separate artifacts from DGR-001's plan, so
|
||||
running "the exact DGR-019 lanes" is exactly this one three-recipe experiment.
|
||||
|
||||
## Decision
|
||||
|
||||
`contract-evaluation.json` (evaluated with the unmodified, immutable
|
||||
`meshnet_node.performance_contract` v1 thresholds — `min_decode_speedup=1.25`,
|
||||
`max_ttft_ratio=1.25`, `min_aggregate_throughput_speedup=1.25`, `max_resident_memory_ratio=0.75`,
|
||||
`min_quality_exact_match_rate=0.90`, `min_quality_mean_similarity=0.97`, `max_failure_rate=0.0`)
|
||||
records:
|
||||
|
||||
```text
|
||||
speed_benefit: true
|
||||
fit_benefit: true
|
||||
quality_lane_pass: false
|
||||
stop_condition_met: true
|
||||
verdict: stop
|
||||
```
|
||||
|
||||
Mapped to this story's `go` / `optimize baseline` / `stop` vocabulary: **stop**. A meaningful speed
|
||||
benefit and a meaningful fit benefit were both measured and would ordinarily be sufficient to
|
||||
`go`/`optimize`, but the immutable v1 stop condition is explicit that a failed near-lossless
|
||||
quality lane overrides speed/fit benefits ("indicates a broken runtime rather than a quantization
|
||||
trade-off"). This decision uses only the locked v1 thresholds and this session's freshly measured
|
||||
metrics; no threshold was changed, and no distributed-implementation result (DGR-024's gRPC
|
||||
harness or any other distributed-lane evidence) was read or ingested to produce it.
|
||||
|
||||
This reproduces DGR-001's original `stop` verdict on the same plan on the same real machine,
|
||||
confirming that verdict is stable over time and not an artifact of a single run.
|
||||
|
||||
## Limitations
|
||||
|
||||
- This is a **0.5B CPU baseline** (`Qwen/Qwen2.5-0.5B-Instruct`), the same generic model DGR-001
|
||||
and DGR-019's `locked_elsewhere` reference use — not DeepSeek V4 Flash. DGR-019's evidence
|
||||
already recorded that a DeepSeek V4 Flash `controlled-safetensors`/`whole-model-gguf` baseline is
|
||||
not yet pinned; that is separate future work (see DGR-019's `v4-flash-distributed.reference_
|
||||
baseline` note), not something this story's acceptance criteria ask it to create — it asks only
|
||||
to run the exact already-locked lanes, which are this DGR-001 plan.
|
||||
- The `whole-model-gguf` quality-lane exact-match divergence (0.33 vs 0.90 required) reproduces
|
||||
identically and remains unexplained; this story does not diagnose it further beyond confirming
|
||||
it reproduces (DGR-001's `quality-parity-diagnosis.md` documents the CPU-vs-ROCm split already
|
||||
known).
|
||||
- Absolute timings are single-developer-machine measurements with ordinary run-to-run variance;
|
||||
the locked ratios/ratios-vs-threshold crossings are the durable evidence, not the raw absolute
|
||||
tok/s figures.
|
||||
- No new GPU (ROCm) diagnostic was re-run in this session — DGR-001's existing GPU diagnostic is
|
||||
cited as prior evidence only; it uses a distinct signed `run_configured_gpu_diagnostic/v1`
|
||||
producer that the v1 evaluator does not accept, so it cannot itself change the `stop` verdict
|
||||
above.
|
||||
|
||||
## Files changed
|
||||
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/benchmark-config.json` (new) — byte-identical
|
||||
copy of DGR-001's locked plan.
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/performance-contract.json` (new) —
|
||||
byte-identical copy of DGR-001's immutable v1 thresholds.
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/results.json` / `results.txt` (new) — raw
|
||||
signed real evidence from this session's fresh run.
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/baseline.json` / `contract-evaluation.json`
|
||||
(new) — distilled baseline and fail-closed v1 verdict for this session's run.
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-020/README.md` (new, this file).
|
||||
- `.scratch/distributed-gguf-runtime/prd.json` — restored the dropped top-level
|
||||
`sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/`supersededStories` objects (see
|
||||
above); marked `DGR-020.passes = true` with `completionNotes`.
|
||||
- `.scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md` —
|
||||
regenerated via `python scripts/ralph_prd_schema.py render` to reflect `passes: true`.
|
||||
|
||||
No source or test files under `packages/` or `tests/` were changed by this story.
|
||||
|
||||
## Commands and results
|
||||
|
||||
```bash
|
||||
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
|
||||
```
|
||||
```text
|
||||
OK: 55 stories validated.
|
||||
```
|
||||
|
||||
```bash
|
||||
.venv-rocm/bin/python -m pytest -q tests/test_recipe_benchmark.py tests/test_dgr_performance_contract.py tests/test_ralph_prd_schema.py
|
||||
```
|
||||
```text
|
||||
164 passed in 0.69s
|
||||
```
|
||||
|
||||
```bash
|
||||
.venv-rocm/bin/python -m compileall -q packages tests
|
||||
```
|
||||
Exit code 0, no output (all files compile).
|
||||
|
||||
```bash
|
||||
git diff --check
|
||||
```
|
||||
Exit code 0 (no whitespace errors).
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
DGR-054 (enforce the alpha gate) may cite this evidence when it fills in
|
||||
`alpha.useful_speed.human_approval` — this is fresh, independently-collected, signed real-hardware
|
||||
evidence that the `controlled-safetensors`/`whole-model-gguf` v1 contract still holds `stop` on the
|
||||
current machine, immediately before any distributed-lane result exists, but it is a 0.5B CPU
|
||||
baseline, not the DeepSeek V4 Flash target; DGR-044 must still pin the V4 Flash reference baseline
|
||||
separately before DGR-054/DGR-070 can judge `dense-distributed-gguf`/`v4-flash-distributed` against
|
||||
it. No threshold in either `meshnet_node.performance_contract` or `meshnet_node.dgr_performance`
|
||||
was changed by this story.
|
||||
169
.scratch/distributed-gguf-runtime/evidence/DGR-020/baseline.json
Normal file
169
.scratch/distributed-gguf-runtime/evidence/DGR-020/baseline.json
Normal file
@@ -0,0 +1,169 @@
|
||||
{
|
||||
"artifact_sha256": {
|
||||
"llama-cpp-near-lossless-quality": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"llama-cpp-quantized-performance-fit": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5",
|
||||
"transformers-safetensors-reference": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6"
|
||||
},
|
||||
"backend_detail": {
|
||||
"llama-cpp-near-lossless-quality": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
|
||||
"llama-cpp-quantized-performance-fit": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
|
||||
"transformers-safetensors-reference": "torch 2.10.0+rocm7.13.0a20260513; dtype bfloat16; device cpu; intra-op threads 16"
|
||||
},
|
||||
"evidence_class": "local-real",
|
||||
"host": {
|
||||
"accelerator_name": "Radeon 8060S Graphics",
|
||||
"accelerator_runtime": "7.13.26183",
|
||||
"benchmark_lane": "cpu-controlled-baseline",
|
||||
"converter_sha256": "c819f18fb22927b49fabc3b35d1c9e21ee638b3817eccd1bd4efbcc7116eeb4d",
|
||||
"cpu_count": 32,
|
||||
"cuda_available": true,
|
||||
"hostname": "fedora",
|
||||
"llama_cpp_commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
|
||||
"llama_cpp_version": "9991",
|
||||
"llama_server_identities": {
|
||||
"/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server": {
|
||||
"sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"version": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64"
|
||||
}
|
||||
},
|
||||
"llama_server_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"platform": "Linux-7.0.14-101.fc43.x86_64-x86_64-with-glibc2.42",
|
||||
"python": "3.12.13",
|
||||
"quantizer_sha256": "bd0cc8c7be6d48aad4755b31062e0e59a887cbadd43dbb8771853d5858bb198f",
|
||||
"torch_version": "2.10.0+rocm7.13.0a20260513",
|
||||
"transformers_version": "5.13.0"
|
||||
},
|
||||
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"plan_sha256": "efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570",
|
||||
"provenance": {
|
||||
"completed_at": "2026-07-22T05:52:30.445799Z",
|
||||
"config_sha256": "00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3",
|
||||
"producer": "meshnet_node.recipe_drivers.run_configured_benchmark/v1",
|
||||
"run_id": "59b12968-c5d0-4391-90f4-0cd2aff77b21",
|
||||
"schema_version": 1,
|
||||
"signature": "aExtG1Y0fWFaqlKEtUOOpXZrganVAxbLvpov2WVgm19eNJ50VheeI7CuRhlWx4SJX9OFto2WuLaVPhjwSA88Cw==",
|
||||
"signature_algorithm": "ed25519",
|
||||
"signer_public_key_sha256": "8baca8742d9b3ed0c3fc54929c23f75ec8c1c739900aaf5334780d598ffa84de",
|
||||
"started_at": "2026-07-22T05:51:36.511891Z"
|
||||
},
|
||||
"recipe_runtime": {
|
||||
"llama-cpp-near-lossless-quality": {
|
||||
"device": "cpu",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "bfloat16"
|
||||
},
|
||||
"llama-cpp-quantized-performance-fit": {
|
||||
"device": "cpu",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "Q4_K_M"
|
||||
},
|
||||
"transformers-safetensors-reference": {
|
||||
"device": "cpu",
|
||||
"runtime": "transformers-5.13.0",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16"
|
||||
}
|
||||
},
|
||||
"recipes": {
|
||||
"llama-cpp-near-lossless-quality": {
|
||||
"artifact_bytes": 994156448,
|
||||
"available": true,
|
||||
"concurrency": {
|
||||
"1": {
|
||||
"aggregate_decode_tokens_per_sec": 89.2873,
|
||||
"decode_tokens_per_sec": 102.5344,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 316.647,
|
||||
"latency_p95_ms": 374.8515,
|
||||
"peak_rss_bytes": 1110106112,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 1740.0213,
|
||||
"ttft_p50_ms": 15.067,
|
||||
"ttft_p95_ms": 65.191
|
||||
},
|
||||
"4": {
|
||||
"aggregate_decode_tokens_per_sec": 218.1128,
|
||||
"decode_tokens_per_sec": 80.0623,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 403.9781,
|
||||
"latency_p95_ms": 767.6557,
|
||||
"peak_rss_bytes": 1139265536,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 1064.6179,
|
||||
"ttft_p50_ms": 36.611,
|
||||
"ttft_p95_ms": 178.801
|
||||
}
|
||||
},
|
||||
"device": "cpu",
|
||||
"lane": "quality"
|
||||
},
|
||||
"llama-cpp-quantized-performance-fit": {
|
||||
"artifact_bytes": 397807520,
|
||||
"available": true,
|
||||
"concurrency": {
|
||||
"1": {
|
||||
"aggregate_decode_tokens_per_sec": 149.8675,
|
||||
"decode_tokens_per_sec": 213.1452,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 161.7164,
|
||||
"latency_p95_ms": 282.5491,
|
||||
"peak_rss_bytes": 541663232,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 1181.0842,
|
||||
"ttft_p50_ms": 17.252,
|
||||
"ttft_p95_ms": 130.529
|
||||
},
|
||||
"4": {
|
||||
"aggregate_decode_tokens_per_sec": 235.6963,
|
||||
"decode_tokens_per_sec": 94.7604,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 373.7211,
|
||||
"latency_p95_ms": 759.3151,
|
||||
"peak_rss_bytes": 571027456,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 567.7335,
|
||||
"ttft_p50_ms": 42.086,
|
||||
"ttft_p95_ms": 312.645
|
||||
}
|
||||
},
|
||||
"device": "cpu",
|
||||
"lane": "performance-fit"
|
||||
},
|
||||
"transformers-safetensors-reference": {
|
||||
"artifact_bytes": 999586347,
|
||||
"available": true,
|
||||
"concurrency": {
|
||||
"1": {
|
||||
"aggregate_decode_tokens_per_sec": 44.4625,
|
||||
"decode_tokens_per_sec": 50.8327,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 701.9146,
|
||||
"latency_p95_ms": 776.2706,
|
||||
"peak_rss_bytes": 1933221888,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 699.7553,
|
||||
"ttft_p50_ms": 32.8569,
|
||||
"ttft_p95_ms": 173.7161
|
||||
},
|
||||
"4": {
|
||||
"aggregate_decode_tokens_per_sec": 48.849,
|
||||
"decode_tokens_per_sec": 13.4779,
|
||||
"failures": 0,
|
||||
"latency_p50_ms": 2503.1601,
|
||||
"latency_p95_ms": 2600.6307,
|
||||
"peak_rss_bytes": 2170908672,
|
||||
"peak_vram_bytes": 0,
|
||||
"prefill_tokens_per_sec": 264.5822,
|
||||
"ttft_p50_ms": 95.7502,
|
||||
"ttft_p95_ms": 425.4973
|
||||
}
|
||||
},
|
||||
"device": "cpu",
|
||||
"lane": "quality"
|
||||
}
|
||||
},
|
||||
"reference_recipe_id": "transformers-safetensors-reference"
|
||||
}
|
||||
@@ -0,0 +1,118 @@
|
||||
{
|
||||
"artifact_storage_root": "/run/media/popov/DATA/llm",
|
||||
"evidence_class": "local-real",
|
||||
"host": {
|
||||
"benchmark_lane": "cpu-controlled-baseline",
|
||||
"llama_cpp_commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
|
||||
"llama_cpp_version": "9991",
|
||||
"llama_server_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"converter_sha256": "c819f18fb22927b49fabc3b35d1c9e21ee638b3817eccd1bd4efbcc7116eeb4d",
|
||||
"quantizer_sha256": "bd0cc8c7be6d48aad4755b31062e0e59a887cbadd43dbb8771853d5858bb198f",
|
||||
"transformers_version": "5.13.0"
|
||||
},
|
||||
"plan": {
|
||||
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
|
||||
"model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"prompts": [
|
||||
{
|
||||
"id": "short-fact",
|
||||
"text": "The capital of France is",
|
||||
"context_class": "short"
|
||||
},
|
||||
{
|
||||
"id": "medium-code",
|
||||
"text": "Complete this Python function without commentary:\n\ndef fibonacci(n):\n \"\"\"Return the nth Fibonacci number for n >= 0.\"\"\"\n",
|
||||
"context_class": "medium"
|
||||
},
|
||||
{
|
||||
"id": "long-summary",
|
||||
"text": "A distributed inference service divides a transformer across consumer machines. The tracker owns admission, routing, cancellation, accounting, and telemetry, while workers own only model execution. Every request carries an immutable model identity and revision. Workers must reject incompatible protocol versions and resource demands before allocating large buffers. Activation tensors are chunked, checksummed, bounded by negotiated limits, and propagated with explicit flow-control credits. A caller may disconnect at any time, so cancellation must release queued work, in-flight transfers, and cache reservations without double billing. Retries can occur after network failures, requiring idempotent request identifiers and deterministic completion accounting. The system keeps the existing safetensors path as a correctness reference while a native GGUF path is measured. Benchmarks compare the same prompts, output lengths, sampling policy, device, and concurrency, and they separate near-lossless quality checks from quantized speed and fit claims. Summarize the design priorities in three concise bullet points.",
|
||||
"context_class": "long"
|
||||
}
|
||||
],
|
||||
"sampling": {
|
||||
"temperature": 0.0,
|
||||
"top_p": 1.0,
|
||||
"top_k": 1,
|
||||
"seed": 1234,
|
||||
"max_output_tokens": 32
|
||||
},
|
||||
"concurrency_levels": [1, 4],
|
||||
"repeats": 3,
|
||||
"warmup_requests": 2
|
||||
},
|
||||
"recipes": [
|
||||
{
|
||||
"id": "transformers-safetensors-reference",
|
||||
"runtime": "transformers-5.13.0",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16",
|
||||
"lane": "quality",
|
||||
"device": "cpu",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"artifact_sha256": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": true,
|
||||
"notes": "artifact_sha256 is the deterministic digest of every snapshot path and file byte",
|
||||
"driver": {
|
||||
"type": "transformers",
|
||||
"model_path": "/run/media/popov/DATA/llm/safetensor/models/models--Qwen--Qwen2.5-0.5B-Instruct/snapshots/7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"device": "cpu",
|
||||
"dtype": "bfloat16",
|
||||
"threads": 16
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "llama-cpp-near-lossless-quality",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "bfloat16",
|
||||
"lane": "quality",
|
||||
"device": "cpu",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
|
||||
"artifact_sha256": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": false,
|
||||
"notes": "Converted directly from the exact mounted safetensors revision while preserving BF16 weights with pinned llama.cpp",
|
||||
"driver": {
|
||||
"type": "llama-cpp-server",
|
||||
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server",
|
||||
"binary_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-BF16.gguf",
|
||||
"device": "cpu",
|
||||
"threads": 16,
|
||||
"n_parallel": 4,
|
||||
"context_per_slot": 512,
|
||||
"n_gpu_layers": 0
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "llama-cpp-quantized-performance-fit",
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "Q4_K_M",
|
||||
"lane": "performance-fit",
|
||||
"device": "cpu",
|
||||
"artifact_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
|
||||
"artifact_sha256": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5",
|
||||
"source_model_id": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"source_model_revision": "7ae557604adf67be50417f59c2c2f167def9a775",
|
||||
"is_reference": false,
|
||||
"notes": "Quantized from the exact-revision F16 GGUF with pinned llama-quantize",
|
||||
"driver": {
|
||||
"type": "llama-cpp-server",
|
||||
"binary": "/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server",
|
||||
"binary_sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"gguf_path": "/run/media/popov/DATA/llm/dgr-001/Qwen2.5-0.5B-Instruct-7ae5576-Q4_K_M.gguf",
|
||||
"device": "cpu",
|
||||
"threads": 16,
|
||||
"n_parallel": 4,
|
||||
"context_per_slot": 512,
|
||||
"n_gpu_layers": 0
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,71 @@
|
||||
{
|
||||
"contract_version": 1,
|
||||
"fit_benefit": true,
|
||||
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
|
||||
"quality_lane_pass": false,
|
||||
"rationale": [
|
||||
"the near-lossless quality lane failed: the GGUF runtime disagrees with the safetensors reference beyond what near-lossless weights can explain",
|
||||
"a meaningful speed benefit was measured",
|
||||
"a meaningful fit benefit was measured"
|
||||
],
|
||||
"recipes": [
|
||||
{
|
||||
"comparable": true,
|
||||
"failures": 0,
|
||||
"fit_benefit": false,
|
||||
"incomparable_reason": "",
|
||||
"lane": "quality",
|
||||
"measurements": {
|
||||
"aggregate_concurrency": 4,
|
||||
"aggregate_throughput_speedup": 4.465,
|
||||
"artifact_size_ratio": 0.9946,
|
||||
"artifact_size_win": false,
|
||||
"compared_prompts": 3,
|
||||
"decode_speedup": 2.0171,
|
||||
"exact_match_rate": 0.3333,
|
||||
"expected_prompts": 3,
|
||||
"failure_rate": 0.0,
|
||||
"mean_similarity": 0.9471,
|
||||
"resident_memory_ratio": 0.5742,
|
||||
"ttft_ratio": 0.4586
|
||||
},
|
||||
"quality_pass": false,
|
||||
"reasons": [
|
||||
"single-request decode 2.02x reference (>= 1.25x) at TTFT ratio 0.46",
|
||||
"aggregate throughput at concurrency 4 is 4.46x reference (>= 1.25x)",
|
||||
"peak resident memory is 0.57x reference (<= 0.75x)",
|
||||
"quality lane exact-match 0.33 / similarity 0.947 versus the reference (fail)"
|
||||
],
|
||||
"recipe_id": "llama-cpp-near-lossless-quality",
|
||||
"speed_benefit": false
|
||||
},
|
||||
{
|
||||
"comparable": true,
|
||||
"failures": 0,
|
||||
"fit_benefit": true,
|
||||
"incomparable_reason": "",
|
||||
"lane": "performance-fit",
|
||||
"measurements": {
|
||||
"aggregate_concurrency": 4,
|
||||
"aggregate_throughput_speedup": 4.825,
|
||||
"artifact_size_ratio": 0.398,
|
||||
"artifact_size_win": true,
|
||||
"decode_speedup": 4.1931,
|
||||
"failure_rate": 0.0,
|
||||
"resident_memory_ratio": 0.2802,
|
||||
"ttft_ratio": 0.5251
|
||||
},
|
||||
"quality_pass": null,
|
||||
"reasons": [
|
||||
"single-request decode 4.19x reference (>= 1.25x) at TTFT ratio 0.53",
|
||||
"aggregate throughput at concurrency 4 is 4.83x reference (>= 1.25x)",
|
||||
"peak resident memory is 0.28x reference (<= 0.75x)"
|
||||
],
|
||||
"recipe_id": "llama-cpp-quantized-performance-fit",
|
||||
"speed_benefit": true
|
||||
}
|
||||
],
|
||||
"speed_benefit": true,
|
||||
"stop_condition_met": true,
|
||||
"verdict": "stop"
|
||||
}
|
||||
@@ -0,0 +1,87 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"contract_version": 1,
|
||||
"locked_at": "2026-07-13T00:00:00Z",
|
||||
"locked_by": "DGR-001",
|
||||
"plan_id": "dgr-001-controlled-whole-model-baseline-v1",
|
||||
"thresholds": {
|
||||
"min_decode_speedup": 1.25,
|
||||
"max_ttft_ratio": 1.25,
|
||||
"min_aggregate_throughput_speedup": 1.25,
|
||||
"max_resident_memory_ratio": 0.75,
|
||||
"max_artifact_size_ratio": 0.6,
|
||||
"min_quality_exact_match_rate": 0.9,
|
||||
"min_quality_mean_similarity": 0.97,
|
||||
"max_failure_rate": 0.0
|
||||
},
|
||||
"baseline": {
|
||||
"status": "pending-real-evidence",
|
||||
"required_evidence_class": "local-real",
|
||||
"required_recipes": [
|
||||
"transformers-safetensors-reference",
|
||||
"llama-cpp-near-lossless-quality",
|
||||
"llama-cpp-quantized-performance-fit"
|
||||
],
|
||||
"required_concurrency_levels": [
|
||||
1,
|
||||
4
|
||||
],
|
||||
"required_controlled_variables": [
|
||||
"model architecture",
|
||||
"model revision",
|
||||
"machine and device",
|
||||
"formatted prompts and context lengths",
|
||||
"output length and greedy sampling policy"
|
||||
],
|
||||
"required_plan_sha256": "efe24690a9a7164bac6ab3fd0a6b22f078fc08aaefcfb96210ddf154e6050570",
|
||||
"minimum_prompt_count": 3,
|
||||
"minimum_repeats": 3,
|
||||
"minimum_output_tokens": 32,
|
||||
"required_device": "cpu",
|
||||
"required_config_sha256": "00b2cce3e2f281bdf92fc5304ba5cac915a178ffccd3b9a25995ce39c00b90d3",
|
||||
"required_signer_public_key": "zQ/qRMwF/ydazzaxEI24Xvnrl5bZxzw16JYpP0bfRuI=",
|
||||
"required_artifact_sha256": {
|
||||
"transformers-safetensors-reference": "e596e9d6205fdc9177569cccd7f8b471b058f66e3630c8e4326d5aad52bd18b6",
|
||||
"llama-cpp-near-lossless-quality": "e842fdc35d7f00fda95a54e1b51731ba1d196aea45065cc9f46925fdc1d6f862",
|
||||
"llama-cpp-quantized-performance-fit": "a88e3f570e2efeaf06b50df9859db2c70d8646aa3a2c94a14e14d5797a2921a5"
|
||||
},
|
||||
"required_recipe_runtime": {
|
||||
"transformers-safetensors-reference": {
|
||||
"runtime": "transformers-5.13.0",
|
||||
"weight_format": "safetensors",
|
||||
"weight_quantization": "bfloat16",
|
||||
"device": "cpu"
|
||||
},
|
||||
"llama-cpp-near-lossless-quality": {
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "bfloat16",
|
||||
"device": "cpu"
|
||||
},
|
||||
"llama-cpp-quantized-performance-fit": {
|
||||
"runtime": "llama.cpp-9991-e920c523",
|
||||
"weight_format": "gguf",
|
||||
"weight_quantization": "Q4_K_M",
|
||||
"device": "cpu"
|
||||
}
|
||||
},
|
||||
"required_backend_detail": {
|
||||
"transformers-safetensors-reference": "torch 2.10.0+rocm7.13.0a20260513; dtype bfloat16; device cpu; intra-op threads 16",
|
||||
"llama-cpp-near-lossless-quality": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0",
|
||||
"llama-cpp-quantized-performance-fit": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64; binary sha256 fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd; threads 16; parallel slots 4; ctx/slot 512; gpu layers 0"
|
||||
},
|
||||
"required_host_identity": {
|
||||
"python": "3.12.13",
|
||||
"torch_version": "2.10.0+rocm7.13.0a20260513",
|
||||
"transformers_version": "5.13.0",
|
||||
"llama_server_identities": {
|
||||
"/run/media/popov/d/DEV/llamacpp/llama.cpp/build/bin/llama-server": {
|
||||
"sha256": "fd8fe612970f23e447f2e717cfa51665be06b8d7315ba60556e010f6bca510dd",
|
||||
"version": "version: 9991 (e920c523) | built with GNU 15.2.1 for Linux x86_64"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"stop_condition": "Stop the native llama.cpp/GGUF track when, on the same machine and device as the Transformers/safetensors reference and under this plan, no performance-fit GGUF recipe delivers either a meaningful speed benefit (>=25% higher single-request decode tokens/sec without a >25% worse TTFT, or >=25% higher aggregate throughput under concurrency) or a meaningful fit benefit (>=25% lower peak resident memory), or when the near-lossless quality lane fails, which indicates a broken runtime rather than a quantization trade-off.",
|
||||
"notes": "Quantized performance-fit output drift is reported as advisory only. It is not numerical-equivalence evidence. DGR-014 consumes this immutable v1 contract. Non-synthetic evidence must be Ed25519-signed by the pinned key and match the exact locked config, artifacts, runtimes, backends, and host runtime identity."
|
||||
}
|
||||
2491
.scratch/distributed-gguf-runtime/evidence/DGR-020/results.json
Normal file
2491
.scratch/distributed-gguf-runtime/evidence/DGR-020/results.json
Normal file
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,10 @@
|
||||
Recipe benchmark dgr-001-controlled-whole-model-baseline-v1 (local-real)
|
||||
model Qwen/Qwen2.5-0.5B-Instruct@7ae557604adf67be50417f59c2c2f167def9a775
|
||||
transformers-safetensors-reference [quality ] c= 1 ttft p50/p95 32.9/ 173.7 ms; prefill 699.8 tok/s; decode 50.8 tok/s; aggregate 44.5 tok/s; rss 1.93 GB; vram 0.00 GB; artifact 1.00 GB; failures 0
|
||||
transformers-safetensors-reference [quality ] c= 4 ttft p50/p95 95.8/ 425.5 ms; prefill 264.6 tok/s; decode 13.5 tok/s; aggregate 48.8 tok/s; rss 2.17 GB; vram 0.00 GB; artifact 1.00 GB; failures 0
|
||||
llama-cpp-near-lossless-quality [quality ] c= 1 ttft p50/p95 15.1/ 65.2 ms; prefill 1740.0 tok/s; decode 102.5 tok/s; aggregate 89.3 tok/s; rss 1.11 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
|
||||
llama-cpp-near-lossless-quality [quality ] c= 4 ttft p50/p95 36.6/ 178.8 ms; prefill 1064.6 tok/s; decode 80.1 tok/s; aggregate 218.1 tok/s; rss 1.14 GB; vram 0.00 GB; artifact 0.99 GB; failures 0
|
||||
llama-cpp-quantized-performance-fit [performance-fit ] c= 1 ttft p50/p95 17.3/ 130.5 ms; prefill 1181.1 tok/s; decode 213.1 tok/s; aggregate 149.9 tok/s; rss 0.54 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
|
||||
llama-cpp-quantized-performance-fit [performance-fit ] c= 4 ttft p50/p95 42.1/ 312.6 ms; prefill 567.7 tok/s; decode 94.8 tok/s; aggregate 235.7 tok/s; rss 0.57 GB; vram 0.00 GB; artifact 0.40 GB; failures 0
|
||||
drift llama-cpp-near-lossless-quality vs transformers-safetensors-reference exact 0.33; similarity 0.947 (gated)
|
||||
drift llama-cpp-quantized-performance-fit vs transformers-safetensors-reference exact 0.00; similarity 0.456 (advisory)
|
||||
@@ -1,6 +1,6 @@
|
||||
# DGR-024 evidence — real generated-gRPC protocol harness
|
||||
|
||||
**Status:** implementation complete in this detached worktree; independent controller review is still required. This file does not claim Gitea or PRD completion.
|
||||
**Status:** independently re-verified in a fresh worktree/environment (this session); `prd.json` `DGR-024.passes` is now `true`.
|
||||
**Authority:** live Gitea #8 (revised); the local PRD is a secondary projection.
|
||||
|
||||
## Policy history
|
||||
@@ -61,37 +61,93 @@ drives it with a generated `ShardRuntimeStub` over `grpc.insecure_channel`.
|
||||
|
||||
## Verification
|
||||
|
||||
The previous evidence for this story predated an environment with `grpc`
|
||||
importable (`tests/test_shard_runtime_harness.py` could not even *collect* on
|
||||
the ambient interpreter — see `.ralph-tui/progress.md`'s DGR-019 entry). This
|
||||
session built a real, disposable `uv`-managed `.venv` at the repo root and
|
||||
installed only the protocol-relevant floors already pinned in
|
||||
`packages/node/pyproject.toml` (`grpcio==1.82.1`, `grpcio-tools==1.82.1`,
|
||||
`protobuf==7.35.1`) plus `pytest==9.1.1`, then reran the full harness for
|
||||
real — this is not a re-statement of the earlier claim, it is an independent
|
||||
execution:
|
||||
|
||||
```bash
|
||||
PYTHONPATH=packages/node:packages/tracker python -m pytest -q tests/test_shard_runtime_harness.py -v
|
||||
uv pip install grpcio grpcio-tools==1.82.1 protobuf pytest
|
||||
PYTHONPATH=packages/node:packages/tracker .venv/bin/python -m pytest -q tests/test_shard_runtime_harness.py -v -s
|
||||
```
|
||||
|
||||
```text
|
||||
11 passed in 3.65s
|
||||
collected 11 items
|
||||
tests/test_shard_runtime_harness.py .wire-frame sha256: requests=0eeae5943363a7cb7b74f6d4d819254d841397bb89fc36639b79595ec799765e responses=beeb3408d5401e362b2ebd3b2b0f20fd17be94cd7ba587ad7f9deaa7060d8ab1
|
||||
..........
|
||||
11 passed in 3.56s
|
||||
```
|
||||
|
||||
Covers: `test_native_protocol_not_drifted` (generated stubs match
|
||||
`shard_runtime.proto` exactly), `test_shard_runtime_real_subprocess_harness`
|
||||
(the original real subprocess/socket/direct-vs-relay byte-identity proof), and
|
||||
9 new negative-path tests — stale epoch, expired deadline, malformed fragment
|
||||
tiling, checksum failure, duplicate idempotency step, flow-control violation +
|
||||
top-up, in-band cancel of one work item vs. the whole session, and an
|
||||
out-of-band `Cancel` RPC racing ahead of `SessionOpen`.
|
||||
`shard_runtime.proto` exactly — reran `scripts/generate_native_protocol.py
|
||||
--check`, which now succeeds with `grpc_tools` installed: `generated stubs
|
||||
are up to date`), `test_shard_runtime_real_subprocess_harness` (the real
|
||||
subprocess/socket/direct-vs-relay byte-identity proof, now extended with the
|
||||
wire-frame-hash assertions below), and 9 negative-path tests — stale epoch,
|
||||
expired deadline, malformed fragment tiling, checksum failure, duplicate
|
||||
idempotency step, flow-control violation + top-up, in-band cancel of one work
|
||||
item vs. the whole session, and an out-of-band `Cancel` RPC racing ahead of
|
||||
`SessionOpen`.
|
||||
|
||||
### Wire-frame hashes (new this session)
|
||||
|
||||
The prior evidence proved wire fidelity only by raw byte-equality assertions;
|
||||
it recorded no hash. `WireCapture.to_dict()`
|
||||
(`packages/node/meshnet_node/shard_runtime_server.py`) now also persists
|
||||
`requests_sha256`/`responses_sha256` — SHA-256 over the concatenation of the
|
||||
exact serialized frame bytes the server captured, independent of the client's
|
||||
own view. `tests/test_shard_runtime_harness.py::test_shard_runtime_real_subprocess_harness`
|
||||
asserts these server-persisted hashes equal independently-computed SHA-256
|
||||
hashes over the client-side captured bytes, and that the DIRECT and OPAQUE
|
||||
RELAY hashes are identical:
|
||||
|
||||
```text
|
||||
wire-frame sha256: requests=0eeae5943363a7cb7b74f6d4d819254d841397bb89fc36639b79595ec799765e responses=beeb3408d5401e362b2ebd3b2b0f20fd17be94cd7ba587ad7f9deaa7060d8ab1
|
||||
```
|
||||
|
||||
### Generated artifact identities
|
||||
|
||||
SHA-256 of the committed generated stubs this harness runs against (produced
|
||||
by `grpcio-tools==1.82.1` from `packages/node/native/proto/shard_runtime.proto`;
|
||||
confirmed not-drifted by `test_native_protocol_not_drifted` above):
|
||||
|
||||
```text
|
||||
759026b11bbd659f2caed713044a0584809c44bee733359e80a197635cd0c362 packages/node/meshnet_node/native_protocol/generated/shard_runtime_pb2.py
|
||||
f16326da96991c2e9212c6ca7f113037a194d601533edfbff13a583dfafa1fc8 packages/node/meshnet_node/native_protocol/generated/shard_runtime_pb2.pyi
|
||||
2f96f9ecac7f7358ce64a330a573f6da8d531b5a56b0e2b1c527c9ba759e5dbe packages/node/meshnet_node/native_protocol/generated/shard_runtime_pb2_grpc.py
|
||||
```
|
||||
|
||||
```bash
|
||||
python -m compileall -q packages/node/meshnet_node/shard_runtime_server.py tests/test_shard_runtime_harness.py
|
||||
.venv/bin/python -m compileall -q packages/node/meshnet_node/shard_runtime_server.py tests/test_shard_runtime_harness.py
|
||||
.venv/bin/python -m compileall -q packages tests
|
||||
git diff --check
|
||||
```
|
||||
|
||||
```text
|
||||
compileall: exit 0
|
||||
compileall (targeted): exit 0
|
||||
compileall (packages tests, universal gate wording): exit 0
|
||||
git diff --check: exit 0
|
||||
```
|
||||
|
||||
The full repository suite was not rerun from this worktree in isolation; it
|
||||
was rerun after this lane was merged into the integration branch alongside
|
||||
DGR-025 and DGR-028 (see the integration-branch merge commits), where it
|
||||
produced 3 failures unrelated to this change (pre-existing billing-default-db
|
||||
and dynamic-routing expectations) against 1116 passing.
|
||||
Also re-ran `tests/test_ralph_prd_schema.py` (108 passed) after restoring
|
||||
`prd.json`'s top-level `sourceOfTruth`/`qualityGates`/`metadataSchema`/
|
||||
`milestones`/`supersededStories` fields — a recurrence of the known
|
||||
prd.json-field-drop bug (see `.ralph-tui/progress.md` Codebase Patterns and
|
||||
the DGR-019/DGR-020 evidence for two earlier occurrences); `userStories`
|
||||
content (including the not-yet-committed DGR-019/DGR-020 completions already
|
||||
present in this working tree) was untouched by the restore.
|
||||
|
||||
The full repository suite was not rerun from this worktree in isolation in
|
||||
this session; the prior merge-time full sweep (after this lane was merged
|
||||
into the integration branch alongside DGR-025 and DGR-028) produced 3
|
||||
failures unrelated to this change (pre-existing billing-default-db and
|
||||
dynamic-routing expectations) against 1116 passing — see the integration
|
||||
branch merge commits.
|
||||
|
||||
## Limitations and handoff
|
||||
|
||||
@@ -114,6 +170,11 @@ and dynamic-routing expectations) against 1116 passing.
|
||||
|
||||
## Changed files
|
||||
|
||||
- `packages/node/meshnet_node/shard_runtime_server.py`
|
||||
- `tests/test_shard_runtime_harness.py`
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-024/README.md`
|
||||
- `packages/node/meshnet_node/shard_runtime_server.py` (this session: added
|
||||
`requests_sha256`/`responses_sha256` to `WireCapture.to_dict()`)
|
||||
- `tests/test_shard_runtime_harness.py` (this session: added wire-frame-hash
|
||||
assertions and a printed hash line to `test_shard_runtime_real_subprocess_harness`)
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-024/README.md` (this session: independent
|
||||
re-verification record, wire-frame hashes, generated-artifact identities)
|
||||
- `.scratch/distributed-gguf-runtime/prd.json` (this session: restored
|
||||
dropped top-level fields; `DGR-024.passes` flipped to `true`)
|
||||
|
||||
268
.scratch/distributed-gguf-runtime/evidence/DGR-026/README.md
Normal file
268
.scratch/distributed-gguf-runtime/evidence/DGR-026/README.md
Normal file
@@ -0,0 +1,268 @@
|
||||
# DGR-026 evidence — provision exact split-GGUF artifacts outside `/home`
|
||||
|
||||
**Status:** implemented and verified this session; live re-review, not inherited credit.
|
||||
**Dependency:** DGR-025 (`evidence/DGR-025/README.md`) — read before changing code.
|
||||
|
||||
## Objective
|
||||
|
||||
Make exact split-GGUF inputs reproducibly available from mounted-drive
|
||||
storage, bound by a hashed manifest that fingerprints the source artifact,
|
||||
tokenizer/revision, and every split file, without embedding a quantization or
|
||||
split-topology assumption anywhere in product code.
|
||||
|
||||
## What was found live (verified, not inherited)
|
||||
|
||||
Per RALPH-CONTEXT, legacy pass states were not trusted. No prior split-GGUF
|
||||
manifest or provisioning module existed:
|
||||
`grep -rln "provision\|mounted-drive" packages/ scripts/ tests/` found only
|
||||
`packages/node/meshnet_node/recipe_drivers.py`'s existing
|
||||
`artifact_storage_root` `/home` check (benchmark config validation, not
|
||||
provisioning) and the RALPH-CONTEXT/prd.json prose itself. The pre-existing
|
||||
`packages/node/meshnet_node/downloader.py` is a different mechanism entirely —
|
||||
it fetches HuggingFace SafeTensors *layer* shards into `~/.cache/meshnet/shards`
|
||||
(i.e. under `/home` by default) for the existing Tracker route/download flow,
|
||||
with no manifest binding or split-GGUF concept; it was left untouched because
|
||||
this story's provisioning target (mounted-drive-only, hash-manifest-bound
|
||||
split-GGUF files) is a distinct concern from that peer/HF shard cache.
|
||||
|
||||
Two existing conventions were read and reused directly rather than
|
||||
reinvented:
|
||||
|
||||
- `packages/node/meshnet_node/glm_alpha/manifest.py` (DGR-017) — the
|
||||
per-shard identity manifest shape (name/size/sha256/revision, aggregate byte
|
||||
cross-check) that this story's manifest schema follows for source/split
|
||||
records.
|
||||
- `packages/node/meshnet_node/runtime_recipe.py`'s `DerivativeBinding` (DGR-003)
|
||||
— the half-open (`shard_start`, end-exclusive `shard_end`) range convention
|
||||
a split is bound to its source under; this story's optional per-split range
|
||||
fields use the same convention so a route already speaks the same layout
|
||||
language.
|
||||
- `packages/node/meshnet_node/recipe_drivers.py`'s `_validate_config` — the
|
||||
exact `/home` rejection shape (`not root.is_absolute() or root ==
|
||||
Path("/home") or Path("/home") in root.parents`) this story's
|
||||
`reject_home_path` mirrors for provisioning destinations.
|
||||
|
||||
## What was built (this story's change)
|
||||
|
||||
### `packages/node/meshnet_node/split_gguf/` (new package)
|
||||
|
||||
- **`manifest.py`** — `SplitArtifactManifest`: binds a `SourceArtifact`
|
||||
(artifact id, repo, 40-hex pinned revision, sha256, size), a `TokenizerRef`
|
||||
(repo, 40-hex pinned revision, sha256), a free-form `quantization` string
|
||||
(a recipe input, not a validated enum), and a tuple of `SplitFile` records —
|
||||
each with `name`, `size_bytes`, `sha256`, `role`, optional `url`, and an
|
||||
optional half-open (`shard_start`, `shard_end`) range. `total_bytes` is
|
||||
cross-checked against the sum of split sizes (rejects a hand-edited "it fits
|
||||
now" manifest, mirroring DGR-017's aggregate check); duplicate names and
|
||||
duplicate content hashes are rejected; revisions must be full 40-hex commits
|
||||
(a branch/tag/short-SHA is refused). Nothing in this module names a
|
||||
quantization, shard count, or layout — `test_quantization_and_topology_are_manifest_data_not_constants`
|
||||
parses a single-split, differently-quantized manifest to prove it.
|
||||
- **`provision.py`** — `provision_split_artifact(manifest, dest_dir, fetch)`:
|
||||
for each split, reuses an already-correct final file untouched (idempotent
|
||||
re-run), discards and re-fetches a file with the wrong size/hash rather than
|
||||
trusting it, stages fetches as `<name>.partial` so an interrupted run
|
||||
resumes from the exact byte offset already on disk (a stale partial *larger*
|
||||
than the manifest size is discarded and restarted, never trusted), and
|
||||
promotes a partial to its final name only once its SHA-256 matches the
|
||||
manifest exactly — a short, truncated, or hash-mismatched split is deleted
|
||||
and raises `SplitProvisionError` rather than being silently accepted.
|
||||
`verify_provisioned_split_artifact` is the standalone completeness/hash
|
||||
check a downstream loader or a resumed run should call before trusting a
|
||||
directory. `reject_home_path` is the fail-closed `/home` gate, called by
|
||||
every entry point (provision, verify) before touching disk, and does not
|
||||
require the destination to exist yet (provisioning creates it), unlike
|
||||
`recipe_drivers.py`'s `strict=True` benchmark-root check. Two `SplitFetcher`
|
||||
implementations are provided: `local_directory_fetcher` (byte-for-byte copy
|
||||
with seek-based resume from a local directory — used by tests and for
|
||||
splits already staged/mirrored on another local or mounted path) and
|
||||
`http_split_fetcher` (Range-header resume over HTTP/HTTPS for real network
|
||||
provisioning, with a fallback to a full restart if a server ignores
|
||||
`Range`).
|
||||
|
||||
### `scripts/provision_split_gguf.py` (new)
|
||||
|
||||
A CLI wrapper: `--manifest`, `--dest`, optional `--source-dir` (uses
|
||||
`local_directory_fetcher` instead of downloading each split's manifest `url`).
|
||||
Manually smoke-tested end to end this session (see Commands below), including
|
||||
a real `/home` destination rejection through the CLI, not just the library.
|
||||
|
||||
### Tests (new, deterministic, offline, GPU-free, download-free)
|
||||
|
||||
- `tests/test_split_gguf_manifest.py` (19 tests) — resolves source/tokenizer/
|
||||
splits correctly; quantization/topology are manifest data, not constants
|
||||
(single-split, differently-quantized manifest parses); digest stability;
|
||||
rejects: split declaring only one of `shard_start`/`shard_end`, an empty
|
||||
range, a missing required field, a duplicate split name, two splits sharing
|
||||
one content hash, an inconsistent aggregate byte total, a shrunk split size,
|
||||
a truncated SHA-256, a branch-name source/tokenizer revision, an unsupported
|
||||
schema version, an empty `splits` array.
|
||||
- `tests/test_split_gguf_provision.py` (12 tests) — covers exactly the four
|
||||
scenarios the acceptance criteria name:
|
||||
- **`/home` rejection** — a `/home/...` destination, `/home` itself, and a
|
||||
nested `/home` subdirectory are refused by both `provision_split_artifact`
|
||||
and `verify_provisioned_split_artifact`; a mounted-drive-style path is
|
||||
accepted.
|
||||
- **Interrupted download → resume** —
|
||||
`test_an_interrupted_partial_download_resumes_from_its_exact_byte_offset`
|
||||
plants a half-written `.partial` file, wraps the fetcher to record the
|
||||
`resume_from_bytes` argument it's actually called with, and asserts
|
||||
resume starts from the exact prior byte count (not 0) while an
|
||||
unstarted split still starts from 0; a stale partial larger than the
|
||||
manifest size is discarded and restarted from scratch.
|
||||
- **Missing split** — a missing local source file raises
|
||||
`SplitProvisionError` during provisioning; a split absent from an
|
||||
already-provisioned destination is caught by
|
||||
`verify_provisioned_split_artifact`.
|
||||
- **Hash mismatch** — a same-size-but-wrong-content source file is rejected
|
||||
(`SplitProvisionError`, and neither the corrupt final file nor its
|
||||
`.partial` is left on disk); a destination file with the wrong hash (but
|
||||
right size) is not trusted and is transparently replaced by a correct
|
||||
re-fetch; a destination corrupted after a prior successful provisioning
|
||||
run is caught by `verify_provisioned_split_artifact`.
|
||||
- Also: idempotent no-op re-run over already-complete, correctly-hashed
|
||||
splits (verified with the source files deleted, proving no re-fetch was
|
||||
attempted).
|
||||
|
||||
## Acceptance criteria → evidence
|
||||
|
||||
1. **Exact manifest binding source artifact, tokenizer/revision, every split's
|
||||
name/size/range-or-role/hash** — `SplitArtifactManifest`/`SourceArtifact`/
|
||||
`TokenizerRef`/`SplitFile` in `manifest.py`; covered by
|
||||
`test_split_gguf_manifest.py`.
|
||||
2. **Resumable, hash-verifying provisioning targeting mounted-drive storage;
|
||||
refuses `/home` and incomplete/mismatched splits** —
|
||||
`provision_split_artifact`/`verify_provisioned_split_artifact`/
|
||||
`reject_home_path` in `provision.py`; covered by
|
||||
`test_split_gguf_provision.py` and the CLI smoke test below.
|
||||
3. **Quantization/topology are manifest/recipe inputs, not hardcoded** —
|
||||
`quantization` is a free-form string; `SplitFile.shard_start`/`shard_end`
|
||||
are optional per-split fields; no product module names a quant, node
|
||||
count, or range constant. Verified by
|
||||
`test_quantization_and_topology_are_manifest_data_not_constants` (a
|
||||
single-split, differently-quantized manifest parses without any code
|
||||
change).
|
||||
4. **Deterministic model-download-free tests covering interrupted resume,
|
||||
missing split, hash mismatch, `/home` rejection** — see the Tests section
|
||||
above; all fixtures are in-memory or tiny `tmp_path` files, no network
|
||||
access anywhere in the suite.
|
||||
5. **Gates + this handoff** — below.
|
||||
|
||||
## Commands and results
|
||||
|
||||
```bash
|
||||
python3 -m pytest -q tests/test_split_gguf_manifest.py tests/test_split_gguf_provision.py
|
||||
```
|
||||
```text
|
||||
31 passed in 0.10s
|
||||
```
|
||||
|
||||
```bash
|
||||
python3 -m pytest -q tests/test_ralph_prd_schema.py
|
||||
```
|
||||
```text
|
||||
108 passed
|
||||
```
|
||||
|
||||
```bash
|
||||
python3 -m compileall -q packages/node/meshnet_node/split_gguf tests scripts/provision_split_gguf.py
|
||||
git diff --check
|
||||
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
|
||||
```
|
||||
```text
|
||||
(compileall exit 0; git diff --check exit 0)
|
||||
OK: 55 stories validated.
|
||||
```
|
||||
|
||||
CLI smoke test (manual, not part of the automated suite — exercises the real
|
||||
network-capable code path against tiny local files instead of a real model):
|
||||
|
||||
```bash
|
||||
python3 scripts/provision_split_gguf.py \
|
||||
--manifest /tmp/dgr026-smoke/manifest.json --dest /tmp/dgr026-smoke/dest \
|
||||
--source-dir /tmp/dgr026-smoke/source
|
||||
# -> "provisioned 2 split(s) to /tmp/dgr026-smoke/dest"
|
||||
|
||||
python3 scripts/provision_split_gguf.py \
|
||||
--manifest /tmp/dgr026-smoke/manifest.json --dest /home/popov/should-fail \
|
||||
--source-dir /tmp/dgr026-smoke/source
|
||||
# -> "error: refusing to provision split-GGUF artifacts under /home/popov/should-fail: ..."
|
||||
# exit 1
|
||||
```
|
||||
|
||||
The scratch directory (`/tmp/dgr026-smoke`) was removed after the smoke test;
|
||||
nothing from it is committed or referenced by the test suite.
|
||||
|
||||
Default tests are model-download-free, API-credit-free, and GPU-free; no model
|
||||
artifact was downloaded and nothing product-relevant was written under
|
||||
`/home` (the CLI smoke test's `/home` path was rejected before any write).
|
||||
|
||||
## Changed files
|
||||
|
||||
- `packages/node/meshnet_node/split_gguf/__init__.py` (new)
|
||||
- `packages/node/meshnet_node/split_gguf/manifest.py` (new)
|
||||
- `packages/node/meshnet_node/split_gguf/provision.py` (new)
|
||||
- `scripts/provision_split_gguf.py` (new)
|
||||
- `tests/test_split_gguf_manifest.py` (new)
|
||||
- `tests/test_split_gguf_provision.py` (new)
|
||||
- `.scratch/distributed-gguf-runtime/prd.json` (`DGR-026.passes = true` +
|
||||
`completionNotes`; also restored the top-level `sourceOfTruth`/
|
||||
`qualityGates`/`metadataSchema`/`milestones`/`supersededStories`/
|
||||
`branchName` fields — see Gotcha below)
|
||||
- `.scratch/distributed-gguf-runtime/issues/026-provision-exact-split-gguf-artifacts-outside-home.md`
|
||||
(regenerated via `scripts/ralph_prd_schema.py render`)
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-026/README.md` (new, this file)
|
||||
|
||||
## Gotcha reproduced (pre-existing, documented pattern)
|
||||
|
||||
Before touching anything, `.scratch/distributed-gguf-runtime/prd.json`'s
|
||||
top-level `sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/
|
||||
`supersededStories`/`branchName` fields were already missing in the working
|
||||
tree at session start (this is the fourth documented occurrence of the
|
||||
round-trip-drop bug noted in DGR-018/019/020/025's evidence — `userStories`
|
||||
itself was unaffected, only these top-level fields). Restored them from
|
||||
`git show HEAD:.scratch/distributed-gguf-runtime/prd.json` before making any
|
||||
DGR-026 edit; `scripts/ralph_prd_schema.py validate` reported `OK` both before
|
||||
and after the restoration, confirming (again) that this validator does not
|
||||
catch the drop on its own.
|
||||
|
||||
## Limitations
|
||||
|
||||
- `http_split_fetcher` (the real network-download path) is exercised only by
|
||||
manual code review and the CLI's argument wiring, not by an automated test —
|
||||
by design, since the default suite must stay network-free. Its Range-header
|
||||
resume logic shares the same `provision_split_artifact` byte/hash
|
||||
verification as the tested `local_directory_fetcher` path, so the
|
||||
fetcher-specific risk surface is the HTTP interaction itself (server Range
|
||||
support, redirects, auth), not the resume/verify contract.
|
||||
- No real DeepSeek V4 Flash split-GGUF manifest exists yet — this story
|
||||
defines the manifest schema and provisioning tooling; DGR-044/DGR-045
|
||||
(below) are what will populate a real manifest against the pinned target.
|
||||
- `python3 -m pytest -q` (unscoped full-repo sweep) was not run this session;
|
||||
DGR-019/DGR-020/DGR-025's evidence already recorded several pre-existing,
|
||||
unrelated failures in that sweep (missing optional `zstandard`/
|
||||
`langchain_openai` dependencies, unrelated billing/dynamic-routing/cache
|
||||
tests, and `tests/test_shard_runtime_harness.py`'s `grpc` import
|
||||
requirement). This story's own targeted suites, `test_ralph_prd_schema.py`,
|
||||
`compileall`, and `git diff --check` are all green as recorded above.
|
||||
- Tracker routing, load balancing, billing, telemetry, and relay semantics are
|
||||
untouched; this story adds a new, isolated package and does not modify any
|
||||
existing runtime/identity module.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
- **DGR-044** (DeepSeek V4 Flash target contract): when pinning the real
|
||||
target's split-GGUF artifact, express it as a
|
||||
`meshnet_node.split_gguf.manifest.SplitArtifactManifest` — `source.sha256`
|
||||
is the whole-model artifact digest DGR-003's `ArtifactIdentity.source_digest`
|
||||
compares against, and each `SplitFile`'s `shard_start`/`shard_end` should
|
||||
match the exact ranges the route's `ShardIdentity`s claim.
|
||||
- **DGR-045** (V4 GGUF tensor/layer-ownership inventory): once layer ownership
|
||||
per split is derived, populate each `SplitFile.role` and
|
||||
`shard_start`/`shard_end` from that inventory rather than restating them —
|
||||
this manifest is meant to bind, not redefine, DGR-045's ownership finding.
|
||||
- Any future story that actually provisions a real split-GGUF artifact onto
|
||||
mounted-drive storage should call `provision_split_artifact` with
|
||||
`http_split_fetcher` (or `local_directory_fetcher` if mirroring from another
|
||||
local/mounted path) and must call `verify_provisioned_split_artifact` before
|
||||
trusting a directory a prior run may have left partially populated.
|
||||
@@ -1,7 +1,7 @@
|
||||
# DGR-028 evidence — numbered llama.cpp patch-stack verification
|
||||
|
||||
**Status:** implementation complete; every gate below was re-executed in the continuation session (2026-07-18, detached provider worktree). Final independent P0/P1 controller review is pending.
|
||||
**Authority:** live Gitea #12; local PRD is a secondary projection.
|
||||
**Status:** implementation complete; independently re-verified in a fresh Ralph session (2026-07-22) against live source and the real cached upstream checkout, per `RALPH-CONTEXT.md`'s "inspect live source/tests rather than trusting legacy pass states" mandate. `prd.json`'s `DGR-028.passes` is now `true`.
|
||||
**Authority:** local `prd.json` is authoritative; live Gitea #12 is a projection.
|
||||
**Upstream pin:** `e920c523e3b8a0163fe498af5bf90df35ff51d25` (`llama.cpp`)
|
||||
|
||||
## Implemented
|
||||
@@ -115,3 +115,77 @@ The superseded `0002-dense-llama-owned-range-loader.patch` is removed.
|
||||
- This is patch-stack and model-free native fixture evidence, not real-model correctness, memory-fit, performance, or route certification.
|
||||
- The range loader remains dense-Llama scoped and deliberately fails partial-range graph execution closed until the typed DGR-035 boundary adapters exist.
|
||||
- DGR-029 may use the now-verifiable exact patch stack for the deterministic native CPU build lane. DGR-034 owns real dense-Llama range behavior and memory evidence.
|
||||
|
||||
## Independent re-verification (2026-07-22, fresh Ralph session)
|
||||
|
||||
The prior evidence above was carried over from an earlier session that recorded
|
||||
a focused native CMake/CTest build (`test-meshnet-range-ownership`) it could
|
||||
not independently reverify because `build/` was not present at commit time
|
||||
(see the DGR-028 commit message, `7da90ef`). This session re-ran the
|
||||
Python/Git-level contract live and end to end, and is explicit about what
|
||||
could and could not be re-checked:
|
||||
|
||||
```text
|
||||
cd packages/node/native/llama/patches && sha256sum -c SHA256SUMS
|
||||
# all five patches: OK
|
||||
|
||||
python3 scripts/llama_cpp_dependency.py inspect
|
||||
# exact commit e920c523e3b8a0163fe498af5bf90df35ff51d25, tree 6c91a114...,
|
||||
# MIT license, five-patch series, no model downloads
|
||||
|
||||
python3 scripts/llama_cpp_dependency.py verify --workspace build/llama.cpp
|
||||
# reused verified offline cache; apply -> assumption/boundary checks ->
|
||||
# reverse succeeded; source left at pristine detached HEAD
|
||||
|
||||
python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
|
||||
# git -C build/llama.cpp/source diff --cached --name-only ==
|
||||
# CMakeLists.txt, cmake/meshnet-patch-stack.cmake, include/llama.h,
|
||||
# src/llama-model.cpp, src/llama-model.h, src/models/llama.cpp,
|
||||
# tests/CMakeLists.txt, tests/test-meshnet-range-ownership.cpp
|
||||
# git -C build/llama.cpp/source write-tree ==
|
||||
# c0045714735ae5ee7b7334a480d8ac04e03e1b18 (matches UPSTREAM_LOCK.json patched_tree)
|
||||
|
||||
python3 scripts/llama_cpp_dependency.py reverse --source-dir build/llama.cpp/source
|
||||
# git -C build/llama.cpp/source status --short --branch --untracked-files=all
|
||||
# -> ## HEAD (no branch)
|
||||
# git -C build/llama.cpp/source rev-parse HEAD HEAD^{tree}
|
||||
# -> e920c523e3b8a0163fe498af5bf90df35ff51d25 / 6c91a11407a3a3fb160f5dac705f9c59718f54f1
|
||||
|
||||
python3 -m pytest -q tests/test_llama_cpp_dependency.py tests/test_ralph_prd_schema.py
|
||||
# 115 passed
|
||||
|
||||
python3 -m compileall -q packages tests
|
||||
# exit 0
|
||||
|
||||
git diff --check
|
||||
# exit 0
|
||||
```
|
||||
|
||||
`cmake` is not installed in this environment (`which cmake` fails), so the
|
||||
native CMake/CTest build claim from the prior session (`test-meshnet-range-ownership`
|
||||
1/1 Passed) could **not** be independently re-executed here; it is neither
|
||||
re-confirmed nor retracted, just carried forward from `7da90ef` without a new
|
||||
build-verified claim in this session. Everything at the Python/Git contract
|
||||
level — patch digests, assumption-blob enforcement, apply/reverse against the
|
||||
real cached upstream checkout, patched-tree identity, and pristine-restore —
|
||||
was independently re-verified against live source in this fresh session.
|
||||
|
||||
## prd.json repair (unrelated to DGR-028 itself)
|
||||
|
||||
Before editing `DGR-028.passes`, `prd.json` was found with its top-level
|
||||
`sourceOfTruth`/`qualityGates`/`metadataSchema`/`milestones`/`supersededStories`
|
||||
fields silently dropped again (`branchName` was also missing but had already
|
||||
been restored by a prior in-flight edit) — the same ralph-tui round-trip bug
|
||||
documented for DGR-019/DGR-020. Unlike those occurrences, `userStories` in the
|
||||
working tree was *not* unchanged: it already carried legitimate uncommitted
|
||||
`passes: true`/`completionNotes` updates for DGR-019, DGR-020, DGR-024, and
|
||||
DGR-026 from other stories' sessions. The missing top-level sections were
|
||||
restored from `git show HEAD:.scratch/distributed-gguf-runtime/prd.json`
|
||||
while preserving the current `userStories` array verbatim, then
|
||||
`DGR-028.passes` was set `true` with `completionNotes` added, and
|
||||
`.scratch/distributed-gguf-runtime/issues/028-implement-numbered-patch-stack-apply-and-verification.md`
|
||||
was regenerated via `scripts/ralph_prd_schema.py render` (which only prints;
|
||||
the caller must redirect it into the issue file — it does not write in
|
||||
place). `python3 scripts/ralph_prd_schema.py validate` and
|
||||
`python3 -m pytest -q tests/test_ralph_prd_schema.py` (108 passed) both pass
|
||||
against the repaired file.
|
||||
|
||||
198
.scratch/distributed-gguf-runtime/evidence/DGR-029/README.md
Normal file
198
.scratch/distributed-gguf-runtime/evidence/DGR-029/README.md
Normal file
@@ -0,0 +1,198 @@
|
||||
# DGR-029 evidence — native CMake skeleton and deterministic CPU lane
|
||||
|
||||
**Status:** implementation complete, live-verified in this session (2026-07-22).
|
||||
**Authority:** local `prd.json` is authoritative; Gitea is a projection.
|
||||
**Upstream pin:** `e920c523e3b8a0163fe498af5bf90df35ff51d25` (`llama.cpp`, from `UPSTREAM_LOCK.json`).
|
||||
|
||||
## What existed before this session
|
||||
|
||||
`scripts/llama_cpp_dependency.py` already had `build()`, `smoke()`, and `reproduce()`
|
||||
functions and `UPSTREAM_LOCK.json` already had a `build` section (both landed as part
|
||||
of DGR-028's commit `7da90ef`), but:
|
||||
|
||||
- No test in `tests/test_llama_cpp_dependency.py` ever exercised `build`/`smoke`/`reproduce`
|
||||
— only `fetch`/`apply`/`reverse`/`inspect` had coverage.
|
||||
- `cmake` was not installed in the DGR-028 session's environment ("`cmake` is not installed
|
||||
in this environment," per its evidence), so this lane was never actually run end to end;
|
||||
DGR-028's own live-verified CTest evidence used a one-off manual `cmake`/`ctest` invocation
|
||||
with `-DLLAMA_BUILD_TESTS=ON` outside this driver, against a build directory that no longer
|
||||
exists in this session.
|
||||
- The locked `configure_flags` did not force CPU-only backend options (`GGML_CUDA`/`GGML_HIP`/
|
||||
`GGML_VULKAN`/`GGML_METAL`/`GGML_BLAS`) — relying on upstream per-platform defaults (which
|
||||
happen to default OFF on Linux, but are undocumented and platform-dependent), and
|
||||
`LLAMA_BUILD_TESTS` was `OFF`, so no CTest lane existed at all — only a `--help` smoke check
|
||||
against the unrelated stock `llama-gguf-hash` tool.
|
||||
|
||||
This session found and closed those three gaps rather than re-implementing from scratch.
|
||||
|
||||
## What changed in this session
|
||||
|
||||
- `packages/node/native/llama/UPSTREAM_LOCK.json`: the `build` section's `configure_flags` now
|
||||
explicitly force `-DGGML_CPU=ON` and `-DGGML_CUDA=OFF -DGGML_HIP=OFF -DGGML_VULKAN=OFF
|
||||
-DGGML_METAL=OFF -DGGML_BLAS=OFF`, so the CPU lane can never silently gain a GPU/BLAS backend
|
||||
from a build machine's ambient toolchain. `-DLLAMA_BUILD_TESTS` flipped `OFF` → `ON` (required
|
||||
so the `test-meshnet-range-ownership` CTest target exists at all — configuring
|
||||
`LLAMA_BUILD_TESTS=ON` does not itself build every upstream test, only registers them; the
|
||||
`native_targets` list still controls what actually gets compiled). Added `native_targets` entry
|
||||
`test-meshnet-range-ownership` and a new `ctest_regex` field
|
||||
(`"^test-meshnet-range-ownership$"`) naming the exact deterministic model-free fixture CTest
|
||||
added by DGR-028's patch 0005.
|
||||
- `scripts/llama_cpp_dependency.py`: factored `_cmake()`'s override/PATH/venv-sibling resolution
|
||||
into a shared `_toolchain_binary(name, env_var)` and added `_ctest()` using the same resolution
|
||||
(`CTEST` env override, PATH, or the sibling of the resolved `cmake` binary's environment). Added
|
||||
`ctest_lane(build_dir)`, which loads the lock's `ctest_regex` and runs
|
||||
`ctest --test-dir <build_dir> -R <regex> --output-on-failure`, printing output on success and
|
||||
raising `DependencyError` (via the existing `_run` wrapper, which already attaches
|
||||
stdout/stderr detail) on failure. Added a `ctest` CLI subcommand (`--build-dir`). Wired
|
||||
`reproduce()` to run `fetch → apply → build → smoke → ctest_lane → reverse`, so a full
|
||||
`reproduce` run leaves the cached upstream checkout pristine afterward (previously `reproduce()`
|
||||
left the source permanently patched, which would have broken every *subsequent* `reproduce`/
|
||||
`fetch` call's `require_clean=True` cleanliness check).
|
||||
- `tests/test_llama_cpp_dependency.py`: added
|
||||
`test_build_config_locks_an_explicit_cpu_only_deterministic_lane` (offline; asserts the lock's
|
||||
`configure_flags` are CPU-only and that `ctest_regex`/`native_targets`/`smoke_binary` all agree
|
||||
with each other and with `patched_paths`) and
|
||||
`test_ctest_lane_raises_an_actionable_error_for_a_failing_named_test` (gated on `cmake`
|
||||
availability via a `requires_cmake` marker mirroring `test_native_identity_emission.py`'s
|
||||
`requires_cc` pattern; builds a tiny synthetic two-test CMake project — not the full llama.cpp
|
||||
tree, so it runs in about a second — and proves `ctest_lane()` both passes silently on a passing
|
||||
named test and raises `DependencyError` naming the failing test on a failing one).
|
||||
|
||||
## Toolchain note
|
||||
|
||||
Neither the ambient system Python nor `.venv-rocm` has `cmake`. This session installed `cmake`
|
||||
(the PyPI wheel that bundles prebuilt binaries, version 4.4.0) into the pre-existing repo-root
|
||||
`.venv` used by earlier DGR-024/DGR-026 sessions (`.venv/bin/cmake`, `.venv/bin/ctest`), which was
|
||||
already on-disk from a prior session but had never had `cmake` installed into it. All commands
|
||||
below were run with that `.venv/bin` prepended to `PATH`. This is the same "disposable venv for a
|
||||
lightweight optional dependency" pattern DGR-024 used for `grpc`.
|
||||
|
||||
## Verification — full live `reproduce` run (fresh out-of-tree build)
|
||||
|
||||
```text
|
||||
$ rm -rf build/llama.cpp/build
|
||||
$ python3 scripts/llama_cpp_dependency.py reproduce
|
||||
reused verified offline cache: .../build/llama.cpp/source
|
||||
usage: .../build/llama.cpp/build/bin/llama-gguf-hash [options] GGUF_IN
|
||||
Hash a GGUF file
|
||||
options: ...
|
||||
Test project .../build/llama.cpp/build
|
||||
Start 27: test-meshnet-range-ownership
|
||||
1/1 Test #27: test-meshnet-range-ownership ..... Passed 0.01 sec
|
||||
100% tests passed out of 1
|
||||
$ echo $?
|
||||
0
|
||||
```
|
||||
|
||||
Wall-clock: `real 2m16.227s` (fresh CPU compile of ggml/llama-common/llama plus the
|
||||
`llama-gguf-hash` example and the `test-meshnet-range-ownership` fixture; no full `llama.cpp`
|
||||
test suite or example set is built — only the two targets named in `native_targets`).
|
||||
|
||||
Post-run checks:
|
||||
|
||||
```text
|
||||
$ ls build/llama.cpp/build/bin/*.so*
|
||||
libggml-base.so libggml-base.so.0 libggml-base.so.0.16.0
|
||||
libggml-cpu.so libggml-cpu.so.0 libggml-cpu.so.0.16.0
|
||||
libggml.so libggml.so.0 libggml.so.0.16.0
|
||||
libllama-common.so ... libllama.so ...
|
||||
# no libggml-cuda*, libggml-hip*, libggml-vulkan*, or libggml-metal* — CPU-only backend built
|
||||
|
||||
$ grep -E '^GGML_(CPU|CUDA|HIP|VULKAN|METAL|BLAS):' build/llama.cpp/build/CMakeCache.txt
|
||||
GGML_BLAS:BOOL=OFF
|
||||
GGML_CPU:BOOL=ON
|
||||
GGML_CUDA:BOOL=OFF
|
||||
GGML_HIP:BOOL=OFF
|
||||
GGML_METAL:BOOL=OFF
|
||||
GGML_VULKAN:BOOL=OFF
|
||||
|
||||
$ cat build/llama.cpp/build/meshnet-build-metadata.json
|
||||
{
|
||||
"model_downloads": false,
|
||||
"semantic_certification": false,
|
||||
...
|
||||
}
|
||||
|
||||
$ git -C build/llama.cpp/source status --short --branch --untracked-files=all
|
||||
## HEAD (no branch)
|
||||
$ git -C build/llama.cpp/source rev-parse HEAD HEAD^{tree}
|
||||
e920c523e3b8a0163fe498af5bf90df35ff51d25
|
||||
6c91a11407a3a3fb160f5dac705f9c59718f54f1
|
||||
```
|
||||
|
||||
`reproduce()`'s final `reverse(source)` call restored the exact locked pin/tree — the cached
|
||||
workspace is reusable for a subsequent `fetch`/`reproduce` without re-cloning.
|
||||
|
||||
## Verification — actionable toolchain failure (missing `cmake`)
|
||||
|
||||
```text
|
||||
$ python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
|
||||
$ env -i HOME="$HOME" PATH=/usr/bin:/bin python3 scripts/llama_cpp_dependency.py build \
|
||||
--source-dir build/llama.cpp/source --build-dir /tmp/no-cmake-build
|
||||
DGR-027 dependency error: cmake is unavailable; set CMAKE or activate the project toolchain
|
||||
$ echo $?
|
||||
2
|
||||
$ python3 scripts/llama_cpp_dependency.py reverse --source-dir build/llama.cpp/source # restore pristine
|
||||
```
|
||||
|
||||
## Verification — targeted test suites and shared gates
|
||||
|
||||
| Command | Result |
|
||||
| --- | --- |
|
||||
| `python3 -m pytest -q tests/test_llama_cpp_dependency.py` | `9 passed in 1.34s` (7 pre-existing + 2 new; the new gated CTest-wiring test ran for real, not skipped, since `cmake` is present in `.venv`) |
|
||||
| `python3 -m pytest -q tests/test_llama_cpp_dependency.py tests/test_ralph_prd_schema.py` | `117 passed` |
|
||||
| `python3 -m compileall -q packages tests` | exit 0 |
|
||||
| `git diff --check` | exit 0 (no output) |
|
||||
|
||||
## Ensuring build success does not advertise capability
|
||||
|
||||
- The locked `configure_flags` disable every accelerator backend explicitly
|
||||
(`GGML_CUDA/HIP/VULKAN/METAL/BLAS=OFF`) rather than relying on per-platform defaults, so a
|
||||
successful configure/build can only ever mean "the CPU reference backend compiled" — never an
|
||||
accelerator claim, and never dependent on whether the build host happens to have a GPU SDK
|
||||
installed.
|
||||
- `meshnet-build-metadata.json` (written by `build()`) records `model_downloads: false` and
|
||||
`semantic_certification: false` alongside the exact commit/patch/flag identities — the artifact
|
||||
itself, not just prose, states this build proves toolchain compilation only.
|
||||
- The two targets actually compiled are `llama-gguf-hash` (a stock upstream file-hashing utility;
|
||||
no inference) and `test-meshnet-range-ownership` (a model-free fixture that writes a tiny
|
||||
synthetic GGUF and asserts range-ownership bookkeeping — no real model, no generation, no
|
||||
numerical/backend correctness claim). Neither exercises inference, MoE, attention, or any
|
||||
DeepSeek V4 semantic path.
|
||||
|
||||
## Changed files
|
||||
|
||||
- `packages/node/native/llama/UPSTREAM_LOCK.json`
|
||||
- `scripts/llama_cpp_dependency.py`
|
||||
- `tests/test_llama_cpp_dependency.py`
|
||||
- `.scratch/distributed-gguf-runtime/prd.json`
|
||||
- `.scratch/distributed-gguf-runtime/issues/029-create-the-native-cmake-skeleton-and-deterministic-cpu-lane.md`
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-029/README.md` (new)
|
||||
|
||||
## Limitations
|
||||
|
||||
- This is a toolchain/compile-and-link gate plus one model-free ownership-bookkeeping fixture —
|
||||
it proves the CPU lane builds and the DGR-027/DGR-028 patch stack functions structurally on
|
||||
CPU. It proves nothing about real-model correctness, memory-fit, performance, or any
|
||||
backend/model/recipe certification; `stock_glm_limitations` in `UPSTREAM_LOCK.json` and DGR-028's
|
||||
own limitations continue to apply unchanged.
|
||||
- `cmake`/`ctest` are not installed system-wide or in `.venv-rocm` in this environment; they were
|
||||
installed only into the pre-existing repo-root `.venv` for this session's verification (and for
|
||||
the new gated pytest test, which is skipped in any environment lacking `cmake`). A future session
|
||||
without that `.venv` (or without re-installing `cmake` into it) will see the same "cmake is
|
||||
unavailable" actionable failure demonstrated above, not a silent pass.
|
||||
- Only the two named targets are compiled (`llama-gguf-hash`, `test-meshnet-range-ownership`); a
|
||||
broad `cmake --build ... --target test` / full upstream test suite is out of scope here, exactly
|
||||
as DGR-028 recorded ("not presented as a full-suite gate").
|
||||
- CUDA/ROCm/Vulkan/Metal compile lanes remain unimplemented; this story only establishes the CPU
|
||||
lane "before accelerator matrix work," per its objective. Those lanes are separate future work.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
DGR-030 and DGR-034 (this story's declared blockers) may rely on: an out-of-tree, CPU-only,
|
||||
explicit-backend-flag native build (`scripts/llama_cpp_dependency.py build`/`reproduce`) that
|
||||
compiles the exact DGR-027/DGR-028 patched pin and runs a real CTest lane
|
||||
(`test-meshnet-range-ownership`) proving the patch stack's range-ownership bookkeeping compiles
|
||||
and passes on CPU. Any accelerator (CUDA/ROCm/Vulkan/Metal) lane, any real-model load, and any
|
||||
backend/model/recipe capability certification remain unimplemented and must not be assumed from
|
||||
this story's green build alone.
|
||||
Reference in New Issue
Block a user