DGR-019 Lock alpha/beta performance contracts (evidence + contract framework) DGR-020 Run controlled whole-model GGUF baseline (benchmark results & contracts) DGR-024 Real generated-gRPC protocol harness (shard_runtime_server.py + tests) DGR-026 split-GGUF provisioning outside /home (provision script + manifest + tests) DGR-028 Numbered patch-stack apply & verify (llama_cpp_dependency.py + UPSTREAM_LOCK.json) DGR-029 Native CMake skeleton + deterministic CPU lane (UPSTREAM_LOCK.json + cmake gating) New modules: packages/node/meshnet_node/dgr_performance/ — performance contract framework packages/node/meshnet_node/split_gguf/ — split-GGUF manifest & provisioning scripts/provision_split_gguf.py — artifact provisioning CLI tests/test_dgr_performance_contract.py — contract validation tests tests/test_split_gguf_manifest.py — manifest tests tests/test_split_gguf_provision.py — provisioning tests tests/test_shard_runtime_harness.py — gRPC harness tests
3.0 KiB
3.0 KiB
DGR-020: Run the controlled whole-model GGUF baseline
- Status / triage: completed;
passes: true - Execution mode:
HITL - Milestone:
M0 - Dependencies:
DGR-019 - Blocks (derived):
DGR-054 - Labels:
area:performance,type:benchmark,priority:p0,gate:hitl,ready-for-human - Evidence class:
real-hardware - Hardware:
required - Model:
generic - Upstream:
no
Objective / description
Fresh Ralph session: read .scratch/distributed-gguf-runtime/RALPH-CONTEXT.md, source issue .scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md, and evidence READMEs for dependencies (DGR-019) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Execute the locked safetensors and whole-model llama.cpp lanes before distributed implementation results can influence the decision.
Acceptance criteria
- Run the exact DGR-019 safetensors and whole-model llama.cpp benchmark lanes with locked prompts, lengths, sampling, concurrency, hardware, and artifact/runtime identities.
- Record raw machine-readable correctness, TTFT, prefill/decode, throughput, latency, memory, artifact-size, failure, and quality-drift metrics without ingesting distributed implementation results.
- Separate quantization/model-fit effects from runtime/kernel effects and preserve failed or unavailable lanes honestly.
- Publish a threshold-based
go,optimize baseline, orstopdecision without changing the locked contract. - Applicable shared quality gates in
prd.jsonpass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
Shared quality gates
- Targeted deterministic tests pass; Python changes also pass
python -m compileall packages tests. git diff --checkpasses.- Default tests are model-download-free, API-credit-free, and GPU-free.
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never
/home. - Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
Evidence handoff
Verified evidence: .scratch/distributed-gguf-runtime/evidence/DGR-020/README.md. Legacy evidence remains provenance only and grants no implementation completion credit.