DGR-020: Run the controlled whole-model GGUF baseline #4

Closed
opened 2026-07-16 20:36:01 +00:00 by popov · 0 comments
Owner

DGR-020: Run the controlled whole-model GGUF baseline

  • Status / triage: completed; passes: true
  • Execution mode: HITL
  • Milestone: M0
  • Dependencies: DGR-019
  • Blocks (derived): DGR-054
  • Labels: area:performance, type:benchmark, priority:p0, gate:hitl, ready-for-human
  • Evidence class: real-hardware
  • Hardware: required
  • Model: generic
  • Upstream: no

Objective / description

Fresh Ralph session: read .scratch/distributed-gguf-runtime/RALPH-CONTEXT.md, source issue .scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md, and evidence READMEs for dependencies (DGR-019) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Execute the locked safetensors and whole-model llama.cpp lanes before distributed implementation results can influence the decision.

Acceptance criteria

  • Run the exact DGR-019 safetensors and whole-model llama.cpp benchmark lanes with locked prompts, lengths, sampling, concurrency, hardware, and artifact/runtime identities.
  • Record raw machine-readable correctness, TTFT, prefill/decode, throughput, latency, memory, artifact-size, failure, and quality-drift metrics without ingesting distributed implementation results.
  • Separate quantization/model-fit effects from runtime/kernel effects and preserve failed or unavailable lanes honestly.
  • Publish a threshold-based go, optimize baseline, or stop decision without changing the locked contract.
  • Applicable shared quality gates in prd.json pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.

Shared quality gates

  • Targeted deterministic tests pass; Python changes also pass python -m compileall packages tests.
  • git diff --check passes.
  • Default tests are model-download-free, API-credit-free, and GPU-free.
  • Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.

Evidence handoff

Verified evidence: .scratch/distributed-gguf-runtime/evidence/DGR-020/README.md. Legacy evidence remains provenance only and grants no implementation completion credit.

<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. --> # DGR-020: Run the controlled whole-model GGUF baseline - **Status / triage:** completed; `passes: true` - **Execution mode:** `HITL` - **Milestone:** `M0` - **Dependencies:** `DGR-019` - **Blocks (derived):** `DGR-054` - **Labels:** `area:performance`, `type:benchmark`, `priority:p0`, `gate:hitl`, `ready-for-human` - **Evidence class:** `real-hardware` - **Hardware:** `required` - **Model:** `generic` - **Upstream:** `no` ## Objective / description Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/020-run-the-controlled-whole-model-gguf-baseline.md`, and evidence READMEs for dependencies (DGR-019) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Execute the locked safetensors and whole-model llama.cpp lanes before distributed implementation results can influence the decision. ## Acceptance criteria - [x] Run the exact DGR-019 safetensors and whole-model llama.cpp benchmark lanes with locked prompts, lengths, sampling, concurrency, hardware, and artifact/runtime identities. - [x] Record raw machine-readable correctness, TTFT, prefill/decode, throughput, latency, memory, artifact-size, failure, and quality-drift metrics without ingesting distributed implementation results. - [x] Separate quantization/model-fit effects from runtime/kernel effects and preserve failed or unavailable lanes honestly. - [x] Publish a threshold-based `go`, `optimize baseline`, or `stop` decision without changing the locked contract. - [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff. ## Shared quality gates - Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`. - `git diff --check` passes. - Default tests are model-download-free, API-credit-free, and GPU-free. - Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit. ## Evidence handoff Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-020/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit. <!-- ralph-story-id: DGR-020 -->
popov added this to the M0 milestone 2026-07-16 20:36:01 +00:00
popov closed this issue 2026-07-22 06:44:22 +00:00
popov added status:completed and removed status:blocked labels 2026-07-22 06:44:23 +00:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: popov/neuron-tai#4