DGR-030 evidence — accelerator build presets and native CI/build matrix
Status: implementation complete, live-verified in this session (2026-07-23).
Authority: local prd.json is authoritative; Gitea is a projection.
Upstream pin: e920c523e3b8a0163fe498af5bf90df35ff51d25 (llama.cpp, unchanged from DGR-027..029).
What existed before this session
DGR-029 locked exactly one build lane — the deterministic CPU-only lane — in
UPSTREAM_LOCK.json's build section, plus scripts/llama_cpp_dependency.py's
build()/smoke()/ctest_lane()/reproduce(). There was no accelerator
preset, no SDK-availability probing, and no matrix runner: only the one CPU
lane existed, and there was no mechanism that could ever advertise a GPU
backend as compiled or capable.
What changed in this session
packages/node/native/llama/UPSTREAM_LOCK.json: added a new top-levelaccelerator_presetsobject with one entry each forcuda(GGML_CUDA),rocm(GGML_HIP),vulkan(GGML_VULKAN), andmetal(GGML_METAL). Each entry names only the one backend flag it flips and ansdk_probe(a binary to resolve onPATH, an optional env-var override, and — for Metal — aplatform_only: "darwin"gate). The existingbuildsection — the deterministic CPU default DGR-029 locked — is untouched.scripts/llama_cpp_dependency.py:_load_lock()now calls a new_verify_accelerator_presets(), which fail-closed-rejects any preset whose named backend flag is notOFFin the CPU default'sconfigure_flags— structurally guaranteeing a preset can only ever add one backend on top of the untouched CPU baseline, never redefine it.accelerator_configure_flags(lock, name)returns a new flag list — the CPU default's ownconfigure_flagslist is never mutated — with exactly the named preset's backend flag flippedONand every other flag (includingGGML_CPU=ON, the fallback ops backend GPU builds still need) left exactly as the CPU default declares it._sdk_probe(probe)/accelerator_status(name, lock)resolve a lane's SDK without ever raising: an absent SDK is returned as{"available": false, "reason": "<binary> is unavailable on PATH"}(or a platform-mismatch reason for Metal), so "unavailable" is data a caller reports, never an exception a caller has to remember to catch.accelerator_build(source, name, build_dir)compiles one lane into its own out-of-treebuild_dir(an isolated directory, never DGR-029's CPUbuild_dir), using the same patched-source verification andnative_targetsas the CPU lane, then writes ameshnet-build-metadata.jsonrecording the exactcommit/commit_tree, per-patch SHA-256 digests, the lane's overriddenconfigure_flags, the resolvedcmake/cxx/SDK-binary versions/paths, and explicitmodel_downloads: false,hardware_execution: false,hardware_certified: false,semantic_certification: falsefields plus anotestating the lane is registered-dark until a real-hardware certification record exists. It never callssmoke()/ctest_lane()— running a binary linked against a real accelerator backend would touch real hardware, which this story deliberately keeps out of scope.- Added
accelerator-status --name <lane>andaccelerator-build --name <lane> --source-dir --build-dirCLI subcommands, mirroring the existingctest/buildsubcommand pattern.
scripts/native_accelerator_matrix.py(new): the native CI/build matrix.run_matrix(workspace)fetches and applies the locked pin/patch stack once, runs the unchanged CPU lane (build → smoke → ctest, exactly DGR-029's contract), then for eachaccelerator_presetsentry either reports{"status": "skipped", "reason": ...}(SDK absent) or compiles it viaaccelerator_buildand reports{"status": "built", ...}— never silently treating a skip as a pass. AnyDependencyErrorfrom a lane (CPU or accelerator) is caught per-lane and reported as{"status": "failed", ...}without aborting the remaining lanes or skipping cleanup.reverse()always runs in afinally, restoring the exact pristine pin/tree regardless of lane outcomes. The CLI prints a JSON report and exits non-zero only if any lane actuallyfailed(askippedlane never fails the run).tests/test_llama_cpp_dependency.py: added 7 new tests —test_accelerator_presets_isolate_one_backend_without_touching_the_cpu_default(every preset flips exactly its own flag and the CPU default list is never mutated),test_accelerator_configure_flags_rejects_an_unknown_lane,test_accelerator_status_reports_unavailable_sdks_without_raising(asserts the exact reason string for cuda/rocm/vulkan/metal absence),test_accelerator_status_honors_an_explicit_sdk_override,test_accelerator_status_rejects_an_unknown_lane,test_accelerator_build_refuses_to_compile_an_unavailable_lane(asserts no build directory is created), and arequires_cmake-gatedtest_accelerator_build_compiles_the_available_lane_with_isolated_evidence, which builds a tiny synthetic CMake project (not the full llama.cpp tree) to proveaccelerator_build's "SDK present" path really configures with the overridden flag, compiles, and writes the registered-dark metadata — in about a second, without a real GPU SDK.tests/test_native_accelerator_matrix.py(new): 3 offline tests exercisingrun_matrix's orchestration withllama_cpp_dependency's fetch/apply/reverse/build/smoke/ctest_lane/accelerator_status/ accelerator_build stubbed out — proving unavailable SDKs are reportedskipped(never a false pass), an available accelerator lane is compiled without ever callingsmoke/ctest_lane, and a lane failure is reported per-lane without aborting sibling lanes or skipping thereverse()cleanup.
Toolchain note
As in DGR-029, neither the ambient system Python nor .venv-rocm has cmake;
this session's .venv also had no cmake (a prior session's install did not
persist). This session ran .venv/bin/python3 -m ensurepip --upgrade (no
pip was present in .venv either) and then
.venv/bin/python3 -m pip install cmake, landing the same PyPI wheel
(cmake==4.4.0) DGR-029 used, at .venv/bin/cmake / .venv/bin/ctest. All
commands below were run with that .venv/bin prepended to PATH. No CUDA,
ROCm, or Vulkan SDK (nvcc, hipcc, glslc) is installed in this
environment, and the host platform is Linux, not darwin — so all four
accelerator lanes are genuinely skipped in this environment's own live run
below, which is real evidence for AC2 ("unavailable SDKs ... explicit
unavailable/skipped lanes"), not a simulated one.
Verification — live native CI/build matrix run
$ rm -rf build/llama.cpp/build build/llama.cpp/build-cuda build/llama.cpp/build-rocm build/llama.cpp/build-vulkan build/llama.cpp/build-metal
$ python3 scripts/native_accelerator_matrix.py
reused verified offline cache: .../build/llama.cpp/source
usage: .../build/llama.cpp/build/bin/llama-gguf-hash [options] GGUF_IN
...
Test project .../build/llama.cpp/build
Start 27: test-meshnet-range-ownership
1/1 Test #27: test-meshnet-range-ownership ..... Passed 0.01 sec
100% tests passed out of 1
{
"failed_lanes": [],
"hardware_certified": false,
"lanes": [
{
"build_dir": ".../build/llama.cpp/build",
"lane": "cpu",
"metadata": {
"cmake": "cmake version 4.4.0",
"commit": "e920c523e3b8a0163fe498af5bf90df35ff51d25",
"commit_tree": "6c91a11407a3a3fb160f5dac705f9c59718f54f1",
"configure_flags": [
"-DCMAKE_BUILD_TYPE=Release", "-DLLAMA_BUILD_TESTS=ON",
"-DLLAMA_BUILD_EXAMPLES=ON", "-DLLAMA_BUILD_SERVER=OFF",
"-DLLAMA_BUILD_TOOLS=OFF", "-DLLAMA_BUILD_APP=OFF", "-DLLAMA_CURL=OFF",
"-DGGML_CPU=ON", "-DGGML_BLAS=OFF", "-DGGML_CUDA=OFF",
"-DGGML_HIP=OFF", "-DGGML_VULKAN=OFF", "-DGGML_METAL=OFF"
],
"cxx": "c++ (GCC) 15.2.1 20260123 (Red Hat 15.2.1-7)",
"model_downloads": false,
"patches": { "...": "... (5 entries, unchanged sha256 digests from DGR-029)" },
"semantic_certification": false
},
"status": "built"
},
{"lane": "cuda", "reason": "nvcc is unavailable on PATH", "status": "skipped"},
{"lane": "rocm", "reason": "hipcc is unavailable on PATH", "status": "skipped"},
{"lane": "vulkan", "reason": "glslc is unavailable on PATH", "status": "skipped"},
{"lane": "metal", "reason": "platform 'linux' is not 'darwin'", "status": "skipped"}
],
"note": "A `built` lane means it compiled with the exact recorded compiler/SDK/upstream-pin/patch-stack/build-option evidence — it never means an accelerator device was exercised. Every backend/model/recipe lane stays registered-dark until a separate real-hardware certification record exists."
}
$ echo $?
0
Wall-clock: real 2m19.797s — matches DGR-029's ~2m16s CPU-lane compile; no
accelerator lane actually compiled in this environment (all four SDKs are
genuinely absent), so this run's added cost over DGR-029's own CPU-only
reproduce() is just the four fast SDK probes.
Post-run checks (source checkout left pristine by the matrix's reverse()):
$ git -C build/llama.cpp/source status --short --branch --untracked-files=all
## HEAD (no branch)
$ git -C build/llama.cpp/source rev-parse HEAD HEAD^{tree}
e920c523e3b8a0163fe498af5bf90df35ff51d25
6c91a11407a3a3fb160f5dac705f9c59718f54f1
$ ls build/llama.cpp/ | grep build
build
Only the CPU lane's build/ directory was created — no build-cuda,
build-rocm, build-vulkan, or build-metal directory exists, because every
accelerator lane was genuinely skipped rather than attempted.
Verification — targeted test suites and shared gates
| Command | Result |
|---|---|
python3 -m pytest -q tests/test_llama_cpp_dependency.py tests/test_native_accelerator_matrix.py |
19 passed (9 pre-existing + 7 new accelerator-lane tests in test_llama_cpp_dependency.py, 3 new in test_native_accelerator_matrix.py; the requires_cmake-gated compile test ran for real, not skipped) |
python3 -m compileall -q packages tests |
exit 0 |
git diff --check -- packages/node/native/llama/UPSTREAM_LOCK.json scripts/llama_cpp_dependency.py tests/test_llama_cpp_dependency.py scripts/native_accelerator_matrix.py tests/test_native_accelerator_matrix.py |
exit 0 |
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json |
OK: 55 stories validated. |
git diff --check against the full working tree separately reports one
pre-existing trailing-whitespace line in .ralph-tui-run.log, which was
already modified before this session started (see the session's initial
git status) and is unrelated to this story's scope; it is excluded above by
naming this story's own changed files explicitly.
python3 -m pytest -q tests/test_ralph_prd_schema.py reports 55 failed, 53 passed in this session (all test_render_issue_markdown_matches_committed_file
drift between prd.json and committed issue Markdown for other stories,
e.g. DGR-053..DGR-071). git stash-ing this session's changes and rerunning
reproduces 56 failed, 52 passed identically — the same 56 failures minus the
one this session's own DGR-030 regeneration fixed, confirming the remaining
55 predate this story and are out of scope to fix here. This session did
regenerate .scratch/distributed-gguf-runtime/issues/030-add-accelerator- build-presets-and-native-ci-matrix.md via
python3 scripts/ralph_prd_schema.py render ... DGR-030 so DGR-030's own
generated issue Markdown matches prd.json byte-for-byte (confirmed by the
test_render_issue_markdown_matches_committed_file[DGR-030] case no longer
appearing in the failure list).
Ensuring build success does not advertise capability
- Every accelerator lane's
meshnet-build-metadata.jsonexplicitly recordshardware_execution: false,hardware_certified: false, andsemantic_certification: false, plus anotestating the lane is registered-dark until a separate real-hardware certification record exists — the same "artifact states this, not just prose" pattern DGR-029 used for the CPU lane'smodel_downloads/semantic_certificationfields. accelerator_buildnever runssmoke()orctest_lane(): it only configures and compiles the exactnative_targetsDGR-029 already locked (llama-gguf-hash,test-meshnet-range-ownership) — no binary linked against a real accelerator backend is ever executed by this story's code._verify_accelerator_presets()structurally refuses any preset whose backend flag is notOFFin the locked CPU default, so a preset can never be defined in a way that redefines (rather than adds one backend on top of) DGR-029's deterministic CPU lane.- The matrix's top-level report always carries
"hardware_certified": falseregardless of how many lanes built, and itsnotefield states this explicitly for any consumer reading only the report, not the per-lane metadata.
Limitations
- This story proves accelerator lanes compile with correct, isolated flags and preserves exact evidence when a lane's SDK is present. It proves nothing about numerical correctness, performance, or any backend/model/ recipe capability on real accelerator hardware — that is explicitly deferred to DGR-041 (capability registration), DGR-053 (real 2-4 stage certification), and DGR-067 (capability matrix certification), all of which remain unimplemented.
- No CUDA, ROCm, or Vulkan SDK, and no macOS/Metal toolchain, is available in
this session's environment, so the "compile an available accelerator lane"
path is proven end-to-end only via the
requires_cmake-gated synthetic- project unit test and the offline matrix-orchestration tests, not via a live compile of the real llama.cpp tree underGGML_CUDA=ON(etc.). A future session with a real SDK installed will exerciseaccelerator_build's real-lane path against the genuine llama.cpp source for the first time; nothing in this story's design assumes that hasn't happened yet. - The accelerator lanes reuse the CPU lane's exact
native_targets(llama-gguf-hash,test-meshnet-range-ownership), so a passing accelerator compile also proves the DGR-027/DGR-028 patch stack's range-ownership code compiles under that backend flag combination — but, per the point above, only structurally; it says nothing about GPU execution correctness. cmake/ctestremain absent system-wide in this environment; this session reinstalled them into.venvexactly as DGR-029 did, and that install does not appear to persist across sessions (this session found.venvwithoutcmakedespite DGR-029's evidence recording its earlier install). A future session without acmake-equipped.venvwill see the same actionable "cmake is unavailable" failure DGR-029 demonstrated, not a silent pass, and the newrequires_cmake-gated tests will be skipped rather than failing.git diff --checkandtests/test_ralph_prd_schema.pyboth carry pre-existing, out-of-scope failures unrelated to this story (see the gates table above); this story's own changed files pass both checks cleanly.
Dependency handoff
DGR-053 (real 2-4 stage certification), DGR-067 (capability matrix
certification), and DGR-068 (packaged releases) may rely on: four isolated,
out-of-tree accelerator build presets (cuda/rocm/vulkan/metal) in
UPSTREAM_LOCK.json's accelerator_presets, each toggling exactly one
backend flag on top of DGR-029's unchanged CPU default; a native CI/build
matrix (scripts/native_accelerator_matrix.py) that compiles every
SDK-available lane with full compiler/SDK/upstream-pin/patch-stack/build-
option evidence and reports SDK-unavailable lanes as explicit skipped
lanes, never a false pass; and a compile-only contract (no lane here ever
runs a binary against real accelerator hardware). Real-hardware execution,
numerical correctness, performance measurement, and backend/model/recipe
certification for any accelerator remain entirely unimplemented and must not
be assumed from any lane's green compile.