Real ShardRuntimeServicer process bound to a real localhost socket, driven by a generated ShardRuntimeStub over grpc.insecure_channel from a separately spawned subprocess. Proves direct-hop and opaque-relay (exact captured request bytes re-sent, no reinterpretation) produce byte-identical server responses, cross-checked against an independent server-side wire capture. Fails closed on the required negative paths: stale route epoch, expired deadline, malformed/non-tiling fragments, checksum failure, exhausted flow-control credit (with in-band top-up), duplicate idempotency steps (acked, not re-applied), and cancel — both in-band CancelSignal (single work item vs whole session) and the out-of-band unary Cancel RPC, including a Cancel that races ahead of SessionOpen. Supersedes the earlier in-memory fake-seam approach for this ticket, which a policy audit rejected under the no-fake-data rule; that code is not reintroduced. Evidence README rewritten to describe the actual files. 11 passed in tests/test_shard_runtime_harness.py.
Distributed GGUF Runtime planning workspace
Specification status: planning artifacts only. No distributed GGUF runtime is implemented. DGR-017 cleanup is complete; no runtime implementation story has completion credit.
prd.jsonis authoritative.
Locked scope
- Existing Meshnet Tracker routing, load balancing, billing, telemetry, relay, and provider semantics are backend-agnostic and are not redesigned. GGUF contributes exact compatibility, range/capacity, queue/load, seam-cost, health/reliability, and certification inputs only.
- The data plane is a standalone project-owned C++ Shard worker with gRPC/Protobuf and a project-owned
ShardEngineboundary. - llama.cpp is fetched at one exact commit into an ignored workspace from an in-repo manifest, then a numbered minimal patch stack is applied. There is no submodule, vendored tree, or permanent-fork dependency.
- llama.cpp owns DeepSeek V4 graphs, mHC, MoE, attention, hash routing, and kernels. Meshnet adds only range-ownership hooks, typed boundary/local-state adapters, worker integration, and parity/certification.
- Quantization and placement are dynamic recipe inputs. The 2–4 and 10+ stage layouts are certification scenarios, never product constants.
- Per-shard Hot KV and V4 CSA/HCA/SWA/indexer/compressor state remain local and keyed by route session/epoch. The WAN seam carries the typed mHC 4×4096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations—not per-layer caches.
- Route changes use cache miss plus re-prefill/restart. There is no WAN KV or V4 auxiliary-cache migration.
- CPU/CUDA/ROCm/Vulkan/Metal compile lanes are planned; only exact real-hardware-certified backend/model/recipe lanes may be advertised.
- Alpha requires correctness and the pre-locked useful-speed gate. MTP is reserved and off for alpha; its ownership contract, implementation, and benchmark are required before beta.
Target identities
- DeepSeek V4 official target SHA:
60d8d70770c6776ff598c94bb586a859a38244f1. - llama.cpp V4 support lineage began at PR 24162 / merge
8c146a8366304c871efc26057cc90370ccf58dad; DGR-027 later pins one exact validated current commit. - V4 scope: 43 main layers plus MTP; mHC 4×4096 boundary; 256 routed + 1 shared experts with six routed active; token IDs required for the first three hash-routed layers.
- Exact split-GGUF artifacts are provisioned to mounted-drive storage with a complete hashed manifest and resumable verification; no model artifact may be placed under
/home.
Navigation
prd.json— sole authoritative 55-story backlog, DGR-017..071.PRD.md— human-readable projection of goals, gates, and all stories.RALPH-CONTEXT.md— mandatory fresh-session context.architecture.md,implementation-strategy.md,milestones.md— design and execution sequence.issues/— generated story specs; files 01..16 are retained legacy artifacts pending DGR-017.evidence/— provenance and future per-story handoffs.