Compare commits
10 Commits
25e53bfeab
...
79c9bbaf63
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
79c9bbaf63 | ||
|
|
d339cfde25 | ||
|
|
27a0d89678 | ||
|
|
8c87fae1ac | ||
|
|
4d530d702c | ||
|
|
7473bb7e44 | ||
|
|
c073826374 | ||
|
|
0c7d475335 | ||
|
|
84d75f4cd2 | ||
|
|
766e480ba5 |
1
.gitignore
vendored
1
.gitignore
vendored
@@ -12,6 +12,7 @@ dist/
|
||||
# Ralph local runtime state
|
||||
.ralph-tui/*
|
||||
!.ralph-tui/config.toml
|
||||
.ralph-lane/
|
||||
|
||||
|
||||
.env
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Distributed GGUF Runtime planning workspace
|
||||
|
||||
> **Specification status:** planning artifacts only. No distributed GGUF runtime is implemented. DGR-017 cleanup is complete; no runtime implementation story has completion credit. `prd.json` is authoritative.
|
||||
> **Implementation status:** DGR-017 through DGR-033 have verified lane evidence, including a fixture-only standalone C++ gRPC worker. These lane checkpoints still require serialized integration and remote publication; they do not claim real model inference. `prd.json` is authoritative.
|
||||
|
||||
|
||||
## Locked scope
|
||||
|
||||
281
.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md
Normal file
281
.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md
Normal file
@@ -0,0 +1,281 @@
|
||||
# DGR-033 evidence — standalone fake C++ gRPC Shard worker
|
||||
|
||||
**Completed:** 2026-07-25 (initial); **repaired:** 2026-07-26 after Codex
|
||||
GPT-5.5 cross-review BLOCK (see "Cross-review repair" below).
|
||||
**Branch:** `ralph/distributed-gguf-opus`
|
||||
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
|
||||
**Dependencies:** DGR-022 (lifecycle/status contract), DGR-024 (real generated
|
||||
gRPC harness + `shard_runtime_server.py` reference semantics), DGR-032
|
||||
(deterministic fake `ShardEngine` semantics).
|
||||
|
||||
## Objective
|
||||
|
||||
Prove the standalone worker process, stream, lifecycle, and supervision shape
|
||||
before any llama.cpp integration: a real C++ executable that serves the whole
|
||||
ShardRuntime lifecycle/stream contract over gRPC using a model-free fake engine,
|
||||
driven end-to-end by Python integration tests over a real socket.
|
||||
|
||||
## What was found live before changing code
|
||||
|
||||
- `packages/node/native/proto/shard_runtime.proto` (DGR-021..023): the single
|
||||
semantic contract. Its `ShardRuntime` service has exactly five RPCs —
|
||||
`GetCapability`, `Health`, `Session` (bidi stream), `Release`, `Cancel`.
|
||||
- `packages/node/meshnet_node/shard_runtime_server.py` (DGR-024): the reference
|
||||
Python servicer. It performs a *bounded real forward* (a CRC over the received
|
||||
bundle bytes) then echoes the chunk, and fails closed on stale epoch, expired
|
||||
deadline, corrupt/mis-tiled fragments, exhausted flow-control credit, duplicate
|
||||
idempotency step, and in-band/out-of-band cancellation, with per-`route_session_id`
|
||||
state kept on the servicer so an out-of-band `Cancel` can reach a live session.
|
||||
**Key finding:** despite the schema labelling the checksum `CRC32C`, this
|
||||
runtime computes it with `zlib.crc32` (standard CRC-32, *not* Castagnoli). The
|
||||
C++ worker mirrors `zlib.crc32` exactly so its checksum acceptance is
|
||||
byte-identical to the existing Python surface (the committed C++ *conformance*
|
||||
test, by contrast, uses true Castagnoli against separately-generated goldens —
|
||||
the two are unrelated code paths).
|
||||
- `packages/node/native/CMakeLists.txt` (DGR-029/030): configures against the
|
||||
ignored `build/native-toolchain` prefix (pinned Protobuf 33.1 + gRPC 1.82.1),
|
||||
always generates both message and service stubs, and registers a C++
|
||||
conformance CTest. There was **no** worker executable and **no** Python
|
||||
worker integration test before this story (confirmed by
|
||||
`ls packages/node/native/worker` → absent, and grep for `shard_worker`).
|
||||
- `packages/node/meshnet_node/fake_shard_engine.py` (DGR-032): the Python fake
|
||||
engine, deliberately *not* wired into the gRPC surface. DGR-033's worker is
|
||||
its native analogue — a separate executable, not a consumer of that module —
|
||||
so both fakes present identical behaviour to a client (deterministic,
|
||||
model-free bounded forward; per-session isolation; fail-closed lifecycle).
|
||||
|
||||
## What was added (this story's change)
|
||||
|
||||
### `packages/node/native/worker/fake_engine.h` (new)
|
||||
|
||||
`meshnet::worker::FakeShardEngine` — a header-only, model-free fixture engine.
|
||||
Its only capability is to validate a `TensorBundle` (fragments tile exactly, the
|
||||
uncompressed CRC-32 matches the declared checksum, the declared payload stays
|
||||
within the negotiated `max_chunk_bytes`) and fold the fragment bytes through a
|
||||
bounded forward. It links, loads, and dispatches to **nothing** — no llama.cpp,
|
||||
no graph execution. Carries `kEvidenceClass = "fixture"` mirroring the Python
|
||||
`FakeShardEngine.EVIDENCE_CLASS` for the later DGR-036 parity check.
|
||||
|
||||
### `packages/node/native/worker/shard_service.{h,cpp}` (new)
|
||||
|
||||
`ShardRuntimeServiceImpl : meshnet::shard::v1::ShardRuntime::Service` — a faithful
|
||||
C++ port of the DGR-024 Python servicer: the same per-`route_session_id`
|
||||
identity/credit/dedup state guarded by a mutex, the same fail-closed negative
|
||||
paths, and the same lifecycle (open → prefill/decode → flow-control top-up →
|
||||
release/cancel). Each per-request response is computed under the lock and written
|
||||
*after* releasing it, so a blocking `Write` can never deadlock the out-of-band
|
||||
`Cancel` RPC that needs the same lock. Bounded messages are enforced two ways: a
|
||||
per-tensor `RESOURCE_EXHAUSTED` app check against `max_chunk_bytes`, plus a hard
|
||||
transport receive ceiling.
|
||||
|
||||
### `packages/node/native/worker/shard_worker_main.cpp` (new)
|
||||
|
||||
The standalone `shard_worker` executable. Binds `MESHNET_SHARD_LISTEN_ADDR`
|
||||
(or an `argv` address), prints one readiness line (`ShardRuntime worker listening
|
||||
on <addr>`), and serves until `SIGTERM`/`SIGINT`. **Graceful shutdown** uses a
|
||||
self-pipe: the async-signal-safe handler writes one byte, a drain thread reads it
|
||||
and calls `server->Shutdown()`, so in-flight sessions finish and the process
|
||||
exits `0` printing `ShardRuntime worker shut down cleanly`. A `--selftest` mode
|
||||
binds an ephemeral port and self-drives capability/health/fragmented-prefill/
|
||||
decode/release over a real loopback gRPC channel, giving a pure-C++ CTest that
|
||||
needs no Python.
|
||||
|
||||
### `packages/node/native/CMakeLists.txt` (modified)
|
||||
|
||||
Adds the `shard_worker` executable (linking only `shard_runtime_grpc` +
|
||||
`gRPC::grpc++` — no llama.cpp) and registers `shard_worker_selftest` as a CTest.
|
||||
|
||||
### `tests/test_native_shard_worker.py` (new)
|
||||
|
||||
18 integration tests that spawn the **real compiled binary** as a subprocess and
|
||||
drive it with the committed generated stubs over a real localhost socket. When
|
||||
the binary is not built they skip (the DGR-029/030 `requires_cmake` gating
|
||||
pattern), locating it via `MESHNET_SHARD_WORKER_BIN` or `build/native/shard_worker`.
|
||||
|
||||
## Acceptance criteria → evidence
|
||||
|
||||
1. **Standalone C++ executable serves the complete lifecycle/stream contract
|
||||
using the fake engine** — `shard_worker` builds and serves all five RPCs; the
|
||||
`shard_worker_selftest` CTest drives open → fragmented prefill → decode →
|
||||
release over real gRPC; the 18 Python tests cover the same against the
|
||||
subprocess.
|
||||
2. **Python integration tests cover startup, health, capability, fragmented
|
||||
prefill, decode, release, cancellation, graceful shutdown** —
|
||||
`test_worker_startup_and_health`, `test_worker_capability`,
|
||||
`test_fragmented_prefill_echoes_reassembled_payload` (3-fragment tiling),
|
||||
`test_decode_step_is_served`, `test_release_is_terminal`,
|
||||
`test_in_band_cancel_of_single_work_item_does_not_end_stream`,
|
||||
`test_in_band_cancel_of_whole_session_is_terminal`,
|
||||
`test_out_of_band_cancel_rpc_races_ahead_of_open`,
|
||||
`test_graceful_shutdown_on_sigterm` (SIGTERM → exit 0 + clean-shutdown line).
|
||||
3. **Bounded messages, deadlines, flow control, independent session
|
||||
cancellation enforced** — `test_bounded_message_is_rejected`
|
||||
(`RESOURCE_EXHAUSTED` on an over-ceiling tensor),
|
||||
`test_expired_deadline_is_rejected`, `test_flow_control_violation_and_topup`,
|
||||
`test_independent_session_cancellation` (cancelling session A leaves session B
|
||||
fully serviceable), plus `test_stale_route_epoch_is_rejected`,
|
||||
`test_duplicate_idempotency_step_is_acked`,
|
||||
`test_malformed_fragment_tiling_is_rejected`.
|
||||
4. **Exposes neither llama.cpp RPC nor arbitrary graph execution** —
|
||||
`ldd build/native/shard_worker` shows no llama/ggml shared libs;
|
||||
`nm -C build/native/shard_worker | grep -icE 'llama_|ggml_'` → `0`; the proto
|
||||
exposes exactly one service with five lifecycle RPCs and no graph-exec entry.
|
||||
5. **Gates + this handoff** — below.
|
||||
|
||||
## Commands and results
|
||||
|
||||
Toolchain (ignored `build/native-toolchain`, pinned Protobuf 33.1 + gRPC 1.82.1):
|
||||
|
||||
```bash
|
||||
bash scripts/bootstrap_native_toolchain.sh "$PWD/build/native-toolchain"
|
||||
# ... gRPC 1.82.1 commit acccf84c0df20487d64101f528e5d426541ca4e5
|
||||
# grpc_cpp_plugin sha256 43705cf26ae9ce98bbcee76b3408f5e171eec746b50bf0dd42dd68d132c6a533
|
||||
```
|
||||
|
||||
Focused out-of-tree CMake build + CTest:
|
||||
|
||||
```bash
|
||||
cmake -S packages/node/native -B build/native -DCMAKE_PREFIX_PATH="$PWD/build/native-toolchain"
|
||||
cmake --build build/native -j"$(nproc)"
|
||||
ctest --test-dir build/native --output-on-failure
|
||||
```
|
||||
```text
|
||||
1/2 Test #1: shard_worker_selftest ............ Passed 0.01 sec
|
||||
2/2 Test #2: shard_protocol_conformance ....... Passed 0.00 sec
|
||||
100% tests passed out of 2
|
||||
```
|
||||
|
||||
Python integration tests against the real binary:
|
||||
|
||||
```bash
|
||||
PYTHONPATH=packages/node:packages/tracker python -m pytest -q tests/test_native_shard_worker.py
|
||||
```
|
||||
```text
|
||||
18 passed in 3.96s
|
||||
```
|
||||
|
||||
AC4 (no llama.cpp / no graph exec):
|
||||
|
||||
```bash
|
||||
ldd build/native/shard_worker | grep -iE 'llama|ggml' # -> (no matches)
|
||||
nm build/native/shard_worker | grep -icE 'llama_|ggml_' # -> 0
|
||||
```
|
||||
|
||||
Shared gates + regression:
|
||||
|
||||
```bash
|
||||
python -m compileall -q packages tests # exit 0
|
||||
git diff --check -- packages/node/native tests/test_native_shard_worker.py # exit 0
|
||||
PYTHONPATH=packages/node:packages/tracker python -m pytest -q \
|
||||
tests/test_shard_runtime_harness.py tests/test_native_shard_protocol.py
|
||||
# -> 61 passed, 2 skipped (DGR-024 harness + native protocol untouched)
|
||||
```
|
||||
|
||||
Toolchain used: `cmake`/`ctest` from the `distributed-gguf-runtime` worktree's
|
||||
`.venv` (PyPI `cmake==4.4.0` wheel — no system cmake exists here, same as
|
||||
DGR-029/030); the Python client uses that venv's `grpcio==1.82.1`,
|
||||
`grpcio-tools==1.82.1`, `protobuf`, `pytest`. `g++ (GCC) 15.2.1`.
|
||||
|
||||
## Limitations
|
||||
|
||||
- This is FIXTURE evidence only. The worker's "forward" is a CRC-over-wire-bytes
|
||||
echo, not real tensor compute; it proves process/stream/lifecycle/supervision
|
||||
shape, nothing about numerical correctness. Real engine binding is DGR-037 and
|
||||
numeric parity is DGR-036/052.
|
||||
- The worker checksum path mirrors the DGR-024 runtime's `zlib.crc32` (standard
|
||||
CRC-32 under a `CRC32C` label). Compressed-tensor tiling/checksum is not
|
||||
independently verified (no zstd decompressor in the fixture) — identical to the
|
||||
DGR-024 limitation.
|
||||
- Default `pytest` runs skip `tests/test_native_shard_worker.py` unless the
|
||||
worker binary is built (or `MESHNET_SHARD_WORKER_BIN` is set); this session
|
||||
built it and ran all 18 for real (results above). Building requires the pinned
|
||||
gRPC C++ toolchain, which is not present by default and must be bootstrapped.
|
||||
- No CUDA/ROCm/GPU, no model download, no network at test time — all default
|
||||
tests are fixture-only and offline.
|
||||
|
||||
## Dependency handoff
|
||||
|
||||
- **DGR-036** (fixture vs real-model parity): the worker's `FakeShardEngine`
|
||||
carries `kEvidenceClass = "fixture"`; diff it against DGR-037's real engine's
|
||||
equivalent marker, and reuse the same lifecycle/stream contract this worker
|
||||
serves to prove behavioural parity before numeric parity.
|
||||
- **DGR-037** (bind llama.cpp): replace `FakeShardEngine`'s bounded forward with
|
||||
the real engine behind the *same* `ShardRuntimeServiceImpl` surface; the
|
||||
service's session/epoch/credit/dedup/cancel machinery and the graceful-shutdown
|
||||
supervision shape are reusable as-is.
|
||||
- **DGR-040** (worker supervision): `shard_worker` already provides the
|
||||
supervision primitives — a readiness line for start detection, `SIGTERM`
|
||||
graceful drain with a clean-exit line, and a `--selftest` liveness probe.
|
||||
A supervisor can start/monitor/restart the process around these.
|
||||
|
||||
## Cross-review repair (2026-07-26)
|
||||
|
||||
An independent Codex GPT-5.5 review BLOCKED the initial implementation. Four
|
||||
root protocol defects in the native worker were fixed in this worktree
|
||||
(`.claude/worktrees/distributed-gguf-opus`); the fake-engine echo semantics and
|
||||
supervision shape are unchanged.
|
||||
|
||||
### Defects fixed
|
||||
|
||||
1. **Activation before SessionOpen bypassed all state.** A chunk/decode whose
|
||||
`route_session_id` had no opened session fell through every `if (state && ...)`
|
||||
guard and was echoed — bypassing lifecycle, cancellation, epoch and
|
||||
flow-control. `SessionState` now carries an `opened` flag set only by a valid
|
||||
`SessionOpen`; chunk and decode fail closed with a terminal
|
||||
`ERROR_CODE_INTERNAL` and end the stream when it is false. A placeholder state
|
||||
created by an out-of-band `Cancel` that races `Open` has `opened == false`, so
|
||||
it can never admit work either.
|
||||
2. **Flow control blindly trusted the peer proposal.** `SessionOpen` copied the
|
||||
proposed `credits/max_inflight/max_chunk_bytes` verbatim into session state and
|
||||
the accepted reply. New `ShardRuntimeServiceImpl::NegotiateFlow` takes the
|
||||
strictest bound of peer-vs-worker for every field (mirroring
|
||||
`negotiate_flow_control` in `native_protocol/codec.py`), stores the negotiated
|
||||
ceilings on the session, and enforces the negotiated per-session
|
||||
`max_chunk_bytes` on every bundle (`FakeShardEngine::Validate` now takes the
|
||||
ceiling as an argument instead of a fixed construction-time value).
|
||||
3. **In-stream `ReleaseSignal` leaked session state.** The stream `release` arm
|
||||
wrote a terminal status but never dropped the session. It now erases the
|
||||
session under the lock before responding, so KV/credits/dedup are freed
|
||||
immediately (the out-of-band `Release` RPC already erased).
|
||||
4. **`SessionOpen` echoed caller identity instead of validating it.** The handshake
|
||||
now rejects an incompatible `schema_version` (`SCHEMA_UNSUPPORTED`), a
|
||||
mismatched model/recipe `Fingerprint` (`FINGERPRINT_MISMATCH`), and a
|
||||
`ShardRange` outside the worker's served range (`SHARD_RANGE_MISMATCH`), each
|
||||
terminal; `SessionAccepted` now reports the worker's own served fingerprint
|
||||
rather than a copy of the caller's.
|
||||
|
||||
### Changed files (repair)
|
||||
|
||||
- `packages/node/native/worker/shard_service.h` — `opened` +
|
||||
`max_prefill_chunk_tokens` on `SessionState`; `NegotiateFlow` decl; engine now
|
||||
default-constructed.
|
||||
- `packages/node/native/worker/shard_service.cpp` — worker-identity constants +
|
||||
fill helpers; `NegotiateFlow`; `SessionOpen` validation/negotiation; fail-closed
|
||||
chunk/decode; per-session `max_chunk_bytes`; in-stream release erase.
|
||||
- `packages/node/native/worker/fake_engine.h` — `Validate(bundle, max_chunk_bytes)`.
|
||||
- `tests/test_native_shard_worker.py` — extended `_open` (schema/fingerprint/range/
|
||||
flow overrides); fixed `test_release_rpc_is_idempotent` for the new erase
|
||||
semantics; added 9 regression tests (chunk/decode before open, flow-control
|
||||
clamp, negotiated-ceiling cap, in-stream release erase, schema/fingerprint/range
|
||||
rejection, worker-fingerprint-not-caller).
|
||||
|
||||
### Re-run gates (real, rebuilt binary)
|
||||
|
||||
Build driven through the pinned `cmake` (Unix Makefiles + `gmake`, gRPC 1.82.1):
|
||||
|
||||
```text
|
||||
cmake --build build/native --parallel 8 -> BUILD_EXIT 0
|
||||
ctest --test-dir build/native --output-on-failure -> 100% (2/2) passed
|
||||
shard_worker_selftest ....... Passed
|
||||
shard_protocol_conformance .. Passed
|
||||
python -m pytest -q tests/test_native_shard_worker.py -> 27 passed
|
||||
python -m pytest -q tests/test_shard_runtime_harness.py \
|
||||
tests/test_native_shard_protocol.py -> 63 passed
|
||||
python -m compileall -q packages tests -> exit 0
|
||||
git diff --check -> clean
|
||||
ldd build/native/shard_worker | grep -iE 'llama|ggml' -> NONE
|
||||
nm -C build/native/shard_worker | grep -cE 'llama_|ggml_' -> 0
|
||||
```
|
||||
|
||||
The worker integration suite grew from 18 to 27 tests; all pass against the
|
||||
freshly compiled binary. No `.ralph-lane` runtime artifacts were touched.
|
||||
94
.scratch/distributed-gguf-runtime/evidence/DGR-034/README.md
Normal file
94
.scratch/distributed-gguf-runtime/evidence/DGR-034/README.md
Normal file
@@ -0,0 +1,94 @@
|
||||
# DGR-034 evidence — dense-Llama range-aware GGUF ownership
|
||||
|
||||
**Status:** implemented and live-verified on 2026-08-01. `prd.json` remains
|
||||
the authority for story state.
|
||||
|
||||
## What changed
|
||||
|
||||
- The pinned llama.cpp patch stack adds `meshnet_owned_layer_start/end` and
|
||||
filters dense-Llama GGUF registration to `blk.N.*` for the requested
|
||||
half-open range. `token_embd.weight` belongs to the head; `output_norm` and
|
||||
`output.weight` (or the tied embedding) belong to the tail.
|
||||
- The load state exposes a C range report derived from the registered model
|
||||
buffers, and a project-owned `meshnet-range-report` tool audits the live
|
||||
registered tensor map. It rejects empty, inverted, out-of-model, missing,
|
||||
outside-range, unexpected, and endpoint-inconsistent loads.
|
||||
- `meshnet_node.range_report` accepts only audited tool output. It makes the
|
||||
range and endpoint flags authoritative from loaded state rather than caller
|
||||
assertions, and fails closed on malformed ownership or byte counts.
|
||||
|
||||
## Real-model memory evidence
|
||||
|
||||
Artifact: `Magistral-Small-2509-Q4_K_M.gguf`, 14,333,911,104 bytes, SHA-256
|
||||
`a17a113480e7f55780ad1d100493c70ac158d1943e578bbdd75acef0872ab7dc`.
|
||||
It stayed on the configured mounted drive; no artifact was downloaded or put
|
||||
under `/home`.
|
||||
|
||||
The direct non-mmap lane proves resident storage tracks owned tensors:
|
||||
|
||||
| Range | Registered tensors | Resident bytes | Process peak RSS |
|
||||
| --- | ---: | ---: | ---: |
|
||||
| `[10, 20)` | 90 | 3,304,898,560 | 3,298,800 KiB |
|
||||
| `[0, 40)` | 363 | 14,326,026,240 | 14,061,632 KiB |
|
||||
|
||||
Raw reports and timings are in `runs/default-mid-a.*` and
|
||||
`runs/default-full-nommap.*`. The middle range is 23.1% of the full
|
||||
resident allocation and owns 24.8% of the registered tensors.
|
||||
|
||||
## Commands and results
|
||||
|
||||
```text
|
||||
python3 scripts/llama_cpp_dependency.py reverse --source-dir build/llama.cpp/source
|
||||
python3 scripts/llama_cpp_dependency.py verify --workspace build/llama.cpp
|
||||
python3 scripts/llama_cpp_dependency.py apply --source-dir build/llama.cpp/source
|
||||
# apply/check/reverse succeeded against e920c523e3b8a0163fe498af5bf90df35ff51d25;
|
||||
# the source was then applied for the focused native checks.
|
||||
|
||||
(cd packages/node/native/llama/patches && sha256sum -c SHA256SUMS)
|
||||
# all six patches: OK
|
||||
|
||||
/home/popov/.hermes/hermes-agent/venv/bin/ctest \
|
||||
--test-dir build/llama.cpp/dgr034-check \
|
||||
-R '^test-meshnet-range-ownership$' --output-on-failure
|
||||
# 1/1 passed
|
||||
|
||||
PYTHONPATH=packages/node MESHNET_RANGE_REPORT_BIN="$PWD/build/llama.cpp/dgr034-check/bin/meshnet-range-report" \
|
||||
/home/popov/.hermes/hermes-agent/venv/bin/pytest -q \
|
||||
tests/test_range_report.py tests/test_meshnet_range_report_tool.py \
|
||||
tests/test_llama_cpp_dependency.py
|
||||
# 56 passed in 0.87s
|
||||
|
||||
PYTHONPATH=packages/node /home/popov/.hermes/hermes-agent/venv/bin/python \
|
||||
-m compileall -q packages tests
|
||||
python3 scripts/ralph_prd_schema.py validate .scratch/distributed-gguf-runtime/prd.json
|
||||
git diff --check && git diff --cached --check
|
||||
# all exit 0; PRD validation: 55 stories validated
|
||||
```
|
||||
|
||||
The model commands used the same `meshnet-range-report` binary with
|
||||
`--no-mmap --no-extra-bufts`, first for `[10,20)` and then `[0,40)`; both
|
||||
returned `ok: true` and their exact output is retained above.
|
||||
|
||||
## Changed files
|
||||
|
||||
- `packages/node/native/llama/PATCH-STACK.md`
|
||||
- `packages/node/native/llama/UPSTREAM_LOCK.json`
|
||||
- `packages/node/native/llama/patches/{series,SHA256SUMS,UPSTREAM-ASSUMPTIONS.json,0006-meshnet-range-report-tool.patch}`
|
||||
- `packages/node/meshnet_node/range_report.py`
|
||||
- `tests/test_range_report.py`
|
||||
- `tests/test_meshnet_range_report_tool.py`
|
||||
- `.scratch/distributed-gguf-runtime/evidence/DGR-034/*`
|
||||
|
||||
## Limitations and dependency handoff
|
||||
|
||||
- The mmap loader can retain broad contiguous file spans when GGUF tensor
|
||||
order places a tail endpoint near the beginning of the artifact; the direct
|
||||
non-mmap lane is the certified resident-memory result. The raw mmap report
|
||||
is retained in `runs/default-head.json` and must not be presented as a
|
||||
physical-RSS saving.
|
||||
- This story proves loading/ownership only. Partial-range graph execution
|
||||
remains fail-closed until DGR-035 provides typed dense boundary adapters.
|
||||
- DGR-037 can bind the worker to `llama_model_meshnet_range_report` or the
|
||||
strict Python consumer; it must use the reported range, not requested range,
|
||||
for capability publication. DGR-051 must add its V4-specific ownership
|
||||
rules separately.
|
||||
@@ -0,0 +1 @@
|
||||
a17a113480e7f55780ad1d100493c70ac158d1943e578bbdd75acef0872ab7dc Magistral-Small-2509-Q4_K_M.gguf
|
||||
@@ -0,0 +1,24 @@
|
||||
{
|
||||
"ok": true,
|
||||
"model": "/run/media/popov/DATA/llm/lmstudio-community/Magistral-Small-2509-GGUF/Magistral-Small-2509-Q4_K_M.gguf",
|
||||
"architecture": "llama",
|
||||
"n_layer": 40,
|
||||
"file_bytes": 14333911104,
|
||||
"requested_range": [0, 40],
|
||||
"reported_range": [0, 40],
|
||||
"mmap": false,
|
||||
"touched": false,
|
||||
"use_extra_bufts": false,
|
||||
"has_token_embeddings": true,
|
||||
"has_output_head": true,
|
||||
"tied_output_head": false,
|
||||
"mapped_bytes": 0,
|
||||
"resident_bytes": 14326026240,
|
||||
"registered_tensors": 363,
|
||||
"registered_bytes": 14326026240,
|
||||
"unexpected_registered_tensors": [],
|
||||
"missing_owned_layers": [],
|
||||
"vm_size_bytes": 14392061952,
|
||||
"vm_rss_bytes": 14387003392,
|
||||
"vm_hwm_bytes": 14399111168
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
elapsed=0:02.48 maxrss_kib=14061632 exit=0
|
||||
@@ -0,0 +1,24 @@
|
||||
{
|
||||
"ok": true,
|
||||
"model": "/run/media/popov/DATA/llm/lmstudio-community/Magistral-Small-2509-GGUF/Magistral-Small-2509-Q4_K_M.gguf",
|
||||
"architecture": "llama",
|
||||
"n_layer": 40,
|
||||
"file_bytes": 14333911104,
|
||||
"requested_range": [0, 10],
|
||||
"reported_range": [0, 10],
|
||||
"mmap": true,
|
||||
"touched": false,
|
||||
"use_extra_bufts": true,
|
||||
"has_token_embeddings": true,
|
||||
"has_output_head": false,
|
||||
"tied_output_head": false,
|
||||
"mapped_bytes": 6219366400,
|
||||
"resident_bytes": 6219366400,
|
||||
"registered_tensors": 91,
|
||||
"registered_bytes": 3771596800,
|
||||
"unexpected_registered_tensors": [],
|
||||
"missing_owned_layers": [],
|
||||
"vm_size_bytes": 16942260224,
|
||||
"vm_rss_bytes": 16937005056,
|
||||
"vm_hwm_bytes": 16947953664
|
||||
}
|
||||
@@ -0,0 +1,24 @@
|
||||
{
|
||||
"ok": true,
|
||||
"model": "/run/media/popov/DATA/llm/lmstudio-community/Magistral-Small-2509-GGUF/Magistral-Small-2509-Q4_K_M.gguf",
|
||||
"architecture": "llama",
|
||||
"n_layer": 40,
|
||||
"file_bytes": 14333911104,
|
||||
"requested_range": [10, 20],
|
||||
"reported_range": [10, 20],
|
||||
"mmap": false,
|
||||
"touched": false,
|
||||
"use_extra_bufts": false,
|
||||
"has_token_embeddings": false,
|
||||
"has_output_head": false,
|
||||
"tied_output_head": false,
|
||||
"mapped_bytes": 0,
|
||||
"resident_bytes": 3304898560,
|
||||
"registered_tensors": 90,
|
||||
"registered_bytes": 3304898560,
|
||||
"unexpected_registered_tensors": [],
|
||||
"missing_owned_layers": [],
|
||||
"vm_size_bytes": 3370934272,
|
||||
"vm_rss_bytes": 3365814272,
|
||||
"vm_hwm_bytes": 3377971200
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
elapsed=0:00.82 maxrss_kib=3298800 exit=0
|
||||
@@ -30,6 +30,9 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
- `git diff --check` passes.
|
||||
- Default tests are model-download-free, API-credit-free, and GPU-free.
|
||||
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
|
||||
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
|
||||
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
|
||||
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
|
||||
|
||||
## Evidence handoff
|
||||
|
||||
|
||||
@@ -30,6 +30,9 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
- `git diff --check` passes.
|
||||
- Default tests are model-download-free, API-credit-free, and GPU-free.
|
||||
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
|
||||
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
|
||||
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
|
||||
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
|
||||
|
||||
## Evidence handoff
|
||||
|
||||
|
||||
@@ -30,6 +30,9 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
- `git diff --check` passes.
|
||||
- Default tests are model-download-free, API-credit-free, and GPU-free.
|
||||
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
|
||||
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
|
||||
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
|
||||
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
|
||||
|
||||
## Evidence handoff
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
|
||||
# DGR-033: Build a standalone fake C++ gRPC Shard worker
|
||||
|
||||
- **Status / triage:** specification only; `ready-for-agent`; `passes: false`
|
||||
- **Status / triage:** completed; `passes: true`
|
||||
- **Execution mode:** `AFK`
|
||||
- **Milestone:** `M1`
|
||||
- **Dependencies:** `DGR-022`, `DGR-024`, `DGR-032`
|
||||
@@ -18,11 +18,11 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] A standalone C++ executable serves the complete lifecycle and stream RPC contract using the fake engine.
|
||||
- [ ] Python integration tests cover startup, health, capability, fragmented prefill, decode, release, cancellation, and graceful shutdown.
|
||||
- [ ] Bounded messages, deadlines, flow control, and independent session cancellation are enforced.
|
||||
- [ ] The worker exposes neither llama.cpp RPC nor arbitrary graph execution.
|
||||
- [ ] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
||||
- [x] A standalone C++ executable serves the complete lifecycle and stream RPC contract using the fake engine.
|
||||
- [x] Python integration tests cover startup, health, capability, fragmented prefill, decode, release, cancellation, and graceful shutdown.
|
||||
- [x] Bounded messages, deadlines, flow control, and independent session cancellation are enforced.
|
||||
- [x] The worker exposes neither llama.cpp RPC nor arbitrary graph execution.
|
||||
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
||||
|
||||
## Shared quality gates
|
||||
|
||||
@@ -36,4 +36,4 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
|
||||
## Evidence handoff
|
||||
|
||||
Write and verify `.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md`. Until every criterion and applicable gate has real evidence, this story remains `passes: false`. Legacy evidence is provenance only, not completion credit.
|
||||
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
|
||||
# DGR-034: Implement dense-Llama range-aware GGUF ownership
|
||||
|
||||
- **Status / triage:** specification only; `ready-for-agent`; `passes: false`
|
||||
- **Status / triage:** completed; `passes: true`
|
||||
- **Execution mode:** `AFK`
|
||||
- **Milestone:** `M2`
|
||||
- **Dependencies:** `DGR-028`, `DGR-029`, `DGR-031`
|
||||
@@ -18,11 +18,11 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] Load only `blk.N.*` tensors in the assigned range, embeddings only at the head, and norm/output or tied output only at the tail.
|
||||
- [ ] Derive authoritative range and endpoint ownership from the loaded engine state.
|
||||
- [ ] Reject invalid/gapped/out-of-model ranges and unexpected required tensors.
|
||||
- [ ] Real-model evidence shows mapped/resident memory scales with owned tensors rather than full artifact size.
|
||||
- [ ] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
||||
- [x] Load only `blk.N.*` tensors in the assigned range, embeddings only at the head, and norm/output or tied output only at the tail.
|
||||
- [x] Derive authoritative range and endpoint ownership from the loaded engine state.
|
||||
- [x] Reject invalid/gapped/out-of-model ranges and unexpected required tensors.
|
||||
- [x] Real-model evidence shows mapped/resident memory scales with owned tensors rather than full artifact size.
|
||||
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
||||
|
||||
## Shared quality gates
|
||||
|
||||
@@ -36,4 +36,4 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
||||
|
||||
## Evidence handoff
|
||||
|
||||
Write and verify `.scratch/distributed-gguf-runtime/evidence/DGR-034/README.md`. Until every criterion and applicable gate has real evidence, this story remains `passes: false`. Legacy evidence is provenance only, not completion credit.
|
||||
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-034/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.
|
||||
|
||||
@@ -1,7 +1,265 @@
|
||||
{
|
||||
"name": "Distributed GGUF Runtime",
|
||||
"description": "Benchmark-gated distributed GGUF Shards using existing Meshnet control-plane routing and a standalone C++ gRPC worker around pinned upstream llama.cpp, targeting DeepSeek V4 Flash without hardcoded quantization or topology.",
|
||||
"branchName": "ralph/distributed-gguf-runtime",
|
||||
"description": "Benchmark-gated distributed GGUF Shards using existing Meshnet control-plane routing and a standalone C++ gRPC worker around pinned upstream llama.cpp, targeting DeepSeek V4 Flash without hardcoded quantization or topology.",
|
||||
"sourceOfTruth": "This prd.json is authoritative. Generated issue Markdown and planning summaries are projections and must not override it. DGR-017 through DGR-033 have verified lane evidence; DGR-034 through DGR-071 remain unimplemented specifications with passes=false. Fixture evidence does not claim real model inference.",
|
||||
"qualityGates": {
|
||||
"universal": [
|
||||
"Targeted deterministic tests pass; Python changes also pass `python -m compileall packages tests`.",
|
||||
"`git diff --check` passes.",
|
||||
"Default tests are model-download-free, API-credit-free, and GPU-free.",
|
||||
"Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit."
|
||||
],
|
||||
"native": [
|
||||
"Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin."
|
||||
],
|
||||
"realModelHardware": [
|
||||
"Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`."
|
||||
],
|
||||
"scope": [
|
||||
"Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed."
|
||||
]
|
||||
},
|
||||
"metadataSchema": {
|
||||
"requiredStoryFields": [
|
||||
"id",
|
||||
"title",
|
||||
"description",
|
||||
"acceptanceCriteria",
|
||||
"priority",
|
||||
"passes",
|
||||
"milestone",
|
||||
"executionMode",
|
||||
"labels",
|
||||
"triage",
|
||||
"evidenceClass",
|
||||
"evidencePath",
|
||||
"hardware",
|
||||
"model",
|
||||
"upstream",
|
||||
"dependsOn",
|
||||
"notes",
|
||||
"blocks"
|
||||
],
|
||||
"optionalStoryFields": [
|
||||
"completionNotes"
|
||||
],
|
||||
"idRange": "DGR-017..DGR-071 inclusive",
|
||||
"triageValues": [
|
||||
"ready-for-agent",
|
||||
"ready-for-human"
|
||||
],
|
||||
"executionModeValues": [
|
||||
"AFK",
|
||||
"HITL"
|
||||
],
|
||||
"evidenceClassValues": [
|
||||
"model-free",
|
||||
"fixture",
|
||||
"real-model",
|
||||
"real-hardware",
|
||||
"release"
|
||||
],
|
||||
"hardwareValues": [
|
||||
"none",
|
||||
"optional",
|
||||
"required"
|
||||
],
|
||||
"upstreamValues": [
|
||||
"yes",
|
||||
"no",
|
||||
"conditional"
|
||||
],
|
||||
"typeDerivation": "A story type is derived from its type:<value> label; gate:<value> stories derive release-gate.",
|
||||
"labelConventions": "Reserved prefixes include type:, priority:, area:, gate:, and ready-for-agent/ready-for-human triage labels; at most one type: and one priority: label are allowed.",
|
||||
"generatedArtifactDisclaimer": "<!-- GENERATED FROM prd.json \u2014 DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->",
|
||||
"dependencyRules": "Dependencies reference existing numerically earlier IDs; graph is acyclic. blocks is mechanically derived from dependsOn.",
|
||||
"authorityRule": "Generated issue files state that prd.json is authoritative and cannot independently claim completion or override it."
|
||||
},
|
||||
"milestones": [
|
||||
{
|
||||
"id": "M0",
|
||||
"name": "Truth and contracts",
|
||||
"stories": "DGR-017..DGR-020",
|
||||
"outcome": "Reconciled legacy truth, canonical metadata, immutable gates, and a controlled whole-model baseline."
|
||||
},
|
||||
{
|
||||
"id": "M1",
|
||||
"name": "Protocol and native substrate",
|
||||
"stories": "DGR-021..DGR-033",
|
||||
"outcome": "Versioned gRPC protocol, exact identities/artifacts, pinned upstream, reproducible builds, ShardEngine, and fake worker."
|
||||
},
|
||||
{
|
||||
"id": "M2",
|
||||
"name": "Dense vertical proof",
|
||||
"stories": "DGR-034..DGR-043",
|
||||
"outcome": "Dense ranged execution, parity, local state, worker integration, and GGUF inputs to existing routing."
|
||||
},
|
||||
{
|
||||
"id": "M3",
|
||||
"name": "DeepSeek V4 Flash alpha",
|
||||
"stories": "DGR-044..DGR-054",
|
||||
"outcome": "Pinned V4 adapter around upstream llama.cpp, real route certification, and pre-locked alpha decision with MTP off."
|
||||
},
|
||||
{
|
||||
"id": "M4",
|
||||
"name": "Performance and beta hardening",
|
||||
"stories": "DGR-055..DGR-067",
|
||||
"outcome": "Batching, backpressure, recovery, scale certification, optimization, MTP, and hardware matrix."
|
||||
},
|
||||
{
|
||||
"id": "M5",
|
||||
"name": "Release and maintenance",
|
||||
"stories": "DGR-068..DGR-071",
|
||||
"outcome": "Reproducible packages, upstream collaboration, beta decision, and sustainable recertification."
|
||||
}
|
||||
],
|
||||
"supersededStories": {
|
||||
"DGR-001": {
|
||||
"newIds": [
|
||||
"DGR-019",
|
||||
"DGR-020",
|
||||
"DGR-054",
|
||||
"DGR-070"
|
||||
],
|
||||
"disposition": "Benchmark scaffold/evidence may be audited; old pass state is void."
|
||||
},
|
||||
"DGR-002": {
|
||||
"newIds": [
|
||||
"DGR-021",
|
||||
"DGR-022",
|
||||
"DGR-023",
|
||||
"DGR-024"
|
||||
],
|
||||
"disposition": "Split protocol, lifecycle, code generation, and fake transport."
|
||||
},
|
||||
"DGR-003": {
|
||||
"newIds": [
|
||||
"DGR-025"
|
||||
],
|
||||
"disposition": "Replaced by exact artifact/runtime compatibility identity."
|
||||
},
|
||||
"DGR-004": {
|
||||
"newIds": [
|
||||
"DGR-027",
|
||||
"DGR-028",
|
||||
"DGR-029",
|
||||
"DGR-030",
|
||||
"DGR-071"
|
||||
],
|
||||
"disposition": "Split provenance, patch stack, builds, and maintenance."
|
||||
},
|
||||
"DGR-005": {
|
||||
"newIds": [
|
||||
"DGR-034",
|
||||
"DGR-045"
|
||||
],
|
||||
"disposition": "Dense and V4 ownership separated."
|
||||
},
|
||||
"DGR-006": {
|
||||
"newIds": [
|
||||
"DGR-031",
|
||||
"DGR-035",
|
||||
"DGR-036",
|
||||
"DGR-046",
|
||||
"DGR-047",
|
||||
"DGR-048",
|
||||
"DGR-049"
|
||||
],
|
||||
"disposition": "Engine, dense boundary, V4 typed boundary, and local-state adapters separated."
|
||||
},
|
||||
"DGR-007": {
|
||||
"newIds": [
|
||||
"DGR-038",
|
||||
"DGR-049"
|
||||
],
|
||||
"disposition": "Replaced by session/epoch-keyed local KV and V4 auxiliary state."
|
||||
},
|
||||
"DGR-008": {
|
||||
"newIds": [
|
||||
"DGR-032",
|
||||
"DGR-033",
|
||||
"DGR-037"
|
||||
],
|
||||
"disposition": "Old implementation/evidence absent; no completion credit transfers."
|
||||
},
|
||||
"DGR-009": {
|
||||
"newIds": [
|
||||
"DGR-040",
|
||||
"DGR-041",
|
||||
"DGR-042",
|
||||
"DGR-043"
|
||||
],
|
||||
"disposition": "Supervision, registration, relay, and routing-input integration separated."
|
||||
},
|
||||
"DGR-010": {
|
||||
"newIds": [
|
||||
"DGR-036",
|
||||
"DGR-039",
|
||||
"DGR-052"
|
||||
],
|
||||
"disposition": "Fixture, dense real acceptance, and V4 parity separated."
|
||||
},
|
||||
"DGR-011": {
|
||||
"newIds": [
|
||||
"DGR-053",
|
||||
"DGR-061",
|
||||
"DGR-062",
|
||||
"DGR-067"
|
||||
],
|
||||
"disposition": "Replaced by scenario-based real 2\u20134, existing-routing 10+, real 10+, and backend certification."
|
||||
},
|
||||
"DGR-012": {
|
||||
"newIds": [
|
||||
"DGR-055",
|
||||
"DGR-056",
|
||||
"DGR-057"
|
||||
],
|
||||
"disposition": "Batching, admission/backpressure, and benchmarking separated."
|
||||
},
|
||||
"DGR-013": {
|
||||
"newIds": [
|
||||
"DGR-058",
|
||||
"DGR-059"
|
||||
],
|
||||
"disposition": "Failure semantics and restart/re-prefill recovery separated."
|
||||
},
|
||||
"DGR-014": {
|
||||
"newIds": [
|
||||
"DGR-019",
|
||||
"DGR-054",
|
||||
"DGR-070"
|
||||
],
|
||||
"disposition": "Replaced by immutable performance, alpha, and beta gates."
|
||||
},
|
||||
"DGR-015": {
|
||||
"newIds": [
|
||||
"DGR-044",
|
||||
"DGR-045",
|
||||
"DGR-046",
|
||||
"DGR-047",
|
||||
"DGR-048",
|
||||
"DGR-049",
|
||||
"DGR-050",
|
||||
"DGR-051",
|
||||
"DGR-052",
|
||||
"DGR-053",
|
||||
"DGR-054",
|
||||
"DGR-060",
|
||||
"DGR-065",
|
||||
"DGR-066",
|
||||
"DGR-067"
|
||||
],
|
||||
"disposition": "Qwen target superseded by DeepSeek V4 Flash; no old completion transfers."
|
||||
},
|
||||
"DGR-016": {
|
||||
"newIds": [
|
||||
"DGR-069",
|
||||
"DGR-071"
|
||||
],
|
||||
"disposition": "Upstream collaboration and ongoing maintenance separated."
|
||||
}
|
||||
},
|
||||
"userStories": [
|
||||
{
|
||||
"id": "DGR-017",
|
||||
@@ -106,7 +364,7 @@
|
||||
"Define controlled safetensors, whole-model GGUF, dense distributed GGUF, and V4 Flash distributed lanes with fixed prompts, context/output lengths, sampling, concurrency, hardware, and metrics.",
|
||||
"Alpha requires correctness plus a human-approved useful-speed threshold; beta adds concurrency, long-context, failure, and sustained-throughput thresholds.",
|
||||
"Separate quantization/model-fit gains from runtime, transport, batching, and kernel gains.",
|
||||
"Treat quants and 2–4/10+ stage counts only as named certification scenarios; no product logic may hardcode them.",
|
||||
"Treat quants and 2\u20134/10+ stage counts only as named certification scenarios; no product logic may hardcode them.",
|
||||
"Lock thresholds and stop conditions in versioned machine-readable data before benchmark result ingestion.",
|
||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||
],
|
||||
@@ -262,12 +520,12 @@
|
||||
"acceptanceCriteria": [
|
||||
"Pin protoc, gRPC, and plugin versions or declare a verified compatible range.",
|
||||
"Generate Python and C++ bindings into out-of-tree build/package locations through documented commands.",
|
||||
"Add Python↔C++ round-trip and descriptor compatibility tests.",
|
||||
"Add Python\u2194C++ round-trip and descriptor compatibility tests.",
|
||||
"A clean checkout regenerates bindings deterministically or fails with an actionable toolchain error.",
|
||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||
],
|
||||
"passes": true,
|
||||
"notes": "Completed from Gitea #7 after controller provisioned and exercised the exact Python/C++ toolchains. Verified deterministic generation, native CMake/CTest, Python↔C++ byte parity, compileall, and diff checks; fixed relative bootstrap prefix resolution.",
|
||||
"notes": "Completed from Gitea #7 after controller provisioned and exercised the exact Python/C++ toolchains. Verified deterministic generation, native CMake/CTest, Python\u2194C++ byte parity, compileall, and diff checks; fixed relative bootstrap prefix resolution.",
|
||||
"completionNotes": "Verified exact grpcio-tools 1.82.1, Protobuf 33.1, Abseil 20250814.1, and gRPC C++ 1.82.1 at commit acccf84c0df20487d64101f528e5d426541ca4e5. Mandatory Python/C++ message and service generation, native CTest, deterministic regeneration, and byte-for-byte Python/C++ parity passed; see evidence/DGR-023/README.md.",
|
||||
"blocks": [
|
||||
"DGR-024",
|
||||
@@ -352,7 +610,7 @@
|
||||
"DGR-041",
|
||||
"DGR-044"
|
||||
],
|
||||
"completionNotes": "Completed 2026-07-17. Verified the live DGR-003-lineage identity core against every criterion: node packages/node/meshnet_node/runtime_recipe.py and the independent tracker packages/tracker/meshnet_tracker/recipe.py (pinned together by tests/data/recipe_fingerprint_vectors.json) fingerprint the source artifact SHA, tokenizer pin, architecture adapter and config digest, boundary/protocol schema versions, backend, weight quant, activation/compute dtypes, and KV dtype/layout under domain-separated digests; shards bind to exact half-open ranges with no topology or quant constants; route/handshake/session checks fail closed with structured mismatch reasons; recipes stay registered-but-dark in the tracker CertificationLedger until a real >=2-distinct-node whole-model distributed forward certifies them. Closed the one open criterion gap (runtime pin/patch stack): new packages/node/meshnet_node/runtime_pin.py derives the runtime_version axis from the DGR-027 lock manifest — exact upstream commit plus a digest over the ordered patch-stack bytes — failing closed on any UPSTREAM_LOCK.json/UPSTREAM_COMMIT/series/SHA256SUMS/patch-byte disagreement, and both identity implementations now reject a moving runtime_version reference. Tests: tests/test_runtime_pin_identity.py (17 passed) plus 196 passing impacted identity/admission/native-emission tests; python -m compileall and git diff --check clean. Also repaired backlog consistency left by prior sessions: added the missing DGR-022/DGR-027 completionNotes, regenerated the DGR-022/025/027 issue projections, and relocated three pre-DGR legacy GLM alpha issue files to issues/legacy/."
|
||||
"completionNotes": "Completed 2026-07-17. Verified the live DGR-003-lineage identity core against every criterion: node packages/node/meshnet_node/runtime_recipe.py and the independent tracker packages/tracker/meshnet_tracker/recipe.py (pinned together by tests/data/recipe_fingerprint_vectors.json) fingerprint the source artifact SHA, tokenizer pin, architecture adapter and config digest, boundary/protocol schema versions, backend, weight quant, activation/compute dtypes, and KV dtype/layout under domain-separated digests; shards bind to exact half-open ranges with no topology or quant constants; route/handshake/session checks fail closed with structured mismatch reasons; recipes stay registered-but-dark in the tracker CertificationLedger until a real >=2-distinct-node whole-model distributed forward certifies them. Closed the one open criterion gap (runtime pin/patch stack): new packages/node/meshnet_node/runtime_pin.py derives the runtime_version axis from the DGR-027 lock manifest \u2014 exact upstream commit plus a digest over the ordered patch-stack bytes \u2014 failing closed on any UPSTREAM_LOCK.json/UPSTREAM_COMMIT/series/SHA256SUMS/patch-byte disagreement, and both identity implementations now reject a moving runtime_version reference. Tests: tests/test_runtime_pin_identity.py (17 passed) plus 196 passing impacted identity/admission/native-emission tests; python -m compileall and git diff --check clean. Also repaired backlog consistency left by prior sessions: added the missing DGR-022/DGR-027 completionNotes, regenerated the DGR-022/025/027 issue projections, and relocated three pre-DGR legacy GLM alpha issue files to issues/legacy/."
|
||||
},
|
||||
{
|
||||
"id": "DGR-026",
|
||||
@@ -419,7 +677,7 @@
|
||||
"Manifest records upstream URL, exact commit, expected source archive/tree hash, license, and retrieval method.",
|
||||
"Fetch tooling verifies identity before use and refuses an unpinned branch/tag.",
|
||||
"Source is fetched into an ignored build workspace; no submodule, vendored source tree, or permanent fork is introduced.",
|
||||
"Offline reuse is supported only after the cached tree’s exact identity is verified.",
|
||||
"Offline reuse is supported only after the cached tree\u2019s exact identity is verified.",
|
||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||
],
|
||||
"passes": true,
|
||||
@@ -656,12 +914,13 @@
|
||||
"The worker exposes neither llama.cpp RPC nor arbitrary graph execution.",
|
||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||
],
|
||||
"passes": false,
|
||||
"passes": true,
|
||||
"notes": "Generated source issue: .scratch/distributed-gguf-runtime/issues/033-build-a-standalone-fake-c-grpc-shard-worker.md; prd.json is authoritative.",
|
||||
"blocks": [
|
||||
"DGR-036",
|
||||
"DGR-040"
|
||||
]
|
||||
],
|
||||
"completionNotes": "Cross-review (Codex GPT-5.5) BLOCK repaired in worktree distributed-gguf-opus. Root protocol defects fixed in the native worker: (1) chunk/decode now fail closed before SessionOpen via a per-session opened flag (terminal ERROR_CODE_INTERNAL), so no activation bypasses lifecycle/cancellation/epoch/flow-control state even when an out-of-band Cancel created placeholder state; (2) flow control is negotiated with strict worker bounds (ShardRuntimeServiceImpl::NegotiateFlow mirrors native_protocol/codec.py negotiate_flow_control) and the negotiated per-session max_chunk_bytes is enforced on every bundle instead of trusting the peer proposal; (3) an in-stream ReleaseSignal now erases session state immediately; (4) SessionOpen rejects incompatible schema, artifact/recipe fingerprint, and shard-range identity and reports the worker own served fingerprint rather than echoing the caller. Nine regression tests added. Real gates on the rebuilt pinned-gRPC binary: cmake --build exit 0; ctest 2/2 passed (shard_worker_selftest, shard_protocol_conformance); tests/test_native_shard_worker.py 27 passed; DGR-024 harness + native protocol 63 passed; compileall exit 0; git diff --check clean; ldd/nm show 0 llama/ggml linkage. Evidence: .scratch/distributed-gguf-runtime/evidence/DGR-033/README.md."
|
||||
},
|
||||
{
|
||||
"id": "DGR-034",
|
||||
@@ -695,13 +954,14 @@
|
||||
"Real-model evidence shows mapped/resident memory scales with owned tensors rather than full artifact size.",
|
||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||
],
|
||||
"passes": false,
|
||||
"passes": true,
|
||||
"notes": "Generated source issue: .scratch/distributed-gguf-runtime/issues/034-implement-dense-llama-range-aware-gguf-ownership.md; prd.json is authoritative.",
|
||||
"blocks": [
|
||||
"DGR-035",
|
||||
"DGR-037",
|
||||
"DGR-051"
|
||||
]
|
||||
],
|
||||
"completionNotes": "Completed by agent"
|
||||
},
|
||||
{
|
||||
"id": "DGR-035",
|
||||
@@ -1158,7 +1418,7 @@
|
||||
"triage": "ready-for-agent",
|
||||
"description": "Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/046-define-the-v4-typed-architecture-boundary-schema.md`, and evidence READMEs for dependencies (DGR-021, DGR-045) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Define the exact cross-stage V4 architecture boundary while keeping per-layer attention and auxiliary caches shard-local.",
|
||||
"acceptanceCriteria": [
|
||||
"Define a versioned named bundle for the mHC 4×4096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations.",
|
||||
"Define a versioned named bundle for the mHC 4\u00d74096 residual boundary, positions, token-ID sideband where required, and schema/cache expectations.",
|
||||
"Explicitly exclude per-layer CSA, HCA, SWA, indexer, compressor, KV, and MTP caches/state from the WAN boundary; those remain local to the owning shard and session/epoch.",
|
||||
"Reserve typed MTP boundary fields but mark MTP execution unsupported and unroutable for alpha.",
|
||||
"Fingerprint independently of quant/topology and fail closed on missing, incompatible, incorrectly shaped, or stale boundary/cache expectations.",
|
||||
@@ -1197,7 +1457,7 @@
|
||||
"triage": "ready-for-agent",
|
||||
"description": "Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/047-adapt-the-upstream-v4-mhc-boundary-for-ranged-ownership.md`, and evidence READMEs for dependencies (DGR-045, DGR-046) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Add range-boundary adapters around upstream llama.cpp V4 mHC execution without reimplementing the V4 graph or kernels.",
|
||||
"acceptanceCriteria": [
|
||||
"Represent and validate the upstream V4 4×4096 mHC boundary without flattening semantic axes.",
|
||||
"Represent and validate the upstream V4 4\u00d74096 mHC boundary without flattening semantic axes.",
|
||||
"Add only head/intermediate/tail range ownership and boundary conversion hooks around the pinned upstream llama.cpp graph.",
|
||||
"Compare deterministic fixture vectors and single-process ranged outputs with upstream whole-model execution.",
|
||||
"Document that llama.cpp owns V4 mHC graph/kernels and that quantized storage does not alter the logical boundary schema.",
|
||||
@@ -1407,7 +1667,7 @@
|
||||
},
|
||||
{
|
||||
"id": "DGR-053",
|
||||
"title": "Certify a real 2–4-stage V4 route",
|
||||
"title": "Certify a real 2\u20134-stage V4 route",
|
||||
"priority": 37,
|
||||
"milestone": "M3",
|
||||
"executionMode": "HITL",
|
||||
@@ -1432,7 +1692,7 @@
|
||||
"triage": "ready-for-human",
|
||||
"description": "Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/053-certify-a-real-2-4-stage-v4-route.md`, and evidence READMEs for dependencies (DGR-030, DGR-043, DGR-052) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Prove real Tracker-selected V4 execution across physical machines before alpha.",
|
||||
"acceptanceCriteria": [
|
||||
"Run one documented 2–4-stage certification scenario using exact compatible artifacts/recipes; the count and chosen quant are evidence inputs, not product constants.",
|
||||
"Run one documented 2\u20134-stage certification scenario using exact compatible artifacts/recipes; the count and chosen quant are evidence inputs, not product constants.",
|
||||
"Actual CPU/GPU work executes on every stage; fake workers do not satisfy acceptance.",
|
||||
"Record parity, TTFT, prefill/decode speed, seam cost, memory, cache/state isolation, cancellation, and cleanup.",
|
||||
"Tracker selection remains dynamic and rejects an injected incompatible backend/recipe.",
|
||||
@@ -1711,7 +1971,7 @@
|
||||
"DGR-058"
|
||||
],
|
||||
"triage": "ready-for-agent",
|
||||
"description": "Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/060-certify-v4-long-context-state-correctness.md`, and evidence READMEs for dependencies (DGR-051, DGR-056, DGR-058) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Prove V4’s KV and auxiliary state remain correct and bounded at long contexts.",
|
||||
"description": "Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`, source issue `.scratch/distributed-gguf-runtime/issues/060-certify-v4-long-context-state-correctness.md`, and evidence READMEs for dependencies (DGR-051, DGR-056, DGR-058) before changing code. Inspect live source/tests rather than trusting legacy pass states. Objective: Prove V4\u2019s KV and auxiliary state remain correct and bounded at long contexts.",
|
||||
"acceptanceCriteria": [
|
||||
"Exercise pre-locked context lengths covering multiple prefill chunks and sustained decode.",
|
||||
"Validate KV plus CSA/HCA/SWA/indexer/compressor state positions across every stage.",
|
||||
@@ -2164,6 +2424,6 @@
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"updatedAt": "2026-07-23T08:09:16.081Z"
|
||||
"updatedAt": "2026-07-23T08:09:17.286Z"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
218
packages/node/meshnet_node/range_report.py
Normal file
218
packages/node/meshnet_node/range_report.py
Normal file
@@ -0,0 +1,218 @@
|
||||
"""Authoritative dense-Llama owned-range reports from the loaded engine state.
|
||||
|
||||
DGR-034 loads only the tensors a shard range owns through the Meshnet
|
||||
owned-range loader (``llama_model_params::meshnet_owned_layer_start/end`` in
|
||||
the pinned llama.cpp patch stack). The project-owned ``meshnet-range-report``
|
||||
native tool runs that load and prints a JSON document derived from the loaded
|
||||
model state — the registered tensor set and the backend buffers — never from
|
||||
caller-asserted values. This module is the strict consumer of that document:
|
||||
it parses it into :class:`OwnedRangeReport` and fails closed on any
|
||||
inconsistency, so a range or endpoint claim that the loaded engine state does
|
||||
not back is rejected before it can reach identity, admission, or routing.
|
||||
|
||||
Ownership contract enforced here (dense Llama only):
|
||||
|
||||
- every registered ``blk.N.*`` tensor lies inside the half-open owned range
|
||||
``[start, end)``, and every layer in that range is present — a gapped or
|
||||
out-of-range registration is rejected;
|
||||
- ``token_embd.weight`` is registered only by the head shard (``start == 0``),
|
||||
or by a tail shard whose model ties the output head to the embedding
|
||||
(``end == n_layer`` and no separate ``output.weight``);
|
||||
- ``output_norm.weight`` and ``output.weight`` are registered only by the
|
||||
tail shard (``end == n_layer``);
|
||||
- any other registered tensor name is unexpected and rejected;
|
||||
- byte counts are consistent: an mmap load maps a file span at least the
|
||||
registered tensor bytes and at most the artifact size; a non-mmap load
|
||||
reports a resident allocation at least the registered tensor bytes.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Mapping
|
||||
|
||||
|
||||
class RangeReportError(ValueError):
|
||||
"""A range report is malformed, or the loaded state breaks ownership."""
|
||||
|
||||
|
||||
_DENSE_ARCHITECTURE = "llama"
|
||||
|
||||
_INT_FIELDS = (
|
||||
"n_layer",
|
||||
"file_bytes",
|
||||
"mapped_bytes",
|
||||
"resident_bytes",
|
||||
"registered_tensors",
|
||||
"registered_bytes",
|
||||
)
|
||||
|
||||
_BOOL_FIELDS = (
|
||||
"mmap",
|
||||
"touched",
|
||||
"has_token_embeddings",
|
||||
"has_output_head",
|
||||
"tied_output_head",
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class OwnedRangeReport:
|
||||
"""One validated owned-range load, derived from loaded engine state.
|
||||
|
||||
``start_layer``/``end_layer`` are the authoritative half-open owned range
|
||||
the engine actually registered (the tool already refused a report whose
|
||||
loaded bounds differ from the requested ones). ``has_token_embeddings`` is
|
||||
true for the head shard, and also for a tail shard on a tied-output model
|
||||
(the embedding tensor *is* its output head); ``tied_output_head``
|
||||
disambiguates those two cases. ``mapped_bytes``/``resident_bytes`` come
|
||||
from the backend buffers: with mmap they are the mapped file span holding
|
||||
the owned tensors, without mmap the resident allocation holding them.
|
||||
"""
|
||||
|
||||
architecture: str
|
||||
n_layer: int
|
||||
start_layer: int
|
||||
end_layer: int
|
||||
has_token_embeddings: bool
|
||||
has_output_head: bool
|
||||
tied_output_head: bool
|
||||
mapped_bytes: int
|
||||
resident_bytes: int
|
||||
registered_tensors: int
|
||||
registered_bytes: int
|
||||
file_bytes: int
|
||||
mmap: bool
|
||||
touched: bool
|
||||
vm_size_bytes: int | None
|
||||
vm_rss_bytes: int | None
|
||||
vm_hwm_bytes: int | None
|
||||
|
||||
@property
|
||||
def is_head(self) -> bool:
|
||||
return self.start_layer == 0
|
||||
|
||||
@property
|
||||
def is_tail(self) -> bool:
|
||||
return self.end_layer == self.n_layer
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if self.architecture != _DENSE_ARCHITECTURE:
|
||||
raise RangeReportError(
|
||||
f"owned-range loading supports dense Llama only, got {self.architecture!r}"
|
||||
)
|
||||
if isinstance(self.n_layer, bool) or self.n_layer < 1:
|
||||
raise RangeReportError("report must record a positive GGUF block count")
|
||||
for name in _INT_FIELDS:
|
||||
value = getattr(self, name)
|
||||
if isinstance(value, bool) or not isinstance(value, int) or value < 0:
|
||||
raise RangeReportError(f"report field {name!r} must be a non-negative integer")
|
||||
for name in _BOOL_FIELDS:
|
||||
if not isinstance(getattr(self, name), bool):
|
||||
raise RangeReportError(f"report field {name!r} must be a boolean")
|
||||
if not 0 <= self.start_layer < self.end_layer <= self.n_layer:
|
||||
raise RangeReportError(
|
||||
f"owned range [{self.start_layer}, {self.end_layer}) is empty or "
|
||||
f"outside the model's {self.n_layer} layers"
|
||||
)
|
||||
if self.tied_output_head and not self.is_tail:
|
||||
raise RangeReportError("a tied output head can only belong to the tail shard")
|
||||
expected_embeddings = self.is_head or self.tied_output_head
|
||||
if self.has_token_embeddings != expected_embeddings:
|
||||
raise RangeReportError(
|
||||
"token-embedding registration disagrees with endpoint ownership: "
|
||||
"embeddings belong to the head shard (or to a tied-output tail)"
|
||||
)
|
||||
if self.has_output_head != self.is_tail:
|
||||
raise RangeReportError(
|
||||
"output-head registration disagrees with endpoint ownership: "
|
||||
"the final norm and output head belong to the tail shard"
|
||||
)
|
||||
if self.registered_tensors < 1 or self.registered_bytes < 1:
|
||||
raise RangeReportError("the owned range registered no tensors")
|
||||
if self.file_bytes < 1:
|
||||
raise RangeReportError("report must record the artifact size")
|
||||
if self.mmap:
|
||||
if self.mapped_bytes < self.registered_bytes:
|
||||
raise RangeReportError(
|
||||
"mapped span undercounts the registered owned tensors"
|
||||
)
|
||||
if self.mapped_bytes > self.file_bytes:
|
||||
raise RangeReportError("mapped span exceeds the artifact size")
|
||||
else:
|
||||
if self.mapped_bytes != 0:
|
||||
raise RangeReportError("a non-mmap load must not claim a mapped span")
|
||||
if self.resident_bytes < self.registered_bytes:
|
||||
raise RangeReportError(
|
||||
"resident allocation undercounts the registered owned tensors"
|
||||
)
|
||||
for name in ("vm_size_bytes", "vm_rss_bytes", "vm_hwm_bytes"):
|
||||
value = getattr(self, name)
|
||||
if value is not None and (
|
||||
isinstance(value, bool) or not isinstance(value, int) or value < 0
|
||||
):
|
||||
raise RangeReportError(f"report field {name!r} must be a non-negative integer or null")
|
||||
|
||||
|
||||
def _require_range(doc: Mapping[str, Any], key: str) -> tuple[int, int]:
|
||||
value = doc.get(key)
|
||||
if (
|
||||
not isinstance(value, (list, tuple))
|
||||
or len(value) != 2
|
||||
or any(isinstance(v, bool) or not isinstance(v, int) for v in value)
|
||||
):
|
||||
raise RangeReportError(f"report field {key!r} must be a [start, end] integer pair")
|
||||
return value[0], value[1]
|
||||
|
||||
|
||||
def parse_owned_range_report(doc: Mapping[str, Any]) -> OwnedRangeReport:
|
||||
"""Parse and validate one ``meshnet-range-report`` JSON document.
|
||||
|
||||
Fails closed: a load the tool rejected (``ok: false``), a requested range
|
||||
the loaded state did not match, a gapped or out-of-range registration, an
|
||||
unexpected registered tensor, and any byte-count inconsistency all raise
|
||||
:class:`RangeReportError` instead of producing a report.
|
||||
"""
|
||||
if not isinstance(doc, Mapping):
|
||||
raise RangeReportError("range report must be a JSON object")
|
||||
if doc.get("ok") is not True:
|
||||
error = doc.get("error")
|
||||
detail = f": {error}" if isinstance(error, str) and error else ""
|
||||
raise RangeReportError(f"the owned-range load was rejected{detail}")
|
||||
|
||||
requested = _require_range(doc, "requested_range")
|
||||
reported = _require_range(doc, "reported_range")
|
||||
if requested != reported:
|
||||
raise RangeReportError(
|
||||
f"reported range {reported} does not match the requested range {requested}; "
|
||||
"ownership must be derived from the loaded engine state"
|
||||
)
|
||||
|
||||
for key in ("unexpected_registered_tensors", "missing_owned_layers"):
|
||||
value = doc.get(key)
|
||||
if not isinstance(value, list):
|
||||
raise RangeReportError(f"report field {key!r} must be a list")
|
||||
if value:
|
||||
raise RangeReportError(
|
||||
f"ownership audit failed: {key} is {value!r}; the registered "
|
||||
"tensor set must exactly cover the owned range and its endpoints"
|
||||
)
|
||||
|
||||
architecture = doc.get("architecture")
|
||||
if not isinstance(architecture, str):
|
||||
raise RangeReportError("report field 'architecture' must be a string")
|
||||
|
||||
fields: dict[str, Any] = {}
|
||||
for name in _INT_FIELDS + _BOOL_FIELDS:
|
||||
if name not in doc:
|
||||
raise RangeReportError(f"range report is missing field {name!r}")
|
||||
fields[name] = doc[name]
|
||||
for name in ("vm_size_bytes", "vm_rss_bytes", "vm_hwm_bytes"):
|
||||
fields[name] = doc.get(name)
|
||||
|
||||
return OwnedRangeReport(
|
||||
architecture=architecture,
|
||||
start_layer=reported[0],
|
||||
end_layer=reported[1],
|
||||
**fields,
|
||||
)
|
||||
@@ -62,6 +62,21 @@ message(STATUS "Pinned gRPC ${gRPC_VERSION}: building ShardRuntime service stubs
|
||||
|
||||
enable_testing()
|
||||
|
||||
# The standalone fake Shard worker (DGR-033): a real gRPC server over the
|
||||
# ShardRuntime service, backed by the model-free FakeShardEngine. It links the
|
||||
# grpc service stubs only — no llama.cpp, no graph-execution entry point.
|
||||
add_executable(shard_worker
|
||||
worker/shard_worker_main.cpp
|
||||
worker/shard_service.cpp)
|
||||
target_include_directories(shard_worker PRIVATE "${CMAKE_CURRENT_SOURCE_DIR}/worker")
|
||||
target_link_libraries(shard_worker PRIVATE shard_runtime_grpc gRPC::grpc++)
|
||||
|
||||
# Pure-C++ CTest: the worker binds an ephemeral port, self-drives the full
|
||||
# lifecycle (capability, health, fragmented prefill, decode, release) over a
|
||||
# real loopback gRPC channel, and exits non-zero on any mismatch. This proves
|
||||
# the worker serves the contract without needing a Python environment.
|
||||
add_test(NAME shard_worker_selftest COMMAND shard_worker --selftest)
|
||||
|
||||
add_executable(shard_protocol_conformance tests/test_shard_protocol_conformance.cpp)
|
||||
target_link_libraries(shard_protocol_conformance PRIVATE shard_runtime_proto)
|
||||
|
||||
|
||||
@@ -28,6 +28,12 @@ One numbered patch per concern (ADR-0024 local seams only):
|
||||
5. `0005-worker-range-report-hook.patch` (worker hooks) exposes the
|
||||
`llama_model_meshnet_range_report` C API the project-owned worker binds to
|
||||
and registers a model-free native fixture test for it.
|
||||
6. `0006-meshnet-range-report-tool.patch` (range reporting) adds the
|
||||
project-owned `meshnet-range-report` tool: it loads one GGUF artifact
|
||||
through the owned-range loader and prints a JSON document derived from the
|
||||
loaded model state — the owned-range report, the registered tensor set
|
||||
audited against the requested ownership, and backend-buffer byte counts.
|
||||
It never builds or runs a compute graph.
|
||||
|
||||
Meshnet routing, Tracker, gRPC, relay, billing, authentication, and telemetry
|
||||
remain outside this directory; the stack is checked for such control-plane
|
||||
|
||||
@@ -10,21 +10,23 @@
|
||||
"method": "git-clone-detached-commit",
|
||||
"workspace": "build/llama.cpp"
|
||||
},
|
||||
"patched_tree": "c0045714735ae5ee7b7334a480d8ac04e03e1b18",
|
||||
"patched_tree": "8f7e87fea6743f0b9744afe44f9e6f9ca3b7d08a",
|
||||
"upstream_license": "MIT",
|
||||
"patch_series": [
|
||||
"0001-cmake-reserve-meshnet-patch-stack-abi-marker.patch",
|
||||
"0002-dense-llama-owned-range-loading.patch",
|
||||
"0003-owned-range-filtered-state-report.patch",
|
||||
"0004-dense-boundary-io-endpoint-guard.patch",
|
||||
"0005-worker-range-report-hook.patch"
|
||||
"0005-worker-range-report-hook.patch",
|
||||
"0006-meshnet-range-report-tool.patch"
|
||||
],
|
||||
"patch_scope": [
|
||||
"Reserved CMake ABI marker only; no execution or model semantics.",
|
||||
"Range loading: dense-Llama owned-range params, validation, and filtered tensor registration with endpoint ownership.",
|
||||
"Filtered state: owned-range report populated from registered tensors and backend buffers, derived never asserted.",
|
||||
"Boundary I/O: endpoint ownership flags and a fail-closed dense graph guard until typed endpoint adapters exist.",
|
||||
"Worker hooks: public C range-report API and the model-free native fixture test the project-owned worker binds to."
|
||||
"Worker hooks: public C range-report API and the model-free native fixture test the project-owned worker binds to.",
|
||||
"Range reporting: project-owned tool that loads one artifact through the owned-range loader and reports derived ownership and buffer-byte state as JSON."
|
||||
],
|
||||
"patch_assumptions": "patches/UPSTREAM-ASSUMPTIONS.json",
|
||||
"build": {
|
||||
@@ -46,7 +48,7 @@
|
||||
"-DGGML_VULKAN=OFF",
|
||||
"-DGGML_METAL=OFF"
|
||||
],
|
||||
"native_targets": ["llama-gguf-hash", "test-meshnet-range-ownership"],
|
||||
"native_targets": ["llama-gguf-hash", "test-meshnet-range-ownership", "meshnet-range-report"],
|
||||
"smoke_binary": "bin/llama-gguf-hash",
|
||||
"smoke_args": ["--help"],
|
||||
"smoke_output_token": "usage",
|
||||
@@ -81,7 +83,9 @@
|
||||
"src/llama-model.h",
|
||||
"src/models/llama.cpp",
|
||||
"tests/CMakeLists.txt",
|
||||
"tests/test-meshnet-range-ownership.cpp"
|
||||
"tests/test-meshnet-range-ownership.cpp",
|
||||
"tools/meshnet-range-report/CMakeLists.txt",
|
||||
"tools/meshnet-range-report/meshnet-range-report.cpp"
|
||||
],
|
||||
"stock_glm_limitations": "This pin may load GLM-5.2 through the dense-MLA compatibility fallback. It does not prove native DSA, IndexShare, MoE semantic correctness, numerical equivalence, performance, or route certification."
|
||||
}
|
||||
|
||||
@@ -0,0 +1,414 @@
|
||||
From: Meshnet <meshnet@invalid>
|
||||
Subject: [PATCH] llama: add dense-Llama owned-range report tool
|
||||
|
||||
Concern: range reporting. Adds the project-owned meshnet-range-report tool:
|
||||
it loads one GGUF artifact through the Meshnet owned-range loader and prints
|
||||
a JSON document derived from the loaded model state — the owned-range
|
||||
report, the registered tensor set audited against the requested ownership,
|
||||
and backend-buffer byte counts (optionally split from repack buffers, plus
|
||||
process resident readings). It never builds or runs a compute graph and
|
||||
never trusts caller-asserted range or endpoint claims.
|
||||
---
|
||||
diff --git a/CMakeLists.txt b/CMakeLists.txt
|
||||
index a9afcff..868793b 100644
|
||||
--- a/CMakeLists.txt
|
||||
+++ b/CMakeLists.txt
|
||||
@@ -281,3 +281,6 @@ configure_file(cmake/llama.pc.in
|
||||
|
||||
install(FILES "${CMAKE_CURRENT_BINARY_DIR}/llama.pc"
|
||||
DESTINATION ${CMAKE_INSTALL_LIBDIR}/pkgconfig)
|
||||
+
|
||||
+# Meshnet-owned owned-range report tool (patch stack, range-report concern).
|
||||
+add_subdirectory(tools/meshnet-range-report)
|
||||
diff --git a/tools/meshnet-range-report/CMakeLists.txt b/tools/meshnet-range-report/CMakeLists.txt
|
||||
new file mode 100644
|
||||
index 000000000..24401007e
|
||||
--- /dev/null
|
||||
+++ b/tools/meshnet-range-report/CMakeLists.txt
|
||||
@@ -0,0 +1,7 @@
|
||||
+# Meshnet-owned dense-Llama owned-range load/report tool.
|
||||
+#
|
||||
+# Built unconditionally with the patched tree: it exercises the Meshnet
|
||||
+# owned-range loader against real GGUF artifacts and reports only state
|
||||
+# derived from the loaded model (registered tensors, backend buffers).
|
||||
+add_executable(meshnet-range-report meshnet-range-report.cpp)
|
||||
+target_link_libraries(meshnet-range-report PRIVATE llama)
|
||||
diff --git a/tools/meshnet-range-report/meshnet-range-report.cpp b/tools/meshnet-range-report/meshnet-range-report.cpp
|
||||
new file mode 100644
|
||||
index 000000000..49a5eb2a0
|
||||
--- /dev/null
|
||||
+++ b/tools/meshnet-range-report/meshnet-range-report.cpp
|
||||
@@ -0,0 +1,373 @@
|
||||
+// Meshnet-owned dense-Llama owned-range load/report tool.
|
||||
+//
|
||||
+// Loads one GGUF artifact through the Meshnet owned-range loader
|
||||
+// (llama_model_params::meshnet_owned_layer_start/end) and prints a single
|
||||
+// JSON report derived from the loaded model state — registered tensors and
|
||||
+// backend buffers, never caller-asserted values. The audit fails closed when
|
||||
+// the registered tensor set disagrees with the requested ownership: every
|
||||
+// registered per-layer tensor must lie inside [start, end), the token
|
||||
+// embedding may be registered only by the head shard (start == 0) or by a
|
||||
+// tail shard whose model ties the output head to the embedding, and the
|
||||
+// final norm plus output head may be registered only by the tail shard
|
||||
+// (end == n_layer).
|
||||
+
|
||||
+#include "ggml.h"
|
||||
+#include "llama.h"
|
||||
+
|
||||
+#include "../../src/llama-model.h"
|
||||
+
|
||||
+#include <cstdint>
|
||||
+#include <cstdio>
|
||||
+#include <cstdlib>
|
||||
+#include <cstring>
|
||||
+#include <set>
|
||||
+#include <string>
|
||||
+#include <sys/stat.h>
|
||||
+#include <vector>
|
||||
+
|
||||
+namespace {
|
||||
+
|
||||
+constexpr int kExitUsage = 2;
|
||||
+constexpr int kExitLoad = 3;
|
||||
+constexpr int kExitAudit = 4;
|
||||
+
|
||||
+std::string g_log_tail;
|
||||
+
|
||||
+void capture_log(enum ggml_log_level level, const char * text, void *) {
|
||||
+ if (level >= GGML_LOG_LEVEL_ERROR) {
|
||||
+ g_log_tail += text;
|
||||
+ if (g_log_tail.size() > 512) {
|
||||
+ g_log_tail.erase(0, g_log_tail.size() - 512);
|
||||
+ }
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+std::string json_escape(const std::string & value) {
|
||||
+ std::string out;
|
||||
+ for (const char c : value) {
|
||||
+ if (c == '"' || c == '\\') {
|
||||
+ out += '\\';
|
||||
+ out += c;
|
||||
+ } else if (c == '\n') {
|
||||
+ out += "\\n";
|
||||
+ } else if (c == '\r') {
|
||||
+ // drop carriage returns from embedded log text
|
||||
+ } else {
|
||||
+ out += c;
|
||||
+ }
|
||||
+ }
|
||||
+ return out;
|
||||
+}
|
||||
+
|
||||
+std::string json_string_array(const std::vector<std::string> & items) {
|
||||
+ std::string out = "[";
|
||||
+ for (size_t i = 0; i < items.size(); ++i) {
|
||||
+ if (i) {
|
||||
+ out += ", ";
|
||||
+ }
|
||||
+ out += "\"" + json_escape(items[i]) + "\"";
|
||||
+ }
|
||||
+ return out + "]";
|
||||
+}
|
||||
+
|
||||
+std::string json_int_array(const std::vector<int> & items) {
|
||||
+ std::string out = "[";
|
||||
+ for (size_t i = 0; i < items.size(); ++i) {
|
||||
+ if (i) {
|
||||
+ out += ", ";
|
||||
+ }
|
||||
+ out += std::to_string(items[i]);
|
||||
+ }
|
||||
+ return out + "]";
|
||||
+}
|
||||
+
|
||||
+int fail(int code, const std::string & error) {
|
||||
+ std::string detail = error;
|
||||
+ if (!g_log_tail.empty()) {
|
||||
+ detail += ": " + g_log_tail;
|
||||
+ }
|
||||
+ std::printf("{\"ok\": false, \"error\": \"%s\"}\n", json_escape(detail).c_str());
|
||||
+ return code;
|
||||
+}
|
||||
+
|
||||
+bool parse_nonnegative(const char * text, int & out) {
|
||||
+ if (text == nullptr || *text == '\0' || *text == '-') {
|
||||
+ return false;
|
||||
+ }
|
||||
+ char * end = nullptr;
|
||||
+ const long value = std::strtol(text, &end, 10);
|
||||
+ if (end == text || *end != '\0' || value > INT32_MAX) {
|
||||
+ return false;
|
||||
+ }
|
||||
+ out = static_cast<int>(value);
|
||||
+ return true;
|
||||
+}
|
||||
+
|
||||
+uint64_t file_size(const std::string & path) {
|
||||
+ struct stat st;
|
||||
+ return ::stat(path.c_str(), &st) == 0 ? static_cast<uint64_t>(st.st_size) : 0;
|
||||
+}
|
||||
+
|
||||
+struct proc_status {
|
||||
+ uint64_t vm_size = 0;
|
||||
+ uint64_t vm_rss = 0;
|
||||
+ uint64_t vm_hwm = 0;
|
||||
+ bool valid = false;
|
||||
+};
|
||||
+
|
||||
+proc_status read_proc_status() {
|
||||
+ proc_status out;
|
||||
+#ifdef __linux__
|
||||
+ FILE * f = std::fopen("/proc/self/status", "r");
|
||||
+ if (!f) {
|
||||
+ return out;
|
||||
+ }
|
||||
+ char line[256];
|
||||
+ while (std::fgets(line, sizeof(line), f)) {
|
||||
+ uint64_t kb = 0;
|
||||
+ if (std::sscanf(line, "VmSize: %lu kB", &kb) == 1) {
|
||||
+ out.vm_size = kb * 1024;
|
||||
+ } else if (std::sscanf(line, "VmRSS: %lu kB", &kb) == 1) {
|
||||
+ out.vm_rss = kb * 1024;
|
||||
+ } else if (std::sscanf(line, "VmHWM: %lu kB", &kb) == 1) {
|
||||
+ out.vm_hwm = kb * 1024;
|
||||
+ }
|
||||
+ }
|
||||
+ std::fclose(f);
|
||||
+ out.valid = true;
|
||||
+#endif
|
||||
+ return out;
|
||||
+}
|
||||
+
|
||||
+void usage(const char * argv0) {
|
||||
+ std::fprintf(stderr,
|
||||
+ "usage: %s --model PATH --start N --end M [--no-mmap] [--no-extra-bufts] [--touch]\n"
|
||||
+ "loads one dense-Llama GGUF through the Meshnet owned-range loader and\n"
|
||||
+ "prints a JSON report derived from the loaded model state\n",
|
||||
+ argv0);
|
||||
+}
|
||||
+
|
||||
+} // namespace
|
||||
+
|
||||
+int main(int argc, char ** argv) {
|
||||
+ std::string model_path;
|
||||
+ int start = -1;
|
||||
+ int end = -1;
|
||||
+ bool use_mmap = true;
|
||||
+ bool use_extra_bufts = true;
|
||||
+ bool touch = false;
|
||||
+
|
||||
+ for (int i = 1; i < argc; ++i) {
|
||||
+ const std::string arg = argv[i];
|
||||
+ if (arg == "--model" && i + 1 < argc) {
|
||||
+ model_path = argv[++i];
|
||||
+ } else if (arg == "--start" && i + 1 < argc) {
|
||||
+ if (!parse_nonnegative(argv[++i], start)) {
|
||||
+ usage(argv[0]);
|
||||
+ return kExitUsage;
|
||||
+ }
|
||||
+ } else if (arg == "--end" && i + 1 < argc) {
|
||||
+ if (!parse_nonnegative(argv[++i], end)) {
|
||||
+ usage(argv[0]);
|
||||
+ return kExitUsage;
|
||||
+ }
|
||||
+ } else if (arg == "--no-mmap") {
|
||||
+ use_mmap = false;
|
||||
+ } else if (arg == "--no-extra-bufts") {
|
||||
+ use_extra_bufts = false;
|
||||
+ } else if (arg == "--touch") {
|
||||
+ touch = true;
|
||||
+ } else {
|
||||
+ usage(argv[0]);
|
||||
+ return kExitUsage;
|
||||
+ }
|
||||
+ }
|
||||
+ if (model_path.empty() || start < 0 || end < 0) {
|
||||
+ usage(argv[0]);
|
||||
+ return kExitUsage;
|
||||
+ }
|
||||
+
|
||||
+ llama_log_set(capture_log, nullptr);
|
||||
+ llama_backend_init();
|
||||
+
|
||||
+ llama_model_params params = llama_model_default_params();
|
||||
+ params.meshnet_owned_layer_start = start;
|
||||
+ params.meshnet_owned_layer_end = end;
|
||||
+ params.use_mmap = use_mmap;
|
||||
+ params.use_extra_bufts = use_extra_bufts;
|
||||
+ params.progress_callback = nullptr;
|
||||
+
|
||||
+ llama_model * model = llama_model_load_from_file(model_path.c_str(), params);
|
||||
+ if (model == nullptr) {
|
||||
+ return fail(kExitLoad, "owned-range load rejected the artifact or range");
|
||||
+ }
|
||||
+
|
||||
+ llama_meshnet_range_report report = {};
|
||||
+ if (!llama_model_meshnet_range_report(model, &report)) {
|
||||
+ llama_model_free(model);
|
||||
+ return fail(kExitLoad, "loaded model carries no owned-range report");
|
||||
+ }
|
||||
+
|
||||
+ char arch_buf[128] = {};
|
||||
+ std::string arch;
|
||||
+ if (llama_model_meta_val_str(model, "general.architecture", arch_buf, sizeof(arch_buf)) >= 0) {
|
||||
+ arch = arch_buf;
|
||||
+ }
|
||||
+ const int n_layer = llama_model_n_layer(model);
|
||||
+ const uint64_t bytes_on_disk = file_size(model_path);
|
||||
+
|
||||
+ // Audit the registered tensor set against the requested ownership.
|
||||
+ const auto & tensors = llama_internal_get_tensor_map(model);
|
||||
+ bool has_embd = false;
|
||||
+ bool has_out_norm = false;
|
||||
+ bool has_out = false;
|
||||
+ std::set<int> owned_layers;
|
||||
+ std::vector<std::string> unexpected;
|
||||
+ uint64_t registered_bytes = 0;
|
||||
+ for (const auto & entry : tensors) {
|
||||
+ const std::string & name = entry.first;
|
||||
+ registered_bytes += ggml_nbytes(entry.second);
|
||||
+ if (name == "token_embd.weight") {
|
||||
+ has_embd = true;
|
||||
+ continue;
|
||||
+ }
|
||||
+ if (name == "output_norm.weight") {
|
||||
+ has_out_norm = true;
|
||||
+ continue;
|
||||
+ }
|
||||
+ if (name == "output.weight") {
|
||||
+ has_out = true;
|
||||
+ continue;
|
||||
+ }
|
||||
+ int block = -1;
|
||||
+ if (std::sscanf(name.c_str(), "blk.%d.", &block) == 1 && block >= 0) {
|
||||
+ owned_layers.insert(block);
|
||||
+ continue;
|
||||
+ }
|
||||
+ unexpected.push_back(name);
|
||||
+ }
|
||||
+
|
||||
+ // A tail shard whose model ties the output head to the token embedding
|
||||
+ // registers token_embd.weight as its output head instead of output.weight.
|
||||
+ const bool tied_tail = end == n_layer && has_embd && !has_out;
|
||||
+ const bool expect_embd = start == 0 || tied_tail;
|
||||
+
|
||||
+ std::vector<int> missing_layers;
|
||||
+ for (int i = start; i < end; ++i) {
|
||||
+ if (!owned_layers.count(i)) {
|
||||
+ missing_layers.push_back(i);
|
||||
+ }
|
||||
+ }
|
||||
+ std::vector<int> outside_layers;
|
||||
+ for (const int block : owned_layers) {
|
||||
+ if (block < start || block >= end) {
|
||||
+ outside_layers.push_back(block);
|
||||
+ }
|
||||
+ }
|
||||
+
|
||||
+ std::vector<std::string> mismatches;
|
||||
+ if (report.start_layer != start || report.end_layer != end) {
|
||||
+ mismatches.push_back("reported range differs from the requested range");
|
||||
+ }
|
||||
+ if (has_embd != expect_embd) {
|
||||
+ mismatches.push_back("token-embedding registration disagrees with endpoint ownership");
|
||||
+ }
|
||||
+ if ((end == n_layer) && !has_out_norm) {
|
||||
+ mismatches.push_back("tail range is missing the final norm");
|
||||
+ }
|
||||
+ if ((end == n_layer) && !has_out && !has_embd) {
|
||||
+ mismatches.push_back("tail range is missing the output head");
|
||||
+ }
|
||||
+ if ((end != n_layer) && (has_out_norm || has_out)) {
|
||||
+ mismatches.push_back("non-tail range registered tail-only tensors");
|
||||
+ }
|
||||
+ if (report.has_token_embeddings != has_embd) {
|
||||
+ mismatches.push_back("reported embedding ownership disagrees with registered tensors");
|
||||
+ }
|
||||
+ if (report.has_output_head != (end == n_layer)) {
|
||||
+ mismatches.push_back("reported output-head ownership disagrees with endpoint ownership");
|
||||
+ }
|
||||
+ if (!missing_layers.empty()) {
|
||||
+ mismatches.push_back("owned range has missing per-layer tensors");
|
||||
+ }
|
||||
+ if (!outside_layers.empty()) {
|
||||
+ mismatches.push_back("registered per-layer tensors lie outside the owned range");
|
||||
+ }
|
||||
+ if (!unexpected.empty()) {
|
||||
+ mismatches.push_back("registered tensors outside the dense-Llama ownership vocabulary");
|
||||
+ }
|
||||
+ if (use_mmap && report.mapped_bytes < registered_bytes) {
|
||||
+ mismatches.push_back("mapped span undercounts the registered tensors");
|
||||
+ }
|
||||
+ if (!use_mmap && report.resident_bytes < registered_bytes) {
|
||||
+ mismatches.push_back("resident allocation undercounts the registered tensors");
|
||||
+ }
|
||||
+
|
||||
+ if (touch) {
|
||||
+ volatile uint64_t sink = 0;
|
||||
+ for (const auto & entry : tensors) {
|
||||
+ const auto * data = static_cast<const volatile uint8_t *>(entry.second->data);
|
||||
+ const size_t nbytes = ggml_nbytes(entry.second);
|
||||
+ for (size_t i = 0; i < nbytes; i += 4096) {
|
||||
+ sink += data[i];
|
||||
+ }
|
||||
+ }
|
||||
+ (void) sink;
|
||||
+ }
|
||||
+
|
||||
+ const proc_status proc = read_proc_status();
|
||||
+
|
||||
+ if (!mismatches.empty()) {
|
||||
+ llama_model_free(model);
|
||||
+ return fail(kExitAudit, "ownership audit failed: " + json_string_array(mismatches));
|
||||
+ }
|
||||
+
|
||||
+ std::printf(
|
||||
+ "{\n"
|
||||
+ " \"ok\": true,\n"
|
||||
+ " \"model\": \"%s\",\n"
|
||||
+ " \"architecture\": \"%s\",\n"
|
||||
+ " \"n_layer\": %d,\n"
|
||||
+ " \"file_bytes\": %llu,\n"
|
||||
+ " \"requested_range\": [%d, %d],\n"
|
||||
+ " \"reported_range\": [%d, %d],\n"
|
||||
+ " \"mmap\": %s,\n"
|
||||
+ " \"touched\": %s,\n"
|
||||
+ " \"use_extra_bufts\": %s,\n"
|
||||
+ " \"has_token_embeddings\": %s,\n"
|
||||
+ " \"has_output_head\": %s,\n"
|
||||
+ " \"tied_output_head\": %s,\n"
|
||||
+ " \"mapped_bytes\": %llu,\n"
|
||||
+ " \"resident_bytes\": %llu,\n"
|
||||
+ " \"registered_tensors\": %d,\n"
|
||||
+ " \"registered_bytes\": %llu,\n"
|
||||
+ " \"unexpected_registered_tensors\": [],\n"
|
||||
+ " \"missing_owned_layers\": [],\n"
|
||||
+ " \"vm_size_bytes\": %llu,\n"
|
||||
+ " \"vm_rss_bytes\": %llu,\n"
|
||||
+ " \"vm_hwm_bytes\": %llu\n"
|
||||
+ "}\n",
|
||||
+ json_escape(model_path).c_str(),
|
||||
+ json_escape(arch).c_str(),
|
||||
+ n_layer,
|
||||
+ (unsigned long long) bytes_on_disk,
|
||||
+ start, end,
|
||||
+ report.start_layer, report.end_layer,
|
||||
+ use_mmap ? "true" : "false",
|
||||
+ touch ? "true" : "false",
|
||||
+ use_extra_bufts ? "true" : "false",
|
||||
+ report.has_token_embeddings ? "true" : "false",
|
||||
+ report.has_output_head ? "true" : "false",
|
||||
+ tied_tail ? "true" : "false",
|
||||
+ (unsigned long long) report.mapped_bytes,
|
||||
+ (unsigned long long) report.resident_bytes,
|
||||
+ (int) tensors.size(),
|
||||
+ (unsigned long long) registered_bytes,
|
||||
+ (unsigned long long) proc.vm_size,
|
||||
+ (unsigned long long) proc.vm_rss,
|
||||
+ (unsigned long long) proc.vm_hwm);
|
||||
+
|
||||
+ llama_model_free(model);
|
||||
+ llama_backend_free();
|
||||
+ return 0;
|
||||
+}
|
||||
@@ -4,3 +4,4 @@
|
||||
4871a37544df658980a01b4f94151a90b609fb144c931b4a814309ee608ebb46 0003-owned-range-filtered-state-report.patch
|
||||
19d451ce259150ffede793c4eb547425375c0fcd97caf326b43e8f1a204f05b6 0004-dense-boundary-io-endpoint-guard.patch
|
||||
cf263357a6a8de193f710836c7c467c38cac7099975303ee2628e0609daf5a47 0005-worker-range-report-hook.patch
|
||||
23b4b8c56243d52ba682f0034022a86bf8ded007885be5b659cf5158ff3eb429 0006-meshnet-range-report-tool.patch
|
||||
|
||||
@@ -112,6 +112,30 @@
|
||||
"llama_internal_get_tensor_map(const llama_model *) in src/llama-model.h",
|
||||
"gguf empty-context writer API: gguf_init_empty, gguf_add_tensor, gguf_write_to_file"
|
||||
]
|
||||
},
|
||||
"0006-meshnet-range-report-tool.patch": {
|
||||
"concern": "range-reporting",
|
||||
"files": {
|
||||
"CMakeLists.txt": {
|
||||
"before": "a9afcffa68bed7cbd8fad39ad9f95ad784251234",
|
||||
"after": "868793b826f565df7f041e7ba55820b5ad744b10"
|
||||
},
|
||||
"tools/meshnet-range-report/CMakeLists.txt": {
|
||||
"before": null,
|
||||
"after": "24401007ee85e217c2741a42c7119fad323ff08a"
|
||||
},
|
||||
"tools/meshnet-range-report/meshnet-range-report.cpp": {
|
||||
"before": null,
|
||||
"after": "49a5eb2a05bf6514e166453ea0e35b8bc9c5fdf6"
|
||||
}
|
||||
},
|
||||
"api_assumptions": [
|
||||
"llama_model_params carries meshnet_owned_layer_start/end, use_mmap, and use_extra_bufts",
|
||||
"llama_model_meshnet_range_report C API and llama_meshnet_range_report fields (patch 0005)",
|
||||
"llama_internal_get_tensor_map(const llama_model *) in src/llama-model.h",
|
||||
"llama_model_meta_val_str and llama_model_n_layer public accessors",
|
||||
"top-level CMakeLists add_subdirectory of a project-owned tool directory after the llama target"
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3,3 +3,4 @@
|
||||
0003-owned-range-filtered-state-report.patch
|
||||
0004-dense-boundary-io-endpoint-guard.patch
|
||||
0005-worker-range-report-hook.patch
|
||||
0006-meshnet-range-report-tool.patch
|
||||
|
||||
165
packages/node/native/worker/fake_engine.h
Normal file
165
packages/node/native/worker/fake_engine.h
Normal file
@@ -0,0 +1,165 @@
|
||||
// Deterministic, model-free fake ShardEngine for the native worker (DGR-033).
|
||||
//
|
||||
// This is the C++ analogue of `meshnet_node.fake_shard_engine.FakeShardEngine`
|
||||
// (DGR-032): a pure fixture that performs a *bounded real forward* over the
|
||||
// bytes it received off the socket and never links, loads, or dispatches to
|
||||
// llama.cpp. It exists to prove the standalone worker process, stream,
|
||||
// lifecycle, and supervision shape before any real engine is bound (DGR-037).
|
||||
//
|
||||
// The "forward" is deliberately transport-verifiable rather than semantic: it
|
||||
// reassembles a tensor's fragments, checks they tile exactly, and derives a
|
||||
// CRC32C over the uncompressed bytes — the same rule the schema's `Checksum`
|
||||
// declares and the same bounded forward the DGR-024 Python surface performs.
|
||||
// Feeding the same bytes back (echo) lets a client prove the payload truly
|
||||
// traversed the wire and returned unmodified; a direct hop and an opaque relay
|
||||
// of the identical frames therefore yield byte-identical responses.
|
||||
//
|
||||
// There is no arbitrary-graph entry point here and no llama.cpp RPC: the engine
|
||||
// only knows how to reassemble/checksum a bundle. That is the whole point of a
|
||||
// fixture worker (acceptance criterion 4).
|
||||
|
||||
#ifndef MESHNET_NATIVE_WORKER_FAKE_ENGINE_H_
|
||||
#define MESHNET_NATIVE_WORKER_FAKE_ENGINE_H_
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstdint>
|
||||
#include <optional>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
#include "shard_runtime.pb.h"
|
||||
|
||||
namespace meshnet::worker {
|
||||
|
||||
namespace sp = ::meshnet::shard::v1;
|
||||
|
||||
// Standard CRC-32 (ISO-HDLC / zlib polynomial 0xEDB88320, reflected).
|
||||
//
|
||||
// The schema's `Checksum` field is labelled CRC32C, but the DGR-024 Python
|
||||
// runtime surface (`shard_runtime_server.py`) computes it with `zlib.crc32`
|
||||
// (standard CRC-32, not the Castagnoli CRC32C). This worker deliberately mirrors
|
||||
// that exact computation so its checksum acceptance is byte-for-byte identical
|
||||
// to the existing Python gRPC surface and to a relayed frame's expectations.
|
||||
inline uint32_t Crc32(const std::string& data, uint32_t seed = 0) {
|
||||
static uint32_t table[256];
|
||||
static bool built = false;
|
||||
if (!built) {
|
||||
for (uint32_t i = 0; i < 256; ++i) {
|
||||
uint32_t c = i;
|
||||
for (int k = 0; k < 8; ++k) {
|
||||
c = (c & 1) ? (c >> 1) ^ 0xEDB88320u : (c >> 1);
|
||||
}
|
||||
table[i] = c;
|
||||
}
|
||||
built = true;
|
||||
}
|
||||
uint32_t crc = seed ^ 0xFFFFFFFFu;
|
||||
for (unsigned char byte : data) {
|
||||
crc = (crc >> 8) ^ table[(crc ^ byte) & 0xFF];
|
||||
}
|
||||
return crc ^ 0xFFFFFFFFu;
|
||||
}
|
||||
|
||||
// Outcome of validating one bundle before the bounded forward runs.
|
||||
struct BundleCheck {
|
||||
// Set when the bundle is malformed/corrupt (maps to PAYLOAD_CORRUPT).
|
||||
std::optional<std::string> corrupt_detail;
|
||||
// Set when the declared payload exceeds the negotiated per-chunk ceiling
|
||||
// (maps to RESOURCE_EXHAUSTED) — the worker refuses unbounded messages.
|
||||
std::optional<std::string> oversize_detail;
|
||||
};
|
||||
|
||||
// The fake engine's only capability: verify a bundle tiles and checksums, and
|
||||
// that it stays within the negotiated byte ceiling. Mirrors `_validate_bundle`
|
||||
// in `shard_runtime_server.py` plus the bounded-message rule DGR-033 adds.
|
||||
class FakeShardEngine {
|
||||
public:
|
||||
// Marker mirroring `FakeShardEngine.EVIDENCE_CLASS` so a future parity check
|
||||
// (DGR-036) can assert this is a fixture, not a real engine.
|
||||
static constexpr const char* kEvidenceClass = "fixture";
|
||||
|
||||
FakeShardEngine() = default;
|
||||
|
||||
// `max_chunk_bytes` is the per-session *negotiated* ceiling (the strictest of
|
||||
// the worker's own limit and the peer's proposal), passed in on every call so
|
||||
// the engine enforces exactly what the SessionOpen handshake settled — never a
|
||||
// value the peer proposed unilaterally.
|
||||
BundleCheck Validate(const sp::TensorBundle& bundle, uint64_t max_chunk_bytes) const {
|
||||
BundleCheck result;
|
||||
for (const auto& tensor : bundle.tensors()) {
|
||||
// Bounded message: a declared payload larger than the ceiling is refused
|
||||
// before any reassembly work is done.
|
||||
if (max_chunk_bytes != 0 && tensor.total_bytes() > max_chunk_bytes) {
|
||||
result.oversize_detail =
|
||||
"tensor '" + tensor.name() + "': declared total_bytes " +
|
||||
std::to_string(tensor.total_bytes()) + " exceeds max_chunk_bytes " +
|
||||
std::to_string(max_chunk_bytes);
|
||||
return result;
|
||||
}
|
||||
|
||||
// Fragments must tile the wire body exactly: no hole, no overlap.
|
||||
std::vector<const sp::TensorFragment*> ordered;
|
||||
ordered.reserve(tensor.fragments_size());
|
||||
for (const auto& fragment : tensor.fragments()) {
|
||||
ordered.push_back(&fragment);
|
||||
}
|
||||
std::sort(ordered.begin(), ordered.end(),
|
||||
[](const sp::TensorFragment* a, const sp::TensorFragment* b) {
|
||||
return a->byte_offset() < b->byte_offset();
|
||||
});
|
||||
uint64_t expected_offset = 0;
|
||||
std::string payload;
|
||||
for (const auto* fragment : ordered) {
|
||||
if (fragment->byte_offset() != expected_offset) {
|
||||
result.corrupt_detail =
|
||||
"tensor '" + tensor.name() + "': fragment at offset " +
|
||||
std::to_string(fragment->byte_offset()) +
|
||||
" does not tile the preceding " + std::to_string(expected_offset) +
|
||||
" bytes (gap or overlap)";
|
||||
return result;
|
||||
}
|
||||
payload.append(fragment->payload());
|
||||
expected_offset += fragment->payload().size();
|
||||
}
|
||||
if (tensor.compression() == sp::COMPRESSION_NONE &&
|
||||
expected_offset != tensor.total_bytes()) {
|
||||
result.corrupt_detail =
|
||||
"tensor '" + tensor.name() + "': fragments cover " +
|
||||
std::to_string(expected_offset) + " bytes, declared total_bytes is " +
|
||||
std::to_string(tensor.total_bytes());
|
||||
return result;
|
||||
}
|
||||
if (tensor.compression() == sp::COMPRESSION_NONE &&
|
||||
tensor.checksum().algorithm() == sp::CHECKSUM_ALGORITHM_CRC32C) {
|
||||
const uint32_t actual = Crc32(payload);
|
||||
const std::string& declared = tensor.checksum().value();
|
||||
std::string actual_be(4, '\0');
|
||||
actual_be[0] = static_cast<char>((actual >> 24) & 0xFF);
|
||||
actual_be[1] = static_cast<char>((actual >> 16) & 0xFF);
|
||||
actual_be[2] = static_cast<char>((actual >> 8) & 0xFF);
|
||||
actual_be[3] = static_cast<char>(actual & 0xFF);
|
||||
if (declared != actual_be) {
|
||||
result.corrupt_detail = "tensor '" + tensor.name() + "': checksum mismatch";
|
||||
return result;
|
||||
}
|
||||
}
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
// Bounded real forward: fold every fragment's payload through CRC32C so the
|
||||
// digest is only reproducible if the payload really traversed the wire.
|
||||
uint32_t BoundedForward(const sp::TensorBundle& bundle) const {
|
||||
uint32_t digest = 0;
|
||||
for (const auto& tensor : bundle.tensors()) {
|
||||
for (const auto& fragment : tensor.fragments()) {
|
||||
digest = Crc32(fragment.payload(), digest);
|
||||
}
|
||||
}
|
||||
return digest;
|
||||
}
|
||||
};
|
||||
|
||||
} // namespace meshnet::worker
|
||||
|
||||
#endif // MESHNET_NATIVE_WORKER_FAKE_ENGINE_H_
|
||||
470
packages/node/native/worker/shard_service.cpp
Normal file
470
packages/node/native/worker/shard_service.cpp
Normal file
@@ -0,0 +1,470 @@
|
||||
#include "shard_service.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <utility>
|
||||
|
||||
namespace meshnet::worker {
|
||||
|
||||
namespace {
|
||||
|
||||
// The exact identity this fixture worker serves. SessionOpen is validated
|
||||
// against these — not echoed back from the caller — so an incompatible peer
|
||||
// fails closed at open rather than being silently accepted with its own claimed
|
||||
// identity. Kept in one place so GetCapability and the open handshake agree.
|
||||
constexpr const char* kModelArtifactDigest = "sha256:native-test-artifact";
|
||||
constexpr const char* kRuntimeRecipeDigest = "sha256:native-test-recipe";
|
||||
constexpr const char* kRecipeId = "native-test";
|
||||
constexpr const char* kRecipeVersion = "1";
|
||||
constexpr const char* kCatalogueVersion = "1";
|
||||
constexpr uint32_t kShardStartLayer = 0;
|
||||
constexpr uint32_t kShardEndLayer = 32;
|
||||
constexpr uint32_t kShardEffectiveStartLayer = 0;
|
||||
|
||||
void FillWorkerFingerprint(sp::Fingerprint* fp) {
|
||||
fp->set_model_artifact_digest(kModelArtifactDigest);
|
||||
fp->set_runtime_recipe_digest(kRuntimeRecipeDigest);
|
||||
fp->set_recipe_id(kRecipeId);
|
||||
fp->set_recipe_version(kRecipeVersion);
|
||||
fp->set_catalogue_version(kCatalogueVersion);
|
||||
}
|
||||
|
||||
void FillWorkerShardRange(sp::ShardRange* range) {
|
||||
range->set_start_layer(kShardStartLayer);
|
||||
range->set_end_layer(kShardEndLayer);
|
||||
range->set_effective_start_layer(kShardEffectiveStartLayer);
|
||||
}
|
||||
|
||||
// Strictest-of-both bound: the smallest positive of `a`/`b`, or `fallback` when
|
||||
// neither is set. Mirrors the `_min` helper in `native_protocol/codec.py`.
|
||||
uint64_t MinPositive(uint64_t a, uint64_t b, uint64_t fallback) {
|
||||
if (a > 0 && b > 0) return std::min(a, b);
|
||||
if (a > 0) return a;
|
||||
if (b > 0) return b;
|
||||
return fallback;
|
||||
}
|
||||
|
||||
int64_t NowUnixNanos() {
|
||||
return std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||
std::chrono::system_clock::now().time_since_epoch())
|
||||
.count();
|
||||
}
|
||||
|
||||
// Build the standard fail response (a terminal-or-not ShardStatus).
|
||||
sp::SessionResponse MakeFail(const std::string& route_session_id, const std::string& work_id,
|
||||
uint64_t step, sp::ErrorCode code, const std::string& detail,
|
||||
bool terminal, bool retryable) {
|
||||
sp::SessionResponse response;
|
||||
sp::ShardStatus* status = response.mutable_status();
|
||||
status->set_work_id(work_id);
|
||||
status->set_route_session_id(route_session_id);
|
||||
status->set_idempotency_step(step);
|
||||
status->set_terminal(terminal);
|
||||
sp::ShardError* error = status->mutable_error();
|
||||
error->set_code(code);
|
||||
error->set_detail(detail);
|
||||
error->set_retryable(retryable);
|
||||
return response;
|
||||
}
|
||||
|
||||
sp::SessionResponse MakeAck(const std::string& work_id, uint64_t step, bool duplicate) {
|
||||
sp::SessionResponse response;
|
||||
sp::Ack* ack = response.mutable_ack();
|
||||
ack->set_work_id(work_id);
|
||||
ack->set_idempotency_step(step);
|
||||
ack->set_duplicate(duplicate);
|
||||
return response;
|
||||
}
|
||||
|
||||
void FillDefaultFlow(sp::FlowControl* fc, const FlowLimits& limits) {
|
||||
fc->set_credits_granted(limits.credits_granted);
|
||||
fc->set_max_inflight_chunks(limits.max_inflight_chunks);
|
||||
fc->set_max_chunk_bytes(limits.max_chunk_bytes);
|
||||
fc->set_max_prefill_chunk_tokens(limits.max_prefill_chunk_tokens);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
grpc::Status ShardRuntimeServiceImpl::GetCapability(grpc::ServerContext*,
|
||||
const sp::CapabilityRequest*,
|
||||
sp::CapabilityReport* response) {
|
||||
response->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
FillWorkerFingerprint(response->mutable_fingerprint());
|
||||
FillWorkerShardRange(response->mutable_shard_range());
|
||||
response->set_backend("grpc-native-cpp");
|
||||
response->set_device("cpu");
|
||||
response->set_validated(true);
|
||||
response->set_detail("bounded real forward passed for fixture artifact");
|
||||
response->set_max_concurrent_sessions(8);
|
||||
response->set_max_context_tokens(131072);
|
||||
FillDefaultFlow(response->mutable_flow_control(), limits_);
|
||||
response->add_accepted_compression(sp::COMPRESSION_NONE);
|
||||
response->add_supported_schema_versions(sp::SCHEMA_VERSION_1);
|
||||
response->set_validated_at_unix_nanos(0);
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
|
||||
grpc::Status ShardRuntimeServiceImpl::Health(grpc::ServerContext*, const sp::HealthRequest*,
|
||||
sp::HealthReport* response) {
|
||||
response->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
response->set_state(sp::SERVING_STATE_SERVING);
|
||||
response->set_active_sessions(1);
|
||||
response->set_queued_chunks(0);
|
||||
response->set_batch_occupancy(0);
|
||||
response->set_kv_pressure(0.0f);
|
||||
response->set_resident_bytes(0);
|
||||
response->set_detail("native fixture worker serving");
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
|
||||
FlowLimits ShardRuntimeServiceImpl::NegotiateFlow(const sp::FlowControl& proposed) const {
|
||||
FlowLimits out;
|
||||
out.max_inflight_chunks = static_cast<uint32_t>(MinPositive(
|
||||
proposed.max_inflight_chunks(), limits_.max_inflight_chunks, limits_.max_inflight_chunks));
|
||||
const uint64_t credits = MinPositive(proposed.credits_granted(), limits_.credits_granted,
|
||||
limits_.credits_granted);
|
||||
out.credits_granted =
|
||||
static_cast<uint32_t>(std::min<uint64_t>(credits, out.max_inflight_chunks));
|
||||
out.max_chunk_bytes =
|
||||
MinPositive(proposed.max_chunk_bytes(), limits_.max_chunk_bytes, limits_.max_chunk_bytes);
|
||||
out.max_prefill_chunk_tokens = static_cast<uint32_t>(MinPositive(
|
||||
proposed.max_prefill_chunk_tokens(), limits_.max_prefill_chunk_tokens,
|
||||
limits_.max_prefill_chunk_tokens));
|
||||
return out;
|
||||
}
|
||||
|
||||
uint32_t ShardRuntimeServiceImpl::MarkCancelled(const std::string& route_session_id,
|
||||
const std::string& work_id) {
|
||||
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||
SessionState& state = sessions_[route_session_id]; // creates on first cancel-before-open
|
||||
if (state.max_inflight == 0) {
|
||||
// Freshly created placeholder for a Cancel that raced ahead of Open.
|
||||
state.credits = limits_.credits_granted;
|
||||
state.max_inflight = limits_.max_inflight_chunks;
|
||||
state.max_chunk_bytes = limits_.max_chunk_bytes;
|
||||
}
|
||||
if (work_id.empty()) {
|
||||
const bool already = state.cancelled_session;
|
||||
state.cancelled_session = true;
|
||||
return already ? 0 : 1;
|
||||
}
|
||||
const bool already = state.cancelled_work.count(work_id) != 0;
|
||||
state.cancelled_work.insert(work_id);
|
||||
return already ? 0 : 1;
|
||||
}
|
||||
|
||||
grpc::Status ShardRuntimeServiceImpl::Session(
|
||||
grpc::ServerContext*,
|
||||
grpc::ServerReaderWriter<sp::SessionResponse, sp::SessionRequest>* stream) {
|
||||
std::string route_session_id;
|
||||
sp::SessionRequest request;
|
||||
|
||||
while (stream->Read(&request)) {
|
||||
switch (request.kind_case()) {
|
||||
case sp::SessionRequest::kOpen: {
|
||||
const sp::SessionOpen& open = request.open();
|
||||
route_session_id = open.route_session_id();
|
||||
|
||||
// Reject an incompatible peer at open rather than mid-generation. The
|
||||
// worker validates the caller's schema, artifact/recipe identity and
|
||||
// requested layer range against its own — it never adopts the caller's
|
||||
// claimed identity.
|
||||
auto reject_open = [&](sp::ErrorCode code, const std::string& detail) {
|
||||
stream->Write(MakeFail(route_session_id, /*work_id=*/"", /*step=*/0, code, detail,
|
||||
/*terminal=*/true, /*retryable=*/false));
|
||||
};
|
||||
if (open.schema_version() != sp::SCHEMA_VERSION_1) {
|
||||
reject_open(sp::ERROR_CODE_SCHEMA_UNSUPPORTED,
|
||||
"worker serves schema version 1 only");
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
const sp::Fingerprint& fp = open.fingerprint();
|
||||
if ((!fp.model_artifact_digest().empty() &&
|
||||
fp.model_artifact_digest() != kModelArtifactDigest) ||
|
||||
(!fp.runtime_recipe_digest().empty() &&
|
||||
fp.runtime_recipe_digest() != kRuntimeRecipeDigest)) {
|
||||
reject_open(sp::ERROR_CODE_FINGERPRINT_MISMATCH,
|
||||
"model artifact or runtime recipe digest does not match this worker");
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
if (open.has_shard_range()) {
|
||||
const sp::ShardRange& r = open.shard_range();
|
||||
const bool within = r.start_layer() >= kShardStartLayer &&
|
||||
r.end_layer() <= kShardEndLayer &&
|
||||
r.start_layer() < r.end_layer() &&
|
||||
r.effective_start_layer() >= r.start_layer() &&
|
||||
r.effective_start_layer() < r.end_layer();
|
||||
if (!within) {
|
||||
reject_open(sp::ERROR_CODE_SHARD_RANGE_MISMATCH,
|
||||
"requested layer range is not served by this worker");
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
}
|
||||
|
||||
// Settle the flow-control window with strict worker bounds, then keep
|
||||
// the negotiated ceilings on the session so every later check enforces
|
||||
// exactly what was agreed — not what the peer proposed.
|
||||
const FlowLimits negotiated =
|
||||
open.has_proposed_flow_control()
|
||||
? NegotiateFlow(open.proposed_flow_control())
|
||||
: limits_;
|
||||
{
|
||||
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||
SessionState state;
|
||||
state.epoch = open.route_epoch();
|
||||
state.credits = negotiated.credits_granted;
|
||||
state.max_inflight = negotiated.max_inflight_chunks;
|
||||
state.max_chunk_bytes = negotiated.max_chunk_bytes;
|
||||
state.max_prefill_chunk_tokens = negotiated.max_prefill_chunk_tokens;
|
||||
state.opened = true;
|
||||
auto it = sessions_.find(route_session_id);
|
||||
if (it != sessions_.end()) {
|
||||
// A prior out-of-band Cancel may have marked this session cancelled
|
||||
// before Open arrived; preserve that so the work still fails closed.
|
||||
state.cancelled_session = it->second.cancelled_session;
|
||||
state.cancelled_work = it->second.cancelled_work;
|
||||
}
|
||||
sessions_[route_session_id] = std::move(state);
|
||||
}
|
||||
sp::SessionResponse response;
|
||||
sp::SessionAccepted* accepted = response.mutable_accepted();
|
||||
accepted->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
accepted->set_route_session_id(open.route_session_id());
|
||||
accepted->set_route_epoch(open.route_epoch());
|
||||
FillDefaultFlow(accepted->mutable_flow_control(), negotiated);
|
||||
if (open.accepted_compression_size() > 0) {
|
||||
for (int c : open.accepted_compression()) {
|
||||
accepted->add_accepted_compression(static_cast<sp::Compression>(c));
|
||||
}
|
||||
} else {
|
||||
accepted->add_accepted_compression(sp::COMPRESSION_NONE);
|
||||
}
|
||||
// Report the fingerprint the worker actually serves, so a mismatch is
|
||||
// visible at open — never a copy of the caller's claimed identity.
|
||||
FillWorkerFingerprint(accepted->mutable_fingerprint());
|
||||
stream->Write(response);
|
||||
break;
|
||||
}
|
||||
|
||||
case sp::SessionRequest::kChunk: {
|
||||
const sp::ActivationChunk& chunk = request.chunk();
|
||||
const sp::Envelope& envelope = chunk.envelope();
|
||||
const std::string work_id = envelope.work_id();
|
||||
const uint64_t step = envelope.idempotency_step();
|
||||
|
||||
// Compute the response under the lock, then write it *after* releasing —
|
||||
// holding the lock across a (possibly blocking) Write would deadlock an
|
||||
// out-of-band Cancel RPC that needs the same lock.
|
||||
sp::SessionResponse response;
|
||||
bool terminate = false;
|
||||
{
|
||||
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||
auto it = sessions_.find(route_session_id);
|
||||
SessionState* state = it != sessions_.end() ? &it->second : nullptr;
|
||||
|
||||
if (state == nullptr || !state->opened) {
|
||||
// Fail closed: an activation before a valid SessionOpen must never
|
||||
// bypass lifecycle, cancellation, epoch or flow-control state.
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_INTERNAL,
|
||||
"activation received before SessionOpen", true, false);
|
||||
terminate = true;
|
||||
} else if (state->cancelled_session || state->cancelled_work.count(work_id)) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_CANCELLED,
|
||||
"work was cancelled", false, false);
|
||||
} else if (envelope.route_epoch() < state->epoch) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_EPOCH_STALE,
|
||||
"stale route epoch", false, false);
|
||||
} else if (envelope.deadline_unix_nanos() != 0 &&
|
||||
NowUnixNanos() > envelope.deadline_unix_nanos()) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_DEADLINE_EXCEEDED,
|
||||
"deadline already passed", false, false);
|
||||
} else if (state->seen_steps.count(step)) {
|
||||
response = MakeAck(work_id, step, /*duplicate=*/true);
|
||||
} else if (state->credits <= 0) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_FLOW_CONTROL_VIOLATION,
|
||||
"no flow-control credit remaining", false, true);
|
||||
} else {
|
||||
const BundleCheck check = engine_.Validate(chunk.bundle(), state->max_chunk_bytes);
|
||||
if (check.oversize_detail) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_RESOURCE_EXHAUSTED,
|
||||
*check.oversize_detail, false, false);
|
||||
} else if (check.corrupt_detail) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_PAYLOAD_CORRUPT,
|
||||
*check.corrupt_detail, false, false);
|
||||
} else {
|
||||
state->seen_steps.insert(step);
|
||||
state->credits -= 1;
|
||||
engine_.BoundedForward(chunk.bundle()); // real bounded forward over wire bytes
|
||||
*response.mutable_chunk() = chunk; // echo the exact bundle back
|
||||
}
|
||||
}
|
||||
}
|
||||
stream->Write(response);
|
||||
if (terminate) {
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
break;
|
||||
}
|
||||
|
||||
case sp::SessionRequest::kDecode: {
|
||||
const sp::DecodeStep& step_msg = request.decode();
|
||||
const std::string work_id = step_msg.work_id();
|
||||
const uint64_t step = step_msg.idempotency_step();
|
||||
|
||||
sp::TensorBundle bundle;
|
||||
if (step_msg.bundle().tensors_size() > 0) {
|
||||
bundle = step_msg.bundle();
|
||||
} else {
|
||||
bundle.set_bundle_version(1);
|
||||
*bundle.add_tensors() = step_msg.tensor();
|
||||
}
|
||||
|
||||
sp::SessionResponse response;
|
||||
bool terminate = false;
|
||||
{
|
||||
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||
auto it = sessions_.find(route_session_id);
|
||||
SessionState* state = it != sessions_.end() ? &it->second : nullptr;
|
||||
|
||||
if (state == nullptr || !state->opened) {
|
||||
// Fail closed: a decode step before a valid SessionOpen must never
|
||||
// bypass lifecycle, cancellation, epoch or flow-control state.
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_INTERNAL,
|
||||
"activation received before SessionOpen", true, false);
|
||||
terminate = true;
|
||||
} else if (state->cancelled_session || state->cancelled_work.count(work_id)) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_CANCELLED,
|
||||
"work was cancelled", false, false);
|
||||
} else if (step_msg.deadline_unix_nanos() != 0 &&
|
||||
NowUnixNanos() > step_msg.deadline_unix_nanos()) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_DEADLINE_EXCEEDED,
|
||||
"deadline already passed", false, false);
|
||||
} else if (state->seen_steps.count(step)) {
|
||||
response = MakeAck(work_id, step, /*duplicate=*/true);
|
||||
} else if (state->credits <= 0) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_FLOW_CONTROL_VIOLATION,
|
||||
"no flow-control credit remaining", false, true);
|
||||
} else {
|
||||
const BundleCheck check = engine_.Validate(bundle, state->max_chunk_bytes);
|
||||
if (check.oversize_detail) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_RESOURCE_EXHAUSTED,
|
||||
*check.oversize_detail, false, false);
|
||||
} else if (check.corrupt_detail) {
|
||||
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_PAYLOAD_CORRUPT,
|
||||
*check.corrupt_detail, false, false);
|
||||
} else {
|
||||
state->seen_steps.insert(step);
|
||||
state->credits -= 1;
|
||||
engine_.BoundedForward(bundle);
|
||||
// No decode response field exists; echo the step back as a
|
||||
// chunk-bearing SessionResponse per the proto's relayed-frame design.
|
||||
sp::ActivationChunk* out = response.mutable_chunk();
|
||||
sp::Envelope* out_env = out->mutable_envelope();
|
||||
out_env->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
out_env->set_work_id(work_id);
|
||||
out_env->set_idempotency_step(step);
|
||||
out_env->set_phase(sp::PHASE_DECODE);
|
||||
sp::PositionSpan* pos = out_env->mutable_position();
|
||||
pos->set_first_position(step_msg.position());
|
||||
pos->set_token_count(1);
|
||||
*out->mutable_bundle() = bundle;
|
||||
}
|
||||
}
|
||||
}
|
||||
stream->Write(response);
|
||||
if (terminate) {
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
break;
|
||||
}
|
||||
|
||||
case sp::SessionRequest::kFlowControl: {
|
||||
const uint32_t topup = request.flow_control().credits_granted();
|
||||
sp::SessionResponse response;
|
||||
{
|
||||
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||
auto it = sessions_.find(route_session_id);
|
||||
sp::FlowControl* fc = response.mutable_flow_control();
|
||||
if (it != sessions_.end()) {
|
||||
SessionState& state = it->second;
|
||||
int64_t granted = std::min<int64_t>(state.credits + topup,
|
||||
static_cast<int64_t>(state.max_inflight));
|
||||
state.credits = granted;
|
||||
fc->set_credits_granted(static_cast<uint32_t>(granted));
|
||||
fc->set_max_inflight_chunks(state.max_inflight);
|
||||
fc->set_max_chunk_bytes(state.max_chunk_bytes);
|
||||
} else {
|
||||
fc->set_credits_granted(topup != 0 ? topup : limits_.credits_granted);
|
||||
fc->set_max_inflight_chunks(limits_.max_inflight_chunks);
|
||||
fc->set_max_chunk_bytes(limits_.max_chunk_bytes);
|
||||
}
|
||||
fc->set_max_prefill_chunk_tokens(limits_.max_prefill_chunk_tokens);
|
||||
}
|
||||
stream->Write(response);
|
||||
break;
|
||||
}
|
||||
|
||||
case sp::SessionRequest::kRelease: {
|
||||
const sp::ReleaseSignal& release = request.release();
|
||||
// An explicit release drops session state immediately (KV, credits,
|
||||
// dedup) instead of holding it for the TTL — the whole point of the
|
||||
// signal. Erase the session this stream opened so its resources are
|
||||
// freed the moment the terminal status is sent.
|
||||
{
|
||||
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||
sessions_.erase(route_session_id);
|
||||
}
|
||||
sp::SessionResponse response;
|
||||
sp::ShardStatus* status = response.mutable_status();
|
||||
status->set_work_id(release.work_id());
|
||||
status->set_route_session_id(release.route_session_id());
|
||||
status->set_terminal(true);
|
||||
stream->Write(response);
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
|
||||
case sp::SessionRequest::kCancel: {
|
||||
const sp::CancelSignal& signal = request.cancel();
|
||||
MarkCancelled(route_session_id, signal.work_id());
|
||||
const bool whole_session = signal.work_id().empty();
|
||||
stream->Write(MakeFail(route_session_id, signal.work_id(), 0, sp::ERROR_CODE_CANCELLED,
|
||||
signal.reason().empty() ? "cancelled" : signal.reason(),
|
||||
whole_session, false));
|
||||
if (whole_session) {
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
break;
|
||||
}
|
||||
|
||||
default: {
|
||||
sp::SessionResponse response;
|
||||
response.mutable_status()->set_terminal(true);
|
||||
stream->Write(response);
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
}
|
||||
}
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
|
||||
grpc::Status ShardRuntimeServiceImpl::Release(grpc::ServerContext*,
|
||||
const sp::ReleaseRequest* request,
|
||||
sp::ReleaseResponse* response) {
|
||||
bool existed;
|
||||
{
|
||||
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||
existed = sessions_.erase(request->route_session_id()) != 0;
|
||||
}
|
||||
response->set_released(existed);
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
|
||||
grpc::Status ShardRuntimeServiceImpl::Cancel(grpc::ServerContext*,
|
||||
const sp::CancelRequest* request,
|
||||
sp::CancelResponse* response) {
|
||||
const uint32_t newly = MarkCancelled(request->route_session_id(), request->work_id());
|
||||
response->set_cancelled_work_items(newly);
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
|
||||
} // namespace meshnet::worker
|
||||
96
packages/node/native/worker/shard_service.h
Normal file
96
packages/node/native/worker/shard_service.h
Normal file
@@ -0,0 +1,96 @@
|
||||
// The native Shard worker's ShardRuntime service (DGR-033).
|
||||
//
|
||||
// A faithful C++ port of `ShardRuntimeServicer` in `shard_runtime_server.py`:
|
||||
// the same per-`route_session_id` identity/credit/dedup state, the same
|
||||
// fail-closed negative paths (stale epoch, expired deadline, corrupt/oversize
|
||||
// payload, exhausted flow-control credit, duplicate idempotency step, in-band
|
||||
// and out-of-band cancellation), and the same lifecycle (open/prefill/decode/
|
||||
// flow-control/release/cancel). The only compute it does is the fake engine's
|
||||
// bounded forward — there is no llama.cpp linkage and no arbitrary-graph RPC.
|
||||
|
||||
#ifndef MESHNET_NATIVE_WORKER_SHARD_SERVICE_H_
|
||||
#define MESHNET_NATIVE_WORKER_SHARD_SERVICE_H_
|
||||
|
||||
#include <cstdint>
|
||||
#include <map>
|
||||
#include <mutex>
|
||||
#include <set>
|
||||
#include <string>
|
||||
|
||||
#include <grpcpp/grpcpp.h>
|
||||
|
||||
#include "fake_engine.h"
|
||||
#include "shard_runtime.grpc.pb.h"
|
||||
#include "shard_runtime.pb.h"
|
||||
|
||||
namespace meshnet::worker {
|
||||
|
||||
namespace sp = ::meshnet::shard::v1;
|
||||
|
||||
struct FlowLimits {
|
||||
uint32_t credits_granted = 16;
|
||||
uint32_t max_inflight_chunks = 16;
|
||||
uint64_t max_chunk_bytes = 4u * 1024u * 1024u;
|
||||
uint32_t max_prefill_chunk_tokens = 512;
|
||||
};
|
||||
|
||||
// Per-route-session identity/credit/dedup state, kept on the servicer instance
|
||||
// (guarded by a lock) so an out-of-band unary Cancel from a different handler
|
||||
// thread can reach a session a concurrent Session stream is still iterating.
|
||||
struct SessionState {
|
||||
uint64_t epoch = 0;
|
||||
int64_t credits = 0;
|
||||
uint32_t max_inflight = 0;
|
||||
uint64_t max_chunk_bytes = 0;
|
||||
uint32_t max_prefill_chunk_tokens = 0;
|
||||
std::set<uint64_t> seen_steps;
|
||||
std::set<std::string> cancelled_work;
|
||||
bool cancelled_session = false;
|
||||
// True only after a valid SessionOpen handshake completed for this
|
||||
// route_session_id. An activation (chunk/decode) that arrives while this is
|
||||
// false fails closed: no work may bypass the lifecycle handshake, even when a
|
||||
// placeholder state already exists from an out-of-band Cancel that raced Open.
|
||||
bool opened = false;
|
||||
};
|
||||
|
||||
class ShardRuntimeServiceImpl final : public sp::ShardRuntime::Service {
|
||||
public:
|
||||
explicit ShardRuntimeServiceImpl(FlowLimits limits) : limits_(limits) {}
|
||||
|
||||
grpc::Status GetCapability(grpc::ServerContext* context,
|
||||
const sp::CapabilityRequest* request,
|
||||
sp::CapabilityReport* response) override;
|
||||
|
||||
grpc::Status Health(grpc::ServerContext* context, const sp::HealthRequest* request,
|
||||
sp::HealthReport* response) override;
|
||||
|
||||
grpc::Status Session(
|
||||
grpc::ServerContext* context,
|
||||
grpc::ServerReaderWriter<sp::SessionResponse, sp::SessionRequest>* stream) override;
|
||||
|
||||
grpc::Status Release(grpc::ServerContext* context, const sp::ReleaseRequest* request,
|
||||
sp::ReleaseResponse* response) override;
|
||||
|
||||
grpc::Status Cancel(grpc::ServerContext* context, const sp::CancelRequest* request,
|
||||
sp::CancelResponse* response) override;
|
||||
|
||||
private:
|
||||
// Returns the number of items newly marked cancelled, creating session state
|
||||
// if the Cancel raced ahead of SessionOpen.
|
||||
uint32_t MarkCancelled(const std::string& route_session_id, const std::string& work_id);
|
||||
|
||||
// Settle a stream's flow-control window against this worker's own limits: the
|
||||
// strictest bound of either peer wins for every field, so a peer can never
|
||||
// raise the worker's ceilings by proposing a larger window. Mirrors
|
||||
// `negotiate_flow_control` in `native_protocol/codec.py`.
|
||||
FlowLimits NegotiateFlow(const sp::FlowControl& proposed) const;
|
||||
|
||||
FlowLimits limits_;
|
||||
FakeShardEngine engine_;
|
||||
std::mutex sessions_mu_;
|
||||
std::map<std::string, SessionState> sessions_;
|
||||
};
|
||||
|
||||
} // namespace meshnet::worker
|
||||
|
||||
#endif // MESHNET_NATIVE_WORKER_SHARD_SERVICE_H_
|
||||
302
packages/node/native/worker/shard_worker_main.cpp
Normal file
302
packages/node/native/worker/shard_worker_main.cpp
Normal file
@@ -0,0 +1,302 @@
|
||||
// Standalone native Shard worker executable (DGR-033).
|
||||
//
|
||||
// Serves the complete ShardRuntime lifecycle/stream contract over real
|
||||
// gRPC/HTTP2 using the model-free FakeShardEngine. It links neither llama.cpp
|
||||
// nor any graph-execution entry point: the only surface it exposes is the
|
||||
// ShardRuntime service defined in shard_runtime.proto.
|
||||
//
|
||||
// Usage:
|
||||
// shard_worker [listen_addr] serve until SIGTERM/SIGINT (graceful drain)
|
||||
// shard_worker --selftest bind an ephemeral port, self-drive the
|
||||
// lifecycle over a real loopback channel, exit
|
||||
//
|
||||
// Environment:
|
||||
// MESHNET_SHARD_LISTEN_ADDR host:port to bind (default localhost:50051)
|
||||
// MESHNET_MAX_CHUNK_BYTES per-chunk byte ceiling the worker enforces
|
||||
//
|
||||
// On a normal run it prints one readiness line — "ShardRuntime worker listening
|
||||
// on <addr>" — once the socket is bound, so a supervisor/harness has a real
|
||||
// readiness signal instead of a sleep.
|
||||
|
||||
#include <atomic>
|
||||
#include <cerrno>
|
||||
#include <csignal>
|
||||
#include <cstdint>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <iostream>
|
||||
#include <memory>
|
||||
#include <string>
|
||||
#include <thread>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <grpcpp/grpcpp.h>
|
||||
|
||||
#include "shard_service.h"
|
||||
#include "shard_runtime.grpc.pb.h"
|
||||
|
||||
namespace {
|
||||
|
||||
namespace sp = ::meshnet::shard::v1;
|
||||
|
||||
// Self-pipe: the signal handler must stay async-signal-safe, so it only writes
|
||||
// one byte; a helper thread reads it and performs the (non-signal-safe) server
|
||||
// Shutdown(). Set once in main() before installing the handler.
|
||||
volatile std::sig_atomic_t g_signal_pipe_write_fd = -1;
|
||||
|
||||
extern "C" void HandleTermination(int /*signum*/) {
|
||||
if (g_signal_pipe_write_fd >= 0) {
|
||||
const char byte = 1;
|
||||
ssize_t rc = ::write(g_signal_pipe_write_fd, &byte, 1);
|
||||
(void)rc; // best-effort; nothing safe to do on failure inside a handler
|
||||
}
|
||||
}
|
||||
|
||||
meshnet::worker::FlowLimits LimitsFromEnv() {
|
||||
meshnet::worker::FlowLimits limits;
|
||||
if (const char* raw = std::getenv("MESHNET_MAX_CHUNK_BYTES")) {
|
||||
char* end = nullptr;
|
||||
const unsigned long long value = std::strtoull(raw, &end, 10);
|
||||
if (end != raw && value > 0) {
|
||||
limits.max_chunk_bytes = static_cast<uint64_t>(value);
|
||||
}
|
||||
}
|
||||
return limits;
|
||||
}
|
||||
|
||||
int RunSelfTest() {
|
||||
meshnet::worker::ShardRuntimeServiceImpl service(LimitsFromEnv());
|
||||
int selected_port = 0;
|
||||
grpc::ServerBuilder builder;
|
||||
builder.AddListeningPort("127.0.0.1:0", grpc::InsecureServerCredentials(), &selected_port);
|
||||
builder.RegisterService(&service);
|
||||
std::unique_ptr<grpc::Server> server(builder.BuildAndStart());
|
||||
if (!server || selected_port == 0) {
|
||||
std::cerr << "selftest: failed to bind ephemeral port\n";
|
||||
return 1;
|
||||
}
|
||||
const std::string target = "127.0.0.1:" + std::to_string(selected_port);
|
||||
auto channel = grpc::CreateChannel(target, grpc::InsecureChannelCredentials());
|
||||
auto stub = sp::ShardRuntime::NewStub(channel);
|
||||
|
||||
int failures = 0;
|
||||
auto check = [&](bool cond, const char* what) {
|
||||
if (!cond) {
|
||||
std::cerr << "selftest FAIL: " << what << "\n";
|
||||
++failures;
|
||||
}
|
||||
};
|
||||
|
||||
// Capability + health.
|
||||
{
|
||||
grpc::ClientContext ctx;
|
||||
sp::CapabilityRequest req;
|
||||
req.set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
sp::CapabilityReport rep;
|
||||
grpc::Status status = stub->GetCapability(&ctx, req, &rep);
|
||||
check(status.ok(), "GetCapability RPC");
|
||||
check(rep.validated(), "capability validated");
|
||||
check(rep.schema_version() == sp::SCHEMA_VERSION_1, "capability schema version");
|
||||
}
|
||||
{
|
||||
grpc::ClientContext ctx;
|
||||
sp::HealthRequest req;
|
||||
req.set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
sp::HealthReport rep;
|
||||
grpc::Status status = stub->Health(&ctx, req, &rep);
|
||||
check(status.ok(), "Health RPC");
|
||||
check(rep.state() == sp::SERVING_STATE_SERVING, "health serving");
|
||||
}
|
||||
|
||||
// A minimal session: open -> fragmented prefill -> decode -> release.
|
||||
{
|
||||
grpc::ClientContext ctx;
|
||||
auto stream = stub->Session(&ctx);
|
||||
|
||||
sp::SessionRequest open;
|
||||
sp::SessionOpen* o = open.mutable_open();
|
||||
o->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
o->set_route_session_id("selftest");
|
||||
o->set_route_epoch(1);
|
||||
sp::FlowControl* fc = o->mutable_proposed_flow_control();
|
||||
fc->set_credits_granted(16);
|
||||
fc->set_max_inflight_chunks(16);
|
||||
fc->set_max_chunk_bytes(4u * 1024u * 1024u);
|
||||
check(stream->Write(open), "write open");
|
||||
|
||||
sp::SessionResponse accepted;
|
||||
check(stream->Read(&accepted), "read accepted");
|
||||
check(accepted.kind_case() == sp::SessionResponse::kAccepted, "accepted kind");
|
||||
|
||||
// Fragmented prefill: two fragments tiling a 6-byte payload.
|
||||
const std::string payload = "ABCDEF";
|
||||
sp::SessionRequest chunk;
|
||||
sp::ActivationChunk* ac = chunk.mutable_chunk();
|
||||
sp::Envelope* env = ac->mutable_envelope();
|
||||
env->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||
env->set_work_id("w1");
|
||||
env->set_route_session_id("selftest");
|
||||
env->set_route_epoch(1);
|
||||
env->set_idempotency_step(1);
|
||||
env->set_phase(sp::PHASE_PREFILL);
|
||||
sp::TensorBundle* bundle = ac->mutable_bundle();
|
||||
bundle->set_bundle_version(1);
|
||||
sp::NamedTensor* tensor = bundle->add_tensors();
|
||||
tensor->set_name("hidden_states");
|
||||
tensor->set_dtype(sp::DTYPE_BFLOAT16);
|
||||
tensor->set_byte_order(sp::BYTE_ORDER_LITTLE_ENDIAN);
|
||||
tensor->set_total_bytes(payload.size());
|
||||
tensor->set_compression(sp::COMPRESSION_NONE);
|
||||
sp::Checksum* cksum = tensor->mutable_checksum();
|
||||
cksum->set_algorithm(sp::CHECKSUM_ALGORITHM_CRC32C);
|
||||
const uint32_t crc = meshnet::worker::Crc32(payload);
|
||||
std::string crc_be(4, '\0');
|
||||
crc_be[0] = static_cast<char>((crc >> 24) & 0xFF);
|
||||
crc_be[1] = static_cast<char>((crc >> 16) & 0xFF);
|
||||
crc_be[2] = static_cast<char>((crc >> 8) & 0xFF);
|
||||
crc_be[3] = static_cast<char>(crc & 0xFF);
|
||||
cksum->set_value(crc_be);
|
||||
sp::TensorFragment* f0 = tensor->add_fragments();
|
||||
f0->set_fragment_index(0);
|
||||
f0->set_fragment_count(2);
|
||||
f0->set_byte_offset(0);
|
||||
f0->set_payload(payload.substr(0, 3));
|
||||
sp::TensorFragment* f1 = tensor->add_fragments();
|
||||
f1->set_fragment_index(1);
|
||||
f1->set_fragment_count(2);
|
||||
f1->set_byte_offset(3);
|
||||
f1->set_payload(payload.substr(3));
|
||||
check(stream->Write(chunk), "write chunk");
|
||||
|
||||
sp::SessionResponse echoed;
|
||||
check(stream->Read(&echoed), "read chunk echo");
|
||||
check(echoed.kind_case() == sp::SessionResponse::kChunk, "chunk echo kind");
|
||||
|
||||
sp::SessionRequest decode;
|
||||
sp::DecodeStep* ds = decode.mutable_decode();
|
||||
ds->set_idempotency_step(2);
|
||||
ds->set_position(1);
|
||||
ds->set_work_id("w2");
|
||||
sp::TensorBundle* dbundle = ds->mutable_bundle();
|
||||
dbundle->set_bundle_version(1);
|
||||
sp::NamedTensor* dt = dbundle->add_tensors();
|
||||
dt->set_name("hidden_states");
|
||||
dt->set_dtype(sp::DTYPE_BFLOAT16);
|
||||
dt->set_byte_order(sp::BYTE_ORDER_LITTLE_ENDIAN);
|
||||
dt->set_total_bytes(payload.size());
|
||||
dt->set_compression(sp::COMPRESSION_NONE);
|
||||
sp::Checksum* dck = dt->mutable_checksum();
|
||||
dck->set_algorithm(sp::CHECKSUM_ALGORITHM_CRC32C);
|
||||
dck->set_value(crc_be);
|
||||
sp::TensorFragment* df = dt->add_fragments();
|
||||
df->set_fragment_index(0);
|
||||
df->set_fragment_count(1);
|
||||
df->set_byte_offset(0);
|
||||
df->set_payload(payload);
|
||||
check(stream->Write(decode), "write decode");
|
||||
|
||||
sp::SessionResponse decode_echo;
|
||||
check(stream->Read(&decode_echo), "read decode echo");
|
||||
check(decode_echo.kind_case() == sp::SessionResponse::kChunk, "decode echo kind");
|
||||
|
||||
sp::SessionRequest release;
|
||||
sp::ReleaseSignal* rs = release.mutable_release();
|
||||
rs->set_route_session_id("selftest");
|
||||
rs->set_work_id("w-final");
|
||||
check(stream->Write(release), "write release");
|
||||
stream->WritesDone();
|
||||
|
||||
sp::SessionResponse terminal;
|
||||
check(stream->Read(&terminal), "read terminal");
|
||||
check(terminal.kind_case() == sp::SessionResponse::kStatus && terminal.status().terminal(),
|
||||
"terminal status");
|
||||
|
||||
grpc::Status status = stream->Finish();
|
||||
check(status.ok(), "stream finish");
|
||||
}
|
||||
|
||||
server->Shutdown();
|
||||
server->Wait();
|
||||
|
||||
if (failures == 0) {
|
||||
std::cout << "selftest: all lifecycle checks passed\n";
|
||||
return 0;
|
||||
}
|
||||
std::cerr << "selftest: " << failures << " check(s) failed\n";
|
||||
return 1;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
GOOGLE_PROTOBUF_VERIFY_VERSION;
|
||||
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
if (std::strcmp(argv[i], "--selftest") == 0) {
|
||||
return RunSelfTest();
|
||||
}
|
||||
}
|
||||
|
||||
std::string listen_addr = "localhost:50051";
|
||||
if (const char* env = std::getenv("MESHNET_SHARD_LISTEN_ADDR")) {
|
||||
listen_addr = env;
|
||||
}
|
||||
if (argc > 1 && argv[1][0] != '-') {
|
||||
listen_addr = argv[1];
|
||||
}
|
||||
|
||||
meshnet::worker::FlowLimits limits = LimitsFromEnv();
|
||||
meshnet::worker::ShardRuntimeServiceImpl service(limits);
|
||||
|
||||
grpc::ServerBuilder builder;
|
||||
int selected_port = 0;
|
||||
builder.AddListeningPort(listen_addr, grpc::InsecureServerCredentials(), &selected_port);
|
||||
// Bounded messages, two layers: a hard transport receive ceiling (never below
|
||||
// 4 MiB so the handshake and normal chunks always fit) plus the finer
|
||||
// app-level per-tensor RESOURCE_EXHAUSTED check the service enforces against
|
||||
// the negotiated max_chunk_bytes. Neither path lets an unbounded frame in.
|
||||
constexpr int kTransportFloor = 4 * 1024 * 1024;
|
||||
const int transport_max = limits.max_chunk_bytes > static_cast<uint64_t>(kTransportFloor)
|
||||
? static_cast<int>(limits.max_chunk_bytes)
|
||||
: kTransportFloor;
|
||||
builder.SetMaxReceiveMessageSize(transport_max);
|
||||
builder.RegisterService(&service);
|
||||
std::unique_ptr<grpc::Server> server(builder.BuildAndStart());
|
||||
if (!server || selected_port == 0) {
|
||||
std::cerr << "failed to bind " << listen_addr << "\n";
|
||||
return 1;
|
||||
}
|
||||
|
||||
int pipe_fds[2];
|
||||
if (::pipe(pipe_fds) != 0) {
|
||||
std::cerr << "failed to create shutdown pipe\n";
|
||||
return 1;
|
||||
}
|
||||
g_signal_pipe_write_fd = pipe_fds[1];
|
||||
|
||||
struct sigaction sa;
|
||||
std::memset(&sa, 0, sizeof(sa));
|
||||
sa.sa_handler = HandleTermination;
|
||||
::sigaction(SIGTERM, &sa, nullptr);
|
||||
::sigaction(SIGINT, &sa, nullptr);
|
||||
|
||||
// Drain thread: wakes on the first termination signal and shuts the server
|
||||
// down gracefully so in-flight sessions finish rather than being severed.
|
||||
std::thread drain([&server, read_fd = pipe_fds[0]]() {
|
||||
char byte = 0;
|
||||
ssize_t rc = 0;
|
||||
do {
|
||||
rc = ::read(read_fd, &byte, 1);
|
||||
} while (rc < 0 && errno == EINTR);
|
||||
server->Shutdown();
|
||||
});
|
||||
|
||||
std::cout << "ShardRuntime worker listening on " << listen_addr << std::endl;
|
||||
|
||||
server->Wait();
|
||||
drain.join();
|
||||
::close(pipe_fds[0]);
|
||||
::close(pipe_fds[1]);
|
||||
std::cout << "ShardRuntime worker shut down cleanly" << std::endl;
|
||||
return 0;
|
||||
}
|
||||
275
tests/test_meshnet_range_report_tool.py
Normal file
275
tests/test_meshnet_range_report_tool.py
Normal file
@@ -0,0 +1,275 @@
|
||||
"""DGR-034: end-to-end owned-range loads through the native report tool.
|
||||
|
||||
Gated on the built ``meshnet-range-report`` binary (the deterministic
|
||||
CPU-only native lane builds it from the pinned, patched llama.cpp tree); in
|
||||
an environment without that build these tests skip rather than fake a pass.
|
||||
When the binary is present they run real loads of a tiny synthetic
|
||||
dense-Llama GGUF — no model download, no GPU — and prove the loader
|
||||
registers exactly the owned tensors, reports ownership derived from the
|
||||
loaded state, and rejects invalid/out-of-model ranges and missing required
|
||||
tensors. The JSON is consumed through ``meshnet_node.range_report`` so the
|
||||
strict project-owned contract is exercised on real tool output.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import struct
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from meshnet_node.range_report import RangeReportError, parse_owned_range_report
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_BINARY = REPO_ROOT / "build" / "llama.cpp" / "build" / "bin" / "meshnet-range-report"
|
||||
|
||||
BINARY = Path(os.environ.get("MESHNET_RANGE_REPORT_BIN", DEFAULT_BINARY))
|
||||
|
||||
requires_range_report_tool = pytest.mark.skipif(
|
||||
not BINARY.is_file(),
|
||||
reason=(
|
||||
"meshnet-range-report is not built; run the deterministic native lane "
|
||||
"(scripts/llama_cpp_dependency.py build) to enable these tests"
|
||||
),
|
||||
)
|
||||
|
||||
# --- Minimal GGUF v3 writer, mirroring the model-free native fixture --------
|
||||
|
||||
K_LAYERS = 4
|
||||
K_EMBD = 8
|
||||
K_FFN = 16
|
||||
K_VOCAB = 16
|
||||
ALIGNMENT = 32
|
||||
|
||||
_GGUF_UINT32 = 4
|
||||
_GGUF_FLOAT32 = 6
|
||||
_GGUF_STRING = 8
|
||||
_GGML_TYPE_F32 = 0
|
||||
|
||||
|
||||
def _gguf_string(value: str) -> bytes:
|
||||
data = value.encode("utf-8")
|
||||
return struct.pack("<Q", len(data)) + data
|
||||
|
||||
|
||||
def _metadata_entries() -> list[tuple[str, int, object]]:
|
||||
return [
|
||||
("general.architecture", _GGUF_STRING, "llama"),
|
||||
("general.alignment", _GGUF_UINT32, ALIGNMENT),
|
||||
("llama.context_length", _GGUF_UINT32, 16),
|
||||
("llama.embedding_length", _GGUF_UINT32, K_EMBD),
|
||||
("llama.block_count", _GGUF_UINT32, K_LAYERS),
|
||||
("llama.feed_forward_length", _GGUF_UINT32, K_FFN),
|
||||
("llama.attention.head_count", _GGUF_UINT32, 2),
|
||||
("llama.attention.head_count_kv", _GGUF_UINT32, 2),
|
||||
("llama.rope.dimension_count", _GGUF_UINT32, 4),
|
||||
("llama.attention.layer_norm_rms_epsilon", _GGUF_FLOAT32, 1.0e-5),
|
||||
("tokenizer.ggml.model", _GGUF_STRING, "no_vocab"),
|
||||
("llama.vocab_size", _GGUF_UINT32, K_VOCAB),
|
||||
]
|
||||
|
||||
|
||||
def _fixture_tensors() -> list[tuple[str, tuple[int, ...]]]:
|
||||
tensors: list[tuple[str, tuple[int, ...]]] = [
|
||||
("token_embd.weight", (K_EMBD, K_VOCAB)),
|
||||
("output_norm.weight", (K_EMBD,)),
|
||||
("output.weight", (K_EMBD, K_VOCAB)),
|
||||
]
|
||||
for layer in range(K_LAYERS):
|
||||
prefix = f"blk.{layer}."
|
||||
tensors += [
|
||||
(prefix + "attn_norm.weight", (K_EMBD,)),
|
||||
(prefix + "attn_q.weight", (K_EMBD, K_EMBD)),
|
||||
(prefix + "attn_k.weight", (K_EMBD, K_EMBD)),
|
||||
(prefix + "attn_v.weight", (K_EMBD, K_EMBD)),
|
||||
(prefix + "attn_output.weight", (K_EMBD, K_EMBD)),
|
||||
(prefix + "ffn_norm.weight", (K_EMBD,)),
|
||||
(prefix + "ffn_gate.weight", (K_EMBD, K_FFN)),
|
||||
(prefix + "ffn_down.weight", (K_FFN, K_EMBD)),
|
||||
(prefix + "ffn_up.weight", (K_EMBD, K_FFN)),
|
||||
]
|
||||
return tensors
|
||||
|
||||
|
||||
def write_dense_llama_gguf(path: Path, *, drop: frozenset[str] = frozenset()) -> Path:
|
||||
"""Write a tiny dense-Llama GGUF; ``drop`` omits tensors (corruption cases)."""
|
||||
kvs = _metadata_entries()
|
||||
tensors = [(name, dims) for name, dims in _fixture_tensors() if name not in drop]
|
||||
|
||||
blob = bytearray()
|
||||
blob += b"GGUF" + struct.pack("<IQQ", 3, len(tensors), len(kvs))
|
||||
for key, vtype, value in kvs:
|
||||
blob += _gguf_string(key)
|
||||
blob += struct.pack("<I", vtype)
|
||||
if vtype == _GGUF_STRING:
|
||||
blob += _gguf_string(value) # type: ignore[arg-type]
|
||||
elif vtype == _GGUF_UINT32:
|
||||
blob += struct.pack("<I", value) # type: ignore[arg-type]
|
||||
elif vtype == _GGUF_FLOAT32:
|
||||
blob += struct.pack("<f", value) # type: ignore[arg-type]
|
||||
else: # pragma: no cover - writer guard
|
||||
raise AssertionError(f"unhandled kv type {vtype}")
|
||||
|
||||
offset = 0
|
||||
infos = bytearray()
|
||||
data = bytearray()
|
||||
for name, dims in tensors:
|
||||
infos += _gguf_string(name)
|
||||
infos += struct.pack("<I", len(dims))
|
||||
for dim in dims:
|
||||
infos += struct.pack("<Q", dim)
|
||||
infos += struct.pack("<IQ", _GGML_TYPE_F32, offset)
|
||||
size = 4
|
||||
for dim in dims:
|
||||
size *= dim
|
||||
assert size % ALIGNMENT == 0
|
||||
data += bytes(size)
|
||||
offset += size
|
||||
|
||||
blob += infos
|
||||
blob += bytes(-len(blob) % ALIGNMENT) # pad header to the data section
|
||||
blob += data
|
||||
path.write_bytes(bytes(blob))
|
||||
return path
|
||||
|
||||
|
||||
# --- Tool driver -------------------------------------------------------------
|
||||
|
||||
LAYER_BYTES = 2624 # 9 registered F32 tensors per layer, see _fixture_tensors
|
||||
EMBD_BYTES = 512
|
||||
OUT_NORM_BYTES = 32
|
||||
OUT_BYTES = 512
|
||||
|
||||
|
||||
def run_tool(model: Path, start: int, end: int, *extra: str) -> tuple[int, dict]:
|
||||
env = dict(os.environ)
|
||||
env["LD_LIBRARY_PATH"] = f"{BINARY.parent}:{env.get('LD_LIBRARY_PATH', '')}"
|
||||
completed = subprocess.run(
|
||||
[
|
||||
str(BINARY),
|
||||
"--model", str(model),
|
||||
"--start", str(start),
|
||||
"--end", str(end),
|
||||
*extra,
|
||||
],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
env=env,
|
||||
timeout=120,
|
||||
)
|
||||
try:
|
||||
doc = json.loads(completed.stdout)
|
||||
except json.JSONDecodeError as exc: # pragma: no cover - diagnostic path
|
||||
raise AssertionError(
|
||||
f"tool did not print a JSON report (exit {completed.returncode}): "
|
||||
f"{completed.stdout!r} {completed.stderr!r}"
|
||||
) from exc
|
||||
return completed.returncode, doc
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def dense_llama_gguf(tmp_path_factory: pytest.TempPathFactory) -> Path:
|
||||
return write_dense_llama_gguf(tmp_path_factory.mktemp("gguf") / "dense-llama.gguf")
|
||||
|
||||
|
||||
@requires_range_report_tool
|
||||
class TestOwnedRangeLoads:
|
||||
def test_middle_range_registers_exactly_its_layers(self, dense_llama_gguf: Path) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 1, 3, "--no-extra-bufts")
|
||||
assert code == 0
|
||||
report = parse_owned_range_report(doc)
|
||||
assert (report.start_layer, report.end_layer) == (1, 3)
|
||||
assert report.registered_tensors == 18
|
||||
assert report.registered_bytes == 2 * LAYER_BYTES
|
||||
# The fixture layers are contiguous in the file, so the pure mmap span
|
||||
# is exactly the owned tensor bytes — scaled down from the artifact.
|
||||
assert report.mapped_bytes == 2 * LAYER_BYTES
|
||||
assert report.mapped_bytes < report.file_bytes
|
||||
|
||||
def test_head_range_owns_embeddings(self, dense_llama_gguf: Path) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 0, 1)
|
||||
assert code == 0
|
||||
report = parse_owned_range_report(doc)
|
||||
assert report.is_head and report.has_token_embeddings
|
||||
assert not report.has_output_head
|
||||
assert report.registered_tensors == 10
|
||||
assert report.registered_bytes == EMBD_BYTES + LAYER_BYTES
|
||||
|
||||
def test_tail_range_owns_norm_and_output(self, dense_llama_gguf: Path) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 3, 4)
|
||||
assert code == 0
|
||||
report = parse_owned_range_report(doc)
|
||||
assert report.is_tail and report.has_output_head
|
||||
assert not report.has_token_embeddings
|
||||
assert report.registered_tensors == 11
|
||||
assert report.registered_bytes == LAYER_BYTES + OUT_NORM_BYTES + OUT_BYTES
|
||||
|
||||
def test_shards_partition_the_whole_model_bytes(self, dense_llama_gguf: Path) -> None:
|
||||
shards = [(0, 1), (1, 3), (3, 4)]
|
||||
registered = []
|
||||
for start, end in shards:
|
||||
code, doc = run_tool(dense_llama_gguf, start, end)
|
||||
assert code == 0
|
||||
registered.append(parse_owned_range_report(doc).registered_bytes)
|
||||
code, doc = run_tool(dense_llama_gguf, 0, 4)
|
||||
assert code == 0
|
||||
whole = parse_owned_range_report(doc)
|
||||
assert whole.registered_tensors == 3 + 9 * K_LAYERS
|
||||
assert sum(registered) == whole.registered_bytes
|
||||
|
||||
def test_non_mmap_load_scales_resident_with_the_range(self, dense_llama_gguf: Path) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 1, 3, "--no-mmap")
|
||||
assert code == 0
|
||||
report = parse_owned_range_report(doc)
|
||||
assert report.mapped_bytes == 0
|
||||
assert report.registered_bytes == 2 * LAYER_BYTES
|
||||
code, doc = run_tool(dense_llama_gguf, 0, 4, "--no-mmap")
|
||||
assert code == 0
|
||||
whole = parse_owned_range_report(doc)
|
||||
assert report.resident_bytes < whole.resident_bytes
|
||||
|
||||
|
||||
@requires_range_report_tool
|
||||
class TestRangeRejection:
|
||||
def test_out_of_model_range_is_refused(self, dense_llama_gguf: Path) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 3, 5)
|
||||
assert code == 3 and doc["ok"] is False
|
||||
with pytest.raises(RangeReportError):
|
||||
parse_owned_range_report(doc)
|
||||
|
||||
def test_empty_range_is_refused(self, dense_llama_gguf: Path) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 2, 2)
|
||||
assert code == 3 and doc["ok"] is False
|
||||
|
||||
def test_inverted_range_is_refused(self, dense_llama_gguf: Path) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 3, 1)
|
||||
assert code == 3 and doc["ok"] is False
|
||||
|
||||
def test_missing_required_owned_tensor_is_refused(self, tmp_path: Path) -> None:
|
||||
corrupted = write_dense_llama_gguf(
|
||||
tmp_path / "missing-tensor.gguf", drop=frozenset({"blk.1.attn_q.weight"})
|
||||
)
|
||||
code, doc = run_tool(corrupted, 0, 2)
|
||||
assert code == 3 and doc["ok"] is False
|
||||
assert "blk.1.attn_q.weight" in doc["error"]
|
||||
|
||||
def test_whole_model_load_still_works_through_the_range_loader(
|
||||
self, dense_llama_gguf: Path
|
||||
) -> None:
|
||||
code, doc = run_tool(dense_llama_gguf, 0, 4)
|
||||
assert code == 0
|
||||
report = parse_owned_range_report(doc)
|
||||
assert report.is_head and report.is_tail
|
||||
assert report.has_token_embeddings and report.has_output_head
|
||||
|
||||
|
||||
def test_tool_binary_gate_points_at_the_locked_build() -> None:
|
||||
# The gate must name the deterministic lane's output, never a downloaded binary.
|
||||
assert DEFAULT_BINARY.name == "meshnet-range-report"
|
||||
assert "llama.cpp" in DEFAULT_BINARY.parts
|
||||
assert DEFAULT_BINARY.parent.name == "bin"
|
||||
assert DEFAULT_BINARY.parent.parent.name == "build"
|
||||
636
tests/test_native_shard_worker.py
Normal file
636
tests/test_native_shard_worker.py
Normal file
@@ -0,0 +1,636 @@
|
||||
"""DGR-033 integration tests for the standalone native C++ Shard worker.
|
||||
|
||||
These tests spawn the *real* compiled ``shard_worker`` executable as a separate
|
||||
OS process, connect to its real localhost socket with the committed generated
|
||||
``ShardRuntimeStub`` stubs, and drive the complete lifecycle/stream contract.
|
||||
There is no in-memory channel, no Python servicer, and no fake transport: the
|
||||
server under test is the C++ binary DGR-033 builds.
|
||||
|
||||
The worker binary is located via ``MESHNET_SHARD_WORKER_BIN`` or the default
|
||||
out-of-tree build path ``build/native/shard_worker``. When it has not been
|
||||
built (a default developer/CI checkout without the pinned gRPC C++ toolchain),
|
||||
every test here is skipped rather than failed — the same ``requires_cmake``
|
||||
gating pattern DGR-029/DGR-030 use for native-build-dependent tests. The
|
||||
session that implemented DGR-033 built the binary and ran these for real; see
|
||||
``evidence/DGR-033/README.md`` for the exact commands and results.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import signal
|
||||
import socket
|
||||
import subprocess
|
||||
import time
|
||||
import zlib
|
||||
|
||||
import grpc
|
||||
import pytest
|
||||
|
||||
REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
|
||||
_PYTHONPATH = os.pathsep.join(
|
||||
[os.path.join(REPO_ROOT, "packages", "node"), os.path.join(REPO_ROOT, "packages", "tracker")]
|
||||
)
|
||||
|
||||
from meshnet_node.native_protocol.generated import ( # noqa: E402
|
||||
shard_runtime_pb2 as pb,
|
||||
shard_runtime_pb2_grpc as pb_grpc,
|
||||
)
|
||||
|
||||
|
||||
def _worker_binary() -> str | None:
|
||||
explicit = os.environ.get("MESHNET_SHARD_WORKER_BIN")
|
||||
if explicit and os.path.exists(explicit):
|
||||
return explicit
|
||||
default = os.path.join(REPO_ROOT, "build", "native", "shard_worker")
|
||||
if os.path.exists(default):
|
||||
return default
|
||||
return None
|
||||
|
||||
|
||||
_WORKER_BIN = _worker_binary()
|
||||
pytestmark = pytest.mark.skipif(
|
||||
_WORKER_BIN is None,
|
||||
reason=(
|
||||
"native shard_worker binary not built; build packages/node/native with the "
|
||||
"pinned gRPC C++ toolchain or set MESHNET_SHARD_WORKER_BIN"
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def _free_port() -> int:
|
||||
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
|
||||
s.bind(("127.0.0.1", 0))
|
||||
port = s.getsockname()[1]
|
||||
s.close()
|
||||
return port
|
||||
|
||||
|
||||
def _start_worker(listen_addr: str, extra_env: dict[str, str] | None = None) -> subprocess.Popen:
|
||||
env = dict(os.environ)
|
||||
env["PYTHONPATH"] = _PYTHONPATH
|
||||
if extra_env:
|
||||
env.update(extra_env)
|
||||
proc = subprocess.Popen(
|
||||
[_WORKER_BIN, listen_addr],
|
||||
cwd=REPO_ROOT,
|
||||
env=env,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.STDOUT,
|
||||
text=True,
|
||||
)
|
||||
deadline = time.time() + 30.0
|
||||
while time.time() < deadline:
|
||||
line = proc.stdout.readline()
|
||||
if not line:
|
||||
if proc.poll() is not None:
|
||||
out, _ = proc.communicate()
|
||||
raise RuntimeError(f"worker exited early:\n{out}")
|
||||
continue
|
||||
if "listening on" in line:
|
||||
return proc
|
||||
raise RuntimeError("worker did not start listening in time")
|
||||
|
||||
|
||||
class _Worker:
|
||||
"""A spawned worker plus a ready channel; also captures stdout on close."""
|
||||
|
||||
def __init__(self, extra_env: dict[str, str] | None = None) -> None:
|
||||
self.port = _free_port()
|
||||
self.addr = f"127.0.0.1:{self.port}"
|
||||
self.proc = _start_worker(self.addr, extra_env)
|
||||
self.channel = grpc.insecure_channel(self.addr)
|
||||
grpc.channel_ready_future(self.channel).result(timeout=15.0)
|
||||
|
||||
def stub(self) -> pb_grpc.ShardRuntimeStub:
|
||||
return pb_grpc.ShardRuntimeStub(self.channel)
|
||||
|
||||
def session(self, requests):
|
||||
call = self.channel.stream_stream(
|
||||
"/meshnet.shard.v1.ShardRuntime/Session",
|
||||
request_serializer=lambda m: m.SerializeToString(),
|
||||
response_deserializer=pb.SessionResponse.FromString,
|
||||
)
|
||||
return list(call(iter(requests)))
|
||||
|
||||
def close(self, *, sig: int = signal.SIGTERM) -> str:
|
||||
self.channel.close()
|
||||
self.proc.send_signal(sig)
|
||||
try:
|
||||
out, _ = self.proc.communicate(timeout=10)
|
||||
except subprocess.TimeoutExpired:
|
||||
self.proc.kill()
|
||||
out, _ = self.proc.communicate()
|
||||
return out or ""
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def worker():
|
||||
w = _Worker()
|
||||
try:
|
||||
yield w
|
||||
finally:
|
||||
if w.proc.poll() is None:
|
||||
w.close()
|
||||
|
||||
|
||||
def _crc32c(payload: bytes) -> bytes:
|
||||
return zlib.crc32(payload).to_bytes(4, "big")
|
||||
|
||||
|
||||
_WORKER_FINGERPRINT = dict(
|
||||
model_artifact_digest="sha256:native-test-artifact",
|
||||
runtime_recipe_digest="sha256:native-test-recipe",
|
||||
recipe_id="native-test",
|
||||
recipe_version="1",
|
||||
catalogue_version="1",
|
||||
)
|
||||
|
||||
|
||||
def _open(
|
||||
*,
|
||||
route_session_id="rs-1",
|
||||
route_epoch=7,
|
||||
credits_granted=16,
|
||||
max_inflight_chunks=16,
|
||||
max_chunk_bytes=4 * 1024 * 1024,
|
||||
schema_version=pb.SCHEMA_VERSION_1,
|
||||
fingerprint=None,
|
||||
shard_range=None,
|
||||
) -> pb.SessionRequest:
|
||||
fp = pb.Fingerprint(**_WORKER_FINGERPRINT) if fingerprint is None else fingerprint
|
||||
sr = (
|
||||
pb.ShardRange(start_layer=0, end_layer=32, effective_start_layer=0)
|
||||
if shard_range is None
|
||||
else shard_range
|
||||
)
|
||||
return pb.SessionRequest(
|
||||
open=pb.SessionOpen(
|
||||
schema_version=schema_version,
|
||||
route_session_id=route_session_id,
|
||||
route_epoch=route_epoch,
|
||||
fingerprint=fp,
|
||||
shard_range=sr,
|
||||
proposed_flow_control=pb.FlowControl(
|
||||
credits_granted=credits_granted,
|
||||
max_inflight_chunks=max_inflight_chunks,
|
||||
max_chunk_bytes=max_chunk_bytes,
|
||||
max_prefill_chunk_tokens=512,
|
||||
),
|
||||
accepted_compression=[pb.COMPRESSION_NONE],
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def _chunk(
|
||||
work_id,
|
||||
payload: bytes,
|
||||
step,
|
||||
*,
|
||||
route_session_id="rs-1",
|
||||
route_epoch=7,
|
||||
deadline_unix_nanos=0,
|
||||
fragments=1,
|
||||
total_bytes=None,
|
||||
) -> pb.SessionRequest:
|
||||
total = len(payload) if total_bytes is None else total_bytes
|
||||
frags = []
|
||||
if fragments == 1:
|
||||
frags = [pb.TensorFragment(fragment_index=0, fragment_count=1, byte_offset=0, payload=payload)]
|
||||
else:
|
||||
# Split into ``fragments`` tiling pieces.
|
||||
size = max(1, len(payload) // fragments)
|
||||
offset = 0
|
||||
idx = 0
|
||||
while offset < len(payload):
|
||||
piece = payload[offset : offset + size] if idx < fragments - 1 else payload[offset:]
|
||||
frags.append(
|
||||
pb.TensorFragment(
|
||||
fragment_index=idx, fragment_count=fragments, byte_offset=offset, payload=piece
|
||||
)
|
||||
)
|
||||
offset += len(piece)
|
||||
idx += 1
|
||||
tensor = pb.NamedTensor(
|
||||
name="hidden_states",
|
||||
shape=[1, 1, 4096],
|
||||
dtype=pb.DTYPE_BFLOAT16,
|
||||
byte_order=pb.BYTE_ORDER_LITTLE_ENDIAN,
|
||||
total_bytes=total,
|
||||
compression=pb.COMPRESSION_NONE,
|
||||
checksum=pb.Checksum(algorithm=pb.CHECKSUM_ALGORITHM_CRC32C, value=_crc32c(payload)),
|
||||
fragments=frags,
|
||||
)
|
||||
bundle = pb.TensorBundle(
|
||||
bundle_version=1,
|
||||
tensors=[tensor],
|
||||
architecture=pb.ARCHITECTURE_TYPE_DENSE,
|
||||
boundary_point="pre_tail_residual",
|
||||
)
|
||||
envelope = pb.Envelope(
|
||||
schema_version=pb.SCHEMA_VERSION_1,
|
||||
work_id=work_id,
|
||||
route_session_id=route_session_id,
|
||||
route_epoch=route_epoch,
|
||||
idempotency_step=step,
|
||||
phase=pb.PHASE_PREFILL,
|
||||
position=pb.PositionSpan(first_position=0, token_count=1),
|
||||
deadline_unix_nanos=deadline_unix_nanos,
|
||||
)
|
||||
return pb.SessionRequest(chunk=pb.ActivationChunk(envelope=envelope, bundle=bundle))
|
||||
|
||||
|
||||
def _decode(work_id, payload: bytes, step, position) -> pb.SessionRequest:
|
||||
tensor = pb.NamedTensor(
|
||||
name="hidden_states",
|
||||
shape=[1, 1, 4096],
|
||||
dtype=pb.DTYPE_BFLOAT16,
|
||||
byte_order=pb.BYTE_ORDER_LITTLE_ENDIAN,
|
||||
total_bytes=len(payload),
|
||||
compression=pb.COMPRESSION_NONE,
|
||||
checksum=pb.Checksum(algorithm=pb.CHECKSUM_ALGORITHM_CRC32C, value=_crc32c(payload)),
|
||||
fragments=[pb.TensorFragment(fragment_index=0, fragment_count=1, byte_offset=0, payload=payload)],
|
||||
)
|
||||
return pb.SessionRequest(
|
||||
decode=pb.DecodeStep(
|
||||
idempotency_step=step,
|
||||
position=position,
|
||||
expected_past_len=position,
|
||||
work_id=work_id,
|
||||
bundle=pb.TensorBundle(bundle_version=1, tensors=[tensor], architecture=pb.ARCHITECTURE_TYPE_DENSE),
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def _release() -> pb.SessionRequest:
|
||||
return pb.SessionRequest(
|
||||
release=pb.ReleaseSignal(route_session_id="rs-1", route_epoch=7, work_id="work-final")
|
||||
)
|
||||
|
||||
|
||||
def _cancel(*, route_session_id="rs-1", work_id="", reason="test cancel") -> pb.SessionRequest:
|
||||
return pb.SessionRequest(
|
||||
cancel=pb.CancelSignal(route_session_id=route_session_id, route_epoch=7, work_id=work_id, reason=reason)
|
||||
)
|
||||
|
||||
|
||||
# --- startup / health / capability ----------------------------------------
|
||||
|
||||
|
||||
def test_worker_startup_and_health(worker):
|
||||
health = worker.stub().Health(pb.HealthRequest(schema_version=pb.SCHEMA_VERSION_1))
|
||||
assert health.state == pb.SERVING_STATE_SERVING
|
||||
assert health.schema_version == pb.SCHEMA_VERSION_1
|
||||
|
||||
|
||||
def test_worker_capability(worker):
|
||||
cap = worker.stub().GetCapability(pb.CapabilityRequest(schema_version=pb.SCHEMA_VERSION_1))
|
||||
assert cap.validated is True
|
||||
assert cap.schema_version == pb.SCHEMA_VERSION_1
|
||||
assert cap.shard_range.end_layer == 32
|
||||
assert pb.SCHEMA_VERSION_1 in cap.supported_schema_versions
|
||||
|
||||
|
||||
# --- fragmented prefill / decode / release ---------------------------------
|
||||
|
||||
|
||||
def test_fragmented_prefill_echoes_reassembled_payload(worker):
|
||||
payload = b"REAL_ACTIVATION_BYTES_prefill_across_three_fragments_1234567890"
|
||||
responses = worker.session([_open(), _chunk("w1", payload, step=1, fragments=3), _release()])
|
||||
assert responses[0].WhichOneof("kind") == "accepted"
|
||||
echoed = responses[1]
|
||||
assert echoed.WhichOneof("kind") == "chunk"
|
||||
got = b"".join(f.payload for f in echoed.chunk.bundle.tensors[0].fragments)
|
||||
assert got == payload
|
||||
assert echoed.chunk.bundle.tensors[0].checksum.value == _crc32c(payload)
|
||||
assert responses[2].status.terminal is True
|
||||
|
||||
|
||||
def test_decode_step_is_served(worker):
|
||||
payload = b"REAL_ACTIVATION_BYTES_decode_step"
|
||||
responses = worker.session([_open(), _decode("w2", payload, step=1, position=1)])
|
||||
echoed = responses[1]
|
||||
assert echoed.WhichOneof("kind") == "chunk"
|
||||
assert echoed.chunk.envelope.phase == pb.PHASE_DECODE
|
||||
assert echoed.chunk.bundle.tensors[0].fragments[0].payload == payload
|
||||
|
||||
|
||||
def test_release_is_terminal(worker):
|
||||
responses = worker.session([_open(), _release()])
|
||||
assert responses[0].WhichOneof("kind") == "accepted"
|
||||
assert responses[1].status.terminal is True
|
||||
|
||||
|
||||
# --- deadlines / flow control / bounded messages ---------------------------
|
||||
|
||||
|
||||
def test_expired_deadline_is_rejected(worker):
|
||||
responses = worker.session([_open(), _chunk("w-late", b"payload", step=1, deadline_unix_nanos=1)])
|
||||
assert responses[1].status.error.code == pb.ERROR_CODE_DEADLINE_EXCEEDED
|
||||
|
||||
|
||||
def test_flow_control_violation_and_topup(worker):
|
||||
responses = worker.session(
|
||||
[
|
||||
_open(credits_granted=1),
|
||||
_chunk("w-a", b"payload-a", step=1),
|
||||
_chunk("w-b", b"payload-b", step=2),
|
||||
pb.SessionRequest(flow_control=pb.FlowControl(credits_granted=5)),
|
||||
_chunk("w-c", b"payload-c", step=3),
|
||||
]
|
||||
)
|
||||
assert responses[1].WhichOneof("kind") == "chunk"
|
||||
assert responses[2].status.error.code == pb.ERROR_CODE_FLOW_CONTROL_VIOLATION
|
||||
assert responses[2].status.error.retryable is True
|
||||
assert responses[3].WhichOneof("kind") == "flow_control"
|
||||
assert responses[3].flow_control.credits_granted >= 5
|
||||
assert responses[4].WhichOneof("kind") == "chunk"
|
||||
|
||||
|
||||
def test_bounded_message_is_rejected():
|
||||
"""A tensor whose declared payload exceeds the negotiated ceiling is refused."""
|
||||
w = _Worker(extra_env={"MESHNET_MAX_CHUNK_BYTES": "64"})
|
||||
try:
|
||||
big = b"x" * 128
|
||||
responses = w.session([_open(), _chunk("w-big", big, step=1, total_bytes=128)])
|
||||
status = responses[1].status
|
||||
assert status.error.code == pb.ERROR_CODE_RESOURCE_EXHAUSTED
|
||||
assert "max_chunk_bytes" in status.error.detail
|
||||
finally:
|
||||
if w.proc.poll() is None:
|
||||
w.close()
|
||||
|
||||
|
||||
def test_malformed_fragment_tiling_is_rejected(worker):
|
||||
# A fragment at a non-zero offset with no predecessor cannot tile.
|
||||
tensor = pb.NamedTensor(
|
||||
name="hidden_states",
|
||||
shape=[1, 1, 4096],
|
||||
dtype=pb.DTYPE_BFLOAT16,
|
||||
byte_order=pb.BYTE_ORDER_LITTLE_ENDIAN,
|
||||
total_bytes=7,
|
||||
compression=pb.COMPRESSION_NONE,
|
||||
checksum=pb.Checksum(algorithm=pb.CHECKSUM_ALGORITHM_CRC32C, value=_crc32c(b"payload")),
|
||||
fragments=[pb.TensorFragment(fragment_index=0, fragment_count=1, byte_offset=5, payload=b"payload")],
|
||||
)
|
||||
bad = pb.SessionRequest(
|
||||
chunk=pb.ActivationChunk(
|
||||
envelope=pb.Envelope(
|
||||
schema_version=pb.SCHEMA_VERSION_1,
|
||||
work_id="w-gap",
|
||||
route_session_id="rs-1",
|
||||
route_epoch=7,
|
||||
idempotency_step=1,
|
||||
),
|
||||
bundle=pb.TensorBundle(bundle_version=1, tensors=[tensor]),
|
||||
)
|
||||
)
|
||||
responses = worker.session([_open(), bad])
|
||||
assert responses[1].status.error.code == pb.ERROR_CODE_PAYLOAD_CORRUPT
|
||||
assert "tile" in responses[1].status.error.detail
|
||||
|
||||
|
||||
def test_stale_route_epoch_is_rejected(worker):
|
||||
responses = worker.session([_open(route_epoch=7), _chunk("w-stale", b"payload", step=1, route_epoch=5)])
|
||||
assert responses[1].status.error.code == pb.ERROR_CODE_EPOCH_STALE
|
||||
|
||||
|
||||
def test_duplicate_idempotency_step_is_acked(worker):
|
||||
chunk = _chunk("w-dup", b"payload", step=1)
|
||||
responses = worker.session([_open(), chunk, chunk])
|
||||
assert responses[1].WhichOneof("kind") == "chunk"
|
||||
assert responses[2].WhichOneof("kind") == "ack"
|
||||
assert responses[2].ack.duplicate is True
|
||||
|
||||
|
||||
# --- cancellation ----------------------------------------------------------
|
||||
|
||||
|
||||
def test_in_band_cancel_of_single_work_item_does_not_end_stream(worker):
|
||||
responses = worker.session(
|
||||
[
|
||||
_open(),
|
||||
_cancel(work_id="work-x"),
|
||||
_chunk("work-x", b"payload", step=1),
|
||||
_chunk("work-y", b"payload", step=2),
|
||||
_release(),
|
||||
]
|
||||
)
|
||||
assert responses[1].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||
assert responses[1].status.terminal is False
|
||||
assert responses[2].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||
assert responses[3].WhichOneof("kind") == "chunk"
|
||||
assert responses[4].status.terminal is True
|
||||
|
||||
|
||||
def test_in_band_cancel_of_whole_session_is_terminal(worker):
|
||||
responses = worker.session([_open(), _cancel(work_id="")])
|
||||
assert responses[1].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||
assert responses[1].status.terminal is True
|
||||
|
||||
|
||||
def test_out_of_band_cancel_rpc_races_ahead_of_open(worker):
|
||||
stub = worker.stub()
|
||||
resp = stub.Cancel(
|
||||
pb.CancelRequest(
|
||||
schema_version=pb.SCHEMA_VERSION_1,
|
||||
route_session_id="rs-precancel",
|
||||
route_epoch=1,
|
||||
work_id="work-precancelled",
|
||||
reason="operator abort",
|
||||
)
|
||||
)
|
||||
assert resp.cancelled_work_items == 1
|
||||
responses = worker.session(
|
||||
[
|
||||
_open(route_session_id="rs-precancel"),
|
||||
_chunk("work-precancelled", b"payload", step=1, route_session_id="rs-precancel"),
|
||||
]
|
||||
)
|
||||
assert responses[1].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||
|
||||
|
||||
def test_release_rpc_is_idempotent(worker):
|
||||
stub = worker.stub()
|
||||
# Open a session WITHOUT an in-stream release so state persists on the
|
||||
# servicer, then drop it out of band twice.
|
||||
worker.session([_open(route_session_id="rs-rel")])
|
||||
first = stub.Release(pb.ReleaseRequest(schema_version=pb.SCHEMA_VERSION_1, route_session_id="rs-rel", route_epoch=7))
|
||||
second = stub.Release(pb.ReleaseRequest(schema_version=pb.SCHEMA_VERSION_1, route_session_id="rs-rel", route_epoch=7))
|
||||
assert first.released is True
|
||||
assert second.released is False # idempotent: nothing left to drop
|
||||
|
||||
|
||||
def test_independent_session_cancellation(worker):
|
||||
# Cancel the whole of session A; session B must remain fully serviceable.
|
||||
a = worker.session([_open(route_session_id="sess-A"), _cancel(route_session_id="sess-A", work_id="")])
|
||||
assert a[1].status.terminal is True
|
||||
b = worker.session(
|
||||
[
|
||||
_open(route_session_id="sess-B"),
|
||||
_chunk("work-b", b"payload-b", step=1, route_session_id="sess-B"),
|
||||
_release(),
|
||||
]
|
||||
)
|
||||
assert b[1].WhichOneof("kind") == "chunk", "cancelling session A must not affect session B"
|
||||
|
||||
|
||||
# --- graceful shutdown -----------------------------------------------------
|
||||
|
||||
|
||||
def test_graceful_shutdown_on_sigterm():
|
||||
w = _Worker()
|
||||
# Confirm it is serving, then send SIGTERM and require a clean drain/exit.
|
||||
assert w.stub().Health(pb.HealthRequest(schema_version=pb.SCHEMA_VERSION_1)).state == pb.SERVING_STATE_SERVING
|
||||
out = w.close(sig=signal.SIGTERM)
|
||||
assert w.proc.returncode == 0, f"worker did not exit cleanly on SIGTERM:\n{out}"
|
||||
assert "shut down cleanly" in out
|
||||
|
||||
|
||||
# --- direct vs opaque relay byte identity ----------------------------------
|
||||
|
||||
|
||||
def test_direct_and_opaque_relay_yield_identical_responses(worker):
|
||||
"""A direct hop and an opaque relay of the exact captured request bytes must
|
||||
produce byte-identical server responses (relays carry frames verbatim)."""
|
||||
payload = b"RELAY_ACTIVATION_BYTES"
|
||||
requests = [_open(), _chunk("w1", payload, step=1), _release()]
|
||||
|
||||
direct_call = worker.channel.stream_stream(
|
||||
"/meshnet.shard.v1.ShardRuntime/Session",
|
||||
request_serializer=lambda m: m.SerializeToString(),
|
||||
response_deserializer=lambda b: b,
|
||||
)
|
||||
direct_resp = list(direct_call(iter(requests)))
|
||||
captured = [m.SerializeToString() for m in requests]
|
||||
|
||||
relay_call = worker.channel.stream_stream(
|
||||
"/meshnet.shard.v1.ShardRuntime/Session",
|
||||
request_serializer=lambda b: b, # raw captured bytes, no reinterpretation
|
||||
response_deserializer=lambda b: b,
|
||||
)
|
||||
relay_resp = list(relay_call(iter(captured)))
|
||||
|
||||
assert len(direct_resp) == len(relay_resp) == 3
|
||||
for i, (d, r) in enumerate(zip(direct_resp, relay_resp)):
|
||||
assert d == r, f"response #{i} differs between direct and opaque relay"
|
||||
|
||||
|
||||
# --- fail-closed before SessionOpen ----------------------------------------
|
||||
|
||||
|
||||
def test_chunk_before_open_is_rejected(worker):
|
||||
# An activation with no preceding SessionOpen must fail closed and end the
|
||||
# stream: no work may bypass the lifecycle handshake.
|
||||
responses = worker.session([_chunk("w-noopen", b"payload", step=1)])
|
||||
assert len(responses) == 1
|
||||
assert responses[0].WhichOneof("kind") == "status"
|
||||
assert responses[0].status.error.code == pb.ERROR_CODE_INTERNAL
|
||||
assert responses[0].status.terminal is True
|
||||
assert "SessionOpen" in responses[0].status.error.detail
|
||||
|
||||
|
||||
def test_decode_before_open_is_rejected(worker):
|
||||
responses = worker.session([_decode("w-noopen", b"payload", step=1, position=0)])
|
||||
assert len(responses) == 1
|
||||
assert responses[0].WhichOneof("kind") == "status"
|
||||
assert responses[0].status.error.code == pb.ERROR_CODE_INTERNAL
|
||||
assert responses[0].status.terminal is True
|
||||
|
||||
|
||||
# --- flow-control negotiation with strict worker bounds --------------------
|
||||
|
||||
|
||||
def test_flow_control_proposal_is_clamped_to_worker_bounds(worker):
|
||||
# A peer proposing a window far above the worker limits must be clamped to
|
||||
# the worker own ceilings, never granted the inflated proposal.
|
||||
responses = worker.session(
|
||||
[_open(credits_granted=9999, max_inflight_chunks=9999, max_chunk_bytes=1073741824)]
|
||||
)
|
||||
fc = responses[0].accepted.flow_control
|
||||
assert fc.max_inflight_chunks == 16
|
||||
assert fc.credits_granted == 16
|
||||
assert fc.max_chunk_bytes == 4 * 1024 * 1024
|
||||
|
||||
|
||||
def test_negotiated_max_chunk_bytes_caps_peer_proposal():
|
||||
# Worker ceiling is 64 bytes; the peer proposes 4 MiB. The negotiated per
|
||||
# session ceiling is the stricter 64, so a 128-byte tensor is refused even
|
||||
# though the peer allowed it — the worker never adopts the peer proposal.
|
||||
w = _Worker(extra_env={"MESHNET_MAX_CHUNK_BYTES": "64"})
|
||||
try:
|
||||
big = b"x" * 128
|
||||
responses = w.session(
|
||||
[_open(max_chunk_bytes=4 * 1024 * 1024), _chunk("w-big", big, step=1, total_bytes=128)]
|
||||
)
|
||||
assert responses[0].accepted.flow_control.max_chunk_bytes == 64
|
||||
assert responses[1].status.error.code == pb.ERROR_CODE_RESOURCE_EXHAUSTED
|
||||
assert "max_chunk_bytes" in responses[1].status.error.detail
|
||||
finally:
|
||||
if w.proc.poll() is None:
|
||||
w.close()
|
||||
|
||||
|
||||
# --- in-stream release erases session state --------------------------------
|
||||
|
||||
|
||||
def test_in_stream_release_erases_session_state(worker):
|
||||
stub = worker.stub()
|
||||
resp = worker.session(
|
||||
[
|
||||
_open(route_session_id="rs-erase"),
|
||||
pb.SessionRequest(
|
||||
release=pb.ReleaseSignal(route_session_id="rs-erase", route_epoch=7, work_id="w-final")
|
||||
),
|
||||
]
|
||||
)
|
||||
assert resp[-1].status.terminal is True
|
||||
# The state is already gone: an out-of-band Release finds nothing to drop.
|
||||
after = stub.Release(
|
||||
pb.ReleaseRequest(schema_version=pb.SCHEMA_VERSION_1, route_session_id="rs-erase", route_epoch=7)
|
||||
)
|
||||
assert after.released is False
|
||||
|
||||
|
||||
# --- SessionOpen identity validation ---------------------------------------
|
||||
|
||||
|
||||
def test_incompatible_schema_is_rejected_at_open(worker):
|
||||
responses = worker.session([_open(schema_version=pb.SCHEMA_VERSION_UNSPECIFIED)])
|
||||
assert responses[0].WhichOneof("kind") == "status"
|
||||
assert responses[0].status.error.code == pb.ERROR_CODE_SCHEMA_UNSUPPORTED
|
||||
assert responses[0].status.terminal is True
|
||||
|
||||
|
||||
def test_incompatible_fingerprint_is_rejected_at_open(worker):
|
||||
bad_fp = pb.Fingerprint(
|
||||
model_artifact_digest="sha256:some-other-model",
|
||||
runtime_recipe_digest="sha256:native-test-recipe",
|
||||
recipe_id="native-test",
|
||||
recipe_version="1",
|
||||
catalogue_version="1",
|
||||
)
|
||||
responses = worker.session([_open(fingerprint=bad_fp)])
|
||||
assert responses[0].WhichOneof("kind") == "status"
|
||||
assert responses[0].status.error.code == pb.ERROR_CODE_FINGERPRINT_MISMATCH
|
||||
assert responses[0].status.terminal is True
|
||||
|
||||
|
||||
def test_shard_range_mismatch_is_rejected_at_open(worker):
|
||||
responses = worker.session(
|
||||
[_open(shard_range=pb.ShardRange(start_layer=0, end_layer=64, effective_start_layer=0))]
|
||||
)
|
||||
assert responses[0].WhichOneof("kind") == "status"
|
||||
assert responses[0].status.error.code == pb.ERROR_CODE_SHARD_RANGE_MISMATCH
|
||||
assert responses[0].status.terminal is True
|
||||
|
||||
|
||||
def test_session_accepted_reports_worker_fingerprint_not_caller(worker):
|
||||
# The caller asserts no fingerprint; SessionAccepted must carry the worker
|
||||
# OWN served identity, not a copy of the caller (empty) fingerprint.
|
||||
responses = worker.session([_open(fingerprint=pb.Fingerprint())])
|
||||
assert responses[0].WhichOneof("kind") == "accepted"
|
||||
accepted = responses[0].accepted
|
||||
assert accepted.fingerprint.model_artifact_digest == "sha256:native-test-artifact"
|
||||
assert accepted.fingerprint.runtime_recipe_digest == "sha256:native-test-recipe"
|
||||
273
tests/test_range_report.py
Normal file
273
tests/test_range_report.py
Normal file
@@ -0,0 +1,273 @@
|
||||
"""DGR-034: strict consumption of owned-range reports from loaded engine state.
|
||||
|
||||
The ``meshnet-range-report`` native tool loads one dense-Llama GGUF through
|
||||
the Meshnet owned-range loader and prints a JSON document derived from the
|
||||
loaded model state. ``meshnet_node.range_report`` is the strict consumer:
|
||||
it must accept exactly the documents that encode the dense-Llama ownership
|
||||
contract and fail closed on everything else — invalid, empty, or
|
||||
out-of-model ranges, endpoint registrations that disagree with the loaded
|
||||
state, gapped or unexpected tensor registrations, and inconsistent byte
|
||||
counts.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
|
||||
from meshnet_node.range_report import (
|
||||
OwnedRangeReport,
|
||||
RangeReportError,
|
||||
parse_owned_range_report,
|
||||
)
|
||||
|
||||
N_LAYER = 40
|
||||
LAYER_BYTES = 300 * 2**20
|
||||
EMBD_BYTES = 360 * 2**20
|
||||
OUT_BYTES = 525 * 2**20
|
||||
FILE_BYTES = 13669 * 2**20
|
||||
|
||||
|
||||
def _doc(**overrides: Any) -> dict[str, Any]:
|
||||
"""A valid middle-range [10, 20) mmap report the consumer must accept."""
|
||||
doc: dict[str, Any] = {
|
||||
"ok": True,
|
||||
"model": "/models/dense.gguf",
|
||||
"architecture": "llama",
|
||||
"n_layer": N_LAYER,
|
||||
"file_bytes": FILE_BYTES,
|
||||
"requested_range": [10, 20],
|
||||
"reported_range": [10, 20],
|
||||
"mmap": True,
|
||||
"touched": False,
|
||||
"use_extra_bufts": True,
|
||||
"has_token_embeddings": False,
|
||||
"has_output_head": False,
|
||||
"tied_output_head": False,
|
||||
"mapped_bytes": 10 * LAYER_BYTES,
|
||||
"resident_bytes": 10 * LAYER_BYTES,
|
||||
"registered_tensors": 90,
|
||||
"registered_bytes": 10 * LAYER_BYTES,
|
||||
"unexpected_registered_tensors": [],
|
||||
"missing_owned_layers": [],
|
||||
"vm_size_bytes": FILE_BYTES + 2**28,
|
||||
"vm_rss_bytes": 2**28,
|
||||
"vm_hwm_bytes": 2**28,
|
||||
}
|
||||
doc.update(overrides)
|
||||
return doc
|
||||
|
||||
|
||||
def _head_doc(**overrides: Any) -> dict[str, Any]:
|
||||
base = _doc(
|
||||
requested_range=[0, 10],
|
||||
reported_range=[0, 10],
|
||||
has_token_embeddings=True,
|
||||
mapped_bytes=10 * LAYER_BYTES + EMBD_BYTES,
|
||||
resident_bytes=10 * LAYER_BYTES + EMBD_BYTES,
|
||||
registered_tensors=91,
|
||||
registered_bytes=10 * LAYER_BYTES + EMBD_BYTES,
|
||||
)
|
||||
base.update(overrides)
|
||||
return base
|
||||
|
||||
|
||||
def _tail_doc(**overrides: Any) -> dict[str, Any]:
|
||||
base = _doc(
|
||||
requested_range=[30, 40],
|
||||
reported_range=[30, 40],
|
||||
has_output_head=True,
|
||||
mapped_bytes=10 * LAYER_BYTES + OUT_BYTES,
|
||||
resident_bytes=10 * LAYER_BYTES + OUT_BYTES,
|
||||
registered_tensors=92,
|
||||
registered_bytes=10 * LAYER_BYTES + OUT_BYTES,
|
||||
)
|
||||
base.update(overrides)
|
||||
return base
|
||||
|
||||
|
||||
class TestAcceptance:
|
||||
def test_middle_range_registers_only_per_layer_tensors(self) -> None:
|
||||
report = parse_owned_range_report(_doc())
|
||||
assert (report.start_layer, report.end_layer) == (10, 20)
|
||||
assert not report.is_head and not report.is_tail
|
||||
assert not report.has_token_embeddings and not report.has_output_head
|
||||
|
||||
def test_head_range_owns_embeddings_only_at_the_head(self) -> None:
|
||||
report = parse_owned_range_report(_head_doc())
|
||||
assert report.is_head and not report.is_tail
|
||||
assert report.has_token_embeddings and not report.has_output_head
|
||||
|
||||
def test_tail_range_owns_norm_and_output_only_at_the_tail(self) -> None:
|
||||
report = parse_owned_range_report(_tail_doc())
|
||||
assert report.is_tail and not report.is_head
|
||||
assert report.has_output_head and not report.has_token_embeddings
|
||||
|
||||
def test_whole_model_range_owns_both_endpoints(self) -> None:
|
||||
report = parse_owned_range_report(
|
||||
_head_doc(
|
||||
requested_range=[0, 40],
|
||||
reported_range=[0, 40],
|
||||
has_output_head=True,
|
||||
mapped_bytes=FILE_BYTES,
|
||||
resident_bytes=FILE_BYTES,
|
||||
registered_tensors=363,
|
||||
registered_bytes=N_LAYER * LAYER_BYTES + EMBD_BYTES + OUT_BYTES,
|
||||
)
|
||||
)
|
||||
assert report.is_head and report.is_tail
|
||||
assert report.has_token_embeddings and report.has_output_head
|
||||
|
||||
def test_tied_output_tail_registers_the_embedding_as_its_output_head(self) -> None:
|
||||
report = parse_owned_range_report(
|
||||
_tail_doc(
|
||||
has_token_embeddings=True,
|
||||
tied_output_head=True,
|
||||
registered_tensors=91,
|
||||
registered_bytes=10 * LAYER_BYTES + EMBD_BYTES,
|
||||
mapped_bytes=10 * LAYER_BYTES + EMBD_BYTES,
|
||||
resident_bytes=10 * LAYER_BYTES + EMBD_BYTES,
|
||||
)
|
||||
)
|
||||
assert report.tied_output_head and report.has_output_head
|
||||
|
||||
def test_non_mmap_load_reports_resident_allocation_only(self) -> None:
|
||||
report = parse_owned_range_report(
|
||||
_doc(mmap=False, mapped_bytes=0, resident_bytes=10 * LAYER_BYTES)
|
||||
)
|
||||
assert report.mapped_bytes == 0
|
||||
assert report.resident_bytes == 10 * LAYER_BYTES
|
||||
|
||||
def test_process_counters_may_be_absent_off_linux(self) -> None:
|
||||
report = parse_owned_range_report(
|
||||
_doc(vm_size_bytes=None, vm_rss_bytes=None, vm_hwm_bytes=None)
|
||||
)
|
||||
assert report.vm_hwm_bytes is None
|
||||
|
||||
|
||||
class TestRangeRejection:
|
||||
def test_rejected_load_fails_closed_with_the_tool_error(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="dense Llama only"):
|
||||
parse_owned_range_report(
|
||||
{"ok": False, "error": "owned-range load rejected the artifact or range: dense Llama only"}
|
||||
)
|
||||
|
||||
def test_reported_range_must_match_the_requested_range(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="loaded engine state"):
|
||||
parse_owned_range_report(_doc(reported_range=[10, 21]))
|
||||
|
||||
def test_out_of_model_range_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="outside the model"):
|
||||
parse_owned_range_report(
|
||||
_doc(requested_range=[30, 41], reported_range=[30, 41], has_output_head=True)
|
||||
)
|
||||
|
||||
def test_empty_range_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="empty or"):
|
||||
parse_owned_range_report(_doc(requested_range=[10, 10], reported_range=[10, 10]))
|
||||
|
||||
def test_inverted_range_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="empty or"):
|
||||
parse_owned_range_report(_doc(requested_range=[20, 10], reported_range=[20, 10]))
|
||||
|
||||
def test_boolean_range_bounds_are_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="integer pair"):
|
||||
parse_owned_range_report(_doc(reported_range=[True, 20]))
|
||||
|
||||
|
||||
class TestEndpointRejection:
|
||||
def test_embeddings_registered_below_the_head_are_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="embeddings belong to the head"):
|
||||
parse_owned_range_report(_doc(has_token_embeddings=True))
|
||||
|
||||
def test_output_head_registered_above_the_tail_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="output head belong to the tail"):
|
||||
parse_owned_range_report(_tail_doc(requested_range=[20, 30], reported_range=[20, 30]))
|
||||
|
||||
def test_tail_without_an_output_head_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="output head belong to the tail"):
|
||||
parse_owned_range_report(_tail_doc(has_output_head=False))
|
||||
|
||||
def test_tied_output_below_the_tail_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="only belong to the tail"):
|
||||
parse_owned_range_report(_doc(tied_output_head=True))
|
||||
|
||||
def test_unexpected_registered_tensors_are_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="unexpected_registered_tensors"):
|
||||
parse_owned_range_report(
|
||||
_doc(unexpected_registered_tensors=["blk.10.attn_q.weight.extra"])
|
||||
)
|
||||
|
||||
def test_missing_owned_layers_are_rejected_as_gaps(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="missing_owned_layers"):
|
||||
parse_owned_range_report(_doc(missing_owned_layers=[12]))
|
||||
|
||||
|
||||
class TestByteCountRejection:
|
||||
def test_mapped_span_must_cover_the_registered_tensors(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="undercounts"):
|
||||
parse_owned_range_report(_doc(mapped_bytes=LAYER_BYTES))
|
||||
|
||||
def test_mapped_span_must_not_exceed_the_artifact(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="exceeds the artifact"):
|
||||
parse_owned_range_report(
|
||||
_tail_doc(mapped_bytes=FILE_BYTES + 1, resident_bytes=FILE_BYTES + 1)
|
||||
)
|
||||
|
||||
def test_non_mmap_load_must_not_claim_a_mapped_span(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="must not claim"):
|
||||
parse_owned_range_report(_doc(mmap=False, mapped_bytes=LAYER_BYTES))
|
||||
|
||||
def test_resident_allocation_must_cover_the_registered_tensors(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="undercounts"):
|
||||
parse_owned_range_report(
|
||||
_doc(mmap=False, mapped_bytes=0, resident_bytes=LAYER_BYTES)
|
||||
)
|
||||
|
||||
def test_an_empty_registration_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="no tensors"):
|
||||
parse_owned_range_report(_doc(registered_tensors=0, registered_bytes=0))
|
||||
|
||||
|
||||
class TestSchemaRejection:
|
||||
def test_wrong_architecture_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="dense Llama only"):
|
||||
parse_owned_range_report(_doc(architecture="qwen2"))
|
||||
|
||||
def test_missing_field_is_rejected(self) -> None:
|
||||
doc = _doc()
|
||||
del doc["mapped_bytes"]
|
||||
with pytest.raises(RangeReportError, match="missing field"):
|
||||
parse_owned_range_report(doc)
|
||||
|
||||
def test_boolean_bytes_are_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="non-negative integer"):
|
||||
parse_owned_range_report(_doc(mapped_bytes=True))
|
||||
|
||||
def test_non_mapping_document_is_rejected(self) -> None:
|
||||
with pytest.raises(RangeReportError, match="JSON object"):
|
||||
parse_owned_range_report(["not", "a", "report"]) # type: ignore[arg-type]
|
||||
|
||||
|
||||
def test_owned_range_report_rejects_direct_construction_outside_the_contract() -> None:
|
||||
with pytest.raises(RangeReportError, match="dense Llama only"):
|
||||
OwnedRangeReport(
|
||||
architecture="qwen2",
|
||||
n_layer=N_LAYER,
|
||||
start_layer=10,
|
||||
end_layer=20,
|
||||
has_token_embeddings=False,
|
||||
has_output_head=False,
|
||||
tied_output_head=False,
|
||||
mapped_bytes=10 * LAYER_BYTES,
|
||||
resident_bytes=10 * LAYER_BYTES,
|
||||
registered_tensors=90,
|
||||
registered_bytes=10 * LAYER_BYTES,
|
||||
file_bytes=FILE_BYTES,
|
||||
mmap=True,
|
||||
touched=False,
|
||||
vm_size_bytes=None,
|
||||
vm_rss_bytes=None,
|
||||
vm_hwm_bytes=None,
|
||||
)
|
||||
Reference in New Issue
Block a user