story: DGR-033 Build a standalone fake C++ gRPC Shard worker
This commit is contained in:
3
.ralph-lane/controller.log
Normal file
3
.ralph-lane/controller.log
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
[2026-07-25T19:19:43+00:00] START lane=opus agent=claude model=opus branch=ralph/distributed-gguf-opus
|
||||||
|
[2026-07-25T19:19:43+00:00] CLAIM DGR-033: Build a standalone fake C++ gRPC Shard worker
|
||||||
|
[2026-07-25T19:19:43+00:00] $ ralph-tui run --prd /run/media/popov/d/DEV/repos/d-popov.com/AI/.claude/worktrees/distributed-gguf-opus/.ralph-lane/current-prd.json --agent claude --model opus --iterations 1 --no-tui --no-setup --verify --cwd /run/media/popov/d/DEV/repos/d-popov.com/AI/.claude/worktrees/distributed-gguf-opus --output-dir /run/media/popov/d/DEV/repos/d-popov.com/AI/.claude/worktrees/distributed-gguf-opus/.ralph-lane/iterations --progress-file /run/media/popov/d/DEV/repos/d-popov.com/AI/.claude/worktrees/distributed-gguf-opus/.ralph-lane/progress.md
|
||||||
1
.ralph-lane/controller.pid
Normal file
1
.ralph-lane/controller.pid
Normal file
@@ -0,0 +1 @@
|
|||||||
|
205508
|
||||||
2210
.ralph-lane/current-prd.json
Normal file
2210
.ralph-lane/current-prd.json
Normal file
File diff suppressed because it is too large
Load Diff
208
.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md
Normal file
208
.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md
Normal file
@@ -0,0 +1,208 @@
|
|||||||
|
# DGR-033 evidence — standalone fake C++ gRPC Shard worker
|
||||||
|
|
||||||
|
**Completed:** 2026-07-25
|
||||||
|
**Branch:** `ralph/distributed-gguf-opus`
|
||||||
|
**Authority:** `.scratch/distributed-gguf-runtime/prd.json`
|
||||||
|
**Dependencies:** DGR-022 (lifecycle/status contract), DGR-024 (real generated
|
||||||
|
gRPC harness + `shard_runtime_server.py` reference semantics), DGR-032
|
||||||
|
(deterministic fake `ShardEngine` semantics).
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
|
||||||
|
Prove the standalone worker process, stream, lifecycle, and supervision shape
|
||||||
|
before any llama.cpp integration: a real C++ executable that serves the whole
|
||||||
|
ShardRuntime lifecycle/stream contract over gRPC using a model-free fake engine,
|
||||||
|
driven end-to-end by Python integration tests over a real socket.
|
||||||
|
|
||||||
|
## What was found live before changing code
|
||||||
|
|
||||||
|
- `packages/node/native/proto/shard_runtime.proto` (DGR-021..023): the single
|
||||||
|
semantic contract. Its `ShardRuntime` service has exactly five RPCs —
|
||||||
|
`GetCapability`, `Health`, `Session` (bidi stream), `Release`, `Cancel`.
|
||||||
|
- `packages/node/meshnet_node/shard_runtime_server.py` (DGR-024): the reference
|
||||||
|
Python servicer. It performs a *bounded real forward* (a CRC over the received
|
||||||
|
bundle bytes) then echoes the chunk, and fails closed on stale epoch, expired
|
||||||
|
deadline, corrupt/mis-tiled fragments, exhausted flow-control credit, duplicate
|
||||||
|
idempotency step, and in-band/out-of-band cancellation, with per-`route_session_id`
|
||||||
|
state kept on the servicer so an out-of-band `Cancel` can reach a live session.
|
||||||
|
**Key finding:** despite the schema labelling the checksum `CRC32C`, this
|
||||||
|
runtime computes it with `zlib.crc32` (standard CRC-32, *not* Castagnoli). The
|
||||||
|
C++ worker mirrors `zlib.crc32` exactly so its checksum acceptance is
|
||||||
|
byte-identical to the existing Python surface (the committed C++ *conformance*
|
||||||
|
test, by contrast, uses true Castagnoli against separately-generated goldens —
|
||||||
|
the two are unrelated code paths).
|
||||||
|
- `packages/node/native/CMakeLists.txt` (DGR-029/030): configures against the
|
||||||
|
ignored `build/native-toolchain` prefix (pinned Protobuf 33.1 + gRPC 1.82.1),
|
||||||
|
always generates both message and service stubs, and registers a C++
|
||||||
|
conformance CTest. There was **no** worker executable and **no** Python
|
||||||
|
worker integration test before this story (confirmed by
|
||||||
|
`ls packages/node/native/worker` → absent, and grep for `shard_worker`).
|
||||||
|
- `packages/node/meshnet_node/fake_shard_engine.py` (DGR-032): the Python fake
|
||||||
|
engine, deliberately *not* wired into the gRPC surface. DGR-033's worker is
|
||||||
|
its native analogue — a separate executable, not a consumer of that module —
|
||||||
|
so both fakes present identical behaviour to a client (deterministic,
|
||||||
|
model-free bounded forward; per-session isolation; fail-closed lifecycle).
|
||||||
|
|
||||||
|
## What was added (this story's change)
|
||||||
|
|
||||||
|
### `packages/node/native/worker/fake_engine.h` (new)
|
||||||
|
|
||||||
|
`meshnet::worker::FakeShardEngine` — a header-only, model-free fixture engine.
|
||||||
|
Its only capability is to validate a `TensorBundle` (fragments tile exactly, the
|
||||||
|
uncompressed CRC-32 matches the declared checksum, the declared payload stays
|
||||||
|
within the negotiated `max_chunk_bytes`) and fold the fragment bytes through a
|
||||||
|
bounded forward. It links, loads, and dispatches to **nothing** — no llama.cpp,
|
||||||
|
no graph execution. Carries `kEvidenceClass = "fixture"` mirroring the Python
|
||||||
|
`FakeShardEngine.EVIDENCE_CLASS` for the later DGR-036 parity check.
|
||||||
|
|
||||||
|
### `packages/node/native/worker/shard_service.{h,cpp}` (new)
|
||||||
|
|
||||||
|
`ShardRuntimeServiceImpl : meshnet::shard::v1::ShardRuntime::Service` — a faithful
|
||||||
|
C++ port of the DGR-024 Python servicer: the same per-`route_session_id`
|
||||||
|
identity/credit/dedup state guarded by a mutex, the same fail-closed negative
|
||||||
|
paths, and the same lifecycle (open → prefill/decode → flow-control top-up →
|
||||||
|
release/cancel). Each per-request response is computed under the lock and written
|
||||||
|
*after* releasing it, so a blocking `Write` can never deadlock the out-of-band
|
||||||
|
`Cancel` RPC that needs the same lock. Bounded messages are enforced two ways: a
|
||||||
|
per-tensor `RESOURCE_EXHAUSTED` app check against `max_chunk_bytes`, plus a hard
|
||||||
|
transport receive ceiling.
|
||||||
|
|
||||||
|
### `packages/node/native/worker/shard_worker_main.cpp` (new)
|
||||||
|
|
||||||
|
The standalone `shard_worker` executable. Binds `MESHNET_SHARD_LISTEN_ADDR`
|
||||||
|
(or an `argv` address), prints one readiness line (`ShardRuntime worker listening
|
||||||
|
on <addr>`), and serves until `SIGTERM`/`SIGINT`. **Graceful shutdown** uses a
|
||||||
|
self-pipe: the async-signal-safe handler writes one byte, a drain thread reads it
|
||||||
|
and calls `server->Shutdown()`, so in-flight sessions finish and the process
|
||||||
|
exits `0` printing `ShardRuntime worker shut down cleanly`. A `--selftest` mode
|
||||||
|
binds an ephemeral port and self-drives capability/health/fragmented-prefill/
|
||||||
|
decode/release over a real loopback gRPC channel, giving a pure-C++ CTest that
|
||||||
|
needs no Python.
|
||||||
|
|
||||||
|
### `packages/node/native/CMakeLists.txt` (modified)
|
||||||
|
|
||||||
|
Adds the `shard_worker` executable (linking only `shard_runtime_grpc` +
|
||||||
|
`gRPC::grpc++` — no llama.cpp) and registers `shard_worker_selftest` as a CTest.
|
||||||
|
|
||||||
|
### `tests/test_native_shard_worker.py` (new)
|
||||||
|
|
||||||
|
18 integration tests that spawn the **real compiled binary** as a subprocess and
|
||||||
|
drive it with the committed generated stubs over a real localhost socket. When
|
||||||
|
the binary is not built they skip (the DGR-029/030 `requires_cmake` gating
|
||||||
|
pattern), locating it via `MESHNET_SHARD_WORKER_BIN` or `build/native/shard_worker`.
|
||||||
|
|
||||||
|
## Acceptance criteria → evidence
|
||||||
|
|
||||||
|
1. **Standalone C++ executable serves the complete lifecycle/stream contract
|
||||||
|
using the fake engine** — `shard_worker` builds and serves all five RPCs; the
|
||||||
|
`shard_worker_selftest` CTest drives open → fragmented prefill → decode →
|
||||||
|
release over real gRPC; the 18 Python tests cover the same against the
|
||||||
|
subprocess.
|
||||||
|
2. **Python integration tests cover startup, health, capability, fragmented
|
||||||
|
prefill, decode, release, cancellation, graceful shutdown** —
|
||||||
|
`test_worker_startup_and_health`, `test_worker_capability`,
|
||||||
|
`test_fragmented_prefill_echoes_reassembled_payload` (3-fragment tiling),
|
||||||
|
`test_decode_step_is_served`, `test_release_is_terminal`,
|
||||||
|
`test_in_band_cancel_of_single_work_item_does_not_end_stream`,
|
||||||
|
`test_in_band_cancel_of_whole_session_is_terminal`,
|
||||||
|
`test_out_of_band_cancel_rpc_races_ahead_of_open`,
|
||||||
|
`test_graceful_shutdown_on_sigterm` (SIGTERM → exit 0 + clean-shutdown line).
|
||||||
|
3. **Bounded messages, deadlines, flow control, independent session
|
||||||
|
cancellation enforced** — `test_bounded_message_is_rejected`
|
||||||
|
(`RESOURCE_EXHAUSTED` on an over-ceiling tensor),
|
||||||
|
`test_expired_deadline_is_rejected`, `test_flow_control_violation_and_topup`,
|
||||||
|
`test_independent_session_cancellation` (cancelling session A leaves session B
|
||||||
|
fully serviceable), plus `test_stale_route_epoch_is_rejected`,
|
||||||
|
`test_duplicate_idempotency_step_is_acked`,
|
||||||
|
`test_malformed_fragment_tiling_is_rejected`.
|
||||||
|
4. **Exposes neither llama.cpp RPC nor arbitrary graph execution** —
|
||||||
|
`ldd build/native/shard_worker` shows no llama/ggml shared libs;
|
||||||
|
`nm -C build/native/shard_worker | grep -icE 'llama_|ggml_'` → `0`; the proto
|
||||||
|
exposes exactly one service with five lifecycle RPCs and no graph-exec entry.
|
||||||
|
5. **Gates + this handoff** — below.
|
||||||
|
|
||||||
|
## Commands and results
|
||||||
|
|
||||||
|
Toolchain (ignored `build/native-toolchain`, pinned Protobuf 33.1 + gRPC 1.82.1):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bash scripts/bootstrap_native_toolchain.sh "$PWD/build/native-toolchain"
|
||||||
|
# ... gRPC 1.82.1 commit acccf84c0df20487d64101f528e5d426541ca4e5
|
||||||
|
# grpc_cpp_plugin sha256 43705cf26ae9ce98bbcee76b3408f5e171eec746b50bf0dd42dd68d132c6a533
|
||||||
|
```
|
||||||
|
|
||||||
|
Focused out-of-tree CMake build + CTest:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cmake -S packages/node/native -B build/native -DCMAKE_PREFIX_PATH="$PWD/build/native-toolchain"
|
||||||
|
cmake --build build/native -j"$(nproc)"
|
||||||
|
ctest --test-dir build/native --output-on-failure
|
||||||
|
```
|
||||||
|
```text
|
||||||
|
1/2 Test #1: shard_worker_selftest ............ Passed 0.01 sec
|
||||||
|
2/2 Test #2: shard_protocol_conformance ....... Passed 0.00 sec
|
||||||
|
100% tests passed out of 2
|
||||||
|
```
|
||||||
|
|
||||||
|
Python integration tests against the real binary:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
PYTHONPATH=packages/node:packages/tracker python -m pytest -q tests/test_native_shard_worker.py
|
||||||
|
```
|
||||||
|
```text
|
||||||
|
18 passed in 3.96s
|
||||||
|
```
|
||||||
|
|
||||||
|
AC4 (no llama.cpp / no graph exec):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ldd build/native/shard_worker | grep -iE 'llama|ggml' # -> (no matches)
|
||||||
|
nm build/native/shard_worker | grep -icE 'llama_|ggml_' # -> 0
|
||||||
|
```
|
||||||
|
|
||||||
|
Shared gates + regression:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m compileall -q packages tests # exit 0
|
||||||
|
git diff --check -- packages/node/native tests/test_native_shard_worker.py # exit 0
|
||||||
|
PYTHONPATH=packages/node:packages/tracker python -m pytest -q \
|
||||||
|
tests/test_shard_runtime_harness.py tests/test_native_shard_protocol.py
|
||||||
|
# -> 61 passed, 2 skipped (DGR-024 harness + native protocol untouched)
|
||||||
|
```
|
||||||
|
|
||||||
|
Toolchain used: `cmake`/`ctest` from the `distributed-gguf-runtime` worktree's
|
||||||
|
`.venv` (PyPI `cmake==4.4.0` wheel — no system cmake exists here, same as
|
||||||
|
DGR-029/030); the Python client uses that venv's `grpcio==1.82.1`,
|
||||||
|
`grpcio-tools==1.82.1`, `protobuf`, `pytest`. `g++ (GCC) 15.2.1`.
|
||||||
|
|
||||||
|
## Limitations
|
||||||
|
|
||||||
|
- This is FIXTURE evidence only. The worker's "forward" is a CRC-over-wire-bytes
|
||||||
|
echo, not real tensor compute; it proves process/stream/lifecycle/supervision
|
||||||
|
shape, nothing about numerical correctness. Real engine binding is DGR-037 and
|
||||||
|
numeric parity is DGR-036/052.
|
||||||
|
- The worker checksum path mirrors the DGR-024 runtime's `zlib.crc32` (standard
|
||||||
|
CRC-32 under a `CRC32C` label). Compressed-tensor tiling/checksum is not
|
||||||
|
independently verified (no zstd decompressor in the fixture) — identical to the
|
||||||
|
DGR-024 limitation.
|
||||||
|
- Default `pytest` runs skip `tests/test_native_shard_worker.py` unless the
|
||||||
|
worker binary is built (or `MESHNET_SHARD_WORKER_BIN` is set); this session
|
||||||
|
built it and ran all 18 for real (results above). Building requires the pinned
|
||||||
|
gRPC C++ toolchain, which is not present by default and must be bootstrapped.
|
||||||
|
- No CUDA/ROCm/GPU, no model download, no network at test time — all default
|
||||||
|
tests are fixture-only and offline.
|
||||||
|
|
||||||
|
## Dependency handoff
|
||||||
|
|
||||||
|
- **DGR-036** (fixture vs real-model parity): the worker's `FakeShardEngine`
|
||||||
|
carries `kEvidenceClass = "fixture"`; diff it against DGR-037's real engine's
|
||||||
|
equivalent marker, and reuse the same lifecycle/stream contract this worker
|
||||||
|
serves to prove behavioural parity before numeric parity.
|
||||||
|
- **DGR-037** (bind llama.cpp): replace `FakeShardEngine`'s bounded forward with
|
||||||
|
the real engine behind the *same* `ShardRuntimeServiceImpl` surface; the
|
||||||
|
service's session/epoch/credit/dedup/cancel machinery and the graceful-shutdown
|
||||||
|
supervision shape are reusable as-is.
|
||||||
|
- **DGR-040** (worker supervision): `shard_worker` already provides the
|
||||||
|
supervision primitives — a readiness line for start detection, `SIGTERM`
|
||||||
|
graceful drain with a clean-exit line, and a `--selftest` liveness probe.
|
||||||
|
A supervisor can start/monitor/restart the process around these.
|
||||||
@@ -1,7 +1,7 @@
|
|||||||
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
|
<!-- GENERATED FROM prd.json — DO NOT EDIT AS AN INDEPENDENT SOURCE. prd.json IS AUTHORITATIVE. -->
|
||||||
# DGR-033: Build a standalone fake C++ gRPC Shard worker
|
# DGR-033: Build a standalone fake C++ gRPC Shard worker
|
||||||
|
|
||||||
- **Status / triage:** specification only; `ready-for-agent`; `passes: false`
|
- **Status / triage:** completed; `passes: true`
|
||||||
- **Execution mode:** `AFK`
|
- **Execution mode:** `AFK`
|
||||||
- **Milestone:** `M1`
|
- **Milestone:** `M1`
|
||||||
- **Dependencies:** `DGR-022`, `DGR-024`, `DGR-032`
|
- **Dependencies:** `DGR-022`, `DGR-024`, `DGR-032`
|
||||||
@@ -18,11 +18,11 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
|||||||
|
|
||||||
## Acceptance criteria
|
## Acceptance criteria
|
||||||
|
|
||||||
- [ ] A standalone C++ executable serves the complete lifecycle and stream RPC contract using the fake engine.
|
- [x] A standalone C++ executable serves the complete lifecycle and stream RPC contract using the fake engine.
|
||||||
- [ ] Python integration tests cover startup, health, capability, fragmented prefill, decode, release, cancellation, and graceful shutdown.
|
- [x] Python integration tests cover startup, health, capability, fragmented prefill, decode, release, cancellation, and graceful shutdown.
|
||||||
- [ ] Bounded messages, deadlines, flow control, and independent session cancellation are enforced.
|
- [x] Bounded messages, deadlines, flow control, and independent session cancellation are enforced.
|
||||||
- [ ] The worker exposes neither llama.cpp RPC nor arbitrary graph execution.
|
- [x] The worker exposes neither llama.cpp RPC nor arbitrary graph execution.
|
||||||
- [ ] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
- [x] Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff.
|
||||||
|
|
||||||
## Shared quality gates
|
## Shared quality gates
|
||||||
|
|
||||||
@@ -30,10 +30,7 @@ Fresh Ralph session: read `.scratch/distributed-gguf-runtime/RALPH-CONTEXT.md`,
|
|||||||
- `git diff --check` passes.
|
- `git diff --check` passes.
|
||||||
- Default tests are model-download-free, API-credit-free, and GPU-free.
|
- Default tests are model-download-free, API-credit-free, and GPU-free.
|
||||||
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
|
- Evidence README records exact changed files, commands/results, limitations, and dependency handoff; no fabricated evidence or inherited completion credit.
|
||||||
- Native changes pass focused out-of-tree CMake build and CTest; patch changes verify clean apply/check/reverse against the exact llama.cpp pin.
|
|
||||||
- Runs are opt-in and record exact artifact/split hashes, runtime/upstream pin, backend/driver, hardware, network, commands, and raw metrics. Model artifacts use configured mounted-drive storage and never `/home`.
|
|
||||||
- Preserve existing Transformers behavior and backend-agnostic Tracker routing/load balancing/billing/relay semantics unless an explicit versioned contract says otherwise. One scoped story commit is expected during execution, but this specification-materialization change is not committed.
|
|
||||||
|
|
||||||
## Evidence handoff
|
## Evidence handoff
|
||||||
|
|
||||||
Write and verify `.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md`. Until every criterion and applicable gate has real evidence, this story remains `passes: false`. Legacy evidence is provenance only, not completion credit.
|
Verified evidence: `.scratch/distributed-gguf-runtime/evidence/DGR-033/README.md`. Legacy evidence remains provenance only and grants no implementation completion credit.
|
||||||
|
|||||||
@@ -656,12 +656,13 @@
|
|||||||
"The worker exposes neither llama.cpp RPC nor arbitrary graph execution.",
|
"The worker exposes neither llama.cpp RPC nor arbitrary graph execution.",
|
||||||
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
"Applicable shared quality gates in `prd.json` pass, and the evidence handoff records exact commands/results, changed files, limitations, and dependency handoff."
|
||||||
],
|
],
|
||||||
"passes": false,
|
"passes": true,
|
||||||
"notes": "Generated source issue: .scratch/distributed-gguf-runtime/issues/033-build-a-standalone-fake-c-grpc-shard-worker.md; prd.json is authoritative.",
|
"notes": "Generated source issue: .scratch/distributed-gguf-runtime/issues/033-build-a-standalone-fake-c-grpc-shard-worker.md; prd.json is authoritative.",
|
||||||
"blocks": [
|
"blocks": [
|
||||||
"DGR-036",
|
"DGR-036",
|
||||||
"DGR-040"
|
"DGR-040"
|
||||||
]
|
],
|
||||||
|
"completionNotes": "Completed by agent"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"id": "DGR-034",
|
"id": "DGR-034",
|
||||||
|
|||||||
@@ -62,6 +62,21 @@ message(STATUS "Pinned gRPC ${gRPC_VERSION}: building ShardRuntime service stubs
|
|||||||
|
|
||||||
enable_testing()
|
enable_testing()
|
||||||
|
|
||||||
|
# The standalone fake Shard worker (DGR-033): a real gRPC server over the
|
||||||
|
# ShardRuntime service, backed by the model-free FakeShardEngine. It links the
|
||||||
|
# grpc service stubs only — no llama.cpp, no graph-execution entry point.
|
||||||
|
add_executable(shard_worker
|
||||||
|
worker/shard_worker_main.cpp
|
||||||
|
worker/shard_service.cpp)
|
||||||
|
target_include_directories(shard_worker PRIVATE "${CMAKE_CURRENT_SOURCE_DIR}/worker")
|
||||||
|
target_link_libraries(shard_worker PRIVATE shard_runtime_grpc gRPC::grpc++)
|
||||||
|
|
||||||
|
# Pure-C++ CTest: the worker binds an ephemeral port, self-drives the full
|
||||||
|
# lifecycle (capability, health, fragmented prefill, decode, release) over a
|
||||||
|
# real loopback gRPC channel, and exits non-zero on any mismatch. This proves
|
||||||
|
# the worker serves the contract without needing a Python environment.
|
||||||
|
add_test(NAME shard_worker_selftest COMMAND shard_worker --selftest)
|
||||||
|
|
||||||
add_executable(shard_protocol_conformance tests/test_shard_protocol_conformance.cpp)
|
add_executable(shard_protocol_conformance tests/test_shard_protocol_conformance.cpp)
|
||||||
target_link_libraries(shard_protocol_conformance PRIVATE shard_runtime_proto)
|
target_link_libraries(shard_protocol_conformance PRIVATE shard_runtime_proto)
|
||||||
|
|
||||||
|
|||||||
164
packages/node/native/worker/fake_engine.h
Normal file
164
packages/node/native/worker/fake_engine.h
Normal file
@@ -0,0 +1,164 @@
|
|||||||
|
// Deterministic, model-free fake ShardEngine for the native worker (DGR-033).
|
||||||
|
//
|
||||||
|
// This is the C++ analogue of `meshnet_node.fake_shard_engine.FakeShardEngine`
|
||||||
|
// (DGR-032): a pure fixture that performs a *bounded real forward* over the
|
||||||
|
// bytes it received off the socket and never links, loads, or dispatches to
|
||||||
|
// llama.cpp. It exists to prove the standalone worker process, stream,
|
||||||
|
// lifecycle, and supervision shape before any real engine is bound (DGR-037).
|
||||||
|
//
|
||||||
|
// The "forward" is deliberately transport-verifiable rather than semantic: it
|
||||||
|
// reassembles a tensor's fragments, checks they tile exactly, and derives a
|
||||||
|
// CRC32C over the uncompressed bytes — the same rule the schema's `Checksum`
|
||||||
|
// declares and the same bounded forward the DGR-024 Python surface performs.
|
||||||
|
// Feeding the same bytes back (echo) lets a client prove the payload truly
|
||||||
|
// traversed the wire and returned unmodified; a direct hop and an opaque relay
|
||||||
|
// of the identical frames therefore yield byte-identical responses.
|
||||||
|
//
|
||||||
|
// There is no arbitrary-graph entry point here and no llama.cpp RPC: the engine
|
||||||
|
// only knows how to reassemble/checksum a bundle. That is the whole point of a
|
||||||
|
// fixture worker (acceptance criterion 4).
|
||||||
|
|
||||||
|
#ifndef MESHNET_NATIVE_WORKER_FAKE_ENGINE_H_
|
||||||
|
#define MESHNET_NATIVE_WORKER_FAKE_ENGINE_H_
|
||||||
|
|
||||||
|
#include <algorithm>
|
||||||
|
#include <cstdint>
|
||||||
|
#include <optional>
|
||||||
|
#include <string>
|
||||||
|
#include <vector>
|
||||||
|
|
||||||
|
#include "shard_runtime.pb.h"
|
||||||
|
|
||||||
|
namespace meshnet::worker {
|
||||||
|
|
||||||
|
namespace sp = ::meshnet::shard::v1;
|
||||||
|
|
||||||
|
// Standard CRC-32 (ISO-HDLC / zlib polynomial 0xEDB88320, reflected).
|
||||||
|
//
|
||||||
|
// The schema's `Checksum` field is labelled CRC32C, but the DGR-024 Python
|
||||||
|
// runtime surface (`shard_runtime_server.py`) computes it with `zlib.crc32`
|
||||||
|
// (standard CRC-32, not the Castagnoli CRC32C). This worker deliberately mirrors
|
||||||
|
// that exact computation so its checksum acceptance is byte-for-byte identical
|
||||||
|
// to the existing Python gRPC surface and to a relayed frame's expectations.
|
||||||
|
inline uint32_t Crc32(const std::string& data, uint32_t seed = 0) {
|
||||||
|
static uint32_t table[256];
|
||||||
|
static bool built = false;
|
||||||
|
if (!built) {
|
||||||
|
for (uint32_t i = 0; i < 256; ++i) {
|
||||||
|
uint32_t c = i;
|
||||||
|
for (int k = 0; k < 8; ++k) {
|
||||||
|
c = (c & 1) ? (c >> 1) ^ 0xEDB88320u : (c >> 1);
|
||||||
|
}
|
||||||
|
table[i] = c;
|
||||||
|
}
|
||||||
|
built = true;
|
||||||
|
}
|
||||||
|
uint32_t crc = seed ^ 0xFFFFFFFFu;
|
||||||
|
for (unsigned char byte : data) {
|
||||||
|
crc = (crc >> 8) ^ table[(crc ^ byte) & 0xFF];
|
||||||
|
}
|
||||||
|
return crc ^ 0xFFFFFFFFu;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Outcome of validating one bundle before the bounded forward runs.
|
||||||
|
struct BundleCheck {
|
||||||
|
// Set when the bundle is malformed/corrupt (maps to PAYLOAD_CORRUPT).
|
||||||
|
std::optional<std::string> corrupt_detail;
|
||||||
|
// Set when the declared payload exceeds the negotiated per-chunk ceiling
|
||||||
|
// (maps to RESOURCE_EXHAUSTED) — the worker refuses unbounded messages.
|
||||||
|
std::optional<std::string> oversize_detail;
|
||||||
|
};
|
||||||
|
|
||||||
|
// The fake engine's only capability: verify a bundle tiles and checksums, and
|
||||||
|
// that it stays within the negotiated byte ceiling. Mirrors `_validate_bundle`
|
||||||
|
// in `shard_runtime_server.py` plus the bounded-message rule DGR-033 adds.
|
||||||
|
class FakeShardEngine {
|
||||||
|
public:
|
||||||
|
// Marker mirroring `FakeShardEngine.EVIDENCE_CLASS` so a future parity check
|
||||||
|
// (DGR-036) can assert this is a fixture, not a real engine.
|
||||||
|
static constexpr const char* kEvidenceClass = "fixture";
|
||||||
|
|
||||||
|
explicit FakeShardEngine(uint64_t max_chunk_bytes) : max_chunk_bytes_(max_chunk_bytes) {}
|
||||||
|
|
||||||
|
BundleCheck Validate(const sp::TensorBundle& bundle) const {
|
||||||
|
BundleCheck result;
|
||||||
|
for (const auto& tensor : bundle.tensors()) {
|
||||||
|
// Bounded message: a declared payload larger than the ceiling is refused
|
||||||
|
// before any reassembly work is done.
|
||||||
|
if (max_chunk_bytes_ != 0 && tensor.total_bytes() > max_chunk_bytes_) {
|
||||||
|
result.oversize_detail =
|
||||||
|
"tensor '" + tensor.name() + "': declared total_bytes " +
|
||||||
|
std::to_string(tensor.total_bytes()) + " exceeds max_chunk_bytes " +
|
||||||
|
std::to_string(max_chunk_bytes_);
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Fragments must tile the wire body exactly: no hole, no overlap.
|
||||||
|
std::vector<const sp::TensorFragment*> ordered;
|
||||||
|
ordered.reserve(tensor.fragments_size());
|
||||||
|
for (const auto& fragment : tensor.fragments()) {
|
||||||
|
ordered.push_back(&fragment);
|
||||||
|
}
|
||||||
|
std::sort(ordered.begin(), ordered.end(),
|
||||||
|
[](const sp::TensorFragment* a, const sp::TensorFragment* b) {
|
||||||
|
return a->byte_offset() < b->byte_offset();
|
||||||
|
});
|
||||||
|
uint64_t expected_offset = 0;
|
||||||
|
std::string payload;
|
||||||
|
for (const auto* fragment : ordered) {
|
||||||
|
if (fragment->byte_offset() != expected_offset) {
|
||||||
|
result.corrupt_detail =
|
||||||
|
"tensor '" + tensor.name() + "': fragment at offset " +
|
||||||
|
std::to_string(fragment->byte_offset()) +
|
||||||
|
" does not tile the preceding " + std::to_string(expected_offset) +
|
||||||
|
" bytes (gap or overlap)";
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
payload.append(fragment->payload());
|
||||||
|
expected_offset += fragment->payload().size();
|
||||||
|
}
|
||||||
|
if (tensor.compression() == sp::COMPRESSION_NONE &&
|
||||||
|
expected_offset != tensor.total_bytes()) {
|
||||||
|
result.corrupt_detail =
|
||||||
|
"tensor '" + tensor.name() + "': fragments cover " +
|
||||||
|
std::to_string(expected_offset) + " bytes, declared total_bytes is " +
|
||||||
|
std::to_string(tensor.total_bytes());
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
if (tensor.compression() == sp::COMPRESSION_NONE &&
|
||||||
|
tensor.checksum().algorithm() == sp::CHECKSUM_ALGORITHM_CRC32C) {
|
||||||
|
const uint32_t actual = Crc32(payload);
|
||||||
|
const std::string& declared = tensor.checksum().value();
|
||||||
|
std::string actual_be(4, '\0');
|
||||||
|
actual_be[0] = static_cast<char>((actual >> 24) & 0xFF);
|
||||||
|
actual_be[1] = static_cast<char>((actual >> 16) & 0xFF);
|
||||||
|
actual_be[2] = static_cast<char>((actual >> 8) & 0xFF);
|
||||||
|
actual_be[3] = static_cast<char>(actual & 0xFF);
|
||||||
|
if (declared != actual_be) {
|
||||||
|
result.corrupt_detail = "tensor '" + tensor.name() + "': checksum mismatch";
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Bounded real forward: fold every fragment's payload through CRC32C so the
|
||||||
|
// digest is only reproducible if the payload really traversed the wire.
|
||||||
|
uint32_t BoundedForward(const sp::TensorBundle& bundle) const {
|
||||||
|
uint32_t digest = 0;
|
||||||
|
for (const auto& tensor : bundle.tensors()) {
|
||||||
|
for (const auto& fragment : tensor.fragments()) {
|
||||||
|
digest = Crc32(fragment.payload(), digest);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return digest;
|
||||||
|
}
|
||||||
|
|
||||||
|
private:
|
||||||
|
uint64_t max_chunk_bytes_;
|
||||||
|
};
|
||||||
|
|
||||||
|
} // namespace meshnet::worker
|
||||||
|
|
||||||
|
#endif // MESHNET_NATIVE_WORKER_FAKE_ENGINE_H_
|
||||||
364
packages/node/native/worker/shard_service.cpp
Normal file
364
packages/node/native/worker/shard_service.cpp
Normal file
@@ -0,0 +1,364 @@
|
|||||||
|
#include "shard_service.h"
|
||||||
|
|
||||||
|
#include <chrono>
|
||||||
|
#include <utility>
|
||||||
|
|
||||||
|
namespace meshnet::worker {
|
||||||
|
|
||||||
|
namespace {
|
||||||
|
|
||||||
|
int64_t NowUnixNanos() {
|
||||||
|
return std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||||
|
std::chrono::system_clock::now().time_since_epoch())
|
||||||
|
.count();
|
||||||
|
}
|
||||||
|
|
||||||
|
// Build the standard fail response (a terminal-or-not ShardStatus).
|
||||||
|
sp::SessionResponse MakeFail(const std::string& route_session_id, const std::string& work_id,
|
||||||
|
uint64_t step, sp::ErrorCode code, const std::string& detail,
|
||||||
|
bool terminal, bool retryable) {
|
||||||
|
sp::SessionResponse response;
|
||||||
|
sp::ShardStatus* status = response.mutable_status();
|
||||||
|
status->set_work_id(work_id);
|
||||||
|
status->set_route_session_id(route_session_id);
|
||||||
|
status->set_idempotency_step(step);
|
||||||
|
status->set_terminal(terminal);
|
||||||
|
sp::ShardError* error = status->mutable_error();
|
||||||
|
error->set_code(code);
|
||||||
|
error->set_detail(detail);
|
||||||
|
error->set_retryable(retryable);
|
||||||
|
return response;
|
||||||
|
}
|
||||||
|
|
||||||
|
sp::SessionResponse MakeAck(const std::string& work_id, uint64_t step, bool duplicate) {
|
||||||
|
sp::SessionResponse response;
|
||||||
|
sp::Ack* ack = response.mutable_ack();
|
||||||
|
ack->set_work_id(work_id);
|
||||||
|
ack->set_idempotency_step(step);
|
||||||
|
ack->set_duplicate(duplicate);
|
||||||
|
return response;
|
||||||
|
}
|
||||||
|
|
||||||
|
void FillDefaultFlow(sp::FlowControl* fc, const FlowLimits& limits) {
|
||||||
|
fc->set_credits_granted(limits.credits_granted);
|
||||||
|
fc->set_max_inflight_chunks(limits.max_inflight_chunks);
|
||||||
|
fc->set_max_chunk_bytes(limits.max_chunk_bytes);
|
||||||
|
fc->set_max_prefill_chunk_tokens(limits.max_prefill_chunk_tokens);
|
||||||
|
}
|
||||||
|
|
||||||
|
} // namespace
|
||||||
|
|
||||||
|
grpc::Status ShardRuntimeServiceImpl::GetCapability(grpc::ServerContext*,
|
||||||
|
const sp::CapabilityRequest*,
|
||||||
|
sp::CapabilityReport* response) {
|
||||||
|
response->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
sp::Fingerprint* fp = response->mutable_fingerprint();
|
||||||
|
fp->set_model_artifact_digest("sha256:native-test-artifact");
|
||||||
|
fp->set_runtime_recipe_digest("sha256:native-test-recipe");
|
||||||
|
fp->set_recipe_id("native-test");
|
||||||
|
fp->set_recipe_version("1");
|
||||||
|
fp->set_catalogue_version("1");
|
||||||
|
sp::ShardRange* range = response->mutable_shard_range();
|
||||||
|
range->set_start_layer(0);
|
||||||
|
range->set_end_layer(32);
|
||||||
|
range->set_effective_start_layer(0);
|
||||||
|
response->set_backend("grpc-native-cpp");
|
||||||
|
response->set_device("cpu");
|
||||||
|
response->set_validated(true);
|
||||||
|
response->set_detail("bounded real forward passed for fixture artifact");
|
||||||
|
response->set_max_concurrent_sessions(8);
|
||||||
|
response->set_max_context_tokens(131072);
|
||||||
|
FillDefaultFlow(response->mutable_flow_control(), limits_);
|
||||||
|
response->add_accepted_compression(sp::COMPRESSION_NONE);
|
||||||
|
response->add_supported_schema_versions(sp::SCHEMA_VERSION_1);
|
||||||
|
response->set_validated_at_unix_nanos(0);
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
|
||||||
|
grpc::Status ShardRuntimeServiceImpl::Health(grpc::ServerContext*, const sp::HealthRequest*,
|
||||||
|
sp::HealthReport* response) {
|
||||||
|
response->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
response->set_state(sp::SERVING_STATE_SERVING);
|
||||||
|
response->set_active_sessions(1);
|
||||||
|
response->set_queued_chunks(0);
|
||||||
|
response->set_batch_occupancy(0);
|
||||||
|
response->set_kv_pressure(0.0f);
|
||||||
|
response->set_resident_bytes(0);
|
||||||
|
response->set_detail("native fixture worker serving");
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
|
||||||
|
uint32_t ShardRuntimeServiceImpl::MarkCancelled(const std::string& route_session_id,
|
||||||
|
const std::string& work_id) {
|
||||||
|
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||||
|
SessionState& state = sessions_[route_session_id]; // creates on first cancel-before-open
|
||||||
|
if (state.max_inflight == 0) {
|
||||||
|
// Freshly created placeholder for a Cancel that raced ahead of Open.
|
||||||
|
state.credits = limits_.credits_granted;
|
||||||
|
state.max_inflight = limits_.max_inflight_chunks;
|
||||||
|
state.max_chunk_bytes = limits_.max_chunk_bytes;
|
||||||
|
}
|
||||||
|
if (work_id.empty()) {
|
||||||
|
const bool already = state.cancelled_session;
|
||||||
|
state.cancelled_session = true;
|
||||||
|
return already ? 0 : 1;
|
||||||
|
}
|
||||||
|
const bool already = state.cancelled_work.count(work_id) != 0;
|
||||||
|
state.cancelled_work.insert(work_id);
|
||||||
|
return already ? 0 : 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
grpc::Status ShardRuntimeServiceImpl::Session(
|
||||||
|
grpc::ServerContext*,
|
||||||
|
grpc::ServerReaderWriter<sp::SessionResponse, sp::SessionRequest>* stream) {
|
||||||
|
std::string route_session_id;
|
||||||
|
sp::SessionRequest request;
|
||||||
|
|
||||||
|
while (stream->Read(&request)) {
|
||||||
|
switch (request.kind_case()) {
|
||||||
|
case sp::SessionRequest::kOpen: {
|
||||||
|
const sp::SessionOpen& open = request.open();
|
||||||
|
route_session_id = open.route_session_id();
|
||||||
|
{
|
||||||
|
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||||
|
SessionState state;
|
||||||
|
state.epoch = open.route_epoch();
|
||||||
|
if (open.has_proposed_flow_control()) {
|
||||||
|
const sp::FlowControl& fc = open.proposed_flow_control();
|
||||||
|
state.credits = fc.credits_granted();
|
||||||
|
state.max_inflight = fc.max_inflight_chunks();
|
||||||
|
state.max_chunk_bytes = fc.max_chunk_bytes();
|
||||||
|
} else {
|
||||||
|
state.credits = limits_.credits_granted;
|
||||||
|
state.max_inflight = limits_.max_inflight_chunks;
|
||||||
|
state.max_chunk_bytes = limits_.max_chunk_bytes;
|
||||||
|
}
|
||||||
|
auto it = sessions_.find(route_session_id);
|
||||||
|
if (it != sessions_.end()) {
|
||||||
|
// A prior out-of-band Cancel may have marked this session cancelled
|
||||||
|
// before Open arrived; preserve that so the work still fails closed.
|
||||||
|
state.cancelled_session = it->second.cancelled_session;
|
||||||
|
state.cancelled_work = it->second.cancelled_work;
|
||||||
|
}
|
||||||
|
sessions_[route_session_id] = std::move(state);
|
||||||
|
}
|
||||||
|
sp::SessionResponse response;
|
||||||
|
sp::SessionAccepted* accepted = response.mutable_accepted();
|
||||||
|
accepted->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
accepted->set_route_session_id(open.route_session_id());
|
||||||
|
accepted->set_route_epoch(open.route_epoch());
|
||||||
|
if (open.has_proposed_flow_control()) {
|
||||||
|
*accepted->mutable_flow_control() = open.proposed_flow_control();
|
||||||
|
} else {
|
||||||
|
FillDefaultFlow(accepted->mutable_flow_control(), limits_);
|
||||||
|
}
|
||||||
|
if (open.accepted_compression_size() > 0) {
|
||||||
|
for (int c : open.accepted_compression()) {
|
||||||
|
accepted->add_accepted_compression(static_cast<sp::Compression>(c));
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
accepted->add_accepted_compression(sp::COMPRESSION_NONE);
|
||||||
|
}
|
||||||
|
*accepted->mutable_fingerprint() = open.fingerprint();
|
||||||
|
stream->Write(response);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case sp::SessionRequest::kChunk: {
|
||||||
|
const sp::ActivationChunk& chunk = request.chunk();
|
||||||
|
const sp::Envelope& envelope = chunk.envelope();
|
||||||
|
const std::string work_id = envelope.work_id();
|
||||||
|
const uint64_t step = envelope.idempotency_step();
|
||||||
|
|
||||||
|
// Compute the response under the lock, then write it *after* releasing —
|
||||||
|
// holding the lock across a (possibly blocking) Write would deadlock an
|
||||||
|
// out-of-band Cancel RPC that needs the same lock.
|
||||||
|
sp::SessionResponse response;
|
||||||
|
{
|
||||||
|
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||||
|
auto it = sessions_.find(route_session_id);
|
||||||
|
SessionState* state = it != sessions_.end() ? &it->second : nullptr;
|
||||||
|
|
||||||
|
if (state && (state->cancelled_session || state->cancelled_work.count(work_id))) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_CANCELLED,
|
||||||
|
"work was cancelled", false, false);
|
||||||
|
} else if (state && envelope.route_epoch() < state->epoch) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_EPOCH_STALE,
|
||||||
|
"stale route epoch", false, false);
|
||||||
|
} else if (envelope.deadline_unix_nanos() != 0 &&
|
||||||
|
NowUnixNanos() > envelope.deadline_unix_nanos()) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_DEADLINE_EXCEEDED,
|
||||||
|
"deadline already passed", false, false);
|
||||||
|
} else if (state && state->seen_steps.count(step)) {
|
||||||
|
response = MakeAck(work_id, step, /*duplicate=*/true);
|
||||||
|
} else if (state && state->credits <= 0) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_FLOW_CONTROL_VIOLATION,
|
||||||
|
"no flow-control credit remaining", false, true);
|
||||||
|
} else {
|
||||||
|
const BundleCheck check = engine_.Validate(chunk.bundle());
|
||||||
|
if (check.oversize_detail) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_RESOURCE_EXHAUSTED,
|
||||||
|
*check.oversize_detail, false, false);
|
||||||
|
} else if (check.corrupt_detail) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_PAYLOAD_CORRUPT,
|
||||||
|
*check.corrupt_detail, false, false);
|
||||||
|
} else {
|
||||||
|
if (state) {
|
||||||
|
state->seen_steps.insert(step);
|
||||||
|
state->credits -= 1;
|
||||||
|
}
|
||||||
|
engine_.BoundedForward(chunk.bundle()); // real bounded forward over wire bytes
|
||||||
|
*response.mutable_chunk() = chunk; // echo the exact bundle back
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
stream->Write(response);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case sp::SessionRequest::kDecode: {
|
||||||
|
const sp::DecodeStep& step_msg = request.decode();
|
||||||
|
const std::string work_id = step_msg.work_id();
|
||||||
|
const uint64_t step = step_msg.idempotency_step();
|
||||||
|
|
||||||
|
sp::TensorBundle bundle;
|
||||||
|
if (step_msg.bundle().tensors_size() > 0) {
|
||||||
|
bundle = step_msg.bundle();
|
||||||
|
} else {
|
||||||
|
bundle.set_bundle_version(1);
|
||||||
|
*bundle.add_tensors() = step_msg.tensor();
|
||||||
|
}
|
||||||
|
|
||||||
|
sp::SessionResponse response;
|
||||||
|
{
|
||||||
|
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||||
|
auto it = sessions_.find(route_session_id);
|
||||||
|
SessionState* state = it != sessions_.end() ? &it->second : nullptr;
|
||||||
|
|
||||||
|
if (state && (state->cancelled_session || state->cancelled_work.count(work_id))) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_CANCELLED,
|
||||||
|
"work was cancelled", false, false);
|
||||||
|
} else if (step_msg.deadline_unix_nanos() != 0 &&
|
||||||
|
NowUnixNanos() > step_msg.deadline_unix_nanos()) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_DEADLINE_EXCEEDED,
|
||||||
|
"deadline already passed", false, false);
|
||||||
|
} else if (state && state->seen_steps.count(step)) {
|
||||||
|
response = MakeAck(work_id, step, /*duplicate=*/true);
|
||||||
|
} else if (state && state->credits <= 0) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_FLOW_CONTROL_VIOLATION,
|
||||||
|
"no flow-control credit remaining", false, true);
|
||||||
|
} else {
|
||||||
|
const BundleCheck check = engine_.Validate(bundle);
|
||||||
|
if (check.oversize_detail) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_RESOURCE_EXHAUSTED,
|
||||||
|
*check.oversize_detail, false, false);
|
||||||
|
} else if (check.corrupt_detail) {
|
||||||
|
response = MakeFail(route_session_id, work_id, step, sp::ERROR_CODE_PAYLOAD_CORRUPT,
|
||||||
|
*check.corrupt_detail, false, false);
|
||||||
|
} else {
|
||||||
|
if (state) {
|
||||||
|
state->seen_steps.insert(step);
|
||||||
|
state->credits -= 1;
|
||||||
|
}
|
||||||
|
engine_.BoundedForward(bundle);
|
||||||
|
// No decode response field exists; echo the step back as a
|
||||||
|
// chunk-bearing SessionResponse per the proto's relayed-frame design.
|
||||||
|
sp::ActivationChunk* out = response.mutable_chunk();
|
||||||
|
sp::Envelope* out_env = out->mutable_envelope();
|
||||||
|
out_env->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
out_env->set_work_id(work_id);
|
||||||
|
out_env->set_idempotency_step(step);
|
||||||
|
out_env->set_phase(sp::PHASE_DECODE);
|
||||||
|
sp::PositionSpan* pos = out_env->mutable_position();
|
||||||
|
pos->set_first_position(step_msg.position());
|
||||||
|
pos->set_token_count(1);
|
||||||
|
*out->mutable_bundle() = bundle;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
stream->Write(response);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case sp::SessionRequest::kFlowControl: {
|
||||||
|
const uint32_t topup = request.flow_control().credits_granted();
|
||||||
|
sp::SessionResponse response;
|
||||||
|
{
|
||||||
|
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||||
|
auto it = sessions_.find(route_session_id);
|
||||||
|
sp::FlowControl* fc = response.mutable_flow_control();
|
||||||
|
if (it != sessions_.end()) {
|
||||||
|
SessionState& state = it->second;
|
||||||
|
int64_t granted = std::min<int64_t>(state.credits + topup,
|
||||||
|
static_cast<int64_t>(state.max_inflight));
|
||||||
|
state.credits = granted;
|
||||||
|
fc->set_credits_granted(static_cast<uint32_t>(granted));
|
||||||
|
fc->set_max_inflight_chunks(state.max_inflight);
|
||||||
|
fc->set_max_chunk_bytes(state.max_chunk_bytes);
|
||||||
|
} else {
|
||||||
|
fc->set_credits_granted(topup != 0 ? topup : limits_.credits_granted);
|
||||||
|
fc->set_max_inflight_chunks(limits_.max_inflight_chunks);
|
||||||
|
fc->set_max_chunk_bytes(limits_.max_chunk_bytes);
|
||||||
|
}
|
||||||
|
fc->set_max_prefill_chunk_tokens(limits_.max_prefill_chunk_tokens);
|
||||||
|
}
|
||||||
|
stream->Write(response);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case sp::SessionRequest::kRelease: {
|
||||||
|
const sp::ReleaseSignal& release = request.release();
|
||||||
|
sp::SessionResponse response;
|
||||||
|
sp::ShardStatus* status = response.mutable_status();
|
||||||
|
status->set_work_id(release.work_id());
|
||||||
|
status->set_route_session_id(release.route_session_id());
|
||||||
|
status->set_terminal(true);
|
||||||
|
stream->Write(response);
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
|
||||||
|
case sp::SessionRequest::kCancel: {
|
||||||
|
const sp::CancelSignal& signal = request.cancel();
|
||||||
|
MarkCancelled(route_session_id, signal.work_id());
|
||||||
|
const bool whole_session = signal.work_id().empty();
|
||||||
|
stream->Write(MakeFail(route_session_id, signal.work_id(), 0, sp::ERROR_CODE_CANCELLED,
|
||||||
|
signal.reason().empty() ? "cancelled" : signal.reason(),
|
||||||
|
whole_session, false));
|
||||||
|
if (whole_session) {
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
default: {
|
||||||
|
sp::SessionResponse response;
|
||||||
|
response.mutable_status()->set_terminal(true);
|
||||||
|
stream->Write(response);
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
|
||||||
|
grpc::Status ShardRuntimeServiceImpl::Release(grpc::ServerContext*,
|
||||||
|
const sp::ReleaseRequest* request,
|
||||||
|
sp::ReleaseResponse* response) {
|
||||||
|
bool existed;
|
||||||
|
{
|
||||||
|
std::lock_guard<std::mutex> lk(sessions_mu_);
|
||||||
|
existed = sessions_.erase(request->route_session_id()) != 0;
|
||||||
|
}
|
||||||
|
response->set_released(existed);
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
|
||||||
|
grpc::Status ShardRuntimeServiceImpl::Cancel(grpc::ServerContext*,
|
||||||
|
const sp::CancelRequest* request,
|
||||||
|
sp::CancelResponse* response) {
|
||||||
|
const uint32_t newly = MarkCancelled(request->route_session_id(), request->work_id());
|
||||||
|
response->set_cancelled_work_items(newly);
|
||||||
|
return grpc::Status::OK;
|
||||||
|
}
|
||||||
|
|
||||||
|
} // namespace meshnet::worker
|
||||||
85
packages/node/native/worker/shard_service.h
Normal file
85
packages/node/native/worker/shard_service.h
Normal file
@@ -0,0 +1,85 @@
|
|||||||
|
// The native Shard worker's ShardRuntime service (DGR-033).
|
||||||
|
//
|
||||||
|
// A faithful C++ port of `ShardRuntimeServicer` in `shard_runtime_server.py`:
|
||||||
|
// the same per-`route_session_id` identity/credit/dedup state, the same
|
||||||
|
// fail-closed negative paths (stale epoch, expired deadline, corrupt/oversize
|
||||||
|
// payload, exhausted flow-control credit, duplicate idempotency step, in-band
|
||||||
|
// and out-of-band cancellation), and the same lifecycle (open/prefill/decode/
|
||||||
|
// flow-control/release/cancel). The only compute it does is the fake engine's
|
||||||
|
// bounded forward — there is no llama.cpp linkage and no arbitrary-graph RPC.
|
||||||
|
|
||||||
|
#ifndef MESHNET_NATIVE_WORKER_SHARD_SERVICE_H_
|
||||||
|
#define MESHNET_NATIVE_WORKER_SHARD_SERVICE_H_
|
||||||
|
|
||||||
|
#include <cstdint>
|
||||||
|
#include <map>
|
||||||
|
#include <mutex>
|
||||||
|
#include <set>
|
||||||
|
#include <string>
|
||||||
|
|
||||||
|
#include <grpcpp/grpcpp.h>
|
||||||
|
|
||||||
|
#include "fake_engine.h"
|
||||||
|
#include "shard_runtime.grpc.pb.h"
|
||||||
|
#include "shard_runtime.pb.h"
|
||||||
|
|
||||||
|
namespace meshnet::worker {
|
||||||
|
|
||||||
|
namespace sp = ::meshnet::shard::v1;
|
||||||
|
|
||||||
|
struct FlowLimits {
|
||||||
|
uint32_t credits_granted = 16;
|
||||||
|
uint32_t max_inflight_chunks = 16;
|
||||||
|
uint64_t max_chunk_bytes = 4u * 1024u * 1024u;
|
||||||
|
uint32_t max_prefill_chunk_tokens = 512;
|
||||||
|
};
|
||||||
|
|
||||||
|
// Per-route-session identity/credit/dedup state, kept on the servicer instance
|
||||||
|
// (guarded by a lock) so an out-of-band unary Cancel from a different handler
|
||||||
|
// thread can reach a session a concurrent Session stream is still iterating.
|
||||||
|
struct SessionState {
|
||||||
|
uint64_t epoch = 0;
|
||||||
|
int64_t credits = 0;
|
||||||
|
uint32_t max_inflight = 0;
|
||||||
|
uint64_t max_chunk_bytes = 0;
|
||||||
|
std::set<uint64_t> seen_steps;
|
||||||
|
std::set<std::string> cancelled_work;
|
||||||
|
bool cancelled_session = false;
|
||||||
|
};
|
||||||
|
|
||||||
|
class ShardRuntimeServiceImpl final : public sp::ShardRuntime::Service {
|
||||||
|
public:
|
||||||
|
explicit ShardRuntimeServiceImpl(FlowLimits limits)
|
||||||
|
: limits_(limits), engine_(limits.max_chunk_bytes) {}
|
||||||
|
|
||||||
|
grpc::Status GetCapability(grpc::ServerContext* context,
|
||||||
|
const sp::CapabilityRequest* request,
|
||||||
|
sp::CapabilityReport* response) override;
|
||||||
|
|
||||||
|
grpc::Status Health(grpc::ServerContext* context, const sp::HealthRequest* request,
|
||||||
|
sp::HealthReport* response) override;
|
||||||
|
|
||||||
|
grpc::Status Session(
|
||||||
|
grpc::ServerContext* context,
|
||||||
|
grpc::ServerReaderWriter<sp::SessionResponse, sp::SessionRequest>* stream) override;
|
||||||
|
|
||||||
|
grpc::Status Release(grpc::ServerContext* context, const sp::ReleaseRequest* request,
|
||||||
|
sp::ReleaseResponse* response) override;
|
||||||
|
|
||||||
|
grpc::Status Cancel(grpc::ServerContext* context, const sp::CancelRequest* request,
|
||||||
|
sp::CancelResponse* response) override;
|
||||||
|
|
||||||
|
private:
|
||||||
|
// Returns the number of items newly marked cancelled, creating session state
|
||||||
|
// if the Cancel raced ahead of SessionOpen.
|
||||||
|
uint32_t MarkCancelled(const std::string& route_session_id, const std::string& work_id);
|
||||||
|
|
||||||
|
FlowLimits limits_;
|
||||||
|
FakeShardEngine engine_;
|
||||||
|
std::mutex sessions_mu_;
|
||||||
|
std::map<std::string, SessionState> sessions_;
|
||||||
|
};
|
||||||
|
|
||||||
|
} // namespace meshnet::worker
|
||||||
|
|
||||||
|
#endif // MESHNET_NATIVE_WORKER_SHARD_SERVICE_H_
|
||||||
302
packages/node/native/worker/shard_worker_main.cpp
Normal file
302
packages/node/native/worker/shard_worker_main.cpp
Normal file
@@ -0,0 +1,302 @@
|
|||||||
|
// Standalone native Shard worker executable (DGR-033).
|
||||||
|
//
|
||||||
|
// Serves the complete ShardRuntime lifecycle/stream contract over real
|
||||||
|
// gRPC/HTTP2 using the model-free FakeShardEngine. It links neither llama.cpp
|
||||||
|
// nor any graph-execution entry point: the only surface it exposes is the
|
||||||
|
// ShardRuntime service defined in shard_runtime.proto.
|
||||||
|
//
|
||||||
|
// Usage:
|
||||||
|
// shard_worker [listen_addr] serve until SIGTERM/SIGINT (graceful drain)
|
||||||
|
// shard_worker --selftest bind an ephemeral port, self-drive the
|
||||||
|
// lifecycle over a real loopback channel, exit
|
||||||
|
//
|
||||||
|
// Environment:
|
||||||
|
// MESHNET_SHARD_LISTEN_ADDR host:port to bind (default localhost:50051)
|
||||||
|
// MESHNET_MAX_CHUNK_BYTES per-chunk byte ceiling the worker enforces
|
||||||
|
//
|
||||||
|
// On a normal run it prints one readiness line — "ShardRuntime worker listening
|
||||||
|
// on <addr>" — once the socket is bound, so a supervisor/harness has a real
|
||||||
|
// readiness signal instead of a sleep.
|
||||||
|
|
||||||
|
#include <atomic>
|
||||||
|
#include <cerrno>
|
||||||
|
#include <csignal>
|
||||||
|
#include <cstdint>
|
||||||
|
#include <cstdlib>
|
||||||
|
#include <cstring>
|
||||||
|
#include <iostream>
|
||||||
|
#include <memory>
|
||||||
|
#include <string>
|
||||||
|
#include <thread>
|
||||||
|
#include <unistd.h>
|
||||||
|
|
||||||
|
#include <grpcpp/grpcpp.h>
|
||||||
|
|
||||||
|
#include "shard_service.h"
|
||||||
|
#include "shard_runtime.grpc.pb.h"
|
||||||
|
|
||||||
|
namespace {
|
||||||
|
|
||||||
|
namespace sp = ::meshnet::shard::v1;
|
||||||
|
|
||||||
|
// Self-pipe: the signal handler must stay async-signal-safe, so it only writes
|
||||||
|
// one byte; a helper thread reads it and performs the (non-signal-safe) server
|
||||||
|
// Shutdown(). Set once in main() before installing the handler.
|
||||||
|
volatile std::sig_atomic_t g_signal_pipe_write_fd = -1;
|
||||||
|
|
||||||
|
extern "C" void HandleTermination(int /*signum*/) {
|
||||||
|
if (g_signal_pipe_write_fd >= 0) {
|
||||||
|
const char byte = 1;
|
||||||
|
ssize_t rc = ::write(g_signal_pipe_write_fd, &byte, 1);
|
||||||
|
(void)rc; // best-effort; nothing safe to do on failure inside a handler
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
meshnet::worker::FlowLimits LimitsFromEnv() {
|
||||||
|
meshnet::worker::FlowLimits limits;
|
||||||
|
if (const char* raw = std::getenv("MESHNET_MAX_CHUNK_BYTES")) {
|
||||||
|
char* end = nullptr;
|
||||||
|
const unsigned long long value = std::strtoull(raw, &end, 10);
|
||||||
|
if (end != raw && value > 0) {
|
||||||
|
limits.max_chunk_bytes = static_cast<uint64_t>(value);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return limits;
|
||||||
|
}
|
||||||
|
|
||||||
|
int RunSelfTest() {
|
||||||
|
meshnet::worker::ShardRuntimeServiceImpl service(LimitsFromEnv());
|
||||||
|
int selected_port = 0;
|
||||||
|
grpc::ServerBuilder builder;
|
||||||
|
builder.AddListeningPort("127.0.0.1:0", grpc::InsecureServerCredentials(), &selected_port);
|
||||||
|
builder.RegisterService(&service);
|
||||||
|
std::unique_ptr<grpc::Server> server(builder.BuildAndStart());
|
||||||
|
if (!server || selected_port == 0) {
|
||||||
|
std::cerr << "selftest: failed to bind ephemeral port\n";
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
const std::string target = "127.0.0.1:" + std::to_string(selected_port);
|
||||||
|
auto channel = grpc::CreateChannel(target, grpc::InsecureChannelCredentials());
|
||||||
|
auto stub = sp::ShardRuntime::NewStub(channel);
|
||||||
|
|
||||||
|
int failures = 0;
|
||||||
|
auto check = [&](bool cond, const char* what) {
|
||||||
|
if (!cond) {
|
||||||
|
std::cerr << "selftest FAIL: " << what << "\n";
|
||||||
|
++failures;
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
// Capability + health.
|
||||||
|
{
|
||||||
|
grpc::ClientContext ctx;
|
||||||
|
sp::CapabilityRequest req;
|
||||||
|
req.set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
sp::CapabilityReport rep;
|
||||||
|
grpc::Status status = stub->GetCapability(&ctx, req, &rep);
|
||||||
|
check(status.ok(), "GetCapability RPC");
|
||||||
|
check(rep.validated(), "capability validated");
|
||||||
|
check(rep.schema_version() == sp::SCHEMA_VERSION_1, "capability schema version");
|
||||||
|
}
|
||||||
|
{
|
||||||
|
grpc::ClientContext ctx;
|
||||||
|
sp::HealthRequest req;
|
||||||
|
req.set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
sp::HealthReport rep;
|
||||||
|
grpc::Status status = stub->Health(&ctx, req, &rep);
|
||||||
|
check(status.ok(), "Health RPC");
|
||||||
|
check(rep.state() == sp::SERVING_STATE_SERVING, "health serving");
|
||||||
|
}
|
||||||
|
|
||||||
|
// A minimal session: open -> fragmented prefill -> decode -> release.
|
||||||
|
{
|
||||||
|
grpc::ClientContext ctx;
|
||||||
|
auto stream = stub->Session(&ctx);
|
||||||
|
|
||||||
|
sp::SessionRequest open;
|
||||||
|
sp::SessionOpen* o = open.mutable_open();
|
||||||
|
o->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
o->set_route_session_id("selftest");
|
||||||
|
o->set_route_epoch(1);
|
||||||
|
sp::FlowControl* fc = o->mutable_proposed_flow_control();
|
||||||
|
fc->set_credits_granted(16);
|
||||||
|
fc->set_max_inflight_chunks(16);
|
||||||
|
fc->set_max_chunk_bytes(4u * 1024u * 1024u);
|
||||||
|
check(stream->Write(open), "write open");
|
||||||
|
|
||||||
|
sp::SessionResponse accepted;
|
||||||
|
check(stream->Read(&accepted), "read accepted");
|
||||||
|
check(accepted.kind_case() == sp::SessionResponse::kAccepted, "accepted kind");
|
||||||
|
|
||||||
|
// Fragmented prefill: two fragments tiling a 6-byte payload.
|
||||||
|
const std::string payload = "ABCDEF";
|
||||||
|
sp::SessionRequest chunk;
|
||||||
|
sp::ActivationChunk* ac = chunk.mutable_chunk();
|
||||||
|
sp::Envelope* env = ac->mutable_envelope();
|
||||||
|
env->set_schema_version(sp::SCHEMA_VERSION_1);
|
||||||
|
env->set_work_id("w1");
|
||||||
|
env->set_route_session_id("selftest");
|
||||||
|
env->set_route_epoch(1);
|
||||||
|
env->set_idempotency_step(1);
|
||||||
|
env->set_phase(sp::PHASE_PREFILL);
|
||||||
|
sp::TensorBundle* bundle = ac->mutable_bundle();
|
||||||
|
bundle->set_bundle_version(1);
|
||||||
|
sp::NamedTensor* tensor = bundle->add_tensors();
|
||||||
|
tensor->set_name("hidden_states");
|
||||||
|
tensor->set_dtype(sp::DTYPE_BFLOAT16);
|
||||||
|
tensor->set_byte_order(sp::BYTE_ORDER_LITTLE_ENDIAN);
|
||||||
|
tensor->set_total_bytes(payload.size());
|
||||||
|
tensor->set_compression(sp::COMPRESSION_NONE);
|
||||||
|
sp::Checksum* cksum = tensor->mutable_checksum();
|
||||||
|
cksum->set_algorithm(sp::CHECKSUM_ALGORITHM_CRC32C);
|
||||||
|
const uint32_t crc = meshnet::worker::Crc32(payload);
|
||||||
|
std::string crc_be(4, '\0');
|
||||||
|
crc_be[0] = static_cast<char>((crc >> 24) & 0xFF);
|
||||||
|
crc_be[1] = static_cast<char>((crc >> 16) & 0xFF);
|
||||||
|
crc_be[2] = static_cast<char>((crc >> 8) & 0xFF);
|
||||||
|
crc_be[3] = static_cast<char>(crc & 0xFF);
|
||||||
|
cksum->set_value(crc_be);
|
||||||
|
sp::TensorFragment* f0 = tensor->add_fragments();
|
||||||
|
f0->set_fragment_index(0);
|
||||||
|
f0->set_fragment_count(2);
|
||||||
|
f0->set_byte_offset(0);
|
||||||
|
f0->set_payload(payload.substr(0, 3));
|
||||||
|
sp::TensorFragment* f1 = tensor->add_fragments();
|
||||||
|
f1->set_fragment_index(1);
|
||||||
|
f1->set_fragment_count(2);
|
||||||
|
f1->set_byte_offset(3);
|
||||||
|
f1->set_payload(payload.substr(3));
|
||||||
|
check(stream->Write(chunk), "write chunk");
|
||||||
|
|
||||||
|
sp::SessionResponse echoed;
|
||||||
|
check(stream->Read(&echoed), "read chunk echo");
|
||||||
|
check(echoed.kind_case() == sp::SessionResponse::kChunk, "chunk echo kind");
|
||||||
|
|
||||||
|
sp::SessionRequest decode;
|
||||||
|
sp::DecodeStep* ds = decode.mutable_decode();
|
||||||
|
ds->set_idempotency_step(2);
|
||||||
|
ds->set_position(1);
|
||||||
|
ds->set_work_id("w2");
|
||||||
|
sp::TensorBundle* dbundle = ds->mutable_bundle();
|
||||||
|
dbundle->set_bundle_version(1);
|
||||||
|
sp::NamedTensor* dt = dbundle->add_tensors();
|
||||||
|
dt->set_name("hidden_states");
|
||||||
|
dt->set_dtype(sp::DTYPE_BFLOAT16);
|
||||||
|
dt->set_byte_order(sp::BYTE_ORDER_LITTLE_ENDIAN);
|
||||||
|
dt->set_total_bytes(payload.size());
|
||||||
|
dt->set_compression(sp::COMPRESSION_NONE);
|
||||||
|
sp::Checksum* dck = dt->mutable_checksum();
|
||||||
|
dck->set_algorithm(sp::CHECKSUM_ALGORITHM_CRC32C);
|
||||||
|
dck->set_value(crc_be);
|
||||||
|
sp::TensorFragment* df = dt->add_fragments();
|
||||||
|
df->set_fragment_index(0);
|
||||||
|
df->set_fragment_count(1);
|
||||||
|
df->set_byte_offset(0);
|
||||||
|
df->set_payload(payload);
|
||||||
|
check(stream->Write(decode), "write decode");
|
||||||
|
|
||||||
|
sp::SessionResponse decode_echo;
|
||||||
|
check(stream->Read(&decode_echo), "read decode echo");
|
||||||
|
check(decode_echo.kind_case() == sp::SessionResponse::kChunk, "decode echo kind");
|
||||||
|
|
||||||
|
sp::SessionRequest release;
|
||||||
|
sp::ReleaseSignal* rs = release.mutable_release();
|
||||||
|
rs->set_route_session_id("selftest");
|
||||||
|
rs->set_work_id("w-final");
|
||||||
|
check(stream->Write(release), "write release");
|
||||||
|
stream->WritesDone();
|
||||||
|
|
||||||
|
sp::SessionResponse terminal;
|
||||||
|
check(stream->Read(&terminal), "read terminal");
|
||||||
|
check(terminal.kind_case() == sp::SessionResponse::kStatus && terminal.status().terminal(),
|
||||||
|
"terminal status");
|
||||||
|
|
||||||
|
grpc::Status status = stream->Finish();
|
||||||
|
check(status.ok(), "stream finish");
|
||||||
|
}
|
||||||
|
|
||||||
|
server->Shutdown();
|
||||||
|
server->Wait();
|
||||||
|
|
||||||
|
if (failures == 0) {
|
||||||
|
std::cout << "selftest: all lifecycle checks passed\n";
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
std::cerr << "selftest: " << failures << " check(s) failed\n";
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
} // namespace
|
||||||
|
|
||||||
|
int main(int argc, char** argv) {
|
||||||
|
GOOGLE_PROTOBUF_VERIFY_VERSION;
|
||||||
|
|
||||||
|
for (int i = 1; i < argc; ++i) {
|
||||||
|
if (std::strcmp(argv[i], "--selftest") == 0) {
|
||||||
|
return RunSelfTest();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
std::string listen_addr = "localhost:50051";
|
||||||
|
if (const char* env = std::getenv("MESHNET_SHARD_LISTEN_ADDR")) {
|
||||||
|
listen_addr = env;
|
||||||
|
}
|
||||||
|
if (argc > 1 && argv[1][0] != '-') {
|
||||||
|
listen_addr = argv[1];
|
||||||
|
}
|
||||||
|
|
||||||
|
meshnet::worker::FlowLimits limits = LimitsFromEnv();
|
||||||
|
meshnet::worker::ShardRuntimeServiceImpl service(limits);
|
||||||
|
|
||||||
|
grpc::ServerBuilder builder;
|
||||||
|
int selected_port = 0;
|
||||||
|
builder.AddListeningPort(listen_addr, grpc::InsecureServerCredentials(), &selected_port);
|
||||||
|
// Bounded messages, two layers: a hard transport receive ceiling (never below
|
||||||
|
// 4 MiB so the handshake and normal chunks always fit) plus the finer
|
||||||
|
// app-level per-tensor RESOURCE_EXHAUSTED check the service enforces against
|
||||||
|
// the negotiated max_chunk_bytes. Neither path lets an unbounded frame in.
|
||||||
|
constexpr int kTransportFloor = 4 * 1024 * 1024;
|
||||||
|
const int transport_max = limits.max_chunk_bytes > static_cast<uint64_t>(kTransportFloor)
|
||||||
|
? static_cast<int>(limits.max_chunk_bytes)
|
||||||
|
: kTransportFloor;
|
||||||
|
builder.SetMaxReceiveMessageSize(transport_max);
|
||||||
|
builder.RegisterService(&service);
|
||||||
|
std::unique_ptr<grpc::Server> server(builder.BuildAndStart());
|
||||||
|
if (!server || selected_port == 0) {
|
||||||
|
std::cerr << "failed to bind " << listen_addr << "\n";
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
int pipe_fds[2];
|
||||||
|
if (::pipe(pipe_fds) != 0) {
|
||||||
|
std::cerr << "failed to create shutdown pipe\n";
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
g_signal_pipe_write_fd = pipe_fds[1];
|
||||||
|
|
||||||
|
struct sigaction sa;
|
||||||
|
std::memset(&sa, 0, sizeof(sa));
|
||||||
|
sa.sa_handler = HandleTermination;
|
||||||
|
::sigaction(SIGTERM, &sa, nullptr);
|
||||||
|
::sigaction(SIGINT, &sa, nullptr);
|
||||||
|
|
||||||
|
// Drain thread: wakes on the first termination signal and shuts the server
|
||||||
|
// down gracefully so in-flight sessions finish rather than being severed.
|
||||||
|
std::thread drain([&server, read_fd = pipe_fds[0]]() {
|
||||||
|
char byte = 0;
|
||||||
|
ssize_t rc = 0;
|
||||||
|
do {
|
||||||
|
rc = ::read(read_fd, &byte, 1);
|
||||||
|
} while (rc < 0 && errno == EINTR);
|
||||||
|
server->Shutdown();
|
||||||
|
});
|
||||||
|
|
||||||
|
std::cout << "ShardRuntime worker listening on " << listen_addr << std::endl;
|
||||||
|
|
||||||
|
server->Wait();
|
||||||
|
drain.join();
|
||||||
|
::close(pipe_fds[0]);
|
||||||
|
::close(pipe_fds[1]);
|
||||||
|
std::cout << "ShardRuntime worker shut down cleanly" << std::endl;
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
498
tests/test_native_shard_worker.py
Normal file
498
tests/test_native_shard_worker.py
Normal file
@@ -0,0 +1,498 @@
|
|||||||
|
"""DGR-033 integration tests for the standalone native C++ Shard worker.
|
||||||
|
|
||||||
|
These tests spawn the *real* compiled ``shard_worker`` executable as a separate
|
||||||
|
OS process, connect to its real localhost socket with the committed generated
|
||||||
|
``ShardRuntimeStub`` stubs, and drive the complete lifecycle/stream contract.
|
||||||
|
There is no in-memory channel, no Python servicer, and no fake transport: the
|
||||||
|
server under test is the C++ binary DGR-033 builds.
|
||||||
|
|
||||||
|
The worker binary is located via ``MESHNET_SHARD_WORKER_BIN`` or the default
|
||||||
|
out-of-tree build path ``build/native/shard_worker``. When it has not been
|
||||||
|
built (a default developer/CI checkout without the pinned gRPC C++ toolchain),
|
||||||
|
every test here is skipped rather than failed — the same ``requires_cmake``
|
||||||
|
gating pattern DGR-029/DGR-030 use for native-build-dependent tests. The
|
||||||
|
session that implemented DGR-033 built the binary and ran these for real; see
|
||||||
|
``evidence/DGR-033/README.md`` for the exact commands and results.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
import signal
|
||||||
|
import socket
|
||||||
|
import subprocess
|
||||||
|
import time
|
||||||
|
import zlib
|
||||||
|
|
||||||
|
import grpc
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
|
||||||
|
_PYTHONPATH = os.pathsep.join(
|
||||||
|
[os.path.join(REPO_ROOT, "packages", "node"), os.path.join(REPO_ROOT, "packages", "tracker")]
|
||||||
|
)
|
||||||
|
|
||||||
|
from meshnet_node.native_protocol.generated import ( # noqa: E402
|
||||||
|
shard_runtime_pb2 as pb,
|
||||||
|
shard_runtime_pb2_grpc as pb_grpc,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _worker_binary() -> str | None:
|
||||||
|
explicit = os.environ.get("MESHNET_SHARD_WORKER_BIN")
|
||||||
|
if explicit and os.path.exists(explicit):
|
||||||
|
return explicit
|
||||||
|
default = os.path.join(REPO_ROOT, "build", "native", "shard_worker")
|
||||||
|
if os.path.exists(default):
|
||||||
|
return default
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
_WORKER_BIN = _worker_binary()
|
||||||
|
pytestmark = pytest.mark.skipif(
|
||||||
|
_WORKER_BIN is None,
|
||||||
|
reason=(
|
||||||
|
"native shard_worker binary not built; build packages/node/native with the "
|
||||||
|
"pinned gRPC C++ toolchain or set MESHNET_SHARD_WORKER_BIN"
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _free_port() -> int:
|
||||||
|
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
|
||||||
|
s.bind(("127.0.0.1", 0))
|
||||||
|
port = s.getsockname()[1]
|
||||||
|
s.close()
|
||||||
|
return port
|
||||||
|
|
||||||
|
|
||||||
|
def _start_worker(listen_addr: str, extra_env: dict[str, str] | None = None) -> subprocess.Popen:
|
||||||
|
env = dict(os.environ)
|
||||||
|
env["PYTHONPATH"] = _PYTHONPATH
|
||||||
|
if extra_env:
|
||||||
|
env.update(extra_env)
|
||||||
|
proc = subprocess.Popen(
|
||||||
|
[_WORKER_BIN, listen_addr],
|
||||||
|
cwd=REPO_ROOT,
|
||||||
|
env=env,
|
||||||
|
stdout=subprocess.PIPE,
|
||||||
|
stderr=subprocess.STDOUT,
|
||||||
|
text=True,
|
||||||
|
)
|
||||||
|
deadline = time.time() + 30.0
|
||||||
|
while time.time() < deadline:
|
||||||
|
line = proc.stdout.readline()
|
||||||
|
if not line:
|
||||||
|
if proc.poll() is not None:
|
||||||
|
out, _ = proc.communicate()
|
||||||
|
raise RuntimeError(f"worker exited early:\n{out}")
|
||||||
|
continue
|
||||||
|
if "listening on" in line:
|
||||||
|
return proc
|
||||||
|
raise RuntimeError("worker did not start listening in time")
|
||||||
|
|
||||||
|
|
||||||
|
class _Worker:
|
||||||
|
"""A spawned worker plus a ready channel; also captures stdout on close."""
|
||||||
|
|
||||||
|
def __init__(self, extra_env: dict[str, str] | None = None) -> None:
|
||||||
|
self.port = _free_port()
|
||||||
|
self.addr = f"127.0.0.1:{self.port}"
|
||||||
|
self.proc = _start_worker(self.addr, extra_env)
|
||||||
|
self.channel = grpc.insecure_channel(self.addr)
|
||||||
|
grpc.channel_ready_future(self.channel).result(timeout=15.0)
|
||||||
|
|
||||||
|
def stub(self) -> pb_grpc.ShardRuntimeStub:
|
||||||
|
return pb_grpc.ShardRuntimeStub(self.channel)
|
||||||
|
|
||||||
|
def session(self, requests):
|
||||||
|
call = self.channel.stream_stream(
|
||||||
|
"/meshnet.shard.v1.ShardRuntime/Session",
|
||||||
|
request_serializer=lambda m: m.SerializeToString(),
|
||||||
|
response_deserializer=pb.SessionResponse.FromString,
|
||||||
|
)
|
||||||
|
return list(call(iter(requests)))
|
||||||
|
|
||||||
|
def close(self, *, sig: int = signal.SIGTERM) -> str:
|
||||||
|
self.channel.close()
|
||||||
|
self.proc.send_signal(sig)
|
||||||
|
try:
|
||||||
|
out, _ = self.proc.communicate(timeout=10)
|
||||||
|
except subprocess.TimeoutExpired:
|
||||||
|
self.proc.kill()
|
||||||
|
out, _ = self.proc.communicate()
|
||||||
|
return out or ""
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def worker():
|
||||||
|
w = _Worker()
|
||||||
|
try:
|
||||||
|
yield w
|
||||||
|
finally:
|
||||||
|
if w.proc.poll() is None:
|
||||||
|
w.close()
|
||||||
|
|
||||||
|
|
||||||
|
def _crc32c(payload: bytes) -> bytes:
|
||||||
|
return zlib.crc32(payload).to_bytes(4, "big")
|
||||||
|
|
||||||
|
|
||||||
|
def _open(*, route_session_id="rs-1", route_epoch=7, credits_granted=16) -> pb.SessionRequest:
|
||||||
|
return pb.SessionRequest(
|
||||||
|
open=pb.SessionOpen(
|
||||||
|
schema_version=pb.SCHEMA_VERSION_1,
|
||||||
|
route_session_id=route_session_id,
|
||||||
|
route_epoch=route_epoch,
|
||||||
|
fingerprint=pb.Fingerprint(
|
||||||
|
model_artifact_digest="sha256:native-test-artifact",
|
||||||
|
runtime_recipe_digest="sha256:native-test-recipe",
|
||||||
|
recipe_id="native-test",
|
||||||
|
recipe_version="1",
|
||||||
|
catalogue_version="1",
|
||||||
|
),
|
||||||
|
shard_range=pb.ShardRange(start_layer=0, end_layer=32, effective_start_layer=0),
|
||||||
|
proposed_flow_control=pb.FlowControl(
|
||||||
|
credits_granted=credits_granted,
|
||||||
|
max_inflight_chunks=16,
|
||||||
|
max_chunk_bytes=4 * 1024 * 1024,
|
||||||
|
max_prefill_chunk_tokens=512,
|
||||||
|
),
|
||||||
|
accepted_compression=[pb.COMPRESSION_NONE],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _chunk(
|
||||||
|
work_id,
|
||||||
|
payload: bytes,
|
||||||
|
step,
|
||||||
|
*,
|
||||||
|
route_session_id="rs-1",
|
||||||
|
route_epoch=7,
|
||||||
|
deadline_unix_nanos=0,
|
||||||
|
fragments=1,
|
||||||
|
total_bytes=None,
|
||||||
|
) -> pb.SessionRequest:
|
||||||
|
total = len(payload) if total_bytes is None else total_bytes
|
||||||
|
frags = []
|
||||||
|
if fragments == 1:
|
||||||
|
frags = [pb.TensorFragment(fragment_index=0, fragment_count=1, byte_offset=0, payload=payload)]
|
||||||
|
else:
|
||||||
|
# Split into ``fragments`` tiling pieces.
|
||||||
|
size = max(1, len(payload) // fragments)
|
||||||
|
offset = 0
|
||||||
|
idx = 0
|
||||||
|
while offset < len(payload):
|
||||||
|
piece = payload[offset : offset + size] if idx < fragments - 1 else payload[offset:]
|
||||||
|
frags.append(
|
||||||
|
pb.TensorFragment(
|
||||||
|
fragment_index=idx, fragment_count=fragments, byte_offset=offset, payload=piece
|
||||||
|
)
|
||||||
|
)
|
||||||
|
offset += len(piece)
|
||||||
|
idx += 1
|
||||||
|
tensor = pb.NamedTensor(
|
||||||
|
name="hidden_states",
|
||||||
|
shape=[1, 1, 4096],
|
||||||
|
dtype=pb.DTYPE_BFLOAT16,
|
||||||
|
byte_order=pb.BYTE_ORDER_LITTLE_ENDIAN,
|
||||||
|
total_bytes=total,
|
||||||
|
compression=pb.COMPRESSION_NONE,
|
||||||
|
checksum=pb.Checksum(algorithm=pb.CHECKSUM_ALGORITHM_CRC32C, value=_crc32c(payload)),
|
||||||
|
fragments=frags,
|
||||||
|
)
|
||||||
|
bundle = pb.TensorBundle(
|
||||||
|
bundle_version=1,
|
||||||
|
tensors=[tensor],
|
||||||
|
architecture=pb.ARCHITECTURE_TYPE_DENSE,
|
||||||
|
boundary_point="pre_tail_residual",
|
||||||
|
)
|
||||||
|
envelope = pb.Envelope(
|
||||||
|
schema_version=pb.SCHEMA_VERSION_1,
|
||||||
|
work_id=work_id,
|
||||||
|
route_session_id=route_session_id,
|
||||||
|
route_epoch=route_epoch,
|
||||||
|
idempotency_step=step,
|
||||||
|
phase=pb.PHASE_PREFILL,
|
||||||
|
position=pb.PositionSpan(first_position=0, token_count=1),
|
||||||
|
deadline_unix_nanos=deadline_unix_nanos,
|
||||||
|
)
|
||||||
|
return pb.SessionRequest(chunk=pb.ActivationChunk(envelope=envelope, bundle=bundle))
|
||||||
|
|
||||||
|
|
||||||
|
def _decode(work_id, payload: bytes, step, position) -> pb.SessionRequest:
|
||||||
|
tensor = pb.NamedTensor(
|
||||||
|
name="hidden_states",
|
||||||
|
shape=[1, 1, 4096],
|
||||||
|
dtype=pb.DTYPE_BFLOAT16,
|
||||||
|
byte_order=pb.BYTE_ORDER_LITTLE_ENDIAN,
|
||||||
|
total_bytes=len(payload),
|
||||||
|
compression=pb.COMPRESSION_NONE,
|
||||||
|
checksum=pb.Checksum(algorithm=pb.CHECKSUM_ALGORITHM_CRC32C, value=_crc32c(payload)),
|
||||||
|
fragments=[pb.TensorFragment(fragment_index=0, fragment_count=1, byte_offset=0, payload=payload)],
|
||||||
|
)
|
||||||
|
return pb.SessionRequest(
|
||||||
|
decode=pb.DecodeStep(
|
||||||
|
idempotency_step=step,
|
||||||
|
position=position,
|
||||||
|
expected_past_len=position,
|
||||||
|
work_id=work_id,
|
||||||
|
bundle=pb.TensorBundle(bundle_version=1, tensors=[tensor], architecture=pb.ARCHITECTURE_TYPE_DENSE),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _release() -> pb.SessionRequest:
|
||||||
|
return pb.SessionRequest(
|
||||||
|
release=pb.ReleaseSignal(route_session_id="rs-1", route_epoch=7, work_id="work-final")
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _cancel(*, route_session_id="rs-1", work_id="", reason="test cancel") -> pb.SessionRequest:
|
||||||
|
return pb.SessionRequest(
|
||||||
|
cancel=pb.CancelSignal(route_session_id=route_session_id, route_epoch=7, work_id=work_id, reason=reason)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# --- startup / health / capability ----------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_worker_startup_and_health(worker):
|
||||||
|
health = worker.stub().Health(pb.HealthRequest(schema_version=pb.SCHEMA_VERSION_1))
|
||||||
|
assert health.state == pb.SERVING_STATE_SERVING
|
||||||
|
assert health.schema_version == pb.SCHEMA_VERSION_1
|
||||||
|
|
||||||
|
|
||||||
|
def test_worker_capability(worker):
|
||||||
|
cap = worker.stub().GetCapability(pb.CapabilityRequest(schema_version=pb.SCHEMA_VERSION_1))
|
||||||
|
assert cap.validated is True
|
||||||
|
assert cap.schema_version == pb.SCHEMA_VERSION_1
|
||||||
|
assert cap.shard_range.end_layer == 32
|
||||||
|
assert pb.SCHEMA_VERSION_1 in cap.supported_schema_versions
|
||||||
|
|
||||||
|
|
||||||
|
# --- fragmented prefill / decode / release ---------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_fragmented_prefill_echoes_reassembled_payload(worker):
|
||||||
|
payload = b"REAL_ACTIVATION_BYTES_prefill_across_three_fragments_1234567890"
|
||||||
|
responses = worker.session([_open(), _chunk("w1", payload, step=1, fragments=3), _release()])
|
||||||
|
assert responses[0].WhichOneof("kind") == "accepted"
|
||||||
|
echoed = responses[1]
|
||||||
|
assert echoed.WhichOneof("kind") == "chunk"
|
||||||
|
got = b"".join(f.payload for f in echoed.chunk.bundle.tensors[0].fragments)
|
||||||
|
assert got == payload
|
||||||
|
assert echoed.chunk.bundle.tensors[0].checksum.value == _crc32c(payload)
|
||||||
|
assert responses[2].status.terminal is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_decode_step_is_served(worker):
|
||||||
|
payload = b"REAL_ACTIVATION_BYTES_decode_step"
|
||||||
|
responses = worker.session([_open(), _decode("w2", payload, step=1, position=1)])
|
||||||
|
echoed = responses[1]
|
||||||
|
assert echoed.WhichOneof("kind") == "chunk"
|
||||||
|
assert echoed.chunk.envelope.phase == pb.PHASE_DECODE
|
||||||
|
assert echoed.chunk.bundle.tensors[0].fragments[0].payload == payload
|
||||||
|
|
||||||
|
|
||||||
|
def test_release_is_terminal(worker):
|
||||||
|
responses = worker.session([_open(), _release()])
|
||||||
|
assert responses[0].WhichOneof("kind") == "accepted"
|
||||||
|
assert responses[1].status.terminal is True
|
||||||
|
|
||||||
|
|
||||||
|
# --- deadlines / flow control / bounded messages ---------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_expired_deadline_is_rejected(worker):
|
||||||
|
responses = worker.session([_open(), _chunk("w-late", b"payload", step=1, deadline_unix_nanos=1)])
|
||||||
|
assert responses[1].status.error.code == pb.ERROR_CODE_DEADLINE_EXCEEDED
|
||||||
|
|
||||||
|
|
||||||
|
def test_flow_control_violation_and_topup(worker):
|
||||||
|
responses = worker.session(
|
||||||
|
[
|
||||||
|
_open(credits_granted=1),
|
||||||
|
_chunk("w-a", b"payload-a", step=1),
|
||||||
|
_chunk("w-b", b"payload-b", step=2),
|
||||||
|
pb.SessionRequest(flow_control=pb.FlowControl(credits_granted=5)),
|
||||||
|
_chunk("w-c", b"payload-c", step=3),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
assert responses[1].WhichOneof("kind") == "chunk"
|
||||||
|
assert responses[2].status.error.code == pb.ERROR_CODE_FLOW_CONTROL_VIOLATION
|
||||||
|
assert responses[2].status.error.retryable is True
|
||||||
|
assert responses[3].WhichOneof("kind") == "flow_control"
|
||||||
|
assert responses[3].flow_control.credits_granted >= 5
|
||||||
|
assert responses[4].WhichOneof("kind") == "chunk"
|
||||||
|
|
||||||
|
|
||||||
|
def test_bounded_message_is_rejected():
|
||||||
|
"""A tensor whose declared payload exceeds the negotiated ceiling is refused."""
|
||||||
|
w = _Worker(extra_env={"MESHNET_MAX_CHUNK_BYTES": "64"})
|
||||||
|
try:
|
||||||
|
big = b"x" * 128
|
||||||
|
responses = w.session([_open(), _chunk("w-big", big, step=1, total_bytes=128)])
|
||||||
|
status = responses[1].status
|
||||||
|
assert status.error.code == pb.ERROR_CODE_RESOURCE_EXHAUSTED
|
||||||
|
assert "max_chunk_bytes" in status.error.detail
|
||||||
|
finally:
|
||||||
|
if w.proc.poll() is None:
|
||||||
|
w.close()
|
||||||
|
|
||||||
|
|
||||||
|
def test_malformed_fragment_tiling_is_rejected(worker):
|
||||||
|
# A fragment at a non-zero offset with no predecessor cannot tile.
|
||||||
|
tensor = pb.NamedTensor(
|
||||||
|
name="hidden_states",
|
||||||
|
shape=[1, 1, 4096],
|
||||||
|
dtype=pb.DTYPE_BFLOAT16,
|
||||||
|
byte_order=pb.BYTE_ORDER_LITTLE_ENDIAN,
|
||||||
|
total_bytes=7,
|
||||||
|
compression=pb.COMPRESSION_NONE,
|
||||||
|
checksum=pb.Checksum(algorithm=pb.CHECKSUM_ALGORITHM_CRC32C, value=_crc32c(b"payload")),
|
||||||
|
fragments=[pb.TensorFragment(fragment_index=0, fragment_count=1, byte_offset=5, payload=b"payload")],
|
||||||
|
)
|
||||||
|
bad = pb.SessionRequest(
|
||||||
|
chunk=pb.ActivationChunk(
|
||||||
|
envelope=pb.Envelope(
|
||||||
|
schema_version=pb.SCHEMA_VERSION_1,
|
||||||
|
work_id="w-gap",
|
||||||
|
route_session_id="rs-1",
|
||||||
|
route_epoch=7,
|
||||||
|
idempotency_step=1,
|
||||||
|
),
|
||||||
|
bundle=pb.TensorBundle(bundle_version=1, tensors=[tensor]),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
responses = worker.session([_open(), bad])
|
||||||
|
assert responses[1].status.error.code == pb.ERROR_CODE_PAYLOAD_CORRUPT
|
||||||
|
assert "tile" in responses[1].status.error.detail
|
||||||
|
|
||||||
|
|
||||||
|
def test_stale_route_epoch_is_rejected(worker):
|
||||||
|
responses = worker.session([_open(route_epoch=7), _chunk("w-stale", b"payload", step=1, route_epoch=5)])
|
||||||
|
assert responses[1].status.error.code == pb.ERROR_CODE_EPOCH_STALE
|
||||||
|
|
||||||
|
|
||||||
|
def test_duplicate_idempotency_step_is_acked(worker):
|
||||||
|
chunk = _chunk("w-dup", b"payload", step=1)
|
||||||
|
responses = worker.session([_open(), chunk, chunk])
|
||||||
|
assert responses[1].WhichOneof("kind") == "chunk"
|
||||||
|
assert responses[2].WhichOneof("kind") == "ack"
|
||||||
|
assert responses[2].ack.duplicate is True
|
||||||
|
|
||||||
|
|
||||||
|
# --- cancellation ----------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_in_band_cancel_of_single_work_item_does_not_end_stream(worker):
|
||||||
|
responses = worker.session(
|
||||||
|
[
|
||||||
|
_open(),
|
||||||
|
_cancel(work_id="work-x"),
|
||||||
|
_chunk("work-x", b"payload", step=1),
|
||||||
|
_chunk("work-y", b"payload", step=2),
|
||||||
|
_release(),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
assert responses[1].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||||
|
assert responses[1].status.terminal is False
|
||||||
|
assert responses[2].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||||
|
assert responses[3].WhichOneof("kind") == "chunk"
|
||||||
|
assert responses[4].status.terminal is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_in_band_cancel_of_whole_session_is_terminal(worker):
|
||||||
|
responses = worker.session([_open(), _cancel(work_id="")])
|
||||||
|
assert responses[1].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||||
|
assert responses[1].status.terminal is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_out_of_band_cancel_rpc_races_ahead_of_open(worker):
|
||||||
|
stub = worker.stub()
|
||||||
|
resp = stub.Cancel(
|
||||||
|
pb.CancelRequest(
|
||||||
|
schema_version=pb.SCHEMA_VERSION_1,
|
||||||
|
route_session_id="rs-precancel",
|
||||||
|
route_epoch=1,
|
||||||
|
work_id="work-precancelled",
|
||||||
|
reason="operator abort",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
assert resp.cancelled_work_items == 1
|
||||||
|
responses = worker.session(
|
||||||
|
[
|
||||||
|
_open(route_session_id="rs-precancel"),
|
||||||
|
_chunk("work-precancelled", b"payload", step=1, route_session_id="rs-precancel"),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
assert responses[1].status.error.code == pb.ERROR_CODE_CANCELLED
|
||||||
|
|
||||||
|
|
||||||
|
def test_release_rpc_is_idempotent(worker):
|
||||||
|
stub = worker.stub()
|
||||||
|
# Open a session so state exists, then release it out of band twice.
|
||||||
|
worker.session([_open(route_session_id="rs-rel"), _release()])
|
||||||
|
# (release signal in-stream does not erase state; the unary Release RPC does)
|
||||||
|
first = stub.Release(pb.ReleaseRequest(schema_version=pb.SCHEMA_VERSION_1, route_session_id="rs-rel", route_epoch=7))
|
||||||
|
second = stub.Release(pb.ReleaseRequest(schema_version=pb.SCHEMA_VERSION_1, route_session_id="rs-rel", route_epoch=7))
|
||||||
|
assert first.released is True
|
||||||
|
assert second.released is False # idempotent: nothing left to drop
|
||||||
|
|
||||||
|
|
||||||
|
def test_independent_session_cancellation(worker):
|
||||||
|
# Cancel the whole of session A; session B must remain fully serviceable.
|
||||||
|
a = worker.session([_open(route_session_id="sess-A"), _cancel(route_session_id="sess-A", work_id="")])
|
||||||
|
assert a[1].status.terminal is True
|
||||||
|
b = worker.session(
|
||||||
|
[
|
||||||
|
_open(route_session_id="sess-B"),
|
||||||
|
_chunk("work-b", b"payload-b", step=1, route_session_id="sess-B"),
|
||||||
|
_release(),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
assert b[1].WhichOneof("kind") == "chunk", "cancelling session A must not affect session B"
|
||||||
|
|
||||||
|
|
||||||
|
# --- graceful shutdown -----------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_graceful_shutdown_on_sigterm():
|
||||||
|
w = _Worker()
|
||||||
|
# Confirm it is serving, then send SIGTERM and require a clean drain/exit.
|
||||||
|
assert w.stub().Health(pb.HealthRequest(schema_version=pb.SCHEMA_VERSION_1)).state == pb.SERVING_STATE_SERVING
|
||||||
|
out = w.close(sig=signal.SIGTERM)
|
||||||
|
assert w.proc.returncode == 0, f"worker did not exit cleanly on SIGTERM:\n{out}"
|
||||||
|
assert "shut down cleanly" in out
|
||||||
|
|
||||||
|
|
||||||
|
# --- direct vs opaque relay byte identity ----------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_direct_and_opaque_relay_yield_identical_responses(worker):
|
||||||
|
"""A direct hop and an opaque relay of the exact captured request bytes must
|
||||||
|
produce byte-identical server responses (relays carry frames verbatim)."""
|
||||||
|
payload = b"RELAY_ACTIVATION_BYTES"
|
||||||
|
requests = [_open(), _chunk("w1", payload, step=1), _release()]
|
||||||
|
|
||||||
|
direct_call = worker.channel.stream_stream(
|
||||||
|
"/meshnet.shard.v1.ShardRuntime/Session",
|
||||||
|
request_serializer=lambda m: m.SerializeToString(),
|
||||||
|
response_deserializer=lambda b: b,
|
||||||
|
)
|
||||||
|
direct_resp = list(direct_call(iter(requests)))
|
||||||
|
captured = [m.SerializeToString() for m in requests]
|
||||||
|
|
||||||
|
relay_call = worker.channel.stream_stream(
|
||||||
|
"/meshnet.shard.v1.ShardRuntime/Session",
|
||||||
|
request_serializer=lambda b: b, # raw captured bytes, no reinterpretation
|
||||||
|
response_deserializer=lambda b: b,
|
||||||
|
)
|
||||||
|
relay_resp = list(relay_call(iter(captured)))
|
||||||
|
|
||||||
|
assert len(direct_resp) == len(relay_resp) == 3
|
||||||
|
for i, (d, r) in enumerate(zip(direct_resp, relay_resp)):
|
||||||
|
assert d == r, f"response #{i} differs between direct and opaque relay"
|
||||||
Reference in New Issue
Block a user