mirror of
https://github.com/baketnk/frame-yap.git
synced 2026-10-06 09:00:05 +02:00
moondream/parakeet-redux was rewritten upstream on 2026-09-23 to a single commit (2bf128600aac), so the pinned revision fad622f no longer resolves. model.safetensors still downloads from the old revision, but config.json, ternary.json, tokenizer.json and README.md return HTTP 404, which makes `install.sh --install-model` fail. The weights, config, ternary map and tokenizer at 2bf1286 have the same size and SHA-256 as the old pins, so the model itself is unchanged. Only the model card README.md differs (8533 -> 8832 bytes), so its pin is updated along with the revision. Fixes #1 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
126 lines
8.3 KiB
Markdown
126 lines
8.3 KiB
Markdown
# Offline worker adapter and backend dispatcher
|
|
|
|
`src/worker.hpp` provides `frameyap::Worker`: call `start(python, script, model,
|
|
threads=2, advanced_debug=false)` explicitly for the legacy Redux script, or
|
|
supply the optional `backend`, `manifest_dir`, `root` arguments for the manifest
|
|
dispatcher. Poll until `ready()`, then `submit(id, pcm)` and poll for one
|
|
`WorkerReply` (text or privacy-safe per-request error). One request at a time;
|
|
no queue, no capture and no input injection. `stop()` discards pending audio,
|
|
terminates/reaps **only its direct child** (TERM, bounded 500 ms, then KILL),
|
|
and is safe to repeat. Destruction stops it. `start()` returns without waiting
|
|
for model load; `poll()` reports warmup/load/crash/timeout/protocol errors by
|
|
throwing, then stops. Launch uses `posix_spawn`, safe with the host's OpenVR/SDL
|
|
threads, rather than running Python setup in a forked multithreaded child.
|
|
Caller must discard stale authorization/results after
|
|
cancellation; this component does not implement focus or delivery policy.
|
|
|
|
The adapter requires an existing absolute, owner-private `$XDG_RUNTIME_DIR`
|
|
(no symlink at the final component), creates its own 0700 `mkdtemp` directory,
|
|
and writes only `clip.raw` with `O_EXCL|O_NOFOLLOW`, mode 0600. Clips are
|
|
3200..320000 finite float samples in `[-1, 1]`, mono 16 kHz, stored as
|
|
little-endian IEEE float32 (0.2..20 s). The microphone adapter saturates finite
|
|
capture/resampler overshoot to this range without rescaling the rest of the
|
|
clip; NaN/infinity are rejected. `Worker::submit` independently rejects samples
|
|
outside this contract before writing a clip or request. Files are unlinked after replies or shutdown, and the
|
|
private directory is removed. Private clips are not encrypted against the
|
|
account owner/root; do not use an untrusted runtime directory. The caller
|
|
should pass a trusted interpreter and script. By default audio/transcripts are
|
|
not logged and child stderr is redirected to `/dev/null`. Request errors report
|
|
only a fixed stage (`audio`, `inference`, `response`) and built-in exception
|
|
category, never the exception message, arbitrary class name, traceback or text.
|
|
The native receiver allowlists those labels before displaying/logging them;
|
|
unknown/legacy errors remain generic.
|
|
|
|
The private pipes use unsigned LE32 payload lengths (1..65536), a one-byte
|
|
message type and, for requests/replies, unsigned LE64 request ID. `T` + ID
|
|
requests reading the fixed clip; `Y` means ready; `F` means load failure
|
|
(`M` for missing/mismatched pinned model or private clip directory, `I` for a
|
|
missing Python dependency, `D` for runtime/model load failure); `R` + ID + UTF-8
|
|
text and `E` + ID + privacy-safe UTF-8 error are replies. Text is at most 4096
|
|
bytes. An unexpected or duplicate reply, wrong ID, extra frame, closed pipe
|
|
or oversized frame stops the worker. A correlated `E` reply or an invalid
|
|
transcript is a **request-level** failure: the Controller drops that clip,
|
|
keeps the ready model loaded and allows Record to retry. A broken protocol,
|
|
worker crash or explicit in-flight cancellation is different and can require a
|
|
reload. Warmup deadline is 120 s, transcription deadline 60 s; `poll()` must
|
|
be called regularly to enforce deadlines. It never initializes a headset or
|
|
starts a recording. There is no auto restart after a process failure.
|
|
|
|
Parakeet Redux is the **model** (`moondream/parakeet-redux`); `moondream` is
|
|
its Python inference package, not a second model or cloud endpoint. Kestrel is
|
|
a native dependency of that local runtime. `python/frameyap/worker.py` lazily
|
|
imports `moondream` only after explicit CLI startup, with HF/Transformers/Datasets
|
|
offline variables and bounded native thread-pool variables set before import. It uses
|
|
`md.photon("moondream/parakeet-redux", model_path=<absolute local directory>,
|
|
device="cpu", cpu_threads=threads)` and persistent
|
|
`transcribe(audio=<numpy float32>, sample_rate=16000)["text"]`.
|
|
Install an **isolated** Python runtime with the
|
|
moondream 2.4.0, kestrel 0.8.0 and compatible CPU dependencies; provide
|
|
preinstalled local weights from revision
|
|
`2bf128600aac4b16946f7ed8372e56117fe5e23b` and pass its directory
|
|
explicitly. The code verifies exact sizes and SHA-256 of weights/config/tokenizer
|
|
against `assets/backends/redux.json` via the shared, offline
|
|
`python/frameyap/model_files.py` schema/verifier before importing model libraries.
|
|
`frameyap --list-models` and `frameyap --check-model redux --model-dir /absolute/model`
|
|
(or `scripts/model-status.py`) read local manifests and optionally hash local
|
|
files; they never load the worker or download weights. For native `--run`,
|
|
`--backend ID`, `--model-store /absolute/store`, and
|
|
`--manifest-dir /absolute/manifests` are wired through the installed launcher
|
|
as explicit overrides; default selection is saved in `config.json` (`redux`). The dispatcher
|
|
`python/frameyap/backend_worker.py` validates a selected manifest and hashes
|
|
its local model, resolves an in-release relative Python/executable launcher
|
|
without a shell, then supervises one child over `frameyap-worker-v1` (ready,
|
|
correlated request/reply and failure frames). The launcher must accept
|
|
`{model_dir}` and `{clip_dir}` (optional `{threads}`); matching manifest and
|
|
pinned weights are **not** a runtime or automatically trusted new backend.
|
|
Only Redux is supplied today. Protocol stdout is isolated at the file-descriptor
|
|
level from third-party diagnostics. Thread limits cover Torch interop/native
|
|
pools and CUDA is not selected. Offline environment
|
|
flags do not prove every third-party internal is unable to access a network.
|
|
Runtime/build/tests perform no downloads; the separate explicit setup utility
|
|
`scripts/fetch-model.py` can provision the public pinned weights.
|
|
|
|
Native-only archives do not bundle the runtime; the explicit
|
|
`install.sh --install-runtime --yes` step pip-installs it into a user-local venv
|
|
(see [packaging](packaging.md) and [third-party notes](third-party.md)). No public runtime bundle has been released.
|
|
Limited ARM64 measurements were taken during development; they are not a claim
|
|
of complete headset acceptance.
|
|
|
|
## Advanced debugging
|
|
|
|
Explicitly set `"advanced_debug": true` in config or enable Settings → Advanced
|
|
debugging. It is off by default. A Settings change restarts the owned worker,
|
|
closes the microphone and discards current work/review; the replacement worker
|
|
warms normally. Turning it off stops detailed capture but **does not delete
|
|
previous logs**. Manual config edits take effect on application restart.
|
|
|
|
With this opt-in the worker receives `--advanced-debug`: native/model stdout and
|
|
stderr, sample counts, full exception tracebacks and recognized transcripts are
|
|
captured. **These logs may contain private speech, transcript text and local
|
|
paths. Inspect/redact before sharing; keep out of Git.** No raw audio archive is
|
|
created; normal temporary clips still expire after replies/cancellation.
|
|
|
|
Files: `$XDG_STATE_HOME/frameyap/worker-debug.log` (fallback
|
|
`~/.local/state/frameyap/worker-debug.log`) and `worker-debug.previous.log`.
|
|
Each debug worker start rotates the current log once; only these two files are
|
|
retained, each bounded to 4 MiB. At the limit a marker is written and further
|
|
output is drained/discarded for that worker session, not allowed to block it.
|
|
The directory is owner-private 0700 and logs are 0600. Unsafe paths, symlinks,
|
|
hardlinks or preexisting permissive files are refused with a visible error;
|
|
there is no fallback to public temporary files. The app's single-instance lock
|
|
is required to avoid competing rotation by multiple workers. Debug output never
|
|
shares the framed protocol stdout. Direct worker CLI use with `--advanced-debug`
|
|
writes to its caller's stderr; the native adapter supplies the private bounded sink.
|
|
|
|
The same opt-in also writes `delivery-debug.log` (rotated to
|
|
`delivery-debug.previous.log` when debugging is enabled, capped near 1 MiB) in
|
|
that directory: paced-delivery metadata only — start/sent byte counts and
|
|
offsets, lease/send errors, finish results and which focus-guard check failed.
|
|
It never contains transcript text.
|
|
|
|
Hardware-free tests run through CTest, including fake-child cancellation, short
|
|
injected warmup/request deadlines, duplicate/stale replies, malformed frames,
|
|
missing/hash-mismatched model files and symlink refusal. Default production
|
|
deadlines remain 120/60 seconds. Python tests never import actual model libraries,
|
|
record a microphone, download assets or initialize OpenVR.
|