Files
bod09andClaude Opus 5.5 bbbd31fa6d Repin Parakeet Redux to the current upstream revision
moondream/parakeet-redux was rewritten upstream on 2026-09-23 to a single
commit (2bf128600aac), so the pinned revision fad622f no longer resolves.
model.safetensors still downloads from the old revision, but config.json,
ternary.json, tokenizer.json and README.md return HTTP 404, which makes
`install.sh --install-model` fail.

The weights, config, ternary map and tokenizer at 2bf1286 have the same
size and SHA-256 as the old pins, so the model itself is unchanged. Only
the model card README.md differs (8533 -> 8832 bytes), so its pin is
updated along with the revision.

Fixes #1

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 19:05:06 +01:00

8.3 KiB

Offline worker adapter and backend dispatcher

src/worker.hpp provides frameyap::Worker: call start(python, script, model, threads=2, advanced_debug=false) explicitly for the legacy Redux script, or supply the optional backend, manifest_dir, root arguments for the manifest dispatcher. Poll until ready(), then submit(id, pcm) and poll for one WorkerReply (text or privacy-safe per-request error). One request at a time; no queue, no capture and no input injection. stop() discards pending audio, terminates/reaps only its direct child (TERM, bounded 500 ms, then KILL), and is safe to repeat. Destruction stops it. start() returns without waiting for model load; poll() reports warmup/load/crash/timeout/protocol errors by throwing, then stops. Launch uses posix_spawn, safe with the host's OpenVR/SDL threads, rather than running Python setup in a forked multithreaded child. Caller must discard stale authorization/results after cancellation; this component does not implement focus or delivery policy.

The adapter requires an existing absolute, owner-private $XDG_RUNTIME_DIR (no symlink at the final component), creates its own 0700 mkdtemp directory, and writes only clip.raw with O_EXCL|O_NOFOLLOW, mode 0600. Clips are 3200..320000 finite float samples in [-1, 1], mono 16 kHz, stored as little-endian IEEE float32 (0.2..20 s). The microphone adapter saturates finite capture/resampler overshoot to this range without rescaling the rest of the clip; NaN/infinity are rejected. Worker::submit independently rejects samples outside this contract before writing a clip or request. Files are unlinked after replies or shutdown, and the private directory is removed. Private clips are not encrypted against the account owner/root; do not use an untrusted runtime directory. The caller should pass a trusted interpreter and script. By default audio/transcripts are not logged and child stderr is redirected to /dev/null. Request errors report only a fixed stage (audio, inference, response) and built-in exception category, never the exception message, arbitrary class name, traceback or text. The native receiver allowlists those labels before displaying/logging them; unknown/legacy errors remain generic.

The private pipes use unsigned LE32 payload lengths (1..65536), a one-byte message type and, for requests/replies, unsigned LE64 request ID. T + ID requests reading the fixed clip; Y means ready; F means load failure (M for missing/mismatched pinned model or private clip directory, I for a missing Python dependency, D for runtime/model load failure); R + ID + UTF-8 text and E + ID + privacy-safe UTF-8 error are replies. Text is at most 4096 bytes. An unexpected or duplicate reply, wrong ID, extra frame, closed pipe or oversized frame stops the worker. A correlated E reply or an invalid transcript is a request-level failure: the Controller drops that clip, keeps the ready model loaded and allows Record to retry. A broken protocol, worker crash or explicit in-flight cancellation is different and can require a reload. Warmup deadline is 120 s, transcription deadline 60 s; poll() must be called regularly to enforce deadlines. It never initializes a headset or starts a recording. There is no auto restart after a process failure.

Parakeet Redux is the model (moondream/parakeet-redux); moondream is its Python inference package, not a second model or cloud endpoint. Kestrel is a native dependency of that local runtime. python/frameyap/worker.py lazily imports moondream only after explicit CLI startup, with HF/Transformers/Datasets offline variables and bounded native thread-pool variables set before import. It uses md.photon("moondream/parakeet-redux", model_path=<absolute local directory>, device="cpu", cpu_threads=threads) and persistent transcribe(audio=<numpy float32>, sample_rate=16000)["text"]. Install an isolated Python runtime with the moondream 2.4.0, kestrel 0.8.0 and compatible CPU dependencies; provide preinstalled local weights from revision 2bf128600aac4b16946f7ed8372e56117fe5e23b and pass its directory explicitly. The code verifies exact sizes and SHA-256 of weights/config/tokenizer against assets/backends/redux.json via the shared, offline python/frameyap/model_files.py schema/verifier before importing model libraries. frameyap --list-models and frameyap --check-model redux --model-dir /absolute/model (or scripts/model-status.py) read local manifests and optionally hash local files; they never load the worker or download weights. For native --run, --backend ID, --model-store /absolute/store, and --manifest-dir /absolute/manifests are wired through the installed launcher as explicit overrides; default selection is saved in config.json (redux). The dispatcher python/frameyap/backend_worker.py validates a selected manifest and hashes its local model, resolves an in-release relative Python/executable launcher without a shell, then supervises one child over frameyap-worker-v1 (ready, correlated request/reply and failure frames). The launcher must accept {model_dir} and {clip_dir} (optional {threads}); matching manifest and pinned weights are not a runtime or automatically trusted new backend. Only Redux is supplied today. Protocol stdout is isolated at the file-descriptor level from third-party diagnostics. Thread limits cover Torch interop/native pools and CUDA is not selected. Offline environment flags do not prove every third-party internal is unable to access a network. Runtime/build/tests perform no downloads; the separate explicit setup utility scripts/fetch-model.py can provision the public pinned weights.

Native-only archives do not bundle the runtime; the explicit install.sh --install-runtime --yes step pip-installs it into a user-local venv (see packaging and third-party notes). No public runtime bundle has been released. Limited ARM64 measurements were taken during development; they are not a claim of complete headset acceptance.

Advanced debugging

Explicitly set "advanced_debug": true in config or enable Settings → Advanced debugging. It is off by default. A Settings change restarts the owned worker, closes the microphone and discards current work/review; the replacement worker warms normally. Turning it off stops detailed capture but does not delete previous logs. Manual config edits take effect on application restart.

With this opt-in the worker receives --advanced-debug: native/model stdout and stderr, sample counts, full exception tracebacks and recognized transcripts are captured. These logs may contain private speech, transcript text and local paths. Inspect/redact before sharing; keep out of Git. No raw audio archive is created; normal temporary clips still expire after replies/cancellation.

Files: $XDG_STATE_HOME/frameyap/worker-debug.log (fallback ~/.local/state/frameyap/worker-debug.log) and worker-debug.previous.log. Each debug worker start rotates the current log once; only these two files are retained, each bounded to 4 MiB. At the limit a marker is written and further output is drained/discarded for that worker session, not allowed to block it. The directory is owner-private 0700 and logs are 0600. Unsafe paths, symlinks, hardlinks or preexisting permissive files are refused with a visible error; there is no fallback to public temporary files. The app's single-instance lock is required to avoid competing rotation by multiple workers. Debug output never shares the framed protocol stdout. Direct worker CLI use with --advanced-debug writes to its caller's stderr; the native adapter supplies the private bounded sink.

The same opt-in also writes delivery-debug.log (rotated to delivery-debug.previous.log when debugging is enabled, capped near 1 MiB) in that directory: paced-delivery metadata only — start/sent byte counts and offsets, lease/send errors, finish results and which focus-guard check failed. It never contains transcript text.

Hardware-free tests run through CTest, including fake-child cancellation, short injected warmup/request deadlines, duplicate/stale replies, malformed frames, missing/hash-mismatched model files and symlink refusal. Default production deadlines remain 120/60 seconds. Python tests never import actual model libraries, record a microphone, download assets or initialize OpenVR.