mirror of
https://github.com/DeeJanuz/frametop.git
synced 2026-10-06 08:00:09 +02:00
Merge hands-migration into experimental
Hand tracking (ft-camd, ft-hands, and their tools) joins the desktop. The hands file and ring move to /run/user/UID/frametop-hands/, which ft-screens' hand cutouts now read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
commit
8bc6739dca
58 files changed
+7682
-54
No files matched your search
@@ -0,0 +1,140 @@
|
||||
# Hands in Frametop: migration plan
|
||||
|
||||
Hand tracking from the headset's own cameras has been built as a separate project, frame-hands (`~/Desktop/Projects/frame-hands` on the Frame, a local git repo with no remote). The plan is to make it a native Frametop component, like `gaze/` and `power/`, instead of a separate module. The work happens on branch `hands-migration` (worktree `frametop-hands/` in the PC workspace) and is merged into `experimental` after it's been tested in the headset.
|
||||
|
||||
Builds from this worktree must sync to their own folder on the Frame, never `~/dev/frametop`: run every script with `FRAME_REPO=/home/steamos/dev/frametop-hands`.
|
||||
|
||||
## Status (2026-09-30)
|
||||
|
||||
Steps 1-5 are done:
|
||||
|
||||
- frame-hands' pending work was committed there (6c63c9e).
|
||||
- Its filtered history was merged under `hands/` (1a76d15), then laid out (`trackd/` to `track/`).
|
||||
- The renames, the Frametop paths, and ft-camd's file capabilities are done. So are `hands/Makefile`, `build.sh`, `run.sh`, the two units, the README, the settings, the installer step, and ft-screens on the shared header.
|
||||
- Built in the dev container on the Frame, and on the 7i.
|
||||
- Checked without the headset:
|
||||
- `ft-handreplay` against frame-hands' `fh-replay`, both x86 with `--cost`, on the whole dim recording and the first 60 s of the bright one: identical summaries and byte-identical depth dumps. The Makefile's own ncnn build is included in that.
|
||||
- `ft-ringplay` into `ft-hands` on the 7i tracked, pinched, and wrote `/run/user/UID/frametop-hands/{hands,gestures}`.
|
||||
- On the Frame, ft-hands in the container finds the calibration through `/run/host/persist`, and ft-camd without its capabilities refuses with a clear message.
|
||||
|
||||
Step 6 has started (2026-09-30 10:30):
|
||||
- `hands/run.sh install` is done, and both services run from this worktree.
|
||||
- The files moved to `/run/user/UID/frametop-hands/`, because `/run/user/UID/frametop` is the desktop session's own runtime folder, deleted at every desktop start.
|
||||
- Until the desktop restarts from a build with this branch's ft-screens, the link `/run/user/UID/frame-hands -> frametop-hands` feeds the running one. It's tmpfs, so it's gone at reboot.
|
||||
|
||||
Found in the headset:
|
||||
- The side cameras were swapped (`HANDS_SWAP_SIDES=1`).
|
||||
- The cutout copy shader lost resolution at `mediump` (now `highp`).
|
||||
- Colour capture isn't reliable (see the README).
|
||||
- Two pinch fixes: one hand no longer pinches both sides, and the palm-down limit stops typing pinches.
|
||||
|
||||
## What frame-hands is today
|
||||
|
||||
| Part | What it is | Size |
|
||||
| --- | --- | --- |
|
||||
| `camd/` | `fh-camd`, the camera broker (C). It borrows XRService's camera DMA-BUFs read-only with `pidfd_getfd`, times them with the `v4l2_dqbuf` tracepoint, and publishes the four IR cameras (and optionally the two colour cameras) to a shared-memory ring. It starts as root and drops to the user after setup. Adapted in part from FrameEyeCameraFeed (MIT, licence file kept). `fh-camprobe` is its discovery and recording probe. | camd 1.1k lines, tp 0.4k, xrcams 0.8k, camprobe 1.1k |
|
||||
| `trackd/` | `fh-tracker` (C++): the tracker, the models on ncnn, the calibration (jsoncpp), the pinch detector, the publisher, and the recorder. Also `fh-replay` (offline replay and scoring), `fh-ringplay` (plays a recording into a ring), and `nettest`. | 2.9k lines |
|
||||
| `include/` | The hands file (`fh_hands.h`, read by ft-screens) and the gestures file (`fh_gestures.h`, pinches). | |
|
||||
| `models/ncnn/` | MediaPipe's palm detector and hand landmark model, from the OpenCV Zoo ONNX ports (Apache-2.0), converted to ncnn in float and int8. | 5.9 MB |
|
||||
| `tools/` | Python analysis: side-camera check, colour calibration check, frame viewer, gesture watcher, depth report, model comparison, int8 calibration, model conversion. | ~1.1k lines |
|
||||
| `tracker/` | The Python prototype of the tracker. Some tools import its `calib.py` and `models.py`. | 1.3k lines |
|
||||
| `probes/`, `notes/`, `re/`, `shim/` | One-off experiments, reverse-engineering notes on SteamVR's passthrough internals, a disassembly (not in git), and a header from an abandoned XRService shim approach. | |
|
||||
| `vendor/`, `captures/` | ncnn and FrameEyeCameraFeed clones, and recordings of the user's hands and room (tens of GB). Neither is in git. | |
|
||||
|
||||
Today it runs by hand: `sudo camd/fh-camd`, then `trackd/fh-tracker`. There are no units and no installer. Files: `/run/frame-hands/ir-ring` (the ring, in a root-owned folder), and `$XDG_RUNTIME_DIR/frame-hands/hands` and `gestures`.
|
||||
|
||||
Frametop already has the consumer side on `experimental`: `screens/handcut.{h,cpp}` cuts the hands out of the screens, with its own copy of the hands file layout, and `screens/handtest.cpp` tries it on a test panel.
|
||||
|
||||
## Where it goes
|
||||
|
||||
A top-level `hands/` folder, laid out like `gaze/`:
|
||||
|
||||
```
|
||||
hands/
|
||||
README.md # from trackd/README.md and camd/README.md
|
||||
build.sh # ft-camd, ft-hands; --tools also builds the replay tools
|
||||
run.sh # install|uninstall|start|stop|restart|status|log
|
||||
frametop-camd.service # user units (templates, @REPO@)
|
||||
frametop-hands.service
|
||||
include/ # fhring.h, fh_hands.h, fh_gestures.h: shared with screens/ and pointer/
|
||||
camd/ # ft-camd: camd.c tp.c xrcams.c, LICENSE.FrameEyeCameraFeed
|
||||
track/ # ft-hands: tracker, nets, calib, pinch, publish, record; replay.cpp
|
||||
# (ft-handreplay) and ringplay.cpp (ft-ringplay) for recordings
|
||||
models/ # the ncnn models, with NOTICE (Apache-2.0, MediaPipe / OpenCV Zoo)
|
||||
tools/ # the Python checks, watch_gestures, depth_report, calib.py, ring.py
|
||||
```
|
||||
|
||||
Left behind in frame-hands, which stays as the lab: the recordings, the Python prototype (the tools that need `calib.py` or `models.py` get a trimmed copy in `hands/tools/`), `probes/`, `notes/`, `re/`, `shim/`, `camprobe`, and `vendor/`. The reverse-engineering notes don't belong in a public repo, and recordings are images of the user's hands and room, so they never go into git.
|
||||
|
||||
## Names
|
||||
|
||||
Programs within 15 characters, `ft-` prefix; files under `frametop`:
|
||||
|
||||
| Now | In Frametop |
|
||||
| --- | --- |
|
||||
| `fh-camd` | `ft-camd` |
|
||||
| `fh-tracker` | `ft-hands` |
|
||||
| `fh-replay`, `fh-ringplay` | `ft-handreplay`, `ft-ringplay` |
|
||||
| `/run/frame-hands/ir-ring` | `/run/user/UID/frametop-hands/cam-ring` |
|
||||
| `$XDG_RUNTIME_DIR/frame-hands/hands`, `gestures` | `/run/user/UID/frametop-hands/hands`, `gestures` |
|
||||
|
||||
The source keeps its `fh_` identifiers and header names (`fh_hands.h`, `fh_gestures.h`, `fhring.h`), and the file formats keep their magic strings, so recordings and tools from frame-hands keep working. Programs, units and runtime paths change.
|
||||
|
||||
## Build
|
||||
|
||||
- `hands/build.sh` builds in the dev container through `scripts/frame.sh --build`, into `hands/build/`, like the other components. `FRAME_BUILDER=pc` can take the ncnn build.
|
||||
- ncnn: fetched at a pinned tag (20260526, as now) into `hands/build/ncnn` and built once, the way `screens/build.sh` fetches the OpenVR header, with frame-hands' options so results match. `NCNN=` points the build at an existing install instead. Every net runs single-threaded (`num_threads = 1`), with the tracker spreading nets over its own pinned threads, so OpenMP could go later.
|
||||
- ft-hands runs in the dev container like ft-pointer and ft-powerd (`distrobox enter dev --`, after `scripts/container-up.sh`). Today's fh-tracker runs on the host and works only because the host happens to have the same `libjsoncpp.so.25` and libgomp as the container. Inside the container the calibration is at `/run/host/persist`, and calib.cpp (and `tools/calib.py`) fall back to it when `/persist` isn't there.
|
||||
- ft-camd has to run on the host (below), so it's linked statically (only libc and libm; `glibc-static` goes into `setup/dev-container.sh`). The host has an older glibc than the container.
|
||||
|
||||
## Running it
|
||||
|
||||
**ft-camd needs privileges**, only while it sets up: `pidfd_getfd` on XRService (the Frame has `ptrace_scope=1`), system-wide tracepoints (`perf_event_paranoid=2`), and the tracepoint files, which are root-only (`/sys/kernel/tracing/events/v4l2/v4l2_dqbuf/{id,format}` are mode 0440). A rootless container's root can't do any of that, so it runs on the host. Two ways:
|
||||
|
||||
- **A. File capabilities (chosen, 2026-09-30).** The installer runs `sudo setcap cap_sys_ptrace,cap_perfmon,cap_dac_read_search+ep hands/build/ft-camd` once. ft-camd then runs as the user, in a user unit `PartOf=steamvr.service`, so it starts and stops with SteamVR, and its ring lives in the user's runtime folder. It drops all capabilities after setup, as it drops root today. Nothing ever runs as root. Writing the file clears its capabilities, so a rebuilt ft-camd needs the setcap again. It changes rarely. `/home` on the Frame is ext4 without `nosuid`, so file capabilities work there.
|
||||
- **B. Root system service**, like the Bluetooth fixes: a root-owned copy in `/var/lib/frametop/`, a unit in `/etc/systemd/system/`. It would have to watch for XRService itself, because a system unit can't follow the user's `steamvr.service`.
|
||||
|
||||
Either way the password is needed once at install, through the same `sudo -S` path the Bluetooth fixes use, and only after asking.
|
||||
|
||||
**ft-hands** is a user unit, `frametop-hands.service`: after `frametop-camd.service`, `PartOf=steamvr.service`, nice 5, model threads on CPUs 5-7 (measured best on 2026-09-29).
|
||||
|
||||
**Settings** in `~/.config/frametop.conf`: `HANDS_SWAP_SIDES=1` and `HANDS_CPUS=5,6,7`, read by ft-hands. It's on while its services are installed (`hands/run.sh install`, `uninstall`), so there's no `HANDS` switch. There's no setting for colour yet. Later, a switch in Frametop Display Settings.
|
||||
|
||||
**Installer:** an optional last step in `install.sh`, off by default, which asks first because it needs sudo.
|
||||
|
||||
## Interfaces
|
||||
|
||||
- `screens/handcut.cpp` includes `hands/include/ft_hands.h` instead of its own copy of the layout, and reads the new path. ft-screens and ft-hands change together on this branch.
|
||||
- Pinches go to the pointer helper. It maps the gestures file and checks the begin and end counters each tick. A begin is a press, an end the release, and the pinch point's movement a drag. In gaze mode, the press lands where you look. The counters mean a quick tap between two ticks isn't missed. The tracker knows nothing about the pointer.
|
||||
|
||||
## Open items that aren't part of the move
|
||||
|
||||
These block shipping hands to other people, not the migration:
|
||||
|
||||
- **The side-camera swap.** After some XRService restarts, fh-camd publishes the two side cameras under each other's names. Today it's caught by hand (`tools/check_sides.py --ring`, then `--swap-sides`). It needs fixing at the source (tell the buffers apart by the `dqbuf` tracepoint's device, the way the colour pair is split), or at least an automatic check at start-up.
|
||||
- **The colour cameras' calibration mapping** (`tools/check_color.py` on a recording with texture).
|
||||
- **Depth when one camera loses the hand.** From the 2026-09-30 replay measurements: drifting 10% per update toward the one-camera guess (`kMonoDepthGain`) makes the depth worse than keeping the last distance. Try 0.02.
|
||||
|
||||
## Public repo
|
||||
|
||||
Frametop is public. **Not pushed to GitHub until the user says it's ready** (user decision, 2026-09-30). When it is, it publishes:
|
||||
|
||||
- The camera borrowing (`pidfd_getfd` on XRService's buffers) and the tracepoint timing. FrameEyeCameraFeed already does the same publicly. Its MIT licence and credit stay with the code.
|
||||
- The models, under Apache-2.0, with a NOTICE.
|
||||
|
||||
It doesn't publish the reverse-engineering notes, the probes, or any recording. They stay in frame-hands.
|
||||
|
||||
## History
|
||||
|
||||
frame-hands' work was committed there first (6c63c9e, its 4th commit). Its history was then filtered to drop what stays behind (`notes/`, `probes/`, `shim/`, `camd/camprobe.c`, the camprobe tools, the Python prototype except `calib.py` and `models.py`, and `.frame-job`) from every commit. It was merged into this branch under `hands/` (a subtree merge), so blame still leads to where each line came from. The renames come after, as their own commits.
|
||||
|
||||
## Steps
|
||||
|
||||
1. In frame-hands: commit the pending work, as its last state before the move (needs the user's OK).
|
||||
2. On this branch: import it under `hands/`, then rename the programs and paths. The behaviour stays identical.
|
||||
3. `hands/build.sh`, `run.sh`, the two units, the README, the settings, and the installer step.
|
||||
4. ft-screens' hand cutouts on the shared header and the new path.
|
||||
5. Check without the headset. `ft-handreplay` on the 2026-09-29 recordings with `--cost` is repeatable, so its summary must match `fh-replay`'s exactly. And `ft-ringplay` into ft-hands must publish the same hands as into fh-tracker.
|
||||
6. In the headset, with the user and after asking: stop fh-camd and fh-tracker, install the services from `~/dev/frametop-hands`, and restart the desktop from this branch so ft-screens reads the new path.
|
||||
7. Pinch into the pointer helper (it can also follow the merge). The gaze work is on the Frame's `~/frametop` main: on 2026-09-30 that branch had 4 commits `experimental` doesn't have, plus uncommitted work in the pointer helper's gaze mode. Build this step on wherever that work lands, not on this branch's older copy.
|
||||
8. Merge into `experimental`. It's checked out in a worktree on the Frame (`~/frametop/.worktrees/experimental`, where the live desktop runs), so the merge happens there, or `experimental` is switched away first.
|
||||
@@ -172,6 +172,17 @@ power/run.sh off | on # the displays off now, or back on
|
||||
power/run.sh log
|
||||
```
|
||||
|
||||
## Hand tracking (experimental)
|
||||
|
||||
Your hands show over the screens: where a tracked hand is between an eye and a screen, ft-screens lets that eye see the room through the screen. The same tracker also detects pinches, for clicking where you look with the gaze pointer (not wired to the pointer yet). It's optional: `hands/run.sh install`, or the last step of `install.sh`.
|
||||
|
||||
- `ft-camd` borrows XRService's camera buffers and publishes the four IR tracking cameras to `/run/user/UID/frametop-hands/cam-ring`. It runs on the host as `frametop-camd.service`, with file capabilities that `hands/run.sh install` sets through sudo, and it drops them once set up. A rebuild clears them: `hands/run.sh caps`.
|
||||
- `ft-hands` runs in the `dev` container as `frametop-hands.service`. It finds and triangulates the hands, and publishes `hands` (read by ft-screens' cutouts) and `gestures` (pinches) next to the ring.
|
||||
- Both start and stop with SteamVR. `hands/run.sh status` and `hands/run.sh log` show how they're doing.
|
||||
- Settings in `~/.config/frametop.conf`: `HANDS_SWAP_SIDES` (after some SteamVR restarts the side cameras' names come out swapped, and hands land beside the holes; `hands/tools/check_sides.py --ring` tells) and `HANDS_CPUS`.
|
||||
|
||||
Details, options, and the recording and replay tools are in [hands/README.md](../hands/README.md).
|
||||
|
||||
## Remote desktop over VNC
|
||||
|
||||
With `REMOTE=1` in the config (`desktops.sh remote on`), the desktop is also served over VNC, for RealVNC Viewer or macOS Screen Sharing. `desktops.sh remote info` prints the address and password.
|
||||
|
||||
@@ -0,0 +1,3 @@
|
||||
# Model sources that tools/convert_models.py downloads; the converted ncnn models are kept
|
||||
models/onnx/
|
||||
models/*.task
|
||||
@@ -0,0 +1,49 @@
|
||||
# Hand tracking, built into build/ (hands/build.sh runs this in the dev container):
|
||||
# make ft-camd (camd/: runs on the host, so linked statically) and ft-hands (track/)
|
||||
# make tools ft-handreplay and ft-ringplay, for recordings
|
||||
# The first build fetches ncnn (NCNN_TAG) and builds it into build/ncnn, which takes a few
|
||||
# minutes. NCNN=DIR uses an ncnn install already built instead.
|
||||
NCNN_TAG = 20260526
|
||||
NCNN ?= build/ncnn/install
|
||||
CFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter
|
||||
CXXFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter -Wno-psabi
|
||||
CXXFLAGS += -std=c++17 -fopenmp -I$(NCNN)/include/ncnn
|
||||
LDLIBS = $(NCNN)/lib/libncnn.a -ljsoncpp -fopenmp -lpthread
|
||||
|
||||
CAMD = camd/camd.c camd/tp.c camd/xrcams.c
|
||||
TRACK = track/calib.cpp track/nets.cpp track/tracker.cpp track/io.cpp track/record.cpp track/pinch.cpp
|
||||
HDR = $(wildcard track/*.h) camd/fhring.h include/fh_hands.h include/fh_gestures.h
|
||||
|
||||
all: build/ft-camd build/ft-hands
|
||||
tools: build/ft-handreplay build/ft-ringplay
|
||||
|
||||
build/ft-camd: $(CAMD) camd/tp.h camd/xrcams.h camd/fhring.h
|
||||
@mkdir -p build
|
||||
$(CC) $(CFLAGS) -static -o $@ $(CAMD) -lm
|
||||
|
||||
build/ft-hands: track/main.cpp $(TRACK) $(HDR) $(NCNN)/lib/libncnn.a
|
||||
@mkdir -p build
|
||||
$(CXX) $(CXXFLAGS) -o $@ track/main.cpp $(TRACK) $(LDLIBS)
|
||||
|
||||
build/ft-handreplay: track/replay.cpp $(TRACK) $(HDR) $(NCNN)/lib/libncnn.a
|
||||
@mkdir -p build
|
||||
$(CXX) $(CXXFLAGS) -o $@ track/replay.cpp $(TRACK) $(LDLIBS)
|
||||
|
||||
build/ft-ringplay: track/ringplay.cpp track/record.h camd/fhring.h
|
||||
@mkdir -p build
|
||||
$(CXX) $(CXXFLAGS) -o $@ track/ringplay.cpp
|
||||
|
||||
# ncnn as frame-hands built it (the models were converted and quantized for it), minus its tools
|
||||
build/ncnn/install/lib/libncnn.a:
|
||||
rm -rf build/ncnn && mkdir -p build/ncnn
|
||||
git clone -q --depth 1 --branch $(NCNN_TAG) -c advice.detachedHead=false https://github.com/Tencent/ncnn.git build/ncnn/src
|
||||
cmake -S build/ncnn/src -B build/ncnn/build -G Ninja -Wno-dev -DCMAKE_BUILD_TYPE=Release \
|
||||
-DCMAKE_INSTALL_PREFIX=$(CURDIR)/build/ncnn/install -DCMAKE_INSTALL_LIBDIR=lib -DNCNN_VULKAN=OFF \
|
||||
-DNCNN_OPENMP=ON -DNCNN_INT8=ON -DNCNN_SIMPLEOCV=ON -DNCNN_BUILD_TOOLS=OFF -DNCNN_BUILD_EXAMPLES=OFF \
|
||||
-DNCNN_BUILD_BENCHMARK=OFF -DNCNN_BUILD_TESTS=OFF -DNCNN_PYTHON=OFF > build/ncnn/cmake.log
|
||||
cmake --build build/ncnn/build --target install > build/ncnn/build.log
|
||||
|
||||
clean:
|
||||
rm -f build/ft-camd build/ft-hands build/ft-handreplay build/ft-ringplay
|
||||
|
||||
.PHONY: all tools clean
|
||||
+159
@@ -0,0 +1,159 @@
|
||||
# Hands (experimental)
|
||||
|
||||
Hand tracking from the headset's own cameras. It serves two things in Frametop:
|
||||
|
||||
- **Hand cutouts:** where your hand is between an eye and a screen, that eye sees the room through the screen (ft-screens, `screens/handcut.cpp`), so your hands show over the screens the way they do on a Vision Pro.
|
||||
- **Pinches:** look at something and pinch to click it, pinch and move to drag, with the eye tracker doing the looking (`gaze/`). The tracker publishes the pinches. The pointer helper doesn't read them yet.
|
||||
|
||||
Two programs, each a user service that starts and stops with SteamVR:
|
||||
|
||||
- `ft-camd` (`camd/`, C) borrows XRService's camera buffers and publishes the four IR tracking cameras' frames to a shared-memory ring. It runs on the host.
|
||||
- `ft-hands` (`track/`, C++) finds hands in those frames with MediaPipe's palm and landmark models on ncnn, triangulates them, and publishes them. It runs in the dev container.
|
||||
|
||||
```
|
||||
hands/run.sh install # build, give ft-camd its capabilities (sudo, once per build), enable
|
||||
hands/run.sh status # the services, and ft-hands' last status lines
|
||||
hands/run.sh log [lines]
|
||||
hands/run.sh restart # after changing a setting
|
||||
hands/run.sh caps # after rebuilding ft-camd (a rebuild clears its capabilities)
|
||||
hands/run.sh uninstall
|
||||
```
|
||||
|
||||
Settings in `~/.config/frametop.conf` (`FT_<name>` in the environment overrides them):
|
||||
|
||||
- `HANDS_SWAP_SIDES=1`: the two side cameras' names are swapped (see ft-camd below). Check with `tools/check_sides.py --ring`.
|
||||
- `HANDS_CPUS=5,6,7`: the CPUs the model threads run on (below).
|
||||
|
||||
Files, all in `/run/user/UID/frametop-hands/` (private to the user; not `/run/user/UID/frametop/`, which the desktop session deletes whenever it starts):
|
||||
|
||||
| File | Written by | Layout | Read by |
|
||||
| --- | --- | --- | --- |
|
||||
| `cam-ring` | ft-camd | `camd/fhring.h` | ft-hands, `tools/ring.py` |
|
||||
| `hands` | ft-hands | `include/fh_hands.h` | ft-screens (`screens/handcut.cpp`) |
|
||||
| `gestures` | ft-hands | `include/fh_gestures.h` | `tools/watch_gestures.py`; the pointer helper, later |
|
||||
|
||||
The source keeps the `fh_` names and magic strings of frame-hands, where this was developed (`~/Desktop/Projects/frame-hands` on the developer's Frame, which keeps the recordings, probes and Python prototype). So its recordings and tools still work.
|
||||
|
||||
## ft-camd
|
||||
|
||||
XRService owns the headset cameras. ft-camd borrows its DMA-BUFs read-only with `pidfd_getfd`, the same way FrameEyeCameraFeed does. It never touches XRService's V4L2 descriptors. `discovery` in `camd/xrcams.c` is adapted from FrameEyeCameraFeed (MIT, see `camd/LICENSE.FrameEyeCameraFeed`).
|
||||
|
||||
Polling buffers for changes can catch a frame while the camera is still writing it. Instead, ft-camd listens to the `v4l2:v4l2_dqbuf` tracepoint, which fires when XRService takes a buffer. It gives the buffer index, the sequence number and the capture timestamp. ft-camd learns which DMA-BUF holds each V4L2 index by watching which buffer changes at each dequeue:
|
||||
|
||||
- Right after XRService allocates its buffers, the mapping is allocation order.
|
||||
- After XRService restarts streaming, the order is shuffled, and the mapping is learned index by index.
|
||||
- The two upper cameras share one run of buffers. For them, only allocation order can tell the cameras apart.
|
||||
- It also re-maps an index on the fly when its buffer holds no new frame.
|
||||
|
||||
**Privileges.** Setting up needs three things. `pidfd_getfd` on XRService needs `CAP_SYS_PTRACE`, because the Frame has `ptrace_scope=1`. The system-wide tracepoint needs `CAP_PERFMON`, because `perf_event_paranoid` is 2. Its format files are root-only, which needs `CAP_DAC_READ_SEARCH`. `hands/run.sh install` gives the binary those capabilities with `sudo setcap`. ft-camd drops them all once it has set up, before it reads a frame, and then runs as you. XRService runs as you too. It also runs under `sudo`, for trying it by hand, and then drops to the user who ran sudo. It reads nothing from the ring's readers.
|
||||
|
||||
The ring is mode 0600, in a folder only you can write. Frame handling:
|
||||
|
||||
- Only complete, bright frames are published. The cameras alternate a normal exposure with a near-black one, so each camera gets 30 of its 60 fps.
|
||||
- A copy torn by the camera overwriting the buffer is dropped.
|
||||
- Each copy takes about 0.1 ms, and a cache sync about 0.15 ms.
|
||||
|
||||
Options:
|
||||
|
||||
- `--with-dark`: also publish the near-black frames, as extra ring cameras flagged `FH_CAM_DARK`. They show only light sources, so they're no use for hands.
|
||||
- `--with-color`: also publish the two Arcturus colour cameras, flagged `FH_CAM_COLOR`. Each is the luma of the 10-bit frame's valid 1972x2464 (the top 8 bits), at half size (`--color-scale 2`: 986x1232) and at most 30 fps (`--color-fps`; the cameras run at 60). Frames that carry the module's warped half-size copy are dropped. Their `capture_ns` is on the colour module's clock (2.2 s off the mono cameras' on 2026-09-29), so line them up with the mono cameras by `dqbuf_ns`. Each frame costs about 0.65 ms of cache sync and 1.1 ms of decoding, so both cameras at 30 fps take about 11% of a core.
|
||||
- The ring holds 8 cameras: 4 mono, plus 4 dark twins or 2 colour cameras.
|
||||
- Colour isn't reliable yet. In the lit-room test of 2026-09-30, the colour cameras kept losing their buffer mapping: 30 frames in a row looked unchanged, the camera relearned, and after 5 relearns ft-camd exited. Each relearn samples all 32 colour buffers, which also made the mono cameras miss frames. Runs with the headset idle (no hands, no cutouts) had none of this, and no half-size copies either, while the failing runs had many. So the passthrough compositor may be writing into the colour buffers while Room View shows. Whether a frame is new is judged on the luma rows only: the chroma after them hardly changes in a lit room. `FT_CAMD_DEBUG=1` prints, at each colour stale frame, how many sampled words changed in every candidate buffer.
|
||||
- `--sensor S`: only the mono cameras whose sensor name contains S.
|
||||
- `--status S`: a status line every S seconds (0: never).
|
||||
|
||||
It exits when XRService exits, or when a camera's buffers keep going stale, which means XRService has reallocated them. The service starts it again, and it attaches to the new buffers.
|
||||
|
||||
**Which camera is which:** video9 is `slam_left`, video13 is `slam_right`, video6 is `upper_left` and video7 is `upper_right`. This was checked by rendering the same view from each camera with the factory calibration. But ft-camd tells the side cameras' buffers apart only by XRService's allocation order, and after some XRService restarts it gets them backwards. Then every hand is seen by one camera only, at the wrong depth, and the hand holes land beside the hands. With the headset on, looking at a room with some texture, `tools/check_sides.py --ring` says whether the names are right (exit 0), swapped (exit 3), or it can't tell (exit 2). When they're swapped, set `HANDS_SWAP_SIDES=1`. The colour cameras are video3 (`arcimx616 0-0010`) and video0 (`0-001a`); which of them is `passthrough_left` in the module's calibration is for `tools/check_color.py` to settle, on a recording with texture in view.
|
||||
|
||||
## ft-hands
|
||||
|
||||
```
|
||||
hands/build/ft-hands # status every 5 s; Ctrl+C to stop
|
||||
hands/build/ft-hands --int8 # the 8-bit models (models/ncnn/*-int8.ncnn.*)
|
||||
```
|
||||
|
||||
Run it in the dev container (`distrobox enter dev -- ...`). It reads the factory calibration from `/persist` (`/run/host/persist` in the container).
|
||||
|
||||
Options:
|
||||
|
||||
- `--threads N`: model threads, pinned to the `--cpus` list. Default 3.
|
||||
- `--cpus LIST`: CPUs for the model threads and the main loop. Default `5,6,7` (`HANDS_CPUS`). SteamOS starts user processes on CPUs 0-4, and XRService's head tracking runs on 2-3. With the headset on, a step took 8.4 ms on 5-7 against 13.2 ms on 2-4, and SteamVR's frame timing didn't change (2026-09-29, three rounds of the same replayed frames).
|
||||
- `--contrast MODE` or `PALM/HAND`: how crops are equalized before the models see them: `clahe[:CLIP]`, `none`, or `stretch` (1st-99th percentile). Default `clahe:2/none`. In the dim recording, CLAHE let the palm search find about 10% more hands, but it made the landmarks jitter more (published median 6.9 mm, against 6.0 mm with plain landmark crops).
|
||||
- `--swap-sides`: swap the two side cameras (`HANDS_SWAP_SIDES`, see ft-camd).
|
||||
- `--seconds N`: stop after N seconds.
|
||||
- `--status S`: how often to print status, in seconds.
|
||||
- `--models DIR`: where the models are.
|
||||
- `--nice N`: niceness. Default 5, so the VR stack wins contested CPUs.
|
||||
- `--no-publish`: don't write the hands and gestures files.
|
||||
- `--record DIR`, `--record-for S`: save every frame set for S seconds (default 120) to `DIR/sets.bin`. That's about 80 MB/s. Sending the tracker SIGUSR1 (`pkill -USR1 -x ft-hands`) starts a recording in `~/.local/share/frametop/hands/rec-<time>` without a restart. Recordings are images of your hands and room: they stay on the headset unless you move them.
|
||||
- `--record-only`: record without tracking or publishing, so it can run beside the live tracker. Give it `--record DIR`, since SIGUSR1 would reach both trackers. With `ft-camd --with-dark`, recordings also hold each camera's newest dark frame as `<name>_dk`, which doubles the rate. With `--with-color`, each colour camera's newest frame is saved with every set, as `color_video<N>`, which adds about 70 MB/s. Run the recorder at normal I/O priority: idle I/O priority stalled a 165 MB/s recording.
|
||||
- `--keep-presence P`: the landmark presence a tracked view needs to stay tracked. New views always need 0.5. Default 0.5. Lowering it to 0.2 barely helped in the bright recording, because lost hands drop to near-zero presence.
|
||||
- `--ring PATH`: read frames from another ring, such as `ft-ringplay`'s.
|
||||
|
||||
The status line also says how often a hand was on each side (by where the wrist is), and why views and hands came and went: views lost (the landmark model stopped seeing the hand), handoff misses (a crop projected from the hand's 3D position found nothing), duplicates, splits (two views disagreed in 3D), and hands created, merged and forgotten.
|
||||
|
||||
### Scheduling
|
||||
|
||||
- Each hand is tracked in its best two cameras, the way MediaPipe tracks: the landmark model runs on a crop placed from the previous landmarks, with no palm detection.
|
||||
- A hand seen in too few cameras is projected into the others through the calibration. Where it lands well inside a camera, that camera gets a crop to try. This is how a hand raised out of the side cameras reaches the upper ones.
|
||||
- The palm detector runs only while fewer than two hands are tracked, at most 5 times a second, on a few zoomed tiles per search. Tiles are picked in proportion to how likely hands are there. Each tile is turned so the expected shoulder-to-hand direction points up.
|
||||
- Frame sets are processed at 30 Hz while a hand moves faster than 0.25 m/s (or a pinch is down or closing), at 15 Hz otherwise, and at 5 Hz while no hand is in view.
|
||||
|
||||
### 3D
|
||||
|
||||
- **Two or more views:** each landmark is triangulated from the camera rays, weighted by the model's presence score. The median ray distance is reported as the residual.
|
||||
- **Pairing views across cameras.** The side cameras sit side by side, so two hands next to each other at the same height fall on the same epipolar lines, and rays to two different hands can nearly meet close to the cameras. That made phantom hands 12-15 cm in front of the eyes, which tore holes through the screens. Each step now scores every way of pairing the views in two cameras and keeps the best. A pair scores well when its rays meet, when each view's apparent size matches the triangulated distance, and when the model calls both the same hand. The size check uses a fixed prior: with the model's average hand, clean pairs measure 0.71-1.51 times the one-view distance, and mismatched pairs mostly far less.
|
||||
- **One view:** depth comes from the model's metric world landmarks, their spread across the palm against the angle it covers in the image, scaled by the user's hand size (learned while two views are available). That distance is off by 10-30% and wanders about 10% between frames, so a hand that drops to one camera keeps its last distance and drifts toward the one-view guess by 10% a frame.
|
||||
- **Smoothing.** The published landmarks go through a One Euro filter: it smooths hard while the hand is still (tracking noise is several mm per frame) and hardly at all while it moves fast. The palm speed that sets the update rate is the filtered one; the raw speed read about 0.25 m/s from noise alone.
|
||||
- **Capsules.** Forearms follow the hand's own axis, and nothing within 12 cm in front of the eyes is published.
|
||||
|
||||
How good the depth is, measured from recordings (2026-09-30, `--depth` below): the two lower cameras see the hands about 77% of the time, a lower and an upper camera 7-12%, and one camera 12-15%. Depth is the noisy direction. With the lower pair, it jitters 4-6 times as much as sideways position (published: 3-7 mm against 1-2 mm). The one-camera guess is a median 2-6 cm off. When a camera drops out, drifting 10% a frame toward that guess is worse than keeping the last distance (after 0.5 s a median 23-30 mm off, against 11-12 mm).
|
||||
|
||||
## Pinch
|
||||
|
||||
ft-hands detects a pinch per hand (`track/pinch.h`) and publishes it to the gestures file. The layout, and how to read it without missing quick taps, is in `include/fh_gestures.h`.
|
||||
|
||||
- A pinch begins when the thumb and index tips come within `--pinch-begin` (default 0.020 m). It ends when they open past `--pinch-end` (0.035 m) for 2 processed frames in a row, or when the hand stays lost for 0.25 s (flagged lost).
|
||||
- The distance comes from MediaPipe's world landmarks: the model's own 3D hand pose, averaged over the hand's views, at the user's hand size. `--pinch-triangulated` uses the triangulated tips instead. On two recordings without deliberate pinches, the world landmarks came under 2 cm in 0.2-1% of frames, against 3.3-4.5% for the triangulated tips. In the dim recording, typing still gave 2 pinches a minute before the palm check below.
|
||||
- No pinch begins while the palm faces down (`--pinch-palm-down MAX`: the palm normal's share of the head's up axis, default 0.6; 1 turns it off), and a close held back that way has to open again before a pinch can begin. Typing curls the thumb onto the index. In the lit recording of 2026-09-30, typing on a keyboard in the lap began 23 pinches in about 2 minutes, all with the palm facing down (0.69-1.00), while the 26 deliberate ones read 0.00-0.50. The limit held back every typing pinch and none of the deliberate ones. Looking down tilts the head frame, which lowers the reading for a hand on a keyboard, so the consumer's gaze check stays the other guard.
|
||||
- A hand a pinch is down on stays with that side until the pinch ends. The left/right call is a running average of the model's, and when it flipped mid-pinch, the other side took the same hand and both sides pinched at once.
|
||||
- The pinch point is midway between the thumb and index tips. A drag is the pinch point now, minus where it was when the pinch began, both turned into the room with the HMD pose at their capture times.
|
||||
- `tools/watch_gestures.py` prints begins, ends and drag offsets live, and `--distance` prints each hand's distance.
|
||||
|
||||
The pointer helper is the natural consumer. Its gaze mode already treats a press as "stop where the gaze put it, drag onto the target, click on release", and "hold still for half a second, then move" as a drag. A pinch begin would be the press, the end the release, and the pinch point's movement the drag.
|
||||
|
||||
## Recordings
|
||||
|
||||
`hands/build.sh --tools` also builds the offline tools.
|
||||
|
||||
`ft-handreplay DIR` runs a recording through the tracker with the live scheduling and reports how well it kept the hands: hands per set, left and right coverage, track lengths, pinches, jitter, and the same reasons as the status line.
|
||||
|
||||
```
|
||||
hands/build/ft-handreplay ~/.local/share/frametop/hands/rec-20260929-120000 --cost --oracle 10 --timeline /tmp/tl.txt
|
||||
```
|
||||
|
||||
- `--cost`: instead of timing the steps, charge each round of model calls what it typically costs live (10 ms landmarks, 18 ms palms), so results repeat exactly.
|
||||
- `--oracle N`: every N-th set, also search every tile of every camera, and report how often the tracker had the hands that full search could find.
|
||||
- `--slow F`: live, the tracker skips sets that arrive while it's busy. Replay counts each step's time times F as busy (default 1; the headset is busier live).
|
||||
- `--timeline FILE`: a line per processed set and hand, with pinch events and distances.
|
||||
- `--cams mono|color|all`: which cameras to track with (default `mono`). `color` tracks with the Arcturus pair alone, for comparing it with the IR cameras on the same recording. It needs a recording made with `ft-camd --with-color`. `--color-left NODE` (`color_video0` or `color_video3`) and `--color-crop subtract|none` say how the module's calibration maps onto the images; `tools/check_color.py` finds out.
|
||||
- `--depth FILE`: a line per hand per processed set for `tools/depth_report.py`, which measures the depth without ground truth: how the hands were seen, the noise along the line of sight against across it, each camera's one-view distance against the triangulated one, and what a camera dropping out would do.
|
||||
- The pinch, contrast and presence options are ft-hands'.
|
||||
|
||||
`ft-ringplay DIR --ring PATH [--from S] [--to S] [--loop]` publishes a recording into a ring file in real time, as ft-camd would, so `ft-hands --ring PATH --no-publish` runs the same frames run after run. It needs no privileges, and it skips the dark frames.
|
||||
|
||||
## Tools
|
||||
|
||||
Python, with NumPy and OpenCV (in the dev container: `python3-numpy`, `python3-opencv`, which `setup/dev-container.sh` installs). Off the Frame, `FRAME_JOB_DEVICE_ROOT` can point at a folder with copies of the headset's calibration files.
|
||||
|
||||
- `tools/check_sides.py --ring` (or a recording): are the side cameras named right?
|
||||
- `tools/check_color.py REC`: how the colour module's calibration maps onto its images.
|
||||
- `tools/show_set.py REC`: a recording's frame sets as images.
|
||||
- `tools/watch_gestures.py [--distance]`: pinches, live.
|
||||
- `tools/depth_report.py DEPTH`: the depth measures above.
|
||||
- `tools/convert_models.py`: how `models/ncnn` was made from the OpenCV Zoo ONNX ports of MediaPipe's models (see `models/NOTICE`).
|
||||
|
||||
## Build
|
||||
|
||||
`hands/build.sh` builds in the dev container on the Frame, into `hands/build/`, with `hands/Makefile`. The first build fetches ncnn at a pinned tag and builds it into `hands/build/ncnn`, which takes a few minutes; `NCNN=DIR` points at an ncnn install already built instead. ft-camd is linked statically, because it runs on the host, which has an older glibc than the container.
|
||||
Executable
+11
@@ -0,0 +1,11 @@
|
||||
#!/usr/bin/env bash
|
||||
# Build hand tracking in the dev container on the Frame, into hands/build/: ft-camd and ft-hands,
|
||||
# and with --tools also ft-handreplay and ft-ringplay. The first build fetches ncnn and builds
|
||||
# it (a few minutes); NCNN=DIR, an ncnn install already on the Frame, skips that.
|
||||
# A rebuilt ft-camd has lost its capabilities: hands/run.sh install sets them again.
|
||||
set -euo pipefail
|
||||
root=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
|
||||
targets=all
|
||||
[ "${1:-}" = --tools ] && targets="all tools"
|
||||
"$root/scripts/sync.sh" >/dev/null
|
||||
exec "$root/scripts/frame.sh" -C hands "make -s ${NCNN:+NCNN=$NCNN} $targets && echo built \$(ls build/ft-* | tr '\n' ' ')"
|
||||
@@ -0,0 +1,21 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2026 Curtis English
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
+1127
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,87 @@
|
||||
/*
|
||||
* fhring - the shared-memory frame ring ft-camd writes and trackers read.
|
||||
*
|
||||
* One file, /run/user/UID/frametop-hands/cam-ring (FH_RING_NAME in the user's runtime
|
||||
* folder; the folder is private to the user), holds a header, then for each camera
|
||||
* a few slots, each a slot header followed by the image rows packed tightly
|
||||
* (stride == width for 8-bit mono). Only complete, bright frames are published.
|
||||
*
|
||||
* Writer, for frame n of a camera: slot = n % nslots
|
||||
* slot.seq = 2n+1; write slot fields and pixels; slot.seq = 2n+2; cam.latest = n
|
||||
* Reader:
|
||||
* n = cam.latest; read slot.seq, expect 2n+2; copy; re-read slot.seq; if it
|
||||
* changed the copy is torn, retry with the new latest.
|
||||
*
|
||||
* All multi-byte fields are little-endian; offsets are fixed so Python can read
|
||||
* them with struct (tools/ring.py mirrors this file).
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <assert.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define FH_RING_MAGIC "FHRING01"
|
||||
#define FH_RING_VERSION 1
|
||||
#define FH_RING_MAX_CAMS 8
|
||||
#define FH_RING_SLOTS 4
|
||||
#define FH_RING_NAME "frametop-hands/cam-ring" /* in /run/user/UID */
|
||||
|
||||
enum {
|
||||
FH_FMT_GREY8 = 0,
|
||||
};
|
||||
|
||||
enum {
|
||||
FH_CAM_DARK = 1u << 0, /* the near-black exposures between this node's */
|
||||
/* normal frames (ft-camd --with-dark) */
|
||||
FH_CAM_COLOR = 1u << 1, /* an Arcturus color camera's luma, downscaled */
|
||||
/* (ft-camd --with-color). Not synced with the */
|
||||
/* mono cameras, and capture_ns is on its own */
|
||||
/* clock: line it up with them by dqbuf_ns */
|
||||
};
|
||||
|
||||
typedef struct {
|
||||
char sensor[32]; /* media entity, e.g. "og01a1b 4-0060" */
|
||||
char name[32]; /* calibration name if known, else sensor slug */
|
||||
int32_t node; /* N of /dev/videoN */
|
||||
uint32_t format; /* FH_FMT_* */
|
||||
uint32_t width;
|
||||
uint32_t height;
|
||||
uint32_t stride; /* bytes per row in the ring */
|
||||
uint32_t nslots;
|
||||
uint64_t slot_offset; /* file offset of slot 0 */
|
||||
uint64_t slot_bytes; /* slot header + image, 64-byte aligned */
|
||||
volatile uint64_t latest; /* newest published frame number, 0 = none yet */
|
||||
uint64_t published; /* frames published */
|
||||
uint64_t dropped; /* dark, stale or torn frames not published */
|
||||
uint32_t flags; /* FH_CAM_* */
|
||||
uint8_t reserved[28];
|
||||
} fh_ring_cam_t; /* 160 bytes */
|
||||
|
||||
typedef struct {
|
||||
volatile uint64_t seq; /* 2n+1 while frame n is written, 2n+2 when done */
|
||||
uint64_t frame; /* n */
|
||||
uint64_t capture_ns; /* V4L2 timestamp (camera clock) */
|
||||
uint64_t dqbuf_ns; /* CLOCK_MONOTONIC when XRService dequeued it */
|
||||
uint64_t publish_ns; /* CLOCK_MONOTONIC when the copy finished */
|
||||
uint32_t v4l2_seq; /* V4L2 sequence number */
|
||||
float mean; /* mean luma on a sparse grid */
|
||||
uint8_t reserved[16];
|
||||
} fh_ring_slot_t; /* 64 bytes, image follows */
|
||||
|
||||
typedef struct {
|
||||
char magic[8]; /* FH_RING_MAGIC */
|
||||
uint32_t version;
|
||||
uint32_t header_bytes; /* sizeof(fh_ring_hdr_t) */
|
||||
uint32_t ncams;
|
||||
uint32_t reserved0;
|
||||
uint64_t file_bytes;
|
||||
int64_t writer_pid;
|
||||
volatile uint64_t heartbeat_ns; /* CLOCK_MONOTONIC, refreshed at least every 0.2 s */
|
||||
uint8_t reserved[16];
|
||||
fh_ring_cam_t cams[FH_RING_MAX_CAMS];
|
||||
} fh_ring_hdr_t;
|
||||
|
||||
static_assert(sizeof(fh_ring_cam_t) == 160, "fh_ring_cam_t layout");
|
||||
static_assert(sizeof(fh_ring_slot_t) == 64, "fh_ring_slot_t layout");
|
||||
static_assert(sizeof(fh_ring_hdr_t) == 64 + 160 * FH_RING_MAX_CAMS, "fh_ring_hdr_t layout");
|
||||
+427
@@ -0,0 +1,427 @@
|
||||
/*
|
||||
* tp - read kernel tracepoints system-wide through perf_event_open.
|
||||
*/
|
||||
|
||||
#define _GNU_SOURCE
|
||||
|
||||
#include "tp.h"
|
||||
|
||||
#include <errno.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/epoll.h>
|
||||
#include <sys/ioctl.h>
|
||||
#include <sys/mman.h>
|
||||
#include <sys/syscall.h>
|
||||
#include <time.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <linux/perf_event.h>
|
||||
|
||||
#ifndef TRACEFS
|
||||
#define TRACEFS "/sys/kernel/tracing/events"
|
||||
#endif
|
||||
#define RING_DATA_PAGES 16
|
||||
|
||||
static void set_err(char *err, size_t n, const char *fmt, ...)
|
||||
{
|
||||
va_list ap;
|
||||
|
||||
va_start(ap, fmt);
|
||||
vsnprintf(err, n, fmt, ap);
|
||||
va_end(ap);
|
||||
}
|
||||
|
||||
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn)
|
||||
{
|
||||
memset(ev, 0, sizeof(*ev));
|
||||
snprintf(ev->system, sizeof(ev->system), "%s", system);
|
||||
snprintf(ev->name, sizeof(ev->name), "%s", name);
|
||||
ev->id = -1;
|
||||
|
||||
char path[256];
|
||||
snprintf(path, sizeof(path), TRACEFS "/%s/%s/format", system, name);
|
||||
|
||||
FILE *f = fopen(path, "r");
|
||||
|
||||
if (!f) {
|
||||
set_err(err, errn, "%s: %s", path, strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
char line[512];
|
||||
|
||||
while (fgets(line, sizeof(line), f)) {
|
||||
|
||||
int id;
|
||||
|
||||
if (sscanf(line, "ID: %d", &id) == 1) {
|
||||
ev->id = id;
|
||||
continue;
|
||||
}
|
||||
|
||||
char *fp = line;
|
||||
|
||||
while (*fp == ' ' || *fp == '\t')
|
||||
fp++;
|
||||
|
||||
if (strncmp(fp, "field:", 6) || ev->nfields >= TP_MAX_FIELDS)
|
||||
continue;
|
||||
|
||||
char *semi = strchr(fp, ';');
|
||||
|
||||
if (!semi)
|
||||
continue;
|
||||
|
||||
/* the field name is the last identifier in the declaration */
|
||||
char decl[256];
|
||||
size_t dl = (size_t)(semi - (fp + 6));
|
||||
|
||||
if (dl >= sizeof(decl))
|
||||
dl = sizeof(decl) - 1;
|
||||
|
||||
memcpy(decl, fp + 6, dl);
|
||||
decl[dl] = 0;
|
||||
|
||||
char *br = strchr(decl, '[');
|
||||
|
||||
if (br)
|
||||
*br = 0;
|
||||
|
||||
char *end = decl + strlen(decl);
|
||||
|
||||
while (end > decl && (end[-1] == ' ' || end[-1] == '\t'))
|
||||
*--end = 0;
|
||||
|
||||
char *start = end;
|
||||
|
||||
while (start > decl && start[-1] != ' ' && start[-1] != '\t' && start[-1] != '*')
|
||||
start--;
|
||||
|
||||
tp_field_t *fd = &ev->fields[ev->nfields];
|
||||
const char *o = strstr(semi, "offset:");
|
||||
const char *s = strstr(semi, "size:");
|
||||
const char *g = strstr(semi, "signed:");
|
||||
|
||||
if (!o || !s)
|
||||
continue;
|
||||
|
||||
snprintf(fd->name, sizeof(fd->name), "%s", start);
|
||||
fd->offset = atoi(o + 7);
|
||||
fd->size = atoi(s + 5);
|
||||
fd->is_signed = g ? atoi(g + 7) != 0 : false;
|
||||
ev->nfields++;
|
||||
}
|
||||
|
||||
fclose(f);
|
||||
|
||||
if (ev->id < 0) {
|
||||
set_err(err, errn, "%s: no ID line", path);
|
||||
return false;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
int tp_field(const tp_event_t *ev, const char *name)
|
||||
{
|
||||
for (int i = 0; i < ev->nfields; i++)
|
||||
if (!strcmp(ev->fields[i].name, name))
|
||||
return i;
|
||||
|
||||
return -1;
|
||||
}
|
||||
|
||||
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen)
|
||||
{
|
||||
if (field < 0 || field >= ev->nfields)
|
||||
return 0;
|
||||
|
||||
const tp_field_t *f = &ev->fields[field];
|
||||
|
||||
if (f->offset < 0 || (uint32_t)(f->offset + f->size) > rawlen)
|
||||
return 0;
|
||||
|
||||
const uint8_t *p = raw + f->offset;
|
||||
|
||||
switch (f->size) {
|
||||
case 1: { uint8_t v; memcpy(&v, p, 1); return f->is_signed ? (int64_t)(int8_t)v : (int64_t)v; }
|
||||
case 2: { uint16_t v; memcpy(&v, p, 2); return f->is_signed ? (int64_t)(int16_t)v : (int64_t)v; }
|
||||
case 4: { uint32_t v; memcpy(&v, p, 4); return f->is_signed ? (int64_t)(int32_t)v : (int64_t)v; }
|
||||
case 8: { uint64_t v; memcpy(&v, p, 8); return (int64_t)v; }
|
||||
default: return 0;
|
||||
}
|
||||
}
|
||||
|
||||
static int online_cpus(int *cpus, int max)
|
||||
{
|
||||
FILE *f = fopen("/sys/devices/system/cpu/online", "r");
|
||||
int n = 0;
|
||||
|
||||
if (!f)
|
||||
return 0;
|
||||
|
||||
char buf[256] = {0};
|
||||
|
||||
if (!fgets(buf, sizeof(buf), f))
|
||||
buf[0] = 0;
|
||||
|
||||
fclose(f);
|
||||
|
||||
for (char *tok = strtok(buf, ",\n"); tok && n < max; tok = strtok(NULL, ",\n")) {
|
||||
|
||||
int a, b;
|
||||
|
||||
if (sscanf(tok, "%d-%d", &a, &b) == 2) {
|
||||
for (int c = a; c <= b && n < max; c++)
|
||||
cpus[n++] = c;
|
||||
} else if (sscanf(tok, "%d", &a) == 1) {
|
||||
cpus[n++] = a;
|
||||
}
|
||||
}
|
||||
|
||||
return n;
|
||||
}
|
||||
|
||||
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn)
|
||||
{
|
||||
memset(tp, 0, sizeof(*tp));
|
||||
tp->epfd = -1;
|
||||
|
||||
if (nevents <= 0 || nevents > TP_MAX_EVENTS) {
|
||||
set_err(err, errn, "bad event count %d", nevents);
|
||||
return false;
|
||||
}
|
||||
|
||||
for (int i = 0; i < nevents; i++)
|
||||
tp->events[i] = events[i];
|
||||
|
||||
tp->nevents = nevents;
|
||||
|
||||
int cpus[TP_MAX_CPUS];
|
||||
tp->ncpu = online_cpus(cpus, TP_MAX_CPUS);
|
||||
|
||||
if (tp->ncpu <= 0) {
|
||||
set_err(err, errn, "no online CPUs found");
|
||||
return false;
|
||||
}
|
||||
|
||||
long page = sysconf(_SC_PAGESIZE);
|
||||
tp->map_len = (size_t)page * (1 + RING_DATA_PAGES);
|
||||
|
||||
tp->epfd = epoll_create1(EPOLL_CLOEXEC);
|
||||
|
||||
if (tp->epfd < 0) {
|
||||
set_err(err, errn, "epoll_create1: %s", strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
for (int c = 0; c < tp->ncpu; c++) {
|
||||
|
||||
tp->ring_fd[c] = -1;
|
||||
|
||||
for (int e = 0; e < nevents; e++) {
|
||||
|
||||
struct perf_event_attr a;
|
||||
memset(&a, 0, sizeof(a));
|
||||
|
||||
a.size = sizeof(a);
|
||||
a.type = PERF_TYPE_TRACEPOINT;
|
||||
a.config = (uint64_t)events[e]->id;
|
||||
a.sample_period = 1;
|
||||
a.sample_type = PERF_SAMPLE_TID | PERF_SAMPLE_TIME | PERF_SAMPLE_CPU | PERF_SAMPLE_RAW;
|
||||
a.wakeup_events = 1;
|
||||
a.use_clockid = 1;
|
||||
a.clockid = CLOCK_MONOTONIC;
|
||||
a.disabled = 1;
|
||||
|
||||
int fd = (int)syscall(SYS_perf_event_open, &a, -1, cpus[c], -1, PERF_FLAG_FD_CLOEXEC);
|
||||
|
||||
if (fd < 0) {
|
||||
set_err(err, errn, "perf_event_open(%s:%s, cpu %d): %s",
|
||||
events[e]->system, events[e]->name, cpus[c], strerror(errno));
|
||||
tp_close(tp);
|
||||
return false;
|
||||
}
|
||||
|
||||
tp->fds[tp->nfds++] = fd;
|
||||
|
||||
if (tp->ring_fd[c] < 0) {
|
||||
|
||||
void *m = mmap(NULL, tp->map_len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
|
||||
if (m == MAP_FAILED) {
|
||||
set_err(err, errn, "mmap perf ring (cpu %d): %s", cpus[c], strerror(errno));
|
||||
tp_close(tp);
|
||||
return false;
|
||||
}
|
||||
|
||||
tp->ring[c] = m;
|
||||
tp->ring_fd[c] = fd;
|
||||
|
||||
struct epoll_event ee = { .events = EPOLLIN, .data.u32 = (uint32_t)c };
|
||||
epoll_ctl(tp->epfd, EPOLL_CTL_ADD, fd, &ee);
|
||||
|
||||
} else if (ioctl(fd, PERF_EVENT_IOC_SET_OUTPUT, tp->ring_fd[c]) < 0) {
|
||||
set_err(err, errn, "PERF_EVENT_IOC_SET_OUTPUT: %s", strerror(errno));
|
||||
tp_close(tp);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
for (int i = 0; i < tp->nfds; i++)
|
||||
ioctl(tp->fds[i], PERF_EVENT_IOC_ENABLE, 0);
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
static void ring_copy(uint8_t *dst, const uint8_t *base, uint64_t size, uint64_t pos, size_t len)
|
||||
{
|
||||
uint64_t off = pos % size;
|
||||
size_t first = (size_t)(size - off);
|
||||
|
||||
if (first >= len) {
|
||||
memcpy(dst, base + off, len);
|
||||
} else {
|
||||
memcpy(dst, base + off, first);
|
||||
memcpy(dst + first, base, len - first);
|
||||
}
|
||||
}
|
||||
|
||||
static int cmp_sample(const void *a, const void *b)
|
||||
{
|
||||
const tp_sample_t *x = a, *y = b;
|
||||
|
||||
return (x->time > y->time) - (x->time < y->time);
|
||||
}
|
||||
|
||||
static void dispatch(tp_t *tp, tp_cb cb, void *ctx)
|
||||
{
|
||||
qsort(tp->pend, tp->npend, sizeof(tp->pend[0]), cmp_sample);
|
||||
|
||||
for (int i = 0; i < tp->npend; i++)
|
||||
cb(ctx, &tp->pend[i]);
|
||||
|
||||
tp->npend = 0;
|
||||
}
|
||||
|
||||
static int drain_ring(tp_t *tp, int c, tp_cb cb, void *ctx)
|
||||
{
|
||||
struct perf_event_mmap_page *pg = tp->ring[c];
|
||||
long page = sysconf(_SC_PAGESIZE);
|
||||
uint64_t off = pg->data_offset ? pg->data_offset : (uint64_t)page;
|
||||
uint64_t size = pg->data_size ? pg->data_size : (uint64_t)page * RING_DATA_PAGES;
|
||||
const uint8_t *base = (const uint8_t *)pg + off;
|
||||
|
||||
uint64_t head = __atomic_load_n(&pg->data_head, __ATOMIC_ACQUIRE);
|
||||
uint64_t tail = pg->data_tail;
|
||||
int n = 0;
|
||||
|
||||
while (tail < head) {
|
||||
|
||||
struct perf_event_header hdr;
|
||||
ring_copy((uint8_t *)&hdr, base, size, tail, sizeof(hdr));
|
||||
|
||||
if (hdr.size < sizeof(hdr))
|
||||
break;
|
||||
|
||||
ring_copy(tp->scratch, base, size, tail, hdr.size);
|
||||
|
||||
const uint8_t *p = tp->scratch + sizeof(hdr);
|
||||
const uint8_t *end = tp->scratch + hdr.size;
|
||||
|
||||
if (hdr.type == PERF_RECORD_LOST && end - p >= 16) {
|
||||
|
||||
uint64_t lost;
|
||||
memcpy(&lost, p + 8, 8);
|
||||
tp->lost += lost;
|
||||
|
||||
} else if (hdr.type == PERF_RECORD_SAMPLE && end - p >= 28) {
|
||||
|
||||
tp_sample_t s;
|
||||
uint32_t v32[2];
|
||||
|
||||
memcpy(v32, p, 8); p += 8;
|
||||
s.pid = v32[0];
|
||||
s.tid = v32[1];
|
||||
memcpy(&s.time, p, 8); p += 8;
|
||||
memcpy(v32, p, 8); p += 8;
|
||||
s.cpu = v32[0];
|
||||
memcpy(&s.rawlen, p, 4); p += 4;
|
||||
s.raw = p;
|
||||
|
||||
if (s.rawlen >= 2 && p + s.rawlen <= end) {
|
||||
|
||||
uint16_t type;
|
||||
memcpy(&type, s.raw, 2);
|
||||
s.ev = NULL;
|
||||
|
||||
for (int e = 0; e < tp->nevents; e++)
|
||||
if (tp->events[e]->id == type)
|
||||
s.ev = tp->events[e];
|
||||
|
||||
if (s.ev && s.rawlen <= TP_MAX_RAW) {
|
||||
|
||||
if (tp->npend == TP_MAX_PENDING)
|
||||
dispatch(tp, cb, ctx);
|
||||
|
||||
memcpy(tp->pend_raw[tp->npend], s.raw, s.rawlen);
|
||||
s.raw = tp->pend_raw[tp->npend];
|
||||
tp->pend[tp->npend++] = s;
|
||||
n++;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
tail += hdr.size;
|
||||
}
|
||||
|
||||
__atomic_store_n(&pg->data_tail, tail, __ATOMIC_RELEASE);
|
||||
|
||||
return n;
|
||||
}
|
||||
|
||||
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx)
|
||||
{
|
||||
struct epoll_event ev[TP_MAX_CPUS];
|
||||
|
||||
if (epoll_wait(tp->epfd, ev, TP_MAX_CPUS, timeout_ms) < 0 && errno != EINTR)
|
||||
return -1;
|
||||
|
||||
/*
|
||||
* Drain every ring, not just the ones that woke us: samples from several
|
||||
* CPUs need to be handled together to keep per-camera order sane.
|
||||
*/
|
||||
int n = 0;
|
||||
|
||||
for (int c = 0; c < tp->ncpu; c++)
|
||||
if (tp->ring[c])
|
||||
n += drain_ring(tp, c, cb, ctx);
|
||||
|
||||
dispatch(tp, cb, ctx);
|
||||
|
||||
return n;
|
||||
}
|
||||
|
||||
void tp_close(tp_t *tp)
|
||||
{
|
||||
for (int i = 0; i < tp->nfds; i++) {
|
||||
ioctl(tp->fds[i], PERF_EVENT_IOC_DISABLE, 0);
|
||||
}
|
||||
|
||||
for (int c = 0; c < tp->ncpu; c++)
|
||||
if (tp->ring[c])
|
||||
munmap(tp->ring[c], tp->map_len);
|
||||
|
||||
for (int i = 0; i < tp->nfds; i++)
|
||||
close(tp->fds[i]);
|
||||
|
||||
if (tp->epfd >= 0)
|
||||
close(tp->epfd);
|
||||
|
||||
tp->nfds = 0;
|
||||
tp->epfd = -1;
|
||||
}
|
||||
@@ -0,0 +1,76 @@
|
||||
/*
|
||||
* tp - read kernel tracepoints system-wide through perf_event_open.
|
||||
*
|
||||
* One perf ring per CPU; every event on that CPU writes into it. Field
|
||||
* offsets come from the tracefs format files, so kernel layout changes don't
|
||||
* silently break parsing. Needs root (or CAP_PERFMON plus tracefs access).
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define TP_MAX_FIELDS 40
|
||||
#define TP_MAX_EVENTS 8
|
||||
#define TP_MAX_CPUS 64
|
||||
#define TP_MAX_PENDING 2048
|
||||
#define TP_MAX_RAW 256
|
||||
|
||||
typedef struct {
|
||||
char name[48];
|
||||
int offset;
|
||||
int size;
|
||||
bool is_signed;
|
||||
} tp_field_t;
|
||||
|
||||
typedef struct {
|
||||
char system[32];
|
||||
char name[48];
|
||||
int id;
|
||||
tp_field_t fields[TP_MAX_FIELDS];
|
||||
int nfields;
|
||||
} tp_event_t;
|
||||
|
||||
typedef struct {
|
||||
const tp_event_t *ev;
|
||||
const uint8_t *raw;
|
||||
uint32_t rawlen;
|
||||
uint64_t time; /* CLOCK_MONOTONIC ns */
|
||||
uint32_t cpu;
|
||||
uint32_t pid;
|
||||
uint32_t tid;
|
||||
} tp_sample_t;
|
||||
|
||||
typedef void (*tp_cb)(void *ctx, const tp_sample_t *s);
|
||||
|
||||
typedef struct {
|
||||
int ncpu;
|
||||
int ring_fd[TP_MAX_CPUS];
|
||||
void *ring[TP_MAX_CPUS];
|
||||
size_t map_len;
|
||||
int fds[TP_MAX_CPUS * TP_MAX_EVENTS];
|
||||
int nfds;
|
||||
int epfd;
|
||||
tp_event_t *events[TP_MAX_EVENTS];
|
||||
int nevents;
|
||||
uint64_t lost;
|
||||
uint8_t scratch[65536];
|
||||
/* samples drained from all rings, sorted by time before dispatch */
|
||||
tp_sample_t pend[TP_MAX_PENDING];
|
||||
uint8_t pend_raw[TP_MAX_PENDING][TP_MAX_RAW];
|
||||
int npend;
|
||||
} tp_t;
|
||||
|
||||
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn);
|
||||
int tp_field(const tp_event_t *ev, const char *name);
|
||||
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen);
|
||||
|
||||
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn);
|
||||
/*
|
||||
* Wait up to timeout_ms, then hand every pending sample to cb in time order,
|
||||
* across all CPUs. Returns samples read, -1 on error.
|
||||
*/
|
||||
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx);
|
||||
void tp_close(tp_t *tp);
|
||||
@@ -0,0 +1,783 @@
|
||||
/*
|
||||
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
|
||||
*
|
||||
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
|
||||
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
|
||||
*
|
||||
* Everything is discovered rather than hardcoded:
|
||||
* - XRService is found by scanning /proc for its cmdline.
|
||||
* - The V4L2 nodes and sensor subdevs it holds open come from /proc/<pid>/fd.
|
||||
* - Each node's geometry comes from VIDIOC_G_FMT on our own handle.
|
||||
* - Each node is traced back to its sensor through MEDIA_IOC_G_TOPOLOGY.
|
||||
* - Buffers are split into queues by allocation order: XRService opens a
|
||||
* sensor subdev, then allocates that camera's buffers.
|
||||
*/
|
||||
|
||||
#define _GNU_SOURCE
|
||||
|
||||
#include "xrcams.h"
|
||||
|
||||
#include <dirent.h>
|
||||
#include <errno.h>
|
||||
#include <fcntl.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/ioctl.h>
|
||||
#include <sys/stat.h>
|
||||
#include <sys/sysmacros.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <linux/media.h>
|
||||
|
||||
#ifndef MEDIA_ENT_F_CAM_SENSOR
|
||||
#define MEDIA_ENT_F_CAM_SENSOR 0x00020001
|
||||
#endif
|
||||
|
||||
#define MAX_FDENTS 4096
|
||||
#define MAX_TOPOS 8
|
||||
|
||||
enum fdkind { FD_DMABUF, FD_SUBDEV_SENSOR, FD_VIDEO };
|
||||
|
||||
typedef struct {
|
||||
int xfd;
|
||||
enum fdkind kind;
|
||||
size_t size;
|
||||
unsigned long ino;
|
||||
char sensor[XR_SENSOR_LEN];
|
||||
char path[64];
|
||||
} fdent_t;
|
||||
|
||||
typedef struct {
|
||||
struct media_v2_entity *ents;
|
||||
struct media_v2_interface *intfs;
|
||||
struct media_v2_pad *pads;
|
||||
struct media_v2_link *links;
|
||||
__u32 nents, nintfs, npads, nlinks;
|
||||
} topo_t;
|
||||
|
||||
static fdent_t fdents[MAX_FDENTS];
|
||||
static int nfdents;
|
||||
static topo_t topos[MAX_TOPOS];
|
||||
static int ntopos;
|
||||
|
||||
static void set_err(char *err, size_t n, const char *fmt, ...)
|
||||
{
|
||||
va_list ap;
|
||||
|
||||
va_start(ap, fmt);
|
||||
vsnprintf(err, n, fmt, ap);
|
||||
va_end(ap);
|
||||
}
|
||||
|
||||
void xr_slugify(const char *in, char *out, size_t n)
|
||||
{
|
||||
size_t i = 0;
|
||||
|
||||
for (; in[i] && i + 1 < n; i++)
|
||||
out[i] = (in[i] == ' ' || in[i] == '/') ? '_' : in[i];
|
||||
|
||||
out[i] = 0;
|
||||
}
|
||||
|
||||
/* --------------------------------------------------- media graph handling */
|
||||
|
||||
static void topo_free_all(void)
|
||||
{
|
||||
for (int i = 0; i < ntopos; i++) {
|
||||
free(topos[i].ents);
|
||||
free(topos[i].intfs);
|
||||
free(topos[i].pads);
|
||||
free(topos[i].links);
|
||||
}
|
||||
|
||||
ntopos = 0;
|
||||
}
|
||||
|
||||
static void topo_load_all(void)
|
||||
{
|
||||
for (int mi = 0; mi < MAX_TOPOS; mi++) {
|
||||
|
||||
char mpath[32];
|
||||
snprintf(mpath, sizeof(mpath), "/dev/media%d", mi);
|
||||
|
||||
int mfd = open(mpath, O_RDWR | O_CLOEXEC);
|
||||
|
||||
if (mfd < 0)
|
||||
continue;
|
||||
|
||||
struct media_v2_topology t;
|
||||
memset(&t, 0, sizeof(t));
|
||||
|
||||
if (ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) < 0) {
|
||||
close(mfd);
|
||||
continue;
|
||||
}
|
||||
|
||||
topo_t *o = &topos[ntopos];
|
||||
memset(o, 0, sizeof(*o));
|
||||
|
||||
o->nents = t.num_entities;
|
||||
o->nintfs = t.num_interfaces;
|
||||
o->npads = t.num_pads;
|
||||
o->nlinks = t.num_links;
|
||||
|
||||
o->ents = calloc(o->nents ? o->nents : 1, sizeof(*o->ents));
|
||||
o->intfs = calloc(o->nintfs ? o->nintfs : 1, sizeof(*o->intfs));
|
||||
o->pads = calloc(o->npads ? o->npads : 1, sizeof(*o->pads));
|
||||
o->links = calloc(o->nlinks ? o->nlinks : 1, sizeof(*o->links));
|
||||
|
||||
t.ptr_entities = (__u64)(uintptr_t)o->ents;
|
||||
t.ptr_interfaces = (__u64)(uintptr_t)o->intfs;
|
||||
t.ptr_pads = (__u64)(uintptr_t)o->pads;
|
||||
t.ptr_links = (__u64)(uintptr_t)o->links;
|
||||
|
||||
bool ok = o->ents && o->intfs && o->pads && o->links &&
|
||||
ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) == 0;
|
||||
close(mfd);
|
||||
|
||||
if (!ok) {
|
||||
free(o->ents); free(o->intfs); free(o->pads); free(o->links);
|
||||
continue;
|
||||
}
|
||||
|
||||
ntopos++;
|
||||
}
|
||||
}
|
||||
|
||||
static struct media_v2_entity *topo_entity(topo_t *t, __u32 id)
|
||||
{
|
||||
for (__u32 i = 0; i < t->nents; i++)
|
||||
if (t->ents[i].id == id)
|
||||
return &t->ents[i];
|
||||
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static struct media_v2_pad *topo_pad(topo_t *t, __u32 id)
|
||||
{
|
||||
for (__u32 i = 0; i < t->npads; i++)
|
||||
if (t->pads[i].id == id)
|
||||
return &t->pads[i];
|
||||
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static __u32 topo_entity_for_devnode(topo_t *t, dev_t rdev)
|
||||
{
|
||||
__u32 intf_id = 0;
|
||||
|
||||
for (__u32 i = 0; i < t->nintfs; i++)
|
||||
if (t->intfs[i].devnode.major == major(rdev) &&
|
||||
t->intfs[i].devnode.minor == minor(rdev)) {
|
||||
intf_id = t->intfs[i].id;
|
||||
break;
|
||||
}
|
||||
|
||||
if (!intf_id)
|
||||
return 0;
|
||||
|
||||
for (__u32 i = 0; i < t->nlinks; i++)
|
||||
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) == MEDIA_LNK_FL_INTERFACE_LINK &&
|
||||
t->links[i].source_id == intf_id)
|
||||
return t->links[i].sink_id;
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
/*
|
||||
* Walk upstream across enabled data links until a sensor is reached. A CSIPHY
|
||||
* carries two sensors on separate (sink, source) pad pairs, so re-enter on the
|
||||
* sink pad paired with the source pad we left through.
|
||||
*/
|
||||
static bool topo_walk_to_sensor(topo_t *t, __u32 ent_id, char *out, size_t outn)
|
||||
{
|
||||
int exit_pad_index = -1;
|
||||
|
||||
for (int hop = 0; hop < 32 && ent_id; hop++) {
|
||||
|
||||
struct media_v2_entity *e = topo_entity(t, ent_id);
|
||||
|
||||
if (!e)
|
||||
return false;
|
||||
|
||||
if (e->function == MEDIA_ENT_F_CAM_SENSOR) {
|
||||
snprintf(out, outn, "%s", e->name);
|
||||
return true;
|
||||
}
|
||||
|
||||
__u32 first_sink = 0, paired = 0;
|
||||
int nsinks = 0;
|
||||
|
||||
for (__u32 p = 0; p < t->npads; p++) {
|
||||
|
||||
if (t->pads[p].entity_id != ent_id || !(t->pads[p].flags & MEDIA_PAD_FL_SINK))
|
||||
continue;
|
||||
|
||||
nsinks++;
|
||||
|
||||
if (!first_sink)
|
||||
first_sink = t->pads[p].id;
|
||||
|
||||
if (exit_pad_index >= 1 && (int)t->pads[p].index == exit_pad_index - 1)
|
||||
paired = t->pads[p].id;
|
||||
}
|
||||
|
||||
__u32 sink_pad = (nsinks == 1) ? first_sink : (paired ? paired : first_sink);
|
||||
|
||||
if (!sink_pad)
|
||||
return false;
|
||||
|
||||
__u32 src_pad = 0;
|
||||
|
||||
for (__u32 i = 0; i < t->nlinks; i++) {
|
||||
|
||||
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) != MEDIA_LNK_FL_DATA_LINK)
|
||||
continue;
|
||||
|
||||
if (!(t->links[i].flags & MEDIA_LNK_FL_ENABLED))
|
||||
continue;
|
||||
|
||||
if (t->links[i].sink_id == sink_pad) {
|
||||
src_pad = t->links[i].source_id;
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
struct media_v2_pad *sp = src_pad ? topo_pad(t, src_pad) : NULL;
|
||||
|
||||
if (!sp)
|
||||
return false;
|
||||
|
||||
ent_id = sp->entity_id;
|
||||
exit_pad_index = (int)sp->index;
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
static bool sensor_for_video(dev_t rdev, char *out, size_t outn)
|
||||
{
|
||||
for (int i = 0; i < ntopos; i++) {
|
||||
|
||||
__u32 ent = topo_entity_for_devnode(&topos[i], rdev);
|
||||
|
||||
if (ent && topo_walk_to_sensor(&topos[i], ent, out, outn))
|
||||
return true;
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
static bool sensor_for_subdev(dev_t rdev, char *out, size_t outn)
|
||||
{
|
||||
for (int i = 0; i < ntopos; i++) {
|
||||
|
||||
__u32 id = topo_entity_for_devnode(&topos[i], rdev);
|
||||
struct media_v2_entity *e = id ? topo_entity(&topos[i], id) : NULL;
|
||||
|
||||
if (e && e->function == MEDIA_ENT_F_CAM_SENSOR) {
|
||||
snprintf(out, outn, "%s", e->name);
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
static const char *role_for_sensor(const char *sensor)
|
||||
{
|
||||
if (strstr(sensor, "og01a1b"))
|
||||
return "tracking"; /* 1056x1024 side fisheye */
|
||||
|
||||
if (strstr(sensor, "og0ve10"))
|
||||
return "tracking"; /* 640x480 upper */
|
||||
|
||||
if (strstr(sensor, "imx616"))
|
||||
return "passthrough"; /* 2464x2464 Arcturus color */
|
||||
|
||||
return "unknown";
|
||||
}
|
||||
|
||||
/* ------------------------------------------------- XRService / proc scan */
|
||||
|
||||
static pid_t find_process(const char *needle)
|
||||
{
|
||||
DIR *d = opendir("/proc");
|
||||
|
||||
if (!d)
|
||||
return 0;
|
||||
|
||||
struct dirent *e;
|
||||
pid_t found = 0;
|
||||
|
||||
while ((e = readdir(d))) {
|
||||
|
||||
if (e->d_name[0] < '0' || e->d_name[0] > '9')
|
||||
continue;
|
||||
|
||||
char path[288];
|
||||
snprintf(path, sizeof(path), "/proc/%s/cmdline", e->d_name);
|
||||
|
||||
FILE *f = fopen(path, "rb");
|
||||
|
||||
if (!f)
|
||||
continue;
|
||||
|
||||
char buf[512] = {0};
|
||||
size_t got = fread(buf, 1, sizeof(buf) - 1, f);
|
||||
fclose(f);
|
||||
|
||||
if (got == 0)
|
||||
continue;
|
||||
|
||||
const char *base = strrchr(buf, '/');
|
||||
base = base ? base + 1 : buf;
|
||||
|
||||
if (strstr(base, needle)) {
|
||||
found = (pid_t)atoi(e->d_name);
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
closedir(d);
|
||||
|
||||
return found;
|
||||
}
|
||||
|
||||
static bool read_dmabuf_size(pid_t pid, int fd, size_t *size, unsigned long *ino)
|
||||
{
|
||||
char path[64];
|
||||
snprintf(path, sizeof(path), "/proc/%d/fdinfo/%d", pid, fd);
|
||||
|
||||
FILE *f = fopen(path, "r");
|
||||
|
||||
if (!f)
|
||||
return false;
|
||||
|
||||
bool have = false;
|
||||
char line[256];
|
||||
|
||||
*ino = 0;
|
||||
|
||||
while (fgets(line, sizeof(line), f)) {
|
||||
|
||||
unsigned long long v;
|
||||
|
||||
if (sscanf(line, "size: %llu", &v) == 1) {
|
||||
*size = (size_t)v;
|
||||
have = true;
|
||||
} else if (sscanf(line, "ino: %llu", &v) == 1) {
|
||||
*ino = (unsigned long)v;
|
||||
}
|
||||
}
|
||||
|
||||
fclose(f);
|
||||
|
||||
return have;
|
||||
}
|
||||
|
||||
static int cmp_int(const void *a, const void *b)
|
||||
{
|
||||
return *(const int *)a - *(const int *)b;
|
||||
}
|
||||
|
||||
static bool scan_xr_fds(pid_t pid, char *err, size_t errn)
|
||||
{
|
||||
char dirpath[64];
|
||||
snprintf(dirpath, sizeof(dirpath), "/proc/%d/fd", pid);
|
||||
|
||||
DIR *d = opendir(dirpath);
|
||||
|
||||
if (!d) {
|
||||
set_err(err, errn, "opendir(%s): %s (are you root?)", dirpath, strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
static int fds[8192];
|
||||
int nfds = 0;
|
||||
struct dirent *e;
|
||||
|
||||
while ((e = readdir(d)) && nfds < (int)(sizeof(fds) / sizeof(fds[0])))
|
||||
if (e->d_name[0] >= '0' && e->d_name[0] <= '9')
|
||||
fds[nfds++] = atoi(e->d_name);
|
||||
|
||||
closedir(d);
|
||||
|
||||
qsort(fds, nfds, sizeof(int), cmp_int);
|
||||
|
||||
nfdents = 0;
|
||||
|
||||
for (int i = 0; i < nfds && nfdents < MAX_FDENTS; i++) {
|
||||
|
||||
char link[64], target[256];
|
||||
snprintf(link, sizeof(link), "/proc/%d/fd/%d", pid, fds[i]);
|
||||
|
||||
ssize_t n = readlink(link, target, sizeof(target) - 1);
|
||||
|
||||
if (n < 0)
|
||||
continue;
|
||||
|
||||
target[n] = 0;
|
||||
|
||||
fdent_t ent;
|
||||
memset(&ent, 0, sizeof(ent));
|
||||
ent.xfd = fds[i];
|
||||
|
||||
if (strstr(target, "dmabuf")) {
|
||||
|
||||
if (!read_dmabuf_size(pid, fds[i], &ent.size, &ent.ino))
|
||||
continue;
|
||||
|
||||
ent.kind = FD_DMABUF;
|
||||
|
||||
} else if (strncmp(target, "/dev/video", 10) == 0) {
|
||||
|
||||
ent.kind = FD_VIDEO;
|
||||
snprintf(ent.path, sizeof(ent.path), "%.63s", target);
|
||||
|
||||
} else if (strncmp(target, "/dev/v4l-subdev", 15) == 0) {
|
||||
|
||||
struct stat st;
|
||||
|
||||
if (stat(target, &st) < 0 || !sensor_for_subdev(st.st_rdev, ent.sensor, sizeof(ent.sensor)))
|
||||
continue;
|
||||
|
||||
ent.kind = FD_SUBDEV_SENSOR;
|
||||
|
||||
} else {
|
||||
continue;
|
||||
}
|
||||
|
||||
fdents[nfdents++] = ent;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------ camera discovery */
|
||||
|
||||
static void probe_cameras(xr_state_t *st)
|
||||
{
|
||||
int seen[64];
|
||||
int nseen = 0;
|
||||
|
||||
for (int i = 0; i < nfdents; i++) {
|
||||
|
||||
if (fdents[i].kind != FD_VIDEO)
|
||||
continue;
|
||||
|
||||
const char *path = fdents[i].path;
|
||||
int node = atoi(path + 10);
|
||||
bool dup = false;
|
||||
|
||||
for (int k = 0; k < nseen; k++)
|
||||
if (seen[k] == node)
|
||||
dup = true;
|
||||
|
||||
if (dup || st->ncameras >= XR_MAX_CAMERAS || nseen >= 64)
|
||||
continue;
|
||||
|
||||
seen[nseen++] = node;
|
||||
|
||||
int fd = open(path, O_RDWR | O_CLOEXEC);
|
||||
|
||||
if (fd < 0)
|
||||
continue;
|
||||
|
||||
struct v4l2_format fmt;
|
||||
memset(&fmt, 0, sizeof(fmt));
|
||||
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE;
|
||||
|
||||
xr_camera_t *c = &st->cameras[st->ncameras];
|
||||
memset(c, 0, sizeof(*c));
|
||||
|
||||
if (ioctl(fd, VIDIOC_G_FMT, &fmt) == 0) {
|
||||
|
||||
c->width = fmt.fmt.pix_mp.width;
|
||||
c->height = fmt.fmt.pix_mp.height;
|
||||
c->pixfmt = fmt.fmt.pix_mp.pixelformat;
|
||||
c->nplanes = fmt.fmt.pix_mp.num_planes;
|
||||
c->bytesperline = fmt.fmt.pix_mp.plane_fmt[0].bytesperline;
|
||||
|
||||
for (unsigned p = 0; p < c->nplanes && p < VIDEO_MAX_PLANES; p++)
|
||||
c->planesize[p] = fmt.fmt.pix_mp.plane_fmt[p].sizeimage;
|
||||
|
||||
} else {
|
||||
|
||||
memset(&fmt, 0, sizeof(fmt));
|
||||
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
|
||||
|
||||
if (ioctl(fd, VIDIOC_G_FMT, &fmt) < 0) {
|
||||
close(fd);
|
||||
continue;
|
||||
}
|
||||
|
||||
c->width = fmt.fmt.pix.width;
|
||||
c->height = fmt.fmt.pix.height;
|
||||
c->pixfmt = fmt.fmt.pix.pixelformat;
|
||||
c->nplanes = 1;
|
||||
c->bytesperline = fmt.fmt.pix.bytesperline;
|
||||
c->planesize[0] = fmt.fmt.pix.sizeimage;
|
||||
}
|
||||
|
||||
struct stat sb;
|
||||
|
||||
if (fstat(fd, &sb) == 0) {
|
||||
c->minor = minor(sb.st_rdev);
|
||||
sensor_for_video(sb.st_rdev, c->sensor, sizeof(c->sensor));
|
||||
}
|
||||
|
||||
close(fd);
|
||||
|
||||
if (!c->sensor[0])
|
||||
snprintf(c->sensor, sizeof(c->sensor), "unknown");
|
||||
|
||||
c->node = node;
|
||||
snprintf(c->path, sizeof(c->path), "%s", path);
|
||||
c->role = role_for_sensor(c->sensor);
|
||||
|
||||
st->ncameras++;
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
* qcom-camss can report bytesperline as the visible width while the VFE
|
||||
* writes a larger aligned pitch. sizeimage is right, so derive the pitch.
|
||||
*/
|
||||
unsigned xr_camera_stride(const xr_camera_t *c)
|
||||
{
|
||||
if (!c->height || !c->planesize[0])
|
||||
return c->bytesperline ? c->bytesperline : c->width;
|
||||
|
||||
double bpp = 1.0;
|
||||
|
||||
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21)
|
||||
bpp = 1.5;
|
||||
|
||||
unsigned s = (unsigned)((double)c->planesize[0] / ((double)c->height * bpp));
|
||||
|
||||
if (s >= c->width && s <= c->width * 4)
|
||||
return s;
|
||||
|
||||
return c->bytesperline ? c->bytesperline : c->width;
|
||||
}
|
||||
|
||||
/*
|
||||
* The Arcturus color cameras (arcimx616) claim 2464x2464 NV12, but measured on
|
||||
* 2026-09-28 their plane 0 holds 10-bit MIPI-packed YUV 4:2:0: 2464 luma rows
|
||||
* then 1232 rows of interleaved UV, each row 2464 packed pixels (3080 bytes)
|
||||
* padded to a 256-byte pitch (3328). Only the first 1972 pixels of a row carry
|
||||
* image; the rest are zero.
|
||||
*/
|
||||
#define IMX616_VALID_WIDTH 1972
|
||||
|
||||
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l)
|
||||
{
|
||||
memset(l, 0, sizeof(*l));
|
||||
l->height = c->height;
|
||||
|
||||
if (c->pixfmt == V4L2_PIX_FMT_NV12 && strstr(c->sensor, "imx616")) {
|
||||
|
||||
unsigned packed = (c->width * 5 + 3) / 4;
|
||||
|
||||
l->fmt = XR_FMT_YUV420_10P;
|
||||
l->pitch = (packed + 255) & ~255u;
|
||||
l->rows = c->height + c->height / 2;
|
||||
l->width = IMX616_VALID_WIDTH < c->width ? IMX616_VALID_WIDTH : c->width;
|
||||
return;
|
||||
}
|
||||
|
||||
l->pitch = xr_camera_stride(c);
|
||||
l->width = c->width < l->pitch ? c->width : l->pitch;
|
||||
|
||||
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21) {
|
||||
l->fmt = XR_FMT_NV12;
|
||||
l->rows = c->height + c->height / 2;
|
||||
} else {
|
||||
l->fmt = XR_FMT_GREY8;
|
||||
l->rows = c->height;
|
||||
}
|
||||
}
|
||||
|
||||
const char *xr_fmt_name(xr_fmt_t f)
|
||||
{
|
||||
switch (f) {
|
||||
case XR_FMT_GREY8: return "grey8";
|
||||
case XR_FMT_NV12: return "nv12";
|
||||
case XR_FMT_YUV420_10P: return "yuv420_10p";
|
||||
}
|
||||
|
||||
return "?";
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------- buffer grouping */
|
||||
|
||||
/*
|
||||
* XRService allocates one udmabuf per plane, plane 0 then plane 1, a whole
|
||||
* queue at a time right after opening the sensor's subdev. Plane 1 matches
|
||||
* VIDIOC_G_FMT exactly; plane 0 has slack, so it is matched with >=.
|
||||
*/
|
||||
static void build_groups(xr_state_t *st)
|
||||
{
|
||||
char current_sensor[XR_SENSOR_LEN] = "";
|
||||
|
||||
for (int i = 0; i < nfdents; i++) {
|
||||
|
||||
if (fdents[i].kind == FD_SUBDEV_SENSOR) {
|
||||
snprintf(current_sensor, sizeof(current_sensor), "%s", fdents[i].sensor);
|
||||
continue;
|
||||
}
|
||||
|
||||
if (fdents[i].kind != FD_DMABUF)
|
||||
continue;
|
||||
|
||||
if (i + 1 >= nfdents || fdents[i + 1].kind != FD_DMABUF)
|
||||
continue;
|
||||
|
||||
size_t s0 = fdents[i].size;
|
||||
size_t s1 = fdents[i + 1].size;
|
||||
bool match = false;
|
||||
|
||||
for (int c = 0; c < st->ncameras; c++) {
|
||||
|
||||
xr_camera_t *cam = &st->cameras[c];
|
||||
|
||||
if (cam->nplanes >= 2 && s1 == cam->planesize[1] && s0 >= cam->planesize[0]) {
|
||||
match = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (!match)
|
||||
continue;
|
||||
|
||||
xr_group_t *g = NULL;
|
||||
|
||||
if (st->ngroups > 0) {
|
||||
|
||||
xr_group_t *last = &st->groups[st->ngroups - 1];
|
||||
|
||||
if (last->planesize[0] == s0 && last->planesize[1] == s1 &&
|
||||
!strcmp(last->sensor, current_sensor))
|
||||
g = last;
|
||||
}
|
||||
|
||||
if (!g) {
|
||||
|
||||
if (st->ngroups >= XR_MAX_GROUPS)
|
||||
break;
|
||||
|
||||
g = &st->groups[st->ngroups++];
|
||||
memset(g, 0, sizeof(*g));
|
||||
g->planesize[0] = s0;
|
||||
g->planesize[1] = s1;
|
||||
snprintf(g->sensor, sizeof(g->sensor), "%s", current_sensor);
|
||||
}
|
||||
|
||||
if (g->nbufs < XR_MAX_RUNBUFS) {
|
||||
g->buf[g->nbufs].xfd = fdents[i].xfd;
|
||||
g->buf[g->nbufs].xfd1 = fdents[i + 1].xfd;
|
||||
g->buf[g->nbufs].size = s0;
|
||||
g->buf[g->nbufs].size1 = s1;
|
||||
g->nbufs++;
|
||||
}
|
||||
|
||||
i++; /* consume the plane 1 descriptor */
|
||||
}
|
||||
|
||||
int keep = 0;
|
||||
|
||||
for (int i = 0; i < st->ngroups; i++)
|
||||
if (st->groups[i].nbufs >= 4)
|
||||
st->groups[keep++] = st->groups[i];
|
||||
|
||||
st->ngroups = keep;
|
||||
|
||||
/*
|
||||
* Bind each run to a camera. The sensor marker alone can be wrong: XRService
|
||||
* sometimes opens another sensor's subdev (e.g. the idle color camera)
|
||||
* between an upper camera's subdev and its buffers, and two upper cameras
|
||||
* can resolve to the same sensor name. So a marker match must also fit the
|
||||
* camera's plane sizes, and each camera takes at most one run.
|
||||
*/
|
||||
for (int pass = 0; pass < 2; pass++)
|
||||
for (int i = 0; i < st->ngroups; i++) {
|
||||
|
||||
xr_group_t *g = &st->groups[i];
|
||||
|
||||
for (int c = 0; c < st->ncameras && !g->cam; c++) {
|
||||
|
||||
xr_camera_t *cam = &st->cameras[c];
|
||||
|
||||
if (pass == 0 && (!g->sensor[0] || strcmp(cam->sensor, g->sensor)))
|
||||
continue;
|
||||
|
||||
if (cam->nplanes < 2 || g->planesize[1] != cam->planesize[1] ||
|
||||
g->planesize[0] < cam->planesize[0])
|
||||
continue;
|
||||
|
||||
bool taken = false;
|
||||
|
||||
for (int k = 0; k < st->ngroups; k++)
|
||||
if (k != i && st->groups[k].cam == cam)
|
||||
taken = true;
|
||||
|
||||
if (!taken)
|
||||
g->cam = cam;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn)
|
||||
{
|
||||
memset(st, 0, sizeof(*st));
|
||||
|
||||
st->pid = find_process(process);
|
||||
|
||||
if (!st->pid) {
|
||||
set_err(err, errn, "%s is not running; start SteamVR on the headset first", process);
|
||||
return false;
|
||||
}
|
||||
|
||||
topo_load_all();
|
||||
|
||||
bool ok = scan_xr_fds(st->pid, err, errn);
|
||||
|
||||
if (ok) {
|
||||
probe_cameras(st);
|
||||
build_groups(st);
|
||||
}
|
||||
|
||||
topo_free_all();
|
||||
|
||||
return ok;
|
||||
}
|
||||
|
||||
void xr_print(const xr_state_t *st, FILE *f)
|
||||
{
|
||||
fprintf(f, "XRService pid %d\n", st->pid);
|
||||
|
||||
for (int i = 0; i < st->ncameras; i++) {
|
||||
|
||||
const xr_camera_t *c = &st->cameras[i];
|
||||
char fcc[5] = {
|
||||
(char)(c->pixfmt & 0xff), (char)((c->pixfmt >> 8) & 0xff),
|
||||
(char)((c->pixfmt >> 16) & 0xff), (char)((c->pixfmt >> 24) & 0xff), 0
|
||||
};
|
||||
|
||||
fprintf(f, " camera %-12s minor %-3u %-16s %ux%u %s pitch %u planes %zu %zu role=%s\n",
|
||||
c->path, c->minor, c->sensor, c->width, c->height, fcc,
|
||||
xr_camera_stride(c), c->planesize[0], c->planesize[1], c->role);
|
||||
}
|
||||
|
||||
for (int i = 0; i < st->ngroups; i++) {
|
||||
|
||||
const xr_group_t *g = &st->groups[i];
|
||||
|
||||
fprintf(f, " queue %d: %d buffers plane0=%zu plane1=%zu fds %d..%d sensor '%s' -> %s\n",
|
||||
i, g->nbufs, g->planesize[0], g->planesize[1],
|
||||
g->buf[0].xfd, g->buf[g->nbufs - 1].xfd1, g->sensor,
|
||||
g->cam ? g->cam->path : "(unbound)");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,82 @@
|
||||
/*
|
||||
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
|
||||
*
|
||||
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
|
||||
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <sys/types.h>
|
||||
|
||||
#include <linux/videodev2.h>
|
||||
|
||||
#define XR_MAX_CAMERAS 16
|
||||
#define XR_MAX_GROUPS 32
|
||||
#define XR_MAX_RUNBUFS 128
|
||||
#define XR_SENSOR_LEN 64
|
||||
|
||||
typedef struct {
|
||||
int node; /* N from /dev/videoN */
|
||||
unsigned minor; /* char device minor, as tracepoints report it */
|
||||
char path[64];
|
||||
unsigned width;
|
||||
unsigned height;
|
||||
unsigned bytesperline;
|
||||
unsigned nplanes;
|
||||
size_t planesize[VIDEO_MAX_PLANES];
|
||||
uint32_t pixfmt;
|
||||
char sensor[XR_SENSOR_LEN]; /* media entity name of the sensor */
|
||||
const char *role;
|
||||
} xr_camera_t;
|
||||
|
||||
typedef struct {
|
||||
int xfd; /* plane 0 descriptor in XRService */
|
||||
int xfd1; /* plane 1 descriptor in XRService */
|
||||
size_t size;
|
||||
size_t size1;
|
||||
} xr_bufref_t;
|
||||
|
||||
/* One run of buffers XRService allocated for a camera queue, in allocation order. */
|
||||
typedef struct {
|
||||
size_t planesize[2];
|
||||
int nbufs;
|
||||
xr_bufref_t buf[XR_MAX_RUNBUFS];
|
||||
char sensor[XR_SENSOR_LEN]; /* from the preceding sensor subdev */
|
||||
xr_camera_t *cam;
|
||||
} xr_group_t;
|
||||
|
||||
typedef struct {
|
||||
pid_t pid;
|
||||
xr_camera_t cameras[XR_MAX_CAMERAS];
|
||||
int ncameras;
|
||||
xr_group_t groups[XR_MAX_GROUPS];
|
||||
int ngroups;
|
||||
} xr_state_t;
|
||||
|
||||
typedef enum {
|
||||
XR_FMT_GREY8, /* 8-bit mono */
|
||||
XR_FMT_NV12, /* 8-bit Y plane then interleaved UV, same pitch */
|
||||
XR_FMT_YUV420_10P /* like NV12, but 10-bit MIPI-packed (4 px in 5 bytes) */
|
||||
} xr_fmt_t;
|
||||
|
||||
/* Where the image really sits in plane 0; V4L2's numbers can be misleading. */
|
||||
typedef struct {
|
||||
xr_fmt_t fmt;
|
||||
unsigned pitch; /* bytes per row */
|
||||
unsigned rows; /* rows in plane 0: luma, plus chroma for YUV */
|
||||
unsigned width; /* valid pixels per row */
|
||||
unsigned height; /* luma rows */
|
||||
} xr_layout_t;
|
||||
|
||||
/* Scan XRService's descriptors and the media graph. Needs root. */
|
||||
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn);
|
||||
unsigned xr_camera_stride(const xr_camera_t *c);
|
||||
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l);
|
||||
const char *xr_fmt_name(xr_fmt_t f);
|
||||
void xr_print(const xr_state_t *st, FILE *f);
|
||||
void xr_slugify(const char *in, char *out, size_t n);
|
||||
@@ -0,0 +1,19 @@
|
||||
# Template: the installer replaces @REPO@ with the repo path on the Frame.
|
||||
[Unit]
|
||||
Description=Frametop camera broker: the headset cameras' frames, for hand tracking
|
||||
Documentation=file://@REPO@/hands/README.md
|
||||
# It borrows XRService's camera buffers, so it comes and goes with SteamVR.
|
||||
After=steamvr.service
|
||||
PartOf=steamvr.service
|
||||
|
||||
[Service]
|
||||
# On the host: the dev container can't reach XRService. Its file capabilities (set by
|
||||
# hands/run.sh install) let it borrow the buffers; it drops them once set up. It exits when
|
||||
# XRService restarts, and comes back to attach to the new one.
|
||||
ExecStart=@REPO@/hands/build/ft-camd --status 60
|
||||
Restart=always
|
||||
RestartSec=5
|
||||
TimeoutStopSec=5
|
||||
|
||||
[Install]
|
||||
WantedBy=steamvr.service
|
||||
@@ -0,0 +1,20 @@
|
||||
# Template: the installer replaces @REPO@ with the repo path on the Frame.
|
||||
[Unit]
|
||||
Description=Frametop hand tracking: hands for the screens' hand cutouts, pinches for the pointer
|
||||
Documentation=file://@REPO@/hands/README.md
|
||||
After=steamvr.service frametop-camd.service
|
||||
Wants=frametop-camd.service
|
||||
PartOf=steamvr.service
|
||||
|
||||
[Service]
|
||||
# In the dev container (it's built against Fedora's libraries). It reads ft-camd's ring and
|
||||
# writes /run/user/UID/frametop-hands/hands and gestures. Settings: HANDS_* in ~/.config/frametop.conf.
|
||||
ExecStartPre=-@REPO@/scripts/container-up.sh
|
||||
ExecStartPre=-/usr/bin/pkill -x ft-hands
|
||||
ExecStart=%h/.local/bin/distrobox enter dev -- @REPO@/hands/build/ft-hands --status 60
|
||||
Restart=always
|
||||
RestartSec=5
|
||||
TimeoutStopSec=5
|
||||
|
||||
[Install]
|
||||
WantedBy=steamvr.service
|
||||
@@ -0,0 +1,66 @@
|
||||
/*
|
||||
* fh_gestures - pinch state ft-hands publishes for input: look at something and pinch to
|
||||
* click it, pinch and move to drag (the Vision Pro model, with the eye tracker doing the
|
||||
* looking). /run/user/UID/frametop-hands/gestures, next to the hands
|
||||
* file, with the same sequence lock (read seq, copy, read seq again; use the copy only if
|
||||
* both reads are the same even number) and the same frame: metres in the head frame at
|
||||
* capture time, OpenVR's HMD frame (+x right, +y up, -z forward).
|
||||
*
|
||||
* One slot per side: pinch[0] is the left hand, pinch[1] the right. A pinch follows the
|
||||
* hand it began on until it ends. It begins when the thumb and index tips close within
|
||||
* begin_m and ends when they open past end_m (the gap between keeps it from flickering),
|
||||
* or when the hand stays lost too long (FH_PINCH_LOST).
|
||||
*
|
||||
* Don't miss short pinches: a reader that polls slower than a quick tap still sees it,
|
||||
* because begins and ends count every pinch. When begins changed, a pinch began at
|
||||
* begin_ns; when ends changed, one ended at end_ns. begins - ends is 1 while pinching.
|
||||
*
|
||||
* Drags: point is where the pinch is now, begin_point where it began. Turn each into the
|
||||
* room with the HMD pose at its capture time (capture_ns, begin_ns) before subtracting,
|
||||
* so turning your head doesn't drag.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <assert.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define FH_GESTURES_MAGIC "FHGEST01"
|
||||
#define FH_GESTURES_VERSION 1
|
||||
|
||||
enum {
|
||||
FH_PINCH_TRACKED = 1u << 0, /* the hand was tracked in this frame */
|
||||
FH_PINCH_DOWN = 1u << 1, /* pinching now */
|
||||
FH_PINCH_LOST = 1u << 2, /* the last pinch ended because the hand was lost */
|
||||
};
|
||||
|
||||
typedef struct {
|
||||
uint32_t flags; /* FH_PINCH_* */
|
||||
uint32_t hand_id; /* fh_hand_t.id of the hand, 0 if none */
|
||||
uint32_t begins; /* pinches begun so far */
|
||||
uint32_t ends; /* pinches ended so far */
|
||||
uint64_t begin_ns; /* capture time (CLOCK_MONOTONIC) the current or */
|
||||
/* last pinch began */
|
||||
uint64_t end_ns; /* ... the last pinch ended */
|
||||
float distance; /* thumb tip to index tip, m, at this user's hand */
|
||||
/* size */
|
||||
float strength; /* 0 open (end_m or more) .. 1 closed (begin_m) */
|
||||
float point[3]; /* between the thumb and index tips */
|
||||
float begin_point[3]; /* point when the current or last pinch began */
|
||||
} fh_pinch_t; /* 64 bytes */
|
||||
|
||||
typedef struct {
|
||||
char magic[8];
|
||||
uint32_t version;
|
||||
uint32_t size;
|
||||
volatile uint64_t seq;
|
||||
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
|
||||
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
|
||||
float begin_m; /* the thresholds in use */
|
||||
float end_m;
|
||||
uint8_t reserved[16];
|
||||
fh_pinch_t pinch[2]; /* [0] left hand, [1] right hand */
|
||||
} fh_gestures_t;
|
||||
|
||||
static_assert(sizeof(fh_pinch_t) == 64, "fh_pinch_t layout");
|
||||
static_assert(sizeof(fh_gestures_t) == 64 + 2 * 64, "fh_gestures_t layout");
|
||||
@@ -0,0 +1,58 @@
|
||||
/*
|
||||
* fh_hands - the tracked-hands file ft-hands publishes for ft-screens' hand cutouts
|
||||
* (/run/user/UID/frametop-hands/hands, directory mode 0700), rewritten in place
|
||||
* under a sequence lock: read seq, copy, read seq again; use the copy only if
|
||||
* both reads are the same even number.
|
||||
*
|
||||
* Positions are metres in the head frame at capture time, which is OpenVR's HMD
|
||||
* frame (+x right, +y up, -z forward). Turn them into the room with the HMD pose
|
||||
* at capture_ns (CLOCK_MONOTONIC). Writer: ft-hands (hands/track/io.cpp).
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <assert.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define FH_HANDS_MAGIC "FHHANDS1"
|
||||
#define FH_HANDS_VERSION 1
|
||||
#define FH_HANDS_MAX_HANDS 2
|
||||
#define FH_HANDS_MAX_CAPSULES 64
|
||||
|
||||
enum {
|
||||
FH_HAND_RIGHT = 1u << 0, /* else the left hand */
|
||||
FH_HAND_STEREO = 1u << 1, /* triangulated from two or more cameras */
|
||||
};
|
||||
|
||||
typedef struct {
|
||||
uint32_t id; /* stays the same while the hand is tracked */
|
||||
uint32_t flags; /* FH_HAND_* */
|
||||
float confidence;
|
||||
float reserved;
|
||||
float pts[21][3]; /* MediaPipe hand landmarks */
|
||||
uint32_t ncapsules; /* this hand's capsules, which follow the */
|
||||
/* previous hands' in capsules[] */
|
||||
} fh_hand_t; /* 272 bytes */
|
||||
|
||||
typedef struct {
|
||||
float a[3], b[3]; /* segment ends */
|
||||
float ra, rb; /* radius at each end */
|
||||
} fh_capsule_t; /* 32 bytes: the hand's shape, to cut out */
|
||||
|
||||
typedef struct {
|
||||
char magic[8];
|
||||
uint32_t version;
|
||||
uint32_t size;
|
||||
volatile uint64_t seq;
|
||||
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
|
||||
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
|
||||
uint32_t nhands;
|
||||
uint32_t ncapsules;
|
||||
uint8_t reserved[16];
|
||||
fh_hand_t hands[FH_HANDS_MAX_HANDS];
|
||||
fh_capsule_t capsules[FH_HANDS_MAX_CAPSULES];
|
||||
} fh_hands_t;
|
||||
|
||||
static_assert(sizeof(fh_hand_t) == 272, "fh_hand_t layout");
|
||||
static_assert(sizeof(fh_capsule_t) == 32, "fh_capsule_t layout");
|
||||
static_assert(sizeof(fh_hands_t) == 64 + 2 * 272 + 64 * 32, "fh_hands_t layout");
|
||||
@@ -0,0 +1,12 @@
|
||||
The models in ncnn/ are converted from the OpenCV Zoo ONNX ports of Google's MediaPipe hand
|
||||
models, by tools/convert_models.py:
|
||||
|
||||
- palm.ncnn.*: palm_detection_mediapipe_2023feb (https://huggingface.co/opencv/palm_detection_mediapipe)
|
||||
- hand.ncnn.*: handpose_estimation_mediapipe_2023feb (https://huggingface.co/opencv/handpose_estimation_mediapipe)
|
||||
|
||||
MediaPipe is Copyright Google LLC. The models and the OpenCV Zoo ports are licensed under the
|
||||
Apache License, Version 2.0 (https://www.apache.org/licenses/LICENSE-2.0).
|
||||
|
||||
Changes made here: converted to ncnn with pnnx, with the palm detector's channel pads
|
||||
rewritten as ncnn Padding layers, and quantized to 8 bits (the *-int8.ncnn.* files) with
|
||||
ncnn's tools.
|
||||
Binary file not shown.
@@ -0,0 +1,79 @@
|
||||
7767517
|
||||
77 90
|
||||
Input in0 0 1 in0
|
||||
Convolution convclip_0 1 1 in0 2 0=24 1=3 3=2 15=1 16=1 5=1 6=648 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 4=1 5=1 6=216 7=24 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_10 1 1 3 4 0=16 1=1 5=1 6=384 8=2
|
||||
Split splitncnn_0 1 2 4 5 6
|
||||
Convolution convclip_1 1 1 6 7 0=64 1=1 5=1 6=1024 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 3=2 15=1 16=1 5=1 6=576 7=64 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_12 1 1 8 9 0=16 1=1 5=1 6=1024 8=2
|
||||
Pooling maxpool2d_1 1 1 5 10 1=2 2=2 5=1
|
||||
BinaryOp add_0 2 1 9 10 11
|
||||
Split splitncnn_1 1 2 11 12 13
|
||||
Convolution convclip_2 1 1 13 14 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 4=1 5=1 6=864 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_14 1 1 15 16 0=16 1=1 5=1 6=1536 8=2
|
||||
BinaryOp add_1 2 1 16 12 17
|
||||
Convolution convclip_3 1 1 17 18 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_16 1 1 19 20 0=24 1=1 5=1 6=2304 8=2
|
||||
Split splitncnn_2 1 2 20 21 22
|
||||
Convolution convclip_4 1 1 22 23 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 4=2 5=1 6=3600 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_18 1 1 24 25 0=24 1=1 5=1 6=3456 8=2
|
||||
BinaryOp add_2 2 1 25 21 26
|
||||
Convolution convclip_5 1 1 26 27 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 3=2 15=1 16=1 5=1 6=1296 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_20 1 1 28 29 0=48 1=1 5=1 6=6912 8=2
|
||||
Split splitncnn_3 1 2 29 30 31
|
||||
Convolution convclip_6 1 1 31 32 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_22 1 1 33 34 0=48 1=1 5=1 6=13824 8=2
|
||||
BinaryOp add_3 2 1 34 30 35
|
||||
Split splitncnn_4 1 2 35 36 37
|
||||
Convolution convclip_7 1 1 37 38 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_24 1 1 39 40 0=48 1=1 5=1 6=13824 8=2
|
||||
BinaryOp add_4 2 1 40 36 41
|
||||
Convolution convclip_8 1 1 41 42 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 4=2 5=1 6=7200 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_26 1 1 43 44 0=64 1=1 5=1 6=18432 8=2
|
||||
Split splitncnn_5 1 2 44 45 46
|
||||
Convolution convclip_9 1 1 46 47 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_28 1 1 48 49 0=64 1=1 5=1 6=24576 8=2
|
||||
BinaryOp add_5 2 1 49 45 50
|
||||
Split splitncnn_6 1 2 50 51 52
|
||||
Convolution convclip_10 1 1 52 53 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_30 1 1 54 55 0=64 1=1 5=1 6=24576 8=2
|
||||
BinaryOp add_6 2 1 55 51 56
|
||||
Convolution convclip_11 1 1 56 57 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 3=2 4=1 15=2 16=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_32 1 1 58 59 0=112 1=1 5=1 6=43008 8=2
|
||||
Split splitncnn_7 1 2 59 60 61
|
||||
Convolution convclip_12 1 1 61 62 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_34 1 1 63 64 0=112 1=1 5=1 6=75264 8=2
|
||||
BinaryOp add_7 2 1 64 60 65
|
||||
Split splitncnn_8 1 2 65 66 67
|
||||
Convolution convclip_13 1 1 67 68 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_36 1 1 69 70 0=112 1=1 5=1 6=75264 8=2
|
||||
BinaryOp add_8 2 1 70 66 71
|
||||
Split splitncnn_9 1 2 71 72 73
|
||||
Convolution convclip_14 1 1 73 74 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_38 1 1 75 76 0=112 1=1 5=1 6=75264 8=2
|
||||
BinaryOp add_9 2 1 76 72 77
|
||||
Convolution convclip_15 1 1 77 78 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 4=1 5=1 6=6048 7=672 8=1 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Pooling gap_0 1 1 79 80 0=1 4=1
|
||||
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
|
||||
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
|
||||
Split splitncnn_10 1 4 82 83 84 85 86
|
||||
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336 8=2
|
||||
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336 8=2
|
||||
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 8=2 9=4
|
||||
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 8=2 9=4
|
||||
Binary file not shown.
@@ -0,0 +1,79 @@
|
||||
7767517
|
||||
77 90
|
||||
Input in0 0 1 in0
|
||||
Convolution convclip_0 1 1 in0 2 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=648 9=3
|
||||
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=216 7=24 9=3
|
||||
Convolution conv_10 1 1 3 4 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=384
|
||||
Split splitncnn_0 1 2 4 5 6
|
||||
Convolution convclip_1 1 1 6 7 0=64 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024 9=3
|
||||
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=576 7=64 9=3
|
||||
Convolution conv_12 1 1 8 9 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
Pooling maxpool2d_1 1 1 5 10 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
BinaryOp add_0 2 1 9 10 11 0=0
|
||||
Split splitncnn_1 1 2 11 12 13
|
||||
Convolution convclip_2 1 1 13 14 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
|
||||
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=864 7=96 9=3
|
||||
Convolution conv_14 1 1 15 16 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
|
||||
BinaryOp add_1 2 1 16 12 17 0=0
|
||||
Convolution convclip_3 1 1 17 18 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
|
||||
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400 7=96 9=3
|
||||
Convolution conv_16 1 1 19 20 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2304
|
||||
Split splitncnn_2 1 2 20 21 22
|
||||
Convolution convclip_4 1 1 22 23 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
|
||||
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3600 7=144 9=3
|
||||
Convolution conv_18 1 1 24 25 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456
|
||||
BinaryOp add_2 2 1 25 21 26 0=0
|
||||
Convolution convclip_5 1 1 26 27 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
|
||||
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=1296 7=144 9=3
|
||||
Convolution conv_20 1 1 28 29 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=6912
|
||||
Split splitncnn_3 1 2 29 30 31
|
||||
Convolution convclip_6 1 1 31 32 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
|
||||
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
|
||||
Convolution conv_22 1 1 33 34 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
|
||||
BinaryOp add_3 2 1 34 30 35 0=0
|
||||
Split splitncnn_4 1 2 35 36 37
|
||||
Convolution convclip_7 1 1 37 38 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
|
||||
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
|
||||
Convolution conv_24 1 1 39 40 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
|
||||
BinaryOp add_4 2 1 40 36 41 0=0
|
||||
Convolution convclip_8 1 1 41 42 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
|
||||
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=7200 7=288 9=3
|
||||
Convolution conv_26 1 1 43 44 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=18432
|
||||
Split splitncnn_5 1 2 44 45 46
|
||||
Convolution convclip_9 1 1 46 47 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
|
||||
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
|
||||
Convolution conv_28 1 1 48 49 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
|
||||
BinaryOp add_5 2 1 49 45 50 0=0
|
||||
Split splitncnn_6 1 2 50 51 52
|
||||
Convolution convclip_10 1 1 52 53 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
|
||||
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
|
||||
Convolution conv_30 1 1 54 55 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
|
||||
BinaryOp add_6 2 1 55 51 56 0=0
|
||||
Convolution convclip_11 1 1 56 57 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
|
||||
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=9600 7=384 9=3
|
||||
Convolution conv_32 1 1 58 59 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=43008
|
||||
Split splitncnn_7 1 2 59 60 61
|
||||
Convolution convclip_12 1 1 61 62 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
|
||||
Convolution conv_34 1 1 63 64 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
|
||||
BinaryOp add_7 2 1 64 60 65 0=0
|
||||
Split splitncnn_8 1 2 65 66 67
|
||||
Convolution convclip_13 1 1 67 68 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
|
||||
Convolution conv_36 1 1 69 70 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
|
||||
BinaryOp add_8 2 1 70 66 71 0=0
|
||||
Split splitncnn_9 1 2 71 72 73
|
||||
Convolution convclip_14 1 1 73 74 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
|
||||
Convolution conv_38 1 1 75 76 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
|
||||
BinaryOp add_9 2 1 76 72 77 0=0
|
||||
Convolution convclip_15 1 1 77 78 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=6048 7=672 9=3
|
||||
Pooling gap_0 1 1 79 80 0=1 4=1
|
||||
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
|
||||
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
|
||||
Split splitncnn_10 1 4 82 83 84 85 86
|
||||
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336
|
||||
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336
|
||||
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 9=4
|
||||
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 9=4
|
||||
Binary file not shown.
@@ -0,0 +1,151 @@
|
||||
7767517
|
||||
149 177
|
||||
Input in0 0 1 in0
|
||||
Convolution padconv_0 1 1 in0 2 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 8=2
|
||||
PReLU prelu_41 1 1 2 3 0=32
|
||||
Split splitncnn_0 1 2 3 4 5
|
||||
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 4=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_12 1 1 6 7 0=32 1=1 5=1 6=1024 8=2
|
||||
BinaryOp add_0 2 1 4 7 8
|
||||
PReLU prelu_42 1 1 8 9 0=32
|
||||
Split splitncnn_1 1 2 9 10 11
|
||||
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 4=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_13 1 1 12 13 0=32 1=1 5=1 6=1024 8=2
|
||||
BinaryOp add_1 2 1 10 13 14
|
||||
PReLU prelu_43 1 1 14 15 0=32
|
||||
Split splitncnn_2 1 2 15 16 17
|
||||
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 4=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_14 1 1 18 19 0=32 1=1 5=1 6=1024 8=2
|
||||
BinaryOp add_2 2 1 16 19 20
|
||||
PReLU prelu_44 1 1 20 21 0=32
|
||||
Split splitncnn_3 1 2 21 22 23
|
||||
Pooling maxpool2d_2 1 1 22 24 1=2 2=2 5=1
|
||||
Padding Pad_16 1 1 24 25 8=32
|
||||
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_15 1 1 26 27 0=64 1=1 5=1 6=2048 8=2
|
||||
BinaryOp add_3 2 1 25 27 28
|
||||
PReLU prelu_45 1 1 28 29 0=64
|
||||
Split splitncnn_4 1 2 29 30 31
|
||||
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_16 1 1 32 33 0=64 1=1 5=1 6=4096 8=2
|
||||
BinaryOp add_4 2 1 30 33 34
|
||||
PReLU prelu_46 1 1 34 35 0=64
|
||||
Split splitncnn_5 1 2 35 36 37
|
||||
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_17 1 1 38 39 0=64 1=1 5=1 6=4096 8=2
|
||||
BinaryOp add_5 2 1 36 39 40
|
||||
PReLU prelu_47 1 1 40 41 0=64
|
||||
Split splitncnn_6 1 2 41 42 43
|
||||
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_18 1 1 44 45 0=64 1=1 5=1 6=4096 8=2
|
||||
BinaryOp add_6 2 1 42 45 46
|
||||
PReLU prelu_48 1 1 46 47 0=64
|
||||
Split splitncnn_7 1 2 47 48 49
|
||||
Pooling maxpool2d_3 1 1 48 50 1=2 2=2 5=1
|
||||
Padding Pad_34 1 1 50 51 8=64
|
||||
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 3=2 4=1 15=2 16=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_19 1 1 52 53 0=128 1=1 5=1 6=8192 8=2
|
||||
BinaryOp add_7 2 1 51 53 54
|
||||
PReLU prelu_49 1 1 54 55 0=128
|
||||
Split splitncnn_8 1 2 55 56 57
|
||||
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_20 1 1 58 59 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_8 2 1 56 59 60
|
||||
PReLU prelu_50 1 1 60 61 0=128
|
||||
Split splitncnn_9 1 2 61 62 63
|
||||
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_21 1 1 64 65 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_9 2 1 62 65 66
|
||||
PReLU prelu_51 1 1 66 67 0=128
|
||||
Split splitncnn_10 1 2 67 68 69
|
||||
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_22 1 1 70 71 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_10 2 1 68 71 72
|
||||
PReLU prelu_52 1 1 72 73 0=128
|
||||
Split splitncnn_11 1 3 73 74 75 76
|
||||
Pooling maxpool2d_4 1 1 75 77 1=2 2=2 5=1
|
||||
Padding Pad_52 1 1 77 78 8=128
|
||||
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 3=2 4=1 15=2 16=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_23 1 1 79 80 0=256 1=1 5=1 6=32768 8=2
|
||||
BinaryOp add_11 2 1 78 80 81
|
||||
PReLU prelu_53 1 1 81 82 0=256
|
||||
Split splitncnn_12 1 2 82 83 84
|
||||
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_24 1 1 85 86 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_12 2 1 83 86 87
|
||||
PReLU prelu_54 1 1 87 88 0=256
|
||||
Split splitncnn_13 1 2 88 89 90
|
||||
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_25 1 1 91 92 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_13 2 1 89 92 93
|
||||
PReLU prelu_55 1 1 93 94 0=256
|
||||
Split splitncnn_14 1 2 94 95 96
|
||||
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_26 1 1 97 98 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_14 2 1 95 98 99
|
||||
PReLU prelu_56 1 1 99 100 0=256
|
||||
Split splitncnn_15 1 3 100 101 102 103
|
||||
Pooling maxpool2d_5 1 1 102 104 1=2 2=2 5=1
|
||||
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 3=2 4=1 15=2 16=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_27 1 1 105 106 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_15 2 1 104 106 107
|
||||
PReLU prelu_57 1 1 107 108 0=256
|
||||
Split splitncnn_16 1 2 108 109 110
|
||||
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_28 1 1 111 112 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_16 2 1 109 112 113
|
||||
PReLU prelu_58 1 1 113 114 0=256
|
||||
Split splitncnn_17 1 2 114 115 116
|
||||
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_29 1 1 117 118 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_17 2 1 115 118 119
|
||||
PReLU prelu_59 1 1 119 120 0=256
|
||||
Split splitncnn_18 1 2 120 121 122
|
||||
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_30 1 1 123 124 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_18 2 1 121 124 125
|
||||
PReLU prelu_60 1 1 125 126 0=256
|
||||
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12
|
||||
Convolution conv_31 1 1 127 128 0=256 1=1 5=1 6=65536 8=2
|
||||
PReLU prelu_61 1 1 128 129 0=256
|
||||
BinaryOp add_19 2 1 101 129 130
|
||||
Split splitncnn_19 1 2 130 131 132
|
||||
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_32 1 1 133 134 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_20 2 1 131 134 135
|
||||
PReLU prelu_62 1 1 135 136 0=256
|
||||
Split splitncnn_20 1 2 136 137 138
|
||||
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_33 1 1 139 140 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_21 2 1 137 140 141
|
||||
PReLU prelu_63 1 1 141 142 0=256
|
||||
Split splitncnn_21 1 3 142 143 144 145
|
||||
Convolution conv_34 1 1 145 146 0=108 1=1 5=1 6=27648 8=2
|
||||
Permute permute_68 1 1 146 147 0=3
|
||||
Reshape reshape_72 1 1 147 148 0=18 1=864
|
||||
Convolution conv_35 1 1 144 149 0=6 1=1 5=1 6=1536 8=2
|
||||
Permute permute_69 1 1 149 150 0=3
|
||||
Reshape reshape_73 1 1 150 151 0=1 1=864
|
||||
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24
|
||||
Convolution conv_36 1 1 152 153 0=128 1=1 5=1 6=32768 8=2
|
||||
PReLU prelu_64 1 1 153 154 0=128
|
||||
BinaryOp add_22 2 1 74 154 155
|
||||
Split splitncnn_22 1 2 155 156 157
|
||||
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_37 1 1 158 159 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_23 2 1 156 159 160
|
||||
PReLU prelu_65 1 1 160 161 0=128
|
||||
Split splitncnn_23 1 2 161 162 163
|
||||
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_38 1 1 164 165 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_24 2 1 162 165 166
|
||||
PReLU prelu_66 1 1 166 167 0=128
|
||||
Split splitncnn_24 1 2 167 168 169
|
||||
Convolution conv_39 1 1 169 170 0=36 1=1 5=1 6=4608 8=2
|
||||
Permute permute_70 1 1 170 171 0=3
|
||||
Reshape reshape_74 1 1 171 172 0=18 1=1152
|
||||
Concat cat_0 2 1 172 148 out0
|
||||
Convolution conv_40 1 1 168 174 0=2 1=1 5=1 6=256 8=2
|
||||
Permute permute_71 1 1 174 175 0=3
|
||||
Reshape reshape_75 1 1 175 176 0=1 1=1152
|
||||
Concat cat_1 2 1 176 151 out1
|
||||
Binary file not shown.
@@ -0,0 +1,151 @@
|
||||
7767517
|
||||
149 177
|
||||
Input in0 0 1 in0
|
||||
Convolution padconv_0 1 1 in0 2 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400
|
||||
PReLU prelu_41 1 1 2 3 0=32
|
||||
Split splitncnn_0 1 2 3 4 5
|
||||
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
|
||||
Convolution conv_12 1 1 6 7 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
BinaryOp add_0 2 1 4 7 8 0=0
|
||||
PReLU prelu_42 1 1 8 9 0=32
|
||||
Split splitncnn_1 1 2 9 10 11
|
||||
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
|
||||
Convolution conv_13 1 1 12 13 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
BinaryOp add_1 2 1 10 13 14 0=0
|
||||
PReLU prelu_43 1 1 14 15 0=32
|
||||
Split splitncnn_2 1 2 15 16 17
|
||||
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
|
||||
Convolution conv_14 1 1 18 19 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
BinaryOp add_2 2 1 16 19 20 0=0
|
||||
PReLU prelu_44 1 1 20 21 0=32
|
||||
Split splitncnn_3 1 2 21 22 23
|
||||
Pooling maxpool2d_2 1 1 22 24 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
Padding Pad_16 1 1 24 25 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=32
|
||||
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=800 7=32
|
||||
Convolution conv_15 1 1 26 27 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2048
|
||||
BinaryOp add_3 2 1 25 27 28 0=0
|
||||
PReLU prelu_45 1 1 28 29 0=64
|
||||
Split splitncnn_4 1 2 29 30 31
|
||||
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
|
||||
Convolution conv_16 1 1 32 33 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
|
||||
BinaryOp add_4 2 1 30 33 34 0=0
|
||||
PReLU prelu_46 1 1 34 35 0=64
|
||||
Split splitncnn_5 1 2 35 36 37
|
||||
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
|
||||
Convolution conv_17 1 1 38 39 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
|
||||
BinaryOp add_5 2 1 36 39 40 0=0
|
||||
PReLU prelu_47 1 1 40 41 0=64
|
||||
Split splitncnn_6 1 2 41 42 43
|
||||
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
|
||||
Convolution conv_18 1 1 44 45 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
|
||||
BinaryOp add_6 2 1 42 45 46 0=0
|
||||
PReLU prelu_48 1 1 46 47 0=64
|
||||
Split splitncnn_7 1 2 47 48 49
|
||||
Pooling maxpool2d_3 1 1 48 50 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
Padding Pad_34 1 1 50 51 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=64
|
||||
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=1600 7=64
|
||||
Convolution conv_19 1 1 52 53 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=8192
|
||||
BinaryOp add_7 2 1 51 53 54 0=0
|
||||
PReLU prelu_49 1 1 54 55 0=128
|
||||
Split splitncnn_8 1 2 55 56 57
|
||||
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_20 1 1 58 59 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_8 2 1 56 59 60 0=0
|
||||
PReLU prelu_50 1 1 60 61 0=128
|
||||
Split splitncnn_9 1 2 61 62 63
|
||||
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_21 1 1 64 65 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_9 2 1 62 65 66 0=0
|
||||
PReLU prelu_51 1 1 66 67 0=128
|
||||
Split splitncnn_10 1 2 67 68 69
|
||||
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_22 1 1 70 71 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_10 2 1 68 71 72 0=0
|
||||
PReLU prelu_52 1 1 72 73 0=128
|
||||
Split splitncnn_11 1 3 73 74 75 76
|
||||
Pooling maxpool2d_4 1 1 75 77 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
Padding Pad_52 1 1 77 78 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=128
|
||||
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=3200 7=128
|
||||
Convolution conv_23 1 1 79 80 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
|
||||
BinaryOp add_11 2 1 78 80 81 0=0
|
||||
PReLU prelu_53 1 1 81 82 0=256
|
||||
Split splitncnn_12 1 2 82 83 84
|
||||
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_24 1 1 85 86 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_12 2 1 83 86 87 0=0
|
||||
PReLU prelu_54 1 1 87 88 0=256
|
||||
Split splitncnn_13 1 2 88 89 90
|
||||
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_25 1 1 91 92 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_13 2 1 89 92 93 0=0
|
||||
PReLU prelu_55 1 1 93 94 0=256
|
||||
Split splitncnn_14 1 2 94 95 96
|
||||
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_26 1 1 97 98 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_14 2 1 95 98 99 0=0
|
||||
PReLU prelu_56 1 1 99 100 0=256
|
||||
Split splitncnn_15 1 3 100 101 102 103
|
||||
Pooling maxpool2d_5 1 1 102 104 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=6400 7=256
|
||||
Convolution conv_27 1 1 105 106 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_15 2 1 104 106 107 0=0
|
||||
PReLU prelu_57 1 1 107 108 0=256
|
||||
Split splitncnn_16 1 2 108 109 110
|
||||
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_28 1 1 111 112 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_16 2 1 109 112 113 0=0
|
||||
PReLU prelu_58 1 1 113 114 0=256
|
||||
Split splitncnn_17 1 2 114 115 116
|
||||
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_29 1 1 117 118 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_17 2 1 115 118 119 0=0
|
||||
PReLU prelu_59 1 1 119 120 0=256
|
||||
Split splitncnn_18 1 2 120 121 122
|
||||
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_30 1 1 123 124 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_18 2 1 121 124 125 0=0
|
||||
PReLU prelu_60 1 1 125 126 0=256
|
||||
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12 6=0
|
||||
Convolution conv_31 1 1 127 128 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
PReLU prelu_61 1 1 128 129 0=256
|
||||
BinaryOp add_19 2 1 101 129 130 0=0
|
||||
Split splitncnn_19 1 2 130 131 132
|
||||
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_32 1 1 133 134 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_20 2 1 131 134 135 0=0
|
||||
PReLU prelu_62 1 1 135 136 0=256
|
||||
Split splitncnn_20 1 2 136 137 138
|
||||
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_33 1 1 139 140 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_21 2 1 137 140 141 0=0
|
||||
PReLU prelu_63 1 1 141 142 0=256
|
||||
Split splitncnn_21 1 3 142 143 144 145
|
||||
Convolution conv_34 1 1 145 146 0=108 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=27648
|
||||
Permute permute_68 1 1 146 147 0=3
|
||||
Reshape reshape_72 1 1 147 148 0=18 1=864
|
||||
Convolution conv_35 1 1 144 149 0=6 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
|
||||
Permute permute_69 1 1 149 150 0=3
|
||||
Reshape reshape_73 1 1 150 151 0=1 1=864
|
||||
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24 6=0
|
||||
Convolution conv_36 1 1 152 153 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
|
||||
PReLU prelu_64 1 1 153 154 0=128
|
||||
BinaryOp add_22 2 1 74 154 155 0=0
|
||||
Split splitncnn_22 1 2 155 156 157
|
||||
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_37 1 1 158 159 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_23 2 1 156 159 160 0=0
|
||||
PReLU prelu_65 1 1 160 161 0=128
|
||||
Split splitncnn_23 1 2 161 162 163
|
||||
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_38 1 1 164 165 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_24 2 1 162 165 166 0=0
|
||||
PReLU prelu_66 1 1 166 167 0=128
|
||||
Split splitncnn_24 1 2 167 168 169
|
||||
Convolution conv_39 1 1 169 170 0=36 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4608
|
||||
Permute permute_70 1 1 170 171 0=3
|
||||
Reshape reshape_74 1 1 171 172 0=18 1=1152
|
||||
Concat cat_0 2 1 172 148 out0 0=0
|
||||
Convolution conv_40 1 1 168 174 0=2 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=256
|
||||
Permute permute_71 1 1 174 175 0=3
|
||||
Reshape reshape_75 1 1 175 176 0=1 1=1152
|
||||
Concat cat_1 2 1 176 151 out1 0=0
|
||||
Executable
+60
@@ -0,0 +1,60 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install, start, stop, or inspect hand tracking on the Frame: ft-camd (the camera broker) and
|
||||
# ft-hands (the tracker), user services that start and stop with SteamVR.
|
||||
# Usage: hands/run.sh install|uninstall
|
||||
# hands/run.sh caps # give ft-camd its capabilities again (a rebuild clears them)
|
||||
# hands/run.sh start|stop|restart|status|log [lines]
|
||||
# install and caps need the password (sudo setcap, once per build of ft-camd). On the Frame,
|
||||
# sudo asks in the terminal. From a PC (or with no terminal), the password comes from
|
||||
# steamos_root_pwd in the repo's .env and is sent to sudo -S on stdin, never on a command line.
|
||||
set -euo pipefail
|
||||
root=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
|
||||
. "$root/scripts/_env.sh"
|
||||
frame="$root/scripts/frame.sh"
|
||||
units="frametop-camd.service frametop-hands.service"
|
||||
# pidfd_getfd on XRService (ptrace_scope=1), system-wide tracepoints, and their root-only
|
||||
# format files. ft-camd drops them all once it has set up.
|
||||
caps=cap_sys_ptrace,cap_perfmon,cap_dac_read_search+ep
|
||||
|
||||
sudo_run() {
|
||||
if [ "$FRAME_LOCAL" = 1 ] && [ -t 0 ]; then
|
||||
sudo bash -c "$1" # asks for the password here
|
||||
return
|
||||
fi
|
||||
local pw
|
||||
pw=$(sed -n 's/^steamos_root_pwd=//p' "$root/.env" 2>/dev/null)
|
||||
pw=${pw#[\"\']}; pw=${pw%[\"\']} # .env values may be quoted
|
||||
[ -n "$pw" ] || { echo "no terminal for sudo, and steamos_root_pwd is missing from $root/.env" >&2; exit 1; }
|
||||
printf '%s\n' "$pw" | on_frame "sudo -S -p '' bash -c $(printf %q "$1")"
|
||||
}
|
||||
|
||||
set_caps() { # only when missing: a rebuild clears them, a reinstall doesn't
|
||||
local bin
|
||||
bin=$(printf %q "$FRAME_REPO/hands/build/ft-camd")
|
||||
if on_frame "getcap $bin | grep -q cap_sys_ptrace"; then
|
||||
echo "ft-camd has its capabilities"
|
||||
return
|
||||
fi
|
||||
sudo_run "setcap $caps $bin && getcap $bin"
|
||||
}
|
||||
|
||||
states="for u in $units; do echo \"\$u: \$(systemctl --user is-active \$u)\"; done"
|
||||
|
||||
case ${1:-status} in
|
||||
install)
|
||||
"$root/hands/build.sh"
|
||||
set_caps
|
||||
for u in $units; do
|
||||
fill_template "$root/hands/$u" | on_frame "mkdir -p ~/.config/systemd/user && cat > ~/.config/systemd/user/$u"
|
||||
done
|
||||
"$frame" --host "set -e; systemctl --user daemon-reload; systemctl --user enable $units
|
||||
if systemctl --user -q is-active steamvr.service; then systemctl --user restart $units; sleep 5; fi
|
||||
$states" ;;
|
||||
caps) set_caps ;;
|
||||
uninstall) "$frame" --host "systemctl --user disable --now $units 2>/dev/null
|
||||
for u in $units; do rm -f ~/.config/systemd/user/\$u; done; systemctl --user daemon-reload; echo removed" ;;
|
||||
start|stop|restart) "$frame" --host "systemctl --user $1 $units; $states" ;;
|
||||
status) "$frame" --host "$states; journalctl --user -u frametop-hands.service --no-pager -o cat -n 4" || true ;;
|
||||
log) "$frame" --host "journalctl --user -u frametop-camd.service -u frametop-hands.service --no-pager -o short -n ${2:-30}" ;;
|
||||
*) echo "usage: $0 install|uninstall|caps|start|stop|restart|status|log [lines]" >&2; exit 2 ;;
|
||||
esac
|
||||
@@ -0,0 +1,197 @@
|
||||
"""Tracking-camera calibration from the headset's factory files.
|
||||
|
||||
/persist/xrservice.json (written by Valve's calibration, loaded by XRService)
|
||||
holds, per camera, Kannala-Brandt fisheye intrinsics ("kb": fx fy cx cy k1-k4,
|
||||
pixel centres at integer coordinates, as in OpenCV's fisheye model) and a pose
|
||||
in the slam_right (Cam0) frame: plus_x/plus_z are the camera axes and position
|
||||
its origin, in mm. /persist/device_config.json gives Cam0's pose in the CAD
|
||||
frame (cv.cad_from_cal, metres) and the head's pose in CAD (head). The CAD frame
|
||||
is +X head-left, +Y up, +Z forward; the head frame is OpenVR's: +x right, +y up,
|
||||
-z forward. Camera frames: +z along the optical axis, +x right and +y down in
|
||||
the image.
|
||||
|
||||
Everything here returns metres in the head frame.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
|
||||
import numpy as np
|
||||
|
||||
XRSERVICE_JSON = '/persist/xrservice.json'
|
||||
DEVICE_JSON = '/persist/device_config.json'
|
||||
# The Arcturus color module's EEPROM: some binary, then its calibration as JSON (world-readable)
|
||||
ARCTURUS_EEPROM = '/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom'
|
||||
ARCTURUS_WIDTH = 1972 # valid pixels per row that XRService's buffers deliver (of 2464)
|
||||
|
||||
|
||||
def _pose(d, scale=1.0):
|
||||
"""4x4 transform from a {plus_x, plus_z, position} pose (child axes in the parent frame)."""
|
||||
x = np.asarray(d['plus_x'], float)
|
||||
z = np.asarray(d['plus_z'], float)
|
||||
y = np.cross(z, x)
|
||||
T = np.eye(4)
|
||||
T[:3, 0], T[:3, 1], T[:3, 2] = x, y, z
|
||||
T[:3, 3] = np.asarray(d['position'], float) * scale
|
||||
return T
|
||||
|
||||
|
||||
class Camera:
|
||||
def __init__(self, name, width, height, kb, head_from_cam):
|
||||
self.name = name
|
||||
self.width, self.height = width, height
|
||||
self.fx, self.fy, self.cx, self.cy = kb['fx'], kb['fy'], kb['cx'], kb['cy']
|
||||
self.k = np.array([kb['k1'], kb['k2'], kb['k3'], kb['k4']])
|
||||
self.head_from_cam = head_from_cam
|
||||
self.R = head_from_cam[:3, :3] # camera axes in the head frame
|
||||
self.origin = head_from_cam[:3, 3] # camera centre in the head frame
|
||||
|
||||
def __repr__(self):
|
||||
return 'Camera(%s %dx%d at %s mm)' % (self.name, self.width, self.height,
|
||||
np.round(self.origin * 1000, 1))
|
||||
|
||||
def _theta_d(self, theta):
|
||||
t2 = theta * theta
|
||||
k1, k2, k3, k4 = self.k
|
||||
return theta * (1 + t2 * (k1 + t2 * (k2 + t2 * (k3 + t2 * k4))))
|
||||
|
||||
def project_cam(self, p):
|
||||
"""Camera-frame points (N,3) -> pixels (N,2). Points behind the lens still map (the lens sees ~180 deg)."""
|
||||
p = np.atleast_2d(p)
|
||||
r = np.hypot(p[:, 0], p[:, 1])
|
||||
theta = np.arctan2(r, p[:, 2])
|
||||
scale = np.where(r > 1e-12, self._theta_d(theta) / np.maximum(r, 1e-12), 0.0)
|
||||
return np.stack([self.fx * p[:, 0] * scale + self.cx, self.fy * p[:, 1] * scale + self.cy], axis=1)
|
||||
|
||||
def unproject(self, uv):
|
||||
"""Pixels (N,2) -> unit rays (N,3) in the camera frame."""
|
||||
uv = np.atleast_2d(np.asarray(uv, float))
|
||||
mx = (uv[:, 0] - self.cx) / self.fx
|
||||
my = (uv[:, 1] - self.cy) / self.fy
|
||||
td = np.hypot(mx, my)
|
||||
theta = td.copy()
|
||||
k1, k2, k3, k4 = self.k
|
||||
for _ in range(8): # Newton on theta_d(theta) = td
|
||||
t2 = theta * theta
|
||||
f = self._theta_d(theta) - td
|
||||
df = 1 + t2 * (3 * k1 + t2 * (5 * k2 + t2 * (7 * k3 + t2 * 9 * k4)))
|
||||
theta = np.clip(theta - f / df, 0.0, np.pi)
|
||||
s = np.where(td > 1e-12, np.sin(theta) / np.maximum(td, 1e-12), 1.0)
|
||||
return np.stack([mx * s, my * s, np.cos(theta)], axis=1)
|
||||
|
||||
def rays(self, uv):
|
||||
"""Pixels -> unit rays in the head frame (all starting at self.origin)."""
|
||||
return self.unproject(uv) @ self.R.T
|
||||
|
||||
def project(self, p_head):
|
||||
"""Head-frame points (N,3) -> pixels (N,2) and depth along the optical axis (N,)."""
|
||||
p = (np.atleast_2d(p_head) - self.origin) @ self.R
|
||||
return self.project_cam(p), p[:, 2]
|
||||
|
||||
def angle_from_axis(self, uv):
|
||||
"""Angle in degrees between each pixel's ray and the optical axis."""
|
||||
return np.degrees(np.arccos(np.clip(self.unproject(uv)[:, 2], -1, 1)))
|
||||
|
||||
|
||||
def device_path(path):
|
||||
"""A headset file such as /persist/xrservice.json. In the dev container the host's / is
|
||||
at /run/host (distrobox doesn't mount /persist); off the Frame, FRAME_JOB_DEVICE_ROOT can
|
||||
point at a folder with copies of them."""
|
||||
root = os.environ.get('FRAME_JOB_DEVICE_ROOT')
|
||||
if root:
|
||||
return root + path
|
||||
if not os.access(path, os.R_OK) and os.access('/run/host' + path, os.R_OK):
|
||||
return '/run/host' + path
|
||||
return path
|
||||
|
||||
|
||||
def load(xrservice=XRSERVICE_JSON, device=DEVICE_JSON):
|
||||
"""{calibration name: Camera} for the tracking cameras, posed in the head frame."""
|
||||
with open(device_path(xrservice)) as f:
|
||||
rig = json.load(f)
|
||||
with open(device_path(device)) as f:
|
||||
dev = json.load(f)
|
||||
cad_from_cam0 = _pose(dev['cv']['cad_from_cal'])
|
||||
head_from_cad = np.linalg.inv(_pose(dev['head']))
|
||||
cams = {}
|
||||
for c in rig['cameras']:
|
||||
kb = next(i for i in c['intrinsics'] if i['cameraModel'] == 'kb')
|
||||
cam0_from_cam = _pose(c['extrinsics'], 1e-3)
|
||||
cams[c['sourceCamera']] = Camera(c['sourceCamera'], c['width'], c['height'], kb,
|
||||
head_from_cad @ cad_from_cam0 @ cam0_from_cam)
|
||||
return cams
|
||||
|
||||
|
||||
def load_color(eeprom=ARCTURUS_EEPROM, device=DEVICE_JSON, scale=2, crop='subtract'):
|
||||
"""{"passthrough_left"/"passthrough_right": Camera} for the Arcturus color cameras, posed in
|
||||
the head frame, for ft-camd --with-color's images (luma at 1/scale size).
|
||||
|
||||
Their calibration is in the CAD frame (mm) with pixel coordinates on the full 2464x2464
|
||||
sensor; each camera also has a cropRegion. crop says how that maps to the delivered
|
||||
image: 'subtract' (image x = sensor x - cropRegion.x) or 'none'. tools/check_color.py
|
||||
tells which fits.
|
||||
"""
|
||||
with open(device_path(eeprom), 'rb') as f:
|
||||
raw = f.read()
|
||||
i = raw.rfind(b'{', 0, raw.find(b'"alignment_method"'))
|
||||
rig, _ = json.JSONDecoder().raw_decode(raw[i:].decode('latin1'))
|
||||
with open(device_path(device)) as f:
|
||||
dev = json.load(f)
|
||||
head_from_cad = np.linalg.inv(_pose(dev['head']))
|
||||
cams = {}
|
||||
for c in rig['cameras']:
|
||||
kb = dict(next(k for k in c['intrinsics'] if k['cameraModel'] == 'kb'))
|
||||
region = c.get('cropRegion', {}) if crop == 'subtract' else {}
|
||||
# integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
|
||||
kb['cx'] = (kb['cx'] - region.get('x', 0) + 0.5) / scale - 0.5
|
||||
kb['cy'] = (kb['cy'] - region.get('y', 0) + 0.5) / scale - 0.5
|
||||
kb['fx'] /= scale
|
||||
kb['fy'] /= scale
|
||||
cams[c['sourceCamera']] = Camera(c['sourceCamera'], ARCTURUS_WIDTH // scale, c['height'] // scale, kb,
|
||||
head_from_cad @ _pose(c['extrinsics'], 1e-3))
|
||||
return cams
|
||||
|
||||
|
||||
def triangulate(origins, dirs, weights=None):
|
||||
"""Least-squares point closest to several rays. Returns (point, rms distance to the rays)."""
|
||||
A = np.zeros((3, 3))
|
||||
b = np.zeros(3)
|
||||
w = np.ones(len(origins)) if weights is None else np.asarray(weights, float)
|
||||
for o, d, wi in zip(origins, dirs, w):
|
||||
P = np.eye(3) - np.outer(d, d)
|
||||
A += wi * P
|
||||
b += wi * P @ o
|
||||
p = np.linalg.solve(A, b)
|
||||
res = [np.linalg.norm((np.eye(3) - np.outer(d, d)) @ (p - o)) for o, d in zip(origins, dirs)]
|
||||
return p, float(np.sqrt(np.mean(np.square(res))))
|
||||
|
||||
|
||||
def triangulate_many(origins, dirs, weights):
|
||||
"""Triangulate K points seen from V cameras at once.
|
||||
|
||||
origins (V,3), dirs (V,K,3) unit rays, weights (V,). Returns points (K,3) and
|
||||
each point's rms distance to its rays (K,).
|
||||
"""
|
||||
P = np.eye(3) - dirs[..., :, None] * dirs[..., None, :] # (V,K,3,3)
|
||||
w = weights[:, None, None, None]
|
||||
A = (w * P).sum(0)
|
||||
b = (w * (P @ origins[:, None, :, None])).sum(0)[..., 0]
|
||||
pts = np.linalg.solve(A, b[..., None])[..., 0]
|
||||
off = pts[None] - origins[:, None, :] # (V,K,3)
|
||||
perp = off - (off * dirs).sum(-1, keepdims=True) * dirs
|
||||
return pts, np.sqrt((perp ** 2).sum(-1).mean(0))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
cams = load()
|
||||
for cam in cams.values():
|
||||
fwd = [float(v) for v in cam.R[:, 2]]
|
||||
print('%-12s at x %+6.1f y %+6.1f z %+6.1f mm, looks %s' % (
|
||||
cam.name, *(cam.origin * 1000),
|
||||
'right' * (fwd[0] > 0.3) + 'left' * (fwd[0] < -0.3) + ' up' * (fwd[1] > 0.3) +
|
||||
' down' * (fwd[1] < -0.3) + ' forward' * (fwd[2] < -0.3) + ' back' * (fwd[2] > 0.3)),
|
||||
np.round(fwd, 2))
|
||||
uv = np.array([[cam.cx + 200, cam.cy - 100], [cam.cx - 0.4 * cam.width, cam.cy + 0.3 * cam.height]])
|
||||
err = np.abs(cam.project_cam(cam.unproject(uv)) - uv).max()
|
||||
assert err < 1e-6, err
|
||||
a, b = cams['slam_left'], cams['slam_right']
|
||||
print('slam baseline %.2f mm' % (1000 * np.linalg.norm(a.origin - b.origin)))
|
||||
@@ -0,0 +1,64 @@
|
||||
"""Which color camera is which, and how their calibration maps onto ft-camd's images.
|
||||
|
||||
usage: python tools/check_color.py REC_DIR [--sets N]
|
||||
|
||||
A recording made with ft-camd --with-color holds color_video<N> frames with each set.
|
||||
This matches features between the two color images and scores every reading of the
|
||||
calibration: which video node is passthrough_left, and whether the calibration's
|
||||
cropRegion is subtracted from x ('subtract') or not ('none'). Only the right reading
|
||||
makes true matches' rays meet in front of both cameras. Then it checks the winner against
|
||||
the side tracking cameras, which tests the CAD-to-head chain shared with them.
|
||||
"""
|
||||
import argparse
|
||||
import itertools
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
|
||||
from tools.check_sides import load_cams, matches, score # noqa: E402
|
||||
from tools.show_set import index, read_set # noqa: E402
|
||||
from tools import calib # noqa: E402
|
||||
|
||||
|
||||
def load_color(crop):
|
||||
return calib.load_color(crop=crop)
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('rec')
|
||||
ap.add_argument('--sets', type=int, default=8)
|
||||
a = ap.parse_args()
|
||||
path = os.path.join(a.rec, 'sets.bin')
|
||||
offs = index(path)
|
||||
sets = [read_set(path, offs[n]) for n in np.linspace(0, len(offs) - 1, a.sets).astype(int)]
|
||||
nodes = sorted(k for k in sets[0] if k.startswith('color_video'))
|
||||
if len(nodes) != 2:
|
||||
sys.exit('need two color_video<N> cameras in the recording (ft-camd --with-color); found %s' % nodes)
|
||||
pairs = [matches(s[nodes[0]][0], s[nodes[1]][0]) for s in sets]
|
||||
print('%d sets, %d matches between %s and %s' % (len(sets), sum(len(p[0]) for p in pairs), *nodes))
|
||||
|
||||
best = None
|
||||
for crop, left in itertools.product(['subtract', 'none'], nodes):
|
||||
cams = load_color(crop)
|
||||
right = nodes[1] if left == nodes[0] else nodes[0]
|
||||
cam = {left: cams['passthrough_left'], right: cams['passthrough_right']}
|
||||
s = np.mean([score(cam[nodes[0]], cam[nodes[1]], ua, ub) for ua, ub in pairs])
|
||||
print(' %s = passthrough_left, crop %-8s: %3.0f%% of matches meet' % (left, crop, 100 * s))
|
||||
if best is None or s > best[0]:
|
||||
best = (s, crop, left, cam)
|
||||
s, crop, left, cam = best
|
||||
print('best: %s = passthrough_left, crop %s (%.0f%%)' % (left, crop, 100 * s))
|
||||
|
||||
mono = load_cams()
|
||||
for node in nodes:
|
||||
for side in ['slam_left', 'slam_right']:
|
||||
ms = [matches(st[node][0], st[side][0]) for st in sets if side in st]
|
||||
sc = np.mean([score(cam[node], mono[side], ua, ub) for ua, ub in ms]) if ms else 0
|
||||
print(' %s vs %-10s: %4d matches, %3.0f%% meet' % (node, side, sum(len(m[0]) for m in ms), 100 * sc))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,128 @@
|
||||
"""Check that the side cameras' images carry the right names (slam_left vs slam_right).
|
||||
|
||||
usage: python tools/check_sides.py REC_DIR [--sets N]
|
||||
python tools/check_sides.py --ring [--sets N] (live, from ft-camd's ring)
|
||||
|
||||
With --ring it exits 0 when the names are right, 3 when they're swapped (run ft-hands
|
||||
with --swap-sides), and 2 when it can't tell (too little texture in view, or the headset
|
||||
isn't worn).
|
||||
|
||||
ft-camd tells the two side cameras' buffers apart by the order XRService allocated them,
|
||||
and after some XRService restarts that order puts each camera's images under the other's
|
||||
name. The tracker then sees every hand in one camera only, at the wrong depth. This
|
||||
matches features between the two images and measures how close each pair's rays pass
|
||||
with the factory calibration, once as named and once swapped: true matches meet in
|
||||
front of both cameras only under the right naming.
|
||||
"""
|
||||
import argparse
|
||||
import os
|
||||
import sys
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
|
||||
from tools.show_set import index, read_set # noqa: E402
|
||||
from tools import calib # noqa: E402
|
||||
|
||||
|
||||
PIPES = {'msm_vfe3_video0': 'slam_left', 'msm_vfe4_video0': 'slam_right'} # as ft-hands maps them
|
||||
|
||||
|
||||
def load_cams():
|
||||
return calib.load()
|
||||
|
||||
|
||||
def matches(a, b):
|
||||
"""Pixel pairs (N,2), (N,2) of ORB matches between two grey images."""
|
||||
clahe = cv2.createCLAHE(2.0, (8, 8))
|
||||
orb = cv2.ORB_create(3000)
|
||||
ka, da = orb.detectAndCompute(clahe.apply(a), None)
|
||||
kb, db = orb.detectAndCompute(clahe.apply(b), None)
|
||||
if da is None or db is None:
|
||||
return np.zeros((0, 2)), np.zeros((0, 2))
|
||||
pairs = cv2.BFMatcher(cv2.NORM_HAMMING).knnMatch(da, db, k=2)
|
||||
good = [p[0] for p in pairs if len(p) == 2 and p[0].distance < 0.75 * p[1].distance]
|
||||
return (np.array([ka[m.queryIdx].pt for m in good]).reshape(-1, 2),
|
||||
np.array([kb[m.trainIdx].pt for m in good]).reshape(-1, 2))
|
||||
|
||||
|
||||
def meet(cam_a, cam_b, ua, ub):
|
||||
"""Per match: closest distance between the two rays (m), and whether they meet in front of both."""
|
||||
ra, rb = cam_a.rays(ua), cam_b.rays(ub)
|
||||
w = cam_b.origin - cam_a.origin
|
||||
n = np.cross(ra, rb)
|
||||
nn = np.linalg.norm(n, axis=1)
|
||||
dist = np.abs(w @ n.T) / np.maximum(nn, 1e-12)
|
||||
# ray parameters at the closest points
|
||||
ta = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, rb.shape), rb), n) / np.maximum(nn ** 2, 1e-12)
|
||||
tb = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, ra.shape), ra), n) / np.maximum(nn ** 2, 1e-12)
|
||||
return dist, (ta > 0.05) & (tb > 0.05)
|
||||
|
||||
|
||||
def score(cam_a, cam_b, ua, ub):
|
||||
"""Share of matches whose rays meet within 1 cm, in front of both cameras."""
|
||||
if len(ua) == 0:
|
||||
return 0.0
|
||||
d, front = meet(cam_a, cam_b, ua, ub)
|
||||
return float(np.mean((d < 0.01) & front))
|
||||
|
||||
|
||||
def recorded_pairs(rec, count):
|
||||
"""(label, slam_left image, slam_right image) from sets spread across a recording."""
|
||||
path = os.path.join(rec, 'sets.bin')
|
||||
offs = index(path)
|
||||
for n in np.linspace(0, len(offs) - 1, count).astype(int):
|
||||
images = read_set(path, offs[n])
|
||||
if 'slam_left' in images and 'slam_right' in images:
|
||||
yield 'set %5d' % n, images['slam_left'][0], images['slam_right'][0]
|
||||
|
||||
|
||||
def live_pairs(count):
|
||||
"""(label, slam_left image, slam_right image) from ft-camd's ring, half a second apart."""
|
||||
import time
|
||||
from tools.ring import Ring
|
||||
ring = Ring()
|
||||
if not ring.alive():
|
||||
sys.exit('ft-camd isn\'t running (no heartbeat)')
|
||||
cams = {}
|
||||
for c in ring.cams:
|
||||
name = PIPES.get(open('/sys/class/video4linux/video%d/name' % c.node).read().strip())
|
||||
if name and not c.name.endswith('-dark'):
|
||||
cams[name] = c
|
||||
for k in range(count):
|
||||
a, b = ring.read(cams['slam_left']), ring.read(cams['slam_right'])
|
||||
if a is not None and b is not None:
|
||||
yield 'frame %2d' % k, a.image, b.image
|
||||
time.sleep(0.5)
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('rec', nargs='?')
|
||||
ap.add_argument('--ring', action='store_true', help='check the live cameras instead of a recording')
|
||||
ap.add_argument('--sets', type=int, default=8, help='how many sets or live frames to check')
|
||||
a = ap.parse_args()
|
||||
if not a.ring and not a.rec:
|
||||
ap.error('give a recording or --ring')
|
||||
cams = load_cams()
|
||||
left, right = cams['slam_left'], cams['slam_right']
|
||||
named = swapped = 0.0
|
||||
n = total_matches = 0
|
||||
for label, img_l, img_r in (live_pairs(a.sets) if a.ring else recorded_pairs(a.rec, a.sets)):
|
||||
ua, ub = matches(img_l, img_r)
|
||||
s_named = score(left, right, ua, ub) # slam_left's image seen by the left camera
|
||||
s_swapped = score(right, left, ua, ub) # ... by the right camera
|
||||
named, swapped, n, total_matches = named + s_named, swapped + s_swapped, n + 1, total_matches + len(ua)
|
||||
print('%s: %4d matches, meeting as named %3.0f%%, swapped %3.0f%%' %
|
||||
(label, len(ua), 100 * s_named, 100 * s_swapped))
|
||||
if n == 0 or total_matches < 100 or abs(named - swapped) / n < 0.2:
|
||||
print('side cameras: can\'t tell (%d matches)' % total_matches)
|
||||
sys.exit(2)
|
||||
print('side cameras: %s (named %.2f, swapped %.2f)' %
|
||||
('as named' if named > swapped else 'SWAPPED', named / n, swapped / n))
|
||||
sys.exit(0 if named > swapped else 3)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,81 @@
|
||||
"""Convert the OpenCV Zoo ONNX ports of MediaPipe's hand models to ncnn.
|
||||
|
||||
The ONNX files are Apache-2.0 ports of MediaPipe's palm detector and hand
|
||||
landmark models (huggingface.co/opencv/palm_detection_mediapipe and
|
||||
huggingface.co/opencv/handpose_estimation_mediapipe). pnnx does the
|
||||
conversion; two fix-ups follow:
|
||||
|
||||
- The palm detector widens channels with ONNX Pad on the channel axis. pnnx
|
||||
emits an ncnn layer called "Pad", which ncnn doesn't have, so rewrite those
|
||||
as ncnn Padding with the channel-end amount (param 8 = behind).
|
||||
- Both models take NHWC input and start with a Permute to NCHW. Drop it, so we
|
||||
can hand ncnn planar CHW Mats straight from the preprocessing step.
|
||||
|
||||
usage: python convert_models.py (writes models/ncnn/{palm,hand}.ncnn.{param,bin})
|
||||
"""
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
ROOT = os.path.join(HERE, '..')
|
||||
PNNX = os.path.join(sys.prefix, 'lib', 'python%d.%d' % sys.version_info[:2], 'site-packages', 'pnnx', 'pnnx')
|
||||
MODELS = [('palm', 'palm_detection_mediapipe_2023feb', 192),
|
||||
('hand', 'handpose_estimation_mediapipe_2023feb', 224)]
|
||||
|
||||
|
||||
def patch(param_text, pnnx_param_text):
|
||||
lines = param_text.splitlines()
|
||||
assert lines[0] == '7767517'
|
||||
nlayers, nblobs = map(int, lines[1].split())
|
||||
body = lines[2:]
|
||||
|
||||
# Channel pads: amounts come from the pnnx graph, which keeps the pads tuple.
|
||||
pads = dict(re.findall(r'^Pad\s+(\S+)\s.*pads=\(0,0,0,0,0,(\d+),0,0\)', pnnx_param_text, re.M))
|
||||
for i, line in enumerate(body):
|
||||
f = line.split()
|
||||
if f[0] == 'Pad':
|
||||
amount = pads[f[1]]
|
||||
body[i] = 'Padding %s %s %s %s %s 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=%s' % (
|
||||
f[1], f[2], f[3], f[4], f[5], amount)
|
||||
|
||||
# Input permute: feed its consumers from in0 instead.
|
||||
perm = next(i for i, line in enumerate(body) if line.split()[0] == 'Permute')
|
||||
f = body[perm].split()
|
||||
assert f[4] == 'in0' and f[6] == '0=4', body[perm]
|
||||
blob = f[5]
|
||||
del body[perm]
|
||||
for i, line in enumerate(body):
|
||||
f = line.split()
|
||||
if f[0] == 'Input':
|
||||
continue
|
||||
nin, nout = int(f[2]), int(f[3])
|
||||
ins = ['in0' if b == blob else b for b in f[4:4 + nin]]
|
||||
body[i] = ' '.join(f[:4] + ins + f[4 + nin:])
|
||||
return '\n'.join(['7767517', '%d %d' % (nlayers - 1, nblobs - 1)] + body) + '\n'
|
||||
|
||||
|
||||
def main():
|
||||
out = os.path.join(ROOT, 'models', 'ncnn')
|
||||
os.makedirs(out, exist_ok=True)
|
||||
for short, name, size in MODELS:
|
||||
src = os.path.join(ROOT, 'models', 'onnx', name + '.onnx')
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
shutil.copy(src, tmp)
|
||||
subprocess.run([PNNX, name + '.onnx', 'inputshape=[1,%d,%d,3]' % (size, size), 'fp16=1'],
|
||||
cwd=tmp, check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
|
||||
with open(os.path.join(tmp, name + '.ncnn.param')) as f:
|
||||
param = f.read()
|
||||
with open(os.path.join(tmp, name + '.pnnx.param')) as f:
|
||||
pparam = f.read()
|
||||
with open(os.path.join(out, short + '.ncnn.param'), 'w') as f:
|
||||
f.write(patch(param, pparam))
|
||||
shutil.copy(os.path.join(tmp, name + '.ncnn.bin'), os.path.join(out, short + '.ncnn.bin'))
|
||||
print('wrote', short)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,256 @@
|
||||
"""How good the tracker's depth is, from a replay's depth dump, without ground truth.
|
||||
|
||||
usage: python3 tools/depth_report.py DEPTH [DEPTH...] [--still M/S]
|
||||
|
||||
DEPTH comes from `hands/build/ft-handreplay DIR --depth DEPTH`. Every measure is split by how the
|
||||
hand was seen: by the two lower cameras ("lower pair"), by a lower and an upper camera on
|
||||
one side ("lower+upper"), or by one camera. Distances are from the head (between the eyes).
|
||||
|
||||
1. How the hands were seen: the share of hand updates in each way, by distance.
|
||||
2. Noise along the line of sight against across it. Each update's palm is compared with a
|
||||
straight line through the two updates before it (ft-handreplay's jitter measure), and the
|
||||
miss is split along the line from the hand's cameras to the palm and across it. Given
|
||||
as a robust sigma per axis, measured (as triangulated) and published (after the One Euro
|
||||
filter), on updates where the published palm moved slower than --still (default 0.15
|
||||
m/s), so the miss is mostly noise and not the hand speeding up. For two cameras,
|
||||
geometry predicts along/across = 2 Z / B: Z the distance, B the cameras' baseline
|
||||
across the line of sight.
|
||||
3. One-camera distance: on two-camera updates, each camera's one-view guess (distance from
|
||||
how big the palm looks, at the user's learned hand size) against the triangulated
|
||||
distance from that camera.
|
||||
4. A camera lost: from two-camera updates, what the tracker would have had if one of the
|
||||
two cameras dropped out there. It keeps the last distance and moves a share of the way
|
||||
to the one-view guess each update (0.1 now, kMonoDepthGain in track/tracker.cpp);
|
||||
also shown with other shares, 0 (keep the distance) and 1 (take each guess), and with
|
||||
the guess first scaled by how far off it was while both cameras saw the hand. Compared with the
|
||||
triangulated distance from that camera, 0.1-2 s after the loss.
|
||||
"""
|
||||
import argparse
|
||||
import collections
|
||||
|
||||
import numpy as np
|
||||
|
||||
BINS = [0.0, 0.35, 0.50, 0.65, 9.0]
|
||||
BIN_NAMES = ['<35 cm', '35-50', '50-65', '65+ cm']
|
||||
MODES = ['lower pair', 'lower+upper', 'one camera']
|
||||
HORIZONS = [0.1, 0.25, 0.5, 1.0, 2.0]
|
||||
# (share of the way toward the one-view guess per update, whether the guess is first scaled by
|
||||
# how far off it was while both cameras saw the hand)
|
||||
GAINS = [(0.0, False), (0.02, False), (0.05, False), (0.1, False), (1.0, False), (0.1, True), (1.0, True)]
|
||||
GAP = 0.1 # s: a longer gap between a hand's updates breaks its run
|
||||
|
||||
|
||||
def load(path):
|
||||
cams, rows = {}, []
|
||||
with open(path) as f:
|
||||
for line in f:
|
||||
w = line.split()
|
||||
if not w:
|
||||
continue
|
||||
if w[0] == '#':
|
||||
if w[1] == 'cam':
|
||||
cams[w[2]] = (np.array([float(x) for x in w[3:6]]), float(w[6]))
|
||||
continue
|
||||
r = {'t': float(w[0]), 'id': int(w[1]), 'side': w[2], 'n': int(w[3]),
|
||||
'cams': w[4], 'res': float(w[5]), 'scale': float(w[6]),
|
||||
'raw': np.array([float(x) for x in w[7:10]]), 'sm': np.array([float(x) for x in w[10:13]]),
|
||||
'views': {}}
|
||||
for k in range(13, len(w), 5):
|
||||
r['views'][w[k]] = np.array([float(x) for x in w[k + 2:k + 5]])
|
||||
rows.append(r)
|
||||
return cams, rows
|
||||
|
||||
|
||||
def mode(r):
|
||||
names = r['cams'].split('+')
|
||||
if r['n'] == 1:
|
||||
return 'one camera'
|
||||
if sorted(names) == ['slam_left', 'slam_right']:
|
||||
return 'lower pair'
|
||||
if len(names) == 2 and all(n.startswith(('slam_', 'upper_')) for n in names) and \
|
||||
names[0].split('_')[1] == names[1].split('_')[1]:
|
||||
return 'lower+upper'
|
||||
return 'other'
|
||||
|
||||
|
||||
def dist_bin(p):
|
||||
return min(np.searchsorted(BINS, np.linalg.norm(p), side='right') - 1, len(BIN_NAMES) - 1)
|
||||
|
||||
|
||||
def tracks(rows):
|
||||
"""A hand's updates, in order, split where they're more than GAP apart."""
|
||||
by_id = collections.defaultdict(list)
|
||||
for r in rows:
|
||||
by_id[r['id']].append(r)
|
||||
for rs in by_id.values():
|
||||
run = [rs[0]]
|
||||
for r in rs[1:]:
|
||||
if r['t'] - run[-1]['t'] >= GAP:
|
||||
yield run
|
||||
run = []
|
||||
run.append(r)
|
||||
yield run
|
||||
|
||||
|
||||
def two_camera_runs(rows):
|
||||
"""Stretches of a hand's updates all seen by the same two cameras."""
|
||||
for run in tracks(rows):
|
||||
seg = []
|
||||
for r in run:
|
||||
ok = r['n'] == 2 and mode(r) != 'other'
|
||||
if ok and seg and r['cams'] == seg[-1]['cams']:
|
||||
seg.append(r)
|
||||
continue
|
||||
if len(seg) > 2:
|
||||
yield seg
|
||||
seg = [r] if ok else []
|
||||
if len(seg) > 2:
|
||||
yield seg
|
||||
|
||||
|
||||
def pct(v, q):
|
||||
return np.percentile(v, q) if len(v) else float('nan')
|
||||
|
||||
|
||||
def seen_share(rows):
|
||||
print('\n1. How the hands were seen (share of hand updates)')
|
||||
count = collections.Counter((mode(r), dist_bin(r['raw'])) for r in rows)
|
||||
total = collections.Counter(dist_bin(r['raw']) for r in rows)
|
||||
print('%-12s' % '' + ''.join('%10s' % b for b in BIN_NAMES) + '%10s' % 'all')
|
||||
for m in MODES + ['other']:
|
||||
cells = [100 * count[m, b] / max(total[b], 1) for b in range(len(BIN_NAMES))]
|
||||
allp = 100 * sum(count[m, b] for b in range(len(BIN_NAMES))) / max(len(rows), 1)
|
||||
print('%-12s' % m + ''.join('%9.0f%%' % c for c in cells) + '%9.0f%%' % allp)
|
||||
print('%-12s' % 'updates' + ''.join('%10d' % total[b] for b in range(len(BIN_NAMES))) + '%10d' % len(rows))
|
||||
res = collections.defaultdict(list)
|
||||
for r in rows:
|
||||
if r['res'] >= 0:
|
||||
res[mode(r)].append(r['res'] * 1000)
|
||||
print('triangulation residual (rms ray miss, median): ' +
|
||||
', '.join('%s %.1f mm' % (m, np.median(v)) for m, v in res.items()))
|
||||
|
||||
|
||||
def noise(rows, cams, still):
|
||||
print('\n2. Noise along the line of sight vs across it (sigma per axis, mm; palm slower than %.2f m/s)' % still)
|
||||
acc = collections.defaultdict(lambda: collections.defaultdict(list))
|
||||
for run in tracks(rows):
|
||||
for a, b, c in zip(run, run[1:], run[2:]):
|
||||
if not (mode(a) == mode(b) == mode(c)) or a['cams'] != b['cams'] or b['cams'] != c['cams']:
|
||||
continue
|
||||
dt0, dt1 = b['t'] - a['t'], c['t'] - b['t']
|
||||
if dt0 < 1e-3 or np.linalg.norm(b['sm'] - a['sm']) / dt0 > still:
|
||||
continue
|
||||
names = b['cams'].split('+')
|
||||
origins = [cams[n][0] for n in names]
|
||||
o = np.mean(origins, axis=0)
|
||||
u = b['raw'] - o
|
||||
z = np.linalg.norm(u)
|
||||
u /= z
|
||||
key = (mode(b), dist_bin(b['raw']))
|
||||
for kind in ('raw', 'sm'):
|
||||
miss = c[kind] - b[kind] - (b[kind] - a[kind]) * (dt1 / dt0)
|
||||
along = miss @ u
|
||||
acc[key][kind + '_along'].append(abs(along))
|
||||
acc[key][kind + '_across'].append(np.linalg.norm(miss - along * u))
|
||||
if len(origins) == 2:
|
||||
base = origins[0] - origins[1]
|
||||
acc[key]['pred'].append(2 * z / np.linalg.norm(base - (base @ u) * u))
|
||||
# |along| is half-normal: sigma = median / 0.674; |across| is Rayleigh (2 axes): sigma = median / 1.177
|
||||
print('%-12s %-7s %6s | %-24s | %-24s | %s' % ('', '', 'n', 'measured along/across', 'published along/across',
|
||||
'ratio measured (geometry)'))
|
||||
for m in MODES:
|
||||
for bi, bn in enumerate(BIN_NAMES):
|
||||
d = acc.get((m, bi))
|
||||
if not d or len(d['raw_along']) < 20:
|
||||
continue
|
||||
s = {k: np.median(v) / (0.674 if k.endswith('along') else 1.177) * 1000
|
||||
for k, v in d.items() if k != 'pred'}
|
||||
pred = '(%.1f)' % np.median(d['pred']) if d['pred'] else ''
|
||||
print('%-12s %-7s %6d | %7.1f / %-5.1f x%-6.1f | %7.1f / %-5.1f x%-6.1f | x%.1f %s' % (
|
||||
m, bn, len(d['raw_along']), s['raw_along'], s['raw_across'], s['raw_along'] / s['raw_across'],
|
||||
s['sm_along'], s['sm_across'], s['sm_along'] / s['sm_across'],
|
||||
s['raw_along'] / s['raw_across'], pred))
|
||||
|
||||
|
||||
def one_camera(rows, cams):
|
||||
print('\n3. One-camera distance vs triangulated, on two-camera updates (error of the one-view guess)')
|
||||
acc = collections.defaultdict(list)
|
||||
for r in rows:
|
||||
if r['n'] != 2 or mode(r) == 'other':
|
||||
continue
|
||||
for name, p in r['views'].items():
|
||||
if np.isnan(p).any():
|
||||
continue
|
||||
o = cams[name][0]
|
||||
truth = np.linalg.norm(r['raw'] - o)
|
||||
acc[name.split('_')[0], dist_bin(r['raw'])].append((np.linalg.norm(p - o) - truth, truth))
|
||||
print('%-8s %-7s %6s %12s %12s %14s %12s' % ('camera', '', 'n', 'median |err|', '90% |err|', 'median |err| %',
|
||||
'bias'))
|
||||
for cam in ('slam', 'upper'):
|
||||
for bi, bn in enumerate(BIN_NAMES):
|
||||
v = acc.get((cam, bi))
|
||||
if not v or len(v) < 20:
|
||||
continue
|
||||
e = np.array([x[0] for x in v])
|
||||
rel = e / np.array([x[1] for x in v])
|
||||
print('%-8s %-7s %6d %9.0f mm %9.0f mm %13.0f%% %+11.0f%%' % (
|
||||
'lower' if cam == 'slam' else 'upper', bn, len(v), 1000 * np.median(abs(e)), 1000 * pct(abs(e), 90),
|
||||
100 * np.median(abs(rel)), 100 * np.median(rel)))
|
||||
|
||||
|
||||
def lost_camera(rows, cams):
|
||||
print('\n4. A camera lost: distance error after the loss (median |err| mm / 90% mm), by how the tracker '
|
||||
'moves toward the one-view guess each update ("scaled": the guess times how far off it was, '
|
||||
'triangulated / guess, median over the last 30 two-camera updates)')
|
||||
errs = collections.defaultdict(list)
|
||||
for run in two_camera_runs(rows):
|
||||
for s in range(1, len(run) - 1, 3):
|
||||
for name in run[s]['views']:
|
||||
o = cams[name][0]
|
||||
guess = lambda r: np.linalg.norm(r['views'][name] - o)
|
||||
ratios = [np.linalg.norm(r['raw'] - o) / guess(r) for r in run[max(0, s - 30):s]
|
||||
if not np.isnan(r['views'][name]).any()]
|
||||
ratio = np.median(ratios) if ratios else 1.0
|
||||
for g, scaled in GAINS:
|
||||
d = np.linalg.norm(run[s - 1]['raw'] - o)
|
||||
h = 0
|
||||
for r in run[s:]:
|
||||
if np.isnan(r['views'][name]).any():
|
||||
break
|
||||
d += g * (guess(r) * (ratio if scaled else 1.0) - d)
|
||||
elapsed = r['t'] - run[s - 1]['t']
|
||||
while h < len(HORIZONS) and elapsed >= HORIZONS[h]:
|
||||
errs[g, scaled, HORIZONS[h], name.split('_')[0]].append(abs(d - np.linalg.norm(r['raw'] - o)))
|
||||
h += 1
|
||||
print('%-8s %-18s' % ('camera', 'toward guess') + ''.join('%14s' % ('%.2g s' % t) for t in HORIZONS))
|
||||
for cam in ('slam', 'upper'):
|
||||
for g, scaled in GAINS:
|
||||
label = {0.0: '0 (keep)', 0.1: '0.1 (now)', 1.0: '1 (guess)'}.get(g, '%g' % g)
|
||||
if scaled:
|
||||
label = '%g scaled' % g
|
||||
cells = []
|
||||
for t in HORIZONS:
|
||||
v = errs.get((g, scaled, t, cam), [])
|
||||
cells.append('%5.0f / %-4.0f' % (1000 * np.median(v), 1000 * pct(v, 90)) if len(v) >= 20 else '%14s' % '-')
|
||||
print('%-8s %-18s' % ('lower' if cam == 'slam' else 'upper', label) + ''.join('%14s' % c for c in cells))
|
||||
n = sum(len(errs.get((0.1, False, HORIZONS[0], c), [])) for c in ('slam', 'upper'))
|
||||
print('(%d simulated losses)' % n)
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('depth', nargs='+')
|
||||
ap.add_argument('--still', type=float, default=0.15, help='m/s: palm speed limit for the noise measure')
|
||||
a = ap.parse_args()
|
||||
for path in a.depth:
|
||||
cams, rows = load(path)
|
||||
print('== %s: %d hand updates, %d hands' % (path, len(rows), len({r['id'] for r in rows})))
|
||||
seen_share(rows)
|
||||
noise(rows, cams, a.still)
|
||||
one_camera(rows, cams)
|
||||
lost_camera(rows, cams)
|
||||
print()
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,80 @@
|
||||
"""Read frames from ft-camd's shared-memory ring (layout: camd/fhring.h)."""
|
||||
import mmap
|
||||
import os
|
||||
import struct
|
||||
import time
|
||||
|
||||
import numpy as np
|
||||
|
||||
RING_FILE = '/run/user/%d/frametop-hands/cam-ring' % os.getuid()
|
||||
MAGIC = b'FHRING01'
|
||||
HDR = struct.Struct('<8sIIIIQqQ16x') # 64 bytes
|
||||
CAM = struct.Struct('<32s32siIIIIIQQQQQ32x') # 160 bytes
|
||||
SLOT = struct.Struct('<QQQQQIf16x') # 64 bytes
|
||||
MAX_CAMS = 8
|
||||
LATEST_OFF = 32 + 32 + 4 * 6 + 8 * 2 # cam.latest within fh_ring_cam_t
|
||||
HEARTBEAT_OFF = 40
|
||||
|
||||
|
||||
class Frame:
|
||||
__slots__ = ('cam', 'frame', 'capture_ns', 'dqbuf_ns', 'publish_ns', 'v4l2_seq', 'mean', 'image')
|
||||
|
||||
def __init__(self, cam, fields, image):
|
||||
self.cam = cam
|
||||
(_, self.frame, self.capture_ns, self.dqbuf_ns, self.publish_ns, self.v4l2_seq, self.mean) = fields
|
||||
self.image = image
|
||||
|
||||
|
||||
class RingCamera:
|
||||
def __init__(self, index, fields):
|
||||
(sensor, name, self.node, self.format, self.width, self.height, self.stride, self.nslots,
|
||||
self.slot_offset, self.slot_bytes, _latest, _pub, _drop) = fields
|
||||
self.index = index
|
||||
self.sensor = sensor.split(b'\0', 1)[0].decode()
|
||||
self.name = name.split(b'\0', 1)[0].decode()
|
||||
self.latest_off = HDR.size + index * CAM.size + LATEST_OFF
|
||||
|
||||
def __repr__(self):
|
||||
return 'RingCamera(video%d %s %dx%d)' % (self.node, self.sensor, self.width, self.height)
|
||||
|
||||
|
||||
class Ring:
|
||||
def __init__(self, path=RING_FILE):
|
||||
fd = os.open(path, os.O_RDONLY)
|
||||
try:
|
||||
self.map = mmap.mmap(fd, 0, mmap.MAP_SHARED, mmap.PROT_READ)
|
||||
finally:
|
||||
os.close(fd)
|
||||
magic, version, hdr_bytes, ncams, _, file_bytes, self.writer_pid, _ = HDR.unpack_from(self.map, 0)
|
||||
if magic != MAGIC or version != 1:
|
||||
raise RuntimeError('%s is not an ft-camd ring (magic %r version %d)' % (path, magic, version))
|
||||
self.cams = [RingCamera(i, CAM.unpack_from(self.map, HDR.size + i * CAM.size)) for i in range(ncams)]
|
||||
|
||||
def heartbeat_ns(self):
|
||||
return struct.unpack_from('<Q', self.map, HEARTBEAT_OFF)[0]
|
||||
|
||||
def alive(self, max_age=1.0):
|
||||
hb = self.heartbeat_ns()
|
||||
return hb != 0 and (time.clock_gettime_ns(time.CLOCK_MONOTONIC) - hb) / 1e9 < max_age
|
||||
|
||||
def latest(self, cam):
|
||||
return struct.unpack_from('<Q', self.map, cam.latest_off)[0]
|
||||
|
||||
def read(self, cam, n=None):
|
||||
"""Copy frame n (default: the newest) of a camera, or None if it's gone or being written."""
|
||||
for _ in range(3):
|
||||
if n is None or n == 0:
|
||||
n = self.latest(cam)
|
||||
if n == 0:
|
||||
return None
|
||||
off = cam.slot_offset + (n % cam.nslots) * cam.slot_bytes
|
||||
fields = SLOT.unpack_from(self.map, off)
|
||||
if fields[0] != 2 * n + 2:
|
||||
return None
|
||||
start = off + SLOT.size
|
||||
image = np.frombuffer(self.map, np.uint8, cam.stride * cam.height, start).reshape(cam.height, cam.stride)
|
||||
image = image[:, :cam.width].copy()
|
||||
if struct.unpack_from('<Q', self.map, off)[0] == fields[0]:
|
||||
return Frame(cam, fields, image)
|
||||
n = None # overwritten while copying: take the newest
|
||||
return None
|
||||
@@ -0,0 +1,129 @@
|
||||
"""Draw frame sets from a recording (ft-hands --record) with what the tracker saw.
|
||||
|
||||
usage: python tools/show_set.py REC_DIR SET [SET...] [--timeline TL] [--out DIR]
|
||||
|
||||
SET is a set index (ft-handreplay's timeline gives them). With --timeline (ft-handreplay
|
||||
--timeline), each camera shows the tracker's views at that set: the crop for the next
|
||||
frame, labelled with the hand and presence. Recordings made with ft-camd --with-dark get
|
||||
a second row: each camera's latest dark frame (<name>_dk), stretched to be visible and
|
||||
labelled with its mean brightness. Recordings made with ft-camd --with-color get a row of
|
||||
the color cameras (color_video<N>). Writes OUT/set_<n>.jpg (default /tmp).
|
||||
"""
|
||||
import argparse
|
||||
import os
|
||||
import struct
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
HDR = struct.Struct('<8sII')
|
||||
CAM = struct.Struct('<16sIIQQ')
|
||||
ORDER = ['slam_left', 'slam_right', 'upper_left', 'upper_right']
|
||||
|
||||
|
||||
def index(path):
|
||||
"""Byte offset of every set in sets.bin."""
|
||||
offs, size = [], os.path.getsize(path)
|
||||
with open(path, 'rb') as f:
|
||||
off = 0
|
||||
while off + HDR.size <= size:
|
||||
f.seek(off)
|
||||
magic, n, nbytes = HDR.unpack(f.read(HDR.size))
|
||||
if magic[:7] != b'FHSET01' or off + nbytes > size:
|
||||
break
|
||||
offs.append(off)
|
||||
off += nbytes
|
||||
return offs
|
||||
|
||||
|
||||
def read_set(path, off):
|
||||
with open(path, 'rb') as f:
|
||||
f.seek(off)
|
||||
_, n, _ = HDR.unpack(f.read(HDR.size))
|
||||
cams = [CAM.unpack(f.read(CAM.size)) for _ in range(n)]
|
||||
out = {}
|
||||
for name, w, h, cap, dq in cams:
|
||||
px = np.frombuffer(f.read(w * h), np.uint8).reshape(h, w)
|
||||
out[name.rstrip(b'\0').decode()] = (px, cap)
|
||||
return out
|
||||
|
||||
|
||||
def views_at(timeline, n):
|
||||
out = []
|
||||
for line in open(timeline):
|
||||
f = line.split()
|
||||
if len(f) > 1 and f[1] == 'view' and int(f[-1]) == n:
|
||||
out.append({'hand': int(f[2]), 'cam': f[3], 'presence': float(f[5]),
|
||||
'c': (float(f[7]), float(f[8])), 'size': float(f[9]), 'rot': float(f[10])})
|
||||
return out
|
||||
|
||||
|
||||
def dark_tile(frame, shape, name):
|
||||
"""A dark frame, stretched from its 1st to 99.5th percentile; black if there's none."""
|
||||
h, w = shape
|
||||
if frame is None:
|
||||
return np.zeros((h, w, 3), np.uint8)
|
||||
px = frame[0]
|
||||
lo, hi = np.percentile(px, (1, 99.5))
|
||||
gain = 255 / max(hi - lo, 1)
|
||||
img = np.clip((px.astype(np.float32) - lo) * gain, 0, 255).astype(np.uint8)
|
||||
img = cv2.cvtColor(cv2.resize(img, (w, h)), cv2.COLOR_GRAY2BGR)
|
||||
cv2.putText(img, '%s_dk mean %.1f, x%.0f' % (name, px.mean(), gain), (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0,
|
||||
(255, 255, 0), 2)
|
||||
return img
|
||||
|
||||
|
||||
def view_tile(px, name, views):
|
||||
"""A frame, CLAHE'd, with the tracker's views on it, 512 px high."""
|
||||
img = cv2.cvtColor(cv2.createCLAHE(2.0, (8, 8)).apply(px), cv2.COLOR_GRAY2BGR)
|
||||
for v in views:
|
||||
if v['cam'] != name:
|
||||
continue
|
||||
c, s, r = v['c'], v['size'], v['rot']
|
||||
box = cv2.boxPoints(((c[0], c[1]), (s, s), np.degrees(r)))
|
||||
col = (0, 255, 0) if v['presence'] >= 0.5 else (0, 0, 255)
|
||||
cv2.polylines(img, [box.astype(np.int32)], True, col, 2)
|
||||
cv2.putText(img, 'h%d %.2f' % (v['hand'], v['presence']), (int(c[0] - s / 2), int(c[1] - s / 2) - 6),
|
||||
cv2.FONT_HERSHEY_SIMPLEX, 0.8, col, 2)
|
||||
cv2.putText(img, name, (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0, (255, 255, 0), 2)
|
||||
scale = 512 / img.shape[0]
|
||||
return cv2.resize(img, (int(img.shape[1] * scale), 512))
|
||||
|
||||
|
||||
def draw(images, views):
|
||||
"""Rows: the mono cameras; their dark frames, if recorded; the color cameras, if recorded."""
|
||||
tiles, dark = [], []
|
||||
for name in ORDER:
|
||||
if name not in images:
|
||||
continue
|
||||
tiles.append(view_tile(images[name][0], name, views))
|
||||
dark.append(dark_tile(images.get(name + '_dk'), tiles[-1].shape[:2], name))
|
||||
rows = [np.hstack(tiles)]
|
||||
if any(k.endswith('_dk') for k in images):
|
||||
rows.append(np.hstack(dark))
|
||||
color = sorted(k for k in images if k.startswith('color_'))
|
||||
if color:
|
||||
rows.append(np.hstack([view_tile(images[k][0], k, views) for k in color]))
|
||||
width = max(r.shape[1] for r in rows)
|
||||
return np.vstack([np.pad(r, ((0, 0), (0, width - r.shape[1]), (0, 0))) for r in rows])
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('rec')
|
||||
ap.add_argument('sets', type=int, nargs='+')
|
||||
ap.add_argument('--timeline')
|
||||
ap.add_argument('--out', default='/tmp')
|
||||
a = ap.parse_args()
|
||||
path = os.path.join(a.rec, 'sets.bin')
|
||||
offs = index(path)
|
||||
for n in a.sets:
|
||||
images = read_set(path, offs[n])
|
||||
views = views_at(a.timeline, n) if a.timeline else []
|
||||
out = os.path.join(a.out, 'set_%05d.jpg' % n)
|
||||
cv2.imwrite(out, draw(images, views), [cv2.IMWRITE_JPEG_QUALITY, 85])
|
||||
print(out)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,84 @@
|
||||
"""Watch the pinch gestures ft-hands publishes, live: begins, ends, and drags.
|
||||
|
||||
usage: python3 tools/watch_gestures.py [--every S] [--distance]
|
||||
|
||||
Prints a line when a pinch begins or ends on either hand. It goes by the counters, so a
|
||||
quick tap between two reads still shows. While a pinch is held, every --every seconds
|
||||
(default 0.1) it prints how far the pinch point has moved since it began, in the head
|
||||
frame (turning your head moves it too; a real consumer turns both points into the room
|
||||
first, see include/fh_gestures.h). --distance also prints each hand's thumb-to-index
|
||||
distance, to see how close a pinch comes to the thresholds.
|
||||
"""
|
||||
import argparse
|
||||
import mmap
|
||||
import os
|
||||
import struct
|
||||
import time
|
||||
|
||||
HDR = struct.Struct('<8sIIQQQff16x') # 64 bytes
|
||||
PINCH = struct.Struct('<IIIIQQff3f3f') # 64 bytes
|
||||
TRACKED, DOWN, LOST = 1, 2, 4
|
||||
SIDES = ('left ', 'right')
|
||||
|
||||
|
||||
def path():
|
||||
return '/run/user/%d/frametop-hands/gestures' % os.getuid()
|
||||
|
||||
|
||||
def read(m):
|
||||
"""(header, [left, right]) under the sequence lock, or None if it's being written."""
|
||||
for _ in range(10):
|
||||
s1 = struct.unpack_from('<Q', m, 16)[0]
|
||||
if s1 % 2 == 0:
|
||||
h = HDR.unpack_from(m, 0)
|
||||
p = [PINCH.unpack_from(m, HDR.size + k * PINCH.size) for k in range(2)]
|
||||
if struct.unpack_from('<Q', m, 16)[0] == s1:
|
||||
return h, p
|
||||
time.sleep(0.0005)
|
||||
return None
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('--every', type=float, default=0.1, help='seconds between drag lines while pinching')
|
||||
ap.add_argument('--distance', action='store_true', help="print each hand's thumb-to-index distance")
|
||||
a = ap.parse_args()
|
||||
with open(path(), 'rb') as f:
|
||||
m = mmap.mmap(f.fileno(), 0, prot=mmap.PROT_READ)
|
||||
first = read(m)
|
||||
if first is None or first[0][0] != b'FHGEST01':
|
||||
raise SystemExit('%s is not an ft-hands gestures file' % path())
|
||||
h, p = first
|
||||
print('thresholds: pinch begins under %.3f m, ends over %.3f m' % (h[6], h[7]))
|
||||
seen = [(q[2], q[3]) for q in p] # begins, ends
|
||||
last_drag = last_dist = 0.0
|
||||
while True:
|
||||
got = read(m)
|
||||
if got:
|
||||
h, p = got
|
||||
now = time.monotonic()
|
||||
for k, q in enumerate(p):
|
||||
flags, hand, begins, ends, begin_ns, end_ns, dist, strength = q[:8]
|
||||
point, begin_point = q[8:11], q[11:14]
|
||||
if begins != seen[k][0]:
|
||||
print('%s pinch BEGIN (#%d, hand %d) at %+.3f %+.3f %+.3f d %.3f' %
|
||||
(SIDES[k], begins, hand, *begin_point, dist), flush=True)
|
||||
if ends != seen[k][1]:
|
||||
held = (end_ns - begin_ns) / 1e9 if end_ns >= begin_ns else 0
|
||||
print('%s pinch %s after %.2f s' % (SIDES[k], 'LOST' if flags & LOST else 'END', held), flush=True)
|
||||
seen[k] = (begins, ends)
|
||||
if flags & DOWN and now - last_drag >= a.every:
|
||||
d = [point[i] - begin_point[i] for i in range(3)]
|
||||
print('%s drag %+6.1f %+6.1f %+6.1f mm (%.0f mm)' %
|
||||
(SIDES[k], *(1000 * x for x in d), 1000 * sum(x * x for x in d) ** 0.5), flush=True)
|
||||
if any(q[0] & DOWN for q in p) and now - last_drag >= a.every:
|
||||
last_drag = now
|
||||
if a.distance and now - last_dist >= 0.2:
|
||||
last_dist = now
|
||||
print(' ' + ' '.join('%s %s' % (SIDES[k].strip(), 'd %.3f s %.2f' % (q[6], q[7]) if q[0] & TRACKED
|
||||
else '-') for k, q in enumerate(p)), flush=True)
|
||||
time.sleep(0.005)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,221 @@
|
||||
#include <cstdlib>
|
||||
#include "calib.h"
|
||||
|
||||
#include <json/json.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <algorithm>
|
||||
|
||||
#include <fstream>
|
||||
#include <iterator>
|
||||
#include <memory>
|
||||
|
||||
namespace {
|
||||
|
||||
double theta_d(const Camera &c, double t) {
|
||||
const double t2 = t * t;
|
||||
return t * (1 + t2 * (c.k[0] + t2 * (c.k[1] + t2 * (c.k[2] + t2 * c.k[3]))));
|
||||
}
|
||||
|
||||
// 4x4 transform (row-major) from a {plus_x, plus_z, position} pose.
|
||||
void pose(const Json::Value &d, double scale, double T[4][4]) {
|
||||
V3 x{d["plus_x"][0].asDouble(), d["plus_x"][1].asDouble(), d["plus_x"][2].asDouble()};
|
||||
V3 z{d["plus_z"][0].asDouble(), d["plus_z"][1].asDouble(), d["plus_z"][2].asDouble()};
|
||||
V3 y{z[1] * x[2] - z[2] * x[1], z[2] * x[0] - z[0] * x[2], z[0] * x[1] - z[1] * x[0]};
|
||||
for (int i = 0; i < 3; ++i) {
|
||||
T[i][0] = x[i], T[i][1] = y[i], T[i][2] = z[i];
|
||||
T[i][3] = d["position"][i].asDouble() * scale;
|
||||
T[3][i] = 0;
|
||||
}
|
||||
T[3][3] = 1;
|
||||
}
|
||||
|
||||
void mul(const double A[4][4], const double B[4][4], double C[4][4]) {
|
||||
for (int i = 0; i < 4; ++i)
|
||||
for (int j = 0; j < 4; ++j) {
|
||||
C[i][j] = 0;
|
||||
for (int k = 0; k < 4; ++k) C[i][j] += A[i][k] * B[k][j];
|
||||
}
|
||||
}
|
||||
|
||||
void invert_rigid(const double A[4][4], double B[4][4]) {
|
||||
for (int i = 0; i < 3; ++i)
|
||||
for (int j = 0; j < 3; ++j) B[i][j] = A[j][i];
|
||||
for (int i = 0; i < 3; ++i) B[i][3] = -(B[i][0] * A[0][3] + B[i][1] * A[1][3] + B[i][2] * A[2][3]);
|
||||
B[3][0] = B[3][1] = B[3][2] = 0, B[3][3] = 1;
|
||||
}
|
||||
|
||||
bool read_json(const char *path, Json::Value &v, std::string &err) {
|
||||
std::ifstream f(path);
|
||||
Json::CharReaderBuilder b;
|
||||
std::string e;
|
||||
if (!f || !Json::parseFromStream(b, f, &v, &e)) {
|
||||
err = std::string(path) + ": " + (f ? e : "can't open");
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
V2 Camera::project_cam(V3 p) const {
|
||||
const double r = std::hypot(p[0], p[1]);
|
||||
const double s = r > 1e-12 ? theta_d(*this, std::atan2(r, p[2])) / r : 0;
|
||||
return {fx * p[0] * s + cx, fy * p[1] * s + cy};
|
||||
}
|
||||
|
||||
V3 Camera::unproject(V2 uv) const {
|
||||
const double mx = (uv[0] - cx) / fx, my = (uv[1] - cy) / fy, td = std::hypot(mx, my);
|
||||
double t = td;
|
||||
for (int i = 0; i < 8; ++i) { // Newton on theta_d(t) = td
|
||||
const double t2 = t * t;
|
||||
const double df = 1 + t2 * (3 * k[0] + t2 * (5 * k[1] + t2 * (7 * k[2] + t2 * 9 * k[3])));
|
||||
t = std::clamp(t - (theta_d(*this, t) - td) / df, 0.0, M_PI);
|
||||
}
|
||||
const double s = td > 1e-12 ? std::sin(t) / td : 1;
|
||||
return {mx * s, my * s, std::cos(t)};
|
||||
}
|
||||
|
||||
V3 Camera::ray(V2 uv) const {
|
||||
const V3 c = unproject(uv);
|
||||
return {R[0][0] * c[0] + R[0][1] * c[1] + R[0][2] * c[2], R[1][0] * c[0] + R[1][1] * c[1] + R[1][2] * c[2],
|
||||
R[2][0] * c[0] + R[2][1] * c[1] + R[2][2] * c[2]};
|
||||
}
|
||||
|
||||
V2 Camera::project(V3 head, double *depth) const {
|
||||
const V3 d = head - origin;
|
||||
const V3 c{R[0][0] * d[0] + R[1][0] * d[1] + R[2][0] * d[2], R[0][1] * d[0] + R[1][1] * d[1] + R[2][1] * d[2],
|
||||
R[0][2] * d[0] + R[1][2] * d[1] + R[2][2] * d[2]};
|
||||
if (depth) *depth = c[2];
|
||||
return project_cam(c);
|
||||
}
|
||||
|
||||
double Camera::off_axis(V2 uv) const { return std::acos(std::clamp(unproject(uv)[2], -1.0, 1.0)) * 180 / M_PI; }
|
||||
|
||||
// A headset file such as /persist/xrservice.json. In the dev container the host's / is at
|
||||
// /run/host (distrobox doesn't mount /persist); off the Frame, FRAME_JOB_DEVICE_ROOT can
|
||||
// point at a folder with copies of them.
|
||||
static std::string device_path(const char *path) {
|
||||
if (const char *root = std::getenv("FRAME_JOB_DEVICE_ROOT")) return std::string(root) + path;
|
||||
const std::string host = std::string("/run/host") + path;
|
||||
return access(path, R_OK) != 0 && access(host.c_str(), R_OK) == 0 ? host : path;
|
||||
}
|
||||
|
||||
bool load_calibration(std::map<std::string, Camera> &out, std::string &err) {
|
||||
Json::Value rig, dev;
|
||||
if (!read_json(device_path("/persist/xrservice.json").c_str(), rig, err) ||
|
||||
!read_json(device_path("/persist/device_config.json").c_str(), dev, err))
|
||||
return false;
|
||||
double cad_from_cam0[4][4], cad_from_head[4][4], head_from_cad[4][4], head_from_cam0[4][4];
|
||||
pose(dev["cv"]["cad_from_cal"], 1.0, cad_from_cam0);
|
||||
pose(dev["head"], 1.0, cad_from_head);
|
||||
invert_rigid(cad_from_head, head_from_cad);
|
||||
mul(head_from_cad, cad_from_cam0, head_from_cam0);
|
||||
for (const Json::Value &c : rig["cameras"]) {
|
||||
Camera cam;
|
||||
cam.name = c["sourceCamera"].asString();
|
||||
cam.width = c["width"].asInt(), cam.height = c["height"].asInt();
|
||||
for (const Json::Value &in : c["intrinsics"]) {
|
||||
if (in["cameraModel"].asString() != "kb") continue;
|
||||
cam.fx = in["fx"].asDouble(), cam.fy = in["fy"].asDouble();
|
||||
cam.cx = in["cx"].asDouble(), cam.cy = in["cy"].asDouble();
|
||||
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
|
||||
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
|
||||
}
|
||||
double cam0_from_cam[4][4], head_from_cam[4][4];
|
||||
pose(c["extrinsics"], 1e-3, cam0_from_cam);
|
||||
mul(head_from_cam0, cam0_from_cam, head_from_cam);
|
||||
for (int i = 0; i < 3; ++i) {
|
||||
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
|
||||
cam.origin[i] = head_from_cam[i][3];
|
||||
}
|
||||
out[cam.name] = cam;
|
||||
}
|
||||
if (out.empty()) err = "no cameras in /persist/xrservice.json";
|
||||
return !out.empty();
|
||||
}
|
||||
|
||||
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
|
||||
const std::string &right_node, bool crop_subtract, int scale, std::string &err) {
|
||||
// The module's EEPROM: some binary, then the calibration as JSON (world-readable)
|
||||
const std::string path = device_path("/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom");
|
||||
std::ifstream f(path, std::ios::binary);
|
||||
const std::string raw((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
||||
const size_t key = raw.find("\"alignment_method\"");
|
||||
const size_t start = key == std::string::npos ? key : raw.rfind('{', key);
|
||||
Json::Value rig, dev;
|
||||
std::string e;
|
||||
std::unique_ptr<Json::CharReader> reader(Json::CharReaderBuilder().newCharReader());
|
||||
if (start == std::string::npos || !reader->parse(raw.data() + start, raw.data() + raw.size(), &rig, &e))
|
||||
return err = path + ": no calibration JSON " + e, false;
|
||||
if (!read_json(device_path("/persist/device_config.json").c_str(), dev, err)) return false;
|
||||
double cad_from_head[4][4], head_from_cad[4][4];
|
||||
pose(dev["head"], 1.0, cad_from_head);
|
||||
invert_rigid(cad_from_head, head_from_cad);
|
||||
constexpr int kValidWidth = 1972; // pixels per row XRService's buffers deliver (of 2464)
|
||||
int n = 0;
|
||||
for (const Json::Value &c : rig["cameras"]) {
|
||||
const std::string source = c["sourceCamera"].asString();
|
||||
const std::string name = source == "passthrough_left" ? left_node : source == "passthrough_right" ? right_node : "";
|
||||
if (name.empty()) continue;
|
||||
Camera cam;
|
||||
cam.name = name;
|
||||
cam.width = kValidWidth / scale, cam.height = c["height"].asInt() / scale;
|
||||
const double dx = crop_subtract ? c["cropRegion"]["x"].asDouble() : 0, dy = crop_subtract ? c["cropRegion"]["y"].asDouble() : 0;
|
||||
for (const Json::Value &in : c["intrinsics"]) {
|
||||
if (in["cameraModel"].asString() != "kb") continue;
|
||||
// integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
|
||||
cam.fx = in["fx"].asDouble() / scale, cam.fy = in["fy"].asDouble() / scale;
|
||||
cam.cx = (in["cx"].asDouble() - dx + 0.5) / scale - 0.5, cam.cy = (in["cy"].asDouble() - dy + 0.5) / scale - 0.5;
|
||||
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
|
||||
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
|
||||
}
|
||||
double cad_from_cam[4][4], head_from_cam[4][4];
|
||||
pose(c["extrinsics"], 1e-3, cad_from_cam);
|
||||
mul(head_from_cad, cad_from_cam, head_from_cam);
|
||||
for (int i = 0; i < 3; ++i) {
|
||||
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
|
||||
cam.origin[i] = head_from_cam[i][3];
|
||||
}
|
||||
out[name] = cam;
|
||||
++n;
|
||||
}
|
||||
if (n != 2) err = path + ": expected passthrough_left and passthrough_right";
|
||||
return n == 2;
|
||||
}
|
||||
|
||||
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms) {
|
||||
double A[3][3] = {}, b[3] = {};
|
||||
for (int v = 0; v < n; ++v) {
|
||||
const V3 &d = dirs[v], &o = origins[v];
|
||||
for (int i = 0; i < 3; ++i)
|
||||
for (int j = 0; j < 3; ++j) {
|
||||
const double P = (i == j ? 1.0 : 0.0) - d[i] * d[j];
|
||||
A[i][j] += weights[v] * P;
|
||||
b[i] += weights[v] * P * o[j];
|
||||
}
|
||||
}
|
||||
// Cramer's rule for the 3x3 system
|
||||
auto det3 = [](const double m[3][3]) {
|
||||
return m[0][0] * (m[1][1] * m[2][2] - m[1][2] * m[2][1]) - m[0][1] * (m[1][0] * m[2][2] - m[1][2] * m[2][0]) +
|
||||
m[0][2] * (m[1][0] * m[2][1] - m[1][1] * m[2][0]);
|
||||
};
|
||||
const double D = det3(A);
|
||||
V3 p{};
|
||||
for (int c = 0; c < 3; ++c) {
|
||||
double M[3][3];
|
||||
for (int i = 0; i < 3; ++i)
|
||||
for (int j = 0; j < 3; ++j) M[i][j] = j == c ? b[i] : A[i][j];
|
||||
p[c] = std::fabs(D) > 1e-18 ? det3(M) / D : 0;
|
||||
}
|
||||
if (rms) {
|
||||
double s = 0;
|
||||
for (int v = 0; v < n; ++v) {
|
||||
const V3 off = p - origins[v];
|
||||
const V3 perp = off - dirs[v] * dot(off, dirs[v]);
|
||||
s += dot(perp, perp);
|
||||
}
|
||||
*rms = std::sqrt(s / n);
|
||||
}
|
||||
return p;
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
// Tracking-camera calibration from the headset's factory files (see tools/calib.py
|
||||
// for the conventions): Kannala-Brandt fisheye intrinsics, and each camera's pose in the
|
||||
// head frame (OpenVR's: +x right, +y up, -z forward), metres.
|
||||
#pragma once
|
||||
|
||||
#include "geom.h"
|
||||
|
||||
#include <map>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
struct Camera {
|
||||
std::string name;
|
||||
int width = 0, height = 0;
|
||||
double fx = 1, fy = 1, cx = 0, cy = 0, k[4] = {};
|
||||
double R[3][3] = {}; // camera axes (columns) in the head frame
|
||||
V3 origin{}; // camera centre in the head frame
|
||||
|
||||
V2 project_cam(V3 p) const; // camera frame -> pixels
|
||||
V3 unproject(V2 uv) const; // pixels -> unit ray, camera frame
|
||||
V3 ray(V2 uv) const; // pixels -> unit ray, head frame
|
||||
V2 project(V3 head, double *depth) const; // head frame -> pixels; depth along the optical axis
|
||||
double off_axis(V2 uv) const; // degrees between the pixel's ray and the axis
|
||||
};
|
||||
|
||||
// Loads /persist/xrservice.json and /persist/device_config.json. Keyed by calibration
|
||||
// name: slam_left, slam_right, upper_left, upper_right.
|
||||
bool load_calibration(std::map<std::string, Camera> &out, std::string &err);
|
||||
|
||||
// The Arcturus color cameras (tools/calib.py load_color has the conventions), for
|
||||
// ft-camd --with-color's images: luma at 1/scale size, recorded as color_video<N>. They're
|
||||
// keyed by those recorded names: left_node is passthrough_left, right_node
|
||||
// passthrough_right. crop_subtract: image x = sensor x - the calibration's cropRegion.x.
|
||||
// tools/check_color.py tells which node is which and which crop reading fits.
|
||||
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
|
||||
const std::string &right_node, bool crop_subtract, int scale, std::string &err);
|
||||
|
||||
// The point closest to several rays (weighted), and its rms distance to them.
|
||||
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms);
|
||||
@@ -0,0 +1,22 @@
|
||||
// Small vector helpers for the tracker.
|
||||
#pragma once
|
||||
|
||||
#include <array>
|
||||
#include <cmath>
|
||||
|
||||
using V2 = std::array<double, 2>;
|
||||
using V3 = std::array<double, 3>;
|
||||
|
||||
inline V3 operator+(V3 a, V3 b) { return {a[0] + b[0], a[1] + b[1], a[2] + b[2]}; }
|
||||
inline V3 operator-(V3 a, V3 b) { return {a[0] - b[0], a[1] - b[1], a[2] - b[2]}; }
|
||||
inline V3 operator*(V3 a, double s) { return {a[0] * s, a[1] * s, a[2] * s}; }
|
||||
inline double dot(V3 a, V3 b) { return a[0] * b[0] + a[1] * b[1] + a[2] * b[2]; }
|
||||
inline double norm(V3 a) { return std::sqrt(dot(a, a)); }
|
||||
inline V3 unit(V3 a) { double n = norm(a); return n > 0 ? a * (1 / n) : a; }
|
||||
|
||||
inline V2 operator+(V2 a, V2 b) { return {a[0] + b[0], a[1] + b[1]}; }
|
||||
inline V2 operator-(V2 a, V2 b) { return {a[0] - b[0], a[1] - b[1]}; }
|
||||
inline V2 operator*(V2 a, double s) { return {a[0] * s, a[1] * s}; }
|
||||
inline double norm(V2 a) { return std::hypot(a[0], a[1]); }
|
||||
|
||||
inline double wrap_angle(double a) { return std::remainder(a, 2 * M_PI); }
|
||||
@@ -0,0 +1,171 @@
|
||||
#include "io.h"
|
||||
|
||||
#include <fcntl.h>
|
||||
#include <sys/mman.h>
|
||||
#include <sys/stat.h>
|
||||
#include <time.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
|
||||
uint64_t mono_ns() {
|
||||
timespec ts;
|
||||
clock_gettime(CLOCK_MONOTONIC, &ts);
|
||||
return uint64_t(ts.tv_sec) * 1'000'000'000 + uint64_t(ts.tv_nsec);
|
||||
}
|
||||
|
||||
int64_t raw_minus_mono_ns() {
|
||||
timespec a, r, b;
|
||||
clock_gettime(CLOCK_MONOTONIC, &a);
|
||||
clock_gettime(CLOCK_MONOTONIC_RAW, &r);
|
||||
clock_gettime(CLOCK_MONOTONIC, &b);
|
||||
const int64_t ma = int64_t(a.tv_sec) * 1'000'000'000 + a.tv_nsec, mb = int64_t(b.tv_sec) * 1'000'000'000 + b.tv_nsec;
|
||||
return int64_t(r.tv_sec) * 1'000'000'000 + r.tv_nsec - (ma + mb) / 2;
|
||||
}
|
||||
|
||||
// ------------------------------------------------------------------------------ ring
|
||||
|
||||
bool Ring::open(const char *path, std::string &err) {
|
||||
const int fd = ::open(path, O_RDONLY | O_CLOEXEC);
|
||||
if (fd < 0) return err = std::string(path) + ": " + std::strerror(errno), false;
|
||||
struct stat st;
|
||||
fstat(fd, &st);
|
||||
len_ = size_t(st.st_size);
|
||||
void *m = len_ >= sizeof(fh_ring_hdr_t) ? mmap(nullptr, len_, PROT_READ, MAP_SHARED, fd, 0) : MAP_FAILED;
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return err = std::string(path) + ": can't map it", false;
|
||||
map_ = static_cast<const uint8_t *>(m);
|
||||
hdr_ = reinterpret_cast<const fh_ring_hdr_t *>(map_);
|
||||
if (std::memcmp(hdr_->magic, FH_RING_MAGIC, 8) || hdr_->version != FH_RING_VERSION || hdr_->file_bytes > len_)
|
||||
return err = std::string(path) + " is not an ft-camd ring", false;
|
||||
return true;
|
||||
}
|
||||
|
||||
bool Ring::alive() const {
|
||||
const uint64_t hb = __atomic_load_n(&hdr_->heartbeat_ns, __ATOMIC_ACQUIRE);
|
||||
return hb && mono_ns() - hb < 1'000'000'000;
|
||||
}
|
||||
|
||||
uint64_t Ring::latest(int i) const { return __atomic_load_n(&hdr_->cams[i].latest, __ATOMIC_ACQUIRE); }
|
||||
|
||||
bool Ring::read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const {
|
||||
const fh_ring_cam_t &c = hdr_->cams[i];
|
||||
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
|
||||
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
|
||||
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
|
||||
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
|
||||
if (seq != 2 * n + 2) return false;
|
||||
std::memcpy(meta, slot, sizeof *meta);
|
||||
out.resize(size_t(c.width) * c.height);
|
||||
for (uint32_t y = 0; y < c.height; ++y)
|
||||
std::memcpy(out.data() + size_t(y) * c.width, slot + sizeof(fh_ring_slot_t) + size_t(y) * c.stride, c.width);
|
||||
__atomic_thread_fence(__ATOMIC_ACQUIRE);
|
||||
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
|
||||
}
|
||||
|
||||
bool Ring::meta(int i, uint64_t n, fh_ring_slot_t *meta) const {
|
||||
const fh_ring_cam_t &c = hdr_->cams[i];
|
||||
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
|
||||
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
|
||||
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
|
||||
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
|
||||
if (seq != 2 * n + 2) return false;
|
||||
std::memcpy(meta, slot, sizeof *meta);
|
||||
__atomic_thread_fence(__ATOMIC_ACQUIRE);
|
||||
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
|
||||
}
|
||||
|
||||
// ------------------------------------------------------------------------- publisher
|
||||
|
||||
namespace {
|
||||
|
||||
// The hand's shape to cut out, as capsules.
|
||||
// Radii are a real hand's half-widths plus a small margin for tracking noise.
|
||||
const int kThumb[][2] = {{0, 1}, {1, 2}, {2, 3}, {3, 4}};
|
||||
const int kFingers[][2] = {{5, 6}, {6, 7}, {7, 8}, {9, 10}, {10, 11}, {11, 12}, {13, 14}, {14, 15}, {15, 16},
|
||||
{17, 18}, {18, 19}, {19, 20}};
|
||||
const int kPalm[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 9}, {9, 13}, {13, 17}, {1, 5}};
|
||||
constexpr double kThumbR = 0.0095, kFinger = 0.0085, kPalmR = 0.015, kArm[2] = {0.028, 0.034}, kArmLen = 0.16,
|
||||
kMargin = 0.004;
|
||||
// Nothing is cut closer than this in front of the eyes (head frame, -z is forward). A
|
||||
// point near the eyes' plane lands far across a screen with a huge radius, so one bad
|
||||
// estimate there tears a hole through it; real hands that close aren't tracked anyway.
|
||||
constexpr double kNear = 0.12;
|
||||
|
||||
// Adds the capsule, clipped to the part at least kNear in front of the eyes.
|
||||
void put(fh_capsule_t *caps, uint32_t &n, V3 a, V3 b, double ra, double rb) {
|
||||
if (n >= FH_HANDS_MAX_CAPSULES) return;
|
||||
const double za = -a[2] - kNear, zb = -b[2] - kNear; // >= 0: far enough in front
|
||||
if (za < 0 && zb < 0) return;
|
||||
if (za < 0 || zb < 0) {
|
||||
const double t = za / (za - zb); // where the segment crosses the near plane
|
||||
const V3 m = a + (b - a) * t;
|
||||
const double rm = ra + (rb - ra) * t;
|
||||
if (za < 0) a = m, ra = rm;
|
||||
else b = m, rb = rm;
|
||||
}
|
||||
fh_capsule_t &c = caps[n++];
|
||||
for (int k = 0; k < 3; ++k) c.a[k] = float(a[k]), c.b[k] = float(b[k]);
|
||||
c.ra = float(ra + kMargin), c.rb = float(rb + kMargin);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
std::string run_dir() {
|
||||
const std::string dir = "/run/user/" + std::to_string(getuid()) + "/frametop-hands";
|
||||
mkdir(dir.c_str(), 0700);
|
||||
return dir;
|
||||
}
|
||||
|
||||
bool Publisher::open(std::string &err) {
|
||||
const std::string path = run_dir() + "/hands";
|
||||
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
|
||||
if (fd < 0 || ftruncate(fd, sizeof(fh_hands_t)) < 0) return err = path + ": " + std::strerror(errno), false;
|
||||
void *m = mmap(nullptr, sizeof(fh_hands_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return err = path + ": can't map it", false;
|
||||
out_ = static_cast<fh_hands_t *>(m);
|
||||
std::memset(out_, 0, sizeof *out_);
|
||||
std::memcpy(out_->magic, FH_HANDS_MAGIC, 8);
|
||||
out_->version = FH_HANDS_VERSION;
|
||||
out_->size = sizeof(fh_hands_t);
|
||||
return true;
|
||||
}
|
||||
|
||||
void Publisher::write(const std::vector<const Hand *> &in, uint64_t capture_ns) {
|
||||
std::vector<const Hand *> hands = in;
|
||||
std::sort(hands.begin(), hands.end(), [](const Hand *a, const Hand *b) { return a->frames > b->frames; });
|
||||
if (hands.size() > FH_HANDS_MAX_HANDS) hands.resize(FH_HANDS_MAX_HANDS);
|
||||
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
uint32_t nc = 0;
|
||||
for (size_t k = 0; k < FH_HANDS_MAX_HANDS; ++k) {
|
||||
fh_hand_t &o = out_->hands[k];
|
||||
std::memset(&o, 0, sizeof o);
|
||||
if (k >= hands.size()) continue;
|
||||
const Hand &h = *hands[k];
|
||||
o.id = uint32_t(h.id);
|
||||
o.flags = (h.right() ? FH_HAND_RIGHT : 0) | (h.nviews >= 2 ? FH_HAND_STEREO : 0);
|
||||
o.confidence = float(std::min(1.0, h.frames / 5.0));
|
||||
for (int i = 0; i < 21; ++i)
|
||||
for (int j = 0; j < 3; ++j) o.pts[i][j] = float(h.smooth[i][j]);
|
||||
const uint32_t first = nc;
|
||||
for (auto &b : kThumb) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kThumbR, kThumbR);
|
||||
for (auto &b : kFingers) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kFinger, kFinger);
|
||||
for (auto &b : kPalm) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kPalmR, kPalmR);
|
||||
// the forearm carries on from the hand's own axis (middle knuckle -> wrist); the
|
||||
// wrist bends, but much less than a guess at where the elbow is gets wrong
|
||||
const V3 wrist = h.smooth[0], d = wrist - h.smooth[9];
|
||||
const double n = norm(d);
|
||||
if (n > 0.02) put(out_->capsules, nc, wrist, wrist + d * (kArmLen / n), kArm[0], kArm[1]);
|
||||
o.ncapsules = nc - first;
|
||||
}
|
||||
for (uint32_t k = nc; k < FH_HANDS_MAX_CAPSULES; ++k) std::memset(&out_->capsules[k], 0, sizeof(fh_capsule_t));
|
||||
out_->capture_ns = capture_ns;
|
||||
out_->publish_ns = mono_ns();
|
||||
out_->nhands = uint32_t(hands.size());
|
||||
out_->ncapsules = nc;
|
||||
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||
}
|
||||
@@ -0,0 +1,52 @@
|
||||
// Frames in from ft-camd's ring (camd/fhring.h), hands out to the hands file
|
||||
// (include/fh_hands.h, read by Frametop's ft-screens).
|
||||
#pragma once
|
||||
|
||||
#include "tracker.h"
|
||||
|
||||
#include <cstdint>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
extern "C" {
|
||||
#include "../camd/fhring.h"
|
||||
#include "../include/fh_hands.h"
|
||||
}
|
||||
|
||||
class Ring {
|
||||
public:
|
||||
bool open(const char *path, std::string &err);
|
||||
bool alive() const; // the writer's heartbeat is fresh
|
||||
int cameras() const { return int(hdr_->ncams); }
|
||||
const fh_ring_cam_t &camera(int i) const { return hdr_->cams[i]; }
|
||||
uint64_t latest(int i) const;
|
||||
// Copy frame n of camera i into out (width x height, tightly packed). False if it's
|
||||
// gone or was being written.
|
||||
bool read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const;
|
||||
// Just frame n's slot header (capture time etc.), without copying the image.
|
||||
bool meta(int i, uint64_t n, fh_ring_slot_t *meta) const;
|
||||
|
||||
private:
|
||||
const uint8_t *map_ = nullptr;
|
||||
const fh_ring_hdr_t *hdr_ = nullptr;
|
||||
size_t len_ = 0;
|
||||
};
|
||||
|
||||
class Publisher {
|
||||
public:
|
||||
bool open(std::string &err);
|
||||
void write(const std::vector<const Hand *> &hands, uint64_t capture_ns);
|
||||
|
||||
private:
|
||||
fh_hands_t *out_ = nullptr;
|
||||
uint64_t seq_ = 0;
|
||||
};
|
||||
|
||||
uint64_t mono_ns();
|
||||
int64_t raw_minus_mono_ns(); // camera timestamps are CLOCK_MONOTONIC_RAW
|
||||
|
||||
// /run/user/UID/frametop-hands, created private to the user if it's missing: where ft-camd's ring
|
||||
// (FH_RING_NAME) and the hands and gestures files live. Not $XDG_RUNTIME_DIR: a terminal in
|
||||
// the Frametop desktop has a private one of its own. And not /run/user/UID/frametop: that is
|
||||
// the desktop session's private runtime folder, which it deletes whenever it starts.
|
||||
std::string run_dir();
|
||||
@@ -0,0 +1,372 @@
|
||||
// ft-hands: hands in 3D from ft-camd's ring, published for Frametop's ft-screens (the hand
|
||||
// cutouts), and pinches for the pointer. It started as a port of frame-hands' Python
|
||||
// prototype: the same scheduling, with the models on a few threads.
|
||||
//
|
||||
// ft-hands [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N]
|
||||
// [--no-publish] [--record DIR] [--swap-sides] ... (--help lists them all)
|
||||
//
|
||||
// Settings in ~/.config/frametop.conf (FT_<name> in the environment overrides them, and
|
||||
// options override both): HANDS_SWAP_SIDES (1: as --swap-sides), HANDS_CPUS (as --cpus).
|
||||
#include "io.h"
|
||||
#include "pinch.h"
|
||||
#include "record.h"
|
||||
|
||||
#include <sched.h>
|
||||
#include <sys/resource.h>
|
||||
#include <sys/stat.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <ctime>
|
||||
#include <memory>
|
||||
|
||||
#include <csignal>
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
#include <fstream>
|
||||
#include <thread>
|
||||
#include <utility>
|
||||
|
||||
namespace {
|
||||
|
||||
volatile std::sig_atomic_t g_stop = 0, g_record = 0;
|
||||
|
||||
// which calibrated camera each capture pipe carries (XRService's fixed routing)
|
||||
const char *camera_for_pipe(int node) {
|
||||
char path[64], name[64] = "";
|
||||
std::snprintf(path, sizeof path, "/sys/class/video4linux/video%d/name", node);
|
||||
std::ifstream f(path);
|
||||
f.getline(name, sizeof name);
|
||||
if (!std::strcmp(name, "msm_vfe3_video0")) return "slam_left";
|
||||
if (!std::strcmp(name, "msm_vfe4_video0")) return "slam_right";
|
||||
if (!std::strcmp(name, "msm_vfe2_video0")) return "upper_left";
|
||||
if (!std::strcmp(name, "msm_vfe2_video1")) return "upper_right";
|
||||
return nullptr;
|
||||
}
|
||||
|
||||
// A setting from ~/.config/frametop.conf, or FT_<key> from the environment; "" if unset.
|
||||
std::string setting(const std::string &key) {
|
||||
if (const char *v = std::getenv(("FT_" + key).c_str())) return v;
|
||||
const char *home = std::getenv("HOME");
|
||||
std::ifstream in(std::string(home ? home : "") + "/.config/frametop.conf");
|
||||
std::string line, value;
|
||||
auto trim = [](std::string s) {
|
||||
s.erase(0, s.find_first_not_of(" \t\"'"));
|
||||
s.erase(s.find_last_not_of(" \t\"'") + 1);
|
||||
return s;
|
||||
};
|
||||
while (std::getline(in, line)) {
|
||||
line = line.substr(0, line.find('#'));
|
||||
const auto eq = line.find('=');
|
||||
if (eq != std::string::npos && trim(line.substr(0, eq)) == key) value = trim(line.substr(eq + 1));
|
||||
}
|
||||
return value;
|
||||
}
|
||||
|
||||
std::vector<int> parse_cpus(const char *p) {
|
||||
std::vector<int> out;
|
||||
while (*p) {
|
||||
char *end;
|
||||
const long c = std::strtol(p, &end, 10);
|
||||
if (end == p) break;
|
||||
out.push_back(int(c));
|
||||
p = *end == ',' ? end + 1 : end;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// Where SIGUSR1 puts recordings: $XDG_DATA_HOME/frametop/hands (~/.local/share/...).
|
||||
std::string recordings_dir() {
|
||||
const char *data = std::getenv("XDG_DATA_HOME"), *home = std::getenv("HOME");
|
||||
std::string dir = data && *data ? data : std::string(home ? home : "") + "/.local/share";
|
||||
for (const char *part : {"/frametop", "/hands"}) mkdir((dir += part).c_str(), 0700);
|
||||
return dir;
|
||||
}
|
||||
|
||||
double cpu_seconds() {
|
||||
rusage r;
|
||||
getrusage(RUSAGE_SELF, &r);
|
||||
return r.ru_utime.tv_sec + r.ru_stime.tv_sec + (r.ru_utime.tv_usec + r.ru_stime.tv_usec) / 1e6;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char **argv) {
|
||||
double seconds = 0, status = 5;
|
||||
int threads = 3, niceness = 5;
|
||||
bool int8 = false, publish = true, track = true, swap_sides = false;
|
||||
std::string models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
|
||||
std::string record, ring_path = "/run/user/" + std::to_string(getuid()) + "/" FH_RING_NAME;
|
||||
// SteamOS starts user processes on CPUs 0-4 and keeps 5-7 (two A720s and the X4) for
|
||||
// SteamVR's compositor, whose threads there run at real-time priority, so they always
|
||||
// win. XRService pins its head tracking to 2-3. frame-hands' probes/core_ab.py
|
||||
// (2026-09-29, headset on, 3 rounds): on 5-7 a step took 8.4 ms against 13.2 on 2-4,
|
||||
// latency 9.6 against 14.1 ms, and the compositor's late frames and CPU/GPU time didn't change.
|
||||
std::vector<int> cpus = {5, 6, 7};
|
||||
if (const auto c = parse_cpus(setting("HANDS_CPUS").c_str()); !c.empty()) cpus = c;
|
||||
swap_sides = setting("HANDS_SWAP_SIDES") == "1";
|
||||
// How crops are equalized. CLAHE helps the palm search find hands (about 10% more in the
|
||||
// dim recording), but makes the landmarks jitter, so they get plain crops.
|
||||
Contrast palm_contrast, hand_contrast{Contrast::None};
|
||||
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
||||
PinchParams pinch_params;
|
||||
double record_for = 120;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
const std::string a = argv[i];
|
||||
const bool more = i + 1 < argc;
|
||||
if (a == "--seconds" && more) seconds = std::atof(argv[++i]);
|
||||
else if (a == "--threads" && more) threads = std::max(1, std::atoi(argv[++i]));
|
||||
else if (a == "--status" && more) status = std::atof(argv[++i]);
|
||||
else if (a == "--models" && more) models = argv[++i];
|
||||
else if (a == "--nice" && more) niceness = std::atoi(argv[++i]);
|
||||
else if (a == "--int8") int8 = true;
|
||||
else if (a == "--no-publish") publish = false;
|
||||
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
|
||||
else if (a == "--pinch-palm-down" && more) pinch_params.palm_down_max = std::atof(argv[++i]);
|
||||
else if (a == "--swap-sides") swap_sides = true;
|
||||
else if (a == "--record-only") track = publish = false;
|
||||
else if (a == "--ring" && more) ring_path = argv[++i];
|
||||
else if (a == "--record" && more) record = argv[++i];
|
||||
else if (a == "--record-for" && more) record_for = std::atof(argv[++i]);
|
||||
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
|
||||
else if (a == "--contrast" && more) {
|
||||
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
|
||||
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
|
||||
} else if (a == "--cpus" && more) {
|
||||
cpus = parse_cpus(argv[++i]);
|
||||
if (cpus.empty()) cpus = {5, 6, 7};
|
||||
}
|
||||
else {
|
||||
std::printf("usage: %s [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] [--no-publish]\n"
|
||||
" [--record DIR] [--record-for S] [--record-only] [--cpus 5,6,7] [--swap-sides]\n"
|
||||
" [--keep-presence P] (0.5) [--ring PATH] (ft-camd's, or ft-ringplay's)\n"
|
||||
" [--pinch-begin M] (0.020) [--pinch-end M] (0.035) [--pinch-triangulated] [--pinch-palm-down MAX] (0.6)\n"
|
||||
" [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch; default clahe:2/none)\n"
|
||||
"Recording saves every frame set for S seconds (120) to DIR/sets.bin, for ft-handreplay; SIGUSR1\n"
|
||||
"starts one in ~/.local/share/frametop/hands/rec-<time>. --record-only records without tracking, so it\n"
|
||||
"can run beside a tracking ft-hands. With ft-camd --with-dark, recordings also get each\n"
|
||||
"camera's newest dark frame, as <name>_dk; with --with-color, the color cameras' as color_video<N>.\n"
|
||||
"Settings in ~/.config/frametop.conf: HANDS_SWAP_SIDES=1, HANDS_CPUS=5,6,7 (FT_<name> overrides).\n",
|
||||
argv[0]);
|
||||
return a == "--help" ? 0 : 1;
|
||||
}
|
||||
}
|
||||
std::setvbuf(stdout, nullptr, _IOLBF, 0); // whole lines to the journal as they come
|
||||
if (nice(niceness) < 0) std::perror("nice"); // the VR stack wins contested CPUs
|
||||
std::signal(SIGINT, [](int) { g_stop = 1; });
|
||||
std::signal(SIGTERM, [](int) { g_stop = 1; });
|
||||
std::signal(SIGUSR1, [](int) { g_record = 1; });
|
||||
|
||||
std::string err;
|
||||
std::map<std::string, Camera> calib;
|
||||
Ring ring;
|
||||
Nets nets;
|
||||
Publisher pub;
|
||||
GesturePublisher gestures;
|
||||
Pinch pinch(pinch_params);
|
||||
std::unique_ptr<Recorder> rec;
|
||||
uint64_t rec_start = 0;
|
||||
auto start_recording = [&](const std::string &dir, std::string &e) {
|
||||
rec = std::make_unique<Recorder>();
|
||||
if (!rec->open(dir, e)) return rec.reset(), false;
|
||||
rec_start = mono_ns();
|
||||
std::printf("recording to %s for %.0f s\n", dir.c_str(), record_for);
|
||||
std::fflush(stdout);
|
||||
return true;
|
||||
};
|
||||
if (!load_calibration(calib, err) || !ring.open(ring_path.c_str(), err) || !nets.load(models, int8, err) ||
|
||||
(publish && (!pub.open(err) || !gestures.open(pinch, err))) || (!record.empty() && !start_recording(record, err))) {
|
||||
std::fprintf(stderr, "%s\n", err.c_str());
|
||||
return 1;
|
||||
}
|
||||
if (!ring.alive()) return std::fprintf(stderr, "ft-camd isn't running (no heartbeat)\n"), 1;
|
||||
|
||||
std::map<std::string, int> index; // calibration name -> ring camera
|
||||
// Recorded only, not tracked: "<name>_dk" (ft-camd --with-dark) and "color_video<N>"
|
||||
// (--with-color; which is left and right is up to tools/check_color.py). Recorded names
|
||||
// hold 15 characters, so "upper_right_dark" wouldn't fit.
|
||||
std::map<std::string, int> dark;
|
||||
std::map<std::string, Camera> used;
|
||||
for (int i = 0; i < ring.cameras(); ++i) {
|
||||
if (ring.camera(i).flags & FH_CAM_COLOR) {
|
||||
dark["color_video" + std::to_string(ring.camera(i).node)] = i;
|
||||
continue;
|
||||
}
|
||||
// ft-camd's cameras by capture pipe; ft-ringplay's (no device) by the name it gives
|
||||
const char *name = camera_for_pipe(ring.camera(i).node);
|
||||
if (!name && ring.camera(i).node < 0) name = ring.camera(i).name;
|
||||
if (!name || !calib.count(name)) continue;
|
||||
if (ring.camera(i).flags & FH_CAM_DARK) dark[std::string(name) + "_dk"] = i;
|
||||
else index[name] = i, used[name] = calib[name];
|
||||
}
|
||||
// ft-camd tells the side cameras' buffers apart by XRService's allocation order, which
|
||||
// some XRService restarts reverse; tools/check_sides.py --ring tells when.
|
||||
if (swap_sides && index.count("slam_left") && index.count("slam_right")) {
|
||||
std::swap(index["slam_left"], index["slam_right"]);
|
||||
if (dark.count("slam_left_dk") && dark.count("slam_right_dk")) std::swap(dark["slam_left_dk"], dark["slam_right_dk"]);
|
||||
std::printf("side cameras swapped (--swap-sides)\n");
|
||||
}
|
||||
std::printf("cameras:");
|
||||
for (auto &[name, i] : index) std::printf(" %s=video%d", name.c_str(), ring.camera(i).node);
|
||||
std::printf(" models: %s%s, %d threads on CPUs", models.c_str(), int8 ? " (int8)" : "", threads);
|
||||
for (int c : cpus) std::printf(" %d", c);
|
||||
std::printf("\n");
|
||||
|
||||
cpu_set_t set; // the main loop too
|
||||
CPU_ZERO(&set);
|
||||
for (int c : cpus) CPU_SET(c, &set);
|
||||
if (sched_setaffinity(0, sizeof set, &set) < 0) std::perror("sched_setaffinity");
|
||||
nets.set_contrast(palm_contrast, hand_contrast);
|
||||
Pool pool(threads, cpus);
|
||||
Tracker tracker(used, nets, pool);
|
||||
tracker.set_keep_presence(keep_presence);
|
||||
std::map<std::string, std::vector<uint8_t>> pixels;
|
||||
std::map<std::string, uint64_t> last;
|
||||
const uint64_t start = mono_ns();
|
||||
uint64_t t_status = start, next_ns = 0;
|
||||
double cpu0 = cpu_seconds();
|
||||
std::vector<double> lat;
|
||||
double hands_sum = 0, resid_sum = 0;
|
||||
int resid_n = 0, left_sets = 0, right_sets = 0, both_sets = 0;
|
||||
|
||||
while (!g_stop && (seconds <= 0 || (mono_ns() - start) / 1e9 < seconds)) {
|
||||
if (!ring.alive()) return std::fprintf(stderr, "ft-camd stopped\n"), 2;
|
||||
// a new frame set: every camera has a newer frame, taken at the same moment
|
||||
std::map<std::string, uint64_t> latest;
|
||||
bool ready = true;
|
||||
for (auto &[name, i] : index) {
|
||||
latest[name] = ring.latest(i);
|
||||
ready = ready && latest[name] > last[name];
|
||||
}
|
||||
if (!ready) {
|
||||
std::this_thread::sleep_for(std::chrono::milliseconds(2));
|
||||
continue;
|
||||
}
|
||||
// not needed at the current rate, and not recorded: skip it without copying images
|
||||
if (track && !rec && !g_record) {
|
||||
uint64_t t0 = UINT64_MAX, t1 = 0;
|
||||
bool ok = true;
|
||||
for (auto &[name, i] : index) {
|
||||
fh_ring_slot_t meta;
|
||||
ok = ok && ring.meta(i, latest[name], &meta);
|
||||
if (ok) t0 = std::min(t0, meta.capture_ns), t1 = std::max(t1, meta.capture_ns);
|
||||
}
|
||||
if (ok && t1 - t0 <= 3'000'000 && t0 < next_ns) {
|
||||
last = latest;
|
||||
continue;
|
||||
}
|
||||
}
|
||||
std::map<std::string, Image> images;
|
||||
std::vector<SetFrame> frames;
|
||||
uint64_t tmin = UINT64_MAX, tmax = 0, dq = 0;
|
||||
bool ok = true;
|
||||
for (auto &[name, i] : index) {
|
||||
fh_ring_slot_t meta;
|
||||
ok = ok && ring.read(i, latest[name], pixels[name], &meta);
|
||||
if (!ok) break;
|
||||
const auto &c = ring.camera(i);
|
||||
images[name] = {pixels[name].data(), int(c.width), int(c.height), int(c.width)};
|
||||
frames.push_back({name, pixels[name].data(), c.width, c.height, meta.capture_ns, meta.dqbuf_ns});
|
||||
tmin = std::min(tmin, meta.capture_ns), tmax = std::max(tmax, meta.capture_ns), dq = std::max(dq, meta.dqbuf_ns);
|
||||
}
|
||||
if (!ok || tmax - tmin > 3'000'000) { // torn, or a camera is a frame behind
|
||||
std::this_thread::sleep_for(std::chrono::milliseconds(1));
|
||||
continue;
|
||||
}
|
||||
last = latest;
|
||||
if (g_record && !rec) {
|
||||
g_record = 0;
|
||||
char name[64];
|
||||
const std::time_t now = std::time(nullptr);
|
||||
std::strftime(name, sizeof name, "rec-%Y%m%d-%H%M%S", std::localtime(&now));
|
||||
const std::string dir = recordings_dir();
|
||||
std::string e;
|
||||
if (!start_recording(dir + "/" + name, e)) std::fprintf(stderr, "%s\n", e.c_str());
|
||||
}
|
||||
if (rec) { // about 80 MB/s; dark frames double that, color frames add 70 MB/s
|
||||
if ((mono_ns() - rec_start) / 1e9 < record_for) {
|
||||
for (auto &[name, i] : dark) { // the newest dark and color frames, as they are
|
||||
fh_ring_slot_t meta;
|
||||
const uint64_t n = ring.latest(i);
|
||||
const auto &c = ring.camera(i);
|
||||
if (n && ring.read(i, n, pixels[name], &meta))
|
||||
frames.push_back({name, pixels[name].data(), c.width, c.height, meta.capture_ns, meta.dqbuf_ns});
|
||||
}
|
||||
rec->add(frames);
|
||||
} else {
|
||||
const size_t n = rec->written(), d = rec->dropped();
|
||||
rec.reset(); // writes out what's queued
|
||||
std::printf("recording done: %zu sets, %zu dropped\n", n, d);
|
||||
std::fflush(stdout);
|
||||
if (!track) break;
|
||||
}
|
||||
}
|
||||
if (!track && rec && status > 0 && (mono_ns() - t_status) / 1e9 >= status) {
|
||||
std::printf("%5.1fs recorded %zu sets, dropped %zu\n", (mono_ns() - start) / 1e9, rec->written(), rec->dropped());
|
||||
std::fflush(stdout);
|
||||
t_status = mono_ns();
|
||||
}
|
||||
if (!track || tmin < next_ns) continue; // not needed yet at the current rate
|
||||
const auto hands = tracker.step(images, int64_t(tmin));
|
||||
const uint64_t capture = uint64_t(int64_t(tmin) - raw_minus_mono_ns()); // CLOCK_MONOTONIC
|
||||
pinch.update(hands, tracker.views_now(), int64_t(capture));
|
||||
// a pinch down or closing gets the full rate, even while the palm holds still
|
||||
next_ns = tmin + uint64_t((std::min(tracker.interval(), pinch.engaged() ? 1 / 30.0 : 1.0) - 0.005) * 1e9);
|
||||
if (publish) pub.write(hands, capture), gestures.write(pinch, capture);
|
||||
for (const Pinch::Event &e : pinch.events) {
|
||||
std::printf("pinch %s %-5s d %.3f m at %+.3f %+.3f %+.3f\n", e.side ? "right" : "left ", e.what, e.distance,
|
||||
e.point[0], e.point[1], e.point[2]);
|
||||
std::fflush(stdout);
|
||||
}
|
||||
lat.push_back((mono_ns() - dq) / 1e6);
|
||||
hands_sum += double(hands.size());
|
||||
bool on_left = false, on_right = false; // by where the wrist is, not the model's label
|
||||
for (const Hand *h : hands) {
|
||||
if (h->residual >= 0) resid_sum += h->residual * 1000, ++resid_n;
|
||||
(h->pts[0][0] < 0 ? on_left : on_right) = true;
|
||||
}
|
||||
left_sets += on_left, right_sets += on_right, both_sets += on_left && on_right;
|
||||
|
||||
const uint64_t now = mono_ns();
|
||||
if (status > 0 && (now - t_status) / 1e9 >= status) {
|
||||
const double dt = (now - t_status) / 1e9, cpu1 = cpu_seconds();
|
||||
const Stats &s = tracker.stats;
|
||||
std::sort(lat.begin(), lat.end());
|
||||
std::printf("%5.1fs %4.1f sets/s hands %.2f views %zu palm %3d calls %4.1f ms/batch hand %3d calls %4.1f ms/batch "
|
||||
"step %4.1f ms latency %4.1f ms resid %.1f mm CPU %3.0f%%\n",
|
||||
(now - start) / 1e9, s.sets / dt, s.sets ? hands_sum / s.sets : 0, tracker.views(), s.palm_calls,
|
||||
s.palm_batches ? s.palm_ms / s.palm_batches : 0, s.hand_calls,
|
||||
s.hand_batches ? s.hand_ms / s.hand_batches : 0, s.sets ? s.step_ms / s.sets : 0,
|
||||
lat.empty() ? 0 : lat[lat.size() / 2], resid_n ? resid_sum / resid_n : 0, 100 * (cpu1 - cpu0) / dt);
|
||||
if (s.sets)
|
||||
std::printf(" sets with a hand: left %2.0f%% right %2.0f%% both %2.0f%% views lost %d, handoff misses %d, "
|
||||
"dups %d, splits %d hands new %d merged %d forgotten %d%s\n",
|
||||
100.0 * left_sets / s.sets, 100.0 * right_sets / s.sets, 100.0 * both_sets / s.sets, s.lost,
|
||||
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten,
|
||||
!rec ? "" : (" recorded " + std::to_string(rec->written()) + " dropped " +
|
||||
std::to_string(rec->dropped())).c_str());
|
||||
std::printf(" pinches: left %u right %u (held back, palm down: %d %d)", pinch.side(0).begins,
|
||||
pinch.side(1).begins, pinch.held_back[0], pinch.held_back[1]);
|
||||
for (int k = 0; k < 2; ++k)
|
||||
if (pinch.side(k).flags & FH_PINCH_TRACKED)
|
||||
std::printf(" %s %s d %.3f", k ? "right" : "left", pinch.side(k).flags & FH_PINCH_DOWN ? "DOWN" : "open",
|
||||
pinch.side(k).distance);
|
||||
std::printf("\n");
|
||||
for (const Hand *h : hands)
|
||||
std::printf(" hand %d %-5s views %d wrist %+.3f %+.3f %+.3f m scale %.2f speed %.2f m/s\n", h->id,
|
||||
h->right() ? "right" : "left", h->nviews, h->pts[0][0], h->pts[0][1], h->pts[0][2], h->scale,
|
||||
h->speed);
|
||||
std::fflush(stdout);
|
||||
tracker.stats = Stats{};
|
||||
t_status = now, cpu0 = cpu1;
|
||||
lat.clear(), hands_sum = 0, resid_sum = 0, resid_n = 0, left_sets = right_sets = both_sets = 0;
|
||||
}
|
||||
}
|
||||
if (publish) {
|
||||
pub.write({}, mono_ns());
|
||||
pinch.release(int64_t(mono_ns())); // a drag in progress ends, as lost
|
||||
gestures.write(pinch, mono_ns());
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,255 @@
|
||||
#include "nets.h"
|
||||
|
||||
#include <mat.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstdlib>
|
||||
#include <numeric>
|
||||
|
||||
namespace {
|
||||
|
||||
constexpr int kPalmSize = 192, kHandSize = 224;
|
||||
const int kRoiLandmarks[] = {0, 1, 2, 3, 5, 6, 9, 10, 13, 14, 17, 18};
|
||||
|
||||
// 2x3 affine taking crop pixels (0..out) to image pixels.
|
||||
void crop_matrix(V2 center, double size, double rotation, int out, float tm[6]) {
|
||||
const double c = std::cos(rotation), s = std::sin(rotation), k = size / out;
|
||||
tm[0] = float(c * k), tm[1] = float(-s * k), tm[3] = float(s * k), tm[4] = float(c * k);
|
||||
tm[2] = float(center[0] - (tm[0] + tm[1]) * out / 2.0);
|
||||
tm[5] = float(center[1] - (tm[3] + tm[4]) * out / 2.0);
|
||||
}
|
||||
|
||||
V2 to_image(const float tm[6], double x, double y) {
|
||||
return {tm[0] * x + tm[1] * y + tm[2], tm[3] * x + tm[4] * y + tm[5]};
|
||||
}
|
||||
|
||||
// OpenCV's CLAHE (4x4 tiles) on a square crop, in place.
|
||||
void clahe(uint8_t *img, int n, double clip_limit) {
|
||||
constexpr int kTiles = 4;
|
||||
const int ts = n / kTiles, area = ts * ts;
|
||||
const int clip = std::max(1, int(clip_limit * area / 256));
|
||||
uint8_t lut[kTiles][kTiles][256];
|
||||
for (int ty = 0; ty < kTiles; ++ty)
|
||||
for (int tx = 0; tx < kTiles; ++tx) {
|
||||
int hist[256] = {};
|
||||
for (int y = ty * ts; y < (ty + 1) * ts; ++y)
|
||||
for (int x = tx * ts; x < (tx + 1) * ts; ++x) ++hist[img[y * n + x]];
|
||||
int excess = 0;
|
||||
for (int &h : hist)
|
||||
if (h > clip) excess += h - clip, h = clip;
|
||||
const int add = excess / 256, residual = excess - add * 256;
|
||||
for (int i = 0; i < 256; ++i) hist[i] += add + (i < residual ? 1 : 0);
|
||||
int sum = 0;
|
||||
const float scale = 255.f / area;
|
||||
for (int i = 0; i < 256; ++i) {
|
||||
sum += hist[i];
|
||||
lut[ty][tx][i] = uint8_t(std::min(255, int(sum * scale + 0.5f)));
|
||||
}
|
||||
}
|
||||
std::vector<uint8_t> out(size_t(n) * n);
|
||||
for (int y = 0; y < n; ++y) {
|
||||
const float fy = (y + 0.5f) / ts - 0.5f;
|
||||
const int y0 = std::clamp(int(std::floor(fy)), 0, kTiles - 1), y1 = std::min(y0 + 1, kTiles - 1);
|
||||
const float wy = std::clamp(fy - y0, 0.f, 1.f);
|
||||
for (int x = 0; x < n; ++x) {
|
||||
const float fx = (x + 0.5f) / ts - 0.5f;
|
||||
const int x0 = std::clamp(int(std::floor(fx)), 0, kTiles - 1), x1 = std::min(x0 + 1, kTiles - 1);
|
||||
const float wx = std::clamp(fx - x0, 0.f, 1.f);
|
||||
const uint8_t v = img[y * n + x];
|
||||
const float top = lut[y0][x0][v] * (1 - wx) + lut[y0][x1][v] * wx;
|
||||
const float bot = lut[y1][x0][v] * (1 - wx) + lut[y1][x1][v] * wx;
|
||||
out[size_t(y) * n + x] = uint8_t(top * (1 - wy) + bot * wy + 0.5f);
|
||||
}
|
||||
}
|
||||
std::copy(out.begin(), out.end(), img);
|
||||
}
|
||||
|
||||
// Linear stretch of the 1st..99th percentile to 0..255, in place.
|
||||
void stretch(uint8_t *img, int n) {
|
||||
int hist[256] = {};
|
||||
const int total = n * n;
|
||||
for (int i = 0; i < total; ++i) ++hist[img[i]];
|
||||
int lo = 0, hi = 255, acc = 0;
|
||||
for (int v = 0; v < 256; ++v)
|
||||
if ((acc += hist[v]) > total / 100) { lo = v; break; }
|
||||
acc = 0;
|
||||
for (int v = 255; v >= 0; --v)
|
||||
if ((acc += hist[v]) > total / 100) { hi = v; break; }
|
||||
if (hi <= lo) return;
|
||||
for (int i = 0; i < total; ++i) img[i] = uint8_t(std::clamp((img[i] - lo) * 255 / (hi - lo), 0, 255));
|
||||
}
|
||||
|
||||
// A crop as the models' input: RGB (the mono plane three times), 0..1.
|
||||
ncnn::Mat crop(const Image &img, const float tm[6], int n, const Contrast &contrast) {
|
||||
std::vector<uint8_t> patch(size_t(n) * n);
|
||||
ncnn::warpaffine_bilinear_c1(img.data, img.width, img.height, img.stride, patch.data(), n, n, n, tm, 0, 0);
|
||||
if (contrast.mode == Contrast::Clahe) clahe(patch.data(), n, contrast.clip);
|
||||
else if (contrast.mode == Contrast::Stretch) stretch(patch.data(), n);
|
||||
ncnn::Mat m = ncnn::Mat::from_pixels(patch.data(), ncnn::Mat::PIXEL_GRAY2RGB, n, n);
|
||||
const float norm[3] = {1 / 255.f, 1 / 255.f, 1 / 255.f};
|
||||
m.substract_mean_normalize(nullptr, norm);
|
||||
return m;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
bool Contrast::parse(const std::string &s, Contrast &out) {
|
||||
if (s == "none") return out.mode = None, true;
|
||||
if (s == "stretch") return out.mode = Stretch, true;
|
||||
if (s.rfind("clahe", 0) == 0) {
|
||||
out.mode = Clahe;
|
||||
out.clip = s.size() > 6 && s[5] == ':' ? std::atof(s.c_str() + 6) : 2.0;
|
||||
return out.clip > 0;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
bool Contrast::parse_pair(const std::string &s, Contrast &palm, Contrast &hand) {
|
||||
const size_t slash = s.find('/');
|
||||
if (slash == std::string::npos) return parse(s, palm) && parse(s, hand);
|
||||
return parse(s.substr(0, slash), palm) && parse(s.substr(slash + 1), hand);
|
||||
}
|
||||
|
||||
namespace {
|
||||
|
||||
bool load_net(ncnn::Net &net, const std::string &base, std::string &err) {
|
||||
net.opt.num_threads = 1;
|
||||
net.opt.use_vulkan_compute = false;
|
||||
net.opt.use_fp16_packed = net.opt.use_fp16_storage = net.opt.use_fp16_arithmetic = true;
|
||||
if (net.load_param((base + ".param").c_str()) || net.load_model((base + ".bin").c_str())) {
|
||||
err = "can't load " + base + ".param/.bin";
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
Roi Palm::roi() const {
|
||||
const V2 a = kp[0], b = kp[2];
|
||||
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(b[1] - a[1]), b[0] - a[0]));
|
||||
const double h = size[1];
|
||||
const V2 shift{-h * -0.5 * std::sin(rot), h * -0.5 * std::cos(rot)};
|
||||
return {center + shift, std::max(size[0], size[1]) * 2.6, rot};
|
||||
}
|
||||
|
||||
Roi roi_from_points(const V2 *p) {
|
||||
const V2 w = p[0];
|
||||
V2 m = (p[5] + p[13]) * 0.5;
|
||||
m = (m + p[9]) * 0.5;
|
||||
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(m[1] - w[1]), m[0] - w[0]));
|
||||
V2 lo{1e9, 1e9}, hi{-1e9, -1e9};
|
||||
for (int i : kRoiLandmarks)
|
||||
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[i][k]), hi[k] = std::max(hi[k], p[i][k]);
|
||||
V2 center = (lo + hi) * 0.5;
|
||||
const double c = std::cos(-rot), s = std::sin(-rot);
|
||||
V2 qlo{1e9, 1e9}, qhi{-1e9, -1e9};
|
||||
for (int i : kRoiLandmarks) {
|
||||
const V2 d = p[i] - center;
|
||||
const V2 q{d[0] * c - d[1] * s, d[0] * s + d[1] * c};
|
||||
for (int k = 0; k < 2; ++k) qlo[k] = std::min(qlo[k], q[k]), qhi[k] = std::max(qhi[k], q[k]);
|
||||
}
|
||||
const V2 mid = (qlo + qhi) * 0.5;
|
||||
const double c2 = std::cos(rot), s2 = std::sin(rot);
|
||||
center = center + V2{mid[0] * c2 - mid[1] * s2, mid[0] * s2 + mid[1] * c2};
|
||||
const double w2 = qhi[0] - qlo[0], h2 = qhi[1] - qlo[1];
|
||||
center = center + V2{-h2 * -0.1 * s2, h2 * -0.1 * c2};
|
||||
return {center, std::max(w2, h2) * 2.0, rot};
|
||||
}
|
||||
|
||||
Roi Landmarks::next_roi() const { return roi_from_points(pts); }
|
||||
|
||||
bool Nets::load(const std::string &dir, bool int8, std::string &err) {
|
||||
const std::string suffix = int8 ? "-int8.ncnn" : ".ncnn";
|
||||
if (!load_net(palm_, dir + "/palm" + suffix, err) || !load_net(hand_, dir + "/hand" + suffix, err)) return false;
|
||||
// SSD anchors of palm_detection_full: strides 8 (2 per cell) and 16 (6 per cell)
|
||||
for (auto [stride, per] : {std::pair{8, 2}, std::pair{16, 6}}) {
|
||||
const int n = kPalmSize / stride;
|
||||
for (int y = 0; y < n; ++y)
|
||||
for (int x = 0; x < n; ++x)
|
||||
for (int k = 0; k < per; ++k) anchors_.push_back({(x + 0.5) / n * kPalmSize, (y + 0.5) / n * kPalmSize});
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
std::vector<Palm> Nets::palms(const Image &img, V2 center, double size, double rotation) const {
|
||||
float tm[6];
|
||||
crop_matrix(center, size, rotation, kPalmSize, tm);
|
||||
ncnn::Extractor ex = palm_.create_extractor();
|
||||
ex.input("in0", crop(img, tm, kPalmSize, palm_contrast_));
|
||||
ncnn::Mat boxes, scores;
|
||||
ex.extract("out0", boxes);
|
||||
ex.extract("out1", scores);
|
||||
const float *raw = boxes, *logit = scores;
|
||||
const int n = int(anchors_.size());
|
||||
const float min_logit = std::log(0.5f / 0.5f); // score 0.5
|
||||
struct Cand { V2 c, s; V2 kp[7]; double score; };
|
||||
std::vector<Cand> cand;
|
||||
for (int i = 0; i < n; ++i) {
|
||||
if (logit[i] <= min_logit) continue;
|
||||
const float *r = raw + i * 18;
|
||||
Cand c;
|
||||
c.c = {r[0] + anchors_[i][0], r[1] + anchors_[i][1]};
|
||||
c.s = {r[2], r[3]};
|
||||
for (int k = 0; k < 7; ++k) c.kp[k] = {r[4 + 2 * k] + anchors_[i][0], r[5 + 2 * k] + anchors_[i][1]};
|
||||
c.score = 1 / (1 + std::exp(-std::clamp(double(logit[i]), -100.0, 100.0)));
|
||||
cand.push_back(c);
|
||||
}
|
||||
// MediaPipe's weighted NMS: overlapping boxes are averaged, weighted by score
|
||||
std::sort(cand.begin(), cand.end(), [](const Cand &a, const Cand &b) { return a.score > b.score; });
|
||||
std::vector<bool> used(cand.size());
|
||||
std::vector<Palm> out;
|
||||
for (size_t i = 0; i < cand.size(); ++i) {
|
||||
if (used[i]) continue;
|
||||
double wsum = 0;
|
||||
Cand acc{};
|
||||
for (size_t j = i; j < cand.size(); ++j) {
|
||||
if (used[j]) continue;
|
||||
const double ix = std::max(0.0, std::min(cand[i].c[0] + cand[i].s[0] / 2, cand[j].c[0] + cand[j].s[0] / 2) -
|
||||
std::max(cand[i].c[0] - cand[i].s[0] / 2, cand[j].c[0] - cand[j].s[0] / 2));
|
||||
const double iy = std::max(0.0, std::min(cand[i].c[1] + cand[i].s[1] / 2, cand[j].c[1] + cand[j].s[1] / 2) -
|
||||
std::max(cand[i].c[1] - cand[i].s[1] / 2, cand[j].c[1] - cand[j].s[1] / 2));
|
||||
const double inter = ix * iy;
|
||||
const double uni = cand[i].s[0] * cand[i].s[1] + cand[j].s[0] * cand[j].s[1] - inter;
|
||||
if (j != i && inter / (uni + 1e-9) <= 0.3) continue;
|
||||
used[j] = true;
|
||||
const double w = cand[j].score;
|
||||
wsum += w;
|
||||
acc.c = acc.c + cand[j].c * w;
|
||||
acc.s = acc.s + cand[j].s * w;
|
||||
for (int k = 0; k < 7; ++k) acc.kp[k] = acc.kp[k] + cand[j].kp[k] * w;
|
||||
}
|
||||
Palm p;
|
||||
const V2 c = acc.c * (1 / wsum);
|
||||
p.center = to_image(tm, c[0], c[1]);
|
||||
p.size = acc.s * (1 / wsum * size / kPalmSize);
|
||||
for (int k = 0; k < 7; ++k) {
|
||||
const V2 q = acc.kp[k] * (1 / wsum);
|
||||
p.kp[k] = to_image(tm, q[0], q[1]);
|
||||
}
|
||||
p.score = cand[i].score;
|
||||
out.push_back(p);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
Landmarks Nets::landmarks(const Image &img, const Roi &roi) const {
|
||||
float tm[6];
|
||||
crop_matrix(roi.center, roi.size, roi.rotation, kHandSize, tm);
|
||||
ncnn::Extractor ex = hand_.create_extractor();
|
||||
ex.input("in0", crop(img, tm, kHandSize, hand_contrast_));
|
||||
ncnn::Mat screen, presence, right, world;
|
||||
ex.extract("out0", screen);
|
||||
ex.extract("out1", presence);
|
||||
ex.extract("out2", right);
|
||||
ex.extract("out3", world);
|
||||
Landmarks lm;
|
||||
const float *s = screen, *w = world;
|
||||
for (int i = 0; i < 21; ++i) {
|
||||
lm.pts[i] = to_image(tm, s[3 * i], s[3 * i + 1]);
|
||||
for (int k = 0; k < 3; ++k) lm.world[i][k] = w[3 * i + k];
|
||||
}
|
||||
lm.presence = presence[0];
|
||||
lm.right = right[0];
|
||||
return lm;
|
||||
}
|
||||
@@ -0,0 +1,65 @@
|
||||
// MediaPipe's palm detector and hand landmark model on ncnn. A crop is a square region of a camera image: centre and size in
|
||||
// pixels, and a rotation that turns the crop's "up" toward the image direction
|
||||
// (sin r, -cos r). Crops are contrast-equalized (CLAHE) before the models see them.
|
||||
// Everything here may run on several threads at once.
|
||||
#pragma once
|
||||
|
||||
#include "geom.h"
|
||||
|
||||
#include <net.h>
|
||||
|
||||
#include <cstdint>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
struct Image {
|
||||
const uint8_t *data = nullptr;
|
||||
int width = 0, height = 0, stride = 0;
|
||||
};
|
||||
|
||||
struct Roi {
|
||||
V2 center{};
|
||||
double size = 0, rotation = 0;
|
||||
};
|
||||
|
||||
struct Palm {
|
||||
V2 center{}, size{};
|
||||
V2 kp[7]{};
|
||||
double score = 0;
|
||||
Roi roi() const; // MediaPipe's hand crop for this palm
|
||||
};
|
||||
|
||||
struct Landmarks {
|
||||
V2 pts[21]{}; // image pixels
|
||||
double world[21][3]{}; // MediaPipe's metric landmarks, hand-centred
|
||||
double presence = 0, right = 0;
|
||||
Roi next_roi() const; // MediaPipe's crop to track the hand in the next frame
|
||||
};
|
||||
|
||||
Roi roi_from_points(const V2 *pts21);
|
||||
|
||||
// How crops are contrast-equalized before the models see them.
|
||||
struct Contrast {
|
||||
enum Mode { Clahe, None, Stretch } mode = Clahe;
|
||||
double clip = 2.0; // Clahe: OpenCV's clip limit (4x4 tiles)
|
||||
// "clahe:2", "none", "stretch" (1st..99th percentile to 0..255)
|
||||
static bool parse(const std::string &s, Contrast &out);
|
||||
// "PALM/HAND" (each as above), or one for both
|
||||
static bool parse_pair(const std::string &s, Contrast &palm, Contrast &hand);
|
||||
};
|
||||
|
||||
class Nets {
|
||||
public:
|
||||
// Loads <dir>/palm.ncnn.* and <dir>/hand.ncnn.*, or the -int8 variants.
|
||||
bool load(const std::string &dir, bool int8, std::string &err);
|
||||
std::vector<Palm> palms(const Image &img, V2 center, double size, double rotation) const;
|
||||
Landmarks landmarks(const Image &img, const Roi &roi) const;
|
||||
// Before any palms()/landmarks(): how the palm search's and the landmark model's crops
|
||||
// are equalized.
|
||||
void set_contrast(const Contrast &palm, const Contrast &hand) { palm_contrast_ = palm, hand_contrast_ = hand; }
|
||||
|
||||
private:
|
||||
Contrast palm_contrast_, hand_contrast_;
|
||||
ncnn::Net palm_, hand_;
|
||||
std::vector<V2> anchors_;
|
||||
};
|
||||
@@ -0,0 +1,165 @@
|
||||
#include "pinch.h"
|
||||
|
||||
#include "io.h"
|
||||
|
||||
#include <fcntl.h>
|
||||
#include <sys/mman.h>
|
||||
#include <sys/stat.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cerrno>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
|
||||
namespace {
|
||||
|
||||
constexpr int kThumbTip = 4, kIndexTip = 8;
|
||||
|
||||
V3 cross(V3 a, V3 b) { return {a[1] * b[2] - a[2] * b[1], a[2] * b[0] - a[0] * b[2], a[0] * b[1] - a[1] * b[0]}; }
|
||||
|
||||
void put3(float out[3], V3 v) {
|
||||
for (int k = 0; k < 3; ++k) out[k] = float(v[k]);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
void Pinch::update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns) {
|
||||
events.clear();
|
||||
// A hand a pinch is down on belongs to that side until it ends. The left/right call is a
|
||||
// running average of the model's, and when it flips mid-pinch the other side would take
|
||||
// the same hand and pinch too (4 times in the 2026-09-30 lit recording).
|
||||
int taken[2] = {0, 0};
|
||||
for (int s = 0; s < 2; ++s)
|
||||
if (side_[s].flags & FH_PINCH_DOWN) taken[s] = follow_[s];
|
||||
for (int s = 0; s < 2; ++s) {
|
||||
fh_pinch_t &o = side_[s];
|
||||
const bool down = o.flags & FH_PINCH_DOWN;
|
||||
// the hand: while down, the one the pinch began on; else the best tracked hand of this side
|
||||
const Hand *h = nullptr;
|
||||
for (const Hand *c : hands) {
|
||||
if (down ? c->id != follow_[s] : c->right() != (s == 1) || c->id == taken[1 - s]) continue;
|
||||
if (!h || c->frames > h->frames) h = c;
|
||||
}
|
||||
world_d[s] = tri_d[s] = palm_down[s] = -1;
|
||||
if (!h) {
|
||||
o.flags &= ~FH_PINCH_TRACKED;
|
||||
if (down && (t_ns - seen_ns_[s]) / 1e9 > p_.grace_s) end(s, t_ns, true);
|
||||
continue;
|
||||
}
|
||||
seen_ns_[s] = t_ns;
|
||||
tri_d[s] = norm(h->pts[kThumbTip] - h->pts[kIndexTip]);
|
||||
const V3 normal = cross(h->smooth[5] - h->smooth[0], h->smooth[17] - h->smooth[0]);
|
||||
palm_down[s] = norm(normal) > 0 ? std::fabs(normal[1]) / norm(normal) : 0;
|
||||
double sum = 0;
|
||||
int n = 0;
|
||||
for (const Seen &v : views) {
|
||||
if (v.hand != h->id) continue;
|
||||
const V3 a{v.lm.world[kThumbTip][0], v.lm.world[kThumbTip][1], v.lm.world[kThumbTip][2]};
|
||||
const V3 b{v.lm.world[kIndexTip][0], v.lm.world[kIndexTip][1], v.lm.world[kIndexTip][2]};
|
||||
sum += norm(a - b), ++n;
|
||||
}
|
||||
if (n) world_d[s] = sum / n * h->scale;
|
||||
const double d = p_.triangulated || world_d[s] < 0 ? tri_d[s] : world_d[s];
|
||||
const V3 point = (h->smooth[kThumbTip] + h->smooth[kIndexTip]) * 0.5;
|
||||
o.flags |= FH_PINCH_TRACKED;
|
||||
o.hand_id = uint32_t(h->id);
|
||||
o.distance = float(d);
|
||||
o.strength = float(std::clamp((p_.end_m - d) / (p_.end_m - p_.begin_m), 0.0, 1.0));
|
||||
put3(o.point, point);
|
||||
if (!down) {
|
||||
// a close held back (palm down) has to open again before a pinch can begin, so
|
||||
// turning the hand with the fingers still closed doesn't start one
|
||||
if (d > p_.end_m) held_[s] = false;
|
||||
if (d < p_.begin_m && !held_[s] && palm_down[s] > p_.palm_down_max) {
|
||||
held_[s] = true;
|
||||
++held_back[s];
|
||||
} else if (d < p_.begin_m && !held_[s]) {
|
||||
o.flags = (o.flags | FH_PINCH_DOWN) & ~FH_PINCH_LOST;
|
||||
++o.begins;
|
||||
o.begin_ns = uint64_t(t_ns);
|
||||
put3(o.begin_point, point);
|
||||
follow_[s] = h->id;
|
||||
open_frames_[s] = 0;
|
||||
events.push_back({s, "begin", t_ns, d, point});
|
||||
}
|
||||
} else if (d > p_.end_m) {
|
||||
if (++open_frames_[s] >= p_.end_frames) end(s, t_ns, false);
|
||||
} else {
|
||||
open_frames_[s] = 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void Pinch::end(int s, int64_t t_ns, bool lost) {
|
||||
fh_pinch_t &o = side_[s];
|
||||
o.flags = (o.flags & ~FH_PINCH_DOWN) | (lost ? FH_PINCH_LOST : 0);
|
||||
++o.ends;
|
||||
o.end_ns = uint64_t(t_ns);
|
||||
follow_[s] = 0;
|
||||
events.push_back({s, lost ? "lost" : "end", t_ns, o.distance, {o.point[0], o.point[1], o.point[2]}});
|
||||
}
|
||||
|
||||
void Pinch::release(int64_t t_ns) {
|
||||
events.clear();
|
||||
for (int s = 0; s < 2; ++s) {
|
||||
side_[s].flags &= ~FH_PINCH_TRACKED;
|
||||
if (side_[s].flags & FH_PINCH_DOWN) end(s, t_ns, true);
|
||||
}
|
||||
}
|
||||
|
||||
bool Pinch::engaged() const {
|
||||
for (const fh_pinch_t &o : side_)
|
||||
if ((o.flags & FH_PINCH_DOWN) || ((o.flags & FH_PINCH_TRACKED) && o.strength > 0.3f)) return true;
|
||||
return false;
|
||||
}
|
||||
|
||||
bool GesturePublisher::open(const Pinch &pinch, std::string &err) {
|
||||
const std::string path = run_dir() + "/gestures";
|
||||
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
|
||||
if (fd < 0 || ftruncate(fd, sizeof(fh_gestures_t)) < 0) return err = path + ": " + std::strerror(errno), false;
|
||||
void *m = mmap(nullptr, sizeof(fh_gestures_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return err = path + ": can't map it", false;
|
||||
out_ = static_cast<fh_gestures_t *>(m);
|
||||
// keep the counters a previous tracker left, so a reader doesn't see them jump back
|
||||
const bool ours = !std::memcmp(out_->magic, FH_GESTURES_MAGIC, 8) && out_->version == FH_GESTURES_VERSION;
|
||||
if (!ours) {
|
||||
std::memset(out_, 0, sizeof *out_);
|
||||
std::memcpy(out_->magic, FH_GESTURES_MAGIC, 8);
|
||||
out_->version = FH_GESTURES_VERSION;
|
||||
out_->size = sizeof(fh_gestures_t);
|
||||
}
|
||||
seq_ = out_->seq / 2 + 1;
|
||||
// a pinch the last tracker left down (it crashed) is over: count its end, as lost
|
||||
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
for (fh_pinch_t &o : out_->pinch)
|
||||
if (o.begins != o.ends) {
|
||||
o.ends = o.begins;
|
||||
o.end_ns = mono_ns();
|
||||
o.flags = (o.flags & ~FH_PINCH_DOWN) | FH_PINCH_LOST;
|
||||
}
|
||||
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||
out_->begin_m = float(pinch.params().begin_m);
|
||||
out_->end_m = float(pinch.params().end_m);
|
||||
return true;
|
||||
}
|
||||
|
||||
void GesturePublisher::write(const Pinch &pinch, uint64_t capture_ns) {
|
||||
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
for (int s = 0; s < 2; ++s) {
|
||||
// counters carry on from what's in the file (a restarted tracker starts its own at 0)
|
||||
const fh_pinch_t &in = pinch.side(s);
|
||||
fh_pinch_t &o = out_->pinch[s];
|
||||
const uint32_t base_b = o.begins - last_begins_[s], base_e = o.ends - last_ends_[s];
|
||||
o = in;
|
||||
o.begins = base_b + in.begins;
|
||||
o.ends = base_e + in.ends;
|
||||
last_begins_[s] = in.begins, last_ends_[s] = in.ends;
|
||||
}
|
||||
out_->capture_ns = capture_ns;
|
||||
out_->publish_ns = mono_ns();
|
||||
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
// Pinch detection for input: look at something and pinch to click, pinch and move to drag.
|
||||
// Per side, from the tracker's hands after each step; published as fh_gestures.h.
|
||||
#pragma once
|
||||
|
||||
#include "tracker.h"
|
||||
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
extern "C" {
|
||||
#include "../include/fh_gestures.h"
|
||||
}
|
||||
|
||||
struct PinchParams {
|
||||
double begin_m = 0.020; // thumb and index tips closer than this: the pinch begins
|
||||
double end_m = 0.035; // further apart than this: it ends (the gap keeps it from flickering)
|
||||
int end_frames = 2; // processed frames in a row past end_m before it ends, so one
|
||||
// noisy frame doesn't drop a drag
|
||||
double grace_s = 0.25; // a pinching hand lost this long ends its pinch (FH_PINCH_LOST)
|
||||
// Where the distance comes from: MediaPipe's world landmarks (the model's own 3D hand
|
||||
// pose, averaged over the hand's views, at the user's hand size), or the tracker's
|
||||
// triangulated tips. The model's pose should hold up better when the fingers hide each
|
||||
// other; tomorrow's recordings will tell.
|
||||
bool triangulated = false;
|
||||
// No pinch begins while the palm faces down more than this (|palm normal . up| in the
|
||||
// head frame; 1 turns it off). Typing curls the thumb onto the index: in the 2026-09-30
|
||||
// lit recording, pinches that began while typing had 0.69-1.00, deliberate ones 0.00-0.50.
|
||||
// Looking down tilts the head frame, which lowers the reading for a hand on a keyboard.
|
||||
double palm_down_max = 0.6;
|
||||
};
|
||||
|
||||
class Pinch {
|
||||
public:
|
||||
explicit Pinch(const PinchParams &p = {}) : p_(p) {}
|
||||
const PinchParams ¶ms() const { return p_; }
|
||||
// After each processed set: the hands out of Tracker::step, the tracker's views (for
|
||||
// the world landmarks) and the capture time.
|
||||
void update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns);
|
||||
// Ends any pinch that's down (as lost), e.g. when the tracker stops.
|
||||
void release(int64_t t_ns);
|
||||
const fh_pinch_t &side(int s) const { return side_[s]; } // 0 left, 1 right
|
||||
// A pinch is down or closing: worth tracking at the full rate.
|
||||
bool engaged() const;
|
||||
|
||||
// What changed in the last update, for logs.
|
||||
struct Event {
|
||||
int side;
|
||||
const char *what; // "begin", "end", "lost"
|
||||
int64_t t_ns;
|
||||
double distance;
|
||||
V3 point;
|
||||
};
|
||||
std::vector<Event> events;
|
||||
// Both distance measures for the last update, per side (-1: no hand), for logs.
|
||||
double world_d[2] = {-1, -1}, tri_d[2] = {-1, -1};
|
||||
double palm_down[2] = {-1, -1}; // |palm normal . up| of each side's hand
|
||||
int held_back[2] = {0, 0}; // pinches that didn't begin because the palm faced down
|
||||
|
||||
private:
|
||||
void end(int s, int64_t t_ns, bool lost);
|
||||
PinchParams p_;
|
||||
fh_pinch_t side_[2]{};
|
||||
int follow_[2] = {0, 0}; // the hand id a pinch follows while down
|
||||
int open_frames_[2] = {0, 0};
|
||||
int64_t seen_ns_[2] = {0, 0};
|
||||
bool held_[2] = {false, false}; // a close held back (palm down) that hasn't opened yet
|
||||
};
|
||||
|
||||
// Writes /run/user/UID/frametop-hands/gestures.
|
||||
class GesturePublisher {
|
||||
public:
|
||||
bool open(const Pinch &pinch, std::string &err);
|
||||
void write(const Pinch &pinch, uint64_t capture_ns);
|
||||
|
||||
private:
|
||||
fh_gestures_t *out_ = nullptr;
|
||||
uint64_t seq_ = 0;
|
||||
uint32_t last_begins_[2] = {0, 0}, last_ends_[2] = {0, 0}; // Pinch's counters last written
|
||||
};
|
||||
@@ -0,0 +1,101 @@
|
||||
#include "record.h"
|
||||
|
||||
#include <sys/stat.h>
|
||||
|
||||
#include <cerrno>
|
||||
#include <cstring>
|
||||
|
||||
namespace {
|
||||
constexpr size_t kMaxQueued = 48; // about 130 MB of sets
|
||||
}
|
||||
|
||||
Recorder::~Recorder() {
|
||||
if (!f_) return;
|
||||
{
|
||||
std::lock_guard<std::mutex> l(mu_);
|
||||
stop_ = true;
|
||||
}
|
||||
wake_.notify_all();
|
||||
thread_.join();
|
||||
std::fclose(f_);
|
||||
}
|
||||
|
||||
bool Recorder::open(const std::string &dir, std::string &err) {
|
||||
if (mkdir(dir.c_str(), 0755) < 0 && errno != EEXIST) return err = dir + ": " + std::strerror(errno), false;
|
||||
const std::string path = dir + "/sets.bin";
|
||||
f_ = std::fopen(path.c_str(), "wbx"); // never overwrite a recording
|
||||
if (!f_) return err = path + ": " + std::strerror(errno), false;
|
||||
thread_ = std::thread(&Recorder::loop, this);
|
||||
return true;
|
||||
}
|
||||
|
||||
void Recorder::add(const std::vector<SetFrame> &frames) {
|
||||
size_t bytes = sizeof(fh_set_hdr_t) + frames.size() * sizeof(fh_set_cam_t);
|
||||
for (const SetFrame &s : frames) bytes += size_t(s.width) * s.height;
|
||||
std::vector<uint8_t> rec(bytes);
|
||||
fh_set_hdr_t h{};
|
||||
std::memcpy(h.magic, FH_SET_MAGIC, 8);
|
||||
h.ncams = uint32_t(frames.size());
|
||||
h.bytes = uint32_t(bytes);
|
||||
std::memcpy(rec.data(), &h, sizeof h);
|
||||
uint8_t *p = rec.data() + sizeof h;
|
||||
for (const SetFrame &s : frames) {
|
||||
fh_set_cam_t c{};
|
||||
std::strncpy(c.name, s.name.c_str(), sizeof c.name - 1);
|
||||
c.width = s.width, c.height = s.height, c.capture_ns = s.capture_ns, c.dqbuf_ns = s.dqbuf_ns;
|
||||
std::memcpy(p, &c, sizeof c);
|
||||
p += sizeof c;
|
||||
}
|
||||
for (const SetFrame &s : frames) {
|
||||
std::memcpy(p, s.px, size_t(s.width) * s.height);
|
||||
p += size_t(s.width) * s.height;
|
||||
}
|
||||
{
|
||||
std::lock_guard<std::mutex> l(mu_);
|
||||
if (queue_.size() >= kMaxQueued) {
|
||||
++dropped_;
|
||||
return;
|
||||
}
|
||||
queue_.push_back(std::move(rec));
|
||||
}
|
||||
wake_.notify_one();
|
||||
}
|
||||
|
||||
void Recorder::loop() {
|
||||
std::unique_lock<std::mutex> l(mu_);
|
||||
for (;;) {
|
||||
wake_.wait(l, [&] { return stop_ || !queue_.empty(); });
|
||||
if (queue_.empty()) return; // stopping, and everything is written
|
||||
std::vector<uint8_t> rec = std::move(queue_.front());
|
||||
queue_.pop_front();
|
||||
l.unlock();
|
||||
const bool ok = std::fwrite(rec.data(), 1, rec.size(), f_) == rec.size();
|
||||
l.lock();
|
||||
ok ? ++written_ : ++dropped_;
|
||||
}
|
||||
}
|
||||
|
||||
SetReader::~SetReader() {
|
||||
if (f_) std::fclose(f_);
|
||||
}
|
||||
|
||||
bool SetReader::open(const std::string &dir, std::string &err) {
|
||||
const std::string path = dir + "/sets.bin";
|
||||
f_ = std::fopen(path.c_str(), "rb");
|
||||
return f_ ? true : (err = path + ": " + std::strerror(errno), false);
|
||||
}
|
||||
|
||||
bool SetReader::next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels) {
|
||||
fh_set_hdr_t h;
|
||||
if (std::fread(&h, sizeof h, 1, f_) != 1 || std::memcmp(h.magic, FH_SET_MAGIC, 8) || h.ncams == 0 || h.ncams > 16)
|
||||
return false;
|
||||
cams.resize(h.ncams);
|
||||
if (std::fread(cams.data(), sizeof(fh_set_cam_t), h.ncams, f_) != h.ncams) return false;
|
||||
pixels.resize(h.ncams);
|
||||
for (uint32_t i = 0; i < h.ncams; ++i) {
|
||||
cams[i].name[sizeof cams[i].name - 1] = 0;
|
||||
pixels[i].resize(size_t(cams[i].width) * cams[i].height);
|
||||
if (std::fread(pixels[i].data(), 1, pixels[i].size(), f_) != pixels[i].size()) return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
@@ -0,0 +1,69 @@
|
||||
// Recordings of frame sets, for replaying live sessions through the tracker offline
|
||||
// (ft-handreplay). A recording is DIR/sets.bin: one record per frame set, each
|
||||
// fh_set_hdr_t, then per camera fh_set_cam_t, then each camera's pixels (w x h, packed)
|
||||
// in the same camera order.
|
||||
#pragma once
|
||||
|
||||
#include <condition_variable>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <deque>
|
||||
#include <mutex>
|
||||
#include <string>
|
||||
#include <thread>
|
||||
#include <vector>
|
||||
|
||||
#define FH_SET_MAGIC "FHSET01"
|
||||
|
||||
struct fh_set_hdr_t {
|
||||
char magic[8];
|
||||
uint32_t ncams;
|
||||
uint32_t bytes; // the whole record, this header included
|
||||
};
|
||||
|
||||
struct fh_set_cam_t {
|
||||
char name[16]; // calibration name, e.g. "slam_left"
|
||||
uint32_t width, height;
|
||||
uint64_t capture_ns; // CLOCK_MONOTONIC_RAW, as the ring has it
|
||||
uint64_t dqbuf_ns; // CLOCK_MONOTONIC
|
||||
};
|
||||
|
||||
struct SetFrame {
|
||||
std::string name;
|
||||
const uint8_t *px;
|
||||
uint32_t width, height;
|
||||
uint64_t capture_ns, dqbuf_ns;
|
||||
};
|
||||
|
||||
// Writes sets on its own thread, so a slow disk never holds up tracking; drops sets
|
||||
// when too many are waiting.
|
||||
class Recorder {
|
||||
public:
|
||||
~Recorder();
|
||||
bool open(const std::string &dir, std::string &err);
|
||||
void add(const std::vector<SetFrame> &frames);
|
||||
size_t written() const { return written_; }
|
||||
size_t dropped() const { return dropped_; }
|
||||
|
||||
private:
|
||||
void loop();
|
||||
FILE *f_ = nullptr;
|
||||
std::thread thread_;
|
||||
std::mutex mu_;
|
||||
std::condition_variable wake_;
|
||||
std::deque<std::vector<uint8_t>> queue_;
|
||||
bool stop_ = false;
|
||||
size_t written_ = 0, dropped_ = 0;
|
||||
};
|
||||
|
||||
// Reads a recording back one set at a time.
|
||||
class SetReader {
|
||||
public:
|
||||
bool open(const std::string &dir, std::string &err);
|
||||
// False at the end (or on a truncated last set).
|
||||
bool next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels);
|
||||
~SetReader();
|
||||
|
||||
private:
|
||||
FILE *f_ = nullptr;
|
||||
};
|
||||
@@ -0,0 +1,352 @@
|
||||
// ft-handreplay: run a recording (ft-hands --record) through the tracker offline, with the
|
||||
// live scheduling, and report how well it kept the hands.
|
||||
//
|
||||
// ft-handreplay DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR]
|
||||
// [--from S] [--to S] [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch)
|
||||
//
|
||||
// --oracle N: every N-th set, also search every tile of every camera (slow), to see
|
||||
// which hands were there to find. Compares that with what the tracker had.
|
||||
// --slow F: the live tracker skips the sets that arrive while it's busy; replay takes
|
||||
// each step's time here times F as the busy time (the headset is busier live).
|
||||
// --cost: instead of timing the steps, charge each round of model calls what it
|
||||
// typically costs live (10 ms landmarks, 18 ms palms): repeatable results.
|
||||
// --timeline: per processed set, a line per hand (time, id, side, views, wrist) and per view
|
||||
// (hand, camera, presence, next crop, set index).
|
||||
// --keep-presence P: landmark presence a tracked view needs to stay (default 0.5, as new ones).
|
||||
// --pinch-begin M, --pinch-end M, --pinch-triangulated, --pinch-palm-down MAX: the pinch detector (track/pinch.h);
|
||||
// the timeline gets its begin/end/lost events and both distance measures per set.
|
||||
// --cams mono|color|all: which cameras to track with (default mono). color and all need a
|
||||
// recording made with ft-camd --with-color; --color-left NODE (color_video0 or
|
||||
// color_video3) and --color-crop subtract|none say how its calibration maps
|
||||
// (tools/check_color.py).
|
||||
// --contrast: how the palm search's and the landmark model's crops are equalized
|
||||
// (default clahe:2/none, as ft-hands).
|
||||
// --poses FILE: per processed set, a line per hand: time, id, the model's left/right call,
|
||||
// views, hand scale, then its 21 world landmarks (the model's own 3D pose, averaged
|
||||
// over its views, times the scale; metres, hand-centred) and its 21 published
|
||||
// points (head frame). For studying gestures (pinch against typing, a fist).
|
||||
// --depth FILE: per processed set, a line per hand for tools/depth_report.py: its views'
|
||||
// cameras, triangulation residual, hand scale, measured and published palm, and
|
||||
// each view's one-view palm (Tracker::single_view at the hand's scale). The
|
||||
// header has each camera's centre and focal length.
|
||||
#include "pinch.h"
|
||||
#include "record.h"
|
||||
#include "tracker.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
#include <map>
|
||||
#include <set>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
namespace {
|
||||
|
||||
struct Track {
|
||||
double first = 0, last = 0;
|
||||
int sets = 0, left = 0;
|
||||
// the last two palm positions (raw, smoothed) and times, for the jitter measure
|
||||
V3 raw[2]{}, sm[2]{};
|
||||
double t[2]{};
|
||||
int line = 0; // updates on the current unbroken run
|
||||
};
|
||||
|
||||
double median(std::vector<double> v) {
|
||||
if (v.empty()) return 0;
|
||||
std::nth_element(v.begin(), v.begin() + v.size() / 2, v.end());
|
||||
return v[v.size() / 2];
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char **argv) {
|
||||
if (argc < 2 || argv[1][0] == '-') {
|
||||
std::printf("usage: %s DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR] [--from S] [--to S]\n", argv[0]);
|
||||
return 1;
|
||||
}
|
||||
const std::string dir = argv[1];
|
||||
int oracle = 0, threads = 2;
|
||||
double slow = 1.0, from = 0, to = 1e9;
|
||||
bool cost = false;
|
||||
Contrast palm_contrast, hand_contrast{Contrast::None}; // as ft-hands's
|
||||
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
||||
PinchParams pinch_params;
|
||||
std::string use = "mono", color_left = "color_video0", color_crop = "subtract";
|
||||
std::string timeline, depth, poses, models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
|
||||
for (int i = 2; i < argc; ++i) {
|
||||
const std::string a = argv[i];
|
||||
const bool more = i + 1 < argc;
|
||||
if (a == "--oracle" && more) oracle = std::atoi(argv[++i]);
|
||||
else if (a == "--slow" && more) slow = std::atof(argv[++i]);
|
||||
else if (a == "--timeline" && more) timeline = argv[++i];
|
||||
else if (a == "--depth" && more) depth = argv[++i];
|
||||
else if (a == "--poses" && more) poses = argv[++i];
|
||||
else if (a == "--threads" && more) threads = std::atoi(argv[++i]);
|
||||
else if (a == "--models" && more) models = argv[++i];
|
||||
else if (a == "--cost") cost = true;
|
||||
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
|
||||
else if (a == "--cams" && more) use = argv[++i];
|
||||
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
|
||||
else if (a == "--pinch-palm-down" && more) pinch_params.palm_down_max = std::atof(argv[++i]);
|
||||
else if (a == "--color-left" && more) color_left = argv[++i];
|
||||
else if (a == "--color-crop" && more) color_crop = argv[++i];
|
||||
else if (a == "--contrast" && more) {
|
||||
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
|
||||
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
|
||||
}
|
||||
else if (a == "--from" && more) from = std::atof(argv[++i]);
|
||||
else if (a == "--to" && more) to = std::atof(argv[++i]);
|
||||
else return std::fprintf(stderr, "unknown option %s\n", a.c_str()), 1;
|
||||
}
|
||||
std::string err;
|
||||
std::map<std::string, Camera> calib;
|
||||
Nets nets;
|
||||
SetReader in;
|
||||
if (!load_calibration(calib, err) || !nets.load(models, false, err) || !in.open(dir, err))
|
||||
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
|
||||
nets.set_contrast(palm_contrast, hand_contrast);
|
||||
FILE *tl = timeline.empty() ? nullptr : std::fopen(timeline.c_str(), "w");
|
||||
|
||||
std::vector<fh_set_cam_t> cams;
|
||||
std::vector<std::vector<uint8_t>> px;
|
||||
if (!in.next(cams, px)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
|
||||
if (use != "mono" && use != "color" && use != "all") return std::fprintf(stderr, "--cams mono|color|all\n"), 1;
|
||||
if (use != "mono") {
|
||||
std::vector<std::string> nodes;
|
||||
for (auto &c : cams)
|
||||
if (std::string(c.name).rfind("color_video", 0) == 0) nodes.push_back(c.name);
|
||||
if (nodes.size() != 2) return std::fprintf(stderr, "%s: no color cameras (ft-camd --with-color)\n", dir.c_str()), 1;
|
||||
const std::string right = nodes[0] == color_left ? nodes[1] : nodes[0];
|
||||
if (!load_color_calibration(calib, color_left, right, color_crop == "subtract", 2, err))
|
||||
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
|
||||
}
|
||||
std::map<std::string, Camera> used;
|
||||
for (auto &c : cams) {
|
||||
const bool color = std::string(c.name).rfind("color_", 0) == 0;
|
||||
if (calib.count(c.name) && (use == "all" || color == (use == "color"))) used[c.name] = calib[c.name];
|
||||
}
|
||||
Pool pool(threads, {2, 3, 4});
|
||||
Tracker tracker(used, nets, pool);
|
||||
tracker.set_keep_presence(keep_presence);
|
||||
FILE *dp = depth.empty() ? nullptr : std::fopen(depth.c_str(), "w");
|
||||
FILE *pp = poses.empty() ? nullptr : std::fopen(poses.c_str(), "w");
|
||||
if (dp)
|
||||
for (auto &[name, c] : used)
|
||||
std::fprintf(dp, "# cam %s %.4f %.4f %.4f %.1f\n", name.c_str(), c.origin[0], c.origin[1], c.origin[2], c.fx);
|
||||
|
||||
uint64_t t0 = 0, busy_until = 0, next_ns = 0, t_prev = 0;
|
||||
int index = -1; // of the set in the recording
|
||||
int nsets = 0, processed = 0, left = 0, right = 0, both = 0, hist[3] = {};
|
||||
std::map<int, Track> tracks;
|
||||
std::vector<const Hand *> last_out;
|
||||
// oracle: sets where a side's hand was findable, and where the tracker had it then
|
||||
int o_sets = 0, o_left = 0, o_right = 0, o_left_hit = 0, o_right_hit = 0, o_left_extra = 0, o_right_extra = 0;
|
||||
std::map<std::string, int> o_by_cam;
|
||||
double busy_ms = 0;
|
||||
// jitter: how far each update's palm is from a straight line through the last two,
|
||||
// mm (steady motion cancels out; what's left is noise and real acceleration)
|
||||
std::vector<double> jit_raw, jit_sm;
|
||||
int near_face = 0, hand_updates = 0; // published palms within 20 cm of the eyes
|
||||
Pinch pinch(pinch_params);
|
||||
double pinch_begin_ts[2] = {0, 0};
|
||||
std::vector<double> pinch_len[2]; // seconds, per side
|
||||
int pinch_lost = 0;
|
||||
do {
|
||||
std::map<std::string, Image> images;
|
||||
// the set's time: the mono cameras' when they're used (the color ones run on another
|
||||
// clock); color frames can repeat across sets, so a set that doesn't move time on is skipped
|
||||
uint64_t t = UINT64_MAX, t_color = UINT64_MAX;
|
||||
for (size_t i = 0; i < cams.size(); ++i) {
|
||||
if (!used.count(cams[i].name)) continue;
|
||||
images[cams[i].name] = {px[i].data(), int(cams[i].width), int(cams[i].height), int(cams[i].width)};
|
||||
uint64_t &ti = std::string(cams[i].name).rfind("color_", 0) == 0 ? t_color : t;
|
||||
ti = std::min(ti, cams[i].capture_ns);
|
||||
}
|
||||
if (t == UINT64_MAX) t = t_color;
|
||||
if (t <= t_prev) {
|
||||
++index;
|
||||
continue;
|
||||
}
|
||||
t_prev = t;
|
||||
if (!t0) t0 = t;
|
||||
const double ts = (t - t0) / 1e9;
|
||||
++index;
|
||||
if (ts < from) continue;
|
||||
if (ts > to) break;
|
||||
++nsets;
|
||||
|
||||
if (t >= busy_until && t >= next_ns) {
|
||||
const auto w0 = std::chrono::steady_clock::now();
|
||||
const Stats before = tracker.stats;
|
||||
const auto out = tracker.step(images, int64_t(t));
|
||||
double ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - w0).count();
|
||||
if (cost) { // repeatable: rounds of model calls at typical live costs, per thread
|
||||
const int hands = tracker.stats.hand_calls - before.hand_calls, palms = tracker.stats.palm_calls - before.palm_calls;
|
||||
ms = (2 + 10.0 * ((hands + threads - 1) / threads) + 18.0 * ((palms + threads - 1) / threads)) / slow;
|
||||
}
|
||||
busy_ms += ms;
|
||||
busy_until = t + uint64_t(ms * slow * 1e6) + 3'000'000; // + the ring hand-off
|
||||
const std::vector<Seen> seen = tracker.views_now();
|
||||
pinch.update(out, seen, int64_t(t));
|
||||
next_ns = t + uint64_t((std::min(tracker.interval(), pinch.engaged() ? 1 / 30.0 : 1.0) - 0.005) * 1e9);
|
||||
for (const Pinch::Event &e : pinch.events) {
|
||||
if (std::string(e.what) == "begin") pinch_begin_ts[e.side] = ts;
|
||||
else pinch_len[e.side].push_back(ts - pinch_begin_ts[e.side]), pinch_lost += std::string(e.what) == "lost";
|
||||
if (tl) std::fprintf(tl, "%.3f pinch %s %s d %.3f point %+.3f %+.3f %+.3f hand %u\n", ts, e.side ? "R" : "L",
|
||||
e.what, e.distance, e.point[0], e.point[1], e.point[2], pinch.side(e.side).hand_id);
|
||||
}
|
||||
if (tl && (pinch.world_d[0] >= 0 || pinch.world_d[1] >= 0)) // both measures, for choosing one
|
||||
std::fprintf(tl, "%.3f pinchd L world %.3f tri %.3f R world %.3f tri %.3f palm %.2f %.2f\n", ts,
|
||||
pinch.world_d[0], pinch.tri_d[0], pinch.world_d[1], pinch.tri_d[1], pinch.palm_down[0],
|
||||
pinch.palm_down[1]);
|
||||
++processed;
|
||||
last_out = out;
|
||||
bool l = false, r = false;
|
||||
for (const Hand *h : out) {
|
||||
(h->pts[0][0] < 0 ? l : r) = true;
|
||||
Track &tr = tracks[h->id];
|
||||
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
|
||||
const V3 raw = palm(h->pts), sm = palm(h->smooth);
|
||||
++hand_updates, near_face += norm(sm) < 0.2;
|
||||
if (tr.line && ts - tr.t[0] >= 0.1) tr.line = 0; // a gap: the line starts over
|
||||
if (tr.line >= 2 && tr.t[0] - tr.t[1] > 1e-3) {
|
||||
const double k = (ts - tr.t[0]) / (tr.t[0] - tr.t[1]);
|
||||
jit_raw.push_back(norm(raw - tr.raw[0] - (tr.raw[0] - tr.raw[1]) * k) * 1000);
|
||||
jit_sm.push_back(norm(sm - tr.sm[0] - (tr.sm[0] - tr.sm[1]) * k) * 1000);
|
||||
}
|
||||
tr.raw[1] = tr.raw[0], tr.sm[1] = tr.sm[0], tr.t[1] = tr.t[0];
|
||||
tr.raw[0] = raw, tr.sm[0] = sm, tr.t[0] = ts;
|
||||
++tr.line;
|
||||
if (!tr.sets) tr.first = ts;
|
||||
tr.last = ts, ++tr.sets, tr.left += h->pts[0][0] < 0;
|
||||
if (tl)
|
||||
std::fprintf(tl, "%.3f %d %s %d %+.3f %+.3f %+.3f\n", ts, h->id, h->pts[0][0] < 0 ? "L" : "R", h->nviews,
|
||||
h->pts[0][0], h->pts[0][1], h->pts[0][2]);
|
||||
if (pp) {
|
||||
double world[21][3] = {};
|
||||
int n = 0;
|
||||
for (const Seen &v : seen)
|
||||
if (v.hand == h->id) {
|
||||
for (int k = 0; k < 21; ++k)
|
||||
for (int j = 0; j < 3; ++j) world[k][j] += v.lm.world[k][j];
|
||||
++n;
|
||||
}
|
||||
std::fprintf(pp, "%.4f %d %s %d %.3f", ts, h->id, h->right() ? "R" : "L", h->nviews, h->scale);
|
||||
for (int k = 0; k < 21; ++k)
|
||||
for (int j = 0; j < 3; ++j) std::fprintf(pp, " %.4f", n ? world[k][j] / n * h->scale : NAN);
|
||||
for (int k = 0; k < 21; ++k)
|
||||
for (int j = 0; j < 3; ++j) std::fprintf(pp, " %.4f", h->smooth[k][j]);
|
||||
std::fputc('\n', pp);
|
||||
}
|
||||
if (dp) {
|
||||
std::vector<const Seen *> vs;
|
||||
for (const Seen &v : seen)
|
||||
if (v.hand == h->id) vs.push_back(&v);
|
||||
std::sort(vs.begin(), vs.end(), [](const Seen *a, const Seen *b) { return a->cam < b->cam; });
|
||||
std::string names;
|
||||
for (const Seen *v : vs) names += (names.empty() ? "" : "+") + v->cam;
|
||||
std::fprintf(dp, "%.4f %d %s %d %s %.4f %.3f %.4f %.4f %.4f %.4f %.4f %.4f", ts, h->id,
|
||||
h->pts[0][0] < 0 ? "L" : "R", h->nviews, names.empty() ? "-" : names.c_str(), h->residual,
|
||||
h->scale, raw[0], raw[1], raw[2], sm[0], sm[1], sm[2]);
|
||||
for (const Seen *v : vs) {
|
||||
V3 mono[21];
|
||||
const bool ok = tracker.single_view(used.at(v->cam), v->lm, h->scale, mono);
|
||||
const V3 p = ok ? palm(mono) : V3{NAN, NAN, NAN};
|
||||
std::fprintf(dp, " %s %.2f %.4f %.4f %.4f", v->cam.c_str(), v->lm.presence, p[0], p[1], p[2]);
|
||||
}
|
||||
std::fputc('\n', dp);
|
||||
}
|
||||
}
|
||||
if (tl && out.empty()) std::fprintf(tl, "%.3f -\n", ts);
|
||||
if (tl)
|
||||
for (const Seen &v : seen)
|
||||
std::fprintf(tl, "%.3f view %d %s presence %.2f roi %.0f %.0f %.0f %.3f set %d\n", ts, v.hand, v.cam.c_str(),
|
||||
v.lm.presence, v.roi.center[0], v.roi.center[1], v.roi.size, v.roi.rotation, index);
|
||||
left += l, right += r, both += l && r;
|
||||
++hist[std::min<size_t>(out.size(), 2)];
|
||||
}
|
||||
|
||||
if (oracle > 0 && nsets % oracle == 0) {
|
||||
const Stats keep = tracker.stats;
|
||||
const auto seen = tracker.exhaustive(images);
|
||||
tracker.stats = keep;
|
||||
bool l = false, r = false;
|
||||
for (const Seen &s : seen) {
|
||||
(s.wrist[0] < 0 ? l : r) = true;
|
||||
++o_by_cam[s.cam + (s.wrist[0] < 0 ? " L" : " R")];
|
||||
}
|
||||
bool tl_ = false, tr_ = false;
|
||||
for (const Hand *h : last_out) (h->pts[0][0] < 0 ? tl_ : tr_) = true;
|
||||
++o_sets;
|
||||
o_left += l, o_right += r;
|
||||
o_left_hit += l && tl_, o_right_hit += r && tr_;
|
||||
o_left_extra += !l && tl_, o_right_extra += !r && tr_;
|
||||
if (tl && ((!l && tl_) || (!r && tr_))) std::fprintf(tl, "%.3f oracle-extra %s%s set %d\n", ts, !l && tl_ ? "L" : "", !r && tr_ ? "R" : "", index);
|
||||
}
|
||||
} while (in.next(cams, px));
|
||||
if (tl) std::fclose(tl);
|
||||
if (dp) std::fclose(dp);
|
||||
if (pp) std::fclose(pp);
|
||||
|
||||
const double secs = nsets > 1 ? nsets / 30.0 : 0;
|
||||
const Stats &s = tracker.stats;
|
||||
std::printf("%s: %d sets (%.0f s), processed %d (%.1f/s), %.1f ms per step\n", dir.c_str(), nsets, secs, processed,
|
||||
processed / std::max(secs, 1e-9), busy_ms / std::max(processed, 1));
|
||||
std::printf("hands per processed set: 0 %.0f%%, 1 %.0f%%, 2 %.0f%%; a hand on the left %.0f%%, right %.0f%%, both %.0f%%\n",
|
||||
100.0 * hist[0] / processed, 100.0 * hist[1] / processed, 100.0 * hist[2] / processed,
|
||||
100.0 * left / processed, 100.0 * right / processed, 100.0 * both / processed);
|
||||
std::vector<double> lens[2];
|
||||
for (auto &[id, tr] : tracks) lens[tr.left * 2 > tr.sets ? 0 : 1].push_back(tr.last - tr.first);
|
||||
for (int k = 0; k < 2; ++k) {
|
||||
double total = 0;
|
||||
for (double d : lens[k]) total += d;
|
||||
std::printf("%s tracks: %zu, median %.1f s, total %.0f s\n", k ? "right" : "left ", lens[k].size(), median(lens[k]), total);
|
||||
}
|
||||
std::printf("views lost %d, handoff misses %d, dups %d, splits %d; hands new %d, merged %d, forgotten %d\n", s.lost,
|
||||
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten);
|
||||
std::printf("model calls: palm %d (%.1f/s), hand %d (%.1f/s)\n", s.palm_calls, s.palm_calls / std::max(secs, 1e-9),
|
||||
s.hand_calls, s.hand_calls / std::max(secs, 1e-9));
|
||||
{
|
||||
std::vector<double> r, step;
|
||||
std::map<int, double> prev;
|
||||
for (auto &[id, x] : s.mono_ratio) {
|
||||
r.push_back(x);
|
||||
if (prev.count(id)) step.push_back(std::fabs(x - prev[id]));
|
||||
prev[id] = x;
|
||||
}
|
||||
std::sort(r.begin(), r.end());
|
||||
std::sort(step.begin(), step.end());
|
||||
if (!r.empty())
|
||||
std::printf("single-view distance / stereo: 10%% %.2f, median %.2f, 90%% %.2f; change between frames median %.3f, 90%% %.3f\n",
|
||||
r[r.size() / 10], r[r.size() / 2], r[r.size() * 9 / 10], step[step.size() / 2], step[step.size() * 9 / 10]);
|
||||
}
|
||||
std::printf("palms within 20 cm of the eyes: %d of %d hand updates\n", near_face, hand_updates);
|
||||
for (int k = 0; k < 2; ++k) std::sort(pinch_len[k].begin(), pinch_len[k].end());
|
||||
std::printf("pinches (%s, %.3f/%.3f m, palm down under %.2f): left %zu (median %.2f s), right %zu (median %.2f s), "
|
||||
"%d ended by losing the hand, held back (palm down) left %d right %d\n",
|
||||
pinch_params.triangulated ? "triangulated tips" : "world landmarks", pinch_params.begin_m, pinch_params.end_m,
|
||||
pinch_params.palm_down_max,
|
||||
pinch_len[0].size(), pinch_len[0].empty() ? 0 : pinch_len[0][pinch_len[0].size() / 2], pinch_len[1].size(),
|
||||
pinch_len[1].empty() ? 0 : pinch_len[1][pinch_len[1].size() / 2], pinch_lost, pinch.held_back[0],
|
||||
pinch.held_back[1]);
|
||||
std::sort(jit_raw.begin(), jit_raw.end());
|
||||
std::sort(jit_sm.begin(), jit_sm.end());
|
||||
if (!jit_raw.empty())
|
||||
std::printf("palm jitter (off a straight line through the last two updates): measured median %.1f mm, 90%% %.1f mm; "
|
||||
"published median %.1f mm, 90%% %.1f mm\n", jit_raw[jit_raw.size() / 2], jit_raw[jit_raw.size() * 9 / 10],
|
||||
jit_sm[jit_sm.size() / 2], jit_sm[jit_sm.size() * 9 / 10]);
|
||||
if (o_sets) {
|
||||
std::printf("oracle, %d sets: a left hand findable in %d, the tracker had it in %d (%.0f%%); right %d, had %d (%.0f%%)\n",
|
||||
o_sets, o_left, o_left_hit, 100.0 * o_left_hit / std::max(o_left, 1), o_right, o_right_hit,
|
||||
100.0 * o_right_hit / std::max(o_right, 1));
|
||||
std::printf(" tracker had a hand the full search didn't find: left %d, right %d\n", o_left_extra, o_right_extra);
|
||||
std::printf(" found by camera:");
|
||||
for (auto &[k, n] : o_by_cam) std::printf(" %s %d", k.c_str(), n);
|
||||
std::printf("\n");
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,233 @@
|
||||
// ft-ringplay: play a recording (ft-hands --record) into a frame ring in real time, the
|
||||
// way ft-camd publishes live cameras, so ft-hands --ring PATH processes the same frames
|
||||
// run after run. For A/B tests of how the tracker runs.
|
||||
//
|
||||
// ft-ringplay DIR --ring PATH [--from S] [--to S] [--loop] [--cpus 0,1]
|
||||
//
|
||||
// Frames are stamped as they're published, so the tracker's latency figures stay
|
||||
// meaningful. Cameras carry their calibration name and no device node (ft-hands maps
|
||||
// them by name), so a recording made with the right names needs no --swap-sides.
|
||||
// Dark frames (<name>_dk) are skipped. Needs no root: the ring is an ordinary file.
|
||||
#include "record.h"
|
||||
|
||||
extern "C" {
|
||||
#include "../camd/fhring.h"
|
||||
}
|
||||
|
||||
#include <fcntl.h>
|
||||
#include <sched.h>
|
||||
#include <sys/mman.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cerrno>
|
||||
#include <csignal>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <ctime>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
namespace {
|
||||
|
||||
volatile std::sig_atomic_t g_stop = 0;
|
||||
|
||||
uint64_t clock_ns(clockid_t id) {
|
||||
timespec ts;
|
||||
clock_gettime(id, &ts);
|
||||
return uint64_t(ts.tv_sec) * 1'000'000'000ull + uint64_t(ts.tv_nsec);
|
||||
}
|
||||
|
||||
// Sets from a recording, reading only the pixels of the cameras that get published.
|
||||
class Reader {
|
||||
public:
|
||||
bool open(const std::string &path) {
|
||||
f_ = std::fopen(path.c_str(), "rb");
|
||||
if (f_) posix_fadvise(fileno(f_), 0, 0, POSIX_FADV_SEQUENTIAL);
|
||||
return f_ != nullptr;
|
||||
}
|
||||
void rewind() { std::fseek(f_, 0, SEEK_SET); }
|
||||
// False at the end. cams: every camera in the set; px[k]: pixels of camera k when
|
||||
// want(name), else left empty.
|
||||
template <class Want>
|
||||
bool next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &px, Want want) {
|
||||
fh_set_hdr_t h;
|
||||
if (std::fread(&h, sizeof h, 1, f_) != 1 || std::memcmp(h.magic, FH_SET_MAGIC, 8) || h.ncams == 0 || h.ncams > 16)
|
||||
return false;
|
||||
cams.resize(h.ncams);
|
||||
if (std::fread(cams.data(), sizeof(fh_set_cam_t), h.ncams, f_) != h.ncams) return false;
|
||||
px.resize(h.ncams);
|
||||
for (uint32_t k = 0; k < h.ncams; ++k) {
|
||||
cams[k].name[sizeof cams[k].name - 1] = 0;
|
||||
const size_t n = size_t(cams[k].width) * cams[k].height;
|
||||
if (want(cams[k].name)) {
|
||||
px[k].resize(n);
|
||||
if (std::fread(px[k].data(), 1, n, f_) != n) return false;
|
||||
} else {
|
||||
px[k].clear();
|
||||
if (std::fseek(f_, long(n), SEEK_CUR)) return false;
|
||||
}
|
||||
}
|
||||
if (++sets_ % 64 == 0) posix_fadvise(fileno(f_), 0, std::ftell(f_), POSIX_FADV_DONTNEED); // RAM is tight
|
||||
return true;
|
||||
}
|
||||
|
||||
private:
|
||||
FILE *f_ = nullptr;
|
||||
uint64_t sets_ = 0;
|
||||
};
|
||||
|
||||
bool is_dark(const char *name) {
|
||||
const size_t n = std::strlen(name);
|
||||
return n > 3 && !std::strcmp(name + n - 3, "_dk");
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char **argv) {
|
||||
if (argc < 2 || argv[1][0] == '-') {
|
||||
std::fprintf(stderr, "usage: %s DIR --ring PATH [--from S] [--to S] [--loop] [--cpus 0,1]\n", argv[0]);
|
||||
return 1;
|
||||
}
|
||||
const std::string dir = argv[1];
|
||||
std::string ring_path;
|
||||
double from = 0, to = 1e9;
|
||||
bool loop = false;
|
||||
std::vector<int> cpus = {0, 1};
|
||||
for (int i = 2; i < argc; ++i) {
|
||||
const std::string a = argv[i];
|
||||
const bool more = i + 1 < argc;
|
||||
if (a == "--ring" && more) ring_path = argv[++i];
|
||||
else if (a == "--from" && more) from = std::atof(argv[++i]);
|
||||
else if (a == "--to" && more) to = std::atof(argv[++i]);
|
||||
else if (a == "--loop") loop = true;
|
||||
else if (a == "--cpus" && more) {
|
||||
cpus.clear();
|
||||
for (char *p = argv[++i]; *p;) {
|
||||
cpus.push_back(int(std::strtol(p, &p, 10)));
|
||||
if (*p == ',') ++p;
|
||||
else if (*p) break;
|
||||
}
|
||||
} else return std::fprintf(stderr, "unknown option %s\n", a.c_str()), 1;
|
||||
}
|
||||
if (ring_path.empty()) return std::fprintf(stderr, "--ring PATH is required\n"), 1;
|
||||
if (!cpus.empty()) {
|
||||
cpu_set_t set;
|
||||
CPU_ZERO(&set);
|
||||
for (int c : cpus) CPU_SET(c, &set);
|
||||
if (sched_setaffinity(0, sizeof set, &set) < 0) std::perror("sched_setaffinity");
|
||||
}
|
||||
std::signal(SIGINT, [](int) { g_stop = 1; });
|
||||
std::signal(SIGTERM, [](int) { g_stop = 1; });
|
||||
|
||||
Reader in;
|
||||
if (!in.open(dir + "/sets.bin")) return std::fprintf(stderr, "%s/sets.bin: %s\n", dir.c_str(), std::strerror(errno)), 1;
|
||||
std::vector<fh_set_cam_t> cams;
|
||||
std::vector<std::vector<uint8_t>> px;
|
||||
auto want = [](const char *name) { return !is_dark(name); };
|
||||
if (!in.next(cams, px, want)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
|
||||
auto set_time = [&](const std::vector<fh_set_cam_t> &cs) { // earliest bright capture, s
|
||||
uint64_t t = UINT64_MAX;
|
||||
for (const auto &c : cs)
|
||||
if (!is_dark(c.name)) t = std::min(t, c.capture_ns);
|
||||
return double(t) * 1e-9;
|
||||
};
|
||||
const double rec0 = set_time(cams);
|
||||
// skip to --from before the ring exists, so a reader never finds it without a heartbeat
|
||||
bool have = true;
|
||||
while (have && set_time(cams) - rec0 < from && !g_stop) have = in.next(cams, px, want);
|
||||
if (!have) return std::fprintf(stderr, "%s: nothing after %.1f s\n", dir.c_str(), from), 1;
|
||||
|
||||
// the ring: the recording's bright cameras, as ft-camd lays them out
|
||||
std::vector<int> pub; // set camera index of each ring camera
|
||||
for (size_t k = 0; k < cams.size() && pub.size() < FH_RING_MAX_CAMS; ++k)
|
||||
if (!is_dark(cams[k].name)) pub.push_back(int(k));
|
||||
size_t len = sizeof(fh_ring_hdr_t);
|
||||
std::vector<uint64_t> offset(pub.size()), slot_bytes(pub.size());
|
||||
for (size_t r = 0; r < pub.size(); ++r) {
|
||||
const fh_set_cam_t &c = cams[pub[r]];
|
||||
slot_bytes[r] = (sizeof(fh_ring_slot_t) + size_t(c.width) * c.height + 63) & ~size_t(63);
|
||||
offset[r] = len;
|
||||
len += FH_RING_SLOTS * slot_bytes[r];
|
||||
}
|
||||
const int fd = ::open(ring_path.c_str(), O_RDWR | O_CREAT | O_TRUNC | O_CLOEXEC, 0600);
|
||||
if (fd < 0 || ftruncate(fd, off_t(len)) < 0)
|
||||
return std::fprintf(stderr, "%s: %s\n", ring_path.c_str(), std::strerror(errno)), 1;
|
||||
void *m = mmap(nullptr, len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return std::fprintf(stderr, "mmap %s: %s\n", ring_path.c_str(), std::strerror(errno)), 1;
|
||||
auto *base = static_cast<uint8_t *>(m);
|
||||
auto *hdr = reinterpret_cast<fh_ring_hdr_t *>(base);
|
||||
for (size_t r = 0; r < pub.size(); ++r) {
|
||||
const fh_set_cam_t &c = cams[pub[r]];
|
||||
fh_ring_cam_t &rc = hdr->cams[r];
|
||||
std::snprintf(rc.sensor, sizeof rc.sensor, "ft-ringplay");
|
||||
std::snprintf(rc.name, sizeof rc.name, "%s", c.name);
|
||||
rc.node = -1;
|
||||
rc.format = FH_FMT_GREY8;
|
||||
rc.width = rc.stride = c.width;
|
||||
rc.height = c.height;
|
||||
rc.nslots = FH_RING_SLOTS;
|
||||
rc.slot_offset = offset[r];
|
||||
rc.slot_bytes = slot_bytes[r];
|
||||
}
|
||||
hdr->version = FH_RING_VERSION;
|
||||
hdr->header_bytes = sizeof(fh_ring_hdr_t);
|
||||
hdr->ncams = uint32_t(pub.size());
|
||||
hdr->file_bytes = len;
|
||||
hdr->writer_pid = getpid();
|
||||
std::memcpy(hdr->magic, FH_RING_MAGIC, 8);
|
||||
__atomic_store_n(&hdr->heartbeat_ns, clock_ns(CLOCK_MONOTONIC), __ATOMIC_RELEASE);
|
||||
std::printf("playing %s into %s:", dir.c_str(), ring_path.c_str());
|
||||
for (int k : pub) std::printf(" %s", cams[k].name);
|
||||
std::printf("\n");
|
||||
std::fflush(stdout);
|
||||
|
||||
uint64_t published = 0, rounds = 0;
|
||||
for (;;) {
|
||||
// one pass over [from, to]: each set goes out at its recorded offset from the first
|
||||
while (have && set_time(cams) - rec0 < from && !g_stop) {
|
||||
__atomic_store_n(&hdr->heartbeat_ns, clock_ns(CLOCK_MONOTONIC), __ATOMIC_RELEASE);
|
||||
have = in.next(cams, px, want);
|
||||
}
|
||||
const double first = set_time(cams);
|
||||
const uint64_t start = clock_ns(CLOCK_MONOTONIC);
|
||||
while (have && !g_stop && set_time(cams) - rec0 <= to) {
|
||||
const uint64_t due = start + uint64_t((set_time(cams) - first) * 1e9);
|
||||
for (uint64_t now = clock_ns(CLOCK_MONOTONIC); now < due && !g_stop; now = clock_ns(CLOCK_MONOTONIC)) {
|
||||
__atomic_store_n(&hdr->heartbeat_ns, now, __ATOMIC_RELEASE);
|
||||
const uint64_t wait = std::min<uint64_t>(due - now, 100'000'000);
|
||||
const timespec ts{time_t(wait / 1'000'000'000), long(wait % 1'000'000'000)};
|
||||
nanosleep(&ts, nullptr);
|
||||
}
|
||||
const uint64_t raw = clock_ns(CLOCK_MONOTONIC_RAW), mono = clock_ns(CLOCK_MONOTONIC);
|
||||
for (size_t r = 0; r < pub.size(); ++r) {
|
||||
fh_ring_cam_t &rc = hdr->cams[r];
|
||||
if (px[pub[r]].size() != size_t(rc.width) * rc.height) continue;
|
||||
const uint64_t n = rc.latest + 1;
|
||||
auto *s = reinterpret_cast<fh_ring_slot_t *>(base + rc.slot_offset + (n % rc.nslots) * rc.slot_bytes);
|
||||
__atomic_store_n(&s->seq, 2 * n + 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
std::memcpy(reinterpret_cast<uint8_t *>(s + 1), px[pub[r]].data(), px[pub[r]].size());
|
||||
s->frame = n;
|
||||
s->capture_ns = raw; // taken now, as a live camera's frame would be
|
||||
s->dqbuf_ns = mono;
|
||||
s->publish_ns = mono;
|
||||
__atomic_store_n(&s->seq, 2 * n + 2, __ATOMIC_RELEASE);
|
||||
__atomic_store_n(&rc.latest, n, __ATOMIC_RELEASE);
|
||||
++rc.published;
|
||||
}
|
||||
__atomic_store_n(&hdr->heartbeat_ns, mono, __ATOMIC_RELEASE);
|
||||
++published;
|
||||
have = in.next(cams, px, want);
|
||||
}
|
||||
++rounds;
|
||||
if (g_stop || !loop) break;
|
||||
in.rewind();
|
||||
have = in.next(cams, px, want);
|
||||
}
|
||||
std::printf("published %llu sets in %llu pass(es)\n", (unsigned long long)published, (unsigned long long)rounds);
|
||||
__atomic_store_n(&hdr->heartbeat_ns, 0, __ATOMIC_RELEASE); // readers see the writer gone
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,609 @@
|
||||
#include "tracker.h"
|
||||
|
||||
#include <pthread.h>
|
||||
#include <sched.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <set>
|
||||
|
||||
namespace {
|
||||
|
||||
// where arms start, head frame: below and slightly behind the eyes
|
||||
const V3 kShoulders[2] = {{0.17, -0.25, 0.08}, {-0.17, -0.25, 0.08}};
|
||||
// landmark pairs across the palm, rigid enough for single-view depth
|
||||
const int kPalmPairs[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 17}, {5, 13}, {9, 17}, {1, 17}, {1, 5}};
|
||||
constexpr double kFastSpeed = 0.25; // m/s
|
||||
constexpr double kSearchInterval = 0.2;
|
||||
// One Euro filter on the published landmarks: still hands are smoothed hard (tracking
|
||||
// noise is a few mm per frame), fast ones barely, so they don't lag.
|
||||
constexpr double kMinCutoff = 2.0; // Hz, a still hand
|
||||
constexpr double kBeta = 30.0; // Hz more per m/s of palm speed
|
||||
constexpr double kSpeedCutoff = 1.5; // Hz, for the palm speed itself
|
||||
// With one view, the hand's distance from the camera comes from how big it looks, which
|
||||
// is off by 10-30% and wanders ~10% between frames. Its direction is exact. So a hand that
|
||||
// was just located keeps its distance, drifting toward the one-view guess by this much a frame.
|
||||
constexpr double kMonoDepthGain = 0.1;
|
||||
// Is a triangulated hand as far from each camera as its apparent size says? With the
|
||||
// model's average hand, clean stereo pairs measure 0.71-1.51 times the one-view distance
|
||||
// (5-95%, median 1.16); pairs of two different hands mostly far less.
|
||||
constexpr double kSizePrior = 1.16, kRatioLo = 0.6, kRatioHi = 1.9;
|
||||
|
||||
double ms_since(std::chrono::steady_clock::time_point t) {
|
||||
return std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t).count();
|
||||
}
|
||||
|
||||
bool is_color(const Camera &c) { return c.name.rfind("color", 0) == 0; } // Arcturus, 145 degree image circle
|
||||
// The wide cameras: the side ones and the color ones
|
||||
bool is_slam(const Camera &c) { return c.name.rfind("slam", 0) == 0 || is_color(c); }
|
||||
|
||||
V2 palm_centre(const Landmarks &lm) { return lm.pts[9]; }
|
||||
|
||||
// Two views in one camera on the same hand: the landmark model puts the same points on
|
||||
// it from both crops, even when the crops differ.
|
||||
bool same_hand(const Landmarks &a, const Landmarks &b, double size) {
|
||||
double d = 0;
|
||||
for (int i = 0; i < 21; ++i) d += norm(a.pts[i] - b.pts[i]) / 21;
|
||||
return norm(palm_centre(a) - palm_centre(b)) < 0.5 * size || d < 0.25 * size;
|
||||
}
|
||||
|
||||
double hand_size(const Landmarks &lm) {
|
||||
double lo[2] = {1e9, 1e9}, hi[2] = {-1e9, -1e9};
|
||||
for (const V2 &p : lm.pts)
|
||||
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[k]), hi[k] = std::max(hi[k], p[k]);
|
||||
return std::max(hi[0] - lo[0], hi[1] - lo[1]);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
// ---------------------------------------------------------------------------- pool
|
||||
|
||||
Pool::Pool(int threads, const std::vector<int> &cpus) {
|
||||
for (int i = 0; i < threads; ++i) threads_.emplace_back(&Pool::loop, this, cpus[i % cpus.size()]);
|
||||
}
|
||||
|
||||
Pool::~Pool() {
|
||||
{
|
||||
std::lock_guard<std::mutex> l(mu_);
|
||||
stop_ = true;
|
||||
}
|
||||
wake_.notify_all();
|
||||
for (auto &t : threads_) t.join();
|
||||
}
|
||||
|
||||
void Pool::loop(int cpu) {
|
||||
cpu_set_t set;
|
||||
CPU_ZERO(&set);
|
||||
CPU_SET(cpu, &set);
|
||||
pthread_setaffinity_np(pthread_self(), sizeof set, &set); // ignored if not allowed
|
||||
std::unique_lock<std::mutex> l(mu_);
|
||||
for (;;) {
|
||||
wake_.wait(l, [&] { return stop_ || (jobs_ && next_ < jobs_->size()); });
|
||||
if (stop_) return;
|
||||
auto &job = (*jobs_)[next_++];
|
||||
l.unlock();
|
||||
job();
|
||||
l.lock();
|
||||
if (++finished_ == jobs_->size()) done_.notify_all();
|
||||
}
|
||||
}
|
||||
|
||||
void Pool::run(std::vector<std::function<void()>> &jobs) {
|
||||
if (jobs.empty()) return;
|
||||
std::unique_lock<std::mutex> l(mu_);
|
||||
jobs_ = &jobs, next_ = 0, finished_ = 0;
|
||||
wake_.notify_all();
|
||||
done_.wait(l, [&] { return finished_ == jobs.size(); });
|
||||
jobs_ = nullptr;
|
||||
}
|
||||
|
||||
// ------------------------------------------------------------------------- tracker
|
||||
|
||||
Tracker::Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views)
|
||||
: nets_(nets), pool_(pool), max_views_(max_views) {
|
||||
for (const auto &[name, cam] : cams) {
|
||||
cams_[name] = &cam;
|
||||
if (is_slam(cam)) {
|
||||
add_tiles(cam, 0.45, 3, 3);
|
||||
add_tiles(cam, 0.65, 2, 2);
|
||||
add_tiles(cam, 1.0, 1, 1); // hands close to the face fill much of the frame
|
||||
} else {
|
||||
add_tiles(cam, 0.6, 3, 2);
|
||||
add_tiles(cam, 1.0, 1, 1);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void Tracker::add_tiles(const Camera &cam, double frac, int gx, int gy) {
|
||||
const double s = frac * std::max(cam.width, cam.height);
|
||||
for (int i = 0; i < gx; ++i)
|
||||
for (int j = 0; j < gy; ++j) {
|
||||
const double x = gx > 1 ? s / 2 + (cam.width - s) * i / (gx - 1) : cam.width / 2.0;
|
||||
const double y = gy > 1 ? s / 2 + (cam.height - s) * j / (gy - 1) : cam.height / 2.0;
|
||||
Tile t{&cam, {x, y}, s, 0, 0};
|
||||
// turn the crop so the expected shoulder-to-hand direction points up
|
||||
const V3 ray = cam.ray(t.center), p = cam.origin + ray * 0.45;
|
||||
const V3 d = unit(p - kShoulders[p[0] > 0 ? 0 : 1]);
|
||||
const V2 a = cam.project(p, nullptr), b = cam.project(p + d * 0.05, nullptr);
|
||||
t.rotation = std::atan2(b[0] - a[0], -(b[1] - a[1]));
|
||||
t.weight = std::max(0.15, dot(ray, unit(V3{0, -0.45, -0.9})));
|
||||
tiles_.push_back(t);
|
||||
}
|
||||
}
|
||||
|
||||
double Tracker::interval() const {
|
||||
double fastest = -1;
|
||||
for (const auto &[id, h] : hands_)
|
||||
if (h.seen_ns == last_ns_) fastest = std::max(fastest, norm(h.dpalm)); // filtered: noise isn't speed
|
||||
return fastest < 0 ? 1 / 5.0 : fastest > kFastSpeed ? 1 / 30.0 : 1 / 15.0;
|
||||
}
|
||||
|
||||
bool Tracker::inside(const Camera &cam, V2 uv) const {
|
||||
const double m = 0.12;
|
||||
return uv[0] >= m * cam.width && uv[0] <= (1 - m) * cam.width && uv[1] >= m * cam.height &&
|
||||
uv[1] <= (1 - m) * cam.height && cam.off_axis(uv) < (is_color(cam) ? 70.0 : is_slam(cam) ? 80.0 : 75.0);
|
||||
}
|
||||
|
||||
void Tracker::run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views) {
|
||||
if (views.empty()) return;
|
||||
const auto t0 = std::chrono::steady_clock::now();
|
||||
std::vector<std::function<void()>> jobs;
|
||||
for (View *v : views) {
|
||||
const Image &img = images.at(v->cam->name);
|
||||
jobs.push_back([this, v, &img] {
|
||||
v->lm = nets_.landmarks(img, v->roi);
|
||||
v->has_lm = true;
|
||||
v->fresh = true;
|
||||
});
|
||||
}
|
||||
pool_.run(jobs);
|
||||
stats.hand_calls += int(views.size());
|
||||
stats.hand_ms += ms_since(t0);
|
||||
++stats.hand_batches;
|
||||
}
|
||||
|
||||
bool Tracker::single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const {
|
||||
V3 rays[21];
|
||||
for (int i = 0; i < 21; ++i) rays[i] = cam.ray(lm.pts[i]);
|
||||
std::vector<std::pair<double, double>> est; // (depth, weight)
|
||||
for (const auto &pr : kPalmPairs) {
|
||||
const int i = pr[0], j = pr[1];
|
||||
const double d = std::hypot(lm.world[i][0] - lm.world[j][0], lm.world[i][1] - lm.world[j][1]) * scale;
|
||||
const double a = std::acos(std::clamp(dot(rays[i], rays[j]), -1.0, 1.0));
|
||||
if (a > 1e-3 && d > 0.01) est.push_back({d / a, d});
|
||||
}
|
||||
if (est.empty()) return false;
|
||||
std::sort(est.begin(), est.end());
|
||||
double total = 0, acc = 0, depth = est.back().first;
|
||||
for (auto &e : est) total += e.second;
|
||||
for (auto &e : est)
|
||||
if ((acc += e.second) >= total / 2) { depth = e.first; break; }
|
||||
double zmean = 0;
|
||||
for (int i = 0; i < 21; ++i) zmean += lm.world[i][2] / 21;
|
||||
for (int i = 0; i < 21; ++i) out[i] = cam.origin + rays[i] * (depth + (lm.world[i][2] - zmean) * scale);
|
||||
return true;
|
||||
}
|
||||
|
||||
// A triangulated hand is in front of each camera, as far as its apparent size says (see
|
||||
// kSizePrior). Returns how far off that is (the sum of |log| ratios), or -1 if implausible.
|
||||
double Tracker::size_misfit(const std::vector<const View *> &views, const V3 *pts) const {
|
||||
double misfit = 0;
|
||||
for (const View *v : views) {
|
||||
V3 mono[21];
|
||||
// along the view's own ray (fisheye: a hand near the image edge is far off the axis)
|
||||
if (dot(pts[9] - v->cam->origin, v->cam->ray(v->lm.pts[9])) < 0.08) return -1;
|
||||
if (!single_view(*v->cam, v->lm, 1.0, mono)) continue;
|
||||
const double r = norm(pts[9] - v->cam->origin) / norm(mono[9] - v->cam->origin);
|
||||
if (r < kRatioLo || r > kRatioHi) return -1;
|
||||
misfit += std::fabs(std::log(r / kSizePrior));
|
||||
}
|
||||
return misfit;
|
||||
}
|
||||
|
||||
bool Tracker::hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns) {
|
||||
views.erase(std::remove_if(views.begin(), views.end(), [](View *v) { return !v->has_lm || !v->fresh; }), views.end());
|
||||
if (views.empty()) return false;
|
||||
if (views.size() >= 2) {
|
||||
const int n = int(views.size());
|
||||
std::vector<V3> origins(n), dirs(n);
|
||||
std::vector<double> w(n), res(21);
|
||||
V3 pts[21];
|
||||
for (int k = 0; k < 21; ++k) {
|
||||
for (int v = 0; v < n; ++v) {
|
||||
origins[v] = views[v]->cam->origin;
|
||||
dirs[v] = views[v]->cam->ray(views[v]->lm.pts[k]);
|
||||
w[v] = views[v]->lm.presence;
|
||||
}
|
||||
pts[k] = triangulate(origins.data(), dirs.data(), w.data(), n, &res[k]);
|
||||
}
|
||||
std::nth_element(res.begin(), res.begin() + 10, res.end());
|
||||
const double residual = res[10];
|
||||
// the views disagree: two different hands; keep the stronger. Rays to two different
|
||||
// hands can pass close to each other near the cameras, so check the distance too.
|
||||
if (residual > 0.03 || size_misfit({views.begin(), views.end()}, pts) < 0) {
|
||||
View *best = *std::max_element(views.begin(), views.end(),
|
||||
[](View *a, View *b) { return a->lm.presence < b->lm.presence; });
|
||||
for (View *v : views)
|
||||
if (v != best) v->hand = -1;
|
||||
++stats.splits;
|
||||
return hand_3d(hand, {best}, t_ns);
|
||||
}
|
||||
// learn how big this user's hand is compared to the model's average hand
|
||||
std::vector<double> t, m;
|
||||
const Landmarks &ref = views[0]->lm;
|
||||
for (const auto &pr : kPalmPairs) {
|
||||
t.push_back(norm(pts[pr[0]] - pts[pr[1]]));
|
||||
m.push_back(norm(V3{ref.world[pr[0]][0], ref.world[pr[0]][1], ref.world[pr[0]][2]} -
|
||||
V3{ref.world[pr[1]][0], ref.world[pr[1]][1], ref.world[pr[1]][2]}));
|
||||
}
|
||||
std::nth_element(t.begin(), t.begin() + t.size() / 2, t.end());
|
||||
std::nth_element(m.begin(), m.begin() + m.size() / 2, m.end());
|
||||
if (m[m.size() / 2] > 0 && residual < 0.008) // only from clean matches
|
||||
hand.scale += 0.1 * (std::clamp(t[t.size() / 2] / m[m.size() / 2], 0.8, 1.6) - hand.scale);
|
||||
std::copy(pts, pts + 21, hand.pts);
|
||||
hand.residual = residual;
|
||||
for (View *v : views) {
|
||||
V3 mono[21];
|
||||
if (!single_view(*v->cam, v->lm, hand.scale, mono)) continue;
|
||||
const V3 o = v->cam->origin;
|
||||
stats.mono_ratio.push_back({hand.id, norm(mono[9] - o) / norm(pts[9] - o)});
|
||||
}
|
||||
} else {
|
||||
const Camera &cam = *views[0]->cam;
|
||||
V3 pts[21];
|
||||
if (!single_view(cam, views[0]->lm, hand.scale, pts)) return false;
|
||||
const double guess = norm(pts[9] - cam.origin);
|
||||
if (hand.has_pts && t_ns - hand.seen_ns < 300'000'000 && guess > 0) {
|
||||
const double was = norm(hand.pts[9] - cam.origin), d = was + kMonoDepthGain * (guess - was);
|
||||
for (V3 &p : pts) p = cam.origin + (p - cam.origin) * (d / guess);
|
||||
}
|
||||
std::copy(pts, pts + 21, hand.pts);
|
||||
hand.residual = -1;
|
||||
}
|
||||
hand.has_pts = true;
|
||||
hand.nviews = int(views.size());
|
||||
for (View *v : views) hand.right_score += 0.2 * (v->lm.right - hand.right_score);
|
||||
return true;
|
||||
}
|
||||
|
||||
// How badly two views in different cameras fit one hand: the rays should meet, each view's
|
||||
// apparent size should match its distance, and the model should call both the same hand
|
||||
// (left or right). Negative if they can't be one hand. Side by side hands sit on the same
|
||||
// epipolar lines of the side cameras, so the distance check is what tells them apart.
|
||||
double Tracker::pair_cost(const View &a, const View &b) const {
|
||||
V3 pts[21];
|
||||
std::vector<double> res(21);
|
||||
for (int k = 0; k < 21; ++k) {
|
||||
const V3 o[2] = {a.cam->origin, b.cam->origin};
|
||||
const V3 d[2] = {a.cam->ray(a.lm.pts[k]), b.cam->ray(b.lm.pts[k])};
|
||||
const double w[2] = {a.lm.presence, b.lm.presence};
|
||||
pts[k] = triangulate(o, d, w, 2, &res[k]);
|
||||
}
|
||||
std::nth_element(res.begin(), res.begin() + 10, res.end());
|
||||
if (res[10] > 0.03) return -1;
|
||||
const double misfit = size_misfit({&a, &b}, pts);
|
||||
return misfit < 0 ? -1 : res[10] / 0.01 + misfit + std::fabs(a.lm.right - b.lm.right);
|
||||
}
|
||||
|
||||
// Which views in two cameras are the same hand: every way of pairing them up (a few views
|
||||
// each), scored with pair_cost. Keeps the hands' pairing unless another is clearly better,
|
||||
// then relabels the views, keeping the longer-tracked hand's id.
|
||||
void Tracker::associate() {
|
||||
constexpr double kPairBonus = 2.0, kBetter = 0.3;
|
||||
std::vector<const Camera *> cams;
|
||||
for (View &v : views_)
|
||||
if (std::find(cams.begin(), cams.end(), v.cam) == cams.end()) cams.push_back(v.cam);
|
||||
std::sort(cams.begin(), cams.end(), [](const Camera *a, const Camera *b) { return a->name < b->name; });
|
||||
for (size_t i = 0; i < cams.size(); ++i)
|
||||
for (size_t j = i + 1; j < cams.size(); ++j) {
|
||||
std::vector<View *> A, B;
|
||||
for (View &v : views_) {
|
||||
if (!v.fresh) continue;
|
||||
if (v.cam == cams[i]) A.push_back(&v);
|
||||
else if (v.cam == cams[j]) B.push_back(&v);
|
||||
}
|
||||
if (A.empty() || B.empty() || A.size() > 3 || B.size() > 3) continue;
|
||||
std::vector<std::vector<double>> c(A.size(), std::vector<double>(B.size()));
|
||||
for (size_t x = 0; x < A.size(); ++x)
|
||||
for (size_t y = 0; y < B.size(); ++y) c[x][y] = pair_cost(*A[x], *B[y]);
|
||||
auto score = [&](const std::vector<int> &m) { // m[x]: A[x]'s partner in B, or -1
|
||||
double s = 0;
|
||||
for (size_t x = 0; x < A.size(); ++x)
|
||||
if (m[x] >= 0 && c[x][m[x]] >= 0) s += c[x][m[x]] - kPairBonus;
|
||||
return s;
|
||||
};
|
||||
std::vector<int> cur(A.size(), -1);
|
||||
for (size_t x = 0; x < A.size(); ++x)
|
||||
for (size_t y = 0; y < B.size(); ++y)
|
||||
if (A[x]->hand == B[y]->hand) cur[x] = int(y);
|
||||
std::vector<int> best = cur, m(A.size(), -1);
|
||||
double best_score = score(cur);
|
||||
const double cur_score = best_score;
|
||||
std::function<void(size_t, unsigned)> walk = [&](size_t x, unsigned used) {
|
||||
if (x == A.size()) {
|
||||
const double sc = score(m);
|
||||
if (sc < best_score) best_score = sc, best = m;
|
||||
return;
|
||||
}
|
||||
m[x] = -1;
|
||||
walk(x + 1, used);
|
||||
for (size_t y = 0; y < B.size(); ++y)
|
||||
if (!(used >> y & 1) && c[x][y] >= 0) {
|
||||
m[x] = int(y);
|
||||
walk(x + 1, used | 1u << y);
|
||||
}
|
||||
m[x] = -1;
|
||||
};
|
||||
walk(0, 0);
|
||||
if (best == cur || best_score > cur_score - kBetter) continue;
|
||||
auto frames = [&](int id) {
|
||||
const auto h = hands_.find(id);
|
||||
return h == hands_.end() ? -1 : h->second.frames;
|
||||
};
|
||||
for (size_t x = 0; x < A.size(); ++x) {
|
||||
if (best[x] < 0) continue;
|
||||
View *a = A[x], *b = B[best[x]];
|
||||
int id = a->hand;
|
||||
bool free = true; // b's hand isn't another A view's
|
||||
for (size_t x2 = 0; x2 < A.size(); ++x2) free = free && (x2 == x || A[x2]->hand != b->hand);
|
||||
if (free && frames(b->hand) > frames(id)) id = b->hand;
|
||||
a->hand = b->hand = id;
|
||||
}
|
||||
// a B view left unpaired that still shares a hand with an A view starts its own
|
||||
for (size_t y = 0; y < B.size(); ++y) {
|
||||
if (std::find(best.begin(), best.end(), int(y)) != best.end()) continue;
|
||||
bool shared = false;
|
||||
for (View *a : A) shared = shared || a->hand == B[y]->hand;
|
||||
if (!shared) continue;
|
||||
B[y]->hand = next_id_++;
|
||||
hands_[B[y]->hand].id = B[y]->hand;
|
||||
++stats.created;
|
||||
}
|
||||
++stats.merged;
|
||||
}
|
||||
}
|
||||
|
||||
std::vector<const Hand *> Tracker::step(const std::map<std::string, Image> &images, int64_t t_ns) {
|
||||
const auto t_step = std::chrono::steady_clock::now();
|
||||
++stats.sets;
|
||||
std::vector<View> live;
|
||||
for (View &v : views_)
|
||||
if (images.count(v.cam->name)) live.push_back(v), live.back().fresh = false;
|
||||
|
||||
// 1. hand-over: give hands with too few views a crop in other cameras
|
||||
for (auto &[id, hand] : hands_) {
|
||||
if (!hand.has_pts) continue;
|
||||
std::set<std::string> have;
|
||||
for (View &v : live)
|
||||
if (v.hand == id) have.insert(v.cam->name);
|
||||
if (int(have.size()) >= max_views_) continue;
|
||||
std::vector<std::pair<double, View>> options;
|
||||
for (auto &[name, cam] : cams_) {
|
||||
if (have.count(name) || !images.count(name)) continue;
|
||||
V2 uv[21];
|
||||
bool front = true;
|
||||
for (int k = 0; k < 21; ++k) {
|
||||
double z;
|
||||
uv[k] = cam->project(hand.pts[k], &z);
|
||||
front = front && z > 0;
|
||||
}
|
||||
const V2 centre = (uv[0] + uv[5] + uv[9] + uv[13] + uv[17]) * 0.2;
|
||||
if (!front || !inside(*cam, centre)) continue;
|
||||
View v{cam, roi_from_points(uv), id, {}, false, 0};
|
||||
options.push_back({cam->off_axis(centre), v});
|
||||
}
|
||||
std::sort(options.begin(), options.end(), [](auto &a, auto &b) { return a.first < b.first; });
|
||||
for (size_t k = 0; k < options.size() && int(have.size() + k) < max_views_; ++k) live.push_back(options[k].second);
|
||||
}
|
||||
|
||||
// 2. the landmark model on each hand's best views, within budget
|
||||
std::map<int, std::vector<View *>> by_hand;
|
||||
for (View &v : live) by_hand[v.hand].push_back(&v);
|
||||
std::vector<View *> chosen;
|
||||
for (auto &[id, vs] : by_hand) {
|
||||
std::sort(vs.begin(), vs.end(), [](View *a, View *b) {
|
||||
if (a->has_lm != b->has_lm) return a->has_lm;
|
||||
return a->cam->off_axis(a->roi.center) < b->cam->off_axis(b->roi.center);
|
||||
});
|
||||
for (int k = 0; k < int(vs.size()) && k < max_views_; ++k) chosen.push_back(vs[k]);
|
||||
}
|
||||
std::stable_sort(chosen.begin(), chosen.end(), [](View *a, View *b) { return a->has_lm > b->has_lm; });
|
||||
if (int(chosen.size()) > hand_budget_) chosen.resize(hand_budget_);
|
||||
run_landmarks(images, chosen);
|
||||
std::vector<View *> kept;
|
||||
for (View *v : chosen)
|
||||
if (v->lm.presence >= (v->frames > 0 ? keep_presence_ : min_presence_)) {
|
||||
v->roi = v->lm.next_roi();
|
||||
++v->frames;
|
||||
kept.push_back(v);
|
||||
} else {
|
||||
++(v->frames > 0 ? stats.lost : stats.handoff_miss);
|
||||
}
|
||||
// the same hand twice in one camera: keep the more confident
|
||||
std::sort(kept.begin(), kept.end(), [](View *a, View *b) { return a->lm.presence > b->lm.presence; });
|
||||
std::vector<View> next;
|
||||
for (View *v : kept) {
|
||||
const double size = hand_size(v->lm);
|
||||
bool dup = false;
|
||||
for (View &o : next) dup = dup || (o.cam == v->cam && same_hand(o.lm, v->lm, size));
|
||||
if (!dup) next.push_back(*v);
|
||||
else ++stats.dups;
|
||||
}
|
||||
views_ = next;
|
||||
|
||||
// 3. search for missing hands
|
||||
std::set<int> tracked;
|
||||
for (View &v : views_) tracked.insert(v.hand);
|
||||
if (tracked.size() < 2 && (t_ns - search_ns_) / 1e9 >= kSearchInterval - 0.01) {
|
||||
search_ns_ = t_ns;
|
||||
const int budget = tracked.empty() ? search_budget_ : std::max(1, search_budget_ - 1);
|
||||
for (Tile &t : tiles_)
|
||||
if (images.count(t.cam->name)) t.credit += t.weight;
|
||||
std::vector<Tile *> picked;
|
||||
for (int b = 0; b < budget; ++b) {
|
||||
Tile *best = nullptr;
|
||||
for (Tile &t : tiles_)
|
||||
if (images.count(t.cam->name) && std::find(picked.begin(), picked.end(), &t) == picked.end() &&
|
||||
(!best || t.credit > best->credit))
|
||||
best = &t;
|
||||
if (!best) break;
|
||||
best->credit = 0;
|
||||
picked.push_back(best);
|
||||
}
|
||||
const auto t0 = std::chrono::steady_clock::now();
|
||||
std::vector<std::vector<Palm>> found(picked.size());
|
||||
std::vector<std::function<void()>> jobs;
|
||||
for (size_t i = 0; i < picked.size(); ++i) {
|
||||
Tile *t = picked[i];
|
||||
const Image &img = images.at(t->cam->name);
|
||||
jobs.push_back([this, t, &img, &found, i] { found[i] = nets_.palms(img, t->center, t->size, t->rotation); });
|
||||
}
|
||||
pool_.run(jobs);
|
||||
stats.palm_calls += int(picked.size());
|
||||
stats.palm_ms += ms_since(t0);
|
||||
++stats.palm_batches;
|
||||
std::vector<View> fresh;
|
||||
for (size_t i = 0; i < picked.size(); ++i)
|
||||
for (const Palm &p : found[i]) {
|
||||
const Roi roi = p.roi();
|
||||
bool near = false;
|
||||
for (auto *list : {&views_, &fresh})
|
||||
for (View &v : *list) near = near || (v.cam == picked[i]->cam && norm(v.roi.center - roi.center) < 0.5 * roi.size);
|
||||
if (!near) fresh.push_back({picked[i]->cam, roi, 0, {}, false, 0});
|
||||
}
|
||||
std::vector<View *> ptrs;
|
||||
for (View &v : fresh) ptrs.push_back(&v);
|
||||
run_landmarks(images, ptrs);
|
||||
for (View &v : fresh)
|
||||
if (v.lm.presence >= min_presence_) {
|
||||
v.roi = v.lm.next_roi();
|
||||
v.frames = 1;
|
||||
views_.push_back(v);
|
||||
}
|
||||
}
|
||||
|
||||
// 4. give new views a hand: the nearest existing hand in 3D, else a new one
|
||||
for (View &v : views_) {
|
||||
if (v.hand > 0 && hands_.count(v.hand)) continue;
|
||||
V3 guess[21];
|
||||
const bool have_guess = single_view(*v.cam, v.lm, 1.0, guess);
|
||||
int best = 0;
|
||||
double dist = 0.12;
|
||||
for (auto &[id, h] : hands_) {
|
||||
if (!h.has_pts) continue;
|
||||
bool same_cam = false;
|
||||
for (View &o : views_) same_cam = same_cam || (o.hand == id && o.cam == v.cam);
|
||||
if (same_cam) continue;
|
||||
const double d = have_guess ? norm(h.pts[9] - guess[9]) : 1e9;
|
||||
if (d < dist) best = id, dist = d;
|
||||
}
|
||||
if (!best) {
|
||||
best = next_id_++;
|
||||
hands_[best].id = best;
|
||||
++stats.created;
|
||||
}
|
||||
v.hand = best;
|
||||
}
|
||||
|
||||
// 5. which views in different cameras are the same hand
|
||||
associate();
|
||||
|
||||
// 6. 3D for every hand seen now; forget hands not seen for a while
|
||||
std::vector<const Hand *> out;
|
||||
for (auto it = hands_.begin(); it != hands_.end();) {
|
||||
Hand &h = it->second;
|
||||
std::vector<View *> vs;
|
||||
for (View &v : views_)
|
||||
if (v.hand == h.id) vs.push_back(&v);
|
||||
if (!vs.empty() && hand_3d(h, vs, t_ns)) {
|
||||
const V3 palm = (h.pts[0] + h.pts[5] + h.pts[9] + h.pts[13] + h.pts[17]) * 0.2;
|
||||
if (h.last_ns && t_ns > h.last_ns)
|
||||
h.speed += 0.5 * (std::min(norm(palm - h.last_palm) / ((t_ns - h.last_ns) / 1e9), 5.0) - h.speed);
|
||||
h.last_ns = t_ns, h.last_palm = palm, h.seen_ns = t_ns;
|
||||
++h.frames;
|
||||
smooth(h, t_ns);
|
||||
out.push_back(&h);
|
||||
++it;
|
||||
} else if (t_ns - h.seen_ns > 300'000'000) {
|
||||
it = hands_.erase(it);
|
||||
++stats.forgotten;
|
||||
} else {
|
||||
++it;
|
||||
}
|
||||
}
|
||||
// views split off by a failed triangulation start over as new hands next frame
|
||||
for (View &v : views_)
|
||||
if (v.hand <= 0) {
|
||||
v.hand = next_id_++;
|
||||
hands_[v.hand].id = v.hand;
|
||||
++stats.created;
|
||||
}
|
||||
last_ns_ = t_ns;
|
||||
stats.step_ms += ms_since(t_step);
|
||||
return out;
|
||||
}
|
||||
|
||||
std::vector<Seen> Tracker::views_now() const {
|
||||
std::vector<Seen> out;
|
||||
for (const View &v : views_) out.push_back({v.cam->name, v.hand, v.roi, v.lm, {}});
|
||||
return out;
|
||||
}
|
||||
|
||||
std::vector<Seen> Tracker::exhaustive(const std::map<std::string, Image> &images) {
|
||||
std::vector<Tile *> tiles;
|
||||
for (Tile &t : tiles_)
|
||||
if (images.count(t.cam->name)) tiles.push_back(&t);
|
||||
std::vector<std::vector<Palm>> found(tiles.size());
|
||||
std::vector<std::function<void()>> jobs;
|
||||
for (size_t i = 0; i < tiles.size(); ++i)
|
||||
jobs.push_back([this, &tiles, &images, &found, i] {
|
||||
const Tile *t = tiles[i];
|
||||
found[i] = nets_.palms(images.at(t->cam->name), t->center, t->size, t->rotation);
|
||||
});
|
||||
pool_.run(jobs);
|
||||
// one crop per palm: tiles overlap, so the same palm turns up several times
|
||||
std::vector<std::pair<double, View>> palms;
|
||||
for (size_t i = 0; i < tiles.size(); ++i)
|
||||
for (const Palm &p : found[i]) palms.push_back({p.score, View{tiles[i]->cam, p.roi(), 0, {}, false, 0}});
|
||||
std::sort(palms.begin(), palms.end(), [](auto &a, auto &b) { return a.first > b.first; });
|
||||
std::vector<View> crops;
|
||||
for (auto &[score, v] : palms) {
|
||||
bool near = false;
|
||||
for (View &o : crops) near = near || (o.cam == v.cam && norm(o.roi.center - v.roi.center) < 0.5 * v.roi.size);
|
||||
if (!near) crops.push_back(v);
|
||||
}
|
||||
std::vector<View *> ptrs;
|
||||
for (View &v : crops) ptrs.push_back(&v);
|
||||
run_landmarks(images, ptrs);
|
||||
std::sort(crops.begin(), crops.end(), [](const View &a, const View &b) { return a.lm.presence > b.lm.presence; });
|
||||
std::vector<Seen> out;
|
||||
for (View &v : crops) {
|
||||
if (v.lm.presence < min_presence_) continue;
|
||||
bool dup = false;
|
||||
for (const Seen &o : out)
|
||||
dup = dup || (o.cam == v.cam->name && norm(palm_centre(o.lm) - palm_centre(v.lm)) < 0.5 * hand_size(v.lm));
|
||||
if (dup) continue;
|
||||
V3 pts[21];
|
||||
Seen s{v.cam->name, 0, v.roi, v.lm, {}};
|
||||
if (single_view(*v.cam, v.lm, 1.0, pts)) s.wrist = pts[0];
|
||||
out.push_back(s);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
void Tracker::smooth(Hand &h, int64_t t_ns) {
|
||||
const double dt = (t_ns - h.smooth_ns) / 1e9;
|
||||
h.smooth_ns = t_ns;
|
||||
if (h.frames <= 1 || dt <= 0 || dt > 0.3) { // new, or back after a gap: start over
|
||||
std::copy(h.pts, h.pts + 21, h.smooth);
|
||||
h.dpalm = {0, 0, 0};
|
||||
return;
|
||||
}
|
||||
auto alpha = [dt](double cutoff) { return 1 / (1 + 1 / (2 * M_PI * cutoff * dt)); };
|
||||
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
|
||||
const V3 d = (palm(h.pts) - palm(h.smooth)) * (1 / dt);
|
||||
h.dpalm = h.dpalm + (d - h.dpalm) * alpha(kSpeedCutoff);
|
||||
// one cutoff for the whole hand, from its palm speed, so its shape stays together
|
||||
const double a = alpha(kMinCutoff + kBeta * norm(h.dpalm));
|
||||
for (int i = 0; i < 21; ++i) h.smooth[i] = h.smooth[i] + (h.pts[i] - h.smooth[i]) * a;
|
||||
}
|
||||
@@ -0,0 +1,131 @@
|
||||
// Multi-camera hand tracking (a port of frame-hands' Python prototype; the scheduling is
|
||||
// described in hands/README.md). All 3D is metres in the head frame.
|
||||
#pragma once
|
||||
|
||||
#include "calib.h"
|
||||
#include "nets.h"
|
||||
|
||||
#include <condition_variable>
|
||||
#include <functional>
|
||||
#include <map>
|
||||
#include <memory>
|
||||
#include <mutex>
|
||||
#include <thread>
|
||||
#include <vector>
|
||||
|
||||
// Runs batches of jobs on a few threads, each pinned to a core.
|
||||
class Pool {
|
||||
public:
|
||||
// One thread per entry of cpus, pinned there (round-robin if threads > cpus).
|
||||
Pool(int threads, const std::vector<int> &cpus);
|
||||
~Pool();
|
||||
void run(std::vector<std::function<void()>> &jobs);
|
||||
|
||||
private:
|
||||
void loop(int cpu);
|
||||
std::vector<std::thread> threads_;
|
||||
std::mutex mu_;
|
||||
std::condition_variable wake_, done_;
|
||||
std::vector<std::function<void()>> *jobs_ = nullptr;
|
||||
size_t next_ = 0, finished_ = 0;
|
||||
bool stop_ = false;
|
||||
};
|
||||
|
||||
struct Hand {
|
||||
int id = 0;
|
||||
V3 pts[21]{}; // as measured this frame; the tracker steers crops by these
|
||||
V3 smooth[21]{}; // filtered (One Euro, see Tracker::smooth): publish these
|
||||
bool has_pts = false;
|
||||
double residual = -1; // rms ray distance of the triangulation (m); -1: one view
|
||||
int nviews = 0;
|
||||
double right_score = 0.5; // the model's right-hand score (these images aren't mirrored)
|
||||
double scale = 1.0; // this user's hand size / the model's world landmarks
|
||||
int64_t seen_ns = 0;
|
||||
int frames = 0;
|
||||
double speed = 0; // palm centre, m/s, smoothed
|
||||
int64_t last_ns = 0;
|
||||
V3 last_palm{};
|
||||
V3 dpalm{}; // the filter's palm velocity, m/s
|
||||
int64_t smooth_ns = 0;
|
||||
bool right() const { return right_score > 0.5; }
|
||||
};
|
||||
|
||||
struct Stats {
|
||||
int palm_calls = 0, hand_calls = 0, sets = 0;
|
||||
double palm_ms = 0, hand_ms = 0, step_ms = 0; // summed batch times
|
||||
int palm_batches = 0, hand_batches = 0;
|
||||
// why views and hands come and go
|
||||
int lost = 0; // a tracked view's landmarks fell below min presence
|
||||
int handoff_miss = 0; // a view projected from the hand's 3D (new camera or retry) found no hand
|
||||
int dups = 0; // the same hand twice in one camera
|
||||
int splits = 0; // a hand's views disagreed in 3D and were split
|
||||
int created = 0, merged = 0, forgotten = 0; // merged: views re-paired across cameras
|
||||
// diagnostics: on stereo frames, each view's single-view palm distance / the stereo one
|
||||
std::vector<std::pair<int, double>> mono_ratio; // (hand id, ratio)
|
||||
};
|
||||
|
||||
// A hand the landmark model found in one camera (Tracker::views_now, Tracker::exhaustive).
|
||||
struct Seen {
|
||||
std::string cam;
|
||||
int hand = 0; // the tracker's hand; 0 in exhaustive()
|
||||
Roi roi;
|
||||
Landmarks lm;
|
||||
V3 wrist{}; // exhaustive(): single-view 3D guess at the model's hand size
|
||||
};
|
||||
|
||||
class Tracker {
|
||||
public:
|
||||
Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views = 2);
|
||||
// images: calibration name -> frame. Returns the hands seen in this set.
|
||||
std::vector<const Hand *> step(const std::map<std::string, Image> &images, int64_t t_ns);
|
||||
// Seconds until the next frame set is worth processing (30 Hz fast hands, 15 Hz slow, 5 Hz none).
|
||||
double interval() const;
|
||||
Stats stats;
|
||||
size_t views() const { return views_.size(); }
|
||||
std::vector<Seen> views_now() const;
|
||||
// Every search tile in every camera, then landmarks on every palm: slow; for checking
|
||||
// what the scheduler misses (ft-handreplay --oracle).
|
||||
std::vector<Seen> exhaustive(const std::map<std::string, Image> &images);
|
||||
// Landmark presence a tracked view needs to stay (new views need min presence, 0.5). In
|
||||
// bright rooms the camera exposes for the room, the hands come out dim, and presence
|
||||
// dips under 0.5 for a frame at a time.
|
||||
void set_keep_presence(double p) { keep_presence_ = p; }
|
||||
// One view's 3D hand: each landmark along its ray, as far as how big the palm looks says
|
||||
// for a hand `scale` times the model's (Hand::scale). False if the palm is degenerate.
|
||||
bool single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const;
|
||||
|
||||
private:
|
||||
struct View {
|
||||
const Camera *cam;
|
||||
Roi roi;
|
||||
int hand = 0; // 0: not assigned yet
|
||||
Landmarks lm;
|
||||
bool has_lm = false;
|
||||
int frames = 0;
|
||||
bool fresh = false; // lm is from this step
|
||||
};
|
||||
struct Tile {
|
||||
const Camera *cam;
|
||||
V2 center;
|
||||
double size, rotation, weight, credit = 0;
|
||||
};
|
||||
void run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views);
|
||||
bool hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns);
|
||||
double pair_cost(const View &a, const View &b) const;
|
||||
double size_misfit(const std::vector<const View *> &views, const V3 *pts) const;
|
||||
void associate();
|
||||
static void smooth(Hand &h, int64_t t_ns);
|
||||
bool inside(const Camera &cam, V2 uv) const;
|
||||
void add_tiles(const Camera &cam, double frac, int gx, int gy);
|
||||
|
||||
std::map<std::string, const Camera *> cams_;
|
||||
const Nets &nets_;
|
||||
Pool &pool_;
|
||||
int max_views_, hand_budget_ = 4, search_budget_ = 3;
|
||||
double min_presence_ = 0.5, keep_presence_ = 0.5;
|
||||
std::vector<View> views_;
|
||||
std::map<int, Hand> hands_;
|
||||
std::vector<Tile> tiles_;
|
||||
int next_id_ = 1;
|
||||
int64_t last_ns_ = 0, search_ns_ = 0;
|
||||
};
|
||||
+17
-10
@@ -1,8 +1,8 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install everything on the Steam Frame: the build container, Frametop (multi-screen
|
||||
# desktop, input relay, universal 3D mouse, settings app), and optionally the Bluetooth
|
||||
# fixes. Run it on the headset in a terminal, from this repo. It's safe to re-run, for
|
||||
# example after `git pull`.
|
||||
# fixes and hand tracking. Run it on the headset in a terminal, from this repo. It's safe
|
||||
# to re-run, for example after `git pull`.
|
||||
# (It also works from a PC over SSH; see "Developing from a PC" in the README.)
|
||||
#
|
||||
# Usage: ./install.sh [--yes] [--no-bluetooth]
|
||||
@@ -45,7 +45,7 @@ else
|
||||
"$root/scripts/sync.sh" >/dev/null
|
||||
fi
|
||||
|
||||
step "1/8 distrobox (container tool, installed in your home folder)"
|
||||
step "1/9 distrobox (container tool, installed in your home folder)"
|
||||
if on_frame 'test -x ~/.local/bin/distrobox'; then
|
||||
echo "already installed: $(on_frame '~/.local/bin/distrobox version | head -1')"
|
||||
else
|
||||
@@ -55,25 +55,25 @@ else
|
||||
cd ~/dev/src/distrobox && ./install --prefix ~/.local'
|
||||
fi
|
||||
|
||||
step "2/8 build container (Fedora 44 'dev', about 1-2 GB the first time)"
|
||||
step "2/9 build container (Fedora 44 'dev', about 1-2 GB the first time)"
|
||||
"$root/setup/dev-container.sh"
|
||||
|
||||
step "3/8 input relay (keeps Bluetooth mice working in SteamVR, device roles, button maps)"
|
||||
step "3/9 input relay (keeps Bluetooth mice working in SteamVR, device roles, button maps)"
|
||||
"$root/desktops.sh" relay install
|
||||
|
||||
step "4/8 3D mouse: SteamVR driver"
|
||||
step "4/9 3D mouse: SteamVR driver"
|
||||
"$root/pointer/driver/build.sh"
|
||||
"$root/pointer/driver/install.sh" install 2>&1 | grep -v xdg-open
|
||||
|
||||
step "5/8 3D mouse: pointer helper service"
|
||||
step "5/9 3D mouse: pointer helper service"
|
||||
"$root/pointer/helper/build.sh"
|
||||
"$root/pointer/helper/run.sh" install
|
||||
|
||||
step "6/8 power service (turns the displays off while the headset isn't used, even on a stand)"
|
||||
step "6/9 power service (turns the displays off while the headset isn't used, even on a stand)"
|
||||
"$root/power/build.sh"
|
||||
"$root/power/run.sh" install
|
||||
|
||||
step "7/8 multi-screen desktop (ft-screens), Frametop Input Settings, and Frametop Display Settings"
|
||||
step "7/9 multi-screen desktop (ft-screens), Frametop Input Settings, and Frametop Display Settings"
|
||||
"$root/screens/build.sh"
|
||||
"$root/desktops.sh" install >/dev/null
|
||||
"$root/input-settings/install.sh"
|
||||
@@ -81,13 +81,20 @@ step "7/8 multi-screen desktop (ft-screens), Frametop Input Settings, and Framet
|
||||
on_frame "sed -i 's/^POINTER=0/POINTER=1/' ~/.config/frametop.conf; grep -q '^POINTER=' ~/.config/frametop.conf || echo 'POINTER=1' >> ~/.config/frametop.conf"
|
||||
echo "the launcher's Desktop entry now opens the multi-screen desktop; 3D mouse on (POINTER=1 in ~/.config/frametop.conf)"
|
||||
|
||||
step "8/8 Bluetooth fixes (optional; they let LE mice and keyboards like the Swiftpoint Z3 reconnect)"
|
||||
step "8/9 Bluetooth fixes (optional; they let LE mice and keyboards like the Swiftpoint Z3 reconnect)"
|
||||
if [ "$bluetooth" = 1 ] && ask "Install the Bluetooth fixes? They need your password (sudo)." n; then
|
||||
"$root/setup/bluetooth/install.sh" install
|
||||
else
|
||||
echo "skipped. Install later with: setup/bluetooth/install.sh install"
|
||||
fi
|
||||
|
||||
step "9/9 hand tracking (optional, experimental: your hands show over the screens)"
|
||||
if ask "Install hand tracking? It needs your password (sudo) to let its camera service read the headset cameras." n; then
|
||||
"$root/hands/run.sh" install
|
||||
else
|
||||
echo "skipped. Install later with: hands/run.sh install"
|
||||
fi
|
||||
|
||||
step "Done"
|
||||
cat <<'EOF'
|
||||
SteamVR has to restart once, to load the 3D mouse driver and to start the input relay
|
||||
|
||||
+29
-39
@@ -1,6 +1,8 @@
|
||||
// Hand cutouts (see handcut.h).
|
||||
#include "handcut.h"
|
||||
|
||||
#include "../hands/include/fh_hands.h"
|
||||
|
||||
#include <EGL/egl.h>
|
||||
#include <EGL/eglext.h>
|
||||
#include <GLES2/gl2.h>
|
||||
@@ -23,10 +25,8 @@
|
||||
namespace handcut {
|
||||
namespace {
|
||||
|
||||
// The hands file (frame-hands/include/fh_hands.h).
|
||||
constexpr char kMagic[8] = {'F', 'H', 'H', 'A', 'N', 'D', 'S', '1'};
|
||||
constexpr size_t kHeader = 64, kHand = 272, kCapsule = 32, kMaxHands = 2, kMaxCapsules = 64;
|
||||
constexpr size_t kFileSize = kHeader + kMaxHands * kHand + kMaxCapsules * kCapsule;
|
||||
// The hands file ft-hands publishes (hands/include/fh_hands.h).
|
||||
constexpr uint32_t kMaxHands = FH_HANDS_MAX_HANDS, kMaxCapsules = FH_HANDS_MAX_CAPSULES;
|
||||
constexpr int64_t kStaleNs = 300'000'000; // hands older than this are gone
|
||||
constexpr int64_t kHistoryNs = 1'000'000'000;
|
||||
constexpr double kNear = 0.12; // metres: nothing closer to an eye than this is cut
|
||||
@@ -59,53 +59,44 @@ bool Hands::Read() {
|
||||
const int64_t now = MonoNs();
|
||||
if (now - lastOpenTry_ < 1'000'000'000) return false;
|
||||
lastOpenTry_ = now;
|
||||
const char *run = std::getenv("XDG_RUNTIME_DIR");
|
||||
const std::string path = std::string(run ? run : "/run/user/" + std::to_string(getuid())) + "/frame-hands/hands";
|
||||
const std::string path = "/run/user/" + std::to_string(getuid()) + "/frametop-hands/hands";
|
||||
fd_ = open(path.c_str(), O_RDONLY | O_CLOEXEC | O_NOFOLLOW);
|
||||
if (fd_ < 0) return false;
|
||||
struct stat st;
|
||||
if (fstat(fd_, &st) < 0 || st.st_uid != getuid() || size_t(st.st_size) < kFileSize) {
|
||||
if (fstat(fd_, &st) < 0 || st.st_uid != getuid() || size_t(st.st_size) < sizeof(fh_hands_t)) {
|
||||
close(fd_), fd_ = -1;
|
||||
return false;
|
||||
}
|
||||
void *m = mmap(nullptr, kFileSize, PROT_READ, MAP_SHARED, fd_, 0);
|
||||
void *m = mmap(nullptr, sizeof(fh_hands_t), PROT_READ, MAP_SHARED, fd_, 0);
|
||||
if (m == MAP_FAILED) {
|
||||
close(fd_), fd_ = -1;
|
||||
return false;
|
||||
}
|
||||
map_ = m;
|
||||
}
|
||||
const auto *p = static_cast<const volatile uint8_t *>(map_);
|
||||
auto u64 = [&](size_t off) { uint64_t v; std::memcpy(&v, const_cast<const uint8_t *>(p) + off, 8); return v; };
|
||||
const uint64_t s1 = __atomic_load_n(reinterpret_cast<const uint64_t *>(const_cast<const uint8_t *>(p) + 16), __ATOMIC_ACQUIRE);
|
||||
const auto *file = static_cast<const fh_hands_t *>(map_);
|
||||
const auto *seq = const_cast<const uint64_t *>(&file->seq);
|
||||
const uint64_t s1 = __atomic_load_n(seq, __ATOMIC_ACQUIRE);
|
||||
if ((s1 & 1) || s1 == seq_) return false;
|
||||
uint8_t copy[kFileSize];
|
||||
std::memcpy(copy, const_cast<const uint8_t *>(p), kFileSize);
|
||||
fh_hands_t copy;
|
||||
std::memcpy(static_cast<void *>(©), map_, sizeof copy);
|
||||
__atomic_thread_fence(__ATOMIC_ACQUIRE);
|
||||
if (u64(16) != s1 || std::memcmp(copy, kMagic, 8) != 0) return false;
|
||||
if (__atomic_load_n(seq, __ATOMIC_RELAXED) != s1 || std::memcmp(copy.magic, FH_HANDS_MAGIC, 8) != 0) return false;
|
||||
seq_ = s1;
|
||||
uint64_t capture, publish;
|
||||
uint32_t nhands, ncaps;
|
||||
std::memcpy(&capture, copy + 24, 8);
|
||||
std::memcpy(&publish, copy + 32, 8);
|
||||
std::memcpy(&nhands, copy + 40, 4);
|
||||
std::memcpy(&ncaps, copy + 44, 4);
|
||||
captureNs_ = int64_t(capture), publishNs_ = int64_t(publish);
|
||||
const uint32_t nhands = std::min<uint32_t>(copy.nhands, kMaxHands), ncaps = std::min<uint32_t>(copy.ncapsules, kMaxCapsules);
|
||||
captureNs_ = int64_t(copy.capture_ns), publishNs_ = int64_t(copy.publish_ns);
|
||||
const Mat head = HeadAt(captureNs_);
|
||||
|
||||
// each hand's palm in the room, and its velocity from the last time it was seen
|
||||
ids_.clear();
|
||||
std::vector<int> owners; // the hand each capsule belongs to, in file order
|
||||
for (uint32_t k = 0; k < std::min<uint32_t>(nhands, kMaxHands); ++k) {
|
||||
const uint8_t *h = copy + kHeader + k * kHand;
|
||||
uint32_t id, n;
|
||||
float pts[21][3];
|
||||
std::memcpy(&id, h, 4);
|
||||
std::memcpy(pts, h + 16, sizeof pts);
|
||||
std::memcpy(&n, h + 268, 4);
|
||||
for (uint32_t k = 0; k < nhands; ++k) {
|
||||
const fh_hand_t &h = copy.hands[k];
|
||||
const uint32_t id = h.id;
|
||||
const auto &pts = h.pts;
|
||||
const int idx = int(ids_.size());
|
||||
ids_.push_back(id);
|
||||
owners.insert(owners.end(), std::min<uint32_t>(n, kMaxCapsules), idx);
|
||||
owners.insert(owners.end(), std::min<uint32_t>(h.ncapsules, kMaxCapsules), idx);
|
||||
double palm[3] = {0, 0, 0};
|
||||
bool ok = true;
|
||||
for (int j : {0, 5, 9, 13, 17}) {
|
||||
@@ -137,19 +128,18 @@ bool Hands::Read() {
|
||||
}
|
||||
for (auto it = motion_.begin(); it != motion_.end();)
|
||||
it = captureNs_ - it->second.ns > kStaleNs ? motion_.erase(it) : std::next(it);
|
||||
if (owners.size() != std::min<uint32_t>(ncaps, kMaxCapsules)) owners.assign(std::min<uint32_t>(ncaps, kMaxCapsules), -1);
|
||||
if (owners.size() != ncaps) owners.assign(ncaps, -1);
|
||||
|
||||
base_.clear(), owner_.clear();
|
||||
for (uint32_t k = 0; k < std::min<uint32_t>(ncaps, kMaxCapsules); ++k) {
|
||||
float f[8];
|
||||
std::memcpy(f, copy + kHeader + kMaxHands * kHand + k * kCapsule, sizeof f);
|
||||
for (uint32_t k = 0; k < ncaps; ++k) {
|
||||
const fh_capsule_t &f = copy.capsules[k];
|
||||
bool ok = true;
|
||||
for (float v : f) ok = ok && std::isfinite(v) && std::fabs(v) < 10;
|
||||
if (!ok || f[6] <= 0 || f[7] <= 0) continue;
|
||||
for (float v : {f.a[0], f.a[1], f.a[2], f.b[0], f.b[1], f.b[2], f.ra, f.rb}) ok = ok && std::isfinite(v) && std::fabs(v) < 10;
|
||||
if (!ok || f.ra <= 0 || f.rb <= 0) continue;
|
||||
Capsule c;
|
||||
Apply(head, f, c.a);
|
||||
Apply(head, f + 3, c.b);
|
||||
c.ra = f[6], c.rb = f[7];
|
||||
Apply(head, f.a, c.a);
|
||||
Apply(head, f.b, c.b);
|
||||
c.ra = f.ra, c.rb = f.rb;
|
||||
base_.push_back(c);
|
||||
owner_.push_back(owners[k]);
|
||||
}
|
||||
@@ -305,7 +295,7 @@ void main() {
|
||||
// The client's pixels, opaque (its alpha is ignored, as IgnoreTextureAlpha did).
|
||||
const char *kCopy = R"(
|
||||
#extension GL_OES_EGL_image_external : require
|
||||
precision mediump float;
|
||||
precision highp float; // mediump (16-bit on Adreno) steps 1.7 texels across a 3440-pixel screen
|
||||
uniform samplerExternalOES tex;
|
||||
varying vec2 uv;
|
||||
void main() { gl_FragColor = vec4(texture2D(tex, uv).rgb, 1.0); })";
|
||||
|
||||
+2
-2
@@ -1,8 +1,8 @@
|
||||
// Hand cutouts: where a tracked hand is between an eye and a screen, that eye sees the
|
||||
// room (Room View) through the screen instead of the screen drawn over the hand.
|
||||
//
|
||||
// frame-hands' tracker (a separate project, ~/Desktop/Projects/frame-hands) publishes
|
||||
// the hands it sees with the headset's cameras to $XDG_RUNTIME_DIR/frame-hands/hands:
|
||||
// ft-hands (hands/) publishes the hands it sees with the headset's cameras to
|
||||
// /run/user/UID/frametop-hands/hands (hands/include/fh_hands.h):
|
||||
// capsules (finger bones, palm, forearm) in the head frame at capture time. Hands turns
|
||||
// them into the room with the head pose at that time. Project() finds where each eye
|
||||
// sees them on a panel, and Renderer draws the panel's client buffer into a side-by-side
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
// ft-handtest: try the hand cutouts without restarting the desktop. Shows a test panel
|
||||
// (a light grid) in front of you as its own overlay; where frame-hands tracks your hands
|
||||
// (a light grid) in front of you as its own overlay; where ft-hands tracks your hands
|
||||
// in front of it, each eye sees through it, like ft-screens' screens with cutouts.
|
||||
//
|
||||
// ft-handtest [--distance m] [--width m] [--seconds s]
|
||||
//
|
||||
// Needs frame-hands' tracker running (it publishes $XDG_RUNTIME_DIR/frame-hands/hands).
|
||||
// Needs hand tracking running (hands/run.sh; ft-hands publishes /run/user/UID/frametop-hands/hands).
|
||||
// Build: screens/build.sh (build/ft-handtest), run in the dev container.
|
||||
#include "handcut.h"
|
||||
|
||||
|
||||
+1
-1
@@ -40,7 +40,7 @@
|
||||
// own; also for flatscreen games, which aren't scene apps).
|
||||
// - during a VR game the screens hide unless the dashboard is open (g_inGames, default),
|
||||
// or stay visible over it; the hotkey still shows them.
|
||||
// - hand cutouts (handcut.cpp): where frame-hands tracks a hand between an eye and a
|
||||
// - hand cutouts (handcut.cpp): where ft-hands (hands/) tracks a hand between an eye and a
|
||||
// screen, that eye sees through the screen (to Room View). Only then is the screen
|
||||
// drawn by us, into a side-by-side buffer (one half per eye); otherwise its client
|
||||
// buffer is shown as is.
|
||||
|
||||
@@ -34,6 +34,8 @@ POINTER_GAZE_HOLD=0.5 # gaze mode: a press held this long (s) without movin
|
||||
POINTER_GAZE_SHOW=1 # gaze mode: the dot shows this long (s) after the mouse moves it; also while a press is held, and a pulse per click
|
||||
GAZE_TRACKER=steam # gaze service: steam = SteamVR's eye tracker | own = our own (frame-eyes' fe-trackd, calibrated in the gaze probe with Own tracker)
|
||||
GAZE_EYE=auto # gaze service: eye bias. auto = each eye weighted by how far off it was at your recent nudges | left | right = that eye counts twice
|
||||
HANDS_SWAP_SIDES=0 # hand tracking (hands/run.sh install): 1 = the side cameras' names are swapped, which some SteamVR restarts cause (hands/tools/check_sides.py --ring tells)
|
||||
HANDS_CPUS=5,6,7 # hand tracking: the CPUs its model threads run on
|
||||
META_DASHBOARD=0 # 1 = a Meta tap on a pass-through keyboard toggles the SteamVR dashboard (pointer mode only)
|
||||
SHARE_KEYS=0 # 1 = keys of keyboards grabbed for the desktop also go to @frametop_keys, for hotkey tools (any local process can listen)
|
||||
DISPLAY_OFF_MIN=0 # minutes without use before frametop-power turns the displays off, even if the headset seems worn (0 = never)
|
||||
|
||||
@@ -18,6 +18,9 @@ packages=(
|
||||
mesa-libgbm-devel wayland-devel vulkan-loader-devel vulkan-headers plasma-wayland-protocols wlroots-devel
|
||||
# ft_pointer SteamVR driver: static C++ runtime (the host has an older glibc)
|
||||
libstdc++-static
|
||||
# hand tracking: ft-hands reads the calibration with jsoncpp; ft-camd runs on the host, linked
|
||||
# statically; the Python tools (hands/tools) need NumPy and OpenCV
|
||||
jsoncpp-devel glibc-static python3-numpy python3-opencv
|
||||
# Frametop Input Settings app (Kirigami, PySide6)
|
||||
python3-pyside6 kf6-kirigami kf6-qqc2-desktop-style qt6-qtwayland breeze-icon-theme plasma-breeze
|
||||
# Frametop remote desktop (VNC bridge through krdp)
|
||||
|
||||
Reference in new issue
Block a user