Compare commits

...
70 Commits
Author SHA1 Message Date
DeeJanuzandClaude Opus 5.5 2d985573c0 Gaze first: abandoned; keep the controller experiment for reference
The last state of the experiment, kept on this branch only:
- input/gazefirst.py, input/steamui.py: the relay reading the controllers from
  vrserver's web socket, switching the compositor binding, and driving the Steam
  UI filter (whose block now expires unless renewed).
- The helper's trigger and bumper holds (steered by the controller's position),
  the laser keeper, and the click gate that waits for the laser.
- pointer/probe/focustest (overlay flag 1 << 4, as an input client too),
  vrsetting (SteamVR settings through vrserver), lasertest --reclaim.

Every press and release Steam sees takes SteamVR out of laser mode, and taking
the laser back breaks clicks and drags. See docs/gaze-controllers.md on
experimental.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 21:11:40 -06:00
DeeJanuzandClaude Opus 5.5 537fdac419 Gaze first: a filter for Steam's UI that drops gamepad input in laser mode
Steam reads the Frame controllers as its own virtual gamepad, so no
SteamVR binding keeps stray bumper and thumbstick presses away from its
UI. input/steam-gamepad-filter.js, evaluated in Steam's SharedJSContext
through its CEF debugger, wraps the gamepad source's OnButtonDown,
OnButtonUp, and OnAnalogPad and drops everything but the Steam button
while the dashboard is in laser mode (mode "auto"), and logs presses
and mode changes. Tested blocking in the headset: nothing reached the
UI; the Steam button and the grips never go through it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 17:55:19 -06:00
DeeJanuzandClaude Opus 5.5 87a63932c0 Gaze first: test results; treadmill from the start, muted compositor binding
From the first headset tests (docs/gaze-first.md, "Test results"):
- SteamVR gives the treadmill path only to a device that hints it when
  it's added, so the driver hints a role that's no hand from Activate
  (a hand still only while shown). With it, our device held the
  dashboard laser with no hand role. The helper doesn't release the
  pointer for having no hand role when POINTER_ROLE is treadmill.
- The helper's global action sets can't mute the compositor (SteamVR
  reports them inactive while the laser mouse has focus), but selecting
  pointer/bindings/vrcompositor_frame_controller_gazefirst.json (the
  stock binding without its trigger and bumper laser entries) through
  vrserver's /input/selectconfig.action does.
- input/vrws.py follows devices that connect later.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 17:23:30 -06:00
DeeJanuzandClaude Opus 5.5 bd06207db9 Gaze first: plan, and tools for the headset tests
docs/gaze-first.md is the plan agreed on 2026-09-30: with gaze on and no
game, the gaze drives the 3D pointer everywhere, either trigger clicks
where you look (tap, precision by hand movement, hold to drag), and the
controllers keep everything but their lasers.

For its four tests:
- The driver takes "role treadmill", and its compositor bindings repeat
  the right hand's under /user/treadmill. They're additive, so nothing
  changes while the device is a hand. POINTER_ROLE accepts treadmill.
- pointer/probe/lasertest follows who has the dashboard laser and the
  controllers' roles, and with --snapback takes the laser back.
- input/vrws.py reads vrserver's web socket with the standard library
  (the host's Python has no websockets module) and prints component
  changes and update rates. Checked: it connects, and lists the
  controllers' bumper, thumbstick axes, and Steam button, and the
  headset's /proximity, which flickers off for 0.3-0.5 s while worn.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 17:00:09 -06:00
DeeJanuzandClaude Opus 5.5 d5c2ec65d7 Take the user, host, and home from the machine, not this Frame
The local RDP login between krdp and the VNC bridge uses the account's
own name instead of "steamos", its certificate the hostname instead of
"steam-frame", and the ft_pointer driver installs under the user's home
instead of /home/steamos. The remote-access address already came from
the tailnet (tailscale0 and tailscaled's local API).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:57:33 -06:00
DeeJanuzandClaude Opus 5.5 1f19732aaf Add Frametop Remote Access, a settings app for the VNC view
A GTK 4 / libadwaita app (remote/ft-remote-settings, on the host's own
Python like the gaze probe): remote access on and off (REMOTE, applied at
once when the desktop allows it), the tailnet name and address to connect
to, and the VNC password shown, copied, or replaced. The password stays
random and made on the Frame, in ~/.config/frametop-remote; none is in
the code.

session/remote-ctl.sh starts, stops, and reports remote access; the
session uses it and leaves a remote-capable marker, since KWin allows the
capture only in a desktop that started with REMOTE=1. The password
between krdp and the VNC bridge (local only, but on krdp's command line)
is now new at every start. The installer adds the app to the menu.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:55:49 -06:00
DeeJanuz 9820e7996f Serve only the desktop's primary screen over VNC
krdp streams every screen, so the VNC screen is now the primary's size and the
FreeRDP window is shifted so the primary fills it (ft-layout remote-view gives
the offset). It resizes and reconnects when the layout changes. With remote
access on, KWin's D-Bus screenshot interface is open too, for scripts that look
at the screens without the headset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 6f5a23c80f2ecd26e6a96c86b102be36d518ee25)
2026-09-30 15:50:46 -06:00
DeeJanuzandClaude Opus 5.5 97f853130d Merge eye-tracking into experimental
Our own eye tracker joins the gaze folder: ft-eyegrab (the root frame
grabber, frametop-eyegrab.service) and ft-eyes, run by the gaze service
when GAZE_TRACKER=own, with its lab tools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:45:38 -06:00
DeeJanuzandClaude Opus 5.5 55d92b411b Input Settings: drop the pointer role choice
The hand role the pointer takes isn't something to choose by hand. The
helper still reads POINTER_ROLE (right by default) for experiments.

Tried and reverted (not committed): in gaze mode, taking the stylus role
by reconnecting. SteamVR gave the device no role at all (hint 5, role
none) and kept the dashboard laser on the held controllers, so the mouse
couldn't click either. The dashboard laser needs a hand role, and a held
Frame controller takes its hand's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:35:08 -06:00
DeeJanuzandClaude Opus 5.5 d666fa031f Gaze first: precision and drag buttons, key combinations, pointer role
Two new actions for mouse buttons, controller buttons, and key
combinations: gaze_precision (hold: the pointer stops where you look and
the button's device steers it, a controller by where it points at
POINTER_PRECISION_GAIN, the mouse by its moves; release: click there) and
gaze_drag (the same with a real press at once, dragging until the
release). The relay sends "precision|gazedrag <source> 1|0" to the
helper, which runs them through the same holds as pinches and grips.

- Key combinations on any keyboard ("key_bindings" in the rules, e.g.
  Ctrl+Alt+G for gaze on/off): the last key isn't typed.
- POINTER_GAZE_MOUSE: in gaze mode the left button is a precision button
  (the default, as before) or clicks right away (direct).
- In gaze mode a moving controller no longer takes the pointer away.
- POINTER_ROLE (right, left, stylus): the driver takes a "role" command,
  so the pointer can stay off the hand holding the precision controller,
  which would otherwise take the role back. Needs the rebuilt driver.
- Input Settings: the actions on the Controllers and Buttons pages; on the
  Gaze page the mouse choice, the role, the precision sliders, and key
  combinations (captured from any keyboard).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:14:57 -06:00
DeeJanuzandClaude Opus 5.5 c2739cb0df Hands: park hand tracking behind ft-handsctl
Hand tracking no longer starts with SteamVR: hands/run.sh install leaves
the units disabled and links hands/ft-handsctl into ~/.local/bin, which
turns it on and off (on | off | status | log | cutouts on|off | gestures).
With it on, hands show through the screens; pinches and grips move the
pointer only with POINTER_HANDS=1, now off by default. ft-camd's service
runs the mono cameras only: while the headset is worn, the colour module
writes just a half-size image into the top-left quarter of its buffers,
which ft-camd can't use yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:08:54 -06:00
DeeJanuzandClaude Opus 5.5 5ca4160a7e Hands: fixes from the second headset test (pinches)
- The palm-down filter is off by default: the user's deliberate pinches,
  hand raised, read 0.90-0.99, like typing.
- Typing is told apart by the keyboard instead: the input relay sends the
  pointer helper "typing" on key presses, and it takes no pinch within
  POINTER_PINCH_TYPING (1 s) of one.
- Grips are still held to hands raised (POINTER_GRIP_BELOW, 0.35 m below
  the eyes); pinches aren't, since the user's own sat 0.35-0.45 m below,
  elbow resting.
- A grip doesn't begin with the thumb on the index tip: that's a pinch
  with the other fingers curled, which was taken for a grip.
- A pinch's point is the index and middle knuckles: the tips' midpoint
  moved 1-2 cm as the pinch opened, dragging every release off its press.
- ft-hands --gesture-log prints what the detectors measure, 10 times a
  second.

Pinches still aren't reliable enough to use; hand tracking is parked for
now in favour of the controllers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:07:58 -06:00
DeeJanuzandClaude Opus 5.5 fbd188f5f9 ft-pointer: without gaze mode, a pinch is a real press
Pressed when the pinch closes, released when it opens, its hand dragging
the pointer in between (at the pinch gain, past the dead zone), like the
mouse's button. That's what the gaze probe's Click practice needs: it
freezes its own gaze dot at the press and drags it by the pointer's
movement until the release. With gaze mode on, the pinch still holds the
press back and clicks on the release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 14:42:59 -06:00
DeeJanuzandClaude Opus 5.5 6d1b04e032 ft-pointer: pinch to click, grip to drag
The pointer helper reads ft-hands' gestures. A pinch holds the pointer
where the gaze put it and clicks on the release; held, the hand nudges
the pointer at half its angle past a 1.5 degree dead zone, which also
teaches the gaze tracker. A grip presses where the pointer is, drags with
the hand, and releases when the hand opens, unless it began more than
30 cm below the eyes (hands on a desk). Hand movement is taken in the
room with the head pose at capture time. Hand use keeps the pointer from
the relay's idle release, as gaze mode does. POINTER_HANDS and friends in
frametop.conf.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 14:31:17 -06:00
DeeJanuzandClaude Opus 5.5 50c14545ed Hands: pick the cameras by the light, and detect grips
ft-hands tracks with the mono IR cameras in dim light and with every
camera (or the colour pair, HANDS_BRIGHT) in bright light, going by the
colour frames' mean brightness with hysteresis and a 2 s hold
(HANDS_CAMERAS=auto, the default; mono, color and all fix it). Colour
frames are placed on the mono cameras' clock by their dequeue time, and a
view in a camera a step lacks waits for that camera's next frame.

A grip (a closed hand) is a second gesture next to the pinch, in version 2
of the gestures file: every fingertip curled toward the wrist, beginning
only on a hand seen open within a second and held up in front. On the
2026-09-30 lit recording that leaves 6 false grips of 14, all with the
hands on the desk; pinch counts are unchanged. ft-handreplay logs grips
and finger curl, watch_gestures.py shows them, and tools/cut_sets.py
copies a few sets out of a recording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 14:31:17 -06:00
DeeJanuzandClaude Opus 5.5 40a2f39b2a ft-camd: keep colour trouble away from the mono cameras, and idle colour
A colour camera probes at most 4 buffers a frame (each probe syncs a
~9 MB buffer's cache, which made the mono cameras miss frames), and one
that goes stale twice in a row is paused (10 s, doubling to 160 s) and
learned again instead of ft-camd exiting. The colour cameras run at 2 fps
until a reader asks for more in frametop-hands/color-fps, which saves
most of their decoding while only their brightness is needed. Each mono
camera's latest near-black frame's mean goes in the ring (dark_mean), a
measure of the room's IR light. The service now starts with --with-color;
HANDS_CAMERAS=mono leaves the colour cameras out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 14:31:16 -06:00
DeeJanuzandClaude Opus 5.5 8bc6739dca Merge hands-migration into experimental
Hand tracking (ft-camd, ft-hands, and their tools) joins the desktop. The
hands file and ring move to /run/user/UID/frametop-hands/, which ft-screens'
hand cutouts now read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 14:02:01 -06:00
DeeJanuzandClaude Opus 5.5 e2abaa06b6 Hands: fixes from the first headset test
- The runtime files move to /run/user/UID/frametop-hands/: the desktop
  session deletes /run/user/UID/frametop at every start.
- The cutout copy shader runs at highp: mediump (16-bit on Adreno)
  stepped 1.7 texels across a 3440-pixel screen.
- One hand no longer pinches both sides after its left/right call flips
  mid-pinch, and --pinch-palm-down (0.6) holds back pinches with the palm
  facing down (typing on a lap keyboard).
- ft-camd judges a colour frame fresh by its luma rows only, and logs
  per-buffer changes at stale colour frames with FT_CAMD_DEBUG=1.
- hands/run.sh caps skips the setcap when ft-camd already has them.
- The replay tool dumps poses (--poses), and its pinch events carry the
  hand id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 14:00:54 -06:00
DeeJanuzandClaude Opus 5.5 4493b789fd Merge floating-windows into experimental
Floating windows join the keyboard and the hand cutouts. Floating panels
don't get cutouts yet: their panel and popups show crops of the client
buffer (texture bounds), which the side-by-side cutout buffer doesn't
match. KWin gets both the spare outputs and our input method, and the
pointer helper's frametop. prefix already covers the float panels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 13:53:28 -06:00
DeeJanuzandClaude Opus 5.5 5714978e65 Add a keyboard for text fields on the desktop
Fixes #7: SteamVR's keyboard never came up for the desktop's apps, and
opening it for our panels doesn't work well on the Frame (it's Steam's own
panel, mounted in the dashboard's scene, it follows the laser between
panels, and it takes the controllers over to SteamVR's laser).

- KWin starts input/ft-textinput as the desktop's input method. It tells
  the input relay when a text field gains or loses focus, and the relay
  asks ft-screens to open or close the keyboard. The session drops the
  QT_IM_MODULE=xim and GTK_IM_MODULE=xim that the gamescope session sets,
  or Qt and GTK apps never report text fields.
- The keyboard is ft-screens' own panel (screens/keyboard.cpp): a US laptop
  layout, typed with a controller's laser or the 3D mouse. It opens 0.7 m
  in front of you, below your eyes and facing you. It has a grab bar to
  move it, a Close key, latching Shift, Ctrl and Alt, and repeat on a held
  key. It's drawn into shared DMA-BUFs, so it doesn't flicker. Its keys
  reach the focused screen as key presses, so every app takes them.
- It steps aside while the Steam menu or Steam's own keyboard is up and
  comes back after. A layout reset closes it, and it doesn't open without
  a head pose.
- Frametop Input Settings has a Keyboard page: open it for every text
  field, only while no keyboard is connected (the default), only from a
  mapped button (the new Open/close keyboard action, for mice and
  controllers), or never. A switch keeps it open until you close it.
- The pointer helper treats every frametop.* overlay as a real panel. The
  keyboard's shared texture reports 0x0 like SteamVR's scene-graph
  controls, and the helper had given it their wide catch radius.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 13:41:47 -06:00
DeeJanuzandClaude Opus 5.5 c8fc6351bb Floating windows: fix the scale crash, the click offset, and placement
- KWin's nested backend makes an output the size it's configured to times its
  scale, so after Meta+scroll every size ft-floatd sent was multiplied again,
  and an odd result disconnected KWin (buffer not divisible by its scale).
  ft-floatd now asks for sizes in the output's scaled terms (kwin_size), and
  asks again after each scale change.
- SteamVR reports mouse positions on a panel with texture bounds in the whole
  texture, not the crop, so clicks on a floating window landed up to ~200 px
  off. The mouse scale is now the buffer's size, as on a screen.
- A floated window starts 30 cm in front of its screen (was 5 cm), so it's
  easy to point at apart from the screen behind it.
- No 1 s wait before a spare turns on (a disabled output never commits), and
  the login splash on the spares isn't taken for floating windows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 10:57:36 -06:00
Nikita Koptelov 4e5131d8ce Start Frametop's SteamVR clients only once SteamVR is up (#6)
The pointer and gaze services need steamvr.service to be running (Requisite=), and ft-pointer, ft-screens, and ft-gaze connect as a background app before switching to overlay, so they never start a vrserver of their own. One started from the dev container never finds the headset, which left a reboot stuck in a loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 51b4e79318)
2026-09-30 10:02:46 -06:00
Nikita Koptelov 2a9fbdebb4 Start Frametop's SteamVR clients only once SteamVR is up (#6)
The pointer and gaze services need steamvr.service to be running (Requisite=), and ft-pointer, ft-screens, and ft-gaze connect as a background app before switching to overlay, so they never start a vrserver of their own. One started from the dev container never finds the headset, which left a reboot stuck in a loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 51b4e79318)
2026-09-30 10:02:46 -06:00
Nikita Koptelov c3375d32d3 Start Frametop's SteamVR clients only once SteamVR is up (#6)
The pointer and gaze services need steamvr.service to be running (Requisite=), and ft-pointer, ft-screens, and ft-gaze connect as a background app before switching to overlay, so they never start a vrserver of their own. One started from the dev container never finds the headset, which left a reboot stuck in a loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 51b4e79318)
2026-09-30 10:02:33 -06:00
Nikita Koptelov 3aea571496 Start Frametop's SteamVR clients only once SteamVR is up (#6)
The pointer and gaze services need steamvr.service to be running (Requisite=), and ft-pointer, ft-screens, and ft-gaze connect as a background app before switching to overlay, so they never start a vrserver of their own. One started from the dev container never finds the headset, which left a reboot stuck in a loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 51b4e79318)
2026-09-30 10:02:33 -06:00
DeeJanuzandClaude Opus 5.5 ed75614f62 Bring our own eye tracker into the gaze folder, run by the gaze service
frame-eyes, a separate project until now, becomes gaze/tracker:
- ft-eyegrab (was fe-bufprobe) copies the eye-camera frames, read-only, out of
  SteamVR's eyetracking process. It runs as the system service
  frametop-eyegrab.service, which gaze/tracker/install.sh installs to
  /etc/frametop with sudo. It keeps only CAP_SYS_PTRACE, CAP_DAC_READ_SEARCH, and
  CAP_CHOWN, and copies frames only while /dev/shm/frametop-eyes-want is fresh,
  holding none of the tracker's buffers otherwise.
- ft-eyes (was fe-trackd) runs under ft-gazed in the dev container, with
  build/venv's pinned numpy and OpenCV: while Eye tracker is Own tracker, or on
  the probe's lease ("eyes SECONDS"). No sudo password or fe-live script at
  run time any more.
- lab/ holds the research tools (ft-eyes-score, -e2e, -record, -replay,
  -session) and findings.md. Recordings live outside the repo, in
  ~/.local/share/frametop/eyes/captures; .gitignore catches stray frame dumps.

Its socket is now @ft_eyes, its output /dev/shm/frametop-eyes-gaze, and its state
~/.local/state/frametop/gaze/eyes. On practice1 -> practice2 the whole live path
(ft-eyes-e2e) gives 1.30 deg median and 3.18 for the worst tenth, as before the
move (1.30, 3.21).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 10:00:07 -06:00
DeeJanuzandClaude Opus 5.5 369736f0ca Read the hands file through the shared header, at its Frametop path
ft-screens' hand cutouts read /run/user/UID/frametop/hands through
hands/include/fh_hands.h instead of their own copy of its offsets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:15:07 -06:00
DeeJanuzandClaude Opus 5.5 499035216c Make hand tracking a Frametop component
- Programs: ft-camd (the camera broker), ft-hands (the tracker), and
  ft-handreplay and ft-ringplay for recordings, built by hands/build.sh
  into hands/build/ with one Makefile. The first build fetches ncnn at
  frame-hands' pinned tag and builds it with the same options.
- ft-camd gets its privileges from file capabilities (CAP_SYS_PTRACE,
  CAP_PERFMON, CAP_DAC_READ_SEARCH) that hands/run.sh install sets with
  sudo, and drops them once set up. It still works under sudo. It runs
  on the host, linked statically, as frametop-camd.service. ft-hands
  runs in the dev container as frametop-hands.service. Both start and
  stop with SteamVR.
- Files move to /run/user/UID/frametop/ (cam-ring, hands, gestures),
  not $XDG_RUNTIME_DIR, which a terminal in the Frametop desktop has
  its own of. SIGUSR1 recordings go to ~/.local/share/frametop/hands.
- The calibration is read through /run/host in the container.
- Settings: HANDS_SWAP_SIDES and HANDS_CPUS in frametop.conf.
- install.sh offers hand tracking as an optional last step.
- The container gets jsoncpp-devel, glibc-static, and NumPy and OpenCV
  for the Python tools.
- tools/ring.py reads the ring, and models/NOTICE credits the
  Apache-2.0 models.

Checked: ft-handreplay gives identical summaries and byte-identical
depth dumps to frame-hands' fh-replay on both 2026-09-29 recordings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:15:07 -06:00
DeeJanuzandClaude Opus 5.5 5565a25444 Let the gaze pointer use our own eye tracker, and weight the eyes
Frametop Input Settings' Gaze page gets two settings, saved in frametop.conf and
read again by ft-gazed when the file changes: Eye tracker (GAZE_TRACKER: SteamVR's
or our own, frame-eyes' fe-trackd) and Eye bias (GAZE_EYE: auto, left, right).

ft-gazed now combines the eyes, each calibrated on its own: SteamVR's set 2 eyes
with the probe's Left eye and Right eye calibrations, or our tracker's eyes as
they come. Without per-eye calibrations, or with --source, it keeps the older
one-source path. A pointer nudge finds its look from the raw gaze the helper
echoes back; with our tracker it goes to fe-trackd as a click.

The bias leans instead of choosing (gazecal.EyeWeights): on 306 live clicks the
eyes' errors partly cancelled, both together 0.65 deg off against 0.96 and 1.11
for either alone. Left or Right counts that eye twice; auto weights each eye by
its RMS miss at its last 20 nudges, since the calibration's fit picked the wrong
eye on SteamVR's test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:11:54 -06:00
DeeJanuz 3440ec8b90 Merge branch 'main' into experimental 2026-09-30 09:11:24 -06:00
DeeJanuzandClaude Opus 5.5 3e3d31c728 Lay out hands/ like Frametop's other components
trackd/ becomes track/, and the calibration helper the Python tools
import moves from the prototype's folder into tools/. The model
development tools (nettest and the scripts that compare it with the
Python models or cut int8 calibration crops) stay in frame-hands with
the prototype they need.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:01:10 -06:00
DeeJanuzandClaude Opus 5.5 1a76d1560b Bring in frame-hands' hand tracking under hands/
The history of frame-hands (~/Desktop/Projects/frame-hands on the
Frame), filtered to what moves: the camera broker (camd/), the tracker
and its offline tools (trackd/), the shared file layouts (include/), the
ncnn models, the analysis tools, and the calibration and model helpers
they import from the Python prototype. The reverse-engineering notes,
probes, camprobe, and the rest of the prototype stay in frame-hands.
Unchanged here: the renames to ft- names and Frametop paths follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 08:59:32 -06:00
DeeJanuzandClaude Opus 5.5 9f6ce5cc59 Plan the move of hand tracking into Frametop
frame-hands (the headset cameras' hand tracking, a separate project so
far) becomes a native component under hands/. docs/hands-migration.md
has what it is, where each part goes, names, build, how ft-camd gets
its privileges (file capabilities), interfaces, open items, and steps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 08:59:27 -06:00
DeeJanuzandClaude Opus 5.5 2031b9ca8f Add colour cameras, pinch gestures, and depth measures
- fh-camd --with-color publishes the Arcturus colour pair (luma, half
  size, 30 fps) to the ring, flagged FH_CAM_COLOR. fh-tracker records
  them, and fh-replay --cams mono|color|all tracks with them, using the
  module's EEPROM calibration (load_color_calibration).
  tools/check_color.py checks which node is left and how the crop maps.
- Pinch detection per hand (trackd/pinch.h), published to
  $XDG_RUNTIME_DIR/frame-hands/gestures (include/fh_gestures.h). It has
  begin and end counters, times, the pinch point, and the begin point
  for drags. tools/watch_gestures.py shows it live, and fh-replay
  reports it.
- fh-replay --depth and tools/depth_report.py measure the depth without
  ground truth: noise along the line of sight against across it, the
  one-camera guess, and a simulated camera loss.
- fh-tracker --swap-sides, --ring, --keep-presence, --record-only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 08:58:30 -06:00
DeeJanuzandClaude Opus 5.5 fc15a8a8a6 Record what's built of floating windows
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:32:03 -06:00
DeeJanuzandClaude Opus 5.5 691b66cd87 Stop ft-screens without an abort
wlroots asserts that nothing still listens to its xdg-shell and decoration
globals when the display goes, so every desktop stop ended with ft-screens
aborting and leaving a core dump.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:31:54 -06:00
DeeJanuzandClaude Opus 5.5 6122b7eb15 Floating windows: full screen, per-window scale, and drags across panels
Full screen fills the window's own panel (the margin drops to zero).
Meta+scroll over a floating window changes its scale at the same size in
pixels. Spare outputs are sized to a multiple of KWin's buffer scale, and
screens to even sizes: an odd buffer at a fractional scale is a protocol
error that disconnected KWin. The 3D mouse's drag lock now crosses onto
other Frametop panels unless the pressed one is being carried.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:31:23 -06:00
DeeJanuzandClaude Opus 5.5 c2e5e861c8 Show floating windows as panels of their own
ft-screens makes a panel for each spare output (frametop.float.N), hidden
until ft-floatd floats a window on it. The panel shows only the window's
rectangle of the buffer at the density of the screen it came from, its
popups and dialogs get small panels over it, pressing its title bar
carries it while KWin's pointer stays put, the corner tab resizes the
window in pixels, and two more buttons close it and put it back on the
desktop. The session adds FLOAT_SLOTS spare outputs to KWin and starts
ft-floatd; ft-layout arranges only the screens' outputs, and the pointer
helper treats the new panels like screens. Not yet tried in the headset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:26:20 -06:00
DeeJanuzandClaude Opus 5.5 87a402c68d Add ft-floatd and the frametop-float KWin script
ft-floatd loads the script into the desktop's KWin, which reports windows
over D-Bus and takes commands through a long poll. Floating a window (the
window menu's Float in VR, Meta+Shift+F, or ft-float) turns on a spare
output sized to the window plus a margin, moves the window onto it, and
tells ft-screens the panel's crop, density, and place; docking puts it
back and turns the spare off. Tested on the headless test desktop; the
panel side in ft-screens comes next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:20:26 -06:00
DeeJanuzandClaude Opus 5.5 fcc8d96246 Record the headless phase 0 results
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:14:22 -06:00
DeeJanuzandClaude Opus 5.5 4859412126 Put the first click after crossing onto a screen where the pointer is
KWin's nested backend ignores the position in wl_pointer.enter, and
wlroots drops a motion to the position it entered at, so KWin kept its old
pointer until the next move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:14:09 -06:00
DeeJanuzandClaude Opus 5.5 42a755b9f4 Add a headless test mode to ft-screens and a throwaway test desktop
ft-screens --no-vr runs without SteamVR and leaves the input relay alone;
--control names its control socket; "toplevels" lists KWin's windows with
their titles and sizes; "input" feeds pointer events as if from a panel.
screens/test/headless.sh starts it with a bare nested KWin next to the
running desktop, and loads KWin scripts, runs apps, and takes screenshots.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:14:09 -06:00
DeeJanuzandClaude Opus 5.5 76f3cdd2dc Ask for one look at a centre dot after the headset was off
With the Own tracker, the probe polls its status each second. After the headset was off
(or the tracker restarted), it shows one dot at the centre until you look at it and press
(S skips): that click resets both eyes' shifts, so the first real clicks aren't 10-17
degrees off. The screen-to-direction geometry the calibration used is shared with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:08:07 -06:00
DeeJanuzandClaude Opus 5.5 0a88a6b7c9 Release a button held on a screen even when the laser lets go between panels
While a button is held on a screen, the laser leaving it no longer takes
KWin's pointer, and an invisible catcher overlay sits on the laser whenever
it's off every panel, so the release reaches KWin at the pointer's last
spot. The pointer helper also reports the mouse's left release as a
backstop. Not yet tested in the headset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:05:54 -06:00
DeeJanuz f17a5a66ab Merge remote-tracking branch 'frame/layouts-headpin' into floating-windows
# Conflicts:
#	README.md
#	display-settings/ft_display_settings.py
#	display-settings/main.qml
#	docs/reference.md
2026-09-29 23:02:12 -06:00
DeeJanuz d5dbe23cae Merge remote-tracking branch 'frame/pointer-ignore' into floating-windows 2026-09-29 23:01:46 -06:00
DeeJanuzandClaude Opus 5.5 389878b9fd Settle the floating-windows design with the user
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:01:46 -06:00
DeeJanuzandClaude Opus 5.5 0af087c776 Plan floating windows: any desktop app in a VR panel of its own
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 23:00:33 -06:00
DeeJanuzandClaude Opus 5.5 3022d7de9d Run the tracker on CPUs 5-7 by default
probes/core_ab.py with the headset on (3 rounds, the same replayed frames in
every block): on 5-7 a step took 8.4 ms against 13.2 ms on 2-4, where XRService's
head tracking also runs, and latency fell from 14.1 to 9.6 ms. The compositor's
late frames and CPU/GPU time per frame didn't change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 22:59:19 -06:00
DeeJanuzandClaude Opus 5.5 c37bee27c9 Keep the Own tracker's calibration in reach, and let it learn after a re-seat
The calibration dots for the Own tracker go out to the Calibration ring angle each way on
an oval, not to the window's corners, which were too far to look at while facing the
centre. The Own tracker learns from drags up to 25 degrees (after taking the headset off
and on, its first clicks were 11-17 degrees off, and the 6-degree limit blocked them). The
probe starts on the Own tracker when it's running, and the calibration header names the
tracker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 22:51:37 -06:00
DeeJanuzandClaude Opus 5.5 355f30f335 Add a live replay harness for the CPU-placement test
fh-ringplay plays a recording into a frame ring in real time, and fh-tracker
--ring reads it (cameras without a device node are mapped by name). With the
same frames in every run, probes/core_ab.py compares the tracker on CPUs 2-4,
on 5-7, and not running, measuring the tracker's step time and latency, the
compositor's late frames and CPU/GPU time, XRService's timing warnings, CPU
temperature and clocks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 22:42:24 -06:00
DeeJanuzandClaude Opus 5.5 067a03ce38 Hand tracking for Frametop's hand cutouts
fh-camd (camd/) borrows XRService's camera buffers and publishes the tracking
cameras' frames to a shared ring. fh-tracker (trackd/) finds hands in them with
MediaPipe's palm and landmark models on ncnn, triangulates them in 3D, and
publishes them for ft-screens. fh-replay replays recordings offline. tracker/ is
the earlier Python version; tools/ and probes/ hold the checks and experiments.

As of this commit: crop contrast defaults to CLAHE for the palm search and plain
crops for the landmarks, --swap-sides works around fh-camd naming the side
cameras backwards after some XRService restarts (tools/check_sides.py detects
it), and --record-only, --with-dark, --cpus and --keep-presence support the
bright-light and CPU-placement tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 22:36:48 -06:00
DeeJanuzandClaude Opus 5.5 b76a8c50e0 Spread the Own tracker's calibration dots over the practice area
Its fit goes wrong past its dots, and the ring (limited by the window's height) never
reached the sides, where frame-eyes' worst practice clicks were. With the Own tracker
the calibration now uses the centre, corners, and side middles of the practice area.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 22:04:06 -06:00
DeeJanuzandClaude Opus 5.5 30a9d15072 Add our own eye tracker as a gaze source in ft-gaze and the probe
ft-gaze reads frame-eyes' /dev/shm/frame-eyes-gaze and reports it as the source "own",
with each eye's gaze, where each lands on the screen, and its slip. The probe gets a
SteamVR / Own tracker toggle. With Own tracker on, it hides SteamVR's gaze and draws a
red dot per eye, its calibration fits fe-trackd's own calibration, and practice clicks
teach fe-trackd instead of the probe's correction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 21:38:17 -06:00
DeeJanuzandClaude Opus 5.5 a31d42b25c Move hand cutouts ahead to where the hands will be
The tracked hands arrive 30-60 ms after the cameras saw them and reach the
displays later still, so holes trailed moving hands. Track each hand's palm
velocity in the room and move its capsules ahead to about when the frame is
on the displays, every tick, so the holes also move smoothly between tracker
updates. Slow hands aren't moved (their velocity is noise). The control
socket gets cutouts predict on|off and cutouts lead <ms>.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:18:40 -06:00
DeeJanuz d888b6c3f1 Merge branch 'pointer-ignore' into experimental 2026-09-29 14:13:36 -06:00
DeeJanuzandClaude Opus 5.5 9c04207bb9 Let the pointer pass through panels you pick, like a performance overlay
A head-locked performance overlay kept catching the 3D mouse. It has no
input method, so SteamVR's laser passes through it, but the helper hit
tests every visible overlay with ComputeOverlayIntersection, and the dot
stuck to it whenever it crossed that corner of the view.

POINTER_IGNORE in frametop.conf now lists overlay keys the helper leaves
out of the collision, comma-separated shell patterns, so "vendor.app*"
covers a whole app, including panels it opens later. The laser starts
just before the cursor point, so an ignored panel nearer to you doesn't
catch it either.

Frametop Input Settings has a new Ignored panels page. It asks the
helper for SteamVR's overlays ("overlays", answered from the list's
thread once vrcmd has run again, even while the pointer is off), groups
them by app, and has a checkbox per panel and one for the whole app.
Frametop's own screens aren't offered, since ignoring one would leave
nothing to click the app on with the mouse. Entries for apps that
aren't open are listed so they can be removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:13:34 -06:00
DeeJanuzandClaude Opus 5.5 094a7b27f0 Don't cut out hand parts right in front of the eyes
A capsule end near the eyes' plane projects far across a screen with a huge
radius, so one bad hand estimate there tore a hole through the screens for a
moment. Clip capsules 12 cm in front of the eye, as the tracker now does too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 11:36:45 -06:00
DeeJanuzandClaude Opus 5.5 b7cbd9fe16 Cut tracked hands out of screens so you see them through it (work in progress)
Where frame-hands tracks a hand between an eye and a screen, that eye sees the
room through the screen. handcut.cpp draws the screen's buffer side by side
(one half per eye) with the hands cut out, only while a hand is in front of it.
The cutouts command turns it on or off. ft-handtest tries it on a test panel.

Not yet tested in the headset.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 10:28:27 -06:00
DeeJanuz ed9542d3b8 Merge branch 'layouts-headpin' into experimental
# Conflicts:
#	README.md
#	display-settings/ft_display_settings.py
#	display-settings/main.qml
#	docs/reference.md
2026-09-29 10:24:30 -06:00
DeeJanuz 0c40b2c6de Merge remote-tracking branch 'origin/main' into experimental 2026-09-29 10:24:04 -06:00
DeeJanuzandClaude Opus 5.5 89522e1894 Let Flatpak apps save and upload files in the desktop
The file picker hands a sandboxed app the host path of the document
portal, which the desktop's private runtime directory moves to
$runtime/doc. Inside the sandbox that path is an empty private folder,
so Brave finished downloads into it and they were lost when the session
cleaned up. Link it to /run/flatpak/doc in each installed app's
runtime folder before Plasma starts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:13:46 -06:00
DeeJanuzandClaude Opus 5.5 65cedb673c Warn in Input Settings when SteamVR hasn't loaded the pointer driver
When SteamVR crashes, safe mode can block the newest add-on, and then it
skips the ft_pointer driver at every start. The cursor still moves (the
pointer helper draws it), but clicks, scrolling and mapped actions go
through the driver's virtual controller, so none of them do anything,
and nothing said why (#4).

Input Settings now shows a warning at the top of every page with the
fix: Manage Add-Ons in SteamVR's settings, unblock ft_pointer, restart
SteamVR. It reads steamvr.vrsettings (~/.config/openvr/config on the
Frame) for blocked_by_safe_mode, a disabled driver, or SteamVR's own
safe mode. When none of those is set but SteamVR is running (the helper
answers) and @ft_pointer doesn't, it says SteamVR runs without the
driver: unblocked but not restarted yet, or not installed. The driver
ignores the "ping" it sends. It checks at startup and every 30 minutes,
since the driver only changes when SteamVR restarts.

Fixes #4

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:04:17 -06:00
DeeJanuzandClaude Opus 5.5 51e7e0fa85 Warn in Input Settings when SteamVR hasn't loaded the pointer driver
When SteamVR crashes, safe mode can block the newest add-on, and then it
skips the ft_pointer driver at every start. The cursor still moves (the
pointer helper draws it), but clicks, scrolling and mapped actions go
through the driver's virtual controller, so none of them do anything,
and nothing said why (#4).

Input Settings now shows a warning at the top of every page with the
fix: Manage Add-Ons in SteamVR's settings, unblock ft_pointer, restart
SteamVR. It reads steamvr.vrsettings (~/.config/openvr/config on the
Frame) for blocked_by_safe_mode, a disabled driver, or SteamVR's own
safe mode. When none of those is set but SteamVR is running (the helper
answers) and @ft_pointer doesn't, it says SteamVR runs without the
driver: unblocked but not restarted yet, or not installed. The driver
ignores the "ping" it sends. It checks at startup and every 30 minutes,
since the driver only changes when SteamVR restarts.

Fixes #4

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:04:17 -06:00
DeeJanuzandClaude Opus 5.5 b99fb97f32 Turn the displays off while the headset isn't used, and keep it awake on a charger
A display mount that covers the proximity sensor makes the headset seem
worn, so SteamVR never turned its displays off and they stayed on all
night. The new power service, ft-powerd (frametop-power.service), goes by
use instead: after DISPLAY_OFF_MIN minutes in which the headset and
controllers didn't move and no input device was used, it turns the
backlight off, and the next movement or input turns it back on.

The new Power tab in Frametop Display Settings sets that time and has a
Stay awake while plugged in switch. The switch sets Steam's own "When
Plugged In and Idle -> Sleep after" to Never through Steam's UI, so the
Frame stays reachable remotely while the power button still works.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 09:34:34 -06:00
DeeJanuzandClaude Opus 5.5 0be802c7e1 Add named layouts and a head pin for screens
Named layouts: Save current arrangement in Frametop Display Settings now
asks for a name, and saved layouts are listed with the presets under
Arrangement, with rename and delete next to the list. A named layout is
the custom arrangement under a name: each screen's place relative to
your head, width, curve, and pin, but not resolution or scale. Using one
copies it into the custom arrangement, so desktop start, Meta+Shift+R,
and Arrange now apply it unchanged; "active" remembers the name, and a
plain capture clears it. ft-layout gains save, use, layouts, rename, and
delete. A layout saved with fewer screens than there are now leaves the
others where they were saved last, or where the preset puts them.

Head pin: ft-screens' pin command takes "head" as well as left and
right, and pins the screen to the headset (device 0) where it is, like a
HUD. A head-pinned screen skips the wrist facing rule and shows whenever
the screens do. Carrying it re-pins it to the head on release, like a
wrist pin, so it can be adjusted in VR. The Visibility tab (now
Visibility & pins) sets each screen's pin: in the room, either wrist, or
your head, and the pin command now rejects anything but left, right, or
head (it used to take anything else as left). This removes the "no HUD"
limit from the docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 22:33:57 -06:00
DeeJanuzandClaude Opus 5.5 f01c84f9a3 Use the other eye when the tracker loses one, and add a headset fit check
SteamVR's combined gaze keeps going on one eye, but it holds the lost
eye's yaw, so the gaze moves half as far sideways as the eyes do.
ft-gaze now reads each eye's tracking uncertainty and raw measurement
from eye-server.mmap. When the tracker loses an eye, ft-gazed takes the
gaze from the other one, plus the offset that eye usually shows
against both, learned while both are seen. On a recording, one eye
alone came out a median 0.8 degrees from both eyes' gaze.

Glances down at the keyboard, past every screen, aren't sent. The
pointer stays put, and eyes lost there don't count as lost.

The gaze probe gets a Headset fit mode. It shows per-eye tracking,
openness and confidence, maps where each eye gets lost, gives hints,
and has a guided check. The settings app opens it from the Gaze page
and shows how often each eye is lost. The probe can also test each eye
alone, and its side panel now collapses to a title bar so the dot
isn't hidden behind it.

Snapping to UI elements is deferred; the mouse drag is the correction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 22:07:41 -06:00
DeeJanuzandClaude Opus 5.5 e60cb2f28a Add a report script for keys that stick or leak
scripts/keys-report.py records what happens to the modifiers, Tab, and
Esc while someone reproduces a key problem: each press, release, and
autorepeat as the relay reads it from a physical keyboard, and as it
comes out of the relay's virtual keyboard to gamescope and SteamVR. It
adds the relay's view of every device (role, grabbed), which programs
have each input node open, and the relay, pointer helper, and desktop
logs for the same time. Other keys show only as "other key", so nothing
typed ends up in the report, and Bluetooth addresses are masked.

For #2: Shift+Tab in the Frametop desktop opened the SteamVR dashboard
and then stayed held, which doesn't happen here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 21:35:16 -06:00
DeeJanuzandClaude Opus 5.5 18aa0fec3a Put clicks where the cursor is on scaled desktop screens
With a screen's scale set to anything but 100%, clicks landed away from
the cursor, further off the further from the top left (#3). ft-screens
hands KWin panel positions in buffer pixels, and KWin's nested Wayland
backend (6.2.5, WaylandInputDevice) adds surface coordinates to its
output's logical position without dividing by the output's scale. At
125% a click at the middle of a 3440x1440 screen, (1720, 720), reached
KWin as logical (1720, 720), pixel (2150, 900).

ft-screens now keeps a scale per screen and divides pointer positions by
it. ft-layout sends each screen's scale, as KWin reports it after
applying, with a new "scale N s" command, whenever it applies scales:
at desktop start and from Frametop Display Settings. Screens default to
1, so an ft-layout that never sends it keeps the old behaviour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 21:33:03 -06:00
DeeJanuzandClaude Opus 5.5 0857a54fab Don't let resting controllers take the laser from the mouse
A controller released the 3D mouse's pointer on a single pose sample
faster than 0.35 m/s or 2 rad/s, once the mouse had been still for
500 ms. The helper polls about every 8 ms, so one noisy sample was
enough: a knock on the desk, or a tracking jump when the headset's
cameras pick a resting controller up again (#1).

A controller now has to stay over the limit for 100 ms in a row, and
only samples with a normal tracking result (Running_OK) count. The new
POINTER_CONTROLLER_PICKUP setting (1 by default, 0.5 to 5) scales both
limits; it's a slider on the Pointer page of Frametop Input Settings
and applies live. A controller picked up for real still gets the laser
back through SteamVR's hand role, which follows its touch sensors. The
release log now records the speed and spin that triggered it, to tune
the default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 21:31:22 -06:00
139 changed files with 21968 additions and 482 deletions

No files matched your search

+6
View File
@@ -5,3 +5,9 @@ target/
captures/
__pycache__/
frametop-report-*.txt
# Eye-camera recordings (biometric) never go in the repo: they live in
# ~/.local/share/frametop/eyes/captures. These catch strays (frame dumps, a lab venv).
*.raw
*.pgm
.venv/
.frame-job.d/
+24 -7
View File
@@ -8,6 +8,8 @@ The mouse shows up as a small dot anchored in the room. It works on the SteamVR
It comes with two settings apps, Frametop Display Settings for the screens and Frametop Input Settings for mice, keyboards, and button mappings, plus fixes that let Bluetooth LE mice and keyboards like the Swiftpoint Z3 reconnect after they sleep.
When you're not wearing the headset, Frametop can turn its displays off and keep it awake on the charger, so you can still reach it remotely. This works even on a stand or mount that makes the headset seem worn.
Frametop is an independent project, not made by or affiliated with Valve.
## Install on the headset
@@ -28,7 +30,7 @@ You need a Steam Frame with an internet connection, a keyboard (Bluetooth, or th
After the restart, Launch a program → Desktop opens the multi-screen desktop, with its screens arranged around where you're facing. Frametop Display Settings and Frametop Input Settings are in the desktop's application menu, under Settings.
If you work in the desktop for long stretches, stop Steam from putting the headset to sleep while it's plugged in: in Steam, open Settings → Power, and under When Plugged In and Idle set Sleep after to Never. By default Steam suspends the Frame after an hour without input, even while it charges. The displays still turn off a few seconds after you take the headset off.
If you work in the desktop for long stretches, or leave the headset on a stand, open Frametop Display Settings → Power. Turn on Stay awake while plugged in: by default Steam puts the Frame to sleep after an hour without input, even while it charges. And choose when the displays turn off while the headset isn't used. SteamVR turns them off a few seconds after you take the headset off, but a stand or mount that covers the proximity sensor inside it makes the headset seem worn, and its displays stay on all night.
### Add a Bluetooth mouse or keyboard
@@ -49,14 +51,26 @@ If you work in the desktop for long stretches, stop Steam from putting the heads
| Click the curve button (next to the bar) | Curves the screen around you, or flattens it |
| Drag the roll button sideways, or scroll on it | Rolls the screen; it snaps level near straight |
| While carrying a screen, sweep its laser across your other controller's ring, then let go | Pins it to that wrist, at its size and distance, as you hold it when you let go; it shows while you see its front. Grab its bar to adjust it (it stays pinned); sweep across the ring again to take it off |
| Set a screen to On your head (Frametop Display Settings, Visibility & pins) | Pins it to your head where it is, like a HUD. Grab its bar to move it; it stays on your head |
| Save current arrangement… (Frametop Display Settings, Layout) | Saves where the screens are, with their sizes and pins, under a name. Pick a saved layout under Arrangement and press Arrange now to switch to it |
| Meta+Shift+R in the desktop | Puts the screens back in their layout (also in the menu as Reset Screen Layout, and mappable to a mouse button) |
| Meta+Shift+H in the desktop | Hides or shows all screens (also in the menu as Hide/Show Screens, and mappable). The Visibility & wrist tab of Frametop Display Settings can instead show them only with the dashboard open, or while you look at your wrist |
| Play a VR game | The screens hide and your controllers stay in the game. Open the SteamVR dashboard, or press Meta+Shift+H, to see and use them. To keep them visible over games, change During VR games on the Visibility & wrist tab; the controllers still stay in the game, and you use the screens with the mouse or the dashboard |
| Meta+Shift+H in the desktop | Hides or shows all screens (also in the menu as Hide/Show Screens, and mappable). The Visibility & pins tab of Frametop Display Settings can instead show them only with the dashboard open, or while you look at your wrist |
| Leave the headset on a stand | Its displays turn off once it has gone unused for the time set in Frametop Display Settings → Power, even if the stand covers its proximity sensor. Pick it up, or use any mouse, keyboard, or button, and they come back on |
| Play a VR game | The screens hide and your controllers stay in the game. Open the SteamVR dashboard, or press Meta+Shift+H, to see and use them. To keep them visible over games, change During VR games on the Visibility & pins tab; the controllers still stay in the game, and you use the screens with the mouse or the dashboard |
You can map the mouse's extra buttons to actions such as Toggle SteamVR dashboard, Recenter pointer, or Head follow on/off on the Buttons page of Frametop Input Settings, and the Frame controllers' buttons on its Controllers page. Pointer speed, dot size, and the rest are on its Pointer page and take effect immediately. Head follow, which is experimental and off by default, makes the pointer come along when you turn your head: it stays put until your head turns past the leash angle, then glides back to its place in your view, and a leash of 0 keeps it fixed in your view. It's only lightly tested and not polished; tuning its settings, or improving how it feels, is open to anyone who wants to take it further.
You can map the mouse's extra buttons to actions such as Toggle SteamVR dashboard, Recenter pointer, or Head follow on/off on the Buttons page of Frametop Input Settings, and the Frame controllers' buttons on its Controllers page. Pointer speed, dot size, and the rest are on its Pointer page and take effect immediately. If a panel you only look at, such as a performance overlay that follows your view, keeps catching the dot, tick it (or its whole app) on the Ignored panels page, and the pointer passes through it. Head follow, which is experimental and off by default, makes the pointer come along when you turn your head: it stays put until your head turns past the leash angle, then glides back to its place in your view, and a leash of 0 keeps it fixed in your view. It's only lightly tested and not polished; tuning its settings, or improving how it feels, is open to anyone who wants to take it further.
Restarting the desktop (Restart desktop in Frametop Display Settings) closes its windows, but background work you started in it, such as servers, tmux sessions, or builds, keeps running.
### Leave the headset on a stand and reach it remotely
To keep the Frame on and connected while you're not wearing it, for SSH, remote desktop, or anything else running on it, open the Power tab in Frametop Display Settings:
- Turn off when unused for: how long the headset can go unused before its displays turn off (Never by default). Unused means the headset and controllers haven't moved and no mouse, keyboard, or button was used. SteamVR normally turns the displays off when its proximity sensor says the headset came off, but a stand or mount that covers the sensor makes the headset seem worn, so the displays stay on all night. This setting doesn't depend on the sensor. Pick the headset up or use any input, and the displays come back on.
- Stay awake while plugged in: stops Steam from putting the Frame to sleep while it charges. By default Steam puts it to sleep after an hour without input, even on the charger, which ends remote sessions. This is Steam's own Settings → Power → When Plugged In and Idle setting, so the power button still puts the Frame to sleep, and Steam's battery setting still applies.
With the displays off, the headset keeps tracking and rendering, so it uses about as much power as in use. Leave it on a charger that keeps up with that: a USB-C PD charger, not a 5 V one.
## Known limitations
This is an early release, tested on one Steam Frame (SteamOS 0.3.0 build 20260922, SteamVR 2.17.10).
@@ -64,11 +78,12 @@ This is an early release, tested on one Steam Frame (SteamOS 0.3.0 build 2026092
- A SteamOS or SteamVR update can break parts of it until Frametop catches up. If something stops working after an update, please report it.
- The first install downloads 1–2 GB for the build container and compiles everything on the headset, which takes several minutes.
- During a VR game you can't show the screens with a controller button, because the game owns the buttons. Open the SteamVR dashboard, press Meta+Shift+H, or use a mapped mouse button instead.
- Flatscreen games aren't detected as games. If your controllers end up working the screens instead of the game, set Controllers on the screens to "Only with the SteamVR dashboard open" (Frametop Display Settings, Visibility & wrist tab).
- Flatscreen games aren't detected as games. If your controllers end up working the screens instead of the game, set Controllers on the screens to "Only with the SteamVR dashboard open" (Frametop Display Settings, Visibility & pins tab).
- Typing follows your last click. A controller click on a panel other than the screens (the dashboard, a Steam app) doesn't move typing there; click it with the mouse, or click a screen to bring typing back.
- The screens don't draw a mouse cursor of their own. The 3D mouse's dot or SteamVR's laser shows where you're pointing.
- On SteamVR's Settings page, the 3D mouse shows a laser beam and a larger hit dot, like a controller. SteamVR doesn't tell other programs where that page is (unlike Steam's pages, such as Library), so the mouse used to miss most of it: clicks went through to a desktop screen behind, and the dot disappeared. As a workaround, on that page only, the laser starts near your eye and SteamVR finds the page itself. See docs/design.md.
- Remote desktop over VNC (`./desktops.sh remote on`) needs Tailscale on the Frame.
- Remote desktop over VNC (Frametop Remote Access in the app menu, or `./desktops.sh remote on`) needs Tailscale on the Frame. It shows the primary screen only. The app turns it on and off, shows the address, and shows, copies, or changes the VNC password. The password is made at random on the Frame and kept in `~/.config/frametop-remote` (only you can read it); VNC limits it to 8 characters, and the tailnet encrypts the connection. Turning it on in a desktop that started with it off takes a desktop restart.
- Turning the displays off on a stand only turns their backlight off. SteamVR has no way for other programs to put the headset in standby, so tracking and rendering keep running, and the headset draws nearly its full power.
## Reporting problems
@@ -92,6 +107,7 @@ cd ~/frametop && git pull && ./install.sh
./desktops.sh uninstall # the launcher's Desktop entry goes back to the stock desktop
./desktops.sh relay uninstall
pointer/helper/run.sh uninstall
power/run.sh uninstall
pointer/driver/install.sh uninstall # then restart SteamVR
input-settings/install.sh uninstall
display-settings/install.sh uninstall
@@ -100,7 +116,7 @@ setup/bluetooth/install.sh uninstall # if you installed the Bluetooth fixes
## How it works
A Plasma session runs nested inside ft-screens (`screens/`), a small Wayland compositor. KWin opens one window per screen, ft-screens sets each window's size, and each frame goes to SteamVR as an overlay without being copied. An input relay (`input/`) keeps Bluetooth mice working in SteamVR and feeds the mouse to the 3D pointer, which drives a virtual SteamVR controller (`pointer/`). [docs/reference.md](docs/reference.md) covers each piece, and [docs/design.md](docs/design.md) explains the design and what we learned about SteamVR on the Frame. [docs/hazards.md](docs/hazards.md) lists known ways the input handling can go wrong.
A Plasma session runs nested inside ft-screens (`screens/`), a small Wayland compositor. KWin opens one window per screen, ft-screens sets each window's size, and each frame goes to SteamVR as an overlay without being copied. An input relay (`input/`) keeps Bluetooth mice working in SteamVR and feeds the mouse to the 3D pointer, which drives a virtual SteamVR controller (`pointer/`). A power service (`power/`) turns the displays off while the headset isn't used. [docs/reference.md](docs/reference.md) covers each piece, and [docs/design.md](docs/design.md) explains the design and what we learned about SteamVR on the Frame. [docs/hazards.md](docs/hazards.md) lists known ways the input handling can go wrong.
| Folder | What it is |
| --- | --- |
@@ -111,6 +127,7 @@ A Plasma session runs nested inside ft-screens (`screens/`), a small Wayland com
| `layout/` | ft-layout: where the screens float, and their sizes. |
| `input/` | The input relay (Bluetooth mice and keyboards, button maps). |
| `pointer/` | The 3D mouse: SteamVR driver, helper service, and a probe tool. |
| `power/` | ft-powerd: turns the displays off while the headset isn't used. |
| `display-settings/`, `input-settings/` | The two settings apps (Kirigami, Python). |
| `setup/` | The build container and the Bluetooth fixes. See [setup/README.md](setup/README.md). |
| `scripts/` | Helpers the installers use. They run commands locally on the Frame, or over SSH from a PC. |
+186 -20
View File
@@ -10,10 +10,16 @@ the dev container:
1920x1080 worth of pixels, rotation for portrait.)
- Visibility (ft-screens): when the screens show (always, only with the SteamVR
dashboard open, while you look at a controller, or only when toggled), the wrist
angle within which a pinned screen shows, and pin or unpin all screens.
- Layout: a preset (curved or flat, rows, distance, gap, height) or the arrangement
captured from where the screens are now, with a preview; arrange now; save the
current arrangement; arrange automatically when the desktop starts.
angle within which a pinned screen shows, and pinning each screen to a wrist or
your head.
- Layout: a preset (curved or flat, rows, distance, gap, height) or a named layout
saved from where the screens are, with a preview; arrange now; save the current
arrangement under a name; rename and delete; arrange automatically when the
desktop starts.
- Power: how long the headset can go unused before ft-powerd turns its displays off
(DISPLAY_OFF_MIN; the service's state comes from its control socket, @ft_powerd),
and whether the Frame stays awake while plugged in, which is Steam's own setting
(steam_settings.py; the value from before is kept as STEAM_SLEEP_AC_BEFORE).
Settings go to ~/.config/frametop.conf and ~/.config/frametop-layout.json. Anything
that touches SteamVR runs layout/ft-layout on the host.
Launch with display-settings/ft-display-settings (host wrapper).
@@ -22,8 +28,9 @@ import os
import shutil
import socket
import sys
import threading
from PySide6.QtCore import Property, QObject, QProcess, QTimer, QUrl, Signal, Slot
from PySide6.QtCore import Property, QObject, QProcess, Qt, QTimer, QUrl, Signal, Slot
from PySide6.QtGui import QGuiApplication, QIcon
from PySide6.QtQml import QQmlApplicationEngine
from PySide6.QtQuickControls2 import QQuickStyle
@@ -32,6 +39,7 @@ HERE = os.path.dirname(os.path.abspath(__file__))
LAYOUT_DIR = os.path.join(HERE, "..", "layout")
sys.path.insert(0, LAYOUT_DIR)
import ft_layout # noqa: E402 (pure Python: the same geometry ft-layout uses)
import steam_settings # noqa: E402
FT_LAYOUT = os.path.join(LAYOUT_DIR, "ft-layout")
DESKTOPS = os.path.join(HERE, "..", "desktops.sh")
@@ -46,6 +54,10 @@ SCREEN_RESOLUTIONS = [(1920, 1080, ""), (2560, 1440, ""), (3840, 2160, "4K"), (2
(2560, 1600, "16:10"), (1080, 1920, "portrait"), (1440, 2560, "portrait"),
(2160, 3840, "portrait 4K")]
FT_SCREENS = "\0ft_screens"
FT_POWERD = "\0ft_powerd"
# Steam's default for "When Plugged In and Idle -> Sleep after", to go back to when
# nothing was saved.
STEAM_SLEEP_AC_DEFAULT = 3600
SCALES = [0.75, 1.0, 1.25, 4 / 3, 1.5, 1.75, 2.0]
ROTATIONS = [("normal", "Landscape"), ("left", "Portrait"), ("right", "Portrait (flipped)")]
@@ -83,7 +95,9 @@ def host_command(*cmd):
class Backend(QObject):
changed = Signal()
busyChanged = Signal()
powerChanged = Signal()
message = Signal(str, bool) # text, is error
_steamDone = Signal(object, object, str) # Steam's sleep settings or None, error or None, what was done
def __init__(self):
super().__init__()
@@ -95,6 +109,15 @@ class Backend(QObject):
self._sock.bind("") # an abstract address ft-screens can reply to
self._sock.settimeout(1.0)
self._started = {} # conf values the running desktop started with
self._psock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
self._psock.bind("") # for ft-powerd's replies
self._psock.settimeout(0.5)
self._powerd = None # ft-powerd's status: (state, seconds unused, timeout seconds); None: not running
self._steam = None # Steam's sleep settings: {"ac": seconds, "battery": seconds}
self._steam_error = ""
self._steam_busy = False
self._steamDone.connect(self._steam_done, Qt.QueuedConnection)
self._pins = [] # each running screen's pin: none | left | right | head
self.poll = QTimer(interval=3000, timeout=self._check_running)
self.poll.start()
self._check_running()
@@ -115,11 +138,23 @@ class Backend(QObject):
# from the container, so ft_layout.nested_env() doesn't work here).
running = os.path.exists(f"/run/user/{os.getuid()}/frametop/wayland-0")
count = self._screens_running() if running and ft_layout.backend() == "screens" else 0
if running != self._running or count != self._running_count:
pins = self._read_pins(count)
if running != self._running or count != self._running_count or pins != self._pins:
if running != self._running or count != self._running_count:
self._started = self._conf() if running else {}
self._running = running
self._running_count = count
self._started = self._conf() if running else {}
self._pins = pins
self.changed.emit()
self._check_powerd()
def _read_pins(self, count):
pins = []
for i in range(count):
reply = self._ask_screens(f"get {i + 1}")
f = reply.split() if reply and reply.startswith("ok") else []
pins.append(f[16] if len(f) > 16 else "none")
return pins
def _ask_screens(self, text):
"""Request/reply to ft-screens; None if it isn't running."""
@@ -370,20 +405,114 @@ class Backend(QObject):
else:
self._ask_screens(f"gesture {v['gesture_hand']} {float(v['gesture_angle']):.1f}")
@Slot(str)
def pinAll(self, hand):
reply = self._ask_screens(f"pin all {hand}") if self._running else None
if reply and reply.startswith("ok"):
self.message.emit(f"All screens ride on your {hand} wrist now; grab a screen's bar to take it off. "
"Save current arrangement keeps it.", False)
else:
self.message.emit(f"Couldn't pin: {reply or 'the desktop is not running'}", True)
@Property("QVariantList", notify=changed)
def pins(self):
return self._pins
@Slot(str, str)
def pin(self, which, where):
"""Pin screen `which` (1-based, or "all") to "left", "right", or "head" as it is
now, or take it off ("none")."""
cmd = f"unpin {which}" if where == "none" else f"pin {which} {where}"
reply = self._ask_screens(cmd) if self._running else None
if not (reply and reply.startswith("ok")):
self.message.emit(f"Couldn't {'unpin' if where == 'none' else 'pin'}: "
f"{reply or 'the desktop is not running'}", True)
elif which == "all" and where != "none":
place = "on your head" if where == "head" else f"on your {where} wrist"
self.message.emit(f"All screens ride {place} now. Save current arrangement (Layout) keeps it.", False)
self._check_running()
# --- power: ft-powerd and Steam's sleep setting ---
def _check_powerd(self):
try:
self._psock.sendto(b"status", FT_POWERD)
reply = self._psock.recv(256).decode().split()
status = (reply[1], float(reply[2]), float(reply[3])) if reply[:1] == ["ok"] else None
except (OSError, IndexError, ValueError):
status = None
if status != self._powerd:
self._powerd = status
self.powerChanged.emit()
@Property("QVariantMap", notify=powerChanged)
def power(self):
try:
off_min = float(ft_layout.read_conf().get("DISPLAY_OFF_MIN") or 0)
except ValueError:
off_min = 0.0
state, unused, _ = self._powerd or ("", 0, 0)
return {"offMinutes": off_min, "service": self._powerd is not None, "state": state, "unused": unused,
"steam": self._steam is not None, "steamBusy": self._steam_busy, "steamError": self._steam_error,
"acSleep": self._steam["ac"] if self._steam else -1,
"batterySleep": self._steam["battery"] if self._steam else -1}
@Slot(float)
def setDisplayOffMinutes(self, minutes):
"""ft-powerd re-reads frametop.conf within 2 s."""
write_conf_value("DISPLAY_OFF_MIN", f"{max(0.0, minutes):g}")
self.powerChanged.emit()
@Slot()
def unpinAll(self):
reply = self._ask_screens("unpin all") if self._running else None
if not (reply and reply.startswith("ok")):
self.message.emit(f"Couldn't unpin: {reply or 'the desktop is not running'}", True)
def displaysOffNow(self):
try:
self._psock.sendto(b"off", FT_POWERD)
reply = self._psock.recv(256).decode()
except OSError:
reply = "error the power service isn't running"
if not reply.startswith("ok"):
self.message.emit(f"Couldn't turn the displays off: {reply.split(' ', 1)[-1]}", True)
self._check_powerd()
def _steam_call(self, what, fn):
"""Runs fn, which talks to Steam (up to a few seconds), off the UI thread, then reads
Steam's sleep settings; _steam_done gets them on the UI thread."""
if self._steam_busy:
return
self._steam_busy = True
self.powerChanged.emit()
def work():
try:
fn()
self._steamDone.emit(steam_settings.sleep_settings(), None, what)
except (steam_settings.SteamUnreachable, OSError, ValueError) as e:
self._steamDone.emit(None, str(e), what)
threading.Thread(target=work, daemon=True).start()
def _steam_done(self, settings, error, what):
self._steam_busy = False
if settings is not None:
self._steam, self._steam_error = settings, ""
else:
self._steam, self._steam_error = None, error
if what:
self.message.emit(f"Couldn't change Steam's sleep setting: {error}", True)
self.powerChanged.emit()
@Slot()
def refreshPower(self):
self._check_powerd()
self._steam_call("", lambda: None)
@Slot(bool)
def setStayAwake(self, on):
"""Steam's "When Plugged In and Idle -> Sleep after" is Never while this is on. The value
from before is kept in frametop.conf and goes back when it's turned off."""
def change():
ac = steam_settings.sleep_settings()["ac"]
if on:
if ac > 0:
write_conf_value("STEAM_SLEEP_AC_BEFORE", str(ac))
steam_settings.set_sleep_setting("system_idle_suspend_ac_sec", 0)
elif ac == 0:
try:
before = int(ft_layout.read_conf().get("STEAM_SLEEP_AC_BEFORE") or STEAM_SLEEP_AC_DEFAULT)
except ValueError:
before = STEAM_SLEEP_AC_DEFAULT
steam_settings.set_sleep_setting("system_idle_suspend_ac_sec", before if before > 0 else STEAM_SLEEP_AC_DEFAULT)
self._steam_call("stay awake" if on else "sleep", change)
@Slot()
def restartDesktop(self):
@@ -401,7 +530,44 @@ class Backend(QObject):
@Slot(str)
def setMode(self, mode):
self._edit_layout(lambda l: l.__setitem__("mode", mode))
def edit(layout):
layout["mode"] = mode
layout.pop("active", None)
self._edit_layout(edit)
@Property("QVariantList", notify=changed)
def layoutNames(self):
return ft_layout.layout_names(ft_layout.load_layout())
@Slot(str)
def useLayout(self, name):
"""A named layout as the arrangement (Arrange now puts the screens there)."""
try:
self._edit_layout(lambda l: ft_layout.use_named(l, name))
except RuntimeError as e:
self.message.emit(str(e), True)
@Slot(str)
def saveLayout(self, name):
try:
ft_layout.check_name(name)
except RuntimeError as e:
return self.message.emit(str(e), True)
self._run(f"Saving the arrangement as {' '.join(name.split())}", "save", name)
@Slot(str, str)
def renameLayout(self, old, new):
try:
self._edit_layout(lambda l: ft_layout.rename_named(l, old, new))
except RuntimeError as e:
self.message.emit(str(e), True)
@Slot(str)
def deleteLayout(self, name):
try:
self._edit_layout(lambda l: ft_layout.delete_named(l, name))
except RuntimeError as e:
self.message.emit(str(e), True)
@Slot(str, "QVariant")
def setPreset(self, key, value):
+298 -31
View File
@@ -12,11 +12,13 @@ Kirigami.ApplicationWindow {
// Pages as tabs across the top (a side drawer was easy to miss).
readonly property var pages: backend.backend === "screens"
? [{ text: "Screens", icon: "video-display", page: screensPage },
{ text: "Layout", icon: "view-grid", page: layoutPage },
{ text: "Visibility & wrist", icon: "view-visible", page: visibilityPage }]
: [{ text: "Screens", icon: "video-display", page: screensPage },
{ text: "Layout", icon: "view-grid", page: layoutPage }]
? [{ name: "screens", text: "Screens", icon: "video-display", page: screensPage },
{ name: "layout", text: "Layout", icon: "view-grid", page: layoutPage },
{ name: "visibility", text: "Visibility & pins", icon: "view-visible", page: visibilityPage },
{ name: "power", text: "Power", icon: "preferences-system-power-management", page: powerPage }]
: [{ name: "screens", text: "Screens", icon: "video-display", page: screensPage },
{ name: "layout", text: "Layout", icon: "view-grid", page: layoutPage },
{ name: "power", text: "Power", icon: "preferences-system-power-management", page: powerPage }]
header: Controls.TabBar {
id: tabs
@@ -29,7 +31,7 @@ Kirigami.ApplicationWindow {
onClicked: root.show(modelData.page)
}
}
Component.onCompleted: currentIndex = ({ layout: 1, visibility: 2 })[startPage] || 0
Component.onCompleted: currentIndex = Math.max(0, root.pages.findIndex(p => p.name === startPage))
}
function show(page) {
@@ -37,8 +39,16 @@ Kirigami.ApplicationWindow {
pageStack.push(page)
}
// FT_DISPLAY_PAGE=layout|visibility opens the app on that page.
pageStack.initialPage: ({ layout: layoutPage, visibility: visibilityPage })[startPage] || screensPage
// FT_DISPLAY_PAGE=layout|visibility|power opens the app on that page.
pageStack.initialPage: ({ layout: layoutPage, visibility: visibilityPage, power: powerPage })[startPage] || screensPage
// "1 hour", "15 minutes", "30 seconds".
function duration(seconds) {
const unit = (n, word) => n + " " + word + (n === 1 ? "" : "s")
if (seconds >= 3600 && seconds % 3600 === 0) return unit(seconds / 3600, "hour")
if (seconds >= 60 && seconds % 60 === 0) return unit(seconds / 60, "minute")
return unit(seconds, "second")
}
Connections {
target: backend
@@ -66,6 +76,89 @@ Kirigami.ApplicationWindow {
]
}
// Save the arrangement under a name, or rename a saved layout.
Kirigami.PromptDialog {
id: nameDialog
property string mode: "save" // save | rename
property string oldName: ""
readonly property var names: backend.layoutNames
readonly property string name: nameField.text.trim().split(/\s+/).join(" ")
readonly property bool taken: name !== oldName && names.indexOf(name) >= 0
readonly property bool ok: name !== "" && !(mode === "rename" && taken)
title: mode === "save" ? "Save the arrangement" : "Rename " + oldName
standardButtons: Kirigami.Dialog.NoButton
function openFor(m, text) {
mode = m
oldName = m === "rename" ? text : ""
nameField.text = text
open()
nameField.forceActiveFocus()
nameField.selectAll()
}
function accept() {
if (!ok) return
close()
if (mode === "save") backend.saveLayout(name)
else if (name !== oldName) backend.renameLayout(oldName, name)
}
ColumnLayout {
Controls.Label {
Layout.fillWidth: true
wrapMode: Text.Wrap
text: nameDialog.mode === "save"
? "Where the screens are now, with their sizes, curves, and pins, under this name:"
: "New name:"
}
Controls.TextField {
id: nameField
Layout.fillWidth: true
maximumLength: 40
onAccepted: nameDialog.accept()
}
Controls.Label {
visible: nameDialog.taken
opacity: 0.7
text: nameDialog.mode === "save" ? "Replaces the saved layout with that name."
: "There's already a layout with that name."
}
}
customFooterActions: [
Kirigami.Action {
text: nameDialog.mode === "save" ? "Save" : "Rename"
icon.name: nameDialog.mode === "save" ? "document-save" : "edit-rename"
enabled: nameDialog.ok
onTriggered: nameDialog.accept()
},
Kirigami.Action {
text: "Cancel"
icon.name: "dialog-cancel"
onTriggered: nameDialog.close()
}
]
}
Kirigami.PromptDialog {
id: deleteDialog
property string name: ""
title: "Delete " + name + "?"
subtitle: "The screens stay where they are; only the saved layout goes."
standardButtons: Kirigami.Dialog.NoButton
customFooterActions: [
Kirigami.Action {
text: "Delete"
icon.name: "edit-delete"
onTriggered: { deleteDialog.close(); backend.deleteLayout(deleteDialog.name) }
},
Kirigami.Action {
text: "Cancel"
icon.name: "dialog-cancel"
onTriggered: deleteDialog.close()
}
]
}
// ---------------------------------------------------------------- Screens
Component {
id: screensPage
@@ -335,6 +428,15 @@ Kirigami.ApplicationWindow {
property var layout: backend.layout
property var preset: layout.preset || {}
property bool hasCustom: (layout.screens || []).some(s => s.pos !== undefined)
// Named layouts: the arrangement is one of them (named) when it came from it, and
// hasn't been placed by hand and saved without a name since.
property var names: backend.layoutNames
property bool fromNamed: names.indexOf(layout.active) >= 0
property bool named: layout.mode === "custom" && fromNamed
property bool unnamed: names.length === 0 || ((hasCustom || layout.mode === "custom") && !fromNamed)
property var choices: [{ text: "Curved around you", value: "arc" }, { text: "Flat wall", value: "flat" }]
.concat(names.map(n => ({ text: n, value: "layout:" + n })))
.concat(unnamed ? [{ text: names.length ? "Unnamed arrangement" : "Saved arrangement", value: "custom" }] : [])
actions: [
Kirigami.Action {
@@ -345,11 +447,12 @@ Kirigami.ApplicationWindow {
onTriggered: backend.arrange()
},
Kirigami.Action {
text: "Save current arrangement"
text: "Save current arrangement…"
icon.name: "document-save"
tooltip: "Use where the screens are now (placed by hand) as the layout"
tooltip: "Save where the screens are now (placed by hand) as a named layout, and use it"
enabled: backend.desktopRunning && backend.busy === ""
onTriggered: backend.capture()
onTriggered: nameDialog.openFor("save", lpage.named ? lpage.layout.active
: "Layout " + (lpage.names.length + 1))
}
]
@@ -366,27 +469,49 @@ Kirigami.ApplicationWindow {
Kirigami.FormLayout {
Layout.fillWidth: true
Controls.ComboBox {
RowLayout {
Kirigami.FormData.label: "Arrangement:"
model: [
{ text: "Curved around you", value: "arc" },
{ text: "Flat wall", value: "flat" },
{ text: "Saved arrangement", value: "custom" }
]
textRole: "text"
valueRole: "value"
currentIndex: lpage.layout.mode === "custom" ? 2 : (lpage.preset.kind === "flat" ? 1 : 0)
onActivated: {
if (currentValue === "custom") backend.setMode("custom")
else backend.setPreset("kind", currentValue)
Controls.ComboBox {
model: lpage.choices
textRole: "text"
valueRole: "value"
currentIndex: lpage.layout.mode !== "custom" ? (lpage.preset.kind === "flat" ? 1 : 0)
: lpage.named ? 2 + lpage.names.indexOf(lpage.layout.active)
: lpage.choices.length - 1
onActivated: {
if (currentValue === "custom") backend.setMode("custom")
else if (currentValue.startsWith("layout:")) backend.useLayout(currentValue.slice(7))
else backend.setPreset("kind", currentValue)
}
}
Controls.ToolButton {
visible: lpage.named
icon.name: "edit-rename"
text: "Rename…"
display: Controls.AbstractButton.IconOnly
Controls.ToolTip.text: text
Controls.ToolTip.visible: hovered
onClicked: nameDialog.openFor("rename", lpage.layout.active)
}
Controls.ToolButton {
visible: lpage.named
icon.name: "edit-delete"
text: "Delete…"
display: Controls.AbstractButton.IconOnly
Controls.ToolTip.text: text
Controls.ToolTip.visible: hovered
onClicked: { deleteDialog.name = lpage.layout.active; deleteDialog.open() }
}
}
Controls.Label {
visible: lpage.layout.mode === "custom"
Kirigami.FormData.label: ""
text: lpage.hasCustom ? "Where the screens were when you saved. Pick a preset to edit."
: "Nothing saved yet: place the screens by hand, then Save current arrangement."
text: lpage.named ? "Where the screens were when you saved it. Arrange now puts them there. "
+ "Save current arrangement updates it or saves a new one."
: lpage.hasCustom ? "Where the screens were when you saved. Save current arrangement "
+ "names it. Pick a preset to edit."
: "Nothing saved yet: place the screens by hand, then Save current arrangement."
opacity: 0.7
wrapMode: Text.Wrap
Layout.maximumWidth: Kirigami.Units.gridUnit * 20
@@ -677,10 +802,28 @@ Kirigami.ApplicationWindow {
}
}
Kirigami.Separator { Kirigami.FormData.isSection: true; Kirigami.FormData.label: "Screens on a wrist" }
Kirigami.Separator { Kirigami.FormData.isSection: true; Kirigami.FormData.label: "Pinned screens" }
Repeater {
model: backend.pins
delegate: Controls.ComboBox {
required property var modelData
required property int index
Kirigami.FormData.label: "Screen " + (index + 1) + ":"
model: [
{ text: "In the room", value: "none" },
{ text: "On the left wrist", value: "left" },
{ text: "On the right wrist", value: "right" },
{ text: "On your head", value: "head" }
]
textRole: "text"
valueRole: "value"
currentIndex: Math.max(0, ["none", "left", "right", "head"].indexOf(modelData))
onActivated: backend.pin(String(index + 1), currentValue)
}
}
RowLayout {
Kirigami.FormData.label: "Show while facing you within:"
Kirigami.FormData.label: "Wrist screens show within:"
Controls.Slider {
id: wrist
from: 20; to: 120; stepSize: 1
@@ -695,17 +838,22 @@ Kirigami.ApplicationWindow {
Controls.Button {
text: "Pin to left wrist"
enabled: backend.desktopRunning
onClicked: backend.pinAll("left")
onClicked: backend.pin("all", "left")
}
Controls.Button {
text: "Pin to right wrist"
enabled: backend.desktopRunning
onClicked: backend.pinAll("right")
onClicked: backend.pin("all", "right")
}
Controls.Button {
text: "Pin to head"
enabled: backend.desktopRunning
onClicked: backend.pin("all", "head")
}
Controls.Button {
text: "Unpin"
enabled: backend.desktopRunning
onClicked: backend.unpinAll()
onClicked: backend.pin("all", "none")
}
}
}
@@ -720,7 +868,126 @@ Kirigami.ApplicationWindow {
+ "then let go: it rides on that wrist at that size and distance, however far away. To adjust a "
+ "pinned screen, grab its bar, move it, and let go (it stays pinned); sweep across the ring to "
+ "take it off. It shows while you see its front within the angle above, and fades out beyond "
+ "it. Save current arrangement (Layout) keeps pins."
+ "it.\n\nPin a screen to your head: choose On your head above. It rides on the headset where it "
+ "is now, like a HUD, and shows whenever the screens do. Grab its bar to move it; it stays on "
+ "your head where you let go. Choosing a pin above keeps the screen where it is now, so place "
+ "it first. Save current arrangement (Layout) keeps pins."
}
}
}
}
// ---------------------------------------------------------------- Power
Component {
id: powerPage
Kirigami.ScrollablePage {
id: ppage
title: "Power"
property var p: backend.power
// The timeout choices, plus a value set by hand in frametop.conf.
property var offChoices: {
const list = [{ text: "Never", value: 0 }].concat([1, 2, 5, 10, 15, 30, 60].map(
m => ({ text: root.duration(m * 60), value: m })))
if (!list.some(c => c.value === p.offMinutes))
list.push({ text: root.duration(Math.round(p.offMinutes * 60)), value: p.offMinutes })
return list
}
Component.onCompleted: backend.refreshPower()
actions: [
Kirigami.Action {
text: "Turn displays off now"
icon.name: "system-suspend"
tooltip: "To try it: they come back on when the headset moves or any input is used"
enabled: ppage.p.service && ppage.p.state === "on"
onTriggered: backend.displaysOffNow()
}
]
header: Kirigami.InlineMessage {
position: Kirigami.InlineMessage.Position.Header
visible: !ppage.p.service
type: Kirigami.MessageType.Warning
text: "The power service (frametop-power) isn't running, so the displays won't turn off on their own. "
+ "It starts with SteamVR once it's installed: power/run.sh install, or run ./install.sh again."
}
ColumnLayout {
spacing: Kirigami.Units.largeSpacing
Kirigami.FormLayout {
Layout.fillWidth: true
Kirigami.Separator { Kirigami.FormData.isSection: true; Kirigami.FormData.label: "Displays" }
Controls.ComboBox {
Kirigami.FormData.label: "Turn off when unused for:"
model: ppage.offChoices
textRole: "text"
valueRole: "value"
Component.onCompleted: currentIndex = Math.max(0, indexOfValue(ppage.p.offMinutes))
onActivated: backend.setDisplayOffMinutes(currentValue)
}
Controls.Label {
text: "Unused means the headset and controllers haven't moved and no mouse, keyboard, or button "
+ "was used. This works even when the headset seems to be worn, like on a display mount "
+ "that covers its proximity sensor. Moving the headset or using any input turns the "
+ "displays back on. Taking the headset off still turns them off within seconds."
opacity: 0.7
font: Kirigami.Theme.smallFont
wrapMode: Text.Wrap
Layout.maximumWidth: Kirigami.Units.gridUnit * 26
}
Controls.Label {
Kirigami.FormData.label: "Now:"
visible: ppage.p.service
text: ppage.p.state === "off" ? "Off. Move the headset or use any input to turn them on."
: ppage.p.state === "away" ? "Off. SteamVR turned them off because the headset isn't being worn."
: ppage.p.offMinutes > 0
? "On, unused for " + (ppage.p.unused < 60 ? Math.floor(ppage.p.unused) + " s"
: Math.floor(ppage.p.unused / 60) + " min " + Math.floor(ppage.p.unused % 60) + " s")
: "On"
}
Kirigami.Separator { Kirigami.FormData.isSection: true; Kirigami.FormData.label: "Sleep" }
Controls.Switch {
id: awake
Kirigami.FormData.label: "While plugged in:"
text: "Stay awake"
checked: ppage.p.acSleep === 0
enabled: ppage.p.steam && !ppage.p.steamBusy
onToggled: {
backend.setStayAwake(checked)
checked = Qt.binding(() => ppage.p.acSleep === 0) // follow what Steam has
}
}
Controls.Label {
text: !ppage.p.steam
? (ppage.p.steamBusy ? "Checking Steam's setting…" : "Couldn't reach Steam: " + ppage.p.steamError)
: "Keeps the Frame awake and connected while it charges, for remote access, downloads, and "
+ "anything else running. This is Steam's own setting (Settings → Power → When Plugged In "
+ "and Idle), so the power button still puts the Frame to sleep. "
+ (ppage.p.acSleep > 0 ? "Now Steam puts it to sleep after " + root.duration(ppage.p.acSleep)
+ " without input, even while it charges. " : "")
+ "On battery, Steam's battery setting still applies ("
+ (ppage.p.batterySleep > 0 ? "sleep after " + root.duration(ppage.p.batterySleep) : "never sleep")
+ ")."
opacity: 0.7
font: Kirigami.Theme.smallFont
wrapMode: Text.Wrap
Layout.maximumWidth: Kirigami.Units.gridUnit * 26
}
}
Controls.Label {
Layout.fillWidth: true
wrapMode: Text.Wrap
opacity: 0.7
text: "With the displays off, the headset keeps tracking and drawing, so it can wake the moment "
+ "it moves. It still uses most of its power, so leave it on a charger that keeps up with it "
+ "in use."
}
}
}
+140
View File
@@ -0,0 +1,140 @@
"""Steam's sleep settings, read and written through Steam's own UI.
Steam, not systemd, puts the Frame to sleep: after "When Plugged In and Idle -> Sleep after"
(an hour by default) without input, even while it charges. That's a Steam client setting,
`system_idle_suspend_ac_sec` (0 = never), with no file or command line to change it. Steam
on the Frame runs with -cef-enable-debugging, so its UI's JavaScript context
(SharedJSContext) is reachable over the Chrome DevTools Protocol on 127.0.0.1:8080. There
`settingsStore.clientSettings` has the current values, and `SteamClient.Settings.SetSetting`
takes a change as a serialized CMsgClientSettings protobuf, which is what Steam's own
Settings -> Power page sends. Standard library only (a minimal WebSocket client).
"""
import base64
import json
import os
import socket
import struct
import urllib.request
CDP_PORT = 8080
# CMsgClientSettings field numbers (Steam's UI bundle maps the names to these).
FIELDS = {"system_idle_suspend_ac_sec": 24004, "system_idle_suspend_battery_sec": 24003}
class SteamUnreachable(Exception):
pass
class _WebSocket:
def __init__(self, url, timeout=5):
host_port, path = url[len("ws://"):].split("/", 1)
host, port = host_port.rsplit(":", 1)
self.sock = socket.create_connection((host, int(port)), timeout=timeout)
key = base64.b64encode(os.urandom(16)).decode()
self.sock.sendall((f"GET /{path} HTTP/1.1\r\nHost: {host_port}\r\nUpgrade: websocket\r\n"
f"Connection: Upgrade\r\nSec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n\r\n").encode())
head = b""
while b"\r\n\r\n" not in head:
chunk = self.sock.recv(4096)
if not chunk:
raise SteamUnreachable("Steam closed the connection")
head += chunk
if b" 101 " not in head.split(b"\r\n", 1)[0]:
raise SteamUnreachable(head.split(b"\r\n", 1)[0].decode(errors="replace"))
self.buf = head.split(b"\r\n\r\n", 1)[1]
def send(self, text):
data, mask = text.encode(), os.urandom(4)
n = len(data)
if n < 126:
head = struct.pack(">BB", 0x81, 0x80 | n)
elif n < 65536:
head = struct.pack(">BBH", 0x81, 0x80 | 126, n)
else:
head = struct.pack(">BBQ", 0x81, 0x80 | 127, n)
self.sock.sendall(head + mask + bytes(b ^ mask[i % 4] for i, b in enumerate(data)))
def _take(self, n):
while len(self.buf) < n:
chunk = self.sock.recv(65536)
if not chunk:
raise SteamUnreachable("Steam closed the connection")
self.buf += chunk
out, self.buf = self.buf[:n], self.buf[n:]
return out
def recv(self):
message = b""
while True:
b0, b1 = self._take(2)
n = b1 & 0x7F
if n == 126:
n = struct.unpack(">H", self._take(2))[0]
elif n == 127:
n = struct.unpack(">Q", self._take(8))[0]
message += self._take(n)
if b0 & 0x80:
return message.decode(errors="replace")
def close(self):
self.sock.close()
def _evaluate(expression):
"""Runs JavaScript in Steam's SharedJSContext and returns its (awaited) value."""
try:
with urllib.request.urlopen(f"http://127.0.0.1:{CDP_PORT}/json", timeout=3) as r:
targets = json.load(r)
except OSError as e:
raise SteamUnreachable(f"Steam isn't reachable on port {CDP_PORT} ({e})") from e
url = next((t["webSocketDebuggerUrl"] for t in targets if t.get("title") == "SharedJSContext"), None)
if not url:
raise SteamUnreachable("Steam's UI isn't running")
try:
ws = _WebSocket(url)
try:
ws.send(json.dumps({"id": 1, "method": "Runtime.evaluate",
"params": {"expression": expression, "awaitPromise": True, "returnByValue": True}}))
while True:
reply = json.loads(ws.recv())
if reply.get("id") == 1:
break
finally:
ws.close()
except OSError as e:
raise SteamUnreachable(str(e)) from e
result = reply.get("result", {})
if "exceptionDetails" in result:
details = result["exceptionDetails"]
raise SteamUnreachable(details.get("exception", {}).get("description") or details.get("text", "error"))
return result.get("result", {}).get("value")
def sleep_settings():
"""{"ac": seconds, "battery": seconds}: when Steam puts the Frame to sleep without input,
plugged in and on battery (0 = never)."""
value = _evaluate("(() => { const c = settingsStore.clientSettings; "
"return {ac: c.system_idle_suspend_ac_sec, battery: c.system_idle_suspend_battery_sec}; })()")
if not isinstance(value, dict) or not all(isinstance(value.get(k), int) for k in ("ac", "battery")):
raise SteamUnreachable("Steam's settings don't have the sleep timeouts")
return value
def set_sleep_setting(name, seconds):
"""Sets one of FIELDS to a whole number of seconds and checks that Steam took it."""
field, seconds = FIELDS[name], int(seconds)
if seconds < 0:
raise ValueError("seconds must be 0 (never) or more")
ok = _evaluate(f"""(async () => {{
const bytes = [];
const varint = n => {{ while (n > 127) {{ bytes.push((n & 127) | 128); n = Math.floor(n / 128); }} bytes.push(n); }};
varint({field} * 8); varint({seconds});
await SteamClient.Settings.SetSetting(btoa(String.fromCharCode(...bytes)));
for (let i = 0; i < 40; i++) {{
if (settingsStore.clientSettings.{name} === {seconds}) return true;
await new Promise(r => setTimeout(r, 50));
}}
return false;
}})()""")
if ok is not True:
raise SteamUnreachable(f"Steam didn't take {name} = {seconds}")
+30 -5
View File
@@ -42,12 +42,18 @@ Wherever ft-screens needs to know where a laser points (showing the controls, th
`ComputeOverlayIntersection` ignores `SetOverlayIntersectionMask`, and a control can't be allowed to cover part of its screen, so the resize tab sits entirely outside the corner.
### Wrist pinning
### Pinning
Pinning started as "bring the screen to your wrist", which doesn't work for big screens, because their centre is far from the edge you bring close. It became aiming: while a screen is carried, the line from the carrying device to its bar is tested against the other hand controllers. Crossing a controller's 6 cm ring arms the pin (leaving past 9 cm, so it doesn't flicker), and crossing it again disarms it. The pin happens on release, with the screen's pose at that moment, so you can arm it and then turn the screen. An earlier version pinned the moment the laser touched the wrist, which left the screen at whatever angle the carrying hand had while pointing there.
A pinned screen's alpha follows the angle between its front and the direction to your head, fully visible inside the wrist angle and fading over the last 10°.
A head pin is the same pin on the headset (device index 0): the screen's transform is relative to the headset, so SteamVR keeps it rigidly in your view with no lag from us. It skips the facing rule, since a screen on your head always faces you the way it did when pinned. There's no aiming gesture for it: the line from the carrying device can't sensibly pass through your own head, and a ring in front of your face would be in the way. So it's set from Frametop Display Settings or `ft-layout`, and it pins the screen where it is. Carrying a head-pinned screen re-pins it on release, like a wrist pin, so it can be adjusted in VR.
### Named layouts
A named layout is the custom arrangement under a name: each screen's pose relative to your head, width, curve, and pin, but not its resolution or scale, which need a desktop restart or belong to KWin. Using one copies it into the custom arrangement, so everything that applies the layout (desktop start, Meta+Shift+R, Arrange now) works unchanged, and `active` remembers which name it came from. Saving without a name (`ft-layout capture`) clears `active`, because the screens have been placed by hand since. Layouts are kept per screen number, so one saved with a different screen count still applies: missing screens keep their last saved place or the preset's.
### Visibility and VR games
`VROverlayFlags_MakeOverlaysInteractiveIfVisible` keeps SteamVR's laser mouse on while an overlay with that flag is visible. Without it, the laser is off whenever the dashboard is closed: the first click on a panel only turns it on, and the laser turns off again as soon as it leaves every panel. With it, controllers work the screens normally, but the laser also takes the controllers away from a VR game.
@@ -64,11 +70,13 @@ SteamVR's dashboard and every overlay it hosts are driven by the vrcompositor `l
Driver poses are in SteamVR's raw tracking space, and client programs work in the standing universe, which on the Frame is about 1.6 m above raw. Mixing them up put the laser's origin 1.6 m above your head. The helper converts using the headset's pose in both spaces every frame.
Frametop's SteamVR clients (ft-pointer, ft-screens, ft-gaze) connect as a background app first and switch to an overlay app only once that works. `VR_Init` as an overlay app starts vrserver itself when none is running, and one started that way from the dev container never finds the headset. At a boot where the gamescope session timed out, systemd dropped `steamvr.service`'s start job, the pointer service (ordered only `After=` it) started anyway, and its vrserver made every SteamVR launch fail with `HmdNotFound`. SteamOS's health check then kept resetting the Steam client and tried to fall back to the previous OS slot. The units also say `Requisite=steamvr.service`, so they don't start at all when SteamVR's start fails.
The driver starts disconnected, because holding the right-hand role while SteamVR starts leaves the Steam UI stuck on its loading icon. It connects when the mouse is used and claims the right hand. SteamVR keeps a hand role reserved for a disconnected device that still asks for it, so the driver switches its role hint between right hand (connected) and opt-out (not connected).
### The cursor
Mouse motion turns into yaw and pitch around an anchor, the head position at the last recenter. A ray from the anchor is tested against every visible overlay with `ComputeOverlayIntersection`. On a hit, the cursor sits on that surface; otherwise it floats at `POINTER_DISTANCE`. Since the anchor isn't your current eye position, a second test runs along your line of sight to the cursor point, and anything nearer wins, so the cursor always lands on what you see under it.
Mouse motion turns into yaw and pitch around an anchor, the head position at the last recenter. A ray from the anchor is tested against every visible overlay with `ComputeOverlayIntersection`. On a hit, the cursor sits on that surface; otherwise it floats at `POINTER_DISTANCE`. Since the anchor isn't your current eye position, a second test runs along your line of sight to the cursor point, and anything nearer wins, so the cursor always lands on what you see under it. Overlays in `POINTER_IGNORE` are left out of both tests. A display-only panel, like a performance overlay locked to your view, has no input method, so SteamVR's laser passes through it, but `ComputeOverlayIntersection` still hits it, and the cursor stuck to it. The laser starts just before the cursor point, so an ignored panel nearer to you doesn't catch it either.
OpenVR has no call to list other programs' overlays, so the helper runs `vrcmd --overlays` in the background. It includes hidden overlays, because a floating window's controls only appear while something hovers the window, and the cursor has to find them immediately.
@@ -83,7 +91,7 @@ A few overlays need special handling:
Head follow is experimental and off by default. It works, but it's only lightly tested, and the feel is mostly a matter of its settings; polishing it is left open. With it on (`POINTER_FOLLOW=1`, or a mouse button mapped to Head follow on/off), the cursor rides on a reference direction, where you were facing when your head last settled, and keeps its offset from it. The mouse can put the cursor anywhere up to `POINTER_FOLLOW_REACH` (70 degrees) from the reference, a corner of your view included. While your head stays within `POINTER_LEASH_DEG` of the reference, nothing moves on its own. Once your head has been past the leash for `POINTER_LEASH_DELAY` (0.2 s, so a glance out and back doesn't count), the reference eases to where you're facing (time constant `POINTER_LEASH_RETURN`, 0.2 s), never falling further behind than the leash, and the cursor ends up back where it was in your view. Then it waits for the leash again. Two earlier versions didn't work out. Moving the reference only while your head pulled at the end of the leash left it up to the leash off after you turned back, and getting it centred again meant overshooting with your head. Easing it toward your facing all the time moved the cursor on every small head movement. A leash of 0 makes the reference your facing direction, so the cursor is locked to your view, and mouse movement shifts it within the view. Head roll is ignored, so tilting your head doesn't swing the cursor around. While the left button is held the cursor stays put in the room, so your head can't nudge a click or a drag. When you let go, it carries on from where it is instead of jumping.
Gaze mode is experimental and off by default (`POINTER_GAZE=1`, the Gaze page of Frametop Input Settings, `gaze/ft-gazectl on`, or a mouse or controller button mapped to Gaze pointer on/off). It's MAGIC pointing (Zhai, Morimoto and Ihde, 1999): the pointer goes where you look, and the mouse does the last bit. The gaze service (`gaze/ft-gazed`) sends the helper the corrected gaze at 90 Hz, and while the gaze has the pointer, the cursor ray is that gaze from the eye. The pointer is aimed at the gaze each frame, not steered toward it, so nothing can pile up. An earlier try in the gaze probe steered the pointer with relative moves, and lost it when the pointer went idle or a controller had the laser. Moving the mouse takes the pointer from the gaze. A left press while the gaze has the pointer isn't sent at once: the pointer stops where the gaze put it, you drag it onto what you meant with the button still down (panels only see it hover), and the release clicks there. Clicking at once clicked wherever the gaze was, often the wrong thing, before you could correct it. A press held still for `POINTER_GAZE_HOLD` (0.5 s) becomes a real press, so drags still work: hold, then move. Outside games the pointer then stays: the mouse going idle doesn't release it. A moving controller still releases it, as without gaze; the mouse is gaze mode's only pointer device for now. The dot shows only while the mouse moves it (`POINTER_GAZE_SHOW`), while a press is held, and as a pulse for each click; otherwise it's transparent, so the laser still lands on it. The gaze moving the pointer doesn't show it: you know where you're looking. Looking more than `POINTER_GAZE_RETAKE` (5 degrees) away from it, with the mouse still, gives it back, so small eye movements around the pointer don't pull it off what you're doing. A mouse nudge of up to `POINTER_GAZE_NUDGE_MAX` (8 degrees) before a click is sent to the gaze service as a lesson: you were looking at where you clicked when the mouse took over, so the nudge is the eye tracker's error there. Using it is what calibrates it. See `gaze/README.md` for the service, the calibration, and what was measured.
Gaze mode is experimental and off by default (`POINTER_GAZE=1`, the Gaze page of Frametop Input Settings, `gaze/ft-gazectl on`, or a mouse or controller button mapped to Gaze pointer on/off). It's MAGIC pointing (Zhai, Morimoto and Ihde, 1999): the pointer goes where you look, and the mouse does the last bit. The gaze service (`gaze/ft-gazed`) sends the helper the corrected gaze at 90 Hz (from one eye while the tracker has lost the other), and while the gaze has the pointer, the cursor ray is that gaze from the eye. The pointer is aimed at the gaze each frame, not steered toward it, so nothing can pile up. An earlier try in the gaze probe steered the pointer with relative moves, and lost it when the pointer went idle or a controller had the laser. Moving the mouse takes the pointer from the gaze. A left press while the gaze has the pointer isn't sent at once: the pointer stops where the gaze put it, you drag it onto what you meant with the button still down (panels only see it hover), and the release clicks there. Clicking at once clicked wherever the gaze was, often the wrong thing, before you could correct it. The drag is the correction. Snapping the pointer onto buttons and links is deferred: it needs accessibility (AT-SPI) on in the Frametop session, where it's off (no registry runs), plus app restarts, and it makes Chromium and Electron apps use more CPU. A press held still for `POINTER_GAZE_HOLD` (0.5 s) becomes a real press, so drags still work: hold, then move. Outside games the pointer then stays: the mouse going idle doesn't release it. A moving controller still releases it, as without gaze; the mouse is gaze mode's only pointer device for now. The dot shows only while the mouse moves it (`POINTER_GAZE_SHOW`), while a press is held, and as a pulse for each click; otherwise it's transparent, so the laser still lands on it. The gaze moving the pointer doesn't show it: you know where you're looking. Looking more than `POINTER_GAZE_RETAKE` (5 degrees) away from it, with the mouse still, gives it back, so small eye movements around the pointer don't pull it off what you're doing. A mouse nudge of up to `POINTER_GAZE_NUDGE_MAX` (8 degrees) before a click is sent to the gaze service as a lesson: you were looking at where you clicked when the mouse took over, so the nudge is the eye tracker's error there. Using it is what calibrates it. See `gaze/README.md` for the service, the calibration, and what was measured.
Replacing a loaded driver's files, as re-running the installer used to do, leaves SteamVR honoring the virtual controller's hand role but not its laser claim: the dashboard pointer stays unassigned until SteamVR restarts. The driver installer now leaves an unchanged driver in place.
@@ -91,7 +99,7 @@ Replacing a loaded driver's files, as re-running the installer used to do, leave
### Handing the laser back and forth
The dashboard follows whichever device summoned it or last pressed its trigger. Frametop adds "last used wins": moving a real controller releases the pointer at once, and the next mouse movement takes the laser back. Small movements don't count; waking needs `POINTER_WAKE_COUNTS` of mouse motion within a second, so desk jitter doesn't steal the laser. While the pointer is awake, a tiny transparent overlay with `MakeOverlaysInteractiveIfVisible` keeps SteamVR's laser mouse on, since otherwise the first click would only switch the laser on.
The dashboard follows whichever device summoned it or last pressed its trigger. Frametop adds "last used wins": moving a real controller releases the pointer, and the next mouse movement takes the laser back. Moving means faster than 0.35 m/s or 2 rad/s (both times `POINTER_CONTROLLER_PICKUP`, 1 by default) for 100 ms in a row, while the controller is tracked normally. A single sample over the limit used to be enough, and controllers resting on a desk took the laser back on a knock or a tracking jump while the mouse was in use. Small movements don't count; waking needs `POINTER_WAKE_COUNTS` of mouse motion within a second, so desk jitter doesn't steal the laser. While the pointer is awake, a tiny transparent overlay with `MakeOverlaysInteractiveIfVisible` keeps SteamVR's laser mouse on, since otherwise the first click would only switch the laser on.
When the headset comes off, SteamVR reports its activity level as idle at once and turns the displays off 5 seconds later (`power.turnOffScreensTimeout`), unless something keeps it awake. An awake pointer did, and so did the helper's `vrcmd` runs: each is a new SteamVR client, and a new client every second kept SteamVR out of standby. The helper now releases the pointer as soon as the headset is idle, ignores the mouse until you're wearing it again, and pauses the overlay list whenever the pointer is off.
@@ -109,6 +117,10 @@ An ungrabbed keyboard reaches both sides at once. In VR, gamescope reads every i
Volume keys must never reach gamescope. With the openvr backend, gamescope sends volume up and down to Steam by moving keyboard focus to Steam for the key and then back to the previously focused surface. When nothing had focus, the one it moves back to is null, and wlroots aborts on a null focus surface (`wlr_seat_keyboard_notify_enter: Assertion 'surface' failed`), which ends the whole VR session. Keyboard focus is often empty while you work in VR, so one press of the headset's volume button could take everything down. gamescope reads the headset's buttons and every keyboard itself (`InputStealer`), as do SteamVR's processes, so the relay has to stop volume keys at the device. Grabbing `gpio-keys` would also take the headset's click button, so the relay remaps the volume entries in each device's keymap (`EVIOCSKEYCODE`) and handles the stand-in codes itself. That fix covers every device at once, including keyboards that aren't grabbed.
Frametop's keyboard opens by itself for a text field on the desktop. The apps run inside the nested KWin, so only KWin knows when a text field has focus, and the way it tells anyone is its input method protocol (`zwp_input_method_v1`): KWin starts one input method program and activates it whenever the focused app turns on text input. `input/ft-textinput` is that program, speaking the Wayland wire protocol directly so it needs nothing but Python on the host. It only reports focus. The gamescope session puts `QT_IM_MODULE=xim` and `GTK_IM_MODULE=xim` in the systemd user environment; with those, Qt and GTK apps use X input methods and never turn on Wayland text input, so the session script drops them.
The keyboard itself is ft-screens' own panel (`screens/keyboard.cpp`). We tried SteamVR's first (`ShowKeyboardForOverlay`), and on the Frame it doesn't fit a desktop. It's Steam's own panel (`valve.steam.gamepadui.keyboard`), which SteamVR mounts in the dashboard's scene, so with the dashboard closed it opened but wasn't drawn. Placing it in the room ourselves (`SetKeyboardTransformAbsolute`) made it show, but SteamVR moves it to whichever overlay the laser goes to and mounts it again, and while it's open the controllers switch to SteamVR's own laser. Our panel is an overlay like the screens' controls: any laser or the 3D mouse clicks it, nothing moves it, and its keys go out as key presses on ft-screens' seat rather than as text handed back to the input method. So nothing typed leaves ft-screens (a socket to the input method could be claimed by any local process, like `@frametop_keys`), apps without text input (X11, Electron) take the keys too, and they mean what the desktop's keyboard layout says. It's drawn on the CPU and uploaded with `SetOverlayRaw` when a key's look changes; the labels come from stb_truetype, so the container needs no text rendering stack.
The Frame controllers can be mapped like mouse buttons, but they aren't input devices on the host: they reach SteamVR over the headset's own radio, and no evdev or hidraw node exists for them. So only a SteamVR client can read them. Overlay apps normally get controller input only while they have input focus, which a background helper never has. SteamVR's experimental global action set priority (`steamvr/globalActionSetPriority`, "Enable global input from overlays") lets an overlay's action set receive input anyway, and takes the inputs it binds from the scene app. Binding every button would take them all from games, so the helper's action manifest puts each button in an action set of its own, and it activates only the sets of mapped buttons. The mapping itself stays in the relay, which does the action, so mice and controllers share one list of actions.
## The desktop session
@@ -117,14 +129,26 @@ The session is modeled on SteamOS's `steamos-nested-desktop` and runs beside it.
The VR launcher starts the session from the Steam client, and the client's environment came along: `LD_LIBRARY_PATH` pointing at Steam's own runtime, whose `libavcodec` has no H.264 decoder, so VLC in the desktop couldn't play most videos, plus the client's overlay and launch settings. The session script drops the client's variables before it starts anything. SteamOS's global Mesa settings (`/usr/share/deckard/mesavars.sh`) stay, and the gamescope session's Vulkan layer (`ENABLE_GAMESCOPE_WSI`) is only kept for the gamescope backend.
Steam, not systemd, suspends the Frame: after `system_idle_suspend_ac_sec` (an hour by default) without input on AC power, it logs `Switching to power state: k_ESystemPowerState_Sleep` and suspends, even while charging. It's a Steam setting (Settings → Power → When Plugged In and Idle → Sleep after), so the README recommends setting it to Never. SteamVR's standby, which turns the displays off when the headset comes off, is separate.
Steam, not systemd, suspends the Frame: after `system_idle_suspend_ac_sec` (an hour by default) without input on AC power, it logs `Switching to power state: k_ESystemPowerState_Sleep` and suspends, even while charging. It's a Steam setting (Settings → Power → When Plugged In and Idle → Sleep after), which the Stay awake while plugged in switch in Frametop Display Settings sets to Never. SteamVR's standby, which turns the displays off when the headset comes off, is separate; see below.
Flatpak apps need `XDG_DATA_DIRS` to include Flatpak's exports, or Plasma opens Discover instead of launching them, so the session sources `/etc/profile.d/flatpak.sh`.
The private runtime directory also moves the session's document portal to `$XDG_RUNTIME_DIR/frametop/doc`, and that broke saving and uploading in Flatpak apps. The file picker (xdg-desktop-portal 1.18.4 on SteamOS) gives a sandboxed app the host path of the file it picked, `/run/user/1000/frametop/doc/ID/NAME`. Inside the sandbox the portal is at `/run/flatpak/doc`, and `/run/user/1000` is a private per-app folder (`.flatpak/APP/xdg-run` in the runtime directory). So Brave created the missing folder there, "finished" the download into it, and the file vanished when the session cleaned up. The session script now links that path to `/run/flatpak/doc` in each installed app's folder before Plasma starts. Upstream xdg-desktop-portal fixed this after 1.22.1 (commit `69ba5e1`) by handing Flatpak apps `/run/flatpak/doc` paths, after which the links go unused.
A podman container's monitor process (conmon) stays in the cgroup of whatever started the container, and `distrobox enter` starts it on demand. When a Frametop service happened to start the `dev` container, stopping that service stopped the container and everything in it, including the desktop's compositor. `scripts/container-up.sh` starts the container in a systemd scope of its own before anything enters it.
Program names stay within 15 characters, because Linux truncates process names there and the scripts find programs with `pgrep -x` and `pkill -x`. That's why the prefix is `ft-`.
## Displays off on a stand
SteamVR decides the headset is off from its proximity sensor, which the driver reads through the DSP, and turns the displays off 5 seconds later. On a display mount that covered the sensor, that never happened: SteamVR kept the headset in use all night (no `entering standby` for device 0 in vrserver.txt, and XRService's user presence stayed at 1), and Steam didn't sleep either, because its idle count treats a present user as active. The battery went from 100% to 12% overnight on a 5 V, 3 A charger, with the headset drawing about 17 W.
There's no client call that puts the headset in standby. The cv driver's `teststandby` debug request (`IVRDebug::DriverDebugRequest`) only answers "Standby unknown hmd" on the Frame. But what the driver does for the displays in standby is write `/sys/class/backlight/ae94000.dsi.0/brightness` ("cv: Set displays off" writes 0, "Set displays on" the old value), and the `video` group can write that file, from the container too. So `ft-powerd` goes by use instead of the sensor and turns the backlight off itself. Tracking and rendering keep running. Turning the backlight off moved the battery current by only about 75 mA (0.5 W), so they're most of the load, but they're also why the displays can wake the moment the headset moves.
Movement is judged within 10-second windows. On the mount, the head pose jittered within 0.5 mm and 0.1 degrees over 20 seconds, and its position drifted 1.7 mm (0.16 degrees) in 4 minutes. Compared with a fixed reference, that drift would count as movement sooner or later and keep the displays on; within 10 seconds it never reaches the 5 mm and 0.5 degree thresholds, and anyone wearing the headset passes them now and then.
Staying awake while charging uses Steam's own setting rather than a logind sleep inhibitor. Steam suspends with `dbus-send ... login1.Manager.Suspend boolean:true`, and a block inhibitor does stop that (`CanSuspend` answers "challenge" while one is held), but it stops the power button too. `system_idle_suspend_ac_sec` is field 24004 of Steam's CMsgClientSettings. In Steam's SharedJSContext, reachable over CDP on port 8080 because Steam runs with `-cef-enable-debugging`, `SteamClient.Settings.SetSetting` takes a change as a base64 protobuf, the way Steam's Power page sends it (0 is never), and `settingsStore.clientSettings` has the current values.
## Approaches we dropped
- WayVR, an existing Wayland desktop for VR. It built and connected to SteamVR on the Frame, but nothing showed in the headset. It has no bindings for the Frame's controllers, and its KDE screen capture needs `xdg-desktop-portal-kde`, which SteamOS doesn't ship.
@@ -138,3 +162,4 @@ Program names stay within 15 characters, because Linux truncates process names t
- Drawing KWin's cursor on the screens.
- Plasma can lose its panels when the number of screens goes down, because they're saved against a screen that no longer exists. Removing `plasma-org.kde.plasma.desktop-appletsrc` and `plasmashellrc` from `~/.config/frametop` brings the default panels back.
- Frame pacing and GPU cost with several busy screens haven't been measured.
- Real standby on a stand, with rendering and tracking paused, not just the backlight off. SteamVR has no call for it, and its activity level follows the proximity sensor.
+209
View File
@@ -0,0 +1,209 @@
# Floating windows (plan)
Status: design settled 2026-09-29 (see "Decisions"); being built on the `floating-windows` branch.
Built so far (2026-09-29; the KWin side tested on the headless test desktop, `screens/test/headless.sh`; nothing yet tried in the headset):
- The catcher (a release off every panel still reaches KWin), and the pointer helper's "up" backstop.
- `float/frametop-float.js` (the KWin script), `float/ft-floatd`, and `float/ft-float`. "Float in VR" is in the window menu under Extensions, and Meta+Shift+F toggles the active window. Floating, docking (back where it came from), closing, full screen, per-window scale (Meta+scroll), popups reported with their rectangles, windows of a floating app floating too, and the notification when every spare is in use.
- ft-screens: a panel per spare output (`frametop.float.N`) with the crop, density, popups and dialogs as small panels over it, title-bar carrying, the corner tab resizing the window, and dock and close buttons. `--spares`, and the commands `float`, `unfloat`, `pose`, `sub`, `minimized`, `carry`.
- The session adds `FLOAT_SLOTS` spares and starts ft-floatd from the desktop's autostart; ft-layout leaves the spares alone.
- The 3D mouse's drag lock crosses onto other Frametop panels (not while carrying one).
Not built yet: phase 2 (the ghost, tear-off by dragging, push-flush docking), phase 3 (launching floating, the Frametop Apps entry and picker, remembered placement), and phase 4. Frametop Apps (decision 10) needs a per-screen hide, which ft-screens doesn't have yet: its hide and show are for all screens.
The goal is to let any desktop app float in VR in a panel of its own, like SteamVR's floating windows, while it stays part of the Frametop desktop. That means drag and drop, the clipboard, and focus keep working between floating windows and the screens.
- There are two ways to get a floating window. Launch the app floating, or drag a desktop window by its title bar off a screen and let go in the air.
- There are two ways to put one back. Push it flush against a screen and let go, or press its "back to desktop" button.
- Files, text, and images drag between any two floating windows, and between floating windows and the screens.
## Decisions
Settled with the user on 2026-09-29. The sections below follow them.
| # | Question | Decision |
|---|---|---|
| 1 | How many windows can float at once | 8 spare outputs by default, configurable (`FLOAT_SLOTS`, and Display Settings); a change needs a desktop restart |
| 2 | Menus and dropdowns | Each floating output has a margin around the window. The panel shows only the window, and each open popup gets a small overlay of its own, cut from the same buffer |
| 3 | Tearing off | Drag the title bar past a screen's edge and let go in the air, with a small dead zone past the edge |
| 4 | Docking by dragging | Push the window flush against a screen (within about 10 cm), with the landing spot highlighted, and let go |
| 5 | Visibility | Floating windows follow the same rules as the screens: the hide hotkey, the visibility modes, and the games rule |
| 6 | Windows a floating app opens | They float too |
| 7 | Launching floating from the headset | One "Frametop Apps" launcher entry with a picker |
| 8 | Build order | The catcher first, as a fix that stands on its own; `pointer-ignore` and `layouts-headpin` merged before phase 1; the hand cutouts stay out |
| 9 | Show Desktop (Meta+D) | Floating windows stay |
| 10 | Frametop Apps and visibility | The entry starts the desktop with each screen hidden on its own (the existing per-screen hide), so only floating windows show. No special mode |
| 11 | Window frame | KWin's title bar and border stay. Frametop's bar, close, and "back to desktop" are extras |
| 12 | Resizing | The window's own edges and Frametop's corner tab both change the size in pixels at the same density; the output follows |
| 13 | Margin | 300 px on each side, configurable |
| 14 | All spares in use | The window opens on the screens, with a notification |
| 15 | Where a floating app's new windows go | Where that app's windows went last time; otherwise to the parent's right, curving around you |
| 16 | Where the code is written | A branch in the PC's clone of the repo |
| 17 | Size when docked by dragging | The current floating size in pixels, shrunk to fit the screen |
| 18 | Bigger text | A scale for each window (KWin's output scale): Meta+scroll over the window, or +/- on its bar. Remembered for each app |
| 19 | Switching to a window you can't see | It's focused, and a glow at the edge of your view points to it. Moving it in front of you is a setting |
| 20 | Full screen | The window fills its own panel. The margin drops to zero while it's full screen, and the panel keeps its size and place |
| 21 | Named layouts | They cover the screens only. Floating windows use the placement remembered for each app |
Also assumed: floating windows get the wrist pin, the head pin, and pass-through (`pointer-ignore`) like screens. Every gesture works with the controllers as well as the 3D mouse. A window launched floating uses the primary screen's density. VNC shows only the primary screen, as now. Anything that restarts the live desktop waits for the user's OK.
## The approach: each floating window gets a KWin output of its own
Drag and drop and the clipboard only work between windows of the same compositor. A Wayland window can't move from one compositor to another. So a floating window has to stay a KWin window.
ft-screens already shows each KWin output as a panel. It sets the output's size with an `xdg_toplevel` configure, and KWin resizes the output to match. So a floating window can get an output of its own, sized to fit it, and ft-screens shows that output as a panel with its own controls. To KWin this is an ordinary desktop with more monitors. Dragging between two floating windows is the same as dragging between two monitors, which KWin already handles. ft-screens already moves the pointer between panels in the middle of a drag: `handle_vr_event` moves pointer focus to another KWin window even while a button is held.
Alternatives we considered:
- **Run floating apps directly on ft-screens.** It's a wlroots compositor, so apps could connect to it and get a panel per window. But they would get no drag and drop or clipboard with desktop apps unless we wrote a bridge. Also, a window that's already on the desktop could never be torn off, because a Wayland client can't change compositors. Rejected.
- **One large hidden "canvas" output.** Every floating window would sit on one big output, and each panel would show a crop of it (`SetOverlayTextureBounds`). That needs only one extra output, with no copies. But an 8K canvas uses about 128 MB per buffer, with two or three buffers in KWin's swapchain. It would also have to repack windows whenever one resized, full screen would fill the whole canvas, and every window would share one scale. This is the fallback if per-window outputs don't work.
- **Screencast single windows** (`zkde_screencast` `stream_window`, over PipeWire). This adds copies and latency, and the window still needs a real place in KWin's layout to receive input. Rejected.
- **SteamOS's own floating windows** (Launch a program from the dashboard). Those apps run in gamescope, outside KWin, so they can't drag and drop with the desktop.
### Where the extra outputs come from: spare outputs
KWin's nested backend opens its outputs at start (`--output-count`). The session starts KWin with the screen count plus `FLOAT_SLOTS` outputs (default 8). Each spare is disabled until it's needed, with `kscreen-doctor` (or in the session's `kwinoutputconfig.json`, so it starts disabled). Floating a window enables a spare, and docking the window disables it again. `FLOAT_SLOTS` limits how many windows can float at once, and changing it means restarting the desktop.
Checked in KWin 6.2.5's source (`src/backends/wayland/`, 2026-09-29):
- Disabling a nested output keeps its host window. `Output::applyChanges` only flips `enabled`, KWin stops rendering it, and Plasma drops its desktop view. So ft-screens keeps the same toplevel, and its screen numbers stay put.
- Each output's host window is titled `KDE Wayland Compositor WL-<n>`, with `- Output disabled` appended while it's disabled (`WaylandOutput::updateWindowTitle`, on every `enabledChanged`). ft-screens reads the title to tell screens (`WL-0` to `WL-<SCREENS-1>`) from spares, and to see a spare turn on and off.
- Pointer positions reach KWin only through motion events: the output's position in the layout plus the position on its window. When ft-screens stops sending motion, KWin's pointer stays put.
- **Virtual outputs don't work.** `createVirtualOutput` makes an output window but never adds it to the backend's `m_outputs`, so `findOutput()` returns null when the pointer enters it, and the next line dereferences it (`Q_ASSERT` is compiled out). A click on such a panel would crash KWin. This rules out the virtual-output fallback (`stream_virtual_output`) without a patched KWin.
ft-screens creates a `screen` for each toplevel in the order they appear, and indexes its settings by that order. Spares come after the screens, so they get indices `SCREENS` and up. Their panels are hidden while their output is disabled.
## How the parts fit together
```
KWin script "frametop-float" ft-floatd (host, Python) ft-screens
window events, moves, menus ── D-Bus ──▶ window ↔ output ↔ panel table ── @ft_screens ──▶ panels, controls,
runs commands ◀─ long poll ─ spare outputs (kscreen-doctor) ◀─ @frametop_float ─ lasers, ghost, catcher
```
- **KWin script `frametop-float`** (JavaScript). A script keeps working across KWin updates. A C++ effect would have to match the host's exact KWin build, and our build container is Fedora, not SteamOS. The script watches windows (`windowAdded`/`windowRemoved`, `interactiveMoveResizeStarted`/`Stepped`/`Finished`, `outputChanged`, `minimizedChanged`, `windowActivated`, `fullScreenChanged`). It runs commands: move a window to an output, set its geometry, put it on all virtual desktops, and restore it. It adds "Float in VR" to the window menu (`registerUserActionsMenu`) and registers a shortcut (`registerShortcut`, Meta+Shift+F). KWin scripts can call D-Bus but can't serve it, so commands come back through a long poll. The script calls ft-floatd's `NextCommand`, which answers when a command is ready, and then the script calls it again. The fallback is loading one-shot scripts through `org.kde.kwin.Scripting`, the way kdotool does. All of these API names are present in the host's KWin 6.2.5.
- **ft-floatd** (Python). The host has dbus-python and PyGObject. It owns `org.frametop.Float` on the session's private bus, and it keeps the table of which window is on which output and panel. It enables and disables spare outputs and sets their size, scale, and position with `kscreen-doctor`, as ft-layout does. It tells ft-screens where each floating window goes and tells the script which window goes where. It also remembers each app's placement and scale, keyed by desktop file name.
- **ft-screens.** `Screen` becomes a panel with a kind: screen or floating window. Floating panels get the same bar, curve, roll, resize tab, and wrist and head pins, plus close, "back to desktop", and scale buttons. New parts are the tear-off ghost, the catcher, popup overlays, carrying a panel during a KWin move, and the dock target highlight. `MAX_SCREENS` goes from 8 to 16. Commands arrive on `@ft_screens`. Events go out to `@frametop_float` from an unbound socket, the same way ft-screens talks to the input relay.
- **Session script.** Adds `FLOAT_SLOTS` to the output count, starts ft-floatd, and enables the KWin script in the session's `kwinrc`.
- **ft-layout.** Arranges only the screens' outputs. Today it arranges everything in `kscreen-doctor -j`, so it has to skip the spares (`WL-<SCREENS>` and up), enabled or not.
- **ft-pointer.** Changes to the drag lock (see "Drag and drop between panels").
- **Frametop Display Settings.** Gets a Floating windows section: slots, margin, and "bring a window in front of you when it's activated".
Program names stay within 15 characters (`ft-floatd`). Overlay keys are `frametop.float.N` and `frametop.float.N.bar`, and so on.
## A floating window
- **Output and margin.** Its output is the window's frame plus a margin on each side (`FLOAT_MARGIN`, default 300 px). KWin keeps a Wayland popup inside its parent's output (`XdgPopupWindow::updateRelativePlacement` uses the output's placement area), so the margin gives menus and dropdowns room past the window's edges. X11 apps place their own menus within the monitor, so the same applies. Enabled spares sit apart from the screens and from each other in KWin's layout, so nothing spills from one to the next. Memory: a 1600 × 1000 window with a 300 px margin is about 14 MB per buffer, 42 MB for three.
- **What the panel shows.** Only the window's frame: ft-screens crops the output's buffer with `SetOverlayTextureBounds` and maps mouse positions through the crop. Each open popup gets a small overlay of its own, cut from the same buffer and placed a few millimetres in front of the window, so the main panel never changes size. KWin tells scripts about popups as windows of their own (`windowAdded` with `popupWindow`), so the script reports their rectangles.
- **Size and scale.** The panel's width is the window's pixel width times the source screen's metres per pixel, so text stays the same size in VR. A window launched floating uses the primary screen's density. Each window also has a scale (KWin's output scale), changed with Meta+scroll over the window or +/- on its bar and remembered for each app. A bigger scale makes the content bigger at the same panel size.
- **Window state.** An ordinary window, not maximized, placed inside its output with the margin around it, and set to show on all virtual desktops. It keeps its title bar and border. Apps that draw their own title bar (GTK, Chromium) keep theirs.
- **Moving.** Press the title bar. KWin starts an interactive move and the script reports it. ft-screens then stops forwarding pointer motion to KWin, so KWin's pointer stays at the press point and the window moves by nothing. Meanwhile ft-screens carries the panel with the pressing device, the same way the bar does today: it follows rigidly, scroll pushes and pulls, and the 3D mouse's right-drag tilts. When the button comes up, KWin gets the release at the press point. The bar under the panel works too.
- **Resizing.** The window's own edges (inside the margin, so KWin's resize works as on the desktop) and Frametop's corner tab both change the window's size in pixels at the same density, so the app lays itself out again. ft-floatd resizes the output to keep the margin, and the panel grows or shrinks around the window's top-left corner. A screen's tab only scales the panel. Resizing is throttled to about 20 updates a second, with a minimum of 320 × 200, like screens.
- **Full screen.** The window fills its own panel: the margin drops to zero while it's full screen, and the output is the panel's size in pixels. The panel keeps its size and place. On leaving full screen, the margin comes back.
- **Buttons.** The close button closes the window. "Back to desktop" docks it where it came from.
- **Minimize.** Minimizing, from the title bar or the taskbar, hides the panel, and restoring it shows the panel again. Floating windows stay in the desktop's taskbar and in Alt+Tab.
- **Activated out of view.** When a floating window is activated (taskbar, Alt+Tab, a notification) and it's more than about 60° from where you're looking, it's focused and a glow at the edge of your view points to it. A setting moves it in front of you instead.
- **New windows.** A dialog of a floating window (`transientFor`) floats in front of its parent. Other new windows of a floating app float too: where that app's windows went last time, otherwise to the parent's right at the same distance, curving around you, and to its left if that's taken. When every spare is in use, the window opens on the screen used last and a notification says so.
- **Show Desktop.** Meta+D leaves floating windows alone.
## Tearing a window off a screen
1. Press a desktop window's title bar and drag it. KWin starts a move, and the script tells ft-floatd, which tells ft-screens: `move-start <output> <window> <rect>`.
2. While the button is held, the laser leaves every Frametop panel by more than a small dead zone (a few centimetres past the edge). Letting go inside the dead zone is an ordinary drop.
3. ft-screens shows a ghost: an overlay showing the screen's live buffer cropped to the window (`SetOverlayTextureBounds`, no copy). It's at the screen's pixel density and distance, on the laser, facing you, with the point you grabbed under the laser. The ghost takes mouse input, so SteamVR's laser lands on it and the release comes to ft-screens.
4. Go back onto a screen before letting go, and the ghost disappears. It's an ordinary move again.
5. Let go on the ghost, and ft-screens releases the button in KWin, which ends the move. It then reports the tear-off to ft-floatd, with the window, the ghost's pose, and the density. ft-floatd enables a spare output and has ft-screens size it to the window plus the margin and put its panel at the ghost's pose. Then it has the script move the window onto that output. The ghost stays until the new panel's first frame at the right size arrives, so nothing blinks.
## Putting it back
- **Button.** "Back to desktop" returns the window to the screen, position, and size it had before it was torn off (saved at tear-off).
- **Dragging.** Carry the floating window, by its title bar or its bar, until the spot you're pointing at is on a screen. Then push it flush with the screen, within about 10 cm of its surface: scroll away with the mouse, or move the controller forward. The screen shows where the window will land, and letting go docks it there at its current size in pixels, shrunk to fit if the screen is smaller. A carried panel keeps its distance, so moving a floating window in front of a screen never docks it by accident.
- Docking disables the output and removes the panel.
## Launching an app floating
- **In the desktop.** Use "Float in VR" in any window's menu, or press Meta+Shift+F for the active window.
- **From a command.** `ft-float run <command>` and `ft-float launch <app.desktop>` start an app and float its first window. ft-floatd records the process it started, and the script matches new windows by PID, including child processes. Some single-instance apps (Firefox, D-Bus-activated apps) open the window from a process that was already running. Those are matched by desktop file name within a few seconds, or by `XDG_ACTIVATION_TOKEN` where the app honors it.
- **From the headset without the desktop open.** One new launcher entry, "Frametop Apps". It starts the Frametop session with each screen hidden on its own (the per-screen hide that already exists), and opens an app picker as a floating window. Picking an app launches it floating. The hide hotkey and visibility modes still apply to everything, and a screen comes back with one click. Everything runs in one session, so dragging between a standalone app and a desktop app works. If the desktop is already running, the entry just opens the picker.
- **The picker.** Either KRunner, floated, or a small Kirigami app like the settings apps, with a grid of apps and their icons.
- **Remembered placement.** Each app's last floating pose, size, and scale, keyed by desktop file name. Named layouts don't include floating windows.
## Drag and drop between panels
KWin handles the protocols: Wayland, X11 through Xwayland, and the portal's file transfer. Frametop has to get the pointer right between panels.
- **Crossing panels.** When the laser moves onto another panel mid-drag, ft-screens gives that panel's KWin window pointer focus. KWin puts its cursor at that output's position, and the drop target gets enter and motion events. Floating windows add nothing new here, but they make gaps between panels the normal case.
- **Gaps (the catcher).** While the laser is between panels, none of our overlays get its events, and ft-screens clears pointer focus on `FT_LEAVE` even with a button held. If you let go in empty space, the release never reaches KWin, and the drag or move stays stuck until the next click. The fix: while a button is held on a Frametop panel and the laser leaves all of them, ft-screens puts an invisible catcher overlay on the laser (the tear-off ghost is the same thing with a picture on it). A release on the catcher releases in KWin wherever the pointer last was. Dropping in a gap cancels, just as dropping outside any window does. The pointer helper also tells ft-screens when the mouse's left button comes up, in case the catcher misses it. This also fixes window moves and drags that end off a panel today.
- **The 3D mouse's drag lock.** While the button is held, the drag lock keeps the cursor at its distance and stops hit tests. So a drag onto a nearer panel passes behind it, and a farther panel works only if SteamVR's laser happens to reach it. The change: while the button is held, keep testing the other panels (not the one pressed on) and move onto a panel the ray meets. Off the edge of the pressed panel, keep that panel's plane, as now, so moves and resizes past the edge still work.
- **Drag icon.** KWin 6 draws the drag icon as part of its scene, so it should show on the panel under the laser (to check). In a gap it stops at the source panel's edge. A later version can show a small drag proxy on the catcher.
- **Flatpak apps.** Dropping files into a sandboxed app goes through the document portal, the same path that the session script's file-picker fix covers. Test it explicitly, for example Dolphin to Brave.
## Things that must keep working
- **Typing follows the last click.** A click on a floating panel counts as a click on the desktop, since the window is a KWin window.
- **Visibility.** Floating windows follow the screens' rules: the hide hotkey, the visibility modes, and hiding during a VR game unless the dashboard is open. Controllers' lasers are off in games.
- **Headset standby.** Nothing new may poll SteamVR with new clients, so no new `vrcmd` loops.
- **The pointer helper's overlay list.** The helper learns about overlays by running `vrcmd --overlays` in the background, so a new floating panel appears in its next listing. Check the delay after a tear-off. If it's too long, ft-screens can send the helper new overlay keys directly.
- **Plasma.** An enabled floating output gets a desktop view (wallpaper) under its window, hidden by the crop. Plasma doesn't add panels to new outputs by default. A floating output must never become primary. With spare outputs, the output count stays the same, which avoids the lost-taskbar problem in design.md's open questions.
- **Restarting the desktop** closes every window, floating ones included. Each app's placement is remembered, so an app launched floating again comes back where it was.
## Plan
### Step 1: the catcher and the branches
- The catcher (see "Gaps"), as a fix that stands on its own, so it can go to `main` by itself.
- Merge `pointer-ignore` and `layouts-headpin`.
### Phase 0: find out
Each item has a pass condition. Items that need a desktop restart with extra outputs wait until the headset is free, or run in the headless test mode from 0.1.
- **0.1 Headless test mode.** `ft-screens --no-vr` runs the compositor without SteamVR. It logs toplevels, titles, and sizes, and answers commands on a separate control socket name. This lets KWin and output experiments run without the headset, and without touching the running desktop.
- **0.2 Outputs.** Start KWin with spare outputs, then disable and re-enable one with `kscreen-doctor`. Pass: the toplevel stays and its title changes (as the source says), a disabled output costs no frames, resizing a spare through configure works, and outputs with gaps between them are accepted. Also: an output larger than its window with the window placed inside, and a popup placed in the margin.
- **0.3 KWin script API on 6.2.5.** Move signals fire for moves from both KWin's title bars and apps' own. `sendClientToScreen` and `frameGeometry` work on another output, the `callDBus` long poll works, and `registerUserActionsMenu` works. Popups show up in `windowAdded` with their geometry. The observers can load into the running desktop through `org.kde.kwin.Scripting` and move nothing, so this is safe while the headset is in use.
- **0.4 Frozen-pointer move.** Pass: a KWin move with no pointer movement doesn't shift the window.
- **0.5 SteamVR's laser.** Find out which overlay gets MouseMove and ButtonUp when a held laser moves from overlay A to overlay B, and when it's released over nothing, for both a controller and the 3D mouse. Pass: an interactive overlay placed on the laser reliably catches the release.
- **0.6 Texture bounds.** Pass: `SetOverlayTextureBounds` crops a DMA-BUF (`SharedTextureHandle`) overlay correctly, mouse positions map to the cropped area, and two overlays can show different crops of one buffer.
#### Results (2026-09-29, headless, KWin 6.2.5)
`screens/test/headless.sh` runs these: ft-screens `--no-vr` with a bare nested KWin next to the running desktop, with an `input` command that feeds pointer events as if from a panel, KWin scripts loaded over D-Bus, and screenshots through ScreenShot2.
- **0.1 passes.** `--no-vr`, `--control`, `toplevels`, and `input` are in ft-screens.
- **0.2 passes.** Disabling a spare with `kscreen-doctor` keeps its toplevel, its title gets `- Output disabled`, and it stops committing; enabling it resumes on the same toplevel. A spare resized (`size`) while disabled comes up at the new size on its first frame, so a tear-off needn't blink. Live resizing works, output scale works (1.5: KWin lays out 1067 × 667 on a 1600 × 1000 buffer), and outputs with gaps between them (x = 5000, 8000, 10000) are accepted. A window placed inside a 1600 × 1200 output with a 300 px margin opens a context menu past its bottom and right edges, into the margin.
- **0.3 mostly passes.** `workspace.screens`, `sendClientToScreen`, `windowList`, `frameGeometry` (set; it applies asynchronously), `callDBus`, `registerShortcut`, `registerUserActionsMenu`, and `readConfig` exist. `windowAdded` reports popups (`popupWindow` true, `transient` true) with their geometry. Move signals fire for KWin's title bars (`interactiveMoveResizeStarted` with `move` true, `Stepped` with the geometry, `Finished`). Not yet checked: apps' own title bars, the `callDBus` long poll, and the window menu entry. `globalThis` isn't defined in KWin's script engine. `print` goes to the journal unless `QT_FORCE_STDERR_LOGGING=1`.
- **0.4: KWin starts the move on the press itself,** before any motion. So ft-screens freezes pointer motion as soon as a press lands in a floating window's title bar band (from the frame and client rectangles ft-floatd sends it), with no round trip. For apps that draw their own title bars, the script reports the move and ft-screens freezes then; the script puts back any few pixels the window slipped before that.
- **Found and fixed:** KWin's nested backend ignores the position in `wl_pointer.enter`, and wlroots drops a motion to the position it entered at, so the first click after crossing onto another screen landed where KWin's pointer had been. ft-screens now enters one unit off.
- **0.5 and 0.6** need SteamVR and the headset.
### Phase 1: float a window from its menu
Build the KWin script, ft-floatd, and floating panels in ft-screens: the margin and popup overlays, controls, moving by the title bar, resizing (edges and tab), scale, full screen, close, and back to desktop. Add drag and drop across panels, with the pointer helper change.
Done when:
- Dolphin and Kate float from "Float in VR".
- A file drags from the floating Dolphin to the floating Kate, to a screen, and back.
- The clipboard works between them.
- A menu near a floating window's edge opens past the edge.
- Closing and docking give the output back.
- Frame pacing and GPU memory are measured with several floating windows.
### Phase 2: tear off and dock by dragging
The ghost and tear-off, and the dock highlight with push-flush docking.
### Phase 3: launch floating
`ft-float run` and `ft-float launch`, window matching, new windows of floating apps, the Frametop Apps launcher entry, the picker, remembered placement, and the notification when every spare is in use.
### Phase 4: polish
The drag proxy, the glow toward a window activated out of view (and the setting to bring it in front), the Display Settings section, and docs (design.md, reference.md, and the Use table in the README).
## Risks
- A SteamOS update can change KWin's script API or its nested backend. The script and the output handling are the parts to recheck after one.
- GPU memory: each floating output has its own swapchain of two or three buffers, including the margin. The Frame has 16 GB shared, with about 4 GB free in normal use (2026-09-29).
- Frame pacing with many panels hasn't been measured (already an open question in design.md). Each output is a separate render pass in KWin.
+158
View File
@@ -0,0 +1,158 @@
# Gaze first: plan
**Abandoned on 2026-09-30.** Steam reads the Frame controllers itself, outside SteamVR's bindings, and every press and release it sees takes SteamVR out of laser mode. Gaze mode stays a mouse feature; `docs/gaze-controllers.md` on `experimental` explains what was tried and why it doesn't work. This branch keeps the plan, the probes, and the code for reference; it isn't merged.
Gaze first makes the eyes the pointer for everything flat in VR, and the controllers its buttons. With gaze on and no game running, the gaze drives Frametop's 3D pointer (the `ft_pointer` driver) everywhere the mouse can go, SteamVR's dashboard included. Either controller's trigger clicks where you look, and the controllers stop showing lasers. The work happens on branch `gaze-first` (worktree `~/frametop/.worktrees/gaze-first`, from `experimental`) and is merged into `experimental` after it's been tested in the headset.
The decisions below were made with the user on 2026-09-30. Nothing is built yet: the first step is a headset session with four tests (see "Tests before building").
## Decisions
Input, with gaze on and outside games:
- The gaze drives `ft_pointer`, so anything the mouse does today works with the gaze, on Frametop's screens, floating windows, the keyboard, and SteamVR's and Steam's panels. The mouse keeps working as it does today in gaze mode.
- The gaze wakes the pointer by itself: the headset is worn and the tracker sees your eyes. No mouse movement is needed.
- Trigger, on either controller:
- A tap clicks where you look.
- Moving the controller within `POINTER_GAZE_HOLD` (0.5 s) of the press is precision: the pointer stops where you looked, and your hand's movement steers it, relative, like tap and drag on the Apple Vision Pro (the controller's position seen from the eye, not where it points). The release clicks there, and the correction goes to the gaze service as a lesson.
- Holding still for 0.5 s makes a real press; then the hand's movement drags 1:1.
- Bumper, on either controller: the same, with the right button.
- Right thumbstick: scrolls at the pointer (vertical, and horizontal when pushed sideways).
Controllers:
- They keep their hand roles and their stock SteamVR bindings: the Steam button (tap for the dashboard, double tap for the room view, hold to recenter, the screenshot chord), gamepad mode (both grips), locomotion, room setup, and every per-hand feature.
- Only the trigger and the bumper are muted from SteamVR while gaze first is on, so they never click or take the laser.
- When something still moves the laser to a controller (a Steam button summoning the dashboard, a grip), the helper moves it straight back to our device.
- Later: mute the grips too, and maybe have turning gaze off switch the controllers to gamepad mode.
Turning gaze on and off:
- Gaze is opt-in and off by default. On or off is remembered across restarts (`POINTER_GAZE`), whatever turned it on or off.
- Any game gets the controllers: a VR game (a scene application) or a flatscreen Steam game. Gaze first is off during one.
- The toggle macro, both thumbstick clicks held 1 s by default, turns gaze on or off. In a game it turns gaze first on for that game only, until the game exits.
- You can record your own toggle macro on the Gaze page of Frametop Input Settings: a chord of buttons on one or both controllers, held together for 0.5 to 2 s. The Steam button can't be part of it. It needs 2 or more buttons, or 1 button held at least 1.5 s. Firing it cancels a gaze click in progress. "Reset to default" brings back the thumbsticks. The one-button "Gaze pointer on/off" mapping and key combinations keep working.
- Gaze can't turn on without a calibration for the tracker in use, or without a working tracker service. Turning it on then opens the calibrator, or points to "Repair eye tracker".
Eye tracker:
- Our own tracker (`gaze/tracker/`) is the default. Choosing SteamVR's tracker on the Gaze page turns ours off. The same calibration rule applies to SteamVR's.
- Its root service, `frametop-eyegrab`, is installed by Frametop's installer (a step that defaults to yes) and stays enabled. Turning our tracker off means no longer asking it for frames: it then holds none of SteamVR's buffers, and nothing needs root.
- The tracker waits for a calibration before tracking, and idles whenever gaze input is off (tracking costs roughly 7 to 36 % of a core).
- If the service is missing or broken, gaze can't turn on, and the Gaze page offers "Repair eye tracker". There's no silent switch to SteamVR's tracker.
- Updates: the gaze service compares the installed `ft-eyegrab` with the build, and when they differ the Gaze page offers "Update eye tracker". Both repair and update ask for the password each time, through polkit (`pkexec`). There's no passwordless sudoers or polkit rule: the build output is writable by the user, so such a rule would let any program running as the user get root.
Calibration, in one head-locked panel:
- Quick check: one centre dot. It's captured by a dwell (about 0.6 s of steady fixation; steadiness, not position, so it works however far off the tracker is), the trigger accepts early, and it closes itself after about 4 s if ignored. It opens when the headset is put on, when our tracker notices the headset slipping (at most once every 2 minutes), and from a mappable action. If the first 3 nudges after it are still more than 2 degrees off, it asks for 5 dots.
- Full calibration: the same panel, about 60 degrees across, running the probe's calibration (21 dots in dark, medium, and bright rounds). Frametop's screens hide while it runs. It opens when turning gaze on finds no calibration. Quitting it leaves gaze off; turning gaze on again reopens it. Resetting the calibration while gaze is on turns gaze off and opens it.
- The dots move with your head, so there's no "keep your head still", and the calibration doesn't depend on where the screens are.
- The gaze probe stays as the lab tool.
## What exists
- Gaze mode (`POINTER_GAZE`, `pointer/helper/ft-pointer.cpp`): the gaze aims the pointer; the mouse's held-back press, precision, hold to drag, and nudge lessons are the model for the trigger.
- Holds in the helper (`struct Hold`): pinches and grips steer by the hand's movement seen from the eye, in the room, from where the eye was when the gesture began (`PoseHistory`). The trigger's steering is the same with the controller's position instead of the hand's.
- `gaze_precision` and `gaze_drag` (relay actions, "precision|gazedrag <source> 1|0" to the helper) steer by the controller's aim. Controller holds switch to position steering.
- Controller buttons through the helper's global action sets (`pointer/helper/vrbuttons.h`), one action set per button, active only for mapped buttons and only outside games.
- ft-screens' `controllers always|outside_games|dashboard` and its `hide`/`show` switch.
- Our tracker's headset-moved detection (`gaze/tracker/eyes_model.py`: a glint slip change over 3 px held 1 s) and its reseat after the frames stop.
- `gaze/tracker/install.sh` already installs `ft-eyegrab` as a system service with sudo.
## What SteamVR does (found 2026-09-30)
- The Frame controller's compositor bindings (`/opt/steamvr/drivers/frame_controller/resources/input/vrcompositor_bindings_frame_controller.json`) put the laser on a controller in only two ways: a button that clicks or switches the laser (trigger and bumper click; trigger, bumper, and grip move the laser to that hand with `switchlaserhand`), or summoning the dashboard with its Steam button. Everything else there doesn't touch the laser.
- The gamepad and laser modes are the `/actions/dualanalog` action set (`ModeSwitch1` and `ModeSwitch2` on the grips), with `dashboard.modalGamepadAndLaser`. Those actions are application-scoped: the mode lives in Steam's UI, and we can't set it from outside.
- The laser doesn't need a hand. The headset's own binding (`/opt/steamvr/drivers/frame_hmd/resources/input/vrcompositor_bindings_frame_hmd.json`) runs the laser from `/user/head/pose/raw`. Our device has only tried the right, left, and stylus roles; the stylus attempt got no role at all.
- A Frame controller in the hand takes its hand's role back through its touch sensors, so a device that needs a hand role loses it whenever both controllers are held. That's why the laser has to live somewhere else.
- vrserver's web socket (`/input/getstate.json` and `request_input_state_updates`, as frame-voice uses) lists each controller's `/input/trigger/click`, `/input/bumper/click`, `/input/grip/click`, `/input/thumbstick/click`, `/input/thumbstick/x` and `y`, `/input/system/click`, and each device's `side`. The headset has `/proximity`, but it flickers off for 0.3 to 0.5 s at a time while worn, so the quick check follows SteamVR's activity level (as the helper already does) rather than the raw sensor. Reading the socket takes nothing from anyone. A controller's path is `/devices/cv/<serial>` instead of `/user/hand/<side>` while our device holds that hand.
## Test results (2026-09-30)
Run with the headset on its stand and the controllers on (`pointer/probe/lasertest`, `input/vrws.py`):
1. **Laser on the treadmill role: works.** Our device held the dashboard laser with no hand role. SteamVR gives a device the `/user/treadmill` path only if it hints treadmill when it's added, so the driver hints a role that's no hand from `Activate`; a hint changed on connecting kept it at `/devices/ft_pointer/ft_pointer_0`. `GetControllerRoleForTrackedDeviceIndex` reports no role for it, so the helper's "no hand role" release skips treadmill.
2. **Muting through the helper's action sets: doesn't work.** SteamVR reported those actions inactive (`vrstatus`: `"active": []`) while its laser mouse had input focus, as frame-voice found on 2026-09-26, and a trigger pull moved the laser to its controller until the release. **Muting through the compositor binding works:** `pointer/bindings/vrcompositor_frame_controller_gazefirst.json` is the stock binding without its trigger and bumper laser entries, selected with `POST /input/selectconfig.action` on vrserver's port 27062 (a JSON body `{"app_key": "openvr.component.vrcompositor", "controller_type": "frame_controller", "url": "file:///..."}`; a form-encoded one gets "Parse failed"). `GET /input/getactions.json?app_key=openvr.component.vrcompositor` shows the choice (`current_binding_url`). With it, the triggers showed no laser. Selecting the stock file puts it back.
3. **Snap back: works** in 11 to 15 ms after a grip or a Steam button summon. A trigger held the laser until its release (0.4 to 1 s), which the muting removes. Our device also takes the laser from "none".
4. **Web socket: works.** Every click, the thumbstick axes, and both thumbsticks clicked together arrive. The headset's `/proximity` flickers off for 0.3 to 0.5 s at a time while worn, so the quick check goes by SteamVR's activity level instead.
Also found: the bumpers, the thumbsticks' movement, and their clicks switch Steam's dashboard into its controller (gamepad) mode, and the laser owner goes to "none"; both grips go back to laser mode. Steam's own Frame controller binding (`steam_vrgamepad_bindings_frame_controller.json`, app `steam.client`) has only haptics, so that input reaches Steam's UI some other way (most likely Steam Input's virtual gamepad), and a SteamVR binding can't mute it.
Decided after the tests: gaze replaces the laser pointer's controls, never the controller mode's. Controller mode takes precedence while it's on, and leaving it gives the laser back to the gaze. Right click is one grip held with a trigger (the bumpers belong to controller mode).
## Tests before building
One headset session, about 30 minutes, with the user wearing the headset. Installing the test driver needs a SteamVR restart, which closes everything in VR, so it's done at the start of the session and only when the user says so.
1. **Laser on the treadmill role.** The driver takes a `role treadmill` command, and its compositor bindings repeat the right hand's under `/user/treadmill`. Pass: after its `switchlaserhand` (`/input/a`), `GetPrimaryDashboardDevice()` is our device, and the pointer clicks Frametop's screens and the dashboard while both controllers are held. Fail: the fallback (below).
2. **Muting.** The helper activates its trigger and bumper action sets on both sides. Pass: a real trigger or bumper, pointed at a panel, neither clicks nor moves the laser, with the dashboard open and closed, and the helper sees the press. Fail: we need our own compositor binding for the Frame controller, switched when gaze first turns on and off.
3. **Snap back.** A Steam button tap on either controller, and a grip. Pass: the helper sees the laser move to a controller and moves it back within about 100 ms, with no stray click.
4. **Web socket.** Update rates for the thumbstick axes and clicks while held, and `/proximity` at don and doff.
Fallback if test 1 fails: gaze input goes straight into ft-screens for Frametop's own panels (the helper already knows which panel it hits and where), with both controllers held. SteamVR's and Steam's panels then take the gaze only while one hand is empty.
## Design by component
### Driver (`pointer/driver`)
- `role treadmill` joins `right`, `left`, and `stylus`. The hint stays OptOut while disconnected, as now.
- `ft_pointer_vrcompositor.json` and `ft_pointer_steam.json` get the same bindings under `/user/treadmill`. They're additive, so nothing changes while the device is a hand.
### Pointer helper (`pointer/helper/ft-pointer.cpp`)
- Gaze first = gaze mode on, the gaze service ready, no game, or a game with the macro's override.
- Waking: in gaze first, fresh gaze with the headset worn wakes the pointer (connect, treadmill role, `switchlaserhand`), with no mouse counts. The "no hand role" release doesn't apply to the treadmill role.
- The laser: while awake in gaze first, if the dashboard's primary device becomes a controller, press `/input/a` on ours again, at most every 100 ms. Last used wins stays off in gaze mode, as now.
- Muting: the trigger and bumper action sets on both sides are active whenever gaze first is on, whatever the relay's mappings say. Their presses go to the built-in state machine, not to the relay.
- Trigger and bumper: a hold with a new source, a controller's position. It starts as a held-back press at the gaze; moving past `POINTER_TRIGGER_DEADZONE` (about 1 degree, seen from the eye; the pull jolts the controller) within `POINTER_GAZE_HOLD` is precision at `POINTER_TRIGGER_GAIN` (0.5); still until `POINTER_GAZE_HOLD` is a real press, and dragging at `POINTER_GAZE_DRAG_GAIN` (1). Release: click, lesson (as with the mouse, up to `POINTER_GAZE_NUDGE_MAX`), or release the press. The existing controller holds (`gaze_precision`, `gaze_drag`) move to position steering too.
- Scroll: "scroll <dx> <dy>" from the relay goes to the driver, as the mouse's wheel does.
- Calibration: while ft-gazed says its panel is up ("calpanel 1|0"), the pointer hides, and a trigger press goes to ft-gazed as "calaccept" instead of clicking.
- The macro's "cancel": drops a held-back press without clicking.
- Headset worn or not goes to ft-gazed ("headset 1|0"), for the quick check.
### Input relay (`input/input-relay.py`)
- A web socket reader for vrserver, in the standard library (the host's Python has no `websockets` module, and the relay needs nothing but Python). It follows the controllers by `side`, since their paths change with the roles.
- The toggle macro: `"gaze_macro": {"buttons": ["left/thumbstick", "right/thumbstick"], "hold": 1.0}` in the rules file. Outside games it writes `POINTER_GAZE` and tells the helper; in a game it toggles the override for that game. It sends the helper "cancel" when it fires.
- Recording: Input Settings asks for "macro record" and gets back the chord and how long it was held, after the checks in "Decisions".
- Scroll: the right thumbstick's axes, while gaze first is on, as "scroll" lines to the helper.
- Games: a Steam game running (a `reaper` process with `SteamLaunch AppId=`, checked every 2 s; to confirm with a flatscreen game) goes to the helper, which already knows scene applications.
- `gaze_toggle` writes `POINTER_GAZE` instead of lasting until a restart.
### Gaze service (`gaze/ft-gazed`)
- `GAZE_TRACKER` defaults to `own`.
- Readiness goes to the helper ("gazeready 1|0"): the tracker's service works (for ours: `frametop-eyegrab` active, and its binary the same as the build) and the tracker in use has a calibration. When gaze is on but not ready, it opens the calibrator, or asks for the repair.
- Our tracker runs only while gaze is on and ready, while a calibration runs, or for the probe's lease.
- The quick check opens on "headset 1" from the helper, on our tracker's jump (at most once every 2 minutes), and on "quickcal" (the mappable action). For our tracker, the dot is a click on `@ft_eyes`, like the probe's one-dot check; for SteamVR's, it's a lesson.
- The full calibration's logic (the dots, rounds, sample rejection, and fits) moves out of the probe into `gazecal.py`, so the probe and the service share it. Quitting writes `POINTER_GAZE=0`.
- Hiding Frametop's screens during a full calibration goes through ft-screens' `hide` and `show`, keeping the user's own switch as it was.
### Calibration panel (`gaze/panel/ft-gazepanel`, new)
A small C++ OpenVR overlay program in the dev container: a head-locked overlay (placed relative to the headset) of a fixed angular size, drawn on the CPU and uploaded with `SetOverlayRaw`, like ft-screens' keyboard, with labels from stb_truetype. ft-gazed drives it over `@ft_gazepanel` (show quick or full, dot at yaw and pitch with its state, background brightness, a line of text, hide). It takes no input: the trigger comes through the helper, and the dwell is ft-gazed's. It's a separate program because the helper is already 2,000 lines, and the panel has nothing to do with pointing.
### Frametop Input Settings, Gaze page
- The gaze switch, blocked with the reason while not ready ("Calibrate first", "Repair eye tracker").
- Eye tracker: ours (default) or SteamVR's.
- Status: the service, the calibration, and when it was made. Buttons for Calibrate, Quick check, and Repair or Update eye tracker (`pkexec gaze/tracker/install.sh`, which then runs without sudo inside).
- The toggle macro: what it is, Record (a 3 s countdown, hold the chord, confirm), and Reset to default.
### Installer (`install.sh`)
- A step for our eye tracker ("It needs your password (sudo)"), defaulting to yes, and the gaze service, which needs no root. Gaze itself stays off.
## Order of work
1. The tests above, then this plan updated with the results.
2. Driver and helper: the treadmill role, muting, snap back, waking by gaze, and the trigger and bumper holds. Gaze still turns on the old way.
3. Relay: the web socket reader, scroll, the macro, games and the override, `POINTER_GAZE` remembered.
4. Gaze service: readiness, idling, the default tracker, the service check, and the quick check's triggers.
5. The calibration panel and calibrator: the quick check, then the full calibration and the move to 5 dots.
6. Input Settings, the installer, and the docs (`README.md`, `docs/design.md`, `docs/reference.md`, `gaze/README.md`).
Each step is tested in the headset before it's merged into `experimental`.
## Risks
- SteamVR may refuse the laser on a treadmill device, or a SteamVR update may change the Frame's compositor bindings.
- The muting relies on "Enable global input from overlays" (`steamvr/globalActionSetPriority`), which SteamVR calls experimental.
- `pkexec` runs a script the user can write. That's acceptable only because every run asks for the password.
- Our tracker's CPU cost while gaze is on, with SteamVR and a busy desktop; it idles otherwise.
- Head-locked panels can be uncomfortable; the panel stays small and short-lived, except for the full calibration.
+140
View File
@@ -0,0 +1,140 @@
# Hands in Frametop: migration plan
Hand tracking from the headset's own cameras has been built as a separate project, frame-hands (`~/Desktop/Projects/frame-hands` on the Frame, a local git repo with no remote). The plan is to make it a native Frametop component, like `gaze/` and `power/`, instead of a separate module. The work happens on branch `hands-migration` (worktree `frametop-hands/` in the PC workspace) and is merged into `experimental` after it's been tested in the headset.
Builds from this worktree must sync to their own folder on the Frame, never `~/dev/frametop`: run every script with `FRAME_REPO=/home/steamos/dev/frametop-hands`.
## Status (2026-09-30)
Steps 1-5 are done:
- frame-hands' pending work was committed there (6c63c9e).
- Its filtered history was merged under `hands/` (1a76d15), then laid out (`trackd/` to `track/`).
- The renames, the Frametop paths, and ft-camd's file capabilities are done. So are `hands/Makefile`, `build.sh`, `run.sh`, the two units, the README, the settings, the installer step, and ft-screens on the shared header.
- Built in the dev container on the Frame, and on the 7i.
- Checked without the headset:
- `ft-handreplay` against frame-hands' `fh-replay`, both x86 with `--cost`, on the whole dim recording and the first 60 s of the bright one: identical summaries and byte-identical depth dumps. The Makefile's own ncnn build is included in that.
- `ft-ringplay` into `ft-hands` on the 7i tracked, pinched, and wrote `/run/user/UID/frametop-hands/{hands,gestures}`.
- On the Frame, ft-hands in the container finds the calibration through `/run/host/persist`, and ft-camd without its capabilities refuses with a clear message.
Step 6 has started (2026-09-30 10:30):
- `hands/run.sh install` is done, and both services run from this worktree.
- The files moved to `/run/user/UID/frametop-hands/`, because `/run/user/UID/frametop` is the desktop session's own runtime folder, deleted at every desktop start.
- Until the desktop restarts from a build with this branch's ft-screens, the link `/run/user/UID/frame-hands -> frametop-hands` feeds the running one. It's tmpfs, so it's gone at reboot.
Found in the headset:
- The side cameras were swapped (`HANDS_SWAP_SIDES=1`).
- The cutout copy shader lost resolution at `mediump` (now `highp`).
- Colour capture isn't reliable (see the README).
- Two pinch fixes: one hand no longer pinches both sides, and the palm-down limit stops typing pinches.
## What frame-hands is today
| Part | What it is | Size |
| --- | --- | --- |
| `camd/` | `fh-camd`, the camera broker (C). It borrows XRService's camera DMA-BUFs read-only with `pidfd_getfd`, times them with the `v4l2_dqbuf` tracepoint, and publishes the four IR cameras (and optionally the two colour cameras) to a shared-memory ring. It starts as root and drops to the user after setup. Adapted in part from FrameEyeCameraFeed (MIT, licence file kept). `fh-camprobe` is its discovery and recording probe. | camd 1.1k lines, tp 0.4k, xrcams 0.8k, camprobe 1.1k |
| `trackd/` | `fh-tracker` (C++): the tracker, the models on ncnn, the calibration (jsoncpp), the pinch detector, the publisher, and the recorder. Also `fh-replay` (offline replay and scoring), `fh-ringplay` (plays a recording into a ring), and `nettest`. | 2.9k lines |
| `include/` | The hands file (`fh_hands.h`, read by ft-screens) and the gestures file (`fh_gestures.h`, pinches). | |
| `models/ncnn/` | MediaPipe's palm detector and hand landmark model, from the OpenCV Zoo ONNX ports (Apache-2.0), converted to ncnn in float and int8. | 5.9 MB |
| `tools/` | Python analysis: side-camera check, colour calibration check, frame viewer, gesture watcher, depth report, model comparison, int8 calibration, model conversion. | ~1.1k lines |
| `tracker/` | The Python prototype of the tracker. Some tools import its `calib.py` and `models.py`. | 1.3k lines |
| `probes/`, `notes/`, `re/`, `shim/` | One-off experiments, reverse-engineering notes on SteamVR's passthrough internals, a disassembly (not in git), and a header from an abandoned XRService shim approach. | |
| `vendor/`, `captures/` | ncnn and FrameEyeCameraFeed clones, and recordings of the user's hands and room (tens of GB). Neither is in git. | |
Today it runs by hand: `sudo camd/fh-camd`, then `trackd/fh-tracker`. There are no units and no installer. Files: `/run/frame-hands/ir-ring` (the ring, in a root-owned folder), and `$XDG_RUNTIME_DIR/frame-hands/hands` and `gestures`.
Frametop already has the consumer side on `experimental`: `screens/handcut.{h,cpp}` cuts the hands out of the screens, with its own copy of the hands file layout, and `screens/handtest.cpp` tries it on a test panel.
## Where it goes
A top-level `hands/` folder, laid out like `gaze/`:
```
hands/
README.md # from trackd/README.md and camd/README.md
build.sh # ft-camd, ft-hands; --tools also builds the replay tools
run.sh # install|uninstall|start|stop|restart|status|log
frametop-camd.service # user units (templates, @REPO@)
frametop-hands.service
include/ # fhring.h, fh_hands.h, fh_gestures.h: shared with screens/ and pointer/
camd/ # ft-camd: camd.c tp.c xrcams.c, LICENSE.FrameEyeCameraFeed
track/ # ft-hands: tracker, nets, calib, pinch, publish, record; replay.cpp
# (ft-handreplay) and ringplay.cpp (ft-ringplay) for recordings
models/ # the ncnn models, with NOTICE (Apache-2.0, MediaPipe / OpenCV Zoo)
tools/ # the Python checks, watch_gestures, depth_report, calib.py, ring.py
```
Left behind in frame-hands, which stays as the lab: the recordings, the Python prototype (the tools that need `calib.py` or `models.py` get a trimmed copy in `hands/tools/`), `probes/`, `notes/`, `re/`, `shim/`, `camprobe`, and `vendor/`. The reverse-engineering notes don't belong in a public repo, and recordings are images of the user's hands and room, so they never go into git.
## Names
Programs within 15 characters, `ft-` prefix; files under `frametop`:
| Now | In Frametop |
| --- | --- |
| `fh-camd` | `ft-camd` |
| `fh-tracker` | `ft-hands` |
| `fh-replay`, `fh-ringplay` | `ft-handreplay`, `ft-ringplay` |
| `/run/frame-hands/ir-ring` | `/run/user/UID/frametop-hands/cam-ring` |
| `$XDG_RUNTIME_DIR/frame-hands/hands`, `gestures` | `/run/user/UID/frametop-hands/hands`, `gestures` |
The source keeps its `fh_` identifiers and header names (`fh_hands.h`, `fh_gestures.h`, `fhring.h`), and the file formats keep their magic strings, so recordings and tools from frame-hands keep working. Programs, units and runtime paths change.
## Build
- `hands/build.sh` builds in the dev container through `scripts/frame.sh --build`, into `hands/build/`, like the other components. `FRAME_BUILDER=pc` can take the ncnn build.
- ncnn: fetched at a pinned tag (20260526, as now) into `hands/build/ncnn` and built once, the way `screens/build.sh` fetches the OpenVR header, with frame-hands' options so results match. `NCNN=` points the build at an existing install instead. Every net runs single-threaded (`num_threads = 1`), with the tracker spreading nets over its own pinned threads, so OpenMP could go later.
- ft-hands runs in the dev container like ft-pointer and ft-powerd (`distrobox enter dev --`, after `scripts/container-up.sh`). Today's fh-tracker runs on the host and works only because the host happens to have the same `libjsoncpp.so.25` and libgomp as the container. Inside the container the calibration is at `/run/host/persist`, and calib.cpp (and `tools/calib.py`) fall back to it when `/persist` isn't there.
- ft-camd has to run on the host (below), so it's linked statically (only libc and libm; `glibc-static` goes into `setup/dev-container.sh`). The host has an older glibc than the container.
## Running it
**ft-camd needs privileges**, only while it sets up: `pidfd_getfd` on XRService (the Frame has `ptrace_scope=1`), system-wide tracepoints (`perf_event_paranoid=2`), and the tracepoint files, which are root-only (`/sys/kernel/tracing/events/v4l2/v4l2_dqbuf/{id,format}` are mode 0440). A rootless container's root can't do any of that, so it runs on the host. Two ways:
- **A. File capabilities (chosen, 2026-09-30).** The installer runs `sudo setcap cap_sys_ptrace,cap_perfmon,cap_dac_read_search+ep hands/build/ft-camd` once. ft-camd then runs as the user, in a user unit `PartOf=steamvr.service`, so it starts and stops with SteamVR, and its ring lives in the user's runtime folder. It drops all capabilities after setup, as it drops root today. Nothing ever runs as root. Writing the file clears its capabilities, so a rebuilt ft-camd needs the setcap again. It changes rarely. `/home` on the Frame is ext4 without `nosuid`, so file capabilities work there.
- **B. Root system service**, like the Bluetooth fixes: a root-owned copy in `/var/lib/frametop/`, a unit in `/etc/systemd/system/`. It would have to watch for XRService itself, because a system unit can't follow the user's `steamvr.service`.
Either way the password is needed once at install, through the same `sudo -S` path the Bluetooth fixes use, and only after asking.
**ft-hands** is a user unit, `frametop-hands.service`: after `frametop-camd.service`, `PartOf=steamvr.service`, nice 5, model threads on CPUs 5-7 (measured best on 2026-09-29).
**Settings** in `~/.config/frametop.conf`: `HANDS_SWAP_SIDES=1` and `HANDS_CPUS=5,6,7`, read by ft-hands. It's on while its services are installed (`hands/run.sh install`, `uninstall`), so there's no `HANDS` switch. There's no setting for colour yet. Later, a switch in Frametop Display Settings.
**Installer:** an optional last step in `install.sh`, off by default, which asks first because it needs sudo.
## Interfaces
- `screens/handcut.cpp` includes `hands/include/ft_hands.h` instead of its own copy of the layout, and reads the new path. ft-screens and ft-hands change together on this branch.
- Pinches go to the pointer helper. It maps the gestures file and checks the begin and end counters each tick. A begin is a press, an end the release, and the pinch point's movement a drag. In gaze mode, the press lands where you look. The counters mean a quick tap between two ticks isn't missed. The tracker knows nothing about the pointer.
## Open items that aren't part of the move
These block shipping hands to other people, not the migration:
- **The side-camera swap.** After some XRService restarts, fh-camd publishes the two side cameras under each other's names. Today it's caught by hand (`tools/check_sides.py --ring`, then `--swap-sides`). It needs fixing at the source (tell the buffers apart by the `dqbuf` tracepoint's device, the way the colour pair is split), or at least an automatic check at start-up.
- **The colour cameras' calibration mapping** (`tools/check_color.py` on a recording with texture).
- **Depth when one camera loses the hand.** From the 2026-09-30 replay measurements: drifting 10% per update toward the one-camera guess (`kMonoDepthGain`) makes the depth worse than keeping the last distance. Try 0.02.
## Public repo
Frametop is public. **Not pushed to GitHub until the user says it's ready** (user decision, 2026-09-30). When it is, it publishes:
- The camera borrowing (`pidfd_getfd` on XRService's buffers) and the tracepoint timing. FrameEyeCameraFeed already does the same publicly. Its MIT licence and credit stay with the code.
- The models, under Apache-2.0, with a NOTICE.
It doesn't publish the reverse-engineering notes, the probes, or any recording. They stay in frame-hands.
## History
frame-hands' work was committed there first (6c63c9e, its 4th commit). Its history was then filtered to drop what stays behind (`notes/`, `probes/`, `shim/`, `camd/camprobe.c`, the camprobe tools, the Python prototype except `calib.py` and `models.py`, and `.frame-job`) from every commit. It was merged into this branch under `hands/` (a subtree merge), so blame still leads to where each line came from. The renames come after, as their own commits.
## Steps
1. In frame-hands: commit the pending work, as its last state before the move (needs the user's OK).
2. On this branch: import it under `hands/`, then rename the programs and paths. The behaviour stays identical.
3. `hands/build.sh`, `run.sh`, the two units, the README, the settings, and the installer step.
4. ft-screens' hand cutouts on the shared header and the new path.
5. Check without the headset. `ft-handreplay` on the 2026-09-29 recordings with `--cost` is repeatable, so its summary must match `fh-replay`'s exactly. And `ft-ringplay` into ft-hands must publish the same hands as into fh-tracker.
6. In the headset, with the user and after asking: stop fh-camd and fh-tracker, install the services from `~/dev/frametop-hands`, and restart the desktop from this branch so ft-screens reads the new path.
7. Pinch into the pointer helper (it can also follow the merge). The gaze work is on the Frame's `~/frametop` main: on 2026-09-30 that branch had 4 commits `experimental` doesn't have, plus uncommitted work in the pointer helper's gaze mode. Build this step on wherever that work lands, not on this branch's older copy.
8. Merge into `experimental`. It's checked out in a worktree on the Frame (`~/frametop/.worktrees/experimental`, where the live desktop runs), so the merge happens there, or `experimental` is switched away first.
+1
View File
@@ -18,6 +18,7 @@ ft-screens drops keys while no screen has focus or the SteamVR dashboard is open
- **A release that never arrives leaves the key held in the desktop.** KWin repeats held keys itself, so a stuck letter repeats and a stuck modifier changes every later key (Ctrl+Alt held turns T into Konsole). Pressing and releasing the key again clears it.
- **A keyboard that disconnects mid-press is one way to get there.** The relay forgets the held key without telling ft-screens. The same goes for the relay restarting while a key is down.
- **To see where a key went,** run `scripts/keys-report.py` and reproduce the problem while it records. It logs the modifiers, Tab, and Esc (no other keys) as the relay reads them and as its virtual keyboard sends them on, with the device roles and grabs, which programs have each keyboard open, and the relay's and desktop's logs.
- **Switching where typing goes waits for keys to come up.** The relay changes a keyboard's grab only while none of its keys are down, so a press and its release go to the same side. A key held for a long time delays the switch until it's let go.
## Typing and grabbed keyboards
+59 -15
View File
@@ -39,7 +39,9 @@ The controls are sized from both the screen's width and its distance from you, f
To pin a screen to a wrist, carry it by its bar and sweep the laser across your other controller. A ring around that controller marks the target, and a dot shows where the laser passes. Crossing the ring arms the pin, and the ring and bar turn blue; crossing it again disarms it. When you let go while armed, the screen rides on that controller at the size, distance, and angle it had, so you can arm the pin first and then turn the screen the way you want. Grab a pinned screen's bar to adjust it; it goes back to the same wrist when you let go unless you disarm it. A pinned screen shows only while you're looking at its front, within the wrist angle, and fades out over the last 10°.
The Visibility & wrist tab of Frametop Display Settings decides when the screens show:
To pin a screen to your head, like a HUD, set it to On your head on the Visibility & pins tab of Frametop Display Settings (or `ft-layout pin N head`). It rides on the headset where it is at that moment, so place it first, and it shows whenever the screens do. Grab its bar to move it; it goes back on your head where you let go. Sweeping across a wrist ring while you carry it moves it to that wrist, and sweeping across again leaves it in the room. The 3D mouse's dot stays in the room, so a head-pinned screen moves away from it when you turn your head, unless head follow is on.
The Visibility & pins tab of Frametop Display Settings decides when the screens show:
- Always. Meta+Shift+H, the Hide/Show Screens menu entry, or a mapped mouse button hides them.
- Only while the SteamVR dashboard is open.
@@ -53,12 +55,16 @@ In the last three modes the hotkey shows the screens anyway. Two more settings o
Input from the lasers reaches KWin through ft-screens' own seat. Keys come from the input relay, from pass-through keyboards and any key a pointer device passes through. Typing follows your last click: after a click on a screen it goes to the desktop, even with the SteamVR dashboard open, and after a mouse click on any other panel (the dashboard, Steam, an app like Spotify) it goes there instead. While it goes to the desktop, the relay grabs pass-through keyboards so gamescope, which reads every keyboard itself, doesn't type them into the Steam app too. A program that watches every keyboard for a hotkey loses a grabbed one; with `SHARE_KEYS=1` in `~/.config/frametop.conf`, their keys also go to `@frametop_keys` for it. That's off by default, since any local process that binds the name first would get everything typed into the desktop. Hidden screens don't take typing.
Frametop's keyboard opens by itself when a text field on the desktop gets focus, and stays open until its Close key, a layout reset, or a mapped button closes it (or, with Keep it open off in Frametop Input Settings, until the text field loses focus). While the Steam menu (the dashboard) or Steam's own keyboard is up, it steps aside, and it comes back where it was when they're gone; one asked for meanwhile appears then. In the "only with the dashboard" visibility mode, the dashboard doesn't count. It doesn't open without a head pose (the headset in standby). It's a panel of keys (a US laptop layout, with Esc where Caps Lock would be, arrows, and a Close key) that ft-screens shows 0.7 m in front of you and below your eyes, facing you. It stays where it opened, and its grab bar (the pill along the top) moves it like a screen's. Type on it with a controller's laser or the 3D mouse. Shift, Ctrl and Alt latch for the next key, and a held key repeats. KWin starts `input/ft-textinput` as the desktop's input method, and KWin activates it whenever the focused app turns on text input for a field. It tells the relay (`textfield 1` or `0`), the relay decides by the Keyboard setting in Frametop Input Settings, and ft-screens opens the keyboard for the screen that has keyboard focus (`vrkeyboard show`, `hide`, or `toggle` from a mapped button). Its keys reach the focused screen as key presses, so it works in every app, but only apps that use Wayland text input (Qt, GTK, Firefox) open it by themselves; Chromium, Electron and X11 apps need the button. The session drops the `QT_IM_MODULE=xim` and `GTK_IM_MODULE=xim` that the gamescope session sets, or Qt and GTK apps wouldn't use Wayland text input either.
KWin's nested backend doesn't undo a screen's scale on pointer input, so ft-screens divides panel positions (in pixels) by it. `ft-layout` sends it each screen's scale as KWin reports it (`scale N s`) whenever it applies scales: at desktop start and from Frametop Display Settings. A scale changed only in Plasma's own display settings is put back to the Frametop layout's the next time `ft-layout` runs.
ft-screens listens for datagrams on the abstract socket `@ft_screens` and replies to the sender:
```
place N x y z yaw pitch roll width N metres curve N radius|on|off
pin N|all left|right [matrix] unpin N|all size N w h
get N screens head state key code value
pin N|all left|right|head [matrix] unpin N|all size N w h
get N screens head state key code value scale N s vrkeyboard show|hide|toggle
visibility always|dashboard|gesture|toggle wrist degrees gesture left|right degrees
hide | show | toggle controllers always|outside_games|dashboard ingames hide|visible
```
@@ -98,17 +104,19 @@ pointer/helper/run.sh status | log | restart
pointer/driver/install.sh probe # devices, hand roles, who owns the dashboard pointer
```
The pointer settings are in `~/.config/frametop.conf`: `POINTER_SENSITIVITY`, `POINTER_IDLE`, `POINTER_WAKE_COUNTS`, `POINTER_DISTANCE`, `POINTER_CURSOR_DEG`, `POINTER_ORIGIN_FRACTION`, `POINTER_ORIGIN_MARGIN`, `POINTER_SCENE_RADIUS`, `POINTER_EDGE_REACH`, `POINTER_LASER_WIDTH`, the head follow settings `POINTER_FOLLOW`, `POINTER_LEASH_DEG`, `POINTER_LEASH_DELAY`, `POINTER_LEASH_RETURN`, and `POINTER_FOLLOW_REACH`, and the gaze mode settings `POINTER_GAZE`, `POINTER_GAZE_RETAKE`, `POINTER_GAZE_NUDGE_MAX`, `POINTER_GAZE_HOLD`, and `POINTER_GAZE_SHOW`. The example config explains each. Frametop Input Settings changes them live; after editing the file by hand, restart the relay or the helper.
The pointer settings are in `~/.config/frametop.conf`: `POINTER_SENSITIVITY`, `POINTER_IDLE`, `POINTER_WAKE_COUNTS`, `POINTER_CONTROLLER_PICKUP`, `POINTER_DISTANCE`, `POINTER_CURSOR_DEG`, `POINTER_ORIGIN_FRACTION`, `POINTER_ORIGIN_MARGIN`, `POINTER_SCENE_RADIUS`, `POINTER_EDGE_REACH`, `POINTER_LASER_WIDTH`, `POINTER_IGNORE`, the head follow settings `POINTER_FOLLOW`, `POINTER_LEASH_DEG`, `POINTER_LEASH_DELAY`, `POINTER_LEASH_RETURN`, and `POINTER_FOLLOW_REACH`, and the gaze mode settings `POINTER_GAZE`, `POINTER_GAZE_RETAKE`, `POINTER_GAZE_NUDGE_MAX`, `POINTER_GAZE_HOLD`, and `POINTER_GAZE_SHOW`, and the gaze service's `GAZE_TRACKER` (SteamVR's eye tracker or our own) and `GAZE_EYE` (the eye bias). The example config explains each. Frametop Input Settings changes them live; after editing the file by hand, restart the relay or the helper (the gaze service reads its two again when the file changes).
## Frametop Input Settings
A Kirigami app with a Python backend, in the Plasma menu under Settings. It runs in the `dev` container and talks to the relay over its control socket, `@frametop_relay`. It has six pages:
A Kirigami app with a Python backend, in the Plasma menu under Settings. It runs in the `dev` container and talks to the relay over its control socket, `@frametop_relay`. It has eight pages:
- Devices lists every USB and Bluetooth mouse and keyboard, with a light that flashes when the device is used. Each device gets a role: 3D pointer (grabbed, drives the pointer; the default for anything with a mouse), Pass through (grabbed only while typing goes to the desktop; the default for keyboards, where a Meta tap toggles the dashboard if `META_DASHBOARD=1` is in `~/.config/frametop.conf`), or Ignore. A device is identified by its Bluetooth address, or its USB ids and name, so all of its input nodes share one role. Forget drops everything saved for a device.
- Buttons maps a pointer device's buttons. Choose Capture a button, press the button or key, then pick an action: a click, back, scroll, toggle dashboard, recenter, pointer on or off, head follow on or off, gaze pointer on or off, faster or slower, pass the key through, or nothing. Devices with saved mappings are listed even while they're asleep.
- Buttons maps a pointer device's buttons. Choose Capture a button, press the button or key, then pick an action: a click, back, scroll, toggle dashboard, recenter, pointer on or off, open or close the keyboard, head follow on or off, gaze pointer on or off, faster or slower, pass the key through, or nothing. Devices with saved mappings are listed even while they're asleep.
- Controllers maps the Frame controllers' buttons (every button but the system button) to the same actions, except passing a key through. Capture a button and press it on a controller, or pick it from the list. The controllers aren't input devices on the host; only SteamVR sees them. So the pointer helper reads them with SteamVR input (`pointer/helper/vrbuttons.h`, `pointer/helper/actions/`) and sends presses to the relay (`vrbtn right/a 1`), which does the mapped action. The helper only takes the buttons that are mapped (the relay tells it with `vrbind`), at an overlay-global priority, and only while no game (scene application) runs, so games keep every button; with In games on (`controller_in_games`), a mapped button is taken from games too. That needs SteamVR's "Enable global input from overlays (Experimental)" setting (`steamvr/globalActionSetPriority`), which the page's Global input switch turns on and off. Mappings are saved as `controller_buttons` in `~/.config/frametop-input.json`.
- Keyboard sets when Frametop's keyboard opens: whenever a text field is selected; only while no pass-through keyboard is connected (the default; keyboards other programs make through uinput, like frame-voice's, don't count); only with a mouse or controller button mapped to Open/close keyboard; or never, which turns the button off too. Keep it open (on by default, `vr_keyboard_persist`) leaves it open after the text field loses focus. The mode is saved as `vr_keyboard` in `~/.config/frametop-input.json`, and the page lists the keyboards that count as connected.
- Pointer has a Head follow switch and sliders for the pointer settings, which apply immediately, and a Recenter button.
- Gaze has the gaze pointer switch (on now and from now on; a mapped button toggles it until the helper restarts), the gaze mode sliders, the gaze service's state (headset, samples per second, whether only one eye is tracked, the calibration, the nudges learned), and Calibrate (opens the gaze probe), Reload calibration, and Forget nudges.
- Ignored panels lists the SteamVR overlays that are showing, grouped by app (the first two parts of the overlay key, such as `sasaken.frame-perf-overlay`), from the pointer helper (`overlays`). Tick a panel, or Ignore the whole app, and the pointer passes through it to what's behind. It's for panels you only look at, like a performance overlay that follows your view. The list is saved as `POINTER_IGNORE` in `~/.config/frametop.conf`: comma-separated overlay keys, where a shell pattern like `vendor.app*` covers a whole app, including panels it opens later. The helper reloads at once. Frametop's own screens aren't listed, and entries for apps that aren't open are listed below, to remove.
- Gaze has the gaze pointer switch (on now and from now on; a mapped button toggles it until the helper restarts), the gaze mode sliders, the gaze service's state (headset, samples per second, how often the tracker is losing each eye, the calibration, the nudges learned), and Calibrate (opens the gaze probe), Check headset fit (opens the probe's Headset fit mode), Reload calibration, and Forget nudges.
- Bluetooth lists paired devices and has Apply Bluetooth fixes, which runs `/etc/steamframe/bt-fixups.sh` through `pkexec`. Pair new devices in Steam.
Device rules are saved in `~/.config/frametop-input.json`. `input-settings/install.sh` installs the menu entry. Its launcher hands podman the real `XDG_RUNTIME_DIR` and user bus and gives the app the session's Wayland socket, because the desktop session runs on a private D-Bus and podman fails on it.
@@ -117,19 +125,24 @@ Device rules are saved in `~/.config/frametop-input.json`. `input-settings/insta
When the desktop starts, its screens arrange themselves around where you're facing. You can move them by hand at any time and put them back with Meta+Shift+R, the Reset Screen Layout menu entry, Arrange now in the app, or a mouse button mapped to Reset desktop screen layout.
The desktop's own screen arrangement follows where the screens are around you, whatever their numbers: a screen you see to the left of another is to its left in Plasma too, so the pointer and dragged windows cross straight to it. Screens one above the other stack, and screens pinned to a wrist come last. It's updated at startup, after arranging or saving the layout, and half a second after you let go of a screen you moved. With the headset off there's no head pose to go by, and the arrangement stays as it was.
The desktop's own screen arrangement follows where the screens are around you, whatever their numbers: a screen you see to the left of another is to its left in Plasma too, so the pointer and dragged windows cross straight to it. Screens one above the other stack, and screens pinned to a wrist or your head come last. It's updated at startup, after arranging or saving the layout, and half a second after you let go of a screen you moved. With the headset off there's no head pose to go by, and the arrangement stays as it was.
Frametop Display Settings has three tabs:
Frametop Display Settings has four tabs (three with the gamescope backend, which has no Visibility & pins):
- Screens: add and remove screens, and set each one's resolution (presets from 1080p to 4K, ultrawide, super ultrawide, portrait, or custom), its width in VR (0.5 to 6 m), its scale, whether it's curved, and whether it has the taskbar. Resolution, width, and curve apply at once. Adding or removing a screen takes a desktop restart, which the app offers.
- Layout: a curve around you, with the screens hinged edge to edge like monitors on a desk and each turned to face you, or a flat wall. Both take rows, distance, gap, and height. Save current arrangement keeps the positions and sizes you set by hand instead. A preview shows the layout from above and from the front, and a switch turns auto-arrange at startup on or off.
- Visibility & wrist: the visibility, game, and controller settings described above, the wrist angle, and buttons to pin all screens to a wrist or unpin them.
- Layout: a curve around you, with the screens hinged edge to edge like monitors on a desk and each turned to face you, or a flat wall. Both take rows, distance, gap, and height. Save current arrangement saves the positions, sizes, curves, and pins you set by hand under a name instead. Named layouts are listed with the presets: pick one and Arrange now to switch to it, and rename or delete it with the buttons next to the list. A layout saved with fewer screens than you have now leaves the others where they were saved last, or where the preset would put them. A preview shows the layout from above and from the front, and a switch turns auto-arrange at startup on or off.
- Visibility & pins: the visibility, game, and controller settings described above, the wrist angle, where each screen is pinned (in the room, a wrist, or your head), and buttons to pin all screens or unpin them.
- Power: when the displays turn off while the headset isn't used, their state now, Turn displays off now (to try it), and Stay awake while plugged in. See [Displays off and sleep](#displays-off-and-sleep).
`layout/ft-layout` does the arranging. It's a Python script that uses only the standard library and runs on the host:
```
layout/ft-layout apply # arrange every screen
layout/ft-layout capture # save the current arrangement and sizes as the layout
layout/ft-layout save NAME # ...under a name too, and use it
layout/ft-layout use NAME # switch to a named layout and arrange the screens in it
layout/ft-layout layouts # list the named layouts (* = in use); rename OLD NEW, delete NAME
layout/ft-layout pin N|all left|right|head # pin as they are now; unpin N|all
layout/ft-layout plan # print the arrangement as JSON (no VR needed)
layout/ft-layout scale # per-screen scale, positions (as the screens are around you), and taskbar screen, to KWin
layout/ft-layout toggle # hide or show all screens
@@ -138,19 +151,50 @@ display-settings/install.sh # menu entries and the Meta+Shift+R and Meta+Shift+H
The layout is stored relative to your head when it's applied. `/tmp/frametop-layout.log` has the run from the last desktop start.
## Displays off and sleep
SteamVR turns the displays off a few seconds after the headset's proximity sensor says it came off. A stand or display mount that covers the sensor makes the headset seem worn, so its displays stay on, and Steam, which then counts someone as present, never puts it to sleep either.
`power/ft-powerd` goes by use instead. It runs in the `dev` container as `frametop-power.service` and starts with SteamVR. Once the headset has gone unused for `DISPLAY_OFF_MIN` minutes (0, the default, is never), it turns the displays' backlight off, and it turns it back on at the next use. Use is any of these:
- The headset, a Frame controller, or the 3D mouse's virtual controller moving more than `DISPLAY_MOVE_MM` (5 mm) or turning more than `DISPLAY_MOVE_DEG` (0.5 degrees) within 10 seconds.
- A key, button, or mouse motion on any input device on the host, including the headset's own buttons and the input relay's virtual mouse and keyboard.
- The headset going back on after SteamVR's own standby, or something else turning the backlight back on.
While SteamVR has the headset in standby, SteamVR owns the displays and ft-powerd waits. The backlight is `/sys/class/backlight/ae94000.dsi.0/brightness`, the same file SteamVR's driver writes for standby. With the backlight off, tracking and rendering keep running, which lets the displays wake the moment the headset moves, but the headset still uses most of its power. ft-powerd puts the backlight back when it stops, and if it was killed with the displays off, the next start does (the value is kept in `~/.cache/frametop/powerd-brightness` meanwhile).
Stay awake while plugged in is Steam's own setting, When Plugged In and Idle → Sleep after (`system_idle_suspend_ac_sec`), set to Never. Frametop Display Settings changes it the way Steam's Settings → Power page does, through Steam's UI on its debugging port (`display-settings/steam_settings.py`), and keeps the value from before in `STEAM_SLEEP_AC_BEFORE` to put back when the switch goes off. The power button still puts the Frame to sleep, and Steam's battery setting still applies.
```
power/build.sh && power/run.sh install
power/run.sh status # "ok on|off|away <seconds unused> <timeout seconds>"
power/run.sh off | on # the displays off now, or back on
power/run.sh log
```
## Hand tracking (experimental)
Your hands show over the screens: where a tracked hand is between an eye and a screen, ft-screens lets that eye see the room through the screen. The same tracker also detects pinches, for clicking where you look with the gaze pointer (not wired to the pointer yet). It's optional: `hands/run.sh install`, or the last step of `install.sh`.
- `ft-camd` borrows XRService's camera buffers and publishes the four IR tracking cameras to `/run/user/UID/frametop-hands/cam-ring`. It runs on the host as `frametop-camd.service`, with file capabilities that `hands/run.sh install` sets through sudo, and it drops them once set up. A rebuild clears them: `hands/run.sh caps`.
- `ft-hands` runs in the `dev` container as `frametop-hands.service`. It finds and triangulates the hands, and publishes `hands` (read by ft-screens' cutouts) and `gestures` (pinches) next to the ring.
- Both start and stop with SteamVR. `hands/run.sh status` and `hands/run.sh log` show how they're doing.
- Settings in `~/.config/frametop.conf`: `HANDS_SWAP_SIDES` (after some SteamVR restarts the side cameras' names come out swapped, and hands land beside the holes; `hands/tools/check_sides.py --ring` tells) and `HANDS_CPUS`.
Details, options, and the recording and replay tools are in [hands/README.md](../hands/README.md).
## Remote desktop over VNC
With `REMOTE=1` in the config (`desktops.sh remote on`), the desktop is also served over VNC, for RealVNC Viewer or macOS Screen Sharing. `desktops.sh remote info` prints the address and password.
With `REMOTE=1` in the config (`desktops.sh remote on`), the desktop's primary screen (the one with the taskbar) is also served over VNC, at that screen's resolution, for RealVNC Viewer or macOS Screen Sharing. `desktops.sh remote info` prints the address and password.
It listens on port 5900 on the Frame's Tailscale address only, not the LAN, so it needs Tailscale on the Frame ([deck-tailscale](https://github.com/tailscale-dev/deck-tailscale)). VNC authentication has no encryption of its own, so viewers warn about it, but the tailnet encrypts the traffic. The password is in `~/.config/frametop-remote/vnc-password` and VNC limits it to 8 characters. To change it, delete that folder and restart the desktop.
No VNC server can capture KWin on SteamOS directly: `krfb` needs `xdg-desktop-portal-kde`, which SteamOS doesn't ship, and `wayvnc` only works with wlroots compositors. So `session/remote-desktop.sh` captures the desktop with KDE's `krdpserver --plasma` on `127.0.0.1:3390`, and `session/vnc-bridge.sh` runs TigerVNC's `Xvnc` on display `:20` with a full-screen FreeRDP client inside it and serves that. Both run in the `dev` container, and the extra hop adds a little latency.
No VNC server can capture KWin on SteamOS directly: `krfb` needs `xdg-desktop-portal-kde`, which SteamOS doesn't ship, and `wayvnc` only works with wlroots compositors. So `session/remote-desktop.sh` captures the desktop with KDE's `krdpserver --plasma` on `127.0.0.1:3390`, and `session/vnc-bridge.sh` runs TigerVNC's `Xvnc` on display `:20` with a FreeRDP client inside it and serves that. Both run in the `dev` container, and the extra hop adds a little latency. krdp streams every screen; the VNC screen is the primary's size, and the FreeRDP window is shifted so the primary fills it (`ft-layout remote-view` gives the offset). krdp's own `--monitor` would stream just one screen, but it maps the pointer as if that screen sat at 0,0, so clicks would miss. When the layout changes, the VNC screen resizes and FreeRDP reconnects within a few seconds.
With remote access on, the nested KWin runs with `KWIN_WAYLAND_NO_PERMISSION_CHECKS=1`, so any app in the Frametop desktop could capture its screen or inject input. This applies only to that desktop, not the stock one. Port 3389 is SteamOS's own `xrdp`, which starts a separate X11 session rather than showing the VR desktop.
With remote access on, the nested KWin runs with `KWIN_WAYLAND_NO_PERMISSION_CHECKS=1` and `KWIN_SCREENSHOT_NO_PERMISSION_CHECKS=1`, so any app in the Frametop desktop could capture its screens or inject input. The second one lets scripts take screenshots through KWin's `org.kde.KWin.ScreenShot2` D-Bus interface. This applies only to that desktop, not the stock one. Port 3389 is SteamOS's own `xrdp`, which starts a separate X11 session rather than showing the VR desktop.
## Limits
- There's no way yet to pin a screen to your head like a HUD.
- A controller button can't show hidden screens; a mapped mouse or keyboard button can.
- KWin's cursor isn't drawn on the screens, because KWin draws it as a host cursor, which ft-screens doesn't render. The 3D mouse's dot and SteamVR's laser dot show where you're pointing.
- The old gamescope backend (`BACKEND=gamescope`) still works, but it gives every screen the same resolution, at most 1920×1080 pixels' worth, and arranging screens borrows the pointer for a few seconds.
+166
View File
@@ -0,0 +1,166 @@
// frametop-float: the KWin side of floating windows (see docs/floating-windows.md). ft-floatd
// loads it into the desktop's KWin over D-Bus (org.kde.kwin.Scripting) and talks to it:
// - events go to ft-floatd as JSON strings (org.frametop.Float.Event), for the windows it
// cares about: floating windows (the ones on a spare output, WL-<screens> and up), their
// popups and dialogs, new windows, and "Float in VR" requests;
// - commands come back through a long poll: the script calls NextCommand, ft-floatd
// answers when it has one (or after a while with nothing), and the script calls again.
// KWin scripts can call D-Bus but can't serve it, hence the poll. Window ids are KWin's
// internalId (a UUID string).
const SERVICE = "org.frametop.Float", PATH = "/Float", IFACE = "org.frametop.Float";
let screens = 0; // outputs WL-0 .. WL-<screens - 1> are screens; the rest are spares
let polling = false;
const watched = {}; // id -> true once its signals are connected
function send(ev) {
callDBus(SERVICE, PATH, IFACE, "Event", JSON.stringify(ev));
}
function outputIndex(o) {
const m = o ? /^WL-(\d+)$/.exec(o.name) : null;
return m ? parseInt(m[1]) : -1;
}
function isSpare(o) {
return screens > 0 && outputIndex(o) >= screens;
}
function rect(g) {
return {x: g.x, y: g.y, w: g.width, h: g.height};
}
function byId(id) {
const all = workspace.windowList();
for (let i = 0; i < all.length; ++i)
if (String(all[i].internalId) === id) return all[i];
return null;
}
function outputByName(name) {
const all = workspace.screens;
for (let i = 0; i < all.length; ++i)
if (all[i].name === name) return all[i];
return null;
}
function info(w) {
const o = w.output;
return {
id: String(w.internalId), pid: w.pid, cls: String(w.resourceClass), app: String(w.desktopFileName),
caption: String(w.caption), output: o ? o.name : "", outputRect: o ? rect(o.geometry) : null,
frame: rect(w.frameGeometry), client: rect(w.clientGeometry), popup: w.popupWindow,
transient: w.transient, parent: w.transientFor ? String(w.transientFor.internalId) : "",
normal: w.normalWindow, dialog: w.dialog, fullScreen: w.fullScreen, minimized: w.minimized,
onAllDesktops: w.onAllDesktops
};
}
function report(type, w) {
const ev = info(w);
ev.ev = type;
send(ev);
}
// Floating windows, and popups and dialogs on a spare output: tell ft-floatd about changes.
function watch(w) {
const id = String(w.internalId);
if (watched[id]) return;
watched[id] = true;
const onSpare = () => isSpare(w.output);
w.frameGeometryChanged.connect(() => { if (onSpare()) report("geometry", w); });
w.outputChanged.connect(() => report("output", w));
w.interactiveMoveResizeStarted.connect(() => {
if (onSpare()) send({ev: "move-start", id: id, move: w.move, resize: w.resize, frame: rect(w.frameGeometry)});
});
w.interactiveMoveResizeFinished.connect(() => { if (onSpare()) report("move-end", w); });
w.fullScreenChanged.connect(() => { if (onSpare()) report("fullscreen", w); });
w.minimizedChanged.connect(() => { if (onSpare()) report("minimized", w); });
w.maximizedChanged.connect(() => {
// A floating window stays an ordinary window: its output is its size plus a margin.
if (onSpare() && w.normalWindow && !w.fullScreen) w.setMaximize(false, false);
});
}
workspace.windowAdded.connect(w => {
watch(w);
report("added", w);
});
workspace.windowRemoved.connect(w => {
send({ev: "removed", id: String(w.internalId)});
delete watched[String(w.internalId)];
});
workspace.windowActivated.connect(w => {
if (w && isSpare(w.output)) send({ev: "activated", id: String(w.internalId)});
});
workspace.windowList().forEach(watch);
function requestFloat(w) {
if (!w || !w.normalWindow || w.popupWindow) return;
report(isSpare(w.output) ? "dock-request" : "float-request", w);
}
registerUserActionsMenu(w => {
if (!w.normalWindow || w.popupWindow) return null;
const floating = isSpare(w.output);
return {
text: floating ? "Back to Desktop" : "Float in VR",
icon: floating ? "window-restore" : "window-new",
triggered: () => requestFloat(w)
};
});
registerShortcut("Frametop Float Window", "Frametop: Float Window in VR (or put it back)", "Meta+Shift+F",
() => requestFloat(workspace.activeWindow));
function run(c) {
const w = c.id ? byId(c.id) : null;
switch (c.cmd) {
case "config":
screens = c.screens;
workspace.windowList().forEach(w => report("window", w));
break;
case "place": { // onto an output, at a frame rectangle (logical, global)
if (!w) break;
const o = outputByName(c.output);
if (!o) break;
if (w.fullScreen && !c.keepFullScreen) w.fullScreen = false;
w.setMaximize(false, false);
workspace.sendClientToScreen(w, o);
w.frameGeometry = {x: c.x, y: c.y, width: c.w, height: c.h};
if (c.onAllDesktops !== undefined) w.onAllDesktops = c.onAllDesktops;
break;
}
case "geometry":
if (w) w.frameGeometry = {x: c.x, y: c.y, width: c.w, height: c.h};
break;
case "close":
if (w) w.closeWindow();
break;
case "activate":
if (w) workspace.activeWindow = w;
break;
case "minimize":
if (w) w.minimized = c.on;
break;
case "info":
if (w) report("window", w);
break;
case "request-active": // ft-float float|dock active: like the shortcut
requestFloat(workspace.activeWindow);
break;
}
}
function poll() {
if (polling) return;
polling = true;
callDBus(SERVICE, PATH, IFACE, "NextCommand", reply => {
polling = false;
if (reply) {
try {
JSON.parse(reply).forEach(run);
} catch (e) {
print("frametop-float: bad command " + reply + ": " + e);
}
}
poll();
});
}
send({ev: "hello"});
poll();
Executable
+25
View File
@@ -0,0 +1,25 @@
#!/usr/bin/env python3
"""ft-float: talk to ft-floatd (floating windows in the Frametop desktop).
ft-float float [ID|active] float a window (default: the active one; for it, a toggle)
ft-float dock [ID|active] put a floating window back on the desktop
ft-float close ID close a window
ft-float list the spare outputs and what floats on them
FT_FLOAT_SOCKET names ft-floatd's socket (default frametop_float).
"""
import os
import socket
import sys
if len(sys.argv) < 2 or sys.argv[1] in ("-h", "--help"):
sys.exit(__doc__)
s = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
s.bind("")
s.settimeout(5)
try:
s.sendto(" ".join(sys.argv[1:]).encode(), "\0" + os.environ.get("FT_FLOAT_SOCKET", "frametop_float"))
reply = s.recv(8192).decode()
except OSError as e:
sys.exit(f"ft-floatd didn't answer ({e}); is the Frametop desktop running?")
print(reply)
sys.exit(0 if reply.startswith("ok") else 1)
+3
View File
@@ -0,0 +1,3 @@
#!/bin/bash
# ft-floatd on the Frame host, inside the Frametop desktop's session (see ft_floatd.py).
exec python3 "$(dirname "$(readlink -f "$0")")/ft_floatd.py" "$@"
+622
View File
@@ -0,0 +1,622 @@
#!/usr/bin/env python3
"""ft-floatd: floating windows for the Frametop desktop (see docs/floating-windows.md).
Runs inside the desktop's Plasma session (its D-Bus and Wayland). It keeps the table of
which window floats on which spare output and panel, and connects three parts:
- the KWin script frametop-float (float/frametop-float.js), which it loads into KWin. The
script sends events over D-Bus (org.frametop.Float.Event) and fetches commands with a
long poll (NextCommand).
- ft-screens, through its control socket (@ft_screens): the spare output's size, and the
floating panel's crop, density, place, and popups.
- kscreen-doctor, to turn spare outputs on and off and place them in KWin's layout.
Commands and ft-screens' events arrive as datagrams on @frametop_float (ft-float is the
command-line side). Replies go to the sender:
float [ID|active] dock [ID|active] close ID list quit (ft-float)
dock N | close N | resize N W H | scale N STEPS (ft-screens, N = its screen)
Spare outputs are WL-<screens> .. WL-<screens + slots - 1>. A floating window's output is its
frame plus a margin on each side (FLOAT_MARGIN pixels), so menus have room; the panel shows
only the window. Spares are placed apart from the screens and from each other in KWin's
layout, all within Xwayland's 32767-pixel limit.
Usage: ft-floatd [--screens N] [--slots N] [--margin PX] [--control NAME] [--socket NAME]
Defaults: FT_SCREEN_COUNT (from the session) or the layout's count, FLOAT_SLOTS and
FLOAT_MARGIN from ~/.config/frametop.conf (8 and 300), @ft_screens, @frametop_float.
"""
import argparse
import json
import math
import os
import re
import socket
import subprocess
import sys
import time
import dbus
import dbus.mainloop.glib
import dbus.service
from gi.repository import GLib
HERE = os.path.dirname(os.path.realpath(__file__))
sys.path.insert(0, os.path.join(HERE, "..", "layout"))
import ft_layout # noqa: E402 (config and layout)
SERVICE = IFACE = "org.frametop.Float"
PATH = "/Float"
SCRIPT = "frametop-float"
POLL_SECONDS = 20 # NextCommand answers empty after this (KWin's D-Bus timeout is 25 s)
SPARE_X, SPARE_CELL = 12000, 5000 # spares in KWin's layout: a grid from here, 4 across
PULL_OUT = 0.3 # a floated window starts this far in front of its screen (metres), clear of it
DEFAULT_MPP = 1.6 / 1920 # metres per pixel when ft-screens can't say (no SteamVR)
DEBUG = os.environ.get("FT_FLOAT_DEBUG") == "1" # log every event from the script
def whole(*scales):
"""The multiple a spare's size in pixels must be at these scales: KWin's nested backend gives
the buffer a whole buffer scale (1.2 -> 2), and a buffer that isn't a multiple of it is a
protocol error that disconnects KWin."""
k = 1
for s in scales:
k = math.lcm(k, max(1, math.ceil(s - 1e-6)))
return k
def kwin_size(px, scale, k):
"""What to ask ft-screens for so a spare comes out at least px pixels, a multiple of k. KWin
makes a nested output the size it's configured to times its scale, rounded. Returns (the
size to ask for, the pixels it comes out); the margin takes the extra pixels."""
n = max(1, math.floor(px / scale))
while True:
exact = n * scale
p = math.floor(exact + 0.5)
if p >= px and p % k == 0 and abs(exact - math.floor(exact) - 0.5) > 1e-6:
return n, p
n += 1
def log(*args):
print(time.strftime("%H:%M:%S"), *args, flush=True)
class Screens:
"""ft-screens' control socket."""
def __init__(self, name):
self.address = "\0" + name
self.sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
self.sock.bind("")
def ask(self, text, quiet=False):
self.sock.settimeout(2.0)
try:
self.sock.sendto(text.encode(), self.address)
reply = self.sock.recv(8192).decode()
except OSError as e:
reply = f"error no answer ({e})"
if not reply.startswith("ok") and not quiet:
log(f"ft-screens: {text}: {reply}")
return reply
def kscreen(*args):
try:
r = subprocess.run(["kscreen-doctor", *args], capture_output=True, text=True, timeout=20)
return r.stdout
except (OSError, subprocess.TimeoutExpired) as e:
log(f"kscreen-doctor: {e}")
return ""
def output_scales():
try:
data = json.loads(kscreen("-j") or "{}")
except ValueError:
return {}
return {o["name"]: float(o.get("scale", 1)) for o in data.get("outputs", []) if o.get("name")}
class Float:
"""A floating window."""
def __init__(self, wid, slot, saved):
self.id = wid
self.slot = slot # Slot
self.saved = saved # where it came from: output, frame, onAllDesktops
self.scale = 1.0 # its output's scale (pixels per logical unit)
self.mpp = DEFAULT_MPP # metres per pixel on its panel
self.frame = None # last frame (logical, global)
self.client = None
self.full = False # full screen: no margin
self.normal = None # its size in pixels when last not full screen
self.unfull_until = 0.0 # left full screen just now (see follow)
self.move_from = None # frame when a title-bar move started (put back after)
self.subs = {} # popup or dialog id -> number on the panel
class Slot:
def __init__(self, k, screens):
self.k = k
self.output = f"WL-{screens + k}"
self.index = screens + k + 1 # ft-screens' number (1-based)
self.pos = (SPARE_X + (k % 4) * SPARE_CELL, (k // 4) * SPARE_CELL)
self.size = None # its output's size in pixels, as last set
self.want = None # the size in pixels asked for (size is at least that)
self.kscale = 1.0 # its output's scale in KWin, as last set
self.window = None # Float
class Daemon:
def __init__(self, args):
self.screens_n = args.screens
self.margin = args.margin
self.slots = [Slot(k, args.screens) for k in range(args.slots)]
self.floats = {} # window id -> Float
self.windows = {} # window id -> last info from the script
self.pending = [] # commands for the script
self.waiter = None # (reply callback, timeout source) while the script waits
self.screens = Screens(args.control)
self.sub_numbers = {} # popup/dialog id -> (window id, number)
self.next_sub = 1
# ------------------------------------------------------------ the script
def command(self, **cmd):
self.pending.append(cmd)
self.flush()
def flush(self):
if not self.waiter or not self.pending:
return
reply, source = self.waiter
self.waiter = None
GLib.source_remove(source)
text, self.pending = json.dumps(self.pending), []
reply(text)
def wait(self, reply):
if self.waiter: # a stale poll (the script reloaded): let it go
old, source = self.waiter
GLib.source_remove(source)
old("")
def timeout():
if self.waiter and self.waiter[0] is reply:
self.waiter = None
reply("")
return False
self.waiter = (reply, GLib.timeout_add_seconds(POLL_SECONDS, timeout))
self.flush()
def load_script(self, bus):
kwin = dbus.Interface(bus.get_object("org.kde.KWin", "/Scripting"), "org.kde.kwin.Scripting")
if kwin.isScriptLoaded(SCRIPT):
kwin.unloadScript(SCRIPT)
sid = int(kwin.loadScript(os.path.join(HERE, "frametop-float.js"), SCRIPT, signature="ss"))
if sid < 0:
raise RuntimeError("KWin didn't load the script")
bus.get_object("org.kde.KWin", f"/Scripting/Script{sid}").run(dbus_interface="org.kde.kwin.Script")
log(f"script loaded ({sid})")
# ------------------------------------------------------------ events from the script
def on_event(self, ev):
kind = ev.get("ev")
if DEBUG:
log("event", {k: v for k, v in ev.items() if k in ("ev", "id", "output", "frame", "fullScreen", "popup",
"parent", "move")})
wid = ev.get("id", "")
if kind == "hello":
self.command(cmd="config", screens=self.screens_n)
# The script reports every window after "config"; spares nothing floats on are off.
GLib.timeout_add(1500, self.disable_unused)
return
if kind == "removed":
self.windows.pop(wid, None)
if wid in self.floats:
log(f"{wid[:9]} closed")
self.release(self.floats.pop(wid))
self.drop_sub(wid)
return
if wid:
self.windows[wid] = ev
f = self.floats.get(wid)
if kind == "float-request":
self.float_window(ev)
elif kind == "dock-request":
if f:
self.dock(f)
elif kind in ("added", "window", "output"):
self.seen(ev)
elif kind == "geometry":
if f:
self.follow(f, ev)
else:
self.sub(ev)
elif kind == "move-start" and f and ev.get("move"):
f.move_from = ev["frame"]
self.screens.ask(f"carry {f.slot.index}", quiet=True)
elif kind == "move-end" and f and f.move_from:
# The panel carried the window; KWin may have slipped it a few pixels first.
m, f.move_from = f.move_from, None
if (m["x"], m["y"]) != (ev["frame"]["x"], ev["frame"]["y"]):
self.command(cmd="geometry", id=f.id, x=m["x"], y=m["y"], w=ev["frame"]["w"], h=ev["frame"]["h"])
elif kind == "fullscreen" and f:
if not ev.get("fullScreen"):
f.unfull_until = time.monotonic() + 1.0
self.follow(f, ev)
elif kind == "minimized" and f:
self.screens.ask(f"minimized {f.slot.index} {1 if ev.get('minimized') else 0}", quiet=True)
def spare_slot(self, output):
for s in self.slots:
if s.output == output:
return s
return None
def seen(self, ev):
"""A window the script told us about: is it somewhere it shouldn't be?"""
wid = ev["id"]
slot = self.spare_slot(ev.get("output", ""))
f = self.floats.get(wid)
if f and f.slot is not slot and ev["ev"] == "output":
# Left its spare (KWin moved it, or docking): it's not floating any more.
if slot is None:
log(f"{wid[:9]} left its floating output")
del self.floats[wid]
self.release(f)
return
if slot is None or f:
return
if ev.get("popup") or (ev.get("transient") and ev.get("parent") in self.floats):
self.sub(ev)
return
if not ev.get("normal"):
return
if slot.window is None:
if ev.get("cls") == "ksplashqml": # the login splash, on every output at first
return
# Floating when ft-floatd (re)started: take it over where it is.
f = Float(wid, slot, None)
slot.window = f
self.floats[wid] = f
f.scale = slot.kscale = output_scales().get(slot.output, 1.0)
log(f"{wid[:9]} ({ev.get('cls')}) already floats on {slot.output}")
self.follow(f, ev)
return
# A new window that opened on a floating window's output: windows of floating apps
# float too; anything else goes to the screens.
if ev["ev"] == "added" and any(o.saved is not None and self.windows.get(o.id, {}).get("pid") == ev.get("pid")
for o in self.floats.values()):
self.float_window(ev)
else:
self.command(cmd="place", id=wid, output="WL-0", x=ev["frame"]["x"] % 400 + 100,
y=ev["frame"]["y"] % 300 + 100, w=ev["frame"]["w"], h=ev["frame"]["h"])
# ------------------------------------------------------------ floating and docking
def sized(self, slot, size, timeout=1.0):
"""Wait (briefly) until KWin has taken the spare's new size, before a scale that needs it."""
end = time.monotonic() + timeout
while time.monotonic() < end:
for line in self.screens.ask("toplevels", quiet=True).splitlines()[1:]:
f = line.split()
if len(f) >= 2 and f[0] == str(slot.index) and f[1] == f"{size[0]}x{size[1]}":
return True
time.sleep(0.03)
log(f"{slot.output} didn't take {size[0]}x{size[1]} in time")
return False
def set_size(self, slot, want, k=None):
"""Size a spare's output to at least want (pixels) at its current scale in KWin."""
if want == slot.want:
return
k = k or whole(slot.kscale)
(w, pw), (h, ph) = kwin_size(want[0], slot.kscale, k), kwin_size(want[1], slot.kscale, k)
slot.want, slot.size = want, (pw, ph)
self.screens.ask(f"size {slot.index} {w} {h}")
def disable_unused(self):
off = [f"output.{s.output}.disable" for s in self.slots if s.window is None]
if off:
kscreen(*off)
return False
def free_slot(self):
for s in self.slots:
if s.window is None:
return s
return None
def screen_mpp(self, output):
"""Metres per pixel on the screen showing this output, from ft-screens."""
m = re.match(r"WL-(\d+)$", output or "")
reply = self.screens.ask("screens", quiet=True)
if m and reply.startswith("ok"):
for part in reply.split()[2:]:
idx, size, metres = part.split(":")
if int(idx) == int(m.group(1)) + 1:
return float(metres) / max(1, int(size.split("x")[0]))
return DEFAULT_MPP
def float_window(self, ev):
wid = ev["id"]
if wid in self.floats:
return
slot = self.free_slot()
if slot is None:
self.notify(f"All {len(self.slots)} floating windows are in use. Put one back on the desktop "
"to float another.")
return
f = Float(wid, slot, {"output": ev["output"], "frame": ev["frame"], "onAllDesktops": ev.get("onAllDesktops")})
slot.window = f
self.floats[wid] = f
scales = output_scales()
f.scale = scales.get(ev["output"], 1.0)
slot.kscale = scales.get(slot.output, slot.kscale)
f.mpp = self.screen_mpp(ev["output"])
fr, s, m = ev["frame"], f.scale, self.margin
w, h = round(fr["w"] * s), round(fr["h"] * s)
log(f"{wid[:9]} ({ev.get('cls')}) floats on {slot.output}: {w}x{h} px, scale {s:g}")
# The spare's size first (while it's off, so its first frame is right; it's sized at
# its old scale, for pixels that suit the new one), then its panel, then turn it on,
# then the window.
slot.want = None
self.set_size(slot, (w + 2 * m, h + 2 * m), whole(slot.kscale, s))
self.screens.ask(f"scale {slot.index} {s:g}") # for pointer positions (KWin's units)
self.set_panel(f, (m, m, w, h), title=round((ev["client"]["y"] - fr["y"]) * s))
self.place_panel(f, ev)
kscreen(f"output.{slot.output}.enable", f"output.{slot.output}.scale.{s:g}",
f"output.{slot.output}.position.{slot.pos[0]},{slot.pos[1]}")
self.rescaled(slot, s)
self.command(cmd="place", id=wid, output=slot.output, x=slot.pos[0] + m / s, y=slot.pos[1] + m / s,
w=fr["w"], h=fr["h"], onAllDesktops=True)
def set_panel(self, f, crop, title=0):
x, y, w, h = crop
self.screens.ask(f"float {f.slot.index} {f.mpp:.7f} {x} {y} {w} {h} {title}")
def place_panel(self, f, ev):
"""Put the panel where the window was on its screen, a little in front of it."""
m = re.match(r"WL-(\d+)$", ev["output"])
reply = self.screens.ask(f"get {int(m.group(1)) + 1}", quiet=True) if m else ""
if not reply.startswith("ok"):
return
g = ft_layout.parse_get(reply)
c, xa, ya, za = g["center"], g["x"], g["y"], g["z"]
out, fr, s = ev["outputRect"], ev["frame"], f.scale
# The window's centre relative to the screen's, in panel pixels, then metres.
dx = ((fr["x"] - out["x"]) + fr["w"] / 2 - out["w"] / 2) * s * f.mpp
dy = ((fr["y"] - out["y"]) + fr["h"] / 2 - out["h"] / 2) * s * f.mpp
p = [c[k] + xa[k] * dx - ya[k] * dy + za[k] * PULL_OUT for k in range(3)]
rows = [f"{xa[k]:.5f} {ya[k]:.5f} {za[k]:.5f} {p[k]:.4f}" for k in range(3)]
self.screens.ask(f"pose {f.slot.index} {' '.join(rows)}")
def follow(self, f, ev):
"""The window moved or resized on its output: crop the panel to it, and keep the
output its size plus the margin. Full screen: no margin, and the output keeps the
window's size from before, so the window fills its own panel."""
slot, s = f.slot, f.scale
fr, cl, out = ev["frame"], ev["client"], ev["outputRect"]
f.frame, f.client = fr, cl
# KWin sizes a window to its output before it reports it full screen; with a margin, a
# window that fills its output is going full screen. (Not just after it left full
# screen: then it fills the output until the output grows back.)
fills = self.margin > 0 and (fr["x"], fr["y"], fr["w"], fr["h"]) == (out["x"], out["y"], out["w"], out["h"])
full = bool(ev.get("fullScreen")) or (fills and time.monotonic() > f.unfull_until)
f.full = full
w, h = round(fr["w"] * s), round(fr["h"] * s)
if not full:
f.normal = (w, h)
m = 0 if full else self.margin
self.set_size(slot, f.normal if full and f.normal else (w + 2 * m, h + 2 * m))
if not full:
x0, y0 = slot.pos[0] + m / s, slot.pos[1] + m / s
if abs(fr["x"] - x0) > 0.5 or abs(fr["y"] - y0) > 0.5:
self.command(cmd="geometry", id=f.id, x=x0, y=y0, w=fr["w"], h=fr["h"])
return # the next geometry event crops the panel
x, y = round((fr["x"] - out["x"]) * s), round((fr["y"] - out["y"]) * s)
self.set_panel(f, (x, y, w, h), title=0 if full else round((cl["y"] - fr["y"]) * s))
def rescale(self, f, steps):
"""Meta+scroll: the window's content bigger or smaller, at the same size in pixels, so its
panel stays the same size (KWin's output scale, in steps of 10%)."""
if not f.frame or f.full or steps == 0:
return
s = min(3.0, max(0.5, round(f.scale * 1.1 ** steps * 20) / 20))
if s == f.scale:
return
w, h = round(f.frame["w"] * f.scale), round(f.frame["h"] * f.scale)
slot, m = f.slot, self.margin
log(f"{f.id[:9]} scale {f.scale:g} -> {s:g}")
f.scale = s
# The output's size in pixels must suit the new scale before KWin draws at it (see whole).
k = whole(slot.kscale, s)
if slot.size[0] % k or slot.size[1] % k:
slot.want = None
self.set_size(slot, slot.size, k)
self.sized(slot, slot.size)
kscreen(f"output.{slot.output}.scale.{s:g}")
self.screens.ask(f"scale {slot.index} {s:g}")
self.rescaled(slot, s)
self.command(cmd="geometry", id=f.id, x=slot.pos[0] + m / s, y=slot.pos[1] + m / s, w=w / s, h=h / s)
def rescaled(self, slot, s):
"""KWin has the spare at scale s now: ask for its size again in the new scale's terms,
or the next configure (any size, or KWin's own) would make it the old size times s."""
if s == slot.kscale:
return
slot.kscale = s
want, slot.want = slot.size, None
self.set_size(slot, want)
def dock(self, f, frame=None):
"""Back where it came from (or onto screen 1 if we don't know)."""
saved = f.saved or {"output": "WL-0", "frame": dict(f.frame or {"x": 100, "y": 100, "w": 800, "h": 600}),
"onAllDesktops": False}
fr = frame or saved["frame"]
log(f"{f.id[:9]} back to {saved['output']}")
self.command(cmd="place", id=f.id, output=saved["output"], x=fr["x"], y=fr["y"], w=fr["w"], h=fr["h"],
onAllDesktops=bool(saved.get("onAllDesktops")))
def release(self, f):
"""Its window left: hide the panel and turn the spare off."""
slot = f.slot
if slot.window is f:
slot.window = None
slot.want = None
for sub_id in list(f.subs):
self.drop_sub(sub_id)
self.screens.ask(f"unfloat {slot.index}", quiet=True)
kscreen(f"output.{slot.output}.disable")
# ------------------------------------------------------------ popups and dialogs
def sub(self, ev):
parent = self.floats.get(ev.get("parent", ""))
if parent is None:
# A popup of a popup: its top-level parent is the floating window.
known = self.sub_numbers.get(ev.get("parent", ""))
parent = self.floats.get(known[0]) if known else None
if parent is None:
return
wid = ev["id"]
if wid not in self.sub_numbers:
self.sub_numbers[wid] = (parent.id, self.next_sub)
parent.subs[wid] = self.next_sub
self.next_sub += 1
n = self.sub_numbers[wid][1]
fr, out, s = ev["frame"], ev["outputRect"], parent.scale
x, y = round((fr["x"] - out["x"]) * s), round((fr["y"] - out["y"]) * s)
self.screens.ask(f"sub {parent.slot.index} {n} {x} {y} {round(fr['w'] * s)} {round(fr['h'] * s)}", quiet=True)
def drop_sub(self, wid):
known = self.sub_numbers.pop(wid, None)
if not known:
return
parent = self.floats.get(known[0])
if parent:
parent.subs.pop(wid, None)
self.screens.ask(f"sub {parent.slot.index} {known[1]} off", quiet=True)
# ------------------------------------------------------------ requests on @frametop_float
def by_panel(self, index):
for s in self.slots:
if s.index == index:
return s.window
return None
def request(self, text):
words = text.split()
if not words:
return "error empty"
cmd, rest = words[0], words[1:]
if cmd == "list":
return "ok " + " ".join(f"{s.output}:{s.window.id if s.window else '-'}" for s in self.slots)
if cmd == "quit":
GLib.idle_add(self.loop.quit)
return "ok"
if cmd in ("float", "dock") and (not rest or rest[0] == "active"):
self.command(cmd="request-active")
return "ok"
if cmd in ("dock", "close", "resize", "scale") and rest and rest[0].isdigit():
f = self.by_panel(int(rest[0]))
if not f:
return f"error no floating window on screen {rest[0]}"
if cmd == "dock":
self.dock(f)
elif cmd == "close":
self.command(cmd="close", id=f.id)
elif cmd == "scale" and len(rest) == 2:
self.rescale(f, int(rest[1]))
elif cmd == "resize" and len(rest) == 3 and f.frame:
w, h = max(320, int(rest[1])), max(200, int(rest[2]))
self.command(cmd="geometry", id=f.id, x=f.frame["x"], y=f.frame["y"], w=w / f.scale, h=h / f.scale)
return "ok"
if cmd in ("float", "dock", "close") and rest:
ev = self.windows.get(rest[0])
if not ev:
return f"error no window {rest[0]}"
if cmd == "float":
self.float_window(ev)
elif cmd == "dock" and rest[0] in self.floats:
self.dock(self.floats[rest[0]])
elif cmd == "close":
self.command(cmd="close", id=rest[0])
return "ok"
return "error unknown command"
def notify(self, text):
log(text)
try:
n = dbus.Interface(dbus.SessionBus().get_object("org.freedesktop.Notifications",
"/org/freedesktop/Notifications"),
"org.freedesktop.Notifications")
n.Notify("Frametop", 0, "window-new", "Floating windows", text, [], {}, 5000)
except dbus.DBusException as e:
log(f"notification: {e.get_dbus_message()}")
class Service(dbus.service.Object):
def __init__(self, bus, daemon):
super().__init__(dbus.service.BusName(SERVICE, bus), PATH)
self.daemon = daemon
@dbus.service.method(IFACE, in_signature="s", out_signature="")
def Event(self, text):
try:
self.daemon.on_event(json.loads(text))
except (ValueError, KeyError, TypeError) as e:
log(f"bad event {text[:200]}: {e!r}")
@dbus.service.method(IFACE, in_signature="", out_signature="s", async_callbacks=("reply", "error"))
def NextCommand(self, reply, error):
self.daemon.wait(reply)
def main():
conf = ft_layout.read_conf()
p = argparse.ArgumentParser(description="Floating windows for the Frametop desktop")
p.add_argument("--screens", type=int, default=int(os.environ.get("FT_SCREEN_COUNT") or 0))
p.add_argument("--slots", type=int, default=int(conf.get("FLOAT_SLOTS") or 8))
p.add_argument("--margin", type=int, default=int(conf.get("FLOAT_MARGIN") or 300))
p.add_argument("--control", default="ft_screens")
p.add_argument("--socket", default="frametop_float")
args = p.parse_args()
if args.screens <= 0:
args.screens = ft_layout.screen_count()
args.slots = max(0, min(16, args.slots))
args.margin = max(0, min(1000, args.margin))
dbus.mainloop.glib.DBusGMainLoop(set_as_default=True)
bus = dbus.SessionBus()
daemon = Daemon(args)
service = Service(bus, daemon) # noqa: F841 (keeps the name)
daemon.loop = GLib.MainLoop()
sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
sock.bind("\0" + args.socket)
sock.setblocking(False)
def readable(*_):
while True:
try:
data, sender = sock.recvfrom(4096)
except BlockingIOError:
return True
reply = daemon.request(data.decode(errors="replace").strip())
if sender:
try:
sock.sendto(reply.encode(), sender)
except OSError:
pass
GLib.io_add_watch(sock.fileno(), GLib.IO_IN, readable)
log(f"{args.screens} screens, {args.slots} floating slots (WL-{args.screens} and up), margin {args.margin} px")
daemon.load_script(bus)
daemon.loop.run()
if __name__ == "__main__":
main()
+36 -5
View File
@@ -4,6 +4,7 @@ The Steam Frame's eye tracking as pointer input: a gaze mode for the 3D mouse (t
- `ft-gaze` (C++, OpenVR, runs in the dev container) reads the eye tracker and prints one JSON line per sample (90 Hz). For each source, it gives the gaze direction relative to the head and the Frametop screen pixel it lands on.
- `gazecal.py` has what the probe and the gaze service share: the correction models, filters, and the reader for SteamVR's eye tracking log.
- `tracker/` is our own eye tracker, an alternative to SteamVR's: `ft-eyes` finds the pupils and glints in the eye-camera frames that `ft-eyegrab` (a small root service) copies out of SteamVR's tracker. See "Our own eye tracker" below.
- `probe/ft-gazeprobe` (GTK 4, host Python) is a fullscreen playground. It runs ft-gaze, draws where you're looking, measures accuracy, and tries out hold-to-adjust clicking with a calibration that learns from your adjustments.
```
@@ -12,15 +13,22 @@ gaze/probe/install.sh # build, and add Frametop Gaze Probe to the app me
gaze/probe/ft-gazeprobe --screen 1
gaze/run.sh install # the gaze service, with SteamVR
gaze/ft-gazectl on # the pointer follows your gaze (off: the mouse alone)
gaze/tracker/install.sh # our own eye tracker's frame grabber (asks for sudo)
```
## Gaze pointer
Gaze as an input method for the whole desktop, without replacing anything of SteamVR's:
- `ft-gazed` (host Python, a user service: `gaze/run.sh install`) runs ft-gaze and corrects its gaze. It uses SteamVR's combined gaze (mmap set 1), which keeps going when the tracker loses one eye; set 2's combined direction is off by half of whatever the lost eye reads, and the tracker does lose an eye for minutes at a time. It drops blinks (both eyes closing), smooths with a fixation lock, and applies the calibration from the probe (`calibration.json`, reloaded when the probe changes it) plus what the pointer has learned since (`pointer-lessons.json`). It sends the result to the pointer helper 90 times a second. It follows SteamVR's eye tracking log, and when the headset goes back on (SteamVR starts its eye model over, and the error moves), older lessons count less, so the first few after relearn the offset.
- `ft-gazed` (host Python, a user service: `gaze/run.sh install`) runs ft-gaze and corrects its gaze. Two settings on the Gaze page of Frametop Input Settings (`GAZE_TRACKER` and `GAZE_EYE` in `~/.config/frametop.conf`, read again when the file changes) pick whose eye tracking it uses and how it weights the eyes:
- **Eye tracker:** SteamVR's (the default), or our own (Own tracker: see "Our own eye tracker" below). The gaze service runs ours while this is on. It keeps its own calibration: calibrate it in the probe with the tracker toggle on Own tracker. The gaze pointer's settings (hand back, nudges, hold to drag, the dot) are the pointer helper's, so they're the same with either.
- **Eye bias:** Auto, Left, or Right. The gaze combines both eyes, each calibrated on its own, because their errors partly cancel: on 306 clicks with our tracker, the eyes' sideways errors were correlated -0.37, and both together were 0.65 degrees off (median) against 0.96 for the left eye alone and 1.11 for the right. So Left or Right leans instead of choosing: that eye counts twice as much as the other (0.03 degrees worse there toward the better eye, 0.13 toward the worse). Auto weights each eye by the inverse square of how far off it was at your last 20 nudges, once each eye has 5, and evenly before that. Each eye's miss is measured before that nudge teaches anything, so each is a fresh test. The calibration's own fit isn't used for this: on SteamVR's test of 2026-09-29, the calibration dots said the left eye was the better one, and new spots said the right. Either eye carries the gaze alone while the other is closed or lost.
With SteamVR, each eye is its own reading (set 2), corrected by its calibration from the probe (the Left eye and Right eye sources) plus what the pointer has taught that eye since. On that test, the two eyes each calibrated and averaged were 1.70 degrees off (median; mean 1.62) against 1.72 (mean 1.84) for SteamVR's combined gaze with its calibration. A calibration from before the probe had the eyes as sources, or `--source`, uses the older path. That path runs on SteamVR's combined gaze (mmap set 1), corrected as a whole. When the tracker loses one eye (its variance for that eye jumps from about 0.001 to 0.02), the gaze comes from the other eye instead: that eye's own reading (set 2) plus what it usually reads against the combined gaze, learned while both eyes are seen, in 10 degree cells of where it looks. Set 1 keeps going on one eye too, but it holds the lost eye's yaw where it was, so the gaze moves half as far sideways as your eyes do. On a recording, one eye alone came out a median 0.8 degrees from both eyes' gaze over a steady look, a little more jittery.
Looks down past the screens (under 20 degrees down, on no Frametop screen: a glance at the keyboard) aren't sent, so the pointer stays where it was instead of following you down, and eyes lost there aren't counted. It drops blinks (both eyes closing or lost), smooths with a fixation lock, and sends the result to the pointer helper 90 times a second. It follows SteamVR's eye tracking log, and when the headset goes back on (SteamVR starts its eye model over, and the error moves), older lessons count less, so the first few after relearn the offset.
- The pointer helper's **gaze mode** (off by default: the Gaze page of Frametop Input Settings, `gaze/ft-gazectl on`, `POINTER_GAZE=1` in `~/.config/frametop.conf`, or a mouse or controller button mapped to "Gaze pointer on/off") works like MAGIC pointing (Zhai et al., 1999). The pointer goes where you look. Move the mouse and it's the mouse's, from where the gaze put it, for the last bit. Look well away (5 degrees) and the gaze takes it back. A press isn't sent at once: the pointer stops where the gaze put it, and if that's wrong, drag it onto what you meant with the button still held; the click happens where you let go. To drag something, hold the press still for half a second first (`POINTER_GAZE_HOLD`), then move. Outside games the pointer stays on while gaze mode is on. The dot only shows while the mouse moves it, while a press is held, and as a pulse when you click.
- **Learning from nudges:** if the mouse took the pointer from the gaze and moved it a little (0.2 to 8 degrees) before you clicked, or you dragged a held press that far, you were nudging it onto what you looked at. The helper sends that as a lesson, from the raw gaze when the mouse took over to where you clicked, and ft-gazed learns it. So using it is what calibrates it. One lesson moves the whole correction by only a third of what it measured (more near where it was taken), since in the first live test one 6 degree lesson moved everything and put the next target 7 degrees off. `ft-gazectl status` shows the lessons, and `ft-gazectl forget` drops them.
- **Learning from nudges:** if the mouse took the pointer from the gaze and moved it a little (0.2 to 8 degrees) before you clicked, or you dragged a held press that far, you were nudging it onto what you looked at. The helper sends that as a lesson, from the raw gaze when the mouse took over to where you clicked, and ft-gazed learns it. So using it is what calibrates it. The raw gaze is one ft-gazed sent, so it also finds when that look was, and what each eye read then. With SteamVR, each eye learns its own error. With our tracker, the look goes to it as a click, like the probe's, and it relearns how the headset sits on your face. After the headset was off, your first nudge and click there resets that (the probe's one-dot check does the same). The helper only sends nudges up to `POINTER_GAZE_NUDGE_MAX` (8 degrees), and right after putting the headset back on our tracker can be further off than that. If so, raise it for a moment, or do the probe's check. One lesson moves the whole correction by only a third of what it measured (more near where it was taken), since in the first live test one 6 degree lesson moved everything and put the next target 7 degrees off. `ft-gazectl status` shows the lessons, and `ft-gazectl forget` drops them.
- Nothing writes to SteamVR, its eye tracker, or its files: ft-gaze maps the eye tracker's shared memory read-only. With no fresh gaze (a blink, the service stopped, the headset off), the pointer stays where it is, and the mouse works as always.
Lessons are logged to `pointer-lessons.jsonl`: the raw gaze, the true direction, the correction at the time, and how far off it was.
@@ -30,13 +38,36 @@ Lessons are logged to `pointer-lessons.jsonl`: the raw gaze, the true direction,
| Source | Where it comes from |
| --- | --- |
| SteamVR action | An `eyetracking` action bound to `/user/head/eyetracking` (`actions/`), read with `IVRInput::GetEyeTrackingDataRelativeToNow`. This is the supported way. |
| mmap set 1, set 2 | `/dev/shm/eye-server.mmap`, which SteamVR's eyetracking process writes for the HMD driver (`driver_cv.so`). It has two sets of per-eye directions in head space. The layout is undocumented (offsets are in `ft-gaze.cpp`) and may change with a SteamVR update. ft-gaze maps it read-only; the file also carries calibration clicks to the tracker and must never be written. |
| mmap set 1, set 2 | `/dev/shm/eye-server.mmap`, which SteamVR's eyetracking process writes for the HMD driver (`driver_cv.so`). It has two sets of per-eye directions in head space: set 1 is filtered, and its two eyes always share one pitch; set 2 is each eye's own reading. After each set come the tracker's variances for each eye, and at the end each eye's raw measurement and its variance (the tracker's confidence in that frame), which ft-gaze passes on for the fit check. |
| Left eye, right eye | Each eye alone, from set 2: calibrate and test them to see what one eye is worth against both. The layout is undocumented (offsets are in `ft-gaze.cpp`) and may change with a SteamVR update. ft-gaze maps it read-only; the file also carries calibration clicks to the tracker and must never be written. |
| Own tracker | Our own tracker (`tracker/`, experimental; see "Our own eye tracker"). It keeps its own calibration, not SteamVR's: the probe's calibration with the tracker toggle on Own tracker fits it (its dots go out to the Calibration ring angle each way, on an oval, since the fit goes wrong past its dots), and practice clicks teach it how far the headset has moved on your face since. After the headset was off, the probe first asks for one look at a centre dot, which resets that. With Own tracker on, the probe hides SteamVR's gaze and draws a red dot where each eye alone puts it, and asks the gaze service to keep the tracker running. The gaze pointer can use it too (Eye tracker: Own tracker, on the Gaze page of Frametop Input Settings). ft-gaze reports it as `own` while it's running, and as `{"ok":0}` otherwise. |
The tracker stops when the headset is off your head. SteamVR also calibrates gaze on its own from laser-mouse clicks, treating each click as a spot you were looking at. That includes mouse clicks through the Frametop pointer, so a click where the pointer's dot isn't what you're looking at teaches SteamVR a wrong sample (it only takes clicks within 5 degrees of your gaze). In the probe, use Enter or Space as the trigger: keys don't go through SteamVR's laser. See `Accept usercal` in `~/.local/share/Steam/logs/eyetracking.txt`.
The tracker stops when the headset is off your head. SteamVR also calibrates gaze on its own from laser-mouse clicks, treating each click as a spot you were looking at. That includes mouse clicks through the Frametop pointer, so a click where the pointer's dot isn't what you're looking at teaches SteamVR a wrong sample (it only takes clicks within 5 degrees of your gaze). In the probe, use Enter or Space as the trigger: keys don't go through SteamVR's laser. See `Accept usercal` in `~/.local/share/Steam/logs/eyetracking.txt`. When the tracker loses an eye, the same log says `CEyePoseUKF L: Large dt` (or `R`) as it starts that eye over.
## Our own eye tracker
`gaze/tracker/` is an eye tracker of our own, because SteamVR's is about 1.5 degrees off after the best correction the gaze service can learn, and what's left is mostly look-to-look noise that no correction on top of its output can remove. Ours processes the eye cameras itself: 0.59 degrees (median) in its best live session against 0.83 for SteamVR's with the probe's correction, and after the headset was taken off and put back without recalibrating, 0.58 once your first clicks had taught it where the headset sat (`tracker/findings.md` has the measurements).
- `ft-eyegrab` (C, root, the system service `frametop-eyegrab.service`) copies the eye-camera frames (512x400, 90 fps per eye) out of the DMA-BUFs SteamVR's `eyetracking` process holds into `/dev/shm/frametop-eyes-cams`, owned by you. It maps them read-only, and it only copies while someone touches `/dev/shm/frametop-eyes-want` (ft-eyes and the recorder do, every second). Otherwise it holds none of the tracker's buffers. Its unit keeps only the capabilities that needs (`CAP_SYS_PTRACE`, `CAP_DAC_READ_SEARCH`, `CAP_CHOWN`). `gaze/tracker/install.sh` builds it and installs it to `/etc/frametop` with sudo, which it asks for (`uninstall`, `status`, and `log` too).
- `ft-eyes` (Python with numpy and OpenCV, in the dev container: `gaze/tracker/build.sh` puts the pinned `requirements.txt` in `gaze/tracker/build/venv`) finds each eye's pupil (dark threshold, closing, ellipse fit) and glint pair (`eyes_pupil.py`), and maps them to a gaze with a quadratic fit per eye (`eyes_model.py`). It follows the headset moving on your face with a per-eye shift, which your clicks teach, and uses the glints only to notice a sudden jump. It publishes the gaze in `/dev/shm/frametop-eyes-gaze` (ft-gaze's source `own`) and takes calibration dots and clicks on `@ft_eyes`. The gaze service runs it while Eye tracker is Own tracker, or while the probe uses it. State (the calibration, each eye's shift, the clicks) is in `~/.local/state/frametop/gaze/eyes/`.
- `lab/` has the tools for improving it on recordings. `ft-eyes-record NAME` (or `ft-eyes-session`, with SteamVR's gaze alongside) records the cameras. `ft-eyes-score` fits and scores on recordings against the probe's practice clicks. `ft-eyes-e2e` runs the whole live path on two recordings (calibrate on one, click through the other). `ft-eyes-replay` plays a recording into a scratch share. Heavy ones run on a PC through `frame-job` (`gaze/tracker/.frame-job`). `lab/py` runs them with that Python (in the dev container on the Frame; frame-job's setup makes the same venv on the PC).
Ground rules, for anyone changing it:
- **Clean room.** Nothing of Valve's goes in: we don't decompile, disassemble, or patch the `eyetracking` binary or its network weights, and we don't copy their code or weights. Its public output (eye-server.mmap, read-only) is fair game as a baseline and as labels, and so are published papers and openly licensed pupil detectors (check each one's license: PuRe, PuReST, ElSe, and ExCuSe are non-commercial only).
- **Root only reads.** ft-eyegrab never writes to, stops, or signals the `eyetracking` process, vrserver, or vrcompositor, never opens `/dev/adsp`, `/dev/cdsp`, or `/dev/spidev0.1`, and never writes to `/dev/shm/eye-server.mmap` (it also carries calibration clicks into SteamVR's tracker), `/opt`, or `/persist`.
- **Eye images are biometric data.** Recordings live outside the repo, in `~/.local/share/frametop/eyes/captures` (0700), and `.gitignore` catches stray frame dumps. They go nowhere but the machine that runs your offline jobs.
- **Mind the headset's budget.** Finding a pupil takes about 0.4 ms a frame while ft-eyes follows it, and 1.4-2.1 ms when it searches the whole frame. Replays, scoring, and training go to a PC.
## Headset fit
The probe's Headset fit mode (Check headset fit on the Gaze page, or `ft-gazeprobe --mode fit`) shows, for each eye, whether the tracker has it, how open it is, and the tracker's confidence in it, and a map of where you looked coloured by how often it lost that eye there. Hints under the maps say which eye gets lost where, and what to try. Enter runs a guided check: dots around the screen, then looking down at the keyboard, up, left and right. R starts over. Adjust the headset while you watch it.
Losing an eye is usually about where you look, not the tracker. On this Frame the left eye was lost 57 to 64 % of the time looking 30 to 50 degrees down (at the keyboard) and the right eye never; at screen height both were seen over 98 % of the time. Looking down, the lids come down over the eyes. That's harmless, since the gaze service ignores looks down past the screens: they show on the maps, but not in the counts or as a problem.
## Probe
The trigger is Enter, Space, or a mouse button. Right-click anywhere in the window (or press the Menu key or Shift+F10) for a menu with Run calibration, Start accuracy test, Calibrate from last test, Reset calibration, the modes, the panel, fullscreen, and Quit. The buttons at the top right show and hide the panel, leave fullscreen, and quit. The keys do the same (Tab, F11, Esc), but only after you click the window once, since Frametop sends typing to the panel you clicked last. If ft-gaze stops, the probe starts it again after 3 s and shows why it stopped. Windowed mode stays on the screen it was on, and a small KWin script tells the probe where the window is, so the dot and targets are still in the right place.
The trigger is Enter, Space, or a mouse button. Right-click anywhere in the window (or press the Menu key or Shift+F10) for a menu with Run calibration, Start accuracy test, Calibrate from last test, Reset calibration, the modes, the panel, fullscreen, and Quit. The buttons at the top right show and hide the panel, leave fullscreen, and quit. The arrow in the panel's title bar collapses it to just that bar, so the dot and targets behind it stay visible; the collapsed bar stays through tests. The keys do the same (Tab, C, F11, Esc), but only after you click the window once, since Frametop sends typing to the panel you clicked last. If ft-gaze stops, the probe starts it again after 3 s and shows why it stopped. Windowed mode stays on the screen it was on, and a small KWin script tells the probe where the window is, so the dot and targets are still in the right place.
- **Run calibration (start here):** the initial calibration, modeled on Apple Vision Pro's eye setup. Face the centre and keep your head still. Look at one dot and press the trigger, then at each of six dots in a circle. That happens in three rounds, and the screen goes dark, then medium, then bright, because pupil size changes with brightness and the tracker's error with it. Each round turns the ring 20 degrees, and the middle round's ring is half the size, so the 21 dots cover the middle, halfway out, and the edge of your view. The ring's size is `Calibration ring` (degrees, 20 by default, less if the window is too small). Error grows toward the edge, and the calibration can only correct as far out as it has seen dots. The current dot is a bright pulsing dot with a point in the middle; finished dots fade to specks, so your eyes don't go back to them. Samples from blinks and from moments when the tracker lost an eye are dropped: openness under half of what it was during that look (not a fixed level, because your lids come down when you look down, and you squint in the bright round), or the angle between the eyes jumping more than 1.5 degrees from its median (that angle depends on how far away you're looking, so only a jump counts). Each dot is measured with medians, so one bad sample can't fail it. A look that lands where the gaze was for another dot of the round is refused as a look at the wrong dot. Mouse clicks don't count during a calibration run or a test: use Enter or Space. If a dot still fails, the message says why and the next try listens longer. After two failures, S (or the menu) skips the dot. Every attempt is logged to `calibration-attempts.jsonl`. At the end it fits every source's calibration from all the dots, replacing what it had learned (quadratic if the model was none). Esc cancels. The run is saved as `calibration-*.json`. With "Test after calibration" on (the default), the accuracy test starts right after, on new spots.
- **Free look:** the gaze dot. The trigger calibrates wherever you're looking (see below).
+367
View File
@@ -0,0 +1,367 @@
"""fitcheck: how well the eye tracker sees each eye, for fitting the headset (ft-gazeprobe's
Headset fit mode).
From each ft-gaze sample it takes, per eye, whether the tracker has that eye (its variance
for the eye's direction, "unc", under EYE_LOST; see gazecal), how open the eye is, and the
tracker's own confidence in its latest measurement of it ("eye" "q": the measurement's
variance, about 2e-5 on a clear view). It keeps that per direction you look in (10 degree
cells, and a few named regions), so a map shows where each eye gets lost, and turns it into
hints.
On the Frame this was written for, the left eye was lost 57-63 % of the time looking 30-50
degrees down (at the keyboard) and the right never; at screen height both were seen over
98 % of the time. Looking down, the lids come down over the eyes, and a glance at the
keyboard isn't where the gaze pointer matters: ft-gazed ignores looks down past the
screens. So those are on the maps, but not in the cards' counts or the hints' warnings.
Directions are head-relative degrees (yaw +left, pitch +up), the combined gaze's.
"""
import math
import statistics
from collections import deque
from gazecal import EYE_FOUND, EYE_LOST
EYES = ("Left eye", "Right eye")
CELL = 10.0
YAW = (-40, 40)
PITCH = (-50, 30)
CLOSED = 0.12 # openness under this: closed (a blink, or squeezed shut)
MIN_REGION = 60 # samples in a region before it's judged (two thirds of a second)
# Named regions, for the hints: (key, words, test on yaw and pitch).
REGIONS = [
("down", "down (at a keyboard or desk)", lambda y, p: p < -20),
("up", "up", lambda y, p: p > 15),
("left", "to the left", lambda y, p: y > 20 and -20 <= p <= 15),
("right", "to the right", lambda y, p: y < -20 and -20 <= p <= 15),
("centre", "straight ahead (screen height)", lambda y, p: abs(y) <= 20 and -20 <= p <= 15),
]
# The guided check: dots on the screen (fractions of its size; the corners stay clear of the
# probe's title bar and toolbar), then prompts to look past it. Seconds each.
GUIDE = [
("dot", (0.5, 0.5), 2.0), ("dot", (0.12, 0.2), 2.0), ("dot", (0.88, 0.2), 2.0),
("dot", (0.88, 0.92), 2.0), ("dot", (0.12, 0.92), 2.0), ("dot", (0.5, 0.92), 2.0),
("look", "Look down at your keyboard", 4.0), ("look", "Look up, above the screen", 3.0),
("look", "Look far to the left", 3.0), ("look", "Look far to the right", 3.0),
("dot", (0.5, 0.5), 2.0),
]
class FitCheck:
def __init__(self):
self.reset()
def reset(self):
self.lost = [False, False]
self.lost_since = [None, None]
self.losses = [0, 0] # times each eye was lost
self.durations = [[], []] # how long each loss lasted (s)
self.cells = [{}, {}] # per eye: (i, j) -> [samples, lost]
self.regions = [{k: [0, 0] for k, _, _ in REGIONS} for _ in EYES]
self.recent = [deque(maxlen=900), deque(maxlen=900)] # (t, lost) for the last 10 s
self.q = [deque(maxlen=180), deque(maxlen=180)] # recent fresh measurement variances
self.open = [0.0, 0.0]
self.unc = [0.0, 0.0]
self.gaze = None
self.samples = 0
self.guide = None # {"start": t, "step": i, "results": [...]}
self.have_eye_data = False
# --- Samples ---
def feed(self, s, now):
m1 = s["src"].get("mmap1") or {}
unc, opens = m1.get("unc"), m1.get("open")
if "hy" not in m1 or not unc or not opens:
return
self.have_eye_data = True
self.samples += 1
eye = s.get("eye") or {}
hy, hp = m1["hy"], m1["hp"]
self.gaze = (hy, hp)
self.unc = list(unc)
for k in (0, 1):
self.open[k] += 0.2 * (opens[k] - self.open[k])
was = self.lost[k]
self.lost[k] = unc[k] > (EYE_FOUND if was else EYE_LOST)
if self.lost[k] and not was:
if not looking_down(hy, hp):
self.losses[k] += 1
self.lost_since[k] = now
elif was and not self.lost[k] and self.lost_since[k] is not None:
if not looking_down(hy, hp):
self.durations[k].append(now - self.lost_since[k])
self.lost_since[k] = None
q = eye.get("q")
if q and (eye.get("new") or [1, 1])[k]:
self.q[k].append(q[k])
closed = [opens[k] < CLOSED for k in (0, 1)]
if all(closed) or all(self.lost):
return # a blink: says nothing about the fit
key = (math.floor(hy / CELL), math.floor(hp / CELL))
for k in (0, 1):
c = self.cells[k].setdefault(key, [0, 0])
c[0] += 1
c[1] += self.lost[k]
for rk, _, test in REGIONS:
if test(hy, hp):
r = self.regions[k][rk]
r[0] += 1
r[1] += self.lost[k]
if not looking_down(hy, hp):
self.recent[k].append((now, self.lost[k]))
g = self.guide
if g and g["step"] < len(GUIDE):
res = g["results"][g["step"]]
res[0] += 1
res[1] += self.lost[0]
res[2] += self.lost[1]
# --- The guided check ---
def toggle_guide(self, now):
if self.guide and self.guide["step"] < len(GUIDE):
self.guide = None
else:
self.guide = {"start": now, "step": 0, "step_start": now, "results": [[0, 0, 0] for _ in GUIDE]}
def guide_step(self, now):
"""The current step (kind, what, seconds left), or None when there's no check running."""
g = self.guide
if not g or g["step"] >= len(GUIDE):
return None
kind, what, secs = GUIDE[g["step"]]
if now - g["step_start"] >= secs:
g["step"] += 1
g["step_start"] = now
return self.guide_step(now)
return kind, what, secs - (now - g["step_start"])
# --- Summaries ---
def status(self, k):
if not self.samples:
return "no data", (0.6, 0.6, 0.6)
if self.lost[k]:
return "LOST", (1.0, 0.35, 0.3)
if self.open[k] < CLOSED:
return "closed", (0.8, 0.8, 0.8)
return "tracking", (0.35, 1.0, 0.5)
def tracked_share(self, k, now, window=10.0):
pts = [lost for t, lost in self.recent[k] if now - t <= window]
return (1 - sum(pts) / len(pts)) if pts else None
def signal(self, k):
"""The tracker's recent confidence in this eye, 0..1 (from its measurement variance:
2e-5 or less is 1, 1e-3 or more is 0), or None."""
if len(self.q[k]) < 10:
return None
q = statistics.median(self.q[k])
return min(1.0, max(0.0, (math.log10(1e-3) - math.log10(max(q, 1e-9))) / (math.log10(1e-3) - math.log10(2e-5))))
def region_share(self, k, key):
n, lost = self.regions[k][key]
return (lost / n) if n >= MIN_REGION else None
def hints(self):
if not self.have_eye_data:
return ["No per-eye data from ft-gaze (it needs SteamVR's eye-server.mmap, and a current build)."]
if self.samples < 3 * MIN_REGION:
return ["Look around slowly (the screen's corners, then down at your keyboard, up, left and right) "
"or press Enter for a guided check."]
out = []
bad = {}
for k in (0, 1):
for key, words, _ in REGIONS:
share = self.region_share(k, key)
if share is not None and share >= 0.15:
bad.setdefault(key, {})[k] = share
for key, words, _ in REGIONS:
if key not in bad:
continue
eyes = bad[key]
if len(eyes) == 2:
if key == "down":
out.append("Both eyes get lost looking down at the keyboard. That's fine: the gaze service "
"ignores looks down past the screens.")
elif key == "centre":
out.append(f"Both eyes get lost looking {words} ({eyes[0]:.0%} and {eyes[1]:.0%} of the time): "
"check the lenses are clean and the headset is on as usual; if it stays like this, "
"the tracker isn't getting a clear view of either eye.")
else:
out.append(f"Both eyes get lost looking {words}: that's past what the tracker covers for your "
"face, not one eye's fit.")
continue
k = next(iter(eyes))
other = self.region_share(1 - k, key)
vs = f", the {EYES[1 - k].lower()} {other:.0%}" if other is not None else ""
line = f"{EYES[k]}: lost {eyes[k]:.0%} of the time looking {words}{vs}."
if key == "down":
line += (" That's fine: glancing at the keyboard, the lids come down over the eyes, and the gaze "
"service ignores looks down past the screens, so the pointer stays put.")
elif key == "centre":
line += (" Even at screen height: clean that lens, and check its distance from your eye and the "
"IPD. Lashes that touch the lens get in the camera's way too.")
else:
line += (" At the edge of your view: try the IPD setting, and centring the headset between your "
"eyes.")
out.append(line)
s0, s1 = self.signal(0), self.signal(1)
if s0 is not None and s1 is not None and abs(s0 - s1) > 0.25:
k = 0 if s0 < s1 else 1
out.append(f"The tracker is less sure of your {EYES[k].lower()} even when it has it "
f"(signal {min(s0, s1):.0%} against {max(s0, s1):.0%}).")
if not out:
out.append("Both eyes are tracked everywhere you've looked so far.")
return out
# --- Drawing (cairo) ---
def draw(self, cr, w, h, text, now):
# Right of the probe's collapsed title bar, under its toolbar (top right).
left = 300
top = 190
text(cr, left, top - 60, "Headset fit", (1, 1, 1), 30)
text(cr, left, top - 28, "Adjust the headset and watch each eye. Enter: guided check. R: start over.",
(0.8, 0.8, 0.8), 18)
card_w = min(560, (w - left - 80) / 2)
mh = max(0, min(card_w * 0.8, h - top - 280 - 200))
for k in (0, 1):
x = left + k * (card_w + 40)
self.draw_card(cr, x, top, card_w, text, now, k)
self.draw_map(cr, x, top + 280, card_w, mh, text, k)
y = top + 280 + (mh + 60 if mh >= 80 else 0)
for line in self.hints()[:4]:
for part in wrap(line, max(40, int((w - left - 40) / 10))):
if y > h - 30:
break
text(cr, left, y, part, (1, 0.95, 0.75), 18)
y += 26
y += 8
step = self.guide_step(now)
g = self.guide
if step:
kind, what, left_s = step
if kind == "dot":
fx, fy = what
x, y = fx * w, fy * h
cr.set_source_rgba(1, 0.85, 0.2, 0.95)
cr.arc(x, y, 14 + 4 * math.sin(now * 6), 0, 2 * math.pi)
cr.fill()
else:
text(cr, w / 2 - 260, h / 2, f"{what} ({left_s:.0f})", (1, 0.85, 0.2), 34)
elif g and g["step"] >= len(GUIDE):
self.draw_guide_results(cr, w, h, text)
def draw_card(self, cr, x, y, cw, text, now, k):
cr.set_source_rgba(1, 1, 1, 0.06)
cr.rectangle(x, y, cw, 230)
cr.fill()
word, col = self.status(k)
text(cr, x + 16, y + 38, EYES[k], (1, 1, 1), 26)
cr.select_font_face("sans")
cr.set_font_size(26)
text(cr, x + cw - 16 - cr.text_extents(word).x_advance, y + 38, word, col, 26)
rows = [("Open", self.open[k]), ("Signal", self.signal(k)), ("Seen, last 10 s", self.tracked_share(k, now))]
yy = y + 70
for label, v in rows:
text(cr, x + 16, yy + 16, label, (0.85, 0.85, 0.85), 17)
bx, bw = x + 170, cw - 250
cr.set_source_rgba(1, 1, 1, 0.12)
cr.rectangle(bx, yy, bw, 20)
cr.fill()
if v is not None:
v = min(1.0, max(0.0, v))
cr.set_source_rgba(*bar_colour(v), 0.9)
cr.rectangle(bx, yy, bw * v, 20)
cr.fill()
text(cr, bx + bw + 10, yy + 16, f"{v:.0%}", (0.9, 0.9, 0.9), 17)
yy += 36
d = self.durations[k]
longest = max(d) if d else 0
n = self.losses[k]
text(cr, x + 16, yy + 22, f"Lost {n} time{'' if n == 1 else 's'}" + (f", longest {longest:.1f} s" if longest >= 0.05 else ""),
(0.85, 0.85, 0.85), 17)
def draw_map(self, cr, x, y, mw, mh, text, k):
"""Where you looked (yaw across, pitch up), each cell coloured by how often this eye
was lost there: green never, red always, dark: not looked there yet."""
if mh < 80:
return
cols = int((YAW[1] - YAW[0]) / CELL)
rows = int((PITCH[1] - PITCH[0]) / CELL)
cw, ch = mw / cols, mh / rows
text(cr, x, y - 8, f"Where the {EYES[k].lower()} gets lost", (0.85, 0.85, 0.85), 17)
for i in range(cols):
yaw_i = math.floor(YAW[1] / CELL) - 1 - i # left of the map is your left (+yaw)
for j in range(rows):
pitch_j = math.floor(PITCH[1] / CELL) - 1 - j
c = self.cells[k].get((yaw_i, pitch_j))
cx, cy = x + i * cw, y + j * ch
if c and c[0] >= 10:
share = c[1] / c[0]
cr.set_source_rgba(*bar_colour(1 - share), 0.75)
else:
cr.set_source_rgba(1, 1, 1, 0.05)
cr.rectangle(cx + 1, cy + 1, cw - 2, ch - 2)
cr.fill()
# Straight ahead, and the gaze now.
def at(yaw, pitch):
return x + (YAW[1] - yaw) / (YAW[1] - YAW[0]) * mw, y + (PITCH[1] - pitch) / (PITCH[1] - PITCH[0]) * mh
cr.set_source_rgba(1, 1, 1, 0.35)
cr.set_line_width(1)
ox, oy = at(0, 0)
cr.move_to(ox - 10, oy)
cr.line_to(ox + 10, oy)
cr.move_to(ox, oy - 10)
cr.line_to(ox, oy + 10)
cr.stroke()
text(cr, x, y + mh + 20, "+ ahead, bottom rows: keyboard", (0.6, 0.6, 0.6), 14)
if self.gaze:
gx, gy = at(max(YAW[0], min(YAW[1], self.gaze[0])), max(PITCH[0], min(PITCH[1], self.gaze[1])))
cr.set_source_rgba(1, 1, 1, 0.95)
cr.arc(gx, gy, 5, 0, 2 * math.pi)
cr.fill()
def draw_guide_results(self, cr, w, h, text):
res = self.guide["results"]
lines = []
for (kind, what, _), (n, l0, l1) in zip(GUIDE, res):
if not n:
continue
name = what if kind == "look" else "dot at {:.0%}, {:.0%}".format(*what)
lines.append(f"{name}: left lost {l0 / n:.0%}, right {l1 / n:.0%}")
y = h / 2 - 20 * len(lines)
text(cr, w / 2 - 300, y - 40, "Guided check", (1, 0.85, 0.2), 26)
for line in lines:
text(cr, w / 2 - 300, y, line, (1, 1, 1), 19)
y += 30
def looking_down(yaw, pitch):
"""A look down at the keyboard: the "down" region, which ft-gazed doesn't send on."""
return pitch < -20
def bar_colour(v):
"""Red (0) through amber to green (1)."""
if v < 0.5:
return 1.0, 0.3 + 0.9 * v, 0.3
return 1.0 - 1.3 * (v - 0.5), 0.75 + 0.25 * (v - 0.5) * 2, 0.35
def wrap(s, width):
words, lines, cur = s.split(), [], ""
for wd in words:
if cur and len(cur) + 1 + len(wd) > width:
lines.append(cur)
cur = wd
else:
cur = f"{cur} {wd}".strip()
if cur:
lines.append(cur)
return lines
+1
View File
@@ -5,6 +5,7 @@ Documentation=file://@REPO@/gaze/README.md
# Needs SteamVR's IPC (ft-gaze is an overlay client); it starts and stops with SteamVR.
After=steamvr.service frametop-pointer.service
PartOf=steamvr.service
Requisite=steamvr.service
[Service]
# Host Python; it runs ft-gaze in the dev container (distrobox enter), which quits when
+148 -10
View File
@@ -8,13 +8,30 @@
//
// {"t":<sample time, CLOCK_MONOTONIC_RAW s>,"age":<ms old when read>,"n":<sample counter>,
// "head":{"yaw":..,"pitch":..,"hit":HIT}, head forward ray (for head nudging)
// "src":{"action":SRC,"mmap1":SRC,"mmap2":SRC}}
// "src":{"action":SRC,"mmap1":SRC,"mmap2":SRC,"left":SRC,"right":SRC,"own":SRC},"eye":EYE}
// SRC = {"hy":..,"hp":..,"hit":HIT} or {"ok":0} hy/hp: gaze direction relative to the
// head, degrees (yaw +left, pitch +up)
// mmap1 adds "open":[l,r] (probably eye openness, 0 in a blink) and "dist" (vergence
// distance, m); both mmap sets add "lr", the angle between the eyes (deg), which
// jumps when the tracker loses an eye, and "eyes":[[hy,hp],[hy,hp]], each eye's own
// direction (left, right), for calibrating the eyes separately.
// direction (left, right), for calibrating the eyes separately, and "unc":[l,r],
// the tracker's uncertainty about each eye's direction (its filter's variance):
// about 0.0005-0.002 while it sees the eye, 0.015-0.03 once it's lost it.
// "left":SRC,"right":SRC each eye's own direction from set 2
// (set 1's eyes always share one pitch, and while it's lost an eye it keeps that
// eye's yaw where it was: set 2 is each eye's own reading). From the head's origin,
// not the eye's.
// "own":SRC our own tracker (gaze/tracker/ft-eyes), from
// /dev/shm/frametop-eyes-gaze; adds "age" (ms since its frame), "eyes":[[hy,hp],[hy,hp]]
// (left, right; null for an eye it doesn't see), "ehit":[HIT,HIT] where each of those
// lands, and "slip":[[x,y],[x,y]] (left, right: each eye's shift in its camera image
// since the calibration, pixels; null until a click has measured it). {"ok":0}
// without the file or when it's over 100 ms old.
// EYE = {"q":[l,r],"m":[[x,y],[x,y]],"new":[l,r]} the tracker's latest measurement of
// each eye before filtering: "m" (camera-relative, undocumented units), "q" its
// variance (about 2e-5 on a clear view of the eye, rising as the lid or lashes get
// in the way), "new" whether it changed since the last sample (it freezes while the
// tracker can't see that eye, and in blinks). "eye" is null without the mmap.
// HIT = {"s":<screen>,"x":..,"y":..,"j":[dx/dhy,dy/dhy,dx/dhp,dy/dhp],"dpp":<deg per px>}
// or null. x, y are pixels on that screen; j is pixels per degree of head-relative
// yaw and pitch there, so a correction in degrees can be turned into pixels and back.
@@ -78,7 +95,13 @@ constexpr size_t kLeft1 = 0x15f, kRight1 = 0x16b; // set 1: unit vectors, head
constexpr size_t kFix1 = 0x18f; // set 1 fixation point: length is the vergence distance (m)
constexpr size_t kLeft2 = 0x19b, kRight2 = 0x1a7; // set 2
constexpr size_t kOpen = 0x1cb; // two floats, 0..1: probably eye openness or confidence
constexpr size_t kNeed = 0x1d3;
// After each set's two directions, six floats: the left eye's variance (three), the
// right's (three; the middle one of each is shared). They jump when an eye is lost.
constexpr size_t kVar1 = 0x177, kVar2 = 0x1b3;
// The measurements the filter is fed: left x, y, right x, y, then the variance of each (left
// x, y, right x, y). An eye's pair stops changing while the tracker can't see it.
constexpr size_t kMeas = 0x1d3;
constexpr size_t kNeed = 0x1f3;
struct EyeFile {
const uint8_t *p = nullptr;
@@ -115,6 +138,7 @@ struct EyeSample {
double t = 0;
Vec3 left1, right1, fix1, left2, right2;
float open[2] = {0, 0};
float var1[6] = {}, var2[6] = {}, meas[8] = {};
};
// A consistent copy: the writer has no seqlock we can use, so read until the counter and
@@ -127,6 +151,9 @@ bool ReadSample(const EyeFile &f, EyeSample &s) {
s.left1 = f.V(kLeft1), s.right1 = f.V(kRight1), s.fix1 = f.V(kFix1);
s.left2 = f.V(kLeft2), s.right2 = f.V(kRight2);
std::memcpy(s.open, f.p + kOpen, sizeof s.open);
std::memcpy(s.var1, f.p + kVar1, sizeof s.var1);
std::memcpy(s.var2, f.p + kVar2, sizeof s.var2);
std::memcpy(s.meas, f.p + kMeas, sizeof s.meas);
std::atomic_thread_fence(std::memory_order_acquire);
if (f.Get<uint32_t>(kCounter) == n0 && f.Get<double>(kTime) == t0) {
s.n = n0, s.t = t0;
@@ -136,6 +163,69 @@ bool ReadSample(const EyeFile &f, EyeSample &s) {
return false;
}
// --- Our own tracker: /dev/shm/frametop-eyes-gaze, written by gaze/tracker/ft-eyes ---
// Layout (ft-eyes' docstring): u32 seq (odd while written), u32 version, f64 t, f32 yaw,
// pitch, u32 flags (bit 0 right eye, 1 left, 2 right slip known, 3 left), u32 n, then f32
// right yaw, pitch, left yaw, pitch; slip right x, y, left x, y; pupils (unused here).
struct OwnSample {
double t = 0;
float yaw = 0, pitch = 0;
uint32_t flags = 0, n = 0;
float eyes[4] = {}, slip[4] = {};
};
class OwnFile {
public:
// Reopened when it appears or is replaced, since ft-eyes may start after us.
bool Read(OwnSample &o) {
const double now = NowRaw();
if (!p_ || now - checked_ > 2.0) Reopen(now);
if (!p_) return false;
for (int attempt = 0; attempt < 4; ++attempt) {
uint32_t s0, s1, version;
std::memcpy(&s0, p_, 4);
if (s0 & 1) continue;
std::atomic_thread_fence(std::memory_order_acquire);
std::memcpy(&version, p_ + 4, 4);
std::memcpy(&o.t, p_ + 8, 8);
std::memcpy(&o.yaw, p_ + 16, 4);
std::memcpy(&o.pitch, p_ + 20, 4);
std::memcpy(&o.flags, p_ + 24, 4);
std::memcpy(&o.n, p_ + 28, 4);
std::memcpy(o.eyes, p_ + 32, sizeof o.eyes);
std::memcpy(o.slip, p_ + 48, sizeof o.slip);
std::atomic_thread_fence(std::memory_order_acquire);
std::memcpy(&s1, p_, 4);
if (s0 == s1) return version == 1;
}
return false;
}
private:
static constexpr size_t kSize = 128;
void Reopen(double now) {
checked_ = now;
struct stat st {};
if (stat("/dev/shm/frametop-eyes-gaze", &st) != 0) return Close();
if (p_ && st.st_ino == ino_) return;
Close();
const int fd = open("/dev/shm/frametop-eyes-gaze", O_RDONLY | O_CLOEXEC);
if (fd < 0) return;
if (fstat(fd, &st) == 0 && size_t(st.st_size) >= kSize) {
void *m = mmap(nullptr, kSize, PROT_READ, MAP_SHARED, fd, 0);
if (m != MAP_FAILED) p_ = static_cast<const uint8_t *>(m), ino_ = st.st_ino;
}
close(fd);
}
void Close() {
if (p_) munmap(const_cast<uint8_t *>(p_), kSize);
p_ = nullptr;
}
const uint8_t *p_ = nullptr;
ino_t ino_ = 0;
double checked_ = -1e9;
};
// --- Screens from ft-screens ---
struct Screen {
int index = 0;
@@ -351,7 +441,11 @@ int main(int argc, char **argv) {
stdinClosed = true;
}).detach();
vr::EVRInitError err = vr::VRInitError_None;
vr::VR_Init(&err, vr::VRApplication_Overlay);
vr::VR_Init(&err, vr::VRApplication_Background);
if (err == vr::VRInitError_None) {
vr::VR_Shutdown();
vr::VR_Init(&err, vr::VRApplication_Overlay);
}
if (err != vr::VRInitError_None) {
std::fprintf(stderr, "ft-gaze: SteamVR: %s\n", vr::VR_GetVRInitErrorAsEnglishDescription(err));
return 1;
@@ -373,10 +467,12 @@ int main(int argc, char **argv) {
const bool haveMmap = eyes.Open();
std::fprintf(stderr, "ft-gaze: eye-server.mmap %s\n", haveMmap ? "open" : "not available");
OwnFile ownFile;
Screens screens;
screens.Start();
PoseHistory history;
uint32_t lastN = 0;
float lastMeas[8] = {};
double lastEmit = 0;
int actionErrors = 0;
vr::EVRInputError lastActionError = vr::VRInputError_None;
@@ -423,7 +519,7 @@ int main(int argc, char **argv) {
lastActionError = ae;
}
std::string m1 = "{\"ok\":0}", m2 = m1;
std::string m1 = "{\"ok\":0}", m2 = m1, left = m1, right = m1, eye = "null";
if (haveMmap) {
// lr: the angle between the two eyes' directions. It's a fraction of a degree
// normally; when the tracker loses one eye (or during a blink) it jumps.
@@ -438,12 +534,53 @@ int main(int argc, char **argv) {
std::snprintf(b, sizeof b, "\"eyes\":[[%.4f,%.4f],[%.4f,%.4f]],", ly, lp, ry, rp);
return std::string(b);
};
char extra[128];
auto unc = [](const float *v) {
char b[64];
std::snprintf(b, sizeof b, "\"unc\":[%.5f,%.5f],", std::max(v[0], v[2]), std::max(v[3], v[5]));
return std::string(b);
};
char extra[256];
std::snprintf(extra, sizeof extra, "\"dist\":%.3f,\"open\":[%.3f,%.3f],\"lr\":%.3f,", Length(s.fix1),
s.open[0], s.open[1], lr(s.left1, s.right1));
m1 = SrcJson(list, headThen, s.left1 + s.right1, extra + eyes(s.left1, s.right1));
m1 = SrcJson(list, headThen, s.left1 + s.right1, extra + eyes(s.left1, s.right1) + unc(s.var1));
std::snprintf(extra, sizeof extra, "\"lr\":%.3f,", lr(s.left2, s.right2));
m2 = SrcJson(list, headThen, s.left2 + s.right2, extra + eyes(s.left2, s.right2));
m2 = SrcJson(list, headThen, s.left2 + s.right2, extra + eyes(s.left2, s.right2) + unc(s.var2));
left = SrcJson(list, headThen, s.left2);
right = SrcJson(list, headThen, s.right2);
const float *m = s.meas;
const bool newL = m[0] != lastMeas[0] || m[1] != lastMeas[1];
const bool newR = m[2] != lastMeas[2] || m[3] != lastMeas[3];
std::memcpy(lastMeas, m, sizeof lastMeas);
std::snprintf(extra, sizeof extra, "{\"q\":[%.3g,%.3g],\"m\":[[%.4f,%.4f],[%.4f,%.4f]],\"new\":[%d,%d]}",
(m[4] + m[5]) / 2, (m[6] + m[7]) / 2, m[0], m[1], m[2], m[3], int(newL), int(newR));
eye = extra;
}
// Our tracker: its own sample time picks the head pose, like the mmap's.
std::string own = "{\"ok\":0}";
OwnSample o;
if (ownFile.Read(o) && now - o.t < 0.1) {
vr::HmdMatrix34_t headOwn = headNow;
history.At(o.t, headOwn);
auto pair = [](bool ok, float a, float b) {
char p[48];
if (!ok) return std::string("null");
std::snprintf(p, sizeof p, "[%.4f,%.4f]", a, b);
return std::string(p);
};
// Stored right eye first; reported left first, like the other sources.
const std::string extra = "\"age\":" + std::to_string(int((now - o.t) * 1000)) +
",\"eyes\":[" + pair(o.flags & 2, o.eyes[2], o.eyes[3]) + "," +
pair(o.flags & 1, o.eyes[0], o.eyes[1]) + "],\"slip\":[" +
pair(o.flags & 8, o.slip[2], o.slip[3]) + "," +
pair(o.flags & 4, o.slip[0], o.slip[1]) + "],";
// Where each eye's own gaze lands (left, right), for drawing them apart.
auto eyeHit = [&](bool ok, float y, float p) {
return ok ? HitJson(list, headOwn, y, p) : std::string("null");
};
const std::string hits = "\"ehit\":[" + eyeHit(o.flags & 2, o.eyes[2], o.eyes[3]) + "," +
eyeHit(o.flags & 1, o.eyes[0], o.eyes[1]) + "],";
own = SrcJson(list, headOwn, Direction(o.yaw, o.pitch), extra + hits);
}
double yaw, pitch;
@@ -451,9 +588,10 @@ int main(int argc, char **argv) {
yaw = std::atan2(-f.x, -f.z) * 180 / M_PI;
pitch = std::asin(std::clamp(f.y, -1.0, 1.0)) * 180 / M_PI;
std::printf("{\"t\":%.5f,\"age\":%.1f,\"n\":%u,\"head\":{\"yaw\":%.4f,\"pitch\":%.4f,\"hit\":%s},"
"\"src\":{\"action\":%s,\"mmap1\":%s,\"mmap2\":%s}}\n",
"\"src\":{\"action\":%s,\"mmap1\":%s,\"mmap2\":%s,\"left\":%s,\"right\":%s,\"own\":%s},"
"\"eye\":%s}\n",
s.t, (now - s.t) * 1000, s.n, yaw, pitch, HitJson(list, headNow, 0, 0).c_str(), action.c_str(),
m1.c_str(), m2.c_str());
m1.c_str(), m2.c_str(), left.c_str(), right.c_str(), own.c_str(), eye.c_str());
if (std::fflush(stdout) != 0) break; // the reader went away
}
+493 -92
View File
@@ -1,32 +1,66 @@
#!/usr/bin/python3
"""ft-gazed: the gaze service. The headset's eye tracking, corrected, for the pointer.
Two settings in ~/.config/frametop.conf (the Gaze page of Frametop Input Settings), read
again when the file changes:
GAZE_TRACKER=steam|own SteamVR's eye tracker (default), or our own (gaze/tracker/ft-eyes,
ft-gaze's source "own"; this service runs it, see below)
GAZE_EYE=auto|left|right the eye bias (gazecal.EyeWeights): auto weights each eye by how far
off it was at your recent nudges; left or right counts that eye twice
as much as the other. Either eye alone carries the gaze when the
other isn't seen.
Runs ft-gaze (in the dev container), and for every eye tracker sample (90 Hz):
1. drops blinks: both eyes' openness under half its running median (each eye its own).
With the default source, mmap set 1, that's all: set 1 is SteamVR's combined gaze,
which keeps going when the tracker loses one eye (its two directions stay together).
Set 2's combined direction is the mean of the eyes' own, so with one eye lost it's
off by half of whatever that eye reads (9 to 14 degrees apart were seen): with set 2,
samples with an eye under its floor, or the angle between the eyes jumping more than
1.5 degrees from its median, are dropped too;
2. smooths it with a fixation lock (the running mean of the current fixation, 1 degree);
3. corrects it: the calibration from ft-gazeprobe (calibration.json, reloaded when the
probe changes it) plus what the pointer's corrections have taught since (LiveCorrection,
saved in pointer-lessons.json);
1. drops blinks: both eyes' openness under half its running median (each eye its own),
or both lost (the tracker's variance for them, ft-gaze's "unc", over EYE_LOST);
Looks down past the screens (pitch under KEYBOARD_PITCH, on no Frametop screen: at the
keyboard, through the gap by the nose) aren't sent, so the pointer stays where it was
instead of following you down; the tracker often loses an eye there (the lids come
down), and that isn't counted as a lost eye either;
2. combines the eyes, each corrected on its own. With SteamVR, that's each eye's own
reading (set 2, ft-gaze's "left" and "right"), corrected by its calibration from
ft-gazeprobe (calibration.json, reloaded when the probe changes it) plus what the
pointer's corrections have taught that eye since (LiveCorrection, saved in
pointer-lessons.json), then weighted by the eye bias. A lost or closed eye drops out.
Our own tracker keeps its own calibration, so its eyes are used as they come;
3. smooths it with a fixation lock (the running mean of the current fixation, 1 degree);
4. sends it to the pointer helper: "gz <yaw> <pitch> <raw yaw> <raw pitch>", head-relative
degrees (yaw +left, pitch +up). The helper uses it only in gaze mode.
Without per-eye calibrations (a calibration from before the probe had the eyes as sources),
or with --source, it's the older path: one source, SteamVR's combined gaze (mmap set 1) by
default, corrected as a whole. There, with one eye lost or closed, the gaze comes from the
other (EyeFallback: that eye's own reading from set 2, plus what it usually reads against
the combined gaze, learned while both are seen). SteamVR's combined gaze (set 1) keeps going
on one eye too, but holds the lost eye's yaw, so it moves half as far sideways as the eyes
do. Before the fallback has learned an eye, set 1 is used as it is; set 2's combined
direction is the mean of the eyes' own (off by half of whatever the lost eye reads), so with
set 2 that sample is dropped, as is one where the angle between the eyes jumps more than 1.5
degrees from its median.
Lessons come back from the helper: when you nudge the gaze-placed pointer with the mouse and
click, it sends "lesson <raw yaw> <raw pitch> <true yaw> <true pitch>": where the raw gaze
was when the mouse took over, and where the pointer was when you clicked (you were looking
there). The gap is the tracker's error there, and it's learned, unless it's more than
LESSON_MAX degrees past the correction (then it wasn't a nudge onto what you looked at).
there). The gap is the tracker's error there. The raw gaze is the one sent, so it also says
when that look was (the history of what was sent), and so what each eye read then:
- SteamVR: each eye learns its own error, unless the gaze was more than LESSON_MAX degrees
past the correction (then it wasn't a nudge onto what you looked at);
- our tracker: the look goes to it as a click ("click T YAW PITCH" on @ft_eyes), as the
probe's clicks do, and it learns how far the headset has moved on your face. That's
what it gets wrong, and after the headset was off, the first click resets it;
- either way, how far off each eye was (before the lesson taught it anything) goes to the
eye bias, for auto.
SteamVR's eye tracking log is followed for the headset going on (its eye model starts over,
and the error moves): lessons from before count less then, so the first few after it
relearn the offset.
Our own tracker (gaze/tracker/ft-eyes) runs here too, in the dev container, while
GAZE_TRACKER=own or the gaze probe asks for it ("eyes SECONDS", a lease the probe renews). It
reads the eye-camera frames the root service frametop-eyegrab copies (gaze/tracker/install.sh),
which copies them only while ft-eyes runs.
Nothing here writes to SteamVR, its eye tracker, or its files: ft-gaze reads the eye
tracker's shared memory read-only.
@@ -34,16 +68,20 @@ Control socket: abstract unix datagram "@ft_gazed":
lesson <rhy> <rhp> <thy> <thp> from the pointer helper (see above)
status reply: one JSON object
forget drop what the lessons taught (the calibration stays)
reload read calibration.json again
reload read calibration.json and the settings again
eyes <seconds> keep our own tracker running that much longer (at most 120),
whatever GAZE_TRACKER says: the probe's lease. Reply: "ok"
Options: --source action|mmap1|mmap2 (default mmap1; set 2 was a little quieter in the probe,
but loses the pointer whenever the tracker loses an eye), -v (a status line every
5 s on stderr), --to NAME (send the gaze to the abstract socket @NAME instead of the pointer
helper; for testing: a helper without gaze mode forwards what it doesn't know to its driver).
Options: --source action|mmap1|mmap2 (the older one-source path with that source, whatever
the settings say; set 2 was a little quieter in the probe, but loses the pointer whenever
the tracker loses an eye), -v (a status line every 5 s on stderr), --to NAME (send the gaze
to the abstract socket @NAME instead of the pointer helper; for testing: a helper without
gaze mode forwards what it doesn't know to its driver).
"""
import argparse
import json
import math
import os
import selectors
import signal
@@ -56,17 +94,34 @@ from collections import deque
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
from gazecal import DEFAULT_MODEL, MODELS, STATE, Correction, Fixation, LiveCorrection, SteamEyeLog # noqa: E402
from gazecal import (DEFAULT_MODEL, EYE_FOUND, EYE_LOST, MODELS, STATE, Correction, EyeFallback, # noqa: E402
EyeWeights, Fixation, LiveCorrection, SteamEyeLog)
REPO = Path(__file__).resolve().parents[1]
HELPER = REPO / "gaze" / "build" / "ft-gaze"
ME = "\0ft_gazed"
POINTER = "\0ft_pointer_helper"
EYES_PROG = REPO / "gaze" / "tracker" / "ft-eyes" # our own tracker
EYES_PYTHON = REPO / "gaze" / "tracker" / "build" / "venv" / "bin" / "python" # numpy, OpenCV (build.sh)
EYES_SOCKET = "\0ft_eyes" # its control socket
EYES_CAMS = Path("/dev/shm/frametop-eyes-cams") # the frames it reads (frametop-eyegrab.service)
CONF = Path.home() / ".config" / "frametop.conf"
CALIBRATION = STATE / "calibration.json"
LESSONS = STATE / "pointer-lessons.json"
LESSON_LOG = STATE / "pointer-lessons.jsonl"
SOURCES = ("action", "mmap1", "mmap2", "left", "right") # the ones with a calibration here
SIDES = ("left", "right") # ft-gaze's order, and the sources for each eye alone
TRACKERS = ("steam", "own")
BIASES = ("auto", "left", "right")
LESSON_MAX = 8.0 # degrees past the correction
OWN_LESSON_MAX = 25.0 # our tracker: after the headset was off, its first clicks can be 10-17 off
HISTORY = 12.0 # seconds of the gaze sent, to find a lesson's look (the helper sends it up to 10 s later)
LOOK = 0.3 # seconds of samples before that moment make the look (the probe's fixation)
RETRY = 3.0 # seconds before starting ft-gaze again
EYES_RETRY = 10.0 # seconds before starting ft-eyes again after it stopped on its own
EYES_LEASE_MAX = 120.0
SETTLE = 0.3 # seconds after an eye is found again before the fallback learns from it
KEYBOARD_PITCH = -20.0 # degrees: gaze under this, on no screen, is a look at the keyboard
class PointerLessons(LiveCorrection):
@@ -85,85 +140,215 @@ def log(msg):
print(f"ft-gazed: {msg}", file=sys.stderr, flush=True)
def read_settings():
"""(tracker, eye bias) from frametop.conf, defaults for anything missing or unknown."""
conf = {}
try:
for line in CONF.read_text().splitlines():
line = line.split("#", 1)[0].strip()
if "=" in line:
k, v = line.split("=", 1)
conf[k.strip()] = v.strip().lower()
except OSError:
pass
tracker = conf.get("GAZE_TRACKER", "steam")
bias = conf.get("GAZE_EYE", "auto")
return tracker if tracker in TRACKERS else "steam", bias if bias in BIASES else "auto"
def mtime(path):
try:
return path.stat().st_mtime
except OSError:
return None
class Service:
def __init__(self, source, verbose, to=POINTER):
self.source, self.verbose, self.to = source, verbose, to
self.override, self.verbose, self.to = source, verbose, to
self.source = source or "mmap1" # the older path's source
STATE.mkdir(parents=True, exist_ok=True)
self.base = Correction()
self.tracker, self.bias = read_settings()
self.conf_mtime = mtime(CONF)
self.models = {name: Correction() for name in SOURCES}
self.mode = DEFAULT_MODEL
self.cal_mtime = None
self.live = PointerLessons()
self.lives = {name: PointerLessons() for name in SOURCES}
self.weights = {t: EyeWeights(self.bias) for t in TRACKERS}
self.dirty = False
self.load_calibration()
self.load_lessons()
self.steam = SteamEyeLog()
self.steam.poll()
self.live.wear_time = self.steam.worn()
self.live.refit(self.base, self.mode)
self.refit()
self.fix = Fixation(radius=1.0)
self.opens = (deque(maxlen=90), deque(maxlen=90)) # left, right
self.vergence = deque(maxlen=90)
self.counts = {"samples": 0, "sent": 0, "blinks": 0, "one_eye": 0, "dropped": 0, "lessons_taken": 0, "refused": 0}
self.fallback = EyeFallback()
self.lost = [False, False]
self.bad_at = [0.0, 0.0] # sample time an eye was last lost or closed
self.counts = {"samples": 0, "sent": 0, "blinks": 0, "one_eye": 0, "one_eye_used": 0, "lost_left": 0,
"lost_right": 0, "looking_down": 0, "dropped": 0, "lessons_taken": 0, "refused": 0}
self.last_sample = 0.0
self.last = None
self.last_kind = None
# What was sent, for finding a lesson's look: (sample time, raw as sent, each eye's reading).
self.history = deque()
self.own = {} # our tracker's last status reply
self.own_at = 0.0
self.eyes_proc = None # ft-eyes, while it runs
self.eyes_until = 0.0 # the probe's lease (monotonic time)
self.eyes_restart_at = 0.0
self.sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM | socket.SOCK_CLOEXEC | socket.SOCK_NONBLOCK)
self.sock.bind(ME)
self.out = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM | socket.SOCK_CLOEXEC | socket.SOCK_NONBLOCK)
# To our tracker, with an address of its own, so its replies don't land on @ft_gazed.
self.eyes_sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM | socket.SOCK_CLOEXEC | socket.SOCK_NONBLOCK)
self.eyes_sock.bind("")
self.sel = selectors.DefaultSelector()
self.sel.register(self.sock, selectors.EVENT_READ, "control")
self.sel.register(self.eyes_sock, selectors.EVENT_READ, "own")
self.proc = None
self.buf = b""
self.restart_at = 0.0
self.running = True
# --- Calibration and lessons ---
@property
def kind(self):
""""own" (our tracker), "eyes" (SteamVR's eyes, each corrected), or "source" (the
older path: one SteamVR source, corrected as a whole)."""
if self.override:
return "source"
if self.tracker == "own":
return "own"
return "eyes" if all(self.models[e].samples for e in SIDES) else "source"
# --- Settings, calibration and lessons ---
def load_settings(self):
self.conf_mtime = mtime(CONF)
tracker, bias = read_settings()
if (tracker, bias) != (self.tracker, self.bias):
log(f"tracker {tracker}, eye bias {bias}" + (f" (--source {self.override} wins)" if self.override else ""))
self.tracker, self.bias = tracker, bias
for w in self.weights.values():
w.bias = bias
self.fix.reset()
def load_calibration(self):
try:
mtime = CALIBRATION.stat().st_mtime
mt = CALIBRATION.stat().st_mtime
d = json.loads(CALIBRATION.read_text())
except (OSError, ValueError):
return
self.cal_mtime = mtime
if self.source in d:
self.base.from_json(d[self.source])
self.cal_mtime = mt
for name in SOURCES:
if name in d:
self.models[name].from_json(d[name])
mode = d.get("_meta", {}).get("model")
self.mode = mode if mode in MODELS else DEFAULT_MODEL
log(f"calibration: {self.mode}, {self.base.samples} samples")
log(f"calibration: {self.mode}, " + ", ".join(f"{n} {self.models[n].samples}" for n in (self.source,) + SIDES)
+ " samples")
def load_lessons(self):
try:
d = json.loads(LESSONS.read_text())
if d.get("source") == self.source:
self.live.samples = d.get("samples", [])[-PointerLessons.KEEP:]
except (OSError, ValueError):
pass
log(f"{len(self.live.samples)} lessons")
d = {}
# The first version kept one source's: {"source": NAME, "samples": [...]}.
sources = d.get("sources") or ({d["source"]: d.get("samples", [])} if "source" in d else {})
for name, samples in sources.items():
if name in self.lives:
self.lives[name].samples = samples[-PointerLessons.KEEP:]
for t, misses in (d.get("misses") or {}).items():
if t in self.weights:
self.weights[t] = EyeWeights(self.bias, misses)
log(", ".join(f"{n} {len(self.lives[n].samples)}" for n in (self.source,) + SIDES) + " lessons")
def save_lessons(self):
tmp = LESSONS.with_suffix(".tmp")
tmp.write_text(json.dumps({"source": self.source, "samples": self.live.samples}))
tmp.write_text(json.dumps({"version": 2, "sources": {n: lv.samples for n, lv in self.lives.items() if lv.samples},
"misses": {t: w.misses for t, w in self.weights.items()}}))
tmp.replace(LESSONS)
self.dirty = False
def correction(self, hy, hp):
by, bp = self.base.get(hy, hp, self.mode)
ly, lp = self.live.get(hy, hp)
def refit(self):
for name, live in self.lives.items():
live.wear_time = self.steam.worn()
live.refit(self.models[name], self.mode)
def correction(self, name, hy, hp):
by, bp = self.models[name].get(hy, hp, self.mode)
ly, lp = self.lives[name].get(hy, hp)
return by + ly, bp + lp
def look(self, ry, rp):
"""When the gaze sent as raw (ry, rp) was last sent, and each eye's median reading over
the LOOK before it: (t, [(yaw, pitch) or None] * 2), or (None, None)."""
key = f"{ry:.3f} {rp:.3f}"
t = next((h[0] for h in reversed(self.history) if h[1] == key), None)
if t is None:
return None, None
eyes = []
for k in (0, 1):
seen = [h[2][k] for h in self.history if t - LOOK <= h[0] <= t and h[2] and h[2][k]]
eyes.append((statistics.median(e[0] for e in seen), statistics.median(e[1] for e in seen)) if seen else None)
return t, eyes
def lesson(self, rhy, rhp, thy, thp):
dy, dp = thy - rhy, thp - rhp # the whole error there
cy, cp = self.correction(rhy, rhp)
left = ((dy - cy) ** 2 + (dp - cp) ** 2) ** 0.5
rec = {"time": time.time(), "source": self.source, "model": self.mode, "raw": [rhy, rhp],
"true": [thy, thp], "correction": [cy, cp], "lesson_deg": left, "wear": self.steam.worn()}
if left > LESSON_MAX:
rec["refused"] = f"more than {LESSON_MAX} deg past the correction"
kind = self.kind
rec = {"time": time.time(), "kind": kind, "raw": [rhy, rhp], "true": [thy, thp], "wear": self.steam.worn()}
if kind == "source":
dy, dp = thy - rhy, thp - rhp # the whole error there
cy, cp = self.correction(self.source, rhy, rhp)
left = math.hypot(dy - cy, dp - cp)
rec.update(source=self.source, model=self.mode, correction=[cy, cp], lesson_deg=left)
if left > LESSON_MAX:
rec["refused"] = f"more than {LESSON_MAX} deg past the correction"
else:
self.lives[self.source].add({"time": rec["time"], "hy": rhy, "hp": rhp, "dy": dy, "dp": dp,
"wy": 1.0, "wp": 1.0, "how": "pointer"}, self.models[self.source], self.mode)
return self.taken(rec)
# The raw gaze sent here is the corrected, combined one: its whole error is left.
left = math.hypot(thy - rhy, thp - rhp)
t, eyes = self.look(rhy, rhp)
weights = self.weights[self.tracker if kind == "own" else "steam"]
rec.update(tracker=self.tracker if kind == "own" else "steam", bias=self.bias, lesson_deg=left, look_t=t,
eyes=eyes, weights=[round(w, 3) for w in weights.weights()])
limit = OWN_LESSON_MAX if kind == "own" else LESSON_MAX
if t is None:
rec["refused"] = "that gaze isn't in the last few seconds sent"
elif left > limit:
rec["refused"] = f"more than {limit} deg off"
if "refused" in rec:
return self.taken(rec)
if kind == "own":
# Its eyes come calibrated: how far off each was is its miss. The click goes to it.
miss = [math.hypot(thy - e[0], thp - e[1]) if e else None for e in eyes]
try:
self.eyes_sock.sendto(f"click {t:.6f} {thy:.4f} {thp:.4f}".encode(), EYES_SOCKET)
except OSError as e:
rec["refused"] = f"our tracker isn't running ({e})"
return self.taken(rec)
else:
miss = []
for name, e in zip(SIDES, eyes):
if not e:
miss.append(None)
continue
cy, cp = self.correction(name, *e)
miss.append(math.hypot(thy - e[0] - cy, thp - e[1] - cp))
self.lives[name].add({"time": rec["time"], "hy": e[0], "hp": e[1], "dy": thy - e[0], "dp": thp - e[1],
"wy": 1.0, "wp": 1.0, "how": "pointer"}, self.models[name], self.mode)
rec["miss"] = miss
weights.add(miss)
return self.taken(rec)
def taken(self, rec):
if "refused" in rec:
self.counts["refused"] += 1
else:
self.live.add({"time": rec["time"], "hy": rhy, "hp": rhp, "dy": dy, "dp": dp, "wy": 1.0, "wp": 1.0,
"how": "pointer"}, self.base, self.mode)
self.counts["lessons_taken"] += 1
self.dirty = True
try:
@@ -214,6 +399,61 @@ class Service:
pass
self.proc = None
def eyes_wanted(self):
return (self.tracker == "own" and not self.override) or time.monotonic() < self.eyes_until
def start_eyes(self):
"""Our own tracker, in the dev container, with build/venv's numpy and OpenCV. Like
ft-gaze, it quits when its stdin closes."""
if not EYES_PYTHON.exists():
log(f"ft-eyes isn't built: run {REPO}/gaze/tracker/build.sh")
self.eyes_restart_at = time.monotonic() + 30
return
env = dict(os.environ)
env["XDG_RUNTIME_DIR"] = f"/run/user/{os.getuid()}"
subprocess.run([str(REPO / "scripts" / "container-up.sh")], env=env, check=False)
distrobox = Path.home() / ".local" / "bin" / "distrobox"
self.eyes_proc = subprocess.Popen([str(distrobox), "enter", "dev", "--", str(EYES_PYTHON), str(EYES_PROG), "-v",
"--watch-stdin"], env=env, stdin=subprocess.PIPE,
stdout=subprocess.DEVNULL, stderr=subprocess.PIPE, start_new_session=True)
os.set_blocking(self.eyes_proc.stderr.fileno(), False)
self.sel.register(self.eyes_proc.stderr, selectors.EVENT_READ, "eyes")
log("ft-eyes started" + ("" if EYES_CAMS.exists() else
f": no {EYES_CAMS} yet (the frame grabber: gaze/tracker/install.sh)"))
def stop_eyes(self):
if not self.eyes_proc:
return
try:
self.sel.unregister(self.eyes_proc.stderr)
except (KeyError, ValueError):
pass
if self.eyes_proc.stdin and not self.eyes_proc.stdin.closed:
self.eyes_proc.stdin.close()
try:
self.eyes_proc.wait(timeout=3)
except subprocess.TimeoutExpired:
try:
os.killpg(self.eyes_proc.pid, signal.SIGTERM)
except ProcessLookupError:
pass
self.eyes_proc = None
self.own = {}
def read_eyes(self):
try:
data = os.read(self.eyes_proc.stderr.fileno(), 65536)
except BlockingIOError:
return
if not data:
log(f"ft-eyes stopped (exit {self.eyes_proc.poll()}); again in {EYES_RETRY:.0f} s if still wanted")
self.stop_eyes()
self.eyes_restart_at = time.monotonic() + EYES_RETRY
return
for line in data.decode("utf-8", "replace").splitlines():
if line.strip() and (self.verbose or "fps" not in line):
log(line)
def read_stdout(self):
try:
data = os.read(self.proc.stdout.fileno(), 65536)
@@ -241,19 +481,13 @@ class Service:
if line.strip():
log(line)
def on_sample(self, s):
src = s["src"].get(self.source) or {}
if "hy" not in src:
return
self.counts["samples"] += 1
self.last_sample = time.monotonic()
m1 = s["src"].get("mmap1") or {}
o = m1.get("open")
lr = src.get("lr", m1.get("lr"))
# Blinks and lost eyes, judged against the last second (see steady_samples: relative,
# because the lids come down looking down, and the vergence depends on distance).
# An eye's floor comes from its good readings, so a lost eye doesn't drag it to 0.
def judge_eyes(self, m1, down):
"""Which eyes (left, right) are closed, from SteamVR's openness (set 1); updates
self.lost from its variances. Blinks and lost eyes are judged against the last second
(see steady_samples: relative, because the lids come down looking down). An eye's
floor comes from its good readings, so a lost eye doesn't drag it to 0."""
low = [False, False]
o = m1.get("open")
if o and len(o) == 2:
for k in (0, 1):
hist = self.opens[k]
@@ -261,25 +495,130 @@ class Service:
floor = max(0.12, 0.5 * statistics.median(good)) if len(good) >= 30 else 0.12
low[k] = o[k] < floor
hist.append(o[k])
if all(low):
unc = m1.get("unc")
if unc and len(unc) == 2:
for k in (0, 1):
self.lost[k] = unc[k] > (EYE_FOUND if self.lost[k] else EYE_LOST)
if not down:
self.counts["lost_left"] += self.lost[0]
self.counts["lost_right"] += self.lost[1]
return low
def on_sample(self, s):
kind = self.kind
if kind != self.last_kind:
log({"own": "our own tracker", "eyes": "SteamVR's eyes, each calibrated",
"source": f"SteamVR's {self.source}, calibrated as a whole"}[kind]
+ (f", eye bias {self.bias}" if kind != "source" else ""))
self.last_kind = kind
self.fix.reset()
if kind == "source":
self.on_source_sample(s)
else:
self.on_eyes_sample(s, kind == "own")
def on_eyes_sample(self, s, own):
m1 = s["src"].get("mmap1") or {}
if own:
src = s["src"].get("own") or {}
if "hy" not in src:
return
eyes = [tuple(e) if e else None for e in (src.get("eyes") or [None, None])]
hp, hit = src["hp"], src.get("hit")
else:
per = [s["src"].get(name) or {} for name in SIDES]
eyes = [(p["hy"], p["hp"]) if "hy" in p else None for p in per]
if not any(eyes):
return
hp, hit = next(e[1] for e in eyes if e), m1.get("hit")
self.counts["samples"] += 1
self.last_sample = time.monotonic()
down = hp < KEYBOARD_PITCH and not hit
low = self.judge_eyes(m1, down)
if down:
self.counts["looking_down"] += 1
return
# Our tracker finds the pupils itself; SteamVR's openness still marks the blinks.
bad = [eyes[k] is None or low[k] or (not own and self.lost[k]) for k in (0, 1)]
if all(bad):
self.counts["blinks"] += 1
return
if any(low):
if any(bad):
self.counts["one_eye"] += 1
if self.source == "mmap2":
jump = (lr is not None and len(self.vergence) >= 30
and abs(lr - statistics.median(self.vergence)) > 1.5)
if lr is not None and not any(low):
self.vergence.append(lr)
if any(low) or jump:
self.counts["one_eye_used"] += 1
seen = [None if bad[k] else eyes[k] for k in (0, 1)]
if own:
corrected = seen
else:
corrected = []
for name, e in zip(SIDES, seen):
c = self.correction(name, *e) if e else None
corrected.append((e[0] + c[0], e[1] + c[1]) if e else None)
gy, gp = self.weights["own" if own else "steam"].combine(corrected)
fy, fp = self.fix(gy, gp, s["t"], 1.0)
self.send(s["t"], fy, fp, fy, fp, seen)
def on_source_sample(self, s):
src = s["src"].get(self.source) or {}
if "hy" not in src:
return
self.counts["samples"] += 1
self.last_sample = time.monotonic()
m1 = s["src"].get("mmap1") or {}
lr = src.get("lr", m1.get("lr"))
down = src["hp"] < KEYBOARD_PITCH and not src.get("hit")
low = self.judge_eyes(m1, down)
if down:
self.counts["looking_down"] += 1
for k in (0, 1):
self.bad_at[k] = s["t"] # the fallback doesn't learn from these either
return
bad = [low[k] or self.lost[k] for k in (0, 1)]
if all(bad):
self.counts["blinks"] += 1
return
hy, hp = src["hy"], src["hp"]
eyes = (s["src"].get("mmap2") or {}).get("eyes")
for k in (0, 1):
if bad[k]:
self.bad_at[k] = s["t"]
if any(bad):
self.counts["one_eye"] += 1
seen = 1 if bad[0] else 0
est = self.fallback.get(seen, eyes[seen][0], eyes[seen][1]) if eyes else None
if est:
hy, hp = est
self.counts["one_eye_used"] += 1
elif self.source == "mmap2":
self.counts["dropped"] += 1
return
else:
if self.source == "mmap2":
jump = (lr is not None and len(self.vergence) >= 30
and abs(lr - statistics.median(self.vergence)) > 1.5)
if lr is not None:
self.vergence.append(lr)
if jump:
self.counts["dropped"] += 1
return
# Learn only once both have been seen for a moment: the tracker's filter starts
# an eye over when it finds it again.
if eyes and s["t"] - max(self.bad_at) > SETTLE:
for k in (0, 1):
self.fallback.update(k, eyes[k][0], eyes[k][1], hy, hp)
# The fixation lock works in degrees here (1 degree per "pixel").
fy, fp = self.fix(src["hy"], src["hp"], s["t"], 1.0)
cy, cp = self.correction(fy, fp)
self.last = (fy + cy, fp + cp, fy, fp)
fy, fp = self.fix(hy, hp, s["t"], 1.0)
cy, cp = self.correction(self.source, fy, fp)
self.send(s["t"], fy + cy, fp + cp, fy, fp, None)
def send(self, t, hy, hp, rhy, rhp, eyes):
self.last = (hy, hp, rhy, rhp)
raw = f"{rhy:.3f} {rhp:.3f}"
self.history.append((t, raw, eyes))
while self.history and self.history[0][0] < t - HISTORY:
self.history.popleft()
try:
self.out.sendto(f"gz {fy + cy:.3f} {fp + cp:.3f} {fy:.3f} {fp:.3f}".encode(), self.to)
self.out.sendto(f"gz {hy:.3f} {hp:.3f} {raw}".encode(), self.to)
self.counts["sent"] += 1
except OSError:
pass # the pointer helper isn't running
@@ -299,19 +638,30 @@ class Service:
rec = self.lesson(*map(float, words[1:]))
reply = "refused" if "refused" in rec else f"ok {rec['lesson_deg']:.2f}"
log(f"lesson {rec['lesson_deg']:.2f} deg at {rec['raw'][0]:+.1f},{rec['raw'][1]:+.1f}"
+ (f", eyes off {', '.join('-' if m is None else f'{m:.2f}' for m in rec['miss'])}"
if rec.get("miss") else "")
+ (f": {rec['refused']}" if "refused" in rec else ""))
except ValueError:
reply = "error bad lesson"
elif words[:1] == ["status"]:
reply = json.dumps(self.status())
elif words[:1] == ["forget"]:
self.live = PointerLessons()
self.live.wear_time = self.steam.worn()
self.lives = {name: PointerLessons() for name in SOURCES}
self.weights = {t: EyeWeights(self.bias) for t in TRACKERS}
self.refit()
self.save_lessons()
reply = "ok"
elif words[:1] == ["eyes"] and len(words) == 2:
try:
secs = min(max(float(words[1]), 0.0), EYES_LEASE_MAX)
self.eyes_until = max(self.eyes_until, time.monotonic() + secs)
reply = "ok"
except ValueError:
reply = "error bad seconds"
elif words[:1] == ["reload"]:
self.load_settings()
self.load_calibration()
self.live.refit(self.base, self.mode)
self.refit()
reply = "ok"
else:
reply = "error unknown command"
@@ -321,27 +671,72 @@ class Service:
except OSError:
pass
def on_own(self):
"""Replies from our tracker: its status (JSON), or a click's "ok ..."/"fail ..."."""
while True:
try:
data = self.eyes_sock.recv(4096).decode("utf-8", "replace")
except (BlockingIOError, OSError):
return
if data.startswith("{"):
try:
self.own, self.own_at = json.loads(data), time.monotonic()
except ValueError:
pass
else:
log(f"our tracker: {data}")
def status(self):
ly, lp = self.live.offset()
return {"source": self.source, "model": self.mode, "calibration_samples": self.base.samples,
"lessons": len(self.live.samples), "lesson_offset": [round(ly, 3), round(lp, 3)],
"ft_gaze": self.proc is not None, "sample_age_s": round(time.monotonic() - self.last_sample, 2)
if self.last_sample else None, "headset_on": self.steam.wearing(), "headset_on_since": self.steam.worn(),
"last": [round(v, 2) for v in self.last] if self.last else None, **self.counts}
kind = self.kind
tracker = "own" if kind == "own" else "steam"
w = self.weights[tracker]
st = {"tracker": tracker, "kind": kind, "source": "own" if kind == "own" else self.source if kind == "source"
else "left+right", "model": self.mode, "eye_bias": self.bias}
if kind == "source":
ly, lp = self.lives[self.source].offset()
st.update(calibration_samples=self.models[self.source].samples, lessons=len(self.lives[self.source].samples),
lesson_offset=[round(ly, 3), round(lp, 3)])
else:
st.update(eye_weights=[round(v, 3) for v in w.weights()], eye_misses=[len(m) for m in w.misses],
eye_rms=[None if r is None else round(r, 2) for r in w.rms()])
if kind == "eyes":
st.update(calibration_samples=min(self.models[e].samples for e in SIDES),
lessons=max(len(self.lives[e].samples) for e in SIDES))
if kind == "own":
own = self.own if time.monotonic() - self.own_at < 5 else {}
cal = own.get("calibration") or {}
st.update(calibration_samples=cal.get("dots", 0), calibration_made=cal.get("made"),
lessons=max(len(m) for m in w.misses), own_running=bool(own),
own_reseat=any(e.get("reseat") for e in own.get("eyes", {}).values()))
st.update(eyes_process=self.eyes_proc is not None, eyegrab=EYES_CAMS.exists())
st.update({"ft_gaze": self.proc is not None, "sample_age_s": round(time.monotonic() - self.last_sample, 2)
if self.last_sample else None, "headset_on": self.steam.wearing(),
"headset_on_since": self.steam.worn(), "last": [round(v, 2) for v in self.last] if self.last else None,
"eyes_lost": self.lost, "fallback_ready": [self.fallback.ready(0), self.fallback.ready(1)],
**self.counts})
return st
def periodic(self):
if self.steam.poll() or self.steam.worn() != self.live.wear_time:
if self.steam.worn() != self.live.wear_time:
if self.steam.poll() or any(lv.wear_time != self.steam.worn() for lv in self.lives.values()):
if any(lv.wear_time != self.steam.worn() for lv in self.lives.values()):
log("headset on again: older lessons count less until new ones come in")
self.live.wear_time = self.steam.worn()
self.live.refit(self.base, self.mode)
try:
mtime = CALIBRATION.stat().st_mtime
except OSError:
mtime = None
if mtime != self.cal_mtime:
self.refit()
if mtime(CALIBRATION) != self.cal_mtime:
self.load_calibration()
self.live.refit(self.base, self.mode)
self.refit()
if mtime(CONF) != self.conf_mtime:
self.load_settings()
want = self.eyes_wanted()
if want and not self.eyes_proc and time.monotonic() >= self.eyes_restart_at:
self.start_eyes()
elif not want and self.eyes_proc:
log("ft-eyes no longer wanted: stopping it")
self.stop_eyes()
if self.eyes_proc:
try:
self.eyes_sock.sendto(b"status", EYES_SOCKET)
except OSError:
pass # not up yet: status() says so once the last answer is old
if self.dirty:
self.save_lessons()
@@ -355,6 +750,10 @@ class Service:
for key, _ in self.sel.select(timeout=0.5):
if key.data == "control":
self.on_control()
elif key.data == "own":
self.on_own()
elif key.data == "eyes" and self.eyes_proc:
self.read_eyes()
elif key.data == "stdout" and self.proc:
self.read_stdout()
elif key.data == "stderr" and self.proc:
@@ -366,13 +765,15 @@ class Service:
log(json.dumps(self.status()))
next_verbose = now + 5
self.stop_helper()
self.stop_eyes()
if self.dirty:
self.save_lessons()
def main():
ap = argparse.ArgumentParser(description="The gaze service: corrected eye tracking for the pointer")
ap.add_argument("--source", choices=["action", "mmap1", "mmap2"], default="mmap1")
ap.add_argument("--source", choices=["action", "mmap1", "mmap2"],
help="the older one-source path with this SteamVR source, whatever the settings say")
ap.add_argument("-v", "--verbose", action="store_true")
ap.add_argument("--to", default="ft_pointer_helper", help="abstract socket to send the gaze to")
args = ap.parse_args()
+145 -3
View File
@@ -2,8 +2,8 @@
The correction models (Correction: the calibration fitted from calibration dots;
LiveCorrection: what clicks teach on the fly, on top of it), the smoothing filters, the
blink and dropout filter for one look at a spot, and SteamEyeLog, which follows SteamVR's
eye tracking log. Angles are head-relative degrees (yaw +left, pitch +up), as ft-gaze
blink and dropout filter for one look at a spot, EyeFallback (the gaze from one eye while
the tracker has lost the other), EyeWeights (how much each eye counts), and SteamEyeLog, which follows SteamVR's eye tracking log. Angles are head-relative degrees (yaw +left, pitch +up), as ft-gaze
reports them.
"""
@@ -391,6 +391,146 @@ class LiveCorrection:
return self.cy[0], self.cp[0]
# The tracker's variance for an eye's direction (ft-gaze's "unc"): 0.0005-0.002 while it
# sees the eye, 0.015-0.03 once it's lost it, falling back through 0.008-0.002 in the 0.1 s
# after it finds it again.
EYE_LOST = 0.004
EYE_FOUND = 0.0025
class EyeFallback:
"""The gaze from one eye, while the tracker has lost the other.
SteamVR's combined gaze (mmap set 1) keeps going with one eye lost, but badly: it holds
the lost eye's yaw where it was and gives it the other eye's pitch, so the gaze moves
half as far sideways as the eyes do (seen: the right eye swung 5 degrees, the combined
gaze 2.5). Set 2's eyes are each eye's own reading. While both are seen, this learns what
each eye reads against the combined gaze (an offset: half the angle between the eyes,
plus how differently the tracker reads each), in 10 degree cells of where that eye
looks, blended over the four nearest; while one is lost, the other eye plus its offset
stands in for the combined gaze. So the rest (fixation lock, calibration, lessons)
carries on as if nothing happened.
On a recording, one eye alone came out 1.1 degrees (median) from both eyes' gaze, 0.8
over a tenth of a second of a steady look, and a little more jittery (0.31-0.37 degrees
against 0.28). Carrying on the offset from just before a loss did no better: what's
left is fast noise, not something particular to that look.
`update` and `get` take head-relative degrees (yaw, pitch)."""
CELL = 10.0
GLOBAL_RATE = 0.01 # per sample: about a second at 90 Hz
CELL_RATE = 0.02 # the least a cell learns per sample, once it has CELL_FULL
CELL_FULL = 30 # samples before a cell counts fully
READY = 45 # samples of both eyes before an eye can stand in
def __init__(self):
self.glob = [None, None] # per eye: [oy, op]
self.seen = [0, 0]
self.cells = [{}, {}] # per eye: (i, j) -> [oy, op, n]
def ready(self, eye):
return self.seen[eye] >= self.READY
def update(self, eye, ey, ep, cy, cp):
oy, op = cy - ey, cp - ep
g = self.glob[eye]
if g is None:
self.glob[eye] = [oy, op]
else:
g[0] += self.GLOBAL_RATE * (oy - g[0])
g[1] += self.GLOBAL_RATE * (op - g[1])
self.seen[eye] += 1
key = (math.floor(ey / self.CELL), math.floor(ep / self.CELL))
c = self.cells[eye].setdefault(key, [oy, op, 0])
c[2] += 1
a = max(1.0 / c[2], self.CELL_RATE)
c[0] += a * (oy - c[0])
c[1] += a * (op - c[1])
def offset(self, eye, ey, ep):
g = self.glob[eye]
if g is None:
return None
# Bilinear over the four cells whose centres surround the point.
fy, fp = ey / self.CELL - 0.5, ep / self.CELL - 0.5
i0, j0 = math.floor(fy), math.floor(fp)
ty, tp = fy - i0, fp - j0
sy = sp = used = 0.0
for di, wi in ((0, 1 - ty), (1, ty)):
for dj, wj in ((0, 1 - tp), (1, tp)):
c = self.cells[eye].get((i0 + di, j0 + dj))
if c:
w = wi * wj * min(1.0, c[2] / self.CELL_FULL)
sy += w * c[0]
sp += w * c[1]
used += w
return sy + (1 - used) * g[0], sp + (1 - used) * g[1]
def get(self, eye, ey, ep):
"""The combined gaze from this eye's reading, or None before it has learned enough."""
if not self.ready(eye):
return None
oy, op = self.offset(eye, ey, ep)
return ey + oy, ep + op
class EyeWeights:
"""How much each eye (0 left, 1 right) counts in the gaze, for ft-gazed's eye bias.
Two eyes beat either one: their errors partly cancel. On 306 live clicks with our own
tracker (gaze/tracker, 2026-09-29) the eyes' sideways errors were correlated -0.37, and
the mean of both was 0.65 degrees off (median), the left eye alone 0.96, the right 1.11.
So a bias leans instead of choosing: "left" or "right" counts that eye LEAN times the
other (on those clicks, 2:1 toward the better eye cost about 0.03 degrees, toward the
worse one about 0.13). "auto" weights each
by the inverse square of its RMS miss at the last KEEP lessons, once both have MIN, and
alike until then. Each miss is measured before its lesson teaches anything, so each is a
fresh test. On SteamVR's own test (2026-09-29) its calibration dots said the left eye was
the better one and new spots said the right, so the misses come from lessons, not the fit.
An eye that isn't seen (None) drops out, and the other carries the gaze alone."""
LEAN = 2.0
KEEP = 20
MIN = 5
FLOOR = 0.3 # degrees: so one lucky run can't give an eye all the weight
STALE = 8.0 # degrees: a miss this big is the headset moved, not the eye's accuracy
def __init__(self, bias="auto", misses=None):
self.bias = bias
self.misses = [list(m) for m in (misses or ([], []))]
def add(self, miss):
"""One lesson's miss per eye (degrees, None where it wasn't seen)."""
if any(m is not None and m > self.STALE for m in miss):
return
for k, m in enumerate(miss):
if m is not None:
self.misses[k] = (self.misses[k] + [m])[-self.KEEP:]
def rms(self):
return [math.sqrt(sum(m * m for m in ms) / len(ms)) if ms else None for ms in self.misses]
def weights(self):
"""(left, right), summing to 1."""
if self.bias in ("left", "right"):
w = [self.LEAN, 1.0] if self.bias == "left" else [1.0, self.LEAN]
elif all(len(ms) >= self.MIN for ms in self.misses):
w = [1.0 / max(r, self.FLOOR) ** 2 for r in self.rms()]
else:
w = [1.0, 1.0]
return w[0] / sum(w), w[1] / sum(w)
def combine(self, eyes):
"""The weighted gaze from [(yaw, pitch) or None, (yaw, pitch) or None], or None."""
w = [wk for wk, e in zip(self.weights(), eyes) if e is not None]
seen = [e for e in eyes if e is not None]
if not seen:
return None
total = sum(w)
return (sum(wk * e[0] for wk, e in zip(w, seen)) / total, sum(wk * e[1] for wk, e in zip(w, seen)) / total)
class SteamEyeLog:
"""Follows SteamVR's eye tracking log (read only) for what moves the raw gaze under a
calibration.
@@ -494,7 +634,8 @@ def cross_validate(points, mode):
def steady_samples(samples, vergence_jump=1.5):
"""The samples of one look at one spot where the tracker had both eyes: none in a blink
(openness under half its median over the samples), and none where the angle between the eyes' directions (`lr`, the
(openness under half its median over the samples), none where it had lost an eye (its
variance over EYE_LOST), and none where the angle between the eyes' directions (`lr`, the
vergence) is more than `vergence_jump` degrees from its median over the samples. The
vergence itself depends on distance (about 2.8 degrees for a screen 1.3 m away, a
fraction of one far off), so only a jump away from what it was during this look means
@@ -513,6 +654,7 @@ def steady_samples(samples, vergence_jump=1.5):
def vergence(smp):
return (smp["src"].get("mmap1") or {}).get("lr", (smp["src"].get("mmap2") or {}).get("lr"))
opened = [smp for smp in opened if max((smp["src"].get("mmap1") or {}).get("unc") or [0]) <= EYE_LOST]
have = [v for v in map(vergence, opened) if v is not None]
if len(have) < 5:
return opened
+372 -41
View File
@@ -2,8 +2,8 @@
"""ft-gazeprobe: a playground for eye tracking as pointer input on the Frametop desktop.
Opens fullscreen on one Frametop screen and shows where the headset's eye tracker says
you're looking, from the three sources ft-gaze reads (SteamVR's eye tracking action and
the two gaze sets in eye-server.mmap). Three modes:
you're looking, from the sources ft-gaze reads (SteamVR's eye tracking action, the two
gaze sets in eye-server.mmap, and each eye alone). The modes:
Free look the gaze dot; the trigger calibrates wherever you are looking.
Accuracy test look at each target and tap the trigger; measures every source's error
@@ -15,6 +15,11 @@ the two gaze sets in eye-server.mmap). Three modes:
it; if it's the wrong one, hold, then glance toward the right one or
move the mouse, and let go on it. Each click teaches the click
corrections (LiveCorrection).
Headset fit how well the tracker sees each eye (fitcheck.py): live per eye, whether
it's tracked, how open it is, and the tracker's confidence, and maps of
where you looked and where each eye got lost, with hints. Enter runs a
guided check (dots around the screen, then down, up, left and right),
R starts over. Adjust the headset while you watch it.
Run calibration: the initial calibration, after Apple Vision Pro's eye setup. One dot,
then six in a circle, in three rounds that go from a dark to a bright screen (pupil size
@@ -22,8 +27,8 @@ changes with brightness, and the tracker's error with it). Look at each highligh
and press the trigger. At the end, each source's calibration is fitted from all 21 dots;
freeze and look refines it on demand after that.
Trigger: Enter, Space, or a mouse button. Tab shows and hides the panel, F11 toggles
fullscreen, Esc quits.
Trigger: Enter, Space, or a mouse button. Tab shows and hides the panel, C collapses it to
its title bar (or its arrow button does), F11 toggles fullscreen, Esc quits.
Freeze and look (the default trigger): the press freezes the dot where the tracker says
you're looking. Look at the frozen dot. After a moment to settle, the probe averages where
@@ -55,6 +60,7 @@ import math
import os
import random
import signal
import socket
import statistics
import subprocess
import sys
@@ -72,16 +78,61 @@ from gi.repository import Adw, Gdk, Gio, GLib, Gtk # noqa: E402
REPO = Path(__file__).resolve().parents[2]
HELPER = REPO / "gaze" / "build" / "ft-gaze"
STATE = Path.home() / ".local" / "state" / "frametop" / "gaze"
SOURCES = ["action", "mmap1", "mmap2"]
SOURCE_NAMES = {"action": "SteamVR action", "mmap1": "mmap set 1", "mmap2": "mmap set 2"}
SOURCE_COLORS = {"action": (0.2, 0.8, 1.0), "mmap1": (1.0, 0.6, 0.1), "mmap2": (0.9, 0.3, 0.9)}
# left and right: each eye alone (set 2's own reading of that eye), calibrated and tested
# like the rest, to see what one eye is worth against both. own: our own tracker
# (gaze/tracker/ft-eyes, which the gaze service runs); it has its own calibration.
SOURCES = ["action", "mmap1", "mmap2", "left", "right", "own"]
SOURCE_NAMES = {"action": "SteamVR action", "mmap1": "mmap set 1", "mmap2": "mmap set 2", "left": "Left eye",
"right": "Right eye", "own": "Own tracker"}
SOURCE_COLORS = {"action": (0.2, 0.8, 1.0), "mmap1": (1.0, 0.6, 0.1), "mmap2": (0.9, 0.3, 0.9),
"left": (0.4, 1.0, 0.6), "right": (1.0, 1.0, 0.4), "own": (1.0, 0.35, 0.35)}
TRIGGER_KEYS = {Gdk.KEY_Return, Gdk.KEY_KP_Enter, Gdk.KEY_space}
class OwnTracker:
"""Talks to our own tracker (gaze/tracker/ft-eyes) over its control socket (@ft_eyes).
It keeps its own calibration: the probe sends it calibration dots and clicks, each with
when you looked and where (head-relative degrees), and it learns from its own frames.
The gaze service (ft-gazed) runs it: `lease` asks it to keep it running a while longer,
whatever the Eye tracker setting says."""
def __init__(self):
self.sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
self.sock.bind(f"\0ft_gazeprobe.{os.getpid()}")
self.sock.settimeout(1.0)
def ask(self, command):
"""The reply line, or "fail ..." if the tracker isn't running or doesn't answer."""
try:
while True: # drop a late reply to an earlier command
self.sock.setblocking(False)
self.sock.recv(4096)
except (BlockingIOError, OSError):
pass
self.sock.settimeout(1.0)
try:
self.sock.sendto(command.encode(), "\0ft_eyes")
return self.sock.recv(4096).decode()
except (ConnectionRefusedError, FileNotFoundError):
return "fail the Own tracker isn't running yet (the gaze service starts it)"
except (socket.timeout, OSError) as e:
return f"fail no answer from the Own tracker ({e})"
def lease(self, seconds=30):
"""Ask the gaze service to keep the Own tracker running for `seconds` more. False if
the gaze service isn't running."""
try:
self.sock.sendto(f"eyes {seconds}".encode(), "\0ft_gazed")
return True
except OSError:
return False
# The math, filters, correction models, and SteamVR log reader are shared with ft-gazed.
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from gazecal import (DEFAULT_MODEL, MODELS, Correction, Fixation, LiveCorrection, OneEuro, # noqa: E402
SteamEyeLog, cross_validate, deg_from_px, px_from_deg, steady_samples)
from fitcheck import FitCheck # noqa: E402
# --- Gaze from ft-gaze ----------------------------------------------------------------
@@ -414,8 +465,12 @@ class Probe(Adw.ApplicationWindow):
self.pointer = None
self.practice = None # click practice: the press being held (see practice_press)
self.recent = deque(maxlen=60) # the last samples, for where you looked at a press
self.fitcheck = FitCheck() # Headset fit: how well the tracker sees each eye
self.models = {s: Correction() for s in SOURCES}
self.own = OwnTracker()
self.reseat_check = False # showing the one-dot check after the headset was off
self.reseat_skipped = False # S skipped this one
self.live = {s: LiveCorrection() for s in SOURCES}
self.steam = SteamEyeLog()
self.steam.poll()
@@ -471,8 +526,8 @@ class Probe(Adw.ApplicationWindow):
# --- The context menu (right-click, the Menu key, or Shift+F10) ---
MODE_KEYS = ["free", "test", "practice", "snap"]
MODE_NAMES = ["Free look", "Accuracy test", "Click practice", "Snap practice"]
MODE_KEYS = ["free", "test", "practice", "snap", "fit"]
MODE_NAMES = ["Free look", "Accuracy test", "Click practice", "Snap practice", "Headset fit"]
def build_menu(self):
def action(name, fn, state=None):
@@ -501,6 +556,8 @@ class Probe(Adw.ApplicationWindow):
self.mode_action.connect("activate", lambda a, v: self.set_mode(v.get_string()))
self.add_action(self.mode_action)
self.panel_action = action("panel", self.on_panel_toggle, GLib.Variant("b", True))
self.collapse_action = action("collapse", lambda a, v: self.set_collapsed(v.get_boolean()),
GLib.Variant("b", False))
self.full_action = action("fullscreen", lambda a, v: self.toggle_fullscreen(), GLib.Variant("b", True))
action("quit", lambda *_: self.close())
@@ -527,6 +584,7 @@ class Probe(Adw.ApplicationWindow):
menu.append_submenu("Correction model", models)
view = Gio.Menu()
view.append("Panel", "win.panel")
view.append("Collapse panel", "win.collapse")
view.append("Fullscreen", "win.fullscreen")
view.append("Quit", "win.quit")
menu.append_section(None, view)
@@ -655,16 +713,29 @@ class Probe(Adw.ApplicationWindow):
return s
def build_panel(self):
box = Gtk.Box(orientation=Gtk.Orientation.VERTICAL, spacing=8, halign=Gtk.Align.START,
valign=Gtk.Align.START, margin_start=40, margin_top=40)
box.add_css_class("probe-panel")
box.set_size_request(460, -1)
panel = Gtk.Box(orientation=Gtk.Orientation.VERTICAL, halign=Gtk.Align.START,
valign=Gtk.Align.START, margin_start=40, margin_top=40)
panel.add_css_class("probe-panel")
title = Gtk.Label(label="Gaze probe", xalign=0)
# The title bar stays when the panel is collapsed, so the gaze dot isn't lost behind it.
head = Gtk.Box(spacing=12)
title = Gtk.Label(label="Gaze probe", xalign=0, hexpand=True)
title.add_css_class("title-2")
box.append(title)
hint = Gtk.Label(label="Trigger: Enter, Space, or click. Right-click for the menu. Tab hides this panel, Esc quits. "
"Click the window once so the keys reach it.", xalign=0, wrap=True)
head.append(title)
self.w_collapse = Gtk.Button(icon_name="pan-up-symbolic", tooltip_text="Collapse the panel (C)")
self.w_collapse.connect("clicked", lambda *_: self.set_collapsed(not self.collapsed))
head.append(self.w_collapse)
panel.append(head)
box = Gtk.Box(orientation=Gtk.Orientation.VERTICAL, spacing=8, margin_top=8)
box.set_size_request(460, -1)
self.panel_body = box
panel.append(box)
# Wrapped labels ask for their whole text on one line, which made the panel as wide
# as the screen; max_width_chars keeps it to about the grid's width.
hint = Gtk.Label(label="Trigger: Enter, Space, or click. Right-click for the menu. Tab hides this panel, "
"C collapses it, Esc quits. Click the window once so the keys reach it.",
xalign=0, wrap=True, max_width_chars=45)
hint.add_css_class("dim-label")
box.append(hint)
@@ -677,12 +748,22 @@ class Probe(Adw.ApplicationWindow):
grid.attach(widget, 1, row, 1, 1)
row += 1
# Which tracker drives the dot. SteamVR's gaze is still logged in the background with
# the Own tracker (for comparing), but not shown.
tracker = Gtk.Box(css_classes=["linked"])
self.w_steamvr = Gtk.ToggleButton(label="SteamVR", active=True)
self.w_own = Gtk.ToggleButton(label="Own tracker", group=self.w_steamvr)
tracker.append(self.w_steamvr)
tracker.append(self.w_own)
self.w_own.connect("toggled", lambda b: self.on_tracker())
add("Tracker", tracker)
self.w_screen = self.dropdown(["Screen 1"])
add("Screen", self.w_screen)
self.w_mode = self.dropdown(self.MODE_NAMES)
add("Mode", self.w_mode)
# mmap set 2 had the least jitter (0.25 deg vs 0.30 for the action, sitting still).
self.w_source = self.dropdown([SOURCE_NAMES[s] for s in SOURCES], 2)
self.w_source.connect("notify::selected", lambda *_: self.sync_tracker())
add("Source", self.w_source)
self.w_filter = self.dropdown(["Raw", "One Euro", "Fixation lock"], 2)
add("Smoothing", self.w_filter)
@@ -740,10 +821,31 @@ class Probe(Adw.ApplicationWindow):
buttons.append(b)
box.append(buttons)
self.w_stats = Gtk.Label(xalign=0, yalign=0, wrap=True, selectable=False)
self.w_stats = Gtk.Label(xalign=0, yalign=0, wrap=True, selectable=False, max_width_chars=50)
self.w_stats.add_css_class("probe-stats")
box.append(self.w_stats)
return box
return panel
@property
def collapsed(self):
return not self.panel_body.get_visible()
def set_collapsed(self, collapsed):
self.panel_body.set_visible(not collapsed)
self.w_collapse.set_icon_name("pan-down-symbolic" if collapsed else "pan-up-symbolic")
self.w_collapse.set_tooltip_text("Expand the panel (C)" if collapsed else "Collapse the panel (C)")
self.collapse_action.set_state(GLib.Variant("b", collapsed))
def under_panel(self, x0, y0, x1, y1, margin=20):
"""Whether a box on the canvas overlaps the panel, collapsed or not."""
if not self.panel.get_visible():
return False
# Collapsed: the title bar. Before the first layout: about the full panel.
px, py, pw, ph = (40, 40, 240, 70) if self.collapsed else (40, 40, 520, 960)
ok, r = self.panel.compute_bounds(self.area)
if ok and r.get_width() > 0:
px, py, pw, ph = r.get_x(), r.get_y(), r.get_width(), r.get_height()
return x0 < px + pw + margin and px - margin < x1 and y0 < py + ph + margin and py - margin < y1
def place(self):
"""Find the Frametop screens and go fullscreen on the chosen one."""
@@ -818,6 +920,30 @@ class Probe(Adw.ApplicationWindow):
n = len(self.live[self.source].samples)
return self.model_mode + (f" + {n} clicks" if n and self.w_snaps.get_active() else "")
def on_tracker(self):
"""The tracker toggle: the Own tracker, or SteamVR's set 2 (its quietest source)."""
own = self.w_own.get_active()
if own != (self.source == "own"):
self.w_source.set_selected(SOURCES.index("own") if own else SOURCES.index("mmap2"))
if own:
if not self.own.lease():
self.on_status("The gaze service isn't running, and it runs the Own tracker "
"(frametop-gaze.service, gaze/run.sh install)")
return
self.lease_at = time.monotonic()
reply = self.own.ask("status")
if reply.startswith("fail"):
self.on_status("Starting the Own tracker (a few seconds)…")
else:
st = json.loads(reply)
self.on_status("Own tracker: " + ("calibrated " + st["calibration"].get("made", "")
if st.get("calibration") else "not calibrated yet: Run calibration"))
def sync_tracker(self):
own = self.source == "own"
if self.w_own.get_active() != own:
(self.w_own if own else self.w_steamvr).set_active(True)
def on_setting(self):
model = self.model_mode
if model != getattr(self, "last_model", model):
@@ -843,6 +969,9 @@ class Probe(Adw.ApplicationWindow):
self.snap = None
elif not self.snap:
self.snap_layout()
if self.mode == "fit" and getattr(self, "last_mode", None) != "fit" and hasattr(self, "collapse_action"):
self.set_collapsed(True) # the fit check needs the room
self.last_mode = self.mode
if hasattr(self, "mode_action"):
self.mode_action.set_state(GLib.Variant("s", self.mode))
self.update_stats()
@@ -860,15 +989,27 @@ class Probe(Adw.ApplicationWindow):
if keyval == Gdk.KEY_Tab:
self.panel.set_visible(not self.panel.get_visible())
return True
if keyval in (Gdk.KEY_c, Gdk.KEY_C):
self.panel.set_visible(True)
self.set_collapsed(not self.collapsed)
return True
if keyval == Gdk.KEY_F11:
self.toggle_fullscreen()
return True
if keyval in (Gdk.KEY_r, Gdk.KEY_R) and self.mode == "fit":
self.fitcheck.reset()
self.area.queue_draw()
return True
if keyval == Gdk.KEY_BackSpace and self.mode == "snap":
self.snap_undo()
return True
if keyval in (Gdk.KEY_s, Gdk.KEY_S) and self.calib:
self.calib_skip()
return True
if keyval in (Gdk.KEY_s, Gdk.KEY_S) and self.reseat_check:
self.reseat_check, self.reseat_skipped = False, True
self.area.queue_draw()
return True
if keyval == Gdk.KEY_Escape:
if self.calib or self.test:
self.calib = self.test = None
@@ -910,12 +1051,69 @@ class Probe(Adw.ApplicationWindow):
if self.steam.poll():
self.on_status("SteamVR's eye tracker just started again: its own calibration started over, "
"so the raw gaze may have moved")
self.poll_own()
self.update_stats()
return True
def poll_own(self):
"""With the Own tracker: after the headset was off (or the tracker restarted), its
next click starts each eye's shift over, and until then the first clicks can be
10-17 degrees off (gaze/tracker/findings.md, session 5). So ask for one look at a centre dot first,
as Varjo's headsets do each time they're put on. Also keeps the lease on the tracker."""
if self.source != "own":
return
if time.monotonic() - getattr(self, "lease_at", 0.0) > 10:
self.lease_at = time.monotonic()
self.own.lease()
if self.calib or self.capture:
return
reply = self.own.ask("status")
if reply.startswith("fail"):
return
st = json.loads(reply)
pending = bool(st.get("calibration")) and any(e.get("reseat") for e in st["eyes"].values())
if not pending:
self.reseat_skipped = False
check = pending and not self.reseat_skipped
if check != self.reseat_check:
self.reseat_check = check
self.area.queue_draw()
def reseat_press(self):
"""The press on the check dot: a click there teaches both eyes' shifts at once."""
if not self.recent:
return
t_end = self.recent[-1]["t"]
w, h = self.canvas_size()
d = self.screen_direction([smp for smp in self.recent if smp["t"] >= t_end - 0.6], (w / 2, h / 2))
if d is None:
self.on_status("no gaze samples on this screen: look at the dot and press again")
return
reply = self.own.ask(f"click {t_end:.6f} {d[0]:.4f} {d[1]:.4f}")
if reply.startswith("ok"):
self.reseat_check = False
self.on_status("Own tracker: checked after the headset was off")
else:
self.on_status(reply[5:] if reply.startswith("fail") else reply)
self.area.queue_draw()
def draw_reseat(self, cr, w, h):
cx, cy = w / 2, h / 2
cr.set_source_rgb(1, 1, 1)
cr.arc(cx, cy, 10, 0, 2 * math.pi)
cr.fill()
cr.set_source_rgb(0.1, 0.1, 0.1)
cr.arc(cx, cy, 3, 0, 2 * math.pi)
cr.fill()
ink = (0.9, 0.9, 0.9)
self.text(cr, 60, 70, "Own tracker: the headset was off, so it may sit differently now", ink, 26)
self.text(cr, 60, 108, self.status if self.status.startswith("no gaze") else
"Look at the dot and press (Enter, Space, or click). S skips.", ink, 20)
def on_sample(self, s):
self.sample = s
self.recent.append(s)
self.fitcheck.feed(s, time.monotonic())
self.rate_count += 1
self.last_arrival = time.monotonic()
t = s["t"]
@@ -931,6 +1129,9 @@ class Probe(Adw.ApplicationWindow):
head = s["head"].get("hit")
ox, oy = self.origin or (0, 0)
self.head = (head["x"] - ox, head["y"] - oy) if head and head["s"] == self.screen else None
# The Own tracker's eyes, each where it alone puts your gaze (left, right).
self.own_eyes = [(h["x"] - ox, h["y"] - oy) if h and h["s"] == self.screen else None
for h in ((s["src"].get("own") or {}).get("ehit") or [])]
r = self.raw.get(self.source)
if r is None:
@@ -965,7 +1166,10 @@ class Probe(Adw.ApplicationWindow):
def correction(self, name, hy, hp, snaps=None):
"""The whole correction at (hy, hp): the calibration, plus what snap and practice
clicks have taught since (when "Learn from clicks" is on)."""
clicks have taught since (when "Learn from clicks" is on). None for the Own tracker:
it learns from the clicks itself (own_click)."""
if name == "own":
return 0.0, 0.0
cy, cp = self.models[name].get(hy, hp, self.model_mode)
if snaps if snaps is not None else self.w_snaps.get_active():
ly, lp = self.live[name].get(hy, hp)
@@ -1094,6 +1298,9 @@ class Probe(Adw.ApplicationWindow):
if self.calib:
self.calib_press()
return
if self.reseat_check:
self.reseat_press()
return
if self.mode == "test":
self.test_press()
return
@@ -1103,6 +1310,10 @@ class Probe(Adw.ApplicationWindow):
if self.mode == "practice":
self.practice_press()
return
if self.mode == "fit":
self.fitcheck.toggle_guide(time.monotonic())
self.area.queue_draw()
return
if self.action == "freeze":
if not self.capture:
self.start_capture()
@@ -1131,7 +1342,7 @@ class Probe(Adw.ApplicationWindow):
return self.cursor
def on_release(self):
if self.mode == "test":
if self.mode in ("test", "fit"):
return
if self.mode == "snap":
self.snap_release()
@@ -1199,7 +1410,7 @@ class Probe(Adw.ApplicationWindow):
self.on_status("no gaze on this screen at the press")
return
self.practice = {"t": time.monotonic(), "F": self.gaze, "fix": fix, "pointer0": self.pointer,
"target": self.targets[0]}
"target": self.targets[0], "t_raw": self.recent[-1]["t"] if self.recent else None}
self.area.queue_draw()
def practice_point(self):
@@ -1240,6 +1451,13 @@ class Probe(Adw.ApplicationWindow):
# You were looking at the release point when you pressed: from the raw gaze to
# it is the whole error there.
oy, op = deg_from_px(fix["j"], px - fix["x"], py - fix["y"])
if name == "own":
rec["sources"][name] = {"hy": fix["hy"], "hp": fix["hp"], "raw": [fix["x"], fix["y"]],
"off": [oy, op], "n": fix["n"]}
if self.source == "own":
rec["own"] = self.own_click(p, fix["hy"] + oy, fix["hp"] + op, math.hypot(oy, op))
verdict = rec["own"]
continue
c = self.correction(name, fix["hy"], fix["hp"], snaps=True)
left = math.hypot(oy - c[0], op - c[1])
rec["sources"][name] = {"hy": fix["hy"], "hp": fix["hp"], "raw": [fix["x"], fix["y"]], "off": [oy, op],
@@ -1256,7 +1474,7 @@ class Probe(Adw.ApplicationWindow):
if learn:
self.save_calibration()
rec["verdict"] = verdict
rec["learned"] = verdict == "learned"
rec["learned"] = verdict == "learned" or verdict.startswith("taught")
self.attempts.append(rec)
self.log("practice.jsonl", rec)
# Drawn for a moment: the drag, from where the gaze put the pointer to where you let go.
@@ -1266,6 +1484,19 @@ class Probe(Adw.ApplicationWindow):
self.new_target()
self.update_stats()
def own_click(self, p, yaw, pitch, off):
"""Teach the Own tracker: you were looking at (yaw, pitch) just before the press.
Every click counts, dragged or not (a click that needed no drag says so too), unless
the drag was too far to be the tracker's error."""
if not self.w_snaps.get_active():
return "clicked (learning is off)"
if off > self.OWN_LEARN_MAX:
return f"dragged {off:.1f} deg: too far to be the tracker's error, not learned"
if p.get("t_raw") is None:
return "no sample time at the press"
reply = self.own.ask(f"click {p['t_raw']:.6f} {yaw:.4f} {pitch:.4f}")
return "taught the Own tracker" if reply.startswith("ok") else reply
def on_motion(self, x, y):
self.pointer = (x, y)
if self.practice:
@@ -1286,8 +1517,8 @@ class Probe(Adw.ApplicationWindow):
last = self.targets[0] if self.targets else None
for _ in range(50):
x, y = random.uniform(margin, w - margin), random.uniform(margin, h - margin)
if self.panel.get_visible() and x < 560 and y < 1000:
continue # not under the panel
if self.under_panel(x - r, y - r, x + r, y + r):
continue
if not last or math.hypot(x - last[0], y - last[1]) > min(w, h) * 0.25:
break
self.targets = [(x, y, r)]
@@ -1323,6 +1554,11 @@ class Probe(Adw.ApplicationWindow):
FLICK = 1.5 # degrees the gaze has to move from where it was at the press to step
REARM = 0.9 # ... and come back within, before the next glance steps again
SNAP_LEARN_MAX = 6.0 # degrees: a lesson bigger than this (after the correction) is a wrong element
# The Own tracker takes bigger ones: after the headset is taken off and put back on, its
# first clicks were 11-17 degrees off, and a 6-degree limit kept them from teaching it
# (gaze/tracker/findings.md, session 5). It keeps the median of an eye's last 5 clicks, so one click
# where you changed your mind does little harm.
OWN_LEARN_MAX = 25.0
def snap_layout(self):
w, h = self.canvas_size()
@@ -1368,8 +1604,8 @@ class Probe(Adw.ApplicationWindow):
corners = [(ox, oy), (ox + gw, oy), (ox, oy + gh), (ox + gw, oy + gh)]
if any(math.hypot(x - cx, y - cy) > radius * 1.05 for x, y in corners):
continue
if self.panel.get_visible() and ox < 560 and oy < 1000:
continue # not under the panel
if self.under_panel(ox, oy, ox + gw, oy + gh):
continue
if any(ox < b[2] + margin and b[0] < ox + gw + margin and oy < b[3] + margin and b[1] < oy + gh + margin
for b in boxes):
continue
@@ -1487,6 +1723,13 @@ class Probe(Adw.ApplicationWindow):
best, best_cost = k, cost
return best
@staticmethod
def steady_for(name, samples):
"""steady_samples, but for the Own tracker its own view of the eyes, not SteamVR's."""
if name != "own":
return steady_samples(samples)
return [smp for smp in samples if all((smp["src"].get("own") or {}).get("eyes") or [None])]
def press_fixation(self, name):
"""Where `name` put your gaze just before the press: the median of the last 300 ms,
without blinks, dropouts, or samples from before an eye movement in that time."""
@@ -1495,7 +1738,7 @@ class Probe(Adw.ApplicationWindow):
t_end = self.recent[-1]["t"]
pts = []
ox, oy = self.origin or (0, 0)
for smp in steady_samples([smp for smp in self.recent if smp["t"] >= t_end - 0.3]):
for smp in self.steady_for(name, [smp for smp in self.recent if smp["t"] >= t_end - 0.3]):
src = smp["src"].get(name) or {}
hit = src.get("hit")
if hit and hit["s"] == self.screen:
@@ -1669,12 +1912,28 @@ class Probe(Adw.ApplicationWindow):
def calib_points(self, rnd):
"""The centre, then six on a ring of `Calibration ring` degrees (as much of it as fits
in the window), turned 20 degrees per round, and half the size in the middle round,
so the fit sees the middle, halfway out, and the edge of your view."""
so the fit sees the middle, halfway out, and the edge of your view.
For the Own tracker: the centre, then eight directions on an oval out to `Calibration
ring` degrees each way (as much as fits across and up the window), turned 20 degrees
per round, and half the size in the middle round. Its fit is quadratic and goes
wrong past its dots, and the ring (limited by the window's height) left the sides
out: in practice2 (gaze/tracker/findings.md) the worst clicks were all past the ring."""
w, h = self.canvas_size()
cx, cy = w / 2, h / 2
r = self.raw.get(self.source)
px_per_deg = 1 / r[5] if r and r[5] else 32.0
radius = min(self.w_ring.get_value() * px_per_deg, w / 2 - 60, h / 2 - 60) * self.RING_SCALE[rnd]
reach = self.w_ring.get_value() * px_per_deg
if self.source == "own":
k = self.RING_SCALE[rnd]
rx = min(reach, cx - 80) * k
up, down = min(reach, cy - 150) * k, min(reach, cy - 80) * k # clear of the text at the top
pts = [(cx, cy)]
for i in range(8):
a = math.radians(-90 + rnd * 20 + i * 45)
pts.append((cx + rx * math.cos(a), cy + (up if math.sin(a) < 0 else down) * math.sin(a)))
return pts
radius = min(reach, w / 2 - 60, h / 2 - 60) * self.RING_SCALE[rnd]
pts = [(cx, cy)]
for k in range(6):
a = math.radians(-90 + rnd * 20 + k * 60)
@@ -1682,6 +1941,13 @@ class Probe(Adw.ApplicationWindow):
return pts
def start_calibration(self):
if self.source == "own":
# A fresh calibration for the Own tracker: it forgets the old one's clicks and
# shifts when this one is fitted.
reply = self.own.ask("calib-start")
if reply.startswith("fail"):
self.on_status(reply[5:])
return
self.test = None
self.results = None
self.capture = None
@@ -1744,7 +2010,9 @@ class Probe(Adw.ApplicationWindow):
"sources": {n: {k: g[k] for k in ("err_deg", "sd_deg", "off_deg", "mean", "n")}
for n, g in entry["sources"].items()}}
verdict = None
if not mine:
if self.source == "own":
verdict = self.own_calib_point(c["samples"], target)
elif not mine:
if len(steady) < 20:
verdict = (f"only {len(steady)} of {len(c['samples'])} samples had both eyes tracked "
"(blinks, or the tracker lost an eye): open your eyes wide and press again")
@@ -1766,11 +2034,44 @@ class Probe(Adw.ApplicationWindow):
c["retry"] = verdict + (" (S skips this dot)" if c["tries"] >= 2 else "")
return
c["tries"] = 0
c.setdefault("seen", {})[c["index"]] = tuple(mine["mean"])
if mine:
c.setdefault("seen", {})[c["index"]] = tuple(mine["mean"])
c["data"].append(entry)
c["done_at"] = now # a short pause on the filled dot, then the next one
GLib.timeout_add(350, self.calib_next)
def own_calib_point(self, samples, target):
"""Send one calibration dot to the Own tracker: the time you looked at it and its
direction. None if accepted, else why not."""
d = self.screen_direction(samples, target)
if d is None or len(samples) < 2:
return "no gaze samples on this screen: look at the dot and press again"
reply = self.own.ask(f"calib-point {samples[0]['t']:.6f} {samples[-1]['t']:.6f} {d[0]:.4f} {d[1]:.4f}")
if reply.startswith("ok"):
return None
return reply[5:] + ": look at the dot and press again"
def screen_direction(self, samples, target):
"""The head-relative direction (yaw, pitch) of a point on the canvas, from the screen
geometry around it: SteamVR's nearby gaze and its pixels-per-degree there (only the
geometry is used, not where SteamVR thinks you looked). None without samples."""
pts = []
ox, oy = self.origin or (0, 0)
for smp in samples:
for name in ("mmap1", "mmap2", "action"):
src = smp["src"].get(name) or {}
hit = src.get("hit")
if hit and hit["s"] == self.screen:
pts.append((hit["x"] - ox, hit["y"] - oy, hit["j"], src["hy"], src["hp"]))
break
if len(pts) < 5:
return None
x = statistics.median(q[0] for q in pts)
y = statistics.median(q[1] for q in pts)
j = [statistics.fmean(q[2][k] for q in pts) for k in range(4)]
oy_, op_ = deg_from_px(j, target[0] - x, target[1] - y)
return statistics.median(q[3] for q in pts) + oy_, statistics.median(q[4] for q in pts) + op_
def other_dot(self, c, g):
"""Was the gaze on another dot of this round rather than the current one?
@@ -1835,17 +2136,22 @@ class Probe(Adw.ApplicationWindow):
self.test_history = []
self.log_points("calibration", data)
self.save_calibration()
own_note = ""
if self.source == "own":
reply = self.own.ask("calib-fit")
own_note = ("Own tracker: " + reply[3:]) if reply.startswith("ok") else ("Own tracker: " + reply)
print(f"calibration: {own_note}", file=sys.stderr, flush=True)
if self.model_mode == "none":
self.last_model = mode # just fitted
self.w_model.set_selected(MODELS.index(mode))
if self.w_autotest.get_active():
# The check on spots the calibration hasn't seen, right away: same head
# position, same session of SteamVR's own calibration.
self.on_status(f"calibrated ({mode}) from {len(data)} dots; now testing it on new spots")
self.on_status(f"calibrated ({mode}) from {len(data)} dots; now testing it on new spots. {own_note}")
GLib.timeout_add(1500, lambda: (self.start_test(), False)[1])
else:
self.panel.set_visible(True)
self.on_status(f"calibrated ({mode}) from {len(data)} dots; Start test to check it")
self.on_status(f"calibrated ({mode}) from {len(data)} dots; Start test to check it. {own_note}")
def draw_calib(self, cr, w, h):
c = self.calib
@@ -1885,9 +2191,9 @@ class Probe(Adw.ApplicationWindow):
cr.arc(x, y, 7, 0, 2 * math.pi)
cr.stroke()
GLib.idle_add(self.area.queue_draw) # the pulse
total = len(self.ROUND_BG) * 7
n = c["round"] * 7 + c["index"] + 1
self.text(cr, 60, 70, f"Calibration: round {c['round'] + 1} of 3 ({self.ROUND_NAMES[c['round']]}), dot {n} of {total}",
total = len(self.ROUND_BG) * len(c["points"])
n = c["round"] * len(c["points"]) + c["index"] + 1
self.text(cr, 60, 70, f"Calibration ({'Own tracker' if self.source == 'own' else 'SteamVR'}): round {c['round'] + 1} of 3 ({self.ROUND_NAMES[c['round']]}), dot {n} of {total}",
ink, 26)
hint = "Face the centre dot, look at the highlighted dot, and press. Esc cancels, S skips a dot."
self.text(cr, 60, 108, c.get("retry") or hint, (0.8, 0.2, 0.1) if c.get("retry") and bright else
@@ -2268,6 +2574,9 @@ class Probe(Adw.ApplicationWindow):
if self.calib:
self.draw_calib(cr, w, h)
return
if self.reseat_check:
self.draw_reseat(cr, w, h)
return
m = self.monitors.get(self.screen)
if m and self.is_fullscreen():
g = m.get_geometry()
@@ -2281,9 +2590,13 @@ class Probe(Adw.ApplicationWindow):
if self.w_cells.get_active() and r and r[5]:
self.degree_grid(cr, w, h, 1 / r[5])
if self.mode == "fit":
self.fitcheck.draw(cr, w, h, self.text, time.monotonic())
if self.fitcheck.guide_step(time.monotonic()):
GLib.idle_add(self.area.queue_draw)
if self.test:
self.draw_test(cr)
if self.results and not self.test and not self.snap:
if self.results and not self.test and not self.snap and self.mode != "fit":
self.draw_results(cr)
if self.snap:
self.draw_snap(cr)
@@ -2318,7 +2631,8 @@ class Probe(Adw.ApplicationWindow):
frozen = self.capture is not None
live = not frozen or self.w_live.get_active()
if self.w_all.get_active() and live:
own = self.source == "own"
if self.w_all.get_active() and live and not own:
for name in SOURCES:
rr = self.raw.get(name)
if rr:
@@ -2335,7 +2649,14 @@ class Probe(Adw.ApplicationWindow):
cr.move_to(x, y - 12)
cr.line_to(x, y + 12)
cr.stroke()
if self.w_rawdot.get_active() and r and live:
if self.w_rawdot.get_active() and own and live:
# Each eye alone: they should meet, with a little jitter, where you look.
for e in getattr(self, "own_eyes", []):
if e:
cr.set_source_rgba(1, 0.25, 0.25, 0.85)
cr.arc(e[0], e[1], 5, 0, 2 * math.pi)
cr.fill()
elif self.w_rawdot.get_active() and r and live:
cr.set_source_rgba(1, 0.3, 0.3, 0.8)
cr.arc(r[0], r[1], 5, 0, 2 * math.pi)
cr.fill()
@@ -2357,7 +2678,7 @@ class Probe(Adw.ApplicationWindow):
res = self.result
if res and now - res["shown"] < 2.5 and "L" in res:
a = max(0.0, 1 - (now - res["shown"]) / 2.5)
ok = res["verdict"] == "learned" or res["verdict"] == "measured"
ok = res["verdict"] in ("learned", "measured") or res["verdict"].startswith("taught")
(fx, fy), (lx, ly) = res["F"], res["L"]
cr.set_source_rgba(*((1, 0.85, 0.2) if ok else (1, 0.35, 0.3)), a)
cr.set_line_width(2)
@@ -2534,12 +2855,22 @@ def region_errors(targets, source, key="cerr_deg"):
def main():
ap = argparse.ArgumentParser(description="Eye tracking playground for the Frametop desktop")
ap.add_argument("--screen", type=int, default=0, help="Frametop screen to open on (default: where it opens)")
ap.add_argument("--mode", choices=Probe.MODE_KEYS, help="start in this mode (fit: Headset fit)")
ap.add_argument("--source", choices=SOURCES,
help="gaze source to start with (default: own if our tracker is running, else mmap2)")
args, rest = ap.parse_known_args()
app = Adw.Application(application_id="dev.frametop.GazeProbe", flags=Gio.ApplicationFlags.NON_UNIQUE)
windows = []
def activate(a):
win = Probe(a, args.screen)
if args.source:
win.w_source.set_selected(SOURCES.index(args.source))
elif not win.own.ask("status").startswith("fail"):
# Our own tracker is running: that's what you're here to test.
win.w_source.set_selected(SOURCES.index("own"))
if args.mode:
win.set_mode(args.mode)
windows.append(win)
win.present()
+14
View File
@@ -0,0 +1,14 @@
# frame-job settings (see `frame-job --help`): offline lab jobs (ft-eyes-score, ft-eyes-e2e)
# run on the 7i. Eye recordings may go there and nowhere else, and never into the repo: they
# live outside it, in ~/.local/share/frametop/eyes/captures, and on the 7i in
# ~/frame-compute/frametop-eyes/data/captures.
NAME=frametop-eyes
RUN_ON=7i
DATA="$HOME/.local/share/frametop/eyes/captures"
RESULTS=""
EXCLUDE=""
# The lab's Python on the 7i: build/venv from requirements.txt, as build.sh makes it on the Frame.
SETUP='cmp -s requirements.txt build/venv/requirements.done || { rm -rf build/venv && python3 -m venv build/venv && build/venv/bin/pip install -q --disable-pip-version-check -r requirements.txt && cp requirements.txt build/venv/requirements.done; }'
VENV=
# Live tools: they read the Frame's eye cameras.
LOCAL_ONLY="ft-eyes ft-eyes-record ft-eyes-session"
+23
View File
@@ -0,0 +1,23 @@
#!/usr/bin/env bash
# Build our own eye tracker on the Frame, in the dev container:
# build/ft-eyegrab the frame grabber. It runs on the host, as root
# (frametop-eyegrab.service, gaze/tracker/install.sh), so this checks it
# only needs glibc symbols the SteamOS host has (2.39; the container has 2.43).
# build/venv Python with numpy and OpenCV (requirements.txt) for ft-eyes and lab/,
# remade when requirements.txt changes.
# Usage: gaze/tracker/build.sh
set -euo pipefail
root=$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)
"$root/scripts/sync.sh" >/dev/null
exec "$root/scripts/frame.sh" -C gaze/tracker 'set -e; mkdir -p build
gcc -std=gnu11 -O2 -Wall -Wextra -pthread -o build/ft-eyegrab ft-eyegrab.c
max=$(objdump -T build/ft-eyegrab | grep -oE "GLIBC_[0-9.]+" | sort -uV | tail -1)
echo "built build/ft-eyegrab, newest glibc symbol: $max"
[ "$(printf "%s\n" "$max" GLIBC_2.39 | sort -V | tail -1)" = GLIBC_2.39 ] || { echo "needs newer glibc than the host has" >&2; exit 1; }
if ! cmp -s requirements.txt build/venv/requirements.done; then
rm -rf build/venv
python3 -m venv build/venv
build/venv/bin/pip install -q --disable-pip-version-check -r requirements.txt
cp requirements.txt build/venv/requirements.done
fi
echo "build/venv: $(build/venv/bin/python -c "import numpy, cv2; print(\"numpy\", numpy.__version__, \"opencv\", cv2.__version__)")"'
+243
View File
@@ -0,0 +1,243 @@
"""The gaze calibration: from pupil and glint positions to head-relative gaze angles.
Shared by ft-eyes (live) and the lab tools (fitting and scoring on recordings). Gaze angles are
ft-gaze's: degrees relative to the head, yaw positive to the left, pitch positive up.
Per eye (0 = right, camera 0; 1 = left, camera 1), three quadratic fits:
pupil pupil centre -> gaze. The one used for output, after the slip correction.
glint pupil minus the glint pair's midpoint -> gaze. Slip moves both alike, so this
holds up when the headset shifts, but it's noisier, and the pair is often gone.
where gaze -> pupil centre: where the pupil sits for a gaze, with no slip.
Slip: wherever the pair is seen, `glint` gives the gaze, `where` says where the pupil should
be, and the difference is how far the eye has moved in the image. Slip changes slowly, so
the median over the last SLIP_WINDOW seconds shifts every frame, glints or not.
That glint estimate is only good to 1-4 px (1-3 degrees), though. Clicks are better: each one
says where the pupil should have been for a known gaze (`where`), so pupil minus that is the
shift. `Shift` keeps the median of the last few, and uses the glints only to notice a sudden
jump (the headset nudged or put back on), until clicks catch up. On practice2 with
practice1's calibration: 1.2 degrees median, against 2.7 with the glints alone (findings.md).
"""
import json
import time
from collections import deque
import numpy as np
SLIP_WINDOW = 30.0
SLIP_MIN = 10 # pair sightings needed before trusting a slip estimate
MIN_CLICKS = 12
SPREAD_MIN = 0.3 # degrees: floor for an eye's fit spread, so one eye can't take all the weight
SHIFT_KEEP = 5 # clicks in the shift estimate
SHIFT_JUMP = 3.0 # a glint slip change this big (px) since the last click is a nudge
JUMP_HOLD = 1.0 # s: ...if it holds this long (a bad glint pair gives a jump that snaps back)
JUMP_MAX = 40.0 # px: bigger is a bad glint pair, not the headset (a re-seat moved 20-30)
JUMP_WINDOW = 10.0 # seconds of glint sightings for noticing a jump
class Quad:
"""Ridged quadratic least squares from 2-D inputs, normalised on the training set."""
def __init__(self, X=None, Y=None, ridge=1e-3):
if X is None:
return
X, Y = np.asarray(X, float), np.asarray(Y, float)
self.m, self.s = X.mean(0), X.std(0) + 1e-9
A = self.terms(X)
R = ridge * np.eye(A.shape[1])
R[0, 0] = 0
self.w = np.linalg.solve(A.T @ A + R, A.T @ Y)
def terms(self, X):
P = (np.atleast_2d(np.asarray(X, float)) - self.m) / self.s
return np.c_[np.ones(len(P)), P, P ** 2, P[:, 0] * P[:, 1]]
def __call__(self, X):
return self.terms(X) @ self.w
def one(self, x, y):
"""Faster for a single point (the live path)."""
px, py = (x - self.m[0]) / self.s[0], (y - self.m[1]) / self.s[1]
t = np.array([1.0, px, py, px * px, py * py, px * py])
return t @ self.w
def to_json(self):
return {"m": self.m.tolist(), "s": self.s.tolist(), "w": self.w.tolist()}
@classmethod
def from_json(cls, d):
q = cls()
q.m, q.s, q.w = (np.array(d[k]) for k in ("m", "s", "w"))
return q
def pair_mid(pair):
return ((pair[0][0] + pair[1][0]) / 2, (pair[0][1] + pair[1][1]) / 2)
class Calibration:
"""The three fits per eye. Build from clicks (ft-eyes-score's features) or load from JSON.
`spread` is each eye's RMS miss (degrees) on the clicks its pupil fit was made from.
`combine` weights the eyes by its inverse square: on practice2 (one headset position,
leave-one-out) that gave 0.75 median against 0.84 for the plain average, since one eye
is usually much better than the other (left 0.73, right 1.32 there)."""
def __init__(self, fits=None, info=None, spread=None):
self.fits = fits or {} # (name, eye) -> Quad
self.info = info or {}
self.spread = spread or {} # eye -> degrees
@classmethod
def fit(cls, clicks, info=None):
truth = lambda cs: np.array([k["truth"] for k in cs]) # noqa: E731
fits, spread = {}, {}
for c in (0, 1):
cs = [k for k in clicks if k["eye"][c] is not None]
if len(cs) < MIN_CLICKS:
continue
fits["pupil", c] = Quad([k["eye"][c]["pupil"] for k in cs], truth(cs))
miss = fits["pupil", c]([k["eye"][c]["pupil"] for k in cs]) - truth(cs)
spread[c] = float(np.sqrt(np.mean(np.sum(miss ** 2, axis=1))))
fits["where", c] = Quad(truth(cs), [k["eye"][c]["pupil"] for k in cs])
gs = [k for k in cs if k["eye"][c]["mid"] is not None]
if len(gs) >= MIN_CLICKS:
fits["glint", c] = Quad([k["eye"][c]["pupil"] - k["eye"][c]["mid"] for k in gs], truth(gs))
return cls(fits, info, spread)
def has(self, name, eye):
return (name, eye) in self.fits
def slip(self, eye, rows):
"""Median slip in pixels from pair sightings `rows` (t, px, py, mx, my), or None."""
if not self.has("glint", eye) or len(rows) < SLIP_MIN:
return None
rows = np.asarray(rows, float)
gaze = self.fits["glint", eye](rows[:, 1:3] - rows[:, 3:5])
return np.median(rows[:, 1:3] - self.fits["where", eye](gaze), axis=0)
def click_shift(self, eye, pupil, truth):
"""The shift a click measures: the pupil, less where it sits for that gaze."""
return np.asarray(pupil, float) - self.fits["where", eye].one(*truth)
def gaze(self, eye, x, y, slip=None):
"""Gaze (yaw, pitch) for a pupil centre, less a slip if there is one."""
if slip is not None:
x, y = x - slip[0], y - slip[1]
return self.fits["pupil", eye].one(x, y)
def weight(self, eye):
s = self.spread.get(eye)
return 1.0 if s is None else 1.0 / max(s, SPREAD_MIN) ** 2
def combine(self, gazes):
"""The weighted mean of {eye: (yaw, pitch)}, or None if empty."""
if not gazes:
return None
w = {c: self.weight(c) for c in gazes}
return sum(w[c] * np.asarray(g, float) for c, g in gazes.items()) / sum(w.values())
def save(self, path):
d = {"version": 1, "info": self.info, "spread": {str(e): s for e, s in self.spread.items()},
"fits": [{"name": n, "eye": e, **q.to_json()} for (n, e), q in self.fits.items()]}
path.parent.mkdir(parents=True, exist_ok=True)
tmp = path.with_suffix(".tmp")
tmp.write_text(json.dumps(d, indent=1))
tmp.replace(path)
@classmethod
def load(cls, path):
d = json.loads(path.read_text())
return cls({(f["name"], f["eye"]): Quad.from_json(f) for f in d["fits"]}, d.get("info"),
{int(e): s for e, s in d.get("spread", {}).items()})
class Shift:
"""Where one eye sits in the image now, relative to the calibration (pixels).
`base` is the median shift the last SHIFT_KEEP clicks measured. `ref` is the glint slip
estimate at the last click; if the glint estimate has since moved more than SHIFT_JUMP
(and less than JUMP_MAX) and stayed there for JUMP_HOLD seconds, the headset moved, and
the change is added until the next click. That click then starts the history over,
since the older ones describe the old position. Live, bad glint pairs made the estimate
leap by up to 68 px for under a second (practice2), hence the hold. `reseat` does the
same for the next click without the glints: the frames stopped (the headset was off),
so the headset may be anywhere now."""
def __init__(self, base=(0.0, 0.0), ref=None):
self.meas = []
self.base = np.asarray(base, float)
self.ref = None if ref is None else np.asarray(ref, float)
self.jump = np.zeros(2)
self.held = None # (change, since when) while a jump waits out JUMP_HOLD
self.reseated = False
def glint(self, g, t):
"""The latest glint slip estimate (or None), at time t (s)."""
if g is None:
return
if self.ref is None:
self.ref = np.asarray(g, float)
d = np.asarray(g, float) - self.ref
size = np.hypot(*d)
if size > JUMP_MAX:
return
if size <= SHIFT_JUMP:
self.jump, self.held = np.zeros(2), None
return
if self.held is None or np.hypot(*(d - self.held[0])) > SHIFT_JUMP:
self.held = (d, t)
elif t - self.held[1] >= JUMP_HOLD:
self.jump = d
def reseat(self):
self.reseated = True
def click(self, d, g=None):
"""A click measured the shift d; g is the glint estimate then."""
if np.any(self.jump) or self.reseated:
self.meas = []
self.reseated = False
self.meas = (self.meas + [np.asarray(d, float)])[-SHIFT_KEEP:]
self.base = np.median(self.meas, axis=0)
self.jump, self.held = np.zeros(2), None
if g is not None:
self.ref = np.asarray(g, float)
@property
def value(self):
return self.base + self.jump
def to_json(self):
return {"base": self.base.tolist(), "ref": None if self.ref is None else self.ref.tolist(),
"meas": [m.tolist() for m in self.meas]}
@classmethod
def from_json(cls, d):
s = cls(d.get("base", (0, 0)), d.get("ref"))
s.meas = [np.asarray(m, float) for m in d.get("meas", [])]
return s
class SlipTracker:
"""Live slip estimate for one eye: pair sightings over the last SLIP_WINDOW seconds,
re-estimated at most every `every` seconds."""
def __init__(self, cal, eye, every=0.5, window=SLIP_WINDOW):
self.cal, self.eye, self.every, self.window = cal, eye, every, window
self.rows = deque()
self.value, self.at = None, 0.0
def add(self, t, pupil, mid):
self.rows.append((t, pupil[0], pupil[1], mid[0], mid[1]))
while self.rows and self.rows[0][0] < t - self.window:
self.rows.popleft()
def get(self, now=None):
now = time.monotonic() if now is None else now
if now - self.at >= self.every:
self.at = now
s = self.cal.slip(self.eye, list(self.rows))
if s is not None:
self.value = s
return self.value
+139
View File
@@ -0,0 +1,139 @@
"""Classic pupil and glint finder for one 512x400 eye-camera frame.
The pupil is a dark blob enclosed by brighter iris and skin. The lens rim and the unlit
background are just as dark, but they touch the image edge, so any dark region that reaches
the edge is dropped. Closing the glints' holes can join the pupil to that background (the
left camera's, when you look more than about 20 degrees left), so when nothing is found
the search runs again with a smaller closing.
"""
import cv2
import numpy as np
DARK = 30 # pupil pixels are below this (the face around it is 40-180)
MIN_AREA = 150 # pupil area range in pixels
MAX_AREA = 20000
MIN_FILL = 0.75 # blob area / fitted-ellipse area
MAX_ASPECT = 3.0 # long / short axis; the steep camera sees a squashed pupil
GLINT = 200 # glints are near-saturated spots on or by the pupil
CLOSE = 7 # px: closes the glints' holes in the pupil (the right eye's need this much)
CLOSE_TIGHT = 3 # px: the retry, keeps a pupil near the dark background apart from it
NEAR = 70 # the windowed search: this many pixels, or 3 pupil radii, around a hint
def find_pupil(frame, near=None):
"""Return dict(x, y, a, b, angle, area, fill, glints) or None.
`near` (a previous result) searches a window around it first (0.4 ms, not 1.4-2.1);
if the pupil isn't wholly inside the window it falls back to the whole frame."""
if near is not None:
r = int(max(NEAR, 3 * near['a']))
x0, y0 = max(int(near['x']) - r, 0), max(int(near['y']) - r, 0)
p = _find_either(frame[y0:int(near['y']) + r, x0:int(near['x']) + r], x0, y0)
if p is not None:
p['glints'] = find_glints(frame, p)
return p
p = _find_either(frame, 0, 0)
if p is not None:
p['glints'] = find_glints(frame, p)
return p
def _find_either(img, ox, oy):
p = _find(img, ox, oy)
return p if p is not None else _find(img, ox, oy, CLOSE_TIGHT)
def _find(img, ox, oy, close=CLOSE):
"""The pupil in `img` (a window at ox, oy of the frame), with frame coordinates. A dark
region touching the window's edge doesn't count: it's background, rim, or cut off."""
f = cv2.GaussianBlur(img, (5, 5), 0)
dark = (f < DARK).astype(np.uint8)
# Glints punch bright holes in the pupil; close them so the blob stays whole.
dark = cv2.morphologyEx(dark, cv2.MORPH_CLOSE, np.ones((close, close), np.uint8))
dark = cv2.morphologyEx(dark, cv2.MORPH_OPEN, np.ones((3, 3), np.uint8))
n, lab, stats, _ = cv2.connectedComponentsWithStats(dark, connectivity=8)
h, w = img.shape
best = None
for i in range(1, n):
x, y, bw, bh, area = stats[i]
if area < MIN_AREA or area > MAX_AREA:
continue
if x <= 1 or y <= 1 or x + bw >= w - 1 or y + bh >= h - 1:
continue
blob = lab[y:y + bh, x:x + bw] == i
cs, _ = cv2.findContours(blob.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)
c = max(cs, key=len)
if len(c) < 5:
continue
(ex, ey), (d1, d2), ang = cv2.fitEllipse(c)
a, b = max(d1, d2) / 2, min(d1, d2) / 2
if b < 3 or a / b > MAX_ASPECT:
continue
fill = area / (np.pi * a * b)
if fill < MIN_FILL or fill > 1.25:
continue
# Prefer the darkest, fullest blob.
score = fill - img[y:y + bh, x:x + bw][blob].mean() / 255
if best is None or score > best[0]:
# fitEllipse's angle is the direction of its first axis (d1); the long axis is
# that one or the one at right angles.
major = np.radians(ang if d1 >= d2 else ang + 90)
best = (score, dict(x=ex + x + ox, y=ey + y + oy, a=a, b=b, angle=ang, major=major,
area=int(area), fill=fill, box=(x + ox, y + oy, bw, bh)))
return best[1] if best else None
GLINT_RING = 140 # a glint sits on dark iris or pupil: its surroundings are below this
GLINT_PAIR = (5, 45) # the two LEDs' reflections: this far apart, in pixels, mostly vertical
def find_glints(frame, p):
"""Small bright spots on dark iris or pupil within 2.5 pupil radii, as (x, y) list.
Bright skin has noise speckle above GLINT too, so a spot counts only if the ring
around it is dark."""
r = int(p['a'] * 2.5) + 6
x0, y0 = max(int(p['x']) - r, 0), max(int(p['y']) - r, 0)
roi = frame[y0:int(p['y']) + r, x0:int(p['x']) + r]
n, lab, stats, cents = cv2.connectedComponentsWithStats((roi >= GLINT).astype(np.uint8))
out = []
for i in range(1, n):
x, y, w, h, area = stats[i]
if not 2 <= area <= 80:
continue
ya, yb, xa, xb = max(y - 4, 0), y + h + 4, max(x - 4, 0), x + w + 4
ring = roi[ya:yb, xa:xb][lab[ya:yb, xa:xb] != i]
if ring.size and np.median(ring) < GLINT_RING:
out.append((cents[i][0] + x0, cents[i][1] + y0))
return out
def glint_pair(p):
"""The two LED reflections as ((x, y) upper, (x, y) lower), or None. Picks the
vertical-ish pair nearest the pupil centre."""
g = p['glints']
best = None
for i in range(len(g)):
for j in range(i + 1, len(g)):
dx, dy = g[j][0] - g[i][0], g[j][1] - g[i][1]
d = np.hypot(dx, dy)
if not GLINT_PAIR[0] <= d <= GLINT_PAIR[1] or abs(dy) < 2 * abs(dx):
continue
mx, my = (g[i][0] + g[j][0]) / 2, (g[i][1] + g[j][1]) / 2
cost = np.hypot(mx - p['x'], my - p['y'])
if best is None or cost < best[0]:
best = (cost, (g[i], g[j]) if dy > 0 else (g[j], g[i]))
return best[1] if best else None
def draw(frame, p, scale=1.0):
img = cv2.cvtColor(cv2.convertScaleAbs(frame, alpha=2.0), cv2.COLOR_GRAY2BGR)
if p:
cv2.ellipse(img, ((p['x'], p['y']), (2 * p['a'], 2 * p['b']), np.degrees(p['major'])),
(0, 255, 0), 1)
for gx, gy in p['glints']:
cv2.circle(img, (int(gx), int(gy)), 3, (0, 0, 255), 1)
if scale != 1.0:
img = cv2.resize(img, None, fx=scale, fy=scale)
return img
+465
View File
@@ -0,0 +1,465 @@
# Findings so far
Our own eye tracker's research notes, newest sections last. They were written while it was a
separate project (frame-eyes), so they use its names: `fe-trackd` is now `ft-eyes`,
`fe-bufprobe` is `ft-eyegrab`, `fe_model`/`fe_pupil` are `eyes_model`/`eyes_pupil`, the
tools are in `lab/` (`fe-score` is `ft-eyes-score`, `fe-replaytest` is `ft-eyes-e2e`,
`fe-record` and `fe-replay` are `ft-eyes-record` and `ft-eyes-replay`), `fe-live` is the
gaze service running ft-eyes, and `captures/NAME` is `~/.local/share/frametop/eyes/captures/NAME`.
These were measured on the Frame on 2026-09-28 (SteamVR eyetracking 2.17.10), and all by
reading only.
## SteamVR's tracker process
- `eyetracking -b CDSP -w .../et_dsp_20250610_03136.weights` runs as the user (steamos),
started by SteamVR. The user is in the `cdsp` and `spidev` groups. `ptrace_scope` is 1,
so reading another process's fds or memory needs root (pidfd_getfd).
- Log: `~/.local/share/Steam/logs/eyetracking.txt`. Component names: `CStereoAdspCams`
(the eye cameras come in through the audio DSP; "Set framerate 72/90"),
`CGazeEstimatorCdsp` / `CDSPGazenet` (the neural net on the compute DSP), and
`CEyePoseUKF L/R` (a filter per eye; "Large dt" means it had no measurement for over
0.4 s and starts that eye over).
- "Failed to grab cdsp input buffer" came up 5,924 times in 6 hours (about 0.3 % of
frames at 90 Hz).
- "Accept usercal" is its passive calibration from quick mouse clicks.
- The eye cameras aren't V4L2 devices, and there's no fastrpc node.
- Open fds that matter:
- `/dev/spidev0.1`: modalias `spi:hid-over-spi`, role unknown.
- Six udmabufs, all `exp_name: udmabuf`: three of 16 MiB (16777216 B) and three of
32 MiB (33554432 B), fds 50, 51, 53, 54, 159 and 169 at the time.
- `/dev/shm/eye-server.mmap`: its output.
- `/dev/input/event0-7`.
- The frames most likely arrive in the udmabufs, which it shares with the DSPs.
## The eye-camera frames (found 2026-09-28, `tools/fe-bufprobe --scan`)
- **The buffers.** The six udmabuf fds are really two buffers, three fds each: a 16 MiB one
(inode 1) and a 32 MiB one (inode 2).
- **The frames.** In the 16 MiB buffer, eight slots sit 0x40000 apart from 0x230000, four
per camera: slots 0-3 are camera 0 and slots 4-7 camera 1. Each slot starts with a small
block (slot 0's holds a table of floats such as 0.00125, 4.655, 90.0, 1.0, possibly
exposure, gain, and frame rate; the others were zero), then a 512x400 8-bit grayscale
frame, row stride 512. The frame starts at slot base + 0x40c0 + 0x40 per slot index,
plus one more 0x40 for camera 1's slots. Found by the dark lens-rim column lining up;
`slot_start()` in fe-bufprobe.
- **What they show.** Infrared images, one camera per eye, dim (mean about 25-40), with a
dark band on one side (the lens rim). One camera saw its eye at a steep angle (squashed
pupil near the image edge, big reflections on the white); the other nearly head-on (a
round pupil with two small glints in it). Which camera is which eye isn't known yet.
- **Timing.**
- About 90 frames a second per camera, the two within about 0.6 ms of each other.
- A frame lands over several milliseconds, in bursts, so a copy taken when the slot
"stops changing" can be half old. The reliable rule: a slot is complete when its camera
starts writing another slot.
- Camera 0 fills its slots in turn (3, 0, 1, 2); camera 1 in a repeating order of eight
(7, 5, 4, 6, 5, 7, 6, 4), never the same slot twice in a row.
- `--rec` stamps each frame when its slot first changed, polling every 0.3 ms, so times
are only good to a few ms.
- The cameras run at 90 fps ("Set framerate 90" in the log). The first recorder copied
a frame as soon as its camera started the next one and got 94 a second in fit1 and 106
in a streaming test. The extras were half-written frames: a frame's last writes can land
after the next frame starts, and a slow poll saw both at once. The recorder now copies a
frame when its camera starts the frame after next (no slot is rewritten sooner than three
frames), ignores late writes to the frame just finished, and saves from a separate
thread. fit1 may hold a few percent of torn frames.
- Eye tracking stops when the headset is off ("HMD off, stopping eye tracking"), so
recordings are empty then.
- **Which camera is which eye** (capture fit1, 2026-09-29, closing one eye at a time):
camera 0 (slots 0-3) is the **right** eye, seen at a steep angle; camera 1 (slots 4-7) is
the **left** eye, seen nearly head-on. Both images have the lens rim dark on the left and
the lit face on the right.
- **Why the left eye is lost looking down:** at the keyboard, camera 1 sees only the upper
lid and lashes. Camera 0 still catches part of the right eye. It's the camera angle, not
the net.
- **The 32 MiB buffer.** Two 48 KiB regions (0x1522000, 0x1532000) that change every frame,
mean bytes about 148, 99 % nonzero. Probably the net's input per eye (crops, maybe not 8-bit
pixels). Not decoded.
- **`tools/fe-session NAME SECONDS`.** Records frames and ft-gaze's samples together, on the
same clock (CLOCK_MONOTONIC_RAW). Each frame's nearest SteamVR sample is a median 3.8 ms
away.
## Our first pupil finder (`tools/fe_pupil.py`, 2026-09-29)
- Threshold dark (< 30), close glint holes, drop dark regions that touch the image edge
(lens rim, background), keep the fullest, darkest ellipse-shaped blob. Glints: spots
>= 200 within 1.5 pupil radii. 3.1 ms a frame on the CPU, unoptimised.
- On fit1 (20 s: open, each eye closed, keyboard, up), pupil found vs SteamVR seeing the
eye:
| Gaze pitch | Right, SteamVR | Right, ours | Left, SteamVR | Left, ours |
| --- | --- | --- | --- | --- |
| below -20 (keyboard) | 79 % | 35 % | 22 % | 24 % |
| -20 to 15 (screens, includes closed-eye time) | 91 % | 89 % | 74 % | 72 % |
| above 15 | 100 % | 100 % | 98 % | 99 % |
It misses the right eye looking down, where the lower lid cuts the pupil. False finds
on closed eyes: 0.4 % right, 2.8 % left.
- A quadratic fit from pupil centre to SteamVR's per-eye gaze, held out by time block:
median 4.8 degrees right, 3.0 left. That isn't an accuracy figure yet. fit1 has few
distinct gaze points, it uses no glints, and SteamVR's per-eye gaze is itself off by
several degrees. It needs a recording against known targets.
## eye-server.mmap (its output)
The file is 324,122 bytes. Only bytes 0x0-0x1f3 are used; the rest is zero. It's packed and
unaligned, so read it with memcpy. Offsets are also in `~/frametop/gaze/ft-gaze.cpp`.
| Offset | What |
| --- | --- |
| 0x38 | u32 sample counter |
| 0x157 | f64 sample time, CLOCK_MONOTONIC_RAW |
| 0x15f, 0x16b | set 1: left and right eye direction (3 f32, head space, -Z forward). Filtered; both eyes always share one pitch; a lost eye keeps its yaw |
| 0x177 | set 1: 6 f32 variances (left 3, right 3; the middle one of each is shared). About 0.0005-0.002 when the eye is seen, 0.015-0.03 when it's lost |
| 0x18f | set 1 fixation point (3 f32; its length is the vergence distance) |
| 0x19b, 0x1a7 | set 2: each eye's own direction |
| 0x1b3 | set 2: 6 f32 variances |
| 0x1cb | 2 f32, 0..1: openness (0 in a blink) |
| 0x1d3 | 8 f32: left measurement x, y; right x, y (camera-relative, freezes while that eye isn't seen); then variance of left x, y, right x, y (about 2e-5 on a clear view, rising as lids or lashes get in the way) |
| 0x0-0x157 | header, plus records that look like the calibration-click channel into the tracker. Never write |
## Accuracy of SteamVR's gaze (this user, this headset)
- **Tonight's practice (71 clicks, 45 minutes):**
- raw error: median 5.1 degrees (3.0 in the 21:23 test; it varies by session);
- corrected at the press: median 1.5, with 1 in 10 past 3.3;
- best smooth correction fitted to the same clicks, each predicted from the rest: 1.7-1.9;
- weighting recent clicks more (half-lives from 20 minutes down to 1) didn't help, so
there's no slow drift to follow.
- **Look-to-look:** two looks within 3 degrees of each other, under 5 minutes apart,
differ by a median 1.15 degrees (3.0 when 5 or more minutes apart). Jitter within one
look is 0.25-0.3.
- **One eye alone (set 2), against both eyes' gaze:** median 0.8 degrees over a steady
look. Per-eye raw errors are large and opposite in yaw: at one spot, left (+8.2, +8.9)
and right (-5.0, +5.2) degrees, both (+1.6, +7.1).
- **Losses:** the left eye was lost 57-64 % of the time looking 30-50 degrees down (at
the keyboard, through the gap by the nose), and the right eye never. At screen height
both were seen over 98 % of the time. Openness looking down: left 0.45, right 0.65.
Harmless for the pointer: ft-gazed ignores looks down past the screens.
## First accuracy test against known targets (practice1, 2026-09-29)
5 minutes, 98 gaze-probe practice clicks, gaze yaw -26..25 and pitch -14..20 degrees. Truth
is SteamVR's raw gaze at the press plus the angle to the release point. `tools/fe-score.py`
fits a quadratic per method and scores each click leave-one-out. Pupil = median centre over
the frames 250-20 ms before the press; no glints yet.
| Method (85 clicks with both pupils found) | Median | 90 % |
| --- | --- | --- |
| SteamVR raw | 6.51 | 11.04 |
| SteamVR + quadratic fit | 1.52 | 3.10 |
| Ours, right pupil only | 0.74 | 1.48 |
| Ours, left pupil only | 0.62 | 1.48 |
| Ours, both pupils averaged | 0.61 | 1.13 |
SteamVR with the probe's live correction: 1.44 median over all 98. The pupil was found
before 88/98 clicks (right) and 90/98 (left). Caveats: one session with the headset
never moved (pupil-only mapping breaks when the headset slips; glints should fix that),
the truth includes the user's own drag precision, and it's offline only.
## Glints and slip (2026-09-29, practice1)
- Two IR LED reflections, a vertical pair 12-19 px apart, sit on the cornea near the pupil.
There's no alternating illumination: the pair is in every frame the geometry allows.
Before a click: right eye 45/88, left 54/90 (the steep right camera loses the pair on
the white when the eye looks across). Bright skin has noise speckle above 200, so a glint
only counts if the ring around it is dark (`find_glints`, `glint_pair` in fe_pupil.py).
- Pupil minus pair midpoint as the feature: 0.84 median, noisier than the pupil alone
(0.61), because the pair's position is noisy.
- Slip method (`tools/fe-score.py`): where the pair is seen, the glint fit gives the gaze,
a fit of gaze to pupil position says where the pupil should be, and the difference is
the slip. The median over the last 30 s shifts every frame, glints or not. In-session:
0.61 median, same as the pupil alone. The estimate stayed within 1-3 px all session.
- Simulated slip (fit on the first 49 clicks, test on the rest shifted 10 px): pupil alone
0.62 -> 2.7-3.1; glint 0.83 unchanged; slip 0.67 unchanged. That checks the math only,
for a pure image shift. A real re-seat also tilts and changes the distance.
- Next test: a second session after taking the headset off and on, scored with
`fe-score.py captures/practice1 captures/practice2`.
## Live tracker (2026-09-29)
- `fe-bufprobe --share` (root) copies each finished frame into `/dev/shm/frame-eyes-cams`
(0600, the user's); `fe-trackd` (user) finds pupils and glints and writes
`/dev/shm/frame-eyes-gaze`; ft-gaze reads that as the source `own`. `tools/fe-live` runs
both. The user side never touches SteamVR's buffers.
- The windowed pupil search gives the same results as the full frame (0.000 px apart on
2000 frames per eye), at 0.4 ms instead of 1.4-2.1; with glints, about 1 ms a frame, 90 fps
per eye.
- Replaying practice1's first minute through fe-trackd: 0.64 median at the 14 clicks
(0.61 offline with the same calibration; in-sample, so a pipeline check, not accuracy).
Sample-to-sample jitter 0.08 degrees; SteamVR's is 0.25-0.3.
- The calibration covers yaw -26..25 and pitch -14..20 degrees (practice1's clicks). Beyond
that the quadratic extrapolates.
## Test 2: a second session after re-seating (practice2, 2026-09-29 15:18)
124 practice clicks over about 4 minutes, headset nudged at about 100 s. The probe stayed on
SteamVR's mmap2 as its source, but recorded our live gaze at every press. Scored with
`fe-score.py captures/practice1 captures/practice2` (median degrees):
| Method | Fit on practice1 | Fit within practice2 (leave-one-out) |
| --- | --- | --- |
| SteamVR raw | 2.86 | 2.86 |
| SteamVR + the probe's live correction | 1.37 | 1.37 |
| SteamVR + quadratic fit | 3.84 | 1.17 |
| Ours, pupil only | 14.31 | 2.86 |
| Ours, glint | 3.31 | 1.54 |
| Ours, slip | 2.68 (live: 2.87) | 1.36 |
- The re-seat moved the eyes 20-30 px in the images: pupil-only goes 14 degrees off. The
slip correction takes that to 2.7, but no further. The rest is partly one offset (yaw
-1.6, pitch +1.2; removing it leaves 1.43), and an offset from the previous 5 clicks
gives 1.31, the same as SteamVR's live correction (1.37).
- Within one headset position (before the nudge, 45 clicks; after it, 68): pupil only
1.01 and 0.99, slip 1.18 and 1.86, SteamVR with the same fit 0.91 and 1.19. So today our
tracker is level with SteamVR within a position, not ahead of it as in practice1 (0.61
against 1.69). One session was not enough to claim a lead.
- The slip estimate adds noise: it's worse than no correction within a position, and it
wandered (the right eye's jumped 18 px near the end, after the clicks). It comes from the
glint fit, which is itself only 1.5-3 degrees good, and the left eye's pair was seen
before only 21 of 124 clicks.
- Live: fe-trackd ran 81-90 fps per eye at 2-3.6 ms a frame alongside VR (1 ms in replay);
ft-gaze's `own` came through on every line, about 37 ms old.
- What would help: a geometric eye model (the eyeball centre from how the pupil ellipse
changes shape, as Swirski's method and Pupil Labs' pye3d do) instead of 2-D regression,
so that headset movement is modelled and not fitted around. practice1 and practice2
together (re-seat plus a nudge) are the benchmark for it.
## The geometric model, and a shift taught by clicks (2026-09-29, practice1 -> practice2)
All offline: calibrate on practice1, score practice2's clicks.
- **Eyeball centre from the pupil ellipses (Swirski-style, weak perspective): worse.** The
centre it finds is steady within a session (a few px per 50 s) and moves between the
sessions about as the slip does, with a rotation radius of about 60 px (10-12 mm). But
it's off from the true shift by up to 8 px (6-7 degrees) on the right eye. Correcting
with it gave 4.3-5.8 median, against 2.7 for the glint slip. The cornea's refraction and
where you happened to look in the window likely bias it.
- **The calibration itself carries over.** The best possible pixel shift per eye, fitted
on practice2's own clicks with one shift per headset position (before and after the
nudge), gives 1.18 held out (shift+scale 1.11, affine 1.07). So practice1's fit is fine
if we know the shift; the problem was only estimating it. One shift for the whole
session gets only 2.7-2.8, because the nudge moved the eyes again.
- **How well each estimate finds that shift (px, right x/y, before the nudge):** true
(-0.7,-28.6), glints (+2.0,-25.6), eyeball centre (-6.9,-22.2). The glints are off by
1-4 px, which is 1-3 degrees; the eyeball centre is worse.
- **Clicks estimate it best.** Each click says where the pupil should have been for a
known gaze, so pupil minus that is the shift. Scored in time order with only earlier
clicks:
| Shift from | Median | 90% |
| --- | --- | --- |
| Glints only, last 10 or 30 s | 2.68-2.69 | 3.78-4.13 |
| Last 3 clicks | 1.29 | 2.80 |
| **Last 5 clicks** | **1.18** | 3.09 |
| Last 5 clicks + glint slip since | 1.30-1.34 | 3.34-3.55 |
| **Last 5 clicks, glints only to catch a jump over 3 px** | **1.22** | **2.33** |
| SteamVR + the probe's live correction (same clicks) | 1.37 | 2.71 |
Adding the glint slip to the clicks' shift adds its noise. Using it only to notice a
nudge (then the shift follows the glints until clicks catch up) keeps the median and
cuts the tail after a nudge. That's `fe_model.Shift`, and `fe-score`'s `clicks` method
reproduces it (1.22 median, 2.33 90%).
- So on this pair of sessions, ours with click correction is slightly ahead of SteamVR
with the probe's click correction: 1.22 against 1.37 median, 2.33 against 2.71 for the
worst tenth. One pair of sessions, so not yet a lead (see Test 2).
- fe-trackd now works this way: the probe's calibration with the source "Own tracker"
sends each dot to fe-trackd (`calib-point`), which fits from its own pupil history
(`calib-fit`, which replaces the old calibration and clears the shifts), and each
practice release sends a `click` that teaches the shift. See fe-trackd's docstring.
- **The live path reproduces it** (`tools/fe-replaytest captures/practice1 captures/practice2`
on the 7i: a scratch fe-trackd, calibrated through `calib-point` from practice1's clicks,
then fed practice2's clicks at the recorded pace and scored on what it published in the
300 ms before each press). 84 of 98 dots accepted (14 had an eye in under 15 frames),
fit 0.59 median. practice2: median 1.21 and 1.23 over two runs (offline 1.22), but 90%
3.02 and 2.86 (offline 2.33). The tail: the first click (19.7, nothing learned yet after
the re-seat), and two clicks at 192 s and 213 s (7.6 and 10.3) with a settled 5-click
shift and no jump. (Offline has the same two, 6.8 and 9.3: see the next section.) The
glint jump restarted the right eye's shift 9 times and the left's 4.
## Wide gaze, the left pupil, and false glint jumps (2026-09-29, practice2)
- **The bad clicks were all past the calibration, or right eye only.** practice1's clicks
reach yaw 25 and pitch 20; practice2's reach 30 and 25. Offline (median / 90%):
| Clicks | Ours | SteamVR + probe |
| --- | --- | --- |
| Inside practice1's range (104) | 1.09 / 1.97 | 1.37 / 2.71 |
| Outside it (18-20) | 1.80 / 5.01 | 1.19 / 2.67 |
| Both eyes (99) | 1.09 / 1.98 | 1.40 / 2.83 |
| Right eye only (23) | 1.66 / 4.05 | 0.99 / 2.47 |
A fit that goes linear past its data, more ridge, or a linear fit didn't help outside.
- **The left eye was lost at every click past about 19 degrees left.** The pupil is still
mid-image (x 293-310 of 512), but the left camera's image is dark from x 0 to about 360,
and the 7 px closing (which heals glint holes) joined the pupil to that background, which
touches the edge, so it was dropped. `fe_pupil` now retries with a 3 px closing when
nothing is found: left eye found in 97% of the frames before practice2's clicks (was
79%; 181 of 217 frames at 20+ degrees, was 16), right eye unchanged, centres moved at
most 0.16 px. The right eye still needs the 7 px (3 px alone: 96% against 98%).
- **Those pupils are accurate.** Within one headset position (practice2 after the nudge,
68 clicks, leave-one-out, so calibrated out there too): left eye alone 0.70 / 1.39 below
19 degrees left and 0.87 / 1.14 beyond; right eye 1.27 / 2.13 and 2.09 / 6.42; both
averaged 0.85 / 1.30 and 1.18 / 5.14. Calibrated where you look, the left eye is our best.
- **But practice1's calibration has 4 clicks per eye beyond 19 degrees**, so the newly seen
left eye extrapolates there (3.64 median alone, right 1.78), and practice1 -> practice2
got a worse tail: 1.20 / 3.20 offline (1.22 / 2.33 when the left eye was simply lost
there). Weighting the eyes, or leaving out an eye or a click past the calibrated range,
didn't recover it. The fix is coverage: the probe's calibration for the Own tracker now
puts its dots on an oval out to the `Calibration ring` angle each way (the ring was
limited by the window's height and never reached the sides). A first try put them at
the practice area's corners, which in a large window were too far to look at while
facing the centre.
- **The glint jumps.** Offline (checked once per click) there were 2 per eye, 3 of 4 real
(the next click found the shift the glints claimed). Live, bad glint pairs made the right
eye's estimate leap by up to 68 px for under a second, 20 times between clicks, and the
shift restarted 8-9 times. `Shift` now takes a jump only once it has held for 1 s
(JUMP_HOLD) and ignores ones over 40 px (JUMP_MAX). Live replay: shift restarts 1 (right)
and 0 (left), biggest leap between clicks 4.9 px, output steps over 10 degrees 86 -> 32.
Offline with the hold: 1.19 / 3.35.
- Live replay with both changes: 94 of 98 calibration dots taken (83-84 before), practice2
1.28 / 3.16. The tail stays until a calibration covers the practice area.
- `fe-score` now reads each capture's own `practice.jsonl` (cut from the probe's log by
`fe-score.py --clicks`, or the first scoring on the Frame), so the 7i can rebuild features.
## Session 3 (2026-09-29 22:04-22:10, live, SteamVR driving)
The calibration and 84 practice clicks went to SteamVR (the probe's tracker toggle was
left on SteamVR), so fe-trackd got no dots or clicks and kept practice1's calibration. The
probe still logged our gaze at 79 presses. SteamVR + the probe's correction: 1.19 median,
2.61 90% (raw 2.18 / 4.30). Ours with practice1's calibration and no clicks: 12.5 median;
with a stand-in for the click shift (the median offset of the previous 5 clicks, in gaze
angles rather than per eye in pixels): 1.38 / 2.38. No frames were recorded.
## Session 4: the Own tracker driving (2026-09-29 22:12-22:22, live)
Fresh calibration from the probe with the Own tracker: 27 dots on an oval out to 20
degrees (30 of 31 attempts accepted; one had no right eye). Then 136 practice clicks, all
with both eyes, each teaching the shift. The probe logged SteamVR at 114 of the presses,
and its correction learned from the same drags, so the comparison is fair:
| At the same 114 presses, to where you let go | Median | 90% |
| --- | --- | --- |
| **Ours** | **0.59** | **1.50** |
| SteamVR + the probe's correction | 0.83 | 2.02 |
| SteamVR raw | 3.49 | 5.21 |
Ours was closer on 76 of 114. Ours at the press (what the dot showed, all 136): 0.67
median, 1.37 90%, and steady from the first 10 clicks (0.71) on; SteamVR's correction
took about 30 clicks to get under 1 degree. No click changed an eye's shift by more than
6 px. The user: "MUCH improved". No frames were recorded, and the headset wasn't nudged or
re-seated, so this is within one position; the cross-session question is still open.
## Session 5: off and on again, no recalibration (2026-09-29 22:26-22:31, live)
Session 4's calibration and shifts, headset taken off and put back on, then 132 practice
clicks with the Own tracker driving. The re-seat moved the eyes about 15 px (right) and
33 px (left) in the images.
- Before the first taught click, the glints had moved the right eye's shift to within
about 4 px and the left's about two thirds of the way. Clicks 1-4 were still 11-17
degrees off, and the probe refused to teach them (its 6-degree limit on a lesson), so
the first taught click was the 5th (5.4 degrees). It restarted each eye's history as
designed; clicks 6, 7, 8: 2.8, 1.3, 0.4.
- After that (clicks 6 on, 127, to where you let go): ours 0.58 median, 1.42 90%, the
same as within one position (session 4: 0.59); SteamVR + the probe's correction 2.94 /
5.73 (SteamVR raw drifted from 3.5 to 4.7 through the session). Ours closer on 117 of 132.
- Changes: the probe lets the Own tracker learn from drags up to 25 degrees
(OWN_LEARN_MAX), and starts on the Own tracker when fe-trackd answers; its calibration
header names the tracker. fe-trackd restarts an eye's shift history at the next click
after frames stop for 3 s (the headset off) and after its own restart, glints or not.
Replay regression (fe-replaytest practice1 practice2): 1.28 / 3.16, as before.
## Weighting the eyes (2026-09-29, practice2 after the nudge)
One headset position, 64 clicks with both eyes, leave-one-out: plain average 0.84 / 1.59;
left eye alone 0.73 / 1.36; right alone 1.32 / 3.05; weighted by each eye's inverse
residual variance on its own calibration 0.75 / 1.26. Not in fe-trackd yet.
## Next steps (2026-09-29, after a literature search; sources in the session report)
Ranked by expected gain for the effort, checked against our own numbers:
1. Weight the eyes by each one's calibration residuals (above: 0.84 -> 0.75 median). S.
2. A one-dot re-seat check when frames come back after a gap (Varjo recalibrates with one
dot at every put-on): one look and press teaches both shifts before the first real
click, instead of 11-17 degree first clicks. S.
3. Our tracker as a source for the Frametop pointer (ft-gazed): session 5 beat SteamVR
across a re-seat. M. Done 2026-09-30 (below).
4. Record frames during live tests (fe-session), so each can be replayed. S (disk: about
2 GB a minute).
5. A less biased glint slip estimate: ours is off by 1-4 px even over hundreds of frames,
so it's bias, not noise; try taking out the part of the glint midpoint that follows the
pupil (regressed on calibration data) before using it. S-M, offline first.
6. Smooth-pursuit calibration (a moving dot): dense labels out to the edge in about 20 s,
for wider coverage. M.
7. Sub-pixel edge ellipse refit with RANSAC for the steep right eye (our weaker eye,
1.32 against 0.73). S-M.
8. Later, if needed: learned pupil segmentation (EllSeg, RITnet: MIT) on the GPU through
ncnn, a 3-D cornea model from the two glints, or a per-user network trained on the
residuals across re-seats. L. Not recommended: the eyeball-centre model (tried, and our
steep camera and +-20 degree range are outside its published conditions). PuRe,
PuReST, ElSe, and ExCuSe are licensed for non-commercial use only.
## Quick wins from the next steps (2026-09-29, late)
- Eye weighting (1): `Calibration.spread` is each eye's RMS miss on its own calibration
dots, and `combine` weights by its inverse square (floor 0.3 degrees). fe-trackd, fe-score,
and so fe-replaytest use it; older calibrations get it from their saved dots. The
22:16 calibration: right 1.48, left 1.08, so the left eye counts about twice as much.
practice1 -> practice2 offline is unchanged (1.20 / 3.35: that tail is extrapolation).
- Re-seat check (2): fe-trackd's status says when the next click will start an eye's shift
over; the probe then shows one centre dot, and a press on it sends that click. Checked
on the 7i (a pending re-seat at start on both eyes, cleared by one click); the probe's
screen for it wasn't seen (the web view didn't connect).
- Recording (4): `tools/fe-record` copies every shared frame (9 s of replay: 1620 frames,
none dropped, all identical to the source); `fe-live --record NAME` runs it alongside.
## The Frametop pointer, and the eyes on live clicks (2026-09-30)
ft-gazed (`~/frametop/gaze`) can now use our tracker: `GAZE_TRACKER=own`, the Eye tracker
setting on the Gaze page of Frametop Input Settings. A mouse nudge before a click reaches
fe-trackd as a click. The nudge's raw gaze is one ft-gazed sent, so ft-gazed finds when
that look was, and fe-trackd keeps 12 s of pupils instead of 5, because the helper sends a
nudge up to 10 s after the look.
`GAZE_EYE` (auto, left, right) weights the eyes there, from each eye's own gaze. Replayed on
the 306 live clicks of sessions 4 and 5 (`clicks.jsonl`: each eye's pupil and shift just
before the click, so each is a fresh test):
- Each eye alone: left 0.96 median (mean 1.17), right 1.11 (1.26). The eyes' RMS misses
were about equal (1.43, 1.46), unlike practice2's leave-one-out (0.73, 1.32).
- Both eyes: 0.65 (0.77) evenly. By the calibration's spread (the 22:16 one: left counts
about twice): 0.63 (0.81). By each eye's RMS miss at its last 5, 10, or 20 clicks: 0.66
(0.79-0.81).
- By share of the right eye: 0.3 gives 0.68, 0.5 gives 0.65, 0.7 gives 0.81.
- The eyes' yaw errors are correlated -0.37: they partly cancel, which is why two eyes
beat either one by a third.
So a bias leans instead of choosing: Left or Right counts that eye twice. Auto starts
even and weights by each eye's RMS miss at its last 20 nudges, once each has 5. On
SteamVR's side the calibration's own fit picked the wrong eye (its dots: left 1.78, right
1.88; new spots: left 2.50, right 1.63), so auto learns from nudges, not the fit.
fe-trackd's own `combine` still uses the spread, for the probe.
## Valve's tracker
It can't be the starting point, legally or practically:
- **No source.** `/opt/steamvr/tools/eyetracking/bin/linuxarm64/eyetracking` is a
stripped aarch64 binary. The paths left in it (`/data/src/eyetracking/eyetracklib/...`)
are Valve's build machine's.
- **The net is just numbers.** `et_dsp_20250610_03136.weights` is 393,600 bytes of raw
floats (about 98,000 parameters), with no header or architecture. The layer layout
lives in the binary and in the program it loads onto the compute DSP (`CDSPGazenet`).
Rebuilding it would mean reverse engineering both.
- **License.** SteamVR is Valve's proprietary software, used under the Steam Subscriber
Agreement. That agreement doesn't allow reverse engineering, decompiling, modifying, or
redistributing it, except where the law allows. `third_party_legal_notices.txt`
covers only the open libraries it uses (Ceres, protobuf, ...), not the tracker. Putting
their code or weights in a GitHub repo would be redistribution. (Not legal advice.)
- **Not much to gain.** Their net is small and tuned to their cameras. Improving it would
need the same thing our own tracker needs: your eye images with known gaze, for
training.
What we can use: its public output (the mmap, read-only), as a baseline and as labels.
Anything published and openly licensed is also fair game: papers and open-source pupil
detectors (check each one's license before using its code).
+36
View File
@@ -0,0 +1,36 @@
# Template: gaze/tracker/install.sh fills in @UID@ and @GID@ (the Frametop user) and installs
# it to /etc/systemd/system.
[Unit]
Description=Frametop eye-camera frames for our own eye tracker (read-only copies from SteamVR's eyetracking)
Documentation=file://@REPO@/gaze/README.md
[Service]
# Idle (no frames copied, none of the tracker's buffers held) until ft-eyes or
# ft-eyes-record touches the want file; see ft-eyegrab.c.
ExecStart=/etc/frametop/ft-eyegrab --share /dev/shm/frametop-eyes-cams --owner @UID@:@GID@ --want /dev/shm/frametop-eyes-want
Restart=on-failure
RestartSec=5
Nice=5
# Root only for what reading another process's buffers needs: CAP_SYS_PTRACE (pidfd_getfd),
# CAP_DAC_READ_SEARCH (its /proc/PID/fd), and CAP_CHOWN (the shared file goes to the user).
CapabilityBoundingSet=CAP_SYS_PTRACE CAP_DAC_READ_SEARCH CAP_CHOWN
AmbientCapabilities=
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
ReadWritePaths=/dev/shm
PrivateNetwork=yes
RestrictAddressFamilies=AF_UNIX
ProtectKernelModules=yes
ProtectKernelTunables=yes
ProtectControlGroups=yes
ProtectClock=yes
ProtectHostname=yes
RestrictNamespaces=yes
RestrictRealtime=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
SystemCallArchitectures=native
[Install]
WantedBy=multi-user.target
+605
View File
@@ -0,0 +1,605 @@
/*
* ft-eyegrab: the eye-camera frames for our own eye tracker (ft-eyes), copied read-only out
* of the DMA-BUFs SteamVR's eyetracking process holds. Runs as root (pidfd_getfd needs
* CAP_SYS_PTRACE; ptrace_scope is 1): the system service frametop-eyegrab.service runs
* --share, installed by gaze/tracker/install.sh. The other modes are for finding the frames
* again after a SteamVR update (run them with sudo).
*
* ft-eyegrab --share PATH [--owner UID:GID] [--want FILE]
* keep the latest frames of both cameras in PATH (shared memory,
* 0600, owned by UID:GID, or the sudo user) for ft-eyes; follows
* the tracker through SteamVR restarts. With --want, only while
* FILE (a regular file owned by that user) was touched in the
* last WANT_FRESH seconds: ft-eyes and ft-eyes-record touch it
* every second, so nothing is copied, and none of the tracker's
* buffers are held, while nobody reads the frames
* ft-eyegrab list the buffers
* ft-eyegrab --scan [N] N snapshots (default 40) about 11 ms apart: which 4 KiB pages
* change, merged into regions, with byte statistics for each
* ft-eyegrab --dump I OFF LEN FILE
* copy LEN bytes at OFF of buffer I (from the list) to FILE
* ft-eyegrab --seq I OFF LEN FRAMES DIR
* FRAMES copies of that region, one each time it changes, to
* DIR/NNNN.raw, with DIR/times.txt (CLOCK_MONOTONIC_RAW)
* ft-eyegrab --rec SECONDS DIR
* every new eye-camera frame for SECONDS: DIR/frames.raw (512x400
* 8-bit frames back to back) and DIR/index.txt, one line per frame:
* "<n> <slot> <camera 0|1> <CLOCK_MONOTONIC_RAW time seen>"
* (lab/ft-eyes-record does the same from the shared frames,
* without root)
*
* The eye frames (found with --scan): in the 16 MiB buffer, eight slots 0x40000 apart from
* 0x230000, four per camera (slots 0-3, 4-7). Each slot starts with a small block, then a
* 512x400 8-bit image at 0x40c0 + 0x40 per slot, and one more 0x40 for the second camera's.
*
* Only reads the tracker's buffers. They are borrowed with pidfd_getfd and mapped PROT_READ; nothing is
* written, and the process isn't stopped or signalled. Reads can tear while the DSP writes.
*/
#define _GNU_SOURCE
#include <dirent.h>
#include <errno.h>
#include <fcntl.h>
#include <pthread.h>
#include <signal.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mman.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <time.h>
#include <unistd.h>
#define MAXBUF 32
#define PAGE 4096
typedef struct {
int xfd, fd;
size_t size;
unsigned long ino;
const uint8_t *p;
} buf_t;
static buf_t bufs[MAXBUF];
static int nbufs;
static double now(void) {
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC_RAW, &ts);
return ts.tv_sec + ts.tv_nsec * 1e-9;
}
static int find_tracker(void) {
DIR *d = opendir("/proc");
struct dirent *e;
int pid = -1;
while (d && (e = readdir(d))) {
char path[300], cmd[512];
if (e->d_name[0] < '0' || e->d_name[0] > '9') continue;
snprintf(path, sizeof path, "/proc/%s/cmdline", e->d_name);
FILE *f = fopen(path, "r");
if (!f) continue;
size_t n = fread(cmd, 1, sizeof cmd - 1, f);
fclose(f);
cmd[n] = 0;
if (strstr(cmd, "tools/eyetracking/bin/") && strstr(cmd, "/eyetracking")) {
pid = atoi(e->d_name);
break;
}
}
if (d) closedir(d);
return pid;
}
static int open_bufs(int pid) {
int pidfd = syscall(SYS_pidfd_open, pid, 0);
if (pidfd < 0) {
perror("pidfd_open");
return -1;
}
char dir[64];
snprintf(dir, sizeof dir, "/proc/%d/fd", pid);
DIR *d = opendir(dir);
struct dirent *e;
while (d && (e = readdir(d)) && nbufs < MAXBUF) {
char link[320], target[256];
if (e->d_name[0] == '.') continue;
snprintf(link, sizeof link, "%s/%s", dir, e->d_name);
ssize_t n = readlink(link, target, sizeof target - 1);
if (n <= 0) continue;
target[n] = 0;
if (strncmp(target, "/dmabuf:", 8) != 0) continue;
int xfd = atoi(e->d_name);
int fd = syscall(SYS_pidfd_getfd, pidfd, xfd, 0);
if (fd < 0) {
fprintf(stderr, "pidfd_getfd %d: %s\n", xfd, strerror(errno));
continue;
}
struct stat st;
fstat(fd, &st);
off_t size = lseek(fd, 0, SEEK_END);
void *p = mmap(NULL, size, PROT_READ, MAP_SHARED, fd, 0);
if (p == MAP_FAILED) {
fprintf(stderr, "mmap fd %d (%lld bytes): %s\n", xfd, (long long)size, strerror(errno));
close(fd);
continue;
}
bufs[nbufs++] = (buf_t){xfd, fd, (size_t)size, (unsigned long)st.st_ino, p};
}
if (d) closedir(d);
close(pidfd);
return nbufs;
}
static uint64_t page_hash(const uint8_t *p) {
const uint64_t *q = (const uint64_t *)p;
uint64_t h = 1469598103934665603ull;
for (int i = 0; i < PAGE / 8; i += 4) h = (h ^ q[i]) * 1099511628211ull; // every 4th word
return h;
}
static void stats(const uint8_t *p, size_t n, double *mean, int *lo, int *hi, double *nonzero) {
uint64_t sum = 0, nz = 0;
int a = 255, b = 0;
for (size_t i = 0; i < n; i++) {
sum += p[i];
nz += p[i] != 0;
if (p[i] < a) a = p[i];
if (p[i] > b) b = p[i];
}
*mean = n ? (double)sum / n : 0;
*lo = a, *hi = b, *nonzero = n ? (double)nz / n : 0;
}
static void scan(int snaps) {
for (int b = 0; b < nbufs; b++) {
size_t pages = bufs[b].size / PAGE;
uint64_t *prev = calloc(pages, 8), *cur = calloc(pages, 8);
int *changes = calloc(pages, sizeof(int));
for (size_t i = 0; i < pages; i++) prev[i] = page_hash(bufs[b].p + i * PAGE);
double t0 = now();
for (int s = 1; s < snaps; s++) {
usleep(11000);
for (size_t i = 0; i < pages; i++) {
cur[i] = page_hash(bufs[b].p + i * PAGE);
if (cur[i] != prev[i]) changes[i]++;
prev[i] = cur[i];
}
}
double dt = now() - t0;
printf("buffer %d (fd %d, %zu bytes, ino %lu): %d snapshots over %.2f s\n", b, bufs[b].xfd, bufs[b].size,
bufs[b].ino, snaps, dt);
// Regions: runs of pages that changed at least once (gaps of up to 2 pages merged).
size_t i = 0;
int regions = 0;
while (i < pages) {
if (!changes[i]) {
i++;
continue;
}
size_t start = i, end = i, gap = 0;
int most = 0;
long total = 0;
for (; i < pages; i++) {
if (changes[i]) {
end = i, gap = 0;
total += changes[i];
if (changes[i] > most) most = changes[i];
} else if (++gap > 2) {
break;
}
}
size_t off = start * PAGE, len = (end - start + 1) * PAGE;
double mean, nz;
int lo, hi;
stats(bufs[b].p + off, len, &mean, &lo, &hi, &nz);
printf(" region 0x%08zx +0x%zx (%zu KiB): changed in up to %d of %d intervals (avg %.1f); "
"bytes mean %.1f min %d max %d nonzero %.0f%%\n",
off, len, len / 1024, most, snaps - 1, (double)total / (end - start + 1), mean, lo, hi, nz * 100);
regions++;
}
if (!regions) {
double mean, nz;
int lo, hi;
stats(bufs[b].p, bufs[b].size, &mean, &lo, &hi, &nz);
printf(" no change; bytes mean %.1f min %d max %d nonzero %.1f%%\n", mean, lo, hi, nz * 100);
}
free(prev), free(cur), free(changes);
}
}
#define EYE_W 512
#define EYE_H 400
#define EYE_SLOTS 8
static size_t slot_start(int k) {
return 0x230000 + (size_t)k * 0x40000 + 0x40c0 + (size_t)k * 0x40 + (k >= 4 ? 0x40 : 0);
}
// A cheap fingerprint of a frame: 256 words spread over it (a new frame changes nearly all).
static uint64_t frame_sig(const uint8_t *p) {
uint64_t h = 1469598103934665603ull, w;
for (int i = 0; i < 256; i++) {
memcpy(&w, p + (size_t)i * (EYE_W * EYE_H / 256), 8);
h = (h ^ w) * 1099511628211ull;
}
return h;
}
static volatile sig_atomic_t stop_rec;
static void on_stop(int sig) { (void)sig; stop_rec = 1; }
// The poller copies finished frames into a ring; a writer thread saves them, so a slow disk
// write never delays the polling.
#define RING 128
static struct {
uint8_t *frames;
int slot[RING];
double time[RING];
size_t head, tail, dropped; // head: next to fill (poller); tail: next to save (writer)
int done;
pthread_mutex_t mu;
pthread_cond_t cv;
FILE *f, *ix;
} ring = {.mu = PTHREAD_MUTEX_INITIALIZER, .cv = PTHREAD_COND_INITIALIZER};
static void *ring_writer(void *arg) {
(void)arg;
size_t fsize = EYE_W * EYE_H, n = 0;
pthread_mutex_lock(&ring.mu);
for (;;) {
while (ring.tail == ring.head && !ring.done) pthread_cond_wait(&ring.cv, &ring.mu);
if (ring.tail == ring.head) break;
size_t i = ring.tail % RING;
pthread_mutex_unlock(&ring.mu);
fwrite(ring.frames + i * fsize, 1, fsize, ring.f);
fprintf(ring.ix, "%zu %d %d %.6f\n", n++, ring.slot[i], ring.slot[i] >= 4, ring.time[i]);
pthread_mutex_lock(&ring.mu);
ring.tail++;
}
pthread_mutex_unlock(&ring.mu);
return NULL;
}
static void ring_put(const uint8_t *frame, int slot, double t) {
size_t fsize = EYE_W * EYE_H;
pthread_mutex_lock(&ring.mu);
int full = ring.head - ring.tail >= RING;
pthread_mutex_unlock(&ring.mu);
if (full) {
ring.dropped++;
return;
}
size_t i = ring.head % RING;
memcpy(ring.frames + i * fsize, frame, fsize);
ring.slot[i] = slot, ring.time[i] = t;
pthread_mutex_lock(&ring.mu);
ring.head++;
pthread_cond_signal(&ring.cv);
pthread_mutex_unlock(&ring.mu);
}
static int eye_buffer(void) {
for (int i = 0; i < nbufs; i++)
if (bufs[i].size == 16777216) return i;
return -1;
}
// Calls done(frame, slot, time) for every complete eye-camera frame until `seconds` pass
// (forever if negative), a stop signal comes, the tracker process goes away, or keep()
// (checked about every 0.25 s, when given) says to stop.
//
// A frame lands over several milliseconds, in bursts, and its last bursts can come after
// the camera has started its next frame. A slot isn't rewritten until at least three frames
// later (camera 0 cycles 3,0,1,2; camera 1 7,5,4,6,5,7,6,4), so a frame is passed on when
// its camera starts the frame after next. Changes to the slot just finished are late bursts,
// not a new frame. A frame's time is when its slot first changed.
static void poll_frames(int b, double seconds, int pid, void (*done)(const uint8_t *, int, double),
int (*keep)(void)) {
uint64_t sig[EYE_SLOTS];
double first[EYE_SLOTS];
int cur[2] = {-1, -1}, prev[2] = {-1, -1};
for (int k = 0; k < EYE_SLOTS; k++) sig[k] = frame_sig(bufs[b].p + slot_start(k)), first[k] = 0;
double start = now(), checked = start, kept = start;
char proc[64];
snprintf(proc, sizeof proc, "/proc/%d", pid);
while ((seconds < 0 || now() - start < seconds) && !stop_rec) {
double t = now();
if (t - checked > 1.0) { // the tracker restarted: its buffers are stale
struct stat st;
if (stat(proc, &st) != 0) return;
checked = t;
}
if (keep && t - kept > 0.25) {
if (!keep()) return;
kept = t;
}
for (int k = 0; k < EYE_SLOTS; k++) {
uint64_t s = frame_sig(bufs[b].p + slot_start(k));
if (s == sig[k]) continue;
sig[k] = s;
int cam = k >= 4;
if (k == cur[cam] || k == prev[cam]) continue; // landing, or a late burst
if (prev[cam] >= 0) done(bufs[b].p + slot_start(prev[cam]), prev[cam], first[prev[cam]]);
prev[cam] = cur[cam];
cur[cam] = k;
first[k] = t;
}
usleep(300);
}
}
static void rec_frame(const uint8_t *frame, int slot, double t) { ring_put(frame, slot, t); }
static int rec(double seconds, const char *dir, int pid) {
int b = eye_buffer();
if (b < 0) {
fprintf(stderr, "no 16 MiB buffer\n");
return 1;
}
// Frames stream to disk (about 37 MB/s), so a long recording doesn't fill memory.
// Ctrl-C or SIGTERM ends it early and keeps what was recorded.
char path[512];
snprintf(path, sizeof path, "%s/frames.raw", dir);
ring.f = fopen(path, "wb");
snprintf(path, sizeof path, "%s/index.txt", dir);
ring.ix = fopen(path, "w");
ring.frames = malloc((size_t)RING * EYE_W * EYE_H);
if (!ring.f || !ring.ix || !ring.frames) {
perror(dir);
return 1;
}
setvbuf(ring.f, NULL, _IOFBF, 4 << 20);
signal(SIGINT, on_stop);
signal(SIGTERM, on_stop);
pthread_t writer;
pthread_create(&writer, NULL, ring_writer, NULL);
double start = now();
poll_frames(b, seconds, pid, rec_frame, NULL);
pthread_mutex_lock(&ring.mu);
ring.done = 1;
pthread_cond_signal(&ring.cv);
pthread_mutex_unlock(&ring.mu);
pthread_join(writer, NULL);
fclose(ring.f), fclose(ring.ix);
printf("%zu frames in %.1f s to %s", ring.head, now() - start, dir);
if (ring.dropped) printf(" (%zu dropped: disk too slow)", ring.dropped);
printf("\n");
free(ring.frames);
return 0;
}
// --- --share: the latest frames in shared memory for the live tracker ---
//
// The file (SHARE_PATH, mode 0600, owned by the --owner user) is a header, then SHARE_SLOTS
// entries per camera. Each entry is a 64-byte head and one 512x400 frame. Frame n of camera
// c goes in entry c * SHARE_SLOTS + n % SHARE_SLOTS. The head's `seq` is odd while it's
// written (read it before and after copying, and retry if it changed or was odd), and
// count[c] is how many frames camera c has published. tracker_pid is 0 while no frames
// come (nobody wants them, or SteamVR's tracker isn't running).
//
// This keeps the tracker's own buffers behind root: the user side only ever sees copies.
#define SHARE_SLOTS 8
#define SHARE_MAGIC 0x31434546u // "FEC1"
#define WANT_FRESH 3.0 // seconds a touch of the --want file lasts
typedef struct {
uint32_t magic, version, width, height, slots, entry_size;
volatile uint64_t count[2];
uint32_t tracker_pid, pad0;
uint8_t pad[16];
} share_head_t;
typedef struct {
volatile uint64_t seq;
double t;
uint64_t n;
uint32_t cam, slot;
uint8_t pad[32];
} share_entry_t;
_Static_assert(sizeof(share_head_t) == 64, "share header");
_Static_assert(sizeof(share_entry_t) == 64, "share entry");
static uint8_t *share;
static const char *want_path;
static uid_t owner_uid = (uid_t)-1;
static gid_t owner_gid = (gid_t)-1;
static void share_frame(const uint8_t *frame, int slot, double t) {
share_head_t *h = (share_head_t *)share;
int cam = slot >= 4;
uint64_t n = h->count[cam];
size_t esize = sizeof(share_entry_t) + EYE_W * EYE_H;
share_entry_t *e = (share_entry_t *)(share + sizeof *h + (cam * SHARE_SLOTS + n % SHARE_SLOTS) * esize);
e->seq++;
__atomic_thread_fence(__ATOMIC_RELEASE);
memcpy((uint8_t *)(e + 1), frame, EYE_W * EYE_H);
e->t = t, e->n = n, e->cam = cam, e->slot = slot;
__atomic_thread_fence(__ATOMIC_RELEASE);
e->seq++;
__atomic_thread_fence(__ATOMIC_RELEASE);
h->count[cam] = n + 1;
}
// Someone reads the frames: the want file was touched lately. It must be a regular file
// (lstat: a link isn't followed) owned by the frames' owner, so no one else can turn this on.
static int wanted(void) {
if (!want_path) return 1;
struct stat st;
if (lstat(want_path, &st) != 0 || !S_ISREG(st.st_mode)) return 0;
if (owner_uid != (uid_t)-1 && st.st_uid != owner_uid) return 0;
struct timespec ts;
clock_gettime(CLOCK_REALTIME, &ts);
double age = (ts.tv_sec - st.st_mtim.tv_sec) + (ts.tv_nsec - st.st_mtim.tv_nsec) * 1e-9;
return age < WANT_FRESH;
}
static void close_bufs(void) {
for (int i = 0; i < nbufs; i++) munmap((void *)bufs[i].p, bufs[i].size), close(bufs[i].fd);
nbufs = 0;
}
static int share_loop(const char *path) {
size_t esize = sizeof(share_entry_t) + EYE_W * EYE_H;
size_t size = sizeof(share_head_t) + 2 * SHARE_SLOTS * esize;
unlink(path);
int fd = open(path, O_RDWR | O_CREAT | O_EXCL | O_NOFOLLOW | O_CLOEXEC, 0600);
if (fd < 0 || ftruncate(fd, size) != 0) {
perror(path);
return 1;
}
if (owner_uid != (uid_t)-1 && fchown(fd, owner_uid, owner_gid) != 0) perror("fchown");
share = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
close(fd);
if (share == MAP_FAILED) {
perror("mmap");
return 1;
}
share_head_t *h = (share_head_t *)share;
*h = (share_head_t){.magic = SHARE_MAGIC, .version = 1, .width = EYE_W, .height = EYE_H,
.slots = SHARE_SLOTS, .entry_size = (uint32_t)esize};
signal(SIGINT, on_stop);
signal(SIGTERM, on_stop);
fprintf(stderr, "ft-eyegrab: sharing frames in %s%s%s\n", path, want_path ? " while wanted by " : "",
want_path ? want_path : "");
int pid = -1, idle = -1, missing = 0;
while (!stop_rec) {
if (!wanted()) {
// Nobody reads the frames: copy nothing, and let go of the tracker's buffers.
if (idle != 1) fprintf(stderr, "ft-eyegrab: idle (nobody wants frames)\n"), idle = 1;
close_bufs();
pid = -1;
h->tracker_pid = 0;
usleep(250000);
continue;
}
if (nbufs == 0 && ((pid = find_tracker()) < 0 || open_bufs(pid) <= 0 || eye_buffer() < 0)) {
// SteamVR's tracker isn't running (yet, or again).
if (!missing) fprintf(stderr, "ft-eyegrab: waiting for SteamVR's eyetracking\n"), missing = 1;
close_bufs();
h->tracker_pid = 0;
sleep(2);
continue;
}
if (idle != 0 || missing) fprintf(stderr, "ft-eyegrab: copying frames from eyetracking %d\n", pid);
idle = 0, missing = 0;
h->tracker_pid = pid;
poll_frames(eye_buffer(), -1, pid, share_frame, wanted);
struct stat st;
char proc[64];
snprintf(proc, sizeof proc, "/proc/%d", pid);
if (!stop_rec && stat(proc, &st) != 0) {
fprintf(stderr, "ft-eyegrab: eyetracking %d went away; waiting for it\n", pid);
close_bufs();
h->tracker_pid = 0;
}
}
close_bufs();
unlink(path);
return 0;
}
static int dump(int b, size_t off, size_t len, const char *file) {
if (b < 0 || b >= nbufs || off + len > bufs[b].size) {
fprintf(stderr, "out of range\n");
return 1;
}
FILE *f = fopen(file, "wb");
if (!f) {
perror(file);
return 1;
}
fwrite(bufs[b].p + off, 1, len, f);
fclose(f);
printf("wrote %zu bytes to %s\n", len, file);
return 0;
}
static int seq(int b, size_t off, size_t len, int frames, const char *dir) {
if (b < 0 || b >= nbufs || off + len > bufs[b].size) {
fprintf(stderr, "out of range\n");
return 1;
}
char path[512];
snprintf(path, sizeof path, "%s/times.txt", dir);
FILE *times = fopen(path, "w");
if (!times) {
perror(path);
return 1;
}
uint8_t *copy = malloc(len);
uint64_t last = 0;
int got = 0;
double start = now();
while (got < frames && now() - start < 30) {
uint64_t h = 0;
for (size_t i = 0; i + PAGE <= len; i += PAGE * 8) h ^= page_hash(bufs[b].p + off + i) + i;
if (h != last) {
last = h;
double t = now();
memcpy(copy, bufs[b].p + off, len);
snprintf(path, sizeof path, "%s/%04d.raw", dir, got);
FILE *f = fopen(path, "wb");
if (f) fwrite(copy, 1, len, f), fclose(f);
fprintf(times, "%d %.6f\n", got, t);
got++;
}
usleep(1000);
}
fclose(times);
free(copy);
printf("%d frames in %s\n", got, dir);
return 0;
}
static void usage(void) {
fprintf(stderr, "usage: ft-eyegrab [--share PATH [--owner UID:GID] [--want FILE] | --scan [N] | --dump I OFF LEN FILE |\n"
" --seq I OFF LEN FRAMES DIR | --rec SECONDS DIR]\n");
}
int main(int argc, char **argv) {
if (argc >= 3 && strcmp(argv[1], "--share") == 0) {
const char *uid = getenv("SUDO_UID"), *gid = getenv("SUDO_GID");
if (uid && gid) owner_uid = (uid_t)atoi(uid), owner_gid = (gid_t)atoi(gid);
for (int i = 3; i < argc; i++) {
unsigned u, g;
if (strcmp(argv[i], "--owner") == 0 && i + 1 < argc && sscanf(argv[i + 1], "%u:%u", &u, &g) == 2) {
owner_uid = u, owner_gid = g, i++;
} else if (strcmp(argv[i], "--want") == 0 && i + 1 < argc) {
want_path = argv[++i];
} else {
usage();
return 2;
}
}
return share_loop(argv[2]);
}
int pid = find_tracker();
if (pid < 0) {
fprintf(stderr, "SteamVR's eyetracking process isn't running\n");
return 1;
}
if (open_bufs(pid) <= 0) {
fprintf(stderr, "no buffers (run as root)\n");
return 1;
}
if (argc >= 2 && strcmp(argv[1], "--scan") == 0) {
scan(argc >= 3 ? atoi(argv[2]) : 40);
} else if (argc == 6 && strcmp(argv[1], "--dump") == 0) {
return dump(atoi(argv[2]), strtoul(argv[3], NULL, 0), strtoul(argv[4], NULL, 0), argv[5]);
} else if (argc == 4 && strcmp(argv[1], "--rec") == 0) {
return rec(atof(argv[2]), argv[3], pid);
} else if (argc == 7 && strcmp(argv[1], "--seq") == 0) {
return seq(atoi(argv[2]), strtoul(argv[3], NULL, 0), strtoul(argv[4], NULL, 0), atoi(argv[5]), argv[6]);
} else if (argc == 1) {
printf("eyetracking pid %d\n", pid);
for (int b = 0; b < nbufs; b++)
printf("buffer %d: fd %d, %zu bytes, ino %lu\n", b, bufs[b].xfd, bufs[b].size, bufs[b].ino);
} else {
usage();
return 2;
}
return 0;
}
+453
View File
@@ -0,0 +1,453 @@
#!/usr/bin/env python3
"""ft-eyes: our own eye tracker, live.
Reads the eye-camera frames that ft-eyegrab (the root service frametop-eyegrab, installed by
gaze/tracker/install.sh) keeps in /dev/shm/frametop-eyes-cams, finds each eye's pupil and
glint pair (eyes_pupil.py), turns them into a gaze with the saved calibration and a running
slip estimate (eyes_model.py), and publishes the result in /dev/shm/frametop-eyes-gaze,
where ft-gaze reads it as the source "own". It touches /dev/shm/frametop-eyes-want every
second, and ft-eyegrab copies frames only while someone does.
The gaze service (gaze/ft-gazed) runs it in the dev container, with --watch-stdin (it quits
when its stdin closes), while Eye tracker is Own tracker or the gaze probe uses it. The
probe calibrates and teaches it; the gaze pointer's nudges reach it as clicks, from ft-gazed.
By hand: distrobox enter dev -- python3 gaze/tracker/ft-eyes -v
Calibration: ~/.local/state/frametop/gaze/eyes/calibration.json, made by the gaze probe's
calibration with the tracker toggle on Own tracker (or lab/ft-eyes-score --save CAPTURE).
Each eye's shift since the calibration (eyes_model.Shift, taught by clicks) is kept in
state.json next to it. After a restart, or frames stopping for GAP seconds (the headset
off), the next click starts that eye's shift over, since the headset may sit differently now.
Control socket: abstract datagram "@ft_eyes"; each command gets one reply line.
status JSON: calibration, and per eye its shift, clicks, whether a
glint jump is applied, and "reseat" (the next click starts
the shift over: the probe asks for a one-dot check then)
calib-start a new calibration: collect dots from now on
calib-point T0 T1 YAW PITCH a dot you looked at from T0 to T1 (CLOCK_MONOTONIC_RAW) in
that direction (head-relative degrees): "ok N0 N1 SD0 SD1"
(frames and spread in px per eye, right first) or "fail WHY"
calib-fit fit the dots, save, and start the shifts and clicks over
click T YAW PITCH a click: you were looking there just before T; teaches the
shift: "ok DX0 DY0 DX1 DY1" (the shift each eye measured)
Environment, for replays (lab/ft-eyes-e2e): FT_EYES_CAMS, FT_EYES_GAZE, FT_EYES_STATE (the
state folder), FT_EYES_SOCKET (the socket's name). With FT_EYES_CAMS set, the want file isn't
touched.
/dev/shm/frametop-eyes-gaze, 128 bytes, little-endian (mirrored in ft-gaze.cpp):
0 u32 seq odd while it's being written: read it before and after, retry if it moved
4 u32 version 1
8 f64 t the newest frame's time, CLOCK_MONOTONIC_RAW seconds
16 f32 yaw, pitch the gaze, head-relative degrees (yaw +left, pitch +up), eyes averaged
24 u32 flags bit 0: right eye in it, 1: left eye in it, 2: right shift from clicks, 3: left
28 u32 n samples published
32 f32 x4 right yaw, pitch, left yaw, pitch (NaN when that eye isn't seen)
48 f32 x4 each eye's shift since the calibration, pixels: right x, y, left x, y
64 f32 x4 pupil centre, pixels: right x, y, left x, y
"""
import json
import math
import mmap
import os
import socket
import struct
import sys
import threading
import time
from collections import deque
from pathlib import Path
import numpy as np
sys.path.insert(0, str(Path(__file__).resolve().parent))
import eyes_model # noqa: E402
import eyes_pupil # noqa: E402
CAMS = os.environ.get("FT_EYES_CAMS", "/dev/shm/frametop-eyes-cams") # overrides for replays
OUT = os.environ.get("FT_EYES_GAZE", "/dev/shm/frametop-eyes-gaze")
WANT = None if "FT_EYES_CAMS" in os.environ else Path("/dev/shm/frametop-eyes-want")
OUT_SIZE = 128
CALIBRATION = Path(os.environ.get("FT_EYES_STATE", Path.home() / ".local/state/frametop/gaze/eyes")) / "calibration.json"
W, H = 512, 400
FRESH = 0.03 # an eye's reading counts toward the output for this long (s)
LOST_EVERY = 3 # while an eye is lost, search the whole frame only every 3rd frame
EYES = ("right", "left")
STATE = CALIBRATION.parent / "state.json"
CLICKS = CALIBRATION.parent / "clicks.jsonl"
SOCKET = "\0" + os.environ.get("FT_EYES_SOCKET", "ft_eyes")
HISTORY = 12.0 # seconds of pupil positions kept per eye, for dots and clicks (the pointer's
# clicks come from ft-gazed up to 10 s after the look)
GAP = 3.0 # s without frames: the headset was off, and may sit differently now
CLICK_BEFORE = 0.3 # a click's frames: the 300 ms before it (like the probe's fixation)
CALIB_MIN = 15 # frames an eye needs in a calibration dot's window
CALIB_SPREAD = 4.0 # px: more than this and the eye moved during the dot
class Cams:
"""The shared frames. Header and entry layout: ft-eyegrab.c, share_head_t/share_entry_t."""
def __init__(self):
fd = os.open(CAMS, os.O_RDONLY)
try:
self.ino = os.fstat(fd).st_ino
self.mm = mmap.mmap(fd, 0, prot=mmap.PROT_READ)
finally:
os.close(fd)
magic, version, w, h, self.slots, self.esize = struct.unpack_from("<6I", self.mm, 0)
if magic != 0x31434546 or version != 1 or (w, h) != (W, H):
raise RuntimeError(f"{CAMS}: unexpected header")
def replaced(self):
"""ft-eyegrab restarted: it makes a new file, and this one is stale."""
try:
return os.stat(CAMS).st_ino != self.ino
except OSError:
return True
def count(self, cam):
return struct.unpack_from("<Q", self.mm, 24 + 8 * cam)[0]
def tracker(self):
return struct.unpack_from("<I", self.mm, 40)[0]
def frame(self, cam, n):
"""Frame n of a camera as (time, array), or None if it was overwritten meanwhile."""
off = 64 + (cam * self.slots + n % self.slots) * self.esize
for _ in range(3):
seq = struct.unpack_from("<Q", self.mm, off)[0]
if seq & 1:
continue
img = np.frombuffer(self.mm, np.uint8, W * H, off + 64).reshape(H, W).copy()
t, got = struct.unpack_from("<dQ", self.mm, off + 8)
if struct.unpack_from("<Q", self.mm, off)[0] == seq and got == n:
return t, img
return None
class Out:
def __init__(self):
fd = os.open(OUT, os.O_RDWR | os.O_CREAT | os.O_NOFOLLOW, 0o600)
try:
os.ftruncate(fd, OUT_SIZE)
self.mm = mmap.mmap(fd, OUT_SIZE)
finally:
os.close(fd)
self.seq = (struct.unpack_from("<I", self.mm, 0)[0] + 1) & ~1 # even: at rest
self.n = 0
def write(self, t, gaze, flags, eyes, slips, pupils):
struct.pack_into("<I", self.mm, 0, self.seq + 1) # odd while writing
self.n += 1
struct.pack_into("<IdffII12f", self.mm, 4, 1, t, gaze[0], gaze[1], flags, self.n,
*eyes, *slips, *pupils)
self.seq = (self.seq + 2) & 0xFFFFFFFE
struct.pack_into("<I", self.mm, 0, self.seq)
class Eye:
def __init__(self, eye):
self.eye = eye
self.cal = None
self.slip = None
self.shift = eyes_model.Shift()
self.history = deque() # (t, x, y, glint mid x, y)
self.last = None # the previous pupil (window hint)
self.gaze = None # (t, yaw, pitch, x, y)
self.lost = 0
self.frames = self.found = 0
self.work = 0.0
self.last_t = None # the previous frame's time
def use(self, cal, shift=None):
self.cal = cal if cal is not None and cal.has("pupil", self.eye) else None
self.slip = eyes_model.SlipTracker(cal, self.eye, window=eyes_model.JUMP_WINDOW) if self.cal else None
self.shift = shift or eyes_model.Shift()
self.gaze = None
def feed(self, t, img):
self.frames += 1
if self.last_t is not None and t - self.last_t > GAP:
self.shift.reseat()
self.last_t = t
if self.last is None:
self.lost += 1
if self.lost % LOST_EVERY:
return
t0 = time.perf_counter()
p = eyes_pupil.find_pupil(img, self.last)
self.last = p
if p is not None:
self.found += 1
self.lost = 0
pair = eyes_pupil.glint_pair(p)
mid = eyes_model.pair_mid(pair) if pair else (math.nan, math.nan)
self.history.append((t, p["x"], p["y"], mid[0], mid[1]))
while self.history and self.history[0][0] < t - HISTORY:
self.history.popleft()
if self.cal:
if pair:
self.slip.add(t, (p["x"], p["y"]), mid)
self.shift.glint(self.slip.get(), t)
g = self.cal.gaze(self.eye, p["x"], p["y"], self.shift.value)
self.gaze = (t, float(g[0]), float(g[1]), p["x"], p["y"])
self.work += time.perf_counter() - t0
def window(self, t0, t1):
"""Median pupil and glint midpoint over [t0, t1], the frame count, and the spread."""
rows = np.array([r for r in self.history if t0 <= r[0] <= t1]).reshape(-1, 5)
if len(rows) == 0:
return None
pupil = np.median(rows[:, 1:3], axis=0)
spread = float(np.median(np.hypot(*(rows[:, 1:3] - pupil).T)))
mids = rows[~np.isnan(rows[:, 3]), 3:5]
mid = np.median(mids, axis=0) if len(mids) >= 3 else None
return dict(pupil=pupil, mid=mid, n=len(rows), spread=spread)
class Tracker:
def __init__(self):
self.eyes = [Eye(0), Eye(1)]
self.cal = None
self.dots = [] # calibration dots so far, in ft-eyes-score's click form
self.calibrating = False
if CALIBRATION.exists():
self.cal = eyes_model.Calibration.load(CALIBRATION)
if not self.cal.spread:
self.cal.spread = spread_from_dots(self.cal)
shifts = {}
try:
d = json.loads(STATE.read_text())
if self.cal and d.get("calibration") == self.cal.info.get("made"):
shifts = {int(k): eyes_model.Shift.from_json(v) for k, v in d.get("shift", {}).items()}
except (OSError, ValueError):
pass
for e in self.eyes:
e.use(self.cal, shifts.get(e.eye))
# Kept from the last run, but the headset may have been off since: the first
# click starts the history over (the saved shift is used until then).
e.shift.reseat()
def save_state(self):
d = {"calibration": self.cal.info.get("made") if self.cal else None,
"shift": {e.eye: e.shift.to_json() for e in self.eyes}}
tmp = STATE.with_suffix(".tmp")
tmp.write_text(json.dumps(d))
tmp.replace(STATE)
def command(self, line):
w = line.split()
if not w:
return "fail empty"
if w[0] == "status":
return json.dumps(self.status())
if w[0] == "calib-start":
self.dots, self.calibrating = [], True
return "ok"
if w[0] == "calib-point" and len(w) == 5:
if not self.calibrating:
return "fail no calibration started"
t0, t1, yaw, pitch = map(float, w[1:])
got = {e.eye: e.window(t0, t1) for e in self.eyes}
for c, g in got.items():
if g is None or g["n"] < CALIB_MIN:
return f"fail the {EYES[c]} eye was seen in only {0 if g is None else g['n']} frames"
if g["spread"] > CALIB_SPREAD:
return f"fail the {EYES[c]} eye moved ({g['spread']:.1f} px)"
self.dots.append(dict(truth=(yaw, pitch), eye={c: dict(pupil=g["pupil"], mid=g["mid"]) for c, g in got.items()}))
return "ok {} {} {:.2f} {:.2f}".format(got[0]["n"], got[1]["n"], got[0]["spread"], got[1]["spread"])
if w[0] == "calib-fit":
if len(self.dots) < eyes_model.MIN_CLICKS:
return f"fail only {len(self.dots)} dots (need {eyes_model.MIN_CLICKS})"
cal = eyes_model.Calibration.fit(self.dots, {"made": time.strftime("%Y-%m-%d %H:%M:%S"),
"dots": len(self.dots), "from": "probe calibration"})
if not all(cal.has(n, c) for n in ("pupil", "where") for c in (0, 1)):
return "fail not enough dots with both eyes"
errs = [float(np.hypot(*(np.mean([cal.gaze(c, *k["eye"][c]["pupil"]) for c in (0, 1)], axis=0)
- k["truth"]))) for k in self.dots]
if CALIBRATION.exists():
CALIBRATION.replace(CALIBRATION.with_name(time.strftime("calibration-%Y%m%d-%H%M%S.json")))
cal.save(CALIBRATION)
self.cal, self.calibrating = cal, False
for e in self.eyes:
e.use(cal)
self.save_state()
with open(CALIBRATION.with_name("calibration-dots.jsonl"), "a") as f:
for k in self.dots:
f.write(json.dumps({"made": cal.info["made"], "truth": k["truth"],
"eye": {c: {"pupil": v["pupil"].tolist(),
"mid": None if v["mid"] is None else v["mid"].tolist()}
for c, v in k["eye"].items()}}) + "\n")
return (f"ok {len(self.dots)} dots, fit median {np.median(errs):.2f} deg, eyes "
+ ", ".join(f"{EYES[c]} {cal.spread[c]:.2f}" for c in sorted(cal.spread)))
if w[0] == "click" and len(w) == 4:
if not self.cal:
return "fail not calibrated"
t, yaw, pitch = map(float, w[1:])
out, rec = [], {"time": time.time(), "t": t, "truth": [yaw, pitch], "eyes": {}}
for e in self.eyes:
g = e.window(t - CLICK_BEFORE, t)
if g is None or g["n"] < 5 or not e.cal:
out += ["nan", "nan"]
continue
d = self.cal.click_shift(e.eye, g["pupil"], (yaw, pitch))
before = e.shift.value.tolist()
e.shift.click(d, e.slip.value if e.slip else None)
rec["eyes"][e.eye] = {"pupil": g["pupil"].tolist(), "measured": d.tolist(), "before": before,
"after": e.shift.value.tolist()}
out += [f"{d[0]:.2f}", f"{d[1]:.2f}"]
self.save_state()
with open(CLICKS, "a") as f:
f.write(json.dumps(rec) + "\n")
return "ok " + " ".join(out)
return f"fail unknown command {w[0]}"
def status(self):
return {"calibration": self.cal.info if self.cal else None, "calibrating": self.calibrating,
"dots": len(self.dots),
"eyes": {EYES[e.eye]: {"shift": e.shift.value.tolist(), "clicks": len(e.shift.meas),
"jump": bool(np.any(e.shift.jump)),
"reseat": e.shift.reseated} for e in self.eyes}}
def spread_from_dots(cal):
"""The eyes' fit spreads for a calibration saved without them, from its dots in
calibration-dots.jsonl ({} if they aren't there: the eyes are then weighted alike)."""
try:
lines = CALIBRATION.with_name("calibration-dots.jsonl").read_text().splitlines()
except OSError:
return {}
dots = []
for line in lines:
d = json.loads(line)
if d.get("made") == cal.info.get("made"):
dots.append(dict(truth=d["truth"], eye={
int(c): dict(pupil=np.array(v["pupil"]), mid=None if v["mid"] is None else np.array(v["mid"]))
for c, v in d["eye"].items()}))
return eyes_model.Calibration.fit(dots).spread if dots else {}
class Want:
"""Touches the want file every second, so ft-eyegrab keeps copying frames."""
def __init__(self):
self.at = 0.0
def __call__(self):
if WANT is None or time.monotonic() - self.at < 1.0:
return
self.at = time.monotonic()
try:
fd = os.open(WANT, os.O_WRONLY | os.O_CREAT | os.O_NOFOLLOW | os.O_CLOEXEC, 0o600)
os.utime(fd)
os.close(fd)
except OSError as e:
print(f"ft-eyes: {WANT}: {e}", file=sys.stderr, flush=True)
def wait_for_cams(sock, tracker, want):
while True:
want()
try:
return Cams()
except (OSError, ValueError, RuntimeError):
serve(sock, tracker)
time.sleep(0.2)
def serve(sock, tracker):
while True:
try:
data, addr = sock.recvfrom(512)
except BlockingIOError:
return
try:
reply = tracker.command(data.decode(errors="replace").strip())
except Exception as ex: # a bad command must not take the tracker down
reply = f"fail {type(ex).__name__}: {ex}"
if addr:
try:
sock.sendto(reply.encode(), addr)
except OSError:
pass
def main():
verbose = "-v" in sys.argv
if "--watch-stdin" in sys.argv:
# Run by ft-gazed through distrobox, which doesn't pass a stop on: quit when our
# stdin (its pipe) closes.
def watch():
while sys.stdin.buffer.read(4096):
pass
os._exit(0)
threading.Thread(target=watch, daemon=True).start()
want = Want()
CALIBRATION.parent.mkdir(parents=True, exist_ok=True)
tracker = Tracker()
eyes = tracker.eyes
out = Out()
sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
sock.bind(SOCKET)
sock.setblocking(False)
cal = tracker.cal
print("ft-eyes: " + (f"calibration from {cal.info.get('made')} ({cal.info.get('from', cal.info.get('capture'))})"
if cal else "not calibrated: run the probe's calibration with the Own tracker")
+ f"; waiting for {CAMS}", file=sys.stderr, flush=True)
cams = wait_for_cams(sock, tracker, want)
print("ft-eyes: frames found, tracking", file=sys.stderr, flush=True)
seen = [cams.count(0), cams.count(1)]
report = time.monotonic()
while True:
serve(sock, tracker)
want()
new = False
for c in (0, 1):
n = cams.count(c)
if n == seen[c]:
continue
seen[c] = n
got = cams.frame(c, n - 1) # only the newest: never fall behind
if got:
eyes[c].feed(*got)
new = True
if new:
latest = max((e.gaze[0] for e in eyes if e.gaze), default=None)
use = [e for e in eyes if e.gaze and latest - e.gaze[0] < FRESH]
if use:
yaw, pitch = tracker.cal.combine({e.eye: e.gaze[1:3] for e in use})
flags, per, shifts, pupils = 0, [], [], []
for i, e in enumerate(eyes):
fresh = e in use
flags |= (1 << i) if fresh else 0
flags |= (4 << i) if e.shift.meas else 0
per += [e.gaze[1], e.gaze[2]] if fresh else [math.nan, math.nan]
shifts += [float(v) for v in e.shift.value]
pupils += [e.gaze[3], e.gaze[4]] if fresh else [math.nan, math.nan]
out.write(latest, (yaw, pitch), flags, per, shifts, pupils)
else:
time.sleep(0.001)
now = time.monotonic()
if now - report >= 5:
if verbose:
parts = []
for e in eyes:
s = e.shift.value
parts.append(f"{EYES[e.eye]} {e.frames / 5:.0f} fps, found {e.found / max(e.frames, 1):.0%}, "
f"{e.work / max(e.frames, 1) * 1000:.2f} ms/frame, shift ({s[0]:+.1f},{s[1]:+.1f})"
f" from {len(e.shift.meas)} clicks" + (", jump" if np.any(e.shift.jump) else ""))
e.frames = e.found = 0
e.work = 0.0
print("ft-eyes: " + "; ".join(parts), file=sys.stderr, flush=True)
report = now
if cams.replaced():
print("ft-eyes: frames went away; waiting", file=sys.stderr, flush=True)
cams = wait_for_cams(sock, tracker, want)
seen = [cams.count(0), cams.count(1)]
if __name__ == "__main__":
try:
main()
except KeyboardInterrupt:
pass
+58
View File
@@ -0,0 +1,58 @@
#!/usr/bin/env bash
# Install (or remove) the frame grabber our own eye tracker needs: ft-eyegrab, as the system
# service frametop-eyegrab.service. It copies the eye-camera frames, read-only, out of
# SteamVR's eyetracking process into /dev/shm/frametop-eyes-cams for ft-eyes, and only while
# ft-eyes wants them. The gaze service (gaze/ft-gazed) runs ft-eyes itself, when Eye tracker
# is Own tracker or the gaze probe uses it.
# Needs host sudo, for the binary (/etc/frametop/ft-eyegrab, root's) and the unit. On the
# Frame, sudo asks for the password in the terminal, or runs SUDO_ASKPASS when that's set.
# From a PC (or with no terminal), the password comes from steamos_root_pwd in the repo's .env
# and is sent to sudo -S on stdin, never on a command line.
# Usage: gaze/tracker/install.sh [install|uninstall|status|log [lines]]
set -euo pipefail
root=$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)
. "$root/scripts/_env.sh"
src=$FRAME_REPO/gaze/tracker
unit=frametop-eyegrab.service
sudo_run() {
if [ "$FRAME_LOCAL" = 1 ] && [ -n "${SUDO_ASKPASS:-}" ]; then
sudo -A bash -c "$1" # SUDO_ASKPASS supplies the password
return
fi
if [ "$FRAME_LOCAL" = 1 ] && [ -t 0 ]; then
sudo bash -c "$1" # asks for the password here
return
fi
local pw
pw=$(sed -n 's/^steamos_root_pwd=//p' "$root/.env" 2>/dev/null)
pw=${pw#[\"\']}; pw=${pw%[\"\']} # .env values may be quoted
[ -n "$pw" ] || { echo "no terminal for sudo, and steamos_root_pwd is missing from $root/.env" >&2; exit 1; }
printf '%s\n' "$pw" | on_frame "sudo -S -p '' bash -c $(printf %q "$1")"
}
case ${1:-install} in
install)
"$root/gaze/tracker/build.sh"
ids=$(on_frame 'echo "$(id -u):$(id -g)"')
fill_template "$root/gaze/tracker/$unit" | sed "s|@UID@|${ids%:*}|g; s|@GID@|${ids#*:}|g" |
on_frame "cat > /tmp/$unit"
sudo_run "set -e
install -D -m 0755 -o root -g root $src/build/ft-eyegrab /etc/frametop/ft-eyegrab
install -D -m 0644 -o root -g root /tmp/$unit /etc/systemd/system/$unit
rm -f /tmp/$unit
systemctl daemon-reload
systemctl enable $unit
systemctl restart $unit
sleep 1
echo \"$unit: \$(systemctl is-active $unit)\""
;;
uninstall)
sudo_run "systemctl disable --now $unit 2>/dev/null
rm -f /etc/systemd/system/$unit /etc/frametop/ft-eyegrab
rmdir /etc/frametop 2>/dev/null; systemctl daemon-reload; echo removed" ;;
status) on_frame "systemctl is-active $unit; ls -l /dev/shm/frametop-eyes-cams 2>/dev/null" || true ;;
log) on_frame "journalctl -u $unit --no-pager -o cat -n ${2:-20}" ;;
*) echo "usage: $0 [install|uninstall|status|log [lines]]" >&2; exit 2 ;;
esac
+31
View File
@@ -0,0 +1,31 @@
"""What the lab tools share: where recordings are kept, and the tracker's modules.
Recordings of the eye cameras are biometric data. They're kept outside the repo, in
~/.local/share/frametop/eyes/captures (FT_EYES_CAPTURES overrides), one folder each, 0700,
and never leave the Frame except for the 7i's copies frame-job makes for offline jobs.
"""
import os
import sys
from pathlib import Path
TRACKER = Path(__file__).resolve().parents[1] # gaze/tracker: ft-eyes, eyes_model, eyes_pupil
LAB = Path(__file__).resolve().parent
CAPTURES = Path(os.environ.get("FT_EYES_CAPTURES", Path.home() / ".local/share/frametop/eyes/captures"))
sys.path.insert(0, str(TRACKER))
def capture(arg):
"""A recording's folder: a path as given, or a bare name in CAPTURES."""
p = Path(arg).expanduser()
return p if p.exists() or os.sep in arg else CAPTURES / arg
def new_capture(name):
"""A new, private recording folder in CAPTURES; exits if it's already there."""
out = CAPTURES / name
if out.exists():
sys.exit(f"{out} exists")
out.mkdir(parents=True)
for d in (CAPTURES, out):
d.chmod(0o700)
return out
+41
View File
@@ -0,0 +1,41 @@
"""Per-frame pupil ellipses through a recording, cached in the capture (numbers only).
ellipses(cap) -> {camera: array of rows (t, x, y, a, b, major, fill, glint mid x, y)}, every
EVERY-th frame of each camera; the glint midpoint is NaN when the pair isn't seen.
"""
import numpy as np
import eyes_lab # noqa: F401 (puts gaze/tracker on the path)
import eyes_pupil
EVERY = 2
VERSION = 2
def load_index(cap):
return np.array([[float(v) for v in l.split()]
for l in (cap / "index.txt").read_text().splitlines() if len(l.split()) == 4])
def ellipses(cap):
cache = cap / f"ellipses-v{VERSION}.npz"
if cache.exists():
z = np.load(cache)
return {0: z["cam0"], 1: z["cam1"]}
idx = load_index(cap)
frames = np.memmap(cap / "frames.raw", dtype=np.uint8, mode="r").reshape(-1, 400, 512)
idx = idx[:len(frames)]
out = {}
for c in (0, 1):
rows, prev = [], None
for i in np.where(idx[:, 2] == c)[0][::EVERY]:
p = eyes_pupil.find_pupil(frames[i], prev)
prev = p
if p is None:
continue
pair = eyes_pupil.glint_pair(p)
mx, my = ((pair[0][0] + pair[1][0]) / 2, (pair[0][1] + pair[1][1]) / 2) if pair else (np.nan, np.nan)
rows.append((idx[i, 3], p["x"], p["y"], p["a"], p["b"], p["major"], p["fill"], mx, my))
out[c] = np.array(rows).reshape(-1, 9)
np.savez(cache, cam0=out[0], cam1=out[1])
return out
+183
View File
@@ -0,0 +1,183 @@
#!/usr/bin/env python3
"""ft-eyes-e2e: the live path end to end on two recordings, without the headset.
Starts a scratch ft-eyes (its own shared memory, socket, and state folder, so the real
calibration is untouched), then:
1. plays CALIB into it and sends each of its practice clicks as a calibration dot (the
300 ms before the press), as the probe's calibration would, and fits;
2. plays TEST into it and, at each of its practice clicks, scores what ft-eyes was
publishing in the 300 ms before the press, then sends the click, as the probe does.
So every TEST click is scored with only earlier data, like ft-eyes-score's `clicks` method.
Usage: frame-job -- lab/py lab/ft-eyes-e2e CALIB TEST [--for S] [--dump FILE]
CALIB and TEST are recordings (a bare name is in eyes_lab.CAPTURES; give full paths for
frame-job to copy them to the 7i). Runs at the recorded pace (about the two recordings'
length). --dump saves everything ft-eyes published during TEST (OUT_FIELDS per row) and each
click's score, as a pickle, for looking into the bad clicks.
"""
import json
import mmap
import os
import pickle
import re
import shutil
import socket
import struct
import subprocess
import sys
import tempfile
import time
from pathlib import Path
import numpy as np
sys.path.insert(0, str(Path(__file__).resolve().parent))
from eyes_lab import LAB, TRACKER, capture # noqa: E402
BEFORE = 0.3 # the probe's fixation window before a press (s)
# ft-eyes' output after its seq and version (ft-eyes' docstring): one row per sample.
OUT_FORMAT = "<dffII12f"
OUT_FIELDS = ("t", "yaw", "pitch", "flags", "n", "r_yaw", "r_pitch", "l_yaw", "l_pitch",
"r_shift_x", "r_shift_y", "l_shift_x", "l_shift_y", "r_pupil_x", "r_pupil_y", "l_pupil_x", "l_pupil_y")
class Run:
def __init__(self):
tag = f"ft-eyes-e2e-{os.getpid()}"
self.state = Path(tempfile.mkdtemp(prefix=tag + "-"))
self.env = dict(os.environ, FT_EYES_CAMS=f"/dev/shm/{tag}-cams", FT_EYES_GAZE=f"/dev/shm/{tag}-gaze",
FT_EYES_STATE=str(self.state), FT_EYES_SOCKET=tag)
self.log = self.state / "ft-eyes.log"
self.trackd = subprocess.Popen([sys.executable, str(TRACKER / "ft-eyes"), "-v"], env=self.env,
stdout=subprocess.DEVNULL, stderr=open(self.log, "w"))
self.sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
self.sock.bind("\0" + tag + "-client")
self.sock.settimeout(2)
self.to = "\0" + tag
self.gaze = None
def cmd(self, line):
for _ in range(50): # ft-eyes may not have bound its socket yet
try:
self.sock.sendto(line.encode(), self.to)
return self.sock.recv(4096).decode()
except ConnectionRefusedError:
time.sleep(0.2)
raise RuntimeError("ft-eyes never answered")
def latest_frame_t(self):
"""The newest replayed frame's time, or None before the replay starts."""
try:
with open(self.env["FT_EYES_CAMS"], "rb") as f:
mm = mmap.mmap(f.fileno(), 0, prot=mmap.PROT_READ)
except (OSError, ValueError):
return None
slots, esize = struct.unpack_from("<2I", mm, 16)
best = None
for cam in (0, 1):
n = struct.unpack_from("<Q", mm, 24 + 8 * cam)[0]
if n:
t = struct.unpack_from("<d", mm, 64 + (cam * slots + (n - 1) % slots) * esize + 8)[0]
best = t if best is None else max(best, t)
return best
def published(self):
"""ft-eyes' newest output (OUT_FIELDS), or None."""
if self.gaze is None:
try:
with open(self.env["FT_EYES_GAZE"], "rb") as f:
self.gaze = mmap.mmap(f.fileno(), 128, prot=mmap.PROT_READ)
except (OSError, ValueError):
return None
for _ in range(3):
seq = struct.unpack_from("<I", self.gaze, 0)[0]
v = struct.unpack_from(OUT_FORMAT, self.gaze, 8)
if not seq & 1 and struct.unpack_from("<I", self.gaze, 0)[0] == seq:
return v
return None
def stage(self, cap, secs, handle):
"""Play `cap` and call handle(click, samples) as each click's press time goes by.
Returns everything ft-eyes published meanwhile."""
clicks = pickle.load(open(cap / "features.pkl", "rb"))["clicks"]
replay = subprocess.Popen([sys.executable, str(LAB / "ft-eyes-replay"), str(cap), self.env["FT_EYES_CAMS"],
"--for", str(secs)], stdout=subprocess.DEVNULL)
todo, samples = list(clicks), []
while replay.poll() is None:
t = self.latest_frame_t()
v = self.published()
if v and v[0] > 0 and (not samples or v[0] != samples[-1][0]):
samples.append(v)
while t and todo and t > todo[0]["t"] + 0.05:
handle(todo.pop(0), samples)
time.sleep(0.003)
return samples
def close(self):
self.trackd.terminate()
self.trackd.wait()
for p in (self.env["FT_EYES_CAMS"], self.env["FT_EYES_GAZE"]):
if os.path.exists(p):
os.unlink(p)
shutil.rmtree(self.state, ignore_errors=True)
def main(argv):
args = [a for i, a in enumerate(argv) if not a.startswith("--") and (i == 0 or argv[i - 1] not in ("--for", "--dump"))]
if len(args) != 2:
sys.exit(__doc__)
calib, test = map(capture, args)
secs = argv[argv.index("--for") + 1] if "--for" in argv else "1e9"
dump = Path(argv[argv.index("--dump") + 1]) if "--dump" in argv else None
run = Run()
try:
print("calib-start:", run.cmd("calib-start"), flush=True)
dots = []
run.stage(calib, secs, lambda k, _s: dots.append(
run.cmd(f"calib-point {k['t'] - BEFORE} {k['t']} {k['truth'][0]} {k['truth'][1]}")))
fails = [d for d in dots if not d.startswith("ok")]
print(f"{calib.name}: {len(dots) - len(fails)} of {len(dots)} clicks taken as dots", flush=True)
whys = [re.sub(r"[\d.]+", "N", f) for f in fails]
for why in sorted(set(whys)):
print(f" {whys.count(why)} x {why}")
print("calib-fit:", run.cmd("calib-fit"), flush=True)
# ft-eyes-replay removes its file at the end; ft-eyes notices within 5 s and waits for the next.
time.sleep(6)
scored = []
def click(k, samples):
s = np.array([v for v in samples if k["t"] - BEFORE <= v[0] <= k["t"]]).reshape(-1, len(OUT_FIELDS))
err = float(np.hypot(*(np.median(s[:, 1:3], axis=0) - k["truth"]))) if len(s) >= 5 else None
reply = run.cmd(f"click {k['t']} {k['truth'][0]} {k['truth'][1]}")
st = json.loads(run.cmd("status"))["eyes"]
scored.append((k["t"], err, reply, [(v["clicks"], v["jump"]) for v in st.values()]))
test_samples = run.stage(test, secs, click)
if dump:
with open(dump, "wb") as f:
pickle.dump({"fields": OUT_FIELDS, "samples": np.array(test_samples, float),
"clicks": [dict(t=x[0], err=x[1], reply=x[2], eyes=x[3]) for x in scored]}, f)
print("dumped to", dump)
e = np.array([x[1] for x in scored if x[1] is not None])
print(f"{test.name}: {len(e)} of {len(scored)} clicks scored live", flush=True)
if len(e):
print(f" median {np.median(e):.2f} deg, 90% {np.percentile(e, 90):.2f}, "
f"after the first 5: median {np.median(e[5:]):.2f}")
print(" clicks ft-eyes refused:", sum(not x[2].startswith("ok") for x in scored))
t0 = scored[0][0] if scored else 0
print(" clicks that started an eye's shift over after a jump: right {}, left {}".format(
*(sum(x[3][c][0] == 1 for x in scored[1:]) for c in (0, 1))))
print(" by time (s): " + ", ".join(
f"{lo}-{lo + 30}: {np.median(b):.2f}" for lo in range(0, 300, 30)
if len(b := [x[1] for x in scored if x[1] is not None and lo <= x[0] - t0 < lo + 30])))
print(" worst: " + ", ".join(
f"{x[0] - t0:.0f}s {x[1]:.1f} (clicks/jump R {x[3][0][0]}/{x[3][0][1]:d} L {x[3][1][0]}/{x[3][1][1]:d})"
for x in sorted((x for x in scored if x[1] is not None), key=lambda x: -x[1])[:10]))
print("status:", run.cmd("status"))
print("ft-eyes' last report:", run.log.read_text().strip().splitlines()[-1:])
finally:
run.close()
if __name__ == "__main__":
main(sys.argv[1:])
+143
View File
@@ -0,0 +1,143 @@
#!/usr/bin/env python3
"""ft-eyes-record: record a live session from the shared frames, so it can be replayed and
scored later (ft-eyes-score, ft-eyes-e2e), without root and alongside ft-eyes.
Reads /dev/shm/frametop-eyes-cams (ft-eyegrab, frametop-eyegrab.service) and writes every
frame of both cameras to a new recording NAME in eyes_lab.CAPTURES: frames.raw, index.txt
("<n> <slot> <camera> <time>", as ft-eyegrab --rec), and clocks.txt (wall clock and
CLOCK_MONOTONIC_RAW, read together). It touches /dev/shm/frametop-eyes-want every second, so
ft-eyegrab copies frames even without ft-eyes. About 2 GB a minute. Stops after SECONDS
(default 900), on Ctrl-C or SIGTERM, or when the disk gets below MIN_FREE_GB. The probe's
clicks are added later, on the Frame: `lab/py lab/ft-eyes-score --clicks NAME`.
Usage: lab/ft-eyes-record NAME [SECONDS] (host Python is enough: no numpy)
"""
import mmap
import os
import shutil
import signal
import struct
import sys
import time
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
from eyes_lab import new_capture # noqa: E402
CAMS = os.environ.get("FT_EYES_CAMS", "/dev/shm/frametop-eyes-cams")
WANT = "/dev/shm/frametop-eyes-want"
W, H = 512, 400
MIN_FREE_GB = 20
class Share:
"""The shared frames. Layout: ft-eyegrab.c, share_head_t/share_entry_t."""
def __init__(self):
fd = os.open(CAMS, os.O_RDONLY)
try:
self.ino = os.fstat(fd).st_ino
self.mm = mmap.mmap(fd, 0, prot=mmap.PROT_READ)
finally:
os.close(fd)
magic, version, w, h, self.slots, self.esize = struct.unpack_from("<6I", self.mm, 0)
if magic != 0x31434546 or version != 1 or (w, h) != (W, H):
raise RuntimeError(f"{CAMS}: unexpected header")
def replaced(self):
try:
return os.stat(CAMS).st_ino != self.ino
except OSError:
return True
def count(self, cam):
return struct.unpack_from("<Q", self.mm, 24 + 8 * cam)[0]
def frame(self, cam, n):
"""(time, slot, bytes) of frame n of a camera, or None if it was overwritten."""
off = 64 + (cam * self.slots + n % self.slots) * self.esize
for _ in range(3):
seq = struct.unpack_from("<Q", self.mm, off)[0]
if seq & 1:
continue
data = self.mm[off + 64:off + 64 + W * H]
t, got, _cam, slot = struct.unpack_from("<dQII", self.mm, off + 8)
if struct.unpack_from("<Q", self.mm, off)[0] == seq and got == n:
return t, slot, data
return None
def touch_want():
"""Tell ft-eyegrab someone wants frames (it idles otherwise)."""
if "FT_EYES_CAMS" in os.environ:
return
try:
fd = os.open(WANT, os.O_WRONLY | os.O_CREAT | os.O_NOFOLLOW | os.O_CLOEXEC, 0o600)
os.utime(fd)
os.close(fd)
except OSError:
pass
def open_share(deadline):
while time.monotonic() < deadline:
touch_want()
try:
return Share()
except (OSError, ValueError, RuntimeError):
time.sleep(0.5)
sys.exit(f"ft-eyes-record: no {CAMS} (is frametop-eyegrab.service running? gaze/tracker/install.sh)")
def main(argv):
if not argv or argv[0].startswith("-"):
sys.exit(__doc__)
secs = float(argv[1]) if len(argv) > 1 else 900.0
out = new_capture(argv[0])
stop = []
for sig in (signal.SIGINT, signal.SIGTERM):
signal.signal(sig, lambda *_: stop.append(1))
share = open_share(time.monotonic() + 10)
(out / "clocks.txt").write_text(f"{time.time()} {time.clock_gettime(time.CLOCK_MONOTONIC_RAW)}\n")
seen = [share.count(0), share.count(1)]
written = dropped = 0
end = time.monotonic() + secs
check = touched = time.monotonic()
with open(out / "frames.raw", "wb") as frames, open(out / "index.txt", "w") as index:
while not stop and time.monotonic() < end:
new = []
for c in (0, 1):
n = share.count(c)
if n - seen[c] > share.slots: # fell behind: those frames are gone
dropped += n - seen[c] - share.slots
seen[c] = n - share.slots
for i in range(seen[c], n):
f = share.frame(c, i)
if f is None:
dropped += 1
else:
new.append((f[0], f[1], c, f[2]))
seen[c] = n
for t, slot, c, data in sorted(new, key=lambda f: f[0]):
frames.write(data)
index.write(f"{written} {slot} {c} {t:.6f}\n")
written += 1
if not new:
time.sleep(0.002)
if time.monotonic() - touched > 1:
touched = time.monotonic()
touch_want()
if time.monotonic() - check > 5:
check = time.monotonic()
if shutil.disk_usage(out).free < MIN_FREE_GB * 1e9:
print(f"ft-eyes-record: under {MIN_FREE_GB} GB free, stopping", file=sys.stderr)
break
if share.replaced():
print("ft-eyes-record: the frame share was restarted; following it", file=sys.stderr)
share = open_share(time.monotonic() + 10)
seen = [share.count(0), share.count(1)]
print(f"ft-eyes-record: {written} frames ({dropped} dropped) in {out}", file=sys.stderr)
if __name__ == "__main__":
main(sys.argv[1:])
+71
View File
@@ -0,0 +1,71 @@
#!/usr/bin/env python3
"""ft-eyes-replay: play a recording into the shared-frame layout, as ft-eyegrab --share
would, at the recorded pace, to test ft-eyes without the headset.
Usage: lab/py lab/ft-eyes-replay NAME [PATH] [--from S] [--for S]
NAME is a recording (a bare name is in eyes_lab.CAPTURES). PATH defaults to
/dev/shm/frametop-eyes-cams-replay; run ft-eyes with FT_EYES_CAMS=PATH (and FT_EYES_GAZE=...
so it doesn't overwrite the live output).
"""
import mmap
import os
import struct
import sys
import time
from pathlib import Path
import numpy as np
sys.path.insert(0, str(Path(__file__).resolve().parent))
from eyes_lab import capture # noqa: E402
W, H, SLOTS = 512, 400, 8
ESIZE = 64 + W * H
def main(argv):
cap = capture(argv[0])
path = argv[1] if len(argv) > 1 and not argv[1].startswith("--") else "/dev/shm/frametop-eyes-cams-replay"
start = float(argv[argv.index("--from") + 1]) if "--from" in argv else 0.0
length = float(argv[argv.index("--for") + 1]) if "--for" in argv else 1e9
idx = np.array([[float(v) for v in l.split()] for l in (cap / "index.txt").read_text().splitlines()
if len(l.split()) == 4])
frames = np.memmap(cap / "frames.raw", dtype=np.uint8, mode="r").reshape(-1, H, W)
size = 64 + 2 * SLOTS * ESIZE
if os.path.exists(path):
os.unlink(path)
fd = os.open(path, os.O_RDWR | os.O_CREAT | os.O_EXCL, 0o600)
os.ftruncate(fd, size)
mm = mmap.mmap(fd, size)
os.close(fd)
struct.pack_into("<6I2QI", mm, 0, 0x31434546, 1, W, H, SLOTS, ESIZE, 0, 0, os.getpid())
count = [0, 0]
t0 = idx[0, 3] + start
wall0 = time.monotonic()
try:
for i in range(len(frames)):
t = idx[i, 3]
if t < t0:
continue
if t - t0 > length:
break
delay = (t - t0) - (time.monotonic() - wall0)
if delay > 0:
time.sleep(delay)
cam = int(idx[i, 2])
n = count[cam]
off = 64 + (cam * SLOTS + n % SLOTS) * ESIZE
seq = struct.unpack_from("<Q", mm, off)[0]
struct.pack_into("<Q", mm, off, seq + 1)
mm[off + 64:off + 64 + W * H] = frames[i].tobytes()
struct.pack_into("<dQII", mm, off + 8, t, n, cam, int(idx[i, 1]))
struct.pack_into("<Q", mm, off, seq + 2)
count[cam] = n + 1
struct.pack_into("<Q", mm, 24 + 8 * cam, n + 1)
finally:
os.unlink(path)
print(f"replayed {sum(count)} frames")
if __name__ == "__main__":
main(sys.argv[1:])
+291
View File
@@ -0,0 +1,291 @@
#!/usr/bin/env python3
"""ft-eyes-score: score our pupil tracker and SteamVR against the gaze probe's practice clicks.
Usage (from gaze/tracker; A and B are recordings: a bare name is in eyes_lab.CAPTURES):
frame-job -- lab/py lab/ft-eyes-score A fit and test on A, leave-one-out
frame-job -- lab/py lab/ft-eyes-score A B fit on A, test on B
lab/py lab/ft-eyes-score --save A fit on A, save it for ft-eyes (on the Frame)
lab/py lab/ft-eyes-score --clicks A... on the Frame: copy each capture's practice
clicks from the probe's log into it (the first scoring does it too)
For frame-job to copy a recording to the 7i, give its full path, e.g.
~/.local/share/frametop/eyes/captures/A.
Each practice click gives a known gaze direction: SteamVR's raw gaze at the press plus the
angle from it to where you released (you were looking there). For each click we take the
frames from just before the press and find each eye's pupil and glint pair.
Methods, each a quadratic fit per eye, both eyes averaged when both are seen:
pupil the pupil centre alone. Breaks when the headset slips on the face.
glint pupil minus the glint pair's midpoint. Slip moves both alike, so this holds up,
but the right eye's pair is often off the cornea.
clicks the pupil centre minus the shift the earlier clicks measured, with the glints
only noticing a sudden jump (eyes_model.Shift). The one ft-eyes uses.
slip the pupil centre minus a slip estimate. Wherever the pair is seen, the glint
method gives the gaze, the fit says where the pupil should be for that gaze, and
the difference is the slip. Slip changes slowly, so the median over the last
30 seconds applies to every frame, with or without glints (see eyes_model.py).
SteamVR gets the same quadratic fit on its raw gaze.
"""
import json
import pickle
import sys
import time
from pathlib import Path
import numpy as np
sys.path.insert(0, str(Path(__file__).resolve().parent))
from eyes_lab import capture # noqa: E402
import eyes_model # noqa: E402
import eyes_pupil # noqa: E402
PRACTICE = Path.home() / ".local/state/frametop/gaze/practice.jsonl"
BEFORE = (0.25, 0.02) # frames from 250 ms to 20 ms before the press
EYES = {0: "right", 1: "left"}
MIN_FRAMES = 5
TRACK_EVERY = 9 # the slip track uses every 9th frame per camera (10 a second)
VERSION = 4 # bump when the features change, to rebuild the caches
# --- Features -------------------------------------------------------------------------
def load_index(cap):
# Skip a half-written last line (the recorder may still be running).
return np.array([[float(v) for v in l.split()]
for l in (cap / "index.txt").read_text().splitlines() if len(l.split()) == 4])
def eye_features(frames):
"""Median pupil centre and glint-pair midpoint over some frames of one eye."""
ps = [p for p in map(eyes_pupil.find_pupil, frames) if p]
if len(ps) < MIN_FRAMES:
return None
pupil = np.median([(p["x"], p["y"]) for p in ps], axis=0)
mids = []
for p in ps:
pair = eyes_pupil.glint_pair(p)
if pair:
mids.append(((pair[0][0] + pair[1][0]) / 2, (pair[0][1] + pair[1][1]) / 2))
mid = np.median(mids, axis=0) if len(mids) >= 3 else None
return dict(pupil=pupil, mid=mid)
def practice_log(cap, t0, t1, wall, mono):
"""The probe's practice records during the capture. The first time (on the Frame), cut
from the probe's log into the capture's practice.jsonl, so the capture carries its own
clicks (the 7i has no probe log)."""
own = cap / "practice.jsonl"
if not own.exists():
if not PRACTICE.exists():
sys.exit(f"{own} is missing: run `lab/py lab/ft-eyes-score --clicks {cap.name}` once on the Frame")
keep = [line for line in open(PRACTICE)
if t0 - 5 < json.loads(line)["time"] - wall + mono < t1 + 60]
own.write_text("".join(keep))
return [json.loads(line) for line in open(own)]
def features(cap):
"""Per-click and per-session features, cached in the capture (numbers only)."""
cache = cap / "features.pkl"
if cache.exists():
f = pickle.loads(cache.read_bytes())
if f.get("version") == VERSION:
return f
wall, mono = map(float, (cap / "clocks.txt").read_text().split())
idx = load_index(cap)
frames = np.memmap(cap / "frames.raw", dtype=np.uint8, mode="r").reshape(-1, 400, 512)
idx = idx[:len(frames)]
t0, t1 = idx[0, 3], idx[-1, 3]
clicks = []
for r in practice_log(cap, t0, t1, wall, mono):
src = r.get("sources", {})
s = src.get("mmap1")
if r.get("mode") != "practice" or not s or "off" not in s:
continue
press = r["time"] - r["held_s"] - wall + mono
if not t0 + BEFORE[0] < press < t1:
continue
k = np.where((idx[:, 3] > press - BEFORE[0]) & (idx[:, 3] < press - BEFORE[1]))[0]
eye = {c: eye_features([frames[i] for i in k if idx[i, 2] == c]) for c in (0, 1)}
# The truth: a source's gaze at the press plus the angle from it to the release
# point. The angle is converted with a local linear fit of the screen, so the
# closer the source, the better: ours when the live tracker was running.
own = src.get("own") if src.get("own", {}).get("off") else None
base = own or s
truth = (base["hy"] + base["off"][0], base["hp"] + base["off"][1])
clicks.append(dict(t=press, truth=truth, truth_from="own" if own else "mmap1",
steam=(s["hy"], s["hp"]),
steam_err=float(np.hypot(s["hy"] - truth[0], s["hp"] - truth[1])),
live_err=float(np.hypot(*own["off"])) if own else None,
press_err=r.get("press_err_deg"), press_source=r.get("source"), eye=eye))
# The slip track: pupil and pair midpoint on a sample of frames through the session.
track = {}
for c in (0, 1):
rows = []
for i in np.where(idx[:, 2] == c)[0][::TRACK_EVERY]:
p = eyes_pupil.find_pupil(frames[i])
pair = p and eyes_pupil.glint_pair(p)
if pair:
rows.append((idx[i, 3], p["x"], p["y"],
(pair[0][0] + pair[1][0]) / 2, (pair[0][1] + pair[1][1]) / 2))
track[c] = np.array(rows).reshape(-1, 5)
f = dict(version=VERSION, clicks=clicks, track=track, span=(t0, t1))
cache.write_bytes(pickle.dumps(f))
return f
# --- Fitting --------------------------------------------------------------------------
class Model:
"""A calibration fitted on some clicks, plus SteamVR's quadratic fit on the same."""
def __init__(self, clicks):
self.cal = eyes_model.Calibration.fit(clicks)
self.steam = eyes_model.Quad([k["steam"] for k in clicks], [k["truth"] for k in clicks])
def slip(self, c, track, t):
"""Median slip (pixels) over the track in the SLIP_WINDOW seconds before t."""
tr = track[c]
if len(tr) == 0:
return None
return self.cal.slip(c, tr[(tr[:, 0] < t) & (tr[:, 0] > t - eyes_model.SLIP_WINDOW)])
def predict(self, method, k, track):
"""Gaze for one click by a method, combining the eyes it has; None if neither."""
cal, out = self.cal, {}
for c in (0, 1):
e = k["eye"][c]
if e is None or not cal.has("pupil", c):
continue
if method == "pupil":
out[c] = cal.gaze(c, *e["pupil"])
elif method == "glint" and cal.has("glint", c) and e["mid"] is not None:
out[c] = cal.fits["glint", c].one(*(e["pupil"] - e["mid"]))
elif method == "slip":
s = self.slip(c, track, k["t"])
if s is not None:
out[c] = cal.gaze(c, *e["pupil"], slip=s)
return cal.combine(out)
# --- Scoring --------------------------------------------------------------------------
METHODS = ("pupil", "glint", "slip")
def report(name, err, total):
err = np.asarray([e for e in err if e is not None])
if len(err) == 0:
print(f" {name:36s} no clicks")
return
print(f" {name:36s} {len(err):3d}/{total} median {np.median(err):5.2f} "
f"mean {err.mean():5.2f} 90% {np.percentile(err, 90):5.2f} deg")
def score(train, test, track, same):
"""Errors per click for each method; leave-one-out when train and test are the same.
"clicks" goes through the test clicks in order, as live: each is predicted with the
shift the earlier ones measured (eyes_model.Shift), then teaches it."""
errs = {m: [] for m in METHODS + ("clicks", "steam")}
model = None if same else Model(train)
shifts = {c: eyes_model.Shift() for c in (0, 1)}
for i, k in enumerate(test):
mdl = Model(train[:i] + train[i + 1:]) if same else model
for m in METHODS:
g = mdl.predict(m, k, track)
errs[m].append(None if g is None else float(np.hypot(*(g - k["truth"]))))
errs["steam"].append(float(np.hypot(*(mdl.steam(k["steam"])[0] - k["truth"]))))
out = {}
for c in (0, 1):
e = k["eye"][c]
if e is None or not mdl.cal.has("pupil", c):
continue
tr = track[c]
# The glint estimate as live would have had it over the last second, for the hold.
for back in (eyes_model.JUMP_HOLD, eyes_model.JUMP_HOLD / 2, 0.0):
te = k["t"] - back
g = mdl.cal.slip(c, tr[(tr[:, 0] < te) & (tr[:, 0] > te - eyes_model.JUMP_WINDOW)]) if len(tr) else None
shifts[c].glint(g, te)
out[c] = mdl.cal.gaze(c, *e["pupil"], slip=shifts[c].value)
shifts[c].click(mdl.cal.click_shift(c, e["pupil"], k["truth"]), g)
errs["clicks"].append(float(np.hypot(*(mdl.cal.combine(out) - k["truth"]))) if out else None)
return errs
def summary(f, label):
cl = f["clicks"]
T = np.array([k["truth"] for k in cl])
print(f"{label}: {len(cl)} clicks over {f['span'][1] - f['span'][0]:.0f} s, gaze yaw "
f"{T[:, 0].min():.0f}..{T[:, 0].max():.0f}, pitch {T[:, 1].min():.0f}..{T[:, 1].max():.0f}")
for c in (0, 1):
n = sum(k["eye"][c] is not None for k in cl)
g = sum(k["eye"][c] is not None and k["eye"][c]["mid"] is not None for k in cl)
print(f" {EYES[c]} eye: pupil before {n} clicks, glint pair before {g}; "
f"slip track {len(f['track'][c])} frames with the pair")
CALIBRATION = Path.home() / ".local/state/frametop/gaze/eyes/calibration.json"
def main(args):
if args and args[0] == "--clicks":
for cap in map(capture, args[1:]):
wall, mono = map(float, (cap / "clocks.txt").read_text().split())
lines = (cap / "index.txt").read_text().splitlines()
ts = [float(l.split()[3]) for l in (lines[0], lines[-1])]
print(f"{cap}: {len(practice_log(cap, *ts, wall, mono))} practice records")
return
if args and args[0] == "--save":
cap = capture(args[1])
f = features(cap)
cal = eyes_model.Calibration.fit(f["clicks"], {"capture": cap.name, "clicks": len(f["clicks"]),
"made": time.strftime("%Y-%m-%d %H:%M")})
cal.save(CALIBRATION)
print(f"saved {CALIBRATION}: {sorted(f'{n} {EYES[e]}' for n, e in cal.fits)}")
return
ca = capture(args[0])
a = features(ca)
summary(a, ca.name)
if len(args) > 1:
cb = capture(args[1])
b = features(cb)
summary(b, cb.name)
print(f"\nFit on {ca.name}, tested on {cb.name}:")
test, track, errs = b["clicks"], b["track"], score(a["clicks"], b["clicks"], b["track"], False)
else:
print("\nLeave-one-out within the session:")
test, track, errs = a["clicks"], a["track"], score(a["clicks"], a["clicks"], a["track"], True)
n = len(test)
report("SteamVR raw", [k["steam_err"] for k in test], n)
steam_press = [k["press_err"] for k in test if k["press_source"] != "own"]
report("SteamVR + probe's live correction", steam_press, len(steam_press))
live = [k["live_err"] for k in test if k["live_err"] is not None]
if live:
report("ours live (ft-eyes, as the probe saw it)", live, n)
own_press = [k["press_err"] for k in test if k["press_source"] == "own"]
report("ours live + probe's live correction", own_press, len(own_press))
report("SteamVR + quadratic fit", errs["steam"], n)
for m in METHODS:
report(f"ours, {m}", errs[m], n)
report("ours, clicks (shift from earlier clicks)", errs["clicks"], n)
# Like for like: the clicks every method scored.
common = [i for i in range(n) if all(errs[m][i] is not None for m in METHODS)]
print(f"\nSame {len(common)} clicks for every method:")
report("SteamVR + quadratic fit", [errs["steam"][i] for i in common], len(common))
for m in METHODS:
report(f"ours, {m}", [errs[m][i] for i in common], len(common))
if len(args) == 1:
for c in (0, 1):
tr = track[c]
if len(tr) < 20:
continue
mdl = Model(a["clicks"])
ss = [mdl.slip(c, track, t) for t in np.linspace(tr[0, 0] + eyes_model.SLIP_WINDOW, tr[-1, 0], 6)]
print(f" {EYES[c]} eye slip estimate through the session (px): "
+ " ".join(f"({s[0]:+.1f},{s[1]:+.1f})" for s in ss if s is not None))
if __name__ == "__main__":
main(sys.argv[1:] or ["practice1"])
+30
View File
@@ -0,0 +1,30 @@
#!/usr/bin/env bash
# ft-eyes-session: record the eye cameras and SteamVR's gaze together, for SECONDS (default 10),
# as a new recording NAME in ~/.local/share/frametop/eyes/captures (FT_EYES_CAPTURES).
# Usage: gaze/tracker/lab/ft-eyes-session NAME [SECONDS]
#
# frames.raw, index.txt, clocks.txt every eye-camera frame (ft-eyes-record, from the shared
# frames of frametop-eyegrab.service; no root)
# gaze.jsonl ft-gaze's samples (gaze/build/ft-gaze)
# Both are timed on CLOCK_MONOTONIC_RAW ("t" in gaze.jsonl, the last column of index.txt).
set -euo pipefail
lab=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
repo=$(cd "$lab/../../.." && pwd)
name=${1:?usage: ft-eyes-session NAME [SECONDS]}
secs=${2:-10}
out=${FT_EYES_CAPTURES:-$HOME/.local/share/frametop/eyes/captures}/$name
[ -e "$out" ] && { echo "$out exists" >&2; exit 1; }
# ft-gaze runs in the dev container; start it first, it takes a moment to connect. It quits
# when its stdin closes (--watch-stdin): killing distrobox doesn't reach it in the container.
# The recorder makes the folder; ft-gaze's output waits for it.
tmp=$(mktemp -d)
sleep $((secs + 3)) | "$HOME/.local/bin/distrobox" enter dev -- "$repo/gaze/build/ft-gaze" \
--watch-stdin > "$tmp/gaze.jsonl" 2> "$tmp/gaze.log" &
gaze=$!
sleep 2
python3 "$lab/ft-eyes-record" "$name" "$secs"
wait $gaze || true
mv "$tmp/gaze.jsonl" "$tmp/gaze.log" "$out/"
rmdir "$tmp"
echo "$(wc -l < "$out/index.txt") frames, $(wc -l < "$out/gaze.jsonl") gaze samples in $out"
+12
View File
@@ -0,0 +1,12 @@
#!/usr/bin/env bash
# Python with numpy and OpenCV for the lab tools: gaze/tracker/build/venv (gaze/tracker/build.sh
# makes it in the dev container on the Frame; frame-job's SETUP makes it on the PC). On the
# Frame's host it runs in the dev container, where it was made.
# Usage: lab/py lab/TOOL [args] (from gaze/tracker)
here=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
py=$here/build/venv/bin/python
if grep -qx 'ID=steamos' /etc/os-release 2>/dev/null; then
exec "$HOME/.local/bin/distrobox" enter dev -- "$py" "$@"
fi
[ -x "$py" ] || { echo "no $py: run gaze/tracker/build.sh (or a frame-job job, whose setup makes it)" >&2; exit 1; }
exec "$py" "$@"
+5
View File
@@ -0,0 +1,5 @@
# ft-eyes and the lab tools (gaze/tracker/build.sh puts them in build/venv, in the dev
# container; frame-job's SETUP does the same on the PC). Fedora's python3-opencv would pull in
# VTK, GDAL, and over a gigabyte of map data; these wheels are about 165 MB.
numpy==2.5.3
opencv-python-headless==5.0.0.93
+3
View File
@@ -0,0 +1,3 @@
# Model sources that tools/convert_models.py downloads; the converted ncnn models are kept
models/onnx/
models/*.task
+49
View File
@@ -0,0 +1,49 @@
# Hand tracking, built into build/ (hands/build.sh runs this in the dev container):
# make ft-camd (camd/: runs on the host, so linked statically) and ft-hands (track/)
# make tools ft-handreplay and ft-ringplay, for recordings
# The first build fetches ncnn (NCNN_TAG) and builds it into build/ncnn, which takes a few
# minutes. NCNN=DIR uses an ncnn install already built instead.
NCNN_TAG = 20260526
NCNN ?= build/ncnn/install
CFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter
CXXFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter -Wno-psabi
CXXFLAGS += -std=c++17 -fopenmp -I$(NCNN)/include/ncnn
LDLIBS = $(NCNN)/lib/libncnn.a -ljsoncpp -fopenmp -lpthread
CAMD = camd/camd.c camd/tp.c camd/xrcams.c
TRACK = track/calib.cpp track/nets.cpp track/tracker.cpp track/io.cpp track/record.cpp track/pinch.cpp
HDR = $(wildcard track/*.h) camd/fhring.h include/fh_hands.h include/fh_gestures.h
all: build/ft-camd build/ft-hands
tools: build/ft-handreplay build/ft-ringplay
build/ft-camd: $(CAMD) camd/tp.h camd/xrcams.h camd/fhring.h
@mkdir -p build
$(CC) $(CFLAGS) -static -o $@ $(CAMD) -lm
build/ft-hands: track/main.cpp $(TRACK) $(HDR) $(NCNN)/lib/libncnn.a
@mkdir -p build
$(CXX) $(CXXFLAGS) -o $@ track/main.cpp $(TRACK) $(LDLIBS)
build/ft-handreplay: track/replay.cpp $(TRACK) $(HDR) $(NCNN)/lib/libncnn.a
@mkdir -p build
$(CXX) $(CXXFLAGS) -o $@ track/replay.cpp $(TRACK) $(LDLIBS)
build/ft-ringplay: track/ringplay.cpp track/record.h camd/fhring.h
@mkdir -p build
$(CXX) $(CXXFLAGS) -o $@ track/ringplay.cpp
# ncnn as frame-hands built it (the models were converted and quantized for it), minus its tools
build/ncnn/install/lib/libncnn.a:
rm -rf build/ncnn && mkdir -p build/ncnn
git clone -q --depth 1 --branch $(NCNN_TAG) -c advice.detachedHead=false https://github.com/Tencent/ncnn.git build/ncnn/src
cmake -S build/ncnn/src -B build/ncnn/build -G Ninja -Wno-dev -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_INSTALL_PREFIX=$(CURDIR)/build/ncnn/install -DCMAKE_INSTALL_LIBDIR=lib -DNCNN_VULKAN=OFF \
-DNCNN_OPENMP=ON -DNCNN_INT8=ON -DNCNN_SIMPLEOCV=ON -DNCNN_BUILD_TOOLS=OFF -DNCNN_BUILD_EXAMPLES=OFF \
-DNCNN_BUILD_BENCHMARK=OFF -DNCNN_BUILD_TESTS=OFF -DNCNN_PYTHON=OFF > build/ncnn/cmake.log
cmake --build build/ncnn/build --target install > build/ncnn/build.log
clean:
rm -f build/ft-camd build/ft-hands build/ft-handreplay build/ft-ringplay
.PHONY: all tools clean
+169
View File
@@ -0,0 +1,169 @@
# Hands (experimental)
Hand tracking from the headset's own cameras. It serves two things in Frametop:
- **Hand cutouts:** where your hand is between an eye and a screen, that eye sees the room through the screen (ft-screens, `screens/handcut.cpp`), so your hands show over the screens the way they do on a Vision Pro.
- **Pinches:** look at something and pinch to click it, pinch and move to drag, with the eye tracker doing the looking (`gaze/`). The tracker publishes the pinches. The pointer helper doesn't read them yet.
Two programs, each a user service that starts and stops with SteamVR:
- `ft-camd` (`camd/`, C) borrows XRService's camera buffers and publishes the four IR tracking cameras' frames to a shared-memory ring. It runs on the host.
- `ft-hands` (`track/`, C++) finds hands in those frames with MediaPipe's palm and landmark models on ncnn, triangulates them, and publishes them. It runs in the dev container.
```
hands/run.sh install # build, give ft-camd its capabilities (sudo, once per build), enable
hands/run.sh status # the services, and ft-hands' last status lines
hands/run.sh log [lines]
hands/run.sh restart # after changing a setting
hands/run.sh caps # after rebuilding ft-camd (a rebuild clears its capabilities)
hands/run.sh uninstall
```
Settings in `~/.config/frametop.conf` (`FT_<name>` in the environment overrides them):
- `HANDS_SWAP_SIDES=1`: the two side cameras' names are swapped (see ft-camd below). Check with `tools/check_sides.py --ring`.
- `HANDS_CPUS=5,6,7`: the CPUs the model threads run on (below).
Files, all in `/run/user/UID/frametop-hands/` (private to the user; not `/run/user/UID/frametop/`, which the desktop session deletes whenever it starts):
| File | Written by | Layout | Read by |
| --- | --- | --- | --- |
| `cam-ring` | ft-camd | `camd/fhring.h` | ft-hands, `tools/ring.py` |
| `hands` | ft-hands | `include/fh_hands.h` | ft-screens (`screens/handcut.cpp`) |
| `gestures` | ft-hands | `include/fh_gestures.h` | `tools/watch_gestures.py`; the pointer helper, later |
The source keeps the `fh_` names and magic strings of frame-hands, where this was developed (`~/Desktop/Projects/frame-hands` on the developer's Frame, which keeps the recordings, probes and Python prototype). So its recordings and tools still work.
## ft-camd
XRService owns the headset cameras. ft-camd borrows its DMA-BUFs read-only with `pidfd_getfd`, the same way FrameEyeCameraFeed does. It never touches XRService's V4L2 descriptors. `discovery` in `camd/xrcams.c` is adapted from FrameEyeCameraFeed (MIT, see `camd/LICENSE.FrameEyeCameraFeed`).
Polling buffers for changes can catch a frame while the camera is still writing it. Instead, ft-camd listens to the `v4l2:v4l2_dqbuf` tracepoint, which fires when XRService takes a buffer. It gives the buffer index, the sequence number and the capture timestamp. ft-camd learns which DMA-BUF holds each V4L2 index by watching which buffer changes at each dequeue:
- Right after XRService allocates its buffers, the mapping is allocation order.
- After XRService restarts streaming, the order is shuffled, and the mapping is learned index by index.
- The two upper cameras share one run of buffers. For them, only allocation order can tell the cameras apart.
- It also re-maps an index on the fly when its buffer holds no new frame.
**Privileges.** Setting up needs three things. `pidfd_getfd` on XRService needs `CAP_SYS_PTRACE`, because the Frame has `ptrace_scope=1`. The system-wide tracepoint needs `CAP_PERFMON`, because `perf_event_paranoid` is 2. Its format files are root-only, which needs `CAP_DAC_READ_SEARCH`. `hands/run.sh install` gives the binary those capabilities with `sudo setcap`. ft-camd drops them all once it has set up, before it reads a frame, and then runs as you. XRService runs as you too. It also runs under `sudo`, for trying it by hand, and then drops to the user who ran sudo. It reads nothing from the ring's readers.
The ring is mode 0600, in a folder only you can write. Frame handling:
- Only complete, bright frames are published. The cameras alternate a normal exposure with a near-black one, so each camera gets 30 of its 60 fps.
- A copy torn by the camera overwriting the buffer is dropped.
- Each copy takes about 0.1 ms, and a cache sync about 0.15 ms.
Options:
- `--with-dark`: also publish the near-black frames, as extra ring cameras flagged `FH_CAM_DARK`. They show only light sources, so they're no use for hands.
- `--with-color` (the service uses it): also publish the two Arcturus colour cameras, flagged `FH_CAM_COLOR`. Each is the luma of the 10-bit frame's valid 1972x2464 (the top 8 bits), at half size (`--color-scale 2`: 986x1232). They run at `--color-idle` (2 fps), enough for ft-hands to tell how bright it is, until a reader asks for more in `/run/user/UID/frametop-hands/color-fps` (ft-hands writes 30 while it tracks or records with them), up to `--color-fps` (30; the cameras run at 60). `HANDS_CAMERAS=mono` leaves them out. Frames that carry the module's warped half-size copy are dropped. Their `capture_ns` is on the colour module's clock (2.2 s off the mono cameras' on 2026-09-29), so line them up with the mono cameras by `dqbuf_ns`. Each frame costs about 0.65 ms of cache sync and 1.1 ms of decoding, so both cameras at 30 fps take about 11% of a core.
- Each mono camera's latest near-black frame's mean goes in the ring (`dark_mean`): a short fixed exposure, so it follows the room's IR light, sunlight above all.
- The ring holds 8 cameras: 4 mono, plus 4 dark twins or 2 colour cameras.
- Colour isn't reliable yet. In the lit-room test of 2026-09-30, the colour cameras kept losing their buffer mapping while the headset was worn: 30 frames in a row looked unchanged, the camera relearned, and after 5 relearns ft-camd exited. Each relearn probed all 32 colour buffers, a whole-buffer cache sync each, which also made the mono cameras miss frames. Runs with the headset idle had none of this. So the passthrough compositor may be writing into the colour buffers while Room View shows. Since then a colour camera never takes the mono ones down: it probes at most 4 buffers a frame, and one that goes stale twice in a row is paused (10 s, doubling up to 160 s) and learned again, without ft-camd exiting. Whether a frame is new is judged on the luma rows only: the chroma after them hardly changes in a lit room. `FT_CAMD_DEBUG=1` prints, at each colour stale frame, how many sampled words changed in every candidate buffer.
- `--sensor S`: only the mono cameras whose sensor name contains S.
- `--status S`: a status line every S seconds (0: never).
It exits when XRService exits, or when a camera's buffers keep going stale, which means XRService has reallocated them. The service starts it again, and it attaches to the new buffers.
**Which camera is which:** video9 is `slam_left`, video13 is `slam_right`, video6 is `upper_left` and video7 is `upper_right`. This was checked by rendering the same view from each camera with the factory calibration. But ft-camd tells the side cameras' buffers apart only by XRService's allocation order, and after some XRService restarts it gets them backwards. Then every hand is seen by one camera only, at the wrong depth, and the hand holes land beside the hands. With the headset on, looking at a room with some texture, `tools/check_sides.py --ring` says whether the names are right (exit 0), swapped (exit 3), or it can't tell (exit 2). When they're swapped, set `HANDS_SWAP_SIDES=1`. The colour cameras are video3 (`arcimx616 0-0010`) and video0 (`0-001a`); which of them is `passthrough_left` in the module's calibration is for `tools/check_color.py` to settle, on a recording with texture in view.
## ft-hands
```
hands/build/ft-hands # status every 5 s; Ctrl+C to stop
hands/build/ft-hands --int8 # the 8-bit models (models/ncnn/*-int8.ncnn.*)
```
Run it in the dev container (`distrobox enter dev -- ...`). It reads the factory calibration from `/persist` (`/run/host/persist` in the container).
Options:
- `--threads N`: model threads, pinned to the `--cpus` list. Default 3.
- `--cpus LIST`: CPUs for the model threads and the main loop. Default `5,6,7` (`HANDS_CPUS`). SteamOS starts user processes on CPUs 0-4, and XRService's head tracking runs on 2-3. With the headset on, a step took 8.4 ms on 5-7 against 13.2 ms on 2-4, and SteamVR's frame timing didn't change (2026-09-29, three rounds of the same replayed frames).
- `--contrast MODE` or `PALM/HAND`: how crops are equalized before the models see them: `clahe[:CLIP]`, `none`, or `stretch` (1st-99th percentile). Default `clahe:2/none`. In the dim recording, CLAHE let the palm search find about 10% more hands, but it made the landmarks jitter more (published median 6.9 mm, against 6.0 mm with plain landmark crops).
- `--swap-sides`: swap the two side cameras (`HANDS_SWAP_SIDES`, see ft-camd).
- `--seconds N`: stop after N seconds.
- `--status S`: how often to print status, in seconds.
- `--models DIR`: where the models are.
- `--nice N`: niceness. Default 5, so the VR stack wins contested CPUs.
- `--no-publish`: don't write the hands and gestures files.
- `--record DIR`, `--record-for S`: save every frame set for S seconds (default 120) to `DIR/sets.bin`. That's about 80 MB/s. Sending the tracker SIGUSR1 (`pkill -USR1 -x ft-hands`) starts a recording in `~/.local/share/frametop/hands/rec-<time>` without a restart. Recordings are images of your hands and room: they stay on the headset unless you move them.
- `--record-only`: record without tracking or publishing, so it can run beside the live tracker. Give it `--record DIR`, since SIGUSR1 would reach both trackers. With `ft-camd --with-dark`, recordings also hold each camera's newest dark frame as `<name>_dk`, which doubles the rate. With `--with-color`, each colour camera's newest frame is saved with every set, as `color_video<N>`, which adds about 70 MB/s. Run the recorder at normal I/O priority: idle I/O priority stalled a 165 MB/s recording.
- `--keep-presence P`: the landmark presence a tracked view needs to stay tracked. New views always need 0.5. Default 0.5. Lowering it to 0.2 barely helped in the bright recording, because lost hands drop to near-zero presence.
- `--ring PATH`: read frames from another ring, such as `ft-ringplay`'s.
- `--cams auto|mono|color|all` (`HANDS_CAMERAS`, default `auto`): which cameras to track with. The mono IR cameras light the hands themselves and track well in dim rooms, but in bright light they expose for the room and the hands come out dark. The colour pair is the other way round. `auto` goes by the colour frames' mean brightness: at `--bright-on` (`HANDS_BRIGHT_ON`, 40) or over for 2 s it tracks with `--bright` (`HANDS_BRIGHT`: `all`, every camera, the default, or `color`), and under `--bright-off` (`HANDS_BRIGHT_OFF`, 25) for 2 s with the mono cameras again. A dim evening room read 9. The switch is logged (`cameras: mono -> all (...)`), and the status line gives the colour level, the mono cameras' ambient IR, and how many steps had colour frames. Colour frames arrive on their own schedule, so a step holds the mono set, the colour pair, or both, and views wait in their camera for its next frame.
- `--color-left NODE` (`HANDS_COLOR_LEFT`, `color_video0`) and `--color-crop subtract|none` (`HANDS_COLOR_CROP`, `subtract`): how the colour module's calibration maps onto the images. Not settled yet: `tools/check_color.py` on a recording with a lit, textured view tells.
- `--grip-begin R`, `--grip-end R`: the grip detector (below).
**Gestures** (`/run/user/UID/frametop-hands/gestures`, `include/fh_gestures.h`), for the pointer helper:
- A pinch: the thumb and index tips within 2 cm, ending past 3.5 cm. Not begun with the palm facing down (`--pinch-palm-down`, 0.6), which is how typing looks.
- A grip, a closed hand: every finger's tip nearer the wrist than 1.2 times its knuckle is (from the model's 3D hand, so hand size doesn't matter), ending when they open past 1.45 on average. It begins only on a hand seen open within the last second (closing it is the gesture), with the palm at most 35 degrees below straight ahead and at least 15 cm in front of the eyes. A grip ends a pinch on the same hand, as lost. In the 2026-09-30 lit recording (no deliberate fists), the checks cut false grips from 14 to 6, all with the hands on the desk while looking down at it; the pointer helper ignores grips that begin more than 30 cm below the eyes, which it can tell and ft-hands can't.
- `tools/watch_gestures.py --distance` shows both live; `ft-handreplay --timeline` logs them and each hand's finger curl.
The status line also says how often a hand was on each side (by where the wrist is), and why views and hands came and went: views lost (the landmark model stopped seeing the hand), handoff misses (a crop projected from the hand's 3D position found nothing), duplicates, splits (two views disagreed in 3D), and hands created, merged and forgotten.
### Scheduling
- Each hand is tracked in its best two cameras, the way MediaPipe tracks: the landmark model runs on a crop placed from the previous landmarks, with no palm detection.
- A hand seen in too few cameras is projected into the others through the calibration. Where it lands well inside a camera, that camera gets a crop to try. This is how a hand raised out of the side cameras reaches the upper ones.
- The palm detector runs only while fewer than two hands are tracked, at most 5 times a second, on a few zoomed tiles per search. Tiles are picked in proportion to how likely hands are there. Each tile is turned so the expected shoulder-to-hand direction points up.
- Frame sets are processed at 30 Hz while a hand moves faster than 0.25 m/s (or a pinch is down or closing), at 15 Hz otherwise, and at 5 Hz while no hand is in view.
### 3D
- **Two or more views:** each landmark is triangulated from the camera rays, weighted by the model's presence score. The median ray distance is reported as the residual.
- **Pairing views across cameras.** The side cameras sit side by side, so two hands next to each other at the same height fall on the same epipolar lines, and rays to two different hands can nearly meet close to the cameras. That made phantom hands 12-15 cm in front of the eyes, which tore holes through the screens. Each step now scores every way of pairing the views in two cameras and keeps the best. A pair scores well when its rays meet, when each view's apparent size matches the triangulated distance, and when the model calls both the same hand. The size check uses a fixed prior: with the model's average hand, clean pairs measure 0.71-1.51 times the one-view distance, and mismatched pairs mostly far less.
- **One view:** depth comes from the model's metric world landmarks, their spread across the palm against the angle it covers in the image, scaled by the user's hand size (learned while two views are available). That distance is off by 10-30% and wanders about 10% between frames, so a hand that drops to one camera keeps its last distance and drifts toward the one-view guess by 10% a frame.
- **Smoothing.** The published landmarks go through a One Euro filter: it smooths hard while the hand is still (tracking noise is several mm per frame) and hardly at all while it moves fast. The palm speed that sets the update rate is the filtered one; the raw speed read about 0.25 m/s from noise alone.
- **Capsules.** Forearms follow the hand's own axis, and nothing within 12 cm in front of the eyes is published.
How good the depth is, measured from recordings (2026-09-30, `--depth` below): the two lower cameras see the hands about 77% of the time, a lower and an upper camera 7-12%, and one camera 12-15%. Depth is the noisy direction. With the lower pair, it jitters 4-6 times as much as sideways position (published: 3-7 mm against 1-2 mm). The one-camera guess is a median 2-6 cm off. When a camera drops out, drifting 10% a frame toward that guess is worse than keeping the last distance (after 0.5 s a median 23-30 mm off, against 11-12 mm).
## Pinch
ft-hands detects a pinch per hand (`track/pinch.h`) and publishes it to the gestures file. The layout, and how to read it without missing quick taps, is in `include/fh_gestures.h`.
- A pinch begins when the thumb and index tips come within `--pinch-begin` (default 0.020 m). It ends when they open past `--pinch-end` (0.035 m) for 2 processed frames in a row, or when the hand stays lost for 0.25 s (flagged lost).
- The distance comes from MediaPipe's world landmarks: the model's own 3D hand pose, averaged over the hand's views, at the user's hand size. `--pinch-triangulated` uses the triangulated tips instead. On two recordings without deliberate pinches, the world landmarks came under 2 cm in 0.2-1% of frames, against 3.3-4.5% for the triangulated tips. In the dim recording, typing still gave 2 pinches a minute before the palm check below.
- No pinch begins while the palm faces down (`--pinch-palm-down MAX`: the palm normal's share of the head's up axis, default 0.6; 1 turns it off), and a close held back that way has to open again before a pinch can begin. Typing curls the thumb onto the index. In the lit recording of 2026-09-30, typing on a keyboard in the lap began 23 pinches in about 2 minutes, all with the palm facing down (0.69-1.00), while the 26 deliberate ones read 0.00-0.50. The limit held back every typing pinch and none of the deliberate ones. Looking down tilts the head frame, which lowers the reading for a hand on a keyboard, so the consumer's gaze check stays the other guard.
- A hand a pinch is down on stays with that side until the pinch ends. The left/right call is a running average of the model's, and when it flipped mid-pinch, the other side took the same hand and both sides pinched at once.
- The pinch point is midway between the thumb and index tips. A drag is the pinch point now, minus where it was when the pinch began, both turned into the room with the HMD pose at their capture times.
- `tools/watch_gestures.py` prints begins, ends and drag offsets live, and `--distance` prints each hand's distance.
The pointer helper is the natural consumer. Its gaze mode already treats a press as "stop where the gaze put it, drag onto the target, click on release", and "hold still for half a second, then move" as a drag. A pinch begin would be the press, the end the release, and the pinch point's movement the drag.
## Recordings
`hands/build.sh --tools` also builds the offline tools.
`ft-handreplay DIR` runs a recording through the tracker with the live scheduling and reports how well it kept the hands: hands per set, left and right coverage, track lengths, pinches, jitter, and the same reasons as the status line.
```
hands/build/ft-handreplay ~/.local/share/frametop/hands/rec-20260929-120000 --cost --oracle 10 --timeline /tmp/tl.txt
```
- `--cost`: instead of timing the steps, charge each round of model calls what it typically costs live (10 ms landmarks, 18 ms palms), so results repeat exactly.
- `--oracle N`: every N-th set, also search every tile of every camera, and report how often the tracker had the hands that full search could find.
- `--slow F`: live, the tracker skips sets that arrive while it's busy. Replay counts each step's time times F as busy (default 1; the headset is busier live).
- `--timeline FILE`: a line per processed set and hand, with pinch events and distances.
- `--cams mono|color|all`: which cameras to track with (default `mono`). `color` tracks with the Arcturus pair alone, for comparing it with the IR cameras on the same recording. It needs a recording made with `ft-camd --with-color`. `--color-left NODE` (`color_video0` or `color_video3`) and `--color-crop subtract|none` say how the module's calibration maps onto the images; `tools/check_color.py` finds out.
- `--depth FILE`: a line per hand per processed set for `tools/depth_report.py`, which measures the depth without ground truth: how the hands were seen, the noise along the line of sight against across it, each camera's one-view distance against the triangulated one, and what a camera dropping out would do.
- The pinch, contrast and presence options are ft-hands'.
`ft-ringplay DIR --ring PATH [--from S] [--to S] [--loop]` publishes a recording into a ring file in real time, as ft-camd would, so `ft-hands --ring PATH --no-publish` runs the same frames run after run. It needs no privileges, and it skips the dark frames.
## Tools
Python, with NumPy and OpenCV (in the dev container: `python3-numpy`, `python3-opencv`, which `setup/dev-container.sh` installs). Off the Frame, `FRAME_JOB_DEVICE_ROOT` can point at a folder with copies of the headset's calibration files.
- `tools/check_sides.py --ring` (or a recording): are the side cameras named right?
- `tools/check_color.py REC`: how the colour module's calibration maps onto its images.
- `tools/show_set.py REC`: a recording's frame sets as images.
- `tools/watch_gestures.py [--distance]`: pinches, live.
- `tools/depth_report.py DEPTH`: the depth measures above.
- `tools/convert_models.py`: how `models/ncnn` was made from the OpenCV Zoo ONNX ports of MediaPipe's models (see `models/NOTICE`).
## Build
`hands/build.sh` builds in the dev container on the Frame, into `hands/build/`, with `hands/Makefile`. The first build fetches ncnn at a pinned tag and builds it into `hands/build/ncnn`, which takes a few minutes; `NCNN=DIR` points at an ncnn install already built instead. ft-camd is linked statically, because it runs on the host, which has an older glibc than the container.
Executable
+11
View File
@@ -0,0 +1,11 @@
#!/usr/bin/env bash
# Build hand tracking in the dev container on the Frame, into hands/build/: ft-camd and ft-hands,
# and with --tools also ft-handreplay and ft-ringplay. The first build fetches ncnn and builds
# it (a few minutes); NCNN=DIR, an ncnn install already on the Frame, skips that.
# A rebuilt ft-camd has lost its capabilities: hands/run.sh install sets them again.
set -euo pipefail
root=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
targets=all
[ "${1:-}" = --tools ] && targets="all tools"
"$root/scripts/sync.sh" >/dev/null
exec "$root/scripts/frame.sh" -C hands "make -s ${NCNN:+NCNN=$NCNN} $targets && echo built \$(ls build/ft-* | tr '\n' ' ')"
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Curtis English
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+1262
View File
File diff suppressed because it is too large. Load diff
+91
View File
@@ -0,0 +1,91 @@
/*
* fhring - the shared-memory frame ring ft-camd writes and trackers read.
*
* One file, /run/user/UID/frametop-hands/cam-ring (FH_RING_NAME in the user's runtime
* folder; the folder is private to the user), holds a header, then for each camera
* a few slots, each a slot header followed by the image rows packed tightly
* (stride == width for 8-bit mono). Only complete, bright frames are published.
*
* Writer, for frame n of a camera: slot = n % nslots
* slot.seq = 2n+1; write slot fields and pixels; slot.seq = 2n+2; cam.latest = n
* Reader:
* n = cam.latest; read slot.seq, expect 2n+2; copy; re-read slot.seq; if it
* changed the copy is torn, retry with the new latest.
*
* All multi-byte fields are little-endian; offsets are fixed so Python can read
* them with struct (tools/ring.py mirrors this file).
*/
#pragma once
#include <assert.h>
#include <stdint.h>
#define FH_RING_MAGIC "FHRING01"
#define FH_RING_VERSION 1
#define FH_RING_MAX_CAMS 8
#define FH_RING_SLOTS 4
#define FH_RING_NAME "frametop-hands/cam-ring" /* in /run/user/UID */
enum {
FH_FMT_GREY8 = 0,
};
enum {
FH_CAM_DARK = 1u << 0, /* the near-black exposures between this node's */
/* normal frames (ft-camd --with-dark) */
FH_CAM_COLOR = 1u << 1, /* an Arcturus color camera's luma, downscaled */
/* (ft-camd --with-color). Not synced with the */
/* mono cameras, and capture_ns is on its own */
/* clock: line it up with them by dqbuf_ns */
};
typedef struct {
char sensor[32]; /* media entity, e.g. "og01a1b 4-0060" */
char name[32]; /* calibration name if known, else sensor slug */
int32_t node; /* N of /dev/videoN */
uint32_t format; /* FH_FMT_* */
uint32_t width;
uint32_t height;
uint32_t stride; /* bytes per row in the ring */
uint32_t nslots;
uint64_t slot_offset; /* file offset of slot 0 */
uint64_t slot_bytes; /* slot header + image, 64-byte aligned */
volatile uint64_t latest; /* newest published frame number, 0 = none yet */
uint64_t published; /* frames published */
uint64_t dropped; /* dark, stale or torn frames not published */
uint32_t flags; /* FH_CAM_* */
float dark_mean; /* mono: mean luma of its latest near-black */
/* frame (a short fixed exposure, so it follows */
/* the room's IR light, sunlight above all); */
/* 0 before the first */
uint8_t reserved[24];
} fh_ring_cam_t; /* 160 bytes */
typedef struct {
volatile uint64_t seq; /* 2n+1 while frame n is written, 2n+2 when done */
uint64_t frame; /* n */
uint64_t capture_ns; /* V4L2 timestamp (camera clock) */
uint64_t dqbuf_ns; /* CLOCK_MONOTONIC when XRService dequeued it */
uint64_t publish_ns; /* CLOCK_MONOTONIC when the copy finished */
uint32_t v4l2_seq; /* V4L2 sequence number */
float mean; /* mean luma on a sparse grid */
uint8_t reserved[16];
} fh_ring_slot_t; /* 64 bytes, image follows */
typedef struct {
char magic[8]; /* FH_RING_MAGIC */
uint32_t version;
uint32_t header_bytes; /* sizeof(fh_ring_hdr_t) */
uint32_t ncams;
uint32_t reserved0;
uint64_t file_bytes;
int64_t writer_pid;
volatile uint64_t heartbeat_ns; /* CLOCK_MONOTONIC, refreshed at least every 0.2 s */
uint8_t reserved[16];
fh_ring_cam_t cams[FH_RING_MAX_CAMS];
} fh_ring_hdr_t;
static_assert(sizeof(fh_ring_cam_t) == 160, "fh_ring_cam_t layout");
static_assert(sizeof(fh_ring_slot_t) == 64, "fh_ring_slot_t layout");
static_assert(sizeof(fh_ring_hdr_t) == 64 + 160 * FH_RING_MAX_CAMS, "fh_ring_hdr_t layout");
+427
View File
@@ -0,0 +1,427 @@
/*
* tp - read kernel tracepoints system-wide through perf_event_open.
*/
#define _GNU_SOURCE
#include "tp.h"
#include <errno.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/epoll.h>
#include <sys/ioctl.h>
#include <sys/mman.h>
#include <sys/syscall.h>
#include <time.h>
#include <unistd.h>
#include <linux/perf_event.h>
#ifndef TRACEFS
#define TRACEFS "/sys/kernel/tracing/events"
#endif
#define RING_DATA_PAGES 16
static void set_err(char *err, size_t n, const char *fmt, ...)
{
va_list ap;
va_start(ap, fmt);
vsnprintf(err, n, fmt, ap);
va_end(ap);
}
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn)
{
memset(ev, 0, sizeof(*ev));
snprintf(ev->system, sizeof(ev->system), "%s", system);
snprintf(ev->name, sizeof(ev->name), "%s", name);
ev->id = -1;
char path[256];
snprintf(path, sizeof(path), TRACEFS "/%s/%s/format", system, name);
FILE *f = fopen(path, "r");
if (!f) {
set_err(err, errn, "%s: %s", path, strerror(errno));
return false;
}
char line[512];
while (fgets(line, sizeof(line), f)) {
int id;
if (sscanf(line, "ID: %d", &id) == 1) {
ev->id = id;
continue;
}
char *fp = line;
while (*fp == ' ' || *fp == '\t')
fp++;
if (strncmp(fp, "field:", 6) || ev->nfields >= TP_MAX_FIELDS)
continue;
char *semi = strchr(fp, ';');
if (!semi)
continue;
/* the field name is the last identifier in the declaration */
char decl[256];
size_t dl = (size_t)(semi - (fp + 6));
if (dl >= sizeof(decl))
dl = sizeof(decl) - 1;
memcpy(decl, fp + 6, dl);
decl[dl] = 0;
char *br = strchr(decl, '[');
if (br)
*br = 0;
char *end = decl + strlen(decl);
while (end > decl && (end[-1] == ' ' || end[-1] == '\t'))
*--end = 0;
char *start = end;
while (start > decl && start[-1] != ' ' && start[-1] != '\t' && start[-1] != '*')
start--;
tp_field_t *fd = &ev->fields[ev->nfields];
const char *o = strstr(semi, "offset:");
const char *s = strstr(semi, "size:");
const char *g = strstr(semi, "signed:");
if (!o || !s)
continue;
snprintf(fd->name, sizeof(fd->name), "%s", start);
fd->offset = atoi(o + 7);
fd->size = atoi(s + 5);
fd->is_signed = g ? atoi(g + 7) != 0 : false;
ev->nfields++;
}
fclose(f);
if (ev->id < 0) {
set_err(err, errn, "%s: no ID line", path);
return false;
}
return true;
}
int tp_field(const tp_event_t *ev, const char *name)
{
for (int i = 0; i < ev->nfields; i++)
if (!strcmp(ev->fields[i].name, name))
return i;
return -1;
}
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen)
{
if (field < 0 || field >= ev->nfields)
return 0;
const tp_field_t *f = &ev->fields[field];
if (f->offset < 0 || (uint32_t)(f->offset + f->size) > rawlen)
return 0;
const uint8_t *p = raw + f->offset;
switch (f->size) {
case 1: { uint8_t v; memcpy(&v, p, 1); return f->is_signed ? (int64_t)(int8_t)v : (int64_t)v; }
case 2: { uint16_t v; memcpy(&v, p, 2); return f->is_signed ? (int64_t)(int16_t)v : (int64_t)v; }
case 4: { uint32_t v; memcpy(&v, p, 4); return f->is_signed ? (int64_t)(int32_t)v : (int64_t)v; }
case 8: { uint64_t v; memcpy(&v, p, 8); return (int64_t)v; }
default: return 0;
}
}
static int online_cpus(int *cpus, int max)
{
FILE *f = fopen("/sys/devices/system/cpu/online", "r");
int n = 0;
if (!f)
return 0;
char buf[256] = {0};
if (!fgets(buf, sizeof(buf), f))
buf[0] = 0;
fclose(f);
for (char *tok = strtok(buf, ",\n"); tok && n < max; tok = strtok(NULL, ",\n")) {
int a, b;
if (sscanf(tok, "%d-%d", &a, &b) == 2) {
for (int c = a; c <= b && n < max; c++)
cpus[n++] = c;
} else if (sscanf(tok, "%d", &a) == 1) {
cpus[n++] = a;
}
}
return n;
}
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn)
{
memset(tp, 0, sizeof(*tp));
tp->epfd = -1;
if (nevents <= 0 || nevents > TP_MAX_EVENTS) {
set_err(err, errn, "bad event count %d", nevents);
return false;
}
for (int i = 0; i < nevents; i++)
tp->events[i] = events[i];
tp->nevents = nevents;
int cpus[TP_MAX_CPUS];
tp->ncpu = online_cpus(cpus, TP_MAX_CPUS);
if (tp->ncpu <= 0) {
set_err(err, errn, "no online CPUs found");
return false;
}
long page = sysconf(_SC_PAGESIZE);
tp->map_len = (size_t)page * (1 + RING_DATA_PAGES);
tp->epfd = epoll_create1(EPOLL_CLOEXEC);
if (tp->epfd < 0) {
set_err(err, errn, "epoll_create1: %s", strerror(errno));
return false;
}
for (int c = 0; c < tp->ncpu; c++) {
tp->ring_fd[c] = -1;
for (int e = 0; e < nevents; e++) {
struct perf_event_attr a;
memset(&a, 0, sizeof(a));
a.size = sizeof(a);
a.type = PERF_TYPE_TRACEPOINT;
a.config = (uint64_t)events[e]->id;
a.sample_period = 1;
a.sample_type = PERF_SAMPLE_TID | PERF_SAMPLE_TIME | PERF_SAMPLE_CPU | PERF_SAMPLE_RAW;
a.wakeup_events = 1;
a.use_clockid = 1;
a.clockid = CLOCK_MONOTONIC;
a.disabled = 1;
int fd = (int)syscall(SYS_perf_event_open, &a, -1, cpus[c], -1, PERF_FLAG_FD_CLOEXEC);
if (fd < 0) {
set_err(err, errn, "perf_event_open(%s:%s, cpu %d): %s",
events[e]->system, events[e]->name, cpus[c], strerror(errno));
tp_close(tp);
return false;
}
tp->fds[tp->nfds++] = fd;
if (tp->ring_fd[c] < 0) {
void *m = mmap(NULL, tp->map_len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
if (m == MAP_FAILED) {
set_err(err, errn, "mmap perf ring (cpu %d): %s", cpus[c], strerror(errno));
tp_close(tp);
return false;
}
tp->ring[c] = m;
tp->ring_fd[c] = fd;
struct epoll_event ee = { .events = EPOLLIN, .data.u32 = (uint32_t)c };
epoll_ctl(tp->epfd, EPOLL_CTL_ADD, fd, &ee);
} else if (ioctl(fd, PERF_EVENT_IOC_SET_OUTPUT, tp->ring_fd[c]) < 0) {
set_err(err, errn, "PERF_EVENT_IOC_SET_OUTPUT: %s", strerror(errno));
tp_close(tp);
return false;
}
}
}
for (int i = 0; i < tp->nfds; i++)
ioctl(tp->fds[i], PERF_EVENT_IOC_ENABLE, 0);
return true;
}
static void ring_copy(uint8_t *dst, const uint8_t *base, uint64_t size, uint64_t pos, size_t len)
{
uint64_t off = pos % size;
size_t first = (size_t)(size - off);
if (first >= len) {
memcpy(dst, base + off, len);
} else {
memcpy(dst, base + off, first);
memcpy(dst + first, base, len - first);
}
}
static int cmp_sample(const void *a, const void *b)
{
const tp_sample_t *x = a, *y = b;
return (x->time > y->time) - (x->time < y->time);
}
static void dispatch(tp_t *tp, tp_cb cb, void *ctx)
{
qsort(tp->pend, tp->npend, sizeof(tp->pend[0]), cmp_sample);
for (int i = 0; i < tp->npend; i++)
cb(ctx, &tp->pend[i]);
tp->npend = 0;
}
static int drain_ring(tp_t *tp, int c, tp_cb cb, void *ctx)
{
struct perf_event_mmap_page *pg = tp->ring[c];
long page = sysconf(_SC_PAGESIZE);
uint64_t off = pg->data_offset ? pg->data_offset : (uint64_t)page;
uint64_t size = pg->data_size ? pg->data_size : (uint64_t)page * RING_DATA_PAGES;
const uint8_t *base = (const uint8_t *)pg + off;
uint64_t head = __atomic_load_n(&pg->data_head, __ATOMIC_ACQUIRE);
uint64_t tail = pg->data_tail;
int n = 0;
while (tail < head) {
struct perf_event_header hdr;
ring_copy((uint8_t *)&hdr, base, size, tail, sizeof(hdr));
if (hdr.size < sizeof(hdr))
break;
ring_copy(tp->scratch, base, size, tail, hdr.size);
const uint8_t *p = tp->scratch + sizeof(hdr);
const uint8_t *end = tp->scratch + hdr.size;
if (hdr.type == PERF_RECORD_LOST && end - p >= 16) {
uint64_t lost;
memcpy(&lost, p + 8, 8);
tp->lost += lost;
} else if (hdr.type == PERF_RECORD_SAMPLE && end - p >= 28) {
tp_sample_t s;
uint32_t v32[2];
memcpy(v32, p, 8); p += 8;
s.pid = v32[0];
s.tid = v32[1];
memcpy(&s.time, p, 8); p += 8;
memcpy(v32, p, 8); p += 8;
s.cpu = v32[0];
memcpy(&s.rawlen, p, 4); p += 4;
s.raw = p;
if (s.rawlen >= 2 && p + s.rawlen <= end) {
uint16_t type;
memcpy(&type, s.raw, 2);
s.ev = NULL;
for (int e = 0; e < tp->nevents; e++)
if (tp->events[e]->id == type)
s.ev = tp->events[e];
if (s.ev && s.rawlen <= TP_MAX_RAW) {
if (tp->npend == TP_MAX_PENDING)
dispatch(tp, cb, ctx);
memcpy(tp->pend_raw[tp->npend], s.raw, s.rawlen);
s.raw = tp->pend_raw[tp->npend];
tp->pend[tp->npend++] = s;
n++;
}
}
}
tail += hdr.size;
}
__atomic_store_n(&pg->data_tail, tail, __ATOMIC_RELEASE);
return n;
}
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx)
{
struct epoll_event ev[TP_MAX_CPUS];
if (epoll_wait(tp->epfd, ev, TP_MAX_CPUS, timeout_ms) < 0 && errno != EINTR)
return -1;
/*
* Drain every ring, not just the ones that woke us: samples from several
* CPUs need to be handled together to keep per-camera order sane.
*/
int n = 0;
for (int c = 0; c < tp->ncpu; c++)
if (tp->ring[c])
n += drain_ring(tp, c, cb, ctx);
dispatch(tp, cb, ctx);
return n;
}
void tp_close(tp_t *tp)
{
for (int i = 0; i < tp->nfds; i++) {
ioctl(tp->fds[i], PERF_EVENT_IOC_DISABLE, 0);
}
for (int c = 0; c < tp->ncpu; c++)
if (tp->ring[c])
munmap(tp->ring[c], tp->map_len);
for (int i = 0; i < tp->nfds; i++)
close(tp->fds[i]);
if (tp->epfd >= 0)
close(tp->epfd);
tp->nfds = 0;
tp->epfd = -1;
}
+76
View File
@@ -0,0 +1,76 @@
/*
* tp - read kernel tracepoints system-wide through perf_event_open.
*
* One perf ring per CPU; every event on that CPU writes into it. Field
* offsets come from the tracefs format files, so kernel layout changes don't
* silently break parsing. Needs root (or CAP_PERFMON plus tracefs access).
*/
#pragma once
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#define TP_MAX_FIELDS 40
#define TP_MAX_EVENTS 8
#define TP_MAX_CPUS 64
#define TP_MAX_PENDING 2048
#define TP_MAX_RAW 256
typedef struct {
char name[48];
int offset;
int size;
bool is_signed;
} tp_field_t;
typedef struct {
char system[32];
char name[48];
int id;
tp_field_t fields[TP_MAX_FIELDS];
int nfields;
} tp_event_t;
typedef struct {
const tp_event_t *ev;
const uint8_t *raw;
uint32_t rawlen;
uint64_t time; /* CLOCK_MONOTONIC ns */
uint32_t cpu;
uint32_t pid;
uint32_t tid;
} tp_sample_t;
typedef void (*tp_cb)(void *ctx, const tp_sample_t *s);
typedef struct {
int ncpu;
int ring_fd[TP_MAX_CPUS];
void *ring[TP_MAX_CPUS];
size_t map_len;
int fds[TP_MAX_CPUS * TP_MAX_EVENTS];
int nfds;
int epfd;
tp_event_t *events[TP_MAX_EVENTS];
int nevents;
uint64_t lost;
uint8_t scratch[65536];
/* samples drained from all rings, sorted by time before dispatch */
tp_sample_t pend[TP_MAX_PENDING];
uint8_t pend_raw[TP_MAX_PENDING][TP_MAX_RAW];
int npend;
} tp_t;
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn);
int tp_field(const tp_event_t *ev, const char *name);
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen);
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn);
/*
* Wait up to timeout_ms, then hand every pending sample to cb in time order,
* across all CPUs. Returns samples read, -1 on error.
*/
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx);
void tp_close(tp_t *tp);
+783
View File
@@ -0,0 +1,783 @@
/*
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
*
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
*
* Everything is discovered rather than hardcoded:
* - XRService is found by scanning /proc for its cmdline.
* - The V4L2 nodes and sensor subdevs it holds open come from /proc/<pid>/fd.
* - Each node's geometry comes from VIDIOC_G_FMT on our own handle.
* - Each node is traced back to its sensor through MEDIA_IOC_G_TOPOLOGY.
* - Buffers are split into queues by allocation order: XRService opens a
* sensor subdev, then allocates that camera's buffers.
*/
#define _GNU_SOURCE
#include "xrcams.h"
#include <dirent.h>
#include <errno.h>
#include <fcntl.h>
#include <stdarg.h>
#include <stdlib.h>
#include <string.h>
#include <sys/ioctl.h>
#include <sys/stat.h>
#include <sys/sysmacros.h>
#include <unistd.h>
#include <linux/media.h>
#ifndef MEDIA_ENT_F_CAM_SENSOR
#define MEDIA_ENT_F_CAM_SENSOR 0x00020001
#endif
#define MAX_FDENTS 4096
#define MAX_TOPOS 8
enum fdkind { FD_DMABUF, FD_SUBDEV_SENSOR, FD_VIDEO };
typedef struct {
int xfd;
enum fdkind kind;
size_t size;
unsigned long ino;
char sensor[XR_SENSOR_LEN];
char path[64];
} fdent_t;
typedef struct {
struct media_v2_entity *ents;
struct media_v2_interface *intfs;
struct media_v2_pad *pads;
struct media_v2_link *links;
__u32 nents, nintfs, npads, nlinks;
} topo_t;
static fdent_t fdents[MAX_FDENTS];
static int nfdents;
static topo_t topos[MAX_TOPOS];
static int ntopos;
static void set_err(char *err, size_t n, const char *fmt, ...)
{
va_list ap;
va_start(ap, fmt);
vsnprintf(err, n, fmt, ap);
va_end(ap);
}
void xr_slugify(const char *in, char *out, size_t n)
{
size_t i = 0;
for (; in[i] && i + 1 < n; i++)
out[i] = (in[i] == ' ' || in[i] == '/') ? '_' : in[i];
out[i] = 0;
}
/* --------------------------------------------------- media graph handling */
static void topo_free_all(void)
{
for (int i = 0; i < ntopos; i++) {
free(topos[i].ents);
free(topos[i].intfs);
free(topos[i].pads);
free(topos[i].links);
}
ntopos = 0;
}
static void topo_load_all(void)
{
for (int mi = 0; mi < MAX_TOPOS; mi++) {
char mpath[32];
snprintf(mpath, sizeof(mpath), "/dev/media%d", mi);
int mfd = open(mpath, O_RDWR | O_CLOEXEC);
if (mfd < 0)
continue;
struct media_v2_topology t;
memset(&t, 0, sizeof(t));
if (ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) < 0) {
close(mfd);
continue;
}
topo_t *o = &topos[ntopos];
memset(o, 0, sizeof(*o));
o->nents = t.num_entities;
o->nintfs = t.num_interfaces;
o->npads = t.num_pads;
o->nlinks = t.num_links;
o->ents = calloc(o->nents ? o->nents : 1, sizeof(*o->ents));
o->intfs = calloc(o->nintfs ? o->nintfs : 1, sizeof(*o->intfs));
o->pads = calloc(o->npads ? o->npads : 1, sizeof(*o->pads));
o->links = calloc(o->nlinks ? o->nlinks : 1, sizeof(*o->links));
t.ptr_entities = (__u64)(uintptr_t)o->ents;
t.ptr_interfaces = (__u64)(uintptr_t)o->intfs;
t.ptr_pads = (__u64)(uintptr_t)o->pads;
t.ptr_links = (__u64)(uintptr_t)o->links;
bool ok = o->ents && o->intfs && o->pads && o->links &&
ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) == 0;
close(mfd);
if (!ok) {
free(o->ents); free(o->intfs); free(o->pads); free(o->links);
continue;
}
ntopos++;
}
}
static struct media_v2_entity *topo_entity(topo_t *t, __u32 id)
{
for (__u32 i = 0; i < t->nents; i++)
if (t->ents[i].id == id)
return &t->ents[i];
return NULL;
}
static struct media_v2_pad *topo_pad(topo_t *t, __u32 id)
{
for (__u32 i = 0; i < t->npads; i++)
if (t->pads[i].id == id)
return &t->pads[i];
return NULL;
}
static __u32 topo_entity_for_devnode(topo_t *t, dev_t rdev)
{
__u32 intf_id = 0;
for (__u32 i = 0; i < t->nintfs; i++)
if (t->intfs[i].devnode.major == major(rdev) &&
t->intfs[i].devnode.minor == minor(rdev)) {
intf_id = t->intfs[i].id;
break;
}
if (!intf_id)
return 0;
for (__u32 i = 0; i < t->nlinks; i++)
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) == MEDIA_LNK_FL_INTERFACE_LINK &&
t->links[i].source_id == intf_id)
return t->links[i].sink_id;
return 0;
}
/*
* Walk upstream across enabled data links until a sensor is reached. A CSIPHY
* carries two sensors on separate (sink, source) pad pairs, so re-enter on the
* sink pad paired with the source pad we left through.
*/
static bool topo_walk_to_sensor(topo_t *t, __u32 ent_id, char *out, size_t outn)
{
int exit_pad_index = -1;
for (int hop = 0; hop < 32 && ent_id; hop++) {
struct media_v2_entity *e = topo_entity(t, ent_id);
if (!e)
return false;
if (e->function == MEDIA_ENT_F_CAM_SENSOR) {
snprintf(out, outn, "%s", e->name);
return true;
}
__u32 first_sink = 0, paired = 0;
int nsinks = 0;
for (__u32 p = 0; p < t->npads; p++) {
if (t->pads[p].entity_id != ent_id || !(t->pads[p].flags & MEDIA_PAD_FL_SINK))
continue;
nsinks++;
if (!first_sink)
first_sink = t->pads[p].id;
if (exit_pad_index >= 1 && (int)t->pads[p].index == exit_pad_index - 1)
paired = t->pads[p].id;
}
__u32 sink_pad = (nsinks == 1) ? first_sink : (paired ? paired : first_sink);
if (!sink_pad)
return false;
__u32 src_pad = 0;
for (__u32 i = 0; i < t->nlinks; i++) {
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) != MEDIA_LNK_FL_DATA_LINK)
continue;
if (!(t->links[i].flags & MEDIA_LNK_FL_ENABLED))
continue;
if (t->links[i].sink_id == sink_pad) {
src_pad = t->links[i].source_id;
break;
}
}
struct media_v2_pad *sp = src_pad ? topo_pad(t, src_pad) : NULL;
if (!sp)
return false;
ent_id = sp->entity_id;
exit_pad_index = (int)sp->index;
}
return false;
}
static bool sensor_for_video(dev_t rdev, char *out, size_t outn)
{
for (int i = 0; i < ntopos; i++) {
__u32 ent = topo_entity_for_devnode(&topos[i], rdev);
if (ent && topo_walk_to_sensor(&topos[i], ent, out, outn))
return true;
}
return false;
}
static bool sensor_for_subdev(dev_t rdev, char *out, size_t outn)
{
for (int i = 0; i < ntopos; i++) {
__u32 id = topo_entity_for_devnode(&topos[i], rdev);
struct media_v2_entity *e = id ? topo_entity(&topos[i], id) : NULL;
if (e && e->function == MEDIA_ENT_F_CAM_SENSOR) {
snprintf(out, outn, "%s", e->name);
return true;
}
}
return false;
}
static const char *role_for_sensor(const char *sensor)
{
if (strstr(sensor, "og01a1b"))
return "tracking"; /* 1056x1024 side fisheye */
if (strstr(sensor, "og0ve10"))
return "tracking"; /* 640x480 upper */
if (strstr(sensor, "imx616"))
return "passthrough"; /* 2464x2464 Arcturus color */
return "unknown";
}
/* ------------------------------------------------- XRService / proc scan */
static pid_t find_process(const char *needle)
{
DIR *d = opendir("/proc");
if (!d)
return 0;
struct dirent *e;
pid_t found = 0;
while ((e = readdir(d))) {
if (e->d_name[0] < '0' || e->d_name[0] > '9')
continue;
char path[288];
snprintf(path, sizeof(path), "/proc/%s/cmdline", e->d_name);
FILE *f = fopen(path, "rb");
if (!f)
continue;
char buf[512] = {0};
size_t got = fread(buf, 1, sizeof(buf) - 1, f);
fclose(f);
if (got == 0)
continue;
const char *base = strrchr(buf, '/');
base = base ? base + 1 : buf;
if (strstr(base, needle)) {
found = (pid_t)atoi(e->d_name);
break;
}
}
closedir(d);
return found;
}
static bool read_dmabuf_size(pid_t pid, int fd, size_t *size, unsigned long *ino)
{
char path[64];
snprintf(path, sizeof(path), "/proc/%d/fdinfo/%d", pid, fd);
FILE *f = fopen(path, "r");
if (!f)
return false;
bool have = false;
char line[256];
*ino = 0;
while (fgets(line, sizeof(line), f)) {
unsigned long long v;
if (sscanf(line, "size: %llu", &v) == 1) {
*size = (size_t)v;
have = true;
} else if (sscanf(line, "ino: %llu", &v) == 1) {
*ino = (unsigned long)v;
}
}
fclose(f);
return have;
}
static int cmp_int(const void *a, const void *b)
{
return *(const int *)a - *(const int *)b;
}
static bool scan_xr_fds(pid_t pid, char *err, size_t errn)
{
char dirpath[64];
snprintf(dirpath, sizeof(dirpath), "/proc/%d/fd", pid);
DIR *d = opendir(dirpath);
if (!d) {
set_err(err, errn, "opendir(%s): %s (are you root?)", dirpath, strerror(errno));
return false;
}
static int fds[8192];
int nfds = 0;
struct dirent *e;
while ((e = readdir(d)) && nfds < (int)(sizeof(fds) / sizeof(fds[0])))
if (e->d_name[0] >= '0' && e->d_name[0] <= '9')
fds[nfds++] = atoi(e->d_name);
closedir(d);
qsort(fds, nfds, sizeof(int), cmp_int);
nfdents = 0;
for (int i = 0; i < nfds && nfdents < MAX_FDENTS; i++) {
char link[64], target[256];
snprintf(link, sizeof(link), "/proc/%d/fd/%d", pid, fds[i]);
ssize_t n = readlink(link, target, sizeof(target) - 1);
if (n < 0)
continue;
target[n] = 0;
fdent_t ent;
memset(&ent, 0, sizeof(ent));
ent.xfd = fds[i];
if (strstr(target, "dmabuf")) {
if (!read_dmabuf_size(pid, fds[i], &ent.size, &ent.ino))
continue;
ent.kind = FD_DMABUF;
} else if (strncmp(target, "/dev/video", 10) == 0) {
ent.kind = FD_VIDEO;
snprintf(ent.path, sizeof(ent.path), "%.63s", target);
} else if (strncmp(target, "/dev/v4l-subdev", 15) == 0) {
struct stat st;
if (stat(target, &st) < 0 || !sensor_for_subdev(st.st_rdev, ent.sensor, sizeof(ent.sensor)))
continue;
ent.kind = FD_SUBDEV_SENSOR;
} else {
continue;
}
fdents[nfdents++] = ent;
}
return true;
}
/* ------------------------------------------------------ camera discovery */
static void probe_cameras(xr_state_t *st)
{
int seen[64];
int nseen = 0;
for (int i = 0; i < nfdents; i++) {
if (fdents[i].kind != FD_VIDEO)
continue;
const char *path = fdents[i].path;
int node = atoi(path + 10);
bool dup = false;
for (int k = 0; k < nseen; k++)
if (seen[k] == node)
dup = true;
if (dup || st->ncameras >= XR_MAX_CAMERAS || nseen >= 64)
continue;
seen[nseen++] = node;
int fd = open(path, O_RDWR | O_CLOEXEC);
if (fd < 0)
continue;
struct v4l2_format fmt;
memset(&fmt, 0, sizeof(fmt));
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE;
xr_camera_t *c = &st->cameras[st->ncameras];
memset(c, 0, sizeof(*c));
if (ioctl(fd, VIDIOC_G_FMT, &fmt) == 0) {
c->width = fmt.fmt.pix_mp.width;
c->height = fmt.fmt.pix_mp.height;
c->pixfmt = fmt.fmt.pix_mp.pixelformat;
c->nplanes = fmt.fmt.pix_mp.num_planes;
c->bytesperline = fmt.fmt.pix_mp.plane_fmt[0].bytesperline;
for (unsigned p = 0; p < c->nplanes && p < VIDEO_MAX_PLANES; p++)
c->planesize[p] = fmt.fmt.pix_mp.plane_fmt[p].sizeimage;
} else {
memset(&fmt, 0, sizeof(fmt));
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
if (ioctl(fd, VIDIOC_G_FMT, &fmt) < 0) {
close(fd);
continue;
}
c->width = fmt.fmt.pix.width;
c->height = fmt.fmt.pix.height;
c->pixfmt = fmt.fmt.pix.pixelformat;
c->nplanes = 1;
c->bytesperline = fmt.fmt.pix.bytesperline;
c->planesize[0] = fmt.fmt.pix.sizeimage;
}
struct stat sb;
if (fstat(fd, &sb) == 0) {
c->minor = minor(sb.st_rdev);
sensor_for_video(sb.st_rdev, c->sensor, sizeof(c->sensor));
}
close(fd);
if (!c->sensor[0])
snprintf(c->sensor, sizeof(c->sensor), "unknown");
c->node = node;
snprintf(c->path, sizeof(c->path), "%s", path);
c->role = role_for_sensor(c->sensor);
st->ncameras++;
}
}
/*
* qcom-camss can report bytesperline as the visible width while the VFE
* writes a larger aligned pitch. sizeimage is right, so derive the pitch.
*/
unsigned xr_camera_stride(const xr_camera_t *c)
{
if (!c->height || !c->planesize[0])
return c->bytesperline ? c->bytesperline : c->width;
double bpp = 1.0;
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21)
bpp = 1.5;
unsigned s = (unsigned)((double)c->planesize[0] / ((double)c->height * bpp));
if (s >= c->width && s <= c->width * 4)
return s;
return c->bytesperline ? c->bytesperline : c->width;
}
/*
* The Arcturus color cameras (arcimx616) claim 2464x2464 NV12, but measured on
* 2026-09-28 their plane 0 holds 10-bit MIPI-packed YUV 4:2:0: 2464 luma rows
* then 1232 rows of interleaved UV, each row 2464 packed pixels (3080 bytes)
* padded to a 256-byte pitch (3328). Only the first 1972 pixels of a row carry
* image; the rest are zero.
*/
#define IMX616_VALID_WIDTH 1972
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l)
{
memset(l, 0, sizeof(*l));
l->height = c->height;
if (c->pixfmt == V4L2_PIX_FMT_NV12 && strstr(c->sensor, "imx616")) {
unsigned packed = (c->width * 5 + 3) / 4;
l->fmt = XR_FMT_YUV420_10P;
l->pitch = (packed + 255) & ~255u;
l->rows = c->height + c->height / 2;
l->width = IMX616_VALID_WIDTH < c->width ? IMX616_VALID_WIDTH : c->width;
return;
}
l->pitch = xr_camera_stride(c);
l->width = c->width < l->pitch ? c->width : l->pitch;
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21) {
l->fmt = XR_FMT_NV12;
l->rows = c->height + c->height / 2;
} else {
l->fmt = XR_FMT_GREY8;
l->rows = c->height;
}
}
const char *xr_fmt_name(xr_fmt_t f)
{
switch (f) {
case XR_FMT_GREY8: return "grey8";
case XR_FMT_NV12: return "nv12";
case XR_FMT_YUV420_10P: return "yuv420_10p";
}
return "?";
}
/* ------------------------------------------------------- buffer grouping */
/*
* XRService allocates one udmabuf per plane, plane 0 then plane 1, a whole
* queue at a time right after opening the sensor's subdev. Plane 1 matches
* VIDIOC_G_FMT exactly; plane 0 has slack, so it is matched with >=.
*/
static void build_groups(xr_state_t *st)
{
char current_sensor[XR_SENSOR_LEN] = "";
for (int i = 0; i < nfdents; i++) {
if (fdents[i].kind == FD_SUBDEV_SENSOR) {
snprintf(current_sensor, sizeof(current_sensor), "%s", fdents[i].sensor);
continue;
}
if (fdents[i].kind != FD_DMABUF)
continue;
if (i + 1 >= nfdents || fdents[i + 1].kind != FD_DMABUF)
continue;
size_t s0 = fdents[i].size;
size_t s1 = fdents[i + 1].size;
bool match = false;
for (int c = 0; c < st->ncameras; c++) {
xr_camera_t *cam = &st->cameras[c];
if (cam->nplanes >= 2 && s1 == cam->planesize[1] && s0 >= cam->planesize[0]) {
match = true;
break;
}
}
if (!match)
continue;
xr_group_t *g = NULL;
if (st->ngroups > 0) {
xr_group_t *last = &st->groups[st->ngroups - 1];
if (last->planesize[0] == s0 && last->planesize[1] == s1 &&
!strcmp(last->sensor, current_sensor))
g = last;
}
if (!g) {
if (st->ngroups >= XR_MAX_GROUPS)
break;
g = &st->groups[st->ngroups++];
memset(g, 0, sizeof(*g));
g->planesize[0] = s0;
g->planesize[1] = s1;
snprintf(g->sensor, sizeof(g->sensor), "%s", current_sensor);
}
if (g->nbufs < XR_MAX_RUNBUFS) {
g->buf[g->nbufs].xfd = fdents[i].xfd;
g->buf[g->nbufs].xfd1 = fdents[i + 1].xfd;
g->buf[g->nbufs].size = s0;
g->buf[g->nbufs].size1 = s1;
g->nbufs++;
}
i++; /* consume the plane 1 descriptor */
}
int keep = 0;
for (int i = 0; i < st->ngroups; i++)
if (st->groups[i].nbufs >= 4)
st->groups[keep++] = st->groups[i];
st->ngroups = keep;
/*
* Bind each run to a camera. The sensor marker alone can be wrong: XRService
* sometimes opens another sensor's subdev (e.g. the idle color camera)
* between an upper camera's subdev and its buffers, and two upper cameras
* can resolve to the same sensor name. So a marker match must also fit the
* camera's plane sizes, and each camera takes at most one run.
*/
for (int pass = 0; pass < 2; pass++)
for (int i = 0; i < st->ngroups; i++) {
xr_group_t *g = &st->groups[i];
for (int c = 0; c < st->ncameras && !g->cam; c++) {
xr_camera_t *cam = &st->cameras[c];
if (pass == 0 && (!g->sensor[0] || strcmp(cam->sensor, g->sensor)))
continue;
if (cam->nplanes < 2 || g->planesize[1] != cam->planesize[1] ||
g->planesize[0] < cam->planesize[0])
continue;
bool taken = false;
for (int k = 0; k < st->ngroups; k++)
if (k != i && st->groups[k].cam == cam)
taken = true;
if (!taken)
g->cam = cam;
}
}
}
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn)
{
memset(st, 0, sizeof(*st));
st->pid = find_process(process);
if (!st->pid) {
set_err(err, errn, "%s is not running; start SteamVR on the headset first", process);
return false;
}
topo_load_all();
bool ok = scan_xr_fds(st->pid, err, errn);
if (ok) {
probe_cameras(st);
build_groups(st);
}
topo_free_all();
return ok;
}
void xr_print(const xr_state_t *st, FILE *f)
{
fprintf(f, "XRService pid %d\n", st->pid);
for (int i = 0; i < st->ncameras; i++) {
const xr_camera_t *c = &st->cameras[i];
char fcc[5] = {
(char)(c->pixfmt & 0xff), (char)((c->pixfmt >> 8) & 0xff),
(char)((c->pixfmt >> 16) & 0xff), (char)((c->pixfmt >> 24) & 0xff), 0
};
fprintf(f, " camera %-12s minor %-3u %-16s %ux%u %s pitch %u planes %zu %zu role=%s\n",
c->path, c->minor, c->sensor, c->width, c->height, fcc,
xr_camera_stride(c), c->planesize[0], c->planesize[1], c->role);
}
for (int i = 0; i < st->ngroups; i++) {
const xr_group_t *g = &st->groups[i];
fprintf(f, " queue %d: %d buffers plane0=%zu plane1=%zu fds %d..%d sensor '%s' -> %s\n",
i, g->nbufs, g->planesize[0], g->planesize[1],
g->buf[0].xfd, g->buf[g->nbufs - 1].xfd1, g->sensor,
g->cam ? g->cam->path : "(unbound)");
}
}
+82
View File
@@ -0,0 +1,82 @@
/*
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
*
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
*/
#pragma once
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
#include <sys/types.h>
#include <linux/videodev2.h>
#define XR_MAX_CAMERAS 16
#define XR_MAX_GROUPS 32
#define XR_MAX_RUNBUFS 128
#define XR_SENSOR_LEN 64
typedef struct {
int node; /* N from /dev/videoN */
unsigned minor; /* char device minor, as tracepoints report it */
char path[64];
unsigned width;
unsigned height;
unsigned bytesperline;
unsigned nplanes;
size_t planesize[VIDEO_MAX_PLANES];
uint32_t pixfmt;
char sensor[XR_SENSOR_LEN]; /* media entity name of the sensor */
const char *role;
} xr_camera_t;
typedef struct {
int xfd; /* plane 0 descriptor in XRService */
int xfd1; /* plane 1 descriptor in XRService */
size_t size;
size_t size1;
} xr_bufref_t;
/* One run of buffers XRService allocated for a camera queue, in allocation order. */
typedef struct {
size_t planesize[2];
int nbufs;
xr_bufref_t buf[XR_MAX_RUNBUFS];
char sensor[XR_SENSOR_LEN]; /* from the preceding sensor subdev */
xr_camera_t *cam;
} xr_group_t;
typedef struct {
pid_t pid;
xr_camera_t cameras[XR_MAX_CAMERAS];
int ncameras;
xr_group_t groups[XR_MAX_GROUPS];
int ngroups;
} xr_state_t;
typedef enum {
XR_FMT_GREY8, /* 8-bit mono */
XR_FMT_NV12, /* 8-bit Y plane then interleaved UV, same pitch */
XR_FMT_YUV420_10P /* like NV12, but 10-bit MIPI-packed (4 px in 5 bytes) */
} xr_fmt_t;
/* Where the image really sits in plane 0; V4L2's numbers can be misleading. */
typedef struct {
xr_fmt_t fmt;
unsigned pitch; /* bytes per row */
unsigned rows; /* rows in plane 0: luma, plus chroma for YUV */
unsigned width; /* valid pixels per row */
unsigned height; /* luma rows */
} xr_layout_t;
/* Scan XRService's descriptors and the media graph. Needs root. */
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn);
unsigned xr_camera_stride(const xr_camera_t *c);
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l);
const char *xr_fmt_name(xr_fmt_t f);
void xr_print(const xr_state_t *st, FILE *f);
void xr_slugify(const char *in, char *out, size_t n);
+22
View File
@@ -0,0 +1,22 @@
# Template: the installer replaces @REPO@ with the repo path on the Frame.
[Unit]
Description=Frametop camera broker: the headset cameras' frames, for hand tracking
Documentation=file://@REPO@/hands/README.md
# It borrows XRService's camera buffers, so it comes and goes with SteamVR.
After=steamvr.service
PartOf=steamvr.service
[Service]
# On the host: the dev container can't reach XRService. Its file capabilities (set by
# hands/run.sh install) let it borrow the buffers; it drops them once set up. It exits when
# XRService restarts, and comes back to attach to the new one.
# Mono cameras only: while the headset is worn the colour module writes only a half-size
# image into the top-left quarter of its buffers (2026-09-30), which ft-camd can't use yet.
# Add --with-color to try the colour cameras (ft-hands then picks them by the light).
ExecStart=@REPO@/hands/build/ft-camd --status 60
Restart=always
RestartSec=5
TimeoutStopSec=5
[Install]
WantedBy=steamvr.service
+20
View File
@@ -0,0 +1,20 @@
# Template: the installer replaces @REPO@ with the repo path on the Frame.
[Unit]
Description=Frametop hand tracking: hands for the screens' hand cutouts, pinches for the pointer
Documentation=file://@REPO@/hands/README.md
After=steamvr.service frametop-camd.service
Wants=frametop-camd.service
PartOf=steamvr.service
[Service]
# In the dev container (it's built against Fedora's libraries). It reads ft-camd's ring and
# writes /run/user/UID/frametop-hands/hands and gestures. Settings: HANDS_* in ~/.config/frametop.conf.
ExecStartPre=-@REPO@/scripts/container-up.sh
ExecStartPre=-/usr/bin/pkill -x ft-hands
ExecStart=%h/.local/bin/distrobox enter dev -- @REPO@/hands/build/ft-hands --status 60
Restart=always
RestartSec=5
TimeoutStopSec=5
[Install]
WantedBy=steamvr.service
+48
View File
@@ -0,0 +1,48 @@
#!/usr/bin/env bash
# ft-handsctl: turn hand tracking on and off by hand, on the Frame. With it on, your hands show
# through Frametop's screens (the hand cutouts); pinches and grips only move the pointer with
# POINTER_HANDS=1. It doesn't start with SteamVR (hands/run.sh install leaves it off), and it
# stops when SteamVR does.
#
# ft-handsctl on | off | status | log [lines]
# ft-handsctl cutouts on|off|state ft-screens' hand cutouts, without stopping tracking
# ft-handsctl gestures watch pinches and grips live (Ctrl+C to stop)
#
# Needs the services installed once: hands/run.sh install (it sets ft-camd's capabilities).
set -euo pipefail
here=$(cd "$(dirname "$(readlink -f "${BASH_SOURCE[0]}")")" && pwd)
units="frametop-camd.service frametop-hands.service"
ask_screens() { # a command to ft-screens' control socket, and its reply
python3 - "$1" <<'EOF'
import socket, sys
s = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
s.bind("")
s.settimeout(2)
try:
s.sendto(sys.argv[1].encode(), "\0ft_screens")
print(s.recv(512).decode())
except OSError as e:
sys.exit("ft-screens didn't answer (is the Frametop desktop running?): %s" % e)
EOF
}
case ${1:-status} in
on)
systemctl --user -q is-active steamvr.service || { echo "SteamVR isn't running" >&2; exit 1; }
systemctl --user start $units
sleep 4
"$0" status ;;
off)
systemctl --user stop $units
echo "hand tracking off" ;;
status)
for u in $units; do echo "$u: $(systemctl --user is-active $u || true)"; done
journalctl --user -u frametop-hands.service --no-pager -o cat -n 40 |
grep -E '^ *[0-9.]+s |sets with a hand|cameras:' | tail -2 | cut -c1-160 || true ;;
log) journalctl --user -u frametop-camd.service -u frametop-hands.service --no-pager -o short -n "${2:-30}" ;;
cutouts)
ask_screens "cutouts ${2:-state}" ;;
gestures) exec python3 "$here/tools/watch_gestures.py" --distance ;;
*) sed -n '2,11p' "$0" | sed 's/^# \{0,1\}//'; exit 2 ;;
esac
+81
View File
@@ -0,0 +1,81 @@
/*
* fh_gestures - hand gestures ft-hands publishes for input (the pointer helper): look at
* something and pinch to click it, or close the hand (a grip) to press and drag it (the
* Vision Pro model, with the eye tracker doing the looking).
* /run/user/UID/frametop-hands/gestures, next to the hands
* file, with the same sequence lock (read seq, copy, read seq again; use the copy only if
* both reads are the same even number) and the same frame: metres in the head frame at
* capture time, OpenVR's HMD frame (+x right, +y up, -z forward).
*
* One slot per side and gesture: pinch[0] and grip[0] are the left hand, [1] the right.
* A gesture follows the hand it began on until it ends.
* Pinch: begins when the thumb and index tips close within begin_m and ends when they
* open past end_m (the gap between keeps it from flickering). point is the index and middle
* knuckles, which don't move as the fingers open and close (the tips' midpoint did).
* Grip: a closed hand. It begins when all four fingers are curled in (each fingertip
* nearer the wrist than grip_begin times its knuckle is) and ends when they open past
* grip_end on average. distance is that average (about 2 open, under 1.2 closed), strength
* 0 open .. 1 closed, and point the palm's centre. A grip ends a pinch on the same hand
* (closing the hand can pass through a pinch on the way), as lost.
* Either ends, as lost (FH_PINCH_LOST), when its hand stays lost too long.
*
* Don't miss short gestures: a reader that polls slower than a quick tap still sees it,
* because begins and ends count every one. When begins changed, one began at begin_ns;
* when ends changed, one ended at end_ns. begins - ends is 1 while it's down.
*
* Drags: point is where the gesture is now, begin_point where it began. Turn each into the
* room with the HMD pose at its capture time (capture_ns, begin_ns) before subtracting,
* so turning your head doesn't drag.
*
* Version 1 had only the pinches (192 bytes); version 2 adds the grips after them.
*/
#pragma once
#include <assert.h>
#include <stdint.h>
#define FH_GESTURES_MAGIC "FHGEST01"
#define FH_GESTURES_VERSION 2
enum {
FH_PINCH_TRACKED = 1u << 0, /* the hand was tracked in this frame */
FH_PINCH_DOWN = 1u << 1, /* the gesture is held now */
FH_PINCH_LOST = 1u << 2, /* the last one ended because the hand was lost */
/* (or, for a pinch, a grip took over) */
};
typedef struct {
uint32_t flags; /* FH_PINCH_* */
uint32_t hand_id; /* fh_hand_t.id of the hand, 0 if none */
uint32_t begins; /* begun so far */
uint32_t ends; /* ended so far */
uint64_t begin_ns; /* capture time (CLOCK_MONOTONIC) the current or */
/* last one began */
uint64_t end_ns; /* ... the last one ended */
float distance; /* pinch: thumb tip to index tip, m, at this user's */
/* hand size. grip: the fingers' mean curl (above) */
float strength; /* 0 open .. 1 closed */
float point[3]; /* pinch: the index and middle knuckles; grip: the */
/* palm's centre */
float begin_point[3]; /* point when the current or last one began */
} fh_pinch_t; /* 64 bytes */
typedef struct {
char magic[8];
uint32_t version;
uint32_t size;
volatile uint64_t seq;
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
float begin_m; /* the pinch thresholds in use */
float end_m;
float grip_begin; /* the grip thresholds in use (curl ratios) */
float grip_end;
uint8_t reserved[8];
fh_pinch_t pinch[2]; /* [0] left hand, [1] right hand */
fh_pinch_t grip[2]; /* version 2 */
} fh_gestures_t;
static_assert(sizeof(fh_pinch_t) == 64, "fh_pinch_t layout");
static_assert(sizeof(fh_gestures_t) == 64 + 4 * 64, "fh_gestures_t layout");
+58
View File
@@ -0,0 +1,58 @@
/*
* fh_hands - the tracked-hands file ft-hands publishes for ft-screens' hand cutouts
* (/run/user/UID/frametop-hands/hands, directory mode 0700), rewritten in place
* under a sequence lock: read seq, copy, read seq again; use the copy only if
* both reads are the same even number.
*
* Positions are metres in the head frame at capture time, which is OpenVR's HMD
* frame (+x right, +y up, -z forward). Turn them into the room with the HMD pose
* at capture_ns (CLOCK_MONOTONIC). Writer: ft-hands (hands/track/io.cpp).
*/
#pragma once
#include <assert.h>
#include <stdint.h>
#define FH_HANDS_MAGIC "FHHANDS1"
#define FH_HANDS_VERSION 1
#define FH_HANDS_MAX_HANDS 2
#define FH_HANDS_MAX_CAPSULES 64
enum {
FH_HAND_RIGHT = 1u << 0, /* else the left hand */
FH_HAND_STEREO = 1u << 1, /* triangulated from two or more cameras */
};
typedef struct {
uint32_t id; /* stays the same while the hand is tracked */
uint32_t flags; /* FH_HAND_* */
float confidence;
float reserved;
float pts[21][3]; /* MediaPipe hand landmarks */
uint32_t ncapsules; /* this hand's capsules, which follow the */
/* previous hands' in capsules[] */
} fh_hand_t; /* 272 bytes */
typedef struct {
float a[3], b[3]; /* segment ends */
float ra, rb; /* radius at each end */
} fh_capsule_t; /* 32 bytes: the hand's shape, to cut out */
typedef struct {
char magic[8];
uint32_t version;
uint32_t size;
volatile uint64_t seq;
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
uint32_t nhands;
uint32_t ncapsules;
uint8_t reserved[16];
fh_hand_t hands[FH_HANDS_MAX_HANDS];
fh_capsule_t capsules[FH_HANDS_MAX_CAPSULES];
} fh_hands_t;
static_assert(sizeof(fh_hand_t) == 272, "fh_hand_t layout");
static_assert(sizeof(fh_capsule_t) == 32, "fh_capsule_t layout");
static_assert(sizeof(fh_hands_t) == 64 + 2 * 272 + 64 * 32, "fh_hands_t layout");
+12
View File
@@ -0,0 +1,12 @@
The models in ncnn/ are converted from the OpenCV Zoo ONNX ports of Google's MediaPipe hand
models, by tools/convert_models.py:
- palm.ncnn.*: palm_detection_mediapipe_2023feb (https://huggingface.co/opencv/palm_detection_mediapipe)
- hand.ncnn.*: handpose_estimation_mediapipe_2023feb (https://huggingface.co/opencv/handpose_estimation_mediapipe)
MediaPipe is Copyright Google LLC. The models and the OpenCV Zoo ports are licensed under the
Apache License, Version 2.0 (https://www.apache.org/licenses/LICENSE-2.0).
Changes made here: converted to ncnn with pnnx, with the palm detector's channel pads
rewritten as ncnn Padding layers, and quantized to 8 bits (the *-int8.ncnn.* files) with
ncnn's tools.
Binary file not shown.
+79
View File
@@ -0,0 +1,79 @@
7767517
77 90
Input in0 0 1 in0
Convolution convclip_0 1 1 in0 2 0=24 1=3 3=2 15=1 16=1 5=1 6=648 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 4=1 5=1 6=216 7=24 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_10 1 1 3 4 0=16 1=1 5=1 6=384 8=2
Split splitncnn_0 1 2 4 5 6
Convolution convclip_1 1 1 6 7 0=64 1=1 5=1 6=1024 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 3=2 15=1 16=1 5=1 6=576 7=64 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_12 1 1 8 9 0=16 1=1 5=1 6=1024 8=2
Pooling maxpool2d_1 1 1 5 10 1=2 2=2 5=1
BinaryOp add_0 2 1 9 10 11
Split splitncnn_1 1 2 11 12 13
Convolution convclip_2 1 1 13 14 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 4=1 5=1 6=864 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_14 1 1 15 16 0=16 1=1 5=1 6=1536 8=2
BinaryOp add_1 2 1 16 12 17
Convolution convclip_3 1 1 17 18 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_16 1 1 19 20 0=24 1=1 5=1 6=2304 8=2
Split splitncnn_2 1 2 20 21 22
Convolution convclip_4 1 1 22 23 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 4=2 5=1 6=3600 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_18 1 1 24 25 0=24 1=1 5=1 6=3456 8=2
BinaryOp add_2 2 1 25 21 26
Convolution convclip_5 1 1 26 27 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 3=2 15=1 16=1 5=1 6=1296 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_20 1 1 28 29 0=48 1=1 5=1 6=6912 8=2
Split splitncnn_3 1 2 29 30 31
Convolution convclip_6 1 1 31 32 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_22 1 1 33 34 0=48 1=1 5=1 6=13824 8=2
BinaryOp add_3 2 1 34 30 35
Split splitncnn_4 1 2 35 36 37
Convolution convclip_7 1 1 37 38 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_24 1 1 39 40 0=48 1=1 5=1 6=13824 8=2
BinaryOp add_4 2 1 40 36 41
Convolution convclip_8 1 1 41 42 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 4=2 5=1 6=7200 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_26 1 1 43 44 0=64 1=1 5=1 6=18432 8=2
Split splitncnn_5 1 2 44 45 46
Convolution convclip_9 1 1 46 47 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_28 1 1 48 49 0=64 1=1 5=1 6=24576 8=2
BinaryOp add_5 2 1 49 45 50
Split splitncnn_6 1 2 50 51 52
Convolution convclip_10 1 1 52 53 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_30 1 1 54 55 0=64 1=1 5=1 6=24576 8=2
BinaryOp add_6 2 1 55 51 56
Convolution convclip_11 1 1 56 57 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 3=2 4=1 15=2 16=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_32 1 1 58 59 0=112 1=1 5=1 6=43008 8=2
Split splitncnn_7 1 2 59 60 61
Convolution convclip_12 1 1 61 62 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_34 1 1 63 64 0=112 1=1 5=1 6=75264 8=2
BinaryOp add_7 2 1 64 60 65
Split splitncnn_8 1 2 65 66 67
Convolution convclip_13 1 1 67 68 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_36 1 1 69 70 0=112 1=1 5=1 6=75264 8=2
BinaryOp add_8 2 1 70 66 71
Split splitncnn_9 1 2 71 72 73
Convolution convclip_14 1 1 73 74 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_38 1 1 75 76 0=112 1=1 5=1 6=75264 8=2
BinaryOp add_9 2 1 76 72 77
Convolution convclip_15 1 1 77 78 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 4=1 5=1 6=6048 7=672 8=1 9=3 -23310=2,0.000000e+00,6.000000e+00
Pooling gap_0 1 1 79 80 0=1 4=1
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
Split splitncnn_10 1 4 82 83 84 85 86
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336 8=2
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336 8=2
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 8=2 9=4
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 8=2 9=4
Binary file not shown.
+79
View File
@@ -0,0 +1,79 @@
7767517
77 90
Input in0 0 1 in0
Convolution convclip_0 1 1 in0 2 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=648 9=3
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=216 7=24 9=3
Convolution conv_10 1 1 3 4 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=384
Split splitncnn_0 1 2 4 5 6
Convolution convclip_1 1 1 6 7 0=64 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024 9=3
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=576 7=64 9=3
Convolution conv_12 1 1 8 9 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
Pooling maxpool2d_1 1 1 5 10 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
BinaryOp add_0 2 1 9 10 11 0=0
Split splitncnn_1 1 2 11 12 13
Convolution convclip_2 1 1 13 14 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=864 7=96 9=3
Convolution conv_14 1 1 15 16 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
BinaryOp add_1 2 1 16 12 17 0=0
Convolution convclip_3 1 1 17 18 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400 7=96 9=3
Convolution conv_16 1 1 19 20 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2304
Split splitncnn_2 1 2 20 21 22
Convolution convclip_4 1 1 22 23 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3600 7=144 9=3
Convolution conv_18 1 1 24 25 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456
BinaryOp add_2 2 1 25 21 26 0=0
Convolution convclip_5 1 1 26 27 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=1296 7=144 9=3
Convolution conv_20 1 1 28 29 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=6912
Split splitncnn_3 1 2 29 30 31
Convolution convclip_6 1 1 31 32 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
Convolution conv_22 1 1 33 34 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
BinaryOp add_3 2 1 34 30 35 0=0
Split splitncnn_4 1 2 35 36 37
Convolution convclip_7 1 1 37 38 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
Convolution conv_24 1 1 39 40 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
BinaryOp add_4 2 1 40 36 41 0=0
Convolution convclip_8 1 1 41 42 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=7200 7=288 9=3
Convolution conv_26 1 1 43 44 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=18432
Split splitncnn_5 1 2 44 45 46
Convolution convclip_9 1 1 46 47 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
Convolution conv_28 1 1 48 49 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
BinaryOp add_5 2 1 49 45 50 0=0
Split splitncnn_6 1 2 50 51 52
Convolution convclip_10 1 1 52 53 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
Convolution conv_30 1 1 54 55 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
BinaryOp add_6 2 1 55 51 56 0=0
Convolution convclip_11 1 1 56 57 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=9600 7=384 9=3
Convolution conv_32 1 1 58 59 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=43008
Split splitncnn_7 1 2 59 60 61
Convolution convclip_12 1 1 61 62 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
Convolution conv_34 1 1 63 64 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
BinaryOp add_7 2 1 64 60 65 0=0
Split splitncnn_8 1 2 65 66 67
Convolution convclip_13 1 1 67 68 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
Convolution conv_36 1 1 69 70 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
BinaryOp add_8 2 1 70 66 71 0=0
Split splitncnn_9 1 2 71 72 73
Convolution convclip_14 1 1 73 74 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
Convolution conv_38 1 1 75 76 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
BinaryOp add_9 2 1 76 72 77 0=0
Convolution convclip_15 1 1 77 78 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=6048 7=672 9=3
Pooling gap_0 1 1 79 80 0=1 4=1
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
Split splitncnn_10 1 4 82 83 84 85 86
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 9=4
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 9=4
Binary file not shown.
+151
View File
@@ -0,0 +1,151 @@
7767517
149 177
Input in0 0 1 in0
Convolution padconv_0 1 1 in0 2 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 8=2
PReLU prelu_41 1 1 2 3 0=32
Split splitncnn_0 1 2 3 4 5
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 4=2 5=1 6=800 7=32 8=101
Convolution conv_12 1 1 6 7 0=32 1=1 5=1 6=1024 8=2
BinaryOp add_0 2 1 4 7 8
PReLU prelu_42 1 1 8 9 0=32
Split splitncnn_1 1 2 9 10 11
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 4=2 5=1 6=800 7=32 8=101
Convolution conv_13 1 1 12 13 0=32 1=1 5=1 6=1024 8=2
BinaryOp add_1 2 1 10 13 14
PReLU prelu_43 1 1 14 15 0=32
Split splitncnn_2 1 2 15 16 17
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 4=2 5=1 6=800 7=32 8=101
Convolution conv_14 1 1 18 19 0=32 1=1 5=1 6=1024 8=2
BinaryOp add_2 2 1 16 19 20
PReLU prelu_44 1 1 20 21 0=32
Split splitncnn_3 1 2 21 22 23
Pooling maxpool2d_2 1 1 22 24 1=2 2=2 5=1
Padding Pad_16 1 1 24 25 8=32
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=800 7=32 8=101
Convolution conv_15 1 1 26 27 0=64 1=1 5=1 6=2048 8=2
BinaryOp add_3 2 1 25 27 28
PReLU prelu_45 1 1 28 29 0=64
Split splitncnn_4 1 2 29 30 31
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
Convolution conv_16 1 1 32 33 0=64 1=1 5=1 6=4096 8=2
BinaryOp add_4 2 1 30 33 34
PReLU prelu_46 1 1 34 35 0=64
Split splitncnn_5 1 2 35 36 37
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
Convolution conv_17 1 1 38 39 0=64 1=1 5=1 6=4096 8=2
BinaryOp add_5 2 1 36 39 40
PReLU prelu_47 1 1 40 41 0=64
Split splitncnn_6 1 2 41 42 43
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
Convolution conv_18 1 1 44 45 0=64 1=1 5=1 6=4096 8=2
BinaryOp add_6 2 1 42 45 46
PReLU prelu_48 1 1 46 47 0=64
Split splitncnn_7 1 2 47 48 49
Pooling maxpool2d_3 1 1 48 50 1=2 2=2 5=1
Padding Pad_34 1 1 50 51 8=64
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 3=2 4=1 15=2 16=2 5=1 6=1600 7=64 8=101
Convolution conv_19 1 1 52 53 0=128 1=1 5=1 6=8192 8=2
BinaryOp add_7 2 1 51 53 54
PReLU prelu_49 1 1 54 55 0=128
Split splitncnn_8 1 2 55 56 57
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_20 1 1 58 59 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_8 2 1 56 59 60
PReLU prelu_50 1 1 60 61 0=128
Split splitncnn_9 1 2 61 62 63
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_21 1 1 64 65 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_9 2 1 62 65 66
PReLU prelu_51 1 1 66 67 0=128
Split splitncnn_10 1 2 67 68 69
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_22 1 1 70 71 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_10 2 1 68 71 72
PReLU prelu_52 1 1 72 73 0=128
Split splitncnn_11 1 3 73 74 75 76
Pooling maxpool2d_4 1 1 75 77 1=2 2=2 5=1
Padding Pad_52 1 1 77 78 8=128
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 3=2 4=1 15=2 16=2 5=1 6=3200 7=128 8=101
Convolution conv_23 1 1 79 80 0=256 1=1 5=1 6=32768 8=2
BinaryOp add_11 2 1 78 80 81
PReLU prelu_53 1 1 81 82 0=256
Split splitncnn_12 1 2 82 83 84
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_24 1 1 85 86 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_12 2 1 83 86 87
PReLU prelu_54 1 1 87 88 0=256
Split splitncnn_13 1 2 88 89 90
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_25 1 1 91 92 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_13 2 1 89 92 93
PReLU prelu_55 1 1 93 94 0=256
Split splitncnn_14 1 2 94 95 96
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_26 1 1 97 98 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_14 2 1 95 98 99
PReLU prelu_56 1 1 99 100 0=256
Split splitncnn_15 1 3 100 101 102 103
Pooling maxpool2d_5 1 1 102 104 1=2 2=2 5=1
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 3=2 4=1 15=2 16=2 5=1 6=6400 7=256 8=101
Convolution conv_27 1 1 105 106 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_15 2 1 104 106 107
PReLU prelu_57 1 1 107 108 0=256
Split splitncnn_16 1 2 108 109 110
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_28 1 1 111 112 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_16 2 1 109 112 113
PReLU prelu_58 1 1 113 114 0=256
Split splitncnn_17 1 2 114 115 116
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_29 1 1 117 118 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_17 2 1 115 118 119
PReLU prelu_59 1 1 119 120 0=256
Split splitncnn_18 1 2 120 121 122
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_30 1 1 123 124 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_18 2 1 121 124 125
PReLU prelu_60 1 1 125 126 0=256
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12
Convolution conv_31 1 1 127 128 0=256 1=1 5=1 6=65536 8=2
PReLU prelu_61 1 1 128 129 0=256
BinaryOp add_19 2 1 101 129 130
Split splitncnn_19 1 2 130 131 132
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_32 1 1 133 134 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_20 2 1 131 134 135
PReLU prelu_62 1 1 135 136 0=256
Split splitncnn_20 1 2 136 137 138
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_33 1 1 139 140 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_21 2 1 137 140 141
PReLU prelu_63 1 1 141 142 0=256
Split splitncnn_21 1 3 142 143 144 145
Convolution conv_34 1 1 145 146 0=108 1=1 5=1 6=27648 8=2
Permute permute_68 1 1 146 147 0=3
Reshape reshape_72 1 1 147 148 0=18 1=864
Convolution conv_35 1 1 144 149 0=6 1=1 5=1 6=1536 8=2
Permute permute_69 1 1 149 150 0=3
Reshape reshape_73 1 1 150 151 0=1 1=864
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24
Convolution conv_36 1 1 152 153 0=128 1=1 5=1 6=32768 8=2
PReLU prelu_64 1 1 153 154 0=128
BinaryOp add_22 2 1 74 154 155
Split splitncnn_22 1 2 155 156 157
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_37 1 1 158 159 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_23 2 1 156 159 160
PReLU prelu_65 1 1 160 161 0=128
Split splitncnn_23 1 2 161 162 163
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_38 1 1 164 165 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_24 2 1 162 165 166
PReLU prelu_66 1 1 166 167 0=128
Split splitncnn_24 1 2 167 168 169
Convolution conv_39 1 1 169 170 0=36 1=1 5=1 6=4608 8=2
Permute permute_70 1 1 170 171 0=3
Reshape reshape_74 1 1 171 172 0=18 1=1152
Concat cat_0 2 1 172 148 out0
Convolution conv_40 1 1 168 174 0=2 1=1 5=1 6=256 8=2
Permute permute_71 1 1 174 175 0=3
Reshape reshape_75 1 1 175 176 0=1 1=1152
Concat cat_1 2 1 176 151 out1
Binary file not shown.
+151
View File
@@ -0,0 +1,151 @@
7767517
149 177
Input in0 0 1 in0
Convolution padconv_0 1 1 in0 2 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400
PReLU prelu_41 1 1 2 3 0=32
Split splitncnn_0 1 2 3 4 5
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
Convolution conv_12 1 1 6 7 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
BinaryOp add_0 2 1 4 7 8 0=0
PReLU prelu_42 1 1 8 9 0=32
Split splitncnn_1 1 2 9 10 11
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
Convolution conv_13 1 1 12 13 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
BinaryOp add_1 2 1 10 13 14 0=0
PReLU prelu_43 1 1 14 15 0=32
Split splitncnn_2 1 2 15 16 17
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
Convolution conv_14 1 1 18 19 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
BinaryOp add_2 2 1 16 19 20 0=0
PReLU prelu_44 1 1 20 21 0=32
Split splitncnn_3 1 2 21 22 23
Pooling maxpool2d_2 1 1 22 24 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
Padding Pad_16 1 1 24 25 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=32
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=800 7=32
Convolution conv_15 1 1 26 27 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2048
BinaryOp add_3 2 1 25 27 28 0=0
PReLU prelu_45 1 1 28 29 0=64
Split splitncnn_4 1 2 29 30 31
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
Convolution conv_16 1 1 32 33 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
BinaryOp add_4 2 1 30 33 34 0=0
PReLU prelu_46 1 1 34 35 0=64
Split splitncnn_5 1 2 35 36 37
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
Convolution conv_17 1 1 38 39 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
BinaryOp add_5 2 1 36 39 40 0=0
PReLU prelu_47 1 1 40 41 0=64
Split splitncnn_6 1 2 41 42 43
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
Convolution conv_18 1 1 44 45 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
BinaryOp add_6 2 1 42 45 46 0=0
PReLU prelu_48 1 1 46 47 0=64
Split splitncnn_7 1 2 47 48 49
Pooling maxpool2d_3 1 1 48 50 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
Padding Pad_34 1 1 50 51 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=64
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=1600 7=64
Convolution conv_19 1 1 52 53 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=8192
BinaryOp add_7 2 1 51 53 54 0=0
PReLU prelu_49 1 1 54 55 0=128
Split splitncnn_8 1 2 55 56 57
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_20 1 1 58 59 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_8 2 1 56 59 60 0=0
PReLU prelu_50 1 1 60 61 0=128
Split splitncnn_9 1 2 61 62 63
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_21 1 1 64 65 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_9 2 1 62 65 66 0=0
PReLU prelu_51 1 1 66 67 0=128
Split splitncnn_10 1 2 67 68 69
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_22 1 1 70 71 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_10 2 1 68 71 72 0=0
PReLU prelu_52 1 1 72 73 0=128
Split splitncnn_11 1 3 73 74 75 76
Pooling maxpool2d_4 1 1 75 77 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
Padding Pad_52 1 1 77 78 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=128
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=3200 7=128
Convolution conv_23 1 1 79 80 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
BinaryOp add_11 2 1 78 80 81 0=0
PReLU prelu_53 1 1 81 82 0=256
Split splitncnn_12 1 2 82 83 84
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_24 1 1 85 86 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_12 2 1 83 86 87 0=0
PReLU prelu_54 1 1 87 88 0=256
Split splitncnn_13 1 2 88 89 90
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_25 1 1 91 92 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_13 2 1 89 92 93 0=0
PReLU prelu_55 1 1 93 94 0=256
Split splitncnn_14 1 2 94 95 96
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_26 1 1 97 98 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_14 2 1 95 98 99 0=0
PReLU prelu_56 1 1 99 100 0=256
Split splitncnn_15 1 3 100 101 102 103
Pooling maxpool2d_5 1 1 102 104 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=6400 7=256
Convolution conv_27 1 1 105 106 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_15 2 1 104 106 107 0=0
PReLU prelu_57 1 1 107 108 0=256
Split splitncnn_16 1 2 108 109 110
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_28 1 1 111 112 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_16 2 1 109 112 113 0=0
PReLU prelu_58 1 1 113 114 0=256
Split splitncnn_17 1 2 114 115 116
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_29 1 1 117 118 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_17 2 1 115 118 119 0=0
PReLU prelu_59 1 1 119 120 0=256
Split splitncnn_18 1 2 120 121 122
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_30 1 1 123 124 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_18 2 1 121 124 125 0=0
PReLU prelu_60 1 1 125 126 0=256
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12 6=0
Convolution conv_31 1 1 127 128 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
PReLU prelu_61 1 1 128 129 0=256
BinaryOp add_19 2 1 101 129 130 0=0
Split splitncnn_19 1 2 130 131 132
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_32 1 1 133 134 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_20 2 1 131 134 135 0=0
PReLU prelu_62 1 1 135 136 0=256
Split splitncnn_20 1 2 136 137 138
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_33 1 1 139 140 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_21 2 1 137 140 141 0=0
PReLU prelu_63 1 1 141 142 0=256
Split splitncnn_21 1 3 142 143 144 145
Convolution conv_34 1 1 145 146 0=108 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=27648
Permute permute_68 1 1 146 147 0=3
Reshape reshape_72 1 1 147 148 0=18 1=864
Convolution conv_35 1 1 144 149 0=6 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
Permute permute_69 1 1 149 150 0=3
Reshape reshape_73 1 1 150 151 0=1 1=864
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24 6=0
Convolution conv_36 1 1 152 153 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
PReLU prelu_64 1 1 153 154 0=128
BinaryOp add_22 2 1 74 154 155 0=0
Split splitncnn_22 1 2 155 156 157
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_37 1 1 158 159 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_23 2 1 156 159 160 0=0
PReLU prelu_65 1 1 160 161 0=128
Split splitncnn_23 1 2 161 162 163
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_38 1 1 164 165 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_24 2 1 162 165 166 0=0
PReLU prelu_66 1 1 166 167 0=128
Split splitncnn_24 1 2 167 168 169
Convolution conv_39 1 1 169 170 0=36 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4608
Permute permute_70 1 1 170 171 0=3
Reshape reshape_74 1 1 171 172 0=18 1=1152
Concat cat_0 2 1 172 148 out0 0=0
Convolution conv_40 1 1 168 174 0=2 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=256
Permute permute_71 1 1 174 175 0=3
Reshape reshape_75 1 1 175 176 0=1 1=1152
Concat cat_1 2 1 176 151 out1 0=0
Executable
+62
View File
@@ -0,0 +1,62 @@
#!/usr/bin/env bash
# Install, start, stop, or inspect hand tracking on the Frame: ft-camd (the camera broker) and
# ft-hands (the tracker), user services that stop with SteamVR. They don't start with it:
# ft-handsctl on|off (on the Frame) or hands/run.sh start|stop.
# Usage: hands/run.sh install|uninstall
# hands/run.sh caps # give ft-camd its capabilities again (a rebuild clears them)
# hands/run.sh start|stop|restart|status|log [lines]
# install and caps need the password (sudo setcap, once per build of ft-camd). On the Frame,
# sudo asks in the terminal. From a PC (or with no terminal), the password comes from
# steamos_root_pwd in the repo's .env and is sent to sudo -S on stdin, never on a command line.
set -euo pipefail
root=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
. "$root/scripts/_env.sh"
frame="$root/scripts/frame.sh"
units="frametop-camd.service frametop-hands.service"
# pidfd_getfd on XRService (ptrace_scope=1), system-wide tracepoints, and their root-only
# format files. ft-camd drops them all once it has set up.
caps=cap_sys_ptrace,cap_perfmon,cap_dac_read_search+ep
sudo_run() {
if [ "$FRAME_LOCAL" = 1 ] && [ -t 0 ]; then
sudo bash -c "$1" # asks for the password here
return
fi
local pw
pw=$(sed -n 's/^steamos_root_pwd=//p' "$root/.env" 2>/dev/null)
pw=${pw#[\"\']}; pw=${pw%[\"\']} # .env values may be quoted
[ -n "$pw" ] || { echo "no terminal for sudo, and steamos_root_pwd is missing from $root/.env" >&2; exit 1; }
printf '%s\n' "$pw" | on_frame "sudo -S -p '' bash -c $(printf %q "$1")"
}
set_caps() { # only when missing: a rebuild clears them, a reinstall doesn't
local bin
bin=$(printf %q "$FRAME_REPO/hands/build/ft-camd")
if on_frame "getcap $bin | grep -q cap_sys_ptrace"; then
echo "ft-camd has its capabilities"
return
fi
sudo_run "setcap $caps $bin && getcap $bin"
}
states="for u in $units; do echo \"\$u: \$(systemctl --user is-active \$u)\"; done"
case ${1:-status} in
install)
"$root/hands/build.sh"
set_caps
for u in $units; do
fill_template "$root/hands/$u" | on_frame "mkdir -p ~/.config/systemd/user && cat > ~/.config/systemd/user/$u"
done
# Installed but not started with SteamVR: ft-handsctl on|off (linked into ~/.local/bin).
"$frame" --host "set -e; systemctl --user daemon-reload; systemctl --user disable $units 2>/dev/null || true
mkdir -p ~/.local/bin && ln -sfn $(printf %q "$FRAME_REPO/hands/ft-handsctl") ~/.local/bin/ft-handsctl
$states; echo 'start it with: ft-handsctl on'" ;;
caps) set_caps ;;
uninstall) "$frame" --host "systemctl --user disable --now $units 2>/dev/null
for u in $units; do rm -f ~/.config/systemd/user/\$u; done; systemctl --user daemon-reload; echo removed" ;;
start|stop|restart) "$frame" --host "systemctl --user $1 $units; $states" ;;
status) "$frame" --host "$states; journalctl --user -u frametop-hands.service --no-pager -o cat -n 4" || true ;;
log) "$frame" --host "journalctl --user -u frametop-camd.service -u frametop-hands.service --no-pager -o short -n ${2:-30}" ;;
*) echo "usage: $0 install|uninstall|caps|start|stop|restart|status|log [lines]" >&2; exit 2 ;;
esac
+197
View File
@@ -0,0 +1,197 @@
"""Tracking-camera calibration from the headset's factory files.
/persist/xrservice.json (written by Valve's calibration, loaded by XRService)
holds, per camera, Kannala-Brandt fisheye intrinsics ("kb": fx fy cx cy k1-k4,
pixel centres at integer coordinates, as in OpenCV's fisheye model) and a pose
in the slam_right (Cam0) frame: plus_x/plus_z are the camera axes and position
its origin, in mm. /persist/device_config.json gives Cam0's pose in the CAD
frame (cv.cad_from_cal, metres) and the head's pose in CAD (head). The CAD frame
is +X head-left, +Y up, +Z forward; the head frame is OpenVR's: +x right, +y up,
-z forward. Camera frames: +z along the optical axis, +x right and +y down in
the image.
Everything here returns metres in the head frame.
"""
import json
import os
import numpy as np
XRSERVICE_JSON = '/persist/xrservice.json'
DEVICE_JSON = '/persist/device_config.json'
# The Arcturus color module's EEPROM: some binary, then its calibration as JSON (world-readable)
ARCTURUS_EEPROM = '/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom'
ARCTURUS_WIDTH = 1972 # valid pixels per row that XRService's buffers deliver (of 2464)
def _pose(d, scale=1.0):
"""4x4 transform from a {plus_x, plus_z, position} pose (child axes in the parent frame)."""
x = np.asarray(d['plus_x'], float)
z = np.asarray(d['plus_z'], float)
y = np.cross(z, x)
T = np.eye(4)
T[:3, 0], T[:3, 1], T[:3, 2] = x, y, z
T[:3, 3] = np.asarray(d['position'], float) * scale
return T
class Camera:
def __init__(self, name, width, height, kb, head_from_cam):
self.name = name
self.width, self.height = width, height
self.fx, self.fy, self.cx, self.cy = kb['fx'], kb['fy'], kb['cx'], kb['cy']
self.k = np.array([kb['k1'], kb['k2'], kb['k3'], kb['k4']])
self.head_from_cam = head_from_cam
self.R = head_from_cam[:3, :3] # camera axes in the head frame
self.origin = head_from_cam[:3, 3] # camera centre in the head frame
def __repr__(self):
return 'Camera(%s %dx%d at %s mm)' % (self.name, self.width, self.height,
np.round(self.origin * 1000, 1))
def _theta_d(self, theta):
t2 = theta * theta
k1, k2, k3, k4 = self.k
return theta * (1 + t2 * (k1 + t2 * (k2 + t2 * (k3 + t2 * k4))))
def project_cam(self, p):
"""Camera-frame points (N,3) -> pixels (N,2). Points behind the lens still map (the lens sees ~180 deg)."""
p = np.atleast_2d(p)
r = np.hypot(p[:, 0], p[:, 1])
theta = np.arctan2(r, p[:, 2])
scale = np.where(r > 1e-12, self._theta_d(theta) / np.maximum(r, 1e-12), 0.0)
return np.stack([self.fx * p[:, 0] * scale + self.cx, self.fy * p[:, 1] * scale + self.cy], axis=1)
def unproject(self, uv):
"""Pixels (N,2) -> unit rays (N,3) in the camera frame."""
uv = np.atleast_2d(np.asarray(uv, float))
mx = (uv[:, 0] - self.cx) / self.fx
my = (uv[:, 1] - self.cy) / self.fy
td = np.hypot(mx, my)
theta = td.copy()
k1, k2, k3, k4 = self.k
for _ in range(8): # Newton on theta_d(theta) = td
t2 = theta * theta
f = self._theta_d(theta) - td
df = 1 + t2 * (3 * k1 + t2 * (5 * k2 + t2 * (7 * k3 + t2 * 9 * k4)))
theta = np.clip(theta - f / df, 0.0, np.pi)
s = np.where(td > 1e-12, np.sin(theta) / np.maximum(td, 1e-12), 1.0)
return np.stack([mx * s, my * s, np.cos(theta)], axis=1)
def rays(self, uv):
"""Pixels -> unit rays in the head frame (all starting at self.origin)."""
return self.unproject(uv) @ self.R.T
def project(self, p_head):
"""Head-frame points (N,3) -> pixels (N,2) and depth along the optical axis (N,)."""
p = (np.atleast_2d(p_head) - self.origin) @ self.R
return self.project_cam(p), p[:, 2]
def angle_from_axis(self, uv):
"""Angle in degrees between each pixel's ray and the optical axis."""
return np.degrees(np.arccos(np.clip(self.unproject(uv)[:, 2], -1, 1)))
def device_path(path):
"""A headset file such as /persist/xrservice.json. In the dev container the host's / is
at /run/host (distrobox doesn't mount /persist); off the Frame, FRAME_JOB_DEVICE_ROOT can
point at a folder with copies of them."""
root = os.environ.get('FRAME_JOB_DEVICE_ROOT')
if root:
return root + path
if not os.access(path, os.R_OK) and os.access('/run/host' + path, os.R_OK):
return '/run/host' + path
return path
def load(xrservice=XRSERVICE_JSON, device=DEVICE_JSON):
"""{calibration name: Camera} for the tracking cameras, posed in the head frame."""
with open(device_path(xrservice)) as f:
rig = json.load(f)
with open(device_path(device)) as f:
dev = json.load(f)
cad_from_cam0 = _pose(dev['cv']['cad_from_cal'])
head_from_cad = np.linalg.inv(_pose(dev['head']))
cams = {}
for c in rig['cameras']:
kb = next(i for i in c['intrinsics'] if i['cameraModel'] == 'kb')
cam0_from_cam = _pose(c['extrinsics'], 1e-3)
cams[c['sourceCamera']] = Camera(c['sourceCamera'], c['width'], c['height'], kb,
head_from_cad @ cad_from_cam0 @ cam0_from_cam)
return cams
def load_color(eeprom=ARCTURUS_EEPROM, device=DEVICE_JSON, scale=2, crop='subtract'):
"""{"passthrough_left"/"passthrough_right": Camera} for the Arcturus color cameras, posed in
the head frame, for ft-camd --with-color's images (luma at 1/scale size).
Their calibration is in the CAD frame (mm) with pixel coordinates on the full 2464x2464
sensor; each camera also has a cropRegion. crop says how that maps to the delivered
image: 'subtract' (image x = sensor x - cropRegion.x) or 'none'. tools/check_color.py
tells which fits.
"""
with open(device_path(eeprom), 'rb') as f:
raw = f.read()
i = raw.rfind(b'{', 0, raw.find(b'"alignment_method"'))
rig, _ = json.JSONDecoder().raw_decode(raw[i:].decode('latin1'))
with open(device_path(device)) as f:
dev = json.load(f)
head_from_cad = np.linalg.inv(_pose(dev['head']))
cams = {}
for c in rig['cameras']:
kb = dict(next(k for k in c['intrinsics'] if k['cameraModel'] == 'kb'))
region = c.get('cropRegion', {}) if crop == 'subtract' else {}
# integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
kb['cx'] = (kb['cx'] - region.get('x', 0) + 0.5) / scale - 0.5
kb['cy'] = (kb['cy'] - region.get('y', 0) + 0.5) / scale - 0.5
kb['fx'] /= scale
kb['fy'] /= scale
cams[c['sourceCamera']] = Camera(c['sourceCamera'], ARCTURUS_WIDTH // scale, c['height'] // scale, kb,
head_from_cad @ _pose(c['extrinsics'], 1e-3))
return cams
def triangulate(origins, dirs, weights=None):
"""Least-squares point closest to several rays. Returns (point, rms distance to the rays)."""
A = np.zeros((3, 3))
b = np.zeros(3)
w = np.ones(len(origins)) if weights is None else np.asarray(weights, float)
for o, d, wi in zip(origins, dirs, w):
P = np.eye(3) - np.outer(d, d)
A += wi * P
b += wi * P @ o
p = np.linalg.solve(A, b)
res = [np.linalg.norm((np.eye(3) - np.outer(d, d)) @ (p - o)) for o, d in zip(origins, dirs)]
return p, float(np.sqrt(np.mean(np.square(res))))
def triangulate_many(origins, dirs, weights):
"""Triangulate K points seen from V cameras at once.
origins (V,3), dirs (V,K,3) unit rays, weights (V,). Returns points (K,3) and
each point's rms distance to its rays (K,).
"""
P = np.eye(3) - dirs[..., :, None] * dirs[..., None, :] # (V,K,3,3)
w = weights[:, None, None, None]
A = (w * P).sum(0)
b = (w * (P @ origins[:, None, :, None])).sum(0)[..., 0]
pts = np.linalg.solve(A, b[..., None])[..., 0]
off = pts[None] - origins[:, None, :] # (V,K,3)
perp = off - (off * dirs).sum(-1, keepdims=True) * dirs
return pts, np.sqrt((perp ** 2).sum(-1).mean(0))
if __name__ == '__main__':
cams = load()
for cam in cams.values():
fwd = [float(v) for v in cam.R[:, 2]]
print('%-12s at x %+6.1f y %+6.1f z %+6.1f mm, looks %s' % (
cam.name, *(cam.origin * 1000),
'right' * (fwd[0] > 0.3) + 'left' * (fwd[0] < -0.3) + ' up' * (fwd[1] > 0.3) +
' down' * (fwd[1] < -0.3) + ' forward' * (fwd[2] < -0.3) + ' back' * (fwd[2] > 0.3)),
np.round(fwd, 2))
uv = np.array([[cam.cx + 200, cam.cy - 100], [cam.cx - 0.4 * cam.width, cam.cy + 0.3 * cam.height]])
err = np.abs(cam.project_cam(cam.unproject(uv)) - uv).max()
assert err < 1e-6, err
a, b = cams['slam_left'], cams['slam_right']
print('slam baseline %.2f mm' % (1000 * np.linalg.norm(a.origin - b.origin)))
+64
View File
@@ -0,0 +1,64 @@
"""Which color camera is which, and how their calibration maps onto ft-camd's images.
usage: python tools/check_color.py REC_DIR [--sets N]
A recording made with ft-camd --with-color holds color_video<N> frames with each set.
This matches features between the two color images and scores every reading of the
calibration: which video node is passthrough_left, and whether the calibration's
cropRegion is subtracted from x ('subtract') or not ('none'). Only the right reading
makes true matches' rays meet in front of both cameras. Then it checks the winner against
the side tracking cameras, which tests the CAD-to-head chain shared with them.
"""
import argparse
import itertools
import os
import sys
import numpy as np
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
from tools.check_sides import load_cams, matches, score # noqa: E402
from tools.show_set import index, read_set # noqa: E402
from tools import calib # noqa: E402
def load_color(crop):
return calib.load_color(crop=crop)
def main():
ap = argparse.ArgumentParser()
ap.add_argument('rec')
ap.add_argument('--sets', type=int, default=8)
a = ap.parse_args()
path = os.path.join(a.rec, 'sets.bin')
offs = index(path)
sets = [read_set(path, offs[n]) for n in np.linspace(0, len(offs) - 1, a.sets).astype(int)]
nodes = sorted(k for k in sets[0] if k.startswith('color_video'))
if len(nodes) != 2:
sys.exit('need two color_video<N> cameras in the recording (ft-camd --with-color); found %s' % nodes)
pairs = [matches(s[nodes[0]][0], s[nodes[1]][0]) for s in sets]
print('%d sets, %d matches between %s and %s' % (len(sets), sum(len(p[0]) for p in pairs), *nodes))
best = None
for crop, left in itertools.product(['subtract', 'none'], nodes):
cams = load_color(crop)
right = nodes[1] if left == nodes[0] else nodes[0]
cam = {left: cams['passthrough_left'], right: cams['passthrough_right']}
s = np.mean([score(cam[nodes[0]], cam[nodes[1]], ua, ub) for ua, ub in pairs])
print(' %s = passthrough_left, crop %-8s: %3.0f%% of matches meet' % (left, crop, 100 * s))
if best is None or s > best[0]:
best = (s, crop, left, cam)
s, crop, left, cam = best
print('best: %s = passthrough_left, crop %s (%.0f%%)' % (left, crop, 100 * s))
mono = load_cams()
for node in nodes:
for side in ['slam_left', 'slam_right']:
ms = [matches(st[node][0], st[side][0]) for st in sets if side in st]
sc = np.mean([score(cam[node], mono[side], ua, ub) for ua, ub in ms]) if ms else 0
print(' %s vs %-10s: %4d matches, %3.0f%% meet' % (node, side, sum(len(m[0]) for m in ms), 100 * sc))
if __name__ == '__main__':
main()
+128
View File
@@ -0,0 +1,128 @@
"""Check that the side cameras' images carry the right names (slam_left vs slam_right).
usage: python tools/check_sides.py REC_DIR [--sets N]
python tools/check_sides.py --ring [--sets N] (live, from ft-camd's ring)
With --ring it exits 0 when the names are right, 3 when they're swapped (run ft-hands
with --swap-sides), and 2 when it can't tell (too little texture in view, or the headset
isn't worn).
ft-camd tells the two side cameras' buffers apart by the order XRService allocated them,
and after some XRService restarts that order puts each camera's images under the other's
name. The tracker then sees every hand in one camera only, at the wrong depth. This
matches features between the two images and measures how close each pair's rays pass
with the factory calibration, once as named and once swapped: true matches meet in
front of both cameras only under the right naming.
"""
import argparse
import os
import sys
import cv2
import numpy as np
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
from tools.show_set import index, read_set # noqa: E402
from tools import calib # noqa: E402
PIPES = {'msm_vfe3_video0': 'slam_left', 'msm_vfe4_video0': 'slam_right'} # as ft-hands maps them
def load_cams():
return calib.load()
def matches(a, b):
"""Pixel pairs (N,2), (N,2) of ORB matches between two grey images."""
clahe = cv2.createCLAHE(2.0, (8, 8))
orb = cv2.ORB_create(3000)
ka, da = orb.detectAndCompute(clahe.apply(a), None)
kb, db = orb.detectAndCompute(clahe.apply(b), None)
if da is None or db is None:
return np.zeros((0, 2)), np.zeros((0, 2))
pairs = cv2.BFMatcher(cv2.NORM_HAMMING).knnMatch(da, db, k=2)
good = [p[0] for p in pairs if len(p) == 2 and p[0].distance < 0.75 * p[1].distance]
return (np.array([ka[m.queryIdx].pt for m in good]).reshape(-1, 2),
np.array([kb[m.trainIdx].pt for m in good]).reshape(-1, 2))
def meet(cam_a, cam_b, ua, ub):
"""Per match: closest distance between the two rays (m), and whether they meet in front of both."""
ra, rb = cam_a.rays(ua), cam_b.rays(ub)
w = cam_b.origin - cam_a.origin
n = np.cross(ra, rb)
nn = np.linalg.norm(n, axis=1)
dist = np.abs(w @ n.T) / np.maximum(nn, 1e-12)
# ray parameters at the closest points
ta = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, rb.shape), rb), n) / np.maximum(nn ** 2, 1e-12)
tb = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, ra.shape), ra), n) / np.maximum(nn ** 2, 1e-12)
return dist, (ta > 0.05) & (tb > 0.05)
def score(cam_a, cam_b, ua, ub):
"""Share of matches whose rays meet within 1 cm, in front of both cameras."""
if len(ua) == 0:
return 0.0
d, front = meet(cam_a, cam_b, ua, ub)
return float(np.mean((d < 0.01) & front))
def recorded_pairs(rec, count):
"""(label, slam_left image, slam_right image) from sets spread across a recording."""
path = os.path.join(rec, 'sets.bin')
offs = index(path)
for n in np.linspace(0, len(offs) - 1, count).astype(int):
images = read_set(path, offs[n])
if 'slam_left' in images and 'slam_right' in images:
yield 'set %5d' % n, images['slam_left'][0], images['slam_right'][0]
def live_pairs(count):
"""(label, slam_left image, slam_right image) from ft-camd's ring, half a second apart."""
import time
from tools.ring import Ring
ring = Ring()
if not ring.alive():
sys.exit('ft-camd isn\'t running (no heartbeat)')
cams = {}
for c in ring.cams:
name = PIPES.get(open('/sys/class/video4linux/video%d/name' % c.node).read().strip())
if name and not c.name.endswith('-dark'):
cams[name] = c
for k in range(count):
a, b = ring.read(cams['slam_left']), ring.read(cams['slam_right'])
if a is not None and b is not None:
yield 'frame %2d' % k, a.image, b.image
time.sleep(0.5)
def main():
ap = argparse.ArgumentParser()
ap.add_argument('rec', nargs='?')
ap.add_argument('--ring', action='store_true', help='check the live cameras instead of a recording')
ap.add_argument('--sets', type=int, default=8, help='how many sets or live frames to check')
a = ap.parse_args()
if not a.ring and not a.rec:
ap.error('give a recording or --ring')
cams = load_cams()
left, right = cams['slam_left'], cams['slam_right']
named = swapped = 0.0
n = total_matches = 0
for label, img_l, img_r in (live_pairs(a.sets) if a.ring else recorded_pairs(a.rec, a.sets)):
ua, ub = matches(img_l, img_r)
s_named = score(left, right, ua, ub) # slam_left's image seen by the left camera
s_swapped = score(right, left, ua, ub) # ... by the right camera
named, swapped, n, total_matches = named + s_named, swapped + s_swapped, n + 1, total_matches + len(ua)
print('%s: %4d matches, meeting as named %3.0f%%, swapped %3.0f%%' %
(label, len(ua), 100 * s_named, 100 * s_swapped))
if n == 0 or total_matches < 100 or abs(named - swapped) / n < 0.2:
print('side cameras: can\'t tell (%d matches)' % total_matches)
sys.exit(2)
print('side cameras: %s (named %.2f, swapped %.2f)' %
('as named' if named > swapped else 'SWAPPED', named / n, swapped / n))
sys.exit(0 if named > swapped else 3)
if __name__ == '__main__':
main()
+81
View File
@@ -0,0 +1,81 @@
"""Convert the OpenCV Zoo ONNX ports of MediaPipe's hand models to ncnn.
The ONNX files are Apache-2.0 ports of MediaPipe's palm detector and hand
landmark models (huggingface.co/opencv/palm_detection_mediapipe and
huggingface.co/opencv/handpose_estimation_mediapipe). pnnx does the
conversion; two fix-ups follow:
- The palm detector widens channels with ONNX Pad on the channel axis. pnnx
emits an ncnn layer called "Pad", which ncnn doesn't have, so rewrite those
as ncnn Padding with the channel-end amount (param 8 = behind).
- Both models take NHWC input and start with a Permute to NCHW. Drop it, so we
can hand ncnn planar CHW Mats straight from the preprocessing step.
usage: python convert_models.py (writes models/ncnn/{palm,hand}.ncnn.{param,bin})
"""
import os
import re
import shutil
import subprocess
import sys
import tempfile
HERE = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.join(HERE, '..')
PNNX = os.path.join(sys.prefix, 'lib', 'python%d.%d' % sys.version_info[:2], 'site-packages', 'pnnx', 'pnnx')
MODELS = [('palm', 'palm_detection_mediapipe_2023feb', 192),
('hand', 'handpose_estimation_mediapipe_2023feb', 224)]
def patch(param_text, pnnx_param_text):
lines = param_text.splitlines()
assert lines[0] == '7767517'
nlayers, nblobs = map(int, lines[1].split())
body = lines[2:]
# Channel pads: amounts come from the pnnx graph, which keeps the pads tuple.
pads = dict(re.findall(r'^Pad\s+(\S+)\s.*pads=\(0,0,0,0,0,(\d+),0,0\)', pnnx_param_text, re.M))
for i, line in enumerate(body):
f = line.split()
if f[0] == 'Pad':
amount = pads[f[1]]
body[i] = 'Padding %s %s %s %s %s 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=%s' % (
f[1], f[2], f[3], f[4], f[5], amount)
# Input permute: feed its consumers from in0 instead.
perm = next(i for i, line in enumerate(body) if line.split()[0] == 'Permute')
f = body[perm].split()
assert f[4] == 'in0' and f[6] == '0=4', body[perm]
blob = f[5]
del body[perm]
for i, line in enumerate(body):
f = line.split()
if f[0] == 'Input':
continue
nin, nout = int(f[2]), int(f[3])
ins = ['in0' if b == blob else b for b in f[4:4 + nin]]
body[i] = ' '.join(f[:4] + ins + f[4 + nin:])
return '\n'.join(['7767517', '%d %d' % (nlayers - 1, nblobs - 1)] + body) + '\n'
def main():
out = os.path.join(ROOT, 'models', 'ncnn')
os.makedirs(out, exist_ok=True)
for short, name, size in MODELS:
src = os.path.join(ROOT, 'models', 'onnx', name + '.onnx')
with tempfile.TemporaryDirectory() as tmp:
shutil.copy(src, tmp)
subprocess.run([PNNX, name + '.onnx', 'inputshape=[1,%d,%d,3]' % (size, size), 'fp16=1'],
cwd=tmp, check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
with open(os.path.join(tmp, name + '.ncnn.param')) as f:
param = f.read()
with open(os.path.join(tmp, name + '.pnnx.param')) as f:
pparam = f.read()
with open(os.path.join(out, short + '.ncnn.param'), 'w') as f:
f.write(patch(param, pparam))
shutil.copy(os.path.join(tmp, name + '.ncnn.bin'), os.path.join(out, short + '.ncnn.bin'))
print('wrote', short)
if __name__ == '__main__':
main()
+51
View File
@@ -0,0 +1,51 @@
"""Copy a few frame sets out of a recording (ft-hands --record) into a small one, to look at
or check elsewhere without moving gigabytes. Plain Python, so it runs on the Frame's host.
usage: python3 tools/cut_sets.py REC_DIR OUT_DIR [--sets N] (8, spread evenly) | [--at I,J,...]
"""
import argparse
import os
import struct
HDR = struct.Struct('<8sII')
def offsets(path):
offs, size = [], os.path.getsize(path)
with open(path, 'rb') as f:
off = 0
while off + HDR.size <= size:
f.seek(off)
magic, _, nbytes = HDR.unpack(f.read(HDR.size))
if magic[:7] != b'FHSET01' or off + nbytes > size:
break
offs.append((off, nbytes))
off += nbytes
return offs
def main():
ap = argparse.ArgumentParser()
ap.add_argument('rec')
ap.add_argument('out')
ap.add_argument('--sets', type=int, default=8)
ap.add_argument('--at', default='')
a = ap.parse_args()
src = os.path.join(a.rec, 'sets.bin')
offs = offsets(src)
if a.at:
pick = [int(i) for i in a.at.split(',')]
else:
n = max(1, min(a.sets, len(offs)))
pick = [round(i * (len(offs) - 1) / max(n - 1, 1)) for i in range(n)]
os.makedirs(a.out, exist_ok=True)
with open(src, 'rb') as f, open(os.path.join(a.out, 'sets.bin'), 'wb') as out:
for i in pick:
off, nbytes = offs[i]
f.seek(off)
out.write(f.read(nbytes))
print('%d of %d sets (%s) -> %s' % (len(pick), len(offs), ','.join(map(str, pick)), a.out))
if __name__ == '__main__':
main()
+256
View File
@@ -0,0 +1,256 @@
"""How good the tracker's depth is, from a replay's depth dump, without ground truth.
usage: python3 tools/depth_report.py DEPTH [DEPTH...] [--still M/S]
DEPTH comes from `hands/build/ft-handreplay DIR --depth DEPTH`. Every measure is split by how the
hand was seen: by the two lower cameras ("lower pair"), by a lower and an upper camera on
one side ("lower+upper"), or by one camera. Distances are from the head (between the eyes).
1. How the hands were seen: the share of hand updates in each way, by distance.
2. Noise along the line of sight against across it. Each update's palm is compared with a
straight line through the two updates before it (ft-handreplay's jitter measure), and the
miss is split along the line from the hand's cameras to the palm and across it. Given
as a robust sigma per axis, measured (as triangulated) and published (after the One Euro
filter), on updates where the published palm moved slower than --still (default 0.15
m/s), so the miss is mostly noise and not the hand speeding up. For two cameras,
geometry predicts along/across = 2 Z / B: Z the distance, B the cameras' baseline
across the line of sight.
3. One-camera distance: on two-camera updates, each camera's one-view guess (distance from
how big the palm looks, at the user's learned hand size) against the triangulated
distance from that camera.
4. A camera lost: from two-camera updates, what the tracker would have had if one of the
two cameras dropped out there. It keeps the last distance and moves a share of the way
to the one-view guess each update (0.1 now, kMonoDepthGain in track/tracker.cpp);
also shown with other shares, 0 (keep the distance) and 1 (take each guess), and with
the guess first scaled by how far off it was while both cameras saw the hand. Compared with the
triangulated distance from that camera, 0.1-2 s after the loss.
"""
import argparse
import collections
import numpy as np
BINS = [0.0, 0.35, 0.50, 0.65, 9.0]
BIN_NAMES = ['<35 cm', '35-50', '50-65', '65+ cm']
MODES = ['lower pair', 'lower+upper', 'one camera']
HORIZONS = [0.1, 0.25, 0.5, 1.0, 2.0]
# (share of the way toward the one-view guess per update, whether the guess is first scaled by
# how far off it was while both cameras saw the hand)
GAINS = [(0.0, False), (0.02, False), (0.05, False), (0.1, False), (1.0, False), (0.1, True), (1.0, True)]
GAP = 0.1 # s: a longer gap between a hand's updates breaks its run
def load(path):
cams, rows = {}, []
with open(path) as f:
for line in f:
w = line.split()
if not w:
continue
if w[0] == '#':
if w[1] == 'cam':
cams[w[2]] = (np.array([float(x) for x in w[3:6]]), float(w[6]))
continue
r = {'t': float(w[0]), 'id': int(w[1]), 'side': w[2], 'n': int(w[3]),
'cams': w[4], 'res': float(w[5]), 'scale': float(w[6]),
'raw': np.array([float(x) for x in w[7:10]]), 'sm': np.array([float(x) for x in w[10:13]]),
'views': {}}
for k in range(13, len(w), 5):
r['views'][w[k]] = np.array([float(x) for x in w[k + 2:k + 5]])
rows.append(r)
return cams, rows
def mode(r):
names = r['cams'].split('+')
if r['n'] == 1:
return 'one camera'
if sorted(names) == ['slam_left', 'slam_right']:
return 'lower pair'
if len(names) == 2 and all(n.startswith(('slam_', 'upper_')) for n in names) and \
names[0].split('_')[1] == names[1].split('_')[1]:
return 'lower+upper'
return 'other'
def dist_bin(p):
return min(np.searchsorted(BINS, np.linalg.norm(p), side='right') - 1, len(BIN_NAMES) - 1)
def tracks(rows):
"""A hand's updates, in order, split where they're more than GAP apart."""
by_id = collections.defaultdict(list)
for r in rows:
by_id[r['id']].append(r)
for rs in by_id.values():
run = [rs[0]]
for r in rs[1:]:
if r['t'] - run[-1]['t'] >= GAP:
yield run
run = []
run.append(r)
yield run
def two_camera_runs(rows):
"""Stretches of a hand's updates all seen by the same two cameras."""
for run in tracks(rows):
seg = []
for r in run:
ok = r['n'] == 2 and mode(r) != 'other'
if ok and seg and r['cams'] == seg[-1]['cams']:
seg.append(r)
continue
if len(seg) > 2:
yield seg
seg = [r] if ok else []
if len(seg) > 2:
yield seg
def pct(v, q):
return np.percentile(v, q) if len(v) else float('nan')
def seen_share(rows):
print('\n1. How the hands were seen (share of hand updates)')
count = collections.Counter((mode(r), dist_bin(r['raw'])) for r in rows)
total = collections.Counter(dist_bin(r['raw']) for r in rows)
print('%-12s' % '' + ''.join('%10s' % b for b in BIN_NAMES) + '%10s' % 'all')
for m in MODES + ['other']:
cells = [100 * count[m, b] / max(total[b], 1) for b in range(len(BIN_NAMES))]
allp = 100 * sum(count[m, b] for b in range(len(BIN_NAMES))) / max(len(rows), 1)
print('%-12s' % m + ''.join('%9.0f%%' % c for c in cells) + '%9.0f%%' % allp)
print('%-12s' % 'updates' + ''.join('%10d' % total[b] for b in range(len(BIN_NAMES))) + '%10d' % len(rows))
res = collections.defaultdict(list)
for r in rows:
if r['res'] >= 0:
res[mode(r)].append(r['res'] * 1000)
print('triangulation residual (rms ray miss, median): ' +
', '.join('%s %.1f mm' % (m, np.median(v)) for m, v in res.items()))
def noise(rows, cams, still):
print('\n2. Noise along the line of sight vs across it (sigma per axis, mm; palm slower than %.2f m/s)' % still)
acc = collections.defaultdict(lambda: collections.defaultdict(list))
for run in tracks(rows):
for a, b, c in zip(run, run[1:], run[2:]):
if not (mode(a) == mode(b) == mode(c)) or a['cams'] != b['cams'] or b['cams'] != c['cams']:
continue
dt0, dt1 = b['t'] - a['t'], c['t'] - b['t']
if dt0 < 1e-3 or np.linalg.norm(b['sm'] - a['sm']) / dt0 > still:
continue
names = b['cams'].split('+')
origins = [cams[n][0] for n in names]
o = np.mean(origins, axis=0)
u = b['raw'] - o
z = np.linalg.norm(u)
u /= z
key = (mode(b), dist_bin(b['raw']))
for kind in ('raw', 'sm'):
miss = c[kind] - b[kind] - (b[kind] - a[kind]) * (dt1 / dt0)
along = miss @ u
acc[key][kind + '_along'].append(abs(along))
acc[key][kind + '_across'].append(np.linalg.norm(miss - along * u))
if len(origins) == 2:
base = origins[0] - origins[1]
acc[key]['pred'].append(2 * z / np.linalg.norm(base - (base @ u) * u))
# |along| is half-normal: sigma = median / 0.674; |across| is Rayleigh (2 axes): sigma = median / 1.177
print('%-12s %-7s %6s | %-24s | %-24s | %s' % ('', '', 'n', 'measured along/across', 'published along/across',
'ratio measured (geometry)'))
for m in MODES:
for bi, bn in enumerate(BIN_NAMES):
d = acc.get((m, bi))
if not d or len(d['raw_along']) < 20:
continue
s = {k: np.median(v) / (0.674 if k.endswith('along') else 1.177) * 1000
for k, v in d.items() if k != 'pred'}
pred = '(%.1f)' % np.median(d['pred']) if d['pred'] else ''
print('%-12s %-7s %6d | %7.1f / %-5.1f x%-6.1f | %7.1f / %-5.1f x%-6.1f | x%.1f %s' % (
m, bn, len(d['raw_along']), s['raw_along'], s['raw_across'], s['raw_along'] / s['raw_across'],
s['sm_along'], s['sm_across'], s['sm_along'] / s['sm_across'],
s['raw_along'] / s['raw_across'], pred))
def one_camera(rows, cams):
print('\n3. One-camera distance vs triangulated, on two-camera updates (error of the one-view guess)')
acc = collections.defaultdict(list)
for r in rows:
if r['n'] != 2 or mode(r) == 'other':
continue
for name, p in r['views'].items():
if np.isnan(p).any():
continue
o = cams[name][0]
truth = np.linalg.norm(r['raw'] - o)
acc[name.split('_')[0], dist_bin(r['raw'])].append((np.linalg.norm(p - o) - truth, truth))
print('%-8s %-7s %6s %12s %12s %14s %12s' % ('camera', '', 'n', 'median |err|', '90% |err|', 'median |err| %',
'bias'))
for cam in ('slam', 'upper'):
for bi, bn in enumerate(BIN_NAMES):
v = acc.get((cam, bi))
if not v or len(v) < 20:
continue
e = np.array([x[0] for x in v])
rel = e / np.array([x[1] for x in v])
print('%-8s %-7s %6d %9.0f mm %9.0f mm %13.0f%% %+11.0f%%' % (
'lower' if cam == 'slam' else 'upper', bn, len(v), 1000 * np.median(abs(e)), 1000 * pct(abs(e), 90),
100 * np.median(abs(rel)), 100 * np.median(rel)))
def lost_camera(rows, cams):
print('\n4. A camera lost: distance error after the loss (median |err| mm / 90% mm), by how the tracker '
'moves toward the one-view guess each update ("scaled": the guess times how far off it was, '
'triangulated / guess, median over the last 30 two-camera updates)')
errs = collections.defaultdict(list)
for run in two_camera_runs(rows):
for s in range(1, len(run) - 1, 3):
for name in run[s]['views']:
o = cams[name][0]
guess = lambda r: np.linalg.norm(r['views'][name] - o)
ratios = [np.linalg.norm(r['raw'] - o) / guess(r) for r in run[max(0, s - 30):s]
if not np.isnan(r['views'][name]).any()]
ratio = np.median(ratios) if ratios else 1.0
for g, scaled in GAINS:
d = np.linalg.norm(run[s - 1]['raw'] - o)
h = 0
for r in run[s:]:
if np.isnan(r['views'][name]).any():
break
d += g * (guess(r) * (ratio if scaled else 1.0) - d)
elapsed = r['t'] - run[s - 1]['t']
while h < len(HORIZONS) and elapsed >= HORIZONS[h]:
errs[g, scaled, HORIZONS[h], name.split('_')[0]].append(abs(d - np.linalg.norm(r['raw'] - o)))
h += 1
print('%-8s %-18s' % ('camera', 'toward guess') + ''.join('%14s' % ('%.2g s' % t) for t in HORIZONS))
for cam in ('slam', 'upper'):
for g, scaled in GAINS:
label = {0.0: '0 (keep)', 0.1: '0.1 (now)', 1.0: '1 (guess)'}.get(g, '%g' % g)
if scaled:
label = '%g scaled' % g
cells = []
for t in HORIZONS:
v = errs.get((g, scaled, t, cam), [])
cells.append('%5.0f / %-4.0f' % (1000 * np.median(v), 1000 * pct(v, 90)) if len(v) >= 20 else '%14s' % '-')
print('%-8s %-18s' % ('lower' if cam == 'slam' else 'upper', label) + ''.join('%14s' % c for c in cells))
n = sum(len(errs.get((0.1, False, HORIZONS[0], c), [])) for c in ('slam', 'upper'))
print('(%d simulated losses)' % n)
def main():
ap = argparse.ArgumentParser()
ap.add_argument('depth', nargs='+')
ap.add_argument('--still', type=float, default=0.15, help='m/s: palm speed limit for the noise measure')
a = ap.parse_args()
for path in a.depth:
cams, rows = load(path)
print('== %s: %d hand updates, %d hands' % (path, len(rows), len({r['id'] for r in rows})))
seen_share(rows)
noise(rows, cams, a.still)
one_camera(rows, cams)
lost_camera(rows, cams)
print()
if __name__ == '__main__':
main()
+80
View File
@@ -0,0 +1,80 @@
"""Read frames from ft-camd's shared-memory ring (layout: camd/fhring.h)."""
import mmap
import os
import struct
import time
import numpy as np
RING_FILE = '/run/user/%d/frametop-hands/cam-ring' % os.getuid()
MAGIC = b'FHRING01'
HDR = struct.Struct('<8sIIIIQqQ16x') # 64 bytes
CAM = struct.Struct('<32s32siIIIIIQQQQQ32x') # 160 bytes
SLOT = struct.Struct('<QQQQQIf16x') # 64 bytes
MAX_CAMS = 8
LATEST_OFF = 32 + 32 + 4 * 6 + 8 * 2 # cam.latest within fh_ring_cam_t
HEARTBEAT_OFF = 40
class Frame:
__slots__ = ('cam', 'frame', 'capture_ns', 'dqbuf_ns', 'publish_ns', 'v4l2_seq', 'mean', 'image')
def __init__(self, cam, fields, image):
self.cam = cam
(_, self.frame, self.capture_ns, self.dqbuf_ns, self.publish_ns, self.v4l2_seq, self.mean) = fields
self.image = image
class RingCamera:
def __init__(self, index, fields):
(sensor, name, self.node, self.format, self.width, self.height, self.stride, self.nslots,
self.slot_offset, self.slot_bytes, _latest, _pub, _drop) = fields
self.index = index
self.sensor = sensor.split(b'\0', 1)[0].decode()
self.name = name.split(b'\0', 1)[0].decode()
self.latest_off = HDR.size + index * CAM.size + LATEST_OFF
def __repr__(self):
return 'RingCamera(video%d %s %dx%d)' % (self.node, self.sensor, self.width, self.height)
class Ring:
def __init__(self, path=RING_FILE):
fd = os.open(path, os.O_RDONLY)
try:
self.map = mmap.mmap(fd, 0, mmap.MAP_SHARED, mmap.PROT_READ)
finally:
os.close(fd)
magic, version, hdr_bytes, ncams, _, file_bytes, self.writer_pid, _ = HDR.unpack_from(self.map, 0)
if magic != MAGIC or version != 1:
raise RuntimeError('%s is not an ft-camd ring (magic %r version %d)' % (path, magic, version))
self.cams = [RingCamera(i, CAM.unpack_from(self.map, HDR.size + i * CAM.size)) for i in range(ncams)]
def heartbeat_ns(self):
return struct.unpack_from('<Q', self.map, HEARTBEAT_OFF)[0]
def alive(self, max_age=1.0):
hb = self.heartbeat_ns()
return hb != 0 and (time.clock_gettime_ns(time.CLOCK_MONOTONIC) - hb) / 1e9 < max_age
def latest(self, cam):
return struct.unpack_from('<Q', self.map, cam.latest_off)[0]
def read(self, cam, n=None):
"""Copy frame n (default: the newest) of a camera, or None if it's gone or being written."""
for _ in range(3):
if n is None or n == 0:
n = self.latest(cam)
if n == 0:
return None
off = cam.slot_offset + (n % cam.nslots) * cam.slot_bytes
fields = SLOT.unpack_from(self.map, off)
if fields[0] != 2 * n + 2:
return None
start = off + SLOT.size
image = np.frombuffer(self.map, np.uint8, cam.stride * cam.height, start).reshape(cam.height, cam.stride)
image = image[:, :cam.width].copy()
if struct.unpack_from('<Q', self.map, off)[0] == fields[0]:
return Frame(cam, fields, image)
n = None # overwritten while copying: take the newest
return None
+129
View File
@@ -0,0 +1,129 @@
"""Draw frame sets from a recording (ft-hands --record) with what the tracker saw.
usage: python tools/show_set.py REC_DIR SET [SET...] [--timeline TL] [--out DIR]
SET is a set index (ft-handreplay's timeline gives them). With --timeline (ft-handreplay
--timeline), each camera shows the tracker's views at that set: the crop for the next
frame, labelled with the hand and presence. Recordings made with ft-camd --with-dark get
a second row: each camera's latest dark frame (<name>_dk), stretched to be visible and
labelled with its mean brightness. Recordings made with ft-camd --with-color get a row of
the color cameras (color_video<N>). Writes OUT/set_<n>.jpg (default /tmp).
"""
import argparse
import os
import struct
import cv2
import numpy as np
HDR = struct.Struct('<8sII')
CAM = struct.Struct('<16sIIQQ')
ORDER = ['slam_left', 'slam_right', 'upper_left', 'upper_right']
def index(path):
"""Byte offset of every set in sets.bin."""
offs, size = [], os.path.getsize(path)
with open(path, 'rb') as f:
off = 0
while off + HDR.size <= size:
f.seek(off)
magic, n, nbytes = HDR.unpack(f.read(HDR.size))
if magic[:7] != b'FHSET01' or off + nbytes > size:
break
offs.append(off)
off += nbytes
return offs
def read_set(path, off):
with open(path, 'rb') as f:
f.seek(off)
_, n, _ = HDR.unpack(f.read(HDR.size))
cams = [CAM.unpack(f.read(CAM.size)) for _ in range(n)]
out = {}
for name, w, h, cap, dq in cams:
px = np.frombuffer(f.read(w * h), np.uint8).reshape(h, w)
out[name.rstrip(b'\0').decode()] = (px, cap)
return out
def views_at(timeline, n):
out = []
for line in open(timeline):
f = line.split()
if len(f) > 1 and f[1] == 'view' and int(f[-1]) == n:
out.append({'hand': int(f[2]), 'cam': f[3], 'presence': float(f[5]),
'c': (float(f[7]), float(f[8])), 'size': float(f[9]), 'rot': float(f[10])})
return out
def dark_tile(frame, shape, name):
"""A dark frame, stretched from its 1st to 99.5th percentile; black if there's none."""
h, w = shape
if frame is None:
return np.zeros((h, w, 3), np.uint8)
px = frame[0]
lo, hi = np.percentile(px, (1, 99.5))
gain = 255 / max(hi - lo, 1)
img = np.clip((px.astype(np.float32) - lo) * gain, 0, 255).astype(np.uint8)
img = cv2.cvtColor(cv2.resize(img, (w, h)), cv2.COLOR_GRAY2BGR)
cv2.putText(img, '%s_dk mean %.1f, x%.0f' % (name, px.mean(), gain), (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0,
(255, 255, 0), 2)
return img
def view_tile(px, name, views):
"""A frame, CLAHE'd, with the tracker's views on it, 512 px high."""
img = cv2.cvtColor(cv2.createCLAHE(2.0, (8, 8)).apply(px), cv2.COLOR_GRAY2BGR)
for v in views:
if v['cam'] != name:
continue
c, s, r = v['c'], v['size'], v['rot']
box = cv2.boxPoints(((c[0], c[1]), (s, s), np.degrees(r)))
col = (0, 255, 0) if v['presence'] >= 0.5 else (0, 0, 255)
cv2.polylines(img, [box.astype(np.int32)], True, col, 2)
cv2.putText(img, 'h%d %.2f' % (v['hand'], v['presence']), (int(c[0] - s / 2), int(c[1] - s / 2) - 6),
cv2.FONT_HERSHEY_SIMPLEX, 0.8, col, 2)
cv2.putText(img, name, (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0, (255, 255, 0), 2)
scale = 512 / img.shape[0]
return cv2.resize(img, (int(img.shape[1] * scale), 512))
def draw(images, views):
"""Rows: the mono cameras; their dark frames, if recorded; the color cameras, if recorded."""
tiles, dark = [], []
for name in ORDER:
if name not in images:
continue
tiles.append(view_tile(images[name][0], name, views))
dark.append(dark_tile(images.get(name + '_dk'), tiles[-1].shape[:2], name))
rows = [np.hstack(tiles)]
if any(k.endswith('_dk') for k in images):
rows.append(np.hstack(dark))
color = sorted(k for k in images if k.startswith('color_'))
if color:
rows.append(np.hstack([view_tile(images[k][0], k, views) for k in color]))
width = max(r.shape[1] for r in rows)
return np.vstack([np.pad(r, ((0, 0), (0, width - r.shape[1]), (0, 0))) for r in rows])
def main():
ap = argparse.ArgumentParser()
ap.add_argument('rec')
ap.add_argument('sets', type=int, nargs='+')
ap.add_argument('--timeline')
ap.add_argument('--out', default='/tmp')
a = ap.parse_args()
path = os.path.join(a.rec, 'sets.bin')
offs = index(path)
for n in a.sets:
images = read_set(path, offs[n])
views = views_at(a.timeline, n) if a.timeline else []
out = os.path.join(a.out, 'set_%05d.jpg' % n)
cv2.imwrite(out, draw(images, views), [cv2.IMWRITE_JPEG_QUALITY, 85])
print(out)
if __name__ == '__main__':
main()
+95
View File
@@ -0,0 +1,95 @@
"""Watch the gestures ft-hands publishes, live: pinches and grips, begins, ends, and drags.
usage: python3 tools/watch_gestures.py [--every S] [--distance]
Prints a line when a pinch or a grip (a closed hand) begins or ends on either hand. It goes
by the counters, so a quick tap between two reads still shows. While one is held, every
--every seconds (default 0.1) it prints how far its point has moved since it began, in the
head frame (turning your head moves it too; a real consumer turns both points into the room
first, see include/fh_gestures.h). --distance also prints each hand's thumb-to-index
distance and finger curl, to see how close a gesture comes to the thresholds.
Version 1 files (pinches only) still work.
"""
import argparse
import mmap
import os
import struct
import time
HDR = struct.Struct('<8sIIQQQffff8x') # 64 bytes
SLOT = struct.Struct('<IIIIQQff3f3f') # 64 bytes
TRACKED, DOWN, LOST = 1, 2, 4
SIDES = ('left ', 'right')
KINDS = ('pinch', 'grip ')
def path():
return '/run/user/%d/frametop-hands/gestures' % os.getuid()
def read(m):
"""(header, [[pinch left, right], [grip left, right]]) under the sequence lock, or None
if it's being written. Version 1 has no grips: they read as all zero."""
for _ in range(10):
s1 = struct.unpack_from('<Q', m, 16)[0]
if s1 % 2 == 0:
h = HDR.unpack_from(m, 0)
kinds = 2 if h[1] >= 2 and len(m) >= HDR.size + 4 * SLOT.size else 1
g = [[SLOT.unpack_from(m, HDR.size + (k * 2 + s) * SLOT.size) for s in range(2)] for k in range(kinds)]
if kinds == 1:
g.append([(0,) * 14, (0,) * 14])
if struct.unpack_from('<Q', m, 16)[0] == s1:
return h, g
time.sleep(0.0005)
return None
def main():
ap = argparse.ArgumentParser()
ap.add_argument('--every', type=float, default=0.1, help='seconds between drag lines while held')
ap.add_argument('--distance', action='store_true', help="print each hand's pinch distance and finger curl")
a = ap.parse_args()
with open(path(), 'rb') as f:
m = mmap.mmap(f.fileno(), 0, prot=mmap.PROT_READ)
first = read(m)
if first is None or first[0][0] != b'FHGEST01':
raise SystemExit('%s is not an ft-hands gestures file' % path())
h, g = first
print('version %d; thresholds: pinch begins under %.3f m, ends over %.3f m; grip begins with every finger '
'curled under %.2f, ends over %.2f' % (h[1], h[6], h[7], h[8], h[9]))
seen = [[(q[2], q[3]) for q in kind] for kind in g] # begins, ends
last_drag = last_dist = 0.0
while True:
got = read(m)
if got:
h, g = got
now = time.monotonic()
for k, kind in enumerate(g):
for s, q in enumerate(kind):
flags, hand, begins, ends, begin_ns, end_ns, dist, strength = q[:8]
point, begin_point = q[8:11], q[11:14]
if begins != seen[k][s][0]:
print('%s %s BEGIN (#%d, hand %d) at %+.3f %+.3f %+.3f %s %.3f' %
(SIDES[s], KINDS[k], begins, hand, *begin_point, 'curl' if k else 'd', dist), flush=True)
if ends != seen[k][s][1]:
held = (end_ns - begin_ns) / 1e9 if end_ns >= begin_ns else 0
print('%s %s %s after %.2f s' % (SIDES[s], KINDS[k], 'LOST' if flags & LOST else 'END', held),
flush=True)
seen[k][s] = (begins, ends)
if flags & DOWN and now - last_drag >= a.every:
d = [point[i] - begin_point[i] for i in range(3)]
print('%s %s drag %+6.1f %+6.1f %+6.1f mm (%.0f mm)' %
(SIDES[s], KINDS[k], *(1000 * x for x in d), 1000 * sum(x * x for x in d) ** 0.5),
flush=True)
if any(q[0] & DOWN for kind in g for q in kind) and now - last_drag >= a.every:
last_drag = now
if a.distance and now - last_dist >= 0.2:
last_dist = now
print(' ' + ' '.join(
'%s %s' % (SIDES[s].strip(), 'd %.3f curl %.2f' % (g[0][s][6], g[1][s][6]) if g[0][s][0] & TRACKED
else '-') for s in range(2)), flush=True)
time.sleep(0.005)
if __name__ == '__main__':
main()
+221
View File
@@ -0,0 +1,221 @@
#include <cstdlib>
#include "calib.h"
#include <json/json.h>
#include <unistd.h>
#include <algorithm>
#include <fstream>
#include <iterator>
#include <memory>
namespace {
double theta_d(const Camera &c, double t) {
const double t2 = t * t;
return t * (1 + t2 * (c.k[0] + t2 * (c.k[1] + t2 * (c.k[2] + t2 * c.k[3]))));
}
// 4x4 transform (row-major) from a {plus_x, plus_z, position} pose.
void pose(const Json::Value &d, double scale, double T[4][4]) {
V3 x{d["plus_x"][0].asDouble(), d["plus_x"][1].asDouble(), d["plus_x"][2].asDouble()};
V3 z{d["plus_z"][0].asDouble(), d["plus_z"][1].asDouble(), d["plus_z"][2].asDouble()};
V3 y{z[1] * x[2] - z[2] * x[1], z[2] * x[0] - z[0] * x[2], z[0] * x[1] - z[1] * x[0]};
for (int i = 0; i < 3; ++i) {
T[i][0] = x[i], T[i][1] = y[i], T[i][2] = z[i];
T[i][3] = d["position"][i].asDouble() * scale;
T[3][i] = 0;
}
T[3][3] = 1;
}
void mul(const double A[4][4], const double B[4][4], double C[4][4]) {
for (int i = 0; i < 4; ++i)
for (int j = 0; j < 4; ++j) {
C[i][j] = 0;
for (int k = 0; k < 4; ++k) C[i][j] += A[i][k] * B[k][j];
}
}
void invert_rigid(const double A[4][4], double B[4][4]) {
for (int i = 0; i < 3; ++i)
for (int j = 0; j < 3; ++j) B[i][j] = A[j][i];
for (int i = 0; i < 3; ++i) B[i][3] = -(B[i][0] * A[0][3] + B[i][1] * A[1][3] + B[i][2] * A[2][3]);
B[3][0] = B[3][1] = B[3][2] = 0, B[3][3] = 1;
}
bool read_json(const char *path, Json::Value &v, std::string &err) {
std::ifstream f(path);
Json::CharReaderBuilder b;
std::string e;
if (!f || !Json::parseFromStream(b, f, &v, &e)) {
err = std::string(path) + ": " + (f ? e : "can't open");
return false;
}
return true;
}
} // namespace
V2 Camera::project_cam(V3 p) const {
const double r = std::hypot(p[0], p[1]);
const double s = r > 1e-12 ? theta_d(*this, std::atan2(r, p[2])) / r : 0;
return {fx * p[0] * s + cx, fy * p[1] * s + cy};
}
V3 Camera::unproject(V2 uv) const {
const double mx = (uv[0] - cx) / fx, my = (uv[1] - cy) / fy, td = std::hypot(mx, my);
double t = td;
for (int i = 0; i < 8; ++i) { // Newton on theta_d(t) = td
const double t2 = t * t;
const double df = 1 + t2 * (3 * k[0] + t2 * (5 * k[1] + t2 * (7 * k[2] + t2 * 9 * k[3])));
t = std::clamp(t - (theta_d(*this, t) - td) / df, 0.0, M_PI);
}
const double s = td > 1e-12 ? std::sin(t) / td : 1;
return {mx * s, my * s, std::cos(t)};
}
V3 Camera::ray(V2 uv) const {
const V3 c = unproject(uv);
return {R[0][0] * c[0] + R[0][1] * c[1] + R[0][2] * c[2], R[1][0] * c[0] + R[1][1] * c[1] + R[1][2] * c[2],
R[2][0] * c[0] + R[2][1] * c[1] + R[2][2] * c[2]};
}
V2 Camera::project(V3 head, double *depth) const {
const V3 d = head - origin;
const V3 c{R[0][0] * d[0] + R[1][0] * d[1] + R[2][0] * d[2], R[0][1] * d[0] + R[1][1] * d[1] + R[2][1] * d[2],
R[0][2] * d[0] + R[1][2] * d[1] + R[2][2] * d[2]};
if (depth) *depth = c[2];
return project_cam(c);
}
double Camera::off_axis(V2 uv) const { return std::acos(std::clamp(unproject(uv)[2], -1.0, 1.0)) * 180 / M_PI; }
// A headset file such as /persist/xrservice.json. In the dev container the host's / is at
// /run/host (distrobox doesn't mount /persist); off the Frame, FRAME_JOB_DEVICE_ROOT can
// point at a folder with copies of them.
static std::string device_path(const char *path) {
if (const char *root = std::getenv("FRAME_JOB_DEVICE_ROOT")) return std::string(root) + path;
const std::string host = std::string("/run/host") + path;
return access(path, R_OK) != 0 && access(host.c_str(), R_OK) == 0 ? host : path;
}
bool load_calibration(std::map<std::string, Camera> &out, std::string &err) {
Json::Value rig, dev;
if (!read_json(device_path("/persist/xrservice.json").c_str(), rig, err) ||
!read_json(device_path("/persist/device_config.json").c_str(), dev, err))
return false;
double cad_from_cam0[4][4], cad_from_head[4][4], head_from_cad[4][4], head_from_cam0[4][4];
pose(dev["cv"]["cad_from_cal"], 1.0, cad_from_cam0);
pose(dev["head"], 1.0, cad_from_head);
invert_rigid(cad_from_head, head_from_cad);
mul(head_from_cad, cad_from_cam0, head_from_cam0);
for (const Json::Value &c : rig["cameras"]) {
Camera cam;
cam.name = c["sourceCamera"].asString();
cam.width = c["width"].asInt(), cam.height = c["height"].asInt();
for (const Json::Value &in : c["intrinsics"]) {
if (in["cameraModel"].asString() != "kb") continue;
cam.fx = in["fx"].asDouble(), cam.fy = in["fy"].asDouble();
cam.cx = in["cx"].asDouble(), cam.cy = in["cy"].asDouble();
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
}
double cam0_from_cam[4][4], head_from_cam[4][4];
pose(c["extrinsics"], 1e-3, cam0_from_cam);
mul(head_from_cam0, cam0_from_cam, head_from_cam);
for (int i = 0; i < 3; ++i) {
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
cam.origin[i] = head_from_cam[i][3];
}
out[cam.name] = cam;
}
if (out.empty()) err = "no cameras in /persist/xrservice.json";
return !out.empty();
}
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
const std::string &right_node, bool crop_subtract, int scale, std::string &err) {
// The module's EEPROM: some binary, then the calibration as JSON (world-readable)
const std::string path = device_path("/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom");
std::ifstream f(path, std::ios::binary);
const std::string raw((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
const size_t key = raw.find("\"alignment_method\"");
const size_t start = key == std::string::npos ? key : raw.rfind('{', key);
Json::Value rig, dev;
std::string e;
std::unique_ptr<Json::CharReader> reader(Json::CharReaderBuilder().newCharReader());
if (start == std::string::npos || !reader->parse(raw.data() + start, raw.data() + raw.size(), &rig, &e))
return err = path + ": no calibration JSON " + e, false;
if (!read_json(device_path("/persist/device_config.json").c_str(), dev, err)) return false;
double cad_from_head[4][4], head_from_cad[4][4];
pose(dev["head"], 1.0, cad_from_head);
invert_rigid(cad_from_head, head_from_cad);
constexpr int kValidWidth = 1972; // pixels per row XRService's buffers deliver (of 2464)
int n = 0;
for (const Json::Value &c : rig["cameras"]) {
const std::string source = c["sourceCamera"].asString();
const std::string name = source == "passthrough_left" ? left_node : source == "passthrough_right" ? right_node : "";
if (name.empty()) continue;
Camera cam;
cam.name = name;
cam.width = kValidWidth / scale, cam.height = c["height"].asInt() / scale;
const double dx = crop_subtract ? c["cropRegion"]["x"].asDouble() : 0, dy = crop_subtract ? c["cropRegion"]["y"].asDouble() : 0;
for (const Json::Value &in : c["intrinsics"]) {
if (in["cameraModel"].asString() != "kb") continue;
// integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
cam.fx = in["fx"].asDouble() / scale, cam.fy = in["fy"].asDouble() / scale;
cam.cx = (in["cx"].asDouble() - dx + 0.5) / scale - 0.5, cam.cy = (in["cy"].asDouble() - dy + 0.5) / scale - 0.5;
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
}
double cad_from_cam[4][4], head_from_cam[4][4];
pose(c["extrinsics"], 1e-3, cad_from_cam);
mul(head_from_cad, cad_from_cam, head_from_cam);
for (int i = 0; i < 3; ++i) {
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
cam.origin[i] = head_from_cam[i][3];
}
out[name] = cam;
++n;
}
if (n != 2) err = path + ": expected passthrough_left and passthrough_right";
return n == 2;
}
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms) {
double A[3][3] = {}, b[3] = {};
for (int v = 0; v < n; ++v) {
const V3 &d = dirs[v], &o = origins[v];
for (int i = 0; i < 3; ++i)
for (int j = 0; j < 3; ++j) {
const double P = (i == j ? 1.0 : 0.0) - d[i] * d[j];
A[i][j] += weights[v] * P;
b[i] += weights[v] * P * o[j];
}
}
// Cramer's rule for the 3x3 system
auto det3 = [](const double m[3][3]) {
return m[0][0] * (m[1][1] * m[2][2] - m[1][2] * m[2][1]) - m[0][1] * (m[1][0] * m[2][2] - m[1][2] * m[2][0]) +
m[0][2] * (m[1][0] * m[2][1] - m[1][1] * m[2][0]);
};
const double D = det3(A);
V3 p{};
for (int c = 0; c < 3; ++c) {
double M[3][3];
for (int i = 0; i < 3; ++i)
for (int j = 0; j < 3; ++j) M[i][j] = j == c ? b[i] : A[i][j];
p[c] = std::fabs(D) > 1e-18 ? det3(M) / D : 0;
}
if (rms) {
double s = 0;
for (int v = 0; v < n; ++v) {
const V3 off = p - origins[v];
const V3 perp = off - dirs[v] * dot(off, dirs[v]);
s += dot(perp, perp);
}
*rms = std::sqrt(s / n);
}
return p;
}
+39
View File
@@ -0,0 +1,39 @@
// Tracking-camera calibration from the headset's factory files (see tools/calib.py
// for the conventions): Kannala-Brandt fisheye intrinsics, and each camera's pose in the
// head frame (OpenVR's: +x right, +y up, -z forward), metres.
#pragma once
#include "geom.h"
#include <map>
#include <string>
#include <vector>
struct Camera {
std::string name;
int width = 0, height = 0;
double fx = 1, fy = 1, cx = 0, cy = 0, k[4] = {};
double R[3][3] = {}; // camera axes (columns) in the head frame
V3 origin{}; // camera centre in the head frame
V2 project_cam(V3 p) const; // camera frame -> pixels
V3 unproject(V2 uv) const; // pixels -> unit ray, camera frame
V3 ray(V2 uv) const; // pixels -> unit ray, head frame
V2 project(V3 head, double *depth) const; // head frame -> pixels; depth along the optical axis
double off_axis(V2 uv) const; // degrees between the pixel's ray and the axis
};
// Loads /persist/xrservice.json and /persist/device_config.json. Keyed by calibration
// name: slam_left, slam_right, upper_left, upper_right.
bool load_calibration(std::map<std::string, Camera> &out, std::string &err);
// The Arcturus color cameras (tools/calib.py load_color has the conventions), for
// ft-camd --with-color's images: luma at 1/scale size, recorded as color_video<N>. They're
// keyed by those recorded names: left_node is passthrough_left, right_node
// passthrough_right. crop_subtract: image x = sensor x - the calibration's cropRegion.x.
// tools/check_color.py tells which node is which and which crop reading fits.
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
const std::string &right_node, bool crop_subtract, int scale, std::string &err);
// The point closest to several rays (weighted), and its rms distance to them.
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms);
+22
View File
@@ -0,0 +1,22 @@
// Small vector helpers for the tracker.
#pragma once
#include <array>
#include <cmath>
using V2 = std::array<double, 2>;
using V3 = std::array<double, 3>;
inline V3 operator+(V3 a, V3 b) { return {a[0] + b[0], a[1] + b[1], a[2] + b[2]}; }
inline V3 operator-(V3 a, V3 b) { return {a[0] - b[0], a[1] - b[1], a[2] - b[2]}; }
inline V3 operator*(V3 a, double s) { return {a[0] * s, a[1] * s, a[2] * s}; }
inline double dot(V3 a, V3 b) { return a[0] * b[0] + a[1] * b[1] + a[2] * b[2]; }
inline double norm(V3 a) { return std::sqrt(dot(a, a)); }
inline V3 unit(V3 a) { double n = norm(a); return n > 0 ? a * (1 / n) : a; }
inline V2 operator+(V2 a, V2 b) { return {a[0] + b[0], a[1] + b[1]}; }
inline V2 operator-(V2 a, V2 b) { return {a[0] - b[0], a[1] - b[1]}; }
inline V2 operator*(V2 a, double s) { return {a[0] * s, a[1] * s}; }
inline double norm(V2 a) { return std::hypot(a[0], a[1]); }
inline double wrap_angle(double a) { return std::remainder(a, 2 * M_PI); }
+171
View File
@@ -0,0 +1,171 @@
#include "io.h"
#include <fcntl.h>
#include <sys/mman.h>
#include <sys/stat.h>
#include <time.h>
#include <unistd.h>
#include <algorithm>
#include <cstdlib>
#include <cstring>
uint64_t mono_ns() {
timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return uint64_t(ts.tv_sec) * 1'000'000'000 + uint64_t(ts.tv_nsec);
}
int64_t raw_minus_mono_ns() {
timespec a, r, b;
clock_gettime(CLOCK_MONOTONIC, &a);
clock_gettime(CLOCK_MONOTONIC_RAW, &r);
clock_gettime(CLOCK_MONOTONIC, &b);
const int64_t ma = int64_t(a.tv_sec) * 1'000'000'000 + a.tv_nsec, mb = int64_t(b.tv_sec) * 1'000'000'000 + b.tv_nsec;
return int64_t(r.tv_sec) * 1'000'000'000 + r.tv_nsec - (ma + mb) / 2;
}
// ------------------------------------------------------------------------------ ring
bool Ring::open(const char *path, std::string &err) {
const int fd = ::open(path, O_RDONLY | O_CLOEXEC);
if (fd < 0) return err = std::string(path) + ": " + std::strerror(errno), false;
struct stat st;
fstat(fd, &st);
len_ = size_t(st.st_size);
void *m = len_ >= sizeof(fh_ring_hdr_t) ? mmap(nullptr, len_, PROT_READ, MAP_SHARED, fd, 0) : MAP_FAILED;
close(fd);
if (m == MAP_FAILED) return err = std::string(path) + ": can't map it", false;
map_ = static_cast<const uint8_t *>(m);
hdr_ = reinterpret_cast<const fh_ring_hdr_t *>(map_);
if (std::memcmp(hdr_->magic, FH_RING_MAGIC, 8) || hdr_->version != FH_RING_VERSION || hdr_->file_bytes > len_)
return err = std::string(path) + " is not an ft-camd ring", false;
return true;
}
bool Ring::alive() const {
const uint64_t hb = __atomic_load_n(&hdr_->heartbeat_ns, __ATOMIC_ACQUIRE);
return hb && mono_ns() - hb < 1'000'000'000;
}
uint64_t Ring::latest(int i) const { return __atomic_load_n(&hdr_->cams[i].latest, __ATOMIC_ACQUIRE); }
bool Ring::read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const {
const fh_ring_cam_t &c = hdr_->cams[i];
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
if (seq != 2 * n + 2) return false;
std::memcpy(meta, slot, sizeof *meta);
out.resize(size_t(c.width) * c.height);
for (uint32_t y = 0; y < c.height; ++y)
std::memcpy(out.data() + size_t(y) * c.width, slot + sizeof(fh_ring_slot_t) + size_t(y) * c.stride, c.width);
__atomic_thread_fence(__ATOMIC_ACQUIRE);
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
}
bool Ring::meta(int i, uint64_t n, fh_ring_slot_t *meta) const {
const fh_ring_cam_t &c = hdr_->cams[i];
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
if (seq != 2 * n + 2) return false;
std::memcpy(meta, slot, sizeof *meta);
__atomic_thread_fence(__ATOMIC_ACQUIRE);
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
}
// ------------------------------------------------------------------------- publisher
namespace {
// The hand's shape to cut out, as capsules.
// Radii are a real hand's half-widths plus a small margin for tracking noise.
const int kThumb[][2] = {{0, 1}, {1, 2}, {2, 3}, {3, 4}};
const int kFingers[][2] = {{5, 6}, {6, 7}, {7, 8}, {9, 10}, {10, 11}, {11, 12}, {13, 14}, {14, 15}, {15, 16},
{17, 18}, {18, 19}, {19, 20}};
const int kPalm[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 9}, {9, 13}, {13, 17}, {1, 5}};
constexpr double kThumbR = 0.0095, kFinger = 0.0085, kPalmR = 0.015, kArm[2] = {0.028, 0.034}, kArmLen = 0.16,
kMargin = 0.004;
// Nothing is cut closer than this in front of the eyes (head frame, -z is forward). A
// point near the eyes' plane lands far across a screen with a huge radius, so one bad
// estimate there tears a hole through it; real hands that close aren't tracked anyway.
constexpr double kNear = 0.12;
// Adds the capsule, clipped to the part at least kNear in front of the eyes.
void put(fh_capsule_t *caps, uint32_t &n, V3 a, V3 b, double ra, double rb) {
if (n >= FH_HANDS_MAX_CAPSULES) return;
const double za = -a[2] - kNear, zb = -b[2] - kNear; // >= 0: far enough in front
if (za < 0 && zb < 0) return;
if (za < 0 || zb < 0) {
const double t = za / (za - zb); // where the segment crosses the near plane
const V3 m = a + (b - a) * t;
const double rm = ra + (rb - ra) * t;
if (za < 0) a = m, ra = rm;
else b = m, rb = rm;
}
fh_capsule_t &c = caps[n++];
for (int k = 0; k < 3; ++k) c.a[k] = float(a[k]), c.b[k] = float(b[k]);
c.ra = float(ra + kMargin), c.rb = float(rb + kMargin);
}
} // namespace
std::string run_dir() {
const std::string dir = "/run/user/" + std::to_string(getuid()) + "/frametop-hands";
mkdir(dir.c_str(), 0700);
return dir;
}
bool Publisher::open(std::string &err) {
const std::string path = run_dir() + "/hands";
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
if (fd < 0 || ftruncate(fd, sizeof(fh_hands_t)) < 0) return err = path + ": " + std::strerror(errno), false;
void *m = mmap(nullptr, sizeof(fh_hands_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
close(fd);
if (m == MAP_FAILED) return err = path + ": can't map it", false;
out_ = static_cast<fh_hands_t *>(m);
std::memset(out_, 0, sizeof *out_);
std::memcpy(out_->magic, FH_HANDS_MAGIC, 8);
out_->version = FH_HANDS_VERSION;
out_->size = sizeof(fh_hands_t);
return true;
}
void Publisher::write(const std::vector<const Hand *> &in, uint64_t capture_ns) {
std::vector<const Hand *> hands = in;
std::sort(hands.begin(), hands.end(), [](const Hand *a, const Hand *b) { return a->frames > b->frames; });
if (hands.size() > FH_HANDS_MAX_HANDS) hands.resize(FH_HANDS_MAX_HANDS);
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
__atomic_thread_fence(__ATOMIC_RELEASE);
uint32_t nc = 0;
for (size_t k = 0; k < FH_HANDS_MAX_HANDS; ++k) {
fh_hand_t &o = out_->hands[k];
std::memset(&o, 0, sizeof o);
if (k >= hands.size()) continue;
const Hand &h = *hands[k];
o.id = uint32_t(h.id);
o.flags = (h.right() ? FH_HAND_RIGHT : 0) | (h.nviews >= 2 ? FH_HAND_STEREO : 0);
o.confidence = float(std::min(1.0, h.frames / 5.0));
for (int i = 0; i < 21; ++i)
for (int j = 0; j < 3; ++j) o.pts[i][j] = float(h.smooth[i][j]);
const uint32_t first = nc;
for (auto &b : kThumb) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kThumbR, kThumbR);
for (auto &b : kFingers) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kFinger, kFinger);
for (auto &b : kPalm) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kPalmR, kPalmR);
// the forearm carries on from the hand's own axis (middle knuckle -> wrist); the
// wrist bends, but much less than a guess at where the elbow is gets wrong
const V3 wrist = h.smooth[0], d = wrist - h.smooth[9];
const double n = norm(d);
if (n > 0.02) put(out_->capsules, nc, wrist, wrist + d * (kArmLen / n), kArm[0], kArm[1]);
o.ncapsules = nc - first;
}
for (uint32_t k = nc; k < FH_HANDS_MAX_CAPSULES; ++k) std::memset(&out_->capsules[k], 0, sizeof(fh_capsule_t));
out_->capture_ns = capture_ns;
out_->publish_ns = mono_ns();
out_->nhands = uint32_t(hands.size());
out_->ncapsules = nc;
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
}
+52
View File
@@ -0,0 +1,52 @@
// Frames in from ft-camd's ring (camd/fhring.h), hands out to the hands file
// (include/fh_hands.h, read by Frametop's ft-screens).
#pragma once
#include "tracker.h"
#include <cstdint>
#include <string>
#include <vector>
extern "C" {
#include "../camd/fhring.h"
#include "../include/fh_hands.h"
}
class Ring {
public:
bool open(const char *path, std::string &err);
bool alive() const; // the writer's heartbeat is fresh
int cameras() const { return int(hdr_->ncams); }
const fh_ring_cam_t &camera(int i) const { return hdr_->cams[i]; }
uint64_t latest(int i) const;
// Copy frame n of camera i into out (width x height, tightly packed). False if it's
// gone or was being written.
bool read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const;
// Just frame n's slot header (capture time etc.), without copying the image.
bool meta(int i, uint64_t n, fh_ring_slot_t *meta) const;
private:
const uint8_t *map_ = nullptr;
const fh_ring_hdr_t *hdr_ = nullptr;
size_t len_ = 0;
};
class Publisher {
public:
bool open(std::string &err);
void write(const std::vector<const Hand *> &hands, uint64_t capture_ns);
private:
fh_hands_t *out_ = nullptr;
uint64_t seq_ = 0;
};
uint64_t mono_ns();
int64_t raw_minus_mono_ns(); // camera timestamps are CLOCK_MONOTONIC_RAW
// /run/user/UID/frametop-hands, created private to the user if it's missing: where ft-camd's ring
// (FH_RING_NAME) and the hands and gestures files live. Not $XDG_RUNTIME_DIR: a terminal in
// the Frametop desktop has a private one of its own. And not /run/user/UID/frametop: that is
// the desktop session's private runtime folder, which it deletes whenever it starts.
std::string run_dir();
+626
View File
@@ -0,0 +1,626 @@
// ft-hands: hands in 3D from ft-camd's ring, published for Frametop's ft-screens (the hand
// cutouts), and pinches and grips for the pointer. It started as a port of frame-hands'
// Python prototype: the same scheduling, with the models on a few threads.
//
// ft-hands [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N]
// [--no-publish] [--record DIR] [--swap-sides] [--cams auto|mono|color|all] ...
// (--help lists them all)
//
// Which cameras (--cams, HANDS_CAMERAS): the four mono IR cameras light the hands with their
// own IR and track well in dim rooms, but in bright light (a sunny room, a window behind the
// hands) they expose for the room and the hands come out dark. The two Arcturus colour
// cameras (the passthrough pair, forward-facing, 145 degrees) are the other way round: dark
// and grainy in a dim room, clear in a lit one. auto (the default) picks by how bright the
// colour cameras' frames are: at HANDS_BRIGHT_ON (mean luma) or over for 2 s, it tracks with
// HANDS_BRIGHT (all: every camera, so hands low at the sides stay in the side cameras; or
// color); under HANDS_BRIGHT_OFF for 2 s, with the mono cameras again. ft-camd runs the colour
// cameras at 2 fps, enough to tell the light, until ft-hands asks for 30
// (/run/user/UID/frametop-hands/color-fps). The colour frames' capture times are on their
// own clock, so they're placed on the mono cameras' by when they were dequeued, less the
// mono cameras' measured delay.
//
// Settings in ~/.config/frametop.conf (FT_<name> in the environment overrides them, and
// options override both): HANDS_SWAP_SIDES (1: as --swap-sides), HANDS_CPUS (as --cpus),
// HANDS_CAMERAS, HANDS_BRIGHT, HANDS_BRIGHT_ON, HANDS_BRIGHT_OFF, HANDS_COLOR_LEFT (which
// colour camera is passthrough_left: color_video0 or color_video3), HANDS_COLOR_CROP
// (subtract or none: tools/check_color.py tells both).
#include "io.h"
#include "pinch.h"
#include "record.h"
#include <sched.h>
#include <sys/resource.h>
#include <sys/stat.h>
#include <unistd.h>
#include <cmath>
#include <ctime>
#include <memory>
#include <csignal>
#include <cstdlib>
#include <cstdio>
#include <cstring>
#include <fstream>
#include <thread>
#include <utility>
namespace {
volatile std::sig_atomic_t g_stop = 0, g_record = 0;
// which calibrated camera each capture pipe carries (XRService's fixed routing)
const char *camera_for_pipe(int node) {
char path[64], name[64] = "";
std::snprintf(path, sizeof path, "/sys/class/video4linux/video%d/name", node);
std::ifstream f(path);
f.getline(name, sizeof name);
if (!std::strcmp(name, "msm_vfe3_video0")) return "slam_left";
if (!std::strcmp(name, "msm_vfe4_video0")) return "slam_right";
if (!std::strcmp(name, "msm_vfe2_video0")) return "upper_left";
if (!std::strcmp(name, "msm_vfe2_video1")) return "upper_right";
return nullptr;
}
// A setting from ~/.config/frametop.conf, or FT_<key> from the environment; "" if unset.
std::string setting(const std::string &key) {
if (const char *v = std::getenv(("FT_" + key).c_str())) return v;
const char *home = std::getenv("HOME");
std::ifstream in(std::string(home ? home : "") + "/.config/frametop.conf");
std::string line, value;
auto trim = [](std::string s) {
s.erase(0, s.find_first_not_of(" \t\"'"));
s.erase(s.find_last_not_of(" \t\"'") + 1);
return s;
};
while (std::getline(in, line)) {
line = line.substr(0, line.find('#'));
const auto eq = line.find('=');
if (eq != std::string::npos && trim(line.substr(0, eq)) == key) value = trim(line.substr(eq + 1));
}
return value;
}
std::vector<int> parse_cpus(const char *p) {
std::vector<int> out;
while (*p) {
char *end;
const long c = std::strtol(p, &end, 10);
if (end == p) break;
out.push_back(int(c));
p = *end == ',' ? end + 1 : end;
}
return out;
}
// Where SIGUSR1 puts recordings: $XDG_DATA_HOME/frametop/hands (~/.local/share/...).
std::string recordings_dir() {
const char *data = std::getenv("XDG_DATA_HOME"), *home = std::getenv("HOME");
std::string dir = data && *data ? data : std::string(home ? home : "") + "/.local/share";
for (const char *part : {"/frametop", "/hands"}) mkdir((dir += part).c_str(), 0700);
return dir;
}
double cpu_seconds() {
rusage r;
getrusage(RUSAGE_SELF, &r);
return r.ru_utime.tv_sec + r.ru_stime.tv_sec + (r.ru_utime.tv_usec + r.ru_stime.tv_usec) / 1e6;
}
enum class Cams { Mono, Color, All };
const char *cams_name(Cams c) { return c == Cams::Mono ? "mono" : c == Cams::Color ? "color" : "all"; }
bool parse_cams(const std::string &s, Cams &out) {
if (s == "mono") out = Cams::Mono;
else if (s == "color") out = Cams::Color;
else if (s == "all") out = Cams::All;
else return false;
return true;
}
// How bright it is, for auto (see the top): the colour frames' mean luma, smoothed over about
// a second, with hysteresis and a 2 s hold each way. No colour frames for 3 s (ft-camd paused
// them, or has none) reads as dim.
struct Lighting {
double on = 40, off = 25;
double level = -1;
bool bright = false;
uint64_t at_ns = 0, since_ns = 0; // the last frame; since when it's wanted the other way
void add(double mean, uint64_t t_ns) {
const double dt = at_ns && t_ns > at_ns ? (t_ns - at_ns) / 1e9 : 1.0;
level = level < 0 ? mean : level + (mean - level) * std::min(1.0, dt / 1.0);
at_ns = t_ns;
}
// True when it switched.
bool update(uint64_t now_ns) {
if (level >= 0 && now_ns - at_ns > 3'000'000'000ull) level = -1;
const bool want = level >= 0 && (bright ? level > off : level >= on);
if (want == bright) return since_ns = 0, false;
if (!since_ns) since_ns = now_ns;
if (now_ns - since_ns < 2'000'000'000ull) return false;
bright = want, since_ns = 0;
return true;
}
};
} // namespace
int main(int argc, char **argv) {
double seconds = 0, status = 5;
int threads = 3, niceness = 5;
bool int8 = false, publish = true, track = true, swap_sides = false;
std::string models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
std::string record, ring_path = "/run/user/" + std::to_string(getuid()) + "/" FH_RING_NAME;
// SteamOS starts user processes on CPUs 0-4 and keeps 5-7 (two A720s and the X4) for
// SteamVR's compositor, whose threads there run at real-time priority, so they always
// win. XRService pins its head tracking to 2-3. frame-hands' probes/core_ab.py
// (2026-09-29, headset on, 3 rounds): on 5-7 a step took 8.4 ms against 13.2 on 2-4,
// latency 9.6 against 14.1 ms, and the compositor's late frames and CPU/GPU time didn't change.
std::vector<int> cpus = {5, 6, 7};
if (const auto c = parse_cpus(setting("HANDS_CPUS").c_str()); !c.empty()) cpus = c;
swap_sides = setting("HANDS_SWAP_SIDES") == "1";
// Which cameras (see the top).
std::string cams_arg = setting("HANDS_CAMERAS"), bright_arg = setting("HANDS_BRIGHT");
std::string color_left = setting("HANDS_COLOR_LEFT"), color_crop = setting("HANDS_COLOR_CROP");
if (cams_arg.empty()) cams_arg = "auto";
if (bright_arg.empty()) bright_arg = "all";
if (color_left.empty()) color_left = "color_video0";
if (color_crop.empty()) color_crop = "subtract";
Lighting light;
if (const std::string v = setting("HANDS_BRIGHT_ON"); !v.empty()) light.on = std::atof(v.c_str());
if (const std::string v = setting("HANDS_BRIGHT_OFF"); !v.empty()) light.off = std::atof(v.c_str());
// How crops are equalized. CLAHE helps the palm search find hands (about 10% more in the
// dim recording), but makes the landmarks jitter, so they get plain crops.
Contrast palm_contrast, hand_contrast{Contrast::None};
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
PinchParams pinch_params;
GripParams grip_params;
bool gesture_log = false; // what the pinch and grip detectors measure, 10 times a second
double record_for = 120;
for (int i = 1; i < argc; ++i) {
const std::string a = argv[i];
const bool more = i + 1 < argc;
if (a == "--seconds" && more) seconds = std::atof(argv[++i]);
else if (a == "--threads" && more) threads = std::max(1, std::atoi(argv[++i]));
else if (a == "--status" && more) status = std::atof(argv[++i]);
else if (a == "--models" && more) models = argv[++i];
else if (a == "--nice" && more) niceness = std::atoi(argv[++i]);
else if (a == "--int8") int8 = true;
else if (a == "--no-publish") publish = false;
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
else if (a == "--pinch-palm-down" && more) pinch_params.palm_down_max = std::atof(argv[++i]);
else if (a == "--grip-begin" && more) grip_params.begin = std::atof(argv[++i]);
else if (a == "--grip-end" && more) grip_params.end = std::atof(argv[++i]);
else if (a == "--gesture-log") gesture_log = true;
else if (a == "--swap-sides") swap_sides = true;
else if (a == "--record-only") track = publish = false;
else if (a == "--ring" && more) ring_path = argv[++i];
else if (a == "--record" && more) record = argv[++i];
else if (a == "--record-for" && more) record_for = std::atof(argv[++i]);
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
else if (a == "--cams" && more) cams_arg = argv[++i];
else if (a == "--bright" && more) bright_arg = argv[++i];
else if (a == "--bright-on" && more) light.on = std::atof(argv[++i]);
else if (a == "--bright-off" && more) light.off = std::atof(argv[++i]);
else if (a == "--color-left" && more) color_left = argv[++i];
else if (a == "--color-crop" && more) color_crop = argv[++i];
else if (a == "--contrast" && more) {
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
} else if (a == "--cpus" && more) {
cpus = parse_cpus(argv[++i]);
if (cpus.empty()) cpus = {5, 6, 7};
}
else {
std::printf("usage: %s [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] [--no-publish]\n"
" [--record DIR] [--record-for S] [--record-only] [--cpus 5,6,7] [--swap-sides]\n"
" [--keep-presence P] (0.5) [--ring PATH] (ft-camd's, or ft-ringplay's)\n"
" [--cams auto|mono|color|all] (auto) [--bright all|color] (all) [--bright-on L] (40) [--bright-off L] (25)\n"
" [--color-left color_video0|color_video3] [--color-crop subtract|none]\n"
" [--pinch-begin M] (0.020) [--pinch-end M] (0.035) [--pinch-triangulated] [--pinch-palm-down MAX] (0.6)\n"
" [--grip-begin R] (1.2) [--grip-end R] (1.45) [--gesture-log]\n"
" [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch; default clahe:2/none)\n"
"Recording saves every frame set for S seconds (120) to DIR/sets.bin, for ft-handreplay; SIGUSR1\n"
"starts one in ~/.local/share/frametop/hands/rec-<time>. --record-only records without tracking, so it\n"
"can run beside a tracking ft-hands. With ft-camd --with-dark, recordings also get each\n"
"camera's newest dark frame, as <name>_dk; with --with-color, the color cameras' as color_video<N>.\n"
"auto picks the cameras by the light (see the top of track/main.cpp).\n"
"Settings in ~/.config/frametop.conf: HANDS_SWAP_SIDES=1, HANDS_CPUS=5,6,7, HANDS_CAMERAS, HANDS_BRIGHT,\n"
"HANDS_BRIGHT_ON, HANDS_BRIGHT_OFF, HANDS_COLOR_LEFT, HANDS_COLOR_CROP (FT_<name> overrides).\n",
argv[0]);
return a == "--help" ? 0 : 1;
}
}
const bool automatic = cams_arg == "auto";
Cams fixed = Cams::Mono, bright_cams = Cams::All;
if ((!automatic && !parse_cams(cams_arg, fixed)) || !parse_cams(bright_arg, bright_cams) || bright_cams == Cams::Mono)
return std::fprintf(stderr, "--cams auto|mono|color|all, --bright all|color\n"), 1;
if (color_crop != "subtract" && color_crop != "none") return std::fprintf(stderr, "--color-crop subtract|none\n"), 1;
std::setvbuf(stdout, nullptr, _IOLBF, 0); // whole lines to the journal as they come
if (nice(niceness) < 0) std::perror("nice"); // the VR stack wins contested CPUs
std::signal(SIGINT, [](int) { g_stop = 1; });
std::signal(SIGTERM, [](int) { g_stop = 1; });
std::signal(SIGUSR1, [](int) { g_record = 1; });
std::string err;
std::map<std::string, Camera> calib;
Ring ring;
Nets nets;
Publisher pub;
GesturePublisher gestures;
Pinch pinch(pinch_params);
Grip grip(grip_params);
std::unique_ptr<Recorder> rec;
uint64_t rec_start = 0;
auto start_recording = [&](const std::string &dir, std::string &e) {
rec = std::make_unique<Recorder>();
if (!rec->open(dir, e)) return rec.reset(), false;
rec_start = mono_ns();
std::printf("recording to %s for %.0f s\n", dir.c_str(), record_for);
std::fflush(stdout);
return true;
};
if (!load_calibration(calib, err) || !ring.open(ring_path.c_str(), err) || !nets.load(models, int8, err) ||
(publish && (!pub.open(err) || !gestures.open(pinch, grip, err))) ||
(!record.empty() && !start_recording(record, err))) {
std::fprintf(stderr, "%s\n", err.c_str());
return 1;
}
if (!ring.alive()) return std::fprintf(stderr, "ft-camd isn't running (no heartbeat)\n"), 1;
std::map<std::string, int> index; // mono calibration name -> ring camera
std::map<std::string, int> color; // colour calibration name (color_video<N>) -> ring camera
// Recorded as they are with each set: "<name>_dk" (ft-camd --with-dark) and "color_video<N>"
// (--with-color). Recorded names hold 15 characters, so "upper_right_dark" wouldn't fit.
std::map<std::string, int> dark;
std::map<std::string, Camera> used;
for (int i = 0; i < ring.cameras(); ++i) {
if (ring.camera(i).flags & FH_CAM_COLOR) {
const std::string name = "color_video" + std::to_string(ring.camera(i).node);
dark[name] = i, color[name] = i;
continue;
}
// ft-camd's cameras by capture pipe; ft-ringplay's (no device) by the name it gives
const char *name = camera_for_pipe(ring.camera(i).node);
if (!name && ring.camera(i).node < 0) name = ring.camera(i).name;
if (!name || !calib.count(name)) continue;
if (ring.camera(i).flags & FH_CAM_DARK) dark[std::string(name) + "_dk"] = i;
else index[name] = i, used[name] = calib[name];
}
// ft-camd tells the side cameras' buffers apart by XRService's allocation order, which
// some XRService restarts reverse; tools/check_sides.py --ring tells when.
if (swap_sides && index.count("slam_left") && index.count("slam_right")) {
std::swap(index["slam_left"], index["slam_right"]);
if (dark.count("slam_left_dk") && dark.count("slam_right_dk")) std::swap(dark["slam_left_dk"], dark["slam_right_dk"]);
std::printf("side cameras swapped (--swap-sides)\n");
}
// The colour pair, calibrated (see the top), unless only the mono cameras are wanted.
if (color.size() == 2 && (automatic || fixed != Cams::Mono)) {
const std::string left = color.count(color_left) ? color_left : color.begin()->first;
const std::string right = color.begin()->first == left ? std::next(color.begin())->first : color.begin()->first;
const int scale = int(std::lround(1972.0 / ring.camera(color[left]).width));
std::map<std::string, Camera> cc;
std::string e;
if (load_color_calibration(cc, left, right, color_crop == "subtract", scale, e)) {
for (auto &[name, cam] : cc) used[name] = cam;
std::printf("colour cameras: %s is passthrough_left, crop %s, 1/%d size\n", left.c_str(), color_crop.c_str(), scale);
} else {
std::fprintf(stderr, "colour cameras left out: %s\n", e.c_str());
color.clear();
}
} else {
color.clear();
}
if (color.empty() && (automatic || fixed != Cams::Mono)) {
if (!automatic) std::printf("no colour cameras (ft-camd --with-color): tracking with the mono ones\n");
fixed = Cams::Mono;
}
const bool switching = automatic && !color.empty();
Cams mode = switching || color.empty() ? Cams::Mono : fixed;
std::printf("cameras:");
for (auto &[name, i] : index) std::printf(" %s=video%d", name.c_str(), ring.camera(i).node);
for (auto &[name, i] : color) std::printf(" %s", name.c_str());
std::printf(" tracking with %s%s models: %s%s, %d threads on CPUs", cams_name(mode),
switching ? " (auto: by the light)" : "", models.c_str(), int8 ? " (int8)" : "", threads);
for (int c : cpus) std::printf(" %d", c);
std::printf("\n");
cpu_set_t set; // the main loop too
CPU_ZERO(&set);
for (int c : cpus) CPU_SET(c, &set);
if (sched_setaffinity(0, sizeof set, &set) < 0) std::perror("sched_setaffinity");
nets.set_contrast(palm_contrast, hand_contrast);
Pool pool(threads, cpus);
Tracker tracker(used, nets, pool);
tracker.set_keep_presence(keep_presence);
std::map<std::string, std::vector<uint8_t>> pixels;
std::map<std::string, uint64_t> last; // per camera: the frame last used
std::map<std::string, uint64_t> lit_seen; // per colour camera: the frame last counted for the light
const uint64_t start = mono_ns();
uint64_t t_status = start, next_ns = 0, t_want = 0, t_glog = 0;
double cpu0 = cpu_seconds();
std::vector<double> lat;
double hands_sum = 0, resid_sum = 0;
int resid_n = 0, left_sets = 0, right_sets = 0, both_sets = 0, color_steps = 0;
// The mono cameras' delay from capture to dequeue (their capture clock is CLOCK_MONOTONIC_RAW),
// to place the colour frames, whose capture clock is their own (see the top).
double mono_delay_ns = -1;
const std::string want_file = run_dir() + "/color-fps";
// The colour pair's newest frames if both are newer than last used and taken together
// (their own clock): their time, on the mono cameras' capture clock, else 0.
auto color_pair = [&](int64_t raw_off, bool copy, std::map<std::string, Image> &images) -> uint64_t {
fh_ring_slot_t meta[2];
std::string names[2];
uint64_t n[2];
int k = 0;
for (auto &[name, i] : color) {
names[k] = name, n[k] = ring.latest(i);
if (n[k] <= last[name] || !ring.meta(i, n[k], &meta[k])) return 0;
++k;
}
if (k != 2 || mono_delay_ns < 0) return 0;
const int64_t apart = int64_t(meta[0].capture_ns) - int64_t(meta[1].capture_ns);
if (std::llabs(apart) > 3'000'000) return 0; // one is a frame ahead: wait for the other
const uint64_t dq = std::min(meta[0].dqbuf_ns, meta[1].dqbuf_ns);
const uint64_t t = uint64_t(int64_t(dq) - int64_t(mono_delay_ns) + raw_off);
if (!copy) return t;
for (int j = 0; j < 2; ++j) {
const int i = color[names[j]];
if (!ring.read(i, n[j], pixels[names[j]], &meta[j])) return 0;
const auto &c = ring.camera(i);
images[names[j]] = {pixels[names[j]].data(), int(c.width), int(c.height), int(c.width)};
}
for (int j = 0; j < 2; ++j) last[names[j]] = n[j];
return t;
};
auto switch_to = [&](Cams to, const char *why) {
if (to == mode) return;
for (auto &[name, cam] : used) {
const bool is_color = color.count(name) > 0;
const bool keep = to == Cams::All || (to == Cams::Color) == is_color;
if (!keep) tracker.drop_camera(name);
}
std::printf("cameras: %s -> %s (%s)\n", cams_name(mode), cams_name(to), why);
mode = to;
};
auto ambient = [&] { // the mono cameras' dark frames: the room's IR light
double sum = 0;
int n = 0;
for (auto &[name, i] : index)
if (ring.camera(i).dark_mean > 0) sum += ring.camera(i).dark_mean, ++n;
return n ? sum / n : -1;
};
while (!g_stop && (seconds <= 0 || (mono_ns() - start) / 1e9 < seconds)) {
if (!ring.alive()) return std::fprintf(stderr, "ft-camd stopped\n"), 2;
const uint64_t now0 = mono_ns();
const int64_t raw_off = raw_minus_mono_ns();
// The light, from the colour frames' brightness (their slot headers only).
for (auto &[name, i] : color) {
fh_ring_slot_t m;
const uint64_t n = ring.latest(i);
if (n && n != lit_seen[name] && ring.meta(i, n, &m)) light.add(m.mean, m.dqbuf_ns), lit_seen[name] = n;
}
if (switching && light.update(now0)) {
char why[96];
std::snprintf(why, sizeof why, "colour frames at %.0f, ambient IR %.1f", light.level, ambient());
switch_to(light.bright ? bright_cams : Cams::Mono, why);
}
// Ask ft-camd for the colour cameras' full rate while tracking or recording with them.
const bool want_color = !color.empty() && (mode != Cams::Mono || rec || g_record);
if (!color.empty() && now0 - t_want > 1'000'000'000) {
t_want = now0;
if (want_color) {
if (FILE *f = std::fopen(want_file.c_str(), "w")) std::fputs("30\n", f), std::fclose(f);
} else {
unlink(want_file.c_str());
}
}
const bool use_mono = mode != Cams::Color, use_color = mode != Cams::Mono;
const bool mono_driven = use_mono || rec || g_record || !track;
std::map<std::string, Image> images;
uint64_t tmin = UINT64_MAX, dq = 0;
if (mono_driven) {
// a new frame set: every mono camera has a newer frame, taken at the same moment
std::map<std::string, uint64_t> latest;
bool ready = true;
for (auto &[name, i] : index) {
latest[name] = ring.latest(i);
ready = ready && latest[name] > last[name];
}
if (!ready) {
std::this_thread::sleep_for(std::chrono::milliseconds(2));
continue;
}
// not needed at the current rate, and not recorded: skip it without copying images
if (track && !rec && !g_record) {
uint64_t t0 = UINT64_MAX, t1 = 0;
bool ok = true;
for (auto &[name, i] : index) {
fh_ring_slot_t meta;
ok = ok && ring.meta(i, latest[name], &meta);
if (ok) t0 = std::min(t0, meta.capture_ns), t1 = std::max(t1, meta.capture_ns);
}
if (ok && t1 - t0 <= 3'000'000 && t0 < next_ns) {
for (auto &[name, i] : index) last[name] = latest[name];
continue;
}
}
std::vector<SetFrame> frames;
uint64_t tmax = 0;
bool ok = true;
for (auto &[name, i] : index) {
fh_ring_slot_t meta;
ok = ok && ring.read(i, latest[name], pixels[name], &meta);
if (!ok) break;
const auto &c = ring.camera(i);
images[name] = {pixels[name].data(), int(c.width), int(c.height), int(c.width)};
frames.push_back({name, pixels[name].data(), c.width, c.height, meta.capture_ns, meta.dqbuf_ns});
tmin = std::min(tmin, meta.capture_ns), tmax = std::max(tmax, meta.capture_ns), dq = std::max(dq, meta.dqbuf_ns);
const double delay = double(int64_t(meta.dqbuf_ns) - (int64_t(meta.capture_ns) - raw_off));
mono_delay_ns = mono_delay_ns < 0 ? delay : mono_delay_ns + 0.02 * (delay - mono_delay_ns);
}
if (!ok || tmax - tmin > 3'000'000) { // torn, or a camera is a frame behind
std::this_thread::sleep_for(std::chrono::milliseconds(1));
continue;
}
for (auto &[name, i] : index) last[name] = latest[name];
if (g_record && !rec) {
g_record = 0;
char name[64];
const std::time_t now = std::time(nullptr);
std::strftime(name, sizeof name, "rec-%Y%m%d-%H%M%S", std::localtime(&now));
const std::string dir = recordings_dir();
std::string e;
if (!start_recording(dir + "/" + name, e)) std::fprintf(stderr, "%s\n", e.c_str());
}
if (rec) { // about 80 MB/s; dark frames double that, color frames add 70 MB/s
if ((mono_ns() - rec_start) / 1e9 < record_for) {
for (auto &[name, i] : dark) { // the newest dark and color frames, as they are
fh_ring_slot_t meta;
const uint64_t n = ring.latest(i);
const auto &c = ring.camera(i);
if (n && ring.read(i, n, pixels[name + "#rec"], &meta))
frames.push_back({name, pixels[name + "#rec"].data(), c.width, c.height, meta.capture_ns,
meta.dqbuf_ns});
}
rec->add(frames);
} else {
const size_t n = rec->written(), d = rec->dropped();
rec.reset(); // writes out what's queued
std::printf("recording done: %zu sets, %zu dropped\n", n, d);
std::fflush(stdout);
if (!track) break;
}
}
if (!track && rec && status > 0 && (mono_ns() - t_status) / 1e9 >= status) {
std::printf("%5.1fs recorded %zu sets, dropped %zu\n", (mono_ns() - start) / 1e9, rec->written(), rec->dropped());
std::fflush(stdout);
t_status = mono_ns();
}
if (!track || tmin < next_ns) continue; // not needed yet at the current rate
if (!use_mono) images.clear(); // colour only, driven by mono while recording
if (use_color && color_pair(raw_off, true, images)) ++color_steps;
if (images.empty()) continue;
} else {
// colour only: a new pair of colour frames
for (auto &[name, i] : index) { // keep the mono cameras' delay current
fh_ring_slot_t meta;
const uint64_t n = ring.latest(i);
if (n && ring.meta(i, n, &meta)) {
const double delay = double(int64_t(meta.dqbuf_ns) - (int64_t(meta.capture_ns) - raw_off));
mono_delay_ns = mono_delay_ns < 0 ? delay : mono_delay_ns + 0.02 * (delay - mono_delay_ns);
}
break;
}
const uint64_t t = color_pair(raw_off, false, images);
if (!t) {
std::this_thread::sleep_for(std::chrono::milliseconds(2));
continue;
}
if (t < next_ns) { // not needed yet: pass it by without copying
for (auto &[name, i] : color) last[name] = ring.latest(i);
continue;
}
if (!color_pair(raw_off, true, images)) continue;
tmin = t, ++color_steps;
for (auto &[name, i] : color) {
fh_ring_slot_t meta;
if (ring.meta(i, last[name], &meta)) dq = std::max(dq, meta.dqbuf_ns);
}
}
const auto hands = tracker.step(images, int64_t(tmin));
const uint64_t capture = uint64_t(int64_t(tmin) - raw_off); // CLOCK_MONOTONIC
const std::vector<Seen> views = tracker.views_now();
grip.update(hands, views, int64_t(capture));
pinch.update(hands, views, int64_t(capture), grip.gripping());
if (gesture_log && capture - t_glog >= 100'000'000) {
t_glog = capture;
for (int k = 0; k < 2; ++k)
if (pinch.world_d[k] >= 0)
std::printf("gesture %s d world %.3f tri %.3f palm-down %.2f curl %.2f%s\n", k ? "right" : "left ",
pinch.world_d[k], pinch.tri_d[k], pinch.palm_down[k], grip.curl[k],
pinch.side(k).flags & FH_PINCH_DOWN ? " PINCH" : grip.side(k).flags & FH_PINCH_DOWN ? " GRIP" : "");
}
// a gesture down or closing gets the full rate, even while the palm holds still
next_ns = tmin + uint64_t((std::min(tracker.interval(), pinch.engaged() || grip.engaged() ? 1 / 30.0 : 1.0) -
0.005) * 1e9);
if (publish) pub.write(hands, capture), gestures.write(pinch, grip, capture);
for (const Pinch::Event &e : grip.events)
std::printf("grip %s %-5s curl %.2f at %+.3f %+.3f %+.3f\n", e.side ? "right" : "left ", e.what, e.distance,
e.point[0], e.point[1], e.point[2]);
for (const Pinch::Event &e : pinch.events)
std::printf("pinch %s %-5s d %.3f m at %+.3f %+.3f %+.3f\n", e.side ? "right" : "left ", e.what, e.distance,
e.point[0], e.point[1], e.point[2]);
lat.push_back((mono_ns() - dq) / 1e6);
hands_sum += double(hands.size());
bool on_left = false, on_right = false; // by where the wrist is, not the model's label
for (const Hand *h : hands) {
if (h->residual >= 0) resid_sum += h->residual * 1000, ++resid_n;
(h->pts[0][0] < 0 ? on_left : on_right) = true;
}
left_sets += on_left, right_sets += on_right, both_sets += on_left && on_right;
const uint64_t now = mono_ns();
if (status > 0 && (now - t_status) / 1e9 >= status) {
const double dt = (now - t_status) / 1e9, cpu1 = cpu_seconds();
const Stats &s = tracker.stats;
std::sort(lat.begin(), lat.end());
std::printf("%5.1fs %4.1f sets/s hands %.2f views %zu palm %3d calls %4.1f ms/batch hand %3d calls %4.1f ms/batch "
"step %4.1f ms latency %4.1f ms resid %.1f mm CPU %3.0f%%\n",
(now - start) / 1e9, s.sets / dt, s.sets ? hands_sum / s.sets : 0, tracker.views(), s.palm_calls,
s.palm_batches ? s.palm_ms / s.palm_batches : 0, s.hand_calls,
s.hand_batches ? s.hand_ms / s.hand_batches : 0, s.sets ? s.step_ms / s.sets : 0,
lat.empty() ? 0 : lat[lat.size() / 2], resid_n ? resid_sum / resid_n : 0, 100 * (cpu1 - cpu0) / dt);
if (!color.empty())
std::printf(" cameras %s%s: colour frames at %.1f (bright at %.0f, dim under %.0f), ambient IR %.1f, "
"%d steps with colour, colour placed %.1f ms after capture\n",
cams_name(mode), switching ? " (auto)" : "", light.level, light.on, light.off, ambient(),
color_steps, mono_delay_ns / 1e6);
if (s.sets)
std::printf(" sets with a hand: left %2.0f%% right %2.0f%% both %2.0f%% views lost %d, handoff misses %d, "
"dups %d, splits %d hands new %d merged %d forgotten %d%s\n",
100.0 * left_sets / s.sets, 100.0 * right_sets / s.sets, 100.0 * both_sets / s.sets, s.lost,
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten,
!rec ? "" : (" recorded " + std::to_string(rec->written()) + " dropped " +
std::to_string(rec->dropped())).c_str());
std::printf(" pinches: left %u right %u (held back, palm down: %d %d) grips: left %u right %u",
pinch.side(0).begins, pinch.side(1).begins, pinch.held_back[0], pinch.held_back[1],
grip.side(0).begins, grip.side(1).begins);
for (int k = 0; k < 2; ++k)
if (pinch.side(k).flags & FH_PINCH_TRACKED)
std::printf(" %s %s d %.3f curl %.2f", k ? "right" : "left",
grip.side(k).flags & FH_PINCH_DOWN ? "GRIP"
: pinch.side(k).flags & FH_PINCH_DOWN ? "PINCH"
: "open",
pinch.side(k).distance, grip.curl[k]);
std::printf("\n");
for (const Hand *h : hands)
std::printf(" hand %d %-5s views %d wrist %+.3f %+.3f %+.3f m scale %.2f speed %.2f m/s\n", h->id,
h->right() ? "right" : "left", h->nviews, h->pts[0][0], h->pts[0][1], h->pts[0][2], h->scale,
h->speed);
std::fflush(stdout);
tracker.stats = Stats{};
t_status = now, cpu0 = cpu1;
lat.clear(), hands_sum = 0, resid_sum = 0, resid_n = 0, left_sets = right_sets = both_sets = 0;
color_steps = 0;
}
}
if (!color.empty()) unlink(want_file.c_str());
if (publish) {
pub.write({}, mono_ns());
pinch.release(int64_t(mono_ns())); // a drag in progress ends, as lost
grip.release(int64_t(mono_ns()));
gestures.write(pinch, grip, mono_ns());
}
return 0;
}
+255
View File
@@ -0,0 +1,255 @@
#include "nets.h"
#include <mat.h>
#include <algorithm>
#include <cstdlib>
#include <numeric>
namespace {
constexpr int kPalmSize = 192, kHandSize = 224;
const int kRoiLandmarks[] = {0, 1, 2, 3, 5, 6, 9, 10, 13, 14, 17, 18};
// 2x3 affine taking crop pixels (0..out) to image pixels.
void crop_matrix(V2 center, double size, double rotation, int out, float tm[6]) {
const double c = std::cos(rotation), s = std::sin(rotation), k = size / out;
tm[0] = float(c * k), tm[1] = float(-s * k), tm[3] = float(s * k), tm[4] = float(c * k);
tm[2] = float(center[0] - (tm[0] + tm[1]) * out / 2.0);
tm[5] = float(center[1] - (tm[3] + tm[4]) * out / 2.0);
}
V2 to_image(const float tm[6], double x, double y) {
return {tm[0] * x + tm[1] * y + tm[2], tm[3] * x + tm[4] * y + tm[5]};
}
// OpenCV's CLAHE (4x4 tiles) on a square crop, in place.
void clahe(uint8_t *img, int n, double clip_limit) {
constexpr int kTiles = 4;
const int ts = n / kTiles, area = ts * ts;
const int clip = std::max(1, int(clip_limit * area / 256));
uint8_t lut[kTiles][kTiles][256];
for (int ty = 0; ty < kTiles; ++ty)
for (int tx = 0; tx < kTiles; ++tx) {
int hist[256] = {};
for (int y = ty * ts; y < (ty + 1) * ts; ++y)
for (int x = tx * ts; x < (tx + 1) * ts; ++x) ++hist[img[y * n + x]];
int excess = 0;
for (int &h : hist)
if (h > clip) excess += h - clip, h = clip;
const int add = excess / 256, residual = excess - add * 256;
for (int i = 0; i < 256; ++i) hist[i] += add + (i < residual ? 1 : 0);
int sum = 0;
const float scale = 255.f / area;
for (int i = 0; i < 256; ++i) {
sum += hist[i];
lut[ty][tx][i] = uint8_t(std::min(255, int(sum * scale + 0.5f)));
}
}
std::vector<uint8_t> out(size_t(n) * n);
for (int y = 0; y < n; ++y) {
const float fy = (y + 0.5f) / ts - 0.5f;
const int y0 = std::clamp(int(std::floor(fy)), 0, kTiles - 1), y1 = std::min(y0 + 1, kTiles - 1);
const float wy = std::clamp(fy - y0, 0.f, 1.f);
for (int x = 0; x < n; ++x) {
const float fx = (x + 0.5f) / ts - 0.5f;
const int x0 = std::clamp(int(std::floor(fx)), 0, kTiles - 1), x1 = std::min(x0 + 1, kTiles - 1);
const float wx = std::clamp(fx - x0, 0.f, 1.f);
const uint8_t v = img[y * n + x];
const float top = lut[y0][x0][v] * (1 - wx) + lut[y0][x1][v] * wx;
const float bot = lut[y1][x0][v] * (1 - wx) + lut[y1][x1][v] * wx;
out[size_t(y) * n + x] = uint8_t(top * (1 - wy) + bot * wy + 0.5f);
}
}
std::copy(out.begin(), out.end(), img);
}
// Linear stretch of the 1st..99th percentile to 0..255, in place.
void stretch(uint8_t *img, int n) {
int hist[256] = {};
const int total = n * n;
for (int i = 0; i < total; ++i) ++hist[img[i]];
int lo = 0, hi = 255, acc = 0;
for (int v = 0; v < 256; ++v)
if ((acc += hist[v]) > total / 100) { lo = v; break; }
acc = 0;
for (int v = 255; v >= 0; --v)
if ((acc += hist[v]) > total / 100) { hi = v; break; }
if (hi <= lo) return;
for (int i = 0; i < total; ++i) img[i] = uint8_t(std::clamp((img[i] - lo) * 255 / (hi - lo), 0, 255));
}
// A crop as the models' input: RGB (the mono plane three times), 0..1.
ncnn::Mat crop(const Image &img, const float tm[6], int n, const Contrast &contrast) {
std::vector<uint8_t> patch(size_t(n) * n);
ncnn::warpaffine_bilinear_c1(img.data, img.width, img.height, img.stride, patch.data(), n, n, n, tm, 0, 0);
if (contrast.mode == Contrast::Clahe) clahe(patch.data(), n, contrast.clip);
else if (contrast.mode == Contrast::Stretch) stretch(patch.data(), n);
ncnn::Mat m = ncnn::Mat::from_pixels(patch.data(), ncnn::Mat::PIXEL_GRAY2RGB, n, n);
const float norm[3] = {1 / 255.f, 1 / 255.f, 1 / 255.f};
m.substract_mean_normalize(nullptr, norm);
return m;
}
} // namespace
bool Contrast::parse(const std::string &s, Contrast &out) {
if (s == "none") return out.mode = None, true;
if (s == "stretch") return out.mode = Stretch, true;
if (s.rfind("clahe", 0) == 0) {
out.mode = Clahe;
out.clip = s.size() > 6 && s[5] == ':' ? std::atof(s.c_str() + 6) : 2.0;
return out.clip > 0;
}
return false;
}
bool Contrast::parse_pair(const std::string &s, Contrast &palm, Contrast &hand) {
const size_t slash = s.find('/');
if (slash == std::string::npos) return parse(s, palm) && parse(s, hand);
return parse(s.substr(0, slash), palm) && parse(s.substr(slash + 1), hand);
}
namespace {
bool load_net(ncnn::Net &net, const std::string &base, std::string &err) {
net.opt.num_threads = 1;
net.opt.use_vulkan_compute = false;
net.opt.use_fp16_packed = net.opt.use_fp16_storage = net.opt.use_fp16_arithmetic = true;
if (net.load_param((base + ".param").c_str()) || net.load_model((base + ".bin").c_str())) {
err = "can't load " + base + ".param/.bin";
return false;
}
return true;
}
} // namespace
Roi Palm::roi() const {
const V2 a = kp[0], b = kp[2];
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(b[1] - a[1]), b[0] - a[0]));
const double h = size[1];
const V2 shift{-h * -0.5 * std::sin(rot), h * -0.5 * std::cos(rot)};
return {center + shift, std::max(size[0], size[1]) * 2.6, rot};
}
Roi roi_from_points(const V2 *p) {
const V2 w = p[0];
V2 m = (p[5] + p[13]) * 0.5;
m = (m + p[9]) * 0.5;
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(m[1] - w[1]), m[0] - w[0]));
V2 lo{1e9, 1e9}, hi{-1e9, -1e9};
for (int i : kRoiLandmarks)
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[i][k]), hi[k] = std::max(hi[k], p[i][k]);
V2 center = (lo + hi) * 0.5;
const double c = std::cos(-rot), s = std::sin(-rot);
V2 qlo{1e9, 1e9}, qhi{-1e9, -1e9};
for (int i : kRoiLandmarks) {
const V2 d = p[i] - center;
const V2 q{d[0] * c - d[1] * s, d[0] * s + d[1] * c};
for (int k = 0; k < 2; ++k) qlo[k] = std::min(qlo[k], q[k]), qhi[k] = std::max(qhi[k], q[k]);
}
const V2 mid = (qlo + qhi) * 0.5;
const double c2 = std::cos(rot), s2 = std::sin(rot);
center = center + V2{mid[0] * c2 - mid[1] * s2, mid[0] * s2 + mid[1] * c2};
const double w2 = qhi[0] - qlo[0], h2 = qhi[1] - qlo[1];
center = center + V2{-h2 * -0.1 * s2, h2 * -0.1 * c2};
return {center, std::max(w2, h2) * 2.0, rot};
}
Roi Landmarks::next_roi() const { return roi_from_points(pts); }
bool Nets::load(const std::string &dir, bool int8, std::string &err) {
const std::string suffix = int8 ? "-int8.ncnn" : ".ncnn";
if (!load_net(palm_, dir + "/palm" + suffix, err) || !load_net(hand_, dir + "/hand" + suffix, err)) return false;
// SSD anchors of palm_detection_full: strides 8 (2 per cell) and 16 (6 per cell)
for (auto [stride, per] : {std::pair{8, 2}, std::pair{16, 6}}) {
const int n = kPalmSize / stride;
for (int y = 0; y < n; ++y)
for (int x = 0; x < n; ++x)
for (int k = 0; k < per; ++k) anchors_.push_back({(x + 0.5) / n * kPalmSize, (y + 0.5) / n * kPalmSize});
}
return true;
}
std::vector<Palm> Nets::palms(const Image &img, V2 center, double size, double rotation) const {
float tm[6];
crop_matrix(center, size, rotation, kPalmSize, tm);
ncnn::Extractor ex = palm_.create_extractor();
ex.input("in0", crop(img, tm, kPalmSize, palm_contrast_));
ncnn::Mat boxes, scores;
ex.extract("out0", boxes);
ex.extract("out1", scores);
const float *raw = boxes, *logit = scores;
const int n = int(anchors_.size());
const float min_logit = std::log(0.5f / 0.5f); // score 0.5
struct Cand { V2 c, s; V2 kp[7]; double score; };
std::vector<Cand> cand;
for (int i = 0; i < n; ++i) {
if (logit[i] <= min_logit) continue;
const float *r = raw + i * 18;
Cand c;
c.c = {r[0] + anchors_[i][0], r[1] + anchors_[i][1]};
c.s = {r[2], r[3]};
for (int k = 0; k < 7; ++k) c.kp[k] = {r[4 + 2 * k] + anchors_[i][0], r[5 + 2 * k] + anchors_[i][1]};
c.score = 1 / (1 + std::exp(-std::clamp(double(logit[i]), -100.0, 100.0)));
cand.push_back(c);
}
// MediaPipe's weighted NMS: overlapping boxes are averaged, weighted by score
std::sort(cand.begin(), cand.end(), [](const Cand &a, const Cand &b) { return a.score > b.score; });
std::vector<bool> used(cand.size());
std::vector<Palm> out;
for (size_t i = 0; i < cand.size(); ++i) {
if (used[i]) continue;
double wsum = 0;
Cand acc{};
for (size_t j = i; j < cand.size(); ++j) {
if (used[j]) continue;
const double ix = std::max(0.0, std::min(cand[i].c[0] + cand[i].s[0] / 2, cand[j].c[0] + cand[j].s[0] / 2) -
std::max(cand[i].c[0] - cand[i].s[0] / 2, cand[j].c[0] - cand[j].s[0] / 2));
const double iy = std::max(0.0, std::min(cand[i].c[1] + cand[i].s[1] / 2, cand[j].c[1] + cand[j].s[1] / 2) -
std::max(cand[i].c[1] - cand[i].s[1] / 2, cand[j].c[1] - cand[j].s[1] / 2));
const double inter = ix * iy;
const double uni = cand[i].s[0] * cand[i].s[1] + cand[j].s[0] * cand[j].s[1] - inter;
if (j != i && inter / (uni + 1e-9) <= 0.3) continue;
used[j] = true;
const double w = cand[j].score;
wsum += w;
acc.c = acc.c + cand[j].c * w;
acc.s = acc.s + cand[j].s * w;
for (int k = 0; k < 7; ++k) acc.kp[k] = acc.kp[k] + cand[j].kp[k] * w;
}
Palm p;
const V2 c = acc.c * (1 / wsum);
p.center = to_image(tm, c[0], c[1]);
p.size = acc.s * (1 / wsum * size / kPalmSize);
for (int k = 0; k < 7; ++k) {
const V2 q = acc.kp[k] * (1 / wsum);
p.kp[k] = to_image(tm, q[0], q[1]);
}
p.score = cand[i].score;
out.push_back(p);
}
return out;
}
Landmarks Nets::landmarks(const Image &img, const Roi &roi) const {
float tm[6];
crop_matrix(roi.center, roi.size, roi.rotation, kHandSize, tm);
ncnn::Extractor ex = hand_.create_extractor();
ex.input("in0", crop(img, tm, kHandSize, hand_contrast_));
ncnn::Mat screen, presence, right, world;
ex.extract("out0", screen);
ex.extract("out1", presence);
ex.extract("out2", right);
ex.extract("out3", world);
Landmarks lm;
const float *s = screen, *w = world;
for (int i = 0; i < 21; ++i) {
lm.pts[i] = to_image(tm, s[3 * i], s[3 * i + 1]);
for (int k = 0; k < 3; ++k) lm.world[i][k] = w[3 * i + k];
}
lm.presence = presence[0];
lm.right = right[0];
return lm;
}
+65
View File
@@ -0,0 +1,65 @@
// MediaPipe's palm detector and hand landmark model on ncnn. A crop is a square region of a camera image: centre and size in
// pixels, and a rotation that turns the crop's "up" toward the image direction
// (sin r, -cos r). Crops are contrast-equalized (CLAHE) before the models see them.
// Everything here may run on several threads at once.
#pragma once
#include "geom.h"
#include <net.h>
#include <cstdint>
#include <string>
#include <vector>
struct Image {
const uint8_t *data = nullptr;
int width = 0, height = 0, stride = 0;
};
struct Roi {
V2 center{};
double size = 0, rotation = 0;
};
struct Palm {
V2 center{}, size{};
V2 kp[7]{};
double score = 0;
Roi roi() const; // MediaPipe's hand crop for this palm
};
struct Landmarks {
V2 pts[21]{}; // image pixels
double world[21][3]{}; // MediaPipe's metric landmarks, hand-centred
double presence = 0, right = 0;
Roi next_roi() const; // MediaPipe's crop to track the hand in the next frame
};
Roi roi_from_points(const V2 *pts21);
// How crops are contrast-equalized before the models see them.
struct Contrast {
enum Mode { Clahe, None, Stretch } mode = Clahe;
double clip = 2.0; // Clahe: OpenCV's clip limit (4x4 tiles)
// "clahe:2", "none", "stretch" (1st..99th percentile to 0..255)
static bool parse(const std::string &s, Contrast &out);
// "PALM/HAND" (each as above), or one for both
static bool parse_pair(const std::string &s, Contrast &palm, Contrast &hand);
};
class Nets {
public:
// Loads <dir>/palm.ncnn.* and <dir>/hand.ncnn.*, or the -int8 variants.
bool load(const std::string &dir, bool int8, std::string &err);
std::vector<Palm> palms(const Image &img, V2 center, double size, double rotation) const;
Landmarks landmarks(const Image &img, const Roi &roi) const;
// Before any palms()/landmarks(): how the palm search's and the landmark model's crops
// are equalized.
void set_contrast(const Contrast &palm, const Contrast &hand) { palm_contrast_ = palm, hand_contrast_ = hand; }
private:
Contrast palm_contrast_, hand_contrast_;
ncnn::Net palm_, hand_;
std::vector<V2> anchors_;
};
+306
View File
@@ -0,0 +1,306 @@
#include "pinch.h"
#include "io.h"
#include <fcntl.h>
#include <sys/mman.h>
#include <sys/stat.h>
#include <unistd.h>
#include <algorithm>
#include <cerrno>
#include <cstdlib>
#include <cstring>
namespace {
constexpr int kThumbTip = 4, kIndexTip = 8;
V3 cross(V3 a, V3 b) { return {a[1] * b[2] - a[2] * b[1], a[2] * b[0] - a[0] * b[2], a[0] * b[1] - a[1] * b[0]}; }
void put3(float out[3], V3 v) {
for (int k = 0; k < 3; ++k) out[k] = float(v[k]);
}
} // namespace
void Pinch::update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns,
const std::vector<int> &gripping) {
events.clear();
auto grips = [&](int id) { return std::find(gripping.begin(), gripping.end(), id) != gripping.end(); };
for (int s = 0; s < 2; ++s) // a grip took this pinch's hand: it's a drag now, not a click
if ((side_[s].flags & FH_PINCH_DOWN) && grips(follow_[s])) end(s, t_ns, true);
// A hand a pinch is down on belongs to that side until it ends. The left/right call is a
// running average of the model's, and when it flips mid-pinch the other side would take
// the same hand and pinch too (4 times in the 2026-09-30 lit recording).
int taken[2] = {0, 0};
for (int s = 0; s < 2; ++s)
if (side_[s].flags & FH_PINCH_DOWN) taken[s] = follow_[s];
for (int s = 0; s < 2; ++s) {
fh_pinch_t &o = side_[s];
const bool down = o.flags & FH_PINCH_DOWN;
// the hand: while down, the one the pinch began on; else the best tracked hand of this side
const Hand *h = nullptr;
for (const Hand *c : hands) {
if (down ? c->id != follow_[s] : c->right() != (s == 1) || c->id == taken[1 - s]) continue;
if (!h || c->frames > h->frames) h = c;
}
world_d[s] = tri_d[s] = palm_down[s] = -1;
if (!h) {
o.flags &= ~FH_PINCH_TRACKED;
if (down && (t_ns - seen_ns_[s]) / 1e9 > p_.grace_s) end(s, t_ns, true);
continue;
}
seen_ns_[s] = t_ns;
tri_d[s] = norm(h->pts[kThumbTip] - h->pts[kIndexTip]);
const V3 normal = cross(h->smooth[5] - h->smooth[0], h->smooth[17] - h->smooth[0]);
palm_down[s] = norm(normal) > 0 ? std::fabs(normal[1]) / norm(normal) : 0;
double sum = 0;
int n = 0;
for (const Seen &v : views) {
if (v.hand != h->id) continue;
const V3 a{v.lm.world[kThumbTip][0], v.lm.world[kThumbTip][1], v.lm.world[kThumbTip][2]};
const V3 b{v.lm.world[kIndexTip][0], v.lm.world[kIndexTip][1], v.lm.world[kIndexTip][2]};
sum += norm(a - b), ++n;
}
if (n) world_d[s] = sum / n * h->scale;
const double d = p_.triangulated || world_d[s] < 0 ? tri_d[s] : world_d[s];
// Where the pinch is, for drags: the index and middle knuckles, which hold still while
// the fingers open and close. The point between the tips moved 1-2 cm as a pinch
// opened, so every release dragged the pointer off what it pressed (headset test,
// 2026-09-30).
const V3 point = (h->smooth[5] + h->smooth[9]) * 0.5;
o.flags |= FH_PINCH_TRACKED;
o.hand_id = uint32_t(h->id);
o.distance = float(d);
o.strength = float(std::clamp((p_.end_m - d) / (p_.end_m - p_.begin_m), 0.0, 1.0));
put3(o.point, point);
if (!down) {
// a close held back (palm down) has to open again before a pinch can begin, so
// turning the hand with the fingers still closed doesn't start one
if (d > p_.end_m) held_[s] = false;
if (grips(h->id)) {
// closed: no pinch until it opens
} else if (d < p_.begin_m && !held_[s] && palm_down[s] > p_.palm_down_max) {
held_[s] = true;
++held_back[s];
} else if (d < p_.begin_m && !held_[s]) {
o.flags = (o.flags | FH_PINCH_DOWN) & ~FH_PINCH_LOST;
++o.begins;
o.begin_ns = uint64_t(t_ns);
put3(o.begin_point, point);
follow_[s] = h->id;
open_frames_[s] = 0;
events.push_back({s, "begin", t_ns, d, point});
}
} else if (d > p_.end_m) {
if (++open_frames_[s] >= p_.end_frames) end(s, t_ns, false);
} else {
open_frames_[s] = 0;
}
}
}
void Pinch::end(int s, int64_t t_ns, bool lost) {
fh_pinch_t &o = side_[s];
o.flags = (o.flags & ~FH_PINCH_DOWN) | (lost ? FH_PINCH_LOST : 0);
++o.ends;
o.end_ns = uint64_t(t_ns);
follow_[s] = 0;
events.push_back({s, lost ? "lost" : "end", t_ns, o.distance, {o.point[0], o.point[1], o.point[2]}});
}
void Pinch::release(int64_t t_ns) {
events.clear();
for (int s = 0; s < 2; ++s) {
side_[s].flags &= ~FH_PINCH_TRACKED;
if (side_[s].flags & FH_PINCH_DOWN) end(s, t_ns, true);
}
}
bool Pinch::engaged() const {
for (const fh_pinch_t &o : side_)
if ((o.flags & FH_PINCH_DOWN) || ((o.flags & FH_PINCH_TRACKED) && o.strength > 0.3f)) return true;
return false;
}
// ---------------------------------------------------------------------------- grip
namespace {
constexpr int kWrist = 0, kFingers[4][2] = {{5, 8}, {9, 12}, {13, 16}, {17, 20}}; // knuckle, tip
// Each finger's curl (see GripParams), and the thumb tip's distance from the index tip (m):
// from the model's world landmarks averaged over the hand's views this step, or the
// tracker's 3D points without any.
bool curls(const Hand &h, const std::vector<Seen> &views, double out[4], double *thumb) {
double sum[4] = {}, gap = 0;
int n = 0;
for (const Seen &v : views) {
if (v.hand != h.id) continue;
auto p = [&](int i) { return V3{v.lm.world[i][0], v.lm.world[i][1], v.lm.world[i][2]}; };
for (int f = 0; f < 4; ++f) {
const double k = norm(p(kFingers[f][0]) - p(kWrist));
sum[f] += k > 1e-4 ? norm(p(kFingers[f][1]) - p(kWrist)) / k : 2;
}
gap += norm(p(4) - p(8)) * h.scale;
++n;
}
*thumb = n ? gap / n : norm(h.smooth[4] - h.smooth[8]);
for (int f = 0; f < 4; ++f) {
if (n) {
out[f] = sum[f] / n;
continue;
}
const double k = norm(h.smooth[kFingers[f][0]] - h.smooth[kWrist]);
if (k < 1e-4) return false;
out[f] = norm(h.smooth[kFingers[f][1]] - h.smooth[kWrist]) / k;
}
return true;
}
V3 palm_centre(const Hand &h) {
return (h.smooth[0] + h.smooth[5] + h.smooth[9] + h.smooth[13] + h.smooth[17]) * 0.2;
}
} // namespace
void Grip::update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns) {
events.clear();
int taken[2] = {0, 0};
for (int s = 0; s < 2; ++s)
if (side_[s].flags & FH_PINCH_DOWN) taken[s] = follow_[s];
for (int s = 0; s < 2; ++s) {
fh_pinch_t &o = side_[s];
const bool down = o.flags & FH_PINCH_DOWN;
const Hand *h = nullptr;
for (const Hand *c : hands) {
if (down ? c->id != follow_[s] : c->right() != (s == 1) || c->id == taken[1 - s]) continue;
if (!h || c->frames > h->frames) h = c;
}
curl[s] = -1;
double f[4], thumb = 1;
if (!h || !curls(*h, views, f, &thumb)) {
o.flags &= ~FH_PINCH_TRACKED;
if (down && (t_ns - seen_ns_[s]) / 1e9 > p_.grace_s) end(s, t_ns, true);
continue;
}
seen_ns_[s] = t_ns;
const double mean = (f[0] + f[1] + f[2] + f[3]) / 4, most = std::max({f[0], f[1], f[2], f[3]});
curl[s] = mean;
const V3 point = palm_centre(*h);
o.flags |= FH_PINCH_TRACKED;
o.hand_id = uint32_t(h->id);
o.distance = float(mean);
o.strength = float(std::clamp((p_.end - mean) / (p_.end - p_.begin), 0.0, 1.0));
put3(o.point, point);
if (!down) {
if (mean > p_.end) open_ns_[s] = t_ns;
const bool ahead = -point[2] >= p_.min_ahead_m &&
std::atan2(-point[1], -point[2]) * 180 / M_PI <= p_.max_down_deg;
if (most < p_.begin && ahead && thumb >= p_.thumb_off_m && open_ns_[s] && (t_ns - open_ns_[s]) / 1e9 <= p_.armed_s) {
o.flags = (o.flags | FH_PINCH_DOWN) & ~FH_PINCH_LOST;
++o.begins;
o.begin_ns = uint64_t(t_ns);
put3(o.begin_point, point);
follow_[s] = h->id;
open_frames_[s] = 0;
events.push_back({s, "begin", t_ns, mean, point});
}
} else if (mean > p_.end) {
if (++open_frames_[s] >= p_.end_frames) {
end(s, t_ns, false);
open_ns_[s] = t_ns;
}
} else {
open_frames_[s] = 0;
}
}
}
void Grip::end(int s, int64_t t_ns, bool lost) {
fh_pinch_t &o = side_[s];
o.flags = (o.flags & ~FH_PINCH_DOWN) | (lost ? FH_PINCH_LOST : 0);
++o.ends;
o.end_ns = uint64_t(t_ns);
follow_[s] = 0;
events.push_back({s, lost ? "lost" : "end", t_ns, o.distance, {o.point[0], o.point[1], o.point[2]}});
}
void Grip::release(int64_t t_ns) {
events.clear();
for (int s = 0; s < 2; ++s) {
side_[s].flags &= ~FH_PINCH_TRACKED;
if (side_[s].flags & FH_PINCH_DOWN) end(s, t_ns, true);
}
}
std::vector<int> Grip::gripping() const {
std::vector<int> out;
for (int s = 0; s < 2; ++s)
if (side_[s].flags & FH_PINCH_DOWN) out.push_back(follow_[s]);
return out;
}
bool Grip::engaged() const {
for (const fh_pinch_t &o : side_)
if ((o.flags & FH_PINCH_DOWN) || ((o.flags & FH_PINCH_TRACKED) && o.strength > 0.3f)) return true;
return false;
}
// ----------------------------------------------------------------------- publisher
bool GesturePublisher::open(const Pinch &pinch, const Grip &grip, std::string &err) {
const std::string path = run_dir() + "/gestures";
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
if (fd < 0 || ftruncate(fd, sizeof(fh_gestures_t)) < 0) return err = path + ": " + std::strerror(errno), false;
void *m = mmap(nullptr, sizeof(fh_gestures_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
close(fd);
if (m == MAP_FAILED) return err = path + ": can't map it", false;
out_ = static_cast<fh_gestures_t *>(m);
// keep the counters a previous tracker left, so a reader doesn't see them jump back
const bool ours = !std::memcmp(out_->magic, FH_GESTURES_MAGIC, 8) && out_->version == FH_GESTURES_VERSION;
if (!ours) {
std::memset(out_, 0, sizeof *out_);
std::memcpy(out_->magic, FH_GESTURES_MAGIC, 8);
out_->version = FH_GESTURES_VERSION;
out_->size = sizeof(fh_gestures_t);
}
seq_ = out_->seq / 2 + 1;
// a gesture the last tracker left down (it crashed) is over: count its end, as lost
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
__atomic_thread_fence(__ATOMIC_RELEASE);
for (fh_pinch_t *slots : {out_->pinch, out_->grip})
for (int s = 0; s < 2; ++s) {
fh_pinch_t &o = slots[s];
if (o.begins == o.ends) continue;
o.ends = o.begins;
o.end_ns = mono_ns();
o.flags = (o.flags & ~FH_PINCH_DOWN) | FH_PINCH_LOST;
}
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
out_->begin_m = float(pinch.params().begin_m);
out_->end_m = float(pinch.params().end_m);
out_->grip_begin = float(grip.params().begin);
out_->grip_end = float(grip.params().end);
return true;
}
void GesturePublisher::write(const Pinch &pinch, const Grip &grip, uint64_t capture_ns) {
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
__atomic_thread_fence(__ATOMIC_RELEASE);
for (int g = 0; g < 2; ++g)
for (int s = 0; s < 2; ++s) {
// counters carry on from what's in the file (a restarted tracker starts its own at 0)
const fh_pinch_t &in = g ? grip.side(s) : pinch.side(s);
fh_pinch_t &o = g ? out_->grip[s] : out_->pinch[s];
const uint32_t base_b = o.begins - last_begins_[g][s], base_e = o.ends - last_ends_[g][s];
o = in;
o.begins = base_b + in.begins;
o.ends = base_e + in.ends;
last_begins_[g][s] = in.begins, last_ends_[g][s] = in.ends;
}
out_->capture_ns = capture_ns;
out_->publish_ns = mono_ns();
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
}
+135
View File
@@ -0,0 +1,135 @@
// Gesture detection for input: look at something and pinch to click it (pinch and move to
// nudge the pointer first), or close the hand to press and drag it. Per side, from the
// tracker's hands after each step; published as fh_gestures.h.
#pragma once
#include "tracker.h"
#include <string>
#include <vector>
extern "C" {
#include "../include/fh_gestures.h"
}
struct PinchParams {
double begin_m = 0.020; // thumb and index tips closer than this: the pinch begins
double end_m = 0.035; // further apart than this: it ends (the gap keeps it from flickering)
int end_frames = 2; // processed frames in a row past end_m before it ends, so one
// noisy frame doesn't drop a drag
double grace_s = 0.25; // a pinching hand lost this long ends its pinch (FH_PINCH_LOST)
// Where the distance comes from: MediaPipe's world landmarks (the model's own 3D hand
// pose, averaged over the hand's views, at the user's hand size), or the tracker's
// triangulated tips. The model's pose should hold up better when the fingers hide each
// other; tomorrow's recordings will tell.
bool triangulated = false;
// No pinch begins while the palm faces down more than this (|palm normal . up| in the
// head frame; 1 turns it off). Typing curls the thumb onto the index: in the 2026-09-30
// lit recording, pinches that began while typing had 0.69-1.00, deliberate ones 0.00-0.50.
// Looking down tilts the head frame, which lowers the reading for a hand on a keyboard.
// Off by default since the first headset test (2026-09-30 14:55): the user's deliberate
// pinches, hand raised in front, read 0.90-0.99 too. The pointer helper now leaves out
// gestures that begin low (hands on a desk), which it can tell with the head's pose.
double palm_down_max = 1.0;
};
class Pinch {
public:
explicit Pinch(const PinchParams &p = {}) : p_(p) {}
const PinchParams &params() const { return p_; }
// After each processed set: the hands out of Tracker::step, the tracker's views (for
// the world landmarks) and the capture time. Hands in `gripping` (their ids) are closed:
// no pinch begins on them, and one that's down on them ends, as lost.
void update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns,
const std::vector<int> &gripping = {});
// Ends any pinch that's down (as lost), e.g. when the tracker stops.
void release(int64_t t_ns);
const fh_pinch_t &side(int s) const { return side_[s]; } // 0 left, 1 right
// A pinch is down or closing: worth tracking at the full rate.
bool engaged() const;
// What changed in the last update, for logs.
struct Event {
int side;
const char *what; // "begin", "end", "lost"
int64_t t_ns;
double distance;
V3 point;
};
std::vector<Event> events;
// Both distance measures for the last update, per side (-1: no hand), for logs.
double world_d[2] = {-1, -1}, tri_d[2] = {-1, -1};
double palm_down[2] = {-1, -1}; // |palm normal . up| of each side's hand
int held_back[2] = {0, 0}; // pinches that didn't begin because the palm faced down
private:
void end(int s, int64_t t_ns, bool lost);
PinchParams p_;
fh_pinch_t side_[2]{};
int follow_[2] = {0, 0}; // the hand id a pinch follows while down
int open_frames_[2] = {0, 0};
int64_t seen_ns_[2] = {0, 0};
bool held_[2] = {false, false}; // a close held back (palm down) that hasn't opened yet
};
struct GripParams {
// How curled a finger is: its tip's distance from the wrist over its knuckle's, from the
// model's world landmarks (so the hand's size doesn't matter). About 1.8-2.0 straight,
// 0.8-1.0 curled into a fist.
double begin = 1.2; // every finger under this: the grip begins
double end = 1.45; // their mean over this: it ends
int end_frames = 2; // processed frames in a row past end before it ends
double grace_s = 0.25; // a gripping hand lost this long ends its grip (FH_PINCH_LOST)
// A grip begins only on a hand seen open (mean over end) within this long: closing the
// hand is the gesture. A hand resting closed (in your lap, on a mouse) never grips.
double armed_s = 1.0;
// ...and only in front of you, where you'd hold a hand up to grab something: the palm no
// more than max_down_deg below straight ahead (head frame) and at least min_ahead_m in
// front of the eyes. Typing curls the fingers like a loose fist: in the 2026-09-30 lit
// recording, typing hands sat about 47 degrees down (14 false grips without this),
// deliberate pinches 5-15.
double max_down_deg = 35;
double min_ahead_m = 0.15;
// ...and not with the thumb on the index fingertip, closer than this (m, the pinch's
// measure): that's a pinch with the other fingers curled, which the first headset test
// took for a grip.
double thumb_off_m = 0.03;
};
// Grip (a closed hand) detection, per side like Pinch: press and drag.
class Grip {
public:
explicit Grip(const GripParams &p = {}) : p_(p) {}
const GripParams &params() const { return p_; }
void update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns);
void release(int64_t t_ns);
const fh_pinch_t &side(int s) const { return side_[s]; }
// The hands gripping now (for Pinch::update).
std::vector<int> gripping() const;
bool engaged() const; // a grip is down or closing: worth tracking at the full rate
std::vector<Pinch::Event> events;
double curl[2] = {-1, -1}; // each side's hand: mean curl, for logs
private:
void end(int s, int64_t t_ns, bool lost);
GripParams p_;
fh_pinch_t side_[2]{};
int follow_[2] = {0, 0};
int open_frames_[2] = {0, 0};
int64_t seen_ns_[2] = {0, 0};
int64_t open_ns_[2] = {0, 0}; // the side's hand was last seen open then
};
// Writes /run/user/UID/frametop-hands/gestures.
class GesturePublisher {
public:
bool open(const Pinch &pinch, const Grip &grip, std::string &err);
void write(const Pinch &pinch, const Grip &grip, uint64_t capture_ns);
private:
fh_gestures_t *out_ = nullptr;
uint64_t seq_ = 0;
// the counters last written, per gesture (0 pinch, 1 grip) and side
uint32_t last_begins_[2][2] = {}, last_ends_[2][2] = {};
};
+101
View File
@@ -0,0 +1,101 @@
#include "record.h"
#include <sys/stat.h>
#include <cerrno>
#include <cstring>
namespace {
constexpr size_t kMaxQueued = 48; // about 130 MB of sets
}
Recorder::~Recorder() {
if (!f_) return;
{
std::lock_guard<std::mutex> l(mu_);
stop_ = true;
}
wake_.notify_all();
thread_.join();
std::fclose(f_);
}
bool Recorder::open(const std::string &dir, std::string &err) {
if (mkdir(dir.c_str(), 0755) < 0 && errno != EEXIST) return err = dir + ": " + std::strerror(errno), false;
const std::string path = dir + "/sets.bin";
f_ = std::fopen(path.c_str(), "wbx"); // never overwrite a recording
if (!f_) return err = path + ": " + std::strerror(errno), false;
thread_ = std::thread(&Recorder::loop, this);
return true;
}
void Recorder::add(const std::vector<SetFrame> &frames) {
size_t bytes = sizeof(fh_set_hdr_t) + frames.size() * sizeof(fh_set_cam_t);
for (const SetFrame &s : frames) bytes += size_t(s.width) * s.height;
std::vector<uint8_t> rec(bytes);
fh_set_hdr_t h{};
std::memcpy(h.magic, FH_SET_MAGIC, 8);
h.ncams = uint32_t(frames.size());
h.bytes = uint32_t(bytes);
std::memcpy(rec.data(), &h, sizeof h);
uint8_t *p = rec.data() + sizeof h;
for (const SetFrame &s : frames) {
fh_set_cam_t c{};
std::strncpy(c.name, s.name.c_str(), sizeof c.name - 1);
c.width = s.width, c.height = s.height, c.capture_ns = s.capture_ns, c.dqbuf_ns = s.dqbuf_ns;
std::memcpy(p, &c, sizeof c);
p += sizeof c;
}
for (const SetFrame &s : frames) {
std::memcpy(p, s.px, size_t(s.width) * s.height);
p += size_t(s.width) * s.height;
}
{
std::lock_guard<std::mutex> l(mu_);
if (queue_.size() >= kMaxQueued) {
++dropped_;
return;
}
queue_.push_back(std::move(rec));
}
wake_.notify_one();
}
void Recorder::loop() {
std::unique_lock<std::mutex> l(mu_);
for (;;) {
wake_.wait(l, [&] { return stop_ || !queue_.empty(); });
if (queue_.empty()) return; // stopping, and everything is written
std::vector<uint8_t> rec = std::move(queue_.front());
queue_.pop_front();
l.unlock();
const bool ok = std::fwrite(rec.data(), 1, rec.size(), f_) == rec.size();
l.lock();
ok ? ++written_ : ++dropped_;
}
}
SetReader::~SetReader() {
if (f_) std::fclose(f_);
}
bool SetReader::open(const std::string &dir, std::string &err) {
const std::string path = dir + "/sets.bin";
f_ = std::fopen(path.c_str(), "rb");
return f_ ? true : (err = path + ": " + std::strerror(errno), false);
}
bool SetReader::next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels) {
fh_set_hdr_t h;
if (std::fread(&h, sizeof h, 1, f_) != 1 || std::memcmp(h.magic, FH_SET_MAGIC, 8) || h.ncams == 0 || h.ncams > 16)
return false;
cams.resize(h.ncams);
if (std::fread(cams.data(), sizeof(fh_set_cam_t), h.ncams, f_) != h.ncams) return false;
pixels.resize(h.ncams);
for (uint32_t i = 0; i < h.ncams; ++i) {
cams[i].name[sizeof cams[i].name - 1] = 0;
pixels[i].resize(size_t(cams[i].width) * cams[i].height);
if (std::fread(pixels[i].data(), 1, pixels[i].size(), f_) != pixels[i].size()) return false;
}
return true;
}
+69
View File
@@ -0,0 +1,69 @@
// Recordings of frame sets, for replaying live sessions through the tracker offline
// (ft-handreplay). A recording is DIR/sets.bin: one record per frame set, each
// fh_set_hdr_t, then per camera fh_set_cam_t, then each camera's pixels (w x h, packed)
// in the same camera order.
#pragma once
#include <condition_variable>
#include <cstdint>
#include <cstdio>
#include <deque>
#include <mutex>
#include <string>
#include <thread>
#include <vector>
#define FH_SET_MAGIC "FHSET01"
struct fh_set_hdr_t {
char magic[8];
uint32_t ncams;
uint32_t bytes; // the whole record, this header included
};
struct fh_set_cam_t {
char name[16]; // calibration name, e.g. "slam_left"
uint32_t width, height;
uint64_t capture_ns; // CLOCK_MONOTONIC_RAW, as the ring has it
uint64_t dqbuf_ns; // CLOCK_MONOTONIC
};
struct SetFrame {
std::string name;
const uint8_t *px;
uint32_t width, height;
uint64_t capture_ns, dqbuf_ns;
};
// Writes sets on its own thread, so a slow disk never holds up tracking; drops sets
// when too many are waiting.
class Recorder {
public:
~Recorder();
bool open(const std::string &dir, std::string &err);
void add(const std::vector<SetFrame> &frames);
size_t written() const { return written_; }
size_t dropped() const { return dropped_; }
private:
void loop();
FILE *f_ = nullptr;
std::thread thread_;
std::mutex mu_;
std::condition_variable wake_;
std::deque<std::vector<uint8_t>> queue_;
bool stop_ = false;
size_t written_ = 0, dropped_ = 0;
};
// Reads a recording back one set at a time.
class SetReader {
public:
bool open(const std::string &dir, std::string &err);
// False at the end (or on a truncated last set).
bool next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels);
~SetReader();
private:
FILE *f_ = nullptr;
};
+374
View File
@@ -0,0 +1,374 @@
// ft-handreplay: run a recording (ft-hands --record) through the tracker offline, with the
// live scheduling, and report how well it kept the hands.
//
// ft-handreplay DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR]
// [--from S] [--to S] [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch)
//
// --oracle N: every N-th set, also search every tile of every camera (slow), to see
// which hands were there to find. Compares that with what the tracker had.
// --slow F: the live tracker skips the sets that arrive while it's busy; replay takes
// each step's time here times F as the busy time (the headset is busier live).
// --cost: instead of timing the steps, charge each round of model calls what it
// typically costs live (10 ms landmarks, 18 ms palms): repeatable results.
// --timeline: per processed set, a line per hand (time, id, side, views, wrist) and per view
// (hand, camera, presence, next crop, set index).
// --keep-presence P: landmark presence a tracked view needs to stay (default 0.5, as new ones).
// --pinch-begin M, --pinch-end M, --pinch-triangulated, --pinch-palm-down MAX: the pinch detector (track/pinch.h);
// the timeline gets its begin/end/lost events and both distance measures per set.
// --grip-begin R, --grip-end R: the grip detector (a closed hand; track/pinch.h); the timeline
// gets its events and each side's finger curl per set.
// --cams mono|color|all: which cameras to track with (default mono). color and all need a
// recording made with ft-camd --with-color; --color-left NODE (color_video0 or
// color_video3) and --color-crop subtract|none say how its calibration maps
// (tools/check_color.py).
// --contrast: how the palm search's and the landmark model's crops are equalized
// (default clahe:2/none, as ft-hands).
// --poses FILE: per processed set, a line per hand: time, id, the model's left/right call,
// views, hand scale, then its 21 world landmarks (the model's own 3D pose, averaged
// over its views, times the scale; metres, hand-centred) and its 21 published
// points (head frame). For studying gestures (pinch against typing, a fist).
// --depth FILE: per processed set, a line per hand for tools/depth_report.py: its views'
// cameras, triangulation residual, hand scale, measured and published palm, and
// each view's one-view palm (Tracker::single_view at the hand's scale). The
// header has each camera's centre and focal length.
#include "pinch.h"
#include "record.h"
#include "tracker.h"
#include <algorithm>
#include <chrono>
#include <cstdio>
#include <cstring>
#include <map>
#include <set>
#include <string>
#include <vector>
namespace {
struct Track {
double first = 0, last = 0;
int sets = 0, left = 0;
// the last two palm positions (raw, smoothed) and times, for the jitter measure
V3 raw[2]{}, sm[2]{};
double t[2]{};
int line = 0; // updates on the current unbroken run
};
double median(std::vector<double> v) {
if (v.empty()) return 0;
std::nth_element(v.begin(), v.begin() + v.size() / 2, v.end());
return v[v.size() / 2];
}
} // namespace
int main(int argc, char **argv) {
if (argc < 2 || argv[1][0] == '-') {
std::printf("usage: %s DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR] [--from S] [--to S]\n", argv[0]);
return 1;
}
const std::string dir = argv[1];
int oracle = 0, threads = 2;
double slow = 1.0, from = 0, to = 1e9;
bool cost = false;
Contrast palm_contrast, hand_contrast{Contrast::None}; // as ft-hands's
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
PinchParams pinch_params;
GripParams grip_params;
std::string use = "mono", color_left = "color_video0", color_crop = "subtract";
std::string timeline, depth, poses, models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
for (int i = 2; i < argc; ++i) {
const std::string a = argv[i];
const bool more = i + 1 < argc;
if (a == "--oracle" && more) oracle = std::atoi(argv[++i]);
else if (a == "--slow" && more) slow = std::atof(argv[++i]);
else if (a == "--timeline" && more) timeline = argv[++i];
else if (a == "--depth" && more) depth = argv[++i];
else if (a == "--poses" && more) poses = argv[++i];
else if (a == "--threads" && more) threads = std::atoi(argv[++i]);
else if (a == "--models" && more) models = argv[++i];
else if (a == "--cost") cost = true;
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
else if (a == "--cams" && more) use = argv[++i];
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
else if (a == "--pinch-palm-down" && more) pinch_params.palm_down_max = std::atof(argv[++i]);
else if (a == "--grip-begin" && more) grip_params.begin = std::atof(argv[++i]);
else if (a == "--grip-end" && more) grip_params.end = std::atof(argv[++i]);
else if (a == "--color-left" && more) color_left = argv[++i];
else if (a == "--color-crop" && more) color_crop = argv[++i];
else if (a == "--contrast" && more) {
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
}
else if (a == "--from" && more) from = std::atof(argv[++i]);
else if (a == "--to" && more) to = std::atof(argv[++i]);
else return std::fprintf(stderr, "unknown option %s\n", a.c_str()), 1;
}
std::string err;
std::map<std::string, Camera> calib;
Nets nets;
SetReader in;
if (!load_calibration(calib, err) || !nets.load(models, false, err) || !in.open(dir, err))
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
nets.set_contrast(palm_contrast, hand_contrast);
FILE *tl = timeline.empty() ? nullptr : std::fopen(timeline.c_str(), "w");
std::vector<fh_set_cam_t> cams;
std::vector<std::vector<uint8_t>> px;
if (!in.next(cams, px)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
if (use != "mono" && use != "color" && use != "all") return std::fprintf(stderr, "--cams mono|color|all\n"), 1;
if (use != "mono") {
std::vector<std::string> nodes;
for (auto &c : cams)
if (std::string(c.name).rfind("color_video", 0) == 0) nodes.push_back(c.name);
if (nodes.size() != 2) return std::fprintf(stderr, "%s: no color cameras (ft-camd --with-color)\n", dir.c_str()), 1;
const std::string right = nodes[0] == color_left ? nodes[1] : nodes[0];
if (!load_color_calibration(calib, color_left, right, color_crop == "subtract", 2, err))
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
}
std::map<std::string, Camera> used;
for (auto &c : cams) {
const bool color = std::string(c.name).rfind("color_", 0) == 0;
if (calib.count(c.name) && (use == "all" || color == (use == "color"))) used[c.name] = calib[c.name];
}
Pool pool(threads, {2, 3, 4});
Tracker tracker(used, nets, pool);
tracker.set_keep_presence(keep_presence);
FILE *dp = depth.empty() ? nullptr : std::fopen(depth.c_str(), "w");
FILE *pp = poses.empty() ? nullptr : std::fopen(poses.c_str(), "w");
if (dp)
for (auto &[name, c] : used)
std::fprintf(dp, "# cam %s %.4f %.4f %.4f %.1f\n", name.c_str(), c.origin[0], c.origin[1], c.origin[2], c.fx);
uint64_t t0 = 0, busy_until = 0, next_ns = 0, t_prev = 0;
int index = -1; // of the set in the recording
int nsets = 0, processed = 0, left = 0, right = 0, both = 0, hist[3] = {};
std::map<int, Track> tracks;
std::vector<const Hand *> last_out;
// oracle: sets where a side's hand was findable, and where the tracker had it then
int o_sets = 0, o_left = 0, o_right = 0, o_left_hit = 0, o_right_hit = 0, o_left_extra = 0, o_right_extra = 0;
std::map<std::string, int> o_by_cam;
double busy_ms = 0;
// jitter: how far each update's palm is from a straight line through the last two,
// mm (steady motion cancels out; what's left is noise and real acceleration)
std::vector<double> jit_raw, jit_sm;
int near_face = 0, hand_updates = 0; // published palms within 20 cm of the eyes
Pinch pinch(pinch_params);
double pinch_begin_ts[2] = {0, 0};
std::vector<double> pinch_len[2]; // seconds, per side
int pinch_lost = 0;
Grip grip(grip_params);
double grip_begin_ts[2] = {0, 0};
std::vector<double> grip_len[2];
do {
std::map<std::string, Image> images;
// the set's time: the mono cameras' when they're used (the color ones run on another
// clock); color frames can repeat across sets, so a set that doesn't move time on is skipped
uint64_t t = UINT64_MAX, t_color = UINT64_MAX;
for (size_t i = 0; i < cams.size(); ++i) {
if (!used.count(cams[i].name)) continue;
images[cams[i].name] = {px[i].data(), int(cams[i].width), int(cams[i].height), int(cams[i].width)};
uint64_t &ti = std::string(cams[i].name).rfind("color_", 0) == 0 ? t_color : t;
ti = std::min(ti, cams[i].capture_ns);
}
if (t == UINT64_MAX) t = t_color;
if (t <= t_prev) {
++index;
continue;
}
t_prev = t;
if (!t0) t0 = t;
const double ts = (t - t0) / 1e9;
++index;
if (ts < from) continue;
if (ts > to) break;
++nsets;
if (t >= busy_until && t >= next_ns) {
const auto w0 = std::chrono::steady_clock::now();
const Stats before = tracker.stats;
const auto out = tracker.step(images, int64_t(t));
double ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - w0).count();
if (cost) { // repeatable: rounds of model calls at typical live costs, per thread
const int hands = tracker.stats.hand_calls - before.hand_calls, palms = tracker.stats.palm_calls - before.palm_calls;
ms = (2 + 10.0 * ((hands + threads - 1) / threads) + 18.0 * ((palms + threads - 1) / threads)) / slow;
}
busy_ms += ms;
busy_until = t + uint64_t(ms * slow * 1e6) + 3'000'000; // + the ring hand-off
const std::vector<Seen> seen = tracker.views_now();
grip.update(out, seen, int64_t(t));
pinch.update(out, seen, int64_t(t), grip.gripping());
next_ns = t + uint64_t((std::min(tracker.interval(), pinch.engaged() || grip.engaged() ? 1 / 30.0 : 1.0) - 0.005) * 1e9);
for (const Pinch::Event &e : grip.events) {
if (std::string(e.what) == "begin") grip_begin_ts[e.side] = ts;
else grip_len[e.side].push_back(ts - grip_begin_ts[e.side]);
if (tl) std::fprintf(tl, "%.3f grip %s %s curl %.2f point %+.3f %+.3f %+.3f hand %u\n", ts, e.side ? "R" : "L",
e.what, e.distance, e.point[0], e.point[1], e.point[2], grip.side(e.side).hand_id);
}
if (tl && (grip.curl[0] >= 0 || grip.curl[1] >= 0))
std::fprintf(tl, "%.3f curl L %.2f R %.2f\n", ts, grip.curl[0], grip.curl[1]);
for (const Pinch::Event &e : pinch.events) {
if (std::string(e.what) == "begin") pinch_begin_ts[e.side] = ts;
else pinch_len[e.side].push_back(ts - pinch_begin_ts[e.side]), pinch_lost += std::string(e.what) == "lost";
if (tl) std::fprintf(tl, "%.3f pinch %s %s d %.3f point %+.3f %+.3f %+.3f hand %u\n", ts, e.side ? "R" : "L",
e.what, e.distance, e.point[0], e.point[1], e.point[2], pinch.side(e.side).hand_id);
}
if (tl && (pinch.world_d[0] >= 0 || pinch.world_d[1] >= 0)) // both measures, for choosing one
std::fprintf(tl, "%.3f pinchd L world %.3f tri %.3f R world %.3f tri %.3f palm %.2f %.2f\n", ts,
pinch.world_d[0], pinch.tri_d[0], pinch.world_d[1], pinch.tri_d[1], pinch.palm_down[0],
pinch.palm_down[1]);
++processed;
last_out = out;
bool l = false, r = false;
for (const Hand *h : out) {
(h->pts[0][0] < 0 ? l : r) = true;
Track &tr = tracks[h->id];
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
const V3 raw = palm(h->pts), sm = palm(h->smooth);
++hand_updates, near_face += norm(sm) < 0.2;
if (tr.line && ts - tr.t[0] >= 0.1) tr.line = 0; // a gap: the line starts over
if (tr.line >= 2 && tr.t[0] - tr.t[1] > 1e-3) {
const double k = (ts - tr.t[0]) / (tr.t[0] - tr.t[1]);
jit_raw.push_back(norm(raw - tr.raw[0] - (tr.raw[0] - tr.raw[1]) * k) * 1000);
jit_sm.push_back(norm(sm - tr.sm[0] - (tr.sm[0] - tr.sm[1]) * k) * 1000);
}
tr.raw[1] = tr.raw[0], tr.sm[1] = tr.sm[0], tr.t[1] = tr.t[0];
tr.raw[0] = raw, tr.sm[0] = sm, tr.t[0] = ts;
++tr.line;
if (!tr.sets) tr.first = ts;
tr.last = ts, ++tr.sets, tr.left += h->pts[0][0] < 0;
if (tl)
std::fprintf(tl, "%.3f %d %s %d %+.3f %+.3f %+.3f\n", ts, h->id, h->pts[0][0] < 0 ? "L" : "R", h->nviews,
h->pts[0][0], h->pts[0][1], h->pts[0][2]);
if (pp) {
double world[21][3] = {};
int n = 0;
for (const Seen &v : seen)
if (v.hand == h->id) {
for (int k = 0; k < 21; ++k)
for (int j = 0; j < 3; ++j) world[k][j] += v.lm.world[k][j];
++n;
}
std::fprintf(pp, "%.4f %d %s %d %.3f", ts, h->id, h->right() ? "R" : "L", h->nviews, h->scale);
for (int k = 0; k < 21; ++k)
for (int j = 0; j < 3; ++j) std::fprintf(pp, " %.4f", n ? world[k][j] / n * h->scale : NAN);
for (int k = 0; k < 21; ++k)
for (int j = 0; j < 3; ++j) std::fprintf(pp, " %.4f", h->smooth[k][j]);
std::fputc('\n', pp);
}
if (dp) {
std::vector<const Seen *> vs;
for (const Seen &v : seen)
if (v.hand == h->id) vs.push_back(&v);
std::sort(vs.begin(), vs.end(), [](const Seen *a, const Seen *b) { return a->cam < b->cam; });
std::string names;
for (const Seen *v : vs) names += (names.empty() ? "" : "+") + v->cam;
std::fprintf(dp, "%.4f %d %s %d %s %.4f %.3f %.4f %.4f %.4f %.4f %.4f %.4f", ts, h->id,
h->pts[0][0] < 0 ? "L" : "R", h->nviews, names.empty() ? "-" : names.c_str(), h->residual,
h->scale, raw[0], raw[1], raw[2], sm[0], sm[1], sm[2]);
for (const Seen *v : vs) {
V3 mono[21];
const bool ok = tracker.single_view(used.at(v->cam), v->lm, h->scale, mono);
const V3 p = ok ? palm(mono) : V3{NAN, NAN, NAN};
std::fprintf(dp, " %s %.2f %.4f %.4f %.4f", v->cam.c_str(), v->lm.presence, p[0], p[1], p[2]);
}
std::fputc('\n', dp);
}
}
if (tl && out.empty()) std::fprintf(tl, "%.3f -\n", ts);
if (tl)
for (const Seen &v : seen)
std::fprintf(tl, "%.3f view %d %s presence %.2f roi %.0f %.0f %.0f %.3f set %d\n", ts, v.hand, v.cam.c_str(),
v.lm.presence, v.roi.center[0], v.roi.center[1], v.roi.size, v.roi.rotation, index);
left += l, right += r, both += l && r;
++hist[std::min<size_t>(out.size(), 2)];
}
if (oracle > 0 && nsets % oracle == 0) {
const Stats keep = tracker.stats;
const auto seen = tracker.exhaustive(images);
tracker.stats = keep;
bool l = false, r = false;
for (const Seen &s : seen) {
(s.wrist[0] < 0 ? l : r) = true;
++o_by_cam[s.cam + (s.wrist[0] < 0 ? " L" : " R")];
}
bool tl_ = false, tr_ = false;
for (const Hand *h : last_out) (h->pts[0][0] < 0 ? tl_ : tr_) = true;
++o_sets;
o_left += l, o_right += r;
o_left_hit += l && tl_, o_right_hit += r && tr_;
o_left_extra += !l && tl_, o_right_extra += !r && tr_;
if (tl && ((!l && tl_) || (!r && tr_))) std::fprintf(tl, "%.3f oracle-extra %s%s set %d\n", ts, !l && tl_ ? "L" : "", !r && tr_ ? "R" : "", index);
}
} while (in.next(cams, px));
if (tl) std::fclose(tl);
if (dp) std::fclose(dp);
if (pp) std::fclose(pp);
const double secs = nsets > 1 ? nsets / 30.0 : 0;
const Stats &s = tracker.stats;
std::printf("%s: %d sets (%.0f s), processed %d (%.1f/s), %.1f ms per step\n", dir.c_str(), nsets, secs, processed,
processed / std::max(secs, 1e-9), busy_ms / std::max(processed, 1));
std::printf("hands per processed set: 0 %.0f%%, 1 %.0f%%, 2 %.0f%%; a hand on the left %.0f%%, right %.0f%%, both %.0f%%\n",
100.0 * hist[0] / processed, 100.0 * hist[1] / processed, 100.0 * hist[2] / processed,
100.0 * left / processed, 100.0 * right / processed, 100.0 * both / processed);
std::vector<double> lens[2];
for (auto &[id, tr] : tracks) lens[tr.left * 2 > tr.sets ? 0 : 1].push_back(tr.last - tr.first);
for (int k = 0; k < 2; ++k) {
double total = 0;
for (double d : lens[k]) total += d;
std::printf("%s tracks: %zu, median %.1f s, total %.0f s\n", k ? "right" : "left ", lens[k].size(), median(lens[k]), total);
}
std::printf("views lost %d, handoff misses %d, dups %d, splits %d; hands new %d, merged %d, forgotten %d\n", s.lost,
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten);
std::printf("model calls: palm %d (%.1f/s), hand %d (%.1f/s)\n", s.palm_calls, s.palm_calls / std::max(secs, 1e-9),
s.hand_calls, s.hand_calls / std::max(secs, 1e-9));
{
std::vector<double> r, step;
std::map<int, double> prev;
for (auto &[id, x] : s.mono_ratio) {
r.push_back(x);
if (prev.count(id)) step.push_back(std::fabs(x - prev[id]));
prev[id] = x;
}
std::sort(r.begin(), r.end());
std::sort(step.begin(), step.end());
if (!r.empty())
std::printf("single-view distance / stereo: 10%% %.2f, median %.2f, 90%% %.2f; change between frames median %.3f, 90%% %.3f\n",
r[r.size() / 10], r[r.size() / 2], r[r.size() * 9 / 10], step[step.size() / 2], step[step.size() * 9 / 10]);
}
std::printf("palms within 20 cm of the eyes: %d of %d hand updates\n", near_face, hand_updates);
for (int k = 0; k < 2; ++k) std::sort(pinch_len[k].begin(), pinch_len[k].end());
std::printf("pinches (%s, %.3f/%.3f m, palm down under %.2f): left %zu (median %.2f s), right %zu (median %.2f s), "
"%d ended by losing the hand, held back (palm down) left %d right %d\n",
pinch_params.triangulated ? "triangulated tips" : "world landmarks", pinch_params.begin_m, pinch_params.end_m,
pinch_params.palm_down_max,
pinch_len[0].size(), pinch_len[0].empty() ? 0 : pinch_len[0][pinch_len[0].size() / 2], pinch_len[1].size(),
pinch_len[1].empty() ? 0 : pinch_len[1][pinch_len[1].size() / 2], pinch_lost, pinch.held_back[0],
pinch.held_back[1]);
for (int k = 0; k < 2; ++k) std::sort(grip_len[k].begin(), grip_len[k].end());
std::printf("grips (curl under %.2f, open over %.2f): left %zu (median %.2f s), right %zu (median %.2f s)\n",
grip_params.begin, grip_params.end, grip_len[0].size(),
grip_len[0].empty() ? 0 : grip_len[0][grip_len[0].size() / 2], grip_len[1].size(),
grip_len[1].empty() ? 0 : grip_len[1][grip_len[1].size() / 2]);
std::sort(jit_raw.begin(), jit_raw.end());
std::sort(jit_sm.begin(), jit_sm.end());
if (!jit_raw.empty())
std::printf("palm jitter (off a straight line through the last two updates): measured median %.1f mm, 90%% %.1f mm; "
"published median %.1f mm, 90%% %.1f mm\n", jit_raw[jit_raw.size() / 2], jit_raw[jit_raw.size() * 9 / 10],
jit_sm[jit_sm.size() / 2], jit_sm[jit_sm.size() * 9 / 10]);
if (o_sets) {
std::printf("oracle, %d sets: a left hand findable in %d, the tracker had it in %d (%.0f%%); right %d, had %d (%.0f%%)\n",
o_sets, o_left, o_left_hit, 100.0 * o_left_hit / std::max(o_left, 1), o_right, o_right_hit,
100.0 * o_right_hit / std::max(o_right, 1));
std::printf(" tracker had a hand the full search didn't find: left %d, right %d\n", o_left_extra, o_right_extra);
std::printf(" found by camera:");
for (auto &[k, n] : o_by_cam) std::printf(" %s %d", k.c_str(), n);
std::printf("\n");
}
return 0;
}
+233
View File
@@ -0,0 +1,233 @@
// ft-ringplay: play a recording (ft-hands --record) into a frame ring in real time, the
// way ft-camd publishes live cameras, so ft-hands --ring PATH processes the same frames
// run after run. For A/B tests of how the tracker runs.
//
// ft-ringplay DIR --ring PATH [--from S] [--to S] [--loop] [--cpus 0,1]
//
// Frames are stamped as they're published, so the tracker's latency figures stay
// meaningful. Cameras carry their calibration name and no device node (ft-hands maps
// them by name), so a recording made with the right names needs no --swap-sides.
// Dark frames (<name>_dk) are skipped. Needs no root: the ring is an ordinary file.
#include "record.h"
extern "C" {
#include "../camd/fhring.h"
}
#include <fcntl.h>
#include <sched.h>
#include <sys/mman.h>
#include <unistd.h>
#include <algorithm>
#include <cerrno>
#include <csignal>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <ctime>
#include <string>
#include <vector>
namespace {
volatile std::sig_atomic_t g_stop = 0;
uint64_t clock_ns(clockid_t id) {
timespec ts;
clock_gettime(id, &ts);
return uint64_t(ts.tv_sec) * 1'000'000'000ull + uint64_t(ts.tv_nsec);
}
// Sets from a recording, reading only the pixels of the cameras that get published.
class Reader {
public:
bool open(const std::string &path) {
f_ = std::fopen(path.c_str(), "rb");
if (f_) posix_fadvise(fileno(f_), 0, 0, POSIX_FADV_SEQUENTIAL);
return f_ != nullptr;
}
void rewind() { std::fseek(f_, 0, SEEK_SET); }
// False at the end. cams: every camera in the set; px[k]: pixels of camera k when
// want(name), else left empty.
template <class Want>
bool next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &px, Want want) {
fh_set_hdr_t h;
if (std::fread(&h, sizeof h, 1, f_) != 1 || std::memcmp(h.magic, FH_SET_MAGIC, 8) || h.ncams == 0 || h.ncams > 16)
return false;
cams.resize(h.ncams);
if (std::fread(cams.data(), sizeof(fh_set_cam_t), h.ncams, f_) != h.ncams) return false;
px.resize(h.ncams);
for (uint32_t k = 0; k < h.ncams; ++k) {
cams[k].name[sizeof cams[k].name - 1] = 0;
const size_t n = size_t(cams[k].width) * cams[k].height;
if (want(cams[k].name)) {
px[k].resize(n);
if (std::fread(px[k].data(), 1, n, f_) != n) return false;
} else {
px[k].clear();
if (std::fseek(f_, long(n), SEEK_CUR)) return false;
}
}
if (++sets_ % 64 == 0) posix_fadvise(fileno(f_), 0, std::ftell(f_), POSIX_FADV_DONTNEED); // RAM is tight
return true;
}
private:
FILE *f_ = nullptr;
uint64_t sets_ = 0;
};
bool is_dark(const char *name) {
const size_t n = std::strlen(name);
return n > 3 && !std::strcmp(name + n - 3, "_dk");
}
} // namespace
int main(int argc, char **argv) {
if (argc < 2 || argv[1][0] == '-') {
std::fprintf(stderr, "usage: %s DIR --ring PATH [--from S] [--to S] [--loop] [--cpus 0,1]\n", argv[0]);
return 1;
}
const std::string dir = argv[1];
std::string ring_path;
double from = 0, to = 1e9;
bool loop = false;
std::vector<int> cpus = {0, 1};
for (int i = 2; i < argc; ++i) {
const std::string a = argv[i];
const bool more = i + 1 < argc;
if (a == "--ring" && more) ring_path = argv[++i];
else if (a == "--from" && more) from = std::atof(argv[++i]);
else if (a == "--to" && more) to = std::atof(argv[++i]);
else if (a == "--loop") loop = true;
else if (a == "--cpus" && more) {
cpus.clear();
for (char *p = argv[++i]; *p;) {
cpus.push_back(int(std::strtol(p, &p, 10)));
if (*p == ',') ++p;
else if (*p) break;
}
} else return std::fprintf(stderr, "unknown option %s\n", a.c_str()), 1;
}
if (ring_path.empty()) return std::fprintf(stderr, "--ring PATH is required\n"), 1;
if (!cpus.empty()) {
cpu_set_t set;
CPU_ZERO(&set);
for (int c : cpus) CPU_SET(c, &set);
if (sched_setaffinity(0, sizeof set, &set) < 0) std::perror("sched_setaffinity");
}
std::signal(SIGINT, [](int) { g_stop = 1; });
std::signal(SIGTERM, [](int) { g_stop = 1; });
Reader in;
if (!in.open(dir + "/sets.bin")) return std::fprintf(stderr, "%s/sets.bin: %s\n", dir.c_str(), std::strerror(errno)), 1;
std::vector<fh_set_cam_t> cams;
std::vector<std::vector<uint8_t>> px;
auto want = [](const char *name) { return !is_dark(name); };
if (!in.next(cams, px, want)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
auto set_time = [&](const std::vector<fh_set_cam_t> &cs) { // earliest bright capture, s
uint64_t t = UINT64_MAX;
for (const auto &c : cs)
if (!is_dark(c.name)) t = std::min(t, c.capture_ns);
return double(t) * 1e-9;
};
const double rec0 = set_time(cams);
// skip to --from before the ring exists, so a reader never finds it without a heartbeat
bool have = true;
while (have && set_time(cams) - rec0 < from && !g_stop) have = in.next(cams, px, want);
if (!have) return std::fprintf(stderr, "%s: nothing after %.1f s\n", dir.c_str(), from), 1;
// the ring: the recording's bright cameras, as ft-camd lays them out
std::vector<int> pub; // set camera index of each ring camera
for (size_t k = 0; k < cams.size() && pub.size() < FH_RING_MAX_CAMS; ++k)
if (!is_dark(cams[k].name)) pub.push_back(int(k));
size_t len = sizeof(fh_ring_hdr_t);
std::vector<uint64_t> offset(pub.size()), slot_bytes(pub.size());
for (size_t r = 0; r < pub.size(); ++r) {
const fh_set_cam_t &c = cams[pub[r]];
slot_bytes[r] = (sizeof(fh_ring_slot_t) + size_t(c.width) * c.height + 63) & ~size_t(63);
offset[r] = len;
len += FH_RING_SLOTS * slot_bytes[r];
}
const int fd = ::open(ring_path.c_str(), O_RDWR | O_CREAT | O_TRUNC | O_CLOEXEC, 0600);
if (fd < 0 || ftruncate(fd, off_t(len)) < 0)
return std::fprintf(stderr, "%s: %s\n", ring_path.c_str(), std::strerror(errno)), 1;
void *m = mmap(nullptr, len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
close(fd);
if (m == MAP_FAILED) return std::fprintf(stderr, "mmap %s: %s\n", ring_path.c_str(), std::strerror(errno)), 1;
auto *base = static_cast<uint8_t *>(m);
auto *hdr = reinterpret_cast<fh_ring_hdr_t *>(base);
for (size_t r = 0; r < pub.size(); ++r) {
const fh_set_cam_t &c = cams[pub[r]];
fh_ring_cam_t &rc = hdr->cams[r];
std::snprintf(rc.sensor, sizeof rc.sensor, "ft-ringplay");
std::snprintf(rc.name, sizeof rc.name, "%s", c.name);
rc.node = -1;
rc.format = FH_FMT_GREY8;
rc.width = rc.stride = c.width;
rc.height = c.height;
rc.nslots = FH_RING_SLOTS;
rc.slot_offset = offset[r];
rc.slot_bytes = slot_bytes[r];
}
hdr->version = FH_RING_VERSION;
hdr->header_bytes = sizeof(fh_ring_hdr_t);
hdr->ncams = uint32_t(pub.size());
hdr->file_bytes = len;
hdr->writer_pid = getpid();
std::memcpy(hdr->magic, FH_RING_MAGIC, 8);
__atomic_store_n(&hdr->heartbeat_ns, clock_ns(CLOCK_MONOTONIC), __ATOMIC_RELEASE);
std::printf("playing %s into %s:", dir.c_str(), ring_path.c_str());
for (int k : pub) std::printf(" %s", cams[k].name);
std::printf("\n");
std::fflush(stdout);
uint64_t published = 0, rounds = 0;
for (;;) {
// one pass over [from, to]: each set goes out at its recorded offset from the first
while (have && set_time(cams) - rec0 < from && !g_stop) {
__atomic_store_n(&hdr->heartbeat_ns, clock_ns(CLOCK_MONOTONIC), __ATOMIC_RELEASE);
have = in.next(cams, px, want);
}
const double first = set_time(cams);
const uint64_t start = clock_ns(CLOCK_MONOTONIC);
while (have && !g_stop && set_time(cams) - rec0 <= to) {
const uint64_t due = start + uint64_t((set_time(cams) - first) * 1e9);
for (uint64_t now = clock_ns(CLOCK_MONOTONIC); now < due && !g_stop; now = clock_ns(CLOCK_MONOTONIC)) {
__atomic_store_n(&hdr->heartbeat_ns, now, __ATOMIC_RELEASE);
const uint64_t wait = std::min<uint64_t>(due - now, 100'000'000);
const timespec ts{time_t(wait / 1'000'000'000), long(wait % 1'000'000'000)};
nanosleep(&ts, nullptr);
}
const uint64_t raw = clock_ns(CLOCK_MONOTONIC_RAW), mono = clock_ns(CLOCK_MONOTONIC);
for (size_t r = 0; r < pub.size(); ++r) {
fh_ring_cam_t &rc = hdr->cams[r];
if (px[pub[r]].size() != size_t(rc.width) * rc.height) continue;
const uint64_t n = rc.latest + 1;
auto *s = reinterpret_cast<fh_ring_slot_t *>(base + rc.slot_offset + (n % rc.nslots) * rc.slot_bytes);
__atomic_store_n(&s->seq, 2 * n + 1, __ATOMIC_RELAXED);
__atomic_thread_fence(__ATOMIC_RELEASE);
std::memcpy(reinterpret_cast<uint8_t *>(s + 1), px[pub[r]].data(), px[pub[r]].size());
s->frame = n;
s->capture_ns = raw; // taken now, as a live camera's frame would be
s->dqbuf_ns = mono;
s->publish_ns = mono;
__atomic_store_n(&s->seq, 2 * n + 2, __ATOMIC_RELEASE);
__atomic_store_n(&rc.latest, n, __ATOMIC_RELEASE);
++rc.published;
}
__atomic_store_n(&hdr->heartbeat_ns, mono, __ATOMIC_RELEASE);
++published;
have = in.next(cams, px, want);
}
++rounds;
if (g_stop || !loop) break;
in.rewind();
have = in.next(cams, px, want);
}
std::printf("published %llu sets in %llu pass(es)\n", (unsigned long long)published, (unsigned long long)rounds);
__atomic_store_n(&hdr->heartbeat_ns, 0, __ATOMIC_RELEASE); // readers see the writer gone
return 0;
}
+621
View File
@@ -0,0 +1,621 @@
#include "tracker.h"
#include <pthread.h>
#include <sched.h>
#include <algorithm>
#include <chrono>
#include <set>
namespace {
// where arms start, head frame: below and slightly behind the eyes
const V3 kShoulders[2] = {{0.17, -0.25, 0.08}, {-0.17, -0.25, 0.08}};
// landmark pairs across the palm, rigid enough for single-view depth
const int kPalmPairs[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 17}, {5, 13}, {9, 17}, {1, 17}, {1, 5}};
constexpr double kFastSpeed = 0.25; // m/s
constexpr double kSearchInterval = 0.2;
// One Euro filter on the published landmarks: still hands are smoothed hard (tracking
// noise is a few mm per frame), fast ones barely, so they don't lag.
constexpr double kMinCutoff = 2.0; // Hz, a still hand
constexpr double kBeta = 30.0; // Hz more per m/s of palm speed
constexpr double kSpeedCutoff = 1.5; // Hz, for the palm speed itself
// With one view, the hand's distance from the camera comes from how big it looks, which
// is off by 10-30% and wanders ~10% between frames. Its direction is exact. So a hand that
// was just located keeps its distance, drifting toward the one-view guess by this much a frame.
constexpr double kMonoDepthGain = 0.1;
// Is a triangulated hand as far from each camera as its apparent size says? With the
// model's average hand, clean stereo pairs measure 0.71-1.51 times the one-view distance
// (5-95%, median 1.16); pairs of two different hands mostly far less.
constexpr double kSizePrior = 1.16, kRatioLo = 0.6, kRatioHi = 1.9;
double ms_since(std::chrono::steady_clock::time_point t) {
return std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t).count();
}
bool is_color(const Camera &c) { return c.name.rfind("color", 0) == 0; } // Arcturus, 145 degree image circle
// The wide cameras: the side ones and the color ones
bool is_slam(const Camera &c) { return c.name.rfind("slam", 0) == 0 || is_color(c); }
V2 palm_centre(const Landmarks &lm) { return lm.pts[9]; }
// Two views in one camera on the same hand: the landmark model puts the same points on
// it from both crops, even when the crops differ.
bool same_hand(const Landmarks &a, const Landmarks &b, double size) {
double d = 0;
for (int i = 0; i < 21; ++i) d += norm(a.pts[i] - b.pts[i]) / 21;
return norm(palm_centre(a) - palm_centre(b)) < 0.5 * size || d < 0.25 * size;
}
double hand_size(const Landmarks &lm) {
double lo[2] = {1e9, 1e9}, hi[2] = {-1e9, -1e9};
for (const V2 &p : lm.pts)
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[k]), hi[k] = std::max(hi[k], p[k]);
return std::max(hi[0] - lo[0], hi[1] - lo[1]);
}
} // namespace
// ---------------------------------------------------------------------------- pool
Pool::Pool(int threads, const std::vector<int> &cpus) {
for (int i = 0; i < threads; ++i) threads_.emplace_back(&Pool::loop, this, cpus[i % cpus.size()]);
}
Pool::~Pool() {
{
std::lock_guard<std::mutex> l(mu_);
stop_ = true;
}
wake_.notify_all();
for (auto &t : threads_) t.join();
}
void Pool::loop(int cpu) {
cpu_set_t set;
CPU_ZERO(&set);
CPU_SET(cpu, &set);
pthread_setaffinity_np(pthread_self(), sizeof set, &set); // ignored if not allowed
std::unique_lock<std::mutex> l(mu_);
for (;;) {
wake_.wait(l, [&] { return stop_ || (jobs_ && next_ < jobs_->size()); });
if (stop_) return;
auto &job = (*jobs_)[next_++];
l.unlock();
job();
l.lock();
if (++finished_ == jobs_->size()) done_.notify_all();
}
}
void Pool::run(std::vector<std::function<void()>> &jobs) {
if (jobs.empty()) return;
std::unique_lock<std::mutex> l(mu_);
jobs_ = &jobs, next_ = 0, finished_ = 0;
wake_.notify_all();
done_.wait(l, [&] { return finished_ == jobs.size(); });
jobs_ = nullptr;
}
// ------------------------------------------------------------------------- tracker
Tracker::Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views)
: nets_(nets), pool_(pool), max_views_(max_views) {
for (const auto &[name, cam] : cams) {
cams_[name] = &cam;
if (is_slam(cam)) {
add_tiles(cam, 0.45, 3, 3);
add_tiles(cam, 0.65, 2, 2);
add_tiles(cam, 1.0, 1, 1); // hands close to the face fill much of the frame
} else {
add_tiles(cam, 0.6, 3, 2);
add_tiles(cam, 1.0, 1, 1);
}
}
}
void Tracker::add_tiles(const Camera &cam, double frac, int gx, int gy) {
const double s = frac * std::max(cam.width, cam.height);
for (int i = 0; i < gx; ++i)
for (int j = 0; j < gy; ++j) {
const double x = gx > 1 ? s / 2 + (cam.width - s) * i / (gx - 1) : cam.width / 2.0;
const double y = gy > 1 ? s / 2 + (cam.height - s) * j / (gy - 1) : cam.height / 2.0;
Tile t{&cam, {x, y}, s, 0, 0};
// turn the crop so the expected shoulder-to-hand direction points up
const V3 ray = cam.ray(t.center), p = cam.origin + ray * 0.45;
const V3 d = unit(p - kShoulders[p[0] > 0 ? 0 : 1]);
const V2 a = cam.project(p, nullptr), b = cam.project(p + d * 0.05, nullptr);
t.rotation = std::atan2(b[0] - a[0], -(b[1] - a[1]));
t.weight = std::max(0.15, dot(ray, unit(V3{0, -0.45, -0.9})));
tiles_.push_back(t);
}
}
double Tracker::interval() const {
double fastest = -1;
for (const auto &[id, h] : hands_)
if (h.seen_ns == last_ns_) fastest = std::max(fastest, norm(h.dpalm)); // filtered: noise isn't speed
return fastest < 0 ? 1 / 5.0 : fastest > kFastSpeed ? 1 / 30.0 : 1 / 15.0;
}
bool Tracker::inside(const Camera &cam, V2 uv) const {
const double m = 0.12;
return uv[0] >= m * cam.width && uv[0] <= (1 - m) * cam.width && uv[1] >= m * cam.height &&
uv[1] <= (1 - m) * cam.height && cam.off_axis(uv) < (is_color(cam) ? 70.0 : is_slam(cam) ? 80.0 : 75.0);
}
void Tracker::run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views) {
if (views.empty()) return;
const auto t0 = std::chrono::steady_clock::now();
std::vector<std::function<void()>> jobs;
for (View *v : views) {
const Image &img = images.at(v->cam->name);
jobs.push_back([this, v, &img] {
v->lm = nets_.landmarks(img, v->roi);
v->has_lm = true;
v->fresh = true;
});
}
pool_.run(jobs);
stats.hand_calls += int(views.size());
stats.hand_ms += ms_since(t0);
++stats.hand_batches;
}
bool Tracker::single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const {
V3 rays[21];
for (int i = 0; i < 21; ++i) rays[i] = cam.ray(lm.pts[i]);
std::vector<std::pair<double, double>> est; // (depth, weight)
for (const auto &pr : kPalmPairs) {
const int i = pr[0], j = pr[1];
const double d = std::hypot(lm.world[i][0] - lm.world[j][0], lm.world[i][1] - lm.world[j][1]) * scale;
const double a = std::acos(std::clamp(dot(rays[i], rays[j]), -1.0, 1.0));
if (a > 1e-3 && d > 0.01) est.push_back({d / a, d});
}
if (est.empty()) return false;
std::sort(est.begin(), est.end());
double total = 0, acc = 0, depth = est.back().first;
for (auto &e : est) total += e.second;
for (auto &e : est)
if ((acc += e.second) >= total / 2) { depth = e.first; break; }
double zmean = 0;
for (int i = 0; i < 21; ++i) zmean += lm.world[i][2] / 21;
for (int i = 0; i < 21; ++i) out[i] = cam.origin + rays[i] * (depth + (lm.world[i][2] - zmean) * scale);
return true;
}
// A triangulated hand is in front of each camera, as far as its apparent size says (see
// kSizePrior). Returns how far off that is (the sum of |log| ratios), or -1 if implausible.
double Tracker::size_misfit(const std::vector<const View *> &views, const V3 *pts) const {
double misfit = 0;
for (const View *v : views) {
V3 mono[21];
// along the view's own ray (fisheye: a hand near the image edge is far off the axis)
if (dot(pts[9] - v->cam->origin, v->cam->ray(v->lm.pts[9])) < 0.08) return -1;
if (!single_view(*v->cam, v->lm, 1.0, mono)) continue;
const double r = norm(pts[9] - v->cam->origin) / norm(mono[9] - v->cam->origin);
if (r < kRatioLo || r > kRatioHi) return -1;
misfit += std::fabs(std::log(r / kSizePrior));
}
return misfit;
}
bool Tracker::hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns) {
views.erase(std::remove_if(views.begin(), views.end(), [](View *v) { return !v->has_lm || !v->fresh; }), views.end());
if (views.empty()) return false;
if (views.size() >= 2) {
const int n = int(views.size());
std::vector<V3> origins(n), dirs(n);
std::vector<double> w(n), res(21);
V3 pts[21];
for (int k = 0; k < 21; ++k) {
for (int v = 0; v < n; ++v) {
origins[v] = views[v]->cam->origin;
dirs[v] = views[v]->cam->ray(views[v]->lm.pts[k]);
w[v] = views[v]->lm.presence;
}
pts[k] = triangulate(origins.data(), dirs.data(), w.data(), n, &res[k]);
}
std::nth_element(res.begin(), res.begin() + 10, res.end());
const double residual = res[10];
// the views disagree: two different hands; keep the stronger. Rays to two different
// hands can pass close to each other near the cameras, so check the distance too.
if (residual > 0.03 || size_misfit({views.begin(), views.end()}, pts) < 0) {
View *best = *std::max_element(views.begin(), views.end(),
[](View *a, View *b) { return a->lm.presence < b->lm.presence; });
for (View *v : views)
if (v != best) v->hand = -1;
++stats.splits;
return hand_3d(hand, {best}, t_ns);
}
// learn how big this user's hand is compared to the model's average hand
std::vector<double> t, m;
const Landmarks &ref = views[0]->lm;
for (const auto &pr : kPalmPairs) {
t.push_back(norm(pts[pr[0]] - pts[pr[1]]));
m.push_back(norm(V3{ref.world[pr[0]][0], ref.world[pr[0]][1], ref.world[pr[0]][2]} -
V3{ref.world[pr[1]][0], ref.world[pr[1]][1], ref.world[pr[1]][2]}));
}
std::nth_element(t.begin(), t.begin() + t.size() / 2, t.end());
std::nth_element(m.begin(), m.begin() + m.size() / 2, m.end());
if (m[m.size() / 2] > 0 && residual < 0.008) // only from clean matches
hand.scale += 0.1 * (std::clamp(t[t.size() / 2] / m[m.size() / 2], 0.8, 1.6) - hand.scale);
std::copy(pts, pts + 21, hand.pts);
hand.residual = residual;
for (View *v : views) {
V3 mono[21];
if (!single_view(*v->cam, v->lm, hand.scale, mono)) continue;
const V3 o = v->cam->origin;
stats.mono_ratio.push_back({hand.id, norm(mono[9] - o) / norm(pts[9] - o)});
}
} else {
const Camera &cam = *views[0]->cam;
V3 pts[21];
if (!single_view(cam, views[0]->lm, hand.scale, pts)) return false;
const double guess = norm(pts[9] - cam.origin);
if (hand.has_pts && t_ns - hand.seen_ns < 300'000'000 && guess > 0) {
const double was = norm(hand.pts[9] - cam.origin), d = was + kMonoDepthGain * (guess - was);
for (V3 &p : pts) p = cam.origin + (p - cam.origin) * (d / guess);
}
std::copy(pts, pts + 21, hand.pts);
hand.residual = -1;
}
hand.has_pts = true;
hand.nviews = int(views.size());
for (View *v : views) hand.right_score += 0.2 * (v->lm.right - hand.right_score);
return true;
}
// How badly two views in different cameras fit one hand: the rays should meet, each view's
// apparent size should match its distance, and the model should call both the same hand
// (left or right). Negative if they can't be one hand. Side by side hands sit on the same
// epipolar lines of the side cameras, so the distance check is what tells them apart.
double Tracker::pair_cost(const View &a, const View &b) const {
V3 pts[21];
std::vector<double> res(21);
for (int k = 0; k < 21; ++k) {
const V3 o[2] = {a.cam->origin, b.cam->origin};
const V3 d[2] = {a.cam->ray(a.lm.pts[k]), b.cam->ray(b.lm.pts[k])};
const double w[2] = {a.lm.presence, b.lm.presence};
pts[k] = triangulate(o, d, w, 2, &res[k]);
}
std::nth_element(res.begin(), res.begin() + 10, res.end());
if (res[10] > 0.03) return -1;
const double misfit = size_misfit({&a, &b}, pts);
return misfit < 0 ? -1 : res[10] / 0.01 + misfit + std::fabs(a.lm.right - b.lm.right);
}
// Which views in two cameras are the same hand: every way of pairing them up (a few views
// each), scored with pair_cost. Keeps the hands' pairing unless another is clearly better,
// then relabels the views, keeping the longer-tracked hand's id.
void Tracker::associate() {
constexpr double kPairBonus = 2.0, kBetter = 0.3;
std::vector<const Camera *> cams;
for (View &v : views_)
if (std::find(cams.begin(), cams.end(), v.cam) == cams.end()) cams.push_back(v.cam);
std::sort(cams.begin(), cams.end(), [](const Camera *a, const Camera *b) { return a->name < b->name; });
for (size_t i = 0; i < cams.size(); ++i)
for (size_t j = i + 1; j < cams.size(); ++j) {
std::vector<View *> A, B;
for (View &v : views_) {
if (!v.fresh) continue;
if (v.cam == cams[i]) A.push_back(&v);
else if (v.cam == cams[j]) B.push_back(&v);
}
if (A.empty() || B.empty() || A.size() > 3 || B.size() > 3) continue;
std::vector<std::vector<double>> c(A.size(), std::vector<double>(B.size()));
for (size_t x = 0; x < A.size(); ++x)
for (size_t y = 0; y < B.size(); ++y) c[x][y] = pair_cost(*A[x], *B[y]);
auto score = [&](const std::vector<int> &m) { // m[x]: A[x]'s partner in B, or -1
double s = 0;
for (size_t x = 0; x < A.size(); ++x)
if (m[x] >= 0 && c[x][m[x]] >= 0) s += c[x][m[x]] - kPairBonus;
return s;
};
std::vector<int> cur(A.size(), -1);
for (size_t x = 0; x < A.size(); ++x)
for (size_t y = 0; y < B.size(); ++y)
if (A[x]->hand == B[y]->hand) cur[x] = int(y);
std::vector<int> best = cur, m(A.size(), -1);
double best_score = score(cur);
const double cur_score = best_score;
std::function<void(size_t, unsigned)> walk = [&](size_t x, unsigned used) {
if (x == A.size()) {
const double sc = score(m);
if (sc < best_score) best_score = sc, best = m;
return;
}
m[x] = -1;
walk(x + 1, used);
for (size_t y = 0; y < B.size(); ++y)
if (!(used >> y & 1) && c[x][y] >= 0) {
m[x] = int(y);
walk(x + 1, used | 1u << y);
}
m[x] = -1;
};
walk(0, 0);
if (best == cur || best_score > cur_score - kBetter) continue;
auto frames = [&](int id) {
const auto h = hands_.find(id);
return h == hands_.end() ? -1 : h->second.frames;
};
for (size_t x = 0; x < A.size(); ++x) {
if (best[x] < 0) continue;
View *a = A[x], *b = B[best[x]];
int id = a->hand;
bool free = true; // b's hand isn't another A view's
for (size_t x2 = 0; x2 < A.size(); ++x2) free = free && (x2 == x || A[x2]->hand != b->hand);
if (free && frames(b->hand) > frames(id)) id = b->hand;
a->hand = b->hand = id;
}
// a B view left unpaired that still shares a hand with an A view starts its own
for (size_t y = 0; y < B.size(); ++y) {
if (std::find(best.begin(), best.end(), int(y)) != best.end()) continue;
bool shared = false;
for (View *a : A) shared = shared || a->hand == B[y]->hand;
if (!shared) continue;
B[y]->hand = next_id_++;
hands_[B[y]->hand].id = B[y]->hand;
++stats.created;
}
++stats.merged;
}
}
std::vector<const Hand *> Tracker::step(const std::map<std::string, Image> &images, int64_t t_ns) {
const auto t_step = std::chrono::steady_clock::now();
++stats.sets;
// Views in cameras without a frame in this set wait, as they are, for their camera's
// next one: the colour cameras run on their own clock, so a set can hold the mono
// cameras, the colour ones, or both (drop_camera ends them when a camera stops being used).
std::vector<View> live, waiting;
for (View &v : views_) {
(images.count(v.cam->name) ? live : waiting).push_back(v);
(images.count(v.cam->name) ? live : waiting).back().fresh = false;
}
// 1. hand-over: give hands with too few views a crop in other cameras
for (auto &[id, hand] : hands_) {
if (!hand.has_pts) continue;
std::set<std::string> have;
for (View &v : live)
if (v.hand == id) have.insert(v.cam->name);
if (int(have.size()) >= max_views_) continue;
std::vector<std::pair<double, View>> options;
for (auto &[name, cam] : cams_) {
if (have.count(name) || !images.count(name)) continue;
V2 uv[21];
bool front = true;
for (int k = 0; k < 21; ++k) {
double z;
uv[k] = cam->project(hand.pts[k], &z);
front = front && z > 0;
}
const V2 centre = (uv[0] + uv[5] + uv[9] + uv[13] + uv[17]) * 0.2;
if (!front || !inside(*cam, centre)) continue;
View v{cam, roi_from_points(uv), id, {}, false, 0};
options.push_back({cam->off_axis(centre), v});
}
std::sort(options.begin(), options.end(), [](auto &a, auto &b) { return a.first < b.first; });
for (size_t k = 0; k < options.size() && int(have.size() + k) < max_views_; ++k) live.push_back(options[k].second);
}
// 2. the landmark model on each hand's best views, within budget
std::map<int, std::vector<View *>> by_hand;
for (View &v : live) by_hand[v.hand].push_back(&v);
std::vector<View *> chosen;
for (auto &[id, vs] : by_hand) {
std::sort(vs.begin(), vs.end(), [](View *a, View *b) {
if (a->has_lm != b->has_lm) return a->has_lm;
return a->cam->off_axis(a->roi.center) < b->cam->off_axis(b->roi.center);
});
for (int k = 0; k < int(vs.size()) && k < max_views_; ++k) chosen.push_back(vs[k]);
}
std::stable_sort(chosen.begin(), chosen.end(), [](View *a, View *b) { return a->has_lm > b->has_lm; });
if (int(chosen.size()) > hand_budget_) chosen.resize(hand_budget_);
run_landmarks(images, chosen);
std::vector<View *> kept;
for (View *v : chosen)
if (v->lm.presence >= (v->frames > 0 ? keep_presence_ : min_presence_)) {
v->roi = v->lm.next_roi();
++v->frames;
kept.push_back(v);
} else {
++(v->frames > 0 ? stats.lost : stats.handoff_miss);
}
// the same hand twice in one camera: keep the more confident
std::sort(kept.begin(), kept.end(), [](View *a, View *b) { return a->lm.presence > b->lm.presence; });
std::vector<View> next;
for (View *v : kept) {
const double size = hand_size(v->lm);
bool dup = false;
for (View &o : next) dup = dup || (o.cam == v->cam && same_hand(o.lm, v->lm, size));
if (!dup) next.push_back(*v);
else ++stats.dups;
}
views_ = next;
views_.insert(views_.end(), waiting.begin(), waiting.end());
// 3. search for missing hands
std::set<int> tracked;
for (View &v : views_) tracked.insert(v.hand);
if (tracked.size() < 2 && (t_ns - search_ns_) / 1e9 >= kSearchInterval - 0.01) {
search_ns_ = t_ns;
const int budget = tracked.empty() ? search_budget_ : std::max(1, search_budget_ - 1);
for (Tile &t : tiles_)
if (images.count(t.cam->name)) t.credit += t.weight;
std::vector<Tile *> picked;
for (int b = 0; b < budget; ++b) {
Tile *best = nullptr;
for (Tile &t : tiles_)
if (images.count(t.cam->name) && std::find(picked.begin(), picked.end(), &t) == picked.end() &&
(!best || t.credit > best->credit))
best = &t;
if (!best) break;
best->credit = 0;
picked.push_back(best);
}
const auto t0 = std::chrono::steady_clock::now();
std::vector<std::vector<Palm>> found(picked.size());
std::vector<std::function<void()>> jobs;
for (size_t i = 0; i < picked.size(); ++i) {
Tile *t = picked[i];
const Image &img = images.at(t->cam->name);
jobs.push_back([this, t, &img, &found, i] { found[i] = nets_.palms(img, t->center, t->size, t->rotation); });
}
pool_.run(jobs);
stats.palm_calls += int(picked.size());
stats.palm_ms += ms_since(t0);
++stats.palm_batches;
std::vector<View> fresh;
for (size_t i = 0; i < picked.size(); ++i)
for (const Palm &p : found[i]) {
const Roi roi = p.roi();
bool near = false;
for (auto *list : {&views_, &fresh})
for (View &v : *list) near = near || (v.cam == picked[i]->cam && norm(v.roi.center - roi.center) < 0.5 * roi.size);
if (!near) fresh.push_back({picked[i]->cam, roi, 0, {}, false, 0});
}
std::vector<View *> ptrs;
for (View &v : fresh) ptrs.push_back(&v);
run_landmarks(images, ptrs);
for (View &v : fresh)
if (v.lm.presence >= min_presence_) {
v.roi = v.lm.next_roi();
v.frames = 1;
views_.push_back(v);
}
}
// 4. give new views a hand: the nearest existing hand in 3D, else a new one
for (View &v : views_) {
if (v.hand > 0 && hands_.count(v.hand)) continue;
V3 guess[21];
const bool have_guess = single_view(*v.cam, v.lm, 1.0, guess);
int best = 0;
double dist = 0.12;
for (auto &[id, h] : hands_) {
if (!h.has_pts) continue;
bool same_cam = false;
for (View &o : views_) same_cam = same_cam || (o.hand == id && o.cam == v.cam);
if (same_cam) continue;
const double d = have_guess ? norm(h.pts[9] - guess[9]) : 1e9;
if (d < dist) best = id, dist = d;
}
if (!best) {
best = next_id_++;
hands_[best].id = best;
++stats.created;
}
v.hand = best;
}
// 5. which views in different cameras are the same hand
associate();
// 6. 3D for every hand seen now; forget hands not seen for a while
std::vector<const Hand *> out;
for (auto it = hands_.begin(); it != hands_.end();) {
Hand &h = it->second;
std::vector<View *> vs;
for (View &v : views_)
if (v.hand == h.id) vs.push_back(&v);
if (!vs.empty() && hand_3d(h, vs, t_ns)) {
const V3 palm = (h.pts[0] + h.pts[5] + h.pts[9] + h.pts[13] + h.pts[17]) * 0.2;
if (h.last_ns && t_ns > h.last_ns)
h.speed += 0.5 * (std::min(norm(palm - h.last_palm) / ((t_ns - h.last_ns) / 1e9), 5.0) - h.speed);
h.last_ns = t_ns, h.last_palm = palm, h.seen_ns = t_ns;
++h.frames;
smooth(h, t_ns);
out.push_back(&h);
++it;
} else if (t_ns - h.seen_ns > 300'000'000) {
it = hands_.erase(it);
++stats.forgotten;
} else {
++it;
}
}
// views split off by a failed triangulation start over as new hands next frame
for (View &v : views_)
if (v.hand <= 0) {
v.hand = next_id_++;
hands_[v.hand].id = v.hand;
++stats.created;
}
last_ns_ = t_ns;
stats.step_ms += ms_since(t_step);
return out;
}
std::vector<Seen> Tracker::views_now() const {
std::vector<Seen> out;
for (const View &v : views_)
if (v.fresh) out.push_back({v.cam->name, v.hand, v.roi, v.lm, {}});
return out;
}
void Tracker::drop_camera(const std::string &name) {
views_.erase(std::remove_if(views_.begin(), views_.end(), [&](const View &v) { return v.cam->name == name; }),
views_.end());
}
std::vector<Seen> Tracker::exhaustive(const std::map<std::string, Image> &images) {
std::vector<Tile *> tiles;
for (Tile &t : tiles_)
if (images.count(t.cam->name)) tiles.push_back(&t);
std::vector<std::vector<Palm>> found(tiles.size());
std::vector<std::function<void()>> jobs;
for (size_t i = 0; i < tiles.size(); ++i)
jobs.push_back([this, &tiles, &images, &found, i] {
const Tile *t = tiles[i];
found[i] = nets_.palms(images.at(t->cam->name), t->center, t->size, t->rotation);
});
pool_.run(jobs);
// one crop per palm: tiles overlap, so the same palm turns up several times
std::vector<std::pair<double, View>> palms;
for (size_t i = 0; i < tiles.size(); ++i)
for (const Palm &p : found[i]) palms.push_back({p.score, View{tiles[i]->cam, p.roi(), 0, {}, false, 0}});
std::sort(palms.begin(), palms.end(), [](auto &a, auto &b) { return a.first > b.first; });
std::vector<View> crops;
for (auto &[score, v] : palms) {
bool near = false;
for (View &o : crops) near = near || (o.cam == v.cam && norm(o.roi.center - v.roi.center) < 0.5 * v.roi.size);
if (!near) crops.push_back(v);
}
std::vector<View *> ptrs;
for (View &v : crops) ptrs.push_back(&v);
run_landmarks(images, ptrs);
std::sort(crops.begin(), crops.end(), [](const View &a, const View &b) { return a.lm.presence > b.lm.presence; });
std::vector<Seen> out;
for (View &v : crops) {
if (v.lm.presence < min_presence_) continue;
bool dup = false;
for (const Seen &o : out)
dup = dup || (o.cam == v.cam->name && norm(palm_centre(o.lm) - palm_centre(v.lm)) < 0.5 * hand_size(v.lm));
if (dup) continue;
V3 pts[21];
Seen s{v.cam->name, 0, v.roi, v.lm, {}};
if (single_view(*v.cam, v.lm, 1.0, pts)) s.wrist = pts[0];
out.push_back(s);
}
return out;
}
void Tracker::smooth(Hand &h, int64_t t_ns) {
const double dt = (t_ns - h.smooth_ns) / 1e9;
h.smooth_ns = t_ns;
if (h.frames <= 1 || dt <= 0 || dt > 0.3) { // new, or back after a gap: start over
std::copy(h.pts, h.pts + 21, h.smooth);
h.dpalm = {0, 0, 0};
return;
}
auto alpha = [dt](double cutoff) { return 1 / (1 + 1 / (2 * M_PI * cutoff * dt)); };
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
const V3 d = (palm(h.pts) - palm(h.smooth)) * (1 / dt);
h.dpalm = h.dpalm + (d - h.dpalm) * alpha(kSpeedCutoff);
// one cutoff for the whole hand, from its palm speed, so its shape stays together
const double a = alpha(kMinCutoff + kBeta * norm(h.dpalm));
for (int i = 0; i < 21; ++i) h.smooth[i] = h.smooth[i] + (h.pts[i] - h.smooth[i]) * a;
}
+134
View File
@@ -0,0 +1,134 @@
// Multi-camera hand tracking (a port of frame-hands' Python prototype; the scheduling is
// described in hands/README.md). All 3D is metres in the head frame.
#pragma once
#include "calib.h"
#include "nets.h"
#include <condition_variable>
#include <functional>
#include <map>
#include <memory>
#include <mutex>
#include <thread>
#include <vector>
// Runs batches of jobs on a few threads, each pinned to a core.
class Pool {
public:
// One thread per entry of cpus, pinned there (round-robin if threads > cpus).
Pool(int threads, const std::vector<int> &cpus);
~Pool();
void run(std::vector<std::function<void()>> &jobs);
private:
void loop(int cpu);
std::vector<std::thread> threads_;
std::mutex mu_;
std::condition_variable wake_, done_;
std::vector<std::function<void()>> *jobs_ = nullptr;
size_t next_ = 0, finished_ = 0;
bool stop_ = false;
};
struct Hand {
int id = 0;
V3 pts[21]{}; // as measured this frame; the tracker steers crops by these
V3 smooth[21]{}; // filtered (One Euro, see Tracker::smooth): publish these
bool has_pts = false;
double residual = -1; // rms ray distance of the triangulation (m); -1: one view
int nviews = 0;
double right_score = 0.5; // the model's right-hand score (these images aren't mirrored)
double scale = 1.0; // this user's hand size / the model's world landmarks
int64_t seen_ns = 0;
int frames = 0;
double speed = 0; // palm centre, m/s, smoothed
int64_t last_ns = 0;
V3 last_palm{};
V3 dpalm{}; // the filter's palm velocity, m/s
int64_t smooth_ns = 0;
bool right() const { return right_score > 0.5; }
};
struct Stats {
int palm_calls = 0, hand_calls = 0, sets = 0;
double palm_ms = 0, hand_ms = 0, step_ms = 0; // summed batch times
int palm_batches = 0, hand_batches = 0;
// why views and hands come and go
int lost = 0; // a tracked view's landmarks fell below min presence
int handoff_miss = 0; // a view projected from the hand's 3D (new camera or retry) found no hand
int dups = 0; // the same hand twice in one camera
int splits = 0; // a hand's views disagreed in 3D and were split
int created = 0, merged = 0, forgotten = 0; // merged: views re-paired across cameras
// diagnostics: on stereo frames, each view's single-view palm distance / the stereo one
std::vector<std::pair<int, double>> mono_ratio; // (hand id, ratio)
};
// A hand the landmark model found in one camera (Tracker::views_now, Tracker::exhaustive).
struct Seen {
std::string cam;
int hand = 0; // the tracker's hand; 0 in exhaustive()
Roi roi;
Landmarks lm;
V3 wrist{}; // exhaustive(): single-view 3D guess at the model's hand size
};
class Tracker {
public:
Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views = 2);
// images: calibration name -> frame. Returns the hands seen in this set.
std::vector<const Hand *> step(const std::map<std::string, Image> &images, int64_t t_ns);
// Seconds until the next frame set is worth processing (30 Hz fast hands, 15 Hz slow, 5 Hz none).
double interval() const;
Stats stats;
size_t views() const { return views_.size(); }
std::vector<Seen> views_now() const; // the views this step updated
// Forget this camera's views, when it stops being tracked with (the lighting switched
// cameras); views in cameras a set lacks otherwise wait for their next frame.
void drop_camera(const std::string &name);
// Every search tile in every camera, then landmarks on every palm: slow; for checking
// what the scheduler misses (ft-handreplay --oracle).
std::vector<Seen> exhaustive(const std::map<std::string, Image> &images);
// Landmark presence a tracked view needs to stay (new views need min presence, 0.5). In
// bright rooms the camera exposes for the room, the hands come out dim, and presence
// dips under 0.5 for a frame at a time.
void set_keep_presence(double p) { keep_presence_ = p; }
// One view's 3D hand: each landmark along its ray, as far as how big the palm looks says
// for a hand `scale` times the model's (Hand::scale). False if the palm is degenerate.
bool single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const;
private:
struct View {
const Camera *cam;
Roi roi;
int hand = 0; // 0: not assigned yet
Landmarks lm;
bool has_lm = false;
int frames = 0;
bool fresh = false; // lm is from this step
};
struct Tile {
const Camera *cam;
V2 center;
double size, rotation, weight, credit = 0;
};
void run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views);
bool hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns);
double pair_cost(const View &a, const View &b) const;
double size_misfit(const std::vector<const View *> &views, const V3 *pts) const;
void associate();
static void smooth(Hand &h, int64_t t_ns);
bool inside(const Camera &cam, V2 uv) const;
void add_tiles(const Camera &cam, double frac, int gx, int gy);
std::map<std::string, const Camera *> cams_;
const Nets &nets_;
Pool &pool_;
int max_views_, hand_budget_ = 4, search_budget_ = 3;
double min_presence_ = 0.5, keep_presence_ = 0.5;
std::vector<View> views_;
std::map<int, Hand> hands_;
std::vector<Tile> tiles_;
int next_id_ = 1;
int64_t last_ns_ = 0, search_ns_ = 0;
};
+343 -7
View File
@@ -10,13 +10,21 @@ to the input relay over its control socket (@frametop_relay):
helper through SteamVR input (@ft_pointer_helper: vrstatus, vrglobal), and a mapped
button is taken from games.
- Pointer: speed, dot size, distance and the rest, applied live.
- Ignored panels: SteamVR overlays the pointer passes through (POINTER_IGNORE), by app or
one by one. The helper lists them (@ft_pointer_helper "overlays").
- Gaze: the pointer's gaze mode (@ft_pointer_helper "gaze") and the gaze service
(gaze/ft-gazed, @ft_gazed: status, forget, reload).
(gaze/ft-gazed, @ft_gazed: status, forget, reload), with its eye tracker (SteamVR's or
our own) and eye bias (GAZE_TRACKER, GAZE_EYE in frametop.conf).
- Bluetooth: paired devices, and re-applying the Bluetooth LE fixes after pairing.
- A warning on every page when SteamVR won't load the ft_pointer driver (blocked after a
crash, disabled, or SteamVR in safe mode) or hasn't loaded it (@ft_pointer doesn't answer
while SteamVR runs): the cursor still moves, but no click lands. Checked at startup and
every 30 minutes.
Rules go to ~/.config/frametop-input.json and pointer settings to
~/.config/frametop.conf; then the relay (and through it the helper) reloads.
Launch with input-settings/ft-input-settings (host wrapper).
"""
import fnmatch
import json
import os
import re
@@ -36,6 +44,10 @@ CONF_PATH = os.path.expanduser("~/.config/frametop.conf")
RELAY = "\0frametop_relay"
HELPER = "\0ft_pointer_helper"
GAZED = "\0ft_gazed"
DRIVER = "\0ft_pointer" # the ft_pointer driver's control socket, bound while SteamVR has it loaded
# SteamVR's settings; older installs keep them under Steam's config.
VRSETTINGS_PATHS = [os.path.expanduser("~/.config/openvr/config/steamvr.vrsettings"),
os.path.expanduser("~/.steam/steam/config/steamvr.vrsettings")]
REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GAZE_PROBE = os.path.join(REPO, "gaze", "probe", "ft-gazeprobe")
BTN_MISC = 0x100
@@ -48,11 +60,22 @@ ACTION_LABELS = {
"scroll_up": "Scroll up", "scroll_down": "Scroll down", "dashboard": "Toggle SteamVR dashboard",
"recenter": "Recenter pointer", "pointer_toggle": "Pointer on/off",
"follow_toggle": "Head follow on/off (experimental)", "gaze_toggle": "Gaze pointer on/off (experimental)",
"gaze_precision": "Gaze precision: hold to steer, release to click",
"gaze_drag": "Gaze drag: press where you look, steer, release",
"sens_up": "Faster pointer",
"sens_down": "Slower pointer", "layout_reset": "Reset desktop screen layout",
"screens_toggle": "Hide/show desktop screens", "key": "Pass through as key",
"screens_toggle": "Hide/show desktop screens", "keyboard_toggle": "Open/close keyboard",
"key": "Pass through as key",
"none": "Do nothing",
}
# When Frametop's keyboard opens ("vr_keyboard" in the rules; the relay's
# VR_KEYBOARD_MODES). "no_keyboard" is the default.
VR_KEYBOARD_MODES = {
"always": "When a text field is selected",
"no_keyboard": "When a text field is selected and no keyboard is connected",
"button": "Only with a mapped mouse or controller button",
"never": "Never",
}
ROLE_LABELS = {"pointer": "3D pointer", "passthrough": "Pass through", "ignore": "Ignore"}
# Frame controller buttons the pointer helper can read (pointer/helper/vrbuttons.h). The
# system button stays SteamVR's.
@@ -71,7 +94,21 @@ GAZE_SETTINGS = [
("POINTER_GAZE_NUDGE_MAX", "Largest nudge to learn", 8, 1, 30, 0.5, "°"),
("POINTER_GAZE_HOLD", "Hold still to drag", 0.5, 0.1, 2.0, 0.05, "s"),
("POINTER_GAZE_SHOW", "Dot shows after moving", 1.0, 0.0, 5.0, 0.1, "s"),
("POINTER_PRECISION_GAIN", "Precision steering", 0.5, 0.1, 2.0, 0.05, "×"),
("POINTER_PRECISION_DEADZONE", "Precision dead zone", 0.3, 0.0, 3.0, 0.1, "°"),
("POINTER_GAZE_DRAG_GAIN", "Drag steering", 1.0, 0.1, 2.0, 0.05, "×"),
]
# In gaze mode, what the mouse's left button does (POINTER_GAZE_MOUSE).
GAZE_MOUSE = {"precision": "Gaze precision: hold to steer with the mouse, release to click",
"direct": "Click right away where the pointer is"}
# The hand role the pointer's virtual controller takes (POINTER_ROLE).
POINTER_ROLES = {"right": "Right hand", "left": "Left hand", "stylus": "Stylus (no hand)"}
# Key combinations ("key_bindings" in the rules): modifiers, either side folded into the left code.
MODIFIER_CODES = {29: 29, 97: 29, 42: 42, 54: 42, 56: 56, 100: 56, 125: 125, 126: 125}
MODIFIER_NAMES = {29: "Ctrl", 42: "Shift", 56: "Alt", 125: "Meta"}
# The gaze service's settings (gaze/ft-gazed): whose eye tracking, and the eye bias.
GAZE_TRACKERS = {"steam": "SteamVR's eye tracker", "own": "our own eye tracker"}
GAZE_EYES = {"auto": "auto", "left": "left eye", "right": "right eye"}
# Pointer settings: key, label, default, min, max, step, unit.
POINTER_SETTINGS = [
("POINTER_SENSITIVITY", "Speed", 0.03, 0.005, 0.12, 0.001, "°/count"),
@@ -87,7 +124,15 @@ POINTER_SETTINGS = [
("POINTER_FOLLOW_REACH", "Head follow reach", 70, 20, 85, 1, "°"),
("POINTER_WAKE_COUNTS", "Movement to wake", 40, 5, 200, 5, "counts"),
("POINTER_IDLE", "Release after idle", 30, 5, 120, 5, "s"),
("POINTER_CONTROLLER_PICKUP", "Controller movement to take over", 1.0, 0.5, 5.0, 0.1, "×"),
]
# Overlay keys (shell patterns) the pointer passes through, comma-separated.
IGNORE_KEY = "POINTER_IGNORE"
def overlay_app(key):
"""The app an overlay key belongs to, by the vendor.app.overlay convention."""
return ".".join(key.split(".")[:2])
def code_names():
@@ -156,6 +201,29 @@ def host(*cmd):
return subprocess.CompletedProcess(cmd, 1, "", str(e))
def driver_block():
"""Why SteamVR won't load the ft_pointer driver: "blocked" (safe mode blocked it after a
crash), "disabled" (turned off in Manage Add-Ons), "safemode" (SteamVR in safe mode, every
add-on off) or "" (nothing stops it). SteamVR reads these at startup."""
for path in VRSETTINGS_PATHS:
if os.path.exists(path):
settings = read_json(path)
break
else:
return ""
section = lambda name: settings.get(name) if isinstance(settings.get(name), dict) else {}
if not isinstance(settings, dict):
return ""
driver = section("driver_ft_pointer")
if driver.get("blocked_by_safe_mode") is True:
return "blocked"
if driver.get("enable") is False:
return "disabled"
if section("steamvr").get("enableSafeMode") is True:
return "safemode"
return ""
class Backend(QObject):
devicesChanged = Signal()
mappingsChanged = Signal()
@@ -163,9 +231,12 @@ class Backend(QObject):
bluetoothChanged = Signal()
controllersChanged = Signal()
gazeChanged = Signal()
driverChanged = Signal()
panelsChanged = Signal()
activity = Signal(str) # device id
captured = Signal(int, str) # code, name
capturedController = Signal(str, str) # button, label
shortcutCaptureChanged = Signal()
message = Signal(str, bool) # text, is error
def __init__(self):
@@ -177,12 +248,16 @@ class Backend(QObject):
self._bluetooth = []
self._relay_ok = False
self._capture_vr = False
self._capture_combo = "" # the action a key combination is being captured for
self._combo_mods = set()
self._vr = {} # the helper's vrstatus, {} when it doesn't answer
self._vr_at = 0.0
self._gaze = {} # ft-gazed's status, {} when it isn't running
self._gaze_prev = None # the status before, for rates
self._gaze_at = 0.0
self._gaze_mode = None # the helper's gaze mode: True, False, None (no answer)
self._driver_block = "" # set by _check_driver
self._panels = None # SteamVR's overlays, from the helper; None until it answers
self.sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
self.sock.bind("") # autobind an abstract address the relay can reply to
self.sock.setblocking(False)
@@ -193,6 +268,11 @@ class Backend(QObject):
self.rewatch = QTimer(interval=50000, timeout=lambda: self._send("watch 60"))
self.rewatch.start()
self.reload_timer = QTimer(singleShot=True, interval=400, timeout=lambda: self._send("reload"))
# The driver only changes with a SteamVR restart, so a rare check is enough. The first one
# waits a moment for the helper's vrstatus answer, which tells us SteamVR is up.
self.driver_timer = QTimer(interval=30 * 60 * 1000, timeout=self._check_driver)
self.driver_timer.start()
QTimer.singleShot(3000, self._check_driver)
self._refresh()
self._send("watch 60")
self.refreshBluetooth()
@@ -224,6 +304,14 @@ class Backend(QObject):
self._gaze_mode = None
self.gazeChanged.emit()
def _check_driver(self):
block = driver_block()
if not block and self._vr and not self._send("ping", DRIVER):
block = "unloaded" # SteamVR runs (the helper answers) without the driver
if block != self._driver_block:
self._driver_block = block
self.driverChanged.emit()
def _read(self):
while True:
try:
@@ -251,6 +339,11 @@ class Backend(QObject):
self._pointer_mode = bool(msg.get("pointer_mode"))
self._relay_ok = True
self.devicesChanged.emit()
elif t == "overlays":
panels = [o for o in msg.get("list", []) if isinstance(o, dict) and isinstance(o.get("key"), str)]
if panels != self._panels:
self._panels = panels
self.panelsChanged.emit()
elif t == "vrstatus":
self._vr = msg
self._vr_at = time.monotonic()
@@ -261,6 +354,17 @@ class Backend(QObject):
self._capture_vr = False
self._send("vrcapture 0")
self.capturedController.emit(msg["code"], CONTROLLER_BUTTONS[msg["code"]])
elif t == "event" and self._capture_combo and msg.get("type") == "key":
self.activity.emit(msg["id"])
code, value = int(msg["code"]), msg["value"]
if code in MODIFIER_CODES:
(self._combo_mods.add if value else self._combo_mods.discard)(MODIFIER_CODES[code])
elif value == 1 and code < BTN_MISC:
self._save_shortcut("+".join(str(c) for c in sorted(self._combo_mods) + [code]),
self._capture_combo)
self._capture_combo = ""
self._combo_mods = set()
self.shortcutCaptureChanged.emit()
elif t == "event":
self.activity.emit(msg["id"])
if (self._capture_id and msg["id"] == self._capture_id and msg["type"] == "key"
@@ -269,6 +373,13 @@ class Backend(QObject):
self._capture_id = ""
self.captured.emit(code, self.codeName(code))
# --- SteamVR driver ---
@Property(str, notify=driverChanged)
def driverBlock(self):
"""driver_block(), or "unloaded": SteamVR runs without the driver (unblocked since it
started, or not installed)."""
return self._driver_block
# --- devices ---
@Property(bool, notify=devicesChanged)
def relayRunning(self):
@@ -288,7 +399,8 @@ class Backend(QObject):
grouped = {}
for n in self._nodes:
d = grouped.setdefault(n["id"], {"id": n["id"], "name": n["name"], "bus": n["bus"], "kinds": [],
"nodes": [], "role": n["role"], "grabbed": False})
"nodes": [], "role": n["role"], "grabbed": False,
"uinput": n.get("uinput", False)})
d["nodes"].append(n["path"])
d["kinds"] = sorted(set(d["kinds"]) | set(n["kinds"]))
d["grabbed"] = d["grabbed"] or n["grabbed"]
@@ -452,6 +564,110 @@ class Backend(QObject):
def recenter(self):
return self._send("recenter", HELPER)
# --- ignored panels ---
@staticmethod
def _ignore_list():
return [p.strip() for p in read_conf().get(IGNORE_KEY, "").split(",") if p.strip()]
def _save_ignore(self, entries, text):
write_conf_value(IGNORE_KEY, ", ".join(entries))
# Only the helper reads it; the relay passes "reload" on only in pointer mode.
self._send("reload", HELPER)
self.panelsChanged.emit()
self.message.emit(text, False)
def _open_panels(self):
"""The helper's overlays, minus Frametop's own: ignoring a screen would leave nothing
to click this app on with the mouse."""
return [o for o in self._panels or [] if not o["key"].startswith("frametop.")]
@Slot()
def refreshPanels(self):
"""Ask the helper for the overlay list; it answers once it has listed them again."""
self._send("overlays", HELPER)
@Property(bool, notify=panelsChanged)
def panelsLoaded(self):
return self._panels is not None
@Property("QVariantList", notify=panelsChanged)
def panelGroups(self):
"""Open overlays by app: {app, title, appIgnored, anyVisible, anyIgnored, panels: [{key,
name, visible, ignoredBy}]}. ignoredBy is the entry that ignores it ("" if none)."""
entries = self._ignore_list()
groups = {}
for o in self._open_panels():
key = o["key"]
app = overlay_app(key)
g = groups.setdefault(app, {"app": app, "title": app, "panels": []})
name = o.get("name") or key
if key == app:
g["title"] = name
g["panels"].append({"key": key, "name": name, "visible": bool(o.get("visible")),
"ignoredBy": next((p for p in entries if fnmatch.fnmatchcase(key, p)), "")})
for g in groups.values():
g["appIgnored"] = g["app"] + "*" in entries
g["anyVisible"] = any(p["visible"] for p in g["panels"])
g["anyIgnored"] = any(p["ignoredBy"] for p in g["panels"])
g["panels"].sort(key=lambda p: (not p["visible"], p["key"]))
return sorted(groups.values(), key=lambda g: (not g["anyVisible"], g["title"].lower()))
@Property("QVariantList", notify=panelsChanged)
def ignoreOrphans(self):
"""Entries that match no open overlay (the app isn't running), so they can be removed."""
keys = [o["key"] for o in self._open_panels()]
return [p for p in self._ignore_list() if not any(fnmatch.fnmatchcase(k, p) for k in keys)]
@Slot(str, bool)
def setPanelIgnored(self, key, on):
entries = [p for p in self._ignore_list() if p != key]
if on:
entries.append(key)
self._save_ignore(entries, f"{key}: " + ("the pointer passes through it" if on else "the pointer lands on it again"))
@Slot(str, bool)
def setAppIgnored(self, app, on):
pattern = app + "*"
entries = [p for p in self._ignore_list() if p != pattern]
if on:
entries.append(pattern)
self._save_ignore(entries, f"{app}: " + ("the pointer passes through all its panels" if on
else "no longer ignored as a whole"))
@Slot(str)
def removeIgnore(self, pattern):
self._save_ignore([p for p in self._ignore_list() if p != pattern], f"{pattern}: no longer ignored")
# --- Frametop's keyboard ---
@Property("QVariantList", constant=True)
def vrKeyboardModes(self):
return [{"value": k, "text": v} for k, v in VR_KEYBOARD_MODES.items()]
@Property(str, notify=mappingsChanged)
def vrKeyboard(self):
mode = read_json(RULES_PATH).get("vr_keyboard")
return mode if mode in VR_KEYBOARD_MODES else "no_keyboard"
@Property(bool, notify=mappingsChanged)
def vrKeyboardPersist(self):
return bool(read_json(RULES_PATH).get("vr_keyboard_persist", True))
@Slot(bool)
def setVrKeyboardPersist(self, on):
rules = read_json(RULES_PATH)
rules["vr_keyboard_persist"] = bool(on)
self._save_rules(rules)
self.message.emit("Keyboard: " + ("stays open until you hide it" if on else "closes with the text field"), False)
@Slot(str)
def setVrKeyboard(self, mode):
if mode not in VR_KEYBOARD_MODES:
return
rules = read_json(RULES_PATH)
rules["vr_keyboard"] = mode
self._save_rules(rules)
self.message.emit(f"Keyboard: {VR_KEYBOARD_MODES[mode].lower()}", False)
# --- controllers ---
@Property("QVariantList", constant=True)
def controllerActions(self):
@@ -540,9 +756,12 @@ class Backend(QObject):
n = status["samples"] - prev[1]["samples"]
status["rate"] = n / (now - prev[0])
status["one_eye_share"] = (status["one_eye"] - prev[1]["one_eye"]) / n if n else 0.0
for key in ("lost_left", "lost_right"):
if key in status and key in prev[1]:
status[key + "_share"] = (status[key] - prev[1][key]) / n if n else 0.0
elif self._gaze:
status.setdefault("rate", self._gaze.get("rate"))
status.setdefault("one_eye_share", self._gaze.get("one_eye_share"))
for key in ("rate", "one_eye_share", "lost_left_share", "lost_right_share"):
status.setdefault(key, self._gaze.get(key))
if not prev or now - prev[0] > 0.5:
self._gaze_prev = (now, status)
self._gaze = status
@@ -573,6 +792,115 @@ class Backend(QObject):
self.reload_timer.start()
self.gazeChanged.emit()
@Property(str, notify=pointerChanged)
def gazeMouse(self):
v = read_conf().get("POINTER_GAZE_MOUSE", "precision")
return v if v in GAZE_MOUSE else "precision"
@Property("QVariantList", constant=True)
def gazeMouseChoices(self):
return [{"value": k, "text": v} for k, v in GAZE_MOUSE.items()]
@Slot(str)
def setGazeMouse(self, mode):
if mode in GAZE_MOUSE:
write_conf_value("POINTER_GAZE_MOUSE", mode)
self.reload_timer.start()
self.pointerChanged.emit()
self.message.emit(f"Mouse in gaze mode: {GAZE_MOUSE[mode].lower()}", False)
@Property(str, notify=pointerChanged)
def pointerRole(self):
v = read_conf().get("POINTER_ROLE", "right")
return v if v in POINTER_ROLES else "right"
@Property("QVariantList", constant=True)
def pointerRoles(self):
return [{"value": k, "text": v} for k, v in POINTER_ROLES.items()]
@Slot(str)
def setPointerRole(self, role):
if role in POINTER_ROLES:
write_conf_value("POINTER_ROLE", role)
self.reload_timer.start()
self.pointerChanged.emit()
self.message.emit(f"Pointer role: {POINTER_ROLES[role].lower()} (from its next wake)", False)
# --- key combinations ("key_bindings") ---
def comboName(self, combo):
parts = [int(c) for c in combo.split("+") if c.isdigit()]
return "+".join(MODIFIER_NAMES.get(c) or self.codeName(c).removeprefix("KEY_").title() for c in parts)
@Property("QVariantList", notify=mappingsChanged)
def keyShortcuts(self):
bound = read_json(RULES_PATH).get("key_bindings", {}) or {}
return [{"combo": c, "label": self.comboName(c), "action": a, "actionLabel": ACTION_LABELS.get(a, a)}
for c, a in sorted(bound.items())]
@Property("QVariantList", constant=True)
def shortcutActions(self):
return [{"value": a, "text": ACTION_LABELS[a]} for a in CONTROLLER_ACTIONS]
@Property(bool, notify=shortcutCaptureChanged)
def capturingShortcut(self):
return bool(self._capture_combo)
@Slot(str)
def startShortcutCapture(self, action):
if action in CONTROLLER_ACTIONS:
self._capture_combo = action
self._combo_mods = set()
self._send("watch 60")
self.shortcutCaptureChanged.emit()
@Slot()
def cancelShortcutCapture(self):
self._capture_combo = ""
self.shortcutCaptureChanged.emit()
def _save_shortcut(self, combo, action):
rules = read_json(RULES_PATH)
rules.setdefault("key_bindings", {})[combo] = action
self._save_rules(rules)
self.message.emit(f"{self.comboName(combo)} → {ACTION_LABELS[action]}", False)
@Slot(str)
def removeShortcut(self, combo):
rules = read_json(RULES_PATH)
(rules.get("key_bindings") or {}).pop(combo, None)
self._save_rules(rules)
self.message.emit(f"{self.comboName(combo)} removed", False)
@Property(str, notify=gazeChanged)
def gazeTracker(self):
"""Whose eye tracking the gaze service uses: "steam" or "own" (GAZE_TRACKER)."""
v = read_conf().get("GAZE_TRACKER", "steam")
return v if v in GAZE_TRACKERS else "steam"
@Property(str, notify=gazeChanged)
def gazeEye(self):
"""The eye bias: "auto", "left" or "right" (GAZE_EYE)."""
v = read_conf().get("GAZE_EYE", "auto")
return v if v in GAZE_EYES else "auto"
@Slot(str)
def setGazeTracker(self, tracker):
if tracker in GAZE_TRACKERS:
self._set_gaze("GAZE_TRACKER", tracker, f"Gaze from {GAZE_TRACKERS[tracker]}")
@Slot(str)
def setGazeEye(self, eye):
if eye in GAZE_EYES:
self._set_gaze("GAZE_EYE", eye, f"Eye bias: {GAZE_EYES[eye]}")
def _set_gaze(self, key, value, done):
"""ft-gazed reads these from frametop.conf; "reload" makes it do so now."""
write_conf_value(key, value)
running = self._send("reload", GAZED)
self.message.emit(done if running else f"{done} (the gaze service isn't running: it takes it when it starts)",
False)
self.gazeChanged.emit()
@Property("QVariantList", notify=pointerChanged)
def gazeSettings(self):
conf = read_conf()
@@ -603,13 +931,21 @@ class Backend(QObject):
@Slot()
def openGazeProbe(self):
"""Calibrate in ft-gazeprobe (a GTK app on the host, fullscreen on a Frametop screen)."""
self._open_probe([], "Opening the gaze probe: calibrate there, then close it")
@Slot()
def openHeadsetFit(self):
"""The probe's Headset fit mode: how well the tracker sees each eye, as you adjust."""
self._open_probe(["--mode", "fit"], "Opening the headset fit check in the gaze probe")
def _open_probe(self, args, done):
runner = ["distrobox-host-exec"] if shutil.which("distrobox-host-exec") else []
env = [f"{k}={os.environ[k]}" for k in ("WAYLAND_DISPLAY", "DISPLAY", "XAUTHORITY", "DBUS_SESSION_BUS_ADDRESS")
if os.environ.get(k)]
try:
subprocess.Popen(runner + ["env"] + env + [GAZE_PROBE], stdin=subprocess.DEVNULL,
subprocess.Popen(runner + ["env"] + env + [GAZE_PROBE] + args, stdin=subprocess.DEVNULL,
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, start_new_session=True)
self.message.emit("Opening the gaze probe: calibrate there, then close it", False)
self.message.emit(done, False)
except OSError as e:
self.message.emit(f"Couldn't open the gaze probe: {e}", True)
+378 -13
View File
@@ -19,19 +19,46 @@ Kirigami.ApplicationWindow {
Kirigami.Action { text: "Devices"; icon.name: "input-mouse"; onTriggered: root.show(devicesPage) },
Kirigami.Action { text: "Buttons"; icon.name: "input-keyboard"; onTriggered: root.show(buttonsPage) },
Kirigami.Action { text: "Controllers"; icon.name: "input-gamepad"; onTriggered: root.show(controllersPage) },
Kirigami.Action { text: "Keyboard"; icon.name: "input-keyboard-virtual"; onTriggered: root.show(keyboardPage) },
Kirigami.Action { text: "Pointer"; icon.name: "transform-move"; onTriggered: root.show(pointerPage) },
Kirigami.Action { text: "Ignored panels"; icon.name: "view-hidden"; onTriggered: root.show(ignorePage) },
Kirigami.Action { text: "Gaze"; icon.name: "view-visible"; onTriggered: root.show(gazePage) },
Kirigami.Action { text: "Bluetooth"; icon.name: "preferences-system-bluetooth"; onTriggered: root.show(bluetoothPage) }
]
}
// Why SteamVR has no ft_pointer driver, and how to get it back. Clicks, scrolling and mapped
// actions all go through that driver, so every page shows this as (part of) its header.
component DriverWarning: Kirigami.InlineMessage {
readonly property string fix: "SteamVR Settings > Startup / Shutdown > Manage Add-Ons"
readonly property string restart: ", then restart SteamVR (or reboot the headset)."
visible: backend.driverBlock !== ""
position: Kirigami.InlineMessage.Position.Header
type: backend.driverBlock === "unloaded" ? Kirigami.MessageType.Warning : Kirigami.MessageType.Error
text: ({
blocked: "SteamVR blocked the Frametop pointer driver (ft_pointer) after a crash, so mouse clicks "
+ "and scrolling do nothing (the cursor still moves). To fix it, open " + fix
+ ", press Unblock next to ft_pointer" + restart,
disabled: "The Frametop pointer driver (ft_pointer) is turned off in SteamVR, so mouse clicks and "
+ "scrolling do nothing (the cursor still moves). To fix it, open " + fix
+ ", turn ft_pointer on" + restart,
safemode: "SteamVR is in safe mode, so it loads no add-ons, the Frametop pointer driver (ft_pointer) "
+ "included: mouse clicks and scrolling do nothing (the cursor still moves). To fix it, turn "
+ "safe mode off in SteamVR and check " + fix + restart,
unloaded: "SteamVR is running without the Frametop pointer driver (ft_pointer), so mouse clicks and "
+ "scrolling do nothing. If you just unblocked it, restart SteamVR (or reboot the headset). "
+ "Otherwise check " + fix + ", or reinstall it with pointer/driver/install.sh."
})[backend.driverBlock] || ""
}
function show(page) {
pageStack.clear()
pageStack.push(page)
}
// FT_INPUT_PAGE=buttons|controllers|pointer|gaze|bluetooth opens the app on that page.
pageStack.initialPage: ({ buttons: buttonsPage, controllers: controllersPage, pointer: pointerPage, gaze: gazePage,
// FT_INPUT_PAGE=buttons|controllers|keyboard|pointer|ignore|gaze|bluetooth opens the app on that page.
pageStack.initialPage: ({ buttons: buttonsPage, controllers: controllersPage, keyboard: keyboardPage,
pointer: pointerPage, ignore: ignorePage, gaze: gazePage,
bluetooth: bluetoothPage })[startPage] || devicesPage
Connections {
@@ -47,13 +74,18 @@ Kirigami.ApplicationWindow {
Kirigami.ScrollablePage {
title: "Devices"
header: Kirigami.InlineMessage {
visible: !backend.relayRunning || !backend.pointerMode
position: Kirigami.InlineMessage.Position.Header
type: backend.relayRunning ? Kirigami.MessageType.Information : Kirigami.MessageType.Error
text: !backend.relayRunning
? "The input relay isn't running (frametop-input-relay.service)."
: "Pointer mode is off (POINTER=0): pointer devices act as a plain mouse."
header: ColumnLayout {
spacing: 0
DriverWarning { Layout.fillWidth: true }
Kirigami.InlineMessage {
Layout.fillWidth: true
visible: !backend.relayRunning || !backend.pointerMode
position: Kirigami.InlineMessage.Position.Header
type: backend.relayRunning ? Kirigami.MessageType.Information : Kirigami.MessageType.Error
text: !backend.relayRunning
? "The input relay isn't running (frametop-input-relay.service)."
: "Pointer mode is off (POINTER=0): pointer devices act as a plain mouse."
}
}
ListView {
@@ -171,6 +203,7 @@ Kirigami.ApplicationWindow {
Kirigami.ScrollablePage {
id: bpage
title: "Buttons"
header: DriverWarning {}
actions: [
Kirigami.Action {
text: "Remove all"
@@ -310,6 +343,7 @@ Kirigami.ApplicationWindow {
Kirigami.ScrollablePage {
id: cpage
title: "Controllers"
header: DriverWarning {}
actions: [
Kirigami.Action {
text: "Remove all"
@@ -473,11 +507,60 @@ Kirigami.ApplicationWindow {
}
}
// ---------------------------------------------------------------- Keyboard
Component {
id: keyboardPage
Kirigami.ScrollablePage {
id: kpage
title: "Keyboard"
// Pass-through keyboards connected now: with "no_keyboard", the keyboard waits for none.
// A program's uinput keyboard (frame-voice's, say) doesn't count.
property var keyboards: backend.devices.filter(d => d.connected && d.role === "passthrough"
&& d.kinds.indexOf("keyboard") >= 0 && !d.uinput)
Kirigami.FormLayout {
Controls.ComboBox {
Kirigami.FormData.label: "Show the keyboard:"
model: backend.vrKeyboardModes
textRole: "text"
valueRole: "value"
Component.onCompleted: currentIndex = indexOfValue(backend.vrKeyboard)
onActivated: backend.setVrKeyboard(currentValue)
}
Controls.Switch {
Kirigami.FormData.label: "Keep it open:"
text: "Until you press its Close key or your keyboard button, not only while the text field has focus"
checked: backend.vrKeyboardPersist
enabled: backend.vrKeyboard !== "never"
onToggled: backend.setVrKeyboardPersist(checked)
}
Controls.Label {
Kirigami.FormData.label: "Keyboards connected:"
text: kpage.keyboards.length ? kpage.keyboards.map(d => d.name).join(", ") : "none"
}
}
footer: Controls.Label {
padding: Kirigami.Units.largeSpacing
wrapMode: Text.Wrap
opacity: 0.7
text: "Frametop's keyboard opens in front of you, below your eyes, and closes with its Close key or a layout reset "
+ "(or when the text field loses focus, with Keep it open off). It steps aside while the Steam "
+ "menu or Steam's own keyboard is up. Type on it with a laser or the 3D mouse. It types into any "
+ "app, but only Qt, GTK and Firefox apps say when a text field is selected; for the rest "
+ "(Chromium, Electron and X11 apps), map Open/close keyboard to a button on the Buttons or "
+ "Controllers page. Keyboards set to Ignore, and keyboards other programs make (like "
+ "frame-voice's), don't count as connected."
}
}
}
// ---------------------------------------------------------------- Pointer
Component {
id: pointerPage
Kirigami.ScrollablePage {
title: "Pointer"
header: DriverWarning {}
actions: [
Kirigami.Action {
text: "Recenter"
@@ -543,12 +626,149 @@ Kirigami.ApplicationWindow {
}
}
// ---------------------------------------------------------------- Ignored panels
Component {
id: ignorePage
Kirigami.ScrollablePage {
id: ipage
title: "Ignored panels"
header: DriverWarning {}
Timer {
// The helper lists SteamVR's panels again for each request.
running: true
repeat: true
triggeredOnStart: true
interval: 3000
onTriggered: backend.refreshPanels()
}
ColumnLayout {
spacing: Kirigami.Units.largeSpacing
Kirigami.InlineMessage {
Layout.fillWidth: true
visible: !backend.controllerStatus.helper
type: Kirigami.MessageType.Error
text: "The pointer helper isn't answering (frametop-pointer.service, needs SteamVR). "
+ "It lists SteamVR's panels and does the ignoring."
}
Controls.Label {
Layout.fillWidth: true
wrapMode: Text.Wrap
text: "The mouse pointer passes through the panels ticked here to whatever is behind them, "
+ "as if they weren't there. Use it for panels you only look at, like a performance "
+ "overlay that follows your view. Ignore a whole app, or only some of its panels. "
+ "Controllers aren't affected."
}
Controls.Switch {
id: showHidden
text: "Also list panels that aren't showing now"
}
Controls.Label {
visible: backend.controllerStatus.helper && !backend.panelsLoaded
text: "Asking the pointer helper for SteamVR's panels…"
opacity: 0.7
}
Repeater {
model: backend.panelGroups
delegate: ColumnLayout {
id: grp
required property var modelData
Layout.fillWidth: true
visible: showHidden.checked || modelData.anyVisible || modelData.anyIgnored
spacing: 0
RowLayout {
Layout.fillWidth: true
Kirigami.Heading {
level: 4
text: grp.modelData.title
}
Controls.Label {
text: grp.modelData.app
opacity: 0.6
}
Item { Layout.fillWidth: true }
Controls.CheckBox {
text: "Ignore the whole app"
checked: grp.modelData.appIgnored
onToggled: backend.setAppIgnored(grp.modelData.app, checked)
}
}
Repeater {
model: grp.modelData.panels
delegate: RowLayout {
id: prow
required property var modelData
// Ignored by the whole app, or by a pattern written in frametop.conf.
readonly property bool byOther: modelData.ignoredBy !== "" && modelData.ignoredBy !== modelData.key
visible: showHidden.checked || modelData.visible || modelData.ignoredBy !== ""
Layout.leftMargin: Kirigami.Units.gridUnit
Controls.CheckBox {
text: prow.modelData.name
checked: prow.modelData.ignoredBy !== ""
enabled: !prow.byOther
onToggled: backend.setPanelIgnored(prow.modelData.key, checked)
}
Controls.Label {
text: prow.modelData.key
opacity: 0.6
}
Controls.Label {
text: (prow.modelData.visible ? "showing" : "hidden")
+ (!prow.byOther ? ""
: prow.modelData.ignoredBy === grp.modelData.app + "*" ? ", whole app ignored"
: ", ignored by " + prow.modelData.ignoredBy)
opacity: 0.6
}
}
}
}
}
Kirigami.Heading {
visible: backend.ignoreOrphans.length > 0
level: 3
text: "Ignored, not open now"
}
Repeater {
model: backend.ignoreOrphans
delegate: RowLayout {
id: orow
required property string modelData
Controls.Label {
text: orow.modelData
Layout.preferredWidth: Kirigami.Units.gridUnit * 16
}
Controls.Button {
text: "Remove"
icon.name: "edit-delete-remove"
onClicked: backend.removeIgnore(orow.modelData)
}
}
}
}
footer: Controls.Label {
padding: Kirigami.Units.largeSpacing
wrapMode: Text.Wrap
opacity: 0.7
text: "Saved as POINTER_IGNORE in ~/.config/frametop.conf, and applied at once. A whole app is "
+ "its key followed by *, which also covers panels it opens later. Frametop's own screens "
+ "aren't listed."
}
}
}
// ---------------------------------------------------------------- Gaze
Component {
id: gazePage
Kirigami.ScrollablePage {
id: gpage
title: "Gaze"
header: DriverWarning {}
property var status: backend.gazeStatus
actions: [
Kirigami.Action {
@@ -557,6 +777,12 @@ Kirigami.ApplicationWindow {
tooltip: "Open the gaze probe to calibrate (fullscreen on a Frametop screen)"
onTriggered: backend.openGazeProbe()
},
Kirigami.Action {
text: "Check headset fit…"
icon.name: "view-visible"
tooltip: "How well the eye tracker sees each eye, and where it loses one, while you adjust the headset"
onTriggered: backend.openHeadsetFit()
},
Kirigami.Action {
text: "Reload calibration"
icon.name: "view-refresh"
@@ -595,6 +821,126 @@ Kirigami.ApplicationWindow {
opacity: 0.7
font: Kirigami.Theme.smallFont
}
ColumnLayout {
Kirigami.FormData.label: "Mouse left button:"
Repeater {
model: backend.gazeMouseChoices
delegate: Controls.RadioButton {
required property var modelData
text: modelData.text
checked: backend.gazeMouse === modelData.value
onToggled: if (checked) backend.setGazeMouse(modelData.value)
}
}
}
Controls.Label {
Layout.maximumWidth: Kirigami.Units.gridUnit * 30
wrapMode: Text.WordWrap
text: "Controllers: map a button to Gaze precision (hold, point the controller to steer, release to "
+ "click) or Gaze drag (the same, pressed at once) on the Controllers page. Gaze pointer on/off "
+ "can go on a controller button, a mouse button, or a key combination below."
opacity: 0.7
font: Kirigami.Theme.smallFont
}
Kirigami.Separator { Kirigami.FormData.isSection: true; Kirigami.FormData.label: "Key combinations" }
Repeater {
model: backend.keyShortcuts
delegate: RowLayout {
required property var modelData
Kirigami.FormData.label: modelData.label + ":"
Controls.Label { text: modelData.actionLabel }
Controls.ToolButton {
icon.name: "edit-delete"
display: Controls.AbstractButton.IconOnly
text: "Remove"
Controls.ToolTip.text: text
Controls.ToolTip.visible: hovered
onClicked: backend.removeShortcut(modelData.combo)
}
}
}
RowLayout {
Kirigami.FormData.label: "New:"
Controls.ComboBox {
id: shortcutAction
model: backend.shortcutActions
textRole: "text"
valueRole: "value"
Component.onCompleted: currentIndex = indexOfValue("gaze_toggle")
Layout.preferredWidth: Kirigami.Units.gridUnit * 16
}
Controls.Button {
text: backend.capturingShortcut ? "Press the keys… (Cancel)" : "Set keys…"
onClicked: backend.capturingShortcut ? backend.cancelShortcutCapture()
: backend.startShortcutCapture(shortcutAction.currentValue)
}
}
Controls.Label {
Layout.maximumWidth: Kirigami.Units.gridUnit * 30
wrapMode: Text.WordWrap
text: "Hold the modifiers (Ctrl, Alt, Shift, Meta), then press the key, on any keyboard. The "
+ "combination's last key isn't typed; the modifiers still reach the app."
opacity: 0.7
font: Kirigami.Theme.smallFont
}
RowLayout {
Kirigami.FormData.label: "Eye tracker:"
Controls.RadioButton {
text: "SteamVR"
checked: backend.gazeTracker === "steam"
onToggled: if (checked) backend.setGazeTracker("steam")
}
Controls.RadioButton {
text: "Own tracker"
checked: backend.gazeTracker === "own"
onToggled: if (checked) backend.setGazeTracker("own")
}
}
Controls.Label {
// The gaze service runs ours (gaze/tracker/ft-eyes); it needs the frame grabber
// (gaze/tracker/install.sh, root) and keeps its own calibration.
visible: backend.gazeTracker === "own" && backend.gazeServiceRunning
&& (!gpage.status.own_running || gpage.status.own_reseat || !gpage.status.calibration_samples)
text: !gpage.status.eyegrab ? "Needs its frame grabber: run gaze/tracker/install.sh (asks for sudo)"
: !gpage.status.own_running ? "Starting…"
: !gpage.status.calibration_samples ? "Not calibrated: use Calibrate… with Own tracker"
: "The headset was off: your first nudge and click resets where it sits"
color: gpage.status.own_running ? Kirigami.Theme.neutralTextColor : Kirigami.Theme.negativeTextColor
font: Kirigami.Theme.smallFont
}
RowLayout {
Kirigami.FormData.label: "Eye bias:"
Controls.RadioButton {
text: "Auto"
checked: backend.gazeEye === "auto"
onToggled: if (checked) backend.setGazeEye("auto")
}
Controls.RadioButton {
text: "Left"
checked: backend.gazeEye === "left"
onToggled: if (checked) backend.setGazeEye("left")
}
Controls.RadioButton {
text: "Right"
checked: backend.gazeEye === "right"
onToggled: if (checked) backend.setGazeEye("right")
}
}
Controls.Label {
// What the bias comes to now; on auto, from how far off each eye was at the last nudges.
property var w: gpage.status.eye_weights
property var n: gpage.status.eye_misses || [0, 0]
property var rms: gpage.status.eye_rms || [null, null]
visible: backend.gazeServiceRunning && w !== undefined
text: w === undefined ? "" : "Left " + Math.round(w[0] * 100) + "% · right " + Math.round(w[1] * 100) + "%"
+ (backend.gazeEye !== "auto" ? ""
: n[0] >= 5 && n[1] >= 5
? " (off by " + rms[0].toFixed(1) + "° and " + rms[1].toFixed(1) + "° at your last nudges)"
: " (even until each eye has 5 nudges: " + n[0] + " and " + n[1] + " so far)")
opacity: 0.7
font: Kirigami.Theme.smallFont
}
Repeater {
model: backend.gazeSettings
delegate: RowLayout {
@@ -645,16 +991,31 @@ Kirigami.ApplicationWindow {
: Math.round(gpage.status.rate) + " samples/s"
+ (gpage.status.one_eye_share > 0.5 ? " · only one eye tracked" : "")
color: gpage.status.one_eye_share > 0.5 ? Kirigami.Theme.neutralTextColor : Kirigami.Theme.textColor
Controls.ToolTip.text: "Only one eye tracked: reseat the headset or check the lenses. It still works, "
+ "probably less precisely."
Controls.ToolTip.text: "Only one eye tracked: the gaze comes from the other eye, a little less "
+ "precisely. Check headset fit… shows where the tracker loses it."
Controls.ToolTip.visible: gpage.status.one_eye_share > 0.5 && ghover.hovered
HoverHandler { id: ghover }
}
Controls.Label {
// Share of the last second's samples where the tracker had lost each eye.
property real lostL: gpage.status.lost_left_share || 0
property real lostR: gpage.status.lost_right_share || 0
visible: backend.gazeServiceRunning && gpage.status.lost_left !== undefined
Kirigami.FormData.label: "Eyes:"
text: lostL < 0.05 && lostR < 0.05 ? "both tracked"
: (lostL >= 0.05 ? "left eye lost " + Math.round(lostL * 100) + "%" : "")
+ (lostL >= 0.05 && lostR >= 0.05 ? " · " : "")
+ (lostR >= 0.05 ? "right eye lost " + Math.round(lostR * 100) + "%" : "")
+ ((lostL >= 0.05) !== (lostR >= 0.05) ? " (the other eye stands in)" : "")
color: lostL >= 0.05 || lostR >= 0.05 ? Kirigami.Theme.neutralTextColor : Kirigami.Theme.textColor
}
Controls.Label {
visible: backend.gazeServiceRunning
Kirigami.FormData.label: "Calibration:"
text: gpage.status.calibration_samples > 0
? gpage.status.calibration_samples + " points (" + gpage.status.model + ")"
? gpage.status.calibration_samples + " points ("
+ (gpage.status.tracker === "own" ? "Own tracker, " + gpage.status.calibration_made
: gpage.status.kind === "eyes" ? gpage.status.model + ", each eye" : gpage.status.model) + ")"
: "none yet: use Calibrate…"
}
Controls.Label {
@@ -674,7 +1035,10 @@ Kirigami.ApplicationWindow {
+ "with the button held and let go there. Hold still to drag: hold a press this long without moving "
+ "to drag something instead. Look away to hand back: how far from the pointer you look before the "
+ "gaze takes it back from the mouse. Largest nudge to learn: bigger mouse moves before a click "
+ "are treated as using the mouse, not correcting the gaze."
+ "are treated as using the mouse, not correcting the gaze. Eye tracker: SteamVR's, or our own "
+ "(gaze/tracker), which keeps its own calibration (Calibrate… with Own tracker). Eye bias: the gaze "
+ "combines both eyes, since their errors partly cancel; Left or Right counts that eye twice as "
+ "much, and Auto weights each by how far off it was at your recent nudges."
}
}
}
@@ -684,6 +1048,7 @@ Kirigami.ApplicationWindow {
id: bluetoothPage
Kirigami.ScrollablePage {
title: "Bluetooth"
header: DriverWarning {}
actions: [
Kirigami.Action { text: "Refresh"; icon.name: "view-refresh"; onTriggered: backend.refreshBluetooth() },
Kirigami.Action {
+92
View File
@@ -0,0 +1,92 @@
#!/usr/bin/env python3
"""ft-textinput: Frametop's input method, which tells the input relay when a text field
on the desktop has keyboard focus.
KWin starts it (--inputmethod, from session/frametop-session.sh) and activates it
whenever the focused app enables text input on a field (text-input v1/v2/v3: Qt, GTK,
and Firefox apps do; Xwayland apps and most Electron apps don't). This only passes that
on: "textfield 1" or "textfield 0" to the input relay (@frametop_relay), which opens or
closes Frametop's keyboard (ft-screens' own key panel), depending on the keyboard setting
in Frametop Input Settings. Typing on it reaches the app as key presses from ft-screens,
not through this input method, so nothing typed passes through here.
Speaks the Wayland wire protocol itself (zwp_input_method_v1 only), so it needs nothing
but Python on the SteamOS host.
"""
import os
import socket
import struct
import sys
RELAY = "\0frametop_relay"
DISPLAY, REGISTRY, INPUT_METHOD = 1, 2, 3 # our object ids
def string(text):
data = text.encode() + b"\0"
return struct.pack("=I", len(data)) + data + b"\0" * (-len(data) % 4)
def read_string(body, at):
"""(text, offset after it) for a string argument at `at`."""
(length,) = struct.unpack_from("=I", body, at)
text = body[at + 4:at + 4 + length - 1].decode(errors="replace")
return text, at + 4 + length + (-length % 4)
def main():
fd = os.environ.pop("WAYLAND_SOCKET", None)
if fd is None:
sys.exit("ft-textinput: no WAYLAND_SOCKET (KWin starts this as its input method)")
conn = socket.socket(fileno=int(fd))
relay = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM | socket.SOCK_NONBLOCK)
def request(obj, opcode, payload=b""):
conn.sendall(struct.pack("=II", obj, (8 + len(payload)) << 16 | opcode) + payload)
def report(on):
try:
relay.sendto(f"textfield {int(on)}".encode(), RELAY)
except OSError:
pass # the relay isn't running
request(DISPLAY, 1, struct.pack("=I", REGISTRY)) # wl_display.get_registry
contexts = set() # KWin makes one per activation
buf = b""
while True:
data = conn.recv(65536)
if not data:
return # KWin went away
buf += data
while len(buf) >= 8:
obj, word = struct.unpack_from("=II", buf)
size, opcode = word >> 16, word & 0xFFFF
if size < 8 or len(buf) < size:
break
body, buf = buf[8:size], buf[size:]
if obj == DISPLAY and opcode == 0: # error
bad, code = struct.unpack_from("=II", body)
message, _ = read_string(body, 8)
sys.exit(f"ft-textinput: Wayland error {code} on object {bad}: {message}")
elif obj == REGISTRY and opcode == 0: # global
(name,) = struct.unpack_from("=I", body)
interface, _ = read_string(body, 4)
if interface == "zwp_input_method_v1":
request(REGISTRY, 0, struct.pack("=I", name) + string(interface) + struct.pack("=II", 1, INPUT_METHOD))
elif obj == INPUT_METHOD and opcode == 0: # activate: a text field has focus
contexts.add(struct.unpack_from("=I", body)[0])
report(True)
elif obj == INPUT_METHOD and opcode == 1: # deactivate: it lost focus
(context,) = struct.unpack_from("=I", body)
if context in contexts:
contexts.discard(context)
request(context, 0) # zwp_input_method_context_v1.destroy
report(bool(contexts))
# Everything else (the contexts' surrounding text, content type, ...) is unused.
if __name__ == "__main__":
try:
main()
except KeyboardInterrupt:
pass
+306
View File
@@ -0,0 +1,306 @@
"""Gaze first, the relay's part (docs/gaze-first.md).
While the pointer helper says gaze first is on ("gazefirst 1" every 5 s: gaze mode, no game,
the headset on; "gazefirst 0" or 12 s of silence ends it), the Frame controllers belong to the
gaze:
- SteamVR: the Frame controller's compositor binding is
pointer/bindings/vrcompositor_frame_controller_gazefirst.json (only the Steam button left),
chosen through vrserver's /input/selectconfig.action, and the stock one again after.
- Steam's UI: input/steam-gamepad-filter.js runs in Steam's SharedJSContext (input/steamui.py)
and drops the controllers' gamepad input (the Steam button still passes). The block lasts 15 s
unless it's renewed, which happens every 5 s, so it ends by itself if the relay dies; and
Steam reloads its UI with SteamVR, which the renewal also covers.
- The trigger and bumper, read from vrserver's web socket (input/vrws.py, which takes nothing
from anyone), go to the helper as "ctrl <left|right> <trigger|bumper> 1|0"; the right
thumbstick scrolls ("scroll <x> <y>", as the mouse's wheel).
Steam still sees the controllers, and SteamVR leaves laser mode after each press and release:
the helper takes the laser back (see "Gaze first" in pointer/helper/ft-pointer.cpp).
The toggle macro works with gaze first on or off: both thumbstick clicks held 1 s turn gaze mode
on or off, and POINTER_GAZE in ~/.config/frametop.conf remembers it. So does `toggle()`, which
the relay's gaze_toggle action uses.
The binding, the filter, and finding new controllers run on a thread of their own: they're HTTP
and web socket round trips that mustn't hold up the relay's mouse and keyboard.
"""
import json
import os
import threading
import time
import urllib.request
import steamui
from vrws import HEADERS, ORIGIN, VrSocket, getstate
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
FILTER_JS = os.path.join(ROOT, "input", "steam-gamepad-filter.js")
GAZE_BINDING = "file://" + os.path.join(ROOT, "pointer", "bindings", "vrcompositor_frame_controller_gazefirst.json")
STOCK_BINDING = "file:///opt/steamvr/drivers/frame_controller/resources/input/vrcompositor_bindings_frame_controller.json"
COMPOSITOR = "openvr.component.vrcompositor"
CONF = os.path.expanduser("~/.config/frametop.conf")
BUTTONS = {"/input/trigger/click": "trigger", "/input/bumper/click": "bumper"}
STICK_CLICK = "/input/thumbstick/click"
MACRO_HOLD = 1.0 # seconds both thumbstick clicks are held to toggle gaze mode
SCROLL_DEADZONE = 0.3
ON_FOR = 12.0 # "gazefirst 1" lasts this long without another
RENEW = 5.0 # the binding and filter are checked (and the filter's block renewed) this often
BLOCK_FOR = 15000 # ms the filter blocks without a renewal
def select_binding(url):
body = json.dumps({"app_key": COMPOSITOR, "controller_type": "frame_controller", "url": url}).encode()
req = urllib.request.Request(f"{ORIGIN}/input/selectconfig.action", data=body,
headers={**HEADERS, "Origin": ORIGIN, "Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=2) as resp:
if not json.load(resp).get("success"):
raise OSError("vrserver refused the binding")
def current_binding():
req = urllib.request.Request(f"{ORIGIN}/input/getactions.json?app_key={COMPOSITOR}", headers=HEADERS)
with urllib.request.urlopen(req, timeout=2) as resp:
return (json.load(resp).get("current_binding_url") or {}).get("frame_controller", "")
def read_gaze(path=CONF):
try:
with open(path) as f:
for line in f:
line = line.split("#", 1)[0].strip()
if line.split("=", 1)[0].strip() == "POINTER_GAZE" and "=" in line:
return line.split("=", 1)[1].strip() not in ("", "0")
except OSError:
pass
return False
def write_gaze(on, path=CONF):
"""POINTER_GAZE=1|0 in the config, keeping every other line as it is."""
try:
with open(path) as f:
lines = f.read().splitlines()
except OSError:
lines = []
value = f"POINTER_GAZE={1 if on else 0}"
for i, line in enumerate(lines):
if line.split("#", 1)[0].split("=", 1)[0].strip() == "POINTER_GAZE":
lines[i] = value
break
else:
lines.append(value)
tmp = path + ".tmp"
with open(tmp, "w") as f:
f.write("\n".join(lines) + "\n")
os.replace(tmp, path)
class GazeFirst:
def __init__(self, send, log):
self.send = send # a command to the pointer helper
self.log = log
self.on = False
self.on_until = 0.0
self.ws = None
self.next_connect = 0.0
self.sides = {} # subscribed device path -> "left" or "right"
self.values = {} # (side, component) -> last value
self.held = set() # (side, button) the helper has as pressed
self.macro_since = None
self.macro_fired = False
self.scrolling = (0.0, 0.0)
# The worker thread's: a wake-up, the devices it found, what it last applied.
self.lock = threading.Lock()
self.wake = threading.Event()
self.found = []
self.applied = None
self.problems = {}
threading.Thread(target=self.work, daemon=True, name="gazefirst").start()
# The relay's select loop: readable while connected to vrserver.
def fileno(self):
return self.ws.fileno()
def message(self, words, now):
""""gazefirst 1|0" from the helper."""
on = words[1] == "1"
self.on_until = now + ON_FOR if on else 0.0
if on != self.on:
self.set_on(on)
def set_on(self, on):
self.on = on
if not on:
self.release_all(cancel=True)
self.log(f"gaze first {'on' if on else 'off'}")
self.wake.set()
def release_all(self, cancel=False):
"""Let go of what the helper holds: as releases (a click), or cancelled (no click)."""
if cancel and self.held:
self.send("ctrlcancel")
else:
for side, button in sorted(self.held):
self.send(f"ctrl {side} {button} 0")
self.held.clear()
if self.scrolling != (0.0, 0.0):
self.send("scroll 0 0")
self.scrolling = (0.0, 0.0)
def toggle(self):
"""Gaze mode on or off, remembered (POINTER_GAZE)."""
on = not read_gaze()
try:
write_gaze(on)
except OSError as e:
self.log(f"gaze mode: can't write {CONF}: {e}")
self.send(f"gaze {'on' if on else 'off'}")
self.log(f"gaze mode {'on' if on else 'off'}")
def timeout(self):
return 0.05 if self.macro_since is not None and not self.macro_fired else 0.5
def tick(self, now):
if self.on and now > self.on_until:
self.set_on(False) # the helper went quiet
if self.ws is None and now >= self.next_connect:
self.next_connect = now + 3.0
try:
self.ws = VrSocket(timeout=1)
self.ws.open(f"frametop_relay_{os.getpid()}")
self.sides = {}
self.wake.set() # find the controllers now
except OSError:
self.ws = None
with self.lock:
found, self.found = self.found, []
for path, side in found:
if self.ws and self.sides.get(path) != side:
try:
self.ws.subscribe(path)
self.sides[path] = side
except OSError:
self.drop()
if self.macro_since is not None and not self.macro_fired and now - self.macro_since >= MACRO_HOLD:
self.macro_fired = True
self.release_all(cancel=True) # a press in progress doesn't click
self.toggle()
def drop(self):
if self.ws:
try:
self.ws.close()
except OSError:
pass
self.ws = None
self.release_all(cancel=True)
def readable(self, now):
try:
while True:
msg = self.ws.recv(timeout=0)
if isinstance(msg, dict) and msg.get("type") == "update_component_states":
side = self.sides.get(msg.get("device"))
if side:
self.update(side, msg.get("components") or {}, now)
if not self.ws.pending():
return # the rest, if any, wakes select again
except OSError:
self.drop()
def update(self, side, components, now):
for name, value in components.items():
key = (side, name)
if self.values.get(key) == value:
continue
self.values[key] = value
button = BUTTONS.get(name)
if button:
down = bool(value)
if down and self.on and (side, button) not in self.held:
self.held.add((side, button))
self.send(f"ctrl {side} {button} 1")
elif not down and (side, button) in self.held:
self.held.discard((side, button))
self.send(f"ctrl {side} {button} 0")
elif name == STICK_CLICK:
both = self.values.get(("left", STICK_CLICK)) and self.values.get(("right", STICK_CLICK))
if both and self.macro_since is None:
self.macro_since, self.macro_fired = now, False
elif not both:
self.macro_since = None
elif side == "right" and name in ("/input/thumbstick/x", "/input/thumbstick/y"):
self.scroll()
def scroll(self):
def shape(v):
v = float(v or 0)
if abs(v) < SCROLL_DEADZONE:
return 0.0
return round((abs(v) - SCROLL_DEADZONE) / (1 - SCROLL_DEADZONE) * (1 if v > 0 else -1), 1)
want = (shape(self.values.get(("right", "/input/thumbstick/x"))),
shape(self.values.get(("right", "/input/thumbstick/y")))) if self.on else (0.0, 0.0)
if want != self.scrolling:
self.scrolling = want
self.send(f"scroll {want[0]:.1f} {want[1]:.1f}")
def shutdown(self):
"""The relay is stopping: the controllers go back to SteamVR and Steam."""
self.on = False
self.release_all(cancel=True)
try:
select_binding(STOCK_BINDING)
except OSError:
pass
try:
steamui.evaluate('window.__frametopGaze && (window.__frametopGaze.mode = "off")', timeout=1)
except (OSError, RuntimeError):
pass
# The worker thread.
def problem(self, what, error=None):
"""Log a failure once, and when it's over (error None)."""
if error is None:
if self.problems.pop(what, None) is not None:
self.log(f"gaze first: {what} works again")
return
if self.problems.get(what) != str(error):
self.problems[what] = str(error)
self.log(f"gaze first: {what}: {error}")
def work(self):
while True:
self.wake.wait(RENEW)
self.wake.clear()
on = self.on
try:
found = [(d["root_path"], d.get("side") or d["root_path"].rsplit("/", 1)[-1])
for d in getstate(timeout=1) if d.get("controller_type") == "frame_controller"]
with self.lock:
self.found = [(p, s) for p, s in found if s in ("left", "right")]
self.problem("vrserver")
except (OSError, ValueError) as e:
self.problem("vrserver", e)
continue
try:
want = GAZE_BINDING if on else STOCK_BINDING
if current_binding() != want:
select_binding(want)
self.log(f"gaze first: {'gaze' if on else 'stock'} controller binding")
self.problem("binding")
except (OSError, ValueError) as e:
self.problem("binding", e)
try:
if on:
with open(FILTER_JS) as f:
steamui.evaluate(f.read())
steamui.evaluate('(() => { const G = window.__frametopGaze; G.mode = "block"; '
f'G.until = Date.now() + {BLOCK_FOR}; return G.mode; }})()')
elif self.applied is not False:
steamui.evaluate('window.__frametopGaze && (window.__frametopGaze.mode = "off")')
self.applied = on
self.problem("Steam's UI")
except (OSError, RuntimeError, ValueError) as e:
self.problem("Steam's UI", e)
+143 -20
View File
@@ -19,13 +19,19 @@ keyboard node for its extra buttons). Roles, from ~/.config/frametop-input.json
ignore not grabbed, only observed for identification in the settings app
Buttons and keys of pointer devices go through a per-device map to actions
(left, right, middle, back, scroll_up, scroll_down, dashboard, recenter,
pointer_toggle, follow_toggle = head follow on or off, gaze_toggle = gaze mode on or off
(the pointer goes where you look; see pointer/helper/ft-pointer.cpp), sens_up, sens_down,
layout_reset = put the desktop screens back in their saved layout, screens_toggle = hide or show the desktop screens, key = pass
through as a key, none).
pointer_toggle, follow_toggle = head follow on or off, gaze_toggle = gaze mode on or off,
remembered as POINTER_GAZE (the pointer goes where you look; see pointer/helper/ft-pointer.cpp
and input/gazefirst.py), gaze_precision = while
held, the pointer stops where you look and the button's device (the mouse, or that
controller's aim) steers it, and the release clicks there, gaze_drag = the same, but pressed
at once, so it drags ("precision|gazedrag mouse|left|right|keyboard 1|0" to the helper), sens_up, sens_down,
layout_reset = put the desktop screens back in their saved layout, screens_toggle = hide or show the desktop screens,
keyboard_toggle = open or close Frametop's keyboard, key = pass through as a key, none).
Frame controller buttons can be mapped too ("controller_buttons": {"right/a": action} in the
rules file; any action but key). The controllers aren't input devices here, only SteamVR sees
rules file; any action but key). So can key combinations on any keyboard ("key_bindings":
{"29+56+34": action}, evdev codes joined by "+", modifiers first and left-hand codes for
either side, here Ctrl+Alt+G): the combination does the action, and its last key isn't typed. The controllers aren't input devices here, only SteamVR sees
them, so the pointer helper reads them with SteamVR input and sends "vrbtn <button> 1|0".
It only takes the buttons the relay tells it to ("vrbind <button>..." to @ft_pointer_helper,
sent on start, reload, and when the helper says "vrhello"), and only while no game runs,
@@ -34,6 +40,8 @@ reaches games; see pointer/helper/vrbuttons.h).
In gaze mode outside games, the helper keeps the pointer ("gazeawake 1", repeated every 5
seconds; "gazeawake 0" or silence ends it): the pointer isn't released when the mouse is idle.
Typing on a keyboard sends the helper "typing" (at most 4 times a second): it takes no hand
pinches right after a key, since typing touches thumb to index like a pinch.
Keys also go to ft-screens (@ft_screens, the Frametop desktop's compositor), which
types them into the desktop screen that has focus: from pass-through keyboards, and
@@ -47,6 +55,16 @@ a grabbed keyboard's keys also go out as "key <code> <value> <device name>" data
It's off by default: any local process that binds that name first gets every key typed
into the desktop.
Frametop's keyboard (ft-screens' key panel): the desktop's input method (input/ft-textinput)
says "textfield 1|0" when a text field on the desktop gains or loses keyboard focus, and
the relay tells ft-screens to open ("vrkeyboard show") or close ("vrkeyboard hide") the
keyboard, depending on "vr_keyboard" in the rules file: "always", "no_keyboard" (the
default: only while no pass-through keyboard is connected; a program's uinput keyboard
doesn't count), "button" (only the keyboard_toggle action opens it), or "never"
(keyboard_toggle does nothing either). With "vr_keyboard_persist" (the default), it stays
open when the text field loses focus, until its Close key, keyboard_toggle, or a layout reset
(ft-layout apply) closes it.
Volume keys, from every device that has them (the headset's own buttons included),
are handled here: wpctl steps the default output. Nothing else may see a volume key,
because gamescope aborts on one when no window has keyboard focus, which ends the
@@ -67,6 +85,8 @@ Control socket (abstract datagram @frametop_relay, JSON replies to the sender):
vrcapture <s> take every controller button for s seconds (0: stop), so the settings
app can capture one; watchers see them as events with id frame_controller
vrbtn, vrhello, gazeawake from the pointer helper (above)
gazefirst 1|0 from the pointer helper: gaze first on or off (input/gazefirst.py)
textfield 1|0 from the desktop's input method (above)
Runs on the Frame host as a user service (frametop-input-relay.service). The
virtual devices are parked in systemd's file descriptor store, so a relay
@@ -92,6 +112,8 @@ import subprocess
import sys
import time
import gazefirst
# Linux input constants (include/uapi/linux/input-event-codes.h, input.h, uinput.h).
EV_SYN, EV_KEY, EV_REL, EV_MSC = 0x00, 0x01, 0x02, 0x04
SYN_REPORT = 0
@@ -155,8 +177,11 @@ def eviocguniq(length):
VIRTUAL_PREFIX = "frametop virtual"
RULES_PATH = os.path.expanduser("~/.config/frametop-input.json")
ACTIONS = ("left", "right", "middle", "back", "scroll_up", "scroll_down", "dashboard", "recenter",
"pointer_toggle", "follow_toggle", "gaze_toggle", "sens_up", "sens_down", "layout_reset", "screens_toggle",
"key", "none")
"pointer_toggle", "follow_toggle", "gaze_toggle", "gaze_precision", "gaze_drag", "sens_up", "sens_down",
"layout_reset", "screens_toggle", "keyboard_toggle", "key", "none")
# Key combinations ("key_bindings"): modifiers, each side's code folded into the left one's.
MODIFIERS = {29: 29, 97: 29, 42: 42, 54: 42, 56: 56, 100: 56, 125: 125, 126: 125}
VR_KEYBOARD_MODES = ("always", "no_keyboard", "button", "never") # when Frametop's keyboard opens
HELPER = "\0ft_pointer_helper"
# Frame controller buttons the pointer helper can read (pointer/helper/vrbuttons.h).
VR_BUTTONS = ("left/view", "left/dpad_up", "left/dpad_down", "left/dpad_left", "left/dpad_right", "left/bumper",
@@ -260,7 +285,9 @@ def read_config(path=os.path.expanduser("~/.config/frametop.conf")):
def read_rules(path=RULES_PATH):
"""{"devices": {id: {"role", "name"}}, "buttons": {id: {"<code>": action}},
"controller_buttons": {"<hand>/<button>": action}}."""
"controller_buttons": {"<hand>/<button>": action}, "controller_in_games": bool,
"key_bindings": {"<code>+<code>...": action},
"vr_keyboard": one of VR_KEYBOARD_MODES, "vr_keyboard_persist": bool}."""
try:
with open(path) as f:
rules = json.load(f)
@@ -269,6 +296,7 @@ def read_rules(path=RULES_PATH):
rules.setdefault("devices", {})
rules.setdefault("buttons", {})
rules.setdefault("controller_buttons", {})
rules.setdefault("key_bindings", {})
return rules
@@ -404,10 +432,17 @@ class Pointer:
self.send(f"scroll 0 {1 if value > 0 else -1}")
self.scroll_until = now + self.SCROLL_PULSE
def action(self, name, value, now):
"""A mapped button: value 1 press, 0 release, 2 autorepeat (ignored)."""
def action(self, name, value, now, source="mouse"):
"""A mapped button: value 1 press, 0 release, 2 autorepeat (ignored). source: what
pressed it (mouse, left, right for a controller, keyboard), for the gaze actions."""
if value == 2:
return
if name in ("gaze_precision", "gaze_drag"):
if value == 1:
self.wake(now)
self.flush()
self.send(f"{'precision' if name == 'gaze_precision' else 'gazedrag'} {source} {value}")
return
driver = self.DRIVER_BUTTONS.get(name)
if driver:
self.wake(now)
@@ -434,9 +469,6 @@ class Pointer:
elif name == "follow_toggle":
self.send("follow toggle") # until the next restart; the setting is POINTER_FOLLOW
log("head follow toggled")
elif name == "gaze_toggle":
self.send("gaze toggle") # until the next restart; the setting is POINTER_GAZE
log("gaze mode toggled")
elif name == "screens_toggle":
try:
self.sock.sendto(b"toggle", SCREENS)
@@ -502,8 +534,9 @@ class Node:
or any other device with volume keys (candidate False, role "volume")."""
def __init__(self, path, fd, name, bus, vendor, product, uniq, is_mouse, is_keyboard,
candidate=True, volume_keys=False, only_volume=False):
candidate=True, volume_keys=False, only_volume=False, uinput=False):
self.path, self.fd, self.name = path, fd, name
self.uinput = uinput # made by a program (frame-voice's keyboard, say), not a real device
self.bus, self.vendor, self.product, self.uniq = bus, vendor, product, uniq
self.is_mouse, self.is_keyboard = is_mouse, is_keyboard
self.candidate, self.volume_keys, self.only_volume = candidate, volume_keys, only_volume
@@ -520,7 +553,7 @@ class Node:
kinds = [k for k, on in (("mouse", self.is_mouse), ("keyboard", self.is_keyboard)) if on]
return {"path": self.path, "name": self.name, "id": self.id, "uniq": self.uniq,
"bus": {BUS_USB: "usb", BUS_BLUETOOTH: "bluetooth"}.get(self.bus, str(self.bus)),
"kinds": kinds, "role": self.role, "grabbed": self.grabbed}
"kinds": kinds, "role": self.role, "grabbed": self.grabbed, "uinput": self.uinput}
def probe(path):
@@ -552,8 +585,10 @@ def probe(path):
volume_keys = bool(keys & VOLUME_CODES) # stand-ins too: kept from before a relay restart
if not (candidate or volume_keys):
raise ValueError
# uinput devices live here (Bluetooth LE ones come through uhid, under virtual/misc).
sysfs = os.path.realpath(f"/sys/class/input/{os.path.basename(path)}/device")
return Node(path, fd, name, bus, vendor, product, uniq, is_mouse, is_keyboard,
candidate, volume_keys, keys <= VOLUME_CODES)
candidate, volume_keys, keys <= VOLUME_CODES, sysfs.startswith("/sys/devices/virtual/input/"))
except (OSError, ValueError):
os.close(fd)
return None
@@ -598,6 +633,7 @@ def main():
load_config()
meta_down = False # Meta pressed with no other key yet: a tap toggles the dashboard
screens_sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM | socket.SOCK_NONBLOCK)
last_typing = 0.0 # the helper was last told of a key then (see "typing" at the top)
def vr_bind(now):
"""Tell the pointer helper which controller buttons to take from games."""
@@ -623,8 +659,68 @@ def main():
reply(addr, {"t": "event", "id": VR_DEVICE, "path": "", "name": "Steam Frame controllers",
"type": "vr", "code": button, "value": value})
action = state["rules"]["controller_buttons"].get(button)
if gaze.on and button.split("/")[-1] in ("trigger", "bumper"):
return # gaze first has them (input/gazefirst.py)
if state["pointer"] and action in ACTIONS and action not in ("key", "none"):
state["pointer"].action(action, value, now)
do_action(action, value, now, button.split("/")[0])
def vr_keyboard_mode():
mode = state["rules"].get("vr_keyboard")
return mode if mode in VR_KEYBOARD_MODES else "no_keyboard"
def vr_keyboard(command):
"""Open or close Frametop's keyboard (ft-screens)."""
try:
screens_sock.sendto(f"vrkeyboard {command}".encode(), SCREENS)
except OSError:
pass # ft-screens not running
def text_field(focused):
"""A text field on the desktop gained or lost keyboard focus (the input method)."""
mode = vr_keyboard_mode()
if not focused:
if not state["rules"].get("vr_keyboard_persist", True):
vr_keyboard("hide") # ft-screens closes it only if it opened it for a text field
elif mode == "always" or (mode == "no_keyboard" and not any(
n.candidate and n.is_keyboard and n.role == "passthrough" and not n.uinput for n in nodes.values())):
vr_keyboard("show")
def do_action(action, value, now, source="mouse"):
"""A mapped mouse or controller button, or key combination (pointer mode only)."""
if action == "gaze_toggle":
if value == 1:
gaze.toggle() # remembered (POINTER_GAZE), like the toggle macro
elif action == "keyboard_toggle":
if value == 1 and vr_keyboard_mode() != "never":
vr_keyboard("toggle")
else:
state["pointer"].action(action, value, now, source)
held_modifiers = set() # on any keyboard, folded (MODIFIERS)
combos_down = {} # key code -> the action its combination started (released with it)
def key_binding(code, value, now):
"""A key from a keyboard: does it complete a key combination ("key_bindings")? True if
it was taken for one (then it isn't typed)."""
if code in MODIFIERS:
(held_modifiers.add if value else held_modifiers.discard)(MODIFIERS[code])
return False
if value == 0 and code in combos_down:
action = combos_down.pop(code)
if state["pointer"]:
do_action(action, 0, now, "keyboard")
return True
if value != 1 or not state["rules"]["key_bindings"]:
return value == 2 and code in combos_down
combo = "+".join(str(c) for c in sorted(held_modifiers) + [code])
action = state["rules"]["key_bindings"].get(combo)
if action not in ACTIONS or action in ("key", "none"):
return False
combos_down[code] = action
if state["pointer"]:
do_action(action, 1, now, "keyboard")
log(f"key combination {combo}: {action}")
return True
def to_screens(code, value):
@@ -789,6 +885,12 @@ def main():
if cmd == "vrhello":
vr_bind(now)
continue
if cmd == "textfield" and len(words) == 2:
text_field(words[1] == "1")
continue
if cmd == "gazefirst" and len(words) == 2:
gaze.message(words, now)
continue
if cmd == "gazeawake" and len(words) == 2:
if state["pointer"]:
state["pointer"].gaze_awake_until = now + 12.0 if words[1] == "1" else 0.0
@@ -832,6 +934,17 @@ def main():
else:
reply(addr, msg)
def to_helper(command):
try:
screens_sock.sendto(command.encode(), HELPER)
except OSError:
pass # helper not running
# Gaze first (input/gazefirst.py): the controllers' trigger, bumper, thumbstick, and the
# toggle macro, from vrserver's web socket; SteamVR's and Steam's muting.
gaze = gazefirst.GazeFirst(to_helper, log)
atexit.register(gaze.shutdown)
vr_bind(time.monotonic()) # a helper that's already running keeps its buttons in step
waiting = False # a keyboard's grab waits for its keys to come up
while True:
@@ -866,11 +979,12 @@ def main():
if added:
apply_roles()
ready, _, _ = select.select(list(nodes) + [control], [], [],
volume.timeout(now, pointer.timeout() if pointer else 0.5))
ready, _, _ = select.select(list(nodes) + [control] + ([gaze] if gaze.ws else []), [], [],
min(gaze.timeout(), volume.timeout(now, pointer.timeout() if pointer else 0.5)))
now = time.monotonic()
if pointer:
pointer.tick(now)
gaze.tick(now)
volume.tick(now)
if state["vr_capture_until"] and now >= state["vr_capture_until"]:
vr_bind(now) # capture over: back to the mapped buttons
@@ -880,6 +994,10 @@ def main():
if fd is control:
handle_control(now)
continue
if fd is gaze:
if gaze.ws:
gaze.readable(now)
continue
node = nodes[fd]
try:
data = os.read(fd, EVENT.size * 64)
@@ -908,7 +1026,12 @@ def main():
if node.role != "pointer":
# Observed only, unless typing goes to the desktop. With META_DASHBOARD=1,
# a Meta tap on any keyboard toggles the dashboard.
if node.role == "passthrough" and etype == EV_KEY and code < BTN_MISC and key_binding(code, value, now):
continue
if node.role == "passthrough" and etype == EV_KEY:
if value == 1 and code < BTN_MISC and now - last_typing >= 0.25:
last_typing = now
screens_sock.sendto(b"typing", HELPER)
to_screens(code, value)
if node.grabbed and code < BTN_MISC and value in (0, 1):
share_key(node, code, value)
@@ -930,7 +1053,7 @@ def main():
if etype == EV_KEY:
action = buttons.get(str(code), DEFAULT_BUTTONS.get(code, "key"))
if pointer and action not in ("key", "none"):
pointer.action(action, value, now)
do_action(action, value, now)
continue
if action == "none":
continue
+106
View File
@@ -0,0 +1,106 @@
// Frametop gaze first: a filter in Steam's UI (SharedJSContext, through Steam's CEF debugger on
// 127.0.0.1:8080) that keeps the Frame controllers' gamepad input away from Steam's UI while
// gaze mode has the dashboard (docs/gaze-first.md). Steam reads the controllers as its own
// virtual gamepad (Steam Input), so no SteamVR binding can do this.
//
// Steam's gamepad input source (webpack module 17900, class E) turns each controller message
// into this.OnButtonDown / OnButtonUp / OnAnalogPad calls, which live on its base class. The
// filter defines wrappers on E's prototype, so both live instances go through them. They look
// up window.__frametopGaze at every call, so evaluating this file again updates the logic in
// place; a wrapper from an older version is replaced. It returns the state as JSON.
//
// The dashboard's mode (laser or gamepad) is SteamVR's, not the UI's: the UI only mirrors it
// (VRFocus, module 84114: J.Instance.ShowGamepadFocusMode, from SteamVR's
// system_panel_interaction_mode).
//
// window.__frametopGaze:
// mode "off" (pass everything), "log" (pass, and log), "block" (drop all but `allow`),
// "auto" (block while the dashboard is in laser mode, pass in gamepad mode)
// until with "block" or "auto": blocking stops at this time (Date.now(), ms) unless it's
// moved on. The input relay (input/gazefirst.py) sets it 15 s ahead every 10 s or so,
// so if the relay dies, the controllers come back to Steam's UI by themselves.
// allow buttons that always pass: the Steam button's guide and quick menu, and Steam's
// own dummy input
// log the last 256 events: [ms, "down"|"up"|"analog", button, controller, passed]
// modes the dashboard's mode changes: [ms, "gamepad"|"laser"] (sampled every 50 ms)
// A release passes if its press did, so nothing stays held when blocking starts.
// Buttons (Steam's enum): 1 OK 2 CANCEL 3 SECONDARY 4 OPTIONS 5/6 bumpers 7/8 triggers 9-12 dpad
// 13 SELECT 14 START 15/16 stick clicks 17/18 stick touch 23-26 rear buttons (24 left grip,
// 26 right grip) 27 STEAM_GUIDE 28 STEAM_QUICK_MENU 29 DUMMY_INPUT.
(() => {
const VERSION = 2;
if (!window.__ftreq) {
const name = Object.keys(window).find(k => k.startsWith("webpackChunk"));
window[name].push([[Symbol("frametop")], {}, r => { window.__ftreq = r; }]);
}
const E = window.__ftreq(17900).E;
const proto = E.prototype;
if (!("HandleControllerInputMessages" in proto)) throw new Error("Steam's gamepad source moved (module 17900)");
const base = Object.getPrototypeOf(proto);
const G = window.__frametopGaze || (window.__frametopGaze = {mode: "log", allow: [27, 28, 29], log: [], held: {}, dropped: 0});
G.modes = G.modes || [];
G.gamepadMode = () => {
try { return !!window.__ftreq(84114).J.Instance?.ShowGamepadFocusMode; } catch (e) { return false; }
};
G.note = (ev, button, controller, passed) => {
G.log.push([Math.round(performance.now()), ev, button, controller, passed]);
if (G.log.length > 256) G.log.shift();
};
G.blocks = button => {
if (G.allow.includes(button) || (G.until && Date.now() > G.until)) return false;
return G.mode === "block" || (G.mode === "auto" && !G.gamepadMode());
};
if (proto.OnButtonDown && proto.OnButtonDown.__frametop !== VERSION) {
delete proto.OnButtonDown;
delete proto.OnButtonUp;
delete proto.OnAnalogPad;
}
if (!Object.prototype.hasOwnProperty.call(proto, "OnButtonDown")) {
proto.OnButtonDown = function (button, controller, ...rest) {
let pass = true;
try {
const g = window.__frametopGaze;
g.inst = this;
pass = !g.blocks(button);
if (g.mode !== "off") g.note("down", button, controller, pass);
if (pass) g.held[controller + ":" + button] = true;
else g.dropped++;
} catch (e) { pass = true; }
if (pass) return base.OnButtonDown.call(this, button, controller, ...rest);
};
proto.OnButtonUp = function (button, controller, ...rest) {
let pass = true;
try {
const g = window.__frametopGaze, key = controller + ":" + button;
pass = !!g.held[key] || !g.blocks(button);
delete g.held[key];
if (g.mode !== "off") g.note("up", button, controller, pass);
} catch (e) { pass = true; }
if (pass) return base.OnButtonUp.call(this, button, controller, ...rest);
};
proto.OnAnalogPad = function (button, x, y, controller, ...rest) {
let pass = true;
try {
const g = window.__frametopGaze;
pass = !g.blocks(button);
if (g.mode !== "off" && (x || y)) g.note("analog", button, controller, pass);
} catch (e) { pass = true; }
if (pass) return base.OnAnalogPad.call(this, button, x, y, controller, ...rest);
};
for (const f of [proto.OnButtonDown, proto.OnButtonUp, proto.OnAnalogPad]) f.__frametop = VERSION;
G.installed = Date.now();
}
if (!G.sampler) {
let last = null;
G.sampler = setInterval(() => {
const m = G.gamepadMode() ? "gamepad" : "laser";
if (m !== last) {
last = m;
G.modes.push([Math.round(performance.now()), m]);
if (G.modes.length > 128) G.modes.shift();
}
}, 50);
}
return JSON.stringify({version: VERSION, mode: G.mode, gamepad: G.gamepadMode(), installed: G.installed,
dropped: G.dropped, modes: G.modes.slice(-10), log: G.log.slice(-10)});
})()
+40
View File
@@ -0,0 +1,40 @@
"""steamui: run JavaScript in Steam's UI (its SharedJSContext) through Steam's CEF debugger on
127.0.0.1:8080, with nothing but the standard library (input/vrws.py's web socket client).
evaluate(expression) the expression's value (JSON-able), or an OSError when Steam or its
UI isn't there, or a RuntimeError when the script threw
"""
import json
import time
import urllib.request
from vrws import VrSocket
DEBUGGER = "http://127.0.0.1:8080"
def evaluate(expression, target="SharedJSContext", timeout=2.0):
with urllib.request.urlopen(f"{DEBUGGER}/json", timeout=timeout) as resp:
pages = json.load(resp)
url = next((p.get("webSocketDebuggerUrl") for p in pages if p.get("title") == target), None)
if not url:
raise OSError(f"Steam's {target} isn't open")
ws = VrSocket(timeout=timeout, url=url)
try:
ws.send(json.dumps({"id": 1, "method": "Runtime.evaluate",
"params": {"expression": expression, "returnByValue": True}}))
deadline = time.monotonic() + timeout
while True:
left = deadline - time.monotonic()
msg = ws.recv(timeout=max(left, 0)) if left > 0 else None
if msg is None:
raise OSError("Steam's UI didn't answer")
if isinstance(msg, dict) and msg.get("id") == 1:
break
finally:
ws.close()
result = msg.get("result") or {}
if "exceptionDetails" in result:
details = result["exceptionDetails"]
raise RuntimeError((details.get("exception") or {}).get("description") or details.get("text", "error"))
return (result.get("result") or {}).get("value")
Executable
+220
View File
@@ -0,0 +1,220 @@
#!/usr/bin/python3
"""vrws: vrserver's local web socket, with nothing but the standard library.
vrserver (SteamVR) serves its controller binding page on 127.0.0.1:27062, and that page
follows the controllers' raw input through a web socket there. Reading it takes nothing
from anyone: whatever the buttons are bound to still happens, and it works whatever app
has focus, in games too. frame-voice reads its push-to-talk button the same way. The host's
Python has no `websockets` module, so this is a small RFC 6455 client of its own.
getstate() the devices SteamVR has now, from /input/getstate.json: each with its
root path (/user/hand/left, /devices/cv/<serial> while our pointer holds
that hand, /user/head), controller type, side, and input components
VrSocket() a connection: open(mailbox), subscribe(device), recv() -> dict or None
Run as a program, it prints component changes as they come, for the gaze-first tests
(docs/gaze-first.md, test 4):
input/vrws.py [--all] [--seconds N] [component ...]
By default it follows the buttons (…/click), the thumbsticks' axes, the trigger's value,
and the headset's proximity; --all shows every component. At the end it prints how often
each one updated.
"""
import base64
import json
import os
import select
import socket
import struct
import sys
import time
import urllib.request
HOST, PORT = "127.0.0.1", 27062
ORIGIN = f"http://{HOST}:{PORT}"
HEADERS = {"Referer": f"{ORIGIN}/dashboard/controllerbinding.html"}
def getstate(timeout=3):
"""The devices SteamVR has now (only the ones with a root path)."""
req = urllib.request.Request(f"{ORIGIN}/input/getstate.json", headers=HEADERS)
with urllib.request.urlopen(req, timeout=timeout) as resp:
return [d for d in json.load(resp).get("devices", []) if d.get("root_path")]
class VrSocket:
"""vrserver's web socket, or with `url` (ws://host:port/path) any other local one, such as
Steam's CEF debugger (input/steamui.py)."""
def __init__(self, timeout=3, url=None):
if url:
hostport, path = url.removeprefix("ws://").split("/", 1)
host, port = hostport.rsplit(":", 1)
extra = ""
else:
host, port, path, hostport = HOST, PORT, "", f"{HOST}:{PORT}"
extra = f"Origin: {ORIGIN}\r\nReferer: {HEADERS['Referer']}\r\n"
self.sock = socket.create_connection((host, int(port)), timeout=timeout)
key = base64.b64encode(os.urandom(16)).decode()
self.sock.sendall((f"GET /{path} HTTP/1.1\r\nHost: {hostport}\r\nUpgrade: websocket\r\n"
f"Connection: Upgrade\r\nSec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n"
f"{extra}\r\n").encode())
head = b""
while b"\r\n\r\n" not in head:
chunk = self.sock.recv(4096)
if not chunk:
raise OSError("vrserver closed the connection during the handshake")
head += chunk
head, self.buf = head.split(b"\r\n\r\n", 1)
status = head.split(b"\r\n", 1)[0]
if b" 101 " not in status + b" ":
raise OSError(f"vrserver refused the web socket: {status.decode(errors='replace')}")
self.sock.settimeout(None)
self.parts = b""
def fileno(self):
return self.sock.fileno()
def close(self):
try:
self._frame(0x8, b"")
except OSError:
pass
self.sock.close()
def _frame(self, opcode, payload):
n = len(payload)
head = bytes([0x80 | opcode])
if n < 126:
head += bytes([0x80 | n])
elif n < 65536:
head += bytes([0x80 | 126]) + struct.pack(">H", n)
else:
head += bytes([0x80 | 127]) + struct.pack(">Q", n)
mask = os.urandom(4)
self.sock.sendall(head + mask + bytes(b ^ mask[i % 4] for i, b in enumerate(payload)))
def send(self, text):
self._frame(0x1, text.encode())
def open(self, mailbox):
self.send(f"mailbox_open {mailbox}")
self.mailbox = mailbox
def subscribe(self, device, on=True):
kind = "request_input_state_updates" if on else "cancel_input_state_updates"
self.send("mailbox_send input_server " +
json.dumps({"type": kind, "device_path": device, "returnAddress": self.mailbox}))
def _read(self, n):
while len(self.buf) < n:
chunk = self.sock.recv(65536)
if not chunk:
raise OSError("vrserver closed the web socket")
self.buf += chunk
out, self.buf = self.buf[:n], self.buf[n:]
return out
def pending(self):
"""A message (or part of one) is already read and waiting."""
return bool(self.buf)
def recv(self, timeout=None):
"""The next text message as parsed JSON (None for one that isn't JSON), or None at
the timeout. Raises OSError when the connection ends."""
if not self.buf and timeout is not None:
if not select.select([self.sock], [], [], timeout)[0]:
return None
while True:
b0, b1 = self._read(2)
opcode, n = b0 & 0x0F, b1 & 0x7F
if n == 126:
n = struct.unpack(">H", self._read(2))[0]
elif n == 127:
n = struct.unpack(">Q", self._read(8))[0]
mask = self._read(4) if b1 & 0x80 else None
data = self._read(n)
if mask:
data = bytes(b ^ mask[i % 4] for i, b in enumerate(data))
if opcode == 0x8:
raise OSError("vrserver closed the web socket")
if opcode == 0x9:
self._frame(0xA, data)
continue
if opcode in (0x1, 0x2, 0x0):
self.parts += data
if not b0 & 0x80:
continue
text, self.parts = self.parts, b""
try:
return json.loads(text)
except ValueError:
return None
def main():
args = [a for a in sys.argv[1:]]
show_all = "--all" in args
seconds = 0.0
if "--seconds" in args:
i = args.index("--seconds")
seconds = float(args[i + 1])
del args[i:i + 2]
wanted = [a for a in args if not a.startswith("--")]
def followed(name):
if show_all:
return True
if wanted:
return name in wanted
return (name.endswith("/click") or name in ("/input/thumbstick/x", "/input/thumbstick/y",
"/input/trigger/value", "/proximity"))
names = {}
ws = VrSocket()
ws.open(f"frametop_vrws_{os.getpid()}")
def follow_new():
"""Subscribe to devices that appeared (vrserver only streams changes, and a device
that connects later, or changes its path with a role, is a new subscription)."""
for d in getstate():
path = d["root_path"]
name = f"{path} ({d.get('controller_type', '?')}{', ' + d['side'] if d.get('side') else ''})"
if names.get(path) != name:
names[path] = name
print(f"{time.monotonic() - start:9.3f} device {name}", flush=True)
ws.subscribe(path)
start = time.monotonic()
follow_new()
next_poll = start + 2
last, counts = {}, {}
try:
while not seconds or time.monotonic() - start < seconds:
if time.monotonic() >= next_poll:
next_poll = time.monotonic() + 2
follow_new()
msg = ws.recv(timeout=0.5)
if not isinstance(msg, dict) or msg.get("type") != "update_component_states":
continue
dev = msg.get("device")
for name, value in (msg.get("components") or {}).items():
if not followed(name):
continue
key = (dev, name)
counts[key] = counts.get(key, 0) + 1
if last.get(key) != value:
last[key] = value
print(f"{time.monotonic() - start:9.3f} {names.get(dev, dev)} {name} = {value}", flush=True)
except KeyboardInterrupt:
pass
finally:
ws.close()
took = max(time.monotonic() - start, 1e-6)
for (dev, name), n in sorted(counts.items()):
print(f"updates: {names.get(dev, dev)} {name}: {n} ({n / took:.1f}/s)")
if __name__ == "__main__":
main()
+25 -12
View File
@@ -1,8 +1,8 @@
#!/usr/bin/env bash
# Install everything on the Steam Frame: the build container, Frametop (multi-screen
# desktop, input relay, universal 3D mouse, settings app), and optionally the Bluetooth
# fixes. Run it on the headset in a terminal, from this repo. It's safe to re-run, for
# example after `git pull`.
# fixes and hand tracking. Run it on the headset in a terminal, from this repo. It's safe
# to re-run, for example after `git pull`.
# (It also works from a PC over SSH; see "Developing from a PC" in the README.)
#
# Usage: ./install.sh [--yes] [--no-bluetooth]
@@ -45,7 +45,7 @@ else
"$root/scripts/sync.sh" >/dev/null
fi
step "1/7 distrobox (container tool, installed in your home folder)"
step "1/9 distrobox (container tool, installed in your home folder)"
if on_frame 'test -x ~/.local/bin/distrobox'; then
echo "already installed: $(on_frame '~/.local/bin/distrobox version | head -1')"
else
@@ -55,44 +55,57 @@ else
cd ~/dev/src/distrobox && ./install --prefix ~/.local'
fi
step "2/7 build container (Fedora 44 'dev', about 1-2 GB the first time)"
step "2/9 build container (Fedora 44 'dev', about 1-2 GB the first time)"
"$root/setup/dev-container.sh"
step "3/7 input relay (keeps Bluetooth mice working in SteamVR, device roles, button maps)"
step "3/9 input relay (keeps Bluetooth mice working in SteamVR, device roles, button maps)"
"$root/desktops.sh" relay install
step "4/7 3D mouse: SteamVR driver"
step "4/9 3D mouse: SteamVR driver"
"$root/pointer/driver/build.sh"
"$root/pointer/driver/install.sh" install 2>&1 | grep -v xdg-open
step "5/7 3D mouse: pointer helper service"
step "5/9 3D mouse: pointer helper service"
"$root/pointer/helper/build.sh"
"$root/pointer/helper/run.sh" install
step "6/7 multi-screen desktop (ft-screens), Frametop Input Settings, and Frametop Display Settings"
step "6/9 power service (turns the displays off while the headset isn't used, even on a stand)"
"$root/power/build.sh"
"$root/power/run.sh" install
step "7/9 multi-screen desktop (ft-screens), Frametop Input Settings, and Frametop Display Settings"
"$root/screens/build.sh"
"$root/desktops.sh" install >/dev/null
"$root/input-settings/install.sh"
"$root/display-settings/install.sh"
"$root/remote/install.sh"
on_frame "sed -i 's/^POINTER=0/POINTER=1/' ~/.config/frametop.conf; grep -q '^POINTER=' ~/.config/frametop.conf || echo 'POINTER=1' >> ~/.config/frametop.conf"
echo "the launcher's Desktop entry now opens the multi-screen desktop; 3D mouse on (POINTER=1 in ~/.config/frametop.conf)"
step "7/7 Bluetooth fixes (optional; they let LE mice and keyboards like the Swiftpoint Z3 reconnect)"
step "8/9 Bluetooth fixes (optional; they let LE mice and keyboards like the Swiftpoint Z3 reconnect)"
if [ "$bluetooth" = 1 ] && ask "Install the Bluetooth fixes? They need your password (sudo)." n; then
"$root/setup/bluetooth/install.sh" install
else
echo "skipped. Install later with: setup/bluetooth/install.sh install"
fi
step "9/9 hand tracking (optional, experimental: your hands show over the screens)"
if ask "Install hand tracking? It needs your password (sudo) to let its camera service read the headset cameras." n; then
"$root/hands/run.sh" install
else
echo "skipped. Install later with: hands/run.sh install"
fi
step "Done"
cat <<'EOF'
SteamVR has to restart once, to load the 3D mouse driver and to start the input relay
before it. Restarting SteamVR closes everything open in VR, including this terminal if
it's in a VR desktop. Rebooting the headset works too.
Recommended: stop Steam from putting the headset to sleep while it's plugged in. In Steam,
open Settings > Power, and under "When Plugged In and Idle" set "Sleep after" to Never.
The displays still turn off when you take the headset off.
Recommended: in Frametop Display Settings > Power, choose when the displays turn off
while the headset isn't used (for a stand or mount that covers its proximity sensor), and
turn on Stay awake while plugged in, so Steam doesn't put the headset to sleep while it
charges. The displays still turn off when you take the headset off.
EOF
if ask "Restart SteamVR now?" n; then
on_frame 'systemctl --user restart steamvr.service'
Loaded 100 of 139 files, more files were not shown because too many files have changed in this diff. Show more