Hand tracking for Frametop's hand cutouts

fh-camd (camd/) borrows XRService's camera buffers and publishes the tracking
cameras' frames to a shared ring. fh-tracker (trackd/) finds hands in them with
MediaPipe's palm and landmark models on ncnn, triangulates them in 3D, and
publishes them for ft-screens. fh-replay replays recordings offline. tracker/ is
the earlier Python version; tools/ and probes/ hold the checks and experiments.

As of this commit: crop contrast defaults to CLAHE for the palm search and plain
crops for the landmarks, --swap-sides works around fh-camd naming the side
cameras backwards after some XRService restarts (tools/check_sides.py detects
it), and --record-only, --with-dark, --cpus and --keep-presence support the
bright-light and CPU-placement tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
DeeJanuzandClaude Opus 5.5 committed 2026-09-29 22:36:48 -06:00
commit 067a03ce38
42 files changed
+6321

No files matched your search

+22
View File
@@ -0,0 +1,22 @@
# Recordings (tens of GB) and the Python venv
captures/
.venv/
__pycache__/
# Upstream clones: ncnn (see trackd/README.md) and FrameEyeCameraFeed (MIT, adapted into camd/)
vendor/
# Disassembly of Valve's vrclient.so, for reverse engineering only
re/*.dis
# Downloadable model sources (tools/convert_models.py); the converted ncnn models are kept
models/onnx/
models/*.task
# Build output
camd/fh-camd
camd/fh-camprobe
trackd/fh-tracker
trackd/fh-replay
trackd/nettest
probes/fh-frametime
probes/mgrvt
probes/ptstate
probes/refprobe
probes/*.bin
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Curtis English
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+15
View File
@@ -0,0 +1,15 @@
CFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter
LDLIBS = -lm
all: fh-camprobe fh-camd
fh-camprobe: camprobe.c tp.c xrcams.c tp.h xrcams.h
$(CC) $(CFLAGS) -o $@ camprobe.c tp.c xrcams.c $(LDLIBS)
fh-camd: camd.c tp.c xrcams.c tp.h xrcams.h fhring.h
$(CC) $(CFLAGS) -o $@ camd.c tp.c xrcams.c $(LDLIBS)
clean:
rm -f fh-camprobe fh-camd
.PHONY: all clean
+56
View File
@@ -0,0 +1,56 @@
# camd
Root-side camera access for frame-hands:
- `fh-camd`: the frame broker. It publishes the four IR tracking cameras to a shared-memory ring that the unprivileged tracker reads.
- `fh-camprobe`: a test tool that records timestamped frames from every camera, including the Arcturus color pair, with CSV logs.
## How it gets frames
XRService owns the headset cameras. Both tools borrow its DMA-BUFs read-only with `pidfd_getfd`, the same way FrameEyeCameraFeed does. They never touch XRService's V4L2 descriptors.
Polling buffers for changes can catch a frame while the camera is still writing it. Instead, they listen to the `v4l2:v4l2_dqbuf` tracepoint, which fires when XRService takes a buffer. It gives the buffer index, the sequence number and the capture timestamp.
The tools learn which DMA-BUF holds each V4L2 index by watching which buffer changes at each dequeue.
- Right after XRService allocates its buffers, the mapping is allocation order.
- After XRService restarts streaming, the order is shuffled, and the mapping is learned index by index.
- The two upper cameras share one run of buffers. For them, only allocation order can tell the cameras apart.
- `fh-camd` also re-maps an index on the fly when its buffer holds no new frame.
## fh-camd
```
make
sudo ./fh-camd # runs until stopped or XRService exits
```
It needs root only to set up: to borrow the buffers (`ptrace_scope=1` blocks `pidfd_getfd`) and to open the root-only tracepoints. Then it drops to the invoking user for good; XRService itself runs as that user. It reads nothing from the ring's readers.
Frames go to `/run/frame-hands/ir-ring`. The file is mode 0600 and owned by the user. It sits in a root-owned directory, so no other account can plant a file or link there. The layout is in `fhring.h`, and `tracker/ring.py` reads it.
- Only complete, bright frames are published. The cameras alternate a normal exposure with a near-black one, so each camera gets 30 of its 60 fps.
- A copy torn by the camera overwriting the buffer is dropped.
- Each copy takes about 0.1 ms, and a cache sync about 0.15 ms.
It exits when XRService exits, or when a camera's buffers keep going stale, which means XRService has reallocated them. Start it again to re-attach.
## fh-camprobe
Wear the headset (cameras only stream while it's worn), then run:
```
sudo ./fh-camprobe # records 15 s
sudo ./fh-camprobe --list # discovery and tracepoints only
```
Output goes to `~/Pictures/framecap/stereo-<time>/`:
- `summary.txt`: delays, buffer mapping, exposure pattern, stereo sync, and torn or stale copies.
- `frames.csv`: one row per frame. `events.csv`: every tracepoint sample.
- `<model>_NNNN_<camera>.pgm`: saved bright pairs. The color cameras are saved raw as `.yuv420_10p` (`tools/decode.py` reads them).
- `plane1_*.bin`: raw plane 1 of a few frames, which may hold sensor metadata.
Which camera is which: video9 is `slam_left`, video13 is `slam_right`, video6 is `upper_left` and video7 is `upper_right`. This was checked by rendering the same view from each camera with the factory calibration. `tracker/live.py` maps them by capture pipe (`/sys/class/video4linux/videoN/name`).
`discovery` in `xrcams.c` is adapted from FrameEyeCameraFeed (MIT, see `LICENSE.FrameEyeCameraFeed`).
+926
View File
@@ -0,0 +1,926 @@
/*
* fh-camd - publish the headset's IR camera frames to unprivileged trackers.
*
* Start it with sudo. As root it:
* 1. finds the mono tracking cameras and XRService's buffer queues (xrcams.c),
* 2. borrows those buffers read-only with pidfd_getfd,
* 3. opens the v4l2_dqbuf tracepoint (tp.c),
* 4. creates the frame ring in /run/frame-hands (fhring.h), owned by the user.
* Then it drops to that user for good. From then on it only learns which
* buffer holds which V4L2 index (as fh-camprobe does), and copies each
* complete bright frame into the ring. It exits when XRService exits or
* reallocates its buffers; start it again (or let systemd) to re-attach.
*
* XRService itself runs as the same user; root is needed only because
* ptrace_scope=1 blocks pidfd_getfd and the tracepoints are root-only.
* Nothing is read from the ring's readers.
*
* Build: make
*/
#define _GNU_SOURCE
#include "fhring.h"
#include "tp.h"
#include "xrcams.h"
#include <errno.h>
#include <fcntl.h>
#include <grp.h>
#include <linux/dma-buf.h>
#include <math.h>
#include <poll.h>
#include <pwd.h>
#include <signal.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/ioctl.h>
#include <sys/mman.h>
#include <sys/prctl.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <time.h>
#include <unistd.h>
#ifndef SYS_pidfd_open
#define SYS_pidfd_open 434
#endif
#ifndef SYS_pidfd_getfd
#define SYS_pidfd_getfd 438
#endif
#define MAX_CAMS FH_RING_MAX_CAMS
#define MAX_SLOTS 64
#define MAX_INDEX 32
#define NSAMP 512
#define STALE_RELEARN 30 /* consecutive unchanged frames: the mapping changed */
#define MAX_RELEARNS 5 /* then assume XRService has new buffers, and exit */
#define RING_DIR FH_RING_DIR
#define RING_FILE FH_RING_PATH
typedef struct {
xr_camera_t *cam;
xr_layout_t lay;
size_t need; /* bytes of plane 0 holding the image */
char slug[32];
/* candidate buffers, in XRService allocation order */
int nslots;
int slot_group[MAX_SLOTS];
int slot_pos[MAX_SLOTS];
int fd[MAX_SLOTS]; /* kept for DMA_BUF_IOCTL_SYNC */
const uint8_t *map[MAX_SLOTS];
uint64_t samp[MAX_SLOTS][NSAMP];
/* index -> buffer learning */
int votes[MAX_INDEX][MAX_SLOTS];
int nobs[MAX_INDEX];
int maxindex;
int learn_events;
bool mapped;
int slot_of[MAX_INDEX]; /* buffer holding each V4L2 index */
int relearns;
bool shared; /* its buffers sit among another camera's */
double recent[8]; /* means of the last frames, for dark detection */
int nrecent;
int stale_run;
uint64_t hist[2][NSAMP]; /* samples of its last two frames (bright and dark) */
fh_ring_cam_t *rc;
uint8_t *ring_slots;
uint64_t frame_no;
fh_ring_cam_t *rc_dark; /* --with-dark: its near-black frames */
uint8_t *ring_dark;
uint64_t dark_no;
uint64_t events, bright, dark, stale, torn, repairs;
uint64_t sync_ns, copy_ns, nsync; /* cost of cache sync and copy */
uint64_t last_bright; /* for the status line */
} cam_t;
static cam_t cams[MAX_CAMS];
static uint64_t filled[XR_MAX_GROUPS][XR_MAX_RUNBUFS]; /* when each buffer last held a frame */
static uint64_t nfilled;
static int ncams;
static xr_state_t xr;
static tp_event_t ev_dqbuf;
static int f_dq_minor, f_dq_index, f_dq_ts, f_dq_seq;
static const char *opt_sensor = "";
static const char *opt_user = NULL;
static double opt_dark = 0.4;
static bool opt_with_dark; /* also publish the near-black frames */
static double opt_status = 10.0;
static volatile sig_atomic_t stop;
/* ------------------------------------------------------------- utilities */
static void die(const char *fmt, ...)
{
va_list ap;
va_start(ap, fmt);
vfprintf(stderr, fmt, ap);
va_end(ap);
fputc('\n', stderr);
exit(1);
}
static uint64_t mono_ns(void)
{
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return (uint64_t)ts.tv_sec * 1000000000ull + (uint64_t)ts.tv_nsec;
}
static void sample_words(const uint8_t *p, size_t len, uint64_t *out)
{
size_t step = (len / NSAMP) & ~(size_t)7;
if (!step)
step = 8;
for (int i = 0; i < NSAMP; i++) {
size_t off = (size_t)i * step;
if (off + 8 > len)
off = len - 8;
memcpy(&out[i], p + off, 8);
}
}
/* Mean of an 8-bit image on a sparse grid. */
static double luma_mean(const cam_t *c, const uint8_t *p)
{
const xr_layout_t *l = &c->lay;
uint64_t sum = 0, n = 0;
for (unsigned y = 0; y < l->height; y += 8)
for (unsigned x = 0; x < l->width; x += 8, n++)
sum += p[(size_t)y * l->pitch + x];
return n ? (double)sum / (double)n : 0.0;
}
/* Make the CPU's view of a camera buffer current; harmless if it already is. */
static void buf_sync(int fd, uint64_t flags)
{
struct dma_buf_sync s = { .flags = flags | DMA_BUF_SYNC_READ };
if (fd >= 0)
ioctl(fd, DMA_BUF_IOCTL_SYNC, &s);
}
static cam_t *cam_by_minor(int64_t minor)
{
for (int i = 0; i < ncams; i++)
if ((int64_t)cams[i].cam->minor == minor)
return &cams[i];
return NULL;
}
static void model_of(const char *sensor, char *out, size_t n)
{
char buf[64];
if (sscanf(sensor, "%63s", buf) != 1)
buf[0] = 0;
snprintf(out, n, "%s", buf);
}
/* ---------------------------------------------------------------- setup */
static bool group_fits(const xr_group_t *g, const xr_camera_t *cam, size_t need)
{
return g->planesize[1] == cam->planesize[1] && g->planesize[0] >= need;
}
/*
* Candidate buffers for a camera: the queue bound to it. The two upper cameras
* share one 32-buffer run, bound to only one of them, so a camera without its
* own run also gets runs bound to cameras of the same model.
*/
static void setup_camera(cam_t *c, xr_camera_t *cam, int pidfd)
{
memset(c, 0, sizeof(*c));
c->cam = cam;
xr_camera_layout(cam, &c->lay);
c->need = (size_t)c->lay.pitch * c->lay.rows;
char sensor[XR_SENSOR_LEN];
xr_slugify(cam->sensor, sensor, sizeof(sensor));
snprintf(c->slug, sizeof(c->slug), "%.20s_video%d", sensor, cam->node);
bool own = false;
for (int g = 0; g < xr.ngroups; g++)
if (xr.groups[g].cam == cam && group_fits(&xr.groups[g], cam, c->need))
own = true;
char model[16];
model_of(cam->sensor, model, sizeof(model));
for (int g = 0; g < xr.ngroups; g++) {
xr_group_t *grp = &xr.groups[g];
if (!group_fits(grp, cam, c->need))
continue;
if (own && grp->cam != cam)
continue;
if (!own) {
char gm[16];
model_of(grp->cam ? grp->cam->sensor : grp->sensor, gm, sizeof(gm));
if (strcmp(gm, model))
continue;
}
for (int b = 0; b < grp->nbufs && c->nslots < MAX_SLOTS; b++) {
int fd = (int)syscall(SYS_pidfd_getfd, pidfd, grp->buf[b].xfd, 0u);
if (fd < 0) {
fprintf(stderr, "pidfd_getfd(%d): %s\n", grp->buf[b].xfd, strerror(errno));
continue;
}
void *m = mmap(NULL, grp->buf[b].size, PROT_READ, MAP_SHARED, fd, 0);
if (m == MAP_FAILED) {
fprintf(stderr, "mmap(xfd %d): %s\n", grp->buf[b].xfd, strerror(errno));
close(fd);
continue;
}
int s = c->nslots++;
c->slot_group[s] = g;
c->slot_pos[s] = b;
c->fd[s] = fd;
c->map[s] = m;
sample_words(c->map[s], c->need, c->samp[s]);
}
}
printf(" %-24s %-12s %ux%u pitch %u, %d candidate buffers%s\n", c->slug, cam->path,
c->lay.width, c->lay.height, c->lay.pitch, c->nslots, own ? "" : " (shared run)");
}
/* ------------------------------------------------------ index -> buffer */
/*
* Every buffer that changed (about as much as the most-changed one) since this
* camera's previous frame gets a vote for the V4L2 index just completed.
*
* XRService queues each index with the same buffer every time, but which
* buffer depends on how it (re)started streaming. Right after allocation it is
* allocation order, so a camera's buffers are a contiguous block (slot = base +
* index, bases at multiples of the queue depth); that is the only way to tell
* apart two cameras sharing one run, since their frames land at the same time.
* After a restart the order is shuffled; a camera with its own run is then
* mapped index by index.
*/
static bool resolve_block(cam_t *c)
{
int depth = c->maxindex + 1;
int best = -1;
double s1 = -1, s2 = -1;
for (int b = 0; b < c->nslots; b++) {
if (c->slot_pos[b] % depth)
continue;
if (b + depth > c->nslots || c->slot_group[b + depth - 1] != c->slot_group[b])
continue;
int hit = 0, obs = 0;
for (int i = 0; i < depth; i++) {
hit += c->votes[i][b + i] < c->nobs[i] ? c->votes[i][b + i] : c->nobs[i];
obs += c->nobs[i];
}
double score = obs ? (double)hit / obs : 0;
if (score > s1) {
s2 = s1; s1 = score; best = b;
} else if (score > s2) {
s2 = score;
}
}
if (s2 < 0)
s2 = 0;
if (best < 0 || s1 < 0.8 || s1 - s2 < 0.3)
return false;
for (int i = 0; i < depth; i++)
c->slot_of[i] = best + i;
printf("%s: queue depth %d, buffers %d-%d in order (score %.2f vs %.2f)\n", c->slug, depth, best,
best + depth - 1, s1, s2);
return true;
}
static bool resolve_each(cam_t *c)
{
int depth = c->maxindex + 1;
bool taken[MAX_SLOTS] = { false };
for (int i = 0; i < depth; i++) {
int best = -1, v1 = 0, v2 = 0;
for (int s = 0; s < c->nslots; s++) {
if (c->votes[i][s] > v1) {
v2 = v1; v1 = c->votes[i][s]; best = s;
} else if (c->votes[i][s] > v2) {
v2 = c->votes[i][s];
}
}
if (c->nobs[i] < 3 || best < 0 || taken[best] || v1 < 0.8 * c->nobs[i] || v2 > 0.3 * c->nobs[i])
return false;
taken[best] = true;
c->slot_of[i] = best;
}
printf("%s: queue depth %d, buffers mapped one by one:", c->slug, depth);
for (int i = 0; i < depth; i++)
printf(" %d", c->slot_of[i]);
printf("\n");
return true;
}
static void reset_learning(cam_t *c)
{
memset(c->votes, 0, sizeof(c->votes));
memset(c->nobs, 0, sizeof(c->nobs));
c->maxindex = 0;
c->learn_events = 0;
c->mapped = false;
}
static void resolve_map(cam_t *c)
{
int depth = c->maxindex + 1;
bool shared = c->slot_group[0] != c->slot_group[c->nslots - 1] || c->nslots > depth;
/* a shuffled shared run can't be split by change votes alone */
c->shared = shared;
c->mapped = resolve_block(c) || (!shared && resolve_each(c));
if (c->mapped) {
fflush(stdout);
} else if (c->learn_events > 20 * depth) {
fprintf(stderr, "%s: can't tell which buffers are this camera's yet; still learning\n", c->slug);
reset_learning(c);
}
}
static void learn(cam_t *c, int index)
{
static uint64_t cur[NSAMP];
int changed[MAX_SLOTS], maxc = 0;
for (int s = 0; s < c->nslots; s++) {
sample_words(c->map[s], c->need, cur);
changed[s] = 0;
for (int i = 0; i < NSAMP; i++)
changed[s] += cur[i] != c->samp[s][i];
memcpy(c->samp[s], cur, sizeof(cur));
if (changed[s] > maxc)
maxc = changed[s];
}
c->nobs[index]++;
c->learn_events++;
if (index > c->maxindex)
c->maxindex = index;
if (maxc >= NSAMP / 50)
for (int s = 0; s < c->nslots; s++)
if (changed[s] * 2 >= maxc)
c->votes[index][s]++;
if (c->learn_events >= 24 && c->learn_events >= 3 * (c->maxindex + 1))
resolve_map(c);
}
/*
* A frame is complete in a buffer when its samples changed since we last saw
* that buffer, down to the bottom rows (a buffer still being written has an
* unchanged tail).
*/
static bool fresh(const uint64_t *cur, const uint64_t *old, int *nchanged)
{
int n = 0, tail = 0;
for (int i = 0; i < NSAMP; i++) {
bool ch = cur[i] != old[i];
n += ch;
tail += ch && i >= NSAMP - NSAMP / 8;
}
*nchanged = n;
return tail > 0;
}
/* How different two sets of samples are: mean absolute difference per byte. */
static double sample_distance(const uint64_t *a, const uint64_t *b)
{
const uint8_t *x = (const uint8_t *)a, *y = (const uint8_t *)b;
uint64_t sum = 0;
for (size_t i = 0; i < NSAMP * 8; i++)
sum += (uint64_t)abs((int)x[i] - (int)y[i]);
return (double)sum / (NSAMP * 8);
}
static uint64_t *filled_at(const cam_t *c, int s)
{
return &filled[c->slot_group[s]][c->slot_pos[s]];
}
/*
* XRService mostly queues each index with the same buffer, but not always, and
* under load the order churns. When the mapped buffer holds no new frame, look
* for the one that does, starting with the buffers filled longest ago (the
* likeliest to be next), and remember it for this index. Each probe costs a
* cache sync of the whole buffer (the camera's DMA isn't cache-coherent), so it
* stops at the first fresh one. Two cameras sharing a run finish frames at the
* same moment, so there it takes the fresher of the first two that looks most
* like this camera's previous frames.
*/
static int find_fresh(cam_t *c, uint64_t *out)
{
static uint64_t cur[NSAMP];
int order[MAX_SLOTS], best = -1, nfresh = 0, n;
double bestd = 1e18;
for (int s = 0; s < c->nslots; s++)
order[s] = s;
for (int i = 1; i < c->nslots; i++) /* insertion sort, oldest first */
for (int j = i; j > 0 && *filled_at(c, order[j]) < *filled_at(c, order[j - 1]); j--) {
int t = order[j]; order[j] = order[j - 1]; order[j - 1] = t;
}
for (int k = 0; k < c->nslots; k++) {
int s = order[k];
buf_sync(c->fd[s], DMA_BUF_SYNC_START);
sample_words(c->map[s], c->need, cur);
buf_sync(c->fd[s], DMA_BUF_SYNC_END);
if (!fresh(cur, c->samp[s], &n))
continue;
double d = c->shared ? fmin(sample_distance(cur, c->hist[0]), sample_distance(cur, c->hist[1])) : 0;
if (d < bestd) {
bestd = d;
best = s;
memcpy(out, cur, sizeof(cur));
}
if (!c->shared || ++nfresh == 2)
break;
}
return best;
}
/* ------------------------------------------------------------- publishing */
/* Copy a frame into ring camera rc (the camera's own, or its dark twin). */
static void publish(cam_t *c, fh_ring_cam_t *rc, uint8_t *slots, uint64_t *frame_no,
const uint8_t *src, uint32_t seq, uint64_t ts, uint64_t evtime, double mean)
{
uint64_t n = ++*frame_no;
fh_ring_slot_t *s = (fh_ring_slot_t *)(slots + (n % rc->nslots) * rc->slot_bytes);
uint8_t *dst = (uint8_t *)(s + 1);
static uint64_t before[NSAMP], after[NSAMP];
__atomic_store_n(&s->seq, 2 * n + 1, __ATOMIC_RELAXED);
__atomic_thread_fence(__ATOMIC_RELEASE);
sample_words(src, c->need, before);
for (unsigned y = 0; y < c->lay.height; y++)
memcpy(dst + (size_t)y * rc->stride, src + (size_t)y * c->lay.pitch, c->lay.width);
sample_words(src, c->need, after);
if (memcmp(before, after, sizeof(before))) {
/* XRService re-queued the buffer and the camera overwrote it mid-copy */
__atomic_store_n(&s->seq, 0, __ATOMIC_RELEASE);
c->torn++;
rc->dropped++;
return;
}
s->frame = n;
s->capture_ns = ts;
s->dqbuf_ns = evtime;
s->publish_ns = mono_ns();
s->v4l2_seq = seq;
s->mean = (float)mean;
__atomic_store_n(&s->seq, 2 * n + 2, __ATOMIC_RELEASE);
__atomic_store_n(&rc->latest, n, __ATOMIC_RELEASE);
rc->published++;
}
static void on_frame(cam_t *c, int64_t index, uint32_t seq, uint64_t ts, uint64_t evtime)
{
c->events++;
if (index < 0 || index >= MAX_INDEX || c->nslots == 0)
return;
if (!c->mapped) {
learn(c, (int)index);
return;
}
if (index > c->maxindex) {
reset_learning(c); /* the queue grew */
return;
}
int slot = c->slot_of[index];
static uint64_t cur[NSAMP];
uint64_t t0 = mono_ns();
buf_sync(c->fd[slot], DMA_BUF_SYNC_START);
c->sync_ns += mono_ns() - t0;
c->nsync++;
sample_words(c->map[slot], c->need, cur);
int nchanged;
int found = -1;
if (!fresh(cur, c->samp[slot], &nchanged)) {
buf_sync(c->fd[slot], DMA_BUF_SYNC_END);
found = find_fresh(c, cur);
if (found >= 0) {
c->slot_of[index] = slot = found;
c->repairs++;
buf_sync(c->fd[slot], DMA_BUF_SYNC_START);
}
}
if (found < 0 && !fresh(cur, c->samp[slot], &nchanged)) {
buf_sync(c->fd[slot], DMA_BUF_SYNC_END);
c->stale++;
c->rc->dropped++;
if (++c->stale_run >= STALE_RELEARN) {
if (++c->relearns > MAX_RELEARNS) {
fprintf(stderr, "%s: its buffers keep going stale; XRService must have new ones\n", c->slug);
exit(3);
}
fprintf(stderr, "%s: %d frames in a row unchanged; relearning its buffers\n", c->slug, c->stale_run);
c->stale_run = 0;
reset_learning(c);
}
return;
}
c->stale_run = 0;
memcpy(c->samp[slot], cur, sizeof(cur));
memcpy(c->hist[1], c->hist[0], sizeof(cur));
memcpy(c->hist[0], cur, sizeof(cur));
*filled_at(c, slot) = ++nfilled;
/*
* The cameras alternate a normal exposure with a near-black one, so judge
* each frame against this camera's recent brightest.
*/
double mean = luma_mean(c, c->map[slot]);
double peak = mean;
c->recent[c->nrecent++ % 8] = mean;
for (int i = 0; i < 8 && i < c->nrecent; i++)
if (c->recent[i] > peak)
peak = c->recent[i];
if (mean < opt_dark * peak) {
c->dark++;
c->rc->dropped++;
if (c->rc_dark)
publish(c, c->rc_dark, c->ring_dark, &c->dark_no, c->map[slot], seq, ts, evtime, mean);
} else {
c->bright++;
uint64_t t1 = mono_ns();
publish(c, c->rc, c->ring_slots, &c->frame_no, c->map[slot], seq, ts, evtime, mean);
c->copy_ns += mono_ns() - t1;
}
buf_sync(c->fd[slot], DMA_BUF_SYNC_END);
}
static void on_sample(void *ctx, const tp_sample_t *s)
{
(void)ctx;
cam_t *c = cam_by_minor(tp_get(s->ev, f_dq_minor, s->raw, s->rawlen));
if (c)
on_frame(c, tp_get(s->ev, f_dq_index, s->raw, s->rawlen),
(uint32_t)tp_get(s->ev, f_dq_seq, s->raw, s->rawlen),
(uint64_t)tp_get(s->ev, f_dq_ts, s->raw, s->rawlen), s->time);
}
/* ------------------------------------------------------------ ring + user */
/* The ring lives in a root-owned directory, so nobody can plant a file or link there. */
static uint8_t *ring_create(uid_t uid, gid_t gid, size_t *len_out)
{
if (mkdir(RING_DIR, 0755) < 0 && errno != EEXIST)
die("mkdir %s: %s", RING_DIR, strerror(errno));
struct stat st;
if (lstat(RING_DIR, &st) < 0 || !S_ISDIR(st.st_mode) || st.st_uid != 0 || (st.st_mode & 022))
die("%s must be a directory owned by root and writable only by root", RING_DIR);
size_t len = sizeof(fh_ring_hdr_t);
int nring = opt_with_dark ? 2 * ncams : ncams;
for (int i = 0; i < nring; i++) {
cam_t *c = &cams[i % ncams];
size_t slot = sizeof(fh_ring_slot_t) + (size_t)c->lay.width * c->lay.height;
len += FH_RING_SLOTS * ((slot + 63) & ~(size_t)63);
}
if (unlink(RING_FILE) < 0 && errno != ENOENT)
die("unlink %s: %s", RING_FILE, strerror(errno));
int fd = open(RING_FILE, O_RDWR | O_CREAT | O_EXCL | O_NOFOLLOW | O_CLOEXEC, 0600);
if (fd < 0)
die("create %s: %s", RING_FILE, strerror(errno));
if (fchown(fd, uid, gid) < 0 || ftruncate(fd, (off_t)len) < 0)
die("prepare %s: %s", RING_FILE, strerror(errno));
uint8_t *m = mmap(NULL, len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
close(fd);
if (m == MAP_FAILED)
die("mmap %s: %s", RING_FILE, strerror(errno));
fh_ring_hdr_t *h = (fh_ring_hdr_t *)m;
size_t off = sizeof(*h);
for (int i = 0; i < nring; i++) {
cam_t *c = &cams[i % ncams];
fh_ring_cam_t *rc = &h->cams[i];
bool dark = i >= ncams;
snprintf(rc->sensor, sizeof(rc->sensor), "%s", c->cam->sensor);
snprintf(rc->name, sizeof(rc->name), "%.26s%s", c->slug, dark ? "-dark" : "");
rc->flags = dark ? FH_CAM_DARK : 0;
rc->node = c->cam->node;
rc->format = FH_FMT_GREY8;
rc->width = c->lay.width;
rc->height = c->lay.height;
rc->stride = c->lay.width;
rc->nslots = FH_RING_SLOTS;
rc->slot_offset = off;
rc->slot_bytes = (sizeof(fh_ring_slot_t) + (size_t)rc->stride * rc->height + 63) & ~(size_t)63;
off += rc->nslots * rc->slot_bytes;
if (dark) {
c->rc_dark = rc;
c->ring_dark = m + rc->slot_offset;
} else {
c->rc = rc;
c->ring_slots = m + rc->slot_offset;
}
}
memcpy(h->magic, FH_RING_MAGIC, 8);
h->version = FH_RING_VERSION;
h->header_bytes = sizeof(*h);
h->ncams = (uint32_t)nring;
h->file_bytes = len;
h->writer_pid = getpid();
h->heartbeat_ns = mono_ns();
*len_out = len;
return m;
}
static void target_user(uid_t *uid, gid_t *gid)
{
if (opt_user) {
struct passwd *pw = getpwnam(opt_user);
if (!pw)
die("unknown user %s", opt_user);
*uid = pw->pw_uid;
*gid = pw->pw_gid;
} else if (getenv("SUDO_UID") && getenv("SUDO_GID")) {
*uid = (uid_t)atoi(getenv("SUDO_UID"));
*gid = (gid_t)atoi(getenv("SUDO_GID"));
} else {
die("run through sudo or pass --user NAME: fh-camd drops root once set up");
}
if (*uid == 0)
die("refusing to keep running as root; pass --user NAME");
}
static void drop_root(uid_t uid, gid_t gid)
{
if (setgroups(0, NULL) < 0 || setresgid(gid, gid, gid) < 0 || setresuid(uid, uid, uid) < 0)
die("dropping root: %s", strerror(errno));
if (setuid(0) == 0 || geteuid() == 0)
die("dropping root failed");
prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0);
}
/* ------------------------------------------------------------------- main */
static void status(double secs)
{
printf("[%.0f s]", secs);
for (int i = 0; i < ncams; i++) {
cam_t *c = &cams[i];
printf(" %s %s%.1f fps (dark %llu stale %llu torn %llu remapped %llu, sync %.2f ms copy %.2f ms)",
c->slug, c->mapped ? "" : "learning ", (double)(c->bright - c->last_bright) / opt_status,
(unsigned long long)c->dark, (unsigned long long)c->stale, (unsigned long long)c->torn,
(unsigned long long)c->repairs,
c->nsync ? c->sync_ns / 1e6 / c->nsync : 0.0, c->bright ? c->copy_ns / 1e6 / c->bright : 0.0);
c->last_bright = c->bright;
}
printf("\n");
fflush(stdout);
}
static void on_signal(int sig)
{
(void)sig;
stop = 1;
}
static void usage(const char *argv0)
{
printf("Usage: sudo %s [options]\n"
" --user NAME user to run as after setup and to own the ring (default: $SUDO_USER)\n"
" --sensor S only cameras whose sensor name contains S (default: all mono cameras)\n"
" --dark R a frame dimmer than R x the camera's recent brightest is dark (default 0.4)\n"
" --with-dark also publish the dark frames, as extra ring cameras flagged FH_CAM_DARK\n"
" --status S print a status line every S seconds, 0 for never (default 10)\n"
"Frames go to " RING_FILE " (layout in fhring.h).\n", argv0);
}
int main(int argc, char **argv)
{
for (int i = 1; i < argc; i++) {
if (!strcmp(argv[i], "--user") && i + 1 < argc)
opt_user = argv[++i];
else if (!strcmp(argv[i], "--sensor") && i + 1 < argc)
opt_sensor = argv[++i];
else if (!strcmp(argv[i], "--dark") && i + 1 < argc)
opt_dark = atof(argv[++i]);
else if (!strcmp(argv[i], "--with-dark"))
opt_with_dark = true;
else if (!strcmp(argv[i], "--status") && i + 1 < argc)
opt_status = atof(argv[++i]);
else {
usage(argv[0]);
return strcmp(argv[i], "--help") ? 1 : 0;
}
}
if (geteuid() != 0)
die("run with sudo: borrowing XRService's buffers and reading tracepoints need root");
uid_t uid;
gid_t gid;
target_user(&uid, &gid);
char err[512];
if (!xr_discover(&xr, "XRService", err, sizeof(err)))
die("%s", err);
if (!tp_event_load(&ev_dqbuf, "v4l2", "v4l2_dqbuf", err, sizeof(err)))
die("tracepoint v4l2:v4l2_dqbuf: %s", err);
f_dq_minor = tp_field(&ev_dqbuf, "minor");
f_dq_index = tp_field(&ev_dqbuf, "index");
f_dq_ts = tp_field(&ev_dqbuf, "timestamp");
f_dq_seq = tp_field(&ev_dqbuf, "sequence");
if (f_dq_minor < 0 || f_dq_index < 0 || f_dq_ts < 0 || f_dq_seq < 0)
die("v4l2_dqbuf lacks the minor/index/timestamp/sequence fields");
int pidfd = (int)syscall(SYS_pidfd_open, xr.pid, 0u);
if (pidfd < 0)
die("pidfd_open(%d): %s", xr.pid, strerror(errno));
printf("XRService pid %d\ncameras:\n", xr.pid);
for (int i = 0; i < xr.ncameras && ncams < MAX_CAMS; i++) {
xr_camera_t *cam = &xr.cameras[i];
xr_layout_t lay;
xr_camera_layout(cam, &lay);
if (lay.fmt == XR_FMT_GREY8 && strstr(cam->sensor, opt_sensor))
setup_camera(&cams[ncams++], cam, pidfd);
}
if (!ncams)
die("no mono camera matches '%s'", opt_sensor);
if (opt_with_dark && 2 * ncams > FH_RING_MAX_CAMS)
die("--with-dark needs a ring camera per dark stream too: pick at most %d cameras with --sensor",
FH_RING_MAX_CAMS / 2);
static tp_t tp; /* large: pending-sample pool */
tp_event_t *evs[1] = { &ev_dqbuf };
if (!tp_open(&tp, evs, 1, err, sizeof(err)))
die("%s", err);
size_t ring_len;
fh_ring_hdr_t *ring = (fh_ring_hdr_t *)ring_create(uid, gid, &ring_len);
drop_root(uid, gid);
printf("ring %s (%.1f MB), running as uid %d\n", RING_FILE, ring_len / 1e6, (int)getuid());
fflush(stdout);
signal(SIGINT, on_signal);
signal(SIGTERM, on_signal);
uint64_t start = mono_ns(), last_status = start;
int rc = 0;
while (!stop) {
if (tp_poll(&tp, 100, on_sample, NULL) < 0) {
rc = 1;
break;
}
uint64_t now = mono_ns();
__atomic_store_n(&ring->heartbeat_ns, now, __ATOMIC_RELEASE);
struct pollfd pf = { .fd = pidfd, .events = POLLIN };
if (poll(&pf, 1, 0) > 0) {
fprintf(stderr, "XRService exited\n");
rc = 2;
break;
}
if (opt_status > 0 && (double)(now - last_status) / 1e9 >= opt_status) {
status((double)(now - start) / 1e9);
last_status = now;
}
}
tp_close(&tp);
ring->heartbeat_ns = 0; /* tells readers the writer is gone */
return rc;
}
+83
View File
@@ -0,0 +1,83 @@
/*
* fhring - the shared-memory frame ring fh-camd writes and trackers read.
*
* One file (FH_RING_PATH) holds a header, then for each camera
* a few slots, each a slot header followed by the image rows packed tightly
* (stride == width for 8-bit mono). Only complete, bright frames are published.
*
* Writer, for frame n of a camera: slot = n % nslots
* slot.seq = 2n+1; write slot fields and pixels; slot.seq = 2n+2; cam.latest = n
* Reader:
* n = cam.latest; read slot.seq, expect 2n+2; copy; re-read slot.seq; if it
* changed the copy is torn, retry with the new latest.
*
* All multi-byte fields are little-endian; offsets are fixed so Python can read
* them with struct (tracker/ring.py mirrors this file).
*/
#pragma once
#include <assert.h>
#include <stdint.h>
#define FH_RING_MAGIC "FHRING01"
#define FH_RING_VERSION 1
#define FH_RING_MAX_CAMS 8
#define FH_RING_SLOTS 4
#define FH_RING_DIR "/run/frame-hands"
#define FH_RING_PATH FH_RING_DIR "/ir-ring"
enum {
FH_FMT_GREY8 = 0,
};
enum {
FH_CAM_DARK = 1u << 0, /* the near-black exposures between this node's */
/* normal frames (fh-camd --with-dark) */
};
typedef struct {
char sensor[32]; /* media entity, e.g. "og01a1b 4-0060" */
char name[32]; /* calibration name if known, else sensor slug */
int32_t node; /* N of /dev/videoN */
uint32_t format; /* FH_FMT_* */
uint32_t width;
uint32_t height;
uint32_t stride; /* bytes per row in the ring */
uint32_t nslots;
uint64_t slot_offset; /* file offset of slot 0 */
uint64_t slot_bytes; /* slot header + image, 64-byte aligned */
volatile uint64_t latest; /* newest published frame number, 0 = none yet */
uint64_t published; /* frames published */
uint64_t dropped; /* dark, stale or torn frames not published */
uint32_t flags; /* FH_CAM_* */
uint8_t reserved[28];
} fh_ring_cam_t; /* 160 bytes */
typedef struct {
volatile uint64_t seq; /* 2n+1 while frame n is written, 2n+2 when done */
uint64_t frame; /* n */
uint64_t capture_ns; /* V4L2 timestamp (camera clock) */
uint64_t dqbuf_ns; /* CLOCK_MONOTONIC when XRService dequeued it */
uint64_t publish_ns; /* CLOCK_MONOTONIC when the copy finished */
uint32_t v4l2_seq; /* V4L2 sequence number */
float mean; /* mean luma on a sparse grid */
uint8_t reserved[16];
} fh_ring_slot_t; /* 64 bytes, image follows */
typedef struct {
char magic[8]; /* FH_RING_MAGIC */
uint32_t version;
uint32_t header_bytes; /* sizeof(fh_ring_hdr_t) */
uint32_t ncams;
uint32_t reserved0;
uint64_t file_bytes;
int64_t writer_pid;
volatile uint64_t heartbeat_ns; /* CLOCK_MONOTONIC, refreshed at least every 0.2 s */
uint8_t reserved[16];
fh_ring_cam_t cams[FH_RING_MAX_CAMS];
} fh_ring_hdr_t;
static_assert(sizeof(fh_ring_cam_t) == 160, "fh_ring_cam_t layout");
static_assert(sizeof(fh_ring_slot_t) == 64, "fh_ring_slot_t layout");
static_assert(sizeof(fh_ring_hdr_t) == 64 + 160 * FH_RING_MAX_CAMS, "fh_ring_hdr_t layout");
+427
View File
@@ -0,0 +1,427 @@
/*
* tp - read kernel tracepoints system-wide through perf_event_open.
*/
#define _GNU_SOURCE
#include "tp.h"
#include <errno.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/epoll.h>
#include <sys/ioctl.h>
#include <sys/mman.h>
#include <sys/syscall.h>
#include <time.h>
#include <unistd.h>
#include <linux/perf_event.h>
#ifndef TRACEFS
#define TRACEFS "/sys/kernel/tracing/events"
#endif
#define RING_DATA_PAGES 16
static void set_err(char *err, size_t n, const char *fmt, ...)
{
va_list ap;
va_start(ap, fmt);
vsnprintf(err, n, fmt, ap);
va_end(ap);
}
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn)
{
memset(ev, 0, sizeof(*ev));
snprintf(ev->system, sizeof(ev->system), "%s", system);
snprintf(ev->name, sizeof(ev->name), "%s", name);
ev->id = -1;
char path[256];
snprintf(path, sizeof(path), TRACEFS "/%s/%s/format", system, name);
FILE *f = fopen(path, "r");
if (!f) {
set_err(err, errn, "%s: %s", path, strerror(errno));
return false;
}
char line[512];
while (fgets(line, sizeof(line), f)) {
int id;
if (sscanf(line, "ID: %d", &id) == 1) {
ev->id = id;
continue;
}
char *fp = line;
while (*fp == ' ' || *fp == '\t')
fp++;
if (strncmp(fp, "field:", 6) || ev->nfields >= TP_MAX_FIELDS)
continue;
char *semi = strchr(fp, ';');
if (!semi)
continue;
/* the field name is the last identifier in the declaration */
char decl[256];
size_t dl = (size_t)(semi - (fp + 6));
if (dl >= sizeof(decl))
dl = sizeof(decl) - 1;
memcpy(decl, fp + 6, dl);
decl[dl] = 0;
char *br = strchr(decl, '[');
if (br)
*br = 0;
char *end = decl + strlen(decl);
while (end > decl && (end[-1] == ' ' || end[-1] == '\t'))
*--end = 0;
char *start = end;
while (start > decl && start[-1] != ' ' && start[-1] != '\t' && start[-1] != '*')
start--;
tp_field_t *fd = &ev->fields[ev->nfields];
const char *o = strstr(semi, "offset:");
const char *s = strstr(semi, "size:");
const char *g = strstr(semi, "signed:");
if (!o || !s)
continue;
snprintf(fd->name, sizeof(fd->name), "%s", start);
fd->offset = atoi(o + 7);
fd->size = atoi(s + 5);
fd->is_signed = g ? atoi(g + 7) != 0 : false;
ev->nfields++;
}
fclose(f);
if (ev->id < 0) {
set_err(err, errn, "%s: no ID line", path);
return false;
}
return true;
}
int tp_field(const tp_event_t *ev, const char *name)
{
for (int i = 0; i < ev->nfields; i++)
if (!strcmp(ev->fields[i].name, name))
return i;
return -1;
}
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen)
{
if (field < 0 || field >= ev->nfields)
return 0;
const tp_field_t *f = &ev->fields[field];
if (f->offset < 0 || (uint32_t)(f->offset + f->size) > rawlen)
return 0;
const uint8_t *p = raw + f->offset;
switch (f->size) {
case 1: { uint8_t v; memcpy(&v, p, 1); return f->is_signed ? (int64_t)(int8_t)v : (int64_t)v; }
case 2: { uint16_t v; memcpy(&v, p, 2); return f->is_signed ? (int64_t)(int16_t)v : (int64_t)v; }
case 4: { uint32_t v; memcpy(&v, p, 4); return f->is_signed ? (int64_t)(int32_t)v : (int64_t)v; }
case 8: { uint64_t v; memcpy(&v, p, 8); return (int64_t)v; }
default: return 0;
}
}
static int online_cpus(int *cpus, int max)
{
FILE *f = fopen("/sys/devices/system/cpu/online", "r");
int n = 0;
if (!f)
return 0;
char buf[256] = {0};
if (!fgets(buf, sizeof(buf), f))
buf[0] = 0;
fclose(f);
for (char *tok = strtok(buf, ",\n"); tok && n < max; tok = strtok(NULL, ",\n")) {
int a, b;
if (sscanf(tok, "%d-%d", &a, &b) == 2) {
for (int c = a; c <= b && n < max; c++)
cpus[n++] = c;
} else if (sscanf(tok, "%d", &a) == 1) {
cpus[n++] = a;
}
}
return n;
}
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn)
{
memset(tp, 0, sizeof(*tp));
tp->epfd = -1;
if (nevents <= 0 || nevents > TP_MAX_EVENTS) {
set_err(err, errn, "bad event count %d", nevents);
return false;
}
for (int i = 0; i < nevents; i++)
tp->events[i] = events[i];
tp->nevents = nevents;
int cpus[TP_MAX_CPUS];
tp->ncpu = online_cpus(cpus, TP_MAX_CPUS);
if (tp->ncpu <= 0) {
set_err(err, errn, "no online CPUs found");
return false;
}
long page = sysconf(_SC_PAGESIZE);
tp->map_len = (size_t)page * (1 + RING_DATA_PAGES);
tp->epfd = epoll_create1(EPOLL_CLOEXEC);
if (tp->epfd < 0) {
set_err(err, errn, "epoll_create1: %s", strerror(errno));
return false;
}
for (int c = 0; c < tp->ncpu; c++) {
tp->ring_fd[c] = -1;
for (int e = 0; e < nevents; e++) {
struct perf_event_attr a;
memset(&a, 0, sizeof(a));
a.size = sizeof(a);
a.type = PERF_TYPE_TRACEPOINT;
a.config = (uint64_t)events[e]->id;
a.sample_period = 1;
a.sample_type = PERF_SAMPLE_TID | PERF_SAMPLE_TIME | PERF_SAMPLE_CPU | PERF_SAMPLE_RAW;
a.wakeup_events = 1;
a.use_clockid = 1;
a.clockid = CLOCK_MONOTONIC;
a.disabled = 1;
int fd = (int)syscall(SYS_perf_event_open, &a, -1, cpus[c], -1, PERF_FLAG_FD_CLOEXEC);
if (fd < 0) {
set_err(err, errn, "perf_event_open(%s:%s, cpu %d): %s",
events[e]->system, events[e]->name, cpus[c], strerror(errno));
tp_close(tp);
return false;
}
tp->fds[tp->nfds++] = fd;
if (tp->ring_fd[c] < 0) {
void *m = mmap(NULL, tp->map_len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
if (m == MAP_FAILED) {
set_err(err, errn, "mmap perf ring (cpu %d): %s", cpus[c], strerror(errno));
tp_close(tp);
return false;
}
tp->ring[c] = m;
tp->ring_fd[c] = fd;
struct epoll_event ee = { .events = EPOLLIN, .data.u32 = (uint32_t)c };
epoll_ctl(tp->epfd, EPOLL_CTL_ADD, fd, &ee);
} else if (ioctl(fd, PERF_EVENT_IOC_SET_OUTPUT, tp->ring_fd[c]) < 0) {
set_err(err, errn, "PERF_EVENT_IOC_SET_OUTPUT: %s", strerror(errno));
tp_close(tp);
return false;
}
}
}
for (int i = 0; i < tp->nfds; i++)
ioctl(tp->fds[i], PERF_EVENT_IOC_ENABLE, 0);
return true;
}
static void ring_copy(uint8_t *dst, const uint8_t *base, uint64_t size, uint64_t pos, size_t len)
{
uint64_t off = pos % size;
size_t first = (size_t)(size - off);
if (first >= len) {
memcpy(dst, base + off, len);
} else {
memcpy(dst, base + off, first);
memcpy(dst + first, base, len - first);
}
}
static int cmp_sample(const void *a, const void *b)
{
const tp_sample_t *x = a, *y = b;
return (x->time > y->time) - (x->time < y->time);
}
static void dispatch(tp_t *tp, tp_cb cb, void *ctx)
{
qsort(tp->pend, tp->npend, sizeof(tp->pend[0]), cmp_sample);
for (int i = 0; i < tp->npend; i++)
cb(ctx, &tp->pend[i]);
tp->npend = 0;
}
static int drain_ring(tp_t *tp, int c, tp_cb cb, void *ctx)
{
struct perf_event_mmap_page *pg = tp->ring[c];
long page = sysconf(_SC_PAGESIZE);
uint64_t off = pg->data_offset ? pg->data_offset : (uint64_t)page;
uint64_t size = pg->data_size ? pg->data_size : (uint64_t)page * RING_DATA_PAGES;
const uint8_t *base = (const uint8_t *)pg + off;
uint64_t head = __atomic_load_n(&pg->data_head, __ATOMIC_ACQUIRE);
uint64_t tail = pg->data_tail;
int n = 0;
while (tail < head) {
struct perf_event_header hdr;
ring_copy((uint8_t *)&hdr, base, size, tail, sizeof(hdr));
if (hdr.size < sizeof(hdr))
break;
ring_copy(tp->scratch, base, size, tail, hdr.size);
const uint8_t *p = tp->scratch + sizeof(hdr);
const uint8_t *end = tp->scratch + hdr.size;
if (hdr.type == PERF_RECORD_LOST && end - p >= 16) {
uint64_t lost;
memcpy(&lost, p + 8, 8);
tp->lost += lost;
} else if (hdr.type == PERF_RECORD_SAMPLE && end - p >= 28) {
tp_sample_t s;
uint32_t v32[2];
memcpy(v32, p, 8); p += 8;
s.pid = v32[0];
s.tid = v32[1];
memcpy(&s.time, p, 8); p += 8;
memcpy(v32, p, 8); p += 8;
s.cpu = v32[0];
memcpy(&s.rawlen, p, 4); p += 4;
s.raw = p;
if (s.rawlen >= 2 && p + s.rawlen <= end) {
uint16_t type;
memcpy(&type, s.raw, 2);
s.ev = NULL;
for (int e = 0; e < tp->nevents; e++)
if (tp->events[e]->id == type)
s.ev = tp->events[e];
if (s.ev && s.rawlen <= TP_MAX_RAW) {
if (tp->npend == TP_MAX_PENDING)
dispatch(tp, cb, ctx);
memcpy(tp->pend_raw[tp->npend], s.raw, s.rawlen);
s.raw = tp->pend_raw[tp->npend];
tp->pend[tp->npend++] = s;
n++;
}
}
}
tail += hdr.size;
}
__atomic_store_n(&pg->data_tail, tail, __ATOMIC_RELEASE);
return n;
}
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx)
{
struct epoll_event ev[TP_MAX_CPUS];
if (epoll_wait(tp->epfd, ev, TP_MAX_CPUS, timeout_ms) < 0 && errno != EINTR)
return -1;
/*
* Drain every ring, not just the ones that woke us: samples from several
* CPUs need to be handled together to keep per-camera order sane.
*/
int n = 0;
for (int c = 0; c < tp->ncpu; c++)
if (tp->ring[c])
n += drain_ring(tp, c, cb, ctx);
dispatch(tp, cb, ctx);
return n;
}
void tp_close(tp_t *tp)
{
for (int i = 0; i < tp->nfds; i++) {
ioctl(tp->fds[i], PERF_EVENT_IOC_DISABLE, 0);
}
for (int c = 0; c < tp->ncpu; c++)
if (tp->ring[c])
munmap(tp->ring[c], tp->map_len);
for (int i = 0; i < tp->nfds; i++)
close(tp->fds[i]);
if (tp->epfd >= 0)
close(tp->epfd);
tp->nfds = 0;
tp->epfd = -1;
}
+76
View File
@@ -0,0 +1,76 @@
/*
* tp - read kernel tracepoints system-wide through perf_event_open.
*
* One perf ring per CPU; every event on that CPU writes into it. Field
* offsets come from the tracefs format files, so kernel layout changes don't
* silently break parsing. Needs root (or CAP_PERFMON plus tracefs access).
*/
#pragma once
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#define TP_MAX_FIELDS 40
#define TP_MAX_EVENTS 8
#define TP_MAX_CPUS 64
#define TP_MAX_PENDING 2048
#define TP_MAX_RAW 256
typedef struct {
char name[48];
int offset;
int size;
bool is_signed;
} tp_field_t;
typedef struct {
char system[32];
char name[48];
int id;
tp_field_t fields[TP_MAX_FIELDS];
int nfields;
} tp_event_t;
typedef struct {
const tp_event_t *ev;
const uint8_t *raw;
uint32_t rawlen;
uint64_t time; /* CLOCK_MONOTONIC ns */
uint32_t cpu;
uint32_t pid;
uint32_t tid;
} tp_sample_t;
typedef void (*tp_cb)(void *ctx, const tp_sample_t *s);
typedef struct {
int ncpu;
int ring_fd[TP_MAX_CPUS];
void *ring[TP_MAX_CPUS];
size_t map_len;
int fds[TP_MAX_CPUS * TP_MAX_EVENTS];
int nfds;
int epfd;
tp_event_t *events[TP_MAX_EVENTS];
int nevents;
uint64_t lost;
uint8_t scratch[65536];
/* samples drained from all rings, sorted by time before dispatch */
tp_sample_t pend[TP_MAX_PENDING];
uint8_t pend_raw[TP_MAX_PENDING][TP_MAX_RAW];
int npend;
} tp_t;
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn);
int tp_field(const tp_event_t *ev, const char *name);
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen);
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn);
/*
* Wait up to timeout_ms, then hand every pending sample to cb in time order,
* across all CPUs. Returns samples read, -1 on error.
*/
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx);
void tp_close(tp_t *tp);
+783
View File
@@ -0,0 +1,783 @@
/*
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
*
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
*
* Everything is discovered rather than hardcoded:
* - XRService is found by scanning /proc for its cmdline.
* - The V4L2 nodes and sensor subdevs it holds open come from /proc/<pid>/fd.
* - Each node's geometry comes from VIDIOC_G_FMT on our own handle.
* - Each node is traced back to its sensor through MEDIA_IOC_G_TOPOLOGY.
* - Buffers are split into queues by allocation order: XRService opens a
* sensor subdev, then allocates that camera's buffers.
*/
#define _GNU_SOURCE
#include "xrcams.h"
#include <dirent.h>
#include <errno.h>
#include <fcntl.h>
#include <stdarg.h>
#include <stdlib.h>
#include <string.h>
#include <sys/ioctl.h>
#include <sys/stat.h>
#include <sys/sysmacros.h>
#include <unistd.h>
#include <linux/media.h>
#ifndef MEDIA_ENT_F_CAM_SENSOR
#define MEDIA_ENT_F_CAM_SENSOR 0x00020001
#endif
#define MAX_FDENTS 4096
#define MAX_TOPOS 8
enum fdkind { FD_DMABUF, FD_SUBDEV_SENSOR, FD_VIDEO };
typedef struct {
int xfd;
enum fdkind kind;
size_t size;
unsigned long ino;
char sensor[XR_SENSOR_LEN];
char path[64];
} fdent_t;
typedef struct {
struct media_v2_entity *ents;
struct media_v2_interface *intfs;
struct media_v2_pad *pads;
struct media_v2_link *links;
__u32 nents, nintfs, npads, nlinks;
} topo_t;
static fdent_t fdents[MAX_FDENTS];
static int nfdents;
static topo_t topos[MAX_TOPOS];
static int ntopos;
static void set_err(char *err, size_t n, const char *fmt, ...)
{
va_list ap;
va_start(ap, fmt);
vsnprintf(err, n, fmt, ap);
va_end(ap);
}
void xr_slugify(const char *in, char *out, size_t n)
{
size_t i = 0;
for (; in[i] && i + 1 < n; i++)
out[i] = (in[i] == ' ' || in[i] == '/') ? '_' : in[i];
out[i] = 0;
}
/* --------------------------------------------------- media graph handling */
static void topo_free_all(void)
{
for (int i = 0; i < ntopos; i++) {
free(topos[i].ents);
free(topos[i].intfs);
free(topos[i].pads);
free(topos[i].links);
}
ntopos = 0;
}
static void topo_load_all(void)
{
for (int mi = 0; mi < MAX_TOPOS; mi++) {
char mpath[32];
snprintf(mpath, sizeof(mpath), "/dev/media%d", mi);
int mfd = open(mpath, O_RDWR | O_CLOEXEC);
if (mfd < 0)
continue;
struct media_v2_topology t;
memset(&t, 0, sizeof(t));
if (ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) < 0) {
close(mfd);
continue;
}
topo_t *o = &topos[ntopos];
memset(o, 0, sizeof(*o));
o->nents = t.num_entities;
o->nintfs = t.num_interfaces;
o->npads = t.num_pads;
o->nlinks = t.num_links;
o->ents = calloc(o->nents ? o->nents : 1, sizeof(*o->ents));
o->intfs = calloc(o->nintfs ? o->nintfs : 1, sizeof(*o->intfs));
o->pads = calloc(o->npads ? o->npads : 1, sizeof(*o->pads));
o->links = calloc(o->nlinks ? o->nlinks : 1, sizeof(*o->links));
t.ptr_entities = (__u64)(uintptr_t)o->ents;
t.ptr_interfaces = (__u64)(uintptr_t)o->intfs;
t.ptr_pads = (__u64)(uintptr_t)o->pads;
t.ptr_links = (__u64)(uintptr_t)o->links;
bool ok = o->ents && o->intfs && o->pads && o->links &&
ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) == 0;
close(mfd);
if (!ok) {
free(o->ents); free(o->intfs); free(o->pads); free(o->links);
continue;
}
ntopos++;
}
}
static struct media_v2_entity *topo_entity(topo_t *t, __u32 id)
{
for (__u32 i = 0; i < t->nents; i++)
if (t->ents[i].id == id)
return &t->ents[i];
return NULL;
}
static struct media_v2_pad *topo_pad(topo_t *t, __u32 id)
{
for (__u32 i = 0; i < t->npads; i++)
if (t->pads[i].id == id)
return &t->pads[i];
return NULL;
}
static __u32 topo_entity_for_devnode(topo_t *t, dev_t rdev)
{
__u32 intf_id = 0;
for (__u32 i = 0; i < t->nintfs; i++)
if (t->intfs[i].devnode.major == major(rdev) &&
t->intfs[i].devnode.minor == minor(rdev)) {
intf_id = t->intfs[i].id;
break;
}
if (!intf_id)
return 0;
for (__u32 i = 0; i < t->nlinks; i++)
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) == MEDIA_LNK_FL_INTERFACE_LINK &&
t->links[i].source_id == intf_id)
return t->links[i].sink_id;
return 0;
}
/*
* Walk upstream across enabled data links until a sensor is reached. A CSIPHY
* carries two sensors on separate (sink, source) pad pairs, so re-enter on the
* sink pad paired with the source pad we left through.
*/
static bool topo_walk_to_sensor(topo_t *t, __u32 ent_id, char *out, size_t outn)
{
int exit_pad_index = -1;
for (int hop = 0; hop < 32 && ent_id; hop++) {
struct media_v2_entity *e = topo_entity(t, ent_id);
if (!e)
return false;
if (e->function == MEDIA_ENT_F_CAM_SENSOR) {
snprintf(out, outn, "%s", e->name);
return true;
}
__u32 first_sink = 0, paired = 0;
int nsinks = 0;
for (__u32 p = 0; p < t->npads; p++) {
if (t->pads[p].entity_id != ent_id || !(t->pads[p].flags & MEDIA_PAD_FL_SINK))
continue;
nsinks++;
if (!first_sink)
first_sink = t->pads[p].id;
if (exit_pad_index >= 1 && (int)t->pads[p].index == exit_pad_index - 1)
paired = t->pads[p].id;
}
__u32 sink_pad = (nsinks == 1) ? first_sink : (paired ? paired : first_sink);
if (!sink_pad)
return false;
__u32 src_pad = 0;
for (__u32 i = 0; i < t->nlinks; i++) {
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) != MEDIA_LNK_FL_DATA_LINK)
continue;
if (!(t->links[i].flags & MEDIA_LNK_FL_ENABLED))
continue;
if (t->links[i].sink_id == sink_pad) {
src_pad = t->links[i].source_id;
break;
}
}
struct media_v2_pad *sp = src_pad ? topo_pad(t, src_pad) : NULL;
if (!sp)
return false;
ent_id = sp->entity_id;
exit_pad_index = (int)sp->index;
}
return false;
}
static bool sensor_for_video(dev_t rdev, char *out, size_t outn)
{
for (int i = 0; i < ntopos; i++) {
__u32 ent = topo_entity_for_devnode(&topos[i], rdev);
if (ent && topo_walk_to_sensor(&topos[i], ent, out, outn))
return true;
}
return false;
}
static bool sensor_for_subdev(dev_t rdev, char *out, size_t outn)
{
for (int i = 0; i < ntopos; i++) {
__u32 id = topo_entity_for_devnode(&topos[i], rdev);
struct media_v2_entity *e = id ? topo_entity(&topos[i], id) : NULL;
if (e && e->function == MEDIA_ENT_F_CAM_SENSOR) {
snprintf(out, outn, "%s", e->name);
return true;
}
}
return false;
}
static const char *role_for_sensor(const char *sensor)
{
if (strstr(sensor, "og01a1b"))
return "tracking"; /* 1056x1024 side fisheye */
if (strstr(sensor, "og0ve10"))
return "tracking"; /* 640x480 upper */
if (strstr(sensor, "imx616"))
return "passthrough"; /* 2464x2464 Arcturus color */
return "unknown";
}
/* ------------------------------------------------- XRService / proc scan */
static pid_t find_process(const char *needle)
{
DIR *d = opendir("/proc");
if (!d)
return 0;
struct dirent *e;
pid_t found = 0;
while ((e = readdir(d))) {
if (e->d_name[0] < '0' || e->d_name[0] > '9')
continue;
char path[288];
snprintf(path, sizeof(path), "/proc/%s/cmdline", e->d_name);
FILE *f = fopen(path, "rb");
if (!f)
continue;
char buf[512] = {0};
size_t got = fread(buf, 1, sizeof(buf) - 1, f);
fclose(f);
if (got == 0)
continue;
const char *base = strrchr(buf, '/');
base = base ? base + 1 : buf;
if (strstr(base, needle)) {
found = (pid_t)atoi(e->d_name);
break;
}
}
closedir(d);
return found;
}
static bool read_dmabuf_size(pid_t pid, int fd, size_t *size, unsigned long *ino)
{
char path[64];
snprintf(path, sizeof(path), "/proc/%d/fdinfo/%d", pid, fd);
FILE *f = fopen(path, "r");
if (!f)
return false;
bool have = false;
char line[256];
*ino = 0;
while (fgets(line, sizeof(line), f)) {
unsigned long long v;
if (sscanf(line, "size: %llu", &v) == 1) {
*size = (size_t)v;
have = true;
} else if (sscanf(line, "ino: %llu", &v) == 1) {
*ino = (unsigned long)v;
}
}
fclose(f);
return have;
}
static int cmp_int(const void *a, const void *b)
{
return *(const int *)a - *(const int *)b;
}
static bool scan_xr_fds(pid_t pid, char *err, size_t errn)
{
char dirpath[64];
snprintf(dirpath, sizeof(dirpath), "/proc/%d/fd", pid);
DIR *d = opendir(dirpath);
if (!d) {
set_err(err, errn, "opendir(%s): %s (are you root?)", dirpath, strerror(errno));
return false;
}
static int fds[8192];
int nfds = 0;
struct dirent *e;
while ((e = readdir(d)) && nfds < (int)(sizeof(fds) / sizeof(fds[0])))
if (e->d_name[0] >= '0' && e->d_name[0] <= '9')
fds[nfds++] = atoi(e->d_name);
closedir(d);
qsort(fds, nfds, sizeof(int), cmp_int);
nfdents = 0;
for (int i = 0; i < nfds && nfdents < MAX_FDENTS; i++) {
char link[64], target[256];
snprintf(link, sizeof(link), "/proc/%d/fd/%d", pid, fds[i]);
ssize_t n = readlink(link, target, sizeof(target) - 1);
if (n < 0)
continue;
target[n] = 0;
fdent_t ent;
memset(&ent, 0, sizeof(ent));
ent.xfd = fds[i];
if (strstr(target, "dmabuf")) {
if (!read_dmabuf_size(pid, fds[i], &ent.size, &ent.ino))
continue;
ent.kind = FD_DMABUF;
} else if (strncmp(target, "/dev/video", 10) == 0) {
ent.kind = FD_VIDEO;
snprintf(ent.path, sizeof(ent.path), "%s", target);
} else if (strncmp(target, "/dev/v4l-subdev", 15) == 0) {
struct stat st;
if (stat(target, &st) < 0 || !sensor_for_subdev(st.st_rdev, ent.sensor, sizeof(ent.sensor)))
continue;
ent.kind = FD_SUBDEV_SENSOR;
} else {
continue;
}
fdents[nfdents++] = ent;
}
return true;
}
/* ------------------------------------------------------ camera discovery */
static void probe_cameras(xr_state_t *st)
{
int seen[64];
int nseen = 0;
for (int i = 0; i < nfdents; i++) {
if (fdents[i].kind != FD_VIDEO)
continue;
const char *path = fdents[i].path;
int node = atoi(path + 10);
bool dup = false;
for (int k = 0; k < nseen; k++)
if (seen[k] == node)
dup = true;
if (dup || st->ncameras >= XR_MAX_CAMERAS || nseen >= 64)
continue;
seen[nseen++] = node;
int fd = open(path, O_RDWR | O_CLOEXEC);
if (fd < 0)
continue;
struct v4l2_format fmt;
memset(&fmt, 0, sizeof(fmt));
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE;
xr_camera_t *c = &st->cameras[st->ncameras];
memset(c, 0, sizeof(*c));
if (ioctl(fd, VIDIOC_G_FMT, &fmt) == 0) {
c->width = fmt.fmt.pix_mp.width;
c->height = fmt.fmt.pix_mp.height;
c->pixfmt = fmt.fmt.pix_mp.pixelformat;
c->nplanes = fmt.fmt.pix_mp.num_planes;
c->bytesperline = fmt.fmt.pix_mp.plane_fmt[0].bytesperline;
for (unsigned p = 0; p < c->nplanes && p < VIDEO_MAX_PLANES; p++)
c->planesize[p] = fmt.fmt.pix_mp.plane_fmt[p].sizeimage;
} else {
memset(&fmt, 0, sizeof(fmt));
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
if (ioctl(fd, VIDIOC_G_FMT, &fmt) < 0) {
close(fd);
continue;
}
c->width = fmt.fmt.pix.width;
c->height = fmt.fmt.pix.height;
c->pixfmt = fmt.fmt.pix.pixelformat;
c->nplanes = 1;
c->bytesperline = fmt.fmt.pix.bytesperline;
c->planesize[0] = fmt.fmt.pix.sizeimage;
}
struct stat sb;
if (fstat(fd, &sb) == 0) {
c->minor = minor(sb.st_rdev);
sensor_for_video(sb.st_rdev, c->sensor, sizeof(c->sensor));
}
close(fd);
if (!c->sensor[0])
snprintf(c->sensor, sizeof(c->sensor), "unknown");
c->node = node;
snprintf(c->path, sizeof(c->path), "%s", path);
c->role = role_for_sensor(c->sensor);
st->ncameras++;
}
}
/*
* qcom-camss can report bytesperline as the visible width while the VFE
* writes a larger aligned pitch. sizeimage is right, so derive the pitch.
*/
unsigned xr_camera_stride(const xr_camera_t *c)
{
if (!c->height || !c->planesize[0])
return c->bytesperline ? c->bytesperline : c->width;
double bpp = 1.0;
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21)
bpp = 1.5;
unsigned s = (unsigned)((double)c->planesize[0] / ((double)c->height * bpp));
if (s >= c->width && s <= c->width * 4)
return s;
return c->bytesperline ? c->bytesperline : c->width;
}
/*
* The Arcturus color cameras (arcimx616) claim 2464x2464 NV12, but measured on
* 2026-09-28 their plane 0 holds 10-bit MIPI-packed YUV 4:2:0: 2464 luma rows
* then 1232 rows of interleaved UV, each row 2464 packed pixels (3080 bytes)
* padded to a 256-byte pitch (3328). Only the first 1972 pixels of a row carry
* image; the rest are zero.
*/
#define IMX616_VALID_WIDTH 1972
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l)
{
memset(l, 0, sizeof(*l));
l->height = c->height;
if (c->pixfmt == V4L2_PIX_FMT_NV12 && strstr(c->sensor, "imx616")) {
unsigned packed = (c->width * 5 + 3) / 4;
l->fmt = XR_FMT_YUV420_10P;
l->pitch = (packed + 255) & ~255u;
l->rows = c->height + c->height / 2;
l->width = IMX616_VALID_WIDTH < c->width ? IMX616_VALID_WIDTH : c->width;
return;
}
l->pitch = xr_camera_stride(c);
l->width = c->width < l->pitch ? c->width : l->pitch;
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21) {
l->fmt = XR_FMT_NV12;
l->rows = c->height + c->height / 2;
} else {
l->fmt = XR_FMT_GREY8;
l->rows = c->height;
}
}
const char *xr_fmt_name(xr_fmt_t f)
{
switch (f) {
case XR_FMT_GREY8: return "grey8";
case XR_FMT_NV12: return "nv12";
case XR_FMT_YUV420_10P: return "yuv420_10p";
}
return "?";
}
/* ------------------------------------------------------- buffer grouping */
/*
* XRService allocates one udmabuf per plane, plane 0 then plane 1, a whole
* queue at a time right after opening the sensor's subdev. Plane 1 matches
* VIDIOC_G_FMT exactly; plane 0 has slack, so it is matched with >=.
*/
static void build_groups(xr_state_t *st)
{
char current_sensor[XR_SENSOR_LEN] = "";
for (int i = 0; i < nfdents; i++) {
if (fdents[i].kind == FD_SUBDEV_SENSOR) {
snprintf(current_sensor, sizeof(current_sensor), "%s", fdents[i].sensor);
continue;
}
if (fdents[i].kind != FD_DMABUF)
continue;
if (i + 1 >= nfdents || fdents[i + 1].kind != FD_DMABUF)
continue;
size_t s0 = fdents[i].size;
size_t s1 = fdents[i + 1].size;
bool match = false;
for (int c = 0; c < st->ncameras; c++) {
xr_camera_t *cam = &st->cameras[c];
if (cam->nplanes >= 2 && s1 == cam->planesize[1] && s0 >= cam->planesize[0]) {
match = true;
break;
}
}
if (!match)
continue;
xr_group_t *g = NULL;
if (st->ngroups > 0) {
xr_group_t *last = &st->groups[st->ngroups - 1];
if (last->planesize[0] == s0 && last->planesize[1] == s1 &&
!strcmp(last->sensor, current_sensor))
g = last;
}
if (!g) {
if (st->ngroups >= XR_MAX_GROUPS)
break;
g = &st->groups[st->ngroups++];
memset(g, 0, sizeof(*g));
g->planesize[0] = s0;
g->planesize[1] = s1;
snprintf(g->sensor, sizeof(g->sensor), "%s", current_sensor);
}
if (g->nbufs < XR_MAX_RUNBUFS) {
g->buf[g->nbufs].xfd = fdents[i].xfd;
g->buf[g->nbufs].xfd1 = fdents[i + 1].xfd;
g->buf[g->nbufs].size = s0;
g->buf[g->nbufs].size1 = s1;
g->nbufs++;
}
i++; /* consume the plane 1 descriptor */
}
int keep = 0;
for (int i = 0; i < st->ngroups; i++)
if (st->groups[i].nbufs >= 4)
st->groups[keep++] = st->groups[i];
st->ngroups = keep;
/*
* Bind each run to a camera. The sensor marker alone can be wrong: XRService
* sometimes opens another sensor's subdev (e.g. the idle color camera)
* between an upper camera's subdev and its buffers, and two upper cameras
* can resolve to the same sensor name. So a marker match must also fit the
* camera's plane sizes, and each camera takes at most one run.
*/
for (int pass = 0; pass < 2; pass++)
for (int i = 0; i < st->ngroups; i++) {
xr_group_t *g = &st->groups[i];
for (int c = 0; c < st->ncameras && !g->cam; c++) {
xr_camera_t *cam = &st->cameras[c];
if (pass == 0 && (!g->sensor[0] || strcmp(cam->sensor, g->sensor)))
continue;
if (cam->nplanes < 2 || g->planesize[1] != cam->planesize[1] ||
g->planesize[0] < cam->planesize[0])
continue;
bool taken = false;
for (int k = 0; k < st->ngroups; k++)
if (k != i && st->groups[k].cam == cam)
taken = true;
if (!taken)
g->cam = cam;
}
}
}
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn)
{
memset(st, 0, sizeof(*st));
st->pid = find_process(process);
if (!st->pid) {
set_err(err, errn, "%s is not running; start SteamVR on the headset first", process);
return false;
}
topo_load_all();
bool ok = scan_xr_fds(st->pid, err, errn);
if (ok) {
probe_cameras(st);
build_groups(st);
}
topo_free_all();
return ok;
}
void xr_print(const xr_state_t *st, FILE *f)
{
fprintf(f, "XRService pid %d\n", st->pid);
for (int i = 0; i < st->ncameras; i++) {
const xr_camera_t *c = &st->cameras[i];
char fcc[5] = {
(char)(c->pixfmt & 0xff), (char)((c->pixfmt >> 8) & 0xff),
(char)((c->pixfmt >> 16) & 0xff), (char)((c->pixfmt >> 24) & 0xff), 0
};
fprintf(f, " camera %-12s minor %-3u %-16s %ux%u %s pitch %u planes %zu %zu role=%s\n",
c->path, c->minor, c->sensor, c->width, c->height, fcc,
xr_camera_stride(c), c->planesize[0], c->planesize[1], c->role);
}
for (int i = 0; i < st->ngroups; i++) {
const xr_group_t *g = &st->groups[i];
fprintf(f, " queue %d: %d buffers plane0=%zu plane1=%zu fds %d..%d sensor '%s' -> %s\n",
i, g->nbufs, g->planesize[0], g->planesize[1],
g->buf[0].xfd, g->buf[g->nbufs - 1].xfd1, g->sensor,
g->cam ? g->cam->path : "(unbound)");
}
}
+82
View File
@@ -0,0 +1,82 @@
/*
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
*
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
*/
#pragma once
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
#include <sys/types.h>
#include <linux/videodev2.h>
#define XR_MAX_CAMERAS 16
#define XR_MAX_GROUPS 32
#define XR_MAX_RUNBUFS 128
#define XR_SENSOR_LEN 64
typedef struct {
int node; /* N from /dev/videoN */
unsigned minor; /* char device minor, as tracepoints report it */
char path[64];
unsigned width;
unsigned height;
unsigned bytesperline;
unsigned nplanes;
size_t planesize[VIDEO_MAX_PLANES];
uint32_t pixfmt;
char sensor[XR_SENSOR_LEN]; /* media entity name of the sensor */
const char *role;
} xr_camera_t;
typedef struct {
int xfd; /* plane 0 descriptor in XRService */
int xfd1; /* plane 1 descriptor in XRService */
size_t size;
size_t size1;
} xr_bufref_t;
/* One run of buffers XRService allocated for a camera queue, in allocation order. */
typedef struct {
size_t planesize[2];
int nbufs;
xr_bufref_t buf[XR_MAX_RUNBUFS];
char sensor[XR_SENSOR_LEN]; /* from the preceding sensor subdev */
xr_camera_t *cam;
} xr_group_t;
typedef struct {
pid_t pid;
xr_camera_t cameras[XR_MAX_CAMERAS];
int ncameras;
xr_group_t groups[XR_MAX_GROUPS];
int ngroups;
} xr_state_t;
typedef enum {
XR_FMT_GREY8, /* 8-bit mono */
XR_FMT_NV12, /* 8-bit Y plane then interleaved UV, same pitch */
XR_FMT_YUV420_10P /* like NV12, but 10-bit MIPI-packed (4 px in 5 bytes) */
} xr_fmt_t;
/* Where the image really sits in plane 0; V4L2's numbers can be misleading. */
typedef struct {
xr_fmt_t fmt;
unsigned pitch; /* bytes per row */
unsigned rows; /* rows in plane 0: luma, plus chroma for YUV */
unsigned width; /* valid pixels per row */
unsigned height; /* luma rows */
} xr_layout_t;
/* Scan XRService's descriptors and the media graph. Needs root. */
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn);
unsigned xr_camera_stride(const xr_camera_t *c);
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l);
const char *xr_fmt_name(xr_fmt_t f);
void xr_print(const xr_state_t *st, FILE *f);
void xr_slugify(const char *in, char *out, size_t n);
+58
View File
@@ -0,0 +1,58 @@
/*
* fh_hands - the tracked-hands file frame-hands' tracker publishes
* ($XDG_RUNTIME_DIR/frame-hands/hands, directory mode 0700), rewritten in place
* under a sequence lock: read seq, copy, read seq again; use the copy only if
* both reads are the same even number.
*
* Positions are metres in the head frame at capture time, which is OpenVR's HMD
* frame (+x right, +y up, -z forward). Turn them into the room with the HMD pose
* at capture_ns (CLOCK_MONOTONIC). Writers: trackd (fh-tracker), tracker/publish.py.
*/
#pragma once
#include <assert.h>
#include <stdint.h>
#define FH_HANDS_MAGIC "FHHANDS1"
#define FH_HANDS_VERSION 1
#define FH_HANDS_MAX_HANDS 2
#define FH_HANDS_MAX_CAPSULES 64
enum {
FH_HAND_RIGHT = 1u << 0, /* else the left hand */
FH_HAND_STEREO = 1u << 1, /* triangulated from two or more cameras */
};
typedef struct {
uint32_t id; /* stays the same while the hand is tracked */
uint32_t flags; /* FH_HAND_* */
float confidence;
float reserved;
float pts[21][3]; /* MediaPipe hand landmarks */
uint32_t ncapsules; /* this hand's capsules, which follow the */
/* previous hands' in capsules[] */
} fh_hand_t; /* 272 bytes */
typedef struct {
float a[3], b[3]; /* segment ends */
float ra, rb; /* radius at each end */
} fh_capsule_t; /* 32 bytes: the hand's shape, to cut out */
typedef struct {
char magic[8];
uint32_t version;
uint32_t size;
volatile uint64_t seq;
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
uint32_t nhands;
uint32_t ncapsules;
uint8_t reserved[16];
fh_hand_t hands[FH_HANDS_MAX_HANDS];
fh_capsule_t capsules[FH_HANDS_MAX_CAPSULES];
} fh_hands_t;
static_assert(sizeof(fh_hand_t) == 272, "fh_hand_t layout");
static_assert(sizeof(fh_capsule_t) == 32, "fh_capsule_t layout");
static_assert(sizeof(fh_hands_t) == 64 + 2 * 272 + 64 * 32, "fh_hands_t layout");
Binary file not shown.
+79
View File
@@ -0,0 +1,79 @@
7767517
77 90
Input in0 0 1 in0
Convolution convclip_0 1 1 in0 2 0=24 1=3 3=2 15=1 16=1 5=1 6=648 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 4=1 5=1 6=216 7=24 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_10 1 1 3 4 0=16 1=1 5=1 6=384 8=2
Split splitncnn_0 1 2 4 5 6
Convolution convclip_1 1 1 6 7 0=64 1=1 5=1 6=1024 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 3=2 15=1 16=1 5=1 6=576 7=64 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_12 1 1 8 9 0=16 1=1 5=1 6=1024 8=2
Pooling maxpool2d_1 1 1 5 10 1=2 2=2 5=1
BinaryOp add_0 2 1 9 10 11
Split splitncnn_1 1 2 11 12 13
Convolution convclip_2 1 1 13 14 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 4=1 5=1 6=864 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_14 1 1 15 16 0=16 1=1 5=1 6=1536 8=2
BinaryOp add_1 2 1 16 12 17
Convolution convclip_3 1 1 17 18 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_16 1 1 19 20 0=24 1=1 5=1 6=2304 8=2
Split splitncnn_2 1 2 20 21 22
Convolution convclip_4 1 1 22 23 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 4=2 5=1 6=3600 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_18 1 1 24 25 0=24 1=1 5=1 6=3456 8=2
BinaryOp add_2 2 1 25 21 26
Convolution convclip_5 1 1 26 27 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 3=2 15=1 16=1 5=1 6=1296 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_20 1 1 28 29 0=48 1=1 5=1 6=6912 8=2
Split splitncnn_3 1 2 29 30 31
Convolution convclip_6 1 1 31 32 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_22 1 1 33 34 0=48 1=1 5=1 6=13824 8=2
BinaryOp add_3 2 1 34 30 35
Split splitncnn_4 1 2 35 36 37
Convolution convclip_7 1 1 37 38 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_24 1 1 39 40 0=48 1=1 5=1 6=13824 8=2
BinaryOp add_4 2 1 40 36 41
Convolution convclip_8 1 1 41 42 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 4=2 5=1 6=7200 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_26 1 1 43 44 0=64 1=1 5=1 6=18432 8=2
Split splitncnn_5 1 2 44 45 46
Convolution convclip_9 1 1 46 47 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_28 1 1 48 49 0=64 1=1 5=1 6=24576 8=2
BinaryOp add_5 2 1 49 45 50
Split splitncnn_6 1 2 50 51 52
Convolution convclip_10 1 1 52 53 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_30 1 1 54 55 0=64 1=1 5=1 6=24576 8=2
BinaryOp add_6 2 1 55 51 56
Convolution convclip_11 1 1 56 57 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 3=2 4=1 15=2 16=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_32 1 1 58 59 0=112 1=1 5=1 6=43008 8=2
Split splitncnn_7 1 2 59 60 61
Convolution convclip_12 1 1 61 62 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_34 1 1 63 64 0=112 1=1 5=1 6=75264 8=2
BinaryOp add_7 2 1 64 60 65
Split splitncnn_8 1 2 65 66 67
Convolution convclip_13 1 1 67 68 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_36 1 1 69 70 0=112 1=1 5=1 6=75264 8=2
BinaryOp add_8 2 1 70 66 71
Split splitncnn_9 1 2 71 72 73
Convolution convclip_14 1 1 73 74 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
Convolution conv_38 1 1 75 76 0=112 1=1 5=1 6=75264 8=2
BinaryOp add_9 2 1 76 72 77
Convolution convclip_15 1 1 77 78 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 4=1 5=1 6=6048 7=672 8=1 9=3 -23310=2,0.000000e+00,6.000000e+00
Pooling gap_0 1 1 79 80 0=1 4=1
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
Split splitncnn_10 1 4 82 83 84 85 86
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336 8=2
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336 8=2
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 8=2 9=4
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 8=2 9=4
Binary file not shown.
+79
View File
@@ -0,0 +1,79 @@
7767517
77 90
Input in0 0 1 in0
Convolution convclip_0 1 1 in0 2 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=648 9=3
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=216 7=24 9=3
Convolution conv_10 1 1 3 4 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=384
Split splitncnn_0 1 2 4 5 6
Convolution convclip_1 1 1 6 7 0=64 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024 9=3
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=576 7=64 9=3
Convolution conv_12 1 1 8 9 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
Pooling maxpool2d_1 1 1 5 10 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
BinaryOp add_0 2 1 9 10 11 0=0
Split splitncnn_1 1 2 11 12 13
Convolution convclip_2 1 1 13 14 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=864 7=96 9=3
Convolution conv_14 1 1 15 16 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
BinaryOp add_1 2 1 16 12 17 0=0
Convolution convclip_3 1 1 17 18 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400 7=96 9=3
Convolution conv_16 1 1 19 20 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2304
Split splitncnn_2 1 2 20 21 22
Convolution convclip_4 1 1 22 23 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3600 7=144 9=3
Convolution conv_18 1 1 24 25 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456
BinaryOp add_2 2 1 25 21 26 0=0
Convolution convclip_5 1 1 26 27 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=1296 7=144 9=3
Convolution conv_20 1 1 28 29 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=6912
Split splitncnn_3 1 2 29 30 31
Convolution convclip_6 1 1 31 32 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
Convolution conv_22 1 1 33 34 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
BinaryOp add_3 2 1 34 30 35 0=0
Split splitncnn_4 1 2 35 36 37
Convolution convclip_7 1 1 37 38 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
Convolution conv_24 1 1 39 40 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
BinaryOp add_4 2 1 40 36 41 0=0
Convolution convclip_8 1 1 41 42 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=7200 7=288 9=3
Convolution conv_26 1 1 43 44 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=18432
Split splitncnn_5 1 2 44 45 46
Convolution convclip_9 1 1 46 47 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
Convolution conv_28 1 1 48 49 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
BinaryOp add_5 2 1 49 45 50 0=0
Split splitncnn_6 1 2 50 51 52
Convolution convclip_10 1 1 52 53 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
Convolution conv_30 1 1 54 55 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
BinaryOp add_6 2 1 55 51 56 0=0
Convolution convclip_11 1 1 56 57 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=9600 7=384 9=3
Convolution conv_32 1 1 58 59 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=43008
Split splitncnn_7 1 2 59 60 61
Convolution convclip_12 1 1 61 62 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
Convolution conv_34 1 1 63 64 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
BinaryOp add_7 2 1 64 60 65 0=0
Split splitncnn_8 1 2 65 66 67
Convolution convclip_13 1 1 67 68 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
Convolution conv_36 1 1 69 70 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
BinaryOp add_8 2 1 70 66 71 0=0
Split splitncnn_9 1 2 71 72 73
Convolution convclip_14 1 1 73 74 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
Convolution conv_38 1 1 75 76 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
BinaryOp add_9 2 1 76 72 77 0=0
Convolution convclip_15 1 1 77 78 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=6048 7=672 9=3
Pooling gap_0 1 1 79 80 0=1 4=1
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
Split splitncnn_10 1 4 82 83 84 85 86
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 9=4
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 9=4
Binary file not shown.
+151
View File
@@ -0,0 +1,151 @@
7767517
149 177
Input in0 0 1 in0
Convolution padconv_0 1 1 in0 2 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 8=2
PReLU prelu_41 1 1 2 3 0=32
Split splitncnn_0 1 2 3 4 5
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 4=2 5=1 6=800 7=32 8=101
Convolution conv_12 1 1 6 7 0=32 1=1 5=1 6=1024 8=2
BinaryOp add_0 2 1 4 7 8
PReLU prelu_42 1 1 8 9 0=32
Split splitncnn_1 1 2 9 10 11
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 4=2 5=1 6=800 7=32 8=101
Convolution conv_13 1 1 12 13 0=32 1=1 5=1 6=1024 8=2
BinaryOp add_1 2 1 10 13 14
PReLU prelu_43 1 1 14 15 0=32
Split splitncnn_2 1 2 15 16 17
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 4=2 5=1 6=800 7=32 8=101
Convolution conv_14 1 1 18 19 0=32 1=1 5=1 6=1024 8=2
BinaryOp add_2 2 1 16 19 20
PReLU prelu_44 1 1 20 21 0=32
Split splitncnn_3 1 2 21 22 23
Pooling maxpool2d_2 1 1 22 24 1=2 2=2 5=1
Padding Pad_16 1 1 24 25 8=32
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=800 7=32 8=101
Convolution conv_15 1 1 26 27 0=64 1=1 5=1 6=2048 8=2
BinaryOp add_3 2 1 25 27 28
PReLU prelu_45 1 1 28 29 0=64
Split splitncnn_4 1 2 29 30 31
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
Convolution conv_16 1 1 32 33 0=64 1=1 5=1 6=4096 8=2
BinaryOp add_4 2 1 30 33 34
PReLU prelu_46 1 1 34 35 0=64
Split splitncnn_5 1 2 35 36 37
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
Convolution conv_17 1 1 38 39 0=64 1=1 5=1 6=4096 8=2
BinaryOp add_5 2 1 36 39 40
PReLU prelu_47 1 1 40 41 0=64
Split splitncnn_6 1 2 41 42 43
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
Convolution conv_18 1 1 44 45 0=64 1=1 5=1 6=4096 8=2
BinaryOp add_6 2 1 42 45 46
PReLU prelu_48 1 1 46 47 0=64
Split splitncnn_7 1 2 47 48 49
Pooling maxpool2d_3 1 1 48 50 1=2 2=2 5=1
Padding Pad_34 1 1 50 51 8=64
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 3=2 4=1 15=2 16=2 5=1 6=1600 7=64 8=101
Convolution conv_19 1 1 52 53 0=128 1=1 5=1 6=8192 8=2
BinaryOp add_7 2 1 51 53 54
PReLU prelu_49 1 1 54 55 0=128
Split splitncnn_8 1 2 55 56 57
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_20 1 1 58 59 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_8 2 1 56 59 60
PReLU prelu_50 1 1 60 61 0=128
Split splitncnn_9 1 2 61 62 63
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_21 1 1 64 65 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_9 2 1 62 65 66
PReLU prelu_51 1 1 66 67 0=128
Split splitncnn_10 1 2 67 68 69
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_22 1 1 70 71 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_10 2 1 68 71 72
PReLU prelu_52 1 1 72 73 0=128
Split splitncnn_11 1 3 73 74 75 76
Pooling maxpool2d_4 1 1 75 77 1=2 2=2 5=1
Padding Pad_52 1 1 77 78 8=128
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 3=2 4=1 15=2 16=2 5=1 6=3200 7=128 8=101
Convolution conv_23 1 1 79 80 0=256 1=1 5=1 6=32768 8=2
BinaryOp add_11 2 1 78 80 81
PReLU prelu_53 1 1 81 82 0=256
Split splitncnn_12 1 2 82 83 84
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_24 1 1 85 86 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_12 2 1 83 86 87
PReLU prelu_54 1 1 87 88 0=256
Split splitncnn_13 1 2 88 89 90
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_25 1 1 91 92 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_13 2 1 89 92 93
PReLU prelu_55 1 1 93 94 0=256
Split splitncnn_14 1 2 94 95 96
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_26 1 1 97 98 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_14 2 1 95 98 99
PReLU prelu_56 1 1 99 100 0=256
Split splitncnn_15 1 3 100 101 102 103
Pooling maxpool2d_5 1 1 102 104 1=2 2=2 5=1
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 3=2 4=1 15=2 16=2 5=1 6=6400 7=256 8=101
Convolution conv_27 1 1 105 106 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_15 2 1 104 106 107
PReLU prelu_57 1 1 107 108 0=256
Split splitncnn_16 1 2 108 109 110
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_28 1 1 111 112 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_16 2 1 109 112 113
PReLU prelu_58 1 1 113 114 0=256
Split splitncnn_17 1 2 114 115 116
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_29 1 1 117 118 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_17 2 1 115 118 119
PReLU prelu_59 1 1 119 120 0=256
Split splitncnn_18 1 2 120 121 122
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_30 1 1 123 124 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_18 2 1 121 124 125
PReLU prelu_60 1 1 125 126 0=256
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12
Convolution conv_31 1 1 127 128 0=256 1=1 5=1 6=65536 8=2
PReLU prelu_61 1 1 128 129 0=256
BinaryOp add_19 2 1 101 129 130
Split splitncnn_19 1 2 130 131 132
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_32 1 1 133 134 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_20 2 1 131 134 135
PReLU prelu_62 1 1 135 136 0=256
Split splitncnn_20 1 2 136 137 138
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
Convolution conv_33 1 1 139 140 0=256 1=1 5=1 6=65536 8=2
BinaryOp add_21 2 1 137 140 141
PReLU prelu_63 1 1 141 142 0=256
Split splitncnn_21 1 3 142 143 144 145
Convolution conv_34 1 1 145 146 0=108 1=1 5=1 6=27648 8=2
Permute permute_68 1 1 146 147 0=3
Reshape reshape_72 1 1 147 148 0=18 1=864
Convolution conv_35 1 1 144 149 0=6 1=1 5=1 6=1536 8=2
Permute permute_69 1 1 149 150 0=3
Reshape reshape_73 1 1 150 151 0=1 1=864
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24
Convolution conv_36 1 1 152 153 0=128 1=1 5=1 6=32768 8=2
PReLU prelu_64 1 1 153 154 0=128
BinaryOp add_22 2 1 74 154 155
Split splitncnn_22 1 2 155 156 157
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_37 1 1 158 159 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_23 2 1 156 159 160
PReLU prelu_65 1 1 160 161 0=128
Split splitncnn_23 1 2 161 162 163
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
Convolution conv_38 1 1 164 165 0=128 1=1 5=1 6=16384 8=2
BinaryOp add_24 2 1 162 165 166
PReLU prelu_66 1 1 166 167 0=128
Split splitncnn_24 1 2 167 168 169
Convolution conv_39 1 1 169 170 0=36 1=1 5=1 6=4608 8=2
Permute permute_70 1 1 170 171 0=3
Reshape reshape_74 1 1 171 172 0=18 1=1152
Concat cat_0 2 1 172 148 out0
Convolution conv_40 1 1 168 174 0=2 1=1 5=1 6=256 8=2
Permute permute_71 1 1 174 175 0=3
Reshape reshape_75 1 1 175 176 0=1 1=1152
Concat cat_1 2 1 176 151 out1
Binary file not shown.
+151
View File
@@ -0,0 +1,151 @@
7767517
149 177
Input in0 0 1 in0
Convolution padconv_0 1 1 in0 2 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400
PReLU prelu_41 1 1 2 3 0=32
Split splitncnn_0 1 2 3 4 5
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
Convolution conv_12 1 1 6 7 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
BinaryOp add_0 2 1 4 7 8 0=0
PReLU prelu_42 1 1 8 9 0=32
Split splitncnn_1 1 2 9 10 11
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
Convolution conv_13 1 1 12 13 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
BinaryOp add_1 2 1 10 13 14 0=0
PReLU prelu_43 1 1 14 15 0=32
Split splitncnn_2 1 2 15 16 17
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
Convolution conv_14 1 1 18 19 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
BinaryOp add_2 2 1 16 19 20 0=0
PReLU prelu_44 1 1 20 21 0=32
Split splitncnn_3 1 2 21 22 23
Pooling maxpool2d_2 1 1 22 24 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
Padding Pad_16 1 1 24 25 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=32
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=800 7=32
Convolution conv_15 1 1 26 27 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2048
BinaryOp add_3 2 1 25 27 28 0=0
PReLU prelu_45 1 1 28 29 0=64
Split splitncnn_4 1 2 29 30 31
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
Convolution conv_16 1 1 32 33 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
BinaryOp add_4 2 1 30 33 34 0=0
PReLU prelu_46 1 1 34 35 0=64
Split splitncnn_5 1 2 35 36 37
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
Convolution conv_17 1 1 38 39 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
BinaryOp add_5 2 1 36 39 40 0=0
PReLU prelu_47 1 1 40 41 0=64
Split splitncnn_6 1 2 41 42 43
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
Convolution conv_18 1 1 44 45 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
BinaryOp add_6 2 1 42 45 46 0=0
PReLU prelu_48 1 1 46 47 0=64
Split splitncnn_7 1 2 47 48 49
Pooling maxpool2d_3 1 1 48 50 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
Padding Pad_34 1 1 50 51 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=64
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=1600 7=64
Convolution conv_19 1 1 52 53 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=8192
BinaryOp add_7 2 1 51 53 54 0=0
PReLU prelu_49 1 1 54 55 0=128
Split splitncnn_8 1 2 55 56 57
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_20 1 1 58 59 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_8 2 1 56 59 60 0=0
PReLU prelu_50 1 1 60 61 0=128
Split splitncnn_9 1 2 61 62 63
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_21 1 1 64 65 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_9 2 1 62 65 66 0=0
PReLU prelu_51 1 1 66 67 0=128
Split splitncnn_10 1 2 67 68 69
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_22 1 1 70 71 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_10 2 1 68 71 72 0=0
PReLU prelu_52 1 1 72 73 0=128
Split splitncnn_11 1 3 73 74 75 76
Pooling maxpool2d_4 1 1 75 77 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
Padding Pad_52 1 1 77 78 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=128
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=3200 7=128
Convolution conv_23 1 1 79 80 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
BinaryOp add_11 2 1 78 80 81 0=0
PReLU prelu_53 1 1 81 82 0=256
Split splitncnn_12 1 2 82 83 84
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_24 1 1 85 86 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_12 2 1 83 86 87 0=0
PReLU prelu_54 1 1 87 88 0=256
Split splitncnn_13 1 2 88 89 90
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_25 1 1 91 92 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_13 2 1 89 92 93 0=0
PReLU prelu_55 1 1 93 94 0=256
Split splitncnn_14 1 2 94 95 96
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_26 1 1 97 98 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_14 2 1 95 98 99 0=0
PReLU prelu_56 1 1 99 100 0=256
Split splitncnn_15 1 3 100 101 102 103
Pooling maxpool2d_5 1 1 102 104 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=6400 7=256
Convolution conv_27 1 1 105 106 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_15 2 1 104 106 107 0=0
PReLU prelu_57 1 1 107 108 0=256
Split splitncnn_16 1 2 108 109 110
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_28 1 1 111 112 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_16 2 1 109 112 113 0=0
PReLU prelu_58 1 1 113 114 0=256
Split splitncnn_17 1 2 114 115 116
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_29 1 1 117 118 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_17 2 1 115 118 119 0=0
PReLU prelu_59 1 1 119 120 0=256
Split splitncnn_18 1 2 120 121 122
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_30 1 1 123 124 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_18 2 1 121 124 125 0=0
PReLU prelu_60 1 1 125 126 0=256
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12 6=0
Convolution conv_31 1 1 127 128 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
PReLU prelu_61 1 1 128 129 0=256
BinaryOp add_19 2 1 101 129 130 0=0
Split splitncnn_19 1 2 130 131 132
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_32 1 1 133 134 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_20 2 1 131 134 135 0=0
PReLU prelu_62 1 1 135 136 0=256
Split splitncnn_20 1 2 136 137 138
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
Convolution conv_33 1 1 139 140 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
BinaryOp add_21 2 1 137 140 141 0=0
PReLU prelu_63 1 1 141 142 0=256
Split splitncnn_21 1 3 142 143 144 145
Convolution conv_34 1 1 145 146 0=108 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=27648
Permute permute_68 1 1 146 147 0=3
Reshape reshape_72 1 1 147 148 0=18 1=864
Convolution conv_35 1 1 144 149 0=6 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
Permute permute_69 1 1 149 150 0=3
Reshape reshape_73 1 1 150 151 0=1 1=864
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24 6=0
Convolution conv_36 1 1 152 153 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
PReLU prelu_64 1 1 153 154 0=128
BinaryOp add_22 2 1 74 154 155 0=0
Split splitncnn_22 1 2 155 156 157
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_37 1 1 158 159 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_23 2 1 156 159 160 0=0
PReLU prelu_65 1 1 160 161 0=128
Split splitncnn_23 1 2 161 162 163
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
Convolution conv_38 1 1 164 165 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
BinaryOp add_24 2 1 162 165 166 0=0
PReLU prelu_66 1 1 166 167 0=128
Split splitncnn_24 1 2 167 168 169
Convolution conv_39 1 1 169 170 0=36 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4608
Permute permute_70 1 1 170 171 0=3
Reshape reshape_74 1 1 171 172 0=18 1=1152
Concat cat_0 2 1 172 148 out0 0=0
Convolution conv_40 1 1 168 174 0=2 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=256
Permute permute_71 1 1 174 175 0=3
Reshape reshape_75 1 1 175 176 0=1 1=1152
Concat cat_1 2 1 176 151 out1 0=0
+129
View File
@@ -0,0 +1,129 @@
"""Check that the side cameras' images carry the right names (slam_left vs slam_right).
usage: python tools/check_sides.py REC_DIR [--sets N]
python tools/check_sides.py --ring [--sets N] (live, from fh-camd's ring)
With --ring it exits 0 when the names are right, 3 when they're swapped (run fh-tracker
with --swap-sides), and 2 when it can't tell (too little texture in view, or the headset
isn't worn).
fh-camd tells the two side cameras' buffers apart by the order XRService allocated them,
and after some XRService restarts that order puts each camera's images under the other's
name. The tracker then sees every hand in one camera only, at the wrong depth. This
matches features between the two images and measures how close each pair's rays pass
with the factory calibration, once as named and once swapped: true matches meet in
front of both cameras only under the right naming.
"""
import argparse
import os
import sys
import cv2
import numpy as np
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
from tools.show_set import index, read_set # noqa: E402
from tracker import calib # noqa: E402
PIPES = {'msm_vfe3_video0': 'slam_left', 'msm_vfe4_video0': 'slam_right'} # as fh-tracker maps them
def load_cams():
root = os.environ.get('FRAME_JOB_DEVICE_ROOT', '') # frame-job's copy of /persist off the Frame
return calib.load(root + calib.XRSERVICE_JSON, root + calib.DEVICE_JSON)
def matches(a, b):
"""Pixel pairs (N,2), (N,2) of ORB matches between two grey images."""
clahe = cv2.createCLAHE(2.0, (8, 8))
orb = cv2.ORB_create(3000)
ka, da = orb.detectAndCompute(clahe.apply(a), None)
kb, db = orb.detectAndCompute(clahe.apply(b), None)
if da is None or db is None:
return np.zeros((0, 2)), np.zeros((0, 2))
pairs = cv2.BFMatcher(cv2.NORM_HAMMING).knnMatch(da, db, k=2)
good = [p[0] for p in pairs if len(p) == 2 and p[0].distance < 0.75 * p[1].distance]
return (np.array([ka[m.queryIdx].pt for m in good]).reshape(-1, 2),
np.array([kb[m.trainIdx].pt for m in good]).reshape(-1, 2))
def meet(cam_a, cam_b, ua, ub):
"""Per match: closest distance between the two rays (m), and whether they meet in front of both."""
ra, rb = cam_a.rays(ua), cam_b.rays(ub)
w = cam_b.origin - cam_a.origin
n = np.cross(ra, rb)
nn = np.linalg.norm(n, axis=1)
dist = np.abs(w @ n.T) / np.maximum(nn, 1e-12)
# ray parameters at the closest points
ta = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, rb.shape), rb), n) / np.maximum(nn ** 2, 1e-12)
tb = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, ra.shape), ra), n) / np.maximum(nn ** 2, 1e-12)
return dist, (ta > 0.05) & (tb > 0.05)
def score(cam_a, cam_b, ua, ub):
"""Share of matches whose rays meet within 1 cm, in front of both cameras."""
if len(ua) == 0:
return 0.0
d, front = meet(cam_a, cam_b, ua, ub)
return float(np.mean((d < 0.01) & front))
def recorded_pairs(rec, count):
"""(label, slam_left image, slam_right image) from sets spread across a recording."""
path = os.path.join(rec, 'sets.bin')
offs = index(path)
for n in np.linspace(0, len(offs) - 1, count).astype(int):
images = read_set(path, offs[n])
if 'slam_left' in images and 'slam_right' in images:
yield 'set %5d' % n, images['slam_left'][0], images['slam_right'][0]
def live_pairs(count):
"""(label, slam_left image, slam_right image) from fh-camd's ring, half a second apart."""
import time
from tracker.ring import Ring
ring = Ring()
if not ring.alive():
sys.exit('fh-camd isn\'t running (no heartbeat)')
cams = {}
for c in ring.cams:
name = PIPES.get(open('/sys/class/video4linux/video%d/name' % c.node).read().strip())
if name and not c.name.endswith('-dark'):
cams[name] = c
for k in range(count):
a, b = ring.read(cams['slam_left']), ring.read(cams['slam_right'])
if a is not None and b is not None:
yield 'frame %2d' % k, a.image, b.image
time.sleep(0.5)
def main():
ap = argparse.ArgumentParser()
ap.add_argument('rec', nargs='?')
ap.add_argument('--ring', action='store_true', help='check the live cameras instead of a recording')
ap.add_argument('--sets', type=int, default=8, help='how many sets or live frames to check')
a = ap.parse_args()
if not a.ring and not a.rec:
ap.error('give a recording or --ring')
cams = load_cams()
left, right = cams['slam_left'], cams['slam_right']
named = swapped = 0.0
n = total_matches = 0
for label, img_l, img_r in (live_pairs(a.sets) if a.ring else recorded_pairs(a.rec, a.sets)):
ua, ub = matches(img_l, img_r)
s_named = score(left, right, ua, ub) # slam_left's image seen by the left camera
s_swapped = score(right, left, ua, ub) # ... by the right camera
named, swapped, n, total_matches = named + s_named, swapped + s_swapped, n + 1, total_matches + len(ua)
print('%s: %4d matches, meeting as named %3.0f%%, swapped %3.0f%%' %
(label, len(ua), 100 * s_named, 100 * s_swapped))
if n == 0 or total_matches < 100 or abs(named - swapped) / n < 0.2:
print('side cameras: can\'t tell (%d matches)' % total_matches)
sys.exit(2)
print('side cameras: %s (named %.2f, swapped %.2f)' %
('as named' if named > swapped else 'SWAPPED', named / n, swapped / n))
sys.exit(0 if named > swapped else 3)
if __name__ == '__main__':
main()
+81
View File
@@ -0,0 +1,81 @@
"""Convert the OpenCV Zoo ONNX ports of MediaPipe's hand models to ncnn.
The ONNX files are Apache-2.0 ports of MediaPipe's palm detector and hand
landmark models (huggingface.co/opencv/palm_detection_mediapipe and
huggingface.co/opencv/handpose_estimation_mediapipe). pnnx does the
conversion; two fix-ups follow:
- The palm detector widens channels with ONNX Pad on the channel axis. pnnx
emits an ncnn layer called "Pad", which ncnn doesn't have, so rewrite those
as ncnn Padding with the channel-end amount (param 8 = behind).
- Both models take NHWC input and start with a Permute to NCHW. Drop it, so we
can hand ncnn planar CHW Mats straight from the preprocessing step.
usage: python convert_models.py (writes models/ncnn/{palm,hand}.ncnn.{param,bin})
"""
import os
import re
import shutil
import subprocess
import sys
import tempfile
HERE = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.join(HERE, '..')
PNNX = os.path.join(sys.prefix, 'lib', 'python%d.%d' % sys.version_info[:2], 'site-packages', 'pnnx', 'pnnx')
MODELS = [('palm', 'palm_detection_mediapipe_2023feb', 192),
('hand', 'handpose_estimation_mediapipe_2023feb', 224)]
def patch(param_text, pnnx_param_text):
lines = param_text.splitlines()
assert lines[0] == '7767517'
nlayers, nblobs = map(int, lines[1].split())
body = lines[2:]
# Channel pads: amounts come from the pnnx graph, which keeps the pads tuple.
pads = dict(re.findall(r'^Pad\s+(\S+)\s.*pads=\(0,0,0,0,0,(\d+),0,0\)', pnnx_param_text, re.M))
for i, line in enumerate(body):
f = line.split()
if f[0] == 'Pad':
amount = pads[f[1]]
body[i] = 'Padding %s %s %s %s %s 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=%s' % (
f[1], f[2], f[3], f[4], f[5], amount)
# Input permute: feed its consumers from in0 instead.
perm = next(i for i, line in enumerate(body) if line.split()[0] == 'Permute')
f = body[perm].split()
assert f[4] == 'in0' and f[6] == '0=4', body[perm]
blob = f[5]
del body[perm]
for i, line in enumerate(body):
f = line.split()
if f[0] == 'Input':
continue
nin, nout = int(f[2]), int(f[3])
ins = ['in0' if b == blob else b for b in f[4:4 + nin]]
body[i] = ' '.join(f[:4] + ins + f[4 + nin:])
return '\n'.join(['7767517', '%d %d' % (nlayers - 1, nblobs - 1)] + body) + '\n'
def main():
out = os.path.join(ROOT, 'models', 'ncnn')
os.makedirs(out, exist_ok=True)
for short, name, size in MODELS:
src = os.path.join(ROOT, 'models', 'onnx', name + '.onnx')
with tempfile.TemporaryDirectory() as tmp:
shutil.copy(src, tmp)
subprocess.run([PNNX, name + '.onnx', 'inputshape=[1,%d,%d,3]' % (size, size), 'fp16=1'],
cwd=tmp, check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
with open(os.path.join(tmp, name + '.ncnn.param')) as f:
param = f.read()
with open(os.path.join(tmp, name + '.pnnx.param')) as f:
pparam = f.read()
with open(os.path.join(out, short + '.ncnn.param'), 'w') as f:
f.write(patch(param, pparam))
shutil.copy(os.path.join(tmp, name + '.ncnn.bin'), os.path.join(out, short + '.ncnn.bin'))
print('wrote', short)
if __name__ == '__main__':
main()
+79
View File
@@ -0,0 +1,79 @@
"""Calibration crops for quantizing the models to int8 (ncnn2table).
Takes the frames of an fh-camprobe capture, cuts the crops the tracker would feed
the models (CLAHE-equalized, as tracker/models.py prepares them), and writes them
as PNGs plus a list per model:
palm: every search tile of every frame (a sample of them)
hand: the hands found in them, each also shifted, scaled and turned a little
usage: python tools/make_int8_calib.py CAPTURE_DIR OUT_DIR
"""
import glob
import os
import random
import sys
import cv2
import numpy as np
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..', 'tracker'))
import calib # noqa: E402
import hands # noqa: E402
import models # noqa: E402
NODES = {'video9': 'slam_left', 'video13': 'slam_right', 'video6': 'upper_left', 'video7': 'upper_right'}
def save(path, patch):
cv2.imwrite(path, cv2.cvtColor(patch, cv2.COLOR_GRAY2BGR))
def main():
cap, out = sys.argv[1], sys.argv[2]
random.seed(1)
cams = calib.load()
eng = models.Engine()
palm, hm = models.PalmDetector(), models.HandLandmarker()
tracker = hands.Tracker(cams, eng)
for d in ('palm', 'hand'):
os.makedirs(os.path.join(out, d), exist_ok=True)
palm_files, hand_files = [], []
for node, name in NODES.items():
tiles = [t for t in tracker.tiles if t.cam.name == name]
for f in sorted(glob.glob(os.path.join(cap, '*_%s.pgm' % node))):
g = cv2.imread(f, cv2.IMREAD_GRAYSCALE)
base = os.path.basename(f)[:-4]
prep = [palm.prepare(g, t.center, t.size, t.rotation) for t in tiles]
outs = eng.run([('palm', p) for p, _ in prep])
rois = []
for k, ((p, ctx), o) in enumerate(zip(prep, outs)):
if random.random() < 0.3:
path = os.path.join(out, 'palm', '%s_t%02d.png' % (base, k))
save(path, p)
palm_files.append(path)
for d in palm.decode(o, ctx):
r = d.roi()
if all(np.linalg.norm(r[0] - q[0]) > 0.5 * r[1] for q in rois):
rois.append(r)
for j, r in enumerate(rois):
p, ctx = hm.prepare(g, r)
if hm.decode(eng.run([('hand', p)])[0], ctx).presence < 0.5:
continue
for v in range(5):
c, s, rot = r
if v:
c = np.asarray(c) + np.random.uniform(-0.08, 0.08, 2) * s
s = s * np.random.uniform(0.87, 1.15)
rot = rot + np.radians(np.random.uniform(-20, 20))
p, _ = hm.prepare(g, (c, s, rot))
path = os.path.join(out, 'hand', '%s_h%d_%d.png' % (base, j, v))
save(path, p)
hand_files.append(path)
for name, files in (('palm', palm_files), ('hand', hand_files)):
with open(os.path.join(out, name + '.txt'), 'w') as fh:
fh.write('\n'.join(files) + '\n')
print(name, len(files), 'crops')
if __name__ == '__main__':
main()
+90
View File
@@ -0,0 +1,90 @@
"""Compare trackd's C++ model code with tracker/models.py on recorded frames.
For each frame, finds a crop with a palm (Python side), runs trackd/nettest on the same
crop, and reports how far apart the palms, ROIs and landmarks are.
usage: python tools/nettest_compare.py CAPTURE_DIR [--int8]
"""
import glob
import os
import subprocess
import sys
import cv2
import numpy as np
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, os.path.join(HERE, '..', 'tracker'))
import calib # noqa: E402
import hands # noqa: E402
import models # noqa: E402
NODES = {'video9': 'slam_left', 'video13': 'slam_right', 'video6': 'upper_left', 'video7': 'upper_right'}
NETTEST = os.path.join(HERE, '..', 'trackd', 'nettest')
MODELS = os.path.join(HERE, '..', 'models', 'ncnn')
def cpp(frame, tile, int8):
out = subprocess.run([NETTEST, MODELS, frame, str(tile.center[0]), str(tile.center[1]), str(tile.size),
str(tile.rotation)] + (['--int8'] if int8 else []), capture_output=True, text=True).stdout
palms, hands_, ms = [], [], {}
for line in out.splitlines():
f = line.split()
if f[0] == 'palm':
palms.append((float(f[1]), np.array([float(f[6]), float(f[7])]), float(f[8]), float(f[9])))
elif f[0] == 'pts':
hands_.append(np.array(f[1:], float).reshape(21, 2))
elif f[0] == 'hand_ms':
ms.setdefault('hand', []).append(float(f[1]))
hands_presence = float(f[3])
ms.setdefault('presence', []).append(hands_presence)
elif f[0] == 'palm_ms':
ms.setdefault('palm', []).append(float(f[1]))
return palms, hands_, ms
def main():
cap = sys.argv[1]
int8 = '--int8' in sys.argv
cams = calib.load()
eng = models.Engine()
palm, hm = models.PalmDetector(), models.HandLandmarker()
tracker = hands.Tracker(cams, eng)
d_roi, d_pts, n_py, n_cpp, pres = [], [], 0, 0, []
times = {'palm': [], 'hand': []}
for node, name in NODES.items():
tiles = [t for t in tracker.tiles if t.cam.name == name]
for f in sorted(glob.glob(os.path.join(cap, '*_%s.pgm' % node)))[::4]:
g = cv2.imread(f, cv2.IMREAD_GRAYSCALE)
for t in tiles:
p, ctx = palm.prepare(g, t.center, t.size, t.rotation)
dets = palm.decode(eng.run([('palm', p)])[0], ctx)
if not dets:
continue
cp, ch, ms = cpp(f, t, int8)
for k in times:
times[k] += ms.get(k, [])
n_py += len(dets)
n_cpp += len(cp)
for d in dets:
r = d.roi()
if not cp:
continue
j = int(np.argmin([np.linalg.norm(c[1] - r[0]) for c in cp]))
d_roi.append(np.linalg.norm(cp[j][1] - r[0]) / r[1])
pp, pctx = hm.prepare(g, r)
lm = hm.decode(eng.run([('hand', pp)])[0], pctx)
if lm.presence >= 0.5 and j < len(ch):
d_pts.append(np.median(np.linalg.norm(ch[j] - lm.pts, axis=1)) / r[1])
pres.append(abs(ms['presence'][j] - lm.presence))
break # one palm tile per frame is enough
print('palms: python %d, c++ %d' % (n_py, n_cpp))
print('roi centre offset: median %.3f of the roi size (max %.3f)' % (np.median(d_roi), np.max(d_roi)))
if d_pts:
print('landmarks: median %.4f of the roi size (max %.4f), presence diff median %.3f' % (
np.median(d_pts), np.max(d_pts), np.median(pres)))
print('c++ time: palm %.1f ms, hand %.1f ms (median, one thread)' % (np.median(times['palm']), np.median(times['hand'])))
if __name__ == '__main__':
main()
+120
View File
@@ -0,0 +1,120 @@
"""Draw frame sets from a recording (fh-tracker --record) with what the tracker saw.
usage: python tools/show_set.py REC_DIR SET [SET...] [--timeline TL] [--out DIR]
SET is a set index (fh-replay's timeline gives them). With --timeline (fh-replay
--timeline), each camera shows the tracker's views at that set: the crop for the next
frame, labelled with the hand and presence. Recordings made with fh-camd --with-dark get
a second row: each camera's latest dark frame (<name>_dk), stretched to be visible and
labelled with its mean brightness. Writes OUT/set_<n>.jpg (default /tmp).
"""
import argparse
import os
import struct
import cv2
import numpy as np
HDR = struct.Struct('<8sII')
CAM = struct.Struct('<16sIIQQ')
ORDER = ['slam_left', 'slam_right', 'upper_left', 'upper_right']
def index(path):
"""Byte offset of every set in sets.bin."""
offs, size = [], os.path.getsize(path)
with open(path, 'rb') as f:
off = 0
while off + HDR.size <= size:
f.seek(off)
magic, n, nbytes = HDR.unpack(f.read(HDR.size))
if magic[:7] != b'FHSET01' or off + nbytes > size:
break
offs.append(off)
off += nbytes
return offs
def read_set(path, off):
with open(path, 'rb') as f:
f.seek(off)
_, n, _ = HDR.unpack(f.read(HDR.size))
cams = [CAM.unpack(f.read(CAM.size)) for _ in range(n)]
out = {}
for name, w, h, cap, dq in cams:
px = np.frombuffer(f.read(w * h), np.uint8).reshape(h, w)
out[name.rstrip(b'\0').decode()] = (px, cap)
return out
def views_at(timeline, n):
out = []
for line in open(timeline):
f = line.split()
if len(f) > 1 and f[1] == 'view' and int(f[-1]) == n:
out.append({'hand': int(f[2]), 'cam': f[3], 'presence': float(f[5]),
'c': (float(f[7]), float(f[8])), 'size': float(f[9]), 'rot': float(f[10])})
return out
def dark_tile(frame, shape, name):
"""A dark frame, stretched from its 1st to 99.5th percentile; black if there's none."""
h, w = shape
if frame is None:
return np.zeros((h, w, 3), np.uint8)
px = frame[0]
lo, hi = np.percentile(px, (1, 99.5))
gain = 255 / max(hi - lo, 1)
img = np.clip((px.astype(np.float32) - lo) * gain, 0, 255).astype(np.uint8)
img = cv2.cvtColor(cv2.resize(img, (w, h)), cv2.COLOR_GRAY2BGR)
cv2.putText(img, '%s_dk mean %.1f, x%.0f' % (name, px.mean(), gain), (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0,
(255, 255, 0), 2)
return img
def draw(images, views):
tiles, dark = [], []
for name in ORDER:
if name not in images:
continue
px = images[name][0]
clahe = cv2.createCLAHE(2.0, (8, 8)).apply(px)
img = cv2.cvtColor(clahe, cv2.COLOR_GRAY2BGR)
for v in views:
if v['cam'] != name:
continue
c, s, r = v['c'], v['size'], v['rot']
box = cv2.boxPoints(((c[0], c[1]), (s, s), np.degrees(r)))
col = (0, 255, 0) if v['presence'] >= 0.5 else (0, 0, 255)
cv2.polylines(img, [box.astype(np.int32)], True, col, 2)
cv2.putText(img, 'h%d %.2f' % (v['hand'], v['presence']), (int(c[0] - s / 2), int(c[1] - s / 2) - 6),
cv2.FONT_HERSHEY_SIMPLEX, 0.8, col, 2)
cv2.putText(img, name, (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0, (255, 255, 0), 2)
scale = 512 / img.shape[0]
tiles.append(cv2.resize(img, (int(img.shape[1] * scale), 512)))
dark.append(dark_tile(images.get(name + '_dk'), tiles[-1].shape[:2], name))
out = np.hstack(tiles)
if any(k.endswith('_dk') for k in images):
out = np.vstack([out, np.hstack(dark)])
return out
def main():
ap = argparse.ArgumentParser()
ap.add_argument('rec')
ap.add_argument('sets', type=int, nargs='+')
ap.add_argument('--timeline')
ap.add_argument('--out', default='/tmp')
a = ap.parse_args()
path = os.path.join(a.rec, 'sets.bin')
offs = index(path)
for n in a.sets:
images = read_set(path, offs[n])
views = views_at(a.timeline, n) if a.timeline else []
out = os.path.join(a.out, 'set_%05d.jpg' % n)
cv2.imwrite(out, draw(images, views), [cv2.IMWRITE_JPEG_QUALITY, 85])
print(out)
if __name__ == '__main__':
main()
+25
View File
@@ -0,0 +1,25 @@
# fh-tracker, the C++ tracker. Needs ncnn built into ../vendor/ncnn/build/install
# (see README.md) and jsoncpp.
NCNN ?= ../vendor/ncnn/build/install
CXXFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter
CXXFLAGS += -std=c++17 -fopenmp -I$(NCNN)/include/ncnn
LDLIBS = $(NCNN)/lib/libncnn.a -ljsoncpp -fopenmp -lpthread
SRC = main.cpp calib.cpp nets.cpp tracker.cpp io.cpp record.cpp
HDR = calib.h geom.h nets.h tracker.h io.h record.h ../camd/fhring.h ../include/fh_hands.h
all: fh-tracker fh-replay
fh-tracker: $(SRC) $(HDR)
$(CXX) $(CXXFLAGS) -o $@ $(SRC) $(LDLIBS)
fh-replay: replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp $(HDR)
$(CXX) $(CXXFLAGS) -o $@ replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp $(LDLIBS)
clean:
rm -f fh-tracker fh-replay nettest
.PHONY: all clean
nettest: nettest.cpp nets.cpp nets.h geom.h
$(CXX) $(CXXFLAGS) -o $@ nettest.cpp nets.cpp $(LDLIBS)
+66
View File
@@ -0,0 +1,66 @@
# trackd
`fh-tracker` is the C++ version of `tracker/live.py`. It uses the same scheduling and the same models. The models run on a few threads, with no Python in the loop. It reads `fh-camd`'s ring and publishes hands to `$XDG_RUNTIME_DIR/frame-hands/hands` (`include/fh_hands.h`), where Frametop's ft-screens picks them up for the hand cutouts.
```
sudo camd/fh-camd &
trackd/fh-tracker # status every 5 s; Ctrl+C to stop
trackd/fh-tracker --int8 # the 8-bit models (models/ncnn/*-int8.ncnn.*)
```
Options:
- `--threads N`: model threads, pinned to the `--cpus` list. Default 3.
- `--cpus LIST`: CPUs for the model threads and the main loop. Default `2,3,4`.
- `--contrast MODE` or `PALM/HAND`: how crops are equalized before the models see them: `clahe[:CLIP]`, `none`, or `stretch` (1st-99th percentile). Default `clahe:2/none`. In the dim recording, CLAHE let the palm search find about 10% more hands, but it made the landmarks jitter more (published median 6.9 mm, against 6.0 mm with plain landmark crops).
- `--swap-sides`: swap the two side cameras (`slam_left`, `slam_right`). fh-camd tells their buffers apart by XRService's allocation order, and after some XRService restarts that order is reversed. Then every hand is seen by one camera only, at the wrong depth, and the hand holes land beside the hands. With the headset on and looking at a room with some texture, `tools/check_sides.py --ring` tells whether the names are right (exit 0), swapped (exit 3), or it can't tell (exit 2).
- `--seconds N`: stop after N seconds.
- `--status S`: how often to print status, in seconds.
- `--models DIR`: where the models are.
- `--nice N`: niceness. Default 5, so the VR stack wins contested CPUs.
- `--no-publish`: don't write the hands file.
- `--record DIR`, `--record-for S`: save every frame set for S seconds (default 120) to `DIR/sets.bin`. That's about 80 MB/s. Sending the tracker SIGUSR1 (`pkill -USR1 -x fh-tracker`) starts a recording in `captures/rec-<time>` without a restart.
- `--record-only`: record without tracking or publishing, so it can run beside the live tracker. Give it `--record DIR`; SIGUSR1 would reach both trackers. With `fh-camd --with-dark`, recordings also hold each camera's newest dark frame as `<name>_dk`, which doubles the rate.
The status line also says how often a hand was on each side (by where the wrist is), and why views and hands came and went: views lost (the landmark model stopped seeing the hand), handoff misses (a crop projected from the hand's 3D position found nothing), duplicates, splits (two views disagreed in 3D), and hands created, merged and forgotten.
## Beyond the Python tracker
fh-tracker started as a port of `tracker/hands.py`. Replaying recordings (below) showed where it went wrong, and it now differs in these ways:
- **Pairing views across cameras.** The side cameras sit side by side, so two hands next to each other at the same height fall on the same epipolar lines, and rays to two different hands can nearly meet close to the cameras. That made phantom hands 12-15 cm in front of the eyes, which tore holes through the screens. Each step now scores every way of pairing the views in two cameras and keeps the best. A pair scores well when its rays meet, when each view's apparent size matches the triangulated distance, and when the model calls both the same hand. The size check uses a fixed prior: with the model's average hand, clean pairs measure 0.71-1.51 times the one-view distance, and mismatched pairs mostly far less. It doesn't use the learned hand size, which bad pairs had corrupted.
- **One view.** From how big the hand looks, its distance is off by 10-30% and wanders about 10% between frames. So a hand that drops to one camera keeps its last distance and drifts toward the one-view guess by 10% a frame.
- **Smoothing.** The published landmarks go through a One Euro filter: it smooths hard while the hand is still (tracking noise is several mm per frame) and hardly at all while it moves fast. The palm speed that sets the update rate (15 or 30 Hz) is the filtered one; the raw speed read about 0.25 m/s from noise alone.
- **Capsules.** Forearms follow the hand's own axis, and nothing within 12 cm in front of the eyes is published.
## Replay
`fh-replay DIR` runs a recording through the tracker with the live scheduling and reports how well it kept the hands: hands per set, left and right coverage, track lengths, and the same reasons as the status line.
```
trackd/fh-replay captures/rec-20260929-120000 --oracle 10 --timeline /tmp/tl.txt
```
- `--oracle N`: every N-th set, also search every tile of every camera, and report how often the tracker had the hands that full search could find.
- `--slow F`: live, the tracker skips sets that arrive while it's busy. Replay counts each step's time times F as busy (default 1; the headset is busier live).
- `--timeline FILE`: a line per processed set and hand.
## Build
ncnn is built from source into `vendor/ncnn/build/install`. The build also provides the `ncnn2table` and `ncnn2int8` quantization tools.
```
git clone --depth 1 --branch 20260526 https://github.com/Tencent/ncnn.git vendor/ncnn
cd vendor/ncnn && mkdir build && cd build
cmake -G Ninja -DCMAKE_BUILD_TYPE=Release -DNCNN_VULKAN=OFF -DNCNN_BUILD_TOOLS=ON -DNCNN_SIMPLEOCV=ON \
-DNCNN_BUILD_EXAMPLES=OFF -DNCNN_BUILD_TESTS=OFF -DCMAKE_INSTALL_PREFIX=$PWD/install ..
nice ninja install
cd ../../../trackd && make # fh-tracker and fh-replay; `make nettest` for the model check
```
It needs jsoncpp for the calibration files (the host has it).
## Checks
- `tools/nettest_compare.py CAPTURE` runs the C++ model code (`nettest`) and `tracker/models.py` on the same crops from a recording, and compares the palms, ROIs and landmarks. Add `--int8` to check the 8-bit models.
- `tools/make_int8_calib.py CAPTURE OUT` cuts the calibration crops that `ncnn2table` needs to quantize the models.
+166
View File
@@ -0,0 +1,166 @@
#include <cstdlib>
#include "calib.h"
#include <json/json.h>
#include <algorithm>
#include <fstream>
namespace {
double theta_d(const Camera &c, double t) {
const double t2 = t * t;
return t * (1 + t2 * (c.k[0] + t2 * (c.k[1] + t2 * (c.k[2] + t2 * c.k[3]))));
}
// 4x4 transform (row-major) from a {plus_x, plus_z, position} pose.
void pose(const Json::Value &d, double scale, double T[4][4]) {
V3 x{d["plus_x"][0].asDouble(), d["plus_x"][1].asDouble(), d["plus_x"][2].asDouble()};
V3 z{d["plus_z"][0].asDouble(), d["plus_z"][1].asDouble(), d["plus_z"][2].asDouble()};
V3 y{z[1] * x[2] - z[2] * x[1], z[2] * x[0] - z[0] * x[2], z[0] * x[1] - z[1] * x[0]};
for (int i = 0; i < 3; ++i) {
T[i][0] = x[i], T[i][1] = y[i], T[i][2] = z[i];
T[i][3] = d["position"][i].asDouble() * scale;
T[3][i] = 0;
}
T[3][3] = 1;
}
void mul(const double A[4][4], const double B[4][4], double C[4][4]) {
for (int i = 0; i < 4; ++i)
for (int j = 0; j < 4; ++j) {
C[i][j] = 0;
for (int k = 0; k < 4; ++k) C[i][j] += A[i][k] * B[k][j];
}
}
void invert_rigid(const double A[4][4], double B[4][4]) {
for (int i = 0; i < 3; ++i)
for (int j = 0; j < 3; ++j) B[i][j] = A[j][i];
for (int i = 0; i < 3; ++i) B[i][3] = -(B[i][0] * A[0][3] + B[i][1] * A[1][3] + B[i][2] * A[2][3]);
B[3][0] = B[3][1] = B[3][2] = 0, B[3][3] = 1;
}
bool read_json(const char *path, Json::Value &v, std::string &err) {
std::ifstream f(path);
Json::CharReaderBuilder b;
std::string e;
if (!f || !Json::parseFromStream(b, f, &v, &e)) {
err = std::string(path) + ": " + (f ? e : "can't open");
return false;
}
return true;
}
} // namespace
V2 Camera::project_cam(V3 p) const {
const double r = std::hypot(p[0], p[1]);
const double s = r > 1e-12 ? theta_d(*this, std::atan2(r, p[2])) / r : 0;
return {fx * p[0] * s + cx, fy * p[1] * s + cy};
}
V3 Camera::unproject(V2 uv) const {
const double mx = (uv[0] - cx) / fx, my = (uv[1] - cy) / fy, td = std::hypot(mx, my);
double t = td;
for (int i = 0; i < 8; ++i) { // Newton on theta_d(t) = td
const double t2 = t * t;
const double df = 1 + t2 * (3 * k[0] + t2 * (5 * k[1] + t2 * (7 * k[2] + t2 * 9 * k[3])));
t = std::clamp(t - (theta_d(*this, t) - td) / df, 0.0, M_PI);
}
const double s = td > 1e-12 ? std::sin(t) / td : 1;
return {mx * s, my * s, std::cos(t)};
}
V3 Camera::ray(V2 uv) const {
const V3 c = unproject(uv);
return {R[0][0] * c[0] + R[0][1] * c[1] + R[0][2] * c[2], R[1][0] * c[0] + R[1][1] * c[1] + R[1][2] * c[2],
R[2][0] * c[0] + R[2][1] * c[1] + R[2][2] * c[2]};
}
V2 Camera::project(V3 head, double *depth) const {
const V3 d = head - origin;
const V3 c{R[0][0] * d[0] + R[1][0] * d[1] + R[2][0] * d[2], R[0][1] * d[0] + R[1][1] * d[1] + R[2][1] * d[2],
R[0][2] * d[0] + R[1][2] * d[1] + R[2][2] * d[2]};
if (depth) *depth = c[2];
return project_cam(c);
}
double Camera::off_axis(V2 uv) const { return std::acos(std::clamp(unproject(uv)[2], -1.0, 1.0)) * 180 / M_PI; }
// Off the Frame (frame-job on the 7i), the /persist files are copies under FRAME_JOB_DEVICE_ROOT.
static std::string device_path(const char *path) {
const char *root = std::getenv("FRAME_JOB_DEVICE_ROOT");
return root ? std::string(root) + path : std::string(path);
}
bool load_calibration(std::map<std::string, Camera> &out, std::string &err) {
Json::Value rig, dev;
if (!read_json(device_path("/persist/xrservice.json").c_str(), rig, err) ||
!read_json(device_path("/persist/device_config.json").c_str(), dev, err))
return false;
double cad_from_cam0[4][4], cad_from_head[4][4], head_from_cad[4][4], head_from_cam0[4][4];
pose(dev["cv"]["cad_from_cal"], 1.0, cad_from_cam0);
pose(dev["head"], 1.0, cad_from_head);
invert_rigid(cad_from_head, head_from_cad);
mul(head_from_cad, cad_from_cam0, head_from_cam0);
for (const Json::Value &c : rig["cameras"]) {
Camera cam;
cam.name = c["sourceCamera"].asString();
cam.width = c["width"].asInt(), cam.height = c["height"].asInt();
for (const Json::Value &in : c["intrinsics"]) {
if (in["cameraModel"].asString() != "kb") continue;
cam.fx = in["fx"].asDouble(), cam.fy = in["fy"].asDouble();
cam.cx = in["cx"].asDouble(), cam.cy = in["cy"].asDouble();
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
}
double cam0_from_cam[4][4], head_from_cam[4][4];
pose(c["extrinsics"], 1e-3, cam0_from_cam);
mul(head_from_cam0, cam0_from_cam, head_from_cam);
for (int i = 0; i < 3; ++i) {
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
cam.origin[i] = head_from_cam[i][3];
}
out[cam.name] = cam;
}
if (out.empty()) err = "no cameras in /persist/xrservice.json";
return !out.empty();
}
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms) {
double A[3][3] = {}, b[3] = {};
for (int v = 0; v < n; ++v) {
const V3 &d = dirs[v], &o = origins[v];
for (int i = 0; i < 3; ++i)
for (int j = 0; j < 3; ++j) {
const double P = (i == j ? 1.0 : 0.0) - d[i] * d[j];
A[i][j] += weights[v] * P;
b[i] += weights[v] * P * o[j];
}
}
// Cramer's rule for the 3x3 system
auto det3 = [](const double m[3][3]) {
return m[0][0] * (m[1][1] * m[2][2] - m[1][2] * m[2][1]) - m[0][1] * (m[1][0] * m[2][2] - m[1][2] * m[2][0]) +
m[0][2] * (m[1][0] * m[2][1] - m[1][1] * m[2][0]);
};
const double D = det3(A);
V3 p{};
for (int c = 0; c < 3; ++c) {
double M[3][3];
for (int i = 0; i < 3; ++i)
for (int j = 0; j < 3; ++j) M[i][j] = j == c ? b[i] : A[i][j];
p[c] = std::fabs(D) > 1e-18 ? det3(M) / D : 0;
}
if (rms) {
double s = 0;
for (int v = 0; v < n; ++v) {
const V3 off = p - origins[v];
const V3 perp = off - dirs[v] * dot(off, dirs[v]);
s += dot(perp, perp);
}
*rms = std::sqrt(s / n);
}
return p;
}
+31
View File
@@ -0,0 +1,31 @@
// Tracking-camera calibration from the headset's factory files (see tracker/calib.py
// for the conventions): Kannala-Brandt fisheye intrinsics, and each camera's pose in the
// head frame (OpenVR's: +x right, +y up, -z forward), metres.
#pragma once
#include "geom.h"
#include <map>
#include <string>
#include <vector>
struct Camera {
std::string name;
int width = 0, height = 0;
double fx = 1, fy = 1, cx = 0, cy = 0, k[4] = {};
double R[3][3] = {}; // camera axes (columns) in the head frame
V3 origin{}; // camera centre in the head frame
V2 project_cam(V3 p) const; // camera frame -> pixels
V3 unproject(V2 uv) const; // pixels -> unit ray, camera frame
V3 ray(V2 uv) const; // pixels -> unit ray, head frame
V2 project(V3 head, double *depth) const; // head frame -> pixels; depth along the optical axis
double off_axis(V2 uv) const; // degrees between the pixel's ray and the axis
};
// Loads /persist/xrservice.json and /persist/device_config.json. Keyed by calibration
// name: slam_left, slam_right, upper_left, upper_right.
bool load_calibration(std::map<std::string, Camera> &out, std::string &err);
// The point closest to several rays (weighted), and its rms distance to them.
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms);
+22
View File
@@ -0,0 +1,22 @@
// Small vector helpers for the tracker.
#pragma once
#include <array>
#include <cmath>
using V2 = std::array<double, 2>;
using V3 = std::array<double, 3>;
inline V3 operator+(V3 a, V3 b) { return {a[0] + b[0], a[1] + b[1], a[2] + b[2]}; }
inline V3 operator-(V3 a, V3 b) { return {a[0] - b[0], a[1] - b[1], a[2] - b[2]}; }
inline V3 operator*(V3 a, double s) { return {a[0] * s, a[1] * s, a[2] * s}; }
inline double dot(V3 a, V3 b) { return a[0] * b[0] + a[1] * b[1] + a[2] * b[2]; }
inline double norm(V3 a) { return std::sqrt(dot(a, a)); }
inline V3 unit(V3 a) { double n = norm(a); return n > 0 ? a * (1 / n) : a; }
inline V2 operator+(V2 a, V2 b) { return {a[0] + b[0], a[1] + b[1]}; }
inline V2 operator-(V2 a, V2 b) { return {a[0] - b[0], a[1] - b[1]}; }
inline V2 operator*(V2 a, double s) { return {a[0] * s, a[1] * s}; }
inline double norm(V2 a) { return std::hypot(a[0], a[1]); }
inline double wrap_angle(double a) { return std::remainder(a, 2 * M_PI); }
+169
View File
@@ -0,0 +1,169 @@
#include "io.h"
#include <fcntl.h>
#include <sys/mman.h>
#include <sys/stat.h>
#include <time.h>
#include <unistd.h>
#include <algorithm>
#include <cstdlib>
#include <cstring>
uint64_t mono_ns() {
timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return uint64_t(ts.tv_sec) * 1'000'000'000 + uint64_t(ts.tv_nsec);
}
int64_t raw_minus_mono_ns() {
timespec a, r, b;
clock_gettime(CLOCK_MONOTONIC, &a);
clock_gettime(CLOCK_MONOTONIC_RAW, &r);
clock_gettime(CLOCK_MONOTONIC, &b);
const int64_t ma = int64_t(a.tv_sec) * 1'000'000'000 + a.tv_nsec, mb = int64_t(b.tv_sec) * 1'000'000'000 + b.tv_nsec;
return int64_t(r.tv_sec) * 1'000'000'000 + r.tv_nsec - (ma + mb) / 2;
}
// ------------------------------------------------------------------------------ ring
bool Ring::open(const char *path, std::string &err) {
const int fd = ::open(path, O_RDONLY | O_CLOEXEC);
if (fd < 0) return err = std::string(path) + ": " + std::strerror(errno), false;
struct stat st;
fstat(fd, &st);
len_ = size_t(st.st_size);
void *m = len_ >= sizeof(fh_ring_hdr_t) ? mmap(nullptr, len_, PROT_READ, MAP_SHARED, fd, 0) : MAP_FAILED;
close(fd);
if (m == MAP_FAILED) return err = std::string(path) + ": can't map it", false;
map_ = static_cast<const uint8_t *>(m);
hdr_ = reinterpret_cast<const fh_ring_hdr_t *>(map_);
if (std::memcmp(hdr_->magic, FH_RING_MAGIC, 8) || hdr_->version != FH_RING_VERSION || hdr_->file_bytes > len_)
return err = std::string(path) + " is not an fh-camd ring", false;
return true;
}
bool Ring::alive() const {
const uint64_t hb = __atomic_load_n(&hdr_->heartbeat_ns, __ATOMIC_ACQUIRE);
return hb && mono_ns() - hb < 1'000'000'000;
}
uint64_t Ring::latest(int i) const { return __atomic_load_n(&hdr_->cams[i].latest, __ATOMIC_ACQUIRE); }
bool Ring::read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const {
const fh_ring_cam_t &c = hdr_->cams[i];
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
if (seq != 2 * n + 2) return false;
std::memcpy(meta, slot, sizeof *meta);
out.resize(size_t(c.width) * c.height);
for (uint32_t y = 0; y < c.height; ++y)
std::memcpy(out.data() + size_t(y) * c.width, slot + sizeof(fh_ring_slot_t) + size_t(y) * c.stride, c.width);
__atomic_thread_fence(__ATOMIC_ACQUIRE);
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
}
bool Ring::meta(int i, uint64_t n, fh_ring_slot_t *meta) const {
const fh_ring_cam_t &c = hdr_->cams[i];
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
if (seq != 2 * n + 2) return false;
std::memcpy(meta, slot, sizeof *meta);
__atomic_thread_fence(__ATOMIC_ACQUIRE);
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
}
// ------------------------------------------------------------------------- publisher
namespace {
// The hand's shape to cut out, as capsules (tracker/publish.py has the same model).
// Radii are a real hand's half-widths plus a small margin for tracking noise.
const int kThumb[][2] = {{0, 1}, {1, 2}, {2, 3}, {3, 4}};
const int kFingers[][2] = {{5, 6}, {6, 7}, {7, 8}, {9, 10}, {10, 11}, {11, 12}, {13, 14}, {14, 15}, {15, 16},
{17, 18}, {18, 19}, {19, 20}};
const int kPalm[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 9}, {9, 13}, {13, 17}, {1, 5}};
constexpr double kThumbR = 0.0095, kFinger = 0.0085, kPalmR = 0.015, kArm[2] = {0.028, 0.034}, kArmLen = 0.16,
kMargin = 0.004;
// Nothing is cut closer than this in front of the eyes (head frame, -z is forward). A
// point near the eyes' plane lands far across a screen with a huge radius, so one bad
// estimate there tears a hole through it; real hands that close aren't tracked anyway.
constexpr double kNear = 0.12;
// Adds the capsule, clipped to the part at least kNear in front of the eyes.
void put(fh_capsule_t *caps, uint32_t &n, V3 a, V3 b, double ra, double rb) {
if (n >= FH_HANDS_MAX_CAPSULES) return;
const double za = -a[2] - kNear, zb = -b[2] - kNear; // >= 0: far enough in front
if (za < 0 && zb < 0) return;
if (za < 0 || zb < 0) {
const double t = za / (za - zb); // where the segment crosses the near plane
const V3 m = a + (b - a) * t;
const double rm = ra + (rb - ra) * t;
if (za < 0) a = m, ra = rm;
else b = m, rb = rm;
}
fh_capsule_t &c = caps[n++];
for (int k = 0; k < 3; ++k) c.a[k] = float(a[k]), c.b[k] = float(b[k]);
c.ra = float(ra + kMargin), c.rb = float(rb + kMargin);
}
} // namespace
bool Publisher::open(std::string &err) {
const char *run = std::getenv("XDG_RUNTIME_DIR");
const std::string dir = std::string(run ? run : "/run/user/" + std::to_string(getuid())) + "/frame-hands";
mkdir(dir.c_str(), 0700);
chmod(dir.c_str(), 0700);
const std::string path = dir + "/hands";
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
if (fd < 0 || ftruncate(fd, sizeof(fh_hands_t)) < 0) return err = path + ": " + std::strerror(errno), false;
void *m = mmap(nullptr, sizeof(fh_hands_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
close(fd);
if (m == MAP_FAILED) return err = path + ": can't map it", false;
out_ = static_cast<fh_hands_t *>(m);
std::memset(out_, 0, sizeof *out_);
std::memcpy(out_->magic, FH_HANDS_MAGIC, 8);
out_->version = FH_HANDS_VERSION;
out_->size = sizeof(fh_hands_t);
return true;
}
void Publisher::write(const std::vector<const Hand *> &in, uint64_t capture_ns) {
std::vector<const Hand *> hands = in;
std::sort(hands.begin(), hands.end(), [](const Hand *a, const Hand *b) { return a->frames > b->frames; });
if (hands.size() > FH_HANDS_MAX_HANDS) hands.resize(FH_HANDS_MAX_HANDS);
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
__atomic_thread_fence(__ATOMIC_RELEASE);
uint32_t nc = 0;
for (size_t k = 0; k < FH_HANDS_MAX_HANDS; ++k) {
fh_hand_t &o = out_->hands[k];
std::memset(&o, 0, sizeof o);
if (k >= hands.size()) continue;
const Hand &h = *hands[k];
o.id = uint32_t(h.id);
o.flags = (h.right() ? FH_HAND_RIGHT : 0) | (h.nviews >= 2 ? FH_HAND_STEREO : 0);
o.confidence = float(std::min(1.0, h.frames / 5.0));
for (int i = 0; i < 21; ++i)
for (int j = 0; j < 3; ++j) o.pts[i][j] = float(h.smooth[i][j]);
const uint32_t first = nc;
for (auto &b : kThumb) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kThumbR, kThumbR);
for (auto &b : kFingers) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kFinger, kFinger);
for (auto &b : kPalm) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kPalmR, kPalmR);
// the forearm carries on from the hand's own axis (middle knuckle -> wrist); the
// wrist bends, but much less than a guess at where the elbow is gets wrong
const V3 wrist = h.smooth[0], d = wrist - h.smooth[9];
const double n = norm(d);
if (n > 0.02) put(out_->capsules, nc, wrist, wrist + d * (kArmLen / n), kArm[0], kArm[1]);
o.ncapsules = nc - first;
}
for (uint32_t k = nc; k < FH_HANDS_MAX_CAPSULES; ++k) std::memset(&out_->capsules[k], 0, sizeof(fh_capsule_t));
out_->capture_ns = capture_ns;
out_->publish_ns = mono_ns();
out_->nhands = uint32_t(hands.size());
out_->ncapsules = nc;
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
}
+46
View File
@@ -0,0 +1,46 @@
// Frames in from fh-camd's ring (camd/fhring.h), hands out to the hands file
// (include/fh_hands.h, read by Frametop's ft-screens).
#pragma once
#include "tracker.h"
#include <cstdint>
#include <string>
#include <vector>
extern "C" {
#include "../camd/fhring.h"
#include "../include/fh_hands.h"
}
class Ring {
public:
bool open(const char *path, std::string &err);
bool alive() const; // the writer's heartbeat is fresh
int cameras() const { return int(hdr_->ncams); }
const fh_ring_cam_t &camera(int i) const { return hdr_->cams[i]; }
uint64_t latest(int i) const;
// Copy frame n of camera i into out (width x height, tightly packed). False if it's
// gone or was being written.
bool read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const;
// Just frame n's slot header (capture time etc.), without copying the image.
bool meta(int i, uint64_t n, fh_ring_slot_t *meta) const;
private:
const uint8_t *map_ = nullptr;
const fh_ring_hdr_t *hdr_ = nullptr;
size_t len_ = 0;
};
class Publisher {
public:
bool open(std::string &err);
void write(const std::vector<const Hand *> &hands, uint64_t capture_ns);
private:
fh_hands_t *out_ = nullptr;
uint64_t seq_ = 0;
};
uint64_t mono_ns();
int64_t raw_minus_mono_ns(); // camera timestamps are CLOCK_MONOTONIC_RAW
+295
View File
@@ -0,0 +1,295 @@
// fh-tracker: hands in 3D from fh-camd's ring, published for Frametop's ft-screens.
// The C++ version of tracker/live.py: the same scheduling, with the models on a few
// threads and no Python in the loop.
//
// fh-tracker [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N]
// [--no-publish] [--record DIR] [--swap-sides] ... (--help lists them all)
#include "io.h"
#include "record.h"
#include <sched.h>
#include <sys/resource.h>
#include <sys/stat.h>
#include <unistd.h>
#include <ctime>
#include <memory>
#include <csignal>
#include <cstdio>
#include <cstring>
#include <fstream>
#include <thread>
#include <utility>
namespace {
volatile std::sig_atomic_t g_stop = 0, g_record = 0;
// which calibrated camera each capture pipe carries (XRService's fixed routing)
const char *camera_for_pipe(int node) {
char path[64], name[64] = "";
std::snprintf(path, sizeof path, "/sys/class/video4linux/video%d/name", node);
std::ifstream f(path);
f.getline(name, sizeof name);
if (!std::strcmp(name, "msm_vfe3_video0")) return "slam_left";
if (!std::strcmp(name, "msm_vfe4_video0")) return "slam_right";
if (!std::strcmp(name, "msm_vfe2_video0")) return "upper_left";
if (!std::strcmp(name, "msm_vfe2_video1")) return "upper_right";
return nullptr;
}
double cpu_seconds() {
rusage r;
getrusage(RUSAGE_SELF, &r);
return r.ru_utime.tv_sec + r.ru_stime.tv_sec + (r.ru_utime.tv_usec + r.ru_stime.tv_usec) / 1e6;
}
} // namespace
int main(int argc, char **argv) {
double seconds = 0, status = 5;
int threads = 3, niceness = 5;
bool int8 = false, publish = true, track = true, swap_sides = false;
std::string models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
std::string record;
// SteamOS starts user processes on CPUs 0-4 (2-4 are the big A720s) and keeps 5-7 (two
// A720s and the X4) for SteamVR's compositor, whose threads there run at real-time
// priority. XRService pins its head tracking to 2-3.
std::vector<int> cpus = {2, 3, 4};
// How crops are equalized. CLAHE helps the palm search find hands (about 10% more in the
// dim recording), but makes the landmarks jitter, so they get plain crops.
Contrast palm_contrast, hand_contrast{Contrast::None};
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
double record_for = 120;
for (int i = 1; i < argc; ++i) {
const std::string a = argv[i];
const bool more = i + 1 < argc;
if (a == "--seconds" && more) seconds = std::atof(argv[++i]);
else if (a == "--threads" && more) threads = std::max(1, std::atoi(argv[++i]));
else if (a == "--status" && more) status = std::atof(argv[++i]);
else if (a == "--models" && more) models = argv[++i];
else if (a == "--nice" && more) niceness = std::atoi(argv[++i]);
else if (a == "--int8") int8 = true;
else if (a == "--no-publish") publish = false;
else if (a == "--swap-sides") swap_sides = true;
else if (a == "--record-only") track = publish = false;
else if (a == "--record" && more) record = argv[++i];
else if (a == "--record-for" && more) record_for = std::atof(argv[++i]);
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
else if (a == "--contrast" && more) {
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
} else if (a == "--cpus" && more) {
cpus.clear();
for (char *p = argv[++i]; *p;) {
cpus.push_back(int(std::strtol(p, &p, 10)));
if (*p == ',') ++p;
else if (*p) break;
}
if (cpus.empty()) cpus = {2, 3, 4};
}
else {
std::printf("usage: %s [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] [--no-publish]\n"
" [--record DIR] [--record-for S] [--record-only] [--cpus 2,3,4] [--swap-sides]\n"
" [--keep-presence P] (0.5)\n"
" [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch; default clahe:2/none)\n"
"Recording saves every frame set for S seconds (120) to DIR/sets.bin, for fh-replay; SIGUSR1\n"
"starts one in captures/rec-<time> next to trackd. --record-only records without tracking, so it\n"
"can run beside a tracking fh-tracker. With fh-camd --with-dark, recordings also get each\n"
"camera's newest dark frame, as <name>_dk.\n",
argv[0]);
return a == "--help" ? 0 : 1;
}
}
if (nice(niceness) < 0) std::perror("nice"); // the VR stack wins contested CPUs
std::signal(SIGINT, [](int) { g_stop = 1; });
std::signal(SIGTERM, [](int) { g_stop = 1; });
std::signal(SIGUSR1, [](int) { g_record = 1; });
std::string err;
std::map<std::string, Camera> calib;
Ring ring;
Nets nets;
Publisher pub;
std::unique_ptr<Recorder> rec;
uint64_t rec_start = 0;
auto start_recording = [&](const std::string &dir, std::string &e) {
rec = std::make_unique<Recorder>();
if (!rec->open(dir, e)) return rec.reset(), false;
rec_start = mono_ns();
std::printf("recording to %s for %.0f s\n", dir.c_str(), record_for);
std::fflush(stdout);
return true;
};
if (!load_calibration(calib, err) || !ring.open(FH_RING_PATH, err) || !nets.load(models, int8, err) ||
(publish && !pub.open(err)) || (!record.empty() && !start_recording(record, err))) {
std::fprintf(stderr, "%s\n", err.c_str());
return 1;
}
if (!ring.alive()) return std::fprintf(stderr, "fh-camd isn't running (no heartbeat)\n"), 1;
std::map<std::string, int> index; // calibration name -> ring camera
// "<name>_dk" -> ring camera (fh-camd --with-dark): recorded only. Recorded names hold 15
// characters, so "upper_right_dark" wouldn't fit.
std::map<std::string, int> dark;
std::map<std::string, Camera> used;
for (int i = 0; i < ring.cameras(); ++i) {
const char *name = camera_for_pipe(ring.camera(i).node);
if (!name || !calib.count(name)) continue;
if (ring.camera(i).flags & FH_CAM_DARK) dark[std::string(name) + "_dk"] = i;
else index[name] = i, used[name] = calib[name];
}
// fh-camd tells the side cameras' buffers apart by XRService's allocation order, which
// some XRService restarts reverse; tools/check_sides.py --ring tells when.
if (swap_sides && index.count("slam_left") && index.count("slam_right")) {
std::swap(index["slam_left"], index["slam_right"]);
if (dark.count("slam_left_dk") && dark.count("slam_right_dk")) std::swap(dark["slam_left_dk"], dark["slam_right_dk"]);
std::printf("side cameras swapped (--swap-sides)\n");
}
std::printf("cameras:");
for (auto &[name, i] : index) std::printf(" %s=video%d", name.c_str(), ring.camera(i).node);
std::printf(" models: %s%s, %d threads on CPUs", models.c_str(), int8 ? " (int8)" : "", threads);
for (int c : cpus) std::printf(" %d", c);
std::printf("\n");
cpu_set_t set; // the main loop too
CPU_ZERO(&set);
for (int c : cpus) CPU_SET(c, &set);
if (sched_setaffinity(0, sizeof set, &set) < 0) std::perror("sched_setaffinity");
nets.set_contrast(palm_contrast, hand_contrast);
Pool pool(threads, cpus);
Tracker tracker(used, nets, pool);
tracker.set_keep_presence(keep_presence);
std::map<std::string, std::vector<uint8_t>> pixels;
std::map<std::string, uint64_t> last;
const uint64_t start = mono_ns();
uint64_t t_status = start, next_ns = 0;
double cpu0 = cpu_seconds();
std::vector<double> lat;
double hands_sum = 0, resid_sum = 0;
int resid_n = 0, left_sets = 0, right_sets = 0, both_sets = 0;
while (!g_stop && (seconds <= 0 || (mono_ns() - start) / 1e9 < seconds)) {
if (!ring.alive()) return std::fprintf(stderr, "fh-camd stopped\n"), 2;
// a new frame set: every camera has a newer frame, taken at the same moment
std::map<std::string, uint64_t> latest;
bool ready = true;
for (auto &[name, i] : index) {
latest[name] = ring.latest(i);
ready = ready && latest[name] > last[name];
}
if (!ready) {
std::this_thread::sleep_for(std::chrono::milliseconds(2));
continue;
}
// not needed at the current rate, and not recorded: skip it without copying images
if (track && !rec && !g_record) {
uint64_t t0 = UINT64_MAX, t1 = 0;
bool ok = true;
for (auto &[name, i] : index) {
fh_ring_slot_t meta;
ok = ok && ring.meta(i, latest[name], &meta);
if (ok) t0 = std::min(t0, meta.capture_ns), t1 = std::max(t1, meta.capture_ns);
}
if (ok && t1 - t0 <= 3'000'000 && t0 < next_ns) {
last = latest;
continue;
}
}
std::map<std::string, Image> images;
std::vector<SetFrame> frames;
uint64_t tmin = UINT64_MAX, tmax = 0, dq = 0;
bool ok = true;
for (auto &[name, i] : index) {
fh_ring_slot_t meta;
ok = ok && ring.read(i, latest[name], pixels[name], &meta);
if (!ok) break;
const auto &c = ring.camera(i);
images[name] = {pixels[name].data(), int(c.width), int(c.height), int(c.width)};
frames.push_back({name, pixels[name].data(), c.width, c.height, meta.capture_ns, meta.dqbuf_ns});
tmin = std::min(tmin, meta.capture_ns), tmax = std::max(tmax, meta.capture_ns), dq = std::max(dq, meta.dqbuf_ns);
}
if (!ok || tmax - tmin > 3'000'000) { // torn, or a camera is a frame behind
std::this_thread::sleep_for(std::chrono::milliseconds(1));
continue;
}
last = latest;
if (g_record && !rec) {
g_record = 0;
char name[64];
const std::time_t now = std::time(nullptr);
std::strftime(name, sizeof name, "rec-%Y%m%d-%H%M%S", std::localtime(&now));
const std::string here = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1);
const std::string dir = here + "../captures";
mkdir(dir.c_str(), 0755);
std::string e;
if (!start_recording(dir + "/" + name, e)) std::fprintf(stderr, "%s\n", e.c_str());
}
if (rec) { // about 80 MB/s, twice that with dark frames
if ((mono_ns() - rec_start) / 1e9 < record_for) {
for (auto &[name, i] : dark) { // the newest dark frame of each camera, as it is
fh_ring_slot_t meta;
const uint64_t n = ring.latest(i);
const auto &c = ring.camera(i);
if (n && ring.read(i, n, pixels[name], &meta))
frames.push_back({name, pixels[name].data(), c.width, c.height, meta.capture_ns, meta.dqbuf_ns});
}
rec->add(frames);
} else {
const size_t n = rec->written(), d = rec->dropped();
rec.reset(); // writes out what's queued
std::printf("recording done: %zu sets, %zu dropped\n", n, d);
std::fflush(stdout);
if (!track) break;
}
}
if (!track && rec && status > 0 && (mono_ns() - t_status) / 1e9 >= status) {
std::printf("%5.1fs recorded %zu sets, dropped %zu\n", (mono_ns() - start) / 1e9, rec->written(), rec->dropped());
std::fflush(stdout);
t_status = mono_ns();
}
if (!track || tmin < next_ns) continue; // not needed yet at the current rate
const auto hands = tracker.step(images, int64_t(tmin));
next_ns = tmin + uint64_t((tracker.interval() - 0.005) * 1e9);
if (publish) pub.write(hands, uint64_t(int64_t(tmin) - raw_minus_mono_ns()));
lat.push_back((mono_ns() - dq) / 1e6);
hands_sum += double(hands.size());
bool on_left = false, on_right = false; // by where the wrist is, not the model's label
for (const Hand *h : hands) {
if (h->residual >= 0) resid_sum += h->residual * 1000, ++resid_n;
(h->pts[0][0] < 0 ? on_left : on_right) = true;
}
left_sets += on_left, right_sets += on_right, both_sets += on_left && on_right;
const uint64_t now = mono_ns();
if (status > 0 && (now - t_status) / 1e9 >= status) {
const double dt = (now - t_status) / 1e9, cpu1 = cpu_seconds();
const Stats &s = tracker.stats;
std::sort(lat.begin(), lat.end());
std::printf("%5.1fs %4.1f sets/s hands %.2f views %zu palm %3d calls %4.1f ms/batch hand %3d calls %4.1f ms/batch "
"step %4.1f ms latency %4.1f ms resid %.1f mm CPU %3.0f%%\n",
(now - start) / 1e9, s.sets / dt, s.sets ? hands_sum / s.sets : 0, tracker.views(), s.palm_calls,
s.palm_batches ? s.palm_ms / s.palm_batches : 0, s.hand_calls,
s.hand_batches ? s.hand_ms / s.hand_batches : 0, s.sets ? s.step_ms / s.sets : 0,
lat.empty() ? 0 : lat[lat.size() / 2], resid_n ? resid_sum / resid_n : 0, 100 * (cpu1 - cpu0) / dt);
if (s.sets)
std::printf(" sets with a hand: left %2.0f%% right %2.0f%% both %2.0f%% views lost %d, handoff misses %d, "
"dups %d, splits %d hands new %d merged %d forgotten %d%s\n",
100.0 * left_sets / s.sets, 100.0 * right_sets / s.sets, 100.0 * both_sets / s.sets, s.lost,
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten,
!rec ? "" : (" recorded " + std::to_string(rec->written()) + " dropped " +
std::to_string(rec->dropped())).c_str());
for (const Hand *h : hands)
std::printf(" hand %d %-5s views %d wrist %+.3f %+.3f %+.3f m scale %.2f speed %.2f m/s\n", h->id,
h->right() ? "right" : "left", h->nviews, h->pts[0][0], h->pts[0][1], h->pts[0][2], h->scale,
h->speed);
std::fflush(stdout);
tracker.stats = Stats{};
t_status = now, cpu0 = cpu1;
lat.clear(), hands_sum = 0, resid_sum = 0, resid_n = 0, left_sets = right_sets = both_sets = 0;
}
}
if (publish) pub.write({}, mono_ns());
return 0;
}
+255
View File
@@ -0,0 +1,255 @@
#include "nets.h"
#include <mat.h>
#include <algorithm>
#include <cstdlib>
#include <numeric>
namespace {
constexpr int kPalmSize = 192, kHandSize = 224;
const int kRoiLandmarks[] = {0, 1, 2, 3, 5, 6, 9, 10, 13, 14, 17, 18};
// 2x3 affine taking crop pixels (0..out) to image pixels.
void crop_matrix(V2 center, double size, double rotation, int out, float tm[6]) {
const double c = std::cos(rotation), s = std::sin(rotation), k = size / out;
tm[0] = float(c * k), tm[1] = float(-s * k), tm[3] = float(s * k), tm[4] = float(c * k);
tm[2] = float(center[0] - (tm[0] + tm[1]) * out / 2.0);
tm[5] = float(center[1] - (tm[3] + tm[4]) * out / 2.0);
}
V2 to_image(const float tm[6], double x, double y) {
return {tm[0] * x + tm[1] * y + tm[2], tm[3] * x + tm[4] * y + tm[5]};
}
// OpenCV's CLAHE (4x4 tiles) on a square crop, in place.
void clahe(uint8_t *img, int n, double clip_limit) {
constexpr int kTiles = 4;
const int ts = n / kTiles, area = ts * ts;
const int clip = std::max(1, int(clip_limit * area / 256));
uint8_t lut[kTiles][kTiles][256];
for (int ty = 0; ty < kTiles; ++ty)
for (int tx = 0; tx < kTiles; ++tx) {
int hist[256] = {};
for (int y = ty * ts; y < (ty + 1) * ts; ++y)
for (int x = tx * ts; x < (tx + 1) * ts; ++x) ++hist[img[y * n + x]];
int excess = 0;
for (int &h : hist)
if (h > clip) excess += h - clip, h = clip;
const int add = excess / 256, residual = excess - add * 256;
for (int i = 0; i < 256; ++i) hist[i] += add + (i < residual ? 1 : 0);
int sum = 0;
const float scale = 255.f / area;
for (int i = 0; i < 256; ++i) {
sum += hist[i];
lut[ty][tx][i] = uint8_t(std::min(255, int(sum * scale + 0.5f)));
}
}
std::vector<uint8_t> out(size_t(n) * n);
for (int y = 0; y < n; ++y) {
const float fy = (y + 0.5f) / ts - 0.5f;
const int y0 = std::clamp(int(std::floor(fy)), 0, kTiles - 1), y1 = std::min(y0 + 1, kTiles - 1);
const float wy = std::clamp(fy - y0, 0.f, 1.f);
for (int x = 0; x < n; ++x) {
const float fx = (x + 0.5f) / ts - 0.5f;
const int x0 = std::clamp(int(std::floor(fx)), 0, kTiles - 1), x1 = std::min(x0 + 1, kTiles - 1);
const float wx = std::clamp(fx - x0, 0.f, 1.f);
const uint8_t v = img[y * n + x];
const float top = lut[y0][x0][v] * (1 - wx) + lut[y0][x1][v] * wx;
const float bot = lut[y1][x0][v] * (1 - wx) + lut[y1][x1][v] * wx;
out[size_t(y) * n + x] = uint8_t(top * (1 - wy) + bot * wy + 0.5f);
}
}
std::copy(out.begin(), out.end(), img);
}
// Linear stretch of the 1st..99th percentile to 0..255, in place.
void stretch(uint8_t *img, int n) {
int hist[256] = {};
const int total = n * n;
for (int i = 0; i < total; ++i) ++hist[img[i]];
int lo = 0, hi = 255, acc = 0;
for (int v = 0; v < 256; ++v)
if ((acc += hist[v]) > total / 100) { lo = v; break; }
acc = 0;
for (int v = 255; v >= 0; --v)
if ((acc += hist[v]) > total / 100) { hi = v; break; }
if (hi <= lo) return;
for (int i = 0; i < total; ++i) img[i] = uint8_t(std::clamp((img[i] - lo) * 255 / (hi - lo), 0, 255));
}
// A crop as the models' input: RGB (the mono plane three times), 0..1.
ncnn::Mat crop(const Image &img, const float tm[6], int n, const Contrast &contrast) {
std::vector<uint8_t> patch(size_t(n) * n);
ncnn::warpaffine_bilinear_c1(img.data, img.width, img.height, img.stride, patch.data(), n, n, n, tm, 0, 0);
if (contrast.mode == Contrast::Clahe) clahe(patch.data(), n, contrast.clip);
else if (contrast.mode == Contrast::Stretch) stretch(patch.data(), n);
ncnn::Mat m = ncnn::Mat::from_pixels(patch.data(), ncnn::Mat::PIXEL_GRAY2RGB, n, n);
const float norm[3] = {1 / 255.f, 1 / 255.f, 1 / 255.f};
m.substract_mean_normalize(nullptr, norm);
return m;
}
} // namespace
bool Contrast::parse(const std::string &s, Contrast &out) {
if (s == "none") return out.mode = None, true;
if (s == "stretch") return out.mode = Stretch, true;
if (s.rfind("clahe", 0) == 0) {
out.mode = Clahe;
out.clip = s.size() > 6 && s[5] == ':' ? std::atof(s.c_str() + 6) : 2.0;
return out.clip > 0;
}
return false;
}
bool Contrast::parse_pair(const std::string &s, Contrast &palm, Contrast &hand) {
const size_t slash = s.find('/');
if (slash == std::string::npos) return parse(s, palm) && parse(s, hand);
return parse(s.substr(0, slash), palm) && parse(s.substr(slash + 1), hand);
}
namespace {
bool load_net(ncnn::Net &net, const std::string &base, std::string &err) {
net.opt.num_threads = 1;
net.opt.use_vulkan_compute = false;
net.opt.use_fp16_packed = net.opt.use_fp16_storage = net.opt.use_fp16_arithmetic = true;
if (net.load_param((base + ".param").c_str()) || net.load_model((base + ".bin").c_str())) {
err = "can't load " + base + ".param/.bin";
return false;
}
return true;
}
} // namespace
Roi Palm::roi() const {
const V2 a = kp[0], b = kp[2];
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(b[1] - a[1]), b[0] - a[0]));
const double h = size[1];
const V2 shift{-h * -0.5 * std::sin(rot), h * -0.5 * std::cos(rot)};
return {center + shift, std::max(size[0], size[1]) * 2.6, rot};
}
Roi roi_from_points(const V2 *p) {
const V2 w = p[0];
V2 m = (p[5] + p[13]) * 0.5;
m = (m + p[9]) * 0.5;
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(m[1] - w[1]), m[0] - w[0]));
V2 lo{1e9, 1e9}, hi{-1e9, -1e9};
for (int i : kRoiLandmarks)
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[i][k]), hi[k] = std::max(hi[k], p[i][k]);
V2 center = (lo + hi) * 0.5;
const double c = std::cos(-rot), s = std::sin(-rot);
V2 qlo{1e9, 1e9}, qhi{-1e9, -1e9};
for (int i : kRoiLandmarks) {
const V2 d = p[i] - center;
const V2 q{d[0] * c - d[1] * s, d[0] * s + d[1] * c};
for (int k = 0; k < 2; ++k) qlo[k] = std::min(qlo[k], q[k]), qhi[k] = std::max(qhi[k], q[k]);
}
const V2 mid = (qlo + qhi) * 0.5;
const double c2 = std::cos(rot), s2 = std::sin(rot);
center = center + V2{mid[0] * c2 - mid[1] * s2, mid[0] * s2 + mid[1] * c2};
const double w2 = qhi[0] - qlo[0], h2 = qhi[1] - qlo[1];
center = center + V2{-h2 * -0.1 * s2, h2 * -0.1 * c2};
return {center, std::max(w2, h2) * 2.0, rot};
}
Roi Landmarks::next_roi() const { return roi_from_points(pts); }
bool Nets::load(const std::string &dir, bool int8, std::string &err) {
const std::string suffix = int8 ? "-int8.ncnn" : ".ncnn";
if (!load_net(palm_, dir + "/palm" + suffix, err) || !load_net(hand_, dir + "/hand" + suffix, err)) return false;
// SSD anchors of palm_detection_full: strides 8 (2 per cell) and 16 (6 per cell)
for (auto [stride, per] : {std::pair{8, 2}, std::pair{16, 6}}) {
const int n = kPalmSize / stride;
for (int y = 0; y < n; ++y)
for (int x = 0; x < n; ++x)
for (int k = 0; k < per; ++k) anchors_.push_back({(x + 0.5) / n * kPalmSize, (y + 0.5) / n * kPalmSize});
}
return true;
}
std::vector<Palm> Nets::palms(const Image &img, V2 center, double size, double rotation) const {
float tm[6];
crop_matrix(center, size, rotation, kPalmSize, tm);
ncnn::Extractor ex = palm_.create_extractor();
ex.input("in0", crop(img, tm, kPalmSize, palm_contrast_));
ncnn::Mat boxes, scores;
ex.extract("out0", boxes);
ex.extract("out1", scores);
const float *raw = boxes, *logit = scores;
const int n = int(anchors_.size());
const float min_logit = std::log(0.5f / 0.5f); // score 0.5
struct Cand { V2 c, s; V2 kp[7]; double score; };
std::vector<Cand> cand;
for (int i = 0; i < n; ++i) {
if (logit[i] <= min_logit) continue;
const float *r = raw + i * 18;
Cand c;
c.c = {r[0] + anchors_[i][0], r[1] + anchors_[i][1]};
c.s = {r[2], r[3]};
for (int k = 0; k < 7; ++k) c.kp[k] = {r[4 + 2 * k] + anchors_[i][0], r[5 + 2 * k] + anchors_[i][1]};
c.score = 1 / (1 + std::exp(-std::clamp(double(logit[i]), -100.0, 100.0)));
cand.push_back(c);
}
// MediaPipe's weighted NMS: overlapping boxes are averaged, weighted by score
std::sort(cand.begin(), cand.end(), [](const Cand &a, const Cand &b) { return a.score > b.score; });
std::vector<bool> used(cand.size());
std::vector<Palm> out;
for (size_t i = 0; i < cand.size(); ++i) {
if (used[i]) continue;
double wsum = 0;
Cand acc{};
for (size_t j = i; j < cand.size(); ++j) {
if (used[j]) continue;
const double ix = std::max(0.0, std::min(cand[i].c[0] + cand[i].s[0] / 2, cand[j].c[0] + cand[j].s[0] / 2) -
std::max(cand[i].c[0] - cand[i].s[0] / 2, cand[j].c[0] - cand[j].s[0] / 2));
const double iy = std::max(0.0, std::min(cand[i].c[1] + cand[i].s[1] / 2, cand[j].c[1] + cand[j].s[1] / 2) -
std::max(cand[i].c[1] - cand[i].s[1] / 2, cand[j].c[1] - cand[j].s[1] / 2));
const double inter = ix * iy;
const double uni = cand[i].s[0] * cand[i].s[1] + cand[j].s[0] * cand[j].s[1] - inter;
if (j != i && inter / (uni + 1e-9) <= 0.3) continue;
used[j] = true;
const double w = cand[j].score;
wsum += w;
acc.c = acc.c + cand[j].c * w;
acc.s = acc.s + cand[j].s * w;
for (int k = 0; k < 7; ++k) acc.kp[k] = acc.kp[k] + cand[j].kp[k] * w;
}
Palm p;
const V2 c = acc.c * (1 / wsum);
p.center = to_image(tm, c[0], c[1]);
p.size = acc.s * (1 / wsum * size / kPalmSize);
for (int k = 0; k < 7; ++k) {
const V2 q = acc.kp[k] * (1 / wsum);
p.kp[k] = to_image(tm, q[0], q[1]);
}
p.score = cand[i].score;
out.push_back(p);
}
return out;
}
Landmarks Nets::landmarks(const Image &img, const Roi &roi) const {
float tm[6];
crop_matrix(roi.center, roi.size, roi.rotation, kHandSize, tm);
ncnn::Extractor ex = hand_.create_extractor();
ex.input("in0", crop(img, tm, kHandSize, hand_contrast_));
ncnn::Mat screen, presence, right, world;
ex.extract("out0", screen);
ex.extract("out1", presence);
ex.extract("out2", right);
ex.extract("out3", world);
Landmarks lm;
const float *s = screen, *w = world;
for (int i = 0; i < 21; ++i) {
lm.pts[i] = to_image(tm, s[3 * i], s[3 * i + 1]);
for (int k = 0; k < 3; ++k) lm.world[i][k] = w[3 * i + k];
}
lm.presence = presence[0];
lm.right = right[0];
return lm;
}
+66
View File
@@ -0,0 +1,66 @@
// MediaPipe's palm detector and hand landmark model on ncnn (see tracker/models.py for
// the conventions). A crop is a square region of a camera image: centre and size in
// pixels, and a rotation that turns the crop's "up" toward the image direction
// (sin r, -cos r). Crops are contrast-equalized (CLAHE) before the models see them.
// Everything here may run on several threads at once.
#pragma once
#include "geom.h"
#include <net.h>
#include <cstdint>
#include <string>
#include <vector>
struct Image {
const uint8_t *data = nullptr;
int width = 0, height = 0, stride = 0;
};
struct Roi {
V2 center{};
double size = 0, rotation = 0;
};
struct Palm {
V2 center{}, size{};
V2 kp[7]{};
double score = 0;
Roi roi() const; // MediaPipe's hand crop for this palm
};
struct Landmarks {
V2 pts[21]{}; // image pixels
double world[21][3]{}; // MediaPipe's metric landmarks, hand-centred
double presence = 0, right = 0;
Roi next_roi() const; // MediaPipe's crop to track the hand in the next frame
};
Roi roi_from_points(const V2 *pts21);
// How crops are contrast-equalized before the models see them.
struct Contrast {
enum Mode { Clahe, None, Stretch } mode = Clahe;
double clip = 2.0; // Clahe: OpenCV's clip limit (4x4 tiles)
// "clahe:2", "none", "stretch" (1st..99th percentile to 0..255)
static bool parse(const std::string &s, Contrast &out);
// "PALM/HAND" (each as above), or one for both
static bool parse_pair(const std::string &s, Contrast &palm, Contrast &hand);
};
class Nets {
public:
// Loads <dir>/palm.ncnn.* and <dir>/hand.ncnn.*, or the -int8 variants.
bool load(const std::string &dir, bool int8, std::string &err);
std::vector<Palm> palms(const Image &img, V2 center, double size, double rotation) const;
Landmarks landmarks(const Image &img, const Roi &roi) const;
// Before any palms()/landmarks(): how the palm search's and the landmark model's crops
// are equalized.
void set_contrast(const Contrast &palm, const Contrast &hand) { palm_contrast_ = palm, hand_contrast_ = hand; }
private:
Contrast palm_contrast_, hand_contrast_;
ncnn::Net palm_, hand_;
std::vector<V2> anchors_;
};
+75
View File
@@ -0,0 +1,75 @@
// Check the C++ model code against tracker/models.py on a recorded frame:
// nettest MODELS_DIR FRAME.pgm cx cy size rotation [--int8]
// Runs the palm detector on that crop, then the landmark model on each palm's ROI, and
// prints what they found; tools/nettest_compare.py runs the Python side on the same input.
#include "nets.h"
#include <algorithm>
#include <chrono>
#include <cstdio>
#include <cstring>
#include <fstream>
#include <string>
#include <vector>
static bool read_pgm(const char *path, std::vector<uint8_t> &px, int &w, int &h) {
std::ifstream f(path, std::ios::binary);
std::string magic;
int maxv;
if (!(f >> magic >> w >> h >> maxv) || magic != "P5") return false;
f.get();
px.resize(size_t(w) * h);
return bool(f.read(reinterpret_cast<char *>(px.data()), px.size()));
}
int main(int argc, char **argv) {
if (argc < 7) return std::fprintf(stderr, "usage: nettest MODELS FRAME.pgm cx cy size rotation [--int8] [--bench N]\n"), 1;
bool int8 = false;
int bench = 0;
for (int i = 7; i < argc; ++i) {
if (!std::strcmp(argv[i], "--int8")) int8 = true;
else if (!std::strcmp(argv[i], "--bench") && i + 1 < argc) bench = std::atoi(argv[++i]);
}
Nets nets;
std::string err;
if (!nets.load(argv[1], int8, err)) return std::fprintf(stderr, "%s\n", err.c_str()), 1;
std::vector<uint8_t> px;
int w, h;
if (!read_pgm(argv[2], px, w, h)) return std::fprintf(stderr, "can't read %s\n", argv[2]), 1;
const Image img{px.data(), w, h, w};
const V2 c{std::atof(argv[3]), std::atof(argv[4])};
if (bench > 0) { // steady-state timing: warm up, then the median of N calls each
const Roi roi{c, std::atof(argv[5]), std::atof(argv[6])};
auto time = [&](auto fn) {
std::vector<double> t;
for (int i = 0; i < bench + 5; ++i) {
const auto t0 = std::chrono::steady_clock::now();
fn();
if (i >= 5) t.push_back(std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t0).count());
}
std::sort(t.begin(), t.end());
return t[t.size() / 2];
};
const double pm = time([&] { nets.palms(img, c, roi.size, roi.rotation); });
const double hm = time([&] { nets.landmarks(img, roi); });
std::printf("bench%s: palm %.2f ms, hand %.2f ms (median of %d, one thread)\n", int8 ? " int8" : "", pm, hm, bench);
return 0;
}
auto t0 = std::chrono::steady_clock::now();
const auto palms = nets.palms(img, c, std::atof(argv[5]), std::atof(argv[6]));
const double palm_ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t0).count();
std::printf("palm_ms %.2f\n", palm_ms);
for (const Palm &p : palms) {
const Roi r = p.roi();
std::printf("palm %.3f center %.1f %.1f roi %.1f %.1f %.1f %.4f\n", p.score, p.center[0], p.center[1],
r.center[0], r.center[1], r.size, r.rotation);
t0 = std::chrono::steady_clock::now();
const Landmarks lm = nets.landmarks(img, r);
const double ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t0).count();
std::printf("hand_ms %.2f presence %.3f right %.3f\n", ms, lm.presence, lm.right);
std::printf("pts");
for (const V2 &q : lm.pts) std::printf(" %.1f %.1f", q[0], q[1]);
std::printf("\n");
}
return 0;
}
+101
View File
@@ -0,0 +1,101 @@
#include "record.h"
#include <sys/stat.h>
#include <cerrno>
#include <cstring>
namespace {
constexpr size_t kMaxQueued = 48; // about 130 MB of sets
}
Recorder::~Recorder() {
if (!f_) return;
{
std::lock_guard<std::mutex> l(mu_);
stop_ = true;
}
wake_.notify_all();
thread_.join();
std::fclose(f_);
}
bool Recorder::open(const std::string &dir, std::string &err) {
if (mkdir(dir.c_str(), 0755) < 0 && errno != EEXIST) return err = dir + ": " + std::strerror(errno), false;
const std::string path = dir + "/sets.bin";
f_ = std::fopen(path.c_str(), "wbx"); // never overwrite a recording
if (!f_) return err = path + ": " + std::strerror(errno), false;
thread_ = std::thread(&Recorder::loop, this);
return true;
}
void Recorder::add(const std::vector<SetFrame> &frames) {
size_t bytes = sizeof(fh_set_hdr_t) + frames.size() * sizeof(fh_set_cam_t);
for (const SetFrame &s : frames) bytes += size_t(s.width) * s.height;
std::vector<uint8_t> rec(bytes);
fh_set_hdr_t h{};
std::memcpy(h.magic, FH_SET_MAGIC, 8);
h.ncams = uint32_t(frames.size());
h.bytes = uint32_t(bytes);
std::memcpy(rec.data(), &h, sizeof h);
uint8_t *p = rec.data() + sizeof h;
for (const SetFrame &s : frames) {
fh_set_cam_t c{};
std::strncpy(c.name, s.name.c_str(), sizeof c.name - 1);
c.width = s.width, c.height = s.height, c.capture_ns = s.capture_ns, c.dqbuf_ns = s.dqbuf_ns;
std::memcpy(p, &c, sizeof c);
p += sizeof c;
}
for (const SetFrame &s : frames) {
std::memcpy(p, s.px, size_t(s.width) * s.height);
p += size_t(s.width) * s.height;
}
{
std::lock_guard<std::mutex> l(mu_);
if (queue_.size() >= kMaxQueued) {
++dropped_;
return;
}
queue_.push_back(std::move(rec));
}
wake_.notify_one();
}
void Recorder::loop() {
std::unique_lock<std::mutex> l(mu_);
for (;;) {
wake_.wait(l, [&] { return stop_ || !queue_.empty(); });
if (queue_.empty()) return; // stopping, and everything is written
std::vector<uint8_t> rec = std::move(queue_.front());
queue_.pop_front();
l.unlock();
const bool ok = std::fwrite(rec.data(), 1, rec.size(), f_) == rec.size();
l.lock();
ok ? ++written_ : ++dropped_;
}
}
SetReader::~SetReader() {
if (f_) std::fclose(f_);
}
bool SetReader::open(const std::string &dir, std::string &err) {
const std::string path = dir + "/sets.bin";
f_ = std::fopen(path.c_str(), "rb");
return f_ ? true : (err = path + ": " + std::strerror(errno), false);
}
bool SetReader::next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels) {
fh_set_hdr_t h;
if (std::fread(&h, sizeof h, 1, f_) != 1 || std::memcmp(h.magic, FH_SET_MAGIC, 8) || h.ncams == 0 || h.ncams > 16)
return false;
cams.resize(h.ncams);
if (std::fread(cams.data(), sizeof(fh_set_cam_t), h.ncams, f_) != h.ncams) return false;
pixels.resize(h.ncams);
for (uint32_t i = 0; i < h.ncams; ++i) {
cams[i].name[sizeof cams[i].name - 1] = 0;
pixels[i].resize(size_t(cams[i].width) * cams[i].height);
if (std::fread(pixels[i].data(), 1, pixels[i].size(), f_) != pixels[i].size()) return false;
}
return true;
}
+69
View File
@@ -0,0 +1,69 @@
// Recordings of frame sets, for replaying live sessions through the tracker offline
// (fh-replay). A recording is DIR/sets.bin: one record per frame set, each
// fh_set_hdr_t, then per camera fh_set_cam_t, then each camera's pixels (w x h, packed)
// in the same camera order.
#pragma once
#include <condition_variable>
#include <cstdint>
#include <cstdio>
#include <deque>
#include <mutex>
#include <string>
#include <thread>
#include <vector>
#define FH_SET_MAGIC "FHSET01"
struct fh_set_hdr_t {
char magic[8];
uint32_t ncams;
uint32_t bytes; // the whole record, this header included
};
struct fh_set_cam_t {
char name[16]; // calibration name, e.g. "slam_left"
uint32_t width, height;
uint64_t capture_ns; // CLOCK_MONOTONIC_RAW, as the ring has it
uint64_t dqbuf_ns; // CLOCK_MONOTONIC
};
struct SetFrame {
std::string name;
const uint8_t *px;
uint32_t width, height;
uint64_t capture_ns, dqbuf_ns;
};
// Writes sets on its own thread, so a slow disk never holds up tracking; drops sets
// when too many are waiting.
class Recorder {
public:
~Recorder();
bool open(const std::string &dir, std::string &err);
void add(const std::vector<SetFrame> &frames);
size_t written() const { return written_; }
size_t dropped() const { return dropped_; }
private:
void loop();
FILE *f_ = nullptr;
std::thread thread_;
std::mutex mu_;
std::condition_variable wake_;
std::deque<std::vector<uint8_t>> queue_;
bool stop_ = false;
size_t written_ = 0, dropped_ = 0;
};
// Reads a recording back one set at a time.
class SetReader {
public:
bool open(const std::string &dir, std::string &err);
// False at the end (or on a truncated last set).
bool next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels);
~SetReader();
private:
FILE *f_ = nullptr;
};
+240
View File
@@ -0,0 +1,240 @@
// fh-replay: run a recording (fh-tracker --record) through the tracker offline, with the
// live scheduling, and report how well it kept the hands.
//
// fh-replay DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR]
// [--from S] [--to S] [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch)
//
// --oracle N: every N-th set, also search every tile of every camera (slow), to see
// which hands were there to find. Compares that with what the tracker had.
// --slow F: the live tracker skips the sets that arrive while it's busy; replay takes
// each step's time here times F as the busy time (the headset is busier live).
// --cost: instead of timing the steps, charge each round of model calls what it
// typically costs live (10 ms landmarks, 18 ms palms): repeatable results.
// --timeline: per processed set, a line per hand (time, id, side, views, wrist) and per view
// (hand, camera, presence, next crop, set index).
// --keep-presence P: landmark presence a tracked view needs to stay (default 0.5, as new ones).
// --contrast: how the palm search's and the landmark model's crops are equalized
// (default clahe:2/none, as fh-tracker).
#include "record.h"
#include "tracker.h"
#include <algorithm>
#include <chrono>
#include <cstdio>
#include <cstring>
#include <map>
#include <set>
#include <string>
#include <vector>
namespace {
struct Track {
double first = 0, last = 0;
int sets = 0, left = 0;
// the last two palm positions (raw, smoothed) and times, for the jitter measure
V3 raw[2]{}, sm[2]{};
double t[2]{};
int line = 0; // updates on the current unbroken run
};
double median(std::vector<double> v) {
if (v.empty()) return 0;
std::nth_element(v.begin(), v.begin() + v.size() / 2, v.end());
return v[v.size() / 2];
}
} // namespace
int main(int argc, char **argv) {
if (argc < 2 || argv[1][0] == '-') {
std::printf("usage: %s DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR] [--from S] [--to S]\n", argv[0]);
return 1;
}
const std::string dir = argv[1];
int oracle = 0, threads = 2;
double slow = 1.0, from = 0, to = 1e9;
bool cost = false;
Contrast palm_contrast, hand_contrast{Contrast::None}; // as fh-tracker's
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
std::string timeline, models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
for (int i = 2; i < argc; ++i) {
const std::string a = argv[i];
const bool more = i + 1 < argc;
if (a == "--oracle" && more) oracle = std::atoi(argv[++i]);
else if (a == "--slow" && more) slow = std::atof(argv[++i]);
else if (a == "--timeline" && more) timeline = argv[++i];
else if (a == "--threads" && more) threads = std::atoi(argv[++i]);
else if (a == "--models" && more) models = argv[++i];
else if (a == "--cost") cost = true;
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
else if (a == "--contrast" && more) {
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
}
else if (a == "--from" && more) from = std::atof(argv[++i]);
else if (a == "--to" && more) to = std::atof(argv[++i]);
else return std::fprintf(stderr, "unknown option %s\n", a.c_str()), 1;
}
std::string err;
std::map<std::string, Camera> calib;
Nets nets;
SetReader in;
if (!load_calibration(calib, err) || !nets.load(models, false, err) || !in.open(dir, err))
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
nets.set_contrast(palm_contrast, hand_contrast);
FILE *tl = timeline.empty() ? nullptr : std::fopen(timeline.c_str(), "w");
std::vector<fh_set_cam_t> cams;
std::vector<std::vector<uint8_t>> px;
if (!in.next(cams, px)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
std::map<std::string, Camera> used;
for (auto &c : cams)
if (calib.count(c.name)) used[c.name] = calib[c.name];
Pool pool(threads, {2, 3, 4});
Tracker tracker(used, nets, pool);
tracker.set_keep_presence(keep_presence);
uint64_t t0 = 0, busy_until = 0, next_ns = 0;
int index = -1; // of the set in the recording
int nsets = 0, processed = 0, left = 0, right = 0, both = 0, hist[3] = {};
std::map<int, Track> tracks;
std::vector<const Hand *> last_out;
// oracle: sets where a side's hand was findable, and where the tracker had it then
int o_sets = 0, o_left = 0, o_right = 0, o_left_hit = 0, o_right_hit = 0, o_left_extra = 0, o_right_extra = 0;
std::map<std::string, int> o_by_cam;
double busy_ms = 0;
// jitter: how far each update's palm is from a straight line through the last two,
// mm (steady motion cancels out; what's left is noise and real acceleration)
std::vector<double> jit_raw, jit_sm;
int near_face = 0, hand_updates = 0; // published palms within 20 cm of the eyes
do {
std::map<std::string, Image> images;
uint64_t t = UINT64_MAX;
for (size_t i = 0; i < cams.size(); ++i) {
if (!used.count(cams[i].name)) continue;
images[cams[i].name] = {px[i].data(), int(cams[i].width), int(cams[i].height), int(cams[i].width)};
t = std::min(t, cams[i].capture_ns);
}
if (!t0) t0 = t;
const double ts = (t - t0) / 1e9;
++index;
if (ts < from) continue;
if (ts > to) break;
++nsets;
if (t >= busy_until && t >= next_ns) {
const auto w0 = std::chrono::steady_clock::now();
const Stats before = tracker.stats;
const auto out = tracker.step(images, int64_t(t));
double ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - w0).count();
if (cost) { // repeatable: rounds of model calls at typical live costs, per thread
const int hands = tracker.stats.hand_calls - before.hand_calls, palms = tracker.stats.palm_calls - before.palm_calls;
ms = (2 + 10.0 * ((hands + threads - 1) / threads) + 18.0 * ((palms + threads - 1) / threads)) / slow;
}
busy_ms += ms;
busy_until = t + uint64_t(ms * slow * 1e6) + 3'000'000; // + the ring hand-off
next_ns = t + uint64_t((tracker.interval() - 0.005) * 1e9);
++processed;
last_out = out;
bool l = false, r = false;
for (const Hand *h : out) {
(h->pts[0][0] < 0 ? l : r) = true;
Track &tr = tracks[h->id];
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
const V3 raw = palm(h->pts), sm = palm(h->smooth);
++hand_updates, near_face += norm(sm) < 0.2;
if (tr.line && ts - tr.t[0] >= 0.1) tr.line = 0; // a gap: the line starts over
if (tr.line >= 2 && tr.t[0] - tr.t[1] > 1e-3) {
const double k = (ts - tr.t[0]) / (tr.t[0] - tr.t[1]);
jit_raw.push_back(norm(raw - tr.raw[0] - (tr.raw[0] - tr.raw[1]) * k) * 1000);
jit_sm.push_back(norm(sm - tr.sm[0] - (tr.sm[0] - tr.sm[1]) * k) * 1000);
}
tr.raw[1] = tr.raw[0], tr.sm[1] = tr.sm[0], tr.t[1] = tr.t[0];
tr.raw[0] = raw, tr.sm[0] = sm, tr.t[0] = ts;
++tr.line;
if (!tr.sets) tr.first = ts;
tr.last = ts, ++tr.sets, tr.left += h->pts[0][0] < 0;
if (tl)
std::fprintf(tl, "%.3f %d %s %d %+.3f %+.3f %+.3f\n", ts, h->id, h->pts[0][0] < 0 ? "L" : "R", h->nviews,
h->pts[0][0], h->pts[0][1], h->pts[0][2]);
}
if (tl && out.empty()) std::fprintf(tl, "%.3f -\n", ts);
if (tl)
for (const Seen &v : tracker.views_now())
std::fprintf(tl, "%.3f view %d %s presence %.2f roi %.0f %.0f %.0f %.3f set %d\n", ts, v.hand, v.cam.c_str(),
v.lm.presence, v.roi.center[0], v.roi.center[1], v.roi.size, v.roi.rotation, index);
left += l, right += r, both += l && r;
++hist[std::min<size_t>(out.size(), 2)];
}
if (oracle > 0 && nsets % oracle == 0) {
const Stats keep = tracker.stats;
const auto seen = tracker.exhaustive(images);
tracker.stats = keep;
bool l = false, r = false;
for (const Seen &s : seen) {
(s.wrist[0] < 0 ? l : r) = true;
++o_by_cam[s.cam + (s.wrist[0] < 0 ? " L" : " R")];
}
bool tl_ = false, tr_ = false;
for (const Hand *h : last_out) (h->pts[0][0] < 0 ? tl_ : tr_) = true;
++o_sets;
o_left += l, o_right += r;
o_left_hit += l && tl_, o_right_hit += r && tr_;
o_left_extra += !l && tl_, o_right_extra += !r && tr_;
if (tl && ((!l && tl_) || (!r && tr_))) std::fprintf(tl, "%.3f oracle-extra %s%s set %d\n", ts, !l && tl_ ? "L" : "", !r && tr_ ? "R" : "", index);
}
} while (in.next(cams, px));
if (tl) std::fclose(tl);
const double secs = nsets > 1 ? nsets / 30.0 : 0;
const Stats &s = tracker.stats;
std::printf("%s: %d sets (%.0f s), processed %d (%.1f/s), %.1f ms per step\n", dir.c_str(), nsets, secs, processed,
processed / std::max(secs, 1e-9), busy_ms / std::max(processed, 1));
std::printf("hands per processed set: 0 %.0f%%, 1 %.0f%%, 2 %.0f%%; a hand on the left %.0f%%, right %.0f%%, both %.0f%%\n",
100.0 * hist[0] / processed, 100.0 * hist[1] / processed, 100.0 * hist[2] / processed,
100.0 * left / processed, 100.0 * right / processed, 100.0 * both / processed);
std::vector<double> lens[2];
for (auto &[id, tr] : tracks) lens[tr.left * 2 > tr.sets ? 0 : 1].push_back(tr.last - tr.first);
for (int k = 0; k < 2; ++k) {
double total = 0;
for (double d : lens[k]) total += d;
std::printf("%s tracks: %zu, median %.1f s, total %.0f s\n", k ? "right" : "left ", lens[k].size(), median(lens[k]), total);
}
std::printf("views lost %d, handoff misses %d, dups %d, splits %d; hands new %d, merged %d, forgotten %d\n", s.lost,
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten);
std::printf("model calls: palm %d (%.1f/s), hand %d (%.1f/s)\n", s.palm_calls, s.palm_calls / std::max(secs, 1e-9),
s.hand_calls, s.hand_calls / std::max(secs, 1e-9));
{
std::vector<double> r, step;
std::map<int, double> prev;
for (auto &[id, x] : s.mono_ratio) {
r.push_back(x);
if (prev.count(id)) step.push_back(std::fabs(x - prev[id]));
prev[id] = x;
}
std::sort(r.begin(), r.end());
std::sort(step.begin(), step.end());
if (!r.empty())
std::printf("single-view distance / stereo: 10%% %.2f, median %.2f, 90%% %.2f; change between frames median %.3f, 90%% %.3f\n",
r[r.size() / 10], r[r.size() / 2], r[r.size() * 9 / 10], step[step.size() / 2], step[step.size() * 9 / 10]);
}
std::printf("palms within 20 cm of the eyes: %d of %d hand updates\n", near_face, hand_updates);
std::sort(jit_raw.begin(), jit_raw.end());
std::sort(jit_sm.begin(), jit_sm.end());
if (!jit_raw.empty())
std::printf("palm jitter (off a straight line through the last two updates): measured median %.1f mm, 90%% %.1f mm; "
"published median %.1f mm, 90%% %.1f mm\n", jit_raw[jit_raw.size() / 2], jit_raw[jit_raw.size() * 9 / 10],
jit_sm[jit_sm.size() / 2], jit_sm[jit_sm.size() * 9 / 10]);
if (o_sets) {
std::printf("oracle, %d sets: a left hand findable in %d, the tracker had it in %d (%.0f%%); right %d, had %d (%.0f%%)\n",
o_sets, o_left, o_left_hit, 100.0 * o_left_hit / std::max(o_left, 1), o_right, o_right_hit,
100.0 * o_right_hit / std::max(o_right, 1));
std::printf(" tracker had a hand the full search didn't find: left %d, right %d\n", o_left_extra, o_right_extra);
std::printf(" found by camera:");
for (auto &[k, n] : o_by_cam) std::printf(" %s %d", k.c_str(), n);
std::printf("\n");
}
return 0;
}
+607
View File
@@ -0,0 +1,607 @@
#include "tracker.h"
#include <pthread.h>
#include <sched.h>
#include <algorithm>
#include <chrono>
#include <set>
namespace {
// where arms start, head frame: below and slightly behind the eyes
const V3 kShoulders[2] = {{0.17, -0.25, 0.08}, {-0.17, -0.25, 0.08}};
// landmark pairs across the palm, rigid enough for single-view depth
const int kPalmPairs[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 17}, {5, 13}, {9, 17}, {1, 17}, {1, 5}};
constexpr double kFastSpeed = 0.25; // m/s
constexpr double kSearchInterval = 0.2;
// One Euro filter on the published landmarks: still hands are smoothed hard (tracking
// noise is a few mm per frame), fast ones barely, so they don't lag.
constexpr double kMinCutoff = 2.0; // Hz, a still hand
constexpr double kBeta = 30.0; // Hz more per m/s of palm speed
constexpr double kSpeedCutoff = 1.5; // Hz, for the palm speed itself
// With one view, the hand's distance from the camera comes from how big it looks, which
// is off by 10-30% and wanders ~10% between frames. Its direction is exact. So a hand that
// was just located keeps its distance, drifting toward the one-view guess by this much a frame.
constexpr double kMonoDepthGain = 0.1;
// Is a triangulated hand as far from each camera as its apparent size says? With the
// model's average hand, clean stereo pairs measure 0.71-1.51 times the one-view distance
// (5-95%, median 1.16); pairs of two different hands mostly far less.
constexpr double kSizePrior = 1.16, kRatioLo = 0.6, kRatioHi = 1.9;
double ms_since(std::chrono::steady_clock::time_point t) {
return std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t).count();
}
bool is_slam(const Camera &c) { return c.name.rfind("slam", 0) == 0; }
V2 palm_centre(const Landmarks &lm) { return lm.pts[9]; }
// Two views in one camera on the same hand: the landmark model puts the same points on
// it from both crops, even when the crops differ.
bool same_hand(const Landmarks &a, const Landmarks &b, double size) {
double d = 0;
for (int i = 0; i < 21; ++i) d += norm(a.pts[i] - b.pts[i]) / 21;
return norm(palm_centre(a) - palm_centre(b)) < 0.5 * size || d < 0.25 * size;
}
double hand_size(const Landmarks &lm) {
double lo[2] = {1e9, 1e9}, hi[2] = {-1e9, -1e9};
for (const V2 &p : lm.pts)
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[k]), hi[k] = std::max(hi[k], p[k]);
return std::max(hi[0] - lo[0], hi[1] - lo[1]);
}
} // namespace
// ---------------------------------------------------------------------------- pool
Pool::Pool(int threads, const std::vector<int> &cpus) {
for (int i = 0; i < threads; ++i) threads_.emplace_back(&Pool::loop, this, cpus[i % cpus.size()]);
}
Pool::~Pool() {
{
std::lock_guard<std::mutex> l(mu_);
stop_ = true;
}
wake_.notify_all();
for (auto &t : threads_) t.join();
}
void Pool::loop(int cpu) {
cpu_set_t set;
CPU_ZERO(&set);
CPU_SET(cpu, &set);
pthread_setaffinity_np(pthread_self(), sizeof set, &set); // ignored if not allowed
std::unique_lock<std::mutex> l(mu_);
for (;;) {
wake_.wait(l, [&] { return stop_ || (jobs_ && next_ < jobs_->size()); });
if (stop_) return;
auto &job = (*jobs_)[next_++];
l.unlock();
job();
l.lock();
if (++finished_ == jobs_->size()) done_.notify_all();
}
}
void Pool::run(std::vector<std::function<void()>> &jobs) {
if (jobs.empty()) return;
std::unique_lock<std::mutex> l(mu_);
jobs_ = &jobs, next_ = 0, finished_ = 0;
wake_.notify_all();
done_.wait(l, [&] { return finished_ == jobs.size(); });
jobs_ = nullptr;
}
// ------------------------------------------------------------------------- tracker
Tracker::Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views)
: nets_(nets), pool_(pool), max_views_(max_views) {
for (const auto &[name, cam] : cams) {
cams_[name] = &cam;
if (is_slam(cam)) {
add_tiles(cam, 0.45, 3, 3);
add_tiles(cam, 0.65, 2, 2);
add_tiles(cam, 1.0, 1, 1); // hands close to the face fill much of the frame
} else {
add_tiles(cam, 0.6, 3, 2);
add_tiles(cam, 1.0, 1, 1);
}
}
}
void Tracker::add_tiles(const Camera &cam, double frac, int gx, int gy) {
const double s = frac * std::max(cam.width, cam.height);
for (int i = 0; i < gx; ++i)
for (int j = 0; j < gy; ++j) {
const double x = gx > 1 ? s / 2 + (cam.width - s) * i / (gx - 1) : cam.width / 2.0;
const double y = gy > 1 ? s / 2 + (cam.height - s) * j / (gy - 1) : cam.height / 2.0;
Tile t{&cam, {x, y}, s, 0, 0};
// turn the crop so the expected shoulder-to-hand direction points up
const V3 ray = cam.ray(t.center), p = cam.origin + ray * 0.45;
const V3 d = unit(p - kShoulders[p[0] > 0 ? 0 : 1]);
const V2 a = cam.project(p, nullptr), b = cam.project(p + d * 0.05, nullptr);
t.rotation = std::atan2(b[0] - a[0], -(b[1] - a[1]));
t.weight = std::max(0.15, dot(ray, unit(V3{0, -0.45, -0.9})));
tiles_.push_back(t);
}
}
double Tracker::interval() const {
double fastest = -1;
for (const auto &[id, h] : hands_)
if (h.seen_ns == last_ns_) fastest = std::max(fastest, norm(h.dpalm)); // filtered: noise isn't speed
return fastest < 0 ? 1 / 5.0 : fastest > kFastSpeed ? 1 / 30.0 : 1 / 15.0;
}
bool Tracker::inside(const Camera &cam, V2 uv) const {
const double m = 0.12;
return uv[0] >= m * cam.width && uv[0] <= (1 - m) * cam.width && uv[1] >= m * cam.height &&
uv[1] <= (1 - m) * cam.height && cam.off_axis(uv) < (is_slam(cam) ? 80.0 : 75.0);
}
void Tracker::run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views) {
if (views.empty()) return;
const auto t0 = std::chrono::steady_clock::now();
std::vector<std::function<void()>> jobs;
for (View *v : views) {
const Image &img = images.at(v->cam->name);
jobs.push_back([this, v, &img] {
v->lm = nets_.landmarks(img, v->roi);
v->has_lm = true;
v->fresh = true;
});
}
pool_.run(jobs);
stats.hand_calls += int(views.size());
stats.hand_ms += ms_since(t0);
++stats.hand_batches;
}
bool Tracker::single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const {
V3 rays[21];
for (int i = 0; i < 21; ++i) rays[i] = cam.ray(lm.pts[i]);
std::vector<std::pair<double, double>> est; // (depth, weight)
for (const auto &pr : kPalmPairs) {
const int i = pr[0], j = pr[1];
const double d = std::hypot(lm.world[i][0] - lm.world[j][0], lm.world[i][1] - lm.world[j][1]) * scale;
const double a = std::acos(std::clamp(dot(rays[i], rays[j]), -1.0, 1.0));
if (a > 1e-3 && d > 0.01) est.push_back({d / a, d});
}
if (est.empty()) return false;
std::sort(est.begin(), est.end());
double total = 0, acc = 0, depth = est.back().first;
for (auto &e : est) total += e.second;
for (auto &e : est)
if ((acc += e.second) >= total / 2) { depth = e.first; break; }
double zmean = 0;
for (int i = 0; i < 21; ++i) zmean += lm.world[i][2] / 21;
for (int i = 0; i < 21; ++i) out[i] = cam.origin + rays[i] * (depth + (lm.world[i][2] - zmean) * scale);
return true;
}
// A triangulated hand is in front of each camera, as far as its apparent size says (see
// kSizePrior). Returns how far off that is (the sum of |log| ratios), or -1 if implausible.
double Tracker::size_misfit(const std::vector<const View *> &views, const V3 *pts) const {
double misfit = 0;
for (const View *v : views) {
V3 mono[21];
// along the view's own ray (fisheye: a hand near the image edge is far off the axis)
if (dot(pts[9] - v->cam->origin, v->cam->ray(v->lm.pts[9])) < 0.08) return -1;
if (!single_view(*v->cam, v->lm, 1.0, mono)) continue;
const double r = norm(pts[9] - v->cam->origin) / norm(mono[9] - v->cam->origin);
if (r < kRatioLo || r > kRatioHi) return -1;
misfit += std::fabs(std::log(r / kSizePrior));
}
return misfit;
}
bool Tracker::hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns) {
views.erase(std::remove_if(views.begin(), views.end(), [](View *v) { return !v->has_lm || !v->fresh; }), views.end());
if (views.empty()) return false;
if (views.size() >= 2) {
const int n = int(views.size());
std::vector<V3> origins(n), dirs(n);
std::vector<double> w(n), res(21);
V3 pts[21];
for (int k = 0; k < 21; ++k) {
for (int v = 0; v < n; ++v) {
origins[v] = views[v]->cam->origin;
dirs[v] = views[v]->cam->ray(views[v]->lm.pts[k]);
w[v] = views[v]->lm.presence;
}
pts[k] = triangulate(origins.data(), dirs.data(), w.data(), n, &res[k]);
}
std::nth_element(res.begin(), res.begin() + 10, res.end());
const double residual = res[10];
// the views disagree: two different hands; keep the stronger. Rays to two different
// hands can pass close to each other near the cameras, so check the distance too.
if (residual > 0.03 || size_misfit({views.begin(), views.end()}, pts) < 0) {
View *best = *std::max_element(views.begin(), views.end(),
[](View *a, View *b) { return a->lm.presence < b->lm.presence; });
for (View *v : views)
if (v != best) v->hand = -1;
++stats.splits;
return hand_3d(hand, {best}, t_ns);
}
// learn how big this user's hand is compared to the model's average hand
std::vector<double> t, m;
const Landmarks &ref = views[0]->lm;
for (const auto &pr : kPalmPairs) {
t.push_back(norm(pts[pr[0]] - pts[pr[1]]));
m.push_back(norm(V3{ref.world[pr[0]][0], ref.world[pr[0]][1], ref.world[pr[0]][2]} -
V3{ref.world[pr[1]][0], ref.world[pr[1]][1], ref.world[pr[1]][2]}));
}
std::nth_element(t.begin(), t.begin() + t.size() / 2, t.end());
std::nth_element(m.begin(), m.begin() + m.size() / 2, m.end());
if (m[m.size() / 2] > 0 && residual < 0.008) // only from clean matches
hand.scale += 0.1 * (std::clamp(t[t.size() / 2] / m[m.size() / 2], 0.8, 1.6) - hand.scale);
std::copy(pts, pts + 21, hand.pts);
hand.residual = residual;
for (View *v : views) {
V3 mono[21];
if (!single_view(*v->cam, v->lm, hand.scale, mono)) continue;
const V3 o = v->cam->origin;
stats.mono_ratio.push_back({hand.id, norm(mono[9] - o) / norm(pts[9] - o)});
}
} else {
const Camera &cam = *views[0]->cam;
V3 pts[21];
if (!single_view(cam, views[0]->lm, hand.scale, pts)) return false;
const double guess = norm(pts[9] - cam.origin);
if (hand.has_pts && t_ns - hand.seen_ns < 300'000'000 && guess > 0) {
const double was = norm(hand.pts[9] - cam.origin), d = was + kMonoDepthGain * (guess - was);
for (V3 &p : pts) p = cam.origin + (p - cam.origin) * (d / guess);
}
std::copy(pts, pts + 21, hand.pts);
hand.residual = -1;
}
hand.has_pts = true;
hand.nviews = int(views.size());
for (View *v : views) hand.right_score += 0.2 * (v->lm.right - hand.right_score);
return true;
}
// How badly two views in different cameras fit one hand: the rays should meet, each view's
// apparent size should match its distance, and the model should call both the same hand
// (left or right). Negative if they can't be one hand. Side by side hands sit on the same
// epipolar lines of the side cameras, so the distance check is what tells them apart.
double Tracker::pair_cost(const View &a, const View &b) const {
V3 pts[21];
std::vector<double> res(21);
for (int k = 0; k < 21; ++k) {
const V3 o[2] = {a.cam->origin, b.cam->origin};
const V3 d[2] = {a.cam->ray(a.lm.pts[k]), b.cam->ray(b.lm.pts[k])};
const double w[2] = {a.lm.presence, b.lm.presence};
pts[k] = triangulate(o, d, w, 2, &res[k]);
}
std::nth_element(res.begin(), res.begin() + 10, res.end());
if (res[10] > 0.03) return -1;
const double misfit = size_misfit({&a, &b}, pts);
return misfit < 0 ? -1 : res[10] / 0.01 + misfit + std::fabs(a.lm.right - b.lm.right);
}
// Which views in two cameras are the same hand: every way of pairing them up (a few views
// each), scored with pair_cost. Keeps the hands' pairing unless another is clearly better,
// then relabels the views, keeping the longer-tracked hand's id.
void Tracker::associate() {
constexpr double kPairBonus = 2.0, kBetter = 0.3;
std::vector<const Camera *> cams;
for (View &v : views_)
if (std::find(cams.begin(), cams.end(), v.cam) == cams.end()) cams.push_back(v.cam);
std::sort(cams.begin(), cams.end(), [](const Camera *a, const Camera *b) { return a->name < b->name; });
for (size_t i = 0; i < cams.size(); ++i)
for (size_t j = i + 1; j < cams.size(); ++j) {
std::vector<View *> A, B;
for (View &v : views_) {
if (!v.fresh) continue;
if (v.cam == cams[i]) A.push_back(&v);
else if (v.cam == cams[j]) B.push_back(&v);
}
if (A.empty() || B.empty() || A.size() > 3 || B.size() > 3) continue;
std::vector<std::vector<double>> c(A.size(), std::vector<double>(B.size()));
for (size_t x = 0; x < A.size(); ++x)
for (size_t y = 0; y < B.size(); ++y) c[x][y] = pair_cost(*A[x], *B[y]);
auto score = [&](const std::vector<int> &m) { // m[x]: A[x]'s partner in B, or -1
double s = 0;
for (size_t x = 0; x < A.size(); ++x)
if (m[x] >= 0 && c[x][m[x]] >= 0) s += c[x][m[x]] - kPairBonus;
return s;
};
std::vector<int> cur(A.size(), -1);
for (size_t x = 0; x < A.size(); ++x)
for (size_t y = 0; y < B.size(); ++y)
if (A[x]->hand == B[y]->hand) cur[x] = int(y);
std::vector<int> best = cur, m(A.size(), -1);
double best_score = score(cur);
const double cur_score = best_score;
std::function<void(size_t, unsigned)> walk = [&](size_t x, unsigned used) {
if (x == A.size()) {
const double sc = score(m);
if (sc < best_score) best_score = sc, best = m;
return;
}
m[x] = -1;
walk(x + 1, used);
for (size_t y = 0; y < B.size(); ++y)
if (!(used >> y & 1) && c[x][y] >= 0) {
m[x] = int(y);
walk(x + 1, used | 1u << y);
}
m[x] = -1;
};
walk(0, 0);
if (best == cur || best_score > cur_score - kBetter) continue;
auto frames = [&](int id) {
const auto h = hands_.find(id);
return h == hands_.end() ? -1 : h->second.frames;
};
for (size_t x = 0; x < A.size(); ++x) {
if (best[x] < 0) continue;
View *a = A[x], *b = B[best[x]];
int id = a->hand;
bool free = true; // b's hand isn't another A view's
for (size_t x2 = 0; x2 < A.size(); ++x2) free = free && (x2 == x || A[x2]->hand != b->hand);
if (free && frames(b->hand) > frames(id)) id = b->hand;
a->hand = b->hand = id;
}
// a B view left unpaired that still shares a hand with an A view starts its own
for (size_t y = 0; y < B.size(); ++y) {
if (std::find(best.begin(), best.end(), int(y)) != best.end()) continue;
bool shared = false;
for (View *a : A) shared = shared || a->hand == B[y]->hand;
if (!shared) continue;
B[y]->hand = next_id_++;
hands_[B[y]->hand].id = B[y]->hand;
++stats.created;
}
++stats.merged;
}
}
std::vector<const Hand *> Tracker::step(const std::map<std::string, Image> &images, int64_t t_ns) {
const auto t_step = std::chrono::steady_clock::now();
++stats.sets;
std::vector<View> live;
for (View &v : views_)
if (images.count(v.cam->name)) live.push_back(v), live.back().fresh = false;
// 1. hand-over: give hands with too few views a crop in other cameras
for (auto &[id, hand] : hands_) {
if (!hand.has_pts) continue;
std::set<std::string> have;
for (View &v : live)
if (v.hand == id) have.insert(v.cam->name);
if (int(have.size()) >= max_views_) continue;
std::vector<std::pair<double, View>> options;
for (auto &[name, cam] : cams_) {
if (have.count(name) || !images.count(name)) continue;
V2 uv[21];
bool front = true;
for (int k = 0; k < 21; ++k) {
double z;
uv[k] = cam->project(hand.pts[k], &z);
front = front && z > 0;
}
const V2 centre = (uv[0] + uv[5] + uv[9] + uv[13] + uv[17]) * 0.2;
if (!front || !inside(*cam, centre)) continue;
View v{cam, roi_from_points(uv), id, {}, false, 0};
options.push_back({cam->off_axis(centre), v});
}
std::sort(options.begin(), options.end(), [](auto &a, auto &b) { return a.first < b.first; });
for (size_t k = 0; k < options.size() && int(have.size() + k) < max_views_; ++k) live.push_back(options[k].second);
}
// 2. the landmark model on each hand's best views, within budget
std::map<int, std::vector<View *>> by_hand;
for (View &v : live) by_hand[v.hand].push_back(&v);
std::vector<View *> chosen;
for (auto &[id, vs] : by_hand) {
std::sort(vs.begin(), vs.end(), [](View *a, View *b) {
if (a->has_lm != b->has_lm) return a->has_lm;
return a->cam->off_axis(a->roi.center) < b->cam->off_axis(b->roi.center);
});
for (int k = 0; k < int(vs.size()) && k < max_views_; ++k) chosen.push_back(vs[k]);
}
std::stable_sort(chosen.begin(), chosen.end(), [](View *a, View *b) { return a->has_lm > b->has_lm; });
if (int(chosen.size()) > hand_budget_) chosen.resize(hand_budget_);
run_landmarks(images, chosen);
std::vector<View *> kept;
for (View *v : chosen)
if (v->lm.presence >= (v->frames > 0 ? keep_presence_ : min_presence_)) {
v->roi = v->lm.next_roi();
++v->frames;
kept.push_back(v);
} else {
++(v->frames > 0 ? stats.lost : stats.handoff_miss);
}
// the same hand twice in one camera: keep the more confident
std::sort(kept.begin(), kept.end(), [](View *a, View *b) { return a->lm.presence > b->lm.presence; });
std::vector<View> next;
for (View *v : kept) {
const double size = hand_size(v->lm);
bool dup = false;
for (View &o : next) dup = dup || (o.cam == v->cam && same_hand(o.lm, v->lm, size));
if (!dup) next.push_back(*v);
else ++stats.dups;
}
views_ = next;
// 3. search for missing hands
std::set<int> tracked;
for (View &v : views_) tracked.insert(v.hand);
if (tracked.size() < 2 && (t_ns - search_ns_) / 1e9 >= kSearchInterval - 0.01) {
search_ns_ = t_ns;
const int budget = tracked.empty() ? search_budget_ : std::max(1, search_budget_ - 1);
for (Tile &t : tiles_)
if (images.count(t.cam->name)) t.credit += t.weight;
std::vector<Tile *> picked;
for (int b = 0; b < budget; ++b) {
Tile *best = nullptr;
for (Tile &t : tiles_)
if (images.count(t.cam->name) && std::find(picked.begin(), picked.end(), &t) == picked.end() &&
(!best || t.credit > best->credit))
best = &t;
if (!best) break;
best->credit = 0;
picked.push_back(best);
}
const auto t0 = std::chrono::steady_clock::now();
std::vector<std::vector<Palm>> found(picked.size());
std::vector<std::function<void()>> jobs;
for (size_t i = 0; i < picked.size(); ++i) {
Tile *t = picked[i];
const Image &img = images.at(t->cam->name);
jobs.push_back([this, t, &img, &found, i] { found[i] = nets_.palms(img, t->center, t->size, t->rotation); });
}
pool_.run(jobs);
stats.palm_calls += int(picked.size());
stats.palm_ms += ms_since(t0);
++stats.palm_batches;
std::vector<View> fresh;
for (size_t i = 0; i < picked.size(); ++i)
for (const Palm &p : found[i]) {
const Roi roi = p.roi();
bool near = false;
for (auto *list : {&views_, &fresh})
for (View &v : *list) near = near || (v.cam == picked[i]->cam && norm(v.roi.center - roi.center) < 0.5 * roi.size);
if (!near) fresh.push_back({picked[i]->cam, roi, 0, {}, false, 0});
}
std::vector<View *> ptrs;
for (View &v : fresh) ptrs.push_back(&v);
run_landmarks(images, ptrs);
for (View &v : fresh)
if (v.lm.presence >= min_presence_) {
v.roi = v.lm.next_roi();
v.frames = 1;
views_.push_back(v);
}
}
// 4. give new views a hand: the nearest existing hand in 3D, else a new one
for (View &v : views_) {
if (v.hand > 0 && hands_.count(v.hand)) continue;
V3 guess[21];
const bool have_guess = single_view(*v.cam, v.lm, 1.0, guess);
int best = 0;
double dist = 0.12;
for (auto &[id, h] : hands_) {
if (!h.has_pts) continue;
bool same_cam = false;
for (View &o : views_) same_cam = same_cam || (o.hand == id && o.cam == v.cam);
if (same_cam) continue;
const double d = have_guess ? norm(h.pts[9] - guess[9]) : 1e9;
if (d < dist) best = id, dist = d;
}
if (!best) {
best = next_id_++;
hands_[best].id = best;
++stats.created;
}
v.hand = best;
}
// 5. which views in different cameras are the same hand
associate();
// 6. 3D for every hand seen now; forget hands not seen for a while
std::vector<const Hand *> out;
for (auto it = hands_.begin(); it != hands_.end();) {
Hand &h = it->second;
std::vector<View *> vs;
for (View &v : views_)
if (v.hand == h.id) vs.push_back(&v);
if (!vs.empty() && hand_3d(h, vs, t_ns)) {
const V3 palm = (h.pts[0] + h.pts[5] + h.pts[9] + h.pts[13] + h.pts[17]) * 0.2;
if (h.last_ns && t_ns > h.last_ns)
h.speed += 0.5 * (std::min(norm(palm - h.last_palm) / ((t_ns - h.last_ns) / 1e9), 5.0) - h.speed);
h.last_ns = t_ns, h.last_palm = palm, h.seen_ns = t_ns;
++h.frames;
smooth(h, t_ns);
out.push_back(&h);
++it;
} else if (t_ns - h.seen_ns > 300'000'000) {
it = hands_.erase(it);
++stats.forgotten;
} else {
++it;
}
}
// views split off by a failed triangulation start over as new hands next frame
for (View &v : views_)
if (v.hand <= 0) {
v.hand = next_id_++;
hands_[v.hand].id = v.hand;
++stats.created;
}
last_ns_ = t_ns;
stats.step_ms += ms_since(t_step);
return out;
}
std::vector<Seen> Tracker::views_now() const {
std::vector<Seen> out;
for (const View &v : views_) out.push_back({v.cam->name, v.hand, v.roi, v.lm, {}});
return out;
}
std::vector<Seen> Tracker::exhaustive(const std::map<std::string, Image> &images) {
std::vector<Tile *> tiles;
for (Tile &t : tiles_)
if (images.count(t.cam->name)) tiles.push_back(&t);
std::vector<std::vector<Palm>> found(tiles.size());
std::vector<std::function<void()>> jobs;
for (size_t i = 0; i < tiles.size(); ++i)
jobs.push_back([this, &tiles, &images, &found, i] {
const Tile *t = tiles[i];
found[i] = nets_.palms(images.at(t->cam->name), t->center, t->size, t->rotation);
});
pool_.run(jobs);
// one crop per palm: tiles overlap, so the same palm turns up several times
std::vector<std::pair<double, View>> palms;
for (size_t i = 0; i < tiles.size(); ++i)
for (const Palm &p : found[i]) palms.push_back({p.score, View{tiles[i]->cam, p.roi(), 0, {}, false, 0}});
std::sort(palms.begin(), palms.end(), [](auto &a, auto &b) { return a.first > b.first; });
std::vector<View> crops;
for (auto &[score, v] : palms) {
bool near = false;
for (View &o : crops) near = near || (o.cam == v.cam && norm(o.roi.center - v.roi.center) < 0.5 * v.roi.size);
if (!near) crops.push_back(v);
}
std::vector<View *> ptrs;
for (View &v : crops) ptrs.push_back(&v);
run_landmarks(images, ptrs);
std::sort(crops.begin(), crops.end(), [](const View &a, const View &b) { return a.lm.presence > b.lm.presence; });
std::vector<Seen> out;
for (View &v : crops) {
if (v.lm.presence < min_presence_) continue;
bool dup = false;
for (const Seen &o : out)
dup = dup || (o.cam == v.cam->name && norm(palm_centre(o.lm) - palm_centre(v.lm)) < 0.5 * hand_size(v.lm));
if (dup) continue;
V3 pts[21];
Seen s{v.cam->name, 0, v.roi, v.lm, {}};
if (single_view(*v.cam, v.lm, 1.0, pts)) s.wrist = pts[0];
out.push_back(s);
}
return out;
}
void Tracker::smooth(Hand &h, int64_t t_ns) {
const double dt = (t_ns - h.smooth_ns) / 1e9;
h.smooth_ns = t_ns;
if (h.frames <= 1 || dt <= 0 || dt > 0.3) { // new, or back after a gap: start over
std::copy(h.pts, h.pts + 21, h.smooth);
h.dpalm = {0, 0, 0};
return;
}
auto alpha = [dt](double cutoff) { return 1 / (1 + 1 / (2 * M_PI * cutoff * dt)); };
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
const V3 d = (palm(h.pts) - palm(h.smooth)) * (1 / dt);
h.dpalm = h.dpalm + (d - h.dpalm) * alpha(kSpeedCutoff);
// one cutoff for the whole hand, from its palm speed, so its shape stays together
const double a = alpha(kMinCutoff + kBeta * norm(h.dpalm));
for (int i = 0; i < 21; ++i) h.smooth[i] = h.smooth[i] + (h.pts[i] - h.smooth[i]) * a;
}
+129
View File
@@ -0,0 +1,129 @@
// Multi-camera hand tracking (a port of tracker/hands.py; the scheduling is described
// there and in tracker/README.md). All 3D is metres in the head frame.
#pragma once
#include "calib.h"
#include "nets.h"
#include <condition_variable>
#include <functional>
#include <map>
#include <memory>
#include <mutex>
#include <thread>
#include <vector>
// Runs batches of jobs on a few threads, each pinned to a core.
class Pool {
public:
// One thread per entry of cpus, pinned there (round-robin if threads > cpus).
Pool(int threads, const std::vector<int> &cpus);
~Pool();
void run(std::vector<std::function<void()>> &jobs);
private:
void loop(int cpu);
std::vector<std::thread> threads_;
std::mutex mu_;
std::condition_variable wake_, done_;
std::vector<std::function<void()>> *jobs_ = nullptr;
size_t next_ = 0, finished_ = 0;
bool stop_ = false;
};
struct Hand {
int id = 0;
V3 pts[21]{}; // as measured this frame; the tracker steers crops by these
V3 smooth[21]{}; // filtered (One Euro, see Tracker::smooth): publish these
bool has_pts = false;
double residual = -1; // rms ray distance of the triangulation (m); -1: one view
int nviews = 0;
double right_score = 0.5; // the model's right-hand score (these images aren't mirrored)
double scale = 1.0; // this user's hand size / the model's world landmarks
int64_t seen_ns = 0;
int frames = 0;
double speed = 0; // palm centre, m/s, smoothed
int64_t last_ns = 0;
V3 last_palm{};
V3 dpalm{}; // the filter's palm velocity, m/s
int64_t smooth_ns = 0;
bool right() const { return right_score > 0.5; }
};
struct Stats {
int palm_calls = 0, hand_calls = 0, sets = 0;
double palm_ms = 0, hand_ms = 0, step_ms = 0; // summed batch times
int palm_batches = 0, hand_batches = 0;
// why views and hands come and go
int lost = 0; // a tracked view's landmarks fell below min presence
int handoff_miss = 0; // a view projected from the hand's 3D (new camera or retry) found no hand
int dups = 0; // the same hand twice in one camera
int splits = 0; // a hand's views disagreed in 3D and were split
int created = 0, merged = 0, forgotten = 0; // merged: views re-paired across cameras
// diagnostics: on stereo frames, each view's single-view palm distance / the stereo one
std::vector<std::pair<int, double>> mono_ratio; // (hand id, ratio)
};
// A hand the landmark model found in one camera (Tracker::views_now, Tracker::exhaustive).
struct Seen {
std::string cam;
int hand = 0; // the tracker's hand; 0 in exhaustive()
Roi roi;
Landmarks lm;
V3 wrist{}; // exhaustive(): single-view 3D guess at the model's hand size
};
class Tracker {
public:
Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views = 2);
// images: calibration name -> frame. Returns the hands seen in this set.
std::vector<const Hand *> step(const std::map<std::string, Image> &images, int64_t t_ns);
// Seconds until the next frame set is worth processing (30 Hz fast hands, 15 Hz slow, 5 Hz none).
double interval() const;
Stats stats;
size_t views() const { return views_.size(); }
std::vector<Seen> views_now() const;
// Every search tile in every camera, then landmarks on every palm: slow; for checking
// what the scheduler misses (fh-replay --oracle).
std::vector<Seen> exhaustive(const std::map<std::string, Image> &images);
// Landmark presence a tracked view needs to stay (new views need min presence, 0.5). In
// bright rooms the camera exposes for the room, the hands come out dim, and presence
// dips under 0.5 for a frame at a time.
void set_keep_presence(double p) { keep_presence_ = p; }
private:
struct View {
const Camera *cam;
Roi roi;
int hand = 0; // 0: not assigned yet
Landmarks lm;
bool has_lm = false;
int frames = 0;
bool fresh = false; // lm is from this step
};
struct Tile {
const Camera *cam;
V2 center;
double size, rotation, weight, credit = 0;
};
void run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views);
bool single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const;
bool hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns);
double pair_cost(const View &a, const View &b) const;
double size_misfit(const std::vector<const View *> &views, const V3 *pts) const;
void associate();
static void smooth(Hand &h, int64_t t_ns);
bool inside(const Camera &cam, V2 uv) const;
void add_tiles(const Camera &cam, double frac, int gx, int gy);
std::map<std::string, const Camera *> cams_;
const Nets &nets_;
Pool &pool_;
int max_views_, hand_budget_ = 4, search_budget_ = 3;
double min_presence_ = 0.5, keep_presence_ = 0.5;
std::vector<View> views_;
std::map<int, Hand> hands_;
std::vector<Tile> tiles_;
int next_id_ = 1;
int64_t last_ns_ = 0, search_ns_ = 0;
};
+151
View File
@@ -0,0 +1,151 @@
"""Tracking-camera calibration from the headset's factory files.
/persist/xrservice.json (written by Valve's calibration, loaded by XRService)
holds, per camera, Kannala-Brandt fisheye intrinsics ("kb": fx fy cx cy k1-k4,
pixel centres at integer coordinates, as in OpenCV's fisheye model) and a pose
in the slam_right (Cam0) frame: plus_x/plus_z are the camera axes and position
its origin, in mm. /persist/device_config.json gives Cam0's pose in the CAD
frame (cv.cad_from_cal, metres) and the head's pose in CAD (head). The CAD frame
is +X head-left, +Y up, +Z forward; the head frame is OpenVR's: +x right, +y up,
-z forward. Camera frames: +z along the optical axis, +x right and +y down in
the image.
Everything here returns metres in the head frame.
"""
import json
import numpy as np
XRSERVICE_JSON = '/persist/xrservice.json'
DEVICE_JSON = '/persist/device_config.json'
def _pose(d, scale=1.0):
"""4x4 transform from a {plus_x, plus_z, position} pose (child axes in the parent frame)."""
x = np.asarray(d['plus_x'], float)
z = np.asarray(d['plus_z'], float)
y = np.cross(z, x)
T = np.eye(4)
T[:3, 0], T[:3, 1], T[:3, 2] = x, y, z
T[:3, 3] = np.asarray(d['position'], float) * scale
return T
class Camera:
def __init__(self, name, width, height, kb, head_from_cam):
self.name = name
self.width, self.height = width, height
self.fx, self.fy, self.cx, self.cy = kb['fx'], kb['fy'], kb['cx'], kb['cy']
self.k = np.array([kb['k1'], kb['k2'], kb['k3'], kb['k4']])
self.head_from_cam = head_from_cam
self.R = head_from_cam[:3, :3] # camera axes in the head frame
self.origin = head_from_cam[:3, 3] # camera centre in the head frame
def __repr__(self):
return 'Camera(%s %dx%d at %s mm)' % (self.name, self.width, self.height,
np.round(self.origin * 1000, 1))
def _theta_d(self, theta):
t2 = theta * theta
k1, k2, k3, k4 = self.k
return theta * (1 + t2 * (k1 + t2 * (k2 + t2 * (k3 + t2 * k4))))
def project_cam(self, p):
"""Camera-frame points (N,3) -> pixels (N,2). Points behind the lens still map (the lens sees ~180 deg)."""
p = np.atleast_2d(p)
r = np.hypot(p[:, 0], p[:, 1])
theta = np.arctan2(r, p[:, 2])
scale = np.where(r > 1e-12, self._theta_d(theta) / np.maximum(r, 1e-12), 0.0)
return np.stack([self.fx * p[:, 0] * scale + self.cx, self.fy * p[:, 1] * scale + self.cy], axis=1)
def unproject(self, uv):
"""Pixels (N,2) -> unit rays (N,3) in the camera frame."""
uv = np.atleast_2d(np.asarray(uv, float))
mx = (uv[:, 0] - self.cx) / self.fx
my = (uv[:, 1] - self.cy) / self.fy
td = np.hypot(mx, my)
theta = td.copy()
k1, k2, k3, k4 = self.k
for _ in range(8): # Newton on theta_d(theta) = td
t2 = theta * theta
f = self._theta_d(theta) - td
df = 1 + t2 * (3 * k1 + t2 * (5 * k2 + t2 * (7 * k3 + t2 * 9 * k4)))
theta = np.clip(theta - f / df, 0.0, np.pi)
s = np.where(td > 1e-12, np.sin(theta) / np.maximum(td, 1e-12), 1.0)
return np.stack([mx * s, my * s, np.cos(theta)], axis=1)
def rays(self, uv):
"""Pixels -> unit rays in the head frame (all starting at self.origin)."""
return self.unproject(uv) @ self.R.T
def project(self, p_head):
"""Head-frame points (N,3) -> pixels (N,2) and depth along the optical axis (N,)."""
p = (np.atleast_2d(p_head) - self.origin) @ self.R
return self.project_cam(p), p[:, 2]
def angle_from_axis(self, uv):
"""Angle in degrees between each pixel's ray and the optical axis."""
return np.degrees(np.arccos(np.clip(self.unproject(uv)[:, 2], -1, 1)))
def load(xrservice=XRSERVICE_JSON, device=DEVICE_JSON):
"""{calibration name: Camera} for the tracking cameras, posed in the head frame."""
with open(xrservice) as f:
rig = json.load(f)
with open(device) as f:
dev = json.load(f)
cad_from_cam0 = _pose(dev['cv']['cad_from_cal'])
head_from_cad = np.linalg.inv(_pose(dev['head']))
cams = {}
for c in rig['cameras']:
kb = next(i for i in c['intrinsics'] if i['cameraModel'] == 'kb')
cam0_from_cam = _pose(c['extrinsics'], 1e-3)
cams[c['sourceCamera']] = Camera(c['sourceCamera'], c['width'], c['height'], kb,
head_from_cad @ cad_from_cam0 @ cam0_from_cam)
return cams
def triangulate(origins, dirs, weights=None):
"""Least-squares point closest to several rays. Returns (point, rms distance to the rays)."""
A = np.zeros((3, 3))
b = np.zeros(3)
w = np.ones(len(origins)) if weights is None else np.asarray(weights, float)
for o, d, wi in zip(origins, dirs, w):
P = np.eye(3) - np.outer(d, d)
A += wi * P
b += wi * P @ o
p = np.linalg.solve(A, b)
res = [np.linalg.norm((np.eye(3) - np.outer(d, d)) @ (p - o)) for o, d in zip(origins, dirs)]
return p, float(np.sqrt(np.mean(np.square(res))))
def triangulate_many(origins, dirs, weights):
"""Triangulate K points seen from V cameras at once.
origins (V,3), dirs (V,K,3) unit rays, weights (V,). Returns points (K,3) and
each point's rms distance to its rays (K,).
"""
P = np.eye(3) - dirs[..., :, None] * dirs[..., None, :] # (V,K,3,3)
w = weights[:, None, None, None]
A = (w * P).sum(0)
b = (w * (P @ origins[:, None, :, None])).sum(0)[..., 0]
pts = np.linalg.solve(A, b[..., None])[..., 0]
off = pts[None] - origins[:, None, :] # (V,K,3)
perp = off - (off * dirs).sum(-1, keepdims=True) * dirs
return pts, np.sqrt((perp ** 2).sum(-1).mean(0))
if __name__ == '__main__':
cams = load()
for cam in cams.values():
fwd = [float(v) for v in cam.R[:, 2]]
print('%-12s at x %+6.1f y %+6.1f z %+6.1f mm, looks %s' % (
cam.name, *(cam.origin * 1000),
'right' * (fwd[0] > 0.3) + 'left' * (fwd[0] < -0.3) + ' up' * (fwd[1] > 0.3) +
' down' * (fwd[1] < -0.3) + ' forward' * (fwd[2] < -0.3) + ' back' * (fwd[2] > 0.3)),
np.round(fwd, 2))
uv = np.array([[cam.cx + 200, cam.cy - 100], [cam.cx - 0.4 * cam.width, cam.cy + 0.3 * cam.height]])
err = np.abs(cam.project_cam(cam.unproject(uv)) - uv).max()
assert err < 1e-6, err
a, b = cams['slam_left'], cams['slam_right']
print('slam baseline %.2f mm' % (1000 * np.linalg.norm(a.origin - b.origin)))
+300
View File
@@ -0,0 +1,300 @@
"""MediaPipe's palm detector and hand landmark model on ncnn, CPU or Vulkan GPU.
The models are the OpenCV Zoo ONNX ports, converted by tools/convert_models.py.
Crops are square regions of a camera image given as (centre, size, rotation)
in pixels and radians; rotation turns the crop's "up" towards the image
direction (sin r, -cos r), as in MediaPipe. Both models take RGB in [0, 1];
mono crops are replicated to three planes.
Preparing crops and decoding outputs happen here; running the networks is an
engine's job: Engine runs them in this process, Pool in worker processes
(ncnn's Python binding holds the GIL while it infers, so threads don't help).
"""
import multiprocessing as mp
import os
from multiprocessing import connection, shared_memory
import cv2
import ncnn
import numpy as np
HERE = os.path.dirname(os.path.abspath(__file__))
MODELS = os.path.join(HERE, '..', 'models', 'ncnn')
# MediaPipe hand_landmarks_to_rect: the palm and finger bases used for the next ROI
ROI_LANDMARKS = [0, 1, 2, 3, 5, 6, 9, 10, 13, 14, 17, 18]
# per model: input size, output blobs and their sizes
SPECS = {'palm': (192, [('out0', 2016 * 18), ('out1', 2016)]),
'hand': (224, [('out0', 63), ('out1', 1), ('out2', 1), ('out3', 63)])}
def load_net(name, gpu, threads=1):
net = ncnn.Net()
net.opt.use_vulkan_compute = gpu
net.opt.num_threads = threads
net.opt.use_fp16_packed = net.opt.use_fp16_storage = net.opt.use_fp16_arithmetic = True
net.load_param(os.path.join(MODELS, name + '.ncnn.param'))
net.load_model(os.path.join(MODELS, name + '.ncnn.bin'))
return net
def infer(net, kind, patch):
"""Run one network on an 8-bit mono crop; returns its outputs, flattened."""
plane = patch.astype(np.float32) * (1.0 / 255.0)
x = np.ascontiguousarray(np.broadcast_to(plane, (3,) + plane.shape))
ex = net.create_extractor()
ex.input('in0', ncnn.Mat(x))
return [np.array(ex.extract(name)[1], np.float32).reshape(-1) for name, _ in SPECS[kind][1]]
class Engine:
"""Runs the networks in this process, one after another."""
def __init__(self, palm_gpu=False, hand_gpu=False):
self.nets = {'palm': load_net('palm', palm_gpu), 'hand': load_net('hand', hand_gpu)}
def run(self, jobs):
"""jobs: [(kind, patch)] -> [outputs]"""
return [infer(self.nets[kind], kind, patch) for kind, patch in jobs]
def close(self):
pass
def _worker(index, shm_name, conn, palm_gpu, hand_gpu, cpu):
if cpu is not None and cpu in os.sched_getaffinity(0):
os.sched_setaffinity(0, {cpu})
nets = {'palm': load_net('palm', palm_gpu), 'hand': load_net('hand', hand_gpu)}
shm = shared_memory.SharedMemory(name=shm_name)
base = index * Pool.STRIDE
while True:
msg = conn.recv()
if msg is None:
break
kind = msg
size = SPECS[kind][0]
patch = np.ndarray((size, size), np.uint8, shm.buf, base)
outs = infer(nets[kind], kind, patch)
off = base + Pool.IN_BYTES
for o in outs:
np.ndarray(o.shape, np.float32, shm.buf, off)[:] = o
off += o.nbytes
conn.send(True)
shm.close()
class Pool:
"""Runs the networks in worker processes pinned to CPUs, several crops at once."""
IN_BYTES = 224 * 224
OUT_BYTES = 4 * (2016 * 18 + 2016)
STRIDE = IN_BYTES + OUT_BYTES
# SteamOS on the Frame keeps user processes on CPUs 0-4 (5-7 carry pinned VR threads);
# 2-4 are the big cores among those
def __init__(self, workers=2, cpus=(2, 3, 4), palm_gpu=False, hand_gpu=False):
ctx = mp.get_context('spawn')
self.shm = shared_memory.SharedMemory(create=True, size=workers * self.STRIDE)
self.conns, self.procs = [], []
for i in range(workers):
a, b = ctx.Pipe()
cpu = cpus[i % len(cpus)] if cpus else None
p = ctx.Process(target=_worker, args=(i, self.shm.name, b, palm_gpu, hand_gpu, cpu), daemon=True)
p.start()
self.conns.append(a)
self.procs.append(p)
def run(self, jobs):
results = [None] * len(jobs)
pending = list(range(len(jobs)))
busy = {} # conn -> job index
free = list(range(len(self.conns)))
while pending or busy:
while pending and free:
w = free.pop()
j = pending.pop(0)
kind, patch = jobs[j]
base = w * self.STRIDE
np.ndarray(patch.shape, np.uint8, self.shm.buf, base)[:] = patch
self.conns[w].send(kind)
busy[self.conns[w]] = (w, j)
for c in connection.wait(list(busy)):
c.recv()
w, j = busy.pop(c)
kind = jobs[j][0]
off = w * self.STRIDE + self.IN_BYTES
outs = []
for _, n in SPECS[kind][1]:
outs.append(np.ndarray((n,), np.float32, self.shm.buf, off).copy())
off += 4 * n
results[j] = outs
free.append(w)
return results
def pids(self):
return [p.pid for p in self.procs]
def close(self):
for c in self.conns:
try:
c.send(None)
except OSError:
pass
for p in self.procs:
p.join(1)
self.shm.close()
self.shm.unlink()
def crop_matrix(center, size, rotation, out):
"""2x3 affine taking crop pixels (0..out) to image pixels."""
c, s = np.cos(rotation), np.sin(rotation)
k = size / out
R = np.array([[c, -s], [s, c]]) * k
t = np.asarray(center, float) - R @ np.array([out / 2.0, out / 2.0])
return np.hstack([R, t[:, None]])
# Local contrast per crop: the IR frames are dim and uneven (the upper cameras
# especially). On the 2026-09-28 capture this found hands in more frames on every
# camera than plain, stretched or gamma-corrected crops.
CLAHE = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(4, 4))
def crop(gray, M, out):
"""8-bit square crop of a mono image, with a crop->image matrix, contrast-equalized."""
patch = cv2.warpAffine(gray, M, (out, out), flags=cv2.INTER_LINEAR | cv2.WARP_INVERSE_MAP,
borderMode=cv2.BORDER_CONSTANT, borderValue=0)
return CLAHE.apply(patch)
def to_image(M, pts):
"""Crop pixels (N,2) -> image pixels."""
return pts @ M[:, :2].T + M[:, 2]
def normalize_angle(a):
return (a + np.pi) % (2 * np.pi) - np.pi
def _anchors():
"""SSD anchors of palm_detection_full: strides 8 (2 per cell) and 16 (6 per cell), 192x192."""
out = []
for stride, per_cell in ((8, 2), (16, 6)):
n = 192 // stride
for y in range(n):
for x in range(n):
out += [((x + 0.5) / n, (y + 0.5) / n)] * per_cell
return np.array(out, np.float32)
class Detection:
"""A palm in image pixels: box centre/size, 7 keypoints, score."""
__slots__ = ('center', 'size', 'keypoints', 'score')
def __init__(self, center, size, keypoints, score):
self.center, self.size, self.keypoints, self.score = center, size, keypoints, score
def roi(self):
"""MediaPipe's hand ROI from a palm: wrist->middle-finger-base sets the rotation, then
shift 0.5 towards the fingers and scale the square box 2.6x."""
(x0, y0), (x1, y1) = self.keypoints[0], self.keypoints[2]
rot = normalize_angle(np.pi / 2 - np.arctan2(-(y1 - y0), x1 - x0))
w, h = (float(v) for v in self.size)
shift = np.array([-h * -0.5 * np.sin(rot), h * -0.5 * np.cos(rot)])
return np.asarray(self.center) + shift, max(w, h) * 2.6, rot
class PalmDetector:
SIZE = 192
def __init__(self, min_score=0.5):
self.anchors = _anchors()
self.min_score = min_score
def prepare(self, gray, center, size, rotation=0.0):
M = crop_matrix(center, size, rotation, self.SIZE)
return crop(gray, M, self.SIZE), (M, size)
def decode(self, outputs, ctx):
"""Palms found in one crop, in image pixels."""
M, size = ctx
raw = outputs[0].reshape(-1, 18)
logit = outputs[1]
keep = np.nonzero(logit > np.log(self.min_score / (1 - self.min_score)))[0]
if not len(keep):
return []
a = self.anchors[keep] * self.SIZE
r = raw[keep]
centers = r[:, 0:2] + a
sizes = r[:, 2:4]
kps = r[:, 4:18].reshape(-1, 7, 2) + a[:, None, :]
scores = 1 / (1 + np.exp(-np.clip(logit[keep], -100, 100)))
return [Detection(to_image(M, c[None])[0], s * size / self.SIZE, to_image(M, k), sc)
for c, s, k, sc in _weighted_nms(centers, sizes, scores, kps)]
def _weighted_nms(centers, sizes, scores, kps, iou_thresh=0.3):
"""MediaPipe's weighted NMS: overlapping boxes are averaged, weighted by score."""
order = np.argsort(-scores)
boxes = np.hstack([centers - sizes / 2, centers + sizes / 2])
area = np.prod(sizes, axis=1)
out = []
while len(order):
i = order[0]
xy0 = np.maximum(boxes[i, :2], boxes[order, :2])
xy1 = np.minimum(boxes[i, 2:], boxes[order, 2:])
inter = np.prod(np.clip(xy1 - xy0, 0, None), axis=1)
iou = inter / (area[i] + area[order] - inter + 1e-9)
group = order[iou > iou_thresh]
w = scores[group][:, None]
out.append(((centers[group] * w).sum(0) / w.sum(), (sizes[group] * w).sum(0) / w.sum(),
(kps[group] * w[:, :, None]).sum(0) / w.sum(), float(scores[i])))
order = order[iou <= iou_thresh]
return out
class Landmarks:
"""21 hand landmarks in image pixels, plus MediaPipe's metric 'world' landmarks."""
__slots__ = ('pts', 'depth', 'world', 'presence', 'right', 'roi')
def __init__(self, pts, depth, world, presence, right, roi):
self.pts, self.depth, self.world, self.presence = pts, depth, world, presence
self.right, self.roi = right, roi
def next_roi(self):
return roi_from_points(self.pts)
def roi_from_points(p):
"""MediaPipe's hand_landmarks_to_rect: the crop to track a hand given its landmarks."""
x0, y0 = p[0]
x1, y1 = ((p[5] + p[13]) / 2 + p[9]) / 2
rot = normalize_angle(np.pi / 2 - np.arctan2(-(y1 - y0), x1 - x0))
sub = p[ROI_LANDMARKS]
center = (sub.min(0) + sub.max(0)) / 2
c, s = np.cos(-rot), np.sin(-rot)
q = (sub - center) @ np.array([[c, -s], [s, c]]).T
lo, hi = q.min(0), q.max(0)
mid = (lo + hi) / 2
c2, s2 = np.cos(rot), np.sin(rot)
center = center + np.array([mid[0] * c2 - mid[1] * s2, mid[0] * s2 + mid[1] * c2])
w, h = hi - lo
center = center + np.array([-h * -0.1 * s2, h * -0.1 * c2])
return center, max(w, h) * 2.0, rot
class HandLandmarker:
SIZE = 224
def prepare(self, gray, roi):
center, size, rotation = roi
M = crop_matrix(center, size, rotation, self.SIZE)
return crop(gray, M, self.SIZE), (M, roi)
def decode(self, outputs, ctx):
M, roi = ctx
screen = outputs[0].reshape(21, 3)
pts = to_image(M, screen[:, :2])
depth = screen[:, 2] * roi[1] / self.SIZE # relative depth, image pixels
return Landmarks(pts, depth, outputs[3].reshape(21, 3), float(outputs[1][0]), float(outputs[2][0]), roi)