mirror of
https://github.com/DeeJanuz/frametop.git
synced 2026-10-06 01:00:06 +02:00
Bring in frame-hands' hand tracking under hands/
The history of frame-hands (~/Desktop/Projects/frame-hands on the Frame), filtered to what moves: the camera broker (camd/), the tracker and its offline tools (trackd/), the shared file layouts (include/), the ncnn models, the analysis tools, and the calibration and model helpers they import from the Python prototype. The reverse-engineering notes, probes, camprobe, and the rest of the prototype stay in frame-hands. Unchanged here: the renames to ft- names and Frametop paths follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
commit
1a76d1560b
49 files changed
+7650
No files matched your search
@@ -0,0 +1,23 @@
|
||||
# Recordings (tens of GB) and the Python venv
|
||||
captures/
|
||||
.venv/
|
||||
__pycache__/
|
||||
# Upstream clones: ncnn (see trackd/README.md) and FrameEyeCameraFeed (MIT, adapted into camd/)
|
||||
vendor/
|
||||
# Disassembly of Valve's vrclient.so, for reverse engineering only
|
||||
re/*.dis
|
||||
# Downloadable model sources (tools/convert_models.py); the converted ncnn models are kept
|
||||
models/onnx/
|
||||
models/*.task
|
||||
# Build output
|
||||
camd/fh-camd
|
||||
camd/fh-camprobe
|
||||
trackd/fh-tracker
|
||||
trackd/fh-replay
|
||||
trackd/fh-ringplay
|
||||
trackd/nettest
|
||||
probes/fh-frametime
|
||||
probes/mgrvt
|
||||
probes/ptstate
|
||||
probes/refprobe
|
||||
probes/*.bin
|
||||
@@ -0,0 +1,21 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2026 Curtis English
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
@@ -0,0 +1,15 @@
|
||||
CFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter
|
||||
LDLIBS = -lm
|
||||
|
||||
all: fh-camprobe fh-camd
|
||||
|
||||
fh-camprobe: camprobe.c tp.c xrcams.c tp.h xrcams.h
|
||||
$(CC) $(CFLAGS) -o $@ camprobe.c tp.c xrcams.c $(LDLIBS)
|
||||
|
||||
fh-camd: camd.c tp.c xrcams.c tp.h xrcams.h fhring.h
|
||||
$(CC) $(CFLAGS) -o $@ camd.c tp.c xrcams.c $(LDLIBS)
|
||||
|
||||
clean:
|
||||
rm -f fh-camprobe fh-camd
|
||||
|
||||
.PHONY: all clean
|
||||
@@ -0,0 +1,63 @@
|
||||
# camd
|
||||
|
||||
Root-side camera access for frame-hands:
|
||||
|
||||
- `fh-camd`: the frame broker. It publishes the four IR tracking cameras, and optionally the Arcturus color pair, to a shared-memory ring that the unprivileged tracker reads.
|
||||
- `fh-camprobe`: a test tool that records timestamped frames from every camera, including the Arcturus color pair, with CSV logs.
|
||||
|
||||
## How it gets frames
|
||||
|
||||
XRService owns the headset cameras. Both tools borrow its DMA-BUFs read-only with `pidfd_getfd`, the same way FrameEyeCameraFeed does. They never touch XRService's V4L2 descriptors.
|
||||
|
||||
Polling buffers for changes can catch a frame while the camera is still writing it. Instead, they listen to the `v4l2:v4l2_dqbuf` tracepoint, which fires when XRService takes a buffer. It gives the buffer index, the sequence number and the capture timestamp.
|
||||
|
||||
The tools learn which DMA-BUF holds each V4L2 index by watching which buffer changes at each dequeue.
|
||||
|
||||
- Right after XRService allocates its buffers, the mapping is allocation order.
|
||||
- After XRService restarts streaming, the order is shuffled, and the mapping is learned index by index.
|
||||
- The two upper cameras share one run of buffers. For them, only allocation order can tell the cameras apart.
|
||||
- `fh-camd` also re-maps an index on the fly when its buffer holds no new frame.
|
||||
|
||||
## fh-camd
|
||||
|
||||
```
|
||||
make
|
||||
sudo ./fh-camd # runs until stopped or XRService exits
|
||||
```
|
||||
|
||||
It needs root only to set up: to borrow the buffers (`ptrace_scope=1` blocks `pidfd_getfd`) and to open the root-only tracepoints. Then it drops to the invoking user for good; XRService itself runs as that user. It reads nothing from the ring's readers.
|
||||
|
||||
Frames go to `/run/frame-hands/ir-ring`. The file is mode 0600 and owned by the user. It sits in a root-owned directory, so no other account can plant a file or link there. The layout is in `fhring.h`, and `tracker/ring.py` reads it.
|
||||
|
||||
- Only complete, bright frames are published. The cameras alternate a normal exposure with a near-black one, so each camera gets 30 of its 60 fps.
|
||||
- A copy torn by the camera overwriting the buffer is dropped.
|
||||
- Each copy takes about 0.1 ms, and a cache sync about 0.15 ms.
|
||||
|
||||
Options:
|
||||
|
||||
- `--with-dark`: also publish the near-black frames, as extra ring cameras flagged `FH_CAM_DARK`. They show only light sources, so they're no use for hands.
|
||||
- `--with-color`: also publish the two Arcturus color cameras, flagged `FH_CAM_COLOR`. Each is the luma of the 10-bit frame's valid 1972x2464 (the top 8 bits), at half size (`--color-scale 2`: 986x1232) and at most 30 fps (`--color-fps`; the cameras run at 60). Frames that carry the module's warped half-size copy are dropped. Their `capture_ns` is on the color module's clock (2.2 s off the mono cameras' on 2026-09-29), so line them up with the mono cameras by `dqbuf_ns`. Each frame costs about 0.65 ms of cache sync and 1.1 ms of decoding, so both cameras at 30 fps take about 11% of a core.
|
||||
- The ring holds 8 cameras: 4 mono, plus 4 dark twins or 2 color cameras.
|
||||
- `--sensor S`: only the mono cameras whose sensor name contains S.
|
||||
|
||||
It exits when XRService exits, or when a camera's buffers keep going stale, which means XRService has reallocated them. Start it again to re-attach.
|
||||
|
||||
## fh-camprobe
|
||||
|
||||
Wear the headset (cameras only stream while it's worn), then run:
|
||||
|
||||
```
|
||||
sudo ./fh-camprobe # records 15 s
|
||||
sudo ./fh-camprobe --list # discovery and tracepoints only
|
||||
```
|
||||
|
||||
Output goes to `~/Pictures/framecap/stereo-<time>/`:
|
||||
|
||||
- `summary.txt`: delays, buffer mapping, exposure pattern, stereo sync, and torn or stale copies.
|
||||
- `frames.csv`: one row per frame. `events.csv`: every tracepoint sample.
|
||||
- `<model>_NNNN_<camera>.pgm`: saved bright pairs. The color cameras are saved raw as `.yuv420_10p` (`tools/decode.py` reads them).
|
||||
- `plane1_*.bin`: raw plane 1 of a few frames, which may hold sensor metadata.
|
||||
|
||||
Which camera is which: video9 is `slam_left`, video13 is `slam_right`, video6 is `upper_left` and video7 is `upper_right`. This was checked by rendering the same view from each camera with the factory calibration. But fh-camd tells the side cameras' buffers apart only by XRService's allocation order, and after some XRService restarts it gets them backwards: check with `tools/check_sides.py --ring` and run the tracker with `--swap-sides` when it says swapped. The color cameras are video3 (`arcimx616 0-0010`) and video0 (`0-001a`); which of them is `passthrough_left` in the module's calibration is for `tools/check_color.py` to settle on a recording with texture in view. `tracker/live.py` maps them by capture pipe (`/sys/class/video4linux/videoN/name`).
|
||||
|
||||
`discovery` in `xrcams.c` is adapted from FrameEyeCameraFeed (MIT, see `LICENSE.FrameEyeCameraFeed`).
|
||||
+1066
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,87 @@
|
||||
/*
|
||||
* fhring - the shared-memory frame ring fh-camd writes and trackers read.
|
||||
*
|
||||
* One file (FH_RING_PATH) holds a header, then for each camera
|
||||
* a few slots, each a slot header followed by the image rows packed tightly
|
||||
* (stride == width for 8-bit mono). Only complete, bright frames are published.
|
||||
*
|
||||
* Writer, for frame n of a camera: slot = n % nslots
|
||||
* slot.seq = 2n+1; write slot fields and pixels; slot.seq = 2n+2; cam.latest = n
|
||||
* Reader:
|
||||
* n = cam.latest; read slot.seq, expect 2n+2; copy; re-read slot.seq; if it
|
||||
* changed the copy is torn, retry with the new latest.
|
||||
*
|
||||
* All multi-byte fields are little-endian; offsets are fixed so Python can read
|
||||
* them with struct (tracker/ring.py mirrors this file).
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <assert.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define FH_RING_MAGIC "FHRING01"
|
||||
#define FH_RING_VERSION 1
|
||||
#define FH_RING_MAX_CAMS 8
|
||||
#define FH_RING_SLOTS 4
|
||||
#define FH_RING_DIR "/run/frame-hands"
|
||||
#define FH_RING_PATH FH_RING_DIR "/ir-ring"
|
||||
|
||||
enum {
|
||||
FH_FMT_GREY8 = 0,
|
||||
};
|
||||
|
||||
enum {
|
||||
FH_CAM_DARK = 1u << 0, /* the near-black exposures between this node's */
|
||||
/* normal frames (fh-camd --with-dark) */
|
||||
FH_CAM_COLOR = 1u << 1, /* an Arcturus color camera's luma, downscaled */
|
||||
/* (fh-camd --with-color). Not synced with the */
|
||||
/* mono cameras, and capture_ns is on its own */
|
||||
/* clock: line it up with them by dqbuf_ns */
|
||||
};
|
||||
|
||||
typedef struct {
|
||||
char sensor[32]; /* media entity, e.g. "og01a1b 4-0060" */
|
||||
char name[32]; /* calibration name if known, else sensor slug */
|
||||
int32_t node; /* N of /dev/videoN */
|
||||
uint32_t format; /* FH_FMT_* */
|
||||
uint32_t width;
|
||||
uint32_t height;
|
||||
uint32_t stride; /* bytes per row in the ring */
|
||||
uint32_t nslots;
|
||||
uint64_t slot_offset; /* file offset of slot 0 */
|
||||
uint64_t slot_bytes; /* slot header + image, 64-byte aligned */
|
||||
volatile uint64_t latest; /* newest published frame number, 0 = none yet */
|
||||
uint64_t published; /* frames published */
|
||||
uint64_t dropped; /* dark, stale or torn frames not published */
|
||||
uint32_t flags; /* FH_CAM_* */
|
||||
uint8_t reserved[28];
|
||||
} fh_ring_cam_t; /* 160 bytes */
|
||||
|
||||
typedef struct {
|
||||
volatile uint64_t seq; /* 2n+1 while frame n is written, 2n+2 when done */
|
||||
uint64_t frame; /* n */
|
||||
uint64_t capture_ns; /* V4L2 timestamp (camera clock) */
|
||||
uint64_t dqbuf_ns; /* CLOCK_MONOTONIC when XRService dequeued it */
|
||||
uint64_t publish_ns; /* CLOCK_MONOTONIC when the copy finished */
|
||||
uint32_t v4l2_seq; /* V4L2 sequence number */
|
||||
float mean; /* mean luma on a sparse grid */
|
||||
uint8_t reserved[16];
|
||||
} fh_ring_slot_t; /* 64 bytes, image follows */
|
||||
|
||||
typedef struct {
|
||||
char magic[8]; /* FH_RING_MAGIC */
|
||||
uint32_t version;
|
||||
uint32_t header_bytes; /* sizeof(fh_ring_hdr_t) */
|
||||
uint32_t ncams;
|
||||
uint32_t reserved0;
|
||||
uint64_t file_bytes;
|
||||
int64_t writer_pid;
|
||||
volatile uint64_t heartbeat_ns; /* CLOCK_MONOTONIC, refreshed at least every 0.2 s */
|
||||
uint8_t reserved[16];
|
||||
fh_ring_cam_t cams[FH_RING_MAX_CAMS];
|
||||
} fh_ring_hdr_t;
|
||||
|
||||
static_assert(sizeof(fh_ring_cam_t) == 160, "fh_ring_cam_t layout");
|
||||
static_assert(sizeof(fh_ring_slot_t) == 64, "fh_ring_slot_t layout");
|
||||
static_assert(sizeof(fh_ring_hdr_t) == 64 + 160 * FH_RING_MAX_CAMS, "fh_ring_hdr_t layout");
|
||||
+427
@@ -0,0 +1,427 @@
|
||||
/*
|
||||
* tp - read kernel tracepoints system-wide through perf_event_open.
|
||||
*/
|
||||
|
||||
#define _GNU_SOURCE
|
||||
|
||||
#include "tp.h"
|
||||
|
||||
#include <errno.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/epoll.h>
|
||||
#include <sys/ioctl.h>
|
||||
#include <sys/mman.h>
|
||||
#include <sys/syscall.h>
|
||||
#include <time.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <linux/perf_event.h>
|
||||
|
||||
#ifndef TRACEFS
|
||||
#define TRACEFS "/sys/kernel/tracing/events"
|
||||
#endif
|
||||
#define RING_DATA_PAGES 16
|
||||
|
||||
static void set_err(char *err, size_t n, const char *fmt, ...)
|
||||
{
|
||||
va_list ap;
|
||||
|
||||
va_start(ap, fmt);
|
||||
vsnprintf(err, n, fmt, ap);
|
||||
va_end(ap);
|
||||
}
|
||||
|
||||
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn)
|
||||
{
|
||||
memset(ev, 0, sizeof(*ev));
|
||||
snprintf(ev->system, sizeof(ev->system), "%s", system);
|
||||
snprintf(ev->name, sizeof(ev->name), "%s", name);
|
||||
ev->id = -1;
|
||||
|
||||
char path[256];
|
||||
snprintf(path, sizeof(path), TRACEFS "/%s/%s/format", system, name);
|
||||
|
||||
FILE *f = fopen(path, "r");
|
||||
|
||||
if (!f) {
|
||||
set_err(err, errn, "%s: %s", path, strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
char line[512];
|
||||
|
||||
while (fgets(line, sizeof(line), f)) {
|
||||
|
||||
int id;
|
||||
|
||||
if (sscanf(line, "ID: %d", &id) == 1) {
|
||||
ev->id = id;
|
||||
continue;
|
||||
}
|
||||
|
||||
char *fp = line;
|
||||
|
||||
while (*fp == ' ' || *fp == '\t')
|
||||
fp++;
|
||||
|
||||
if (strncmp(fp, "field:", 6) || ev->nfields >= TP_MAX_FIELDS)
|
||||
continue;
|
||||
|
||||
char *semi = strchr(fp, ';');
|
||||
|
||||
if (!semi)
|
||||
continue;
|
||||
|
||||
/* the field name is the last identifier in the declaration */
|
||||
char decl[256];
|
||||
size_t dl = (size_t)(semi - (fp + 6));
|
||||
|
||||
if (dl >= sizeof(decl))
|
||||
dl = sizeof(decl) - 1;
|
||||
|
||||
memcpy(decl, fp + 6, dl);
|
||||
decl[dl] = 0;
|
||||
|
||||
char *br = strchr(decl, '[');
|
||||
|
||||
if (br)
|
||||
*br = 0;
|
||||
|
||||
char *end = decl + strlen(decl);
|
||||
|
||||
while (end > decl && (end[-1] == ' ' || end[-1] == '\t'))
|
||||
*--end = 0;
|
||||
|
||||
char *start = end;
|
||||
|
||||
while (start > decl && start[-1] != ' ' && start[-1] != '\t' && start[-1] != '*')
|
||||
start--;
|
||||
|
||||
tp_field_t *fd = &ev->fields[ev->nfields];
|
||||
const char *o = strstr(semi, "offset:");
|
||||
const char *s = strstr(semi, "size:");
|
||||
const char *g = strstr(semi, "signed:");
|
||||
|
||||
if (!o || !s)
|
||||
continue;
|
||||
|
||||
snprintf(fd->name, sizeof(fd->name), "%s", start);
|
||||
fd->offset = atoi(o + 7);
|
||||
fd->size = atoi(s + 5);
|
||||
fd->is_signed = g ? atoi(g + 7) != 0 : false;
|
||||
ev->nfields++;
|
||||
}
|
||||
|
||||
fclose(f);
|
||||
|
||||
if (ev->id < 0) {
|
||||
set_err(err, errn, "%s: no ID line", path);
|
||||
return false;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
int tp_field(const tp_event_t *ev, const char *name)
|
||||
{
|
||||
for (int i = 0; i < ev->nfields; i++)
|
||||
if (!strcmp(ev->fields[i].name, name))
|
||||
return i;
|
||||
|
||||
return -1;
|
||||
}
|
||||
|
||||
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen)
|
||||
{
|
||||
if (field < 0 || field >= ev->nfields)
|
||||
return 0;
|
||||
|
||||
const tp_field_t *f = &ev->fields[field];
|
||||
|
||||
if (f->offset < 0 || (uint32_t)(f->offset + f->size) > rawlen)
|
||||
return 0;
|
||||
|
||||
const uint8_t *p = raw + f->offset;
|
||||
|
||||
switch (f->size) {
|
||||
case 1: { uint8_t v; memcpy(&v, p, 1); return f->is_signed ? (int64_t)(int8_t)v : (int64_t)v; }
|
||||
case 2: { uint16_t v; memcpy(&v, p, 2); return f->is_signed ? (int64_t)(int16_t)v : (int64_t)v; }
|
||||
case 4: { uint32_t v; memcpy(&v, p, 4); return f->is_signed ? (int64_t)(int32_t)v : (int64_t)v; }
|
||||
case 8: { uint64_t v; memcpy(&v, p, 8); return (int64_t)v; }
|
||||
default: return 0;
|
||||
}
|
||||
}
|
||||
|
||||
static int online_cpus(int *cpus, int max)
|
||||
{
|
||||
FILE *f = fopen("/sys/devices/system/cpu/online", "r");
|
||||
int n = 0;
|
||||
|
||||
if (!f)
|
||||
return 0;
|
||||
|
||||
char buf[256] = {0};
|
||||
|
||||
if (!fgets(buf, sizeof(buf), f))
|
||||
buf[0] = 0;
|
||||
|
||||
fclose(f);
|
||||
|
||||
for (char *tok = strtok(buf, ",\n"); tok && n < max; tok = strtok(NULL, ",\n")) {
|
||||
|
||||
int a, b;
|
||||
|
||||
if (sscanf(tok, "%d-%d", &a, &b) == 2) {
|
||||
for (int c = a; c <= b && n < max; c++)
|
||||
cpus[n++] = c;
|
||||
} else if (sscanf(tok, "%d", &a) == 1) {
|
||||
cpus[n++] = a;
|
||||
}
|
||||
}
|
||||
|
||||
return n;
|
||||
}
|
||||
|
||||
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn)
|
||||
{
|
||||
memset(tp, 0, sizeof(*tp));
|
||||
tp->epfd = -1;
|
||||
|
||||
if (nevents <= 0 || nevents > TP_MAX_EVENTS) {
|
||||
set_err(err, errn, "bad event count %d", nevents);
|
||||
return false;
|
||||
}
|
||||
|
||||
for (int i = 0; i < nevents; i++)
|
||||
tp->events[i] = events[i];
|
||||
|
||||
tp->nevents = nevents;
|
||||
|
||||
int cpus[TP_MAX_CPUS];
|
||||
tp->ncpu = online_cpus(cpus, TP_MAX_CPUS);
|
||||
|
||||
if (tp->ncpu <= 0) {
|
||||
set_err(err, errn, "no online CPUs found");
|
||||
return false;
|
||||
}
|
||||
|
||||
long page = sysconf(_SC_PAGESIZE);
|
||||
tp->map_len = (size_t)page * (1 + RING_DATA_PAGES);
|
||||
|
||||
tp->epfd = epoll_create1(EPOLL_CLOEXEC);
|
||||
|
||||
if (tp->epfd < 0) {
|
||||
set_err(err, errn, "epoll_create1: %s", strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
for (int c = 0; c < tp->ncpu; c++) {
|
||||
|
||||
tp->ring_fd[c] = -1;
|
||||
|
||||
for (int e = 0; e < nevents; e++) {
|
||||
|
||||
struct perf_event_attr a;
|
||||
memset(&a, 0, sizeof(a));
|
||||
|
||||
a.size = sizeof(a);
|
||||
a.type = PERF_TYPE_TRACEPOINT;
|
||||
a.config = (uint64_t)events[e]->id;
|
||||
a.sample_period = 1;
|
||||
a.sample_type = PERF_SAMPLE_TID | PERF_SAMPLE_TIME | PERF_SAMPLE_CPU | PERF_SAMPLE_RAW;
|
||||
a.wakeup_events = 1;
|
||||
a.use_clockid = 1;
|
||||
a.clockid = CLOCK_MONOTONIC;
|
||||
a.disabled = 1;
|
||||
|
||||
int fd = (int)syscall(SYS_perf_event_open, &a, -1, cpus[c], -1, PERF_FLAG_FD_CLOEXEC);
|
||||
|
||||
if (fd < 0) {
|
||||
set_err(err, errn, "perf_event_open(%s:%s, cpu %d): %s",
|
||||
events[e]->system, events[e]->name, cpus[c], strerror(errno));
|
||||
tp_close(tp);
|
||||
return false;
|
||||
}
|
||||
|
||||
tp->fds[tp->nfds++] = fd;
|
||||
|
||||
if (tp->ring_fd[c] < 0) {
|
||||
|
||||
void *m = mmap(NULL, tp->map_len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
|
||||
if (m == MAP_FAILED) {
|
||||
set_err(err, errn, "mmap perf ring (cpu %d): %s", cpus[c], strerror(errno));
|
||||
tp_close(tp);
|
||||
return false;
|
||||
}
|
||||
|
||||
tp->ring[c] = m;
|
||||
tp->ring_fd[c] = fd;
|
||||
|
||||
struct epoll_event ee = { .events = EPOLLIN, .data.u32 = (uint32_t)c };
|
||||
epoll_ctl(tp->epfd, EPOLL_CTL_ADD, fd, &ee);
|
||||
|
||||
} else if (ioctl(fd, PERF_EVENT_IOC_SET_OUTPUT, tp->ring_fd[c]) < 0) {
|
||||
set_err(err, errn, "PERF_EVENT_IOC_SET_OUTPUT: %s", strerror(errno));
|
||||
tp_close(tp);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
for (int i = 0; i < tp->nfds; i++)
|
||||
ioctl(tp->fds[i], PERF_EVENT_IOC_ENABLE, 0);
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
static void ring_copy(uint8_t *dst, const uint8_t *base, uint64_t size, uint64_t pos, size_t len)
|
||||
{
|
||||
uint64_t off = pos % size;
|
||||
size_t first = (size_t)(size - off);
|
||||
|
||||
if (first >= len) {
|
||||
memcpy(dst, base + off, len);
|
||||
} else {
|
||||
memcpy(dst, base + off, first);
|
||||
memcpy(dst + first, base, len - first);
|
||||
}
|
||||
}
|
||||
|
||||
static int cmp_sample(const void *a, const void *b)
|
||||
{
|
||||
const tp_sample_t *x = a, *y = b;
|
||||
|
||||
return (x->time > y->time) - (x->time < y->time);
|
||||
}
|
||||
|
||||
static void dispatch(tp_t *tp, tp_cb cb, void *ctx)
|
||||
{
|
||||
qsort(tp->pend, tp->npend, sizeof(tp->pend[0]), cmp_sample);
|
||||
|
||||
for (int i = 0; i < tp->npend; i++)
|
||||
cb(ctx, &tp->pend[i]);
|
||||
|
||||
tp->npend = 0;
|
||||
}
|
||||
|
||||
static int drain_ring(tp_t *tp, int c, tp_cb cb, void *ctx)
|
||||
{
|
||||
struct perf_event_mmap_page *pg = tp->ring[c];
|
||||
long page = sysconf(_SC_PAGESIZE);
|
||||
uint64_t off = pg->data_offset ? pg->data_offset : (uint64_t)page;
|
||||
uint64_t size = pg->data_size ? pg->data_size : (uint64_t)page * RING_DATA_PAGES;
|
||||
const uint8_t *base = (const uint8_t *)pg + off;
|
||||
|
||||
uint64_t head = __atomic_load_n(&pg->data_head, __ATOMIC_ACQUIRE);
|
||||
uint64_t tail = pg->data_tail;
|
||||
int n = 0;
|
||||
|
||||
while (tail < head) {
|
||||
|
||||
struct perf_event_header hdr;
|
||||
ring_copy((uint8_t *)&hdr, base, size, tail, sizeof(hdr));
|
||||
|
||||
if (hdr.size < sizeof(hdr))
|
||||
break;
|
||||
|
||||
ring_copy(tp->scratch, base, size, tail, hdr.size);
|
||||
|
||||
const uint8_t *p = tp->scratch + sizeof(hdr);
|
||||
const uint8_t *end = tp->scratch + hdr.size;
|
||||
|
||||
if (hdr.type == PERF_RECORD_LOST && end - p >= 16) {
|
||||
|
||||
uint64_t lost;
|
||||
memcpy(&lost, p + 8, 8);
|
||||
tp->lost += lost;
|
||||
|
||||
} else if (hdr.type == PERF_RECORD_SAMPLE && end - p >= 28) {
|
||||
|
||||
tp_sample_t s;
|
||||
uint32_t v32[2];
|
||||
|
||||
memcpy(v32, p, 8); p += 8;
|
||||
s.pid = v32[0];
|
||||
s.tid = v32[1];
|
||||
memcpy(&s.time, p, 8); p += 8;
|
||||
memcpy(v32, p, 8); p += 8;
|
||||
s.cpu = v32[0];
|
||||
memcpy(&s.rawlen, p, 4); p += 4;
|
||||
s.raw = p;
|
||||
|
||||
if (s.rawlen >= 2 && p + s.rawlen <= end) {
|
||||
|
||||
uint16_t type;
|
||||
memcpy(&type, s.raw, 2);
|
||||
s.ev = NULL;
|
||||
|
||||
for (int e = 0; e < tp->nevents; e++)
|
||||
if (tp->events[e]->id == type)
|
||||
s.ev = tp->events[e];
|
||||
|
||||
if (s.ev && s.rawlen <= TP_MAX_RAW) {
|
||||
|
||||
if (tp->npend == TP_MAX_PENDING)
|
||||
dispatch(tp, cb, ctx);
|
||||
|
||||
memcpy(tp->pend_raw[tp->npend], s.raw, s.rawlen);
|
||||
s.raw = tp->pend_raw[tp->npend];
|
||||
tp->pend[tp->npend++] = s;
|
||||
n++;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
tail += hdr.size;
|
||||
}
|
||||
|
||||
__atomic_store_n(&pg->data_tail, tail, __ATOMIC_RELEASE);
|
||||
|
||||
return n;
|
||||
}
|
||||
|
||||
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx)
|
||||
{
|
||||
struct epoll_event ev[TP_MAX_CPUS];
|
||||
|
||||
if (epoll_wait(tp->epfd, ev, TP_MAX_CPUS, timeout_ms) < 0 && errno != EINTR)
|
||||
return -1;
|
||||
|
||||
/*
|
||||
* Drain every ring, not just the ones that woke us: samples from several
|
||||
* CPUs need to be handled together to keep per-camera order sane.
|
||||
*/
|
||||
int n = 0;
|
||||
|
||||
for (int c = 0; c < tp->ncpu; c++)
|
||||
if (tp->ring[c])
|
||||
n += drain_ring(tp, c, cb, ctx);
|
||||
|
||||
dispatch(tp, cb, ctx);
|
||||
|
||||
return n;
|
||||
}
|
||||
|
||||
void tp_close(tp_t *tp)
|
||||
{
|
||||
for (int i = 0; i < tp->nfds; i++) {
|
||||
ioctl(tp->fds[i], PERF_EVENT_IOC_DISABLE, 0);
|
||||
}
|
||||
|
||||
for (int c = 0; c < tp->ncpu; c++)
|
||||
if (tp->ring[c])
|
||||
munmap(tp->ring[c], tp->map_len);
|
||||
|
||||
for (int i = 0; i < tp->nfds; i++)
|
||||
close(tp->fds[i]);
|
||||
|
||||
if (tp->epfd >= 0)
|
||||
close(tp->epfd);
|
||||
|
||||
tp->nfds = 0;
|
||||
tp->epfd = -1;
|
||||
}
|
||||
@@ -0,0 +1,76 @@
|
||||
/*
|
||||
* tp - read kernel tracepoints system-wide through perf_event_open.
|
||||
*
|
||||
* One perf ring per CPU; every event on that CPU writes into it. Field
|
||||
* offsets come from the tracefs format files, so kernel layout changes don't
|
||||
* silently break parsing. Needs root (or CAP_PERFMON plus tracefs access).
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define TP_MAX_FIELDS 40
|
||||
#define TP_MAX_EVENTS 8
|
||||
#define TP_MAX_CPUS 64
|
||||
#define TP_MAX_PENDING 2048
|
||||
#define TP_MAX_RAW 256
|
||||
|
||||
typedef struct {
|
||||
char name[48];
|
||||
int offset;
|
||||
int size;
|
||||
bool is_signed;
|
||||
} tp_field_t;
|
||||
|
||||
typedef struct {
|
||||
char system[32];
|
||||
char name[48];
|
||||
int id;
|
||||
tp_field_t fields[TP_MAX_FIELDS];
|
||||
int nfields;
|
||||
} tp_event_t;
|
||||
|
||||
typedef struct {
|
||||
const tp_event_t *ev;
|
||||
const uint8_t *raw;
|
||||
uint32_t rawlen;
|
||||
uint64_t time; /* CLOCK_MONOTONIC ns */
|
||||
uint32_t cpu;
|
||||
uint32_t pid;
|
||||
uint32_t tid;
|
||||
} tp_sample_t;
|
||||
|
||||
typedef void (*tp_cb)(void *ctx, const tp_sample_t *s);
|
||||
|
||||
typedef struct {
|
||||
int ncpu;
|
||||
int ring_fd[TP_MAX_CPUS];
|
||||
void *ring[TP_MAX_CPUS];
|
||||
size_t map_len;
|
||||
int fds[TP_MAX_CPUS * TP_MAX_EVENTS];
|
||||
int nfds;
|
||||
int epfd;
|
||||
tp_event_t *events[TP_MAX_EVENTS];
|
||||
int nevents;
|
||||
uint64_t lost;
|
||||
uint8_t scratch[65536];
|
||||
/* samples drained from all rings, sorted by time before dispatch */
|
||||
tp_sample_t pend[TP_MAX_PENDING];
|
||||
uint8_t pend_raw[TP_MAX_PENDING][TP_MAX_RAW];
|
||||
int npend;
|
||||
} tp_t;
|
||||
|
||||
bool tp_event_load(tp_event_t *ev, const char *system, const char *name, char *err, size_t errn);
|
||||
int tp_field(const tp_event_t *ev, const char *name);
|
||||
int64_t tp_get(const tp_event_t *ev, int field, const uint8_t *raw, uint32_t rawlen);
|
||||
|
||||
bool tp_open(tp_t *tp, tp_event_t **events, int nevents, char *err, size_t errn);
|
||||
/*
|
||||
* Wait up to timeout_ms, then hand every pending sample to cb in time order,
|
||||
* across all CPUs. Returns samples read, -1 on error.
|
||||
*/
|
||||
int tp_poll(tp_t *tp, int timeout_ms, tp_cb cb, void *ctx);
|
||||
void tp_close(tp_t *tp);
|
||||
@@ -0,0 +1,783 @@
|
||||
/*
|
||||
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
|
||||
*
|
||||
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
|
||||
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
|
||||
*
|
||||
* Everything is discovered rather than hardcoded:
|
||||
* - XRService is found by scanning /proc for its cmdline.
|
||||
* - The V4L2 nodes and sensor subdevs it holds open come from /proc/<pid>/fd.
|
||||
* - Each node's geometry comes from VIDIOC_G_FMT on our own handle.
|
||||
* - Each node is traced back to its sensor through MEDIA_IOC_G_TOPOLOGY.
|
||||
* - Buffers are split into queues by allocation order: XRService opens a
|
||||
* sensor subdev, then allocates that camera's buffers.
|
||||
*/
|
||||
|
||||
#define _GNU_SOURCE
|
||||
|
||||
#include "xrcams.h"
|
||||
|
||||
#include <dirent.h>
|
||||
#include <errno.h>
|
||||
#include <fcntl.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/ioctl.h>
|
||||
#include <sys/stat.h>
|
||||
#include <sys/sysmacros.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <linux/media.h>
|
||||
|
||||
#ifndef MEDIA_ENT_F_CAM_SENSOR
|
||||
#define MEDIA_ENT_F_CAM_SENSOR 0x00020001
|
||||
#endif
|
||||
|
||||
#define MAX_FDENTS 4096
|
||||
#define MAX_TOPOS 8
|
||||
|
||||
enum fdkind { FD_DMABUF, FD_SUBDEV_SENSOR, FD_VIDEO };
|
||||
|
||||
typedef struct {
|
||||
int xfd;
|
||||
enum fdkind kind;
|
||||
size_t size;
|
||||
unsigned long ino;
|
||||
char sensor[XR_SENSOR_LEN];
|
||||
char path[64];
|
||||
} fdent_t;
|
||||
|
||||
typedef struct {
|
||||
struct media_v2_entity *ents;
|
||||
struct media_v2_interface *intfs;
|
||||
struct media_v2_pad *pads;
|
||||
struct media_v2_link *links;
|
||||
__u32 nents, nintfs, npads, nlinks;
|
||||
} topo_t;
|
||||
|
||||
static fdent_t fdents[MAX_FDENTS];
|
||||
static int nfdents;
|
||||
static topo_t topos[MAX_TOPOS];
|
||||
static int ntopos;
|
||||
|
||||
static void set_err(char *err, size_t n, const char *fmt, ...)
|
||||
{
|
||||
va_list ap;
|
||||
|
||||
va_start(ap, fmt);
|
||||
vsnprintf(err, n, fmt, ap);
|
||||
va_end(ap);
|
||||
}
|
||||
|
||||
void xr_slugify(const char *in, char *out, size_t n)
|
||||
{
|
||||
size_t i = 0;
|
||||
|
||||
for (; in[i] && i + 1 < n; i++)
|
||||
out[i] = (in[i] == ' ' || in[i] == '/') ? '_' : in[i];
|
||||
|
||||
out[i] = 0;
|
||||
}
|
||||
|
||||
/* --------------------------------------------------- media graph handling */
|
||||
|
||||
static void topo_free_all(void)
|
||||
{
|
||||
for (int i = 0; i < ntopos; i++) {
|
||||
free(topos[i].ents);
|
||||
free(topos[i].intfs);
|
||||
free(topos[i].pads);
|
||||
free(topos[i].links);
|
||||
}
|
||||
|
||||
ntopos = 0;
|
||||
}
|
||||
|
||||
static void topo_load_all(void)
|
||||
{
|
||||
for (int mi = 0; mi < MAX_TOPOS; mi++) {
|
||||
|
||||
char mpath[32];
|
||||
snprintf(mpath, sizeof(mpath), "/dev/media%d", mi);
|
||||
|
||||
int mfd = open(mpath, O_RDWR | O_CLOEXEC);
|
||||
|
||||
if (mfd < 0)
|
||||
continue;
|
||||
|
||||
struct media_v2_topology t;
|
||||
memset(&t, 0, sizeof(t));
|
||||
|
||||
if (ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) < 0) {
|
||||
close(mfd);
|
||||
continue;
|
||||
}
|
||||
|
||||
topo_t *o = &topos[ntopos];
|
||||
memset(o, 0, sizeof(*o));
|
||||
|
||||
o->nents = t.num_entities;
|
||||
o->nintfs = t.num_interfaces;
|
||||
o->npads = t.num_pads;
|
||||
o->nlinks = t.num_links;
|
||||
|
||||
o->ents = calloc(o->nents ? o->nents : 1, sizeof(*o->ents));
|
||||
o->intfs = calloc(o->nintfs ? o->nintfs : 1, sizeof(*o->intfs));
|
||||
o->pads = calloc(o->npads ? o->npads : 1, sizeof(*o->pads));
|
||||
o->links = calloc(o->nlinks ? o->nlinks : 1, sizeof(*o->links));
|
||||
|
||||
t.ptr_entities = (__u64)(uintptr_t)o->ents;
|
||||
t.ptr_interfaces = (__u64)(uintptr_t)o->intfs;
|
||||
t.ptr_pads = (__u64)(uintptr_t)o->pads;
|
||||
t.ptr_links = (__u64)(uintptr_t)o->links;
|
||||
|
||||
bool ok = o->ents && o->intfs && o->pads && o->links &&
|
||||
ioctl(mfd, MEDIA_IOC_G_TOPOLOGY, &t) == 0;
|
||||
close(mfd);
|
||||
|
||||
if (!ok) {
|
||||
free(o->ents); free(o->intfs); free(o->pads); free(o->links);
|
||||
continue;
|
||||
}
|
||||
|
||||
ntopos++;
|
||||
}
|
||||
}
|
||||
|
||||
static struct media_v2_entity *topo_entity(topo_t *t, __u32 id)
|
||||
{
|
||||
for (__u32 i = 0; i < t->nents; i++)
|
||||
if (t->ents[i].id == id)
|
||||
return &t->ents[i];
|
||||
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static struct media_v2_pad *topo_pad(topo_t *t, __u32 id)
|
||||
{
|
||||
for (__u32 i = 0; i < t->npads; i++)
|
||||
if (t->pads[i].id == id)
|
||||
return &t->pads[i];
|
||||
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static __u32 topo_entity_for_devnode(topo_t *t, dev_t rdev)
|
||||
{
|
||||
__u32 intf_id = 0;
|
||||
|
||||
for (__u32 i = 0; i < t->nintfs; i++)
|
||||
if (t->intfs[i].devnode.major == major(rdev) &&
|
||||
t->intfs[i].devnode.minor == minor(rdev)) {
|
||||
intf_id = t->intfs[i].id;
|
||||
break;
|
||||
}
|
||||
|
||||
if (!intf_id)
|
||||
return 0;
|
||||
|
||||
for (__u32 i = 0; i < t->nlinks; i++)
|
||||
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) == MEDIA_LNK_FL_INTERFACE_LINK &&
|
||||
t->links[i].source_id == intf_id)
|
||||
return t->links[i].sink_id;
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
/*
|
||||
* Walk upstream across enabled data links until a sensor is reached. A CSIPHY
|
||||
* carries two sensors on separate (sink, source) pad pairs, so re-enter on the
|
||||
* sink pad paired with the source pad we left through.
|
||||
*/
|
||||
static bool topo_walk_to_sensor(topo_t *t, __u32 ent_id, char *out, size_t outn)
|
||||
{
|
||||
int exit_pad_index = -1;
|
||||
|
||||
for (int hop = 0; hop < 32 && ent_id; hop++) {
|
||||
|
||||
struct media_v2_entity *e = topo_entity(t, ent_id);
|
||||
|
||||
if (!e)
|
||||
return false;
|
||||
|
||||
if (e->function == MEDIA_ENT_F_CAM_SENSOR) {
|
||||
snprintf(out, outn, "%s", e->name);
|
||||
return true;
|
||||
}
|
||||
|
||||
__u32 first_sink = 0, paired = 0;
|
||||
int nsinks = 0;
|
||||
|
||||
for (__u32 p = 0; p < t->npads; p++) {
|
||||
|
||||
if (t->pads[p].entity_id != ent_id || !(t->pads[p].flags & MEDIA_PAD_FL_SINK))
|
||||
continue;
|
||||
|
||||
nsinks++;
|
||||
|
||||
if (!first_sink)
|
||||
first_sink = t->pads[p].id;
|
||||
|
||||
if (exit_pad_index >= 1 && (int)t->pads[p].index == exit_pad_index - 1)
|
||||
paired = t->pads[p].id;
|
||||
}
|
||||
|
||||
__u32 sink_pad = (nsinks == 1) ? first_sink : (paired ? paired : first_sink);
|
||||
|
||||
if (!sink_pad)
|
||||
return false;
|
||||
|
||||
__u32 src_pad = 0;
|
||||
|
||||
for (__u32 i = 0; i < t->nlinks; i++) {
|
||||
|
||||
if ((t->links[i].flags & MEDIA_LNK_FL_LINK_TYPE) != MEDIA_LNK_FL_DATA_LINK)
|
||||
continue;
|
||||
|
||||
if (!(t->links[i].flags & MEDIA_LNK_FL_ENABLED))
|
||||
continue;
|
||||
|
||||
if (t->links[i].sink_id == sink_pad) {
|
||||
src_pad = t->links[i].source_id;
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
struct media_v2_pad *sp = src_pad ? topo_pad(t, src_pad) : NULL;
|
||||
|
||||
if (!sp)
|
||||
return false;
|
||||
|
||||
ent_id = sp->entity_id;
|
||||
exit_pad_index = (int)sp->index;
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
static bool sensor_for_video(dev_t rdev, char *out, size_t outn)
|
||||
{
|
||||
for (int i = 0; i < ntopos; i++) {
|
||||
|
||||
__u32 ent = topo_entity_for_devnode(&topos[i], rdev);
|
||||
|
||||
if (ent && topo_walk_to_sensor(&topos[i], ent, out, outn))
|
||||
return true;
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
static bool sensor_for_subdev(dev_t rdev, char *out, size_t outn)
|
||||
{
|
||||
for (int i = 0; i < ntopos; i++) {
|
||||
|
||||
__u32 id = topo_entity_for_devnode(&topos[i], rdev);
|
||||
struct media_v2_entity *e = id ? topo_entity(&topos[i], id) : NULL;
|
||||
|
||||
if (e && e->function == MEDIA_ENT_F_CAM_SENSOR) {
|
||||
snprintf(out, outn, "%s", e->name);
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
static const char *role_for_sensor(const char *sensor)
|
||||
{
|
||||
if (strstr(sensor, "og01a1b"))
|
||||
return "tracking"; /* 1056x1024 side fisheye */
|
||||
|
||||
if (strstr(sensor, "og0ve10"))
|
||||
return "tracking"; /* 640x480 upper */
|
||||
|
||||
if (strstr(sensor, "imx616"))
|
||||
return "passthrough"; /* 2464x2464 Arcturus color */
|
||||
|
||||
return "unknown";
|
||||
}
|
||||
|
||||
/* ------------------------------------------------- XRService / proc scan */
|
||||
|
||||
static pid_t find_process(const char *needle)
|
||||
{
|
||||
DIR *d = opendir("/proc");
|
||||
|
||||
if (!d)
|
||||
return 0;
|
||||
|
||||
struct dirent *e;
|
||||
pid_t found = 0;
|
||||
|
||||
while ((e = readdir(d))) {
|
||||
|
||||
if (e->d_name[0] < '0' || e->d_name[0] > '9')
|
||||
continue;
|
||||
|
||||
char path[288];
|
||||
snprintf(path, sizeof(path), "/proc/%s/cmdline", e->d_name);
|
||||
|
||||
FILE *f = fopen(path, "rb");
|
||||
|
||||
if (!f)
|
||||
continue;
|
||||
|
||||
char buf[512] = {0};
|
||||
size_t got = fread(buf, 1, sizeof(buf) - 1, f);
|
||||
fclose(f);
|
||||
|
||||
if (got == 0)
|
||||
continue;
|
||||
|
||||
const char *base = strrchr(buf, '/');
|
||||
base = base ? base + 1 : buf;
|
||||
|
||||
if (strstr(base, needle)) {
|
||||
found = (pid_t)atoi(e->d_name);
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
closedir(d);
|
||||
|
||||
return found;
|
||||
}
|
||||
|
||||
static bool read_dmabuf_size(pid_t pid, int fd, size_t *size, unsigned long *ino)
|
||||
{
|
||||
char path[64];
|
||||
snprintf(path, sizeof(path), "/proc/%d/fdinfo/%d", pid, fd);
|
||||
|
||||
FILE *f = fopen(path, "r");
|
||||
|
||||
if (!f)
|
||||
return false;
|
||||
|
||||
bool have = false;
|
||||
char line[256];
|
||||
|
||||
*ino = 0;
|
||||
|
||||
while (fgets(line, sizeof(line), f)) {
|
||||
|
||||
unsigned long long v;
|
||||
|
||||
if (sscanf(line, "size: %llu", &v) == 1) {
|
||||
*size = (size_t)v;
|
||||
have = true;
|
||||
} else if (sscanf(line, "ino: %llu", &v) == 1) {
|
||||
*ino = (unsigned long)v;
|
||||
}
|
||||
}
|
||||
|
||||
fclose(f);
|
||||
|
||||
return have;
|
||||
}
|
||||
|
||||
static int cmp_int(const void *a, const void *b)
|
||||
{
|
||||
return *(const int *)a - *(const int *)b;
|
||||
}
|
||||
|
||||
static bool scan_xr_fds(pid_t pid, char *err, size_t errn)
|
||||
{
|
||||
char dirpath[64];
|
||||
snprintf(dirpath, sizeof(dirpath), "/proc/%d/fd", pid);
|
||||
|
||||
DIR *d = opendir(dirpath);
|
||||
|
||||
if (!d) {
|
||||
set_err(err, errn, "opendir(%s): %s (are you root?)", dirpath, strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
static int fds[8192];
|
||||
int nfds = 0;
|
||||
struct dirent *e;
|
||||
|
||||
while ((e = readdir(d)) && nfds < (int)(sizeof(fds) / sizeof(fds[0])))
|
||||
if (e->d_name[0] >= '0' && e->d_name[0] <= '9')
|
||||
fds[nfds++] = atoi(e->d_name);
|
||||
|
||||
closedir(d);
|
||||
|
||||
qsort(fds, nfds, sizeof(int), cmp_int);
|
||||
|
||||
nfdents = 0;
|
||||
|
||||
for (int i = 0; i < nfds && nfdents < MAX_FDENTS; i++) {
|
||||
|
||||
char link[64], target[256];
|
||||
snprintf(link, sizeof(link), "/proc/%d/fd/%d", pid, fds[i]);
|
||||
|
||||
ssize_t n = readlink(link, target, sizeof(target) - 1);
|
||||
|
||||
if (n < 0)
|
||||
continue;
|
||||
|
||||
target[n] = 0;
|
||||
|
||||
fdent_t ent;
|
||||
memset(&ent, 0, sizeof(ent));
|
||||
ent.xfd = fds[i];
|
||||
|
||||
if (strstr(target, "dmabuf")) {
|
||||
|
||||
if (!read_dmabuf_size(pid, fds[i], &ent.size, &ent.ino))
|
||||
continue;
|
||||
|
||||
ent.kind = FD_DMABUF;
|
||||
|
||||
} else if (strncmp(target, "/dev/video", 10) == 0) {
|
||||
|
||||
ent.kind = FD_VIDEO;
|
||||
snprintf(ent.path, sizeof(ent.path), "%s", target);
|
||||
|
||||
} else if (strncmp(target, "/dev/v4l-subdev", 15) == 0) {
|
||||
|
||||
struct stat st;
|
||||
|
||||
if (stat(target, &st) < 0 || !sensor_for_subdev(st.st_rdev, ent.sensor, sizeof(ent.sensor)))
|
||||
continue;
|
||||
|
||||
ent.kind = FD_SUBDEV_SENSOR;
|
||||
|
||||
} else {
|
||||
continue;
|
||||
}
|
||||
|
||||
fdents[nfdents++] = ent;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------ camera discovery */
|
||||
|
||||
static void probe_cameras(xr_state_t *st)
|
||||
{
|
||||
int seen[64];
|
||||
int nseen = 0;
|
||||
|
||||
for (int i = 0; i < nfdents; i++) {
|
||||
|
||||
if (fdents[i].kind != FD_VIDEO)
|
||||
continue;
|
||||
|
||||
const char *path = fdents[i].path;
|
||||
int node = atoi(path + 10);
|
||||
bool dup = false;
|
||||
|
||||
for (int k = 0; k < nseen; k++)
|
||||
if (seen[k] == node)
|
||||
dup = true;
|
||||
|
||||
if (dup || st->ncameras >= XR_MAX_CAMERAS || nseen >= 64)
|
||||
continue;
|
||||
|
||||
seen[nseen++] = node;
|
||||
|
||||
int fd = open(path, O_RDWR | O_CLOEXEC);
|
||||
|
||||
if (fd < 0)
|
||||
continue;
|
||||
|
||||
struct v4l2_format fmt;
|
||||
memset(&fmt, 0, sizeof(fmt));
|
||||
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE;
|
||||
|
||||
xr_camera_t *c = &st->cameras[st->ncameras];
|
||||
memset(c, 0, sizeof(*c));
|
||||
|
||||
if (ioctl(fd, VIDIOC_G_FMT, &fmt) == 0) {
|
||||
|
||||
c->width = fmt.fmt.pix_mp.width;
|
||||
c->height = fmt.fmt.pix_mp.height;
|
||||
c->pixfmt = fmt.fmt.pix_mp.pixelformat;
|
||||
c->nplanes = fmt.fmt.pix_mp.num_planes;
|
||||
c->bytesperline = fmt.fmt.pix_mp.plane_fmt[0].bytesperline;
|
||||
|
||||
for (unsigned p = 0; p < c->nplanes && p < VIDEO_MAX_PLANES; p++)
|
||||
c->planesize[p] = fmt.fmt.pix_mp.plane_fmt[p].sizeimage;
|
||||
|
||||
} else {
|
||||
|
||||
memset(&fmt, 0, sizeof(fmt));
|
||||
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
|
||||
|
||||
if (ioctl(fd, VIDIOC_G_FMT, &fmt) < 0) {
|
||||
close(fd);
|
||||
continue;
|
||||
}
|
||||
|
||||
c->width = fmt.fmt.pix.width;
|
||||
c->height = fmt.fmt.pix.height;
|
||||
c->pixfmt = fmt.fmt.pix.pixelformat;
|
||||
c->nplanes = 1;
|
||||
c->bytesperline = fmt.fmt.pix.bytesperline;
|
||||
c->planesize[0] = fmt.fmt.pix.sizeimage;
|
||||
}
|
||||
|
||||
struct stat sb;
|
||||
|
||||
if (fstat(fd, &sb) == 0) {
|
||||
c->minor = minor(sb.st_rdev);
|
||||
sensor_for_video(sb.st_rdev, c->sensor, sizeof(c->sensor));
|
||||
}
|
||||
|
||||
close(fd);
|
||||
|
||||
if (!c->sensor[0])
|
||||
snprintf(c->sensor, sizeof(c->sensor), "unknown");
|
||||
|
||||
c->node = node;
|
||||
snprintf(c->path, sizeof(c->path), "%s", path);
|
||||
c->role = role_for_sensor(c->sensor);
|
||||
|
||||
st->ncameras++;
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
* qcom-camss can report bytesperline as the visible width while the VFE
|
||||
* writes a larger aligned pitch. sizeimage is right, so derive the pitch.
|
||||
*/
|
||||
unsigned xr_camera_stride(const xr_camera_t *c)
|
||||
{
|
||||
if (!c->height || !c->planesize[0])
|
||||
return c->bytesperline ? c->bytesperline : c->width;
|
||||
|
||||
double bpp = 1.0;
|
||||
|
||||
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21)
|
||||
bpp = 1.5;
|
||||
|
||||
unsigned s = (unsigned)((double)c->planesize[0] / ((double)c->height * bpp));
|
||||
|
||||
if (s >= c->width && s <= c->width * 4)
|
||||
return s;
|
||||
|
||||
return c->bytesperline ? c->bytesperline : c->width;
|
||||
}
|
||||
|
||||
/*
|
||||
* The Arcturus color cameras (arcimx616) claim 2464x2464 NV12, but measured on
|
||||
* 2026-09-28 their plane 0 holds 10-bit MIPI-packed YUV 4:2:0: 2464 luma rows
|
||||
* then 1232 rows of interleaved UV, each row 2464 packed pixels (3080 bytes)
|
||||
* padded to a 256-byte pitch (3328). Only the first 1972 pixels of a row carry
|
||||
* image; the rest are zero.
|
||||
*/
|
||||
#define IMX616_VALID_WIDTH 1972
|
||||
|
||||
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l)
|
||||
{
|
||||
memset(l, 0, sizeof(*l));
|
||||
l->height = c->height;
|
||||
|
||||
if (c->pixfmt == V4L2_PIX_FMT_NV12 && strstr(c->sensor, "imx616")) {
|
||||
|
||||
unsigned packed = (c->width * 5 + 3) / 4;
|
||||
|
||||
l->fmt = XR_FMT_YUV420_10P;
|
||||
l->pitch = (packed + 255) & ~255u;
|
||||
l->rows = c->height + c->height / 2;
|
||||
l->width = IMX616_VALID_WIDTH < c->width ? IMX616_VALID_WIDTH : c->width;
|
||||
return;
|
||||
}
|
||||
|
||||
l->pitch = xr_camera_stride(c);
|
||||
l->width = c->width < l->pitch ? c->width : l->pitch;
|
||||
|
||||
if (c->pixfmt == V4L2_PIX_FMT_NV12 || c->pixfmt == V4L2_PIX_FMT_NV21) {
|
||||
l->fmt = XR_FMT_NV12;
|
||||
l->rows = c->height + c->height / 2;
|
||||
} else {
|
||||
l->fmt = XR_FMT_GREY8;
|
||||
l->rows = c->height;
|
||||
}
|
||||
}
|
||||
|
||||
const char *xr_fmt_name(xr_fmt_t f)
|
||||
{
|
||||
switch (f) {
|
||||
case XR_FMT_GREY8: return "grey8";
|
||||
case XR_FMT_NV12: return "nv12";
|
||||
case XR_FMT_YUV420_10P: return "yuv420_10p";
|
||||
}
|
||||
|
||||
return "?";
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------- buffer grouping */
|
||||
|
||||
/*
|
||||
* XRService allocates one udmabuf per plane, plane 0 then plane 1, a whole
|
||||
* queue at a time right after opening the sensor's subdev. Plane 1 matches
|
||||
* VIDIOC_G_FMT exactly; plane 0 has slack, so it is matched with >=.
|
||||
*/
|
||||
static void build_groups(xr_state_t *st)
|
||||
{
|
||||
char current_sensor[XR_SENSOR_LEN] = "";
|
||||
|
||||
for (int i = 0; i < nfdents; i++) {
|
||||
|
||||
if (fdents[i].kind == FD_SUBDEV_SENSOR) {
|
||||
snprintf(current_sensor, sizeof(current_sensor), "%s", fdents[i].sensor);
|
||||
continue;
|
||||
}
|
||||
|
||||
if (fdents[i].kind != FD_DMABUF)
|
||||
continue;
|
||||
|
||||
if (i + 1 >= nfdents || fdents[i + 1].kind != FD_DMABUF)
|
||||
continue;
|
||||
|
||||
size_t s0 = fdents[i].size;
|
||||
size_t s1 = fdents[i + 1].size;
|
||||
bool match = false;
|
||||
|
||||
for (int c = 0; c < st->ncameras; c++) {
|
||||
|
||||
xr_camera_t *cam = &st->cameras[c];
|
||||
|
||||
if (cam->nplanes >= 2 && s1 == cam->planesize[1] && s0 >= cam->planesize[0]) {
|
||||
match = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (!match)
|
||||
continue;
|
||||
|
||||
xr_group_t *g = NULL;
|
||||
|
||||
if (st->ngroups > 0) {
|
||||
|
||||
xr_group_t *last = &st->groups[st->ngroups - 1];
|
||||
|
||||
if (last->planesize[0] == s0 && last->planesize[1] == s1 &&
|
||||
!strcmp(last->sensor, current_sensor))
|
||||
g = last;
|
||||
}
|
||||
|
||||
if (!g) {
|
||||
|
||||
if (st->ngroups >= XR_MAX_GROUPS)
|
||||
break;
|
||||
|
||||
g = &st->groups[st->ngroups++];
|
||||
memset(g, 0, sizeof(*g));
|
||||
g->planesize[0] = s0;
|
||||
g->planesize[1] = s1;
|
||||
snprintf(g->sensor, sizeof(g->sensor), "%s", current_sensor);
|
||||
}
|
||||
|
||||
if (g->nbufs < XR_MAX_RUNBUFS) {
|
||||
g->buf[g->nbufs].xfd = fdents[i].xfd;
|
||||
g->buf[g->nbufs].xfd1 = fdents[i + 1].xfd;
|
||||
g->buf[g->nbufs].size = s0;
|
||||
g->buf[g->nbufs].size1 = s1;
|
||||
g->nbufs++;
|
||||
}
|
||||
|
||||
i++; /* consume the plane 1 descriptor */
|
||||
}
|
||||
|
||||
int keep = 0;
|
||||
|
||||
for (int i = 0; i < st->ngroups; i++)
|
||||
if (st->groups[i].nbufs >= 4)
|
||||
st->groups[keep++] = st->groups[i];
|
||||
|
||||
st->ngroups = keep;
|
||||
|
||||
/*
|
||||
* Bind each run to a camera. The sensor marker alone can be wrong: XRService
|
||||
* sometimes opens another sensor's subdev (e.g. the idle color camera)
|
||||
* between an upper camera's subdev and its buffers, and two upper cameras
|
||||
* can resolve to the same sensor name. So a marker match must also fit the
|
||||
* camera's plane sizes, and each camera takes at most one run.
|
||||
*/
|
||||
for (int pass = 0; pass < 2; pass++)
|
||||
for (int i = 0; i < st->ngroups; i++) {
|
||||
|
||||
xr_group_t *g = &st->groups[i];
|
||||
|
||||
for (int c = 0; c < st->ncameras && !g->cam; c++) {
|
||||
|
||||
xr_camera_t *cam = &st->cameras[c];
|
||||
|
||||
if (pass == 0 && (!g->sensor[0] || strcmp(cam->sensor, g->sensor)))
|
||||
continue;
|
||||
|
||||
if (cam->nplanes < 2 || g->planesize[1] != cam->planesize[1] ||
|
||||
g->planesize[0] < cam->planesize[0])
|
||||
continue;
|
||||
|
||||
bool taken = false;
|
||||
|
||||
for (int k = 0; k < st->ngroups; k++)
|
||||
if (k != i && st->groups[k].cam == cam)
|
||||
taken = true;
|
||||
|
||||
if (!taken)
|
||||
g->cam = cam;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn)
|
||||
{
|
||||
memset(st, 0, sizeof(*st));
|
||||
|
||||
st->pid = find_process(process);
|
||||
|
||||
if (!st->pid) {
|
||||
set_err(err, errn, "%s is not running; start SteamVR on the headset first", process);
|
||||
return false;
|
||||
}
|
||||
|
||||
topo_load_all();
|
||||
|
||||
bool ok = scan_xr_fds(st->pid, err, errn);
|
||||
|
||||
if (ok) {
|
||||
probe_cameras(st);
|
||||
build_groups(st);
|
||||
}
|
||||
|
||||
topo_free_all();
|
||||
|
||||
return ok;
|
||||
}
|
||||
|
||||
void xr_print(const xr_state_t *st, FILE *f)
|
||||
{
|
||||
fprintf(f, "XRService pid %d\n", st->pid);
|
||||
|
||||
for (int i = 0; i < st->ncameras; i++) {
|
||||
|
||||
const xr_camera_t *c = &st->cameras[i];
|
||||
char fcc[5] = {
|
||||
(char)(c->pixfmt & 0xff), (char)((c->pixfmt >> 8) & 0xff),
|
||||
(char)((c->pixfmt >> 16) & 0xff), (char)((c->pixfmt >> 24) & 0xff), 0
|
||||
};
|
||||
|
||||
fprintf(f, " camera %-12s minor %-3u %-16s %ux%u %s pitch %u planes %zu %zu role=%s\n",
|
||||
c->path, c->minor, c->sensor, c->width, c->height, fcc,
|
||||
xr_camera_stride(c), c->planesize[0], c->planesize[1], c->role);
|
||||
}
|
||||
|
||||
for (int i = 0; i < st->ngroups; i++) {
|
||||
|
||||
const xr_group_t *g = &st->groups[i];
|
||||
|
||||
fprintf(f, " queue %d: %d buffers plane0=%zu plane1=%zu fds %d..%d sensor '%s' -> %s\n",
|
||||
i, g->nbufs, g->planesize[0], g->planesize[1],
|
||||
g->buf[0].xfd, g->buf[g->nbufs - 1].xfd1, g->sensor,
|
||||
g->cam ? g->cam->path : "(unbound)");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,82 @@
|
||||
/*
|
||||
* xrcams - find the headset cameras and the DMA-BUF queues XRService feeds them.
|
||||
*
|
||||
* Adapted from framecap.c in FrameEyeCameraFeed (vendor/FrameEyeCameraFeed),
|
||||
* MIT License, Copyright (c) 2026 Curtis English. See LICENSE.FrameEyeCameraFeed.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <sys/types.h>
|
||||
|
||||
#include <linux/videodev2.h>
|
||||
|
||||
#define XR_MAX_CAMERAS 16
|
||||
#define XR_MAX_GROUPS 32
|
||||
#define XR_MAX_RUNBUFS 128
|
||||
#define XR_SENSOR_LEN 64
|
||||
|
||||
typedef struct {
|
||||
int node; /* N from /dev/videoN */
|
||||
unsigned minor; /* char device minor, as tracepoints report it */
|
||||
char path[64];
|
||||
unsigned width;
|
||||
unsigned height;
|
||||
unsigned bytesperline;
|
||||
unsigned nplanes;
|
||||
size_t planesize[VIDEO_MAX_PLANES];
|
||||
uint32_t pixfmt;
|
||||
char sensor[XR_SENSOR_LEN]; /* media entity name of the sensor */
|
||||
const char *role;
|
||||
} xr_camera_t;
|
||||
|
||||
typedef struct {
|
||||
int xfd; /* plane 0 descriptor in XRService */
|
||||
int xfd1; /* plane 1 descriptor in XRService */
|
||||
size_t size;
|
||||
size_t size1;
|
||||
} xr_bufref_t;
|
||||
|
||||
/* One run of buffers XRService allocated for a camera queue, in allocation order. */
|
||||
typedef struct {
|
||||
size_t planesize[2];
|
||||
int nbufs;
|
||||
xr_bufref_t buf[XR_MAX_RUNBUFS];
|
||||
char sensor[XR_SENSOR_LEN]; /* from the preceding sensor subdev */
|
||||
xr_camera_t *cam;
|
||||
} xr_group_t;
|
||||
|
||||
typedef struct {
|
||||
pid_t pid;
|
||||
xr_camera_t cameras[XR_MAX_CAMERAS];
|
||||
int ncameras;
|
||||
xr_group_t groups[XR_MAX_GROUPS];
|
||||
int ngroups;
|
||||
} xr_state_t;
|
||||
|
||||
typedef enum {
|
||||
XR_FMT_GREY8, /* 8-bit mono */
|
||||
XR_FMT_NV12, /* 8-bit Y plane then interleaved UV, same pitch */
|
||||
XR_FMT_YUV420_10P /* like NV12, but 10-bit MIPI-packed (4 px in 5 bytes) */
|
||||
} xr_fmt_t;
|
||||
|
||||
/* Where the image really sits in plane 0; V4L2's numbers can be misleading. */
|
||||
typedef struct {
|
||||
xr_fmt_t fmt;
|
||||
unsigned pitch; /* bytes per row */
|
||||
unsigned rows; /* rows in plane 0: luma, plus chroma for YUV */
|
||||
unsigned width; /* valid pixels per row */
|
||||
unsigned height; /* luma rows */
|
||||
} xr_layout_t;
|
||||
|
||||
/* Scan XRService's descriptors and the media graph. Needs root. */
|
||||
bool xr_discover(xr_state_t *st, const char *process, char *err, size_t errn);
|
||||
unsigned xr_camera_stride(const xr_camera_t *c);
|
||||
void xr_camera_layout(const xr_camera_t *c, xr_layout_t *l);
|
||||
const char *xr_fmt_name(xr_fmt_t f);
|
||||
void xr_print(const xr_state_t *st, FILE *f);
|
||||
void xr_slugify(const char *in, char *out, size_t n);
|
||||
@@ -0,0 +1,66 @@
|
||||
/*
|
||||
* fh_gestures - pinch state frame-hands' tracker publishes for input: look at something
|
||||
* and pinch to click it, pinch and move to drag (the Vision Pro model, with the eye
|
||||
* tracker doing the looking). $XDG_RUNTIME_DIR/frame-hands/gestures, next to the hands
|
||||
* file, with the same sequence lock (read seq, copy, read seq again; use the copy only if
|
||||
* both reads are the same even number) and the same frame: metres in the head frame at
|
||||
* capture time, OpenVR's HMD frame (+x right, +y up, -z forward).
|
||||
*
|
||||
* One slot per side: pinch[0] is the left hand, pinch[1] the right. A pinch follows the
|
||||
* hand it began on until it ends. It begins when the thumb and index tips close within
|
||||
* begin_m and ends when they open past end_m (the gap between keeps it from flickering),
|
||||
* or when the hand stays lost too long (FH_PINCH_LOST).
|
||||
*
|
||||
* Don't miss short pinches: a reader that polls slower than a quick tap still sees it,
|
||||
* because begins and ends count every pinch. When begins changed, a pinch began at
|
||||
* begin_ns; when ends changed, one ended at end_ns. begins - ends is 1 while pinching.
|
||||
*
|
||||
* Drags: point is where the pinch is now, begin_point where it began. Turn each into the
|
||||
* room with the HMD pose at its capture time (capture_ns, begin_ns) before subtracting,
|
||||
* so turning your head doesn't drag.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <assert.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define FH_GESTURES_MAGIC "FHGEST01"
|
||||
#define FH_GESTURES_VERSION 1
|
||||
|
||||
enum {
|
||||
FH_PINCH_TRACKED = 1u << 0, /* the hand was tracked in this frame */
|
||||
FH_PINCH_DOWN = 1u << 1, /* pinching now */
|
||||
FH_PINCH_LOST = 1u << 2, /* the last pinch ended because the hand was lost */
|
||||
};
|
||||
|
||||
typedef struct {
|
||||
uint32_t flags; /* FH_PINCH_* */
|
||||
uint32_t hand_id; /* fh_hand_t.id of the hand, 0 if none */
|
||||
uint32_t begins; /* pinches begun so far */
|
||||
uint32_t ends; /* pinches ended so far */
|
||||
uint64_t begin_ns; /* capture time (CLOCK_MONOTONIC) the current or */
|
||||
/* last pinch began */
|
||||
uint64_t end_ns; /* ... the last pinch ended */
|
||||
float distance; /* thumb tip to index tip, m, at this user's hand */
|
||||
/* size */
|
||||
float strength; /* 0 open (end_m or more) .. 1 closed (begin_m) */
|
||||
float point[3]; /* between the thumb and index tips */
|
||||
float begin_point[3]; /* point when the current or last pinch began */
|
||||
} fh_pinch_t; /* 64 bytes */
|
||||
|
||||
typedef struct {
|
||||
char magic[8];
|
||||
uint32_t version;
|
||||
uint32_t size;
|
||||
volatile uint64_t seq;
|
||||
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
|
||||
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
|
||||
float begin_m; /* the thresholds in use */
|
||||
float end_m;
|
||||
uint8_t reserved[16];
|
||||
fh_pinch_t pinch[2]; /* [0] left hand, [1] right hand */
|
||||
} fh_gestures_t;
|
||||
|
||||
static_assert(sizeof(fh_pinch_t) == 64, "fh_pinch_t layout");
|
||||
static_assert(sizeof(fh_gestures_t) == 64 + 2 * 64, "fh_gestures_t layout");
|
||||
@@ -0,0 +1,58 @@
|
||||
/*
|
||||
* fh_hands - the tracked-hands file frame-hands' tracker publishes
|
||||
* ($XDG_RUNTIME_DIR/frame-hands/hands, directory mode 0700), rewritten in place
|
||||
* under a sequence lock: read seq, copy, read seq again; use the copy only if
|
||||
* both reads are the same even number.
|
||||
*
|
||||
* Positions are metres in the head frame at capture time, which is OpenVR's HMD
|
||||
* frame (+x right, +y up, -z forward). Turn them into the room with the HMD pose
|
||||
* at capture_ns (CLOCK_MONOTONIC). Writers: trackd (fh-tracker), tracker/publish.py.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <assert.h>
|
||||
#include <stdint.h>
|
||||
|
||||
#define FH_HANDS_MAGIC "FHHANDS1"
|
||||
#define FH_HANDS_VERSION 1
|
||||
#define FH_HANDS_MAX_HANDS 2
|
||||
#define FH_HANDS_MAX_CAPSULES 64
|
||||
|
||||
enum {
|
||||
FH_HAND_RIGHT = 1u << 0, /* else the left hand */
|
||||
FH_HAND_STEREO = 1u << 1, /* triangulated from two or more cameras */
|
||||
};
|
||||
|
||||
typedef struct {
|
||||
uint32_t id; /* stays the same while the hand is tracked */
|
||||
uint32_t flags; /* FH_HAND_* */
|
||||
float confidence;
|
||||
float reserved;
|
||||
float pts[21][3]; /* MediaPipe hand landmarks */
|
||||
uint32_t ncapsules; /* this hand's capsules, which follow the */
|
||||
/* previous hands' in capsules[] */
|
||||
} fh_hand_t; /* 272 bytes */
|
||||
|
||||
typedef struct {
|
||||
float a[3], b[3]; /* segment ends */
|
||||
float ra, rb; /* radius at each end */
|
||||
} fh_capsule_t; /* 32 bytes: the hand's shape, to cut out */
|
||||
|
||||
typedef struct {
|
||||
char magic[8];
|
||||
uint32_t version;
|
||||
uint32_t size;
|
||||
volatile uint64_t seq;
|
||||
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
|
||||
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
|
||||
uint32_t nhands;
|
||||
uint32_t ncapsules;
|
||||
uint8_t reserved[16];
|
||||
fh_hand_t hands[FH_HANDS_MAX_HANDS];
|
||||
fh_capsule_t capsules[FH_HANDS_MAX_CAPSULES];
|
||||
} fh_hands_t;
|
||||
|
||||
static_assert(sizeof(fh_hand_t) == 272, "fh_hand_t layout");
|
||||
static_assert(sizeof(fh_capsule_t) == 32, "fh_capsule_t layout");
|
||||
static_assert(sizeof(fh_hands_t) == 64 + 2 * 272 + 64 * 32, "fh_hands_t layout");
|
||||
Binary file not shown.
@@ -0,0 +1,79 @@
|
||||
7767517
|
||||
77 90
|
||||
Input in0 0 1 in0
|
||||
Convolution convclip_0 1 1 in0 2 0=24 1=3 3=2 15=1 16=1 5=1 6=648 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 4=1 5=1 6=216 7=24 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_10 1 1 3 4 0=16 1=1 5=1 6=384 8=2
|
||||
Split splitncnn_0 1 2 4 5 6
|
||||
Convolution convclip_1 1 1 6 7 0=64 1=1 5=1 6=1024 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 3=2 15=1 16=1 5=1 6=576 7=64 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_12 1 1 8 9 0=16 1=1 5=1 6=1024 8=2
|
||||
Pooling maxpool2d_1 1 1 5 10 1=2 2=2 5=1
|
||||
BinaryOp add_0 2 1 9 10 11
|
||||
Split splitncnn_1 1 2 11 12 13
|
||||
Convolution convclip_2 1 1 13 14 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 4=1 5=1 6=864 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_14 1 1 15 16 0=16 1=1 5=1 6=1536 8=2
|
||||
BinaryOp add_1 2 1 16 12 17
|
||||
Convolution convclip_3 1 1 17 18 0=96 1=1 5=1 6=1536 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 7=96 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_16 1 1 19 20 0=24 1=1 5=1 6=2304 8=2
|
||||
Split splitncnn_2 1 2 20 21 22
|
||||
Convolution convclip_4 1 1 22 23 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 4=2 5=1 6=3600 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_18 1 1 24 25 0=24 1=1 5=1 6=3456 8=2
|
||||
BinaryOp add_2 2 1 25 21 26
|
||||
Convolution convclip_5 1 1 26 27 0=144 1=1 5=1 6=3456 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 3=2 15=1 16=1 5=1 6=1296 7=144 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_20 1 1 28 29 0=48 1=1 5=1 6=6912 8=2
|
||||
Split splitncnn_3 1 2 29 30 31
|
||||
Convolution convclip_6 1 1 31 32 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_22 1 1 33 34 0=48 1=1 5=1 6=13824 8=2
|
||||
BinaryOp add_3 2 1 34 30 35
|
||||
Split splitncnn_4 1 2 35 36 37
|
||||
Convolution convclip_7 1 1 37 38 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 4=1 5=1 6=2592 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_24 1 1 39 40 0=48 1=1 5=1 6=13824 8=2
|
||||
BinaryOp add_4 2 1 40 36 41
|
||||
Convolution convclip_8 1 1 41 42 0=288 1=1 5=1 6=13824 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 4=2 5=1 6=7200 7=288 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_26 1 1 43 44 0=64 1=1 5=1 6=18432 8=2
|
||||
Split splitncnn_5 1 2 44 45 46
|
||||
Convolution convclip_9 1 1 46 47 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_28 1 1 48 49 0=64 1=1 5=1 6=24576 8=2
|
||||
BinaryOp add_5 2 1 49 45 50
|
||||
Split splitncnn_6 1 2 50 51 52
|
||||
Convolution convclip_10 1 1 52 53 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 4=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_30 1 1 54 55 0=64 1=1 5=1 6=24576 8=2
|
||||
BinaryOp add_6 2 1 55 51 56
|
||||
Convolution convclip_11 1 1 56 57 0=384 1=1 5=1 6=24576 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 3=2 4=1 15=2 16=2 5=1 6=9600 7=384 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_32 1 1 58 59 0=112 1=1 5=1 6=43008 8=2
|
||||
Split splitncnn_7 1 2 59 60 61
|
||||
Convolution convclip_12 1 1 61 62 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_34 1 1 63 64 0=112 1=1 5=1 6=75264 8=2
|
||||
BinaryOp add_7 2 1 64 60 65
|
||||
Split splitncnn_8 1 2 65 66 67
|
||||
Convolution convclip_13 1 1 67 68 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_36 1 1 69 70 0=112 1=1 5=1 6=75264 8=2
|
||||
BinaryOp add_8 2 1 70 66 71
|
||||
Split splitncnn_9 1 2 71 72 73
|
||||
Convolution convclip_14 1 1 73 74 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 4=2 5=1 6=16800 7=672 8=101 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Convolution conv_38 1 1 75 76 0=112 1=1 5=1 6=75264 8=2
|
||||
BinaryOp add_9 2 1 76 72 77
|
||||
Convolution convclip_15 1 1 77 78 0=672 1=1 5=1 6=75264 8=102 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 4=1 5=1 6=6048 7=672 8=1 9=3 -23310=2,0.000000e+00,6.000000e+00
|
||||
Pooling gap_0 1 1 79 80 0=1 4=1
|
||||
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
|
||||
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
|
||||
Split splitncnn_10 1 4 82 83 84 85 86
|
||||
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336 8=2
|
||||
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336 8=2
|
||||
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 8=2 9=4
|
||||
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 8=2 9=4
|
||||
Binary file not shown.
@@ -0,0 +1,79 @@
|
||||
7767517
|
||||
77 90
|
||||
Input in0 0 1 in0
|
||||
Convolution convclip_0 1 1 in0 2 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=648 9=3
|
||||
ConvolutionDepthWise convdwclip_0 1 1 2 3 0=24 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=216 7=24 9=3
|
||||
Convolution conv_10 1 1 3 4 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=384
|
||||
Split splitncnn_0 1 2 4 5 6
|
||||
Convolution convclip_1 1 1 6 7 0=64 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024 9=3
|
||||
ConvolutionDepthWise convdwclip_1 1 1 7 8 0=64 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=576 7=64 9=3
|
||||
Convolution conv_12 1 1 8 9 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
Pooling maxpool2d_1 1 1 5 10 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
BinaryOp add_0 2 1 9 10 11 0=0
|
||||
Split splitncnn_1 1 2 11 12 13
|
||||
Convolution convclip_2 1 1 13 14 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
|
||||
ConvolutionDepthWise convdwclip_2 1 1 14 15 0=96 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=864 7=96 9=3
|
||||
Convolution conv_14 1 1 15 16 0=16 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
|
||||
BinaryOp add_1 2 1 16 12 17 0=0
|
||||
Convolution convclip_3 1 1 17 18 0=96 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536 9=3
|
||||
ConvolutionDepthWise convdwclip_3 1 1 18 19 0=96 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400 7=96 9=3
|
||||
Convolution conv_16 1 1 19 20 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2304
|
||||
Split splitncnn_2 1 2 20 21 22
|
||||
Convolution convclip_4 1 1 22 23 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
|
||||
ConvolutionDepthWise convdwclip_4 1 1 23 24 0=144 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3600 7=144 9=3
|
||||
Convolution conv_18 1 1 24 25 0=24 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456
|
||||
BinaryOp add_2 2 1 25 21 26 0=0
|
||||
Convolution convclip_5 1 1 26 27 0=144 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=3456 9=3
|
||||
ConvolutionDepthWise convdwclip_5 1 1 27 28 0=144 1=3 -23310=2,0.0,6.0 11=3 12=1 13=2 14=0 15=1 16=1 2=1 3=2 4=0 5=1 6=1296 7=144 9=3
|
||||
Convolution conv_20 1 1 28 29 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=6912
|
||||
Split splitncnn_3 1 2 29 30 31
|
||||
Convolution convclip_6 1 1 31 32 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
|
||||
ConvolutionDepthWise convdwclip_6 1 1 32 33 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
|
||||
Convolution conv_22 1 1 33 34 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
|
||||
BinaryOp add_3 2 1 34 30 35 0=0
|
||||
Split splitncnn_4 1 2 35 36 37
|
||||
Convolution convclip_7 1 1 37 38 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
|
||||
ConvolutionDepthWise convdwclip_7 1 1 38 39 0=288 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=2592 7=288 9=3
|
||||
Convolution conv_24 1 1 39 40 0=48 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824
|
||||
BinaryOp add_4 2 1 40 36 41 0=0
|
||||
Convolution convclip_8 1 1 41 42 0=288 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=13824 9=3
|
||||
ConvolutionDepthWise convdwclip_8 1 1 42 43 0=288 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=7200 7=288 9=3
|
||||
Convolution conv_26 1 1 43 44 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=18432
|
||||
Split splitncnn_5 1 2 44 45 46
|
||||
Convolution convclip_9 1 1 46 47 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
|
||||
ConvolutionDepthWise convdwclip_9 1 1 47 48 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
|
||||
Convolution conv_28 1 1 48 49 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
|
||||
BinaryOp add_5 2 1 49 45 50 0=0
|
||||
Split splitncnn_6 1 2 50 51 52
|
||||
Convolution convclip_10 1 1 52 53 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
|
||||
ConvolutionDepthWise convdwclip_10 1 1 53 54 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=9600 7=384 9=3
|
||||
Convolution conv_30 1 1 54 55 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576
|
||||
BinaryOp add_6 2 1 55 51 56 0=0
|
||||
Convolution convclip_11 1 1 56 57 0=384 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=24576 9=3
|
||||
ConvolutionDepthWise convdwclip_11 1 1 57 58 0=384 1=5 -23310=2,0.0,6.0 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=9600 7=384 9=3
|
||||
Convolution conv_32 1 1 58 59 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=43008
|
||||
Split splitncnn_7 1 2 59 60 61
|
||||
Convolution convclip_12 1 1 61 62 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_12 1 1 62 63 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
|
||||
Convolution conv_34 1 1 63 64 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
|
||||
BinaryOp add_7 2 1 64 60 65 0=0
|
||||
Split splitncnn_8 1 2 65 66 67
|
||||
Convolution convclip_13 1 1 67 68 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_13 1 1 68 69 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
|
||||
Convolution conv_36 1 1 69 70 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
|
||||
BinaryOp add_8 2 1 70 66 71 0=0
|
||||
Split splitncnn_9 1 2 71 72 73
|
||||
Convolution convclip_14 1 1 73 74 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_14 1 1 74 75 0=672 1=5 -23310=2,0.0,6.0 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=16800 7=672 9=3
|
||||
Convolution conv_38 1 1 75 76 0=112 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264
|
||||
BinaryOp add_9 2 1 76 72 77 0=0
|
||||
Convolution convclip_15 1 1 77 78 0=672 1=1 -23310=2,0.0,6.0 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=75264 9=3
|
||||
ConvolutionDepthWise convdwclip_15 1 1 78 79 0=672 1=3 -23310=2,0.0,6.0 11=3 12=1 13=1 14=1 2=1 3=1 4=1 5=1 6=6048 7=672 9=3
|
||||
Pooling gap_0 1 1 79 80 0=1 4=1
|
||||
Reshape reshape_45 1 1 80 81 0=1 1=1 2=-1
|
||||
Squeeze squeeze_78 1 1 81 82 -23303=2,1,2
|
||||
Split splitncnn_10 1 4 82 83 84 85 86
|
||||
InnerProduct linear_42 1 1 84 out0 0=63 1=1 2=42336
|
||||
InnerProduct linear_43 1 1 83 out3 0=63 1=1 2=42336
|
||||
InnerProduct fcsigmoid_0 1 1 85 out2 0=1 1=1 2=672 9=4
|
||||
InnerProduct fcsigmoid_1 1 1 86 out1 0=1 1=1 2=672 9=4
|
||||
Binary file not shown.
@@ -0,0 +1,151 @@
|
||||
7767517
|
||||
149 177
|
||||
Input in0 0 1 in0
|
||||
Convolution padconv_0 1 1 in0 2 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=2400 8=2
|
||||
PReLU prelu_41 1 1 2 3 0=32
|
||||
Split splitncnn_0 1 2 3 4 5
|
||||
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 4=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_12 1 1 6 7 0=32 1=1 5=1 6=1024 8=2
|
||||
BinaryOp add_0 2 1 4 7 8
|
||||
PReLU prelu_42 1 1 8 9 0=32
|
||||
Split splitncnn_1 1 2 9 10 11
|
||||
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 4=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_13 1 1 12 13 0=32 1=1 5=1 6=1024 8=2
|
||||
BinaryOp add_1 2 1 10 13 14
|
||||
PReLU prelu_43 1 1 14 15 0=32
|
||||
Split splitncnn_2 1 2 15 16 17
|
||||
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 4=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_14 1 1 18 19 0=32 1=1 5=1 6=1024 8=2
|
||||
BinaryOp add_2 2 1 16 19 20
|
||||
PReLU prelu_44 1 1 20 21 0=32
|
||||
Split splitncnn_3 1 2 21 22 23
|
||||
Pooling maxpool2d_2 1 1 22 24 1=2 2=2 5=1
|
||||
Padding Pad_16 1 1 24 25 8=32
|
||||
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 3=2 4=1 15=2 16=2 5=1 6=800 7=32 8=101
|
||||
Convolution conv_15 1 1 26 27 0=64 1=1 5=1 6=2048 8=2
|
||||
BinaryOp add_3 2 1 25 27 28
|
||||
PReLU prelu_45 1 1 28 29 0=64
|
||||
Split splitncnn_4 1 2 29 30 31
|
||||
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_16 1 1 32 33 0=64 1=1 5=1 6=4096 8=2
|
||||
BinaryOp add_4 2 1 30 33 34
|
||||
PReLU prelu_46 1 1 34 35 0=64
|
||||
Split splitncnn_5 1 2 35 36 37
|
||||
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_17 1 1 38 39 0=64 1=1 5=1 6=4096 8=2
|
||||
BinaryOp add_5 2 1 36 39 40
|
||||
PReLU prelu_47 1 1 40 41 0=64
|
||||
Split splitncnn_6 1 2 41 42 43
|
||||
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 4=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_18 1 1 44 45 0=64 1=1 5=1 6=4096 8=2
|
||||
BinaryOp add_6 2 1 42 45 46
|
||||
PReLU prelu_48 1 1 46 47 0=64
|
||||
Split splitncnn_7 1 2 47 48 49
|
||||
Pooling maxpool2d_3 1 1 48 50 1=2 2=2 5=1
|
||||
Padding Pad_34 1 1 50 51 8=64
|
||||
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 3=2 4=1 15=2 16=2 5=1 6=1600 7=64 8=101
|
||||
Convolution conv_19 1 1 52 53 0=128 1=1 5=1 6=8192 8=2
|
||||
BinaryOp add_7 2 1 51 53 54
|
||||
PReLU prelu_49 1 1 54 55 0=128
|
||||
Split splitncnn_8 1 2 55 56 57
|
||||
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_20 1 1 58 59 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_8 2 1 56 59 60
|
||||
PReLU prelu_50 1 1 60 61 0=128
|
||||
Split splitncnn_9 1 2 61 62 63
|
||||
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_21 1 1 64 65 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_9 2 1 62 65 66
|
||||
PReLU prelu_51 1 1 66 67 0=128
|
||||
Split splitncnn_10 1 2 67 68 69
|
||||
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_22 1 1 70 71 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_10 2 1 68 71 72
|
||||
PReLU prelu_52 1 1 72 73 0=128
|
||||
Split splitncnn_11 1 3 73 74 75 76
|
||||
Pooling maxpool2d_4 1 1 75 77 1=2 2=2 5=1
|
||||
Padding Pad_52 1 1 77 78 8=128
|
||||
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 3=2 4=1 15=2 16=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_23 1 1 79 80 0=256 1=1 5=1 6=32768 8=2
|
||||
BinaryOp add_11 2 1 78 80 81
|
||||
PReLU prelu_53 1 1 81 82 0=256
|
||||
Split splitncnn_12 1 2 82 83 84
|
||||
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_24 1 1 85 86 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_12 2 1 83 86 87
|
||||
PReLU prelu_54 1 1 87 88 0=256
|
||||
Split splitncnn_13 1 2 88 89 90
|
||||
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_25 1 1 91 92 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_13 2 1 89 92 93
|
||||
PReLU prelu_55 1 1 93 94 0=256
|
||||
Split splitncnn_14 1 2 94 95 96
|
||||
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_26 1 1 97 98 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_14 2 1 95 98 99
|
||||
PReLU prelu_56 1 1 99 100 0=256
|
||||
Split splitncnn_15 1 3 100 101 102 103
|
||||
Pooling maxpool2d_5 1 1 102 104 1=2 2=2 5=1
|
||||
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 3=2 4=1 15=2 16=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_27 1 1 105 106 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_15 2 1 104 106 107
|
||||
PReLU prelu_57 1 1 107 108 0=256
|
||||
Split splitncnn_16 1 2 108 109 110
|
||||
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_28 1 1 111 112 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_16 2 1 109 112 113
|
||||
PReLU prelu_58 1 1 113 114 0=256
|
||||
Split splitncnn_17 1 2 114 115 116
|
||||
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_29 1 1 117 118 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_17 2 1 115 118 119
|
||||
PReLU prelu_59 1 1 119 120 0=256
|
||||
Split splitncnn_18 1 2 120 121 122
|
||||
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_30 1 1 123 124 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_18 2 1 121 124 125
|
||||
PReLU prelu_60 1 1 125 126 0=256
|
||||
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12
|
||||
Convolution conv_31 1 1 127 128 0=256 1=1 5=1 6=65536 8=2
|
||||
PReLU prelu_61 1 1 128 129 0=256
|
||||
BinaryOp add_19 2 1 101 129 130
|
||||
Split splitncnn_19 1 2 130 131 132
|
||||
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_32 1 1 133 134 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_20 2 1 131 134 135
|
||||
PReLU prelu_62 1 1 135 136 0=256
|
||||
Split splitncnn_20 1 2 136 137 138
|
||||
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 4=2 5=1 6=6400 7=256 8=101
|
||||
Convolution conv_33 1 1 139 140 0=256 1=1 5=1 6=65536 8=2
|
||||
BinaryOp add_21 2 1 137 140 141
|
||||
PReLU prelu_63 1 1 141 142 0=256
|
||||
Split splitncnn_21 1 3 142 143 144 145
|
||||
Convolution conv_34 1 1 145 146 0=108 1=1 5=1 6=27648 8=2
|
||||
Permute permute_68 1 1 146 147 0=3
|
||||
Reshape reshape_72 1 1 147 148 0=18 1=864
|
||||
Convolution conv_35 1 1 144 149 0=6 1=1 5=1 6=1536 8=2
|
||||
Permute permute_69 1 1 149 150 0=3
|
||||
Reshape reshape_73 1 1 150 151 0=1 1=864
|
||||
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24
|
||||
Convolution conv_36 1 1 152 153 0=128 1=1 5=1 6=32768 8=2
|
||||
PReLU prelu_64 1 1 153 154 0=128
|
||||
BinaryOp add_22 2 1 74 154 155
|
||||
Split splitncnn_22 1 2 155 156 157
|
||||
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_37 1 1 158 159 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_23 2 1 156 159 160
|
||||
PReLU prelu_65 1 1 160 161 0=128
|
||||
Split splitncnn_23 1 2 161 162 163
|
||||
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 4=2 5=1 6=3200 7=128 8=101
|
||||
Convolution conv_38 1 1 164 165 0=128 1=1 5=1 6=16384 8=2
|
||||
BinaryOp add_24 2 1 162 165 166
|
||||
PReLU prelu_66 1 1 166 167 0=128
|
||||
Split splitncnn_24 1 2 167 168 169
|
||||
Convolution conv_39 1 1 169 170 0=36 1=1 5=1 6=4608 8=2
|
||||
Permute permute_70 1 1 170 171 0=3
|
||||
Reshape reshape_74 1 1 171 172 0=18 1=1152
|
||||
Concat cat_0 2 1 172 148 out0
|
||||
Convolution conv_40 1 1 168 174 0=2 1=1 5=1 6=256 8=2
|
||||
Permute permute_71 1 1 174 175 0=3
|
||||
Reshape reshape_75 1 1 175 176 0=1 1=1152
|
||||
Concat cat_1 2 1 176 151 out1
|
||||
Binary file not shown.
@@ -0,0 +1,151 @@
|
||||
7767517
|
||||
149 177
|
||||
Input in0 0 1 in0
|
||||
Convolution padconv_0 1 1 in0 2 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=2400
|
||||
PReLU prelu_41 1 1 2 3 0=32
|
||||
Split splitncnn_0 1 2 3 4 5
|
||||
ConvolutionDepthWise convdw_76 1 1 5 6 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
|
||||
Convolution conv_12 1 1 6 7 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
BinaryOp add_0 2 1 4 7 8 0=0
|
||||
PReLU prelu_42 1 1 8 9 0=32
|
||||
Split splitncnn_1 1 2 9 10 11
|
||||
ConvolutionDepthWise convdw_77 1 1 11 12 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
|
||||
Convolution conv_13 1 1 12 13 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
BinaryOp add_1 2 1 10 13 14 0=0
|
||||
PReLU prelu_43 1 1 14 15 0=32
|
||||
Split splitncnn_2 1 2 15 16 17
|
||||
ConvolutionDepthWise convdw_78 1 1 17 18 0=32 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=800 7=32
|
||||
Convolution conv_14 1 1 18 19 0=32 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1024
|
||||
BinaryOp add_2 2 1 16 19 20 0=0
|
||||
PReLU prelu_44 1 1 20 21 0=32
|
||||
Split splitncnn_3 1 2 21 22 23
|
||||
Pooling maxpool2d_2 1 1 22 24 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
Padding Pad_16 1 1 24 25 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=32
|
||||
ConvolutionDepthWise padconvdw_0 1 1 23 26 0=32 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=800 7=32
|
||||
Convolution conv_15 1 1 26 27 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=2048
|
||||
BinaryOp add_3 2 1 25 27 28 0=0
|
||||
PReLU prelu_45 1 1 28 29 0=64
|
||||
Split splitncnn_4 1 2 29 30 31
|
||||
ConvolutionDepthWise convdw_80 1 1 31 32 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
|
||||
Convolution conv_16 1 1 32 33 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
|
||||
BinaryOp add_4 2 1 30 33 34 0=0
|
||||
PReLU prelu_46 1 1 34 35 0=64
|
||||
Split splitncnn_5 1 2 35 36 37
|
||||
ConvolutionDepthWise convdw_81 1 1 37 38 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
|
||||
Convolution conv_17 1 1 38 39 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
|
||||
BinaryOp add_5 2 1 36 39 40 0=0
|
||||
PReLU prelu_47 1 1 40 41 0=64
|
||||
Split splitncnn_6 1 2 41 42 43
|
||||
ConvolutionDepthWise convdw_82 1 1 43 44 0=64 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=1600 7=64
|
||||
Convolution conv_18 1 1 44 45 0=64 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4096
|
||||
BinaryOp add_6 2 1 42 45 46 0=0
|
||||
PReLU prelu_48 1 1 46 47 0=64
|
||||
Split splitncnn_7 1 2 47 48 49
|
||||
Pooling maxpool2d_3 1 1 48 50 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
Padding Pad_34 1 1 50 51 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=64
|
||||
ConvolutionDepthWise padconvdw_1 1 1 49 52 0=64 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=1600 7=64
|
||||
Convolution conv_19 1 1 52 53 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=8192
|
||||
BinaryOp add_7 2 1 51 53 54 0=0
|
||||
PReLU prelu_49 1 1 54 55 0=128
|
||||
Split splitncnn_8 1 2 55 56 57
|
||||
ConvolutionDepthWise convdw_84 1 1 57 58 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_20 1 1 58 59 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_8 2 1 56 59 60 0=0
|
||||
PReLU prelu_50 1 1 60 61 0=128
|
||||
Split splitncnn_9 1 2 61 62 63
|
||||
ConvolutionDepthWise convdw_85 1 1 63 64 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_21 1 1 64 65 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_9 2 1 62 65 66 0=0
|
||||
PReLU prelu_51 1 1 66 67 0=128
|
||||
Split splitncnn_10 1 2 67 68 69
|
||||
ConvolutionDepthWise convdw_86 1 1 69 70 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_22 1 1 70 71 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_10 2 1 68 71 72 0=0
|
||||
PReLU prelu_52 1 1 72 73 0=128
|
||||
Split splitncnn_11 1 3 73 74 75 76
|
||||
Pooling maxpool2d_4 1 1 75 77 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
Padding Pad_52 1 1 77 78 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=128
|
||||
ConvolutionDepthWise padconvdw_2 1 1 76 79 0=128 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=3200 7=128
|
||||
Convolution conv_23 1 1 79 80 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
|
||||
BinaryOp add_11 2 1 78 80 81 0=0
|
||||
PReLU prelu_53 1 1 81 82 0=256
|
||||
Split splitncnn_12 1 2 82 83 84
|
||||
ConvolutionDepthWise convdw_88 1 1 84 85 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_24 1 1 85 86 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_12 2 1 83 86 87 0=0
|
||||
PReLU prelu_54 1 1 87 88 0=256
|
||||
Split splitncnn_13 1 2 88 89 90
|
||||
ConvolutionDepthWise convdw_89 1 1 90 91 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_25 1 1 91 92 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_13 2 1 89 92 93 0=0
|
||||
PReLU prelu_55 1 1 93 94 0=256
|
||||
Split splitncnn_14 1 2 94 95 96
|
||||
ConvolutionDepthWise convdw_90 1 1 96 97 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_26 1 1 97 98 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_14 2 1 95 98 99 0=0
|
||||
PReLU prelu_56 1 1 99 100 0=256
|
||||
Split splitncnn_15 1 3 100 101 102 103
|
||||
Pooling maxpool2d_5 1 1 102 104 0=0 1=2 11=2 12=2 13=0 2=2 3=0 5=1
|
||||
ConvolutionDepthWise padconvdw_3 1 1 103 105 0=256 1=5 11=5 12=1 13=2 14=1 15=2 16=2 2=1 3=2 4=1 5=1 6=6400 7=256
|
||||
Convolution conv_27 1 1 105 106 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_15 2 1 104 106 107 0=0
|
||||
PReLU prelu_57 1 1 107 108 0=256
|
||||
Split splitncnn_16 1 2 108 109 110
|
||||
ConvolutionDepthWise convdw_92 1 1 110 111 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_28 1 1 111 112 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_16 2 1 109 112 113 0=0
|
||||
PReLU prelu_58 1 1 113 114 0=256
|
||||
Split splitncnn_17 1 2 114 115 116
|
||||
ConvolutionDepthWise convdw_93 1 1 116 117 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_29 1 1 117 118 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_17 2 1 115 118 119 0=0
|
||||
PReLU prelu_59 1 1 119 120 0=256
|
||||
Split splitncnn_18 1 2 120 121 122
|
||||
ConvolutionDepthWise convdw_94 1 1 122 123 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_30 1 1 123 124 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_18 2 1 121 124 125 0=0
|
||||
PReLU prelu_60 1 1 125 126 0=256
|
||||
Interp interpolate_0 1 1 126 127 0=2 3=12 4=12 6=0
|
||||
Convolution conv_31 1 1 127 128 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
PReLU prelu_61 1 1 128 129 0=256
|
||||
BinaryOp add_19 2 1 101 129 130 0=0
|
||||
Split splitncnn_19 1 2 130 131 132
|
||||
ConvolutionDepthWise convdw_95 1 1 132 133 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_32 1 1 133 134 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_20 2 1 131 134 135 0=0
|
||||
PReLU prelu_62 1 1 135 136 0=256
|
||||
Split splitncnn_20 1 2 136 137 138
|
||||
ConvolutionDepthWise convdw_96 1 1 138 139 0=256 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=6400 7=256
|
||||
Convolution conv_33 1 1 139 140 0=256 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=65536
|
||||
BinaryOp add_21 2 1 137 140 141 0=0
|
||||
PReLU prelu_63 1 1 141 142 0=256
|
||||
Split splitncnn_21 1 3 142 143 144 145
|
||||
Convolution conv_34 1 1 145 146 0=108 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=27648
|
||||
Permute permute_68 1 1 146 147 0=3
|
||||
Reshape reshape_72 1 1 147 148 0=18 1=864
|
||||
Convolution conv_35 1 1 144 149 0=6 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=1536
|
||||
Permute permute_69 1 1 149 150 0=3
|
||||
Reshape reshape_73 1 1 150 151 0=1 1=864
|
||||
Interp interpolate_1 1 1 143 152 0=2 3=24 4=24 6=0
|
||||
Convolution conv_36 1 1 152 153 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=32768
|
||||
PReLU prelu_64 1 1 153 154 0=128
|
||||
BinaryOp add_22 2 1 74 154 155 0=0
|
||||
Split splitncnn_22 1 2 155 156 157
|
||||
ConvolutionDepthWise convdw_97 1 1 157 158 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_37 1 1 158 159 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_23 2 1 156 159 160 0=0
|
||||
PReLU prelu_65 1 1 160 161 0=128
|
||||
Split splitncnn_23 1 2 161 162 163
|
||||
ConvolutionDepthWise convdw_98 1 1 163 164 0=128 1=5 11=5 12=1 13=1 14=2 2=1 3=1 4=2 5=1 6=3200 7=128
|
||||
Convolution conv_38 1 1 164 165 0=128 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=16384
|
||||
BinaryOp add_24 2 1 162 165 166 0=0
|
||||
PReLU prelu_66 1 1 166 167 0=128
|
||||
Split splitncnn_24 1 2 167 168 169
|
||||
Convolution conv_39 1 1 169 170 0=36 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=4608
|
||||
Permute permute_70 1 1 170 171 0=3
|
||||
Reshape reshape_74 1 1 171 172 0=18 1=1152
|
||||
Concat cat_0 2 1 172 148 out0 0=0
|
||||
Convolution conv_40 1 1 168 174 0=2 1=1 11=1 12=1 13=1 14=0 2=1 3=1 4=0 5=1 6=256
|
||||
Permute permute_71 1 1 174 175 0=3
|
||||
Reshape reshape_75 1 1 175 176 0=1 1=1152
|
||||
Concat cat_1 2 1 176 151 out1 0=0
|
||||
@@ -0,0 +1,65 @@
|
||||
"""Which color camera is which, and how their calibration maps onto fh-camd's images.
|
||||
|
||||
usage: python tools/check_color.py REC_DIR [--sets N]
|
||||
|
||||
A recording made with fh-camd --with-color holds color_video<N> frames with each set.
|
||||
This matches features between the two color images and scores every reading of the
|
||||
calibration: which video node is passthrough_left, and whether the calibration's
|
||||
cropRegion is subtracted from x ('subtract') or not ('none'). Only the right reading
|
||||
makes true matches' rays meet in front of both cameras. Then it checks the winner against
|
||||
the side tracking cameras, which tests the CAD-to-head chain shared with them.
|
||||
"""
|
||||
import argparse
|
||||
import itertools
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
|
||||
from tools.check_sides import load_cams, matches, score # noqa: E402
|
||||
from tools.show_set import index, read_set # noqa: E402
|
||||
from tracker import calib # noqa: E402
|
||||
|
||||
|
||||
def load_color(crop):
|
||||
root = os.environ.get('FRAME_JOB_DEVICE_ROOT', '') # frame-job's copy of the device files
|
||||
return calib.load_color(root + calib.ARCTURUS_EEPROM, root + calib.DEVICE_JSON, crop=crop)
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('rec')
|
||||
ap.add_argument('--sets', type=int, default=8)
|
||||
a = ap.parse_args()
|
||||
path = os.path.join(a.rec, 'sets.bin')
|
||||
offs = index(path)
|
||||
sets = [read_set(path, offs[n]) for n in np.linspace(0, len(offs) - 1, a.sets).astype(int)]
|
||||
nodes = sorted(k for k in sets[0] if k.startswith('color_video'))
|
||||
if len(nodes) != 2:
|
||||
sys.exit('need two color_video<N> cameras in the recording (fh-camd --with-color); found %s' % nodes)
|
||||
pairs = [matches(s[nodes[0]][0], s[nodes[1]][0]) for s in sets]
|
||||
print('%d sets, %d matches between %s and %s' % (len(sets), sum(len(p[0]) for p in pairs), *nodes))
|
||||
|
||||
best = None
|
||||
for crop, left in itertools.product(['subtract', 'none'], nodes):
|
||||
cams = load_color(crop)
|
||||
right = nodes[1] if left == nodes[0] else nodes[0]
|
||||
cam = {left: cams['passthrough_left'], right: cams['passthrough_right']}
|
||||
s = np.mean([score(cam[nodes[0]], cam[nodes[1]], ua, ub) for ua, ub in pairs])
|
||||
print(' %s = passthrough_left, crop %-8s: %3.0f%% of matches meet' % (left, crop, 100 * s))
|
||||
if best is None or s > best[0]:
|
||||
best = (s, crop, left, cam)
|
||||
s, crop, left, cam = best
|
||||
print('best: %s = passthrough_left, crop %s (%.0f%%)' % (left, crop, 100 * s))
|
||||
|
||||
mono = load_cams()
|
||||
for node in nodes:
|
||||
for side in ['slam_left', 'slam_right']:
|
||||
ms = [matches(st[node][0], st[side][0]) for st in sets if side in st]
|
||||
sc = np.mean([score(cam[node], mono[side], ua, ub) for ua, ub in ms]) if ms else 0
|
||||
print(' %s vs %-10s: %4d matches, %3.0f%% meet' % (node, side, sum(len(m[0]) for m in ms), 100 * sc))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,129 @@
|
||||
"""Check that the side cameras' images carry the right names (slam_left vs slam_right).
|
||||
|
||||
usage: python tools/check_sides.py REC_DIR [--sets N]
|
||||
python tools/check_sides.py --ring [--sets N] (live, from fh-camd's ring)
|
||||
|
||||
With --ring it exits 0 when the names are right, 3 when they're swapped (run fh-tracker
|
||||
with --swap-sides), and 2 when it can't tell (too little texture in view, or the headset
|
||||
isn't worn).
|
||||
|
||||
fh-camd tells the two side cameras' buffers apart by the order XRService allocated them,
|
||||
and after some XRService restarts that order puts each camera's images under the other's
|
||||
name. The tracker then sees every hand in one camera only, at the wrong depth. This
|
||||
matches features between the two images and measures how close each pair's rays pass
|
||||
with the factory calibration, once as named and once swapped: true matches meet in
|
||||
front of both cameras only under the right naming.
|
||||
"""
|
||||
import argparse
|
||||
import os
|
||||
import sys
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
|
||||
from tools.show_set import index, read_set # noqa: E402
|
||||
from tracker import calib # noqa: E402
|
||||
|
||||
|
||||
PIPES = {'msm_vfe3_video0': 'slam_left', 'msm_vfe4_video0': 'slam_right'} # as fh-tracker maps them
|
||||
|
||||
|
||||
def load_cams():
|
||||
root = os.environ.get('FRAME_JOB_DEVICE_ROOT', '') # frame-job's copy of /persist off the Frame
|
||||
return calib.load(root + calib.XRSERVICE_JSON, root + calib.DEVICE_JSON)
|
||||
|
||||
|
||||
def matches(a, b):
|
||||
"""Pixel pairs (N,2), (N,2) of ORB matches between two grey images."""
|
||||
clahe = cv2.createCLAHE(2.0, (8, 8))
|
||||
orb = cv2.ORB_create(3000)
|
||||
ka, da = orb.detectAndCompute(clahe.apply(a), None)
|
||||
kb, db = orb.detectAndCompute(clahe.apply(b), None)
|
||||
if da is None or db is None:
|
||||
return np.zeros((0, 2)), np.zeros((0, 2))
|
||||
pairs = cv2.BFMatcher(cv2.NORM_HAMMING).knnMatch(da, db, k=2)
|
||||
good = [p[0] for p in pairs if len(p) == 2 and p[0].distance < 0.75 * p[1].distance]
|
||||
return (np.array([ka[m.queryIdx].pt for m in good]).reshape(-1, 2),
|
||||
np.array([kb[m.trainIdx].pt for m in good]).reshape(-1, 2))
|
||||
|
||||
|
||||
def meet(cam_a, cam_b, ua, ub):
|
||||
"""Per match: closest distance between the two rays (m), and whether they meet in front of both."""
|
||||
ra, rb = cam_a.rays(ua), cam_b.rays(ub)
|
||||
w = cam_b.origin - cam_a.origin
|
||||
n = np.cross(ra, rb)
|
||||
nn = np.linalg.norm(n, axis=1)
|
||||
dist = np.abs(w @ n.T) / np.maximum(nn, 1e-12)
|
||||
# ray parameters at the closest points
|
||||
ta = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, rb.shape), rb), n) / np.maximum(nn ** 2, 1e-12)
|
||||
tb = np.einsum('ij,ij->i', np.cross(np.broadcast_to(w, ra.shape), ra), n) / np.maximum(nn ** 2, 1e-12)
|
||||
return dist, (ta > 0.05) & (tb > 0.05)
|
||||
|
||||
|
||||
def score(cam_a, cam_b, ua, ub):
|
||||
"""Share of matches whose rays meet within 1 cm, in front of both cameras."""
|
||||
if len(ua) == 0:
|
||||
return 0.0
|
||||
d, front = meet(cam_a, cam_b, ua, ub)
|
||||
return float(np.mean((d < 0.01) & front))
|
||||
|
||||
|
||||
def recorded_pairs(rec, count):
|
||||
"""(label, slam_left image, slam_right image) from sets spread across a recording."""
|
||||
path = os.path.join(rec, 'sets.bin')
|
||||
offs = index(path)
|
||||
for n in np.linspace(0, len(offs) - 1, count).astype(int):
|
||||
images = read_set(path, offs[n])
|
||||
if 'slam_left' in images and 'slam_right' in images:
|
||||
yield 'set %5d' % n, images['slam_left'][0], images['slam_right'][0]
|
||||
|
||||
|
||||
def live_pairs(count):
|
||||
"""(label, slam_left image, slam_right image) from fh-camd's ring, half a second apart."""
|
||||
import time
|
||||
from tracker.ring import Ring
|
||||
ring = Ring()
|
||||
if not ring.alive():
|
||||
sys.exit('fh-camd isn\'t running (no heartbeat)')
|
||||
cams = {}
|
||||
for c in ring.cams:
|
||||
name = PIPES.get(open('/sys/class/video4linux/video%d/name' % c.node).read().strip())
|
||||
if name and not c.name.endswith('-dark'):
|
||||
cams[name] = c
|
||||
for k in range(count):
|
||||
a, b = ring.read(cams['slam_left']), ring.read(cams['slam_right'])
|
||||
if a is not None and b is not None:
|
||||
yield 'frame %2d' % k, a.image, b.image
|
||||
time.sleep(0.5)
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('rec', nargs='?')
|
||||
ap.add_argument('--ring', action='store_true', help='check the live cameras instead of a recording')
|
||||
ap.add_argument('--sets', type=int, default=8, help='how many sets or live frames to check')
|
||||
a = ap.parse_args()
|
||||
if not a.ring and not a.rec:
|
||||
ap.error('give a recording or --ring')
|
||||
cams = load_cams()
|
||||
left, right = cams['slam_left'], cams['slam_right']
|
||||
named = swapped = 0.0
|
||||
n = total_matches = 0
|
||||
for label, img_l, img_r in (live_pairs(a.sets) if a.ring else recorded_pairs(a.rec, a.sets)):
|
||||
ua, ub = matches(img_l, img_r)
|
||||
s_named = score(left, right, ua, ub) # slam_left's image seen by the left camera
|
||||
s_swapped = score(right, left, ua, ub) # ... by the right camera
|
||||
named, swapped, n, total_matches = named + s_named, swapped + s_swapped, n + 1, total_matches + len(ua)
|
||||
print('%s: %4d matches, meeting as named %3.0f%%, swapped %3.0f%%' %
|
||||
(label, len(ua), 100 * s_named, 100 * s_swapped))
|
||||
if n == 0 or total_matches < 100 or abs(named - swapped) / n < 0.2:
|
||||
print('side cameras: can\'t tell (%d matches)' % total_matches)
|
||||
sys.exit(2)
|
||||
print('side cameras: %s (named %.2f, swapped %.2f)' %
|
||||
('as named' if named > swapped else 'SWAPPED', named / n, swapped / n))
|
||||
sys.exit(0 if named > swapped else 3)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,81 @@
|
||||
"""Convert the OpenCV Zoo ONNX ports of MediaPipe's hand models to ncnn.
|
||||
|
||||
The ONNX files are Apache-2.0 ports of MediaPipe's palm detector and hand
|
||||
landmark models (huggingface.co/opencv/palm_detection_mediapipe and
|
||||
huggingface.co/opencv/handpose_estimation_mediapipe). pnnx does the
|
||||
conversion; two fix-ups follow:
|
||||
|
||||
- The palm detector widens channels with ONNX Pad on the channel axis. pnnx
|
||||
emits an ncnn layer called "Pad", which ncnn doesn't have, so rewrite those
|
||||
as ncnn Padding with the channel-end amount (param 8 = behind).
|
||||
- Both models take NHWC input and start with a Permute to NCHW. Drop it, so we
|
||||
can hand ncnn planar CHW Mats straight from the preprocessing step.
|
||||
|
||||
usage: python convert_models.py (writes models/ncnn/{palm,hand}.ncnn.{param,bin})
|
||||
"""
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
ROOT = os.path.join(HERE, '..')
|
||||
PNNX = os.path.join(sys.prefix, 'lib', 'python%d.%d' % sys.version_info[:2], 'site-packages', 'pnnx', 'pnnx')
|
||||
MODELS = [('palm', 'palm_detection_mediapipe_2023feb', 192),
|
||||
('hand', 'handpose_estimation_mediapipe_2023feb', 224)]
|
||||
|
||||
|
||||
def patch(param_text, pnnx_param_text):
|
||||
lines = param_text.splitlines()
|
||||
assert lines[0] == '7767517'
|
||||
nlayers, nblobs = map(int, lines[1].split())
|
||||
body = lines[2:]
|
||||
|
||||
# Channel pads: amounts come from the pnnx graph, which keeps the pads tuple.
|
||||
pads = dict(re.findall(r'^Pad\s+(\S+)\s.*pads=\(0,0,0,0,0,(\d+),0,0\)', pnnx_param_text, re.M))
|
||||
for i, line in enumerate(body):
|
||||
f = line.split()
|
||||
if f[0] == 'Pad':
|
||||
amount = pads[f[1]]
|
||||
body[i] = 'Padding %s %s %s %s %s 0=0 1=0 2=0 3=0 4=0 5=0.000000e+00 7=0 8=%s' % (
|
||||
f[1], f[2], f[3], f[4], f[5], amount)
|
||||
|
||||
# Input permute: feed its consumers from in0 instead.
|
||||
perm = next(i for i, line in enumerate(body) if line.split()[0] == 'Permute')
|
||||
f = body[perm].split()
|
||||
assert f[4] == 'in0' and f[6] == '0=4', body[perm]
|
||||
blob = f[5]
|
||||
del body[perm]
|
||||
for i, line in enumerate(body):
|
||||
f = line.split()
|
||||
if f[0] == 'Input':
|
||||
continue
|
||||
nin, nout = int(f[2]), int(f[3])
|
||||
ins = ['in0' if b == blob else b for b in f[4:4 + nin]]
|
||||
body[i] = ' '.join(f[:4] + ins + f[4 + nin:])
|
||||
return '\n'.join(['7767517', '%d %d' % (nlayers - 1, nblobs - 1)] + body) + '\n'
|
||||
|
||||
|
||||
def main():
|
||||
out = os.path.join(ROOT, 'models', 'ncnn')
|
||||
os.makedirs(out, exist_ok=True)
|
||||
for short, name, size in MODELS:
|
||||
src = os.path.join(ROOT, 'models', 'onnx', name + '.onnx')
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
shutil.copy(src, tmp)
|
||||
subprocess.run([PNNX, name + '.onnx', 'inputshape=[1,%d,%d,3]' % (size, size), 'fp16=1'],
|
||||
cwd=tmp, check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
|
||||
with open(os.path.join(tmp, name + '.ncnn.param')) as f:
|
||||
param = f.read()
|
||||
with open(os.path.join(tmp, name + '.pnnx.param')) as f:
|
||||
pparam = f.read()
|
||||
with open(os.path.join(out, short + '.ncnn.param'), 'w') as f:
|
||||
f.write(patch(param, pparam))
|
||||
shutil.copy(os.path.join(tmp, name + '.ncnn.bin'), os.path.join(out, short + '.ncnn.bin'))
|
||||
print('wrote', short)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,256 @@
|
||||
"""How good the tracker's depth is, from a replay's depth dump, without ground truth.
|
||||
|
||||
usage: python3 tools/depth_report.py DEPTH [DEPTH...] [--still M/S]
|
||||
|
||||
DEPTH comes from `trackd/fh-replay DIR --depth DEPTH`. Every measure is split by how the
|
||||
hand was seen: by the two lower cameras ("lower pair"), by a lower and an upper camera on
|
||||
one side ("lower+upper"), or by one camera. Distances are from the head (between the eyes).
|
||||
|
||||
1. How the hands were seen: the share of hand updates in each way, by distance.
|
||||
2. Noise along the line of sight against across it. Each update's palm is compared with a
|
||||
straight line through the two updates before it (fh-replay's jitter measure), and the
|
||||
miss is split along the line from the hand's cameras to the palm and across it. Given
|
||||
as a robust sigma per axis, measured (as triangulated) and published (after the One Euro
|
||||
filter), on updates where the published palm moved slower than --still (default 0.15
|
||||
m/s), so the miss is mostly noise and not the hand speeding up. For two cameras,
|
||||
geometry predicts along/across = 2 Z / B: Z the distance, B the cameras' baseline
|
||||
across the line of sight.
|
||||
3. One-camera distance: on two-camera updates, each camera's one-view guess (distance from
|
||||
how big the palm looks, at the user's learned hand size) against the triangulated
|
||||
distance from that camera.
|
||||
4. A camera lost: from two-camera updates, what the tracker would have had if one of the
|
||||
two cameras dropped out there. It keeps the last distance and moves a share of the way
|
||||
to the one-view guess each update (0.1 now, kMonoDepthGain in trackd/tracker.cpp);
|
||||
also shown with other shares, 0 (keep the distance) and 1 (take each guess), and with
|
||||
the guess first scaled by how far off it was while both cameras saw the hand. Compared with the
|
||||
triangulated distance from that camera, 0.1-2 s after the loss.
|
||||
"""
|
||||
import argparse
|
||||
import collections
|
||||
|
||||
import numpy as np
|
||||
|
||||
BINS = [0.0, 0.35, 0.50, 0.65, 9.0]
|
||||
BIN_NAMES = ['<35 cm', '35-50', '50-65', '65+ cm']
|
||||
MODES = ['lower pair', 'lower+upper', 'one camera']
|
||||
HORIZONS = [0.1, 0.25, 0.5, 1.0, 2.0]
|
||||
# (share of the way toward the one-view guess per update, whether the guess is first scaled by
|
||||
# how far off it was while both cameras saw the hand)
|
||||
GAINS = [(0.0, False), (0.02, False), (0.05, False), (0.1, False), (1.0, False), (0.1, True), (1.0, True)]
|
||||
GAP = 0.1 # s: a longer gap between a hand's updates breaks its run
|
||||
|
||||
|
||||
def load(path):
|
||||
cams, rows = {}, []
|
||||
with open(path) as f:
|
||||
for line in f:
|
||||
w = line.split()
|
||||
if not w:
|
||||
continue
|
||||
if w[0] == '#':
|
||||
if w[1] == 'cam':
|
||||
cams[w[2]] = (np.array([float(x) for x in w[3:6]]), float(w[6]))
|
||||
continue
|
||||
r = {'t': float(w[0]), 'id': int(w[1]), 'side': w[2], 'n': int(w[3]),
|
||||
'cams': w[4], 'res': float(w[5]), 'scale': float(w[6]),
|
||||
'raw': np.array([float(x) for x in w[7:10]]), 'sm': np.array([float(x) for x in w[10:13]]),
|
||||
'views': {}}
|
||||
for k in range(13, len(w), 5):
|
||||
r['views'][w[k]] = np.array([float(x) for x in w[k + 2:k + 5]])
|
||||
rows.append(r)
|
||||
return cams, rows
|
||||
|
||||
|
||||
def mode(r):
|
||||
names = r['cams'].split('+')
|
||||
if r['n'] == 1:
|
||||
return 'one camera'
|
||||
if sorted(names) == ['slam_left', 'slam_right']:
|
||||
return 'lower pair'
|
||||
if len(names) == 2 and all(n.startswith(('slam_', 'upper_')) for n in names) and \
|
||||
names[0].split('_')[1] == names[1].split('_')[1]:
|
||||
return 'lower+upper'
|
||||
return 'other'
|
||||
|
||||
|
||||
def dist_bin(p):
|
||||
return min(np.searchsorted(BINS, np.linalg.norm(p), side='right') - 1, len(BIN_NAMES) - 1)
|
||||
|
||||
|
||||
def tracks(rows):
|
||||
"""A hand's updates, in order, split where they're more than GAP apart."""
|
||||
by_id = collections.defaultdict(list)
|
||||
for r in rows:
|
||||
by_id[r['id']].append(r)
|
||||
for rs in by_id.values():
|
||||
run = [rs[0]]
|
||||
for r in rs[1:]:
|
||||
if r['t'] - run[-1]['t'] >= GAP:
|
||||
yield run
|
||||
run = []
|
||||
run.append(r)
|
||||
yield run
|
||||
|
||||
|
||||
def two_camera_runs(rows):
|
||||
"""Stretches of a hand's updates all seen by the same two cameras."""
|
||||
for run in tracks(rows):
|
||||
seg = []
|
||||
for r in run:
|
||||
ok = r['n'] == 2 and mode(r) != 'other'
|
||||
if ok and seg and r['cams'] == seg[-1]['cams']:
|
||||
seg.append(r)
|
||||
continue
|
||||
if len(seg) > 2:
|
||||
yield seg
|
||||
seg = [r] if ok else []
|
||||
if len(seg) > 2:
|
||||
yield seg
|
||||
|
||||
|
||||
def pct(v, q):
|
||||
return np.percentile(v, q) if len(v) else float('nan')
|
||||
|
||||
|
||||
def seen_share(rows):
|
||||
print('\n1. How the hands were seen (share of hand updates)')
|
||||
count = collections.Counter((mode(r), dist_bin(r['raw'])) for r in rows)
|
||||
total = collections.Counter(dist_bin(r['raw']) for r in rows)
|
||||
print('%-12s' % '' + ''.join('%10s' % b for b in BIN_NAMES) + '%10s' % 'all')
|
||||
for m in MODES + ['other']:
|
||||
cells = [100 * count[m, b] / max(total[b], 1) for b in range(len(BIN_NAMES))]
|
||||
allp = 100 * sum(count[m, b] for b in range(len(BIN_NAMES))) / max(len(rows), 1)
|
||||
print('%-12s' % m + ''.join('%9.0f%%' % c for c in cells) + '%9.0f%%' % allp)
|
||||
print('%-12s' % 'updates' + ''.join('%10d' % total[b] for b in range(len(BIN_NAMES))) + '%10d' % len(rows))
|
||||
res = collections.defaultdict(list)
|
||||
for r in rows:
|
||||
if r['res'] >= 0:
|
||||
res[mode(r)].append(r['res'] * 1000)
|
||||
print('triangulation residual (rms ray miss, median): ' +
|
||||
', '.join('%s %.1f mm' % (m, np.median(v)) for m, v in res.items()))
|
||||
|
||||
|
||||
def noise(rows, cams, still):
|
||||
print('\n2. Noise along the line of sight vs across it (sigma per axis, mm; palm slower than %.2f m/s)' % still)
|
||||
acc = collections.defaultdict(lambda: collections.defaultdict(list))
|
||||
for run in tracks(rows):
|
||||
for a, b, c in zip(run, run[1:], run[2:]):
|
||||
if not (mode(a) == mode(b) == mode(c)) or a['cams'] != b['cams'] or b['cams'] != c['cams']:
|
||||
continue
|
||||
dt0, dt1 = b['t'] - a['t'], c['t'] - b['t']
|
||||
if dt0 < 1e-3 or np.linalg.norm(b['sm'] - a['sm']) / dt0 > still:
|
||||
continue
|
||||
names = b['cams'].split('+')
|
||||
origins = [cams[n][0] for n in names]
|
||||
o = np.mean(origins, axis=0)
|
||||
u = b['raw'] - o
|
||||
z = np.linalg.norm(u)
|
||||
u /= z
|
||||
key = (mode(b), dist_bin(b['raw']))
|
||||
for kind in ('raw', 'sm'):
|
||||
miss = c[kind] - b[kind] - (b[kind] - a[kind]) * (dt1 / dt0)
|
||||
along = miss @ u
|
||||
acc[key][kind + '_along'].append(abs(along))
|
||||
acc[key][kind + '_across'].append(np.linalg.norm(miss - along * u))
|
||||
if len(origins) == 2:
|
||||
base = origins[0] - origins[1]
|
||||
acc[key]['pred'].append(2 * z / np.linalg.norm(base - (base @ u) * u))
|
||||
# |along| is half-normal: sigma = median / 0.674; |across| is Rayleigh (2 axes): sigma = median / 1.177
|
||||
print('%-12s %-7s %6s | %-24s | %-24s | %s' % ('', '', 'n', 'measured along/across', 'published along/across',
|
||||
'ratio measured (geometry)'))
|
||||
for m in MODES:
|
||||
for bi, bn in enumerate(BIN_NAMES):
|
||||
d = acc.get((m, bi))
|
||||
if not d or len(d['raw_along']) < 20:
|
||||
continue
|
||||
s = {k: np.median(v) / (0.674 if k.endswith('along') else 1.177) * 1000
|
||||
for k, v in d.items() if k != 'pred'}
|
||||
pred = '(%.1f)' % np.median(d['pred']) if d['pred'] else ''
|
||||
print('%-12s %-7s %6d | %7.1f / %-5.1f x%-6.1f | %7.1f / %-5.1f x%-6.1f | x%.1f %s' % (
|
||||
m, bn, len(d['raw_along']), s['raw_along'], s['raw_across'], s['raw_along'] / s['raw_across'],
|
||||
s['sm_along'], s['sm_across'], s['sm_along'] / s['sm_across'],
|
||||
s['raw_along'] / s['raw_across'], pred))
|
||||
|
||||
|
||||
def one_camera(rows, cams):
|
||||
print('\n3. One-camera distance vs triangulated, on two-camera updates (error of the one-view guess)')
|
||||
acc = collections.defaultdict(list)
|
||||
for r in rows:
|
||||
if r['n'] != 2 or mode(r) == 'other':
|
||||
continue
|
||||
for name, p in r['views'].items():
|
||||
if np.isnan(p).any():
|
||||
continue
|
||||
o = cams[name][0]
|
||||
truth = np.linalg.norm(r['raw'] - o)
|
||||
acc[name.split('_')[0], dist_bin(r['raw'])].append((np.linalg.norm(p - o) - truth, truth))
|
||||
print('%-8s %-7s %6s %12s %12s %14s %12s' % ('camera', '', 'n', 'median |err|', '90% |err|', 'median |err| %',
|
||||
'bias'))
|
||||
for cam in ('slam', 'upper'):
|
||||
for bi, bn in enumerate(BIN_NAMES):
|
||||
v = acc.get((cam, bi))
|
||||
if not v or len(v) < 20:
|
||||
continue
|
||||
e = np.array([x[0] for x in v])
|
||||
rel = e / np.array([x[1] for x in v])
|
||||
print('%-8s %-7s %6d %9.0f mm %9.0f mm %13.0f%% %+11.0f%%' % (
|
||||
'lower' if cam == 'slam' else 'upper', bn, len(v), 1000 * np.median(abs(e)), 1000 * pct(abs(e), 90),
|
||||
100 * np.median(abs(rel)), 100 * np.median(rel)))
|
||||
|
||||
|
||||
def lost_camera(rows, cams):
|
||||
print('\n4. A camera lost: distance error after the loss (median |err| mm / 90% mm), by how the tracker '
|
||||
'moves toward the one-view guess each update ("scaled": the guess times how far off it was, '
|
||||
'triangulated / guess, median over the last 30 two-camera updates)')
|
||||
errs = collections.defaultdict(list)
|
||||
for run in two_camera_runs(rows):
|
||||
for s in range(1, len(run) - 1, 3):
|
||||
for name in run[s]['views']:
|
||||
o = cams[name][0]
|
||||
guess = lambda r: np.linalg.norm(r['views'][name] - o)
|
||||
ratios = [np.linalg.norm(r['raw'] - o) / guess(r) for r in run[max(0, s - 30):s]
|
||||
if not np.isnan(r['views'][name]).any()]
|
||||
ratio = np.median(ratios) if ratios else 1.0
|
||||
for g, scaled in GAINS:
|
||||
d = np.linalg.norm(run[s - 1]['raw'] - o)
|
||||
h = 0
|
||||
for r in run[s:]:
|
||||
if np.isnan(r['views'][name]).any():
|
||||
break
|
||||
d += g * (guess(r) * (ratio if scaled else 1.0) - d)
|
||||
elapsed = r['t'] - run[s - 1]['t']
|
||||
while h < len(HORIZONS) and elapsed >= HORIZONS[h]:
|
||||
errs[g, scaled, HORIZONS[h], name.split('_')[0]].append(abs(d - np.linalg.norm(r['raw'] - o)))
|
||||
h += 1
|
||||
print('%-8s %-18s' % ('camera', 'toward guess') + ''.join('%14s' % ('%.2g s' % t) for t in HORIZONS))
|
||||
for cam in ('slam', 'upper'):
|
||||
for g, scaled in GAINS:
|
||||
label = {0.0: '0 (keep)', 0.1: '0.1 (now)', 1.0: '1 (guess)'}.get(g, '%g' % g)
|
||||
if scaled:
|
||||
label = '%g scaled' % g
|
||||
cells = []
|
||||
for t in HORIZONS:
|
||||
v = errs.get((g, scaled, t, cam), [])
|
||||
cells.append('%5.0f / %-4.0f' % (1000 * np.median(v), 1000 * pct(v, 90)) if len(v) >= 20 else '%14s' % '-')
|
||||
print('%-8s %-18s' % ('lower' if cam == 'slam' else 'upper', label) + ''.join('%14s' % c for c in cells))
|
||||
n = sum(len(errs.get((0.1, False, HORIZONS[0], c), [])) for c in ('slam', 'upper'))
|
||||
print('(%d simulated losses)' % n)
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('depth', nargs='+')
|
||||
ap.add_argument('--still', type=float, default=0.15, help='m/s: palm speed limit for the noise measure')
|
||||
a = ap.parse_args()
|
||||
for path in a.depth:
|
||||
cams, rows = load(path)
|
||||
print('== %s: %d hand updates, %d hands' % (path, len(rows), len({r['id'] for r in rows})))
|
||||
seen_share(rows)
|
||||
noise(rows, cams, a.still)
|
||||
one_camera(rows, cams)
|
||||
lost_camera(rows, cams)
|
||||
print()
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,79 @@
|
||||
"""Calibration crops for quantizing the models to int8 (ncnn2table).
|
||||
|
||||
Takes the frames of an fh-camprobe capture, cuts the crops the tracker would feed
|
||||
the models (CLAHE-equalized, as tracker/models.py prepares them), and writes them
|
||||
as PNGs plus a list per model:
|
||||
palm: every search tile of every frame (a sample of them)
|
||||
hand: the hands found in them, each also shifted, scaled and turned a little
|
||||
|
||||
usage: python tools/make_int8_calib.py CAPTURE_DIR OUT_DIR
|
||||
"""
|
||||
import glob
|
||||
import os
|
||||
import random
|
||||
import sys
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..', 'tracker'))
|
||||
import calib # noqa: E402
|
||||
import hands # noqa: E402
|
||||
import models # noqa: E402
|
||||
|
||||
NODES = {'video9': 'slam_left', 'video13': 'slam_right', 'video6': 'upper_left', 'video7': 'upper_right'}
|
||||
|
||||
|
||||
def save(path, patch):
|
||||
cv2.imwrite(path, cv2.cvtColor(patch, cv2.COLOR_GRAY2BGR))
|
||||
|
||||
|
||||
def main():
|
||||
cap, out = sys.argv[1], sys.argv[2]
|
||||
random.seed(1)
|
||||
cams = calib.load()
|
||||
eng = models.Engine()
|
||||
palm, hm = models.PalmDetector(), models.HandLandmarker()
|
||||
tracker = hands.Tracker(cams, eng)
|
||||
for d in ('palm', 'hand'):
|
||||
os.makedirs(os.path.join(out, d), exist_ok=True)
|
||||
palm_files, hand_files = [], []
|
||||
for node, name in NODES.items():
|
||||
tiles = [t for t in tracker.tiles if t.cam.name == name]
|
||||
for f in sorted(glob.glob(os.path.join(cap, '*_%s.pgm' % node))):
|
||||
g = cv2.imread(f, cv2.IMREAD_GRAYSCALE)
|
||||
base = os.path.basename(f)[:-4]
|
||||
prep = [palm.prepare(g, t.center, t.size, t.rotation) for t in tiles]
|
||||
outs = eng.run([('palm', p) for p, _ in prep])
|
||||
rois = []
|
||||
for k, ((p, ctx), o) in enumerate(zip(prep, outs)):
|
||||
if random.random() < 0.3:
|
||||
path = os.path.join(out, 'palm', '%s_t%02d.png' % (base, k))
|
||||
save(path, p)
|
||||
palm_files.append(path)
|
||||
for d in palm.decode(o, ctx):
|
||||
r = d.roi()
|
||||
if all(np.linalg.norm(r[0] - q[0]) > 0.5 * r[1] for q in rois):
|
||||
rois.append(r)
|
||||
for j, r in enumerate(rois):
|
||||
p, ctx = hm.prepare(g, r)
|
||||
if hm.decode(eng.run([('hand', p)])[0], ctx).presence < 0.5:
|
||||
continue
|
||||
for v in range(5):
|
||||
c, s, rot = r
|
||||
if v:
|
||||
c = np.asarray(c) + np.random.uniform(-0.08, 0.08, 2) * s
|
||||
s = s * np.random.uniform(0.87, 1.15)
|
||||
rot = rot + np.radians(np.random.uniform(-20, 20))
|
||||
p, _ = hm.prepare(g, (c, s, rot))
|
||||
path = os.path.join(out, 'hand', '%s_h%d_%d.png' % (base, j, v))
|
||||
save(path, p)
|
||||
hand_files.append(path)
|
||||
for name, files in (('palm', palm_files), ('hand', hand_files)):
|
||||
with open(os.path.join(out, name + '.txt'), 'w') as fh:
|
||||
fh.write('\n'.join(files) + '\n')
|
||||
print(name, len(files), 'crops')
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,90 @@
|
||||
"""Compare trackd's C++ model code with tracker/models.py on recorded frames.
|
||||
|
||||
For each frame, finds a crop with a palm (Python side), runs trackd/nettest on the same
|
||||
crop, and reports how far apart the palms, ROIs and landmarks are.
|
||||
|
||||
usage: python tools/nettest_compare.py CAPTURE_DIR [--int8]
|
||||
"""
|
||||
import glob
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
sys.path.insert(0, os.path.join(HERE, '..', 'tracker'))
|
||||
import calib # noqa: E402
|
||||
import hands # noqa: E402
|
||||
import models # noqa: E402
|
||||
|
||||
NODES = {'video9': 'slam_left', 'video13': 'slam_right', 'video6': 'upper_left', 'video7': 'upper_right'}
|
||||
NETTEST = os.path.join(HERE, '..', 'trackd', 'nettest')
|
||||
MODELS = os.path.join(HERE, '..', 'models', 'ncnn')
|
||||
|
||||
|
||||
def cpp(frame, tile, int8):
|
||||
out = subprocess.run([NETTEST, MODELS, frame, str(tile.center[0]), str(tile.center[1]), str(tile.size),
|
||||
str(tile.rotation)] + (['--int8'] if int8 else []), capture_output=True, text=True).stdout
|
||||
palms, hands_, ms = [], [], {}
|
||||
for line in out.splitlines():
|
||||
f = line.split()
|
||||
if f[0] == 'palm':
|
||||
palms.append((float(f[1]), np.array([float(f[6]), float(f[7])]), float(f[8]), float(f[9])))
|
||||
elif f[0] == 'pts':
|
||||
hands_.append(np.array(f[1:], float).reshape(21, 2))
|
||||
elif f[0] == 'hand_ms':
|
||||
ms.setdefault('hand', []).append(float(f[1]))
|
||||
hands_presence = float(f[3])
|
||||
ms.setdefault('presence', []).append(hands_presence)
|
||||
elif f[0] == 'palm_ms':
|
||||
ms.setdefault('palm', []).append(float(f[1]))
|
||||
return palms, hands_, ms
|
||||
|
||||
|
||||
def main():
|
||||
cap = sys.argv[1]
|
||||
int8 = '--int8' in sys.argv
|
||||
cams = calib.load()
|
||||
eng = models.Engine()
|
||||
palm, hm = models.PalmDetector(), models.HandLandmarker()
|
||||
tracker = hands.Tracker(cams, eng)
|
||||
d_roi, d_pts, n_py, n_cpp, pres = [], [], 0, 0, []
|
||||
times = {'palm': [], 'hand': []}
|
||||
for node, name in NODES.items():
|
||||
tiles = [t for t in tracker.tiles if t.cam.name == name]
|
||||
for f in sorted(glob.glob(os.path.join(cap, '*_%s.pgm' % node)))[::4]:
|
||||
g = cv2.imread(f, cv2.IMREAD_GRAYSCALE)
|
||||
for t in tiles:
|
||||
p, ctx = palm.prepare(g, t.center, t.size, t.rotation)
|
||||
dets = palm.decode(eng.run([('palm', p)])[0], ctx)
|
||||
if not dets:
|
||||
continue
|
||||
cp, ch, ms = cpp(f, t, int8)
|
||||
for k in times:
|
||||
times[k] += ms.get(k, [])
|
||||
n_py += len(dets)
|
||||
n_cpp += len(cp)
|
||||
for d in dets:
|
||||
r = d.roi()
|
||||
if not cp:
|
||||
continue
|
||||
j = int(np.argmin([np.linalg.norm(c[1] - r[0]) for c in cp]))
|
||||
d_roi.append(np.linalg.norm(cp[j][1] - r[0]) / r[1])
|
||||
pp, pctx = hm.prepare(g, r)
|
||||
lm = hm.decode(eng.run([('hand', pp)])[0], pctx)
|
||||
if lm.presence >= 0.5 and j < len(ch):
|
||||
d_pts.append(np.median(np.linalg.norm(ch[j] - lm.pts, axis=1)) / r[1])
|
||||
pres.append(abs(ms['presence'][j] - lm.presence))
|
||||
break # one palm tile per frame is enough
|
||||
print('palms: python %d, c++ %d' % (n_py, n_cpp))
|
||||
print('roi centre offset: median %.3f of the roi size (max %.3f)' % (np.median(d_roi), np.max(d_roi)))
|
||||
if d_pts:
|
||||
print('landmarks: median %.4f of the roi size (max %.4f), presence diff median %.3f' % (
|
||||
np.median(d_pts), np.max(d_pts), np.median(pres)))
|
||||
print('c++ time: palm %.1f ms, hand %.1f ms (median, one thread)' % (np.median(times['palm']), np.median(times['hand'])))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,129 @@
|
||||
"""Draw frame sets from a recording (fh-tracker --record) with what the tracker saw.
|
||||
|
||||
usage: python tools/show_set.py REC_DIR SET [SET...] [--timeline TL] [--out DIR]
|
||||
|
||||
SET is a set index (fh-replay's timeline gives them). With --timeline (fh-replay
|
||||
--timeline), each camera shows the tracker's views at that set: the crop for the next
|
||||
frame, labelled with the hand and presence. Recordings made with fh-camd --with-dark get
|
||||
a second row: each camera's latest dark frame (<name>_dk), stretched to be visible and
|
||||
labelled with its mean brightness. Recordings made with fh-camd --with-color get a row of
|
||||
the color cameras (color_video<N>). Writes OUT/set_<n>.jpg (default /tmp).
|
||||
"""
|
||||
import argparse
|
||||
import os
|
||||
import struct
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
HDR = struct.Struct('<8sII')
|
||||
CAM = struct.Struct('<16sIIQQ')
|
||||
ORDER = ['slam_left', 'slam_right', 'upper_left', 'upper_right']
|
||||
|
||||
|
||||
def index(path):
|
||||
"""Byte offset of every set in sets.bin."""
|
||||
offs, size = [], os.path.getsize(path)
|
||||
with open(path, 'rb') as f:
|
||||
off = 0
|
||||
while off + HDR.size <= size:
|
||||
f.seek(off)
|
||||
magic, n, nbytes = HDR.unpack(f.read(HDR.size))
|
||||
if magic[:7] != b'FHSET01' or off + nbytes > size:
|
||||
break
|
||||
offs.append(off)
|
||||
off += nbytes
|
||||
return offs
|
||||
|
||||
|
||||
def read_set(path, off):
|
||||
with open(path, 'rb') as f:
|
||||
f.seek(off)
|
||||
_, n, _ = HDR.unpack(f.read(HDR.size))
|
||||
cams = [CAM.unpack(f.read(CAM.size)) for _ in range(n)]
|
||||
out = {}
|
||||
for name, w, h, cap, dq in cams:
|
||||
px = np.frombuffer(f.read(w * h), np.uint8).reshape(h, w)
|
||||
out[name.rstrip(b'\0').decode()] = (px, cap)
|
||||
return out
|
||||
|
||||
|
||||
def views_at(timeline, n):
|
||||
out = []
|
||||
for line in open(timeline):
|
||||
f = line.split()
|
||||
if len(f) > 1 and f[1] == 'view' and int(f[-1]) == n:
|
||||
out.append({'hand': int(f[2]), 'cam': f[3], 'presence': float(f[5]),
|
||||
'c': (float(f[7]), float(f[8])), 'size': float(f[9]), 'rot': float(f[10])})
|
||||
return out
|
||||
|
||||
|
||||
def dark_tile(frame, shape, name):
|
||||
"""A dark frame, stretched from its 1st to 99.5th percentile; black if there's none."""
|
||||
h, w = shape
|
||||
if frame is None:
|
||||
return np.zeros((h, w, 3), np.uint8)
|
||||
px = frame[0]
|
||||
lo, hi = np.percentile(px, (1, 99.5))
|
||||
gain = 255 / max(hi - lo, 1)
|
||||
img = np.clip((px.astype(np.float32) - lo) * gain, 0, 255).astype(np.uint8)
|
||||
img = cv2.cvtColor(cv2.resize(img, (w, h)), cv2.COLOR_GRAY2BGR)
|
||||
cv2.putText(img, '%s_dk mean %.1f, x%.0f' % (name, px.mean(), gain), (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0,
|
||||
(255, 255, 0), 2)
|
||||
return img
|
||||
|
||||
|
||||
def view_tile(px, name, views):
|
||||
"""A frame, CLAHE'd, with the tracker's views on it, 512 px high."""
|
||||
img = cv2.cvtColor(cv2.createCLAHE(2.0, (8, 8)).apply(px), cv2.COLOR_GRAY2BGR)
|
||||
for v in views:
|
||||
if v['cam'] != name:
|
||||
continue
|
||||
c, s, r = v['c'], v['size'], v['rot']
|
||||
box = cv2.boxPoints(((c[0], c[1]), (s, s), np.degrees(r)))
|
||||
col = (0, 255, 0) if v['presence'] >= 0.5 else (0, 0, 255)
|
||||
cv2.polylines(img, [box.astype(np.int32)], True, col, 2)
|
||||
cv2.putText(img, 'h%d %.2f' % (v['hand'], v['presence']), (int(c[0] - s / 2), int(c[1] - s / 2) - 6),
|
||||
cv2.FONT_HERSHEY_SIMPLEX, 0.8, col, 2)
|
||||
cv2.putText(img, name, (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0, (255, 255, 0), 2)
|
||||
scale = 512 / img.shape[0]
|
||||
return cv2.resize(img, (int(img.shape[1] * scale), 512))
|
||||
|
||||
|
||||
def draw(images, views):
|
||||
"""Rows: the mono cameras; their dark frames, if recorded; the color cameras, if recorded."""
|
||||
tiles, dark = [], []
|
||||
for name in ORDER:
|
||||
if name not in images:
|
||||
continue
|
||||
tiles.append(view_tile(images[name][0], name, views))
|
||||
dark.append(dark_tile(images.get(name + '_dk'), tiles[-1].shape[:2], name))
|
||||
rows = [np.hstack(tiles)]
|
||||
if any(k.endswith('_dk') for k in images):
|
||||
rows.append(np.hstack(dark))
|
||||
color = sorted(k for k in images if k.startswith('color_'))
|
||||
if color:
|
||||
rows.append(np.hstack([view_tile(images[k][0], k, views) for k in color]))
|
||||
width = max(r.shape[1] for r in rows)
|
||||
return np.vstack([np.pad(r, ((0, 0), (0, width - r.shape[1]), (0, 0))) for r in rows])
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('rec')
|
||||
ap.add_argument('sets', type=int, nargs='+')
|
||||
ap.add_argument('--timeline')
|
||||
ap.add_argument('--out', default='/tmp')
|
||||
a = ap.parse_args()
|
||||
path = os.path.join(a.rec, 'sets.bin')
|
||||
offs = index(path)
|
||||
for n in a.sets:
|
||||
images = read_set(path, offs[n])
|
||||
views = views_at(a.timeline, n) if a.timeline else []
|
||||
out = os.path.join(a.out, 'set_%05d.jpg' % n)
|
||||
cv2.imwrite(out, draw(images, views), [cv2.IMWRITE_JPEG_QUALITY, 85])
|
||||
print(out)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,85 @@
|
||||
"""Watch the pinch gestures fh-tracker publishes, live: begins, ends, and drags.
|
||||
|
||||
usage: python3 tools/watch_gestures.py [--every S] [--distance]
|
||||
|
||||
Prints a line when a pinch begins or ends on either hand. It goes by the counters, so a
|
||||
quick tap between two reads still shows. While a pinch is held, every --every seconds
|
||||
(default 0.1) it prints how far the pinch point has moved since it began, in the head
|
||||
frame (turning your head moves it too; a real consumer turns both points into the room
|
||||
first, see include/fh_gestures.h). --distance also prints each hand's thumb-to-index
|
||||
distance, to see how close a pinch comes to the thresholds.
|
||||
"""
|
||||
import argparse
|
||||
import mmap
|
||||
import os
|
||||
import struct
|
||||
import time
|
||||
|
||||
HDR = struct.Struct('<8sIIQQQff16x') # 64 bytes
|
||||
PINCH = struct.Struct('<IIIIQQff3f3f') # 64 bytes
|
||||
TRACKED, DOWN, LOST = 1, 2, 4
|
||||
SIDES = ('left ', 'right')
|
||||
|
||||
|
||||
def path():
|
||||
run = os.environ.get('XDG_RUNTIME_DIR', '/run/user/%d' % os.getuid())
|
||||
return os.path.join(run, 'frame-hands', 'gestures')
|
||||
|
||||
|
||||
def read(m):
|
||||
"""(header, [left, right]) under the sequence lock, or None if it's being written."""
|
||||
for _ in range(10):
|
||||
s1 = struct.unpack_from('<Q', m, 16)[0]
|
||||
if s1 % 2 == 0:
|
||||
h = HDR.unpack_from(m, 0)
|
||||
p = [PINCH.unpack_from(m, HDR.size + k * PINCH.size) for k in range(2)]
|
||||
if struct.unpack_from('<Q', m, 16)[0] == s1:
|
||||
return h, p
|
||||
time.sleep(0.0005)
|
||||
return None
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument('--every', type=float, default=0.1, help='seconds between drag lines while pinching')
|
||||
ap.add_argument('--distance', action='store_true', help="print each hand's thumb-to-index distance")
|
||||
a = ap.parse_args()
|
||||
with open(path(), 'rb') as f:
|
||||
m = mmap.mmap(f.fileno(), 0, prot=mmap.PROT_READ)
|
||||
first = read(m)
|
||||
if first is None or first[0][0] != b'FHGEST01':
|
||||
raise SystemExit('%s is not an fh-tracker gestures file' % path())
|
||||
h, p = first
|
||||
print('thresholds: pinch begins under %.3f m, ends over %.3f m' % (h[6], h[7]))
|
||||
seen = [(q[2], q[3]) for q in p] # begins, ends
|
||||
last_drag = last_dist = 0.0
|
||||
while True:
|
||||
got = read(m)
|
||||
if got:
|
||||
h, p = got
|
||||
now = time.monotonic()
|
||||
for k, q in enumerate(p):
|
||||
flags, hand, begins, ends, begin_ns, end_ns, dist, strength = q[:8]
|
||||
point, begin_point = q[8:11], q[11:14]
|
||||
if begins != seen[k][0]:
|
||||
print('%s pinch BEGIN (#%d, hand %d) at %+.3f %+.3f %+.3f d %.3f' %
|
||||
(SIDES[k], begins, hand, *begin_point, dist), flush=True)
|
||||
if ends != seen[k][1]:
|
||||
held = (end_ns - begin_ns) / 1e9 if end_ns >= begin_ns else 0
|
||||
print('%s pinch %s after %.2f s' % (SIDES[k], 'LOST' if flags & LOST else 'END', held), flush=True)
|
||||
seen[k] = (begins, ends)
|
||||
if flags & DOWN and now - last_drag >= a.every:
|
||||
d = [point[i] - begin_point[i] for i in range(3)]
|
||||
print('%s drag %+6.1f %+6.1f %+6.1f mm (%.0f mm)' %
|
||||
(SIDES[k], *(1000 * x for x in d), 1000 * sum(x * x for x in d) ** 0.5), flush=True)
|
||||
if any(q[0] & DOWN for q in p) and now - last_drag >= a.every:
|
||||
last_drag = now
|
||||
if a.distance and now - last_dist >= 0.2:
|
||||
last_dist = now
|
||||
print(' ' + ' '.join('%s %s' % (SIDES[k].strip(), 'd %.3f s %.2f' % (q[6], q[7]) if q[0] & TRACKED
|
||||
else '-') for k, q in enumerate(p)), flush=True)
|
||||
time.sleep(0.005)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,28 @@
|
||||
# fh-tracker, the C++ tracker. Needs ncnn built into ../vendor/ncnn/build/install
|
||||
# (see README.md) and jsoncpp.
|
||||
NCNN ?= ../vendor/ncnn/build/install
|
||||
CXXFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter
|
||||
CXXFLAGS += -std=c++17 -fopenmp -I$(NCNN)/include/ncnn
|
||||
LDLIBS = $(NCNN)/lib/libncnn.a -ljsoncpp -fopenmp -lpthread
|
||||
|
||||
SRC = main.cpp calib.cpp nets.cpp tracker.cpp io.cpp record.cpp pinch.cpp
|
||||
HDR = calib.h geom.h nets.h tracker.h io.h record.h pinch.h ../camd/fhring.h ../include/fh_hands.h ../include/fh_gestures.h
|
||||
|
||||
all: fh-tracker fh-replay fh-ringplay
|
||||
|
||||
fh-tracker: $(SRC) $(HDR)
|
||||
$(CXX) $(CXXFLAGS) -o $@ $(SRC) $(LDLIBS)
|
||||
|
||||
fh-replay: replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp pinch.cpp io.cpp $(HDR)
|
||||
$(CXX) $(CXXFLAGS) -o $@ replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp pinch.cpp io.cpp $(LDLIBS)
|
||||
|
||||
fh-ringplay: ringplay.cpp record.h ../camd/fhring.h
|
||||
$(CXX) $(CXXFLAGS) -o $@ ringplay.cpp
|
||||
|
||||
clean:
|
||||
rm -f fh-tracker fh-replay fh-ringplay nettest
|
||||
|
||||
.PHONY: all clean
|
||||
|
||||
nettest: nettest.cpp nets.cpp nets.h geom.h
|
||||
$(CXX) $(CXXFLAGS) -o $@ nettest.cpp nets.cpp $(LDLIBS)
|
||||
@@ -0,0 +1,86 @@
|
||||
# trackd
|
||||
|
||||
`fh-tracker` is the C++ version of `tracker/live.py`. It uses the same scheduling and the same models. The models run on a few threads, with no Python in the loop. It reads `fh-camd`'s ring and publishes hands to `$XDG_RUNTIME_DIR/frame-hands/hands` (`include/fh_hands.h`), where Frametop's ft-screens picks them up for the hand cutouts.
|
||||
|
||||
```
|
||||
sudo camd/fh-camd &
|
||||
trackd/fh-tracker # status every 5 s; Ctrl+C to stop
|
||||
trackd/fh-tracker --int8 # the 8-bit models (models/ncnn/*-int8.ncnn.*)
|
||||
```
|
||||
|
||||
Options:
|
||||
|
||||
- `--threads N`: model threads, pinned to the `--cpus` list. Default 3.
|
||||
- `--cpus LIST`: CPUs for the model threads and the main loop. Default `5,6,7`. With the headset on, that ran a step in 8.4 ms against 13.2 ms on `2,3,4`, where XRService's head tracking also runs, and SteamVR's frame timing didn't change (`probes/core_ab.py`, 2026-09-29).
|
||||
- `--contrast MODE` or `PALM/HAND`: how crops are equalized before the models see them: `clahe[:CLIP]`, `none`, or `stretch` (1st-99th percentile). Default `clahe:2/none`. In the dim recording, CLAHE let the palm search find about 10% more hands, but it made the landmarks jitter more (published median 6.9 mm, against 6.0 mm with plain landmark crops).
|
||||
- `--swap-sides`: swap the two side cameras (`slam_left`, `slam_right`). fh-camd tells their buffers apart by XRService's allocation order, and after some XRService restarts that order is reversed. Then every hand is seen by one camera only, at the wrong depth, and the hand holes land beside the hands. With the headset on and looking at a room with some texture, `tools/check_sides.py --ring` tells whether the names are right (exit 0), swapped (exit 3), or it can't tell (exit 2).
|
||||
- `--seconds N`: stop after N seconds.
|
||||
- `--status S`: how often to print status, in seconds.
|
||||
- `--models DIR`: where the models are.
|
||||
- `--nice N`: niceness. Default 5, so the VR stack wins contested CPUs.
|
||||
- `--no-publish`: don't write the hands file.
|
||||
- `--record DIR`, `--record-for S`: save every frame set for S seconds (default 120) to `DIR/sets.bin`. That's about 80 MB/s. Sending the tracker SIGUSR1 (`pkill -USR1 -x fh-tracker`) starts a recording in `captures/rec-<time>` without a restart.
|
||||
- `--record-only`: record without tracking or publishing, so it can run beside the live tracker. Give it `--record DIR`; SIGUSR1 would reach both trackers. With `fh-camd --with-dark`, recordings also hold each camera's newest dark frame as `<name>_dk`, which doubles the rate.
|
||||
- `--keep-presence P`: the landmark presence a tracked view needs to stay tracked. New views always need 0.5. Default 0.5. Lowering it to 0.2 barely helped in the bright recording, because lost hands drop to near-zero presence.
|
||||
- `--ring PATH`: read frames from another ring, such as `fh-ringplay`'s.
|
||||
- Recording from a ring with color cameras (`fh-camd --with-color`) also saves each color camera's newest frame with every set, as `color_video<N>`. That adds about 70 MB/s. Run the recorder at normal I/O priority (not under `frame-job`, whose `ionice -c 3` stalled a 165 MB/s recording).
|
||||
|
||||
The status line also says how often a hand was on each side (by where the wrist is), and why views and hands came and went: views lost (the landmark model stopped seeing the hand), handoff misses (a crop projected from the hand's 3D position found nothing), duplicates, splits (two views disagreed in 3D), and hands created, merged and forgotten.
|
||||
|
||||
## Beyond the Python tracker
|
||||
|
||||
fh-tracker started as a port of `tracker/hands.py`. Replaying recordings (below) showed where it went wrong, and it now differs in these ways:
|
||||
|
||||
- **Pairing views across cameras.** The side cameras sit side by side, so two hands next to each other at the same height fall on the same epipolar lines, and rays to two different hands can nearly meet close to the cameras. That made phantom hands 12-15 cm in front of the eyes, which tore holes through the screens. Each step now scores every way of pairing the views in two cameras and keeps the best. A pair scores well when its rays meet, when each view's apparent size matches the triangulated distance, and when the model calls both the same hand. The size check uses a fixed prior: with the model's average hand, clean pairs measure 0.71-1.51 times the one-view distance, and mismatched pairs mostly far less. It doesn't use the learned hand size, which bad pairs had corrupted.
|
||||
- **One view.** From how big the hand looks, its distance is off by 10-30% and wanders about 10% between frames. So a hand that drops to one camera keeps its last distance and drifts toward the one-view guess by 10% a frame.
|
||||
- **Smoothing.** The published landmarks go through a One Euro filter: it smooths hard while the hand is still (tracking noise is several mm per frame) and hardly at all while it moves fast. The palm speed that sets the update rate (15 or 30 Hz) is the filtered one; the raw speed read about 0.25 m/s from noise alone.
|
||||
- **Capsules.** Forearms follow the hand's own axis, and nothing within 12 cm in front of the eyes is published.
|
||||
|
||||
## Pinch
|
||||
|
||||
For input, the Vision Pro way: look at something and pinch to click, pinch and move to drag, with the eye tracker doing the looking. fh-tracker detects a pinch per hand (`pinch.h`) and publishes it to `$XDG_RUNTIME_DIR/frame-hands/gestures`, next to the hands file. The layout, and how to read it without missing quick taps, is in `include/fh_gestures.h`.
|
||||
|
||||
- A pinch begins when the thumb and index tips come within `--pinch-begin` (default 0.020 m). It ends when they open past `--pinch-end` (0.035 m) for 2 processed frames in a row, or when the hand stays lost for 0.25 s (flagged lost). While a pinch is down or closing, the tracker runs at the full 30 Hz.
|
||||
- The distance comes from MediaPipe's world landmarks (the model's own 3D hand pose, averaged over the hand's views, at the user's hand size). `--pinch-triangulated` uses the triangulated tips instead. On the two recordings without deliberate pinches, the world landmarks came under 2 cm in 0.2-1% of frames, against 3.3-4.5% for the triangulated tips. Typing still gave 2 pinches a minute, so a consumer should only act on a pinch while the gaze is on a target.
|
||||
- The pinch point is midway between the thumb and index tips. A drag is the pinch point now, minus where it was when the pinch began, both turned into the room with the HMD pose at their capture times.
|
||||
- `tools/watch_gestures.py` prints begins, ends and drag offsets live, and `--distance` prints each hand's distance. `fh-replay` runs the same detector and reports pinch counts. Its `--timeline` gets each begin, end and lost event, and both distance measures per set.
|
||||
|
||||
Frametop's pointer helper is the natural consumer. Its gaze mode already treats a press as "stop where the gaze put it, drag onto the target, click on release", and "hold still for half a second, then move" as a drag. A pinch begin would be the press, the end the release, and the pinch point's movement the drag.
|
||||
|
||||
## Replay
|
||||
|
||||
`fh-replay DIR` runs a recording through the tracker with the live scheduling and reports how well it kept the hands: hands per set, left and right coverage, track lengths, and the same reasons as the status line.
|
||||
|
||||
```
|
||||
trackd/fh-replay captures/rec-20260929-120000 --oracle 10 --timeline /tmp/tl.txt
|
||||
```
|
||||
|
||||
- `--oracle N`: every N-th set, also search every tile of every camera, and report how often the tracker had the hands that full search could find.
|
||||
- `--slow F`: live, the tracker skips sets that arrive while it's busy. Replay counts each step's time times F as busy (default 1; the headset is busier live).
|
||||
- `--timeline FILE`: a line per processed set and hand.
|
||||
- `--cams mono|color|all`: which cameras to track with (default `mono`). `color` tracks with the Arcturus pair alone, for comparing it with the IR cameras on the same recording. It needs a recording made with `fh-camd --with-color`. `--color-left NODE` (`color_video0` or `color_video3`) and `--color-crop subtract|none` say how the module's calibration maps onto the images; `tools/check_color.py` finds out. Color frames repeat across sets (the newest one is saved with each), and repeats are skipped.
|
||||
- `--depth FILE`: a line per hand per processed set for `tools/depth_report.py`, which measures the depth without ground truth. It reports how the hands were seen (by the two lower cameras, a lower and an upper one, or one camera), the noise along the line of sight against across it, each camera's one-view distance against the triangulated one, and what a camera dropping out would do to the distance.
|
||||
|
||||
## Playing a recording live
|
||||
|
||||
`fh-ringplay DIR --ring PATH [--from S] [--to S] [--loop]` publishes a recording into a ring file in real time, as fh-camd would, so `fh-tracker --ring PATH --no-publish` runs the same frames run after run. It needs no root, and it skips the dark frames. `probes/core_ab.py` uses it to compare CPU placements (A: CPUs 2-4, B: 5-7, C: no tracker), with the headset on so SteamVR's compositor is running.
|
||||
|
||||
## Build
|
||||
|
||||
ncnn is built from source into `vendor/ncnn/build/install`. The build also provides the `ncnn2table` and `ncnn2int8` quantization tools.
|
||||
|
||||
```
|
||||
git clone --depth 1 --branch 20260526 https://github.com/Tencent/ncnn.git vendor/ncnn
|
||||
cd vendor/ncnn && mkdir build && cd build
|
||||
cmake -G Ninja -DCMAKE_BUILD_TYPE=Release -DNCNN_VULKAN=OFF -DNCNN_BUILD_TOOLS=ON -DNCNN_SIMPLEOCV=ON \
|
||||
-DNCNN_BUILD_EXAMPLES=OFF -DNCNN_BUILD_TESTS=OFF -DCMAKE_INSTALL_PREFIX=$PWD/install ..
|
||||
nice ninja install
|
||||
cd ../../../trackd && make # fh-tracker and fh-replay; `make nettest` for the model check
|
||||
```
|
||||
|
||||
It needs jsoncpp for the calibration files (the host has it).
|
||||
|
||||
## Checks
|
||||
|
||||
- `tools/nettest_compare.py CAPTURE` runs the C++ model code (`nettest`) and `tracker/models.py` on the same crops from a recording, and compares the palms, ROIs and landmarks. Add `--int8` to check the 8-bit models.
|
||||
- `tools/make_int8_calib.py CAPTURE OUT` cuts the calibration crops that `ncnn2table` needs to quantize the models.
|
||||
@@ -0,0 +1,217 @@
|
||||
#include <cstdlib>
|
||||
#include "calib.h"
|
||||
|
||||
#include <json/json.h>
|
||||
|
||||
#include <algorithm>
|
||||
|
||||
#include <fstream>
|
||||
#include <iterator>
|
||||
#include <memory>
|
||||
|
||||
namespace {
|
||||
|
||||
double theta_d(const Camera &c, double t) {
|
||||
const double t2 = t * t;
|
||||
return t * (1 + t2 * (c.k[0] + t2 * (c.k[1] + t2 * (c.k[2] + t2 * c.k[3]))));
|
||||
}
|
||||
|
||||
// 4x4 transform (row-major) from a {plus_x, plus_z, position} pose.
|
||||
void pose(const Json::Value &d, double scale, double T[4][4]) {
|
||||
V3 x{d["plus_x"][0].asDouble(), d["plus_x"][1].asDouble(), d["plus_x"][2].asDouble()};
|
||||
V3 z{d["plus_z"][0].asDouble(), d["plus_z"][1].asDouble(), d["plus_z"][2].asDouble()};
|
||||
V3 y{z[1] * x[2] - z[2] * x[1], z[2] * x[0] - z[0] * x[2], z[0] * x[1] - z[1] * x[0]};
|
||||
for (int i = 0; i < 3; ++i) {
|
||||
T[i][0] = x[i], T[i][1] = y[i], T[i][2] = z[i];
|
||||
T[i][3] = d["position"][i].asDouble() * scale;
|
||||
T[3][i] = 0;
|
||||
}
|
||||
T[3][3] = 1;
|
||||
}
|
||||
|
||||
void mul(const double A[4][4], const double B[4][4], double C[4][4]) {
|
||||
for (int i = 0; i < 4; ++i)
|
||||
for (int j = 0; j < 4; ++j) {
|
||||
C[i][j] = 0;
|
||||
for (int k = 0; k < 4; ++k) C[i][j] += A[i][k] * B[k][j];
|
||||
}
|
||||
}
|
||||
|
||||
void invert_rigid(const double A[4][4], double B[4][4]) {
|
||||
for (int i = 0; i < 3; ++i)
|
||||
for (int j = 0; j < 3; ++j) B[i][j] = A[j][i];
|
||||
for (int i = 0; i < 3; ++i) B[i][3] = -(B[i][0] * A[0][3] + B[i][1] * A[1][3] + B[i][2] * A[2][3]);
|
||||
B[3][0] = B[3][1] = B[3][2] = 0, B[3][3] = 1;
|
||||
}
|
||||
|
||||
bool read_json(const char *path, Json::Value &v, std::string &err) {
|
||||
std::ifstream f(path);
|
||||
Json::CharReaderBuilder b;
|
||||
std::string e;
|
||||
if (!f || !Json::parseFromStream(b, f, &v, &e)) {
|
||||
err = std::string(path) + ": " + (f ? e : "can't open");
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
V2 Camera::project_cam(V3 p) const {
|
||||
const double r = std::hypot(p[0], p[1]);
|
||||
const double s = r > 1e-12 ? theta_d(*this, std::atan2(r, p[2])) / r : 0;
|
||||
return {fx * p[0] * s + cx, fy * p[1] * s + cy};
|
||||
}
|
||||
|
||||
V3 Camera::unproject(V2 uv) const {
|
||||
const double mx = (uv[0] - cx) / fx, my = (uv[1] - cy) / fy, td = std::hypot(mx, my);
|
||||
double t = td;
|
||||
for (int i = 0; i < 8; ++i) { // Newton on theta_d(t) = td
|
||||
const double t2 = t * t;
|
||||
const double df = 1 + t2 * (3 * k[0] + t2 * (5 * k[1] + t2 * (7 * k[2] + t2 * 9 * k[3])));
|
||||
t = std::clamp(t - (theta_d(*this, t) - td) / df, 0.0, M_PI);
|
||||
}
|
||||
const double s = td > 1e-12 ? std::sin(t) / td : 1;
|
||||
return {mx * s, my * s, std::cos(t)};
|
||||
}
|
||||
|
||||
V3 Camera::ray(V2 uv) const {
|
||||
const V3 c = unproject(uv);
|
||||
return {R[0][0] * c[0] + R[0][1] * c[1] + R[0][2] * c[2], R[1][0] * c[0] + R[1][1] * c[1] + R[1][2] * c[2],
|
||||
R[2][0] * c[0] + R[2][1] * c[1] + R[2][2] * c[2]};
|
||||
}
|
||||
|
||||
V2 Camera::project(V3 head, double *depth) const {
|
||||
const V3 d = head - origin;
|
||||
const V3 c{R[0][0] * d[0] + R[1][0] * d[1] + R[2][0] * d[2], R[0][1] * d[0] + R[1][1] * d[1] + R[2][1] * d[2],
|
||||
R[0][2] * d[0] + R[1][2] * d[1] + R[2][2] * d[2]};
|
||||
if (depth) *depth = c[2];
|
||||
return project_cam(c);
|
||||
}
|
||||
|
||||
double Camera::off_axis(V2 uv) const { return std::acos(std::clamp(unproject(uv)[2], -1.0, 1.0)) * 180 / M_PI; }
|
||||
|
||||
// Off the Frame (frame-job on the 7i), the /persist files are copies under FRAME_JOB_DEVICE_ROOT.
|
||||
static std::string device_path(const char *path) {
|
||||
const char *root = std::getenv("FRAME_JOB_DEVICE_ROOT");
|
||||
return root ? std::string(root) + path : std::string(path);
|
||||
}
|
||||
|
||||
bool load_calibration(std::map<std::string, Camera> &out, std::string &err) {
|
||||
Json::Value rig, dev;
|
||||
if (!read_json(device_path("/persist/xrservice.json").c_str(), rig, err) ||
|
||||
!read_json(device_path("/persist/device_config.json").c_str(), dev, err))
|
||||
return false;
|
||||
double cad_from_cam0[4][4], cad_from_head[4][4], head_from_cad[4][4], head_from_cam0[4][4];
|
||||
pose(dev["cv"]["cad_from_cal"], 1.0, cad_from_cam0);
|
||||
pose(dev["head"], 1.0, cad_from_head);
|
||||
invert_rigid(cad_from_head, head_from_cad);
|
||||
mul(head_from_cad, cad_from_cam0, head_from_cam0);
|
||||
for (const Json::Value &c : rig["cameras"]) {
|
||||
Camera cam;
|
||||
cam.name = c["sourceCamera"].asString();
|
||||
cam.width = c["width"].asInt(), cam.height = c["height"].asInt();
|
||||
for (const Json::Value &in : c["intrinsics"]) {
|
||||
if (in["cameraModel"].asString() != "kb") continue;
|
||||
cam.fx = in["fx"].asDouble(), cam.fy = in["fy"].asDouble();
|
||||
cam.cx = in["cx"].asDouble(), cam.cy = in["cy"].asDouble();
|
||||
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
|
||||
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
|
||||
}
|
||||
double cam0_from_cam[4][4], head_from_cam[4][4];
|
||||
pose(c["extrinsics"], 1e-3, cam0_from_cam);
|
||||
mul(head_from_cam0, cam0_from_cam, head_from_cam);
|
||||
for (int i = 0; i < 3; ++i) {
|
||||
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
|
||||
cam.origin[i] = head_from_cam[i][3];
|
||||
}
|
||||
out[cam.name] = cam;
|
||||
}
|
||||
if (out.empty()) err = "no cameras in /persist/xrservice.json";
|
||||
return !out.empty();
|
||||
}
|
||||
|
||||
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
|
||||
const std::string &right_node, bool crop_subtract, int scale, std::string &err) {
|
||||
// The module's EEPROM: some binary, then the calibration as JSON (world-readable)
|
||||
const std::string path = device_path("/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom");
|
||||
std::ifstream f(path, std::ios::binary);
|
||||
const std::string raw((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
||||
const size_t key = raw.find("\"alignment_method\"");
|
||||
const size_t start = key == std::string::npos ? key : raw.rfind('{', key);
|
||||
Json::Value rig, dev;
|
||||
std::string e;
|
||||
std::unique_ptr<Json::CharReader> reader(Json::CharReaderBuilder().newCharReader());
|
||||
if (start == std::string::npos || !reader->parse(raw.data() + start, raw.data() + raw.size(), &rig, &e))
|
||||
return err = path + ": no calibration JSON " + e, false;
|
||||
if (!read_json(device_path("/persist/device_config.json").c_str(), dev, err)) return false;
|
||||
double cad_from_head[4][4], head_from_cad[4][4];
|
||||
pose(dev["head"], 1.0, cad_from_head);
|
||||
invert_rigid(cad_from_head, head_from_cad);
|
||||
constexpr int kValidWidth = 1972; // pixels per row XRService's buffers deliver (of 2464)
|
||||
int n = 0;
|
||||
for (const Json::Value &c : rig["cameras"]) {
|
||||
const std::string source = c["sourceCamera"].asString();
|
||||
const std::string name = source == "passthrough_left" ? left_node : source == "passthrough_right" ? right_node : "";
|
||||
if (name.empty()) continue;
|
||||
Camera cam;
|
||||
cam.name = name;
|
||||
cam.width = kValidWidth / scale, cam.height = c["height"].asInt() / scale;
|
||||
const double dx = crop_subtract ? c["cropRegion"]["x"].asDouble() : 0, dy = crop_subtract ? c["cropRegion"]["y"].asDouble() : 0;
|
||||
for (const Json::Value &in : c["intrinsics"]) {
|
||||
if (in["cameraModel"].asString() != "kb") continue;
|
||||
// integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
|
||||
cam.fx = in["fx"].asDouble() / scale, cam.fy = in["fy"].asDouble() / scale;
|
||||
cam.cx = (in["cx"].asDouble() - dx + 0.5) / scale - 0.5, cam.cy = (in["cy"].asDouble() - dy + 0.5) / scale - 0.5;
|
||||
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
|
||||
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
|
||||
}
|
||||
double cad_from_cam[4][4], head_from_cam[4][4];
|
||||
pose(c["extrinsics"], 1e-3, cad_from_cam);
|
||||
mul(head_from_cad, cad_from_cam, head_from_cam);
|
||||
for (int i = 0; i < 3; ++i) {
|
||||
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
|
||||
cam.origin[i] = head_from_cam[i][3];
|
||||
}
|
||||
out[name] = cam;
|
||||
++n;
|
||||
}
|
||||
if (n != 2) err = path + ": expected passthrough_left and passthrough_right";
|
||||
return n == 2;
|
||||
}
|
||||
|
||||
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms) {
|
||||
double A[3][3] = {}, b[3] = {};
|
||||
for (int v = 0; v < n; ++v) {
|
||||
const V3 &d = dirs[v], &o = origins[v];
|
||||
for (int i = 0; i < 3; ++i)
|
||||
for (int j = 0; j < 3; ++j) {
|
||||
const double P = (i == j ? 1.0 : 0.0) - d[i] * d[j];
|
||||
A[i][j] += weights[v] * P;
|
||||
b[i] += weights[v] * P * o[j];
|
||||
}
|
||||
}
|
||||
// Cramer's rule for the 3x3 system
|
||||
auto det3 = [](const double m[3][3]) {
|
||||
return m[0][0] * (m[1][1] * m[2][2] - m[1][2] * m[2][1]) - m[0][1] * (m[1][0] * m[2][2] - m[1][2] * m[2][0]) +
|
||||
m[0][2] * (m[1][0] * m[2][1] - m[1][1] * m[2][0]);
|
||||
};
|
||||
const double D = det3(A);
|
||||
V3 p{};
|
||||
for (int c = 0; c < 3; ++c) {
|
||||
double M[3][3];
|
||||
for (int i = 0; i < 3; ++i)
|
||||
for (int j = 0; j < 3; ++j) M[i][j] = j == c ? b[i] : A[i][j];
|
||||
p[c] = std::fabs(D) > 1e-18 ? det3(M) / D : 0;
|
||||
}
|
||||
if (rms) {
|
||||
double s = 0;
|
||||
for (int v = 0; v < n; ++v) {
|
||||
const V3 off = p - origins[v];
|
||||
const V3 perp = off - dirs[v] * dot(off, dirs[v]);
|
||||
s += dot(perp, perp);
|
||||
}
|
||||
*rms = std::sqrt(s / n);
|
||||
}
|
||||
return p;
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
// Tracking-camera calibration from the headset's factory files (see tracker/calib.py
|
||||
// for the conventions): Kannala-Brandt fisheye intrinsics, and each camera's pose in the
|
||||
// head frame (OpenVR's: +x right, +y up, -z forward), metres.
|
||||
#pragma once
|
||||
|
||||
#include "geom.h"
|
||||
|
||||
#include <map>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
struct Camera {
|
||||
std::string name;
|
||||
int width = 0, height = 0;
|
||||
double fx = 1, fy = 1, cx = 0, cy = 0, k[4] = {};
|
||||
double R[3][3] = {}; // camera axes (columns) in the head frame
|
||||
V3 origin{}; // camera centre in the head frame
|
||||
|
||||
V2 project_cam(V3 p) const; // camera frame -> pixels
|
||||
V3 unproject(V2 uv) const; // pixels -> unit ray, camera frame
|
||||
V3 ray(V2 uv) const; // pixels -> unit ray, head frame
|
||||
V2 project(V3 head, double *depth) const; // head frame -> pixels; depth along the optical axis
|
||||
double off_axis(V2 uv) const; // degrees between the pixel's ray and the axis
|
||||
};
|
||||
|
||||
// Loads /persist/xrservice.json and /persist/device_config.json. Keyed by calibration
|
||||
// name: slam_left, slam_right, upper_left, upper_right.
|
||||
bool load_calibration(std::map<std::string, Camera> &out, std::string &err);
|
||||
|
||||
// The Arcturus color cameras (tracker/calib.py load_color has the conventions), for
|
||||
// fh-camd --with-color's images: luma at 1/scale size, recorded as color_video<N>. They're
|
||||
// keyed by those recorded names: left_node is passthrough_left, right_node
|
||||
// passthrough_right. crop_subtract: image x = sensor x - the calibration's cropRegion.x.
|
||||
// tools/check_color.py tells which node is which and which crop reading fits.
|
||||
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
|
||||
const std::string &right_node, bool crop_subtract, int scale, std::string &err);
|
||||
|
||||
// The point closest to several rays (weighted), and its rms distance to them.
|
||||
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms);
|
||||
@@ -0,0 +1,22 @@
|
||||
// Small vector helpers for the tracker.
|
||||
#pragma once
|
||||
|
||||
#include <array>
|
||||
#include <cmath>
|
||||
|
||||
using V2 = std::array<double, 2>;
|
||||
using V3 = std::array<double, 3>;
|
||||
|
||||
inline V3 operator+(V3 a, V3 b) { return {a[0] + b[0], a[1] + b[1], a[2] + b[2]}; }
|
||||
inline V3 operator-(V3 a, V3 b) { return {a[0] - b[0], a[1] - b[1], a[2] - b[2]}; }
|
||||
inline V3 operator*(V3 a, double s) { return {a[0] * s, a[1] * s, a[2] * s}; }
|
||||
inline double dot(V3 a, V3 b) { return a[0] * b[0] + a[1] * b[1] + a[2] * b[2]; }
|
||||
inline double norm(V3 a) { return std::sqrt(dot(a, a)); }
|
||||
inline V3 unit(V3 a) { double n = norm(a); return n > 0 ? a * (1 / n) : a; }
|
||||
|
||||
inline V2 operator+(V2 a, V2 b) { return {a[0] + b[0], a[1] + b[1]}; }
|
||||
inline V2 operator-(V2 a, V2 b) { return {a[0] - b[0], a[1] - b[1]}; }
|
||||
inline V2 operator*(V2 a, double s) { return {a[0] * s, a[1] * s}; }
|
||||
inline double norm(V2 a) { return std::hypot(a[0], a[1]); }
|
||||
|
||||
inline double wrap_angle(double a) { return std::remainder(a, 2 * M_PI); }
|
||||
@@ -0,0 +1,169 @@
|
||||
#include "io.h"
|
||||
|
||||
#include <fcntl.h>
|
||||
#include <sys/mman.h>
|
||||
#include <sys/stat.h>
|
||||
#include <time.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
|
||||
uint64_t mono_ns() {
|
||||
timespec ts;
|
||||
clock_gettime(CLOCK_MONOTONIC, &ts);
|
||||
return uint64_t(ts.tv_sec) * 1'000'000'000 + uint64_t(ts.tv_nsec);
|
||||
}
|
||||
|
||||
int64_t raw_minus_mono_ns() {
|
||||
timespec a, r, b;
|
||||
clock_gettime(CLOCK_MONOTONIC, &a);
|
||||
clock_gettime(CLOCK_MONOTONIC_RAW, &r);
|
||||
clock_gettime(CLOCK_MONOTONIC, &b);
|
||||
const int64_t ma = int64_t(a.tv_sec) * 1'000'000'000 + a.tv_nsec, mb = int64_t(b.tv_sec) * 1'000'000'000 + b.tv_nsec;
|
||||
return int64_t(r.tv_sec) * 1'000'000'000 + r.tv_nsec - (ma + mb) / 2;
|
||||
}
|
||||
|
||||
// ------------------------------------------------------------------------------ ring
|
||||
|
||||
bool Ring::open(const char *path, std::string &err) {
|
||||
const int fd = ::open(path, O_RDONLY | O_CLOEXEC);
|
||||
if (fd < 0) return err = std::string(path) + ": " + std::strerror(errno), false;
|
||||
struct stat st;
|
||||
fstat(fd, &st);
|
||||
len_ = size_t(st.st_size);
|
||||
void *m = len_ >= sizeof(fh_ring_hdr_t) ? mmap(nullptr, len_, PROT_READ, MAP_SHARED, fd, 0) : MAP_FAILED;
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return err = std::string(path) + ": can't map it", false;
|
||||
map_ = static_cast<const uint8_t *>(m);
|
||||
hdr_ = reinterpret_cast<const fh_ring_hdr_t *>(map_);
|
||||
if (std::memcmp(hdr_->magic, FH_RING_MAGIC, 8) || hdr_->version != FH_RING_VERSION || hdr_->file_bytes > len_)
|
||||
return err = std::string(path) + " is not an fh-camd ring", false;
|
||||
return true;
|
||||
}
|
||||
|
||||
bool Ring::alive() const {
|
||||
const uint64_t hb = __atomic_load_n(&hdr_->heartbeat_ns, __ATOMIC_ACQUIRE);
|
||||
return hb && mono_ns() - hb < 1'000'000'000;
|
||||
}
|
||||
|
||||
uint64_t Ring::latest(int i) const { return __atomic_load_n(&hdr_->cams[i].latest, __ATOMIC_ACQUIRE); }
|
||||
|
||||
bool Ring::read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const {
|
||||
const fh_ring_cam_t &c = hdr_->cams[i];
|
||||
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
|
||||
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
|
||||
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
|
||||
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
|
||||
if (seq != 2 * n + 2) return false;
|
||||
std::memcpy(meta, slot, sizeof *meta);
|
||||
out.resize(size_t(c.width) * c.height);
|
||||
for (uint32_t y = 0; y < c.height; ++y)
|
||||
std::memcpy(out.data() + size_t(y) * c.width, slot + sizeof(fh_ring_slot_t) + size_t(y) * c.stride, c.width);
|
||||
__atomic_thread_fence(__ATOMIC_ACQUIRE);
|
||||
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
|
||||
}
|
||||
|
||||
bool Ring::meta(int i, uint64_t n, fh_ring_slot_t *meta) const {
|
||||
const fh_ring_cam_t &c = hdr_->cams[i];
|
||||
if (!n || c.slot_offset + c.nslots * c.slot_bytes > len_) return false;
|
||||
const uint8_t *slot = map_ + c.slot_offset + (n % c.nslots) * c.slot_bytes;
|
||||
const auto *s = reinterpret_cast<const fh_ring_slot_t *>(slot);
|
||||
const uint64_t seq = __atomic_load_n(&s->seq, __ATOMIC_ACQUIRE);
|
||||
if (seq != 2 * n + 2) return false;
|
||||
std::memcpy(meta, slot, sizeof *meta);
|
||||
__atomic_thread_fence(__ATOMIC_ACQUIRE);
|
||||
return __atomic_load_n(&s->seq, __ATOMIC_RELAXED) == seq;
|
||||
}
|
||||
|
||||
// ------------------------------------------------------------------------- publisher
|
||||
|
||||
namespace {
|
||||
|
||||
// The hand's shape to cut out, as capsules (tracker/publish.py has the same model).
|
||||
// Radii are a real hand's half-widths plus a small margin for tracking noise.
|
||||
const int kThumb[][2] = {{0, 1}, {1, 2}, {2, 3}, {3, 4}};
|
||||
const int kFingers[][2] = {{5, 6}, {6, 7}, {7, 8}, {9, 10}, {10, 11}, {11, 12}, {13, 14}, {14, 15}, {15, 16},
|
||||
{17, 18}, {18, 19}, {19, 20}};
|
||||
const int kPalm[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 9}, {9, 13}, {13, 17}, {1, 5}};
|
||||
constexpr double kThumbR = 0.0095, kFinger = 0.0085, kPalmR = 0.015, kArm[2] = {0.028, 0.034}, kArmLen = 0.16,
|
||||
kMargin = 0.004;
|
||||
// Nothing is cut closer than this in front of the eyes (head frame, -z is forward). A
|
||||
// point near the eyes' plane lands far across a screen with a huge radius, so one bad
|
||||
// estimate there tears a hole through it; real hands that close aren't tracked anyway.
|
||||
constexpr double kNear = 0.12;
|
||||
|
||||
// Adds the capsule, clipped to the part at least kNear in front of the eyes.
|
||||
void put(fh_capsule_t *caps, uint32_t &n, V3 a, V3 b, double ra, double rb) {
|
||||
if (n >= FH_HANDS_MAX_CAPSULES) return;
|
||||
const double za = -a[2] - kNear, zb = -b[2] - kNear; // >= 0: far enough in front
|
||||
if (za < 0 && zb < 0) return;
|
||||
if (za < 0 || zb < 0) {
|
||||
const double t = za / (za - zb); // where the segment crosses the near plane
|
||||
const V3 m = a + (b - a) * t;
|
||||
const double rm = ra + (rb - ra) * t;
|
||||
if (za < 0) a = m, ra = rm;
|
||||
else b = m, rb = rm;
|
||||
}
|
||||
fh_capsule_t &c = caps[n++];
|
||||
for (int k = 0; k < 3; ++k) c.a[k] = float(a[k]), c.b[k] = float(b[k]);
|
||||
c.ra = float(ra + kMargin), c.rb = float(rb + kMargin);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
bool Publisher::open(std::string &err) {
|
||||
const char *run = std::getenv("XDG_RUNTIME_DIR");
|
||||
const std::string dir = std::string(run ? run : "/run/user/" + std::to_string(getuid())) + "/frame-hands";
|
||||
mkdir(dir.c_str(), 0700);
|
||||
chmod(dir.c_str(), 0700);
|
||||
const std::string path = dir + "/hands";
|
||||
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
|
||||
if (fd < 0 || ftruncate(fd, sizeof(fh_hands_t)) < 0) return err = path + ": " + std::strerror(errno), false;
|
||||
void *m = mmap(nullptr, sizeof(fh_hands_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return err = path + ": can't map it", false;
|
||||
out_ = static_cast<fh_hands_t *>(m);
|
||||
std::memset(out_, 0, sizeof *out_);
|
||||
std::memcpy(out_->magic, FH_HANDS_MAGIC, 8);
|
||||
out_->version = FH_HANDS_VERSION;
|
||||
out_->size = sizeof(fh_hands_t);
|
||||
return true;
|
||||
}
|
||||
|
||||
void Publisher::write(const std::vector<const Hand *> &in, uint64_t capture_ns) {
|
||||
std::vector<const Hand *> hands = in;
|
||||
std::sort(hands.begin(), hands.end(), [](const Hand *a, const Hand *b) { return a->frames > b->frames; });
|
||||
if (hands.size() > FH_HANDS_MAX_HANDS) hands.resize(FH_HANDS_MAX_HANDS);
|
||||
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
uint32_t nc = 0;
|
||||
for (size_t k = 0; k < FH_HANDS_MAX_HANDS; ++k) {
|
||||
fh_hand_t &o = out_->hands[k];
|
||||
std::memset(&o, 0, sizeof o);
|
||||
if (k >= hands.size()) continue;
|
||||
const Hand &h = *hands[k];
|
||||
o.id = uint32_t(h.id);
|
||||
o.flags = (h.right() ? FH_HAND_RIGHT : 0) | (h.nviews >= 2 ? FH_HAND_STEREO : 0);
|
||||
o.confidence = float(std::min(1.0, h.frames / 5.0));
|
||||
for (int i = 0; i < 21; ++i)
|
||||
for (int j = 0; j < 3; ++j) o.pts[i][j] = float(h.smooth[i][j]);
|
||||
const uint32_t first = nc;
|
||||
for (auto &b : kThumb) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kThumbR, kThumbR);
|
||||
for (auto &b : kFingers) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kFinger, kFinger);
|
||||
for (auto &b : kPalm) put(out_->capsules, nc, h.smooth[b[0]], h.smooth[b[1]], kPalmR, kPalmR);
|
||||
// the forearm carries on from the hand's own axis (middle knuckle -> wrist); the
|
||||
// wrist bends, but much less than a guess at where the elbow is gets wrong
|
||||
const V3 wrist = h.smooth[0], d = wrist - h.smooth[9];
|
||||
const double n = norm(d);
|
||||
if (n > 0.02) put(out_->capsules, nc, wrist, wrist + d * (kArmLen / n), kArm[0], kArm[1]);
|
||||
o.ncapsules = nc - first;
|
||||
}
|
||||
for (uint32_t k = nc; k < FH_HANDS_MAX_CAPSULES; ++k) std::memset(&out_->capsules[k], 0, sizeof(fh_capsule_t));
|
||||
out_->capture_ns = capture_ns;
|
||||
out_->publish_ns = mono_ns();
|
||||
out_->nhands = uint32_t(hands.size());
|
||||
out_->ncapsules = nc;
|
||||
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||
}
|
||||
@@ -0,0 +1,46 @@
|
||||
// Frames in from fh-camd's ring (camd/fhring.h), hands out to the hands file
|
||||
// (include/fh_hands.h, read by Frametop's ft-screens).
|
||||
#pragma once
|
||||
|
||||
#include "tracker.h"
|
||||
|
||||
#include <cstdint>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
extern "C" {
|
||||
#include "../camd/fhring.h"
|
||||
#include "../include/fh_hands.h"
|
||||
}
|
||||
|
||||
class Ring {
|
||||
public:
|
||||
bool open(const char *path, std::string &err);
|
||||
bool alive() const; // the writer's heartbeat is fresh
|
||||
int cameras() const { return int(hdr_->ncams); }
|
||||
const fh_ring_cam_t &camera(int i) const { return hdr_->cams[i]; }
|
||||
uint64_t latest(int i) const;
|
||||
// Copy frame n of camera i into out (width x height, tightly packed). False if it's
|
||||
// gone or was being written.
|
||||
bool read(int i, uint64_t n, std::vector<uint8_t> &out, fh_ring_slot_t *meta) const;
|
||||
// Just frame n's slot header (capture time etc.), without copying the image.
|
||||
bool meta(int i, uint64_t n, fh_ring_slot_t *meta) const;
|
||||
|
||||
private:
|
||||
const uint8_t *map_ = nullptr;
|
||||
const fh_ring_hdr_t *hdr_ = nullptr;
|
||||
size_t len_ = 0;
|
||||
};
|
||||
|
||||
class Publisher {
|
||||
public:
|
||||
bool open(std::string &err);
|
||||
void write(const std::vector<const Hand *> &hands, uint64_t capture_ns);
|
||||
|
||||
private:
|
||||
fh_hands_t *out_ = nullptr;
|
||||
uint64_t seq_ = 0;
|
||||
};
|
||||
|
||||
uint64_t mono_ns();
|
||||
int64_t raw_minus_mono_ns(); // camera timestamps are CLOCK_MONOTONIC_RAW
|
||||
@@ -0,0 +1,331 @@
|
||||
// fh-tracker: hands in 3D from fh-camd's ring, published for Frametop's ft-screens.
|
||||
// The C++ version of tracker/live.py: the same scheduling, with the models on a few
|
||||
// threads and no Python in the loop.
|
||||
//
|
||||
// fh-tracker [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N]
|
||||
// [--no-publish] [--record DIR] [--swap-sides] ... (--help lists them all)
|
||||
#include "io.h"
|
||||
#include "pinch.h"
|
||||
#include "record.h"
|
||||
|
||||
#include <sched.h>
|
||||
#include <sys/resource.h>
|
||||
#include <sys/stat.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <ctime>
|
||||
#include <memory>
|
||||
|
||||
#include <csignal>
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
#include <fstream>
|
||||
#include <thread>
|
||||
#include <utility>
|
||||
|
||||
namespace {
|
||||
|
||||
volatile std::sig_atomic_t g_stop = 0, g_record = 0;
|
||||
|
||||
// which calibrated camera each capture pipe carries (XRService's fixed routing)
|
||||
const char *camera_for_pipe(int node) {
|
||||
char path[64], name[64] = "";
|
||||
std::snprintf(path, sizeof path, "/sys/class/video4linux/video%d/name", node);
|
||||
std::ifstream f(path);
|
||||
f.getline(name, sizeof name);
|
||||
if (!std::strcmp(name, "msm_vfe3_video0")) return "slam_left";
|
||||
if (!std::strcmp(name, "msm_vfe4_video0")) return "slam_right";
|
||||
if (!std::strcmp(name, "msm_vfe2_video0")) return "upper_left";
|
||||
if (!std::strcmp(name, "msm_vfe2_video1")) return "upper_right";
|
||||
return nullptr;
|
||||
}
|
||||
|
||||
double cpu_seconds() {
|
||||
rusage r;
|
||||
getrusage(RUSAGE_SELF, &r);
|
||||
return r.ru_utime.tv_sec + r.ru_stime.tv_sec + (r.ru_utime.tv_usec + r.ru_stime.tv_usec) / 1e6;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char **argv) {
|
||||
double seconds = 0, status = 5;
|
||||
int threads = 3, niceness = 5;
|
||||
bool int8 = false, publish = true, track = true, swap_sides = false;
|
||||
std::string models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
|
||||
std::string record, ring_path = FH_RING_PATH;
|
||||
// SteamOS starts user processes on CPUs 0-4 and keeps 5-7 (two A720s and the X4) for
|
||||
// SteamVR's compositor, whose threads there run at real-time priority, so they always
|
||||
// win. XRService pins its head tracking to 2-3. probes/core_ab.py (2026-09-29, headset
|
||||
// on, 3 rounds): on 5-7 a step took 8.4 ms against 13.2 on 2-4, latency 9.6 against
|
||||
// 14.1 ms, and the compositor's late frames and CPU/GPU time didn't change.
|
||||
std::vector<int> cpus = {5, 6, 7};
|
||||
// How crops are equalized. CLAHE helps the palm search find hands (about 10% more in the
|
||||
// dim recording), but makes the landmarks jitter, so they get plain crops.
|
||||
Contrast palm_contrast, hand_contrast{Contrast::None};
|
||||
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
||||
PinchParams pinch_params;
|
||||
double record_for = 120;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
const std::string a = argv[i];
|
||||
const bool more = i + 1 < argc;
|
||||
if (a == "--seconds" && more) seconds = std::atof(argv[++i]);
|
||||
else if (a == "--threads" && more) threads = std::max(1, std::atoi(argv[++i]));
|
||||
else if (a == "--status" && more) status = std::atof(argv[++i]);
|
||||
else if (a == "--models" && more) models = argv[++i];
|
||||
else if (a == "--nice" && more) niceness = std::atoi(argv[++i]);
|
||||
else if (a == "--int8") int8 = true;
|
||||
else if (a == "--no-publish") publish = false;
|
||||
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
|
||||
else if (a == "--swap-sides") swap_sides = true;
|
||||
else if (a == "--record-only") track = publish = false;
|
||||
else if (a == "--ring" && more) ring_path = argv[++i];
|
||||
else if (a == "--record" && more) record = argv[++i];
|
||||
else if (a == "--record-for" && more) record_for = std::atof(argv[++i]);
|
||||
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
|
||||
else if (a == "--contrast" && more) {
|
||||
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
|
||||
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
|
||||
} else if (a == "--cpus" && more) {
|
||||
cpus.clear();
|
||||
for (char *p = argv[++i]; *p;) {
|
||||
cpus.push_back(int(std::strtol(p, &p, 10)));
|
||||
if (*p == ',') ++p;
|
||||
else if (*p) break;
|
||||
}
|
||||
if (cpus.empty()) cpus = {5, 6, 7};
|
||||
}
|
||||
else {
|
||||
std::printf("usage: %s [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] [--no-publish]\n"
|
||||
" [--record DIR] [--record-for S] [--record-only] [--cpus 5,6,7] [--swap-sides]\n"
|
||||
" [--keep-presence P] (0.5) [--ring PATH] (fh-camd's, or fh-ringplay's)\n"
|
||||
" [--pinch-begin M] (0.020) [--pinch-end M] (0.035) [--pinch-triangulated]\n"
|
||||
" [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch; default clahe:2/none)\n"
|
||||
"Recording saves every frame set for S seconds (120) to DIR/sets.bin, for fh-replay; SIGUSR1\n"
|
||||
"starts one in captures/rec-<time> next to trackd. --record-only records without tracking, so it\n"
|
||||
"can run beside a tracking fh-tracker. With fh-camd --with-dark, recordings also get each\n"
|
||||
"camera's newest dark frame, as <name>_dk; with --with-color, the color cameras' as color_video<N>.\n",
|
||||
argv[0]);
|
||||
return a == "--help" ? 0 : 1;
|
||||
}
|
||||
}
|
||||
if (nice(niceness) < 0) std::perror("nice"); // the VR stack wins contested CPUs
|
||||
std::signal(SIGINT, [](int) { g_stop = 1; });
|
||||
std::signal(SIGTERM, [](int) { g_stop = 1; });
|
||||
std::signal(SIGUSR1, [](int) { g_record = 1; });
|
||||
|
||||
std::string err;
|
||||
std::map<std::string, Camera> calib;
|
||||
Ring ring;
|
||||
Nets nets;
|
||||
Publisher pub;
|
||||
GesturePublisher gestures;
|
||||
Pinch pinch(pinch_params);
|
||||
std::unique_ptr<Recorder> rec;
|
||||
uint64_t rec_start = 0;
|
||||
auto start_recording = [&](const std::string &dir, std::string &e) {
|
||||
rec = std::make_unique<Recorder>();
|
||||
if (!rec->open(dir, e)) return rec.reset(), false;
|
||||
rec_start = mono_ns();
|
||||
std::printf("recording to %s for %.0f s\n", dir.c_str(), record_for);
|
||||
std::fflush(stdout);
|
||||
return true;
|
||||
};
|
||||
if (!load_calibration(calib, err) || !ring.open(ring_path.c_str(), err) || !nets.load(models, int8, err) ||
|
||||
(publish && (!pub.open(err) || !gestures.open(pinch, err))) || (!record.empty() && !start_recording(record, err))) {
|
||||
std::fprintf(stderr, "%s\n", err.c_str());
|
||||
return 1;
|
||||
}
|
||||
if (!ring.alive()) return std::fprintf(stderr, "fh-camd isn't running (no heartbeat)\n"), 1;
|
||||
|
||||
std::map<std::string, int> index; // calibration name -> ring camera
|
||||
// Recorded only, not tracked: "<name>_dk" (fh-camd --with-dark) and "color_video<N>"
|
||||
// (--with-color; which is left and right is up to tools/check_color.py). Recorded names
|
||||
// hold 15 characters, so "upper_right_dark" wouldn't fit.
|
||||
std::map<std::string, int> dark;
|
||||
std::map<std::string, Camera> used;
|
||||
for (int i = 0; i < ring.cameras(); ++i) {
|
||||
if (ring.camera(i).flags & FH_CAM_COLOR) {
|
||||
dark["color_video" + std::to_string(ring.camera(i).node)] = i;
|
||||
continue;
|
||||
}
|
||||
// fh-camd's cameras by capture pipe; fh-ringplay's (no device) by the name it gives
|
||||
const char *name = camera_for_pipe(ring.camera(i).node);
|
||||
if (!name && ring.camera(i).node < 0) name = ring.camera(i).name;
|
||||
if (!name || !calib.count(name)) continue;
|
||||
if (ring.camera(i).flags & FH_CAM_DARK) dark[std::string(name) + "_dk"] = i;
|
||||
else index[name] = i, used[name] = calib[name];
|
||||
}
|
||||
// fh-camd tells the side cameras' buffers apart by XRService's allocation order, which
|
||||
// some XRService restarts reverse; tools/check_sides.py --ring tells when.
|
||||
if (swap_sides && index.count("slam_left") && index.count("slam_right")) {
|
||||
std::swap(index["slam_left"], index["slam_right"]);
|
||||
if (dark.count("slam_left_dk") && dark.count("slam_right_dk")) std::swap(dark["slam_left_dk"], dark["slam_right_dk"]);
|
||||
std::printf("side cameras swapped (--swap-sides)\n");
|
||||
}
|
||||
std::printf("cameras:");
|
||||
for (auto &[name, i] : index) std::printf(" %s=video%d", name.c_str(), ring.camera(i).node);
|
||||
std::printf(" models: %s%s, %d threads on CPUs", models.c_str(), int8 ? " (int8)" : "", threads);
|
||||
for (int c : cpus) std::printf(" %d", c);
|
||||
std::printf("\n");
|
||||
|
||||
cpu_set_t set; // the main loop too
|
||||
CPU_ZERO(&set);
|
||||
for (int c : cpus) CPU_SET(c, &set);
|
||||
if (sched_setaffinity(0, sizeof set, &set) < 0) std::perror("sched_setaffinity");
|
||||
nets.set_contrast(palm_contrast, hand_contrast);
|
||||
Pool pool(threads, cpus);
|
||||
Tracker tracker(used, nets, pool);
|
||||
tracker.set_keep_presence(keep_presence);
|
||||
std::map<std::string, std::vector<uint8_t>> pixels;
|
||||
std::map<std::string, uint64_t> last;
|
||||
const uint64_t start = mono_ns();
|
||||
uint64_t t_status = start, next_ns = 0;
|
||||
double cpu0 = cpu_seconds();
|
||||
std::vector<double> lat;
|
||||
double hands_sum = 0, resid_sum = 0;
|
||||
int resid_n = 0, left_sets = 0, right_sets = 0, both_sets = 0;
|
||||
|
||||
while (!g_stop && (seconds <= 0 || (mono_ns() - start) / 1e9 < seconds)) {
|
||||
if (!ring.alive()) return std::fprintf(stderr, "fh-camd stopped\n"), 2;
|
||||
// a new frame set: every camera has a newer frame, taken at the same moment
|
||||
std::map<std::string, uint64_t> latest;
|
||||
bool ready = true;
|
||||
for (auto &[name, i] : index) {
|
||||
latest[name] = ring.latest(i);
|
||||
ready = ready && latest[name] > last[name];
|
||||
}
|
||||
if (!ready) {
|
||||
std::this_thread::sleep_for(std::chrono::milliseconds(2));
|
||||
continue;
|
||||
}
|
||||
// not needed at the current rate, and not recorded: skip it without copying images
|
||||
if (track && !rec && !g_record) {
|
||||
uint64_t t0 = UINT64_MAX, t1 = 0;
|
||||
bool ok = true;
|
||||
for (auto &[name, i] : index) {
|
||||
fh_ring_slot_t meta;
|
||||
ok = ok && ring.meta(i, latest[name], &meta);
|
||||
if (ok) t0 = std::min(t0, meta.capture_ns), t1 = std::max(t1, meta.capture_ns);
|
||||
}
|
||||
if (ok && t1 - t0 <= 3'000'000 && t0 < next_ns) {
|
||||
last = latest;
|
||||
continue;
|
||||
}
|
||||
}
|
||||
std::map<std::string, Image> images;
|
||||
std::vector<SetFrame> frames;
|
||||
uint64_t tmin = UINT64_MAX, tmax = 0, dq = 0;
|
||||
bool ok = true;
|
||||
for (auto &[name, i] : index) {
|
||||
fh_ring_slot_t meta;
|
||||
ok = ok && ring.read(i, latest[name], pixels[name], &meta);
|
||||
if (!ok) break;
|
||||
const auto &c = ring.camera(i);
|
||||
images[name] = {pixels[name].data(), int(c.width), int(c.height), int(c.width)};
|
||||
frames.push_back({name, pixels[name].data(), c.width, c.height, meta.capture_ns, meta.dqbuf_ns});
|
||||
tmin = std::min(tmin, meta.capture_ns), tmax = std::max(tmax, meta.capture_ns), dq = std::max(dq, meta.dqbuf_ns);
|
||||
}
|
||||
if (!ok || tmax - tmin > 3'000'000) { // torn, or a camera is a frame behind
|
||||
std::this_thread::sleep_for(std::chrono::milliseconds(1));
|
||||
continue;
|
||||
}
|
||||
last = latest;
|
||||
if (g_record && !rec) {
|
||||
g_record = 0;
|
||||
char name[64];
|
||||
const std::time_t now = std::time(nullptr);
|
||||
std::strftime(name, sizeof name, "rec-%Y%m%d-%H%M%S", std::localtime(&now));
|
||||
const std::string here = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1);
|
||||
const std::string dir = here + "../captures";
|
||||
mkdir(dir.c_str(), 0755);
|
||||
std::string e;
|
||||
if (!start_recording(dir + "/" + name, e)) std::fprintf(stderr, "%s\n", e.c_str());
|
||||
}
|
||||
if (rec) { // about 80 MB/s; dark frames double that, color frames add 70 MB/s
|
||||
if ((mono_ns() - rec_start) / 1e9 < record_for) {
|
||||
for (auto &[name, i] : dark) { // the newest dark and color frames, as they are
|
||||
fh_ring_slot_t meta;
|
||||
const uint64_t n = ring.latest(i);
|
||||
const auto &c = ring.camera(i);
|
||||
if (n && ring.read(i, n, pixels[name], &meta))
|
||||
frames.push_back({name, pixels[name].data(), c.width, c.height, meta.capture_ns, meta.dqbuf_ns});
|
||||
}
|
||||
rec->add(frames);
|
||||
} else {
|
||||
const size_t n = rec->written(), d = rec->dropped();
|
||||
rec.reset(); // writes out what's queued
|
||||
std::printf("recording done: %zu sets, %zu dropped\n", n, d);
|
||||
std::fflush(stdout);
|
||||
if (!track) break;
|
||||
}
|
||||
}
|
||||
if (!track && rec && status > 0 && (mono_ns() - t_status) / 1e9 >= status) {
|
||||
std::printf("%5.1fs recorded %zu sets, dropped %zu\n", (mono_ns() - start) / 1e9, rec->written(), rec->dropped());
|
||||
std::fflush(stdout);
|
||||
t_status = mono_ns();
|
||||
}
|
||||
if (!track || tmin < next_ns) continue; // not needed yet at the current rate
|
||||
const auto hands = tracker.step(images, int64_t(tmin));
|
||||
const uint64_t capture = uint64_t(int64_t(tmin) - raw_minus_mono_ns()); // CLOCK_MONOTONIC
|
||||
pinch.update(hands, tracker.views_now(), int64_t(capture));
|
||||
// a pinch down or closing gets the full rate, even while the palm holds still
|
||||
next_ns = tmin + uint64_t((std::min(tracker.interval(), pinch.engaged() ? 1 / 30.0 : 1.0) - 0.005) * 1e9);
|
||||
if (publish) pub.write(hands, capture), gestures.write(pinch, capture);
|
||||
for (const Pinch::Event &e : pinch.events) {
|
||||
std::printf("pinch %s %-5s d %.3f m at %+.3f %+.3f %+.3f\n", e.side ? "right" : "left ", e.what, e.distance,
|
||||
e.point[0], e.point[1], e.point[2]);
|
||||
std::fflush(stdout);
|
||||
}
|
||||
lat.push_back((mono_ns() - dq) / 1e6);
|
||||
hands_sum += double(hands.size());
|
||||
bool on_left = false, on_right = false; // by where the wrist is, not the model's label
|
||||
for (const Hand *h : hands) {
|
||||
if (h->residual >= 0) resid_sum += h->residual * 1000, ++resid_n;
|
||||
(h->pts[0][0] < 0 ? on_left : on_right) = true;
|
||||
}
|
||||
left_sets += on_left, right_sets += on_right, both_sets += on_left && on_right;
|
||||
|
||||
const uint64_t now = mono_ns();
|
||||
if (status > 0 && (now - t_status) / 1e9 >= status) {
|
||||
const double dt = (now - t_status) / 1e9, cpu1 = cpu_seconds();
|
||||
const Stats &s = tracker.stats;
|
||||
std::sort(lat.begin(), lat.end());
|
||||
std::printf("%5.1fs %4.1f sets/s hands %.2f views %zu palm %3d calls %4.1f ms/batch hand %3d calls %4.1f ms/batch "
|
||||
"step %4.1f ms latency %4.1f ms resid %.1f mm CPU %3.0f%%\n",
|
||||
(now - start) / 1e9, s.sets / dt, s.sets ? hands_sum / s.sets : 0, tracker.views(), s.palm_calls,
|
||||
s.palm_batches ? s.palm_ms / s.palm_batches : 0, s.hand_calls,
|
||||
s.hand_batches ? s.hand_ms / s.hand_batches : 0, s.sets ? s.step_ms / s.sets : 0,
|
||||
lat.empty() ? 0 : lat[lat.size() / 2], resid_n ? resid_sum / resid_n : 0, 100 * (cpu1 - cpu0) / dt);
|
||||
if (s.sets)
|
||||
std::printf(" sets with a hand: left %2.0f%% right %2.0f%% both %2.0f%% views lost %d, handoff misses %d, "
|
||||
"dups %d, splits %d hands new %d merged %d forgotten %d%s\n",
|
||||
100.0 * left_sets / s.sets, 100.0 * right_sets / s.sets, 100.0 * both_sets / s.sets, s.lost,
|
||||
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten,
|
||||
!rec ? "" : (" recorded " + std::to_string(rec->written()) + " dropped " +
|
||||
std::to_string(rec->dropped())).c_str());
|
||||
std::printf(" pinches: left %u right %u", pinch.side(0).begins, pinch.side(1).begins);
|
||||
for (int k = 0; k < 2; ++k)
|
||||
if (pinch.side(k).flags & FH_PINCH_TRACKED)
|
||||
std::printf(" %s %s d %.3f", k ? "right" : "left", pinch.side(k).flags & FH_PINCH_DOWN ? "DOWN" : "open",
|
||||
pinch.side(k).distance);
|
||||
std::printf("\n");
|
||||
for (const Hand *h : hands)
|
||||
std::printf(" hand %d %-5s views %d wrist %+.3f %+.3f %+.3f m scale %.2f speed %.2f m/s\n", h->id,
|
||||
h->right() ? "right" : "left", h->nviews, h->pts[0][0], h->pts[0][1], h->pts[0][2], h->scale,
|
||||
h->speed);
|
||||
std::fflush(stdout);
|
||||
tracker.stats = Stats{};
|
||||
t_status = now, cpu0 = cpu1;
|
||||
lat.clear(), hands_sum = 0, resid_sum = 0, resid_n = 0, left_sets = right_sets = both_sets = 0;
|
||||
}
|
||||
}
|
||||
if (publish) {
|
||||
pub.write({}, mono_ns());
|
||||
pinch.release(int64_t(mono_ns())); // a drag in progress ends, as lost
|
||||
gestures.write(pinch, mono_ns());
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,255 @@
|
||||
#include "nets.h"
|
||||
|
||||
#include <mat.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstdlib>
|
||||
#include <numeric>
|
||||
|
||||
namespace {
|
||||
|
||||
constexpr int kPalmSize = 192, kHandSize = 224;
|
||||
const int kRoiLandmarks[] = {0, 1, 2, 3, 5, 6, 9, 10, 13, 14, 17, 18};
|
||||
|
||||
// 2x3 affine taking crop pixels (0..out) to image pixels.
|
||||
void crop_matrix(V2 center, double size, double rotation, int out, float tm[6]) {
|
||||
const double c = std::cos(rotation), s = std::sin(rotation), k = size / out;
|
||||
tm[0] = float(c * k), tm[1] = float(-s * k), tm[3] = float(s * k), tm[4] = float(c * k);
|
||||
tm[2] = float(center[0] - (tm[0] + tm[1]) * out / 2.0);
|
||||
tm[5] = float(center[1] - (tm[3] + tm[4]) * out / 2.0);
|
||||
}
|
||||
|
||||
V2 to_image(const float tm[6], double x, double y) {
|
||||
return {tm[0] * x + tm[1] * y + tm[2], tm[3] * x + tm[4] * y + tm[5]};
|
||||
}
|
||||
|
||||
// OpenCV's CLAHE (4x4 tiles) on a square crop, in place.
|
||||
void clahe(uint8_t *img, int n, double clip_limit) {
|
||||
constexpr int kTiles = 4;
|
||||
const int ts = n / kTiles, area = ts * ts;
|
||||
const int clip = std::max(1, int(clip_limit * area / 256));
|
||||
uint8_t lut[kTiles][kTiles][256];
|
||||
for (int ty = 0; ty < kTiles; ++ty)
|
||||
for (int tx = 0; tx < kTiles; ++tx) {
|
||||
int hist[256] = {};
|
||||
for (int y = ty * ts; y < (ty + 1) * ts; ++y)
|
||||
for (int x = tx * ts; x < (tx + 1) * ts; ++x) ++hist[img[y * n + x]];
|
||||
int excess = 0;
|
||||
for (int &h : hist)
|
||||
if (h > clip) excess += h - clip, h = clip;
|
||||
const int add = excess / 256, residual = excess - add * 256;
|
||||
for (int i = 0; i < 256; ++i) hist[i] += add + (i < residual ? 1 : 0);
|
||||
int sum = 0;
|
||||
const float scale = 255.f / area;
|
||||
for (int i = 0; i < 256; ++i) {
|
||||
sum += hist[i];
|
||||
lut[ty][tx][i] = uint8_t(std::min(255, int(sum * scale + 0.5f)));
|
||||
}
|
||||
}
|
||||
std::vector<uint8_t> out(size_t(n) * n);
|
||||
for (int y = 0; y < n; ++y) {
|
||||
const float fy = (y + 0.5f) / ts - 0.5f;
|
||||
const int y0 = std::clamp(int(std::floor(fy)), 0, kTiles - 1), y1 = std::min(y0 + 1, kTiles - 1);
|
||||
const float wy = std::clamp(fy - y0, 0.f, 1.f);
|
||||
for (int x = 0; x < n; ++x) {
|
||||
const float fx = (x + 0.5f) / ts - 0.5f;
|
||||
const int x0 = std::clamp(int(std::floor(fx)), 0, kTiles - 1), x1 = std::min(x0 + 1, kTiles - 1);
|
||||
const float wx = std::clamp(fx - x0, 0.f, 1.f);
|
||||
const uint8_t v = img[y * n + x];
|
||||
const float top = lut[y0][x0][v] * (1 - wx) + lut[y0][x1][v] * wx;
|
||||
const float bot = lut[y1][x0][v] * (1 - wx) + lut[y1][x1][v] * wx;
|
||||
out[size_t(y) * n + x] = uint8_t(top * (1 - wy) + bot * wy + 0.5f);
|
||||
}
|
||||
}
|
||||
std::copy(out.begin(), out.end(), img);
|
||||
}
|
||||
|
||||
// Linear stretch of the 1st..99th percentile to 0..255, in place.
|
||||
void stretch(uint8_t *img, int n) {
|
||||
int hist[256] = {};
|
||||
const int total = n * n;
|
||||
for (int i = 0; i < total; ++i) ++hist[img[i]];
|
||||
int lo = 0, hi = 255, acc = 0;
|
||||
for (int v = 0; v < 256; ++v)
|
||||
if ((acc += hist[v]) > total / 100) { lo = v; break; }
|
||||
acc = 0;
|
||||
for (int v = 255; v >= 0; --v)
|
||||
if ((acc += hist[v]) > total / 100) { hi = v; break; }
|
||||
if (hi <= lo) return;
|
||||
for (int i = 0; i < total; ++i) img[i] = uint8_t(std::clamp((img[i] - lo) * 255 / (hi - lo), 0, 255));
|
||||
}
|
||||
|
||||
// A crop as the models' input: RGB (the mono plane three times), 0..1.
|
||||
ncnn::Mat crop(const Image &img, const float tm[6], int n, const Contrast &contrast) {
|
||||
std::vector<uint8_t> patch(size_t(n) * n);
|
||||
ncnn::warpaffine_bilinear_c1(img.data, img.width, img.height, img.stride, patch.data(), n, n, n, tm, 0, 0);
|
||||
if (contrast.mode == Contrast::Clahe) clahe(patch.data(), n, contrast.clip);
|
||||
else if (contrast.mode == Contrast::Stretch) stretch(patch.data(), n);
|
||||
ncnn::Mat m = ncnn::Mat::from_pixels(patch.data(), ncnn::Mat::PIXEL_GRAY2RGB, n, n);
|
||||
const float norm[3] = {1 / 255.f, 1 / 255.f, 1 / 255.f};
|
||||
m.substract_mean_normalize(nullptr, norm);
|
||||
return m;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
bool Contrast::parse(const std::string &s, Contrast &out) {
|
||||
if (s == "none") return out.mode = None, true;
|
||||
if (s == "stretch") return out.mode = Stretch, true;
|
||||
if (s.rfind("clahe", 0) == 0) {
|
||||
out.mode = Clahe;
|
||||
out.clip = s.size() > 6 && s[5] == ':' ? std::atof(s.c_str() + 6) : 2.0;
|
||||
return out.clip > 0;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
bool Contrast::parse_pair(const std::string &s, Contrast &palm, Contrast &hand) {
|
||||
const size_t slash = s.find('/');
|
||||
if (slash == std::string::npos) return parse(s, palm) && parse(s, hand);
|
||||
return parse(s.substr(0, slash), palm) && parse(s.substr(slash + 1), hand);
|
||||
}
|
||||
|
||||
namespace {
|
||||
|
||||
bool load_net(ncnn::Net &net, const std::string &base, std::string &err) {
|
||||
net.opt.num_threads = 1;
|
||||
net.opt.use_vulkan_compute = false;
|
||||
net.opt.use_fp16_packed = net.opt.use_fp16_storage = net.opt.use_fp16_arithmetic = true;
|
||||
if (net.load_param((base + ".param").c_str()) || net.load_model((base + ".bin").c_str())) {
|
||||
err = "can't load " + base + ".param/.bin";
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
Roi Palm::roi() const {
|
||||
const V2 a = kp[0], b = kp[2];
|
||||
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(b[1] - a[1]), b[0] - a[0]));
|
||||
const double h = size[1];
|
||||
const V2 shift{-h * -0.5 * std::sin(rot), h * -0.5 * std::cos(rot)};
|
||||
return {center + shift, std::max(size[0], size[1]) * 2.6, rot};
|
||||
}
|
||||
|
||||
Roi roi_from_points(const V2 *p) {
|
||||
const V2 w = p[0];
|
||||
V2 m = (p[5] + p[13]) * 0.5;
|
||||
m = (m + p[9]) * 0.5;
|
||||
const double rot = wrap_angle(M_PI / 2 - std::atan2(-(m[1] - w[1]), m[0] - w[0]));
|
||||
V2 lo{1e9, 1e9}, hi{-1e9, -1e9};
|
||||
for (int i : kRoiLandmarks)
|
||||
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[i][k]), hi[k] = std::max(hi[k], p[i][k]);
|
||||
V2 center = (lo + hi) * 0.5;
|
||||
const double c = std::cos(-rot), s = std::sin(-rot);
|
||||
V2 qlo{1e9, 1e9}, qhi{-1e9, -1e9};
|
||||
for (int i : kRoiLandmarks) {
|
||||
const V2 d = p[i] - center;
|
||||
const V2 q{d[0] * c - d[1] * s, d[0] * s + d[1] * c};
|
||||
for (int k = 0; k < 2; ++k) qlo[k] = std::min(qlo[k], q[k]), qhi[k] = std::max(qhi[k], q[k]);
|
||||
}
|
||||
const V2 mid = (qlo + qhi) * 0.5;
|
||||
const double c2 = std::cos(rot), s2 = std::sin(rot);
|
||||
center = center + V2{mid[0] * c2 - mid[1] * s2, mid[0] * s2 + mid[1] * c2};
|
||||
const double w2 = qhi[0] - qlo[0], h2 = qhi[1] - qlo[1];
|
||||
center = center + V2{-h2 * -0.1 * s2, h2 * -0.1 * c2};
|
||||
return {center, std::max(w2, h2) * 2.0, rot};
|
||||
}
|
||||
|
||||
Roi Landmarks::next_roi() const { return roi_from_points(pts); }
|
||||
|
||||
bool Nets::load(const std::string &dir, bool int8, std::string &err) {
|
||||
const std::string suffix = int8 ? "-int8.ncnn" : ".ncnn";
|
||||
if (!load_net(palm_, dir + "/palm" + suffix, err) || !load_net(hand_, dir + "/hand" + suffix, err)) return false;
|
||||
// SSD anchors of palm_detection_full: strides 8 (2 per cell) and 16 (6 per cell)
|
||||
for (auto [stride, per] : {std::pair{8, 2}, std::pair{16, 6}}) {
|
||||
const int n = kPalmSize / stride;
|
||||
for (int y = 0; y < n; ++y)
|
||||
for (int x = 0; x < n; ++x)
|
||||
for (int k = 0; k < per; ++k) anchors_.push_back({(x + 0.5) / n * kPalmSize, (y + 0.5) / n * kPalmSize});
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
std::vector<Palm> Nets::palms(const Image &img, V2 center, double size, double rotation) const {
|
||||
float tm[6];
|
||||
crop_matrix(center, size, rotation, kPalmSize, tm);
|
||||
ncnn::Extractor ex = palm_.create_extractor();
|
||||
ex.input("in0", crop(img, tm, kPalmSize, palm_contrast_));
|
||||
ncnn::Mat boxes, scores;
|
||||
ex.extract("out0", boxes);
|
||||
ex.extract("out1", scores);
|
||||
const float *raw = boxes, *logit = scores;
|
||||
const int n = int(anchors_.size());
|
||||
const float min_logit = std::log(0.5f / 0.5f); // score 0.5
|
||||
struct Cand { V2 c, s; V2 kp[7]; double score; };
|
||||
std::vector<Cand> cand;
|
||||
for (int i = 0; i < n; ++i) {
|
||||
if (logit[i] <= min_logit) continue;
|
||||
const float *r = raw + i * 18;
|
||||
Cand c;
|
||||
c.c = {r[0] + anchors_[i][0], r[1] + anchors_[i][1]};
|
||||
c.s = {r[2], r[3]};
|
||||
for (int k = 0; k < 7; ++k) c.kp[k] = {r[4 + 2 * k] + anchors_[i][0], r[5 + 2 * k] + anchors_[i][1]};
|
||||
c.score = 1 / (1 + std::exp(-std::clamp(double(logit[i]), -100.0, 100.0)));
|
||||
cand.push_back(c);
|
||||
}
|
||||
// MediaPipe's weighted NMS: overlapping boxes are averaged, weighted by score
|
||||
std::sort(cand.begin(), cand.end(), [](const Cand &a, const Cand &b) { return a.score > b.score; });
|
||||
std::vector<bool> used(cand.size());
|
||||
std::vector<Palm> out;
|
||||
for (size_t i = 0; i < cand.size(); ++i) {
|
||||
if (used[i]) continue;
|
||||
double wsum = 0;
|
||||
Cand acc{};
|
||||
for (size_t j = i; j < cand.size(); ++j) {
|
||||
if (used[j]) continue;
|
||||
const double ix = std::max(0.0, std::min(cand[i].c[0] + cand[i].s[0] / 2, cand[j].c[0] + cand[j].s[0] / 2) -
|
||||
std::max(cand[i].c[0] - cand[i].s[0] / 2, cand[j].c[0] - cand[j].s[0] / 2));
|
||||
const double iy = std::max(0.0, std::min(cand[i].c[1] + cand[i].s[1] / 2, cand[j].c[1] + cand[j].s[1] / 2) -
|
||||
std::max(cand[i].c[1] - cand[i].s[1] / 2, cand[j].c[1] - cand[j].s[1] / 2));
|
||||
const double inter = ix * iy;
|
||||
const double uni = cand[i].s[0] * cand[i].s[1] + cand[j].s[0] * cand[j].s[1] - inter;
|
||||
if (j != i && inter / (uni + 1e-9) <= 0.3) continue;
|
||||
used[j] = true;
|
||||
const double w = cand[j].score;
|
||||
wsum += w;
|
||||
acc.c = acc.c + cand[j].c * w;
|
||||
acc.s = acc.s + cand[j].s * w;
|
||||
for (int k = 0; k < 7; ++k) acc.kp[k] = acc.kp[k] + cand[j].kp[k] * w;
|
||||
}
|
||||
Palm p;
|
||||
const V2 c = acc.c * (1 / wsum);
|
||||
p.center = to_image(tm, c[0], c[1]);
|
||||
p.size = acc.s * (1 / wsum * size / kPalmSize);
|
||||
for (int k = 0; k < 7; ++k) {
|
||||
const V2 q = acc.kp[k] * (1 / wsum);
|
||||
p.kp[k] = to_image(tm, q[0], q[1]);
|
||||
}
|
||||
p.score = cand[i].score;
|
||||
out.push_back(p);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
Landmarks Nets::landmarks(const Image &img, const Roi &roi) const {
|
||||
float tm[6];
|
||||
crop_matrix(roi.center, roi.size, roi.rotation, kHandSize, tm);
|
||||
ncnn::Extractor ex = hand_.create_extractor();
|
||||
ex.input("in0", crop(img, tm, kHandSize, hand_contrast_));
|
||||
ncnn::Mat screen, presence, right, world;
|
||||
ex.extract("out0", screen);
|
||||
ex.extract("out1", presence);
|
||||
ex.extract("out2", right);
|
||||
ex.extract("out3", world);
|
||||
Landmarks lm;
|
||||
const float *s = screen, *w = world;
|
||||
for (int i = 0; i < 21; ++i) {
|
||||
lm.pts[i] = to_image(tm, s[3 * i], s[3 * i + 1]);
|
||||
for (int k = 0; k < 3; ++k) lm.world[i][k] = w[3 * i + k];
|
||||
}
|
||||
lm.presence = presence[0];
|
||||
lm.right = right[0];
|
||||
return lm;
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
// MediaPipe's palm detector and hand landmark model on ncnn (see tracker/models.py for
|
||||
// the conventions). A crop is a square region of a camera image: centre and size in
|
||||
// pixels, and a rotation that turns the crop's "up" toward the image direction
|
||||
// (sin r, -cos r). Crops are contrast-equalized (CLAHE) before the models see them.
|
||||
// Everything here may run on several threads at once.
|
||||
#pragma once
|
||||
|
||||
#include "geom.h"
|
||||
|
||||
#include <net.h>
|
||||
|
||||
#include <cstdint>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
struct Image {
|
||||
const uint8_t *data = nullptr;
|
||||
int width = 0, height = 0, stride = 0;
|
||||
};
|
||||
|
||||
struct Roi {
|
||||
V2 center{};
|
||||
double size = 0, rotation = 0;
|
||||
};
|
||||
|
||||
struct Palm {
|
||||
V2 center{}, size{};
|
||||
V2 kp[7]{};
|
||||
double score = 0;
|
||||
Roi roi() const; // MediaPipe's hand crop for this palm
|
||||
};
|
||||
|
||||
struct Landmarks {
|
||||
V2 pts[21]{}; // image pixels
|
||||
double world[21][3]{}; // MediaPipe's metric landmarks, hand-centred
|
||||
double presence = 0, right = 0;
|
||||
Roi next_roi() const; // MediaPipe's crop to track the hand in the next frame
|
||||
};
|
||||
|
||||
Roi roi_from_points(const V2 *pts21);
|
||||
|
||||
// How crops are contrast-equalized before the models see them.
|
||||
struct Contrast {
|
||||
enum Mode { Clahe, None, Stretch } mode = Clahe;
|
||||
double clip = 2.0; // Clahe: OpenCV's clip limit (4x4 tiles)
|
||||
// "clahe:2", "none", "stretch" (1st..99th percentile to 0..255)
|
||||
static bool parse(const std::string &s, Contrast &out);
|
||||
// "PALM/HAND" (each as above), or one for both
|
||||
static bool parse_pair(const std::string &s, Contrast &palm, Contrast &hand);
|
||||
};
|
||||
|
||||
class Nets {
|
||||
public:
|
||||
// Loads <dir>/palm.ncnn.* and <dir>/hand.ncnn.*, or the -int8 variants.
|
||||
bool load(const std::string &dir, bool int8, std::string &err);
|
||||
std::vector<Palm> palms(const Image &img, V2 center, double size, double rotation) const;
|
||||
Landmarks landmarks(const Image &img, const Roi &roi) const;
|
||||
// Before any palms()/landmarks(): how the palm search's and the landmark model's crops
|
||||
// are equalized.
|
||||
void set_contrast(const Contrast &palm, const Contrast &hand) { palm_contrast_ = palm, hand_contrast_ = hand; }
|
||||
|
||||
private:
|
||||
Contrast palm_contrast_, hand_contrast_;
|
||||
ncnn::Net palm_, hand_;
|
||||
std::vector<V2> anchors_;
|
||||
};
|
||||
@@ -0,0 +1,75 @@
|
||||
// Check the C++ model code against tracker/models.py on a recorded frame:
|
||||
// nettest MODELS_DIR FRAME.pgm cx cy size rotation [--int8]
|
||||
// Runs the palm detector on that crop, then the landmark model on each palm's ROI, and
|
||||
// prints what they found; tools/nettest_compare.py runs the Python side on the same input.
|
||||
#include "nets.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
#include <fstream>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
static bool read_pgm(const char *path, std::vector<uint8_t> &px, int &w, int &h) {
|
||||
std::ifstream f(path, std::ios::binary);
|
||||
std::string magic;
|
||||
int maxv;
|
||||
if (!(f >> magic >> w >> h >> maxv) || magic != "P5") return false;
|
||||
f.get();
|
||||
px.resize(size_t(w) * h);
|
||||
return bool(f.read(reinterpret_cast<char *>(px.data()), px.size()));
|
||||
}
|
||||
|
||||
int main(int argc, char **argv) {
|
||||
if (argc < 7) return std::fprintf(stderr, "usage: nettest MODELS FRAME.pgm cx cy size rotation [--int8] [--bench N]\n"), 1;
|
||||
bool int8 = false;
|
||||
int bench = 0;
|
||||
for (int i = 7; i < argc; ++i) {
|
||||
if (!std::strcmp(argv[i], "--int8")) int8 = true;
|
||||
else if (!std::strcmp(argv[i], "--bench") && i + 1 < argc) bench = std::atoi(argv[++i]);
|
||||
}
|
||||
Nets nets;
|
||||
std::string err;
|
||||
if (!nets.load(argv[1], int8, err)) return std::fprintf(stderr, "%s\n", err.c_str()), 1;
|
||||
std::vector<uint8_t> px;
|
||||
int w, h;
|
||||
if (!read_pgm(argv[2], px, w, h)) return std::fprintf(stderr, "can't read %s\n", argv[2]), 1;
|
||||
const Image img{px.data(), w, h, w};
|
||||
const V2 c{std::atof(argv[3]), std::atof(argv[4])};
|
||||
if (bench > 0) { // steady-state timing: warm up, then the median of N calls each
|
||||
const Roi roi{c, std::atof(argv[5]), std::atof(argv[6])};
|
||||
auto time = [&](auto fn) {
|
||||
std::vector<double> t;
|
||||
for (int i = 0; i < bench + 5; ++i) {
|
||||
const auto t0 = std::chrono::steady_clock::now();
|
||||
fn();
|
||||
if (i >= 5) t.push_back(std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t0).count());
|
||||
}
|
||||
std::sort(t.begin(), t.end());
|
||||
return t[t.size() / 2];
|
||||
};
|
||||
const double pm = time([&] { nets.palms(img, c, roi.size, roi.rotation); });
|
||||
const double hm = time([&] { nets.landmarks(img, roi); });
|
||||
std::printf("bench%s: palm %.2f ms, hand %.2f ms (median of %d, one thread)\n", int8 ? " int8" : "", pm, hm, bench);
|
||||
return 0;
|
||||
}
|
||||
auto t0 = std::chrono::steady_clock::now();
|
||||
const auto palms = nets.palms(img, c, std::atof(argv[5]), std::atof(argv[6]));
|
||||
const double palm_ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t0).count();
|
||||
std::printf("palm_ms %.2f\n", palm_ms);
|
||||
for (const Palm &p : palms) {
|
||||
const Roi r = p.roi();
|
||||
std::printf("palm %.3f center %.1f %.1f roi %.1f %.1f %.1f %.4f\n", p.score, p.center[0], p.center[1],
|
||||
r.center[0], r.center[1], r.size, r.rotation);
|
||||
t0 = std::chrono::steady_clock::now();
|
||||
const Landmarks lm = nets.landmarks(img, r);
|
||||
const double ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t0).count();
|
||||
std::printf("hand_ms %.2f presence %.3f right %.3f\n", ms, lm.presence, lm.right);
|
||||
std::printf("pts");
|
||||
for (const V2 &q : lm.pts) std::printf(" %.1f %.1f", q[0], q[1]);
|
||||
std::printf("\n");
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,153 @@
|
||||
#include "pinch.h"
|
||||
|
||||
#include "io.h"
|
||||
|
||||
#include <fcntl.h>
|
||||
#include <sys/mman.h>
|
||||
#include <sys/stat.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cerrno>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
|
||||
namespace {
|
||||
|
||||
constexpr int kThumbTip = 4, kIndexTip = 8;
|
||||
|
||||
void put3(float out[3], V3 v) {
|
||||
for (int k = 0; k < 3; ++k) out[k] = float(v[k]);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
void Pinch::update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns) {
|
||||
events.clear();
|
||||
for (int s = 0; s < 2; ++s) {
|
||||
fh_pinch_t &o = side_[s];
|
||||
const bool down = o.flags & FH_PINCH_DOWN;
|
||||
// the hand: while down, the one the pinch began on; else the best tracked hand of this side
|
||||
const Hand *h = nullptr;
|
||||
for (const Hand *c : hands) {
|
||||
if (down ? c->id != follow_[s] : c->right() != (s == 1)) continue;
|
||||
if (!h || c->frames > h->frames) h = c;
|
||||
}
|
||||
world_d[s] = tri_d[s] = -1;
|
||||
if (!h) {
|
||||
o.flags &= ~FH_PINCH_TRACKED;
|
||||
if (down && (t_ns - seen_ns_[s]) / 1e9 > p_.grace_s) end(s, t_ns, true);
|
||||
continue;
|
||||
}
|
||||
seen_ns_[s] = t_ns;
|
||||
tri_d[s] = norm(h->pts[kThumbTip] - h->pts[kIndexTip]);
|
||||
double sum = 0;
|
||||
int n = 0;
|
||||
for (const Seen &v : views) {
|
||||
if (v.hand != h->id) continue;
|
||||
const V3 a{v.lm.world[kThumbTip][0], v.lm.world[kThumbTip][1], v.lm.world[kThumbTip][2]};
|
||||
const V3 b{v.lm.world[kIndexTip][0], v.lm.world[kIndexTip][1], v.lm.world[kIndexTip][2]};
|
||||
sum += norm(a - b), ++n;
|
||||
}
|
||||
if (n) world_d[s] = sum / n * h->scale;
|
||||
const double d = p_.triangulated || world_d[s] < 0 ? tri_d[s] : world_d[s];
|
||||
const V3 point = (h->smooth[kThumbTip] + h->smooth[kIndexTip]) * 0.5;
|
||||
o.flags |= FH_PINCH_TRACKED;
|
||||
o.hand_id = uint32_t(h->id);
|
||||
o.distance = float(d);
|
||||
o.strength = float(std::clamp((p_.end_m - d) / (p_.end_m - p_.begin_m), 0.0, 1.0));
|
||||
put3(o.point, point);
|
||||
if (!down) {
|
||||
if (d < p_.begin_m) {
|
||||
o.flags = (o.flags | FH_PINCH_DOWN) & ~FH_PINCH_LOST;
|
||||
++o.begins;
|
||||
o.begin_ns = uint64_t(t_ns);
|
||||
put3(o.begin_point, point);
|
||||
follow_[s] = h->id;
|
||||
open_frames_[s] = 0;
|
||||
events.push_back({s, "begin", t_ns, d, point});
|
||||
}
|
||||
} else if (d > p_.end_m) {
|
||||
if (++open_frames_[s] >= p_.end_frames) end(s, t_ns, false);
|
||||
} else {
|
||||
open_frames_[s] = 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void Pinch::end(int s, int64_t t_ns, bool lost) {
|
||||
fh_pinch_t &o = side_[s];
|
||||
o.flags = (o.flags & ~FH_PINCH_DOWN) | (lost ? FH_PINCH_LOST : 0);
|
||||
++o.ends;
|
||||
o.end_ns = uint64_t(t_ns);
|
||||
follow_[s] = 0;
|
||||
events.push_back({s, lost ? "lost" : "end", t_ns, o.distance, {o.point[0], o.point[1], o.point[2]}});
|
||||
}
|
||||
|
||||
void Pinch::release(int64_t t_ns) {
|
||||
events.clear();
|
||||
for (int s = 0; s < 2; ++s) {
|
||||
side_[s].flags &= ~FH_PINCH_TRACKED;
|
||||
if (side_[s].flags & FH_PINCH_DOWN) end(s, t_ns, true);
|
||||
}
|
||||
}
|
||||
|
||||
bool Pinch::engaged() const {
|
||||
for (const fh_pinch_t &o : side_)
|
||||
if ((o.flags & FH_PINCH_DOWN) || ((o.flags & FH_PINCH_TRACKED) && o.strength > 0.3f)) return true;
|
||||
return false;
|
||||
}
|
||||
|
||||
bool GesturePublisher::open(const Pinch &pinch, std::string &err) {
|
||||
const char *run = std::getenv("XDG_RUNTIME_DIR");
|
||||
const std::string dir = std::string(run ? run : "/run/user/" + std::to_string(getuid())) + "/frame-hands";
|
||||
mkdir(dir.c_str(), 0700);
|
||||
chmod(dir.c_str(), 0700);
|
||||
const std::string path = dir + "/gestures";
|
||||
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
|
||||
if (fd < 0 || ftruncate(fd, sizeof(fh_gestures_t)) < 0) return err = path + ": " + std::strerror(errno), false;
|
||||
void *m = mmap(nullptr, sizeof(fh_gestures_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return err = path + ": can't map it", false;
|
||||
out_ = static_cast<fh_gestures_t *>(m);
|
||||
// keep the counters a previous tracker left, so a reader doesn't see them jump back
|
||||
const bool ours = !std::memcmp(out_->magic, FH_GESTURES_MAGIC, 8) && out_->version == FH_GESTURES_VERSION;
|
||||
if (!ours) {
|
||||
std::memset(out_, 0, sizeof *out_);
|
||||
std::memcpy(out_->magic, FH_GESTURES_MAGIC, 8);
|
||||
out_->version = FH_GESTURES_VERSION;
|
||||
out_->size = sizeof(fh_gestures_t);
|
||||
}
|
||||
seq_ = out_->seq / 2 + 1;
|
||||
// a pinch the last tracker left down (it crashed) is over: count its end, as lost
|
||||
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
for (fh_pinch_t &o : out_->pinch)
|
||||
if (o.begins != o.ends) {
|
||||
o.ends = o.begins;
|
||||
o.end_ns = mono_ns();
|
||||
o.flags = (o.flags & ~FH_PINCH_DOWN) | FH_PINCH_LOST;
|
||||
}
|
||||
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||
out_->begin_m = float(pinch.params().begin_m);
|
||||
out_->end_m = float(pinch.params().end_m);
|
||||
return true;
|
||||
}
|
||||
|
||||
void GesturePublisher::write(const Pinch &pinch, uint64_t capture_ns) {
|
||||
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
for (int s = 0; s < 2; ++s) {
|
||||
// counters carry on from what's in the file (a restarted tracker starts its own at 0)
|
||||
const fh_pinch_t &in = pinch.side(s);
|
||||
fh_pinch_t &o = out_->pinch[s];
|
||||
const uint32_t base_b = o.begins - last_begins_[s], base_e = o.ends - last_ends_[s];
|
||||
o = in;
|
||||
o.begins = base_b + in.begins;
|
||||
o.ends = base_e + in.ends;
|
||||
last_begins_[s] = in.begins, last_ends_[s] = in.ends;
|
||||
}
|
||||
out_->capture_ns = capture_ns;
|
||||
out_->publish_ns = mono_ns();
|
||||
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||
}
|
||||
@@ -0,0 +1,71 @@
|
||||
// Pinch detection for input: look at something and pinch to click, pinch and move to drag.
|
||||
// Per side, from the tracker's hands after each step; published as fh_gestures.h.
|
||||
#pragma once
|
||||
|
||||
#include "tracker.h"
|
||||
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
extern "C" {
|
||||
#include "../include/fh_gestures.h"
|
||||
}
|
||||
|
||||
struct PinchParams {
|
||||
double begin_m = 0.020; // thumb and index tips closer than this: the pinch begins
|
||||
double end_m = 0.035; // further apart than this: it ends (the gap keeps it from flickering)
|
||||
int end_frames = 2; // processed frames in a row past end_m before it ends, so one
|
||||
// noisy frame doesn't drop a drag
|
||||
double grace_s = 0.25; // a pinching hand lost this long ends its pinch (FH_PINCH_LOST)
|
||||
// Where the distance comes from: MediaPipe's world landmarks (the model's own 3D hand
|
||||
// pose, averaged over the hand's views, at the user's hand size), or the tracker's
|
||||
// triangulated tips. The model's pose should hold up better when the fingers hide each
|
||||
// other; tomorrow's recordings will tell.
|
||||
bool triangulated = false;
|
||||
};
|
||||
|
||||
class Pinch {
|
||||
public:
|
||||
explicit Pinch(const PinchParams &p = {}) : p_(p) {}
|
||||
const PinchParams ¶ms() const { return p_; }
|
||||
// After each processed set: the hands out of Tracker::step, the tracker's views (for
|
||||
// the world landmarks) and the capture time.
|
||||
void update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns);
|
||||
// Ends any pinch that's down (as lost), e.g. when the tracker stops.
|
||||
void release(int64_t t_ns);
|
||||
const fh_pinch_t &side(int s) const { return side_[s]; } // 0 left, 1 right
|
||||
// A pinch is down or closing: worth tracking at the full rate.
|
||||
bool engaged() const;
|
||||
|
||||
// What changed in the last update, for logs.
|
||||
struct Event {
|
||||
int side;
|
||||
const char *what; // "begin", "end", "lost"
|
||||
int64_t t_ns;
|
||||
double distance;
|
||||
V3 point;
|
||||
};
|
||||
std::vector<Event> events;
|
||||
// Both distance measures for the last update, per side (-1: no hand), for logs.
|
||||
double world_d[2] = {-1, -1}, tri_d[2] = {-1, -1};
|
||||
|
||||
private:
|
||||
void end(int s, int64_t t_ns, bool lost);
|
||||
PinchParams p_;
|
||||
fh_pinch_t side_[2]{};
|
||||
int follow_[2] = {0, 0}; // the hand id a pinch follows while down
|
||||
int open_frames_[2] = {0, 0};
|
||||
int64_t seen_ns_[2] = {0, 0};
|
||||
};
|
||||
|
||||
// Writes $XDG_RUNTIME_DIR/frame-hands/gestures.
|
||||
class GesturePublisher {
|
||||
public:
|
||||
bool open(const Pinch &pinch, std::string &err);
|
||||
void write(const Pinch &pinch, uint64_t capture_ns);
|
||||
|
||||
private:
|
||||
fh_gestures_t *out_ = nullptr;
|
||||
uint64_t seq_ = 0;
|
||||
uint32_t last_begins_[2] = {0, 0}, last_ends_[2] = {0, 0}; // Pinch's counters last written
|
||||
};
|
||||
@@ -0,0 +1,101 @@
|
||||
#include "record.h"
|
||||
|
||||
#include <sys/stat.h>
|
||||
|
||||
#include <cerrno>
|
||||
#include <cstring>
|
||||
|
||||
namespace {
|
||||
constexpr size_t kMaxQueued = 48; // about 130 MB of sets
|
||||
}
|
||||
|
||||
Recorder::~Recorder() {
|
||||
if (!f_) return;
|
||||
{
|
||||
std::lock_guard<std::mutex> l(mu_);
|
||||
stop_ = true;
|
||||
}
|
||||
wake_.notify_all();
|
||||
thread_.join();
|
||||
std::fclose(f_);
|
||||
}
|
||||
|
||||
bool Recorder::open(const std::string &dir, std::string &err) {
|
||||
if (mkdir(dir.c_str(), 0755) < 0 && errno != EEXIST) return err = dir + ": " + std::strerror(errno), false;
|
||||
const std::string path = dir + "/sets.bin";
|
||||
f_ = std::fopen(path.c_str(), "wbx"); // never overwrite a recording
|
||||
if (!f_) return err = path + ": " + std::strerror(errno), false;
|
||||
thread_ = std::thread(&Recorder::loop, this);
|
||||
return true;
|
||||
}
|
||||
|
||||
void Recorder::add(const std::vector<SetFrame> &frames) {
|
||||
size_t bytes = sizeof(fh_set_hdr_t) + frames.size() * sizeof(fh_set_cam_t);
|
||||
for (const SetFrame &s : frames) bytes += size_t(s.width) * s.height;
|
||||
std::vector<uint8_t> rec(bytes);
|
||||
fh_set_hdr_t h{};
|
||||
std::memcpy(h.magic, FH_SET_MAGIC, 8);
|
||||
h.ncams = uint32_t(frames.size());
|
||||
h.bytes = uint32_t(bytes);
|
||||
std::memcpy(rec.data(), &h, sizeof h);
|
||||
uint8_t *p = rec.data() + sizeof h;
|
||||
for (const SetFrame &s : frames) {
|
||||
fh_set_cam_t c{};
|
||||
std::strncpy(c.name, s.name.c_str(), sizeof c.name - 1);
|
||||
c.width = s.width, c.height = s.height, c.capture_ns = s.capture_ns, c.dqbuf_ns = s.dqbuf_ns;
|
||||
std::memcpy(p, &c, sizeof c);
|
||||
p += sizeof c;
|
||||
}
|
||||
for (const SetFrame &s : frames) {
|
||||
std::memcpy(p, s.px, size_t(s.width) * s.height);
|
||||
p += size_t(s.width) * s.height;
|
||||
}
|
||||
{
|
||||
std::lock_guard<std::mutex> l(mu_);
|
||||
if (queue_.size() >= kMaxQueued) {
|
||||
++dropped_;
|
||||
return;
|
||||
}
|
||||
queue_.push_back(std::move(rec));
|
||||
}
|
||||
wake_.notify_one();
|
||||
}
|
||||
|
||||
void Recorder::loop() {
|
||||
std::unique_lock<std::mutex> l(mu_);
|
||||
for (;;) {
|
||||
wake_.wait(l, [&] { return stop_ || !queue_.empty(); });
|
||||
if (queue_.empty()) return; // stopping, and everything is written
|
||||
std::vector<uint8_t> rec = std::move(queue_.front());
|
||||
queue_.pop_front();
|
||||
l.unlock();
|
||||
const bool ok = std::fwrite(rec.data(), 1, rec.size(), f_) == rec.size();
|
||||
l.lock();
|
||||
ok ? ++written_ : ++dropped_;
|
||||
}
|
||||
}
|
||||
|
||||
SetReader::~SetReader() {
|
||||
if (f_) std::fclose(f_);
|
||||
}
|
||||
|
||||
bool SetReader::open(const std::string &dir, std::string &err) {
|
||||
const std::string path = dir + "/sets.bin";
|
||||
f_ = std::fopen(path.c_str(), "rb");
|
||||
return f_ ? true : (err = path + ": " + std::strerror(errno), false);
|
||||
}
|
||||
|
||||
bool SetReader::next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels) {
|
||||
fh_set_hdr_t h;
|
||||
if (std::fread(&h, sizeof h, 1, f_) != 1 || std::memcmp(h.magic, FH_SET_MAGIC, 8) || h.ncams == 0 || h.ncams > 16)
|
||||
return false;
|
||||
cams.resize(h.ncams);
|
||||
if (std::fread(cams.data(), sizeof(fh_set_cam_t), h.ncams, f_) != h.ncams) return false;
|
||||
pixels.resize(h.ncams);
|
||||
for (uint32_t i = 0; i < h.ncams; ++i) {
|
||||
cams[i].name[sizeof cams[i].name - 1] = 0;
|
||||
pixels[i].resize(size_t(cams[i].width) * cams[i].height);
|
||||
if (std::fread(pixels[i].data(), 1, pixels[i].size(), f_) != pixels[i].size()) return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
@@ -0,0 +1,69 @@
|
||||
// Recordings of frame sets, for replaying live sessions through the tracker offline
|
||||
// (fh-replay). A recording is DIR/sets.bin: one record per frame set, each
|
||||
// fh_set_hdr_t, then per camera fh_set_cam_t, then each camera's pixels (w x h, packed)
|
||||
// in the same camera order.
|
||||
#pragma once
|
||||
|
||||
#include <condition_variable>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <deque>
|
||||
#include <mutex>
|
||||
#include <string>
|
||||
#include <thread>
|
||||
#include <vector>
|
||||
|
||||
#define FH_SET_MAGIC "FHSET01"
|
||||
|
||||
struct fh_set_hdr_t {
|
||||
char magic[8];
|
||||
uint32_t ncams;
|
||||
uint32_t bytes; // the whole record, this header included
|
||||
};
|
||||
|
||||
struct fh_set_cam_t {
|
||||
char name[16]; // calibration name, e.g. "slam_left"
|
||||
uint32_t width, height;
|
||||
uint64_t capture_ns; // CLOCK_MONOTONIC_RAW, as the ring has it
|
||||
uint64_t dqbuf_ns; // CLOCK_MONOTONIC
|
||||
};
|
||||
|
||||
struct SetFrame {
|
||||
std::string name;
|
||||
const uint8_t *px;
|
||||
uint32_t width, height;
|
||||
uint64_t capture_ns, dqbuf_ns;
|
||||
};
|
||||
|
||||
// Writes sets on its own thread, so a slow disk never holds up tracking; drops sets
|
||||
// when too many are waiting.
|
||||
class Recorder {
|
||||
public:
|
||||
~Recorder();
|
||||
bool open(const std::string &dir, std::string &err);
|
||||
void add(const std::vector<SetFrame> &frames);
|
||||
size_t written() const { return written_; }
|
||||
size_t dropped() const { return dropped_; }
|
||||
|
||||
private:
|
||||
void loop();
|
||||
FILE *f_ = nullptr;
|
||||
std::thread thread_;
|
||||
std::mutex mu_;
|
||||
std::condition_variable wake_;
|
||||
std::deque<std::vector<uint8_t>> queue_;
|
||||
bool stop_ = false;
|
||||
size_t written_ = 0, dropped_ = 0;
|
||||
};
|
||||
|
||||
// Reads a recording back one set at a time.
|
||||
class SetReader {
|
||||
public:
|
||||
bool open(const std::string &dir, std::string &err);
|
||||
// False at the end (or on a truncated last set).
|
||||
bool next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &pixels);
|
||||
~SetReader();
|
||||
|
||||
private:
|
||||
FILE *f_ = nullptr;
|
||||
};
|
||||
@@ -0,0 +1,324 @@
|
||||
// fh-replay: run a recording (fh-tracker --record) through the tracker offline, with the
|
||||
// live scheduling, and report how well it kept the hands.
|
||||
//
|
||||
// fh-replay DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR]
|
||||
// [--from S] [--to S] [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch)
|
||||
//
|
||||
// --oracle N: every N-th set, also search every tile of every camera (slow), to see
|
||||
// which hands were there to find. Compares that with what the tracker had.
|
||||
// --slow F: the live tracker skips the sets that arrive while it's busy; replay takes
|
||||
// each step's time here times F as the busy time (the headset is busier live).
|
||||
// --cost: instead of timing the steps, charge each round of model calls what it
|
||||
// typically costs live (10 ms landmarks, 18 ms palms): repeatable results.
|
||||
// --timeline: per processed set, a line per hand (time, id, side, views, wrist) and per view
|
||||
// (hand, camera, presence, next crop, set index).
|
||||
// --keep-presence P: landmark presence a tracked view needs to stay (default 0.5, as new ones).
|
||||
// --pinch-begin M, --pinch-end M, --pinch-triangulated: the pinch detector (trackd/pinch.h);
|
||||
// the timeline gets its begin/end/lost events and both distance measures per set.
|
||||
// --cams mono|color|all: which cameras to track with (default mono). color and all need a
|
||||
// recording made with fh-camd --with-color; --color-left NODE (color_video0 or
|
||||
// color_video3) and --color-crop subtract|none say how its calibration maps
|
||||
// (tools/check_color.py).
|
||||
// --contrast: how the palm search's and the landmark model's crops are equalized
|
||||
// (default clahe:2/none, as fh-tracker).
|
||||
// --depth FILE: per processed set, a line per hand for tools/depth_report.py: its views'
|
||||
// cameras, triangulation residual, hand scale, measured and published palm, and
|
||||
// each view's one-view palm (Tracker::single_view at the hand's scale). The
|
||||
// header has each camera's centre and focal length.
|
||||
#include "pinch.h"
|
||||
#include "record.h"
|
||||
#include "tracker.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
#include <map>
|
||||
#include <set>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
namespace {
|
||||
|
||||
struct Track {
|
||||
double first = 0, last = 0;
|
||||
int sets = 0, left = 0;
|
||||
// the last two palm positions (raw, smoothed) and times, for the jitter measure
|
||||
V3 raw[2]{}, sm[2]{};
|
||||
double t[2]{};
|
||||
int line = 0; // updates on the current unbroken run
|
||||
};
|
||||
|
||||
double median(std::vector<double> v) {
|
||||
if (v.empty()) return 0;
|
||||
std::nth_element(v.begin(), v.begin() + v.size() / 2, v.end());
|
||||
return v[v.size() / 2];
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char **argv) {
|
||||
if (argc < 2 || argv[1][0] == '-') {
|
||||
std::printf("usage: %s DIR [--oracle N] [--slow F] [--timeline FILE] [--threads N] [--models DIR] [--from S] [--to S]\n", argv[0]);
|
||||
return 1;
|
||||
}
|
||||
const std::string dir = argv[1];
|
||||
int oracle = 0, threads = 2;
|
||||
double slow = 1.0, from = 0, to = 1e9;
|
||||
bool cost = false;
|
||||
Contrast palm_contrast, hand_contrast{Contrast::None}; // as fh-tracker's
|
||||
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
||||
PinchParams pinch_params;
|
||||
std::string use = "mono", color_left = "color_video0", color_crop = "subtract";
|
||||
std::string timeline, depth, models =std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
|
||||
for (int i = 2; i < argc; ++i) {
|
||||
const std::string a = argv[i];
|
||||
const bool more = i + 1 < argc;
|
||||
if (a == "--oracle" && more) oracle = std::atoi(argv[++i]);
|
||||
else if (a == "--slow" && more) slow = std::atof(argv[++i]);
|
||||
else if (a == "--timeline" && more) timeline = argv[++i];
|
||||
else if (a == "--depth" && more) depth = argv[++i];
|
||||
else if (a == "--threads" && more) threads = std::atoi(argv[++i]);
|
||||
else if (a == "--models" && more) models = argv[++i];
|
||||
else if (a == "--cost") cost = true;
|
||||
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
|
||||
else if (a == "--cams" && more) use = argv[++i];
|
||||
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
|
||||
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
|
||||
else if (a == "--color-left" && more) color_left = argv[++i];
|
||||
else if (a == "--color-crop" && more) color_crop = argv[++i];
|
||||
else if (a == "--contrast" && more) {
|
||||
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
|
||||
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
|
||||
}
|
||||
else if (a == "--from" && more) from = std::atof(argv[++i]);
|
||||
else if (a == "--to" && more) to = std::atof(argv[++i]);
|
||||
else return std::fprintf(stderr, "unknown option %s\n", a.c_str()), 1;
|
||||
}
|
||||
std::string err;
|
||||
std::map<std::string, Camera> calib;
|
||||
Nets nets;
|
||||
SetReader in;
|
||||
if (!load_calibration(calib, err) || !nets.load(models, false, err) || !in.open(dir, err))
|
||||
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
|
||||
nets.set_contrast(palm_contrast, hand_contrast);
|
||||
FILE *tl = timeline.empty() ? nullptr : std::fopen(timeline.c_str(), "w");
|
||||
|
||||
std::vector<fh_set_cam_t> cams;
|
||||
std::vector<std::vector<uint8_t>> px;
|
||||
if (!in.next(cams, px)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
|
||||
if (use != "mono" && use != "color" && use != "all") return std::fprintf(stderr, "--cams mono|color|all\n"), 1;
|
||||
if (use != "mono") {
|
||||
std::vector<std::string> nodes;
|
||||
for (auto &c : cams)
|
||||
if (std::string(c.name).rfind("color_video", 0) == 0) nodes.push_back(c.name);
|
||||
if (nodes.size() != 2) return std::fprintf(stderr, "%s: no color cameras (fh-camd --with-color)\n", dir.c_str()), 1;
|
||||
const std::string right = nodes[0] == color_left ? nodes[1] : nodes[0];
|
||||
if (!load_color_calibration(calib, color_left, right, color_crop == "subtract", 2, err))
|
||||
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
|
||||
}
|
||||
std::map<std::string, Camera> used;
|
||||
for (auto &c : cams) {
|
||||
const bool color = std::string(c.name).rfind("color_", 0) == 0;
|
||||
if (calib.count(c.name) && (use == "all" || color == (use == "color"))) used[c.name] = calib[c.name];
|
||||
}
|
||||
Pool pool(threads, {2, 3, 4});
|
||||
Tracker tracker(used, nets, pool);
|
||||
tracker.set_keep_presence(keep_presence);
|
||||
FILE *dp = depth.empty() ? nullptr : std::fopen(depth.c_str(), "w");
|
||||
if (dp)
|
||||
for (auto &[name, c] : used)
|
||||
std::fprintf(dp, "# cam %s %.4f %.4f %.4f %.1f\n", name.c_str(), c.origin[0], c.origin[1], c.origin[2], c.fx);
|
||||
|
||||
uint64_t t0 = 0, busy_until = 0, next_ns = 0, t_prev = 0;
|
||||
int index = -1; // of the set in the recording
|
||||
int nsets = 0, processed = 0, left = 0, right = 0, both = 0, hist[3] = {};
|
||||
std::map<int, Track> tracks;
|
||||
std::vector<const Hand *> last_out;
|
||||
// oracle: sets where a side's hand was findable, and where the tracker had it then
|
||||
int o_sets = 0, o_left = 0, o_right = 0, o_left_hit = 0, o_right_hit = 0, o_left_extra = 0, o_right_extra = 0;
|
||||
std::map<std::string, int> o_by_cam;
|
||||
double busy_ms = 0;
|
||||
// jitter: how far each update's palm is from a straight line through the last two,
|
||||
// mm (steady motion cancels out; what's left is noise and real acceleration)
|
||||
std::vector<double> jit_raw, jit_sm;
|
||||
int near_face = 0, hand_updates = 0; // published palms within 20 cm of the eyes
|
||||
Pinch pinch(pinch_params);
|
||||
double pinch_begin_ts[2] = {0, 0};
|
||||
std::vector<double> pinch_len[2]; // seconds, per side
|
||||
int pinch_lost = 0;
|
||||
do {
|
||||
std::map<std::string, Image> images;
|
||||
// the set's time: the mono cameras' when they're used (the color ones run on another
|
||||
// clock); color frames can repeat across sets, so a set that doesn't move time on is skipped
|
||||
uint64_t t = UINT64_MAX, t_color = UINT64_MAX;
|
||||
for (size_t i = 0; i < cams.size(); ++i) {
|
||||
if (!used.count(cams[i].name)) continue;
|
||||
images[cams[i].name] = {px[i].data(), int(cams[i].width), int(cams[i].height), int(cams[i].width)};
|
||||
uint64_t &ti = std::string(cams[i].name).rfind("color_", 0) == 0 ? t_color : t;
|
||||
ti = std::min(ti, cams[i].capture_ns);
|
||||
}
|
||||
if (t == UINT64_MAX) t = t_color;
|
||||
if (t <= t_prev) {
|
||||
++index;
|
||||
continue;
|
||||
}
|
||||
t_prev = t;
|
||||
if (!t0) t0 = t;
|
||||
const double ts = (t - t0) / 1e9;
|
||||
++index;
|
||||
if (ts < from) continue;
|
||||
if (ts > to) break;
|
||||
++nsets;
|
||||
|
||||
if (t >= busy_until && t >= next_ns) {
|
||||
const auto w0 = std::chrono::steady_clock::now();
|
||||
const Stats before = tracker.stats;
|
||||
const auto out = tracker.step(images, int64_t(t));
|
||||
double ms = std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - w0).count();
|
||||
if (cost) { // repeatable: rounds of model calls at typical live costs, per thread
|
||||
const int hands = tracker.stats.hand_calls - before.hand_calls, palms = tracker.stats.palm_calls - before.palm_calls;
|
||||
ms = (2 + 10.0 * ((hands + threads - 1) / threads) + 18.0 * ((palms + threads - 1) / threads)) / slow;
|
||||
}
|
||||
busy_ms += ms;
|
||||
busy_until = t + uint64_t(ms * slow * 1e6) + 3'000'000; // + the ring hand-off
|
||||
const std::vector<Seen> seen = tracker.views_now();
|
||||
pinch.update(out, seen, int64_t(t));
|
||||
next_ns = t + uint64_t((std::min(tracker.interval(), pinch.engaged() ? 1 / 30.0 : 1.0) - 0.005) * 1e9);
|
||||
for (const Pinch::Event &e : pinch.events) {
|
||||
if (std::string(e.what) == "begin") pinch_begin_ts[e.side] = ts;
|
||||
else pinch_len[e.side].push_back(ts - pinch_begin_ts[e.side]), pinch_lost += std::string(e.what) == "lost";
|
||||
if (tl) std::fprintf(tl, "%.3f pinch %s %s d %.3f point %+.3f %+.3f %+.3f\n", ts, e.side ? "R" : "L", e.what,
|
||||
e.distance, e.point[0], e.point[1], e.point[2]);
|
||||
}
|
||||
if (tl && (pinch.world_d[0] >= 0 || pinch.world_d[1] >= 0)) // both measures, for choosing one
|
||||
std::fprintf(tl, "%.3f pinchd L world %.3f tri %.3f R world %.3f tri %.3f\n", ts, pinch.world_d[0],
|
||||
pinch.tri_d[0], pinch.world_d[1], pinch.tri_d[1]);
|
||||
++processed;
|
||||
last_out = out;
|
||||
bool l = false, r = false;
|
||||
for (const Hand *h : out) {
|
||||
(h->pts[0][0] < 0 ? l : r) = true;
|
||||
Track &tr = tracks[h->id];
|
||||
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
|
||||
const V3 raw = palm(h->pts), sm = palm(h->smooth);
|
||||
++hand_updates, near_face += norm(sm) < 0.2;
|
||||
if (tr.line && ts - tr.t[0] >= 0.1) tr.line = 0; // a gap: the line starts over
|
||||
if (tr.line >= 2 && tr.t[0] - tr.t[1] > 1e-3) {
|
||||
const double k = (ts - tr.t[0]) / (tr.t[0] - tr.t[1]);
|
||||
jit_raw.push_back(norm(raw - tr.raw[0] - (tr.raw[0] - tr.raw[1]) * k) * 1000);
|
||||
jit_sm.push_back(norm(sm - tr.sm[0] - (tr.sm[0] - tr.sm[1]) * k) * 1000);
|
||||
}
|
||||
tr.raw[1] = tr.raw[0], tr.sm[1] = tr.sm[0], tr.t[1] = tr.t[0];
|
||||
tr.raw[0] = raw, tr.sm[0] = sm, tr.t[0] = ts;
|
||||
++tr.line;
|
||||
if (!tr.sets) tr.first = ts;
|
||||
tr.last = ts, ++tr.sets, tr.left += h->pts[0][0] < 0;
|
||||
if (tl)
|
||||
std::fprintf(tl, "%.3f %d %s %d %+.3f %+.3f %+.3f\n", ts, h->id, h->pts[0][0] < 0 ? "L" : "R", h->nviews,
|
||||
h->pts[0][0], h->pts[0][1], h->pts[0][2]);
|
||||
if (dp) {
|
||||
std::vector<const Seen *> vs;
|
||||
for (const Seen &v : seen)
|
||||
if (v.hand == h->id) vs.push_back(&v);
|
||||
std::sort(vs.begin(), vs.end(), [](const Seen *a, const Seen *b) { return a->cam < b->cam; });
|
||||
std::string names;
|
||||
for (const Seen *v : vs) names += (names.empty() ? "" : "+") + v->cam;
|
||||
std::fprintf(dp, "%.4f %d %s %d %s %.4f %.3f %.4f %.4f %.4f %.4f %.4f %.4f", ts, h->id,
|
||||
h->pts[0][0] < 0 ? "L" : "R", h->nviews, names.empty() ? "-" : names.c_str(), h->residual,
|
||||
h->scale, raw[0], raw[1], raw[2], sm[0], sm[1], sm[2]);
|
||||
for (const Seen *v : vs) {
|
||||
V3 mono[21];
|
||||
const bool ok = tracker.single_view(used.at(v->cam), v->lm, h->scale, mono);
|
||||
const V3 p = ok ? palm(mono) : V3{NAN, NAN, NAN};
|
||||
std::fprintf(dp, " %s %.2f %.4f %.4f %.4f", v->cam.c_str(), v->lm.presence, p[0], p[1], p[2]);
|
||||
}
|
||||
std::fputc('\n', dp);
|
||||
}
|
||||
}
|
||||
if (tl && out.empty()) std::fprintf(tl, "%.3f -\n", ts);
|
||||
if (tl)
|
||||
for (const Seen &v : seen)
|
||||
std::fprintf(tl, "%.3f view %d %s presence %.2f roi %.0f %.0f %.0f %.3f set %d\n", ts, v.hand, v.cam.c_str(),
|
||||
v.lm.presence, v.roi.center[0], v.roi.center[1], v.roi.size, v.roi.rotation, index);
|
||||
left += l, right += r, both += l && r;
|
||||
++hist[std::min<size_t>(out.size(), 2)];
|
||||
}
|
||||
|
||||
if (oracle > 0 && nsets % oracle == 0) {
|
||||
const Stats keep = tracker.stats;
|
||||
const auto seen = tracker.exhaustive(images);
|
||||
tracker.stats = keep;
|
||||
bool l = false, r = false;
|
||||
for (const Seen &s : seen) {
|
||||
(s.wrist[0] < 0 ? l : r) = true;
|
||||
++o_by_cam[s.cam + (s.wrist[0] < 0 ? " L" : " R")];
|
||||
}
|
||||
bool tl_ = false, tr_ = false;
|
||||
for (const Hand *h : last_out) (h->pts[0][0] < 0 ? tl_ : tr_) = true;
|
||||
++o_sets;
|
||||
o_left += l, o_right += r;
|
||||
o_left_hit += l && tl_, o_right_hit += r && tr_;
|
||||
o_left_extra += !l && tl_, o_right_extra += !r && tr_;
|
||||
if (tl && ((!l && tl_) || (!r && tr_))) std::fprintf(tl, "%.3f oracle-extra %s%s set %d\n", ts, !l && tl_ ? "L" : "", !r && tr_ ? "R" : "", index);
|
||||
}
|
||||
} while (in.next(cams, px));
|
||||
if (tl) std::fclose(tl);
|
||||
if (dp) std::fclose(dp);
|
||||
|
||||
const double secs = nsets > 1 ? nsets / 30.0 : 0;
|
||||
const Stats &s = tracker.stats;
|
||||
std::printf("%s: %d sets (%.0f s), processed %d (%.1f/s), %.1f ms per step\n", dir.c_str(), nsets, secs, processed,
|
||||
processed / std::max(secs, 1e-9), busy_ms / std::max(processed, 1));
|
||||
std::printf("hands per processed set: 0 %.0f%%, 1 %.0f%%, 2 %.0f%%; a hand on the left %.0f%%, right %.0f%%, both %.0f%%\n",
|
||||
100.0 * hist[0] / processed, 100.0 * hist[1] / processed, 100.0 * hist[2] / processed,
|
||||
100.0 * left / processed, 100.0 * right / processed, 100.0 * both / processed);
|
||||
std::vector<double> lens[2];
|
||||
for (auto &[id, tr] : tracks) lens[tr.left * 2 > tr.sets ? 0 : 1].push_back(tr.last - tr.first);
|
||||
for (int k = 0; k < 2; ++k) {
|
||||
double total = 0;
|
||||
for (double d : lens[k]) total += d;
|
||||
std::printf("%s tracks: %zu, median %.1f s, total %.0f s\n", k ? "right" : "left ", lens[k].size(), median(lens[k]), total);
|
||||
}
|
||||
std::printf("views lost %d, handoff misses %d, dups %d, splits %d; hands new %d, merged %d, forgotten %d\n", s.lost,
|
||||
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten);
|
||||
std::printf("model calls: palm %d (%.1f/s), hand %d (%.1f/s)\n", s.palm_calls, s.palm_calls / std::max(secs, 1e-9),
|
||||
s.hand_calls, s.hand_calls / std::max(secs, 1e-9));
|
||||
{
|
||||
std::vector<double> r, step;
|
||||
std::map<int, double> prev;
|
||||
for (auto &[id, x] : s.mono_ratio) {
|
||||
r.push_back(x);
|
||||
if (prev.count(id)) step.push_back(std::fabs(x - prev[id]));
|
||||
prev[id] = x;
|
||||
}
|
||||
std::sort(r.begin(), r.end());
|
||||
std::sort(step.begin(), step.end());
|
||||
if (!r.empty())
|
||||
std::printf("single-view distance / stereo: 10%% %.2f, median %.2f, 90%% %.2f; change between frames median %.3f, 90%% %.3f\n",
|
||||
r[r.size() / 10], r[r.size() / 2], r[r.size() * 9 / 10], step[step.size() / 2], step[step.size() * 9 / 10]);
|
||||
}
|
||||
std::printf("palms within 20 cm of the eyes: %d of %d hand updates\n", near_face, hand_updates);
|
||||
for (int k = 0; k < 2; ++k) std::sort(pinch_len[k].begin(), pinch_len[k].end());
|
||||
std::printf("pinches (%s, %.3f/%.3f m): left %zu (median %.2f s), right %zu (median %.2f s), %d ended by losing the hand\n",
|
||||
pinch_params.triangulated ? "triangulated tips" : "world landmarks", pinch_params.begin_m, pinch_params.end_m,
|
||||
pinch_len[0].size(), pinch_len[0].empty() ? 0 : pinch_len[0][pinch_len[0].size() / 2], pinch_len[1].size(),
|
||||
pinch_len[1].empty() ? 0 : pinch_len[1][pinch_len[1].size() / 2], pinch_lost);
|
||||
std::sort(jit_raw.begin(), jit_raw.end());
|
||||
std::sort(jit_sm.begin(), jit_sm.end());
|
||||
if (!jit_raw.empty())
|
||||
std::printf("palm jitter (off a straight line through the last two updates): measured median %.1f mm, 90%% %.1f mm; "
|
||||
"published median %.1f mm, 90%% %.1f mm\n", jit_raw[jit_raw.size() / 2], jit_raw[jit_raw.size() * 9 / 10],
|
||||
jit_sm[jit_sm.size() / 2], jit_sm[jit_sm.size() * 9 / 10]);
|
||||
if (o_sets) {
|
||||
std::printf("oracle, %d sets: a left hand findable in %d, the tracker had it in %d (%.0f%%); right %d, had %d (%.0f%%)\n",
|
||||
o_sets, o_left, o_left_hit, 100.0 * o_left_hit / std::max(o_left, 1), o_right, o_right_hit,
|
||||
100.0 * o_right_hit / std::max(o_right, 1));
|
||||
std::printf(" tracker had a hand the full search didn't find: left %d, right %d\n", o_left_extra, o_right_extra);
|
||||
std::printf(" found by camera:");
|
||||
for (auto &[k, n] : o_by_cam) std::printf(" %s %d", k.c_str(), n);
|
||||
std::printf("\n");
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,233 @@
|
||||
// fh-ringplay: play a recording (fh-tracker --record) into a frame ring in real time, the
|
||||
// way fh-camd publishes live cameras, so fh-tracker --ring PATH processes the same frames
|
||||
// run after run. For A/B tests of how the tracker runs (probes/core-ab.sh).
|
||||
//
|
||||
// fh-ringplay DIR --ring PATH [--from S] [--to S] [--loop] [--cpus 0,1]
|
||||
//
|
||||
// Frames are stamped as they're published, so the tracker's latency figures stay
|
||||
// meaningful. Cameras carry their calibration name and no device node (fh-tracker maps
|
||||
// them by name), so a recording made with the right names needs no --swap-sides.
|
||||
// Dark frames (<name>_dk) are skipped. Needs no root: the ring is an ordinary file.
|
||||
#include "record.h"
|
||||
|
||||
extern "C" {
|
||||
#include "../camd/fhring.h"
|
||||
}
|
||||
|
||||
#include <fcntl.h>
|
||||
#include <sched.h>
|
||||
#include <sys/mman.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cerrno>
|
||||
#include <csignal>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <ctime>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
namespace {
|
||||
|
||||
volatile std::sig_atomic_t g_stop = 0;
|
||||
|
||||
uint64_t clock_ns(clockid_t id) {
|
||||
timespec ts;
|
||||
clock_gettime(id, &ts);
|
||||
return uint64_t(ts.tv_sec) * 1'000'000'000ull + uint64_t(ts.tv_nsec);
|
||||
}
|
||||
|
||||
// Sets from a recording, reading only the pixels of the cameras that get published.
|
||||
class Reader {
|
||||
public:
|
||||
bool open(const std::string &path) {
|
||||
f_ = std::fopen(path.c_str(), "rb");
|
||||
if (f_) posix_fadvise(fileno(f_), 0, 0, POSIX_FADV_SEQUENTIAL);
|
||||
return f_ != nullptr;
|
||||
}
|
||||
void rewind() { std::fseek(f_, 0, SEEK_SET); }
|
||||
// False at the end. cams: every camera in the set; px[k]: pixels of camera k when
|
||||
// want(name), else left empty.
|
||||
template <class Want>
|
||||
bool next(std::vector<fh_set_cam_t> &cams, std::vector<std::vector<uint8_t>> &px, Want want) {
|
||||
fh_set_hdr_t h;
|
||||
if (std::fread(&h, sizeof h, 1, f_) != 1 || std::memcmp(h.magic, FH_SET_MAGIC, 8) || h.ncams == 0 || h.ncams > 16)
|
||||
return false;
|
||||
cams.resize(h.ncams);
|
||||
if (std::fread(cams.data(), sizeof(fh_set_cam_t), h.ncams, f_) != h.ncams) return false;
|
||||
px.resize(h.ncams);
|
||||
for (uint32_t k = 0; k < h.ncams; ++k) {
|
||||
cams[k].name[sizeof cams[k].name - 1] = 0;
|
||||
const size_t n = size_t(cams[k].width) * cams[k].height;
|
||||
if (want(cams[k].name)) {
|
||||
px[k].resize(n);
|
||||
if (std::fread(px[k].data(), 1, n, f_) != n) return false;
|
||||
} else {
|
||||
px[k].clear();
|
||||
if (std::fseek(f_, long(n), SEEK_CUR)) return false;
|
||||
}
|
||||
}
|
||||
if (++sets_ % 64 == 0) posix_fadvise(fileno(f_), 0, std::ftell(f_), POSIX_FADV_DONTNEED); // RAM is tight
|
||||
return true;
|
||||
}
|
||||
|
||||
private:
|
||||
FILE *f_ = nullptr;
|
||||
uint64_t sets_ = 0;
|
||||
};
|
||||
|
||||
bool is_dark(const char *name) {
|
||||
const size_t n = std::strlen(name);
|
||||
return n > 3 && !std::strcmp(name + n - 3, "_dk");
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char **argv) {
|
||||
if (argc < 2 || argv[1][0] == '-') {
|
||||
std::fprintf(stderr, "usage: %s DIR --ring PATH [--from S] [--to S] [--loop] [--cpus 0,1]\n", argv[0]);
|
||||
return 1;
|
||||
}
|
||||
const std::string dir = argv[1];
|
||||
std::string ring_path;
|
||||
double from = 0, to = 1e9;
|
||||
bool loop = false;
|
||||
std::vector<int> cpus = {0, 1};
|
||||
for (int i = 2; i < argc; ++i) {
|
||||
const std::string a = argv[i];
|
||||
const bool more = i + 1 < argc;
|
||||
if (a == "--ring" && more) ring_path = argv[++i];
|
||||
else if (a == "--from" && more) from = std::atof(argv[++i]);
|
||||
else if (a == "--to" && more) to = std::atof(argv[++i]);
|
||||
else if (a == "--loop") loop = true;
|
||||
else if (a == "--cpus" && more) {
|
||||
cpus.clear();
|
||||
for (char *p = argv[++i]; *p;) {
|
||||
cpus.push_back(int(std::strtol(p, &p, 10)));
|
||||
if (*p == ',') ++p;
|
||||
else if (*p) break;
|
||||
}
|
||||
} else return std::fprintf(stderr, "unknown option %s\n", a.c_str()), 1;
|
||||
}
|
||||
if (ring_path.empty()) return std::fprintf(stderr, "--ring PATH is required\n"), 1;
|
||||
if (!cpus.empty()) {
|
||||
cpu_set_t set;
|
||||
CPU_ZERO(&set);
|
||||
for (int c : cpus) CPU_SET(c, &set);
|
||||
if (sched_setaffinity(0, sizeof set, &set) < 0) std::perror("sched_setaffinity");
|
||||
}
|
||||
std::signal(SIGINT, [](int) { g_stop = 1; });
|
||||
std::signal(SIGTERM, [](int) { g_stop = 1; });
|
||||
|
||||
Reader in;
|
||||
if (!in.open(dir + "/sets.bin")) return std::fprintf(stderr, "%s/sets.bin: %s\n", dir.c_str(), std::strerror(errno)), 1;
|
||||
std::vector<fh_set_cam_t> cams;
|
||||
std::vector<std::vector<uint8_t>> px;
|
||||
auto want = [](const char *name) { return !is_dark(name); };
|
||||
if (!in.next(cams, px, want)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
|
||||
auto set_time = [&](const std::vector<fh_set_cam_t> &cs) { // earliest bright capture, s
|
||||
uint64_t t = UINT64_MAX;
|
||||
for (const auto &c : cs)
|
||||
if (!is_dark(c.name)) t = std::min(t, c.capture_ns);
|
||||
return double(t) * 1e-9;
|
||||
};
|
||||
const double rec0 = set_time(cams);
|
||||
// skip to --from before the ring exists, so a reader never finds it without a heartbeat
|
||||
bool have = true;
|
||||
while (have && set_time(cams) - rec0 < from && !g_stop) have = in.next(cams, px, want);
|
||||
if (!have) return std::fprintf(stderr, "%s: nothing after %.1f s\n", dir.c_str(), from), 1;
|
||||
|
||||
// the ring: the recording's bright cameras, as fh-camd lays them out
|
||||
std::vector<int> pub; // set camera index of each ring camera
|
||||
for (size_t k = 0; k < cams.size() && pub.size() < FH_RING_MAX_CAMS; ++k)
|
||||
if (!is_dark(cams[k].name)) pub.push_back(int(k));
|
||||
size_t len = sizeof(fh_ring_hdr_t);
|
||||
std::vector<uint64_t> offset(pub.size()), slot_bytes(pub.size());
|
||||
for (size_t r = 0; r < pub.size(); ++r) {
|
||||
const fh_set_cam_t &c = cams[pub[r]];
|
||||
slot_bytes[r] = (sizeof(fh_ring_slot_t) + size_t(c.width) * c.height + 63) & ~size_t(63);
|
||||
offset[r] = len;
|
||||
len += FH_RING_SLOTS * slot_bytes[r];
|
||||
}
|
||||
const int fd = ::open(ring_path.c_str(), O_RDWR | O_CREAT | O_TRUNC | O_CLOEXEC, 0600);
|
||||
if (fd < 0 || ftruncate(fd, off_t(len)) < 0)
|
||||
return std::fprintf(stderr, "%s: %s\n", ring_path.c_str(), std::strerror(errno)), 1;
|
||||
void *m = mmap(nullptr, len, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||
close(fd);
|
||||
if (m == MAP_FAILED) return std::fprintf(stderr, "mmap %s: %s\n", ring_path.c_str(), std::strerror(errno)), 1;
|
||||
auto *base = static_cast<uint8_t *>(m);
|
||||
auto *hdr = reinterpret_cast<fh_ring_hdr_t *>(base);
|
||||
for (size_t r = 0; r < pub.size(); ++r) {
|
||||
const fh_set_cam_t &c = cams[pub[r]];
|
||||
fh_ring_cam_t &rc = hdr->cams[r];
|
||||
std::snprintf(rc.sensor, sizeof rc.sensor, "fh-ringplay");
|
||||
std::snprintf(rc.name, sizeof rc.name, "%s", c.name);
|
||||
rc.node = -1;
|
||||
rc.format = FH_FMT_GREY8;
|
||||
rc.width = rc.stride = c.width;
|
||||
rc.height = c.height;
|
||||
rc.nslots = FH_RING_SLOTS;
|
||||
rc.slot_offset = offset[r];
|
||||
rc.slot_bytes = slot_bytes[r];
|
||||
}
|
||||
hdr->version = FH_RING_VERSION;
|
||||
hdr->header_bytes = sizeof(fh_ring_hdr_t);
|
||||
hdr->ncams = uint32_t(pub.size());
|
||||
hdr->file_bytes = len;
|
||||
hdr->writer_pid = getpid();
|
||||
std::memcpy(hdr->magic, FH_RING_MAGIC, 8);
|
||||
__atomic_store_n(&hdr->heartbeat_ns, clock_ns(CLOCK_MONOTONIC), __ATOMIC_RELEASE);
|
||||
std::printf("playing %s into %s:", dir.c_str(), ring_path.c_str());
|
||||
for (int k : pub) std::printf(" %s", cams[k].name);
|
||||
std::printf("\n");
|
||||
std::fflush(stdout);
|
||||
|
||||
uint64_t published = 0, rounds = 0;
|
||||
for (;;) {
|
||||
// one pass over [from, to]: each set goes out at its recorded offset from the first
|
||||
while (have && set_time(cams) - rec0 < from && !g_stop) {
|
||||
__atomic_store_n(&hdr->heartbeat_ns, clock_ns(CLOCK_MONOTONIC), __ATOMIC_RELEASE);
|
||||
have = in.next(cams, px, want);
|
||||
}
|
||||
const double first = set_time(cams);
|
||||
const uint64_t start = clock_ns(CLOCK_MONOTONIC);
|
||||
while (have && !g_stop && set_time(cams) - rec0 <= to) {
|
||||
const uint64_t due = start + uint64_t((set_time(cams) - first) * 1e9);
|
||||
for (uint64_t now = clock_ns(CLOCK_MONOTONIC); now < due && !g_stop; now = clock_ns(CLOCK_MONOTONIC)) {
|
||||
__atomic_store_n(&hdr->heartbeat_ns, now, __ATOMIC_RELEASE);
|
||||
const uint64_t wait = std::min<uint64_t>(due - now, 100'000'000);
|
||||
const timespec ts{time_t(wait / 1'000'000'000), long(wait % 1'000'000'000)};
|
||||
nanosleep(&ts, nullptr);
|
||||
}
|
||||
const uint64_t raw = clock_ns(CLOCK_MONOTONIC_RAW), mono = clock_ns(CLOCK_MONOTONIC);
|
||||
for (size_t r = 0; r < pub.size(); ++r) {
|
||||
fh_ring_cam_t &rc = hdr->cams[r];
|
||||
if (px[pub[r]].size() != size_t(rc.width) * rc.height) continue;
|
||||
const uint64_t n = rc.latest + 1;
|
||||
auto *s = reinterpret_cast<fh_ring_slot_t *>(base + rc.slot_offset + (n % rc.nslots) * rc.slot_bytes);
|
||||
__atomic_store_n(&s->seq, 2 * n + 1, __ATOMIC_RELAXED);
|
||||
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||
std::memcpy(reinterpret_cast<uint8_t *>(s + 1), px[pub[r]].data(), px[pub[r]].size());
|
||||
s->frame = n;
|
||||
s->capture_ns = raw; // taken now, as a live camera's frame would be
|
||||
s->dqbuf_ns = mono;
|
||||
s->publish_ns = mono;
|
||||
__atomic_store_n(&s->seq, 2 * n + 2, __ATOMIC_RELEASE);
|
||||
__atomic_store_n(&rc.latest, n, __ATOMIC_RELEASE);
|
||||
++rc.published;
|
||||
}
|
||||
__atomic_store_n(&hdr->heartbeat_ns, mono, __ATOMIC_RELEASE);
|
||||
++published;
|
||||
have = in.next(cams, px, want);
|
||||
}
|
||||
++rounds;
|
||||
if (g_stop || !loop) break;
|
||||
in.rewind();
|
||||
have = in.next(cams, px, want);
|
||||
}
|
||||
std::printf("published %llu sets in %llu pass(es)\n", (unsigned long long)published, (unsigned long long)rounds);
|
||||
__atomic_store_n(&hdr->heartbeat_ns, 0, __ATOMIC_RELEASE); // readers see the writer gone
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,609 @@
|
||||
#include "tracker.h"
|
||||
|
||||
#include <pthread.h>
|
||||
#include <sched.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <set>
|
||||
|
||||
namespace {
|
||||
|
||||
// where arms start, head frame: below and slightly behind the eyes
|
||||
const V3 kShoulders[2] = {{0.17, -0.25, 0.08}, {-0.17, -0.25, 0.08}};
|
||||
// landmark pairs across the palm, rigid enough for single-view depth
|
||||
const int kPalmPairs[][2] = {{0, 5}, {0, 9}, {0, 13}, {0, 17}, {5, 17}, {5, 13}, {9, 17}, {1, 17}, {1, 5}};
|
||||
constexpr double kFastSpeed = 0.25; // m/s
|
||||
constexpr double kSearchInterval = 0.2;
|
||||
// One Euro filter on the published landmarks: still hands are smoothed hard (tracking
|
||||
// noise is a few mm per frame), fast ones barely, so they don't lag.
|
||||
constexpr double kMinCutoff = 2.0; // Hz, a still hand
|
||||
constexpr double kBeta = 30.0; // Hz more per m/s of palm speed
|
||||
constexpr double kSpeedCutoff = 1.5; // Hz, for the palm speed itself
|
||||
// With one view, the hand's distance from the camera comes from how big it looks, which
|
||||
// is off by 10-30% and wanders ~10% between frames. Its direction is exact. So a hand that
|
||||
// was just located keeps its distance, drifting toward the one-view guess by this much a frame.
|
||||
constexpr double kMonoDepthGain = 0.1;
|
||||
// Is a triangulated hand as far from each camera as its apparent size says? With the
|
||||
// model's average hand, clean stereo pairs measure 0.71-1.51 times the one-view distance
|
||||
// (5-95%, median 1.16); pairs of two different hands mostly far less.
|
||||
constexpr double kSizePrior = 1.16, kRatioLo = 0.6, kRatioHi = 1.9;
|
||||
|
||||
double ms_since(std::chrono::steady_clock::time_point t) {
|
||||
return std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t).count();
|
||||
}
|
||||
|
||||
bool is_color(const Camera &c) { return c.name.rfind("color", 0) == 0; } // Arcturus, 145 degree image circle
|
||||
// The wide cameras: the side ones and the color ones
|
||||
bool is_slam(const Camera &c) { return c.name.rfind("slam", 0) == 0 || is_color(c); }
|
||||
|
||||
V2 palm_centre(const Landmarks &lm) { return lm.pts[9]; }
|
||||
|
||||
// Two views in one camera on the same hand: the landmark model puts the same points on
|
||||
// it from both crops, even when the crops differ.
|
||||
bool same_hand(const Landmarks &a, const Landmarks &b, double size) {
|
||||
double d = 0;
|
||||
for (int i = 0; i < 21; ++i) d += norm(a.pts[i] - b.pts[i]) / 21;
|
||||
return norm(palm_centre(a) - palm_centre(b)) < 0.5 * size || d < 0.25 * size;
|
||||
}
|
||||
|
||||
double hand_size(const Landmarks &lm) {
|
||||
double lo[2] = {1e9, 1e9}, hi[2] = {-1e9, -1e9};
|
||||
for (const V2 &p : lm.pts)
|
||||
for (int k = 0; k < 2; ++k) lo[k] = std::min(lo[k], p[k]), hi[k] = std::max(hi[k], p[k]);
|
||||
return std::max(hi[0] - lo[0], hi[1] - lo[1]);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
// ---------------------------------------------------------------------------- pool
|
||||
|
||||
Pool::Pool(int threads, const std::vector<int> &cpus) {
|
||||
for (int i = 0; i < threads; ++i) threads_.emplace_back(&Pool::loop, this, cpus[i % cpus.size()]);
|
||||
}
|
||||
|
||||
Pool::~Pool() {
|
||||
{
|
||||
std::lock_guard<std::mutex> l(mu_);
|
||||
stop_ = true;
|
||||
}
|
||||
wake_.notify_all();
|
||||
for (auto &t : threads_) t.join();
|
||||
}
|
||||
|
||||
void Pool::loop(int cpu) {
|
||||
cpu_set_t set;
|
||||
CPU_ZERO(&set);
|
||||
CPU_SET(cpu, &set);
|
||||
pthread_setaffinity_np(pthread_self(), sizeof set, &set); // ignored if not allowed
|
||||
std::unique_lock<std::mutex> l(mu_);
|
||||
for (;;) {
|
||||
wake_.wait(l, [&] { return stop_ || (jobs_ && next_ < jobs_->size()); });
|
||||
if (stop_) return;
|
||||
auto &job = (*jobs_)[next_++];
|
||||
l.unlock();
|
||||
job();
|
||||
l.lock();
|
||||
if (++finished_ == jobs_->size()) done_.notify_all();
|
||||
}
|
||||
}
|
||||
|
||||
void Pool::run(std::vector<std::function<void()>> &jobs) {
|
||||
if (jobs.empty()) return;
|
||||
std::unique_lock<std::mutex> l(mu_);
|
||||
jobs_ = &jobs, next_ = 0, finished_ = 0;
|
||||
wake_.notify_all();
|
||||
done_.wait(l, [&] { return finished_ == jobs.size(); });
|
||||
jobs_ = nullptr;
|
||||
}
|
||||
|
||||
// ------------------------------------------------------------------------- tracker
|
||||
|
||||
Tracker::Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views)
|
||||
: nets_(nets), pool_(pool), max_views_(max_views) {
|
||||
for (const auto &[name, cam] : cams) {
|
||||
cams_[name] = &cam;
|
||||
if (is_slam(cam)) {
|
||||
add_tiles(cam, 0.45, 3, 3);
|
||||
add_tiles(cam, 0.65, 2, 2);
|
||||
add_tiles(cam, 1.0, 1, 1); // hands close to the face fill much of the frame
|
||||
} else {
|
||||
add_tiles(cam, 0.6, 3, 2);
|
||||
add_tiles(cam, 1.0, 1, 1);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void Tracker::add_tiles(const Camera &cam, double frac, int gx, int gy) {
|
||||
const double s = frac * std::max(cam.width, cam.height);
|
||||
for (int i = 0; i < gx; ++i)
|
||||
for (int j = 0; j < gy; ++j) {
|
||||
const double x = gx > 1 ? s / 2 + (cam.width - s) * i / (gx - 1) : cam.width / 2.0;
|
||||
const double y = gy > 1 ? s / 2 + (cam.height - s) * j / (gy - 1) : cam.height / 2.0;
|
||||
Tile t{&cam, {x, y}, s, 0, 0};
|
||||
// turn the crop so the expected shoulder-to-hand direction points up
|
||||
const V3 ray = cam.ray(t.center), p = cam.origin + ray * 0.45;
|
||||
const V3 d = unit(p - kShoulders[p[0] > 0 ? 0 : 1]);
|
||||
const V2 a = cam.project(p, nullptr), b = cam.project(p + d * 0.05, nullptr);
|
||||
t.rotation = std::atan2(b[0] - a[0], -(b[1] - a[1]));
|
||||
t.weight = std::max(0.15, dot(ray, unit(V3{0, -0.45, -0.9})));
|
||||
tiles_.push_back(t);
|
||||
}
|
||||
}
|
||||
|
||||
double Tracker::interval() const {
|
||||
double fastest = -1;
|
||||
for (const auto &[id, h] : hands_)
|
||||
if (h.seen_ns == last_ns_) fastest = std::max(fastest, norm(h.dpalm)); // filtered: noise isn't speed
|
||||
return fastest < 0 ? 1 / 5.0 : fastest > kFastSpeed ? 1 / 30.0 : 1 / 15.0;
|
||||
}
|
||||
|
||||
bool Tracker::inside(const Camera &cam, V2 uv) const {
|
||||
const double m = 0.12;
|
||||
return uv[0] >= m * cam.width && uv[0] <= (1 - m) * cam.width && uv[1] >= m * cam.height &&
|
||||
uv[1] <= (1 - m) * cam.height && cam.off_axis(uv) < (is_color(cam) ? 70.0 : is_slam(cam) ? 80.0 : 75.0);
|
||||
}
|
||||
|
||||
void Tracker::run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views) {
|
||||
if (views.empty()) return;
|
||||
const auto t0 = std::chrono::steady_clock::now();
|
||||
std::vector<std::function<void()>> jobs;
|
||||
for (View *v : views) {
|
||||
const Image &img = images.at(v->cam->name);
|
||||
jobs.push_back([this, v, &img] {
|
||||
v->lm = nets_.landmarks(img, v->roi);
|
||||
v->has_lm = true;
|
||||
v->fresh = true;
|
||||
});
|
||||
}
|
||||
pool_.run(jobs);
|
||||
stats.hand_calls += int(views.size());
|
||||
stats.hand_ms += ms_since(t0);
|
||||
++stats.hand_batches;
|
||||
}
|
||||
|
||||
bool Tracker::single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const {
|
||||
V3 rays[21];
|
||||
for (int i = 0; i < 21; ++i) rays[i] = cam.ray(lm.pts[i]);
|
||||
std::vector<std::pair<double, double>> est; // (depth, weight)
|
||||
for (const auto &pr : kPalmPairs) {
|
||||
const int i = pr[0], j = pr[1];
|
||||
const double d = std::hypot(lm.world[i][0] - lm.world[j][0], lm.world[i][1] - lm.world[j][1]) * scale;
|
||||
const double a = std::acos(std::clamp(dot(rays[i], rays[j]), -1.0, 1.0));
|
||||
if (a > 1e-3 && d > 0.01) est.push_back({d / a, d});
|
||||
}
|
||||
if (est.empty()) return false;
|
||||
std::sort(est.begin(), est.end());
|
||||
double total = 0, acc = 0, depth = est.back().first;
|
||||
for (auto &e : est) total += e.second;
|
||||
for (auto &e : est)
|
||||
if ((acc += e.second) >= total / 2) { depth = e.first; break; }
|
||||
double zmean = 0;
|
||||
for (int i = 0; i < 21; ++i) zmean += lm.world[i][2] / 21;
|
||||
for (int i = 0; i < 21; ++i) out[i] = cam.origin + rays[i] * (depth + (lm.world[i][2] - zmean) * scale);
|
||||
return true;
|
||||
}
|
||||
|
||||
// A triangulated hand is in front of each camera, as far as its apparent size says (see
|
||||
// kSizePrior). Returns how far off that is (the sum of |log| ratios), or -1 if implausible.
|
||||
double Tracker::size_misfit(const std::vector<const View *> &views, const V3 *pts) const {
|
||||
double misfit = 0;
|
||||
for (const View *v : views) {
|
||||
V3 mono[21];
|
||||
// along the view's own ray (fisheye: a hand near the image edge is far off the axis)
|
||||
if (dot(pts[9] - v->cam->origin, v->cam->ray(v->lm.pts[9])) < 0.08) return -1;
|
||||
if (!single_view(*v->cam, v->lm, 1.0, mono)) continue;
|
||||
const double r = norm(pts[9] - v->cam->origin) / norm(mono[9] - v->cam->origin);
|
||||
if (r < kRatioLo || r > kRatioHi) return -1;
|
||||
misfit += std::fabs(std::log(r / kSizePrior));
|
||||
}
|
||||
return misfit;
|
||||
}
|
||||
|
||||
bool Tracker::hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns) {
|
||||
views.erase(std::remove_if(views.begin(), views.end(), [](View *v) { return !v->has_lm || !v->fresh; }), views.end());
|
||||
if (views.empty()) return false;
|
||||
if (views.size() >= 2) {
|
||||
const int n = int(views.size());
|
||||
std::vector<V3> origins(n), dirs(n);
|
||||
std::vector<double> w(n), res(21);
|
||||
V3 pts[21];
|
||||
for (int k = 0; k < 21; ++k) {
|
||||
for (int v = 0; v < n; ++v) {
|
||||
origins[v] = views[v]->cam->origin;
|
||||
dirs[v] = views[v]->cam->ray(views[v]->lm.pts[k]);
|
||||
w[v] = views[v]->lm.presence;
|
||||
}
|
||||
pts[k] = triangulate(origins.data(), dirs.data(), w.data(), n, &res[k]);
|
||||
}
|
||||
std::nth_element(res.begin(), res.begin() + 10, res.end());
|
||||
const double residual = res[10];
|
||||
// the views disagree: two different hands; keep the stronger. Rays to two different
|
||||
// hands can pass close to each other near the cameras, so check the distance too.
|
||||
if (residual > 0.03 || size_misfit({views.begin(), views.end()}, pts) < 0) {
|
||||
View *best = *std::max_element(views.begin(), views.end(),
|
||||
[](View *a, View *b) { return a->lm.presence < b->lm.presence; });
|
||||
for (View *v : views)
|
||||
if (v != best) v->hand = -1;
|
||||
++stats.splits;
|
||||
return hand_3d(hand, {best}, t_ns);
|
||||
}
|
||||
// learn how big this user's hand is compared to the model's average hand
|
||||
std::vector<double> t, m;
|
||||
const Landmarks &ref = views[0]->lm;
|
||||
for (const auto &pr : kPalmPairs) {
|
||||
t.push_back(norm(pts[pr[0]] - pts[pr[1]]));
|
||||
m.push_back(norm(V3{ref.world[pr[0]][0], ref.world[pr[0]][1], ref.world[pr[0]][2]} -
|
||||
V3{ref.world[pr[1]][0], ref.world[pr[1]][1], ref.world[pr[1]][2]}));
|
||||
}
|
||||
std::nth_element(t.begin(), t.begin() + t.size() / 2, t.end());
|
||||
std::nth_element(m.begin(), m.begin() + m.size() / 2, m.end());
|
||||
if (m[m.size() / 2] > 0 && residual < 0.008) // only from clean matches
|
||||
hand.scale += 0.1 * (std::clamp(t[t.size() / 2] / m[m.size() / 2], 0.8, 1.6) - hand.scale);
|
||||
std::copy(pts, pts + 21, hand.pts);
|
||||
hand.residual = residual;
|
||||
for (View *v : views) {
|
||||
V3 mono[21];
|
||||
if (!single_view(*v->cam, v->lm, hand.scale, mono)) continue;
|
||||
const V3 o = v->cam->origin;
|
||||
stats.mono_ratio.push_back({hand.id, norm(mono[9] - o) / norm(pts[9] - o)});
|
||||
}
|
||||
} else {
|
||||
const Camera &cam = *views[0]->cam;
|
||||
V3 pts[21];
|
||||
if (!single_view(cam, views[0]->lm, hand.scale, pts)) return false;
|
||||
const double guess = norm(pts[9] - cam.origin);
|
||||
if (hand.has_pts && t_ns - hand.seen_ns < 300'000'000 && guess > 0) {
|
||||
const double was = norm(hand.pts[9] - cam.origin), d = was + kMonoDepthGain * (guess - was);
|
||||
for (V3 &p : pts) p = cam.origin + (p - cam.origin) * (d / guess);
|
||||
}
|
||||
std::copy(pts, pts + 21, hand.pts);
|
||||
hand.residual = -1;
|
||||
}
|
||||
hand.has_pts = true;
|
||||
hand.nviews = int(views.size());
|
||||
for (View *v : views) hand.right_score += 0.2 * (v->lm.right - hand.right_score);
|
||||
return true;
|
||||
}
|
||||
|
||||
// How badly two views in different cameras fit one hand: the rays should meet, each view's
|
||||
// apparent size should match its distance, and the model should call both the same hand
|
||||
// (left or right). Negative if they can't be one hand. Side by side hands sit on the same
|
||||
// epipolar lines of the side cameras, so the distance check is what tells them apart.
|
||||
double Tracker::pair_cost(const View &a, const View &b) const {
|
||||
V3 pts[21];
|
||||
std::vector<double> res(21);
|
||||
for (int k = 0; k < 21; ++k) {
|
||||
const V3 o[2] = {a.cam->origin, b.cam->origin};
|
||||
const V3 d[2] = {a.cam->ray(a.lm.pts[k]), b.cam->ray(b.lm.pts[k])};
|
||||
const double w[2] = {a.lm.presence, b.lm.presence};
|
||||
pts[k] = triangulate(o, d, w, 2, &res[k]);
|
||||
}
|
||||
std::nth_element(res.begin(), res.begin() + 10, res.end());
|
||||
if (res[10] > 0.03) return -1;
|
||||
const double misfit = size_misfit({&a, &b}, pts);
|
||||
return misfit < 0 ? -1 : res[10] / 0.01 + misfit + std::fabs(a.lm.right - b.lm.right);
|
||||
}
|
||||
|
||||
// Which views in two cameras are the same hand: every way of pairing them up (a few views
|
||||
// each), scored with pair_cost. Keeps the hands' pairing unless another is clearly better,
|
||||
// then relabels the views, keeping the longer-tracked hand's id.
|
||||
void Tracker::associate() {
|
||||
constexpr double kPairBonus = 2.0, kBetter = 0.3;
|
||||
std::vector<const Camera *> cams;
|
||||
for (View &v : views_)
|
||||
if (std::find(cams.begin(), cams.end(), v.cam) == cams.end()) cams.push_back(v.cam);
|
||||
std::sort(cams.begin(), cams.end(), [](const Camera *a, const Camera *b) { return a->name < b->name; });
|
||||
for (size_t i = 0; i < cams.size(); ++i)
|
||||
for (size_t j = i + 1; j < cams.size(); ++j) {
|
||||
std::vector<View *> A, B;
|
||||
for (View &v : views_) {
|
||||
if (!v.fresh) continue;
|
||||
if (v.cam == cams[i]) A.push_back(&v);
|
||||
else if (v.cam == cams[j]) B.push_back(&v);
|
||||
}
|
||||
if (A.empty() || B.empty() || A.size() > 3 || B.size() > 3) continue;
|
||||
std::vector<std::vector<double>> c(A.size(), std::vector<double>(B.size()));
|
||||
for (size_t x = 0; x < A.size(); ++x)
|
||||
for (size_t y = 0; y < B.size(); ++y) c[x][y] = pair_cost(*A[x], *B[y]);
|
||||
auto score = [&](const std::vector<int> &m) { // m[x]: A[x]'s partner in B, or -1
|
||||
double s = 0;
|
||||
for (size_t x = 0; x < A.size(); ++x)
|
||||
if (m[x] >= 0 && c[x][m[x]] >= 0) s += c[x][m[x]] - kPairBonus;
|
||||
return s;
|
||||
};
|
||||
std::vector<int> cur(A.size(), -1);
|
||||
for (size_t x = 0; x < A.size(); ++x)
|
||||
for (size_t y = 0; y < B.size(); ++y)
|
||||
if (A[x]->hand == B[y]->hand) cur[x] = int(y);
|
||||
std::vector<int> best = cur, m(A.size(), -1);
|
||||
double best_score = score(cur);
|
||||
const double cur_score = best_score;
|
||||
std::function<void(size_t, unsigned)> walk = [&](size_t x, unsigned used) {
|
||||
if (x == A.size()) {
|
||||
const double sc = score(m);
|
||||
if (sc < best_score) best_score = sc, best = m;
|
||||
return;
|
||||
}
|
||||
m[x] = -1;
|
||||
walk(x + 1, used);
|
||||
for (size_t y = 0; y < B.size(); ++y)
|
||||
if (!(used >> y & 1) && c[x][y] >= 0) {
|
||||
m[x] = int(y);
|
||||
walk(x + 1, used | 1u << y);
|
||||
}
|
||||
m[x] = -1;
|
||||
};
|
||||
walk(0, 0);
|
||||
if (best == cur || best_score > cur_score - kBetter) continue;
|
||||
auto frames = [&](int id) {
|
||||
const auto h = hands_.find(id);
|
||||
return h == hands_.end() ? -1 : h->second.frames;
|
||||
};
|
||||
for (size_t x = 0; x < A.size(); ++x) {
|
||||
if (best[x] < 0) continue;
|
||||
View *a = A[x], *b = B[best[x]];
|
||||
int id = a->hand;
|
||||
bool free = true; // b's hand isn't another A view's
|
||||
for (size_t x2 = 0; x2 < A.size(); ++x2) free = free && (x2 == x || A[x2]->hand != b->hand);
|
||||
if (free && frames(b->hand) > frames(id)) id = b->hand;
|
||||
a->hand = b->hand = id;
|
||||
}
|
||||
// a B view left unpaired that still shares a hand with an A view starts its own
|
||||
for (size_t y = 0; y < B.size(); ++y) {
|
||||
if (std::find(best.begin(), best.end(), int(y)) != best.end()) continue;
|
||||
bool shared = false;
|
||||
for (View *a : A) shared = shared || a->hand == B[y]->hand;
|
||||
if (!shared) continue;
|
||||
B[y]->hand = next_id_++;
|
||||
hands_[B[y]->hand].id = B[y]->hand;
|
||||
++stats.created;
|
||||
}
|
||||
++stats.merged;
|
||||
}
|
||||
}
|
||||
|
||||
std::vector<const Hand *> Tracker::step(const std::map<std::string, Image> &images, int64_t t_ns) {
|
||||
const auto t_step = std::chrono::steady_clock::now();
|
||||
++stats.sets;
|
||||
std::vector<View> live;
|
||||
for (View &v : views_)
|
||||
if (images.count(v.cam->name)) live.push_back(v), live.back().fresh = false;
|
||||
|
||||
// 1. hand-over: give hands with too few views a crop in other cameras
|
||||
for (auto &[id, hand] : hands_) {
|
||||
if (!hand.has_pts) continue;
|
||||
std::set<std::string> have;
|
||||
for (View &v : live)
|
||||
if (v.hand == id) have.insert(v.cam->name);
|
||||
if (int(have.size()) >= max_views_) continue;
|
||||
std::vector<std::pair<double, View>> options;
|
||||
for (auto &[name, cam] : cams_) {
|
||||
if (have.count(name) || !images.count(name)) continue;
|
||||
V2 uv[21];
|
||||
bool front = true;
|
||||
for (int k = 0; k < 21; ++k) {
|
||||
double z;
|
||||
uv[k] = cam->project(hand.pts[k], &z);
|
||||
front = front && z > 0;
|
||||
}
|
||||
const V2 centre = (uv[0] + uv[5] + uv[9] + uv[13] + uv[17]) * 0.2;
|
||||
if (!front || !inside(*cam, centre)) continue;
|
||||
View v{cam, roi_from_points(uv), id, {}, false, 0};
|
||||
options.push_back({cam->off_axis(centre), v});
|
||||
}
|
||||
std::sort(options.begin(), options.end(), [](auto &a, auto &b) { return a.first < b.first; });
|
||||
for (size_t k = 0; k < options.size() && int(have.size() + k) < max_views_; ++k) live.push_back(options[k].second);
|
||||
}
|
||||
|
||||
// 2. the landmark model on each hand's best views, within budget
|
||||
std::map<int, std::vector<View *>> by_hand;
|
||||
for (View &v : live) by_hand[v.hand].push_back(&v);
|
||||
std::vector<View *> chosen;
|
||||
for (auto &[id, vs] : by_hand) {
|
||||
std::sort(vs.begin(), vs.end(), [](View *a, View *b) {
|
||||
if (a->has_lm != b->has_lm) return a->has_lm;
|
||||
return a->cam->off_axis(a->roi.center) < b->cam->off_axis(b->roi.center);
|
||||
});
|
||||
for (int k = 0; k < int(vs.size()) && k < max_views_; ++k) chosen.push_back(vs[k]);
|
||||
}
|
||||
std::stable_sort(chosen.begin(), chosen.end(), [](View *a, View *b) { return a->has_lm > b->has_lm; });
|
||||
if (int(chosen.size()) > hand_budget_) chosen.resize(hand_budget_);
|
||||
run_landmarks(images, chosen);
|
||||
std::vector<View *> kept;
|
||||
for (View *v : chosen)
|
||||
if (v->lm.presence >= (v->frames > 0 ? keep_presence_ : min_presence_)) {
|
||||
v->roi = v->lm.next_roi();
|
||||
++v->frames;
|
||||
kept.push_back(v);
|
||||
} else {
|
||||
++(v->frames > 0 ? stats.lost : stats.handoff_miss);
|
||||
}
|
||||
// the same hand twice in one camera: keep the more confident
|
||||
std::sort(kept.begin(), kept.end(), [](View *a, View *b) { return a->lm.presence > b->lm.presence; });
|
||||
std::vector<View> next;
|
||||
for (View *v : kept) {
|
||||
const double size = hand_size(v->lm);
|
||||
bool dup = false;
|
||||
for (View &o : next) dup = dup || (o.cam == v->cam && same_hand(o.lm, v->lm, size));
|
||||
if (!dup) next.push_back(*v);
|
||||
else ++stats.dups;
|
||||
}
|
||||
views_ = next;
|
||||
|
||||
// 3. search for missing hands
|
||||
std::set<int> tracked;
|
||||
for (View &v : views_) tracked.insert(v.hand);
|
||||
if (tracked.size() < 2 && (t_ns - search_ns_) / 1e9 >= kSearchInterval - 0.01) {
|
||||
search_ns_ = t_ns;
|
||||
const int budget = tracked.empty() ? search_budget_ : std::max(1, search_budget_ - 1);
|
||||
for (Tile &t : tiles_)
|
||||
if (images.count(t.cam->name)) t.credit += t.weight;
|
||||
std::vector<Tile *> picked;
|
||||
for (int b = 0; b < budget; ++b) {
|
||||
Tile *best = nullptr;
|
||||
for (Tile &t : tiles_)
|
||||
if (images.count(t.cam->name) && std::find(picked.begin(), picked.end(), &t) == picked.end() &&
|
||||
(!best || t.credit > best->credit))
|
||||
best = &t;
|
||||
if (!best) break;
|
||||
best->credit = 0;
|
||||
picked.push_back(best);
|
||||
}
|
||||
const auto t0 = std::chrono::steady_clock::now();
|
||||
std::vector<std::vector<Palm>> found(picked.size());
|
||||
std::vector<std::function<void()>> jobs;
|
||||
for (size_t i = 0; i < picked.size(); ++i) {
|
||||
Tile *t = picked[i];
|
||||
const Image &img = images.at(t->cam->name);
|
||||
jobs.push_back([this, t, &img, &found, i] { found[i] = nets_.palms(img, t->center, t->size, t->rotation); });
|
||||
}
|
||||
pool_.run(jobs);
|
||||
stats.palm_calls += int(picked.size());
|
||||
stats.palm_ms += ms_since(t0);
|
||||
++stats.palm_batches;
|
||||
std::vector<View> fresh;
|
||||
for (size_t i = 0; i < picked.size(); ++i)
|
||||
for (const Palm &p : found[i]) {
|
||||
const Roi roi = p.roi();
|
||||
bool near = false;
|
||||
for (auto *list : {&views_, &fresh})
|
||||
for (View &v : *list) near = near || (v.cam == picked[i]->cam && norm(v.roi.center - roi.center) < 0.5 * roi.size);
|
||||
if (!near) fresh.push_back({picked[i]->cam, roi, 0, {}, false, 0});
|
||||
}
|
||||
std::vector<View *> ptrs;
|
||||
for (View &v : fresh) ptrs.push_back(&v);
|
||||
run_landmarks(images, ptrs);
|
||||
for (View &v : fresh)
|
||||
if (v.lm.presence >= min_presence_) {
|
||||
v.roi = v.lm.next_roi();
|
||||
v.frames = 1;
|
||||
views_.push_back(v);
|
||||
}
|
||||
}
|
||||
|
||||
// 4. give new views a hand: the nearest existing hand in 3D, else a new one
|
||||
for (View &v : views_) {
|
||||
if (v.hand > 0 && hands_.count(v.hand)) continue;
|
||||
V3 guess[21];
|
||||
const bool have_guess = single_view(*v.cam, v.lm, 1.0, guess);
|
||||
int best = 0;
|
||||
double dist = 0.12;
|
||||
for (auto &[id, h] : hands_) {
|
||||
if (!h.has_pts) continue;
|
||||
bool same_cam = false;
|
||||
for (View &o : views_) same_cam = same_cam || (o.hand == id && o.cam == v.cam);
|
||||
if (same_cam) continue;
|
||||
const double d = have_guess ? norm(h.pts[9] - guess[9]) : 1e9;
|
||||
if (d < dist) best = id, dist = d;
|
||||
}
|
||||
if (!best) {
|
||||
best = next_id_++;
|
||||
hands_[best].id = best;
|
||||
++stats.created;
|
||||
}
|
||||
v.hand = best;
|
||||
}
|
||||
|
||||
// 5. which views in different cameras are the same hand
|
||||
associate();
|
||||
|
||||
// 6. 3D for every hand seen now; forget hands not seen for a while
|
||||
std::vector<const Hand *> out;
|
||||
for (auto it = hands_.begin(); it != hands_.end();) {
|
||||
Hand &h = it->second;
|
||||
std::vector<View *> vs;
|
||||
for (View &v : views_)
|
||||
if (v.hand == h.id) vs.push_back(&v);
|
||||
if (!vs.empty() && hand_3d(h, vs, t_ns)) {
|
||||
const V3 palm = (h.pts[0] + h.pts[5] + h.pts[9] + h.pts[13] + h.pts[17]) * 0.2;
|
||||
if (h.last_ns && t_ns > h.last_ns)
|
||||
h.speed += 0.5 * (std::min(norm(palm - h.last_palm) / ((t_ns - h.last_ns) / 1e9), 5.0) - h.speed);
|
||||
h.last_ns = t_ns, h.last_palm = palm, h.seen_ns = t_ns;
|
||||
++h.frames;
|
||||
smooth(h, t_ns);
|
||||
out.push_back(&h);
|
||||
++it;
|
||||
} else if (t_ns - h.seen_ns > 300'000'000) {
|
||||
it = hands_.erase(it);
|
||||
++stats.forgotten;
|
||||
} else {
|
||||
++it;
|
||||
}
|
||||
}
|
||||
// views split off by a failed triangulation start over as new hands next frame
|
||||
for (View &v : views_)
|
||||
if (v.hand <= 0) {
|
||||
v.hand = next_id_++;
|
||||
hands_[v.hand].id = v.hand;
|
||||
++stats.created;
|
||||
}
|
||||
last_ns_ = t_ns;
|
||||
stats.step_ms += ms_since(t_step);
|
||||
return out;
|
||||
}
|
||||
|
||||
std::vector<Seen> Tracker::views_now() const {
|
||||
std::vector<Seen> out;
|
||||
for (const View &v : views_) out.push_back({v.cam->name, v.hand, v.roi, v.lm, {}});
|
||||
return out;
|
||||
}
|
||||
|
||||
std::vector<Seen> Tracker::exhaustive(const std::map<std::string, Image> &images) {
|
||||
std::vector<Tile *> tiles;
|
||||
for (Tile &t : tiles_)
|
||||
if (images.count(t.cam->name)) tiles.push_back(&t);
|
||||
std::vector<std::vector<Palm>> found(tiles.size());
|
||||
std::vector<std::function<void()>> jobs;
|
||||
for (size_t i = 0; i < tiles.size(); ++i)
|
||||
jobs.push_back([this, &tiles, &images, &found, i] {
|
||||
const Tile *t = tiles[i];
|
||||
found[i] = nets_.palms(images.at(t->cam->name), t->center, t->size, t->rotation);
|
||||
});
|
||||
pool_.run(jobs);
|
||||
// one crop per palm: tiles overlap, so the same palm turns up several times
|
||||
std::vector<std::pair<double, View>> palms;
|
||||
for (size_t i = 0; i < tiles.size(); ++i)
|
||||
for (const Palm &p : found[i]) palms.push_back({p.score, View{tiles[i]->cam, p.roi(), 0, {}, false, 0}});
|
||||
std::sort(palms.begin(), palms.end(), [](auto &a, auto &b) { return a.first > b.first; });
|
||||
std::vector<View> crops;
|
||||
for (auto &[score, v] : palms) {
|
||||
bool near = false;
|
||||
for (View &o : crops) near = near || (o.cam == v.cam && norm(o.roi.center - v.roi.center) < 0.5 * v.roi.size);
|
||||
if (!near) crops.push_back(v);
|
||||
}
|
||||
std::vector<View *> ptrs;
|
||||
for (View &v : crops) ptrs.push_back(&v);
|
||||
run_landmarks(images, ptrs);
|
||||
std::sort(crops.begin(), crops.end(), [](const View &a, const View &b) { return a.lm.presence > b.lm.presence; });
|
||||
std::vector<Seen> out;
|
||||
for (View &v : crops) {
|
||||
if (v.lm.presence < min_presence_) continue;
|
||||
bool dup = false;
|
||||
for (const Seen &o : out)
|
||||
dup = dup || (o.cam == v.cam->name && norm(palm_centre(o.lm) - palm_centre(v.lm)) < 0.5 * hand_size(v.lm));
|
||||
if (dup) continue;
|
||||
V3 pts[21];
|
||||
Seen s{v.cam->name, 0, v.roi, v.lm, {}};
|
||||
if (single_view(*v.cam, v.lm, 1.0, pts)) s.wrist = pts[0];
|
||||
out.push_back(s);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
void Tracker::smooth(Hand &h, int64_t t_ns) {
|
||||
const double dt = (t_ns - h.smooth_ns) / 1e9;
|
||||
h.smooth_ns = t_ns;
|
||||
if (h.frames <= 1 || dt <= 0 || dt > 0.3) { // new, or back after a gap: start over
|
||||
std::copy(h.pts, h.pts + 21, h.smooth);
|
||||
h.dpalm = {0, 0, 0};
|
||||
return;
|
||||
}
|
||||
auto alpha = [dt](double cutoff) { return 1 / (1 + 1 / (2 * M_PI * cutoff * dt)); };
|
||||
auto palm = [](const V3 *p) { return (p[0] + p[5] + p[9] + p[13] + p[17]) * 0.2; };
|
||||
const V3 d = (palm(h.pts) - palm(h.smooth)) * (1 / dt);
|
||||
h.dpalm = h.dpalm + (d - h.dpalm) * alpha(kSpeedCutoff);
|
||||
// one cutoff for the whole hand, from its palm speed, so its shape stays together
|
||||
const double a = alpha(kMinCutoff + kBeta * norm(h.dpalm));
|
||||
for (int i = 0; i < 21; ++i) h.smooth[i] = h.smooth[i] + (h.pts[i] - h.smooth[i]) * a;
|
||||
}
|
||||
@@ -0,0 +1,131 @@
|
||||
// Multi-camera hand tracking (a port of tracker/hands.py; the scheduling is described
|
||||
// there and in tracker/README.md). All 3D is metres in the head frame.
|
||||
#pragma once
|
||||
|
||||
#include "calib.h"
|
||||
#include "nets.h"
|
||||
|
||||
#include <condition_variable>
|
||||
#include <functional>
|
||||
#include <map>
|
||||
#include <memory>
|
||||
#include <mutex>
|
||||
#include <thread>
|
||||
#include <vector>
|
||||
|
||||
// Runs batches of jobs on a few threads, each pinned to a core.
|
||||
class Pool {
|
||||
public:
|
||||
// One thread per entry of cpus, pinned there (round-robin if threads > cpus).
|
||||
Pool(int threads, const std::vector<int> &cpus);
|
||||
~Pool();
|
||||
void run(std::vector<std::function<void()>> &jobs);
|
||||
|
||||
private:
|
||||
void loop(int cpu);
|
||||
std::vector<std::thread> threads_;
|
||||
std::mutex mu_;
|
||||
std::condition_variable wake_, done_;
|
||||
std::vector<std::function<void()>> *jobs_ = nullptr;
|
||||
size_t next_ = 0, finished_ = 0;
|
||||
bool stop_ = false;
|
||||
};
|
||||
|
||||
struct Hand {
|
||||
int id = 0;
|
||||
V3 pts[21]{}; // as measured this frame; the tracker steers crops by these
|
||||
V3 smooth[21]{}; // filtered (One Euro, see Tracker::smooth): publish these
|
||||
bool has_pts = false;
|
||||
double residual = -1; // rms ray distance of the triangulation (m); -1: one view
|
||||
int nviews = 0;
|
||||
double right_score = 0.5; // the model's right-hand score (these images aren't mirrored)
|
||||
double scale = 1.0; // this user's hand size / the model's world landmarks
|
||||
int64_t seen_ns = 0;
|
||||
int frames = 0;
|
||||
double speed = 0; // palm centre, m/s, smoothed
|
||||
int64_t last_ns = 0;
|
||||
V3 last_palm{};
|
||||
V3 dpalm{}; // the filter's palm velocity, m/s
|
||||
int64_t smooth_ns = 0;
|
||||
bool right() const { return right_score > 0.5; }
|
||||
};
|
||||
|
||||
struct Stats {
|
||||
int palm_calls = 0, hand_calls = 0, sets = 0;
|
||||
double palm_ms = 0, hand_ms = 0, step_ms = 0; // summed batch times
|
||||
int palm_batches = 0, hand_batches = 0;
|
||||
// why views and hands come and go
|
||||
int lost = 0; // a tracked view's landmarks fell below min presence
|
||||
int handoff_miss = 0; // a view projected from the hand's 3D (new camera or retry) found no hand
|
||||
int dups = 0; // the same hand twice in one camera
|
||||
int splits = 0; // a hand's views disagreed in 3D and were split
|
||||
int created = 0, merged = 0, forgotten = 0; // merged: views re-paired across cameras
|
||||
// diagnostics: on stereo frames, each view's single-view palm distance / the stereo one
|
||||
std::vector<std::pair<int, double>> mono_ratio; // (hand id, ratio)
|
||||
};
|
||||
|
||||
// A hand the landmark model found in one camera (Tracker::views_now, Tracker::exhaustive).
|
||||
struct Seen {
|
||||
std::string cam;
|
||||
int hand = 0; // the tracker's hand; 0 in exhaustive()
|
||||
Roi roi;
|
||||
Landmarks lm;
|
||||
V3 wrist{}; // exhaustive(): single-view 3D guess at the model's hand size
|
||||
};
|
||||
|
||||
class Tracker {
|
||||
public:
|
||||
Tracker(const std::map<std::string, Camera> &cams, const Nets &nets, Pool &pool, int max_views = 2);
|
||||
// images: calibration name -> frame. Returns the hands seen in this set.
|
||||
std::vector<const Hand *> step(const std::map<std::string, Image> &images, int64_t t_ns);
|
||||
// Seconds until the next frame set is worth processing (30 Hz fast hands, 15 Hz slow, 5 Hz none).
|
||||
double interval() const;
|
||||
Stats stats;
|
||||
size_t views() const { return views_.size(); }
|
||||
std::vector<Seen> views_now() const;
|
||||
// Every search tile in every camera, then landmarks on every palm: slow; for checking
|
||||
// what the scheduler misses (fh-replay --oracle).
|
||||
std::vector<Seen> exhaustive(const std::map<std::string, Image> &images);
|
||||
// Landmark presence a tracked view needs to stay (new views need min presence, 0.5). In
|
||||
// bright rooms the camera exposes for the room, the hands come out dim, and presence
|
||||
// dips under 0.5 for a frame at a time.
|
||||
void set_keep_presence(double p) { keep_presence_ = p; }
|
||||
// One view's 3D hand: each landmark along its ray, as far as how big the palm looks says
|
||||
// for a hand `scale` times the model's (Hand::scale). False if the palm is degenerate.
|
||||
bool single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const;
|
||||
|
||||
private:
|
||||
struct View {
|
||||
const Camera *cam;
|
||||
Roi roi;
|
||||
int hand = 0; // 0: not assigned yet
|
||||
Landmarks lm;
|
||||
bool has_lm = false;
|
||||
int frames = 0;
|
||||
bool fresh = false; // lm is from this step
|
||||
};
|
||||
struct Tile {
|
||||
const Camera *cam;
|
||||
V2 center;
|
||||
double size, rotation, weight, credit = 0;
|
||||
};
|
||||
void run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views);
|
||||
bool hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns);
|
||||
double pair_cost(const View &a, const View &b) const;
|
||||
double size_misfit(const std::vector<const View *> &views, const V3 *pts) const;
|
||||
void associate();
|
||||
static void smooth(Hand &h, int64_t t_ns);
|
||||
bool inside(const Camera &cam, V2 uv) const;
|
||||
void add_tiles(const Camera &cam, double frac, int gx, int gy);
|
||||
|
||||
std::map<std::string, const Camera *> cams_;
|
||||
const Nets &nets_;
|
||||
Pool &pool_;
|
||||
int max_views_, hand_budget_ = 4, search_budget_ = 3;
|
||||
double min_presence_ = 0.5, keep_presence_ = 0.5;
|
||||
std::vector<View> views_;
|
||||
std::map<int, Hand> hands_;
|
||||
std::vector<Tile> tiles_;
|
||||
int next_id_ = 1;
|
||||
int64_t last_ns_ = 0, search_ns_ = 0;
|
||||
};
|
||||
@@ -0,0 +1,184 @@
|
||||
"""Tracking-camera calibration from the headset's factory files.
|
||||
|
||||
/persist/xrservice.json (written by Valve's calibration, loaded by XRService)
|
||||
holds, per camera, Kannala-Brandt fisheye intrinsics ("kb": fx fy cx cy k1-k4,
|
||||
pixel centres at integer coordinates, as in OpenCV's fisheye model) and a pose
|
||||
in the slam_right (Cam0) frame: plus_x/plus_z are the camera axes and position
|
||||
its origin, in mm. /persist/device_config.json gives Cam0's pose in the CAD
|
||||
frame (cv.cad_from_cal, metres) and the head's pose in CAD (head). The CAD frame
|
||||
is +X head-left, +Y up, +Z forward; the head frame is OpenVR's: +x right, +y up,
|
||||
-z forward. Camera frames: +z along the optical axis, +x right and +y down in
|
||||
the image.
|
||||
|
||||
Everything here returns metres in the head frame.
|
||||
"""
|
||||
import json
|
||||
|
||||
import numpy as np
|
||||
|
||||
XRSERVICE_JSON = '/persist/xrservice.json'
|
||||
DEVICE_JSON = '/persist/device_config.json'
|
||||
# The Arcturus color module's EEPROM: some binary, then its calibration as JSON (world-readable)
|
||||
ARCTURUS_EEPROM = '/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom'
|
||||
ARCTURUS_WIDTH = 1972 # valid pixels per row that XRService's buffers deliver (of 2464)
|
||||
|
||||
|
||||
def _pose(d, scale=1.0):
|
||||
"""4x4 transform from a {plus_x, plus_z, position} pose (child axes in the parent frame)."""
|
||||
x = np.asarray(d['plus_x'], float)
|
||||
z = np.asarray(d['plus_z'], float)
|
||||
y = np.cross(z, x)
|
||||
T = np.eye(4)
|
||||
T[:3, 0], T[:3, 1], T[:3, 2] = x, y, z
|
||||
T[:3, 3] = np.asarray(d['position'], float) * scale
|
||||
return T
|
||||
|
||||
|
||||
class Camera:
|
||||
def __init__(self, name, width, height, kb, head_from_cam):
|
||||
self.name = name
|
||||
self.width, self.height = width, height
|
||||
self.fx, self.fy, self.cx, self.cy = kb['fx'], kb['fy'], kb['cx'], kb['cy']
|
||||
self.k = np.array([kb['k1'], kb['k2'], kb['k3'], kb['k4']])
|
||||
self.head_from_cam = head_from_cam
|
||||
self.R = head_from_cam[:3, :3] # camera axes in the head frame
|
||||
self.origin = head_from_cam[:3, 3] # camera centre in the head frame
|
||||
|
||||
def __repr__(self):
|
||||
return 'Camera(%s %dx%d at %s mm)' % (self.name, self.width, self.height,
|
||||
np.round(self.origin * 1000, 1))
|
||||
|
||||
def _theta_d(self, theta):
|
||||
t2 = theta * theta
|
||||
k1, k2, k3, k4 = self.k
|
||||
return theta * (1 + t2 * (k1 + t2 * (k2 + t2 * (k3 + t2 * k4))))
|
||||
|
||||
def project_cam(self, p):
|
||||
"""Camera-frame points (N,3) -> pixels (N,2). Points behind the lens still map (the lens sees ~180 deg)."""
|
||||
p = np.atleast_2d(p)
|
||||
r = np.hypot(p[:, 0], p[:, 1])
|
||||
theta = np.arctan2(r, p[:, 2])
|
||||
scale = np.where(r > 1e-12, self._theta_d(theta) / np.maximum(r, 1e-12), 0.0)
|
||||
return np.stack([self.fx * p[:, 0] * scale + self.cx, self.fy * p[:, 1] * scale + self.cy], axis=1)
|
||||
|
||||
def unproject(self, uv):
|
||||
"""Pixels (N,2) -> unit rays (N,3) in the camera frame."""
|
||||
uv = np.atleast_2d(np.asarray(uv, float))
|
||||
mx = (uv[:, 0] - self.cx) / self.fx
|
||||
my = (uv[:, 1] - self.cy) / self.fy
|
||||
td = np.hypot(mx, my)
|
||||
theta = td.copy()
|
||||
k1, k2, k3, k4 = self.k
|
||||
for _ in range(8): # Newton on theta_d(theta) = td
|
||||
t2 = theta * theta
|
||||
f = self._theta_d(theta) - td
|
||||
df = 1 + t2 * (3 * k1 + t2 * (5 * k2 + t2 * (7 * k3 + t2 * 9 * k4)))
|
||||
theta = np.clip(theta - f / df, 0.0, np.pi)
|
||||
s = np.where(td > 1e-12, np.sin(theta) / np.maximum(td, 1e-12), 1.0)
|
||||
return np.stack([mx * s, my * s, np.cos(theta)], axis=1)
|
||||
|
||||
def rays(self, uv):
|
||||
"""Pixels -> unit rays in the head frame (all starting at self.origin)."""
|
||||
return self.unproject(uv) @ self.R.T
|
||||
|
||||
def project(self, p_head):
|
||||
"""Head-frame points (N,3) -> pixels (N,2) and depth along the optical axis (N,)."""
|
||||
p = (np.atleast_2d(p_head) - self.origin) @ self.R
|
||||
return self.project_cam(p), p[:, 2]
|
||||
|
||||
def angle_from_axis(self, uv):
|
||||
"""Angle in degrees between each pixel's ray and the optical axis."""
|
||||
return np.degrees(np.arccos(np.clip(self.unproject(uv)[:, 2], -1, 1)))
|
||||
|
||||
|
||||
def load(xrservice=XRSERVICE_JSON, device=DEVICE_JSON):
|
||||
"""{calibration name: Camera} for the tracking cameras, posed in the head frame."""
|
||||
with open(xrservice) as f:
|
||||
rig = json.load(f)
|
||||
with open(device) as f:
|
||||
dev = json.load(f)
|
||||
cad_from_cam0 = _pose(dev['cv']['cad_from_cal'])
|
||||
head_from_cad = np.linalg.inv(_pose(dev['head']))
|
||||
cams = {}
|
||||
for c in rig['cameras']:
|
||||
kb = next(i for i in c['intrinsics'] if i['cameraModel'] == 'kb')
|
||||
cam0_from_cam = _pose(c['extrinsics'], 1e-3)
|
||||
cams[c['sourceCamera']] = Camera(c['sourceCamera'], c['width'], c['height'], kb,
|
||||
head_from_cad @ cad_from_cam0 @ cam0_from_cam)
|
||||
return cams
|
||||
|
||||
|
||||
def load_color(eeprom=ARCTURUS_EEPROM, device=DEVICE_JSON, scale=2, crop='subtract'):
|
||||
"""{"passthrough_left"/"passthrough_right": Camera} for the Arcturus color cameras, posed in
|
||||
the head frame, for fh-camd --with-color's images (luma at 1/scale size).
|
||||
|
||||
Their calibration is in the CAD frame (mm) with pixel coordinates on the full 2464x2464
|
||||
sensor; each camera also has a cropRegion. crop says how that maps to the delivered
|
||||
image: 'subtract' (image x = sensor x - cropRegion.x) or 'none'. tools/check_color.py
|
||||
tells which fits.
|
||||
"""
|
||||
with open(eeprom, 'rb') as f:
|
||||
raw = f.read()
|
||||
i = raw.rfind(b'{', 0, raw.find(b'"alignment_method"'))
|
||||
rig, _ = json.JSONDecoder().raw_decode(raw[i:].decode('latin1'))
|
||||
with open(device) as f:
|
||||
dev = json.load(f)
|
||||
head_from_cad = np.linalg.inv(_pose(dev['head']))
|
||||
cams = {}
|
||||
for c in rig['cameras']:
|
||||
kb = dict(next(k for k in c['intrinsics'] if k['cameraModel'] == 'kb'))
|
||||
region = c.get('cropRegion', {}) if crop == 'subtract' else {}
|
||||
# integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
|
||||
kb['cx'] = (kb['cx'] - region.get('x', 0) + 0.5) / scale - 0.5
|
||||
kb['cy'] = (kb['cy'] - region.get('y', 0) + 0.5) / scale - 0.5
|
||||
kb['fx'] /= scale
|
||||
kb['fy'] /= scale
|
||||
cams[c['sourceCamera']] = Camera(c['sourceCamera'], ARCTURUS_WIDTH // scale, c['height'] // scale, kb,
|
||||
head_from_cad @ _pose(c['extrinsics'], 1e-3))
|
||||
return cams
|
||||
|
||||
|
||||
def triangulate(origins, dirs, weights=None):
|
||||
"""Least-squares point closest to several rays. Returns (point, rms distance to the rays)."""
|
||||
A = np.zeros((3, 3))
|
||||
b = np.zeros(3)
|
||||
w = np.ones(len(origins)) if weights is None else np.asarray(weights, float)
|
||||
for o, d, wi in zip(origins, dirs, w):
|
||||
P = np.eye(3) - np.outer(d, d)
|
||||
A += wi * P
|
||||
b += wi * P @ o
|
||||
p = np.linalg.solve(A, b)
|
||||
res = [np.linalg.norm((np.eye(3) - np.outer(d, d)) @ (p - o)) for o, d in zip(origins, dirs)]
|
||||
return p, float(np.sqrt(np.mean(np.square(res))))
|
||||
|
||||
|
||||
def triangulate_many(origins, dirs, weights):
|
||||
"""Triangulate K points seen from V cameras at once.
|
||||
|
||||
origins (V,3), dirs (V,K,3) unit rays, weights (V,). Returns points (K,3) and
|
||||
each point's rms distance to its rays (K,).
|
||||
"""
|
||||
P = np.eye(3) - dirs[..., :, None] * dirs[..., None, :] # (V,K,3,3)
|
||||
w = weights[:, None, None, None]
|
||||
A = (w * P).sum(0)
|
||||
b = (w * (P @ origins[:, None, :, None])).sum(0)[..., 0]
|
||||
pts = np.linalg.solve(A, b[..., None])[..., 0]
|
||||
off = pts[None] - origins[:, None, :] # (V,K,3)
|
||||
perp = off - (off * dirs).sum(-1, keepdims=True) * dirs
|
||||
return pts, np.sqrt((perp ** 2).sum(-1).mean(0))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
cams = load()
|
||||
for cam in cams.values():
|
||||
fwd = [float(v) for v in cam.R[:, 2]]
|
||||
print('%-12s at x %+6.1f y %+6.1f z %+6.1f mm, looks %s' % (
|
||||
cam.name, *(cam.origin * 1000),
|
||||
'right' * (fwd[0] > 0.3) + 'left' * (fwd[0] < -0.3) + ' up' * (fwd[1] > 0.3) +
|
||||
' down' * (fwd[1] < -0.3) + ' forward' * (fwd[2] < -0.3) + ' back' * (fwd[2] > 0.3)),
|
||||
np.round(fwd, 2))
|
||||
uv = np.array([[cam.cx + 200, cam.cy - 100], [cam.cx - 0.4 * cam.width, cam.cy + 0.3 * cam.height]])
|
||||
err = np.abs(cam.project_cam(cam.unproject(uv)) - uv).max()
|
||||
assert err < 1e-6, err
|
||||
a, b = cams['slam_left'], cams['slam_right']
|
||||
print('slam baseline %.2f mm' % (1000 * np.linalg.norm(a.origin - b.origin)))
|
||||
@@ -0,0 +1,300 @@
|
||||
"""MediaPipe's palm detector and hand landmark model on ncnn, CPU or Vulkan GPU.
|
||||
|
||||
The models are the OpenCV Zoo ONNX ports, converted by tools/convert_models.py.
|
||||
Crops are square regions of a camera image given as (centre, size, rotation)
|
||||
in pixels and radians; rotation turns the crop's "up" towards the image
|
||||
direction (sin r, -cos r), as in MediaPipe. Both models take RGB in [0, 1];
|
||||
mono crops are replicated to three planes.
|
||||
|
||||
Preparing crops and decoding outputs happen here; running the networks is an
|
||||
engine's job: Engine runs them in this process, Pool in worker processes
|
||||
(ncnn's Python binding holds the GIL while it infers, so threads don't help).
|
||||
"""
|
||||
import multiprocessing as mp
|
||||
import os
|
||||
from multiprocessing import connection, shared_memory
|
||||
|
||||
import cv2
|
||||
import ncnn
|
||||
import numpy as np
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
MODELS = os.path.join(HERE, '..', 'models', 'ncnn')
|
||||
|
||||
# MediaPipe hand_landmarks_to_rect: the palm and finger bases used for the next ROI
|
||||
ROI_LANDMARKS = [0, 1, 2, 3, 5, 6, 9, 10, 13, 14, 17, 18]
|
||||
|
||||
# per model: input size, output blobs and their sizes
|
||||
SPECS = {'palm': (192, [('out0', 2016 * 18), ('out1', 2016)]),
|
||||
'hand': (224, [('out0', 63), ('out1', 1), ('out2', 1), ('out3', 63)])}
|
||||
|
||||
|
||||
def load_net(name, gpu, threads=1):
|
||||
net = ncnn.Net()
|
||||
net.opt.use_vulkan_compute = gpu
|
||||
net.opt.num_threads = threads
|
||||
net.opt.use_fp16_packed = net.opt.use_fp16_storage = net.opt.use_fp16_arithmetic = True
|
||||
net.load_param(os.path.join(MODELS, name + '.ncnn.param'))
|
||||
net.load_model(os.path.join(MODELS, name + '.ncnn.bin'))
|
||||
return net
|
||||
|
||||
|
||||
def infer(net, kind, patch):
|
||||
"""Run one network on an 8-bit mono crop; returns its outputs, flattened."""
|
||||
plane = patch.astype(np.float32) * (1.0 / 255.0)
|
||||
x = np.ascontiguousarray(np.broadcast_to(plane, (3,) + plane.shape))
|
||||
ex = net.create_extractor()
|
||||
ex.input('in0', ncnn.Mat(x))
|
||||
return [np.array(ex.extract(name)[1], np.float32).reshape(-1) for name, _ in SPECS[kind][1]]
|
||||
|
||||
|
||||
class Engine:
|
||||
"""Runs the networks in this process, one after another."""
|
||||
|
||||
def __init__(self, palm_gpu=False, hand_gpu=False):
|
||||
self.nets = {'palm': load_net('palm', palm_gpu), 'hand': load_net('hand', hand_gpu)}
|
||||
|
||||
def run(self, jobs):
|
||||
"""jobs: [(kind, patch)] -> [outputs]"""
|
||||
return [infer(self.nets[kind], kind, patch) for kind, patch in jobs]
|
||||
|
||||
def close(self):
|
||||
pass
|
||||
|
||||
|
||||
def _worker(index, shm_name, conn, palm_gpu, hand_gpu, cpu):
|
||||
if cpu is not None and cpu in os.sched_getaffinity(0):
|
||||
os.sched_setaffinity(0, {cpu})
|
||||
nets = {'palm': load_net('palm', palm_gpu), 'hand': load_net('hand', hand_gpu)}
|
||||
shm = shared_memory.SharedMemory(name=shm_name)
|
||||
base = index * Pool.STRIDE
|
||||
while True:
|
||||
msg = conn.recv()
|
||||
if msg is None:
|
||||
break
|
||||
kind = msg
|
||||
size = SPECS[kind][0]
|
||||
patch = np.ndarray((size, size), np.uint8, shm.buf, base)
|
||||
outs = infer(nets[kind], kind, patch)
|
||||
off = base + Pool.IN_BYTES
|
||||
for o in outs:
|
||||
np.ndarray(o.shape, np.float32, shm.buf, off)[:] = o
|
||||
off += o.nbytes
|
||||
conn.send(True)
|
||||
shm.close()
|
||||
|
||||
|
||||
class Pool:
|
||||
"""Runs the networks in worker processes pinned to CPUs, several crops at once."""
|
||||
IN_BYTES = 224 * 224
|
||||
OUT_BYTES = 4 * (2016 * 18 + 2016)
|
||||
STRIDE = IN_BYTES + OUT_BYTES
|
||||
|
||||
# SteamOS on the Frame keeps user processes on CPUs 0-4 (5-7 carry pinned VR threads);
|
||||
# 2-4 are the big cores among those
|
||||
def __init__(self, workers=2, cpus=(2, 3, 4), palm_gpu=False, hand_gpu=False):
|
||||
ctx = mp.get_context('spawn')
|
||||
self.shm = shared_memory.SharedMemory(create=True, size=workers * self.STRIDE)
|
||||
self.conns, self.procs = [], []
|
||||
for i in range(workers):
|
||||
a, b = ctx.Pipe()
|
||||
cpu = cpus[i % len(cpus)] if cpus else None
|
||||
p = ctx.Process(target=_worker, args=(i, self.shm.name, b, palm_gpu, hand_gpu, cpu), daemon=True)
|
||||
p.start()
|
||||
self.conns.append(a)
|
||||
self.procs.append(p)
|
||||
|
||||
def run(self, jobs):
|
||||
results = [None] * len(jobs)
|
||||
pending = list(range(len(jobs)))
|
||||
busy = {} # conn -> job index
|
||||
free = list(range(len(self.conns)))
|
||||
while pending or busy:
|
||||
while pending and free:
|
||||
w = free.pop()
|
||||
j = pending.pop(0)
|
||||
kind, patch = jobs[j]
|
||||
base = w * self.STRIDE
|
||||
np.ndarray(patch.shape, np.uint8, self.shm.buf, base)[:] = patch
|
||||
self.conns[w].send(kind)
|
||||
busy[self.conns[w]] = (w, j)
|
||||
for c in connection.wait(list(busy)):
|
||||
c.recv()
|
||||
w, j = busy.pop(c)
|
||||
kind = jobs[j][0]
|
||||
off = w * self.STRIDE + self.IN_BYTES
|
||||
outs = []
|
||||
for _, n in SPECS[kind][1]:
|
||||
outs.append(np.ndarray((n,), np.float32, self.shm.buf, off).copy())
|
||||
off += 4 * n
|
||||
results[j] = outs
|
||||
free.append(w)
|
||||
return results
|
||||
|
||||
def pids(self):
|
||||
return [p.pid for p in self.procs]
|
||||
|
||||
def close(self):
|
||||
for c in self.conns:
|
||||
try:
|
||||
c.send(None)
|
||||
except OSError:
|
||||
pass
|
||||
for p in self.procs:
|
||||
p.join(1)
|
||||
self.shm.close()
|
||||
self.shm.unlink()
|
||||
|
||||
|
||||
def crop_matrix(center, size, rotation, out):
|
||||
"""2x3 affine taking crop pixels (0..out) to image pixels."""
|
||||
c, s = np.cos(rotation), np.sin(rotation)
|
||||
k = size / out
|
||||
R = np.array([[c, -s], [s, c]]) * k
|
||||
t = np.asarray(center, float) - R @ np.array([out / 2.0, out / 2.0])
|
||||
return np.hstack([R, t[:, None]])
|
||||
|
||||
|
||||
# Local contrast per crop: the IR frames are dim and uneven (the upper cameras
|
||||
# especially). On the 2026-09-28 capture this found hands in more frames on every
|
||||
# camera than plain, stretched or gamma-corrected crops.
|
||||
CLAHE = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(4, 4))
|
||||
|
||||
|
||||
def crop(gray, M, out):
|
||||
"""8-bit square crop of a mono image, with a crop->image matrix, contrast-equalized."""
|
||||
patch = cv2.warpAffine(gray, M, (out, out), flags=cv2.INTER_LINEAR | cv2.WARP_INVERSE_MAP,
|
||||
borderMode=cv2.BORDER_CONSTANT, borderValue=0)
|
||||
return CLAHE.apply(patch)
|
||||
|
||||
|
||||
def to_image(M, pts):
|
||||
"""Crop pixels (N,2) -> image pixels."""
|
||||
return pts @ M[:, :2].T + M[:, 2]
|
||||
|
||||
|
||||
def normalize_angle(a):
|
||||
return (a + np.pi) % (2 * np.pi) - np.pi
|
||||
|
||||
|
||||
def _anchors():
|
||||
"""SSD anchors of palm_detection_full: strides 8 (2 per cell) and 16 (6 per cell), 192x192."""
|
||||
out = []
|
||||
for stride, per_cell in ((8, 2), (16, 6)):
|
||||
n = 192 // stride
|
||||
for y in range(n):
|
||||
for x in range(n):
|
||||
out += [((x + 0.5) / n, (y + 0.5) / n)] * per_cell
|
||||
return np.array(out, np.float32)
|
||||
|
||||
|
||||
class Detection:
|
||||
"""A palm in image pixels: box centre/size, 7 keypoints, score."""
|
||||
__slots__ = ('center', 'size', 'keypoints', 'score')
|
||||
|
||||
def __init__(self, center, size, keypoints, score):
|
||||
self.center, self.size, self.keypoints, self.score = center, size, keypoints, score
|
||||
|
||||
def roi(self):
|
||||
"""MediaPipe's hand ROI from a palm: wrist->middle-finger-base sets the rotation, then
|
||||
shift 0.5 towards the fingers and scale the square box 2.6x."""
|
||||
(x0, y0), (x1, y1) = self.keypoints[0], self.keypoints[2]
|
||||
rot = normalize_angle(np.pi / 2 - np.arctan2(-(y1 - y0), x1 - x0))
|
||||
w, h = (float(v) for v in self.size)
|
||||
shift = np.array([-h * -0.5 * np.sin(rot), h * -0.5 * np.cos(rot)])
|
||||
return np.asarray(self.center) + shift, max(w, h) * 2.6, rot
|
||||
|
||||
|
||||
class PalmDetector:
|
||||
SIZE = 192
|
||||
|
||||
def __init__(self, min_score=0.5):
|
||||
self.anchors = _anchors()
|
||||
self.min_score = min_score
|
||||
|
||||
def prepare(self, gray, center, size, rotation=0.0):
|
||||
M = crop_matrix(center, size, rotation, self.SIZE)
|
||||
return crop(gray, M, self.SIZE), (M, size)
|
||||
|
||||
def decode(self, outputs, ctx):
|
||||
"""Palms found in one crop, in image pixels."""
|
||||
M, size = ctx
|
||||
raw = outputs[0].reshape(-1, 18)
|
||||
logit = outputs[1]
|
||||
keep = np.nonzero(logit > np.log(self.min_score / (1 - self.min_score)))[0]
|
||||
if not len(keep):
|
||||
return []
|
||||
a = self.anchors[keep] * self.SIZE
|
||||
r = raw[keep]
|
||||
centers = r[:, 0:2] + a
|
||||
sizes = r[:, 2:4]
|
||||
kps = r[:, 4:18].reshape(-1, 7, 2) + a[:, None, :]
|
||||
scores = 1 / (1 + np.exp(-np.clip(logit[keep], -100, 100)))
|
||||
return [Detection(to_image(M, c[None])[0], s * size / self.SIZE, to_image(M, k), sc)
|
||||
for c, s, k, sc in _weighted_nms(centers, sizes, scores, kps)]
|
||||
|
||||
|
||||
def _weighted_nms(centers, sizes, scores, kps, iou_thresh=0.3):
|
||||
"""MediaPipe's weighted NMS: overlapping boxes are averaged, weighted by score."""
|
||||
order = np.argsort(-scores)
|
||||
boxes = np.hstack([centers - sizes / 2, centers + sizes / 2])
|
||||
area = np.prod(sizes, axis=1)
|
||||
out = []
|
||||
while len(order):
|
||||
i = order[0]
|
||||
xy0 = np.maximum(boxes[i, :2], boxes[order, :2])
|
||||
xy1 = np.minimum(boxes[i, 2:], boxes[order, 2:])
|
||||
inter = np.prod(np.clip(xy1 - xy0, 0, None), axis=1)
|
||||
iou = inter / (area[i] + area[order] - inter + 1e-9)
|
||||
group = order[iou > iou_thresh]
|
||||
w = scores[group][:, None]
|
||||
out.append(((centers[group] * w).sum(0) / w.sum(), (sizes[group] * w).sum(0) / w.sum(),
|
||||
(kps[group] * w[:, :, None]).sum(0) / w.sum(), float(scores[i])))
|
||||
order = order[iou <= iou_thresh]
|
||||
return out
|
||||
|
||||
|
||||
class Landmarks:
|
||||
"""21 hand landmarks in image pixels, plus MediaPipe's metric 'world' landmarks."""
|
||||
__slots__ = ('pts', 'depth', 'world', 'presence', 'right', 'roi')
|
||||
|
||||
def __init__(self, pts, depth, world, presence, right, roi):
|
||||
self.pts, self.depth, self.world, self.presence = pts, depth, world, presence
|
||||
self.right, self.roi = right, roi
|
||||
|
||||
def next_roi(self):
|
||||
return roi_from_points(self.pts)
|
||||
|
||||
|
||||
def roi_from_points(p):
|
||||
"""MediaPipe's hand_landmarks_to_rect: the crop to track a hand given its landmarks."""
|
||||
x0, y0 = p[0]
|
||||
x1, y1 = ((p[5] + p[13]) / 2 + p[9]) / 2
|
||||
rot = normalize_angle(np.pi / 2 - np.arctan2(-(y1 - y0), x1 - x0))
|
||||
sub = p[ROI_LANDMARKS]
|
||||
center = (sub.min(0) + sub.max(0)) / 2
|
||||
c, s = np.cos(-rot), np.sin(-rot)
|
||||
q = (sub - center) @ np.array([[c, -s], [s, c]]).T
|
||||
lo, hi = q.min(0), q.max(0)
|
||||
mid = (lo + hi) / 2
|
||||
c2, s2 = np.cos(rot), np.sin(rot)
|
||||
center = center + np.array([mid[0] * c2 - mid[1] * s2, mid[0] * s2 + mid[1] * c2])
|
||||
w, h = hi - lo
|
||||
center = center + np.array([-h * -0.1 * s2, h * -0.1 * c2])
|
||||
return center, max(w, h) * 2.0, rot
|
||||
|
||||
|
||||
class HandLandmarker:
|
||||
SIZE = 224
|
||||
|
||||
def prepare(self, gray, roi):
|
||||
center, size, rotation = roi
|
||||
M = crop_matrix(center, size, rotation, self.SIZE)
|
||||
return crop(gray, M, self.SIZE), (M, roi)
|
||||
|
||||
def decode(self, outputs, ctx):
|
||||
M, roi = ctx
|
||||
screen = outputs[0].reshape(21, 3)
|
||||
pts = to_image(M, screen[:, :2])
|
||||
depth = screen[:, 2] * roi[1] / self.SIZE # relative depth, image pixels
|
||||
return Landmarks(pts, depth, outputs[3].reshape(21, 3), float(outputs[1][0]), float(outputs[2][0]), roi)
|
||||
Reference in new issue
Block a user