mirror of
https://github.com/DeeJanuz/frametop.git
synced 2026-10-06 04:04:09 +02:00
Add colour cameras, pinch gestures, and depth measures
- fh-camd --with-color publishes the Arcturus colour pair (luma, half size, 30 fps) to the ring, flagged FH_CAM_COLOR. fh-tracker records them, and fh-replay --cams mono|color|all tracks with them, using the module's EEPROM calibration (load_color_calibration). tools/check_color.py checks which node is left and how the crop maps. - Pinch detection per hand (trackd/pinch.h), published to $XDG_RUNTIME_DIR/frame-hands/gestures (include/fh_gestures.h). It has begin and end counters, times, the pinch point, and the begin point for drags. tools/watch_gestures.py shows it live, and fh-replay reports it. - fh-replay --depth and tools/depth_report.py measure the depth without ground truth: noise along the line of sight against across it, the one-camera guess, and a simulated camera loss. - fh-tracker --swap-sides, --ring, --keep-presence, --record-only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
1 parent
3022d7de9d
commit
2031b9ca8f
19 files changed
+1141
-60
No files matched your search
+9
-2
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
Root-side camera access for frame-hands:
|
Root-side camera access for frame-hands:
|
||||||
|
|
||||||
- `fh-camd`: the frame broker. It publishes the four IR tracking cameras to a shared-memory ring that the unprivileged tracker reads.
|
- `fh-camd`: the frame broker. It publishes the four IR tracking cameras, and optionally the Arcturus color pair, to a shared-memory ring that the unprivileged tracker reads.
|
||||||
- `fh-camprobe`: a test tool that records timestamped frames from every camera, including the Arcturus color pair, with CSV logs.
|
- `fh-camprobe`: a test tool that records timestamped frames from every camera, including the Arcturus color pair, with CSV logs.
|
||||||
|
|
||||||
## How it gets frames
|
## How it gets frames
|
||||||
@@ -33,6 +33,13 @@ Frames go to `/run/frame-hands/ir-ring`. The file is mode 0600 and owned by the
|
|||||||
- A copy torn by the camera overwriting the buffer is dropped.
|
- A copy torn by the camera overwriting the buffer is dropped.
|
||||||
- Each copy takes about 0.1 ms, and a cache sync about 0.15 ms.
|
- Each copy takes about 0.1 ms, and a cache sync about 0.15 ms.
|
||||||
|
|
||||||
|
Options:
|
||||||
|
|
||||||
|
- `--with-dark`: also publish the near-black frames, as extra ring cameras flagged `FH_CAM_DARK`. They show only light sources, so they're no use for hands.
|
||||||
|
- `--with-color`: also publish the two Arcturus color cameras, flagged `FH_CAM_COLOR`. Each is the luma of the 10-bit frame's valid 1972x2464 (the top 8 bits), at half size (`--color-scale 2`: 986x1232) and at most 30 fps (`--color-fps`; the cameras run at 60). Frames that carry the module's warped half-size copy are dropped. Their `capture_ns` is on the color module's clock (2.2 s off the mono cameras' on 2026-09-29), so line them up with the mono cameras by `dqbuf_ns`. Each frame costs about 0.65 ms of cache sync and 1.1 ms of decoding, so both cameras at 30 fps take about 11% of a core.
|
||||||
|
- The ring holds 8 cameras: 4 mono, plus 4 dark twins or 2 color cameras.
|
||||||
|
- `--sensor S`: only the mono cameras whose sensor name contains S.
|
||||||
|
|
||||||
It exits when XRService exits, or when a camera's buffers keep going stale, which means XRService has reallocated them. Start it again to re-attach.
|
It exits when XRService exits, or when a camera's buffers keep going stale, which means XRService has reallocated them. Start it again to re-attach.
|
||||||
|
|
||||||
## fh-camprobe
|
## fh-camprobe
|
||||||
@@ -51,6 +58,6 @@ Output goes to `~/Pictures/framecap/stereo-<time>/`:
|
|||||||
- `<model>_NNNN_<camera>.pgm`: saved bright pairs. The color cameras are saved raw as `.yuv420_10p` (`tools/decode.py` reads them).
|
- `<model>_NNNN_<camera>.pgm`: saved bright pairs. The color cameras are saved raw as `.yuv420_10p` (`tools/decode.py` reads them).
|
||||||
- `plane1_*.bin`: raw plane 1 of a few frames, which may hold sensor metadata.
|
- `plane1_*.bin`: raw plane 1 of a few frames, which may hold sensor metadata.
|
||||||
|
|
||||||
Which camera is which: video9 is `slam_left`, video13 is `slam_right`, video6 is `upper_left` and video7 is `upper_right`. This was checked by rendering the same view from each camera with the factory calibration. `tracker/live.py` maps them by capture pipe (`/sys/class/video4linux/videoN/name`).
|
Which camera is which: video9 is `slam_left`, video13 is `slam_right`, video6 is `upper_left` and video7 is `upper_right`. This was checked by rendering the same view from each camera with the factory calibration. But fh-camd tells the side cameras' buffers apart only by XRService's allocation order, and after some XRService restarts it gets them backwards: check with `tools/check_sides.py --ring` and run the tracker with `--swap-sides` when it says swapped. The color cameras are video3 (`arcimx616 0-0010`) and video0 (`0-001a`); which of them is `passthrough_left` in the module's calibration is for `tools/check_color.py` to settle on a recording with texture in view. `tracker/live.py` maps them by capture pipe (`/sys/class/video4linux/videoN/name`).
|
||||||
|
|
||||||
`discovery` in `xrcams.c` is adapted from FrameEyeCameraFeed (MIT, see `LICENSE.FrameEyeCameraFeed`).
|
`discovery` in `xrcams.c` is adapted from FrameEyeCameraFeed (MIT, see `LICENSE.FrameEyeCameraFeed`).
|
||||||
+155
-15
@@ -56,6 +56,7 @@
|
|||||||
#define MAX_SLOTS 64
|
#define MAX_SLOTS 64
|
||||||
#define MAX_INDEX 32
|
#define MAX_INDEX 32
|
||||||
#define NSAMP 512
|
#define NSAMP 512
|
||||||
|
#define SEAM_LIMIT 2.0 /* seam_score above this: the frame carries the half-size copy */
|
||||||
#define STALE_RELEARN 30 /* consecutive unchanged frames: the mapping changed */
|
#define STALE_RELEARN 30 /* consecutive unchanged frames: the mapping changed */
|
||||||
#define MAX_RELEARNS 5 /* then assume XRService has new buffers, and exit */
|
#define MAX_RELEARNS 5 /* then assume XRService has new buffers, and exit */
|
||||||
#define RING_DIR FH_RING_DIR
|
#define RING_DIR FH_RING_DIR
|
||||||
@@ -94,6 +95,10 @@ typedef struct {
|
|||||||
uint8_t *ring_slots;
|
uint8_t *ring_slots;
|
||||||
uint64_t frame_no;
|
uint64_t frame_no;
|
||||||
fh_ring_cam_t *rc_dark; /* --with-dark: its near-black frames */
|
fh_ring_cam_t *rc_dark; /* --with-dark: its near-black frames */
|
||||||
|
bool color; /* Arcturus color: luma, downscaled, published */
|
||||||
|
unsigned out_w, out_h; /* image size in the ring */
|
||||||
|
uint64_t last_pub_ns; /* --color-fps pacing */
|
||||||
|
uint64_t paced, seams; /* color frames skipped: pacing, the half-size copy */
|
||||||
uint8_t *ring_dark;
|
uint8_t *ring_dark;
|
||||||
uint64_t dark_no;
|
uint64_t dark_no;
|
||||||
|
|
||||||
@@ -115,6 +120,9 @@ static const char *opt_sensor = "";
|
|||||||
static const char *opt_user = NULL;
|
static const char *opt_user = NULL;
|
||||||
static double opt_dark = 0.4;
|
static double opt_dark = 0.4;
|
||||||
static bool opt_with_dark; /* also publish the near-black frames */
|
static bool opt_with_dark; /* also publish the near-black frames */
|
||||||
|
static bool opt_with_color; /* also publish the Arcturus color cameras */
|
||||||
|
static unsigned opt_color_scale = 2; /* ... at 1/N size */
|
||||||
|
static double opt_color_fps = 30; /* ... at most this often */
|
||||||
static double opt_status = 10.0;
|
static double opt_status = 10.0;
|
||||||
|
|
||||||
static volatile sig_atomic_t stop;
|
static volatile sig_atomic_t stop;
|
||||||
@@ -218,6 +226,9 @@ static void setup_camera(cam_t *c, xr_camera_t *cam, int pidfd)
|
|||||||
c->cam = cam;
|
c->cam = cam;
|
||||||
xr_camera_layout(cam, &c->lay);
|
xr_camera_layout(cam, &c->lay);
|
||||||
c->need = (size_t)c->lay.pitch * c->lay.rows;
|
c->need = (size_t)c->lay.pitch * c->lay.rows;
|
||||||
|
c->color = c->lay.fmt == XR_FMT_YUV420_10P;
|
||||||
|
c->out_w = c->color ? c->lay.width / opt_color_scale : c->lay.width;
|
||||||
|
c->out_h = c->color ? c->lay.height / opt_color_scale : c->lay.height;
|
||||||
|
|
||||||
char sensor[XR_SENSOR_LEN];
|
char sensor[XR_SENSOR_LEN];
|
||||||
xr_slugify(cam->sensor, sensor, sizeof(sensor));
|
xr_slugify(cam->sensor, sensor, sizeof(sensor));
|
||||||
@@ -518,6 +529,60 @@ static int find_fresh(cam_t *c, uint64_t *out)
|
|||||||
|
|
||||||
/* ------------------------------------------------------------- publishing */
|
/* ------------------------------------------------------------- publishing */
|
||||||
|
|
||||||
|
/*
|
||||||
|
* A color frame's luma into the ring: the top 8 bits of every Nth pixel of every Nth row
|
||||||
|
* (MIPI RAW10 packs 4 pixels in 5 bytes, the high bytes first). Returns the mean of a
|
||||||
|
* sparse grid of the output.
|
||||||
|
*/
|
||||||
|
static double decode_luma(const cam_t *c, const uint8_t *src, uint8_t *dst, unsigned stride)
|
||||||
|
{
|
||||||
|
unsigned s = opt_color_scale;
|
||||||
|
uint64_t sum = 0, n = 0;
|
||||||
|
|
||||||
|
for (unsigned y = 0; y < c->out_h; y++) {
|
||||||
|
|
||||||
|
const uint8_t *row = src + (size_t)y * s * c->lay.pitch;
|
||||||
|
uint8_t *out = dst + (size_t)y * stride;
|
||||||
|
|
||||||
|
for (unsigned x = 0; x < c->out_w; x++) {
|
||||||
|
unsigned sx = x * s;
|
||||||
|
out[x] = row[(sx >> 2) * 5 + (sx & 3)];
|
||||||
|
}
|
||||||
|
|
||||||
|
if (y % 8 == 0)
|
||||||
|
for (unsigned x = 0; x < c->out_w; x += 8, n++)
|
||||||
|
sum += out[x];
|
||||||
|
}
|
||||||
|
|
||||||
|
return n ? (double)sum / (double)n : 0.0;
|
||||||
|
}
|
||||||
|
|
||||||
|
/*
|
||||||
|
* The color module sometimes writes a warped half-size copy of the image into the
|
||||||
|
* top-left quarter of its buffers. Its bottom edge is a seam between the middle rows
|
||||||
|
* in the left half: this is ~1 for a clean frame and well above for one carrying the
|
||||||
|
* copy (from fh-camprobe --seam).
|
||||||
|
*/
|
||||||
|
static double seam_score(const cam_t *c, const uint8_t *p)
|
||||||
|
{
|
||||||
|
const xr_layout_t *l = &c->lay;
|
||||||
|
unsigned r = l->height / 2 - 1;
|
||||||
|
double across = 0, below = 0;
|
||||||
|
int n = 0;
|
||||||
|
|
||||||
|
for (unsigned x = 8; x < l->width / 2 - 8; x += 2, n++) {
|
||||||
|
|
||||||
|
size_t off = (x / 4) * 5 + (x % 4);
|
||||||
|
int a = p[(size_t)r * l->pitch + off], b = p[(size_t)(r + 1) * l->pitch + off];
|
||||||
|
int d = p[(size_t)(r + 2) * l->pitch + off];
|
||||||
|
|
||||||
|
across += abs(a - b);
|
||||||
|
below += abs(b - d);
|
||||||
|
}
|
||||||
|
|
||||||
|
return n ? (across / n + 0.5) / (below / n + 0.5) : 0;
|
||||||
|
}
|
||||||
|
|
||||||
/* Copy a frame into ring camera rc (the camera's own, or its dark twin). */
|
/* Copy a frame into ring camera rc (the camera's own, or its dark twin). */
|
||||||
static void publish(cam_t *c, fh_ring_cam_t *rc, uint8_t *slots, uint64_t *frame_no,
|
static void publish(cam_t *c, fh_ring_cam_t *rc, uint8_t *slots, uint64_t *frame_no,
|
||||||
const uint8_t *src, uint32_t seq, uint64_t ts, uint64_t evtime, double mean)
|
const uint8_t *src, uint32_t seq, uint64_t ts, uint64_t evtime, double mean)
|
||||||
@@ -532,8 +597,11 @@ static void publish(cam_t *c, fh_ring_cam_t *rc, uint8_t *slots, uint64_t *frame
|
|||||||
|
|
||||||
sample_words(src, c->need, before);
|
sample_words(src, c->need, before);
|
||||||
|
|
||||||
for (unsigned y = 0; y < c->lay.height; y++)
|
if (c->color)
|
||||||
memcpy(dst + (size_t)y * rc->stride, src + (size_t)y * c->lay.pitch, c->lay.width);
|
mean = decode_luma(c, src, dst, rc->stride);
|
||||||
|
else
|
||||||
|
for (unsigned y = 0; y < c->lay.height; y++)
|
||||||
|
memcpy(dst + (size_t)y * rc->stride, src + (size_t)y * c->lay.pitch, c->lay.width);
|
||||||
|
|
||||||
sample_words(src, c->need, after);
|
sample_words(src, c->need, after);
|
||||||
|
|
||||||
@@ -576,6 +644,12 @@ static void on_frame(cam_t *c, int64_t index, uint32_t seq, uint64_t ts, uint64_
|
|||||||
|
|
||||||
int slot = c->slot_of[index];
|
int slot = c->slot_of[index];
|
||||||
|
|
||||||
|
/* color runs at 60 fps: skip frames early enough that --color-fps holds, before any sync */
|
||||||
|
if (c->color && opt_color_fps > 0 && evtime - c->last_pub_ns < (uint64_t)(1e9 / opt_color_fps) - 3000000) {
|
||||||
|
c->paced++;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
static uint64_t cur[NSAMP];
|
static uint64_t cur[NSAMP];
|
||||||
uint64_t t0 = mono_ns();
|
uint64_t t0 = mono_ns();
|
||||||
|
|
||||||
@@ -619,6 +693,21 @@ static void on_frame(cam_t *c, int64_t index, uint32_t seq, uint64_t ts, uint64_
|
|||||||
memcpy(c->hist[0], cur, sizeof(cur));
|
memcpy(c->hist[0], cur, sizeof(cur));
|
||||||
*filled_at(c, slot) = ++nfilled;
|
*filled_at(c, slot) = ++nfilled;
|
||||||
|
|
||||||
|
if (c->color) { /* no dark frames here; skip the ones carrying the half-size copy */
|
||||||
|
if (seam_score(c, c->map[slot]) > SEAM_LIMIT) {
|
||||||
|
c->seams++;
|
||||||
|
c->rc->dropped++;
|
||||||
|
} else {
|
||||||
|
uint64_t t1 = mono_ns();
|
||||||
|
publish(c, c->rc, c->ring_slots, &c->frame_no, c->map[slot], seq, ts, evtime, 0);
|
||||||
|
c->copy_ns += mono_ns() - t1;
|
||||||
|
c->bright++;
|
||||||
|
c->last_pub_ns = evtime;
|
||||||
|
}
|
||||||
|
buf_sync(c->fd[slot], DMA_BUF_SYNC_END);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
/*
|
/*
|
||||||
* The cameras alternate a normal exposure with a near-black one, so judge
|
* The cameras alternate a normal exposure with a near-black one, so judge
|
||||||
* each frame against this camera's recent brightest.
|
* each frame against this camera's recent brightest.
|
||||||
@@ -661,6 +750,29 @@ static void on_sample(void *ctx, const tp_sample_t *s)
|
|||||||
|
|
||||||
/* ------------------------------------------------------------ ring + user */
|
/* ------------------------------------------------------------ ring + user */
|
||||||
|
|
||||||
|
/*
|
||||||
|
* Ring cameras: every camera (mono and color) at its own index, then with --with-dark a
|
||||||
|
* dark twin for each mono camera. Returns how many, with which camera each one shows.
|
||||||
|
*/
|
||||||
|
static int ring_layout(int cam_of[], bool dark_of[], int max)
|
||||||
|
{
|
||||||
|
int n = 0;
|
||||||
|
|
||||||
|
for (int i = 0; i < ncams; i++) {
|
||||||
|
if (n < max) { cam_of[n] = i; dark_of[n] = false; }
|
||||||
|
n++;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (opt_with_dark)
|
||||||
|
for (int i = 0; i < ncams; i++)
|
||||||
|
if (!cams[i].color) {
|
||||||
|
if (n < max) { cam_of[n] = i; dark_of[n] = true; }
|
||||||
|
n++;
|
||||||
|
}
|
||||||
|
|
||||||
|
return n;
|
||||||
|
}
|
||||||
|
|
||||||
/* The ring lives in a root-owned directory, so nobody can plant a file or link there. */
|
/* The ring lives in a root-owned directory, so nobody can plant a file or link there. */
|
||||||
static uint8_t *ring_create(uid_t uid, gid_t gid, size_t *len_out)
|
static uint8_t *ring_create(uid_t uid, gid_t gid, size_t *len_out)
|
||||||
{
|
{
|
||||||
@@ -673,11 +785,13 @@ static uint8_t *ring_create(uid_t uid, gid_t gid, size_t *len_out)
|
|||||||
die("%s must be a directory owned by root and writable only by root", RING_DIR);
|
die("%s must be a directory owned by root and writable only by root", RING_DIR);
|
||||||
|
|
||||||
size_t len = sizeof(fh_ring_hdr_t);
|
size_t len = sizeof(fh_ring_hdr_t);
|
||||||
int nring = opt_with_dark ? 2 * ncams : ncams;
|
int cam_of[FH_RING_MAX_CAMS];
|
||||||
|
bool dark_of[FH_RING_MAX_CAMS];
|
||||||
|
int nring = ring_layout(cam_of, dark_of, FH_RING_MAX_CAMS);
|
||||||
|
|
||||||
for (int i = 0; i < nring; i++) {
|
for (int i = 0; i < nring; i++) {
|
||||||
cam_t *c = &cams[i % ncams];
|
cam_t *c = &cams[cam_of[i]];
|
||||||
size_t slot = sizeof(fh_ring_slot_t) + (size_t)c->lay.width * c->lay.height;
|
size_t slot = sizeof(fh_ring_slot_t) + (size_t)c->out_w * c->out_h;
|
||||||
len += FH_RING_SLOTS * ((slot + 63) & ~(size_t)63);
|
len += FH_RING_SLOTS * ((slot + 63) & ~(size_t)63);
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -703,18 +817,18 @@ static uint8_t *ring_create(uid_t uid, gid_t gid, size_t *len_out)
|
|||||||
|
|
||||||
for (int i = 0; i < nring; i++) {
|
for (int i = 0; i < nring; i++) {
|
||||||
|
|
||||||
cam_t *c = &cams[i % ncams];
|
cam_t *c = &cams[cam_of[i]];
|
||||||
fh_ring_cam_t *rc = &h->cams[i];
|
fh_ring_cam_t *rc = &h->cams[i];
|
||||||
bool dark = i >= ncams;
|
bool dark = dark_of[i];
|
||||||
|
|
||||||
snprintf(rc->sensor, sizeof(rc->sensor), "%s", c->cam->sensor);
|
snprintf(rc->sensor, sizeof(rc->sensor), "%s", c->cam->sensor);
|
||||||
snprintf(rc->name, sizeof(rc->name), "%.26s%s", c->slug, dark ? "-dark" : "");
|
snprintf(rc->name, sizeof(rc->name), "%.26s%s", c->slug, dark ? "-dark" : "");
|
||||||
rc->flags = dark ? FH_CAM_DARK : 0;
|
rc->flags = dark ? FH_CAM_DARK : c->color ? FH_CAM_COLOR : 0;
|
||||||
rc->node = c->cam->node;
|
rc->node = c->cam->node;
|
||||||
rc->format = FH_FMT_GREY8;
|
rc->format = FH_FMT_GREY8;
|
||||||
rc->width = c->lay.width;
|
rc->width = c->out_w;
|
||||||
rc->height = c->lay.height;
|
rc->height = c->out_h;
|
||||||
rc->stride = c->lay.width;
|
rc->stride = c->out_w;
|
||||||
rc->nslots = FH_RING_SLOTS;
|
rc->nslots = FH_RING_SLOTS;
|
||||||
rc->slot_offset = off;
|
rc->slot_offset = off;
|
||||||
rc->slot_bytes = (sizeof(fh_ring_slot_t) + (size_t)rc->stride * rc->height + 63) & ~(size_t)63;
|
rc->slot_bytes = (sizeof(fh_ring_slot_t) + (size_t)rc->stride * rc->height + 63) & ~(size_t)63;
|
||||||
@@ -779,6 +893,15 @@ static void status(double secs)
|
|||||||
|
|
||||||
for (int i = 0; i < ncams; i++) {
|
for (int i = 0; i < ncams; i++) {
|
||||||
cam_t *c = &cams[i];
|
cam_t *c = &cams[i];
|
||||||
|
if (c->color) {
|
||||||
|
printf(" %s %s%.1f fps (paced %llu half-size copy %llu stale %llu torn %llu, sync %.2f ms decode %.2f ms)",
|
||||||
|
c->slug, c->mapped ? "" : "learning ", (double)(c->bright - c->last_bright) / opt_status,
|
||||||
|
(unsigned long long)c->paced, (unsigned long long)c->seams, (unsigned long long)c->stale,
|
||||||
|
(unsigned long long)c->torn, c->nsync ? c->sync_ns / 1e6 / c->nsync : 0.0,
|
||||||
|
c->bright ? c->copy_ns / 1e6 / c->bright : 0.0);
|
||||||
|
c->last_bright = c->bright;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
printf(" %s %s%.1f fps (dark %llu stale %llu torn %llu remapped %llu, sync %.2f ms copy %.2f ms)",
|
printf(" %s %s%.1f fps (dark %llu stale %llu torn %llu remapped %llu, sync %.2f ms copy %.2f ms)",
|
||||||
c->slug, c->mapped ? "" : "learning ", (double)(c->bright - c->last_bright) / opt_status,
|
c->slug, c->mapped ? "" : "learning ", (double)(c->bright - c->last_bright) / opt_status,
|
||||||
(unsigned long long)c->dark, (unsigned long long)c->stale, (unsigned long long)c->torn,
|
(unsigned long long)c->dark, (unsigned long long)c->stale, (unsigned long long)c->torn,
|
||||||
@@ -804,6 +927,9 @@ static void usage(const char *argv0)
|
|||||||
" --sensor S only cameras whose sensor name contains S (default: all mono cameras)\n"
|
" --sensor S only cameras whose sensor name contains S (default: all mono cameras)\n"
|
||||||
" --dark R a frame dimmer than R x the camera's recent brightest is dark (default 0.4)\n"
|
" --dark R a frame dimmer than R x the camera's recent brightest is dark (default 0.4)\n"
|
||||||
" --with-dark also publish the dark frames, as extra ring cameras flagged FH_CAM_DARK\n"
|
" --with-dark also publish the dark frames, as extra ring cameras flagged FH_CAM_DARK\n"
|
||||||
|
" --with-color also publish the Arcturus color cameras' luma, flagged FH_CAM_COLOR\n"
|
||||||
|
" --color-scale N ... at 1/N size (default 2: 986x1232)\n"
|
||||||
|
" --color-fps F ... at most F frames a second (default 30; they run at 60)\n"
|
||||||
" --status S print a status line every S seconds, 0 for never (default 10)\n"
|
" --status S print a status line every S seconds, 0 for never (default 10)\n"
|
||||||
"Frames go to " RING_FILE " (layout in fhring.h).\n", argv0);
|
"Frames go to " RING_FILE " (layout in fhring.h).\n", argv0);
|
||||||
}
|
}
|
||||||
@@ -820,6 +946,12 @@ int main(int argc, char **argv)
|
|||||||
opt_dark = atof(argv[++i]);
|
opt_dark = atof(argv[++i]);
|
||||||
else if (!strcmp(argv[i], "--with-dark"))
|
else if (!strcmp(argv[i], "--with-dark"))
|
||||||
opt_with_dark = true;
|
opt_with_dark = true;
|
||||||
|
else if (!strcmp(argv[i], "--with-color"))
|
||||||
|
opt_with_color = true;
|
||||||
|
else if (!strcmp(argv[i], "--color-scale") && i + 1 < argc && atoi(argv[i + 1]) >= 1)
|
||||||
|
opt_color_scale = (unsigned)atoi(argv[++i]);
|
||||||
|
else if (!strcmp(argv[i], "--color-fps") && i + 1 < argc)
|
||||||
|
opt_color_fps = atof(argv[++i]);
|
||||||
else if (!strcmp(argv[i], "--status") && i + 1 < argc)
|
else if (!strcmp(argv[i], "--status") && i + 1 < argc)
|
||||||
opt_status = atof(argv[++i]);
|
opt_status = atof(argv[++i]);
|
||||||
else {
|
else {
|
||||||
@@ -858,6 +990,8 @@ int main(int argc, char **argv)
|
|||||||
|
|
||||||
printf("XRService pid %d\ncameras:\n", xr.pid);
|
printf("XRService pid %d\ncameras:\n", xr.pid);
|
||||||
|
|
||||||
|
int nmono = 0;
|
||||||
|
|
||||||
for (int i = 0; i < xr.ncameras && ncams < MAX_CAMS; i++) {
|
for (int i = 0; i < xr.ncameras && ncams < MAX_CAMS; i++) {
|
||||||
|
|
||||||
xr_camera_t *cam = &xr.cameras[i];
|
xr_camera_t *cam = &xr.cameras[i];
|
||||||
@@ -866,15 +1000,21 @@ int main(int argc, char **argv)
|
|||||||
xr_camera_layout(cam, &lay);
|
xr_camera_layout(cam, &lay);
|
||||||
|
|
||||||
if (lay.fmt == XR_FMT_GREY8 && strstr(cam->sensor, opt_sensor))
|
if (lay.fmt == XR_FMT_GREY8 && strstr(cam->sensor, opt_sensor))
|
||||||
|
setup_camera(&cams[ncams++], cam, pidfd), nmono++;
|
||||||
|
else if (lay.fmt == XR_FMT_YUV420_10P && opt_with_color)
|
||||||
setup_camera(&cams[ncams++], cam, pidfd);
|
setup_camera(&cams[ncams++], cam, pidfd);
|
||||||
}
|
}
|
||||||
|
|
||||||
if (!ncams)
|
if (!nmono)
|
||||||
die("no mono camera matches '%s'", opt_sensor);
|
die("no mono camera matches '%s'", opt_sensor);
|
||||||
|
|
||||||
if (opt_with_dark && 2 * ncams > FH_RING_MAX_CAMS)
|
int cam_of[FH_RING_MAX_CAMS];
|
||||||
die("--with-dark needs a ring camera per dark stream too: pick at most %d cameras with --sensor",
|
bool dark_of[FH_RING_MAX_CAMS];
|
||||||
FH_RING_MAX_CAMS / 2);
|
int nring = ring_layout(cam_of, dark_of, FH_RING_MAX_CAMS);
|
||||||
|
|
||||||
|
if (nring > FH_RING_MAX_CAMS)
|
||||||
|
die("%d ring cameras (mono, color, dark twins) but the ring holds %d: drop --with-dark or --with-color, "
|
||||||
|
"or pick cameras with --sensor", nring, FH_RING_MAX_CAMS);
|
||||||
|
|
||||||
static tp_t tp; /* large: pending-sample pool */
|
static tp_t tp; /* large: pending-sample pool */
|
||||||
tp_event_t *evs[1] = { &ev_dqbuf };
|
tp_event_t *evs[1] = { &ev_dqbuf };
|
||||||
|
|||||||
@@ -34,6 +34,10 @@ enum {
|
|||||||
enum {
|
enum {
|
||||||
FH_CAM_DARK = 1u << 0, /* the near-black exposures between this node's */
|
FH_CAM_DARK = 1u << 0, /* the near-black exposures between this node's */
|
||||||
/* normal frames (fh-camd --with-dark) */
|
/* normal frames (fh-camd --with-dark) */
|
||||||
|
FH_CAM_COLOR = 1u << 1, /* an Arcturus color camera's luma, downscaled */
|
||||||
|
/* (fh-camd --with-color). Not synced with the */
|
||||||
|
/* mono cameras, and capture_ns is on its own */
|
||||||
|
/* clock: line it up with them by dqbuf_ns */
|
||||||
};
|
};
|
||||||
|
|
||||||
typedef struct {
|
typedef struct {
|
||||||
|
|||||||
@@ -0,0 +1,66 @@
|
|||||||
|
/*
|
||||||
|
* fh_gestures - pinch state frame-hands' tracker publishes for input: look at something
|
||||||
|
* and pinch to click it, pinch and move to drag (the Vision Pro model, with the eye
|
||||||
|
* tracker doing the looking). $XDG_RUNTIME_DIR/frame-hands/gestures, next to the hands
|
||||||
|
* file, with the same sequence lock (read seq, copy, read seq again; use the copy only if
|
||||||
|
* both reads are the same even number) and the same frame: metres in the head frame at
|
||||||
|
* capture time, OpenVR's HMD frame (+x right, +y up, -z forward).
|
||||||
|
*
|
||||||
|
* One slot per side: pinch[0] is the left hand, pinch[1] the right. A pinch follows the
|
||||||
|
* hand it began on until it ends. It begins when the thumb and index tips close within
|
||||||
|
* begin_m and ends when they open past end_m (the gap between keeps it from flickering),
|
||||||
|
* or when the hand stays lost too long (FH_PINCH_LOST).
|
||||||
|
*
|
||||||
|
* Don't miss short pinches: a reader that polls slower than a quick tap still sees it,
|
||||||
|
* because begins and ends count every pinch. When begins changed, a pinch began at
|
||||||
|
* begin_ns; when ends changed, one ended at end_ns. begins - ends is 1 while pinching.
|
||||||
|
*
|
||||||
|
* Drags: point is where the pinch is now, begin_point where it began. Turn each into the
|
||||||
|
* room with the HMD pose at its capture time (capture_ns, begin_ns) before subtracting,
|
||||||
|
* so turning your head doesn't drag.
|
||||||
|
*/
|
||||||
|
|
||||||
|
#pragma once
|
||||||
|
|
||||||
|
#include <assert.h>
|
||||||
|
#include <stdint.h>
|
||||||
|
|
||||||
|
#define FH_GESTURES_MAGIC "FHGEST01"
|
||||||
|
#define FH_GESTURES_VERSION 1
|
||||||
|
|
||||||
|
enum {
|
||||||
|
FH_PINCH_TRACKED = 1u << 0, /* the hand was tracked in this frame */
|
||||||
|
FH_PINCH_DOWN = 1u << 1, /* pinching now */
|
||||||
|
FH_PINCH_LOST = 1u << 2, /* the last pinch ended because the hand was lost */
|
||||||
|
};
|
||||||
|
|
||||||
|
typedef struct {
|
||||||
|
uint32_t flags; /* FH_PINCH_* */
|
||||||
|
uint32_t hand_id; /* fh_hand_t.id of the hand, 0 if none */
|
||||||
|
uint32_t begins; /* pinches begun so far */
|
||||||
|
uint32_t ends; /* pinches ended so far */
|
||||||
|
uint64_t begin_ns; /* capture time (CLOCK_MONOTONIC) the current or */
|
||||||
|
/* last pinch began */
|
||||||
|
uint64_t end_ns; /* ... the last pinch ended */
|
||||||
|
float distance; /* thumb tip to index tip, m, at this user's hand */
|
||||||
|
/* size */
|
||||||
|
float strength; /* 0 open (end_m or more) .. 1 closed (begin_m) */
|
||||||
|
float point[3]; /* between the thumb and index tips */
|
||||||
|
float begin_point[3]; /* point when the current or last pinch began */
|
||||||
|
} fh_pinch_t; /* 64 bytes */
|
||||||
|
|
||||||
|
typedef struct {
|
||||||
|
char magic[8];
|
||||||
|
uint32_t version;
|
||||||
|
uint32_t size;
|
||||||
|
volatile uint64_t seq;
|
||||||
|
uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */
|
||||||
|
uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */
|
||||||
|
float begin_m; /* the thresholds in use */
|
||||||
|
float end_m;
|
||||||
|
uint8_t reserved[16];
|
||||||
|
fh_pinch_t pinch[2]; /* [0] left hand, [1] right hand */
|
||||||
|
} fh_gestures_t;
|
||||||
|
|
||||||
|
static_assert(sizeof(fh_pinch_t) == 64, "fh_pinch_t layout");
|
||||||
|
static_assert(sizeof(fh_gestures_t) == 64 + 2 * 64, "fh_gestures_t layout");
|
||||||
@@ -0,0 +1,65 @@
|
|||||||
|
"""Which color camera is which, and how their calibration maps onto fh-camd's images.
|
||||||
|
|
||||||
|
usage: python tools/check_color.py REC_DIR [--sets N]
|
||||||
|
|
||||||
|
A recording made with fh-camd --with-color holds color_video<N> frames with each set.
|
||||||
|
This matches features between the two color images and scores every reading of the
|
||||||
|
calibration: which video node is passthrough_left, and whether the calibration's
|
||||||
|
cropRegion is subtracted from x ('subtract') or not ('none'). Only the right reading
|
||||||
|
makes true matches' rays meet in front of both cameras. Then it checks the winner against
|
||||||
|
the side tracking cameras, which tests the CAD-to-head chain shared with them.
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import itertools
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
|
||||||
|
from tools.check_sides import load_cams, matches, score # noqa: E402
|
||||||
|
from tools.show_set import index, read_set # noqa: E402
|
||||||
|
from tracker import calib # noqa: E402
|
||||||
|
|
||||||
|
|
||||||
|
def load_color(crop):
|
||||||
|
root = os.environ.get('FRAME_JOB_DEVICE_ROOT', '') # frame-job's copy of the device files
|
||||||
|
return calib.load_color(root + calib.ARCTURUS_EEPROM, root + calib.DEVICE_JSON, crop=crop)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument('rec')
|
||||||
|
ap.add_argument('--sets', type=int, default=8)
|
||||||
|
a = ap.parse_args()
|
||||||
|
path = os.path.join(a.rec, 'sets.bin')
|
||||||
|
offs = index(path)
|
||||||
|
sets = [read_set(path, offs[n]) for n in np.linspace(0, len(offs) - 1, a.sets).astype(int)]
|
||||||
|
nodes = sorted(k for k in sets[0] if k.startswith('color_video'))
|
||||||
|
if len(nodes) != 2:
|
||||||
|
sys.exit('need two color_video<N> cameras in the recording (fh-camd --with-color); found %s' % nodes)
|
||||||
|
pairs = [matches(s[nodes[0]][0], s[nodes[1]][0]) for s in sets]
|
||||||
|
print('%d sets, %d matches between %s and %s' % (len(sets), sum(len(p[0]) for p in pairs), *nodes))
|
||||||
|
|
||||||
|
best = None
|
||||||
|
for crop, left in itertools.product(['subtract', 'none'], nodes):
|
||||||
|
cams = load_color(crop)
|
||||||
|
right = nodes[1] if left == nodes[0] else nodes[0]
|
||||||
|
cam = {left: cams['passthrough_left'], right: cams['passthrough_right']}
|
||||||
|
s = np.mean([score(cam[nodes[0]], cam[nodes[1]], ua, ub) for ua, ub in pairs])
|
||||||
|
print(' %s = passthrough_left, crop %-8s: %3.0f%% of matches meet' % (left, crop, 100 * s))
|
||||||
|
if best is None or s > best[0]:
|
||||||
|
best = (s, crop, left, cam)
|
||||||
|
s, crop, left, cam = best
|
||||||
|
print('best: %s = passthrough_left, crop %s (%.0f%%)' % (left, crop, 100 * s))
|
||||||
|
|
||||||
|
mono = load_cams()
|
||||||
|
for node in nodes:
|
||||||
|
for side in ['slam_left', 'slam_right']:
|
||||||
|
ms = [matches(st[node][0], st[side][0]) for st in sets if side in st]
|
||||||
|
sc = np.mean([score(cam[node], mono[side], ua, ub) for ua, ub in ms]) if ms else 0
|
||||||
|
print(' %s vs %-10s: %4d matches, %3.0f%% meet' % (node, side, sum(len(m[0]) for m in ms), 100 * sc))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
@@ -0,0 +1,256 @@
|
|||||||
|
"""How good the tracker's depth is, from a replay's depth dump, without ground truth.
|
||||||
|
|
||||||
|
usage: python3 tools/depth_report.py DEPTH [DEPTH...] [--still M/S]
|
||||||
|
|
||||||
|
DEPTH comes from `trackd/fh-replay DIR --depth DEPTH`. Every measure is split by how the
|
||||||
|
hand was seen: by the two lower cameras ("lower pair"), by a lower and an upper camera on
|
||||||
|
one side ("lower+upper"), or by one camera. Distances are from the head (between the eyes).
|
||||||
|
|
||||||
|
1. How the hands were seen: the share of hand updates in each way, by distance.
|
||||||
|
2. Noise along the line of sight against across it. Each update's palm is compared with a
|
||||||
|
straight line through the two updates before it (fh-replay's jitter measure), and the
|
||||||
|
miss is split along the line from the hand's cameras to the palm and across it. Given
|
||||||
|
as a robust sigma per axis, measured (as triangulated) and published (after the One Euro
|
||||||
|
filter), on updates where the published palm moved slower than --still (default 0.15
|
||||||
|
m/s), so the miss is mostly noise and not the hand speeding up. For two cameras,
|
||||||
|
geometry predicts along/across = 2 Z / B: Z the distance, B the cameras' baseline
|
||||||
|
across the line of sight.
|
||||||
|
3. One-camera distance: on two-camera updates, each camera's one-view guess (distance from
|
||||||
|
how big the palm looks, at the user's learned hand size) against the triangulated
|
||||||
|
distance from that camera.
|
||||||
|
4. A camera lost: from two-camera updates, what the tracker would have had if one of the
|
||||||
|
two cameras dropped out there. It keeps the last distance and moves a share of the way
|
||||||
|
to the one-view guess each update (0.1 now, kMonoDepthGain in trackd/tracker.cpp);
|
||||||
|
also shown with other shares, 0 (keep the distance) and 1 (take each guess), and with
|
||||||
|
the guess first scaled by how far off it was while both cameras saw the hand. Compared with the
|
||||||
|
triangulated distance from that camera, 0.1-2 s after the loss.
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import collections
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
BINS = [0.0, 0.35, 0.50, 0.65, 9.0]
|
||||||
|
BIN_NAMES = ['<35 cm', '35-50', '50-65', '65+ cm']
|
||||||
|
MODES = ['lower pair', 'lower+upper', 'one camera']
|
||||||
|
HORIZONS = [0.1, 0.25, 0.5, 1.0, 2.0]
|
||||||
|
# (share of the way toward the one-view guess per update, whether the guess is first scaled by
|
||||||
|
# how far off it was while both cameras saw the hand)
|
||||||
|
GAINS = [(0.0, False), (0.02, False), (0.05, False), (0.1, False), (1.0, False), (0.1, True), (1.0, True)]
|
||||||
|
GAP = 0.1 # s: a longer gap between a hand's updates breaks its run
|
||||||
|
|
||||||
|
|
||||||
|
def load(path):
|
||||||
|
cams, rows = {}, []
|
||||||
|
with open(path) as f:
|
||||||
|
for line in f:
|
||||||
|
w = line.split()
|
||||||
|
if not w:
|
||||||
|
continue
|
||||||
|
if w[0] == '#':
|
||||||
|
if w[1] == 'cam':
|
||||||
|
cams[w[2]] = (np.array([float(x) for x in w[3:6]]), float(w[6]))
|
||||||
|
continue
|
||||||
|
r = {'t': float(w[0]), 'id': int(w[1]), 'side': w[2], 'n': int(w[3]),
|
||||||
|
'cams': w[4], 'res': float(w[5]), 'scale': float(w[6]),
|
||||||
|
'raw': np.array([float(x) for x in w[7:10]]), 'sm': np.array([float(x) for x in w[10:13]]),
|
||||||
|
'views': {}}
|
||||||
|
for k in range(13, len(w), 5):
|
||||||
|
r['views'][w[k]] = np.array([float(x) for x in w[k + 2:k + 5]])
|
||||||
|
rows.append(r)
|
||||||
|
return cams, rows
|
||||||
|
|
||||||
|
|
||||||
|
def mode(r):
|
||||||
|
names = r['cams'].split('+')
|
||||||
|
if r['n'] == 1:
|
||||||
|
return 'one camera'
|
||||||
|
if sorted(names) == ['slam_left', 'slam_right']:
|
||||||
|
return 'lower pair'
|
||||||
|
if len(names) == 2 and all(n.startswith(('slam_', 'upper_')) for n in names) and \
|
||||||
|
names[0].split('_')[1] == names[1].split('_')[1]:
|
||||||
|
return 'lower+upper'
|
||||||
|
return 'other'
|
||||||
|
|
||||||
|
|
||||||
|
def dist_bin(p):
|
||||||
|
return min(np.searchsorted(BINS, np.linalg.norm(p), side='right') - 1, len(BIN_NAMES) - 1)
|
||||||
|
|
||||||
|
|
||||||
|
def tracks(rows):
|
||||||
|
"""A hand's updates, in order, split where they're more than GAP apart."""
|
||||||
|
by_id = collections.defaultdict(list)
|
||||||
|
for r in rows:
|
||||||
|
by_id[r['id']].append(r)
|
||||||
|
for rs in by_id.values():
|
||||||
|
run = [rs[0]]
|
||||||
|
for r in rs[1:]:
|
||||||
|
if r['t'] - run[-1]['t'] >= GAP:
|
||||||
|
yield run
|
||||||
|
run = []
|
||||||
|
run.append(r)
|
||||||
|
yield run
|
||||||
|
|
||||||
|
|
||||||
|
def two_camera_runs(rows):
|
||||||
|
"""Stretches of a hand's updates all seen by the same two cameras."""
|
||||||
|
for run in tracks(rows):
|
||||||
|
seg = []
|
||||||
|
for r in run:
|
||||||
|
ok = r['n'] == 2 and mode(r) != 'other'
|
||||||
|
if ok and seg and r['cams'] == seg[-1]['cams']:
|
||||||
|
seg.append(r)
|
||||||
|
continue
|
||||||
|
if len(seg) > 2:
|
||||||
|
yield seg
|
||||||
|
seg = [r] if ok else []
|
||||||
|
if len(seg) > 2:
|
||||||
|
yield seg
|
||||||
|
|
||||||
|
|
||||||
|
def pct(v, q):
|
||||||
|
return np.percentile(v, q) if len(v) else float('nan')
|
||||||
|
|
||||||
|
|
||||||
|
def seen_share(rows):
|
||||||
|
print('\n1. How the hands were seen (share of hand updates)')
|
||||||
|
count = collections.Counter((mode(r), dist_bin(r['raw'])) for r in rows)
|
||||||
|
total = collections.Counter(dist_bin(r['raw']) for r in rows)
|
||||||
|
print('%-12s' % '' + ''.join('%10s' % b for b in BIN_NAMES) + '%10s' % 'all')
|
||||||
|
for m in MODES + ['other']:
|
||||||
|
cells = [100 * count[m, b] / max(total[b], 1) for b in range(len(BIN_NAMES))]
|
||||||
|
allp = 100 * sum(count[m, b] for b in range(len(BIN_NAMES))) / max(len(rows), 1)
|
||||||
|
print('%-12s' % m + ''.join('%9.0f%%' % c for c in cells) + '%9.0f%%' % allp)
|
||||||
|
print('%-12s' % 'updates' + ''.join('%10d' % total[b] for b in range(len(BIN_NAMES))) + '%10d' % len(rows))
|
||||||
|
res = collections.defaultdict(list)
|
||||||
|
for r in rows:
|
||||||
|
if r['res'] >= 0:
|
||||||
|
res[mode(r)].append(r['res'] * 1000)
|
||||||
|
print('triangulation residual (rms ray miss, median): ' +
|
||||||
|
', '.join('%s %.1f mm' % (m, np.median(v)) for m, v in res.items()))
|
||||||
|
|
||||||
|
|
||||||
|
def noise(rows, cams, still):
|
||||||
|
print('\n2. Noise along the line of sight vs across it (sigma per axis, mm; palm slower than %.2f m/s)' % still)
|
||||||
|
acc = collections.defaultdict(lambda: collections.defaultdict(list))
|
||||||
|
for run in tracks(rows):
|
||||||
|
for a, b, c in zip(run, run[1:], run[2:]):
|
||||||
|
if not (mode(a) == mode(b) == mode(c)) or a['cams'] != b['cams'] or b['cams'] != c['cams']:
|
||||||
|
continue
|
||||||
|
dt0, dt1 = b['t'] - a['t'], c['t'] - b['t']
|
||||||
|
if dt0 < 1e-3 or np.linalg.norm(b['sm'] - a['sm']) / dt0 > still:
|
||||||
|
continue
|
||||||
|
names = b['cams'].split('+')
|
||||||
|
origins = [cams[n][0] for n in names]
|
||||||
|
o = np.mean(origins, axis=0)
|
||||||
|
u = b['raw'] - o
|
||||||
|
z = np.linalg.norm(u)
|
||||||
|
u /= z
|
||||||
|
key = (mode(b), dist_bin(b['raw']))
|
||||||
|
for kind in ('raw', 'sm'):
|
||||||
|
miss = c[kind] - b[kind] - (b[kind] - a[kind]) * (dt1 / dt0)
|
||||||
|
along = miss @ u
|
||||||
|
acc[key][kind + '_along'].append(abs(along))
|
||||||
|
acc[key][kind + '_across'].append(np.linalg.norm(miss - along * u))
|
||||||
|
if len(origins) == 2:
|
||||||
|
base = origins[0] - origins[1]
|
||||||
|
acc[key]['pred'].append(2 * z / np.linalg.norm(base - (base @ u) * u))
|
||||||
|
# |along| is half-normal: sigma = median / 0.674; |across| is Rayleigh (2 axes): sigma = median / 1.177
|
||||||
|
print('%-12s %-7s %6s | %-24s | %-24s | %s' % ('', '', 'n', 'measured along/across', 'published along/across',
|
||||||
|
'ratio measured (geometry)'))
|
||||||
|
for m in MODES:
|
||||||
|
for bi, bn in enumerate(BIN_NAMES):
|
||||||
|
d = acc.get((m, bi))
|
||||||
|
if not d or len(d['raw_along']) < 20:
|
||||||
|
continue
|
||||||
|
s = {k: np.median(v) / (0.674 if k.endswith('along') else 1.177) * 1000
|
||||||
|
for k, v in d.items() if k != 'pred'}
|
||||||
|
pred = '(%.1f)' % np.median(d['pred']) if d['pred'] else ''
|
||||||
|
print('%-12s %-7s %6d | %7.1f / %-5.1f x%-6.1f | %7.1f / %-5.1f x%-6.1f | x%.1f %s' % (
|
||||||
|
m, bn, len(d['raw_along']), s['raw_along'], s['raw_across'], s['raw_along'] / s['raw_across'],
|
||||||
|
s['sm_along'], s['sm_across'], s['sm_along'] / s['sm_across'],
|
||||||
|
s['raw_along'] / s['raw_across'], pred))
|
||||||
|
|
||||||
|
|
||||||
|
def one_camera(rows, cams):
|
||||||
|
print('\n3. One-camera distance vs triangulated, on two-camera updates (error of the one-view guess)')
|
||||||
|
acc = collections.defaultdict(list)
|
||||||
|
for r in rows:
|
||||||
|
if r['n'] != 2 or mode(r) == 'other':
|
||||||
|
continue
|
||||||
|
for name, p in r['views'].items():
|
||||||
|
if np.isnan(p).any():
|
||||||
|
continue
|
||||||
|
o = cams[name][0]
|
||||||
|
truth = np.linalg.norm(r['raw'] - o)
|
||||||
|
acc[name.split('_')[0], dist_bin(r['raw'])].append((np.linalg.norm(p - o) - truth, truth))
|
||||||
|
print('%-8s %-7s %6s %12s %12s %14s %12s' % ('camera', '', 'n', 'median |err|', '90% |err|', 'median |err| %',
|
||||||
|
'bias'))
|
||||||
|
for cam in ('slam', 'upper'):
|
||||||
|
for bi, bn in enumerate(BIN_NAMES):
|
||||||
|
v = acc.get((cam, bi))
|
||||||
|
if not v or len(v) < 20:
|
||||||
|
continue
|
||||||
|
e = np.array([x[0] for x in v])
|
||||||
|
rel = e / np.array([x[1] for x in v])
|
||||||
|
print('%-8s %-7s %6d %9.0f mm %9.0f mm %13.0f%% %+11.0f%%' % (
|
||||||
|
'lower' if cam == 'slam' else 'upper', bn, len(v), 1000 * np.median(abs(e)), 1000 * pct(abs(e), 90),
|
||||||
|
100 * np.median(abs(rel)), 100 * np.median(rel)))
|
||||||
|
|
||||||
|
|
||||||
|
def lost_camera(rows, cams):
|
||||||
|
print('\n4. A camera lost: distance error after the loss (median |err| mm / 90% mm), by how the tracker '
|
||||||
|
'moves toward the one-view guess each update ("scaled": the guess times how far off it was, '
|
||||||
|
'triangulated / guess, median over the last 30 two-camera updates)')
|
||||||
|
errs = collections.defaultdict(list)
|
||||||
|
for run in two_camera_runs(rows):
|
||||||
|
for s in range(1, len(run) - 1, 3):
|
||||||
|
for name in run[s]['views']:
|
||||||
|
o = cams[name][0]
|
||||||
|
guess = lambda r: np.linalg.norm(r['views'][name] - o)
|
||||||
|
ratios = [np.linalg.norm(r['raw'] - o) / guess(r) for r in run[max(0, s - 30):s]
|
||||||
|
if not np.isnan(r['views'][name]).any()]
|
||||||
|
ratio = np.median(ratios) if ratios else 1.0
|
||||||
|
for g, scaled in GAINS:
|
||||||
|
d = np.linalg.norm(run[s - 1]['raw'] - o)
|
||||||
|
h = 0
|
||||||
|
for r in run[s:]:
|
||||||
|
if np.isnan(r['views'][name]).any():
|
||||||
|
break
|
||||||
|
d += g * (guess(r) * (ratio if scaled else 1.0) - d)
|
||||||
|
elapsed = r['t'] - run[s - 1]['t']
|
||||||
|
while h < len(HORIZONS) and elapsed >= HORIZONS[h]:
|
||||||
|
errs[g, scaled, HORIZONS[h], name.split('_')[0]].append(abs(d - np.linalg.norm(r['raw'] - o)))
|
||||||
|
h += 1
|
||||||
|
print('%-8s %-18s' % ('camera', 'toward guess') + ''.join('%14s' % ('%.2g s' % t) for t in HORIZONS))
|
||||||
|
for cam in ('slam', 'upper'):
|
||||||
|
for g, scaled in GAINS:
|
||||||
|
label = {0.0: '0 (keep)', 0.1: '0.1 (now)', 1.0: '1 (guess)'}.get(g, '%g' % g)
|
||||||
|
if scaled:
|
||||||
|
label = '%g scaled' % g
|
||||||
|
cells = []
|
||||||
|
for t in HORIZONS:
|
||||||
|
v = errs.get((g, scaled, t, cam), [])
|
||||||
|
cells.append('%5.0f / %-4.0f' % (1000 * np.median(v), 1000 * pct(v, 90)) if len(v) >= 20 else '%14s' % '-')
|
||||||
|
print('%-8s %-18s' % ('lower' if cam == 'slam' else 'upper', label) + ''.join('%14s' % c for c in cells))
|
||||||
|
n = sum(len(errs.get((0.1, False, HORIZONS[0], c), [])) for c in ('slam', 'upper'))
|
||||||
|
print('(%d simulated losses)' % n)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument('depth', nargs='+')
|
||||||
|
ap.add_argument('--still', type=float, default=0.15, help='m/s: palm speed limit for the noise measure')
|
||||||
|
a = ap.parse_args()
|
||||||
|
for path in a.depth:
|
||||||
|
cams, rows = load(path)
|
||||||
|
print('== %s: %d hand updates, %d hands' % (path, len(rows), len({r['id'] for r in rows})))
|
||||||
|
seen_share(rows)
|
||||||
|
noise(rows, cams, a.still)
|
||||||
|
one_camera(rows, cams)
|
||||||
|
lost_camera(rows, cams)
|
||||||
|
print()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
+28
-19
@@ -6,7 +6,8 @@ SET is a set index (fh-replay's timeline gives them). With --timeline (fh-replay
|
|||||||
--timeline), each camera shows the tracker's views at that set: the crop for the next
|
--timeline), each camera shows the tracker's views at that set: the crop for the next
|
||||||
frame, labelled with the hand and presence. Recordings made with fh-camd --with-dark get
|
frame, labelled with the hand and presence. Recordings made with fh-camd --with-dark get
|
||||||
a second row: each camera's latest dark frame (<name>_dk), stretched to be visible and
|
a second row: each camera's latest dark frame (<name>_dk), stretched to be visible and
|
||||||
labelled with its mean brightness. Writes OUT/set_<n>.jpg (default /tmp).
|
labelled with its mean brightness. Recordings made with fh-camd --with-color get a row of
|
||||||
|
the color cameras (color_video<N>). Writes OUT/set_<n>.jpg (default /tmp).
|
||||||
"""
|
"""
|
||||||
import argparse
|
import argparse
|
||||||
import os
|
import os
|
||||||
@@ -72,31 +73,39 @@ def dark_tile(frame, shape, name):
|
|||||||
return img
|
return img
|
||||||
|
|
||||||
|
|
||||||
|
def view_tile(px, name, views):
|
||||||
|
"""A frame, CLAHE'd, with the tracker's views on it, 512 px high."""
|
||||||
|
img = cv2.cvtColor(cv2.createCLAHE(2.0, (8, 8)).apply(px), cv2.COLOR_GRAY2BGR)
|
||||||
|
for v in views:
|
||||||
|
if v['cam'] != name:
|
||||||
|
continue
|
||||||
|
c, s, r = v['c'], v['size'], v['rot']
|
||||||
|
box = cv2.boxPoints(((c[0], c[1]), (s, s), np.degrees(r)))
|
||||||
|
col = (0, 255, 0) if v['presence'] >= 0.5 else (0, 0, 255)
|
||||||
|
cv2.polylines(img, [box.astype(np.int32)], True, col, 2)
|
||||||
|
cv2.putText(img, 'h%d %.2f' % (v['hand'], v['presence']), (int(c[0] - s / 2), int(c[1] - s / 2) - 6),
|
||||||
|
cv2.FONT_HERSHEY_SIMPLEX, 0.8, col, 2)
|
||||||
|
cv2.putText(img, name, (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0, (255, 255, 0), 2)
|
||||||
|
scale = 512 / img.shape[0]
|
||||||
|
return cv2.resize(img, (int(img.shape[1] * scale), 512))
|
||||||
|
|
||||||
|
|
||||||
def draw(images, views):
|
def draw(images, views):
|
||||||
|
"""Rows: the mono cameras; their dark frames, if recorded; the color cameras, if recorded."""
|
||||||
tiles, dark = [], []
|
tiles, dark = [], []
|
||||||
for name in ORDER:
|
for name in ORDER:
|
||||||
if name not in images:
|
if name not in images:
|
||||||
continue
|
continue
|
||||||
px = images[name][0]
|
tiles.append(view_tile(images[name][0], name, views))
|
||||||
clahe = cv2.createCLAHE(2.0, (8, 8)).apply(px)
|
|
||||||
img = cv2.cvtColor(clahe, cv2.COLOR_GRAY2BGR)
|
|
||||||
for v in views:
|
|
||||||
if v['cam'] != name:
|
|
||||||
continue
|
|
||||||
c, s, r = v['c'], v['size'], v['rot']
|
|
||||||
box = cv2.boxPoints(((c[0], c[1]), (s, s), np.degrees(r)))
|
|
||||||
col = (0, 255, 0) if v['presence'] >= 0.5 else (0, 0, 255)
|
|
||||||
cv2.polylines(img, [box.astype(np.int32)], True, col, 2)
|
|
||||||
cv2.putText(img, 'h%d %.2f' % (v['hand'], v['presence']), (int(c[0] - s / 2), int(c[1] - s / 2) - 6),
|
|
||||||
cv2.FONT_HERSHEY_SIMPLEX, 0.8, col, 2)
|
|
||||||
cv2.putText(img, name, (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1.0, (255, 255, 0), 2)
|
|
||||||
scale = 512 / img.shape[0]
|
|
||||||
tiles.append(cv2.resize(img, (int(img.shape[1] * scale), 512)))
|
|
||||||
dark.append(dark_tile(images.get(name + '_dk'), tiles[-1].shape[:2], name))
|
dark.append(dark_tile(images.get(name + '_dk'), tiles[-1].shape[:2], name))
|
||||||
out = np.hstack(tiles)
|
rows = [np.hstack(tiles)]
|
||||||
if any(k.endswith('_dk') for k in images):
|
if any(k.endswith('_dk') for k in images):
|
||||||
out = np.vstack([out, np.hstack(dark)])
|
rows.append(np.hstack(dark))
|
||||||
return out
|
color = sorted(k for k in images if k.startswith('color_'))
|
||||||
|
if color:
|
||||||
|
rows.append(np.hstack([view_tile(images[k][0], k, views) for k in color]))
|
||||||
|
width = max(r.shape[1] for r in rows)
|
||||||
|
return np.vstack([np.pad(r, ((0, 0), (0, width - r.shape[1]), (0, 0))) for r in rows])
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
|
|||||||
@@ -0,0 +1,85 @@
|
|||||||
|
"""Watch the pinch gestures fh-tracker publishes, live: begins, ends, and drags.
|
||||||
|
|
||||||
|
usage: python3 tools/watch_gestures.py [--every S] [--distance]
|
||||||
|
|
||||||
|
Prints a line when a pinch begins or ends on either hand. It goes by the counters, so a
|
||||||
|
quick tap between two reads still shows. While a pinch is held, every --every seconds
|
||||||
|
(default 0.1) it prints how far the pinch point has moved since it began, in the head
|
||||||
|
frame (turning your head moves it too; a real consumer turns both points into the room
|
||||||
|
first, see include/fh_gestures.h). --distance also prints each hand's thumb-to-index
|
||||||
|
distance, to see how close a pinch comes to the thresholds.
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import mmap
|
||||||
|
import os
|
||||||
|
import struct
|
||||||
|
import time
|
||||||
|
|
||||||
|
HDR = struct.Struct('<8sIIQQQff16x') # 64 bytes
|
||||||
|
PINCH = struct.Struct('<IIIIQQff3f3f') # 64 bytes
|
||||||
|
TRACKED, DOWN, LOST = 1, 2, 4
|
||||||
|
SIDES = ('left ', 'right')
|
||||||
|
|
||||||
|
|
||||||
|
def path():
|
||||||
|
run = os.environ.get('XDG_RUNTIME_DIR', '/run/user/%d' % os.getuid())
|
||||||
|
return os.path.join(run, 'frame-hands', 'gestures')
|
||||||
|
|
||||||
|
|
||||||
|
def read(m):
|
||||||
|
"""(header, [left, right]) under the sequence lock, or None if it's being written."""
|
||||||
|
for _ in range(10):
|
||||||
|
s1 = struct.unpack_from('<Q', m, 16)[0]
|
||||||
|
if s1 % 2 == 0:
|
||||||
|
h = HDR.unpack_from(m, 0)
|
||||||
|
p = [PINCH.unpack_from(m, HDR.size + k * PINCH.size) for k in range(2)]
|
||||||
|
if struct.unpack_from('<Q', m, 16)[0] == s1:
|
||||||
|
return h, p
|
||||||
|
time.sleep(0.0005)
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument('--every', type=float, default=0.1, help='seconds between drag lines while pinching')
|
||||||
|
ap.add_argument('--distance', action='store_true', help="print each hand's thumb-to-index distance")
|
||||||
|
a = ap.parse_args()
|
||||||
|
with open(path(), 'rb') as f:
|
||||||
|
m = mmap.mmap(f.fileno(), 0, prot=mmap.PROT_READ)
|
||||||
|
first = read(m)
|
||||||
|
if first is None or first[0][0] != b'FHGEST01':
|
||||||
|
raise SystemExit('%s is not an fh-tracker gestures file' % path())
|
||||||
|
h, p = first
|
||||||
|
print('thresholds: pinch begins under %.3f m, ends over %.3f m' % (h[6], h[7]))
|
||||||
|
seen = [(q[2], q[3]) for q in p] # begins, ends
|
||||||
|
last_drag = last_dist = 0.0
|
||||||
|
while True:
|
||||||
|
got = read(m)
|
||||||
|
if got:
|
||||||
|
h, p = got
|
||||||
|
now = time.monotonic()
|
||||||
|
for k, q in enumerate(p):
|
||||||
|
flags, hand, begins, ends, begin_ns, end_ns, dist, strength = q[:8]
|
||||||
|
point, begin_point = q[8:11], q[11:14]
|
||||||
|
if begins != seen[k][0]:
|
||||||
|
print('%s pinch BEGIN (#%d, hand %d) at %+.3f %+.3f %+.3f d %.3f' %
|
||||||
|
(SIDES[k], begins, hand, *begin_point, dist), flush=True)
|
||||||
|
if ends != seen[k][1]:
|
||||||
|
held = (end_ns - begin_ns) / 1e9 if end_ns >= begin_ns else 0
|
||||||
|
print('%s pinch %s after %.2f s' % (SIDES[k], 'LOST' if flags & LOST else 'END', held), flush=True)
|
||||||
|
seen[k] = (begins, ends)
|
||||||
|
if flags & DOWN and now - last_drag >= a.every:
|
||||||
|
d = [point[i] - begin_point[i] for i in range(3)]
|
||||||
|
print('%s drag %+6.1f %+6.1f %+6.1f mm (%.0f mm)' %
|
||||||
|
(SIDES[k], *(1000 * x for x in d), 1000 * sum(x * x for x in d) ** 0.5), flush=True)
|
||||||
|
if any(q[0] & DOWN for q in p) and now - last_drag >= a.every:
|
||||||
|
last_drag = now
|
||||||
|
if a.distance and now - last_dist >= 0.2:
|
||||||
|
last_dist = now
|
||||||
|
print(' ' + ' '.join('%s %s' % (SIDES[k].strip(), 'd %.3f s %.2f' % (q[6], q[7]) if q[0] & TRACKED
|
||||||
|
else '-') for k, q in enumerate(p)), flush=True)
|
||||||
|
time.sleep(0.005)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
+4
-4
@@ -5,16 +5,16 @@ CXXFLAGS ?= -O2 -g -Wall -Wextra -Wno-unused-parameter
|
|||||||
CXXFLAGS += -std=c++17 -fopenmp -I$(NCNN)/include/ncnn
|
CXXFLAGS += -std=c++17 -fopenmp -I$(NCNN)/include/ncnn
|
||||||
LDLIBS = $(NCNN)/lib/libncnn.a -ljsoncpp -fopenmp -lpthread
|
LDLIBS = $(NCNN)/lib/libncnn.a -ljsoncpp -fopenmp -lpthread
|
||||||
|
|
||||||
SRC = main.cpp calib.cpp nets.cpp tracker.cpp io.cpp record.cpp
|
SRC = main.cpp calib.cpp nets.cpp tracker.cpp io.cpp record.cpp pinch.cpp
|
||||||
HDR = calib.h geom.h nets.h tracker.h io.h record.h ../camd/fhring.h ../include/fh_hands.h
|
HDR = calib.h geom.h nets.h tracker.h io.h record.h pinch.h ../camd/fhring.h ../include/fh_hands.h ../include/fh_gestures.h
|
||||||
|
|
||||||
all: fh-tracker fh-replay fh-ringplay
|
all: fh-tracker fh-replay fh-ringplay
|
||||||
|
|
||||||
fh-tracker: $(SRC) $(HDR)
|
fh-tracker: $(SRC) $(HDR)
|
||||||
$(CXX) $(CXXFLAGS) -o $@ $(SRC) $(LDLIBS)
|
$(CXX) $(CXXFLAGS) -o $@ $(SRC) $(LDLIBS)
|
||||||
|
|
||||||
fh-replay: replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp $(HDR)
|
fh-replay: replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp pinch.cpp io.cpp $(HDR)
|
||||||
$(CXX) $(CXXFLAGS) -o $@ replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp $(LDLIBS)
|
$(CXX) $(CXXFLAGS) -o $@ replay.cpp calib.cpp nets.cpp tracker.cpp record.cpp pinch.cpp io.cpp $(LDLIBS)
|
||||||
|
|
||||||
fh-ringplay: ringplay.cpp record.h ../camd/fhring.h
|
fh-ringplay: ringplay.cpp record.h ../camd/fhring.h
|
||||||
$(CXX) $(CXXFLAGS) -o $@ ringplay.cpp
|
$(CXX) $(CXXFLAGS) -o $@ ringplay.cpp
|
||||||
|
|||||||
@@ -23,6 +23,7 @@ Options:
|
|||||||
- `--record-only`: record without tracking or publishing, so it can run beside the live tracker. Give it `--record DIR`; SIGUSR1 would reach both trackers. With `fh-camd --with-dark`, recordings also hold each camera's newest dark frame as `<name>_dk`, which doubles the rate.
|
- `--record-only`: record without tracking or publishing, so it can run beside the live tracker. Give it `--record DIR`; SIGUSR1 would reach both trackers. With `fh-camd --with-dark`, recordings also hold each camera's newest dark frame as `<name>_dk`, which doubles the rate.
|
||||||
- `--keep-presence P`: the landmark presence a tracked view needs to stay tracked. New views always need 0.5. Default 0.5. Lowering it to 0.2 barely helped in the bright recording, because lost hands drop to near-zero presence.
|
- `--keep-presence P`: the landmark presence a tracked view needs to stay tracked. New views always need 0.5. Default 0.5. Lowering it to 0.2 barely helped in the bright recording, because lost hands drop to near-zero presence.
|
||||||
- `--ring PATH`: read frames from another ring, such as `fh-ringplay`'s.
|
- `--ring PATH`: read frames from another ring, such as `fh-ringplay`'s.
|
||||||
|
- Recording from a ring with color cameras (`fh-camd --with-color`) also saves each color camera's newest frame with every set, as `color_video<N>`. That adds about 70 MB/s. Run the recorder at normal I/O priority (not under `frame-job`, whose `ionice -c 3` stalled a 165 MB/s recording).
|
||||||
|
|
||||||
The status line also says how often a hand was on each side (by where the wrist is), and why views and hands came and went: views lost (the landmark model stopped seeing the hand), handoff misses (a crop projected from the hand's 3D position found nothing), duplicates, splits (two views disagreed in 3D), and hands created, merged and forgotten.
|
The status line also says how often a hand was on each side (by where the wrist is), and why views and hands came and went: views lost (the landmark model stopped seeing the hand), handoff misses (a crop projected from the hand's 3D position found nothing), duplicates, splits (two views disagreed in 3D), and hands created, merged and forgotten.
|
||||||
|
|
||||||
@@ -35,6 +36,17 @@ fh-tracker started as a port of `tracker/hands.py`. Replaying recordings (below)
|
|||||||
- **Smoothing.** The published landmarks go through a One Euro filter: it smooths hard while the hand is still (tracking noise is several mm per frame) and hardly at all while it moves fast. The palm speed that sets the update rate (15 or 30 Hz) is the filtered one; the raw speed read about 0.25 m/s from noise alone.
|
- **Smoothing.** The published landmarks go through a One Euro filter: it smooths hard while the hand is still (tracking noise is several mm per frame) and hardly at all while it moves fast. The palm speed that sets the update rate (15 or 30 Hz) is the filtered one; the raw speed read about 0.25 m/s from noise alone.
|
||||||
- **Capsules.** Forearms follow the hand's own axis, and nothing within 12 cm in front of the eyes is published.
|
- **Capsules.** Forearms follow the hand's own axis, and nothing within 12 cm in front of the eyes is published.
|
||||||
|
|
||||||
|
## Pinch
|
||||||
|
|
||||||
|
For input, the Vision Pro way: look at something and pinch to click, pinch and move to drag, with the eye tracker doing the looking. fh-tracker detects a pinch per hand (`pinch.h`) and publishes it to `$XDG_RUNTIME_DIR/frame-hands/gestures`, next to the hands file. The layout, and how to read it without missing quick taps, is in `include/fh_gestures.h`.
|
||||||
|
|
||||||
|
- A pinch begins when the thumb and index tips come within `--pinch-begin` (default 0.020 m). It ends when they open past `--pinch-end` (0.035 m) for 2 processed frames in a row, or when the hand stays lost for 0.25 s (flagged lost). While a pinch is down or closing, the tracker runs at the full 30 Hz.
|
||||||
|
- The distance comes from MediaPipe's world landmarks (the model's own 3D hand pose, averaged over the hand's views, at the user's hand size). `--pinch-triangulated` uses the triangulated tips instead. On the two recordings without deliberate pinches, the world landmarks came under 2 cm in 0.2-1% of frames, against 3.3-4.5% for the triangulated tips. Typing still gave 2 pinches a minute, so a consumer should only act on a pinch while the gaze is on a target.
|
||||||
|
- The pinch point is midway between the thumb and index tips. A drag is the pinch point now, minus where it was when the pinch began, both turned into the room with the HMD pose at their capture times.
|
||||||
|
- `tools/watch_gestures.py` prints begins, ends and drag offsets live, and `--distance` prints each hand's distance. `fh-replay` runs the same detector and reports pinch counts. Its `--timeline` gets each begin, end and lost event, and both distance measures per set.
|
||||||
|
|
||||||
|
Frametop's pointer helper is the natural consumer. Its gaze mode already treats a press as "stop where the gaze put it, drag onto the target, click on release", and "hold still for half a second, then move" as a drag. A pinch begin would be the press, the end the release, and the pinch point's movement the drag.
|
||||||
|
|
||||||
## Replay
|
## Replay
|
||||||
|
|
||||||
`fh-replay DIR` runs a recording through the tracker with the live scheduling and reports how well it kept the hands: hands per set, left and right coverage, track lengths, and the same reasons as the status line.
|
`fh-replay DIR` runs a recording through the tracker with the live scheduling and reports how well it kept the hands: hands per set, left and right coverage, track lengths, and the same reasons as the status line.
|
||||||
@@ -46,6 +58,8 @@ trackd/fh-replay captures/rec-20260929-120000 --oracle 10 --timeline /tmp/tl.txt
|
|||||||
- `--oracle N`: every N-th set, also search every tile of every camera, and report how often the tracker had the hands that full search could find.
|
- `--oracle N`: every N-th set, also search every tile of every camera, and report how often the tracker had the hands that full search could find.
|
||||||
- `--slow F`: live, the tracker skips sets that arrive while it's busy. Replay counts each step's time times F as busy (default 1; the headset is busier live).
|
- `--slow F`: live, the tracker skips sets that arrive while it's busy. Replay counts each step's time times F as busy (default 1; the headset is busier live).
|
||||||
- `--timeline FILE`: a line per processed set and hand.
|
- `--timeline FILE`: a line per processed set and hand.
|
||||||
|
- `--cams mono|color|all`: which cameras to track with (default `mono`). `color` tracks with the Arcturus pair alone, for comparing it with the IR cameras on the same recording. It needs a recording made with `fh-camd --with-color`. `--color-left NODE` (`color_video0` or `color_video3`) and `--color-crop subtract|none` say how the module's calibration maps onto the images; `tools/check_color.py` finds out. Color frames repeat across sets (the newest one is saved with each), and repeats are skipped.
|
||||||
|
- `--depth FILE`: a line per hand per processed set for `tools/depth_report.py`, which measures the depth without ground truth. It reports how the hands were seen (by the two lower cameras, a lower and an upper one, or one camera), the noise along the line of sight against across it, each camera's one-view distance against the triangulated one, and what a camera dropping out would do to the distance.
|
||||||
|
|
||||||
## Playing a recording live
|
## Playing a recording live
|
||||||
|
|
||||||
|
|||||||
@@ -6,6 +6,8 @@
|
|||||||
#include <algorithm>
|
#include <algorithm>
|
||||||
|
|
||||||
#include <fstream>
|
#include <fstream>
|
||||||
|
#include <iterator>
|
||||||
|
#include <memory>
|
||||||
|
|
||||||
namespace {
|
namespace {
|
||||||
|
|
||||||
@@ -129,6 +131,55 @@ bool load_calibration(std::map<std::string, Camera> &out, std::string &err) {
|
|||||||
return !out.empty();
|
return !out.empty();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
|
||||||
|
const std::string &right_node, bool crop_subtract, int scale, std::string &err) {
|
||||||
|
// The module's EEPROM: some binary, then the calibration as JSON (world-readable)
|
||||||
|
const std::string path = device_path("/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom");
|
||||||
|
std::ifstream f(path, std::ios::binary);
|
||||||
|
const std::string raw((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
||||||
|
const size_t key = raw.find("\"alignment_method\"");
|
||||||
|
const size_t start = key == std::string::npos ? key : raw.rfind('{', key);
|
||||||
|
Json::Value rig, dev;
|
||||||
|
std::string e;
|
||||||
|
std::unique_ptr<Json::CharReader> reader(Json::CharReaderBuilder().newCharReader());
|
||||||
|
if (start == std::string::npos || !reader->parse(raw.data() + start, raw.data() + raw.size(), &rig, &e))
|
||||||
|
return err = path + ": no calibration JSON " + e, false;
|
||||||
|
if (!read_json(device_path("/persist/device_config.json").c_str(), dev, err)) return false;
|
||||||
|
double cad_from_head[4][4], head_from_cad[4][4];
|
||||||
|
pose(dev["head"], 1.0, cad_from_head);
|
||||||
|
invert_rigid(cad_from_head, head_from_cad);
|
||||||
|
constexpr int kValidWidth = 1972; // pixels per row XRService's buffers deliver (of 2464)
|
||||||
|
int n = 0;
|
||||||
|
for (const Json::Value &c : rig["cameras"]) {
|
||||||
|
const std::string source = c["sourceCamera"].asString();
|
||||||
|
const std::string name = source == "passthrough_left" ? left_node : source == "passthrough_right" ? right_node : "";
|
||||||
|
if (name.empty()) continue;
|
||||||
|
Camera cam;
|
||||||
|
cam.name = name;
|
||||||
|
cam.width = kValidWidth / scale, cam.height = c["height"].asInt() / scale;
|
||||||
|
const double dx = crop_subtract ? c["cropRegion"]["x"].asDouble() : 0, dy = crop_subtract ? c["cropRegion"]["y"].asDouble() : 0;
|
||||||
|
for (const Json::Value &in : c["intrinsics"]) {
|
||||||
|
if (in["cameraModel"].asString() != "kb") continue;
|
||||||
|
// integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
|
||||||
|
cam.fx = in["fx"].asDouble() / scale, cam.fy = in["fy"].asDouble() / scale;
|
||||||
|
cam.cx = (in["cx"].asDouble() - dx + 0.5) / scale - 0.5, cam.cy = (in["cy"].asDouble() - dy + 0.5) / scale - 0.5;
|
||||||
|
cam.k[0] = in["k1"].asDouble(), cam.k[1] = in["k2"].asDouble();
|
||||||
|
cam.k[2] = in["k3"].asDouble(), cam.k[3] = in["k4"].asDouble();
|
||||||
|
}
|
||||||
|
double cad_from_cam[4][4], head_from_cam[4][4];
|
||||||
|
pose(c["extrinsics"], 1e-3, cad_from_cam);
|
||||||
|
mul(head_from_cad, cad_from_cam, head_from_cam);
|
||||||
|
for (int i = 0; i < 3; ++i) {
|
||||||
|
for (int j = 0; j < 3; ++j) cam.R[i][j] = head_from_cam[i][j];
|
||||||
|
cam.origin[i] = head_from_cam[i][3];
|
||||||
|
}
|
||||||
|
out[name] = cam;
|
||||||
|
++n;
|
||||||
|
}
|
||||||
|
if (n != 2) err = path + ": expected passthrough_left and passthrough_right";
|
||||||
|
return n == 2;
|
||||||
|
}
|
||||||
|
|
||||||
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms) {
|
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms) {
|
||||||
double A[3][3] = {}, b[3] = {};
|
double A[3][3] = {}, b[3] = {};
|
||||||
for (int v = 0; v < n; ++v) {
|
for (int v = 0; v < n; ++v) {
|
||||||
|
|||||||
@@ -27,5 +27,13 @@ struct Camera {
|
|||||||
// name: slam_left, slam_right, upper_left, upper_right.
|
// name: slam_left, slam_right, upper_left, upper_right.
|
||||||
bool load_calibration(std::map<std::string, Camera> &out, std::string &err);
|
bool load_calibration(std::map<std::string, Camera> &out, std::string &err);
|
||||||
|
|
||||||
|
// The Arcturus color cameras (tracker/calib.py load_color has the conventions), for
|
||||||
|
// fh-camd --with-color's images: luma at 1/scale size, recorded as color_video<N>. They're
|
||||||
|
// keyed by those recorded names: left_node is passthrough_left, right_node
|
||||||
|
// passthrough_right. crop_subtract: image x = sensor x - the calibration's cropRegion.x.
|
||||||
|
// tools/check_color.py tells which node is which and which crop reading fits.
|
||||||
|
bool load_color_calibration(std::map<std::string, Camera> &out, const std::string &left_node,
|
||||||
|
const std::string &right_node, bool crop_subtract, int scale, std::string &err);
|
||||||
|
|
||||||
// The point closest to several rays (weighted), and its rms distance to them.
|
// The point closest to several rays (weighted), and its rms distance to them.
|
||||||
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms);
|
V3 triangulate(const V3 *origins, const V3 *dirs, const double *weights, int n, double *rms);
|
||||||
+40
-9
@@ -5,6 +5,7 @@
|
|||||||
// fh-tracker [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N]
|
// fh-tracker [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N]
|
||||||
// [--no-publish] [--record DIR] [--swap-sides] ... (--help lists them all)
|
// [--no-publish] [--record DIR] [--swap-sides] ... (--help lists them all)
|
||||||
#include "io.h"
|
#include "io.h"
|
||||||
|
#include "pinch.h"
|
||||||
#include "record.h"
|
#include "record.h"
|
||||||
|
|
||||||
#include <sched.h>
|
#include <sched.h>
|
||||||
@@ -63,6 +64,7 @@ int main(int argc, char **argv) {
|
|||||||
// dim recording), but makes the landmarks jitter, so they get plain crops.
|
// dim recording), but makes the landmarks jitter, so they get plain crops.
|
||||||
Contrast palm_contrast, hand_contrast{Contrast::None};
|
Contrast palm_contrast, hand_contrast{Contrast::None};
|
||||||
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
||||||
|
PinchParams pinch_params;
|
||||||
double record_for = 120;
|
double record_for = 120;
|
||||||
for (int i = 1; i < argc; ++i) {
|
for (int i = 1; i < argc; ++i) {
|
||||||
const std::string a = argv[i];
|
const std::string a = argv[i];
|
||||||
@@ -74,6 +76,9 @@ int main(int argc, char **argv) {
|
|||||||
else if (a == "--nice" && more) niceness = std::atoi(argv[++i]);
|
else if (a == "--nice" && more) niceness = std::atoi(argv[++i]);
|
||||||
else if (a == "--int8") int8 = true;
|
else if (a == "--int8") int8 = true;
|
||||||
else if (a == "--no-publish") publish = false;
|
else if (a == "--no-publish") publish = false;
|
||||||
|
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
|
||||||
|
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
|
||||||
|
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
|
||||||
else if (a == "--swap-sides") swap_sides = true;
|
else if (a == "--swap-sides") swap_sides = true;
|
||||||
else if (a == "--record-only") track = publish = false;
|
else if (a == "--record-only") track = publish = false;
|
||||||
else if (a == "--ring" && more) ring_path = argv[++i];
|
else if (a == "--ring" && more) ring_path = argv[++i];
|
||||||
@@ -96,11 +101,12 @@ int main(int argc, char **argv) {
|
|||||||
std::printf("usage: %s [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] [--no-publish]\n"
|
std::printf("usage: %s [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] [--no-publish]\n"
|
||||||
" [--record DIR] [--record-for S] [--record-only] [--cpus 5,6,7] [--swap-sides]\n"
|
" [--record DIR] [--record-for S] [--record-only] [--cpus 5,6,7] [--swap-sides]\n"
|
||||||
" [--keep-presence P] (0.5) [--ring PATH] (fh-camd's, or fh-ringplay's)\n"
|
" [--keep-presence P] (0.5) [--ring PATH] (fh-camd's, or fh-ringplay's)\n"
|
||||||
|
" [--pinch-begin M] (0.020) [--pinch-end M] (0.035) [--pinch-triangulated]\n"
|
||||||
" [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch; default clahe:2/none)\n"
|
" [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch; default clahe:2/none)\n"
|
||||||
"Recording saves every frame set for S seconds (120) to DIR/sets.bin, for fh-replay; SIGUSR1\n"
|
"Recording saves every frame set for S seconds (120) to DIR/sets.bin, for fh-replay; SIGUSR1\n"
|
||||||
"starts one in captures/rec-<time> next to trackd. --record-only records without tracking, so it\n"
|
"starts one in captures/rec-<time> next to trackd. --record-only records without tracking, so it\n"
|
||||||
"can run beside a tracking fh-tracker. With fh-camd --with-dark, recordings also get each\n"
|
"can run beside a tracking fh-tracker. With fh-camd --with-dark, recordings also get each\n"
|
||||||
"camera's newest dark frame, as <name>_dk.\n",
|
"camera's newest dark frame, as <name>_dk; with --with-color, the color cameras' as color_video<N>.\n",
|
||||||
argv[0]);
|
argv[0]);
|
||||||
return a == "--help" ? 0 : 1;
|
return a == "--help" ? 0 : 1;
|
||||||
}
|
}
|
||||||
@@ -115,6 +121,8 @@ int main(int argc, char **argv) {
|
|||||||
Ring ring;
|
Ring ring;
|
||||||
Nets nets;
|
Nets nets;
|
||||||
Publisher pub;
|
Publisher pub;
|
||||||
|
GesturePublisher gestures;
|
||||||
|
Pinch pinch(pinch_params);
|
||||||
std::unique_ptr<Recorder> rec;
|
std::unique_ptr<Recorder> rec;
|
||||||
uint64_t rec_start = 0;
|
uint64_t rec_start = 0;
|
||||||
auto start_recording = [&](const std::string &dir, std::string &e) {
|
auto start_recording = [&](const std::string &dir, std::string &e) {
|
||||||
@@ -126,18 +134,23 @@ int main(int argc, char **argv) {
|
|||||||
return true;
|
return true;
|
||||||
};
|
};
|
||||||
if (!load_calibration(calib, err) || !ring.open(ring_path.c_str(), err) || !nets.load(models, int8, err) ||
|
if (!load_calibration(calib, err) || !ring.open(ring_path.c_str(), err) || !nets.load(models, int8, err) ||
|
||||||
(publish && !pub.open(err)) || (!record.empty() && !start_recording(record, err))) {
|
(publish && (!pub.open(err) || !gestures.open(pinch, err))) || (!record.empty() && !start_recording(record, err))) {
|
||||||
std::fprintf(stderr, "%s\n", err.c_str());
|
std::fprintf(stderr, "%s\n", err.c_str());
|
||||||
return 1;
|
return 1;
|
||||||
}
|
}
|
||||||
if (!ring.alive()) return std::fprintf(stderr, "fh-camd isn't running (no heartbeat)\n"), 1;
|
if (!ring.alive()) return std::fprintf(stderr, "fh-camd isn't running (no heartbeat)\n"), 1;
|
||||||
|
|
||||||
std::map<std::string, int> index; // calibration name -> ring camera
|
std::map<std::string, int> index; // calibration name -> ring camera
|
||||||
// "<name>_dk" -> ring camera (fh-camd --with-dark): recorded only. Recorded names hold 15
|
// Recorded only, not tracked: "<name>_dk" (fh-camd --with-dark) and "color_video<N>"
|
||||||
// characters, so "upper_right_dark" wouldn't fit.
|
// (--with-color; which is left and right is up to tools/check_color.py). Recorded names
|
||||||
|
// hold 15 characters, so "upper_right_dark" wouldn't fit.
|
||||||
std::map<std::string, int> dark;
|
std::map<std::string, int> dark;
|
||||||
std::map<std::string, Camera> used;
|
std::map<std::string, Camera> used;
|
||||||
for (int i = 0; i < ring.cameras(); ++i) {
|
for (int i = 0; i < ring.cameras(); ++i) {
|
||||||
|
if (ring.camera(i).flags & FH_CAM_COLOR) {
|
||||||
|
dark["color_video" + std::to_string(ring.camera(i).node)] = i;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
// fh-camd's cameras by capture pipe; fh-ringplay's (no device) by the name it gives
|
// fh-camd's cameras by capture pipe; fh-ringplay's (no device) by the name it gives
|
||||||
const char *name = camera_for_pipe(ring.camera(i).node);
|
const char *name = camera_for_pipe(ring.camera(i).node);
|
||||||
if (!name && ring.camera(i).node < 0) name = ring.camera(i).name;
|
if (!name && ring.camera(i).node < 0) name = ring.camera(i).name;
|
||||||
@@ -231,9 +244,9 @@ int main(int argc, char **argv) {
|
|||||||
std::string e;
|
std::string e;
|
||||||
if (!start_recording(dir + "/" + name, e)) std::fprintf(stderr, "%s\n", e.c_str());
|
if (!start_recording(dir + "/" + name, e)) std::fprintf(stderr, "%s\n", e.c_str());
|
||||||
}
|
}
|
||||||
if (rec) { // about 80 MB/s, twice that with dark frames
|
if (rec) { // about 80 MB/s; dark frames double that, color frames add 70 MB/s
|
||||||
if ((mono_ns() - rec_start) / 1e9 < record_for) {
|
if ((mono_ns() - rec_start) / 1e9 < record_for) {
|
||||||
for (auto &[name, i] : dark) { // the newest dark frame of each camera, as it is
|
for (auto &[name, i] : dark) { // the newest dark and color frames, as they are
|
||||||
fh_ring_slot_t meta;
|
fh_ring_slot_t meta;
|
||||||
const uint64_t n = ring.latest(i);
|
const uint64_t n = ring.latest(i);
|
||||||
const auto &c = ring.camera(i);
|
const auto &c = ring.camera(i);
|
||||||
@@ -256,8 +269,16 @@ int main(int argc, char **argv) {
|
|||||||
}
|
}
|
||||||
if (!track || tmin < next_ns) continue; // not needed yet at the current rate
|
if (!track || tmin < next_ns) continue; // not needed yet at the current rate
|
||||||
const auto hands = tracker.step(images, int64_t(tmin));
|
const auto hands = tracker.step(images, int64_t(tmin));
|
||||||
next_ns = tmin + uint64_t((tracker.interval() - 0.005) * 1e9);
|
const uint64_t capture = uint64_t(int64_t(tmin) - raw_minus_mono_ns()); // CLOCK_MONOTONIC
|
||||||
if (publish) pub.write(hands, uint64_t(int64_t(tmin) - raw_minus_mono_ns()));
|
pinch.update(hands, tracker.views_now(), int64_t(capture));
|
||||||
|
// a pinch down or closing gets the full rate, even while the palm holds still
|
||||||
|
next_ns = tmin + uint64_t((std::min(tracker.interval(), pinch.engaged() ? 1 / 30.0 : 1.0) - 0.005) * 1e9);
|
||||||
|
if (publish) pub.write(hands, capture), gestures.write(pinch, capture);
|
||||||
|
for (const Pinch::Event &e : pinch.events) {
|
||||||
|
std::printf("pinch %s %-5s d %.3f m at %+.3f %+.3f %+.3f\n", e.side ? "right" : "left ", e.what, e.distance,
|
||||||
|
e.point[0], e.point[1], e.point[2]);
|
||||||
|
std::fflush(stdout);
|
||||||
|
}
|
||||||
lat.push_back((mono_ns() - dq) / 1e6);
|
lat.push_back((mono_ns() - dq) / 1e6);
|
||||||
hands_sum += double(hands.size());
|
hands_sum += double(hands.size());
|
||||||
bool on_left = false, on_right = false; // by where the wrist is, not the model's label
|
bool on_left = false, on_right = false; // by where the wrist is, not the model's label
|
||||||
@@ -285,6 +306,12 @@ int main(int argc, char **argv) {
|
|||||||
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten,
|
s.handoff_miss, s.dups, s.splits, s.created, s.merged, s.forgotten,
|
||||||
!rec ? "" : (" recorded " + std::to_string(rec->written()) + " dropped " +
|
!rec ? "" : (" recorded " + std::to_string(rec->written()) + " dropped " +
|
||||||
std::to_string(rec->dropped())).c_str());
|
std::to_string(rec->dropped())).c_str());
|
||||||
|
std::printf(" pinches: left %u right %u", pinch.side(0).begins, pinch.side(1).begins);
|
||||||
|
for (int k = 0; k < 2; ++k)
|
||||||
|
if (pinch.side(k).flags & FH_PINCH_TRACKED)
|
||||||
|
std::printf(" %s %s d %.3f", k ? "right" : "left", pinch.side(k).flags & FH_PINCH_DOWN ? "DOWN" : "open",
|
||||||
|
pinch.side(k).distance);
|
||||||
|
std::printf("\n");
|
||||||
for (const Hand *h : hands)
|
for (const Hand *h : hands)
|
||||||
std::printf(" hand %d %-5s views %d wrist %+.3f %+.3f %+.3f m scale %.2f speed %.2f m/s\n", h->id,
|
std::printf(" hand %d %-5s views %d wrist %+.3f %+.3f %+.3f m scale %.2f speed %.2f m/s\n", h->id,
|
||||||
h->right() ? "right" : "left", h->nviews, h->pts[0][0], h->pts[0][1], h->pts[0][2], h->scale,
|
h->right() ? "right" : "left", h->nviews, h->pts[0][0], h->pts[0][1], h->pts[0][2], h->scale,
|
||||||
@@ -295,6 +322,10 @@ int main(int argc, char **argv) {
|
|||||||
lat.clear(), hands_sum = 0, resid_sum = 0, resid_n = 0, left_sets = right_sets = both_sets = 0;
|
lat.clear(), hands_sum = 0, resid_sum = 0, resid_n = 0, left_sets = right_sets = both_sets = 0;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
if (publish) pub.write({}, mono_ns());
|
if (publish) {
|
||||||
|
pub.write({}, mono_ns());
|
||||||
|
pinch.release(int64_t(mono_ns())); // a drag in progress ends, as lost
|
||||||
|
gestures.write(pinch, mono_ns());
|
||||||
|
}
|
||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
@@ -0,0 +1,153 @@
|
|||||||
|
#include "pinch.h"
|
||||||
|
|
||||||
|
#include "io.h"
|
||||||
|
|
||||||
|
#include <fcntl.h>
|
||||||
|
#include <sys/mman.h>
|
||||||
|
#include <sys/stat.h>
|
||||||
|
#include <unistd.h>
|
||||||
|
|
||||||
|
#include <algorithm>
|
||||||
|
#include <cerrno>
|
||||||
|
#include <cstdlib>
|
||||||
|
#include <cstring>
|
||||||
|
|
||||||
|
namespace {
|
||||||
|
|
||||||
|
constexpr int kThumbTip = 4, kIndexTip = 8;
|
||||||
|
|
||||||
|
void put3(float out[3], V3 v) {
|
||||||
|
for (int k = 0; k < 3; ++k) out[k] = float(v[k]);
|
||||||
|
}
|
||||||
|
|
||||||
|
} // namespace
|
||||||
|
|
||||||
|
void Pinch::update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns) {
|
||||||
|
events.clear();
|
||||||
|
for (int s = 0; s < 2; ++s) {
|
||||||
|
fh_pinch_t &o = side_[s];
|
||||||
|
const bool down = o.flags & FH_PINCH_DOWN;
|
||||||
|
// the hand: while down, the one the pinch began on; else the best tracked hand of this side
|
||||||
|
const Hand *h = nullptr;
|
||||||
|
for (const Hand *c : hands) {
|
||||||
|
if (down ? c->id != follow_[s] : c->right() != (s == 1)) continue;
|
||||||
|
if (!h || c->frames > h->frames) h = c;
|
||||||
|
}
|
||||||
|
world_d[s] = tri_d[s] = -1;
|
||||||
|
if (!h) {
|
||||||
|
o.flags &= ~FH_PINCH_TRACKED;
|
||||||
|
if (down && (t_ns - seen_ns_[s]) / 1e9 > p_.grace_s) end(s, t_ns, true);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
seen_ns_[s] = t_ns;
|
||||||
|
tri_d[s] = norm(h->pts[kThumbTip] - h->pts[kIndexTip]);
|
||||||
|
double sum = 0;
|
||||||
|
int n = 0;
|
||||||
|
for (const Seen &v : views) {
|
||||||
|
if (v.hand != h->id) continue;
|
||||||
|
const V3 a{v.lm.world[kThumbTip][0], v.lm.world[kThumbTip][1], v.lm.world[kThumbTip][2]};
|
||||||
|
const V3 b{v.lm.world[kIndexTip][0], v.lm.world[kIndexTip][1], v.lm.world[kIndexTip][2]};
|
||||||
|
sum += norm(a - b), ++n;
|
||||||
|
}
|
||||||
|
if (n) world_d[s] = sum / n * h->scale;
|
||||||
|
const double d = p_.triangulated || world_d[s] < 0 ? tri_d[s] : world_d[s];
|
||||||
|
const V3 point = (h->smooth[kThumbTip] + h->smooth[kIndexTip]) * 0.5;
|
||||||
|
o.flags |= FH_PINCH_TRACKED;
|
||||||
|
o.hand_id = uint32_t(h->id);
|
||||||
|
o.distance = float(d);
|
||||||
|
o.strength = float(std::clamp((p_.end_m - d) / (p_.end_m - p_.begin_m), 0.0, 1.0));
|
||||||
|
put3(o.point, point);
|
||||||
|
if (!down) {
|
||||||
|
if (d < p_.begin_m) {
|
||||||
|
o.flags = (o.flags | FH_PINCH_DOWN) & ~FH_PINCH_LOST;
|
||||||
|
++o.begins;
|
||||||
|
o.begin_ns = uint64_t(t_ns);
|
||||||
|
put3(o.begin_point, point);
|
||||||
|
follow_[s] = h->id;
|
||||||
|
open_frames_[s] = 0;
|
||||||
|
events.push_back({s, "begin", t_ns, d, point});
|
||||||
|
}
|
||||||
|
} else if (d > p_.end_m) {
|
||||||
|
if (++open_frames_[s] >= p_.end_frames) end(s, t_ns, false);
|
||||||
|
} else {
|
||||||
|
open_frames_[s] = 0;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
void Pinch::end(int s, int64_t t_ns, bool lost) {
|
||||||
|
fh_pinch_t &o = side_[s];
|
||||||
|
o.flags = (o.flags & ~FH_PINCH_DOWN) | (lost ? FH_PINCH_LOST : 0);
|
||||||
|
++o.ends;
|
||||||
|
o.end_ns = uint64_t(t_ns);
|
||||||
|
follow_[s] = 0;
|
||||||
|
events.push_back({s, lost ? "lost" : "end", t_ns, o.distance, {o.point[0], o.point[1], o.point[2]}});
|
||||||
|
}
|
||||||
|
|
||||||
|
void Pinch::release(int64_t t_ns) {
|
||||||
|
events.clear();
|
||||||
|
for (int s = 0; s < 2; ++s) {
|
||||||
|
side_[s].flags &= ~FH_PINCH_TRACKED;
|
||||||
|
if (side_[s].flags & FH_PINCH_DOWN) end(s, t_ns, true);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
bool Pinch::engaged() const {
|
||||||
|
for (const fh_pinch_t &o : side_)
|
||||||
|
if ((o.flags & FH_PINCH_DOWN) || ((o.flags & FH_PINCH_TRACKED) && o.strength > 0.3f)) return true;
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
bool GesturePublisher::open(const Pinch &pinch, std::string &err) {
|
||||||
|
const char *run = std::getenv("XDG_RUNTIME_DIR");
|
||||||
|
const std::string dir = std::string(run ? run : "/run/user/" + std::to_string(getuid())) + "/frame-hands";
|
||||||
|
mkdir(dir.c_str(), 0700);
|
||||||
|
chmod(dir.c_str(), 0700);
|
||||||
|
const std::string path = dir + "/gestures";
|
||||||
|
const int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600);
|
||||||
|
if (fd < 0 || ftruncate(fd, sizeof(fh_gestures_t)) < 0) return err = path + ": " + std::strerror(errno), false;
|
||||||
|
void *m = mmap(nullptr, sizeof(fh_gestures_t), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
|
||||||
|
close(fd);
|
||||||
|
if (m == MAP_FAILED) return err = path + ": can't map it", false;
|
||||||
|
out_ = static_cast<fh_gestures_t *>(m);
|
||||||
|
// keep the counters a previous tracker left, so a reader doesn't see them jump back
|
||||||
|
const bool ours = !std::memcmp(out_->magic, FH_GESTURES_MAGIC, 8) && out_->version == FH_GESTURES_VERSION;
|
||||||
|
if (!ours) {
|
||||||
|
std::memset(out_, 0, sizeof *out_);
|
||||||
|
std::memcpy(out_->magic, FH_GESTURES_MAGIC, 8);
|
||||||
|
out_->version = FH_GESTURES_VERSION;
|
||||||
|
out_->size = sizeof(fh_gestures_t);
|
||||||
|
}
|
||||||
|
seq_ = out_->seq / 2 + 1;
|
||||||
|
// a pinch the last tracker left down (it crashed) is over: count its end, as lost
|
||||||
|
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||||
|
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||||
|
for (fh_pinch_t &o : out_->pinch)
|
||||||
|
if (o.begins != o.ends) {
|
||||||
|
o.ends = o.begins;
|
||||||
|
o.end_ns = mono_ns();
|
||||||
|
o.flags = (o.flags & ~FH_PINCH_DOWN) | FH_PINCH_LOST;
|
||||||
|
}
|
||||||
|
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||||
|
out_->begin_m = float(pinch.params().begin_m);
|
||||||
|
out_->end_m = float(pinch.params().end_m);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
void GesturePublisher::write(const Pinch &pinch, uint64_t capture_ns) {
|
||||||
|
__atomic_store_n(&out_->seq, 2 * ++seq_ - 1, __ATOMIC_RELAXED);
|
||||||
|
__atomic_thread_fence(__ATOMIC_RELEASE);
|
||||||
|
for (int s = 0; s < 2; ++s) {
|
||||||
|
// counters carry on from what's in the file (a restarted tracker starts its own at 0)
|
||||||
|
const fh_pinch_t &in = pinch.side(s);
|
||||||
|
fh_pinch_t &o = out_->pinch[s];
|
||||||
|
const uint32_t base_b = o.begins - last_begins_[s], base_e = o.ends - last_ends_[s];
|
||||||
|
o = in;
|
||||||
|
o.begins = base_b + in.begins;
|
||||||
|
o.ends = base_e + in.ends;
|
||||||
|
last_begins_[s] = in.begins, last_ends_[s] = in.ends;
|
||||||
|
}
|
||||||
|
out_->capture_ns = capture_ns;
|
||||||
|
out_->publish_ns = mono_ns();
|
||||||
|
__atomic_store_n(&out_->seq, 2 * seq_, __ATOMIC_RELEASE);
|
||||||
|
}
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
// Pinch detection for input: look at something and pinch to click, pinch and move to drag.
|
||||||
|
// Per side, from the tracker's hands after each step; published as fh_gestures.h.
|
||||||
|
#pragma once
|
||||||
|
|
||||||
|
#include "tracker.h"
|
||||||
|
|
||||||
|
#include <string>
|
||||||
|
#include <vector>
|
||||||
|
|
||||||
|
extern "C" {
|
||||||
|
#include "../include/fh_gestures.h"
|
||||||
|
}
|
||||||
|
|
||||||
|
struct PinchParams {
|
||||||
|
double begin_m = 0.020; // thumb and index tips closer than this: the pinch begins
|
||||||
|
double end_m = 0.035; // further apart than this: it ends (the gap keeps it from flickering)
|
||||||
|
int end_frames = 2; // processed frames in a row past end_m before it ends, so one
|
||||||
|
// noisy frame doesn't drop a drag
|
||||||
|
double grace_s = 0.25; // a pinching hand lost this long ends its pinch (FH_PINCH_LOST)
|
||||||
|
// Where the distance comes from: MediaPipe's world landmarks (the model's own 3D hand
|
||||||
|
// pose, averaged over the hand's views, at the user's hand size), or the tracker's
|
||||||
|
// triangulated tips. The model's pose should hold up better when the fingers hide each
|
||||||
|
// other; tomorrow's recordings will tell.
|
||||||
|
bool triangulated = false;
|
||||||
|
};
|
||||||
|
|
||||||
|
class Pinch {
|
||||||
|
public:
|
||||||
|
explicit Pinch(const PinchParams &p = {}) : p_(p) {}
|
||||||
|
const PinchParams ¶ms() const { return p_; }
|
||||||
|
// After each processed set: the hands out of Tracker::step, the tracker's views (for
|
||||||
|
// the world landmarks) and the capture time.
|
||||||
|
void update(const std::vector<const Hand *> &hands, const std::vector<Seen> &views, int64_t t_ns);
|
||||||
|
// Ends any pinch that's down (as lost), e.g. when the tracker stops.
|
||||||
|
void release(int64_t t_ns);
|
||||||
|
const fh_pinch_t &side(int s) const { return side_[s]; } // 0 left, 1 right
|
||||||
|
// A pinch is down or closing: worth tracking at the full rate.
|
||||||
|
bool engaged() const;
|
||||||
|
|
||||||
|
// What changed in the last update, for logs.
|
||||||
|
struct Event {
|
||||||
|
int side;
|
||||||
|
const char *what; // "begin", "end", "lost"
|
||||||
|
int64_t t_ns;
|
||||||
|
double distance;
|
||||||
|
V3 point;
|
||||||
|
};
|
||||||
|
std::vector<Event> events;
|
||||||
|
// Both distance measures for the last update, per side (-1: no hand), for logs.
|
||||||
|
double world_d[2] = {-1, -1}, tri_d[2] = {-1, -1};
|
||||||
|
|
||||||
|
private:
|
||||||
|
void end(int s, int64_t t_ns, bool lost);
|
||||||
|
PinchParams p_;
|
||||||
|
fh_pinch_t side_[2]{};
|
||||||
|
int follow_[2] = {0, 0}; // the hand id a pinch follows while down
|
||||||
|
int open_frames_[2] = {0, 0};
|
||||||
|
int64_t seen_ns_[2] = {0, 0};
|
||||||
|
};
|
||||||
|
|
||||||
|
// Writes $XDG_RUNTIME_DIR/frame-hands/gestures.
|
||||||
|
class GesturePublisher {
|
||||||
|
public:
|
||||||
|
bool open(const Pinch &pinch, std::string &err);
|
||||||
|
void write(const Pinch &pinch, uint64_t capture_ns);
|
||||||
|
|
||||||
|
private:
|
||||||
|
fh_gestures_t *out_ = nullptr;
|
||||||
|
uint64_t seq_ = 0;
|
||||||
|
uint32_t last_begins_[2] = {0, 0}, last_ends_[2] = {0, 0}; // Pinch's counters last written
|
||||||
|
};
|
||||||
+92
-8
@@ -13,8 +13,19 @@
|
|||||||
// --timeline: per processed set, a line per hand (time, id, side, views, wrist) and per view
|
// --timeline: per processed set, a line per hand (time, id, side, views, wrist) and per view
|
||||||
// (hand, camera, presence, next crop, set index).
|
// (hand, camera, presence, next crop, set index).
|
||||||
// --keep-presence P: landmark presence a tracked view needs to stay (default 0.5, as new ones).
|
// --keep-presence P: landmark presence a tracked view needs to stay (default 0.5, as new ones).
|
||||||
|
// --pinch-begin M, --pinch-end M, --pinch-triangulated: the pinch detector (trackd/pinch.h);
|
||||||
|
// the timeline gets its begin/end/lost events and both distance measures per set.
|
||||||
|
// --cams mono|color|all: which cameras to track with (default mono). color and all need a
|
||||||
|
// recording made with fh-camd --with-color; --color-left NODE (color_video0 or
|
||||||
|
// color_video3) and --color-crop subtract|none say how its calibration maps
|
||||||
|
// (tools/check_color.py).
|
||||||
// --contrast: how the palm search's and the landmark model's crops are equalized
|
// --contrast: how the palm search's and the landmark model's crops are equalized
|
||||||
// (default clahe:2/none, as fh-tracker).
|
// (default clahe:2/none, as fh-tracker).
|
||||||
|
// --depth FILE: per processed set, a line per hand for tools/depth_report.py: its views'
|
||||||
|
// cameras, triangulation residual, hand scale, measured and published palm, and
|
||||||
|
// each view's one-view palm (Tracker::single_view at the hand's scale). The
|
||||||
|
// header has each camera's centre and focal length.
|
||||||
|
#include "pinch.h"
|
||||||
#include "record.h"
|
#include "record.h"
|
||||||
#include "tracker.h"
|
#include "tracker.h"
|
||||||
|
|
||||||
@@ -57,17 +68,26 @@ int main(int argc, char **argv) {
|
|||||||
bool cost = false;
|
bool cost = false;
|
||||||
Contrast palm_contrast, hand_contrast{Contrast::None}; // as fh-tracker's
|
Contrast palm_contrast, hand_contrast{Contrast::None}; // as fh-tracker's
|
||||||
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
double keep_presence = 0.5; // landmark presence a tracked view needs to stay
|
||||||
std::string timeline, models = std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
|
PinchParams pinch_params;
|
||||||
|
std::string use = "mono", color_left = "color_video0", color_crop = "subtract";
|
||||||
|
std::string timeline, depth, models =std::string(argv[0]).substr(0, std::string(argv[0]).rfind('/') + 1) + "../models/ncnn";
|
||||||
for (int i = 2; i < argc; ++i) {
|
for (int i = 2; i < argc; ++i) {
|
||||||
const std::string a = argv[i];
|
const std::string a = argv[i];
|
||||||
const bool more = i + 1 < argc;
|
const bool more = i + 1 < argc;
|
||||||
if (a == "--oracle" && more) oracle = std::atoi(argv[++i]);
|
if (a == "--oracle" && more) oracle = std::atoi(argv[++i]);
|
||||||
else if (a == "--slow" && more) slow = std::atof(argv[++i]);
|
else if (a == "--slow" && more) slow = std::atof(argv[++i]);
|
||||||
else if (a == "--timeline" && more) timeline = argv[++i];
|
else if (a == "--timeline" && more) timeline = argv[++i];
|
||||||
|
else if (a == "--depth" && more) depth = argv[++i];
|
||||||
else if (a == "--threads" && more) threads = std::atoi(argv[++i]);
|
else if (a == "--threads" && more) threads = std::atoi(argv[++i]);
|
||||||
else if (a == "--models" && more) models = argv[++i];
|
else if (a == "--models" && more) models = argv[++i];
|
||||||
else if (a == "--cost") cost = true;
|
else if (a == "--cost") cost = true;
|
||||||
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
|
else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]);
|
||||||
|
else if (a == "--cams" && more) use = argv[++i];
|
||||||
|
else if (a == "--pinch-begin" && more) pinch_params.begin_m = std::atof(argv[++i]);
|
||||||
|
else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]);
|
||||||
|
else if (a == "--pinch-triangulated") pinch_params.triangulated = true;
|
||||||
|
else if (a == "--color-left" && more) color_left = argv[++i];
|
||||||
|
else if (a == "--color-crop" && more) color_crop = argv[++i];
|
||||||
else if (a == "--contrast" && more) {
|
else if (a == "--contrast" && more) {
|
||||||
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
|
if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast))
|
||||||
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
|
return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1;
|
||||||
@@ -88,14 +108,30 @@ int main(int argc, char **argv) {
|
|||||||
std::vector<fh_set_cam_t> cams;
|
std::vector<fh_set_cam_t> cams;
|
||||||
std::vector<std::vector<uint8_t>> px;
|
std::vector<std::vector<uint8_t>> px;
|
||||||
if (!in.next(cams, px)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
|
if (!in.next(cams, px)) return std::fprintf(stderr, "%s: no sets\n", dir.c_str()), 1;
|
||||||
|
if (use != "mono" && use != "color" && use != "all") return std::fprintf(stderr, "--cams mono|color|all\n"), 1;
|
||||||
|
if (use != "mono") {
|
||||||
|
std::vector<std::string> nodes;
|
||||||
|
for (auto &c : cams)
|
||||||
|
if (std::string(c.name).rfind("color_video", 0) == 0) nodes.push_back(c.name);
|
||||||
|
if (nodes.size() != 2) return std::fprintf(stderr, "%s: no color cameras (fh-camd --with-color)\n", dir.c_str()), 1;
|
||||||
|
const std::string right = nodes[0] == color_left ? nodes[1] : nodes[0];
|
||||||
|
if (!load_color_calibration(calib, color_left, right, color_crop == "subtract", 2, err))
|
||||||
|
return std::fprintf(stderr, "%s\n", err.c_str()), 1;
|
||||||
|
}
|
||||||
std::map<std::string, Camera> used;
|
std::map<std::string, Camera> used;
|
||||||
for (auto &c : cams)
|
for (auto &c : cams) {
|
||||||
if (calib.count(c.name)) used[c.name] = calib[c.name];
|
const bool color = std::string(c.name).rfind("color_", 0) == 0;
|
||||||
|
if (calib.count(c.name) && (use == "all" || color == (use == "color"))) used[c.name] = calib[c.name];
|
||||||
|
}
|
||||||
Pool pool(threads, {2, 3, 4});
|
Pool pool(threads, {2, 3, 4});
|
||||||
Tracker tracker(used, nets, pool);
|
Tracker tracker(used, nets, pool);
|
||||||
tracker.set_keep_presence(keep_presence);
|
tracker.set_keep_presence(keep_presence);
|
||||||
|
FILE *dp = depth.empty() ? nullptr : std::fopen(depth.c_str(), "w");
|
||||||
|
if (dp)
|
||||||
|
for (auto &[name, c] : used)
|
||||||
|
std::fprintf(dp, "# cam %s %.4f %.4f %.4f %.1f\n", name.c_str(), c.origin[0], c.origin[1], c.origin[2], c.fx);
|
||||||
|
|
||||||
uint64_t t0 = 0, busy_until = 0, next_ns = 0;
|
uint64_t t0 = 0, busy_until = 0, next_ns = 0, t_prev = 0;
|
||||||
int index = -1; // of the set in the recording
|
int index = -1; // of the set in the recording
|
||||||
int nsets = 0, processed = 0, left = 0, right = 0, both = 0, hist[3] = {};
|
int nsets = 0, processed = 0, left = 0, right = 0, both = 0, hist[3] = {};
|
||||||
std::map<int, Track> tracks;
|
std::map<int, Track> tracks;
|
||||||
@@ -108,14 +144,27 @@ int main(int argc, char **argv) {
|
|||||||
// mm (steady motion cancels out; what's left is noise and real acceleration)
|
// mm (steady motion cancels out; what's left is noise and real acceleration)
|
||||||
std::vector<double> jit_raw, jit_sm;
|
std::vector<double> jit_raw, jit_sm;
|
||||||
int near_face = 0, hand_updates = 0; // published palms within 20 cm of the eyes
|
int near_face = 0, hand_updates = 0; // published palms within 20 cm of the eyes
|
||||||
|
Pinch pinch(pinch_params);
|
||||||
|
double pinch_begin_ts[2] = {0, 0};
|
||||||
|
std::vector<double> pinch_len[2]; // seconds, per side
|
||||||
|
int pinch_lost = 0;
|
||||||
do {
|
do {
|
||||||
std::map<std::string, Image> images;
|
std::map<std::string, Image> images;
|
||||||
uint64_t t = UINT64_MAX;
|
// the set's time: the mono cameras' when they're used (the color ones run on another
|
||||||
|
// clock); color frames can repeat across sets, so a set that doesn't move time on is skipped
|
||||||
|
uint64_t t = UINT64_MAX, t_color = UINT64_MAX;
|
||||||
for (size_t i = 0; i < cams.size(); ++i) {
|
for (size_t i = 0; i < cams.size(); ++i) {
|
||||||
if (!used.count(cams[i].name)) continue;
|
if (!used.count(cams[i].name)) continue;
|
||||||
images[cams[i].name] = {px[i].data(), int(cams[i].width), int(cams[i].height), int(cams[i].width)};
|
images[cams[i].name] = {px[i].data(), int(cams[i].width), int(cams[i].height), int(cams[i].width)};
|
||||||
t = std::min(t, cams[i].capture_ns);
|
uint64_t &ti = std::string(cams[i].name).rfind("color_", 0) == 0 ? t_color : t;
|
||||||
|
ti = std::min(ti, cams[i].capture_ns);
|
||||||
}
|
}
|
||||||
|
if (t == UINT64_MAX) t = t_color;
|
||||||
|
if (t <= t_prev) {
|
||||||
|
++index;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
t_prev = t;
|
||||||
if (!t0) t0 = t;
|
if (!t0) t0 = t;
|
||||||
const double ts = (t - t0) / 1e9;
|
const double ts = (t - t0) / 1e9;
|
||||||
++index;
|
++index;
|
||||||
@@ -134,7 +183,18 @@ int main(int argc, char **argv) {
|
|||||||
}
|
}
|
||||||
busy_ms += ms;
|
busy_ms += ms;
|
||||||
busy_until = t + uint64_t(ms * slow * 1e6) + 3'000'000; // + the ring hand-off
|
busy_until = t + uint64_t(ms * slow * 1e6) + 3'000'000; // + the ring hand-off
|
||||||
next_ns = t + uint64_t((tracker.interval() - 0.005) * 1e9);
|
const std::vector<Seen> seen = tracker.views_now();
|
||||||
|
pinch.update(out, seen, int64_t(t));
|
||||||
|
next_ns = t + uint64_t((std::min(tracker.interval(), pinch.engaged() ? 1 / 30.0 : 1.0) - 0.005) * 1e9);
|
||||||
|
for (const Pinch::Event &e : pinch.events) {
|
||||||
|
if (std::string(e.what) == "begin") pinch_begin_ts[e.side] = ts;
|
||||||
|
else pinch_len[e.side].push_back(ts - pinch_begin_ts[e.side]), pinch_lost += std::string(e.what) == "lost";
|
||||||
|
if (tl) std::fprintf(tl, "%.3f pinch %s %s d %.3f point %+.3f %+.3f %+.3f\n", ts, e.side ? "R" : "L", e.what,
|
||||||
|
e.distance, e.point[0], e.point[1], e.point[2]);
|
||||||
|
}
|
||||||
|
if (tl && (pinch.world_d[0] >= 0 || pinch.world_d[1] >= 0)) // both measures, for choosing one
|
||||||
|
std::fprintf(tl, "%.3f pinchd L world %.3f tri %.3f R world %.3f tri %.3f\n", ts, pinch.world_d[0],
|
||||||
|
pinch.tri_d[0], pinch.world_d[1], pinch.tri_d[1]);
|
||||||
++processed;
|
++processed;
|
||||||
last_out = out;
|
last_out = out;
|
||||||
bool l = false, r = false;
|
bool l = false, r = false;
|
||||||
@@ -158,10 +218,28 @@ int main(int argc, char **argv) {
|
|||||||
if (tl)
|
if (tl)
|
||||||
std::fprintf(tl, "%.3f %d %s %d %+.3f %+.3f %+.3f\n", ts, h->id, h->pts[0][0] < 0 ? "L" : "R", h->nviews,
|
std::fprintf(tl, "%.3f %d %s %d %+.3f %+.3f %+.3f\n", ts, h->id, h->pts[0][0] < 0 ? "L" : "R", h->nviews,
|
||||||
h->pts[0][0], h->pts[0][1], h->pts[0][2]);
|
h->pts[0][0], h->pts[0][1], h->pts[0][2]);
|
||||||
|
if (dp) {
|
||||||
|
std::vector<const Seen *> vs;
|
||||||
|
for (const Seen &v : seen)
|
||||||
|
if (v.hand == h->id) vs.push_back(&v);
|
||||||
|
std::sort(vs.begin(), vs.end(), [](const Seen *a, const Seen *b) { return a->cam < b->cam; });
|
||||||
|
std::string names;
|
||||||
|
for (const Seen *v : vs) names += (names.empty() ? "" : "+") + v->cam;
|
||||||
|
std::fprintf(dp, "%.4f %d %s %d %s %.4f %.3f %.4f %.4f %.4f %.4f %.4f %.4f", ts, h->id,
|
||||||
|
h->pts[0][0] < 0 ? "L" : "R", h->nviews, names.empty() ? "-" : names.c_str(), h->residual,
|
||||||
|
h->scale, raw[0], raw[1], raw[2], sm[0], sm[1], sm[2]);
|
||||||
|
for (const Seen *v : vs) {
|
||||||
|
V3 mono[21];
|
||||||
|
const bool ok = tracker.single_view(used.at(v->cam), v->lm, h->scale, mono);
|
||||||
|
const V3 p = ok ? palm(mono) : V3{NAN, NAN, NAN};
|
||||||
|
std::fprintf(dp, " %s %.2f %.4f %.4f %.4f", v->cam.c_str(), v->lm.presence, p[0], p[1], p[2]);
|
||||||
|
}
|
||||||
|
std::fputc('\n', dp);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
if (tl && out.empty()) std::fprintf(tl, "%.3f -\n", ts);
|
if (tl && out.empty()) std::fprintf(tl, "%.3f -\n", ts);
|
||||||
if (tl)
|
if (tl)
|
||||||
for (const Seen &v : tracker.views_now())
|
for (const Seen &v : seen)
|
||||||
std::fprintf(tl, "%.3f view %d %s presence %.2f roi %.0f %.0f %.0f %.3f set %d\n", ts, v.hand, v.cam.c_str(),
|
std::fprintf(tl, "%.3f view %d %s presence %.2f roi %.0f %.0f %.0f %.3f set %d\n", ts, v.hand, v.cam.c_str(),
|
||||||
v.lm.presence, v.roi.center[0], v.roi.center[1], v.roi.size, v.roi.rotation, index);
|
v.lm.presence, v.roi.center[0], v.roi.center[1], v.roi.size, v.roi.rotation, index);
|
||||||
left += l, right += r, both += l && r;
|
left += l, right += r, both += l && r;
|
||||||
@@ -187,6 +265,7 @@ int main(int argc, char **argv) {
|
|||||||
}
|
}
|
||||||
} while (in.next(cams, px));
|
} while (in.next(cams, px));
|
||||||
if (tl) std::fclose(tl);
|
if (tl) std::fclose(tl);
|
||||||
|
if (dp) std::fclose(dp);
|
||||||
|
|
||||||
const double secs = nsets > 1 ? nsets / 30.0 : 0;
|
const double secs = nsets > 1 ? nsets / 30.0 : 0;
|
||||||
const Stats &s = tracker.stats;
|
const Stats &s = tracker.stats;
|
||||||
@@ -221,6 +300,11 @@ int main(int argc, char **argv) {
|
|||||||
r[r.size() / 10], r[r.size() / 2], r[r.size() * 9 / 10], step[step.size() / 2], step[step.size() * 9 / 10]);
|
r[r.size() / 10], r[r.size() / 2], r[r.size() * 9 / 10], step[step.size() / 2], step[step.size() * 9 / 10]);
|
||||||
}
|
}
|
||||||
std::printf("palms within 20 cm of the eyes: %d of %d hand updates\n", near_face, hand_updates);
|
std::printf("palms within 20 cm of the eyes: %d of %d hand updates\n", near_face, hand_updates);
|
||||||
|
for (int k = 0; k < 2; ++k) std::sort(pinch_len[k].begin(), pinch_len[k].end());
|
||||||
|
std::printf("pinches (%s, %.3f/%.3f m): left %zu (median %.2f s), right %zu (median %.2f s), %d ended by losing the hand\n",
|
||||||
|
pinch_params.triangulated ? "triangulated tips" : "world landmarks", pinch_params.begin_m, pinch_params.end_m,
|
||||||
|
pinch_len[0].size(), pinch_len[0].empty() ? 0 : pinch_len[0][pinch_len[0].size() / 2], pinch_len[1].size(),
|
||||||
|
pinch_len[1].empty() ? 0 : pinch_len[1][pinch_len[1].size() / 2], pinch_lost);
|
||||||
std::sort(jit_raw.begin(), jit_raw.end());
|
std::sort(jit_raw.begin(), jit_raw.end());
|
||||||
std::sort(jit_sm.begin(), jit_sm.end());
|
std::sort(jit_sm.begin(), jit_sm.end());
|
||||||
if (!jit_raw.empty())
|
if (!jit_raw.empty())
|
||||||
|
|||||||
+4
-2
@@ -33,7 +33,9 @@ double ms_since(std::chrono::steady_clock::time_point t) {
|
|||||||
return std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t).count();
|
return std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t).count();
|
||||||
}
|
}
|
||||||
|
|
||||||
bool is_slam(const Camera &c) { return c.name.rfind("slam", 0) == 0; }
|
bool is_color(const Camera &c) { return c.name.rfind("color", 0) == 0; } // Arcturus, 145 degree image circle
|
||||||
|
// The wide cameras: the side ones and the color ones
|
||||||
|
bool is_slam(const Camera &c) { return c.name.rfind("slam", 0) == 0 || is_color(c); }
|
||||||
|
|
||||||
V2 palm_centre(const Landmarks &lm) { return lm.pts[9]; }
|
V2 palm_centre(const Landmarks &lm) { return lm.pts[9]; }
|
||||||
|
|
||||||
@@ -139,7 +141,7 @@ double Tracker::interval() const {
|
|||||||
bool Tracker::inside(const Camera &cam, V2 uv) const {
|
bool Tracker::inside(const Camera &cam, V2 uv) const {
|
||||||
const double m = 0.12;
|
const double m = 0.12;
|
||||||
return uv[0] >= m * cam.width && uv[0] <= (1 - m) * cam.width && uv[1] >= m * cam.height &&
|
return uv[0] >= m * cam.width && uv[0] <= (1 - m) * cam.width && uv[1] >= m * cam.height &&
|
||||||
uv[1] <= (1 - m) * cam.height && cam.off_axis(uv) < (is_slam(cam) ? 80.0 : 75.0);
|
uv[1] <= (1 - m) * cam.height && cam.off_axis(uv) < (is_color(cam) ? 70.0 : is_slam(cam) ? 80.0 : 75.0);
|
||||||
}
|
}
|
||||||
|
|
||||||
void Tracker::run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views) {
|
void Tracker::run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views) {
|
||||||
|
|||||||
+3
-1
@@ -90,6 +90,9 @@ public:
|
|||||||
// bright rooms the camera exposes for the room, the hands come out dim, and presence
|
// bright rooms the camera exposes for the room, the hands come out dim, and presence
|
||||||
// dips under 0.5 for a frame at a time.
|
// dips under 0.5 for a frame at a time.
|
||||||
void set_keep_presence(double p) { keep_presence_ = p; }
|
void set_keep_presence(double p) { keep_presence_ = p; }
|
||||||
|
// One view's 3D hand: each landmark along its ray, as far as how big the palm looks says
|
||||||
|
// for a hand `scale` times the model's (Hand::scale). False if the palm is degenerate.
|
||||||
|
bool single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const;
|
||||||
|
|
||||||
private:
|
private:
|
||||||
struct View {
|
struct View {
|
||||||
@@ -107,7 +110,6 @@ private:
|
|||||||
double size, rotation, weight, credit = 0;
|
double size, rotation, weight, credit = 0;
|
||||||
};
|
};
|
||||||
void run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views);
|
void run_landmarks(const std::map<std::string, Image> &images, std::vector<View *> &views);
|
||||||
bool single_view(const Camera &cam, const Landmarks &lm, double scale, V3 out[21]) const;
|
|
||||||
bool hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns);
|
bool hand_3d(Hand &hand, std::vector<View *> views, int64_t t_ns);
|
||||||
double pair_cost(const View &a, const View &b) const;
|
double pair_cost(const View &a, const View &b) const;
|
||||||
double size_misfit(const std::vector<const View *> &views, const V3 *pts) const;
|
double size_misfit(const std::vector<const View *> &views, const V3 *pts) const;
|
||||||
|
|||||||
@@ -18,6 +18,9 @@ import numpy as np
|
|||||||
|
|
||||||
XRSERVICE_JSON = '/persist/xrservice.json'
|
XRSERVICE_JSON = '/persist/xrservice.json'
|
||||||
DEVICE_JSON = '/persist/device_config.json'
|
DEVICE_JSON = '/persist/device_config.json'
|
||||||
|
# The Arcturus color module's EEPROM: some binary, then its calibration as JSON (world-readable)
|
||||||
|
ARCTURUS_EEPROM = '/sys/devices/platform/soc@0/ac15000.cci/i2c-0/0-0050/eeprom'
|
||||||
|
ARCTURUS_WIDTH = 1972 # valid pixels per row that XRService's buffers deliver (of 2464)
|
||||||
|
|
||||||
|
|
||||||
def _pose(d, scale=1.0):
|
def _pose(d, scale=1.0):
|
||||||
@@ -105,6 +108,36 @@ def load(xrservice=XRSERVICE_JSON, device=DEVICE_JSON):
|
|||||||
return cams
|
return cams
|
||||||
|
|
||||||
|
|
||||||
|
def load_color(eeprom=ARCTURUS_EEPROM, device=DEVICE_JSON, scale=2, crop='subtract'):
|
||||||
|
"""{"passthrough_left"/"passthrough_right": Camera} for the Arcturus color cameras, posed in
|
||||||
|
the head frame, for fh-camd --with-color's images (luma at 1/scale size).
|
||||||
|
|
||||||
|
Their calibration is in the CAD frame (mm) with pixel coordinates on the full 2464x2464
|
||||||
|
sensor; each camera also has a cropRegion. crop says how that maps to the delivered
|
||||||
|
image: 'subtract' (image x = sensor x - cropRegion.x) or 'none'. tools/check_color.py
|
||||||
|
tells which fits.
|
||||||
|
"""
|
||||||
|
with open(eeprom, 'rb') as f:
|
||||||
|
raw = f.read()
|
||||||
|
i = raw.rfind(b'{', 0, raw.find(b'"alignment_method"'))
|
||||||
|
rig, _ = json.JSONDecoder().raw_decode(raw[i:].decode('latin1'))
|
||||||
|
with open(device) as f:
|
||||||
|
dev = json.load(f)
|
||||||
|
head_from_cad = np.linalg.inv(_pose(dev['head']))
|
||||||
|
cams = {}
|
||||||
|
for c in rig['cameras']:
|
||||||
|
kb = dict(next(k for k in c['intrinsics'] if k['cameraModel'] == 'kb'))
|
||||||
|
region = c.get('cropRegion', {}) if crop == 'subtract' else {}
|
||||||
|
# integer pixel centres: sensor u -> image (u - crop + 0.5) / scale - 0.5
|
||||||
|
kb['cx'] = (kb['cx'] - region.get('x', 0) + 0.5) / scale - 0.5
|
||||||
|
kb['cy'] = (kb['cy'] - region.get('y', 0) + 0.5) / scale - 0.5
|
||||||
|
kb['fx'] /= scale
|
||||||
|
kb['fy'] /= scale
|
||||||
|
cams[c['sourceCamera']] = Camera(c['sourceCamera'], ARCTURUS_WIDTH // scale, c['height'] // scale, kb,
|
||||||
|
head_from_cad @ _pose(c['extrinsics'], 1e-3))
|
||||||
|
return cams
|
||||||
|
|
||||||
|
|
||||||
def triangulate(origins, dirs, weights=None):
|
def triangulate(origins, dirs, weights=None):
|
||||||
"""Least-squares point closest to several rays. Returns (point, rms distance to the rays)."""
|
"""Least-squares point closest to several rays. Returns (point, rms distance to the rays)."""
|
||||||
A = np.zeros((3, 3))
|
A = np.zeros((3, 3))
|
||||||
|
|||||||
Reference in new issue
Block a user