From 50c14545ed3f396b10c92414d970864e6929287e Mon Sep 17 00:00:00 2001 From: DeeJanuz <45082401+DeeJanuz@users.noreply.github.com> Date: Wed, 30 Sep 2026 14:31:17 -0600 Subject: [PATCH] Hands: pick the cameras by the light, and detect grips ft-hands tracks with the mono IR cameras in dim light and with every camera (or the colour pair, HANDS_BRIGHT) in bright light, going by the colour frames' mean brightness with hysteresis and a 2 s hold (HANDS_CAMERAS=auto, the default; mono, color and all fix it). Colour frames are placed on the mono cameras' clock by their dequeue time, and a view in a camera a step lacks waits for that camera's next frame. A grip (a closed hand) is a second gesture next to the pinch, in version 2 of the gestures file: every fingertip curled toward the wrist, beginning only on a hand seen open within a second and held up in front. On the 2026-09-30 lit recording that leaves 6 false grips of 14, all with the hands on the desk; pinch counts are unchanged. ft-handreplay logs grips and finger curl, watch_gestures.py shows them, and tools/cut_sets.py copies a few sets out of a recording. Co-Authored-By: Claude Opus 5.5 --- hands/README.md | 14 +- hands/include/fh_gestures.h | 66 +++--- hands/tools/cut_sets.py | 51 ++++ hands/tools/watch_gestures.py | 79 ++++--- hands/track/main.cpp | 428 ++++++++++++++++++++++++++-------- hands/track/pinch.cpp | 168 +++++++++++-- hands/track/pinch.h | 63 ++++- hands/track/replay.cpp | 26 ++- hands/track/tracker.cpp | 20 +- hands/track/tracker.h | 5 +- session/frametop.conf.example | 11 + 11 files changed, 746 insertions(+), 185 deletions(-) create mode 100644 hands/tools/cut_sets.py diff --git a/hands/README.md b/hands/README.md index d0c2649..3b3c877 100644 --- a/hands/README.md +++ b/hands/README.md @@ -56,9 +56,10 @@ The ring is mode 0600, in a folder only you can write. Frame handling: Options: - `--with-dark`: also publish the near-black frames, as extra ring cameras flagged `FH_CAM_DARK`. They show only light sources, so they're no use for hands. -- `--with-color`: also publish the two Arcturus colour cameras, flagged `FH_CAM_COLOR`. Each is the luma of the 10-bit frame's valid 1972x2464 (the top 8 bits), at half size (`--color-scale 2`: 986x1232) and at most 30 fps (`--color-fps`; the cameras run at 60). Frames that carry the module's warped half-size copy are dropped. Their `capture_ns` is on the colour module's clock (2.2 s off the mono cameras' on 2026-09-29), so line them up with the mono cameras by `dqbuf_ns`. Each frame costs about 0.65 ms of cache sync and 1.1 ms of decoding, so both cameras at 30 fps take about 11% of a core. +- `--with-color` (the service uses it): also publish the two Arcturus colour cameras, flagged `FH_CAM_COLOR`. Each is the luma of the 10-bit frame's valid 1972x2464 (the top 8 bits), at half size (`--color-scale 2`: 986x1232). They run at `--color-idle` (2 fps), enough for ft-hands to tell how bright it is, until a reader asks for more in `/run/user/UID/frametop-hands/color-fps` (ft-hands writes 30 while it tracks or records with them), up to `--color-fps` (30; the cameras run at 60). `HANDS_CAMERAS=mono` leaves them out. Frames that carry the module's warped half-size copy are dropped. Their `capture_ns` is on the colour module's clock (2.2 s off the mono cameras' on 2026-09-29), so line them up with the mono cameras by `dqbuf_ns`. Each frame costs about 0.65 ms of cache sync and 1.1 ms of decoding, so both cameras at 30 fps take about 11% of a core. +- Each mono camera's latest near-black frame's mean goes in the ring (`dark_mean`): a short fixed exposure, so it follows the room's IR light, sunlight above all. - The ring holds 8 cameras: 4 mono, plus 4 dark twins or 2 colour cameras. -- Colour isn't reliable yet. In the lit-room test of 2026-09-30, the colour cameras kept losing their buffer mapping: 30 frames in a row looked unchanged, the camera relearned, and after 5 relearns ft-camd exited. Each relearn samples all 32 colour buffers, which also made the mono cameras miss frames. Runs with the headset idle (no hands, no cutouts) had none of this, and no half-size copies either, while the failing runs had many. So the passthrough compositor may be writing into the colour buffers while Room View shows. Whether a frame is new is judged on the luma rows only: the chroma after them hardly changes in a lit room. `FT_CAMD_DEBUG=1` prints, at each colour stale frame, how many sampled words changed in every candidate buffer. +- Colour isn't reliable yet. In the lit-room test of 2026-09-30, the colour cameras kept losing their buffer mapping while the headset was worn: 30 frames in a row looked unchanged, the camera relearned, and after 5 relearns ft-camd exited. Each relearn probed all 32 colour buffers, a whole-buffer cache sync each, which also made the mono cameras miss frames. Runs with the headset idle had none of this. So the passthrough compositor may be writing into the colour buffers while Room View shows. Since then a colour camera never takes the mono ones down: it probes at most 4 buffers a frame, and one that goes stale twice in a row is paused (10 s, doubling up to 160 s) and learned again, without ft-camd exiting. Whether a frame is new is judged on the luma rows only: the chroma after them hardly changes in a lit room. `FT_CAMD_DEBUG=1` prints, at each colour stale frame, how many sampled words changed in every candidate buffer. - `--sensor S`: only the mono cameras whose sensor name contains S. - `--status S`: a status line every S seconds (0: never). @@ -90,6 +91,15 @@ Options: - `--record-only`: record without tracking or publishing, so it can run beside the live tracker. Give it `--record DIR`, since SIGUSR1 would reach both trackers. With `ft-camd --with-dark`, recordings also hold each camera's newest dark frame as `_dk`, which doubles the rate. With `--with-color`, each colour camera's newest frame is saved with every set, as `color_video`, which adds about 70 MB/s. Run the recorder at normal I/O priority: idle I/O priority stalled a 165 MB/s recording. - `--keep-presence P`: the landmark presence a tracked view needs to stay tracked. New views always need 0.5. Default 0.5. Lowering it to 0.2 barely helped in the bright recording, because lost hands drop to near-zero presence. - `--ring PATH`: read frames from another ring, such as `ft-ringplay`'s. +- `--cams auto|mono|color|all` (`HANDS_CAMERAS`, default `auto`): which cameras to track with. The mono IR cameras light the hands themselves and track well in dim rooms, but in bright light they expose for the room and the hands come out dark. The colour pair is the other way round. `auto` goes by the colour frames' mean brightness: at `--bright-on` (`HANDS_BRIGHT_ON`, 40) or over for 2 s it tracks with `--bright` (`HANDS_BRIGHT`: `all`, every camera, the default, or `color`), and under `--bright-off` (`HANDS_BRIGHT_OFF`, 25) for 2 s with the mono cameras again. A dim evening room read 9. The switch is logged (`cameras: mono -> all (...)`), and the status line gives the colour level, the mono cameras' ambient IR, and how many steps had colour frames. Colour frames arrive on their own schedule, so a step holds the mono set, the colour pair, or both, and views wait in their camera for its next frame. +- `--color-left NODE` (`HANDS_COLOR_LEFT`, `color_video0`) and `--color-crop subtract|none` (`HANDS_COLOR_CROP`, `subtract`): how the colour module's calibration maps onto the images. Not settled yet: `tools/check_color.py` on a recording with a lit, textured view tells. +- `--grip-begin R`, `--grip-end R`: the grip detector (below). + +**Gestures** (`/run/user/UID/frametop-hands/gestures`, `include/fh_gestures.h`), for the pointer helper: + +- A pinch: the thumb and index tips within 2 cm, ending past 3.5 cm. Not begun with the palm facing down (`--pinch-palm-down`, 0.6), which is how typing looks. +- A grip, a closed hand: every finger's tip nearer the wrist than 1.2 times its knuckle is (from the model's 3D hand, so hand size doesn't matter), ending when they open past 1.45 on average. It begins only on a hand seen open within the last second (closing it is the gesture), with the palm at most 35 degrees below straight ahead and at least 15 cm in front of the eyes. A grip ends a pinch on the same hand, as lost. In the 2026-09-30 lit recording (no deliberate fists), the checks cut false grips from 14 to 6, all with the hands on the desk while looking down at it; the pointer helper ignores grips that begin more than 30 cm below the eyes, which it can tell and ft-hands can't. +- `tools/watch_gestures.py --distance` shows both live; `ft-handreplay --timeline` logs them and each hand's finger curl. The status line also says how often a hand was on each side (by where the wrist is), and why views and hands came and went: views lost (the landmark model stopped seeing the hand), handoff misses (a crop projected from the hand's 3D position found nothing), duplicates, splits (two views disagreed in 3D), and hands created, merged and forgotten. diff --git a/hands/include/fh_gestures.h b/hands/include/fh_gestures.h index bb198dd..5a843fa 100644 --- a/hands/include/fh_gestures.h +++ b/hands/include/fh_gestures.h @@ -1,23 +1,32 @@ /* - * fh_gestures - pinch state ft-hands publishes for input: look at something and pinch to - * click it, pinch and move to drag (the Vision Pro model, with the eye tracker doing the - * looking). /run/user/UID/frametop-hands/gestures, next to the hands + * fh_gestures - hand gestures ft-hands publishes for input (the pointer helper): look at + * something and pinch to click it, or close the hand (a grip) to press and drag it (the + * Vision Pro model, with the eye tracker doing the looking). + * /run/user/UID/frametop-hands/gestures, next to the hands * file, with the same sequence lock (read seq, copy, read seq again; use the copy only if * both reads are the same even number) and the same frame: metres in the head frame at * capture time, OpenVR's HMD frame (+x right, +y up, -z forward). * - * One slot per side: pinch[0] is the left hand, pinch[1] the right. A pinch follows the - * hand it began on until it ends. It begins when the thumb and index tips close within - * begin_m and ends when they open past end_m (the gap between keeps it from flickering), - * or when the hand stays lost too long (FH_PINCH_LOST). + * One slot per side and gesture: pinch[0] and grip[0] are the left hand, [1] the right. + * A gesture follows the hand it began on until it ends. + * Pinch: begins when the thumb and index tips close within begin_m and ends when they + * open past end_m (the gap between keeps it from flickering). point is between the tips. + * Grip: a closed hand. It begins when all four fingers are curled in (each fingertip + * nearer the wrist than grip_begin times its knuckle is) and ends when they open past + * grip_end on average. distance is that average (about 2 open, under 1.2 closed), strength + * 0 open .. 1 closed, and point the palm's centre. A grip ends a pinch on the same hand + * (closing the hand can pass through a pinch on the way), as lost. + * Either ends, as lost (FH_PINCH_LOST), when its hand stays lost too long. * - * Don't miss short pinches: a reader that polls slower than a quick tap still sees it, - * because begins and ends count every pinch. When begins changed, a pinch began at - * begin_ns; when ends changed, one ended at end_ns. begins - ends is 1 while pinching. + * Don't miss short gestures: a reader that polls slower than a quick tap still sees it, + * because begins and ends count every one. When begins changed, one began at begin_ns; + * when ends changed, one ended at end_ns. begins - ends is 1 while it's down. * - * Drags: point is where the pinch is now, begin_point where it began. Turn each into the + * Drags: point is where the gesture is now, begin_point where it began. Turn each into the * room with the HMD pose at its capture time (capture_ns, begin_ns) before subtracting, * so turning your head doesn't drag. + * + * Version 1 had only the pinches (192 bytes); version 2 adds the grips after them. */ #pragma once @@ -26,27 +35,29 @@ #include #define FH_GESTURES_MAGIC "FHGEST01" -#define FH_GESTURES_VERSION 1 +#define FH_GESTURES_VERSION 2 enum { FH_PINCH_TRACKED = 1u << 0, /* the hand was tracked in this frame */ - FH_PINCH_DOWN = 1u << 1, /* pinching now */ - FH_PINCH_LOST = 1u << 2, /* the last pinch ended because the hand was lost */ + FH_PINCH_DOWN = 1u << 1, /* the gesture is held now */ + FH_PINCH_LOST = 1u << 2, /* the last one ended because the hand was lost */ + /* (or, for a pinch, a grip took over) */ }; typedef struct { uint32_t flags; /* FH_PINCH_* */ uint32_t hand_id; /* fh_hand_t.id of the hand, 0 if none */ - uint32_t begins; /* pinches begun so far */ - uint32_t ends; /* pinches ended so far */ + uint32_t begins; /* begun so far */ + uint32_t ends; /* ended so far */ uint64_t begin_ns; /* capture time (CLOCK_MONOTONIC) the current or */ - /* last pinch began */ - uint64_t end_ns; /* ... the last pinch ended */ - float distance; /* thumb tip to index tip, m, at this user's hand */ - /* size */ - float strength; /* 0 open (end_m or more) .. 1 closed (begin_m) */ - float point[3]; /* between the thumb and index tips */ - float begin_point[3]; /* point when the current or last pinch began */ + /* last one began */ + uint64_t end_ns; /* ... the last one ended */ + float distance; /* pinch: thumb tip to index tip, m, at this user's */ + /* hand size. grip: the fingers' mean curl (above) */ + float strength; /* 0 open .. 1 closed */ + float point[3]; /* pinch: between the thumb and index tips; grip: */ + /* the palm's centre */ + float begin_point[3]; /* point when the current or last one began */ } fh_pinch_t; /* 64 bytes */ typedef struct { @@ -56,11 +67,14 @@ typedef struct { volatile uint64_t seq; uint64_t capture_ns; /* CLOCK_MONOTONIC when the cameras took the frames */ uint64_t publish_ns; /* CLOCK_MONOTONIC when this was written */ - float begin_m; /* the thresholds in use */ + float begin_m; /* the pinch thresholds in use */ float end_m; - uint8_t reserved[16]; + float grip_begin; /* the grip thresholds in use (curl ratios) */ + float grip_end; + uint8_t reserved[8]; fh_pinch_t pinch[2]; /* [0] left hand, [1] right hand */ + fh_pinch_t grip[2]; /* version 2 */ } fh_gestures_t; static_assert(sizeof(fh_pinch_t) == 64, "fh_pinch_t layout"); -static_assert(sizeof(fh_gestures_t) == 64 + 2 * 64, "fh_gestures_t layout"); +static_assert(sizeof(fh_gestures_t) == 64 + 4 * 64, "fh_gestures_t layout"); diff --git a/hands/tools/cut_sets.py b/hands/tools/cut_sets.py new file mode 100644 index 0000000..2005667 --- /dev/null +++ b/hands/tools/cut_sets.py @@ -0,0 +1,51 @@ +"""Copy a few frame sets out of a recording (ft-hands --record) into a small one, to look at +or check elsewhere without moving gigabytes. Plain Python, so it runs on the Frame's host. + +usage: python3 tools/cut_sets.py REC_DIR OUT_DIR [--sets N] (8, spread evenly) | [--at I,J,...] +""" +import argparse +import os +import struct + +HDR = struct.Struct('<8sII') + + +def offsets(path): + offs, size = [], os.path.getsize(path) + with open(path, 'rb') as f: + off = 0 + while off + HDR.size <= size: + f.seek(off) + magic, _, nbytes = HDR.unpack(f.read(HDR.size)) + if magic[:7] != b'FHSET01' or off + nbytes > size: + break + offs.append((off, nbytes)) + off += nbytes + return offs + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument('rec') + ap.add_argument('out') + ap.add_argument('--sets', type=int, default=8) + ap.add_argument('--at', default='') + a = ap.parse_args() + src = os.path.join(a.rec, 'sets.bin') + offs = offsets(src) + if a.at: + pick = [int(i) for i in a.at.split(',')] + else: + n = max(1, min(a.sets, len(offs))) + pick = [round(i * (len(offs) - 1) / max(n - 1, 1)) for i in range(n)] + os.makedirs(a.out, exist_ok=True) + with open(src, 'rb') as f, open(os.path.join(a.out, 'sets.bin'), 'wb') as out: + for i in pick: + off, nbytes = offs[i] + f.seek(off) + out.write(f.read(nbytes)) + print('%d of %d sets (%s) -> %s' % (len(pick), len(offs), ','.join(map(str, pick)), a.out)) + + +if __name__ == '__main__': + main() diff --git a/hands/tools/watch_gestures.py b/hands/tools/watch_gestures.py index 8a14b90..c1ce16d 100644 --- a/hands/tools/watch_gestures.py +++ b/hands/tools/watch_gestures.py @@ -1,13 +1,14 @@ -"""Watch the pinch gestures ft-hands publishes, live: begins, ends, and drags. +"""Watch the gestures ft-hands publishes, live: pinches and grips, begins, ends, and drags. usage: python3 tools/watch_gestures.py [--every S] [--distance] -Prints a line when a pinch begins or ends on either hand. It goes by the counters, so a -quick tap between two reads still shows. While a pinch is held, every --every seconds -(default 0.1) it prints how far the pinch point has moved since it began, in the head -frame (turning your head moves it too; a real consumer turns both points into the room +Prints a line when a pinch or a grip (a closed hand) begins or ends on either hand. It goes +by the counters, so a quick tap between two reads still shows. While one is held, every +--every seconds (default 0.1) it prints how far its point has moved since it began, in the +head frame (turning your head moves it too; a real consumer turns both points into the room first, see include/fh_gestures.h). --distance also prints each hand's thumb-to-index -distance, to see how close a pinch comes to the thresholds. +distance and finger curl, to see how close a gesture comes to the thresholds. +Version 1 files (pinches only) still work. """ import argparse import mmap @@ -15,10 +16,11 @@ import os import struct import time -HDR = struct.Struct('<8sIIQQQff16x') # 64 bytes -PINCH = struct.Struct('= 2 and len(m) >= HDR.size + 4 * SLOT.size else 1 + g = [[SLOT.unpack_from(m, HDR.size + (k * 2 + s) * SLOT.size) for s in range(2)] for k in range(kinds)] + if kinds == 1: + g.append([(0,) * 14, (0,) * 14]) if struct.unpack_from('= begin_ns else 0 - print('%s pinch %s after %.2f s' % (SIDES[k], 'LOST' if flags & LOST else 'END', held), flush=True) - seen[k] = (begins, ends) - if flags & DOWN and now - last_drag >= a.every: - d = [point[i] - begin_point[i] for i in range(3)] - print('%s drag %+6.1f %+6.1f %+6.1f mm (%.0f mm)' % - (SIDES[k], *(1000 * x for x in d), 1000 * sum(x * x for x in d) ** 0.5), flush=True) - if any(q[0] & DOWN for q in p) and now - last_drag >= a.every: + for k, kind in enumerate(g): + for s, q in enumerate(kind): + flags, hand, begins, ends, begin_ns, end_ns, dist, strength = q[:8] + point, begin_point = q[8:11], q[11:14] + if begins != seen[k][s][0]: + print('%s %s BEGIN (#%d, hand %d) at %+.3f %+.3f %+.3f %s %.3f' % + (SIDES[s], KINDS[k], begins, hand, *begin_point, 'curl' if k else 'd', dist), flush=True) + if ends != seen[k][s][1]: + held = (end_ns - begin_ns) / 1e9 if end_ns >= begin_ns else 0 + print('%s %s %s after %.2f s' % (SIDES[s], KINDS[k], 'LOST' if flags & LOST else 'END', held), + flush=True) + seen[k][s] = (begins, ends) + if flags & DOWN and now - last_drag >= a.every: + d = [point[i] - begin_point[i] for i in range(3)] + print('%s %s drag %+6.1f %+6.1f %+6.1f mm (%.0f mm)' % + (SIDES[s], KINDS[k], *(1000 * x for x in d), 1000 * sum(x * x for x in d) ** 0.5), + flush=True) + if any(q[0] & DOWN for kind in g for q in kind) and now - last_drag >= a.every: last_drag = now if a.distance and now - last_dist >= 0.2: last_dist = now - print(' ' + ' '.join('%s %s' % (SIDES[k].strip(), 'd %.3f s %.2f' % (q[6], q[7]) if q[0] & TRACKED - else '-') for k, q in enumerate(p)), flush=True) + print(' ' + ' '.join( + '%s %s' % (SIDES[s].strip(), 'd %.3f curl %.2f' % (g[0][s][6], g[1][s][6]) if g[0][s][0] & TRACKED + else '-') for s in range(2)), flush=True) time.sleep(0.005) diff --git a/hands/track/main.cpp b/hands/track/main.cpp index 49549a8..40aedf0 100644 --- a/hands/track/main.cpp +++ b/hands/track/main.cpp @@ -1,12 +1,29 @@ // ft-hands: hands in 3D from ft-camd's ring, published for Frametop's ft-screens (the hand -// cutouts), and pinches for the pointer. It started as a port of frame-hands' Python -// prototype: the same scheduling, with the models on a few threads. +// cutouts), and pinches and grips for the pointer. It started as a port of frame-hands' +// Python prototype: the same scheduling, with the models on a few threads. // // ft-hands [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] -// [--no-publish] [--record DIR] [--swap-sides] ... (--help lists them all) +// [--no-publish] [--record DIR] [--swap-sides] [--cams auto|mono|color|all] ... +// (--help lists them all) +// +// Which cameras (--cams, HANDS_CAMERAS): the four mono IR cameras light the hands with their +// own IR and track well in dim rooms, but in bright light (a sunny room, a window behind the +// hands) they expose for the room and the hands come out dark. The two Arcturus colour +// cameras (the passthrough pair, forward-facing, 145 degrees) are the other way round: dark +// and grainy in a dim room, clear in a lit one. auto (the default) picks by how bright the +// colour cameras' frames are: at HANDS_BRIGHT_ON (mean luma) or over for 2 s, it tracks with +// HANDS_BRIGHT (all: every camera, so hands low at the sides stay in the side cameras; or +// color); under HANDS_BRIGHT_OFF for 2 s, with the mono cameras again. ft-camd runs the colour +// cameras at 2 fps, enough to tell the light, until ft-hands asks for 30 +// (/run/user/UID/frametop-hands/color-fps). The colour frames' capture times are on their +// own clock, so they're placed on the mono cameras' by when they were dequeued, less the +// mono cameras' measured delay. // // Settings in ~/.config/frametop.conf (FT_ in the environment overrides them, and -// options override both): HANDS_SWAP_SIDES (1: as --swap-sides), HANDS_CPUS (as --cpus). +// options override both): HANDS_SWAP_SIDES (1: as --swap-sides), HANDS_CPUS (as --cpus), +// HANDS_CAMERAS, HANDS_BRIGHT, HANDS_BRIGHT_ON, HANDS_BRIGHT_OFF, HANDS_COLOR_LEFT (which +// colour camera is passthrough_left: color_video0 or color_video3), HANDS_COLOR_CROP +// (subtract or none: tools/check_color.py tells both). #include "io.h" #include "pinch.h" #include "record.h" @@ -16,10 +33,12 @@ #include #include +#include #include #include #include +#include #include #include #include @@ -88,6 +107,44 @@ double cpu_seconds() { return r.ru_utime.tv_sec + r.ru_stime.tv_sec + (r.ru_utime.tv_usec + r.ru_stime.tv_usec) / 1e6; } +enum class Cams { Mono, Color, All }; + +const char *cams_name(Cams c) { return c == Cams::Mono ? "mono" : c == Cams::Color ? "color" : "all"; } + +bool parse_cams(const std::string &s, Cams &out) { + if (s == "mono") out = Cams::Mono; + else if (s == "color") out = Cams::Color; + else if (s == "all") out = Cams::All; + else return false; + return true; +} + +// How bright it is, for auto (see the top): the colour frames' mean luma, smoothed over about +// a second, with hysteresis and a 2 s hold each way. No colour frames for 3 s (ft-camd paused +// them, or has none) reads as dim. +struct Lighting { + double on = 40, off = 25; + double level = -1; + bool bright = false; + uint64_t at_ns = 0, since_ns = 0; // the last frame; since when it's wanted the other way + + void add(double mean, uint64_t t_ns) { + const double dt = at_ns && t_ns > at_ns ? (t_ns - at_ns) / 1e9 : 1.0; + level = level < 0 ? mean : level + (mean - level) * std::min(1.0, dt / 1.0); + at_ns = t_ns; + } + // True when it switched. + bool update(uint64_t now_ns) { + if (level >= 0 && now_ns - at_ns > 3'000'000'000ull) level = -1; + const bool want = level >= 0 && (bright ? level > off : level >= on); + if (want == bright) return since_ns = 0, false; + if (!since_ns) since_ns = now_ns; + if (now_ns - since_ns < 2'000'000'000ull) return false; + bright = want, since_ns = 0; + return true; + } +}; + } // namespace int main(int argc, char **argv) { @@ -104,11 +161,22 @@ int main(int argc, char **argv) { std::vector cpus = {5, 6, 7}; if (const auto c = parse_cpus(setting("HANDS_CPUS").c_str()); !c.empty()) cpus = c; swap_sides = setting("HANDS_SWAP_SIDES") == "1"; + // Which cameras (see the top). + std::string cams_arg = setting("HANDS_CAMERAS"), bright_arg = setting("HANDS_BRIGHT"); + std::string color_left = setting("HANDS_COLOR_LEFT"), color_crop = setting("HANDS_COLOR_CROP"); + if (cams_arg.empty()) cams_arg = "auto"; + if (bright_arg.empty()) bright_arg = "all"; + if (color_left.empty()) color_left = "color_video0"; + if (color_crop.empty()) color_crop = "subtract"; + Lighting light; + if (const std::string v = setting("HANDS_BRIGHT_ON"); !v.empty()) light.on = std::atof(v.c_str()); + if (const std::string v = setting("HANDS_BRIGHT_OFF"); !v.empty()) light.off = std::atof(v.c_str()); // How crops are equalized. CLAHE helps the palm search find hands (about 10% more in the // dim recording), but makes the landmarks jitter, so they get plain crops. Contrast palm_contrast, hand_contrast{Contrast::None}; double keep_presence = 0.5; // landmark presence a tracked view needs to stay PinchParams pinch_params; + GripParams grip_params; double record_for = 120; for (int i = 1; i < argc; ++i) { const std::string a = argv[i]; @@ -124,12 +192,20 @@ int main(int argc, char **argv) { else if (a == "--pinch-end" && more) pinch_params.end_m = std::atof(argv[++i]); else if (a == "--pinch-triangulated") pinch_params.triangulated = true; else if (a == "--pinch-palm-down" && more) pinch_params.palm_down_max = std::atof(argv[++i]); + else if (a == "--grip-begin" && more) grip_params.begin = std::atof(argv[++i]); + else if (a == "--grip-end" && more) grip_params.end = std::atof(argv[++i]); else if (a == "--swap-sides") swap_sides = true; else if (a == "--record-only") track = publish = false; else if (a == "--ring" && more) ring_path = argv[++i]; else if (a == "--record" && more) record = argv[++i]; else if (a == "--record-for" && more) record_for = std::atof(argv[++i]); else if (a == "--keep-presence" && more) keep_presence = std::atof(argv[++i]); + else if (a == "--cams" && more) cams_arg = argv[++i]; + else if (a == "--bright" && more) bright_arg = argv[++i]; + else if (a == "--bright-on" && more) light.on = std::atof(argv[++i]); + else if (a == "--bright-off" && more) light.off = std::atof(argv[++i]); + else if (a == "--color-left" && more) color_left = argv[++i]; + else if (a == "--color-crop" && more) color_crop = argv[++i]; else if (a == "--contrast" && more) { if (!Contrast::parse_pair(argv[++i], palm_contrast, hand_contrast)) return std::fprintf(stderr, "--contrast MODE or PALM/HAND, each clahe[:CLIP]|none|stretch\n"), 1; @@ -141,17 +217,27 @@ int main(int argc, char **argv) { std::printf("usage: %s [--seconds N] [--threads N] [--int8] [--status S] [--models DIR] [--nice N] [--no-publish]\n" " [--record DIR] [--record-for S] [--record-only] [--cpus 5,6,7] [--swap-sides]\n" " [--keep-presence P] (0.5) [--ring PATH] (ft-camd's, or ft-ringplay's)\n" + " [--cams auto|mono|color|all] (auto) [--bright all|color] (all) [--bright-on L] (40) [--bright-off L] (25)\n" + " [--color-left color_video0|color_video3] [--color-crop subtract|none]\n" " [--pinch-begin M] (0.020) [--pinch-end M] (0.035) [--pinch-triangulated] [--pinch-palm-down MAX] (0.6)\n" + " [--grip-begin R] (1.2) [--grip-end R] (1.45)\n" " [--contrast MODE|PALM/HAND] (clahe[:CLIP], none, stretch; default clahe:2/none)\n" "Recording saves every frame set for S seconds (120) to DIR/sets.bin, for ft-handreplay; SIGUSR1\n" "starts one in ~/.local/share/frametop/hands/rec-