Files
daniel-lynch--ovrplugin-ope…/docs/research/ghost-fix-2026-06-27.md
T
Daniel LynchandClaude Opus 4.8 a72a79ad29 Initial public release: OVRPlugin→OpenXR interoperability shim
An independent reimplementation of Meta's libOVRPlugin ABI on top of OpenXR, so
VrApi-era Meta Quest VR titles can run on non-Meta OpenXR runtimes (Monado,
Steam Frame) instead of being locked to Meta hardware. Original code only — no
Meta/Epic/Capcom binaries, headers, or assets. Includes a desktop harness that
drives the shim against Monado headless.

Scope/legal: interoperability; entitlement handling is out of scope. See README
for the legal/scope section and docs/ for the research trail and design notes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D6sFYGXZPsq3v7xtcDES6g
2026-06-29 00:48:48 -04:00

196 lines
15 KiB
Markdown

# P4 passthru — NATIVE ground-truth (2026-06-27, title screen, seated)
Captured via `debug.re4vr.passthru=1` (real libOVRPlugin owns the session through our shim).
Boots clean to the title screen, no ghost — so these are the CORRECT values to match.
## Eye layer desc (CalculateEyeLayerDesc2)
- size: **1440 x 1584** per eye
- FovL: U 1.111 D 1.192 L 0.933 R 1.000 (tangents)
- FovR: U 1.111 D 1.192 L 1.000 R 0.933
- => asymmetric (U≠D) AND canted stereo (per-eye L/R mirrored: inner edge = 1.000, outer = 0.933)
## EndFrame4 submit
- 1 layer, id=1 (first 3 frames) then id=2, **flags=0x4**, pose = all-zero (0,0,0)/(0,0,0,0)
- flags 0x4 = bit2 (NOT HeadLocked=0x1)
## Eye poses (GetNodePoseState3, head tilted on table)
- eye0 (L): pos=(-0.0934, 1.2170, 0.2354) quat=(0.0314, 0.1109, -0.0207, -0.9931)
- eye1 (R): pos=(-0.0271, 1.2203, 0.2503) quat=(same as L)
- both eyes share orientation; offset L->R ≈ (+0.066, +0.003, +0.015) m in world (IPD ~66mm)
## OUR shim values (captured passthru=0, title screen) — DIFF RESULT
- eye size: **1440x1584** — MATCHES native (the 1728x1900 was an old supersample config).
- FOV: ours via OpenXR xrLocateViews, logged as angles (rad). Steady-state L.fov(l,r,u,d)=
(-0.750,0.785,0.838,-0.873) -> tangents U1.110 D1.190 L0.932 R0.991 — MATCHES native
(U1.111 D1.192 L0.933 R1.000). The game does NOT call GetNodeFrustum2; FOV comes from
CalculateEyeLayerDesc2 in BOTH paths.
- IPD: 0.065-0.068 — MATCHES native ~0.066.
- eye orientation: e0.q==e1.q==head.q in both — MATCHES.
- submit flags: eye-fov LayerId flags=0x4 — MATCHES native (0x14 later = menu-quad layer
present, same game logic).
## CONCLUSION (title screen)
Static eye geometry (size/fov/ipd/orientation/flags) is IDENTICAL native vs shim at the title
screen. The seated in-game hand-deform ghost is therefore NOT a static-geometry mismatch.
ONE anomaly: our shim's FIRST ~5s report a WIDER outer FOV (L outer angle -0.855, tan 1.149)
that settles to native's -0.750/0.933 — a cold-start transient (matches title-ghost-is-coldload
-reprojection-judder; self-resolves warm). Not the persistent in-game ghost.
## CONFIRMED: passthru (native) has ZERO ghost — hands + title (user, 2026-06-27)
=> the ghost is in OUR shim's path, not the game/headset.
## *** ACTUAL ROOT CAUSE: DROPPED FRAMES -> compositor reprojection multiples (2026-06-27 PM) ***
The submit-ordering theory below was WRONG (submithook=2 present-on-submit AND submithook=3
+full device-wait-idle completion BOTH still ghost). Added always-on ANOMALY logging of the
OpenXR runtime RESPONSES (not our inputs, which all check out): locateViews validity,
WaitSwapchainImage timeout, xrEndFrame errors, frame-pacing. In-game ghost capture result:
- frame-pacing: 42 MISS / 3986 frames (~1%), gaps 21-137ms vs 13.9ms period
- locateViews invalid: 0 WaitSwapchain timeout: 0 xrEndFrame err: 0
=> the ONLY fault is DROPPED FRAMES. We miss the display deadline; the compositor reprojects the
held frame to fill the gap; during motion that = the flashing multiples. Scales exactly with user's
report: 21ms gap (1 drop)=mild, 137ms (~10 drops)=violent; load/position-dependent; non-deterministic;
clean dumped eye textures (pixels fine, just late). User CONFIRMED dump "no ghost" was the dump's
crawling fps freezing reprojection, not a fix — consistent.
KEY OPEN Q: are the drops OURS (shim latency, e.g. per-frame xrr_vk_flush_wait) or the game's load
(native hits them too but its compositor rides them out cleanly)? Added ENDFRAME-PACE log at
ovrp_EndFrame4 ENTRY (core.c) that runs in BOTH modes -> compare NATIVE vs shim gap rate.
- if NATIVE also hits 137ms gaps clean -> fix = match native compositor frame-timing/handling
(ExtraLatencyMode / phase sync / how late frames are submitted), NOT eliminate hitches.
- if NATIVE smooth -> our shim adds the latency -> reduce it (flush-wait is prime suspect).
NOTE for Steam Frame: passthru is Meta-only (vrapi); the OpenXR shim is the only cross-platform
path, so this drop/repro fix is what matters for portability.
## *** SUCCESS: game-thread pacing FIXED the ghost (2026-06-27, user-confirmed) ***
User: title ghost GONE, seated ghost GONE, standing black-flash GONE, "way better... feels like a
great success." Logs: BC (render-thread stall) 148ms->~3ms (FIXED); frame-pacing drops 0.79%
(180/22862, was ~1%+10%-in-motion, native 0.34%); WAITPACE steady 13.92ms. The ROOT CAUSE was: our
shim left the game thread UNPACED (WaitToBeginFrame no-op) and paced the render thread instead, which
desynced UE's internal pipeline and stalled the render thread during motion -> dropped frames ->
compositor judder = the "ghost". Fix = pace the game thread (xrWaitFrame in xrr_wait_frame) + FIFO
frameState handoff to render thread. Keep ffr=-1 (game FFR) + blackcount=0. RESIDUAL (tunable):
occasional "BeginFrame too many times" -> XR_FRAME_DISCARDED = brief hiccup; from the ring DROPPING
oldest frameState when render falls behind (desyncs wait/begin 1:1). FIX: block the game thread for
ring space (back-pressure) instead of dropping. THEN: strip diagnostic scaffolding (FLOOP/STALL/
JUDDER/WAITPACE/burst-dump/submithook leftovers), make ffr=-1+blackcount=0 defaults, commit.
## (attempt that became the fix) game-thread pacing rework
FLOOP trace proved the loop: WAIT(N) on GAME thread (tid A) runs 1 frame ahead of BEGIN/END(N-1)
on RENDER thread (tid B); steady BEGIN->END ~1ms. Old shim: WaitToBeginFrame=no-op, xrWaitFrame on
RENDER thread inside ovrp_BeginFrame4 (a workaround for "BeginFrame too many times"). REWORK (matches
vrapi + OpenXR's recommended pipelined model): xrWaitFrame now runs in xrr_wait_frame on the GAME
thread (blocks=paces it), NOT under g_xrlock; the frameState is handed to the render thread via a 1:1
FIFO ring (g_fsRing/g_fsHead/g_fsTail + g_fsCond). xrr_begin_frame pops the frameState (cond_timedwait
20ms) instead of calling xrWaitFrame; xrBeginFrame/xrEndFrame stay on the render thread. Theory: the
game thread was unpaced (no-op wait) so UE's internal pipeline desynced and the render thread stalled
~85ms during motion -> dropped frames -> judder. Pacing the game thread should fix it. RISK: wait/
begin must stay 1:1 to the runtime; if a begin is rejected (inFrame) after a wait was pushed, could
desync -> watch for xrWaitFrame/xrBeginFrame xrfail. Keep ffr=-1 + blackcount=0 (both help). Testing.
## *** REFINED: BC (UE render thread) BLOCKS ~85ms at only ~14ms GPU = architectural (2026-06-27 latest) ***
After FFR=-1 (applied ffr=1, GPU fed~13.7ms) AND blackcount/lumagate off (luma readback gone):
ghost PERSISTS. STALL split now: BC (begin_frame->end_frame = UE render) = 47-101ms while GPU does
only ~14ms of work => UE's RENDER THREAD is BLOCKING ~85ms, not computing. A(end->begin)~0.3-1.8ms,
our end-frame phases <25ms. FFR helped (BC was 5ms one run) but BC hitch is non-deterministic and
returns. Native at the SAME ~14ms GPU + same FFR = smooth (0.34% 1-frame drops). So the cause is
NOT GPU load, NOT our per-frame overhead, NOT submit-timing — it's UE's render thread intermittently
stalling under our OpenXR frame-loop. RULED OUT for the stall: xrWaitFrame (<25ms, 1 hit),
lock-wait, flush, xrEndFrame, WaitSwapchainImage (all <25ms). The block is in UE's own
render-recording window (begin_frame return -> EndFrame4 call).
HYPOTHESIS (user's "vrapi lazy / openxr eager"): our frame loop paces the RENDER thread (xrWaitFrame
in ovrp_BeginFrame4) and makes the GAME thread's ovrp_WaitToBeginFrame a NO-OP — opposite of vrapi,
which paces the GAME thread. So the game thread runs unpaced/eager and the render-thread back-pressure
(xrWaitFrame + frames-in-flight + compositor holding our 3 swapchain images during reproj) makes UE's
render thread block intermittently => drops => judder. Comment at xrr_wait_frame says game-thread
pacing was tried and caused "BeginFrame too many times" (XR_ERROR_CALL_ORDER_INVALID) -> they
worked around by render-thread pacing. The REAL fix is likely an architectural frame-loop rework:
pace the GAME thread (like vrapi) with correct xrWaitFrame/Begin/End 1:1:1 ordering across the two
threads. Non-trivial, real risk. Other cheap-ish probes: more swapchain images (UE frames-in-flight
stall if compositor holds our 3); check UE's own RHI frame-pacing/dynamic-res CVars.
## (helped, secondary) full-res render (FFR forced OFF) raised GPU load -> drop judder
Stall localization (per-phase + gap-split probes in xr_runtime.c) showed the 105-250ms stalls are
NOT in any of our blocking calls (xrWaitFrame/lock/flush/xrEndFrame/WaitSwapchainImage all <25ms,
A=end->begin ~0.3-1.4ms) — they're in BC = begin_frame->end_frame = UE's own render. Matched-motion:
native = 0.34% drops (all 1-frame); shim = more drops + 137-264ms stalls. So OUR shim makes UE's
render hitch. WHY: GPU-TIME log shows `fed=14ms gameLevel=1 dynamic=0 -> applied ffr=0` — the GAME
requests foveation (TiledMultiRes level 1, which native honors) but our shim had debug.re4vr.ffr=0
FORCING foveation OFF -> UE renders FULL RES -> GPU pinned at ~14ms (right at the 13.9ms/72Hz budget,
zero headroom) -> any head-motion load spike pushes GPU over budget -> render thread stalls on GPU ->
dropped frame -> compositor timewarp judder = the ghost. Native applies the game's FFR -> headroom ->
smooth. FIX (quality-neutral, matches native): debug.re4vr.ffr=-1 (game-driven) so we apply the
game's requested foveation level. Testing now. If confirmed, make ffr=-1 (game-driven) the default
in code (not 0). FFR maps game TiledMultiRes -> XR_FB_foveation (xr_runtime.c apply_foveation /
foveation_entrypoints, ~L1766+).
## *** VISUALLY CONFIRMED: whole-frame TEMPORAL JUDDER (2026-06-27 late) ***
Pulled the user's on-device recordings (/sdcard/Oculus/VideoShots/*.mp4, 30fps mono). Blending 3
consecutive frames (ImageMagick -evaluate-sequence mean) of the title screen shows the "Resident
Evil" banner + candle flames DOUBLED — two sharp offset copies (diagonal shift), WHOLE frame, not
just close objects. = real temporal judder (two distinct positions), not motion blur, not stereo.
Matches user: "blend frames shows it / not just close objects / mild-violent / random."
Tooling that works: adb pull the VideoShots mp4 (adb screenrecord gives 0 bytes — Quest blocks the
VR surface); ffmpeg extract frames; `compare -metric MAE` to find motion spikes; `convert
-evaluate-sequence mean` to blend & reveal judder; Read the PNG to view it.
Key: frame-pacing shows we DO present ~72fps (not half-rate), yet consecutive frames land at TWO
positions => the predicted-display-time or submitted pose ALTERNATES/jitters frame-to-frame and the
compositor timewarp snaps between spots. Added JUDDER probe: per-frame predictedDisplayTime delta in
xrr_begin_frame (after xrWaitFrame) — steady ~13.9ms = pose source; alternating/jittery = timing.
Mechanism candidate: our frame loop splits ovrp_WaitToBeginFrame(game thread, no-op) from
xrWaitFrame+LocateViews+Begin (render thread, in xrr_begin_frame) — this nonstandard pacing can give
the compositor jittery predicted times => judder. Native (vrapi) is phase-locked => smooth.
## (SUPERSEDED) render/composite SUBMIT-ORDERING race theory
Chain of elimination, all by in-MOTION data (static title was a red herring — geometry matches
statically; ghost only shows in motion):
- FOV: render (CalculateEyeLayerDesc2) == composite (g_xr.views) == native, per-frame, stable.
(Earlier "inner-fov mismatch" was MY arithmetic error: tan(0.785 rad)=1.000, not 0.991.)
- IPD ~0.065, eye orientation == head, head == eye-mid, no LAYER MISMATCH (stage==acquired),
full viewport, 3-image swapchain. ALL geometry/composition intrinsics correct.
- DECISIVE: a 60-frame eye-texture burst dump (debug.re4vr.dump=N) showed EVERY frame CLEAN
(single hand) AND the user saw NO ghost while the dump ran. The dump adds a per-frame GPU
fence-wait (synchronous readback) that stalls the game thread enough that UE's eye render
(on its own RHI-thread queue, NOT our s_queue — qwait was a no-op, confirming separate queue)
lands before we release+xrEndFrame. => the ghost is: WE RELEASE THE SWAPCHAIN / xrEndFrame
BEFORE UE SUBMITS ITS EYE RENDER. OpenXR then syncs the compositor against incomplete/previous
content => flashing per-eye double on fast/close motion. NOT geometry, NOT depth, NOT reproject.
- NEXT: test the built submit-hook (debug.re4vr.submithook=2 = present/release-on-submit; patches
UE's global vkQueueSubmit PFN) — it orders our release AFTER UE's eye submit. Was refuted for
BLACK but the ghost is a different artifact. If it fixes the ghost, refine to minimize latency.
If not, try a completion fence injected at the hooked submit, or device-wait before release.
## Full call census (PTC, native returns) — 41 PT_FWD'd fns, all match our shim's returns
All getters return success/same values in native and shim. Only diff: GetMixedRealityInitialized
native=1 vs ours=0 (MR irrelevant to eye render). GetSystemDisplayFrequency2/PerfMetrics have
pre-init transient failures (-1002/-1008) then succeed — same as ours. CONCLUSION: the game makes
the same calls and gets the same answers in both modes => the ghost is NOT a getter-return diff;
it's in COMPOSITION (our OpenXR layer submit vs native vrapi compositor), which is not a game call.
## *** KEY DIFFERENCE: native submits eye-fov with ReverseZ depth reprojection ***
Native EndFrame4 eye-fov layer flags=**0x4 = ovrpLayerSubmitFlag_ReverseZ** (1<<2). NOT NoDepth(0x8).
=> native composites with DEPTH-BASED POSITIONAL TIMEWARP, reverse-Z convention. And the game DOES
request a depth buffer: our CalculateEyeLayerDesc2 logged depthFormat=10. Depth-aware reprojection
is exactly what corrects close-object parallax under head translation — its absence = "deform/swim
on close objects" = the seated hand ghost. PRIME SUSPECT.
Our shim CAN chain XrCompositionLayerDepthInfoKHR (xr_runtime.c build_composition L896, reverse-z via
g_depthRevZ; depth swapchain created L2012 gated on `depth_wanted()` + game depthFormat). But memory
says "depth on didn't fix it" — so VERIFY whether the game actually RENDERS valid depth into OUR
depth swapchain (UE only renders depth if GetLayerTexture2 returns a depth handle AND its RHI targets
it). If our depth image is empty/garbage, positional timewarp is a no-op (or worse) => ghost persists
even with depth=1. That's the next probe.
## NEXT probe: verify our depth reprojection is functionally live (passthru=0, depth=1)
1. "setup_layer: DEPTH swapchain ..." present? (swapchain created)
2. does GetLayerTexture2 return a depth handle to the game (outDepthTex non-null path)?
3. is depth chained each frame in build_composition (diag==0 && !pipeline && depthSwapchain)?
4. is the depth IMAGE actually written by UE (dump min/max; all-1.0 or all-0 = not rendered)?
If depth never reaches our swapchain -> that's the fix (wire UE's depth -> our depth image), and it
would explain native(ReverseZ)=clean vs ours=ghost.
## (superseded) NEXT: in-game passthru (the actual repro)
Title doesn't exercise the seated hand-deform. Run passthru=1, load save, play SEATED:
(a) does native eliminate the hand-deform? If yes -> ghost is in our submission/render path,
not geometry (since geometry matches). Capture in-game native EndFrame4/pose/fov, diff vs
our in-game STEREO/views/HEADvsEYE logs at the same moment.
(b) if native ALSO deforms -> not our shim's fault (game/headset).