Added support for VR Frame Interpolation

This commit is contained in:
iChris4 committed 2026-09-10 21:33:34 +02:00
1 parent c8eaa52727
commit 6467c6390c
24 files changed
+1028 -130

No files matched your search

+56 -8
View File
@@ -29,7 +29,7 @@ Standalone launches remain opt-in. `Config.toml` is created with the following d
enabled = false
required = false
mirror_view = "normal"
eager_frame_heartbeat = false
frame_interpolation_fps = 0
render_scale = 1.0
world_units_per_meter = 500.0
hud_distance_meters = 2.0
@@ -66,12 +66,27 @@ image is chosen, so the setting can always be changed back.
Eye mirror modes retain the last eye image when a desktop frame has no new XR packet, so they
do not alternate with the normal camera. `"none"` also stays black between XR packets.
`eager_frame_heartbeat` is live in **F10 > VR > Eager Frame Heartbeat** and defaults to `false`.
With it off, the XR thread waits for new game frames and repeats the last valid image during
stalls, with a 50 ms keep-alive interval. With it on, the thread repeats at headset display
deadlines while waiting for new rendering, matching the eager behavior. Compare both settings
during the same race to check smoothness with your OpenXR runtime. Both retain the protection
against black frames during pauses and window dragging.
**F10 > VR > VR frame interpolation (experimental)** offers **Off, Auto, 72, 90, 120** and
applies immediately. `frame_interpolation_fps` stores `0` for Off (the default), `1` for Auto,
or the selected rate. The earlier `frame_interpolation = true` checkbox migrates to Auto.
Auto renders at the headset's display deadlines; the numbered choices cap the rate of new
stereo frames. They do not change the headset's physical refresh setting. For VDXR with Virtual
Desktop set to 90 Hz, select Auto or 90. The menu shows both the detected headset rate and the
rate of newly rendered VR frames, excluding repeated images. Menus and other virtual-screen
scenes continue at the game's rate; assess interpolation during an immersive race.
VR interpolation is independent of **Graphics > Race frame interpolation**. The simulation,
physics, audio and VI remain at 60 Hz. Scene motion is delayed by one game frame (about 16.7 ms)
to interpolate between known transforms; each rendered eye pair uses a fresh predicted head
pose. This needs enough GPU headroom to render both eyes at the target rate, and carries the
desktop interpolator's experimental artifacts, especially for unmatched or changing geometry.
Refresh detection uses `XR_FB_display_refresh_rate` when available and the OpenXR predicted
display period otherwise. Interpolation requires `XR_KHR_win32_convert_performance_counter_time`
to relate those display deadlines to the game's clock; the menu reports if it is unavailable.
The old Eager Frame Heartbeat option has been removed and existing `eager_frame_heartbeat`
settings are ignored. Completed rendering wakes the XR thread immediately. A 50 ms keep-alive
still protects pauses and window dragging without eager repeats during rendering.
`render_scale` scales the per-eye size recommended by the OpenXR runtime.
`world_units_per_meter` controls the scale of headset translation in the game world.
@@ -158,8 +173,16 @@ short-lived immutable stereo packet. Each sealed GX frame and immersive packet c
policy-generation tag; a mismatch is rendered in mono and the acquired XR frame is canceled, so an
asynchronous menu/race transition cannot replay race transforms over unsafe content.
With VR interpolation enabled, Aurora retains each sealed race's command stream and matched
previous/current transform uniforms. New OpenXR packets wake the frame worker between game
frames. It interpolates at the requested display time, then applies that packet's head pose and
the scene anchor to both eyes. Native offscreen effects and the 2D HUD retain their game-frame
updates. A mid-frame EFB readback invalidates retained GPU data; a policy-tag mismatch rejects
the replay. Missing matches use current transforms, and stalls clamp at the last known pose
instead of extrapolating. The ordinary desktop interpolation settings remain independent.
The D3D12 pacing thread retains the last completed projection or virtual-screen layer and
resubmits it during stalls (or missed display deadlines with eager heartbeat enabled), including while moving the desktop window,
resubmits it during stalls, including while moving the desktop window,
pausing, or minimizing. The scene freezes until rendering resumes; the compositor can still
reproject the retained image for head movement. Repeated layers keep their original render poses
and field of view, paired with the new compositor display time. Two pairs of eye swapchains keep
@@ -215,6 +238,31 @@ The Vulkan path intentionally does not create an unrelated Vulkan device or use
a workaround. It accepts a future explicit Dawn native context, including external queue locking,
so it can be enabled once Aurora exposes those handles safely.
### Interpolation validation
The GX tests cover retained transform endpoints with desktop interpolation off and continuous
sampling at 72/90/120 Hz. `mkw_frame_interpolation_pacing_tests` covers fixed-rate scheduling,
live changes, stalls and configuration migration; `mkw_openxr_replay_tests` exercises swapchain
ownership and retained-layer submission without a headset.
For a Windows GPU check, configure Aurora with its tests enabled and
`AURORA_GPU_SMOKE_TESTS=ON`, then build/run `stereo_frame_worker_smoke`. This feeds the actual
renderer a 60 Hz GX stream and an independent 90 Hz stereo provider. The development check
produced 359 new stereo submissions in 4 seconds (89.7 FPS). This verifies submission cadence,
not full-race performance or visual quality on a headset. Pass a draw count, for example
`stereo_frame_worker_smoke 2000`, to stress uniform preparation and renderer/producer overlap;
`stereo_frame_worker_smoke 2000 0` checks native stereo with interpolation Off.
The test checks that the producer stays above 55 FPS as well as checking headset submissions;
replaying an old scene more often must not hide a slowed simulation. Validate actual races in VDXR at 90 Hz with
Auto/90 selected, including race entry/exit, first person, recentering and pauses.
Stereo uniform calculations use cached CPU memory, followed by a single write into the upload
buffer. Reading or modifying matrices directly in D3D12 upload memory can be extremely slow,
especially with many character draws; see Microsoft's [Map guidance](https://learn.microsoft.com/en-us/windows/win32/api/d3d12/nf-d3d12-id3d12resource-map).
Retained interpolation reserves eye ranges at seal time and fills them once at the headset sample
time. VR interpolation also releases the producer after sealing so eye encoding can overlap the
next game frame, as it does with desktop interpolation.
## Current limitations
- Only the project's supported PAL `RMCP01` translation has race instrumentation addresses.
+7
View File
@@ -127,6 +127,9 @@ typedef struct {
// Appended to preserve the frameToken/eyes prefix used by older providers.
AuroraStereoFrameMode mode;
uint64_t contentTag;
// Predicted display time converted to std::chrono::steady_clock nanoseconds.
// Zero disables temporal interpolation for this packet.
uint64_t displayTimeNanos;
} AuroraStereoFrame;
/**
@@ -136,6 +139,10 @@ typedef struct {
*/
typedef bool (*AuroraStereoFrameProvider)(uint32_t logicalFrame, AuroraStereoFrame* frame, void* userdata);
// Wake retained stereo replay after publishing a packet. The provider remains
// non-blocking, and all OpenXR calls stay on the application's pacing thread.
void aurora_notify_stereo_frame();
typedef struct {
const char* appName;
const char* userPath;
+5
View File
@@ -72,6 +72,11 @@ void aurora_get_frame_interpolation_diagnostics(AuroraFrameInterpolationDiagnost
void aurora_set_frame_interpolation_fps(uint32_t targetFps);
uint32_t aurora_get_frame_interpolation_fps();
// Independent from desktop interpolation: replay captured race transforms at
// each headset deadline, leaving guest simulation and VI timing at 60 Hz.
void aurora_set_stereo_frame_interpolation(bool enabled);
bool aurora_get_stereo_frame_interpolation();
// Newly encountered GX pipelines compile on the bounded worker queue. Draws whose pipeline is not
// ready are skipped rather than stalling submission, and pick it up once compilation finishes.
void aurora_set_skip_unready_pipelines(bool enabled);
+124 -6
View File
@@ -9,6 +9,7 @@
#include "imgui.hpp"
#include "stereo.hpp"
#include "stereo_mirror.hpp"
#include "stereo_interpolation.hpp"
#include "webgpu/gpu.hpp"
#include <webgpu/webgpu_cpp.h>
#endif
@@ -272,6 +273,7 @@ struct FrameWorkerState {
bool started = false;
bool stop = false;
bool jobPending = false;
bool stereoPending = false;
// Written with jobPending and copied by the worker under this mutex. They
// belong to that exact queued frame, not to the producer's next frame.
uint64_t contentTag = AURORA_STEREO_CONTENT_TAG_UNKNOWN;
@@ -326,6 +328,7 @@ bool frame_worker_requested() noexcept {
// Returns false when a stop request was observed mid-cycle.
bool run_frame_worker_cycle(gfx::SealedFrame& sealedFrame, uint64_t contentTag,
const StereoSceneAnchor& sceneAnchor) noexcept;
void run_retained_stereo_frame(gfx::SealedFrame& sealedFrame) noexcept;
#endif
void frame_worker_main() noexcept {
@@ -343,12 +346,18 @@ void frame_worker_main() noexcept {
for (;;) {
uint64_t contentTag = AURORA_STEREO_CONTENT_TAG_UNKNOWN;
StereoSceneAnchor sceneAnchor{};
bool stereoOnly = false;
{
std::unique_lock lock(g_frameWorker.mutex);
g_frameWorker.cv.wait(lock, [] { return g_frameWorker.stop || g_frameWorker.jobPending; });
g_frameWorker.cv.wait(
lock, [] { return g_frameWorker.stop || g_frameWorker.jobPending || g_frameWorker.stereoPending; });
if (g_frameWorker.stop) {
break;
}
stereoOnly = !g_frameWorker.jobPending;
g_frameWorker.stereoPending = false;
if (stereoOnly)
g_frameWorker.ready.store(false, std::memory_order_release);
contentTag = g_frameWorker.contentTag;
g_frameWorker.contentTag = AURORA_STEREO_CONTENT_TAG_UNKNOWN;
sceneAnchor = g_frameWorker.sceneAnchor;
@@ -359,6 +368,19 @@ void frame_worker_main() noexcept {
// The CPU already decoded the sealed frame at its GX boundary; the worker only owns
// encode/submit/present, so it never touches the producer's next FIFO buffer.
#ifdef AURORA_ENABLE_GX
if (stereoOnly) {
run_retained_stereo_frame(sealedFrame);
{
std::lock_guard lock(g_frameWorker.mutex);
// The producer can queue its next seal after observing the preceding
// DONE but before this idle replay claims the worker. Do not publish
// that newer job as done before it has actually run.
if (!g_frameWorker.jobPending)
g_frameWorker.ready.store(true, std::memory_order_release);
}
g_frameWorker.cv.notify_all();
continue;
}
if (!run_frame_worker_cycle(sealedFrame, contentTag, sceneAnchor)) {
break;
}
@@ -388,6 +410,7 @@ void ensure_frame_worker_started() noexcept {
}
g_frameWorker.stop = false;
g_frameWorker.jobPending = false;
g_frameWorker.stereoPending = false;
g_frameWorker.contentTag = AURORA_STEREO_CONTENT_TAG_UNKNOWN;
g_frameWorker.sceneAnchor = {};
g_frameWorker.sealed.store(true, std::memory_order_release);
@@ -1662,8 +1685,62 @@ struct SealedFrameContext {
uint32_t logicalFrame = 0;
bool interpolationActive = false;
bool replayInterpolatedFrames = false;
std::optional<AuroraStereoFrame> stereoInput;
bool retainStereo = false;
};
// Worker-owned scene state. A separate buffer generation check protects against
// synchronous EFB submissions overwriting the retained frame's GPU data.
struct RetainedStereoContext {
uint64_t contentTag = AURORA_STEREO_CONTENT_TAG_UNKNOWN;
uint64_t boundary = 0;
uint64_t interval = 0;
uint32_t logicalFrame = 0;
StereoSceneAnchor anchor;
StereoSceneAnchor previousAnchor;
bool continuous = false;
} g_retainedStereo;
gfx::StereoReplayFrame interpolated_stereo_frame(const AuroraStereoFrame& input, float& weight) {
const auto& retained = g_retainedStereo;
weight = retained.continuous
? stereo::interpolation_weight(input.displayTimeNanos, retained.boundary, retained.interval)
: 1.0f;
auto anchor = retained.anchor;
if (weight < 1.0f && anchor.active && retained.previousAnchor.active) {
Mat3x4<float> previous, current, result;
std::memcpy(&previous, retained.previousAnchor.anchorFromScene.data(), sizeof(previous));
std::memcpy(&current, anchor.anchorFromScene.data(), sizeof(current));
if (gx::interpolate_transform(previous, current, weight, result)) {
std::memcpy(anchor.anchorFromScene.data(), &result, sizeof(result));
}
}
return make_stereo_replay_frame(input, anchor);
}
void run_retained_stereo_frame(gfx::SealedFrame& sealedFrame) noexcept {
std::lock_guard gpuLock(g_rendererGpuMutex);
if (!gx::stereo_frame_interpolation_active() || !gfx::has_late_stereo_replay(sealedFrame))
return;
const auto input = request_stereo_frame(g_retainedStereo.logicalFrame, g_retainedStereo.contentTag);
if (!input || input->mode != AURORA_STEREO_FRAME_IMMERSIVE_REPLAY)
return;
float weight;
auto replay = interpolated_stereo_frame(*input, weight);
auto encoder = g_device.CreateCommandEncoder();
if (!gfx::prepare_late_stereo_replay(sealedFrame, encoder, replay, weight))
return;
for (uint32_t eye = 0; eye < AURORA_STEREO_EYE_COUNT; ++eye) {
gfx::render_stereo_eye(sealedFrame, encoder, replay, eye, false);
}
const auto sink = run_stereo_sink(encoder, input->frameToken, g_retainedStereo.logicalFrame, input->mode);
const auto buffer = encoder.Finish();
std::lock_guard submitLock(g_queueSubmitMutex);
g_queue.Submit(1, &buffer);
if (sink && sink->submitted)
sink->submitted(sink->frame, sink->userdata);
}
// Phase 1: everything that touches producer-shared renderer state. Needs g_rendererGpuMutex and
// a FIFO already drained into the recorded pass list.
void seal_frame_locked(gfx::SealedFrame& sealedFrame, SealedFrameContext& ctx, uint64_t contentTag,
@@ -1680,6 +1757,7 @@ void seal_frame_locked(gfx::SealedFrame& sealedFrame, SealedFrameContext& ctx, u
// pre-first-frame UINT32_MAX value to logical frame zero.
ctx.logicalFrame = gfx::current_frame() + 1;
if (const auto stereoInput = request_stereo_frame(ctx.logicalFrame, contentTag)) {
ctx.stereoInput = stereoInput;
ctx.stereoFrameToken = stereoInput->frameToken;
ctx.stereoFrameMode = stereoInput->mode;
ctx.stereoReplay = make_stereo_replay_frame(*stereoInput, sceneAnchor);
@@ -1727,6 +1805,16 @@ void seal_frame_locked(gfx::SealedFrame& sealedFrame, SealedFrameContext& ctx, u
// Detach the recorded passes. From here the producer's list is empty and the
// encode phase reads only worker-private state.
gfx::seal_frame(sealedFrame);
ctx.retainStereo = gx::stereo_frame_interpolation_active() && gfx::has_late_stereo_replay(sealedFrame);
const bool continuous = ctx.retainStereo && g_retainedStereo.contentTag == contentTag &&
g_retainedStereo.interval != 0 && ctx.scheduleIntervalNanos != 0 &&
ctx.scheduleBaseNanos > g_retainedStereo.boundary &&
ctx.scheduleBaseNanos - g_retainedStereo.boundary <= ctx.scheduleIntervalNanos * 3 / 2 &&
g_retainedStereo.anchor.active == sceneAnchor.active;
const auto previousAnchor = g_retainedStereo.anchor;
g_retainedStereo = {contentTag, ctx.scheduleBaseNanos, ctx.scheduleIntervalNanos,
ctx.logicalFrame, sceneAnchor, previousAnchor,
continuous};
gfx::expire_bind_group_cache();
}
@@ -1801,7 +1889,7 @@ std::vector<PresentationJob> encode_sealed_frame(gfx::SealedFrame& sealedFrame,
// A demanded CPU-visible EFB readback submits a prefix of the frame, so replaying the resumed
// stream would mutate an already-rendered EFB. Render once, then duplicate into the slots.
gfx::render(sealedFrame, encoder, -1, !immersiveReplay);
gfx::render(sealedFrame, encoder, -1, !immersiveReplay && !ctx.retainStereo);
// The copy targets now hold this frame's resolves, so queue their readbacks on the same encoder;
// completion is harvested in gfx::after_submit, never waited on here.
gfx::efb_ram::encode_async_downloads(encoder);
@@ -1828,8 +1916,14 @@ std::vector<PresentationJob> encode_sealed_frame(gfx::SealedFrame& sealedFrame,
// previous frame's; the interpolated slots above necessarily mirror the
// previous frame, having been encoded before this replay.
if (immersiveReplay) {
if (ctx.retainStereo && ctx.stereoInput) {
float weight;
ctx.stereoReplay = interpolated_stereo_frame(*ctx.stereoInput, weight);
gfx::prepare_late_stereo_replay(sealedFrame, encoder, *ctx.stereoReplay, weight);
}
for (uint32_t eye = 0; eye < AURORA_STEREO_EYE_COUNT; ++eye) {
gfx::render_stereo_eye(sealedFrame, encoder, *ctx.stereoReplay, eye, eye + 1 == AURORA_STEREO_EYE_COUNT);
gfx::render_stereo_eye(sealedFrame, encoder, *ctx.stereoReplay, eye,
!ctx.retainStereo && eye + 1 == AURORA_STEREO_EYE_COUNT);
}
}
@@ -1978,8 +2072,8 @@ void record_frame_telemetry() {
FrameMarkNamed("Aurora frame");
}
// One complete frame-worker cycle. The scene encode only leaves the renderer mutex when
// interpolation actually inserts slots; otherwise both phases publish together.
// One complete frame-worker cycle. Desktop and headset interpolation both
// release the producer after sealing, before encoding their extra scene views.
bool run_frame_worker_cycle(gfx::SealedFrame& sealedFrame, uint64_t contentTag,
const StereoSceneAnchor& sceneAnchor) noexcept {
ZoneScopedN("Frame worker cycle");
@@ -1990,7 +2084,7 @@ bool run_frame_worker_cycle(gfx::SealedFrame& sealedFrame, uint64_t contentTag,
{
std::lock_guard gpuLock(g_rendererGpuMutex);
seal_frame_locked(sealedFrame, ctx, contentTag, sceneAnchor);
overlapEncode = ctx.interpolationActive;
overlapEncode = ctx.interpolationActive || ctx.retainStereo;
if (!overlapEncode) {
presentationJobs = encode_sealed_frame(sealedFrame, ctx);
}
@@ -2284,6 +2378,30 @@ void aurora_set_frame_worker_wait_callback(AuroraFrameWorkerWaitCallback callbac
void aurora_set_stereo_frame_provider(AuroraStereoFrameProvider provider, void* userdata) {
aurora::set_stereo_frame_provider(provider, userdata);
}
void aurora_notify_stereo_frame() {
#ifdef AURORA_ENABLE_GX
std::lock_guard lock(aurora::g_frameWorker.mutex);
if (aurora::g_frameWorker.started && aurora::gx::stereo_frame_interpolation_active()) {
aurora::g_frameWorker.stereoPending = true;
aurora::g_frameWorker.cv.notify_one();
}
#endif
}
void aurora_set_stereo_frame_interpolation(bool enabled) {
#ifdef AURORA_ENABLE_GX
std::lock_guard lock(aurora::g_frameWorker.mutex);
aurora::gx::detail::g_stereoFrameInterpolation.store(enabled, std::memory_order_release);
if (!enabled)
aurora::g_frameWorker.stereoPending = false;
#endif
}
bool aurora_get_stereo_frame_interpolation() {
#ifdef AURORA_ENABLE_GX
return aurora::gx::stereo_frame_interpolation_active();
#else
return false;
#endif
}
void aurora_wait_for_frame_worker() { aurora::wait_for_frame_worker(); }
bool aurora_wait_for_frame_worker_for(uint32_t timeoutMicros) {
return aurora::wait_for_frame_worker_for(std::chrono::microseconds(timeoutMicros));
+223 -67
View File
@@ -296,8 +296,29 @@ static void recycle_render_passes(std::vector<RenderPass>& passes) noexcept {
passes.clear();
}
struct LateStereoUniform {
gx::UniformReplayLayout layout;
Viewport viewport;
Range current;
Range previous;
std::array<Range, AURORA_STEREO_EYE_COUNT> eyes;
};
struct LateStereoData {
std::vector<LateStereoUniform> uniforms;
std::vector<uint8_t> sources;
std::vector<uint8_t> uploadBytes;
ClipRect displayRegion{};
stereo_replay::HudScreen hudScreen{};
uint64_t generation = 0;
uint32_t uploadOffset = 0;
uint32_t uploadSize = 0;
};
// Advanced for every upload, including synchronous mid-frame EFB readbacks.
static std::atomic_uint64_t g_replayBufferGeneration{0};
static LateStereoData g_pendingLateStereo;
struct SealedFrameData {
std::vector<RenderPass> passes;
LateStereoData stereo;
};
SealedFrame::SealedFrame() : m_data(std::make_unique<SealedFrameData>()) {}
@@ -1274,7 +1295,79 @@ static stereo_replay::HudScreen stereo_hud_screen() noexcept {
};
}
static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame) noexcept {
// Shared by the normal seal and headset-deadline replay. Head transforms are
// composed after scene interpolation, so free look never inherits its delay.
static void write_stereo_uniform(std::span<uint8_t> uniform, const gx::UniformReplayLayout& layout,
const StereoReplayEye& eye, const Mat4x4<float>& gameProjection,
const Viewport& drawViewport, ClipRect displayRegion,
const stereo_replay::HudScreen& hudScreen) noexcept {
if (layout.perspective) {
const auto projection = stereo_replay::compose_projection(eye.projection, gameProjection);
std::memcpy(uniform.data() + layout.projectionOffset, &projection, sizeof(projection));
for (uint32_t matrix = 0; matrix < layout.positionMatrixCount; ++matrix) {
if ((layout.positionMatrixMask & (1u << matrix)) == 0) {
continue;
}
const size_t offset = layout.positionOffset + matrix * sizeof(Mat3x4<float>);
Mat3x4<float> source;
std::memcpy(&source, uniform.data() + offset, sizeof(source));
const auto transformed = stereo_replay::compose_affine(eye.viewFromScene, source);
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
}
for (uint32_t matrix = 0; matrix < layout.normalMatrixCount; ++matrix) {
const size_t offset = layout.normalOffset + matrix * sizeof(Mat3x4<float>);
Mat3x4<float> source;
std::memcpy(&source, uniform.data() + offset, sizeof(source));
const auto transformed = stereo_replay::compose_normal(eye.viewFromScene, source);
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
}
} else {
// 2D content reaches the eye entirely through its projection: the
// draw's own position matrices lay the element out in screen space.
// The screen rectangle is built in the VR-neutral view space, so this
// path uses viewFromCenter, not viewFromScene: folding the anchor in
// would leave the screen behind at the camera the anchor replaced.
// First lift viewport-local NDC into displayed-frame NDC; replay will
// use a full-eye viewport so sub-pane elements are not transformed by
// the recorded viewport a second time.
const auto ndcRemap = stereo_replay::make_hud_ndc_remap(
drawViewport.left, drawViewport.top, drawViewport.width, drawViewport.height,
static_cast<float>(displayRegion.x), static_cast<float>(displayRegion.y),
static_cast<float>(displayRegion.width), static_cast<float>(displayRegion.height));
const auto projection = stereo_replay::compose_hud_screen_projection(eye.projection, eye.viewFromCenter, hudScreen,
gameProjection, ndcRemap);
std::memcpy(uniform.data() + layout.projectionOffset, &projection, sizeof(projection));
}
if (displayRegion.width > 0 && displayRegion.height > 0) {
float renderSize[2];
float logicalSize[2];
std::memcpy(renderSize, uniform.data() + 8, sizeof(renderSize));
std::memcpy(logicalSize, uniform.data() + 16, sizeof(logicalSize));
if (layout.perspective) {
renderSize[0] *= static_cast<float>(eye.target.size.width) / static_cast<float>(displayRegion.width);
renderSize[1] *= static_cast<float>(eye.target.size.height) / static_cast<float>(displayRegion.height);
} else {
// Point/line expansion and GX's pixel-center correction now operate
// in the full eye viewport. Recover the complete logical frame size
// from this draw's logical-to-render scale.
if (renderSize[0] != 0.0f) {
logicalSize[0] *= static_cast<float>(displayRegion.width) / renderSize[0];
}
if (renderSize[1] != 0.0f) {
logicalSize[1] *= static_cast<float>(displayRegion.height) / renderSize[1];
}
renderSize[0] = static_cast<float>(eye.target.size.width);
renderSize[1] = static_cast<float>(eye.target.size.height);
std::memcpy(uniform.data() + 16, logicalSize, sizeof(logicalSize));
}
std::memcpy(uniform.data() + 8, renderSize, sizeof(renderSize));
}
}
static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame,
LateStereoData* history = nullptr) noexcept {
const StereoDisplaySource displaySource = stereo_display_source(g_renderPasses);
const ClipRect displayRegion = displaySource.region;
// This is the producer-side preparation path; eye replay can query the pure
@@ -1345,6 +1438,14 @@ static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame)
return false;
}
if (history != nullptr) {
history->uniforms.reserve(replayPerspectiveDrawCount + replayHudScreenDrawCount);
history->sources.reserve(requiredBytes);
history->displayRegion = displayRegion;
history->hudScreen = hudScreen;
history->uploadOffset =
static_cast<uint32_t>(AURORA_ALIGN(g_uniforms.size(), g_cachedLimits.minUniformBufferOffsetAlignment));
}
Viewport drawViewport{
.left = static_cast<float>(displayRegion.x),
.top = static_cast<float>(displayRegion.y),
@@ -1353,6 +1454,10 @@ static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame)
.znear = 0.0f,
.zfar = 1.0f,
};
// D3D12 upload heaps can be write-combined. Read each source once, and do
// all read/modify/write operations in cached CPU memory before uploading.
std::array<uint8_t, gx::MaxUniformSize> sourceUniform;
std::array<uint8_t, gx::MaxUniformSize> eyeUniform;
for (auto& pass : g_renderPasses) {
if (!pass.efbTarget) {
continue;
@@ -1368,9 +1473,9 @@ static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame)
}
auto& draw = command.data.draw.gx;
const auto& layout = draw.uniformReplayLayout;
std::memcpy(sourceUniform.data(), g_uniforms.data() + draw.uniformRange.offset, draw.uniformRange.size);
Mat4x4<float> gameProjection;
std::memcpy(&gameProjection, g_uniforms.data() + draw.uniformRange.offset + layout.projectionOffset,
sizeof(gameProjection));
std::memcpy(&gameProjection, sourceUniform.data() + layout.projectionOffset, sizeof(gameProjection));
// Only a genuinely affine projection carries its NDC position in its clip
// position, which is what the virtual screen reprojection consumes. GX
// tracks the projection type separately from the matrix, so a 2D draw
@@ -1379,76 +1484,40 @@ static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame)
if (!layout.perspective && !stereo_replay::is_orthographic_projection(gameProjection)) {
continue;
}
LateStereoUniform* saved = nullptr;
if (history != nullptr) {
saved = &history->uniforms.emplace_back();
saved->layout = layout;
saved->viewport = drawViewport;
const auto save = [&](const uint8_t* source, uint32_t size) -> Range {
if (size == 0)
return {};
const Range copy{static_cast<uint32_t>(history->sources.size()), size};
history->sources.insert(history->sources.end(), source, source + size);
return copy;
};
saved->current = save(sourceUniform.data(), draw.uniformRange.size);
saved->previous = save(g_uniforms.data() + draw.previousUniformRange.offset, draw.previousUniformRange.size);
}
for (uint32_t eyeIndex = 0; eyeIndex < AURORA_STEREO_EYE_COUNT; ++eyeIndex) {
auto [uniform, range] = copy_uniform(draw.uniformRange);
auto [uniform, range] = map_uniform(draw.uniformRange.size);
draw.stereoUniformRanges[eyeIndex] = range;
const auto& eye = stereoFrame.eyes[eyeIndex];
if (layout.perspective) {
const auto projection = stereo_replay::compose_projection(eye.projection, gameProjection);
std::memcpy(uniform.data() + layout.projectionOffset, &projection, sizeof(projection));
for (uint32_t matrix = 0; matrix < layout.positionMatrixCount; ++matrix) {
if ((layout.positionMatrixMask & (1u << matrix)) == 0) {
if (saved != nullptr) {
// The late replay uploads these ranges at the actual display time.
// Preparing another eye pair here would immediately be overwritten.
saved->eyes[eyeIndex] = range;
continue;
}
const size_t offset = layout.positionOffset + matrix * sizeof(Mat3x4<float>);
Mat3x4<float> source;
std::memcpy(&source, uniform.data() + offset, sizeof(source));
const auto transformed = stereo_replay::compose_affine(eye.viewFromScene, source);
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
}
for (uint32_t matrix = 0; matrix < layout.normalMatrixCount; ++matrix) {
const size_t offset = layout.normalOffset + matrix * sizeof(Mat3x4<float>);
Mat3x4<float> source;
std::memcpy(&source, uniform.data() + offset, sizeof(source));
const auto transformed = stereo_replay::compose_normal(eye.viewFromScene, source);
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
}
} else {
// 2D content reaches the eye entirely through its projection: the
// draw's own position matrices lay the element out in screen space.
// The screen rectangle is built in the VR-neutral view space, so this
// path uses viewFromCenter, not viewFromScene: folding the anchor in
// would leave the screen behind at the camera the anchor replaced.
// First lift viewport-local NDC into displayed-frame NDC; replay will
// use a full-eye viewport so sub-pane elements are not transformed by
// the recorded viewport a second time.
const auto ndcRemap = stereo_replay::make_hud_ndc_remap(
drawViewport.left, drawViewport.top, drawViewport.width, drawViewport.height,
static_cast<float>(displayRegion.x), static_cast<float>(displayRegion.y),
static_cast<float>(displayRegion.width), static_cast<float>(displayRegion.height));
const auto projection = stereo_replay::compose_hud_screen_projection(
eye.projection, eye.viewFromCenter, hudScreen, gameProjection, ndcRemap);
std::memcpy(uniform.data() + layout.projectionOffset, &projection, sizeof(projection));
}
if (displayRegion.width > 0 && displayRegion.height > 0) {
float renderSize[2];
float logicalSize[2];
std::memcpy(renderSize, uniform.data() + 8, sizeof(renderSize));
std::memcpy(logicalSize, uniform.data() + 16, sizeof(logicalSize));
if (layout.perspective) {
renderSize[0] *= static_cast<float>(eye.target.size.width) / static_cast<float>(displayRegion.width);
renderSize[1] *= static_cast<float>(eye.target.size.height) / static_cast<float>(displayRegion.height);
} else {
// Point/line expansion and GX's pixel-center correction now operate
// in the full eye viewport. Recover the complete logical frame size
// from this draw's logical-to-render scale.
if (renderSize[0] != 0.0f) {
logicalSize[0] *= static_cast<float>(displayRegion.width) / renderSize[0];
}
if (renderSize[1] != 0.0f) {
logicalSize[1] *= static_cast<float>(displayRegion.height) / renderSize[1];
}
renderSize[0] = static_cast<float>(eye.target.size.width);
renderSize[1] = static_cast<float>(eye.target.size.height);
std::memcpy(uniform.data() + 16, logicalSize, sizeof(logicalSize));
}
std::memcpy(uniform.data() + 8, renderSize, sizeof(renderSize));
const auto& eye = stereoFrame.eyes[eyeIndex];
std::memcpy(eyeUniform.data(), sourceUniform.data(), range.size);
write_stereo_uniform({eyeUniform.data(), range.size}, layout, eye, gameProjection, drawViewport, displayRegion,
hudScreen);
std::memcpy(uniform.data(), eyeUniform.data(), range.size);
}
}
}
if (history != nullptr) {
history->uploadSize = static_cast<uint32_t>(g_uniforms.size()) - history->uploadOffset;
}
return true;
}
@@ -1464,7 +1533,19 @@ static bool end_batch_impl(const wgpu::CommandEncoder& cmd, bool advanceFrame,
// interpolation tasks pointing into it would dangle. Tie the clear to the rotation itself.
gx::drop_pending_frame_interpolation_uniforms();
}
const bool stereoPrepared = stereoFrame == nullptr || prepare_stereo_replay_uniforms(*stereoFrame);
g_pendingLateStereo = {};
++g_replayBufferGeneration;
const bool captureStereo =
advanceFrame && gx::stereo_frame_interpolation_active() && gx::frame_interpolation_replay_safe();
const StereoReplayFrame placeholder{};
const bool stereoPrepared = (stereoFrame == nullptr && !captureStereo) ||
prepare_stereo_replay_uniforms(stereoFrame != nullptr ? *stereoFrame : placeholder,
captureStereo ? &g_pendingLateStereo : nullptr);
if (captureStereo && stereoPrepared) {
g_pendingLateStereo.generation = g_replayBufferGeneration.load(std::memory_order_acquire);
} else {
g_pendingLateStereo = {};
}
g_uniforms.append_zeroes(gx::MaxUniformSize); // Pad the end of the buffer
uint64_t bufferOffset = 0;
const auto writeBuffer = [&](ByteBuffer& buf, wgpu::Buffer& out, uint64_t size, std::string_view label) {
@@ -1735,6 +1816,7 @@ void seal_frame(SealedFrame& out) noexcept {
// The encode that could still have been holding these has completed: the
// producer joins the worker's DONE phase before it seals another frame.
g_retiredBindGroups.clear();
out.data().stereo = std::move(g_pendingLateStereo);
auto& passes = out.data().passes;
// The previous cycle already recycled these, so this normally just hands the empty vector, its
// capacity included, back to the producer.
@@ -1752,6 +1834,80 @@ void render(SealedFrame& frame, wgpu::CommandEncoder& cmd, int32_t interpolatedF
});
}
bool has_late_stereo_replay(const SealedFrame& frame) noexcept {
const auto& data = frame.data().stereo;
return data.generation != 0 && data.generation == g_replayBufferGeneration.load(std::memory_order_acquire) &&
!data.uniforms.empty();
}
bool prepare_late_stereo_replay(SealedFrame& frame, wgpu::CommandEncoder& cmd, const StereoReplayFrame& stereoFrame,
float weight) {
if (!has_late_stereo_replay(frame))
return false;
auto& data = frame.data().stereo;
data.uploadBytes.resize(data.uploadSize);
auto* bytes = data.uploadBytes.data();
weight = std::clamp(weight, 0.0f, 1.0f);
for (const auto& saved : data.uniforms) {
// Interpolate once; both eyes share exactly the same scene sample.
std::span<uint8_t> uniform{bytes + saved.eyes[0].offset - data.uploadOffset, saved.current.size};
std::memcpy(uniform.data(), data.sources.data() + saved.current.offset, uniform.size());
const auto& layout = saved.layout;
if (layout.perspective && saved.previous.size == saved.current.size && weight < 1.0f) {
const auto* previous = data.sources.data() + saved.previous.offset;
const auto interpolateMatrices = [&](uint32_t offset, uint32_t count, uint32_t mask) {
for (uint32_t matrix = 0; matrix < count; ++matrix) {
if ((mask & (1u << matrix)) == 0)
continue;
const size_t at = offset + matrix * sizeof(Mat3x4<float>);
Mat3x4<float> before, current, result;
std::memcpy(&before, previous + at, sizeof(before));
std::memcpy(&current, uniform.data() + at, sizeof(current));
const bool valid = layout.indexedMatrices ? gx::interpolate_indexed_transform(before, current, weight, result)
: gx::interpolate_transform(before, current, weight, result);
if (valid)
std::memcpy(uniform.data() + at, &result, sizeof(result));
}
};
interpolateMatrices(layout.positionOffset, layout.positionMatrixCount, layout.positionMatrixMask);
interpolateMatrices(layout.normalOffset, layout.normalMatrixCount, layout.positionMatrixMask);
// Interpolate the game depth mapping before applying the HMD frustum.
for (size_t component = 0; component < 16; ++component) {
const size_t at = layout.projectionOffset + component * sizeof(float);
float before, current;
std::memcpy(&before, previous + at, sizeof(float));
std::memcpy(&current, uniform.data() + at, sizeof(float));
const float value = before + (current - before) * weight;
std::memcpy(uniform.data() + at, &value, sizeof(float));
}
}
Mat4x4<float> projection;
std::memcpy(&projection, uniform.data() + layout.projectionOffset, sizeof(projection));
for (uint32_t eye = 1; eye < AURORA_STEREO_EYE_COUNT; ++eye) {
std::memcpy(bytes + saved.eyes[eye].offset - data.uploadOffset, uniform.data(), uniform.size());
}
for (uint32_t eye = 0; eye < AURORA_STEREO_EYE_COUNT; ++eye) {
write_stereo_uniform({bytes + saved.eyes[eye].offset - data.uploadOffset, saved.current.size}, layout,
stereoFrame.eyes[eye], projection, saved.viewport, data.displayRegion, data.hudScreen);
}
}
// Never interpolate in mapped upload memory: write-combined pages make CPU
// reads expensive even for values that were just written there.
const wgpu::BufferDescriptor descriptor{
.label = "Headset interpolation uniforms",
.usage = wgpu::BufferUsage::CopySrc,
.size = data.uploadSize,
.mappedAtCreation = true,
};
auto upload = g_device.CreateBuffer(&descriptor);
std::memcpy(upload.GetMappedRange(), bytes, data.uploadSize);
upload.Unmap();
// Command-buffer ordering keeps these writes after the preceding eye pair,
// without mutating any producer staging memory or desktop uniforms.
cmd.CopyBufferToBuffer(upload, 0, g_uniformBuffer, data.uploadOffset, data.uploadSize);
return true;
}
void render_stereo_eye(SealedFrame& frame, wgpu::CommandEncoder& cmd, const StereoReplayFrame& stereoFrame,
uint32_t eye, bool finalize) {
CHECK(eye < AURORA_STEREO_EYE_COUNT, "invalid stereo eye {}", eye);
+6
View File
@@ -343,6 +343,12 @@ private:
// called with the renderer GPU mutex held; see SealedFrame.
void seal_frame(SealedFrame& out) noexcept;
// Retained replay owns CPU transform endpoints and reserved eye-uniform ranges.
// Call with the renderer mutex held: a mid-frame EFB submission invalidates it.
bool has_late_stereo_replay(const SealedFrame& frame) noexcept;
bool prepare_late_stereo_replay(SealedFrame& frame, wgpu::CommandEncoder& cmd, const StereoReplayFrame& stereoFrame,
float weight);
// Encode a sealed frame. Never touches the producer-visible recording state,
// so this may run concurrently with the producer's FIFO drains.
void render(SealedFrame& frame, wgpu::CommandEncoder& cmd, int32_t interpolatedFrame = -1, bool finalize = true);
+1
View File
@@ -2313,6 +2313,7 @@ static void handle_draw_unmerged(GXPrimitive prim, GXVtxFmt fmt, u16 vtxCount, g
.uniformRange = uniformRanges.current,
.interpolatedUniformRanges = uniformRanges.interpolated,
.stereoUniformRanges = {},
.previousUniformRange = uniformRanges.previous,
.uniformReplayLayout = uniformRanges.replayLayout,
.vtxCount = vtxCount,
.indexCount = numIndices,
+39 -7
View File
@@ -26,6 +26,7 @@
namespace aurora::gx {
namespace detail {
std::atomic_uint32_t g_frameInterpolationFps{0};
std::atomic_bool g_stereoFrameInterpolation{false};
} // namespace detail
namespace {
@@ -790,9 +791,12 @@ void report_producer_paced(bool paced) noexcept {
void begin_frame_interpolation() noexcept {
const uint32_t targetFps = frame_interpolation_fps();
if (targetFps != s_previousInterpolationFps) {
static bool previousStereo = false;
const bool stereo = stereo_frame_interpolation_active();
if (targetFps != s_previousInterpolationFps || stereo != previousStereo) {
recycle_transform_entries(s_previousFrameTransforms);
s_previousInterpolationFps = targetFps;
previousStereo = stereo;
}
// Latch the slot count for the frame starting here; see s_activeInterpolationSamples
// for why it cannot move again until the seal.
@@ -825,8 +829,11 @@ void begin_frame_interpolation() noexcept {
void finalize_frame_interpolation() noexcept {
// A frame reported late seals without inserted slots, so the encode phase renders
// the native frame only. Its transforms still seed the next frame's matching.
if (s_dropInterpolationAtSeal.exchange(false, std::memory_order_acq_rel)) {
const bool late = s_dropInterpolationAtSeal.exchange(false, std::memory_order_acq_rel);
if (late) {
s_diagLateSealDrops.fetch_add(1, std::memory_order_relaxed);
}
if (late && !stereo_frame_interpolation_active()) {
s_hasInterpolatedFrame.store(false, std::memory_order_release);
s_pendingUniformInterpolations.clear();
retire_frame_transforms();
@@ -1078,7 +1085,7 @@ void finalize_frame_interpolation() noexcept {
[](size_t previousIndex) { return previousIndex != SIZE_MAX; }));
// Interpolation never pauses on match quality: an unmatched draw just renders its
// end-frame state, while a ratio gate flapped the whole output cadence instead.
const bool eligible = frame_interpolation_fps() != 0;
const bool eligible = frame_interpolation_fps() != 0 && !late;
// Overlay observability: the live match ratio, and how often the scene sits in
// low-match territory where inserted slots mostly duplicate draws.
@@ -1127,7 +1134,7 @@ void finalize_frame_interpolation() noexcept {
}
}
if (eligible) {
if (eligible || stereo_frame_interpolation_active()) {
// Prepare each matched pair once: every sample of a draw shares the same
// previous/current matrices. A flat vector keeps the sample tasks parallel.
std::vector<PreparedTransformInterpolation> preparedTransforms(s_currentFrameTransforms.size());
@@ -1445,9 +1452,14 @@ void extend_interpolation_draw(uint16_t usedPnMtxMask) noexcept {
snapshot.usedMatrixMask |= addedSlots;
}
std::array<gfx::Range, MaxInterpolatedFrames> record_interpolation_draw(
const FrameInterpolationDrawIdentity& identity, const Mat4x4<float>& projection,
uint16_t usedPnMtxMask, const InterpolatedUniformLayout& uniformLayout) noexcept {
std::array<gfx::Range, MaxInterpolatedFrames> record_interpolation_draw(const FrameInterpolationDrawIdentity& identity,
const Mat4x4<float>& projection,
uint16_t usedPnMtxMask,
const InterpolatedUniformLayout& uniformLayout,
gfx::Range* previousUniform) noexcept {
if (previousUniform != nullptr) {
*previousUniform = {};
}
FrameTransformSnapshot snapshot{
.projection = projection,
.usedMatrixMask = usedPnMtxMask,
@@ -1517,6 +1529,26 @@ std::array<gfx::Range, MaxInterpolatedFrames> record_interpolation_draw(
});
interpolatedRanges[sample] = interpolatedRange;
}
// Keep the matched previous endpoint, in the current palette's layout, for
// arbitrary headset display times. It shares desktop matching and cut guards.
if (previousUniform != nullptr && stereo_frame_interpolation_active()) {
auto [buffer, range] = gfx::map_uniform(uniformLayout.uniformSize);
std::memcpy(buffer.data(), uniformLayout.sourceUniformData, uniformLayout.uniformSize);
s_pendingUniformInterpolations.push_back({
.currentTransformIndex = currentTransformIndex,
.sourceUniformData = uniformLayout.sourceUniformData,
.uniformData = buffer.data(),
.uniformSize = uniformLayout.uniformSize,
.projectionOffset = uniformLayout.projectionOffset,
.positionOffset = uniformLayout.positionOffset,
.normalOffset = uniformLayout.normalOffset,
.currentMatrix = uniformLayout.currentMatrix,
.numerator = 0,
.denominator = 1,
.indexedMatrices = uniformLayout.indexedMatrices,
});
*previousUniform = range;
}
}
return interpolatedRanges;
}
+12 -4
View File
@@ -39,13 +39,19 @@ namespace detail {
// Defined in frame_interpolation.cpp, exposed so the early-outs below stay inline: the shipped
// build compiles shards without LTO, so a cross-TU call would land on every draw in the frame.
extern std::atomic_uint32_t g_frameInterpolationFps;
extern std::atomic_bool g_stereoFrameInterpolation;
} // namespace detail
// 0 when interpolation is disabled; otherwise the configured target (120/180/240).
inline uint32_t frame_interpolation_fps() noexcept {
return detail::g_frameInterpolationFps.load(std::memory_order_acquire);
}
inline bool frame_interpolation_active() noexcept { return frame_interpolation_fps() != 0; }
inline bool stereo_frame_interpolation_active() noexcept {
return detail::g_stereoFrameInterpolation.load(std::memory_order_acquire);
}
inline bool frame_interpolation_active() noexcept {
return frame_interpolation_fps() != 0 || stereo_frame_interpolation_active();
}
void set_frame_interpolation_fps(uint32_t targetFps) noexcept;
@@ -67,9 +73,11 @@ bool frame_interpolation_replay_safe() noexcept;
// Records one perspective draw and maps its intermediate uniform copies, returning the mapped
// range per slot (empty when the draw has no counterpart). Called by build_uniform.
std::array<gfx::Range, MaxInterpolatedFrames> record_interpolation_draw(
const FrameInterpolationDrawIdentity& identity, const Mat4x4<float>& projection,
uint16_t usedPnMtxMask, const InterpolatedUniformLayout& uniformLayout) noexcept;
std::array<gfx::Range, MaxInterpolatedFrames> record_interpolation_draw(const FrameInterpolationDrawIdentity& identity,
const Mat4x4<float>& projection,
uint16_t usedPnMtxMask,
const InterpolatedUniformLayout& uniformLayout,
gfx::Range* previousUniform = nullptr) noexcept;
// Folds a merged draw back into the snapshot of the draw it joined. aurora renders merged
// primitives through the first one's uniform block, so without this the merged-in bones tear.
+1
View File
@@ -14,6 +14,7 @@ struct DrawData {
gfx::Range uniformRange;
std::array<gfx::Range, MaxInterpolatedFrames> interpolatedUniformRanges;
std::array<gfx::Range, AURORA_STEREO_EYE_COUNT> stereoUniformRanges;
gfx::Range previousUniformRange;
UniformReplayLayout uniformReplayLayout;
uint32_t vtxCount;
uint32_t indexCount;
+6 -2
View File
@@ -767,10 +767,11 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
.positionMatrixCount = layout.postexCount,
.normalMatrixCount = layout.nrmCount,
.perspective = perspective,
.indexedMatrices = info.indexAttr.test(GX_VA_PNMTXIDX),
.nativeEfbEffect = !perspective && samplesRecentEfbCopy && (samplesReducedEfbCopy || blends),
};
if (!perspective || frame_interpolation_fps() == 0) {
if (!perspective || !frame_interpolation_active()) {
g_gxState.stateDirty = false;
return {
.current = range,
@@ -779,6 +780,7 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
};
}
gfx::Range previousUniform{};
const auto interpolatedRanges = record_interpolation_draw(
drawIdentity, effectiveProj, usedPnMtxMask,
InterpolatedUniformLayout{
@@ -790,11 +792,13 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
// A compacted position region holds the current matrix at slot 0.
.currentMatrix = layout.absolutePosRegion ? std::min<size_t>(g_gxState.currentPnMtx, MaxPnMtx - 1) : 0,
.indexedMatrices = info.indexAttr.test(GX_VA_PNMTXIDX),
});
},
&previousUniform);
g_gxState.stateDirty = false;
return {
.current = range,
.interpolated = interpolatedRanges,
.previous = previousUniform,
.replayLayout = replayLayout,
};
}
+2
View File
@@ -13,6 +13,7 @@ struct UniformReplayLayout {
uint8_t positionMatrixCount = 0;
uint8_t normalMatrixCount = 0;
bool perspective = false;
bool indexedMatrices = false;
// A 2D draw compositing the framebuffer back over itself: bloom, blur and the
// rest of the native post-processing chain. It belongs to the rendered image,
// not to the game's 2D layer, so it must stay where the game aimed it.
@@ -22,6 +23,7 @@ struct UniformReplayLayout {
struct UniformRanges {
gfx::Range current;
std::array<gfx::Range, MaxInterpolatedFrames> interpolated;
gfx::Range previous;
UniformReplayLayout replayLayout;
};
+17
View File
@@ -0,0 +1,17 @@
#pragma once
#include <algorithm>
#include <cstdint>
namespace aurora::stereo {
// The current scene belongs to the VI presentation boundary. Delaying scene
// motion by one guest interval lets every display sample fall between known
// endpoints; head tracking is applied afterwards at its predicted display time.
inline float interpolation_weight(uint64_t displayTime, uint64_t boundary, uint64_t interval) noexcept {
if (displayTime == 0 || boundary == 0 || interval == 0)
return 1.0f;
if (displayTime <= boundary)
return 0.0f;
return static_cast<float>(std::min(static_cast<double>(displayTime - boundary) / static_cast<double>(interval), 1.0));
}
} // namespace aurora::stereo
+9
View File
@@ -1,6 +1,14 @@
include(FetchContent)
include(GoogleTest)
option(AURORA_GPU_SMOKE_TESTS "Build opt-in tests requiring a desktop GPU" OFF)
if (AURORA_GPU_SMOKE_TESTS AND AURORA_ENABLE_GX AND WIN32)
add_executable(stereo_frame_worker_smoke stereo_frame_worker_smoke.cpp)
target_include_directories(stereo_frame_worker_smoke PRIVATE ../lib)
target_link_libraries(stereo_frame_worker_smoke PRIVATE aurora::core aurora::gx aurora::main aurora::vi
dawn::dawncpp_headers)
endif ()
if (NOT TARGET gtest)
FetchContent_Declare(googletest
URL https://github.com/google/googletest/archive/refs/tags/v1.17.0.tar.gz
@@ -18,6 +26,7 @@ if (AURORA_ENABLE_GX)
gx_fifo_test.cpp
gx_test_stubs.cpp
stereo_replay_test.cpp
stereo_interpolation_test.cpp
stereo_mirror_test.cpp
texture_bind_group_cache_key_test.cpp
../lib/gfx/efb_ram_encoder.cpp
+48 -1
View File
@@ -200,6 +200,48 @@ static u32 read_be32_at(const std::vector<u8>& bytes, size_t offset) {
(static_cast<u32>(bytes[offset + 2]) << 8) | static_cast<u32>(bytes[offset + 3]);
}
TEST_F(GXFifoTest, VrKeepsMatchedEndpointsWithDesktopInterpolationOff) {
struct Reset {
~Reset() {
aurora::gx::detail::g_stereoFrameInterpolation.store(false);
aurora::gx::set_frame_interpolation_fps(0);
aurora::gx::begin_frame_interpolation();
}
} reset;
aurora::gx::set_frame_interpolation_fps(0);
aurora::gx::detail::g_stereoFrameInterpolation.store(true);
const auto info = aurora::gx::build_shader_info({});
gxState().currentPnMtx = 0;
gxState().pnMtx[0].pos = {{1, 0, 0, 0}, {0, 1, 0, 0}, {0, 0, 1, -50}};
gxState().pnMtx[0].nrm = {{1, 0, 0, 0}, {0, 1, 0, 0}, {0, 0, 1, 0}};
const auto build = [&](float x, aurora::HashType identity, bool split = false) {
aurora::gx::begin_frame_interpolation();
aurora::gfx::testing::reset_uniform_allocations();
gxState().pnMtx[0].pos.m0[3] = x;
if (split)
aurora::gx::mark_frame_interpolation_replay_unsafe();
const auto result = aurora::gx::build_uniform(info, 0, {}, {identity, identity, 7}, true);
aurora::gx::finalize_frame_interpolation();
return result;
};
EXPECT_EQ(build(10, 100).previous.size, 0u); // Warm-up.
auto uniforms = build(20, 100);
ASSERT_NE(uniforms.previous.size, 0u);
const auto readX = [&](aurora::gfx::Range range) {
const auto& bytes = aurora::gfx::testing::uniform_allocation(range.offset);
float x;
std::memcpy(&x, bytes.data() + uniforms.replayLayout.positionOffset + 3 * sizeof(float), sizeof(x));
return x;
};
EXPECT_FLOAT_EQ(readX(uniforms.previous), 10);
EXPECT_FLOAT_EQ(readX(uniforms.current), 20);
EXPECT_EQ(aurora::gx::interpolated_frame_count(), 0u); // Desktop remains off.
EXPECT_TRUE(std::all_of(uniforms.interpolated.begin(), uniforms.interpolated.end(),
[](auto range) { return range.size == 0; }));
EXPECT_EQ(build(30, 200).previous.size, 0u); // Unmatched draws use current transforms.
EXPECT_EQ(build(40, 200, true).previous.size, 0u); // Readback split invalidates replay.
}
TEST(FrameInterpolationContract, RequiresStablePerspectiveDrawSequence) {
const auto resetInterpolation = [] {
aurora::gx::set_frame_interpolation_fps(0);
@@ -473,8 +515,13 @@ TEST(FrameInterpolationContract, IndexedPaletteHistoryKeepsAbsoluteVertexSlots)
std::array<uint8_t, uniformSize> changedSource{};
aurora::gx::begin_frame_interpolation();
const auto changedRanges = recordFrame(changedTopology, 91.0f, 9.0f, changedSource);
EXPECT_EQ(changedRanges[0].size, 0u);
aurora::gx::finalize_frame_interpolation();
// Palette borrowing may reserve a range before matching is resolved. The
// safety contract is that a changed topology never reuses the old transforms.
if (changedRanges[0].size != 0) {
const auto& unchanged = aurora::gfx::testing::uniform_allocation(changedRanges[0].offset);
EXPECT_EQ(std::memcmp(unchanged.data(), changedSource.data(), uniformSize), 0);
}
aurora::gx::set_frame_interpolation_fps(0);
aurora::gx::begin_frame_interpolation();
@@ -0,0 +1,162 @@
// Optional GPU smoke test: a 60 Hz GX producer with an independent 90 Hz
// compositor. Exercises the real frame worker, eye replay and submission sink.
#include <aurora/aurora.h>
#include <aurora/gfx.h>
#include <aurora/main.h>
#include <dolphin/gx.h>
#include <dolphin/mtx.h>
#include "stereo.hpp"
#include <algorithm>
#include <atomic>
#include <chrono>
#include <condition_variable>
#include <cstdio>
#include <cstdlib>
#include <filesystem>
#include <mutex>
#include <thread>
using Clock = std::chrono::steady_clock;
static std::mutex packetMutex;
static std::condition_variable packetCv;
static AuroraStereoFrame packet{};
static bool available = false;
static bool stop = false;
static uint64_t completed = 0;
static std::atomic_uint32_t submitted{0};
static bool Provide(uint32_t, AuroraStereoFrame* output, void*) {
std::lock_guard lock(packetMutex);
if (!available)
return false;
*output = packet;
available = false;
return true;
}
static bool Encode(wgpu::CommandEncoder&, const aurora::stereo::SinkFrame&, void*) noexcept { return true; }
static void Submitted(const aurora::stereo::SinkFrame& frame, void*) noexcept {
std::lock_guard lock(packetMutex);
completed = frame.frameToken;
++submitted;
packetCv.notify_all();
}
static void Log(AuroraLogLevel level, const char* module, const char* message, unsigned int length) {
if (level >= LOG_WARNING)
std::fprintf(stderr, "%s: %.*s\n", module, static_cast<int>(length), message);
}
int main(int argc, char** argv) {
// Extra distinct draws expose CPU uniform/replay costs that a single triangle
// cannot exercise. Keep the eye targets small to isolate that regression.
const unsigned drawCount = argc > 1 ? std::max(1, std::atoi(argv[1])) : 1;
const bool interpolate = argc < 3 || std::atoi(argv[2]) != 0;
std::filesystem::create_directories("stereo-smoke-cache");
AuroraConfig config{};
config.appName = "Aurora VR interpolation smoke";
config.userPath = ".";
config.cachePath = "stereo-smoke-cache";
config.desiredBackend = BACKEND_D3D12;
config.windowWidth = 160;
config.windowHeight = 120;
config.hasWindowPosition = true;
config.windowPosX = -30000;
config.windowPosY = -30000;
config.logCallback = Log;
config.logLevel = LOG_WARNING;
config.xrInterop = true;
aurora_initialize(argc, argv, &config);
aurora_set_frame_interpolation_fps(0);
aurora_set_stereo_frame_interpolation(interpolate);
aurora_set_stereo_frame_provider(Provide, nullptr);
aurora::stereo::set_sink(Encode, Submitted, nullptr);
std::thread compositor([] {
const auto start = Clock::now();
for (uint64_t token = 1;; ++token) {
const auto deadline = start + std::chrono::nanoseconds(token * 1'000'000'000 / 90);
std::this_thread::sleep_until(deadline);
{
std::lock_guard lock(packetMutex);
if (stop)
return;
packet = {};
packet.frameToken = token;
packet.contentTag = 42;
packet.displayTimeNanos =
std::chrono::duration_cast<std::chrono::nanoseconds>(deadline.time_since_epoch()).count();
for (auto& eye : packet.eyes) {
eye.width = 160;
eye.height = 120;
eye.projection[0] = eye.projection[5] = 1;
eye.projection[10] = -1;
eye.projection[11] = -1;
eye.projection[14] = -1;
eye.viewFromCenter[0] = eye.viewFromCenter[5] = eye.viewFromCenter[10] = 1;
}
available = true;
}
aurora_notify_stereo_frame();
std::unique_lock lock(packetMutex);
packetCv.wait(lock, [&] { return stop || completed == token; });
if (stop)
return;
}
});
const auto start = Clock::now();
for (uint64_t frame = 0; frame < 240; ++frame) {
const auto boundary = start + std::chrono::nanoseconds((frame + 1) * 1'000'000'000 / 60);
std::this_thread::sleep_until(boundary);
aurora_update();
if (!aurora_begin_frame())
continue;
aurora_set_present_schedule(
std::chrono::duration_cast<std::chrono::nanoseconds>(boundary.time_since_epoch()).count(), 16'666'667);
Mtx44 projection{{1, 0, 0, 0}, {0, 1, 0, 0}, {0, 0, -1, -1}, {0, 0, -1, 0}};
Mtx transform{{1, 0, 0, static_cast<float>(frame % 60) * 0.01f}, {0, 1, 0, 0}, {0, 0, 1, -3}};
GXSetProjection(projection, GX_PERSPECTIVE);
GXSetCurrentMtx(GX_PNMTX0);
GXSetViewport(0, 0, 160, 120, 0, 1);
GXSetScissor(0, 0, 160, 120);
GXClearVtxDesc();
GXSetVtxDesc(GX_VA_POS, GX_DIRECT);
GXSetVtxAttrFmt(GX_VTXFMT0, GX_VA_POS, GX_POS_XYZ, GX_F32, 0);
GXSetNumTexGens(0);
GXSetNumChans(0);
GXSetNumTevStages(1);
GXSetTevOrder(GX_TEVSTAGE0, GX_TEXCOORD_NULL, GX_TEXMAP_NULL, GX_COLOR_NULL);
GXSetTevOp(GX_TEVSTAGE0, GX_PASSCLR);
for (unsigned draw = 0; draw < drawCount; ++draw) {
transform[1][3] = static_cast<float>(draw % 20) * 0.01f;
GXLoadPosMtxImm(transform, GX_PNMTX0);
GXBegin(GX_TRIANGLES, GX_VTXFMT0, 3);
GXPosition3f32(-1 + static_cast<float>(draw) * 0.0001f, -1, 0);
GXPosition3f32(1, -1, 0);
GXPosition3f32(0, 1, 0);
GXEnd();
}
aurora_end_frame_tagged(42);
if (drawCount > 1 && (frame + 1) % 60 == 0) {
std::printf("Completed %llu producer frames in %.2f s\n", static_cast<unsigned long long>(frame + 1),
std::chrono::duration<double>(Clock::now() - start).count());
std::fflush(stdout);
}
}
{
std::lock_guard lock(packetMutex);
stop = true;
packetCv.notify_all();
}
compositor.join();
aurora_set_stereo_frame_interpolation(false);
aurora_quiesce_frame_worker();
aurora_set_stereo_frame_provider(nullptr, nullptr);
aurora::stereo::set_sink(nullptr, nullptr);
const double elapsed = std::chrono::duration<double>(Clock::now() - start).count();
const double fps = submitted.load() / elapsed;
std::printf("%u draws: producer %.1f FPS; %u stereo submissions in %.2f s (%.1f FPS)\n", drawCount, 240 / elapsed,
submitted.load(), elapsed, fps);
aurora_shutdown();
// More headset submissions must not come at the expense of simulation speed.
return 240 / elapsed > 55 && fps > (interpolate ? 85 : 55) && fps < (interpolate ? 100 : 65) ? 0 : 1;
}
@@ -0,0 +1,32 @@
#include "stereo_interpolation.hpp"
#include <gtest/gtest.h>
#include <cmath>
TEST(StereoInterpolation, ContinuousMotionAcross60HzScenesAtHeadsetRates) {
constexpr uint64_t interval = 16'666'667;
// A camera/object moving one unit per guest frame must advance uniformly,
// even at 72/90 Hz where many samples are neither midpoints nor endpoints.
for (uint64_t hz : {72u, 90u, 120u}) {
double previousPosition = -1;
for (uint64_t sample = 1; sample <= hz; ++sample) {
const uint64_t displayTime = 1'000'000'000 + sample * 1'000'000'000 / hz;
const uint64_t scene = (displayTime - 1'000'000'000) / interval;
const uint64_t boundary = 1'000'000'000 + scene * interval;
const float weight = aurora::stereo::interpolation_weight(displayTime, boundary, interval);
const double position = static_cast<double>(scene) + weight;
if (sample > 1)
EXPECT_NEAR(position - previousPosition, 1'000'000'000.0 / hz / interval, 1e-5);
previousPosition = position;
}
}
}
TEST(StereoInterpolation, MissingTimingAndStallsDoNotExtrapolate) {
using aurora::stereo::interpolation_weight;
EXPECT_FLOAT_EQ(interpolation_weight(0, 100, 10), 1);
EXPECT_FLOAT_EQ(interpolation_weight(105, 0, 10), 1);
EXPECT_FLOAT_EQ(interpolation_weight(105, 100, 0), 1);
EXPECT_FLOAT_EQ(interpolation_weight(95, 100, 10), 0);
EXPECT_FLOAT_EQ(interpolation_weight(105, 100, 10), 0.5);
EXPECT_FLOAT_EQ(interpolation_weight(500, 100, 10), 1);
}
+5
View File
@@ -333,6 +333,11 @@ set_target_properties(mkw_platform PROPERTIES UNITY_BUILD OFF)
# Keep these independent from Aurora's BUILD_TESTING option: they validate the
# project's host-platform contracts, not Aurora's third-party test suite.
enable_testing()
add_executable(mkw_frame_interpolation_pacing_tests tests/frame_interpolation_pacing_tests.cpp)
target_include_directories(mkw_frame_interpolation_pacing_tests PRIVATE "${CMAKE_CURRENT_LIST_DIR}/include")
target_link_libraries(mkw_frame_interpolation_pacing_tests PRIVATE mkw::toml11)
target_compile_features(mkw_frame_interpolation_pacing_tests PRIVATE cxx_std_20)
add_test(NAME mkw_frame_interpolation_pacing_tests COMMAND mkw_frame_interpolation_pacing_tests)
add_executable(mkw_platform_paths_tests "${CMAKE_CURRENT_LIST_DIR}/tests/platform_paths_tests.cpp")
target_link_libraries(mkw_platform_paths_tests PRIVATE mkw_platform)
target_compile_features(mkw_platform_paths_tests PRIVATE cxx_std_17)
+16 -9
View File
@@ -20,6 +20,7 @@
#include <vector>
#include <toml.hpp>
#include "platform/host_platform.h"
#include "vr/frame_interpolation_pacing.h"
#ifdef _WIN32
#ifndef WIN32_LEAN_AND_MEAN
#define WIN32_LEAN_AND_MEAN
@@ -57,7 +58,7 @@ struct RuntimeUserConfig {
std::optional<bool> vrStopAtDisplayCopy;
std::optional<bool> vrSkipCopyClears;
std::optional<std::string> vrMirrorView;
std::optional<bool> vrEagerFrameHeartbeat;
std::optional<uint32_t> vrFrameInterpolationFps;
std::optional<bool> vrFirstPerson;
std::optional<float> vrFirstPersonUnitsPerMeter;
std::optional<float> vrFirstPersonHeadUpMeters;
@@ -380,8 +381,8 @@ inline void EnsureConfigFile() {
"# \"right\" mirror the headset's eyes, and \"none\" blacks the window\n"
"# out. Changeable live from the F10 menu.\n"
"mirror_view = \"normal\"\n"
"# Repeat at headset cadence (true), or only during stalls (false). Live.\n"
"eager_frame_heartbeat = false\n"
"# VR interpolation: 0 = Off, 1 = Auto, or 72/90/120 FPS. Live.\n"
"frame_interpolation_fps = 0\n"
"render_scale = 1.0\n"
"world_units_per_meter = 500.0\n"
"hud_distance_meters = 2.0\n"
@@ -642,7 +643,12 @@ inline RuntimeUserConfig ParseConfigDocument(const toml::value& document) {
value && IsSupportedVrMirrorView(*value)) {
config.vrMirrorView = *value;
}
config.vrEagerFrameHeartbeat = FindConfigValue<bool>(document, "vr", "eager_frame_heartbeat");
config.vrFrameInterpolationFps = FindConfigValue<uint32_t>(document, "vr", "frame_interpolation_fps");
if (!config.vrFrameInterpolationFps) {
// Migrate the initial experimental checkbox to Auto.
if (const auto legacy = FindConfigValue<bool>(document, "vr", "frame_interpolation"))
config.vrFrameInterpolationFps = *legacy ? 1u : 0u;
}
if (auto value = FindConfigInt(document, "vr", "first_person_hidden_model");
value && *value >= -1 && *value <= 31) {
config.vrFirstPersonHiddenModel = static_cast<int32_t>(*value);
@@ -944,9 +950,10 @@ inline bool SetVrMirrorView(std::string value) {
return WriteSetting("vr", "mirror_view", FormatString(value));
}
inline bool SetVrEagerFrameHeartbeat(bool value) {
Mutable().vrEagerFrameHeartbeat = value;
return WriteSetting("vr", "eager_frame_heartbeat", value ? "true" : "false");
inline bool SetVrFrameInterpolationFps(uint32_t value) {
value = mkw::vr::NormalizeFrameInterpolationFps(value);
Mutable().vrFrameInterpolationFps = value;
return WriteSetting("vr", "frame_interpolation_fps", std::to_string(value));
}
inline bool SetVrFirstPersonRotation(std::string value) {
@@ -1281,8 +1288,8 @@ inline std::string VrMirrorView(std::string fallback = kVrMirrorViewDefault) {
return value && IsSupportedVrMirrorView(*value) ? *value : std::move(fallback);
}
inline bool VrEagerFrameHeartbeat() {
return Get().vrEagerFrameHeartbeat.value_or(false);
inline uint32_t VrFrameInterpolationFps() {
return mkw::vr::NormalizeFrameInterpolationFps(Get().vrFrameInterpolationFps.value_or(0));
}
inline std::string VrFirstPersonRotation(std::string fallback = kVrFirstPersonRotationDefault) {
@@ -0,0 +1,41 @@
// SPDX-License-Identifier: GPL-3.0-or-later
#pragma once
#include <algorithm>
#include <cstdint>
namespace mkw::vr {
inline uint32_t NormalizeFrameInterpolationFps(uint32_t value) noexcept {
return value == 1 || value == 72 || value == 90 || value == 120 ? value : 0;
}
// A render-rate ceiling on the compositor's own display-time grid. Auto (1)
// renders every tick. Fixed targets cannot increase the physical refresh rate.
class FrameInterpolationPacing {
public:
bool ShouldRender(int64_t display_time, uint32_t target) noexcept {
if (target != target_ || display_time <= last_time_) {
next_time_ = 0;
target_ = target;
}
last_time_ = display_time;
if (target == 0 || target == 1) {
next_time_ = 0;
return true;
}
if (next_time_ != 0 && display_time + 1'000 < next_time_) return false;
const int64_t interval = 1'000'000'000 / target;
if (next_time_ == 0 || display_time - next_time_ > interval * 2) {
next_time_ = display_time + interval;
} else {
next_time_ += interval;
}
return true;
}
void Reset() noexcept { *this = {}; }
private:
int64_t next_time_ = 0;
int64_t last_time_ = 0;
uint32_t target_ = 0;
};
} // namespace mkw::vr
+9 -3
View File
@@ -53,8 +53,14 @@ void OpenXRRequestRecenter() noexcept;
// once per published frame.
void OpenXRSetLeanBackDegrees(float degrees) noexcept;
// Live pacing choice. Disabled: submit fresh frames promptly and repeat during
// stalls. Enabled: repeat at each headset display deadline while waiting.
void OpenXRSetEagerFrameHeartbeat(bool enabled) noexcept;
// Live scene interpolation at the headset's own display deadlines.
// 0 = Off, 1 = Auto, otherwise 72/90/120 as a rendering-rate ceiling.
void OpenXRSetFrameInterpolationFps(uint32_t target) noexcept;
bool OpenXRFrameInterpolationAvailable() noexcept;
struct OpenXRFrameTiming {
float headset_hz = 0;
float rendered_fps = 0; // Newly rendered pairs; excludes retained-layer repeats.
};
OpenXRFrameTiming OpenXRGetFrameTiming() noexcept;
} // namespace mkw::vr
+23 -7
View File
@@ -117,7 +117,13 @@ float g_vrFirstPersonHeadUp = RuntimeConfigFile::VrFirstPersonHeadUpMeters();
float g_vrFirstPersonHeadForward = RuntimeConfigFile::VrFirstPersonHeadForwardMeters();
float g_vrFirstPersonHeadRight = RuntimeConfigFile::VrFirstPersonHeadRightMeters();
bool g_vrFirstPersonHideDriver = RuntimeConfigFile::VrFirstPersonHideDriver();
bool g_vrEagerFrameHeartbeat = RuntimeConfigFile::VrEagerFrameHeartbeat();
constexpr std::array<uint32_t, 5> kVrInterpolationFps{0, 1, 72, 90, 120};
constexpr std::array<const char*, 5> kVrInterpolationLabels{"Off", "Auto", "72", "90", "120"};
int g_vrFrameInterpolationMode = [] {
const auto value = RuntimeConfigFile::VrFrameInterpolationFps();
return static_cast<int>(std::find(kVrInterpolationFps.begin(), kVrInterpolationFps.end(), value) -
kVrInterpolationFps.begin());
}();
int g_vrFirstPersonHiddenModel = RuntimeConfigFile::VrFirstPersonHiddenModel();
// Config spellings and menu labels for the desktop mirror, index-matched to
// AuroraStereoMirrorView so the combo selection converts to either directly.
@@ -931,15 +937,25 @@ void DrawVrSettings() {
"same desktop image, so the eye choices only differ from Normal during a race.");
}
if (ImGui::Checkbox("Eager Frame Heartbeat", &g_vrEagerFrameHeartbeat)) {
mkw::vr::OpenXRSetEagerFrameHeartbeat(g_vrEagerFrameHeartbeat);
RuntimeConfigFile::SetVrEagerFrameHeartbeat(g_vrEagerFrameHeartbeat);
if (ImGui::Combo("VR frame interpolation (experimental)", &g_vrFrameInterpolationMode,
kVrInterpolationLabels.data(), static_cast<int>(kVrInterpolationLabels.size()))) {
const auto target = kVrInterpolationFps[static_cast<size_t>(g_vrFrameInterpolationMode)];
mkw::vr::OpenXRSetFrameInterpolationFps(target);
RuntimeConfigFile::SetVrFrameInterpolationFps(target);
}
if (ImGui::IsItemHovered()) {
ImGui::SetTooltip(
"On: repeat the last frame at headset refresh rate while waiting for the game. "
"Off (default): follow the game's frame rate, with repeats during stalls. "
"Compare both during a race to check headset smoothness. Applies immediately.");
"Auto matches the headset refresh rate. 72, 90 and 120 cap the scene rendering rate; "
"set the headset's refresh rate in Virtual Desktop or your VR runtime. "
"The game stays at 60 Hz. Adds one game frame of scene latency; head tracking stays current. "
"Needs GPU headroom and may show interpolation artifacts. Applies immediately.");
}
const auto xrTiming = mkw::vr::OpenXRGetFrameTiming();
if (mkw::vr::OpenXRIsRunning()) {
ImGui::TextDisabled("Headset: %.1f Hz | New VR frames: %.1f FPS", xrTiming.headset_hz, xrTiming.rendered_fps);
if (g_vrFrameInterpolationMode != 0 && !mkw::vr::OpenXRFrameInterpolationAvailable()) {
ImGui::TextWrapped("The OpenXR runtime does not provide the clock conversion needed for interpolation.");
}
}
// Like the mirror above and unlike the enable toggle, these two apply to the
+131 -16
View File
@@ -11,6 +11,7 @@
#include "vr/mkw_vr_first_person.h"
#include "vr/mkw_vr_policy.h"
#include "vr/mkw_vr_instrumentation.h"
#include <aurora/gfx.h>
#include <algorithm>
#include <array>
@@ -267,11 +268,21 @@ public:
config.engine_name = "Aurora";
config.resolution_scale = RuntimeConfigFile::VrRenderScale(1.0f);
config.required_extensions = {"XR_KHR_D3D12_enable"};
config.optional_extensions = {"XR_KHR_win32_convert_performance_counter_time", "XR_FB_display_refresh_rate"};
if (!runtime_->Initialize(config)) {
SetError("OpenXR instance initialization failed: " + runtime_->LastError().message);
ResetPreparedObjects();
return OpenXRStartupResult::Unavailable;
}
const auto& extensions = runtime_->EnabledExtensions();
if (std::find(extensions.begin(), extensions.end(),
"XR_KHR_win32_convert_performance_counter_time") != extensions.end()) {
runtime_->LoadFunction("xrConvertTimeToWin32PerformanceCounterKHR", &convert_display_time_);
}
if (std::find(extensions.begin(), extensions.end(), "XR_FB_display_refresh_rate") != extensions.end()) {
runtime_->LoadFunction("xrGetDisplayRefreshRateFB", &get_display_refresh_rate_);
}
interpolation_available_.store(convert_display_time_ != nullptr, std::memory_order_release);
if (!backend_->QueryGraphicsRequirements(*runtime_)) {
SetError(backend_->LastError());
ResetPreparedObjects();
@@ -304,6 +315,10 @@ public:
}
stop_.store(false, std::memory_order_release);
{
std::lock_guard lock(interpolation_mutex_);
interpolation_stopping_ = false;
}
teardown_requested_.store(false, std::memory_order_release);
WithdrawPublishedFrame();
aurora_set_stereo_frame_provider(&OpenXRIntegration::ProvideStereoFrame, this);
@@ -325,6 +340,12 @@ public:
void Shutdown() noexcept {
teardown_requested_.store(false, std::memory_order_release);
// Stop idle replays before draining; no new worker job may race provider removal.
{
std::lock_guard lock(interpolation_mutex_);
interpolation_stopping_ = true;
aurora_set_stereo_frame_interpolation(false);
}
if (pacing_thread_.joinable()) {
// Registration changes are only safe while no sealed frame is in
// flight. The caller invokes us before Aurora teardown.
@@ -354,6 +375,11 @@ public:
backend_.reset();
runtime_.reset();
prepared_ = false;
convert_display_time_ = nullptr;
get_display_refresh_rate_ = nullptr;
headset_hz_.store(0, std::memory_order_relaxed);
rendered_fps_.store(0, std::memory_order_relaxed);
interpolation_available_.store(false, std::memory_order_release);
ResetTrackingOrigin();
applied_session_run_serial_ = 0;
session_was_active_ = false;
@@ -365,8 +391,16 @@ public:
recenter_requested_.store(true, std::memory_order_release);
}
void SetEagerFrameHeartbeat(bool enabled) noexcept {
eager_frame_heartbeat_.store(enabled, std::memory_order_relaxed);
void SetFrameInterpolationFps(uint32_t target) noexcept {
frame_interpolation_fps_.store(NormalizeFrameInterpolationFps(target), std::memory_order_relaxed);
}
OpenXRFrameTiming FrameTiming() const noexcept {
return {headset_hz_.load(std::memory_order_relaxed), rendered_fps_.load(std::memory_order_relaxed)};
}
bool FrameInterpolationAvailable() const noexcept {
return interpolation_available_.load(std::memory_order_acquire);
}
void SetLeanBackDegrees(float degrees) noexcept {
@@ -463,6 +497,9 @@ private:
break;
}
if (!session_active) {
SetInterpolationActive(false);
interpolation_pacing_.Reset();
rendered_fps_.store(0, std::memory_order_relaxed);
WaitForStopOrDelay(std::chrono::milliseconds(5));
continue;
}
@@ -495,6 +532,11 @@ private:
presentation.quad_distance_meters = policy.config.hud_distance_meters;
presentation.quad_width_meters = policy.config.hud_width_meters;
// Updating this on the owner thread also confines retained replay to
// validated race content. The provider checks policy tags again.
const uint32_t interpolation_target = frame_interpolation_fps_.load(std::memory_order_relaxed);
SetInterpolationActive(immersive && FrameInterpolationAvailable() && interpolation_target != 0);
OpenXRD3D12Frame frame{};
const OpenXRD3D12BeginStatus begin = backend_->BeginFrame(presentation, frame);
if (begin == OpenXRD3D12BeginStatus::SessionNotRunning) {
@@ -511,6 +553,7 @@ private:
break;
}
UpdateFrameTiming(frame.xr_frame);
// Both of these read this frame's located head pose and must run
// before FinishFrame submits a layer built from it.
ServiceRecenterRequest();
@@ -524,6 +567,15 @@ private:
continue;
}
if (aurora_get_stereo_frame_interpolation() &&
!interpolation_pacing_.ShouldRender(frame.xr_frame.predicted_display_time, interpolation_target)) {
if (!backend_->TryCancelPendingFrame(frame) || !backend_->FinishFrame(frame, false)) {
SetError(backend_->LastError());
fatal = true;
}
continue;
}
{
std::lock_guard lock(published_mutex_);
// First person renders at life-size scale, third person at the
@@ -534,6 +586,7 @@ private:
policy.content_tag);
published_.store(&published_frame_, std::memory_order_release);
}
aurora_notify_stereo_frame();
OpenXRD3D12SubmissionStatus submission = OpenXRD3D12SubmissionStatus::Timeout;
bool canceled_before_encode = false;
@@ -541,15 +594,9 @@ private:
std::chrono::steady_clock::now() + std::chrono::milliseconds(50);
while (!stop_.load(std::memory_order_acquire) &&
submission == OpenXRD3D12SubmissionStatus::Timeout) {
// Eager mode fills missed headset slots. With it off, completion
// wakes us immediately at the game's cadence, while a 50 ms
// keep-alive still protects pauses and window drags from black.
// Read each iteration so the toggle also works during a stall.
const auto wait_ms = eager_frame_heartbeat_.load(std::memory_order_relaxed)
? static_cast<uint32_t>(std::clamp<XrDuration>(
frame.xr_frame.predicted_display_period / 1'000'000 - 4, 0, 12))
: 50u;
submission = backend_->WaitForSubmission(frame, wait_ms);
// Fresh rendering wakes us immediately. A 50 ms keep-alive
// protects stalls without issuing eager repeats during GPU work.
submission = backend_->WaitForSubmission(frame, 50);
if (submission == OpenXRD3D12SubmissionStatus::Timeout) {
// A pause, minimized window, or guest stall may leave no GX
// frame to consume this packet. Withdraw it, then cancel the
@@ -592,7 +639,10 @@ private:
} else if (!submit) {
SetError("Aurora's D3D12 stereo copy failed; continuing on the desktop mirror");
fatal = true;
} else if (immersive && !immersive_submission_logged) {
} else {
++timing_submissions_;
}
if (submit && !fatal && immersive && !immersive_submission_logged) {
immersive_submission_logged = true;
RT_LOG(RT_TAG_RUNTIME)
<< "[mkw-vr] first immersive packet consumed and submitted as "
@@ -601,6 +651,7 @@ private:
}
}
SetInterpolationActive(false);
running_.store(false, std::memory_order_release);
MkwVRPolicySetSessionActive(false);
if (!stop_.load(std::memory_order_acquire)) {
@@ -621,6 +672,7 @@ private:
destination = {};
destination.frameToken = source.xr_frame.serial;
destination.contentTag = content_tag;
destination.displayTimeNanos = DisplayTimeNanos(source.xr_frame.predicted_display_time);
destination.mode = immersive ? AURORA_STEREO_FRAME_IMMERSIVE_REPLAY
: AURORA_STEREO_FRAME_VIRTUAL_SCREEN;
for (uint32_t eye = 0; eye < kOpenXREyeCount; ++eye) {
@@ -701,6 +753,42 @@ private:
virtual_screen_pose_valid_ = false;
}
void SetInterpolationActive(bool active) noexcept {
std::lock_guard lock(interpolation_mutex_);
aurora_set_stereo_frame_interpolation(active && !interpolation_stopping_);
}
void UpdateFrameTiming(const OpenXRFrame& frame) noexcept {
float hz = 0;
if (get_display_refresh_rate_ == nullptr ||
XR_FAILED(get_display_refresh_rate_(runtime_->Session(), &hz)) || !(hz > 0)) {
if (frame.predicted_display_period > 0)
hz = static_cast<float>(1.0e9 / static_cast<double>(frame.predicted_display_period));
}
headset_hz_.store(hz, std::memory_order_relaxed);
const auto now = std::chrono::steady_clock::now();
const float elapsed = std::chrono::duration<float>(now - timing_start_).count();
if (elapsed >= 1.0f) {
rendered_fps_.store(static_cast<float>(timing_submissions_) / elapsed, std::memory_order_relaxed);
timing_start_ = now;
timing_submissions_ = 0;
}
}
uint64_t DisplayTimeNanos(XrTime display_time) noexcept {
if (convert_display_time_ == nullptr) return 0;
LARGE_INTEGER display_counter{}, counter{}, frequency{};
if (XR_FAILED(convert_display_time_(runtime_->Instance(), display_time, &display_counter)) ||
!QueryPerformanceFrequency(&frequency) || frequency.QuadPart <= 0 ||
!QueryPerformanceCounter(&counter)) return 0;
const auto now = std::chrono::steady_clock::now();
const auto now_ns = std::chrono::duration_cast<std::chrono::nanoseconds>(now.time_since_epoch()).count();
const auto delta = static_cast<int64_t>(
(static_cast<double>(display_counter.QuadPart) - static_cast<double>(counter.QuadPart)) *
1.0e9 / static_cast<double>(frequency.QuadPart));
return now_ns + delta > 0 ? static_cast<uint64_t>(now_ns + delta) : 0;
}
void WaitForStopOrDelay(std::chrono::milliseconds delay) {
std::unique_lock lock(stop_mutex_);
stop_cv_.wait_for(lock, delay,
@@ -729,7 +817,18 @@ private:
std::atomic_bool teardown_requested_{false};
std::atomic_bool recenter_requested_{false};
std::atomic<float> lean_back_degrees_{RuntimeConfigFile::VrLeanBackDegrees()};
std::atomic_bool eager_frame_heartbeat_{RuntimeConfigFile::VrEagerFrameHeartbeat()};
std::atomic_uint32_t frame_interpolation_fps_{RuntimeConfigFile::VrFrameInterpolationFps()};
std::atomic_bool interpolation_available_{false};
std::mutex interpolation_mutex_;
bool interpolation_stopping_ = true;
FrameInterpolationPacing interpolation_pacing_;
std::atomic<float> headset_hz_{0};
std::atomic<float> rendered_fps_{0};
std::chrono::steady_clock::time_point timing_start_ = std::chrono::steady_clock::now();
uint32_t timing_submissions_ = 0;
PFN_xrGetDisplayRefreshRateFB get_display_refresh_rate_ = nullptr;
using ConvertDisplayTime = XrResult (XRAPI_PTR*)(XrInstance, XrTime, LARGE_INTEGER*);
ConvertDisplayTime convert_display_time_ = nullptr;
std::atomic<PublishedFrame*> published_{nullptr};
PublishedFrame published_frame_{};
std::mutex published_mutex_;
@@ -818,11 +917,27 @@ void OpenXRSetLeanBackDegrees(float degrees) noexcept {
#endif
}
void OpenXRSetEagerFrameHeartbeat(bool enabled) noexcept {
void OpenXRSetFrameInterpolationFps(uint32_t target) noexcept {
#if defined(MKW_ENABLE_OPENXR) && defined(_WIN32)
OpenXRIntegration::Get().SetEagerFrameHeartbeat(enabled);
OpenXRIntegration::Get().SetFrameInterpolationFps(target);
#else
(void)enabled;
(void)target;
#endif
}
OpenXRFrameTiming OpenXRGetFrameTiming() noexcept {
#if defined(MKW_ENABLE_OPENXR) && defined(_WIN32)
return OpenXRIntegration::Get().FrameTiming();
#else
return {};
#endif
}
bool OpenXRFrameInterpolationAvailable() noexcept {
#if defined(MKW_ENABLE_OPENXR) && defined(_WIN32)
return OpenXRIntegration::Get().FrameInterpolationAvailable();
#else
return false;
#endif
}
@@ -0,0 +1,53 @@
// SPDX-License-Identifier: GPL-3.0-or-later
#include "vr/frame_interpolation_pacing.h"
#include "runtime_config.h"
#include <cmath>
#include <cstdlib>
#include <iostream>
using mkw::vr::FrameInterpolationPacing;
static void Require(bool condition) {
if (!condition) std::abort();
}
int main() {
for (uint32_t target : {0u, 1u, 72u, 90u, 120u}) {
std::istringstream input("[vr]\nframe_interpolation_fps = " + std::to_string(target) + "\n");
const auto config = RuntimeConfigFile::ParseConfig(input);
Require(config.vrFrameInterpolationFps == target);
}
std::istringstream legacy("[vr]\nframe_interpolation = true\neager_frame_heartbeat = true\n");
Require(RuntimeConfigFile::ParseConfig(legacy).vrFrameInterpolationFps == 1);
std::istringstream explicitOff("[vr]\nframe_interpolation_fps = 0\nframe_interpolation = true\n");
Require(RuntimeConfigFile::ParseConfig(explicitOff).vrFrameInterpolationFps == 0);
std::istringstream missing("[vr]\n");
Require(!RuntimeConfigFile::ParseConfig(missing).vrFrameInterpolationFps.has_value());
for (uint32_t headset : {72u, 90u, 120u}) {
for (uint32_t target : {1u, 72u, 90u, 120u}) {
FrameInterpolationPacing pacing;
uint32_t rendered = 0;
// Ten seconds on the actual display grid, including noninteger ratios.
for (uint32_t frame = 0; frame < headset * 10; ++frame) {
const auto time = 1'000'000'000ll + static_cast<int64_t>(frame) * 1'000'000'000ll / headset;
if (pacing.ShouldRender(time, target)) ++rendered;
}
const auto expected = std::min(headset, target == 1 ? headset : target) * 10;
Require(std::abs(static_cast<int>(rendered) - static_cast<int>(expected)) <= 1);
}
}
FrameInterpolationPacing pacing;
Require(pacing.ShouldRender(1'000'000'000, 72));
Require(!pacing.ShouldRender(1'008'333'333, 72));
Require(pacing.ShouldRender(1'008'333'334, 120)); // Live target change.
Require(pacing.ShouldRender(5'000'000'000, 120)); // Stall: no catch-up burst.
Require(!pacing.ShouldRender(5'000'000'001, 120));
Require(pacing.ShouldRender(1'000'000, 120)); // New session clock.
pacing.Reset();
Require(pacing.ShouldRender(1'000'001, 120));
Require(mkw::vr::NormalizeFrameInterpolationFps(90) == 90);
Require(mkw::vr::NormalizeFrameInterpolationFps(1) == 1);
Require(mkw::vr::NormalizeFrameInterpolationFps(60) == 0);
Require(mkw::vr::NormalizeFrameInterpolationFps(UINT32_MAX) == 0);
std::cout << "VR interpolation pacing tests passed\n";
}