mirror of
https://github.com/mitch030504/Wiicompiled_VR_Frame.git
synced 2026-10-06 01:00:14 +02:00
Added 2D Virtual Screen for HUD and Ortho Elements
This commit is contained in:
1 parent
60316f0115
commit
7898a76a22
24 files changed
+1392
-809
No files matched your search
@@ -25,6 +25,8 @@ required = false
|
||||
render_scale = 1.0
|
||||
world_units_per_meter = 500.0
|
||||
hud_distance_meters = 2.0
|
||||
hud_width_meters = 2.4
|
||||
hud_virtual_screen = true
|
||||
stop_at_display_copy = true
|
||||
skip_copy_clears = true
|
||||
```
|
||||
@@ -39,7 +41,10 @@ or graphics-binding failure is logged and the game continues in ordinary desktop
|
||||
|
||||
`render_scale` scales the per-eye size recommended by the OpenXR runtime.
|
||||
`world_units_per_meter` controls the scale of headset translation in the game world.
|
||||
`hud_distance_meters` controls the distance of the head-locked virtual screen.
|
||||
`hud_distance_meters` and `hud_width_meters` place and size the virtual screen. They are read at
|
||||
launch and govern both the menu screen and the in-race 2D screen, so 2D content keeps its place
|
||||
across the transition. `hud_virtual_screen` decides whether the race's 2D layer uses that screen;
|
||||
it is live and can be flipped from the F10 settings bar.
|
||||
`stop_at_display_copy` ends eye replay at the final `GXCopyDisp`, matching the frame shown on the
|
||||
desktop. `skip_copy_clears` independently suppresses the EFB reset performed after a copy. Both
|
||||
default on and can be changed live from the F10 settings bar for diagnostics.
|
||||
@@ -56,8 +61,8 @@ The runtime deliberately fails safe instead of guessing which Mario Kart camera
|
||||
virtual screen. Session/runtime loss safely tears down XR and continues on the desktop mirror.
|
||||
|
||||
Aurora records the original GX frame once and replays it for both OpenXR eyes. Perspective GX draws
|
||||
receive asymmetric headset projections; orthographic and unclassified draws retain their original
|
||||
GX transforms during immersive replay. Menus and unsafe whole scenes use the virtual-screen path.
|
||||
receive asymmetric headset projections, while the game's 2D layer goes on a fixed virtual screen
|
||||
(see below). Menus and unsafe whole scenes use the virtual-screen path.
|
||||
Head pose is sampled by the OpenXR pacing thread, while Aurora's frame worker consumes a
|
||||
short-lived immutable stereo packet. Each sealed GX frame and immersive packet carry the same
|
||||
policy-generation tag; a mismatch is rendered in mono and the acquired XR frame is canceled, so an
|
||||
@@ -66,6 +71,35 @@ minimized window can withdraw an unencoded packet and end that compositor frame
|
||||
encoded-work stall requests teardown at Aurora's next safe producer boundary. All OpenXR session
|
||||
and swapchain calls remain on their owning thread.
|
||||
|
||||
## The race's 2D layer
|
||||
|
||||
The minimap, race position, item roulette, lap times and the rest of the game's orthographic layer
|
||||
would otherwise be stretched across each eye's entire field of view. With `hud_virtual_screen` on
|
||||
they are instead placed on a rectangle fixed in the recorded camera's own frame, `hud_distance_meters`
|
||||
ahead of it and `hud_width_meters` across, its height following the aspect ratio the game is
|
||||
presenting at. The screen stays where the camera puts it, so looking around moves the view across it
|
||||
rather than dragging it along.
|
||||
|
||||
An orthographic GX projection is affine, so the draw's clip position is already its position on the
|
||||
flat frame. Replay folds three further steps into that same projection matrix, one per eye: the
|
||||
draw viewport into full-frame coordinates, the frame position onto the screen rectangle, and the
|
||||
screen through that eye's view and OpenXR frustum. The draw's own position matrices are left alone.
|
||||
|
||||
Depth uses the equivalent of DolphinXR's Exact Screen Depth path. A replay-only shader variant
|
||||
carries the draw's original GX depth through a flat-interpolated value and explicitly writes it at
|
||||
the fragment, including the draw's recorded viewport depth range. The reprojected geometry itself
|
||||
is parked at mid-depth for clipping. This avoids the view-dependent perspective-divide rounding
|
||||
that otherwise breaks equal-depth `LEQUAL` ordering and causes overlapping menu/HUD elements to
|
||||
z-fight.
|
||||
|
||||
Two classes of draw are deliberately left on their recorded transforms: native framebuffer effects
|
||||
(bloom and the rest of the post-processing chain, recognised by sampling a freshly produced,
|
||||
reduced or blended-back EFB copy), which belong to the rendered image rather than to the game's 2D
|
||||
layer, and any draw whose matrix is not actually affine. Retained one-shot EFB bakes such as Mario
|
||||
Kart Wii's minimap are treated as game art and remain eligible for the screen. A reprojected 2D draw
|
||||
uses the full eye viewport and scissor because its recorded rectangle no longer describes where it
|
||||
ended up; its original viewport is folded into the projection instead.
|
||||
|
||||
## Backend status
|
||||
|
||||
| Backend | Status |
|
||||
|
||||
@@ -95,6 +95,17 @@ bool aurora_get_stereo_stop_at_display_copy();
|
||||
void aurora_set_stereo_skip_copy_clears(bool enabled);
|
||||
bool aurora_get_stereo_skip_copy_clears();
|
||||
|
||||
// Places orthographic GX draws (menus, HUD, 2D overlays) on a fixed virtual
|
||||
// screen during immersive replay instead of stretching them across the whole
|
||||
// eye viewport. The screen hangs `distance` world units straight ahead of the
|
||||
// game camera and is `width` world units across, its height following the
|
||||
// aspect ratio the game is presenting at. It stays put in the camera's frame,
|
||||
// so looking around moves the view across it rather than dragging it along.
|
||||
// Also live; a cleared flag or a non-positive size leaves 2D content on its
|
||||
// recorded GX transforms.
|
||||
void aurora_set_stereo_hud_screen(bool enabled, float width, float distance);
|
||||
bool aurora_get_stereo_hud_screen_enabled();
|
||||
|
||||
// Guest-RAM write tracking. `generation` changes whenever guest RAM covering a host range was
|
||||
// written (or returns AURORA_GUEST_WRITE_UNTRACKED); `notify` reports writes aurora made itself.
|
||||
#define AURORA_GUEST_WRITE_UNTRACKED UINT64_MAX
|
||||
|
||||
+123
-172
@@ -82,6 +82,7 @@ struct StereoSinkRegistration {
|
||||
#endif
|
||||
std::mutex g_stereoRegistrationMutex;
|
||||
StereoProviderRegistration g_stereoProvider;
|
||||
std::atomic_bool g_stereoProviderActive{false};
|
||||
#ifdef AURORA_ENABLE_GX
|
||||
StereoSinkRegistration g_stereoSink;
|
||||
#endif
|
||||
@@ -129,8 +130,7 @@ void wait_until_precise(PresentClock::time_point deadline) noexcept {
|
||||
static thread_local HighResolutionTimer timer;
|
||||
if (timer.handle != nullptr && timerDeadline > now) {
|
||||
const auto remaining100ns =
|
||||
std::chrono::duration_cast<std::chrono::duration<int64_t, std::ratio<1, 10000000>>>(
|
||||
timerDeadline - now);
|
||||
std::chrono::duration_cast<std::chrono::duration<int64_t, std::ratio<1, 10000000>>>(timerDeadline - now);
|
||||
LARGE_INTEGER due{};
|
||||
due.QuadPart = -std::max<int64_t>(remaining100ns.count(), 1);
|
||||
if (::SetWaitableTimerEx(timer.handle, &due, 0, nullptr, nullptr, nullptr, 0) != FALSE) {
|
||||
@@ -151,14 +151,9 @@ void wait_until_precise(PresentClock::time_point deadline) noexcept {
|
||||
}
|
||||
}
|
||||
|
||||
void record_successful_present(bool, uint32_t,
|
||||
std::chrono::nanoseconds,
|
||||
std::chrono::nanoseconds,
|
||||
std::chrono::nanoseconds,
|
||||
std::chrono::nanoseconds,
|
||||
std::chrono::nanoseconds,
|
||||
std::chrono::nanoseconds, bool duplicated,
|
||||
uint32_t, std::chrono::nanoseconds) noexcept {
|
||||
void record_successful_present(bool, uint32_t, std::chrono::nanoseconds, std::chrono::nanoseconds,
|
||||
std::chrono::nanoseconds, std::chrono::nanoseconds, std::chrono::nanoseconds,
|
||||
std::chrono::nanoseconds, bool duplicated, uint32_t, std::chrono::nanoseconds) noexcept {
|
||||
const auto now = PresentClock::now();
|
||||
std::chrono::nanoseconds interval{};
|
||||
{
|
||||
@@ -191,8 +186,7 @@ AuroraPresentTiming snapshot_present_timing() noexcept {
|
||||
for (size_t i = 0; i < g_presentTimingSampleCount; ++i) {
|
||||
const auto& sample = g_presentTimingSamples[i];
|
||||
if (sample.presentedAt >= cutoff && sample.interval.count() > 0) {
|
||||
milliseconds.push_back(
|
||||
std::chrono::duration<double, std::milli>(sample.interval).count());
|
||||
milliseconds.push_back(std::chrono::duration<double, std::milli>(sample.interval).count());
|
||||
if (!sample.duplicated) {
|
||||
++newMotionSamples;
|
||||
}
|
||||
@@ -217,8 +211,7 @@ AuroraPresentTiming snapshot_present_timing() noexcept {
|
||||
// Duplicated presentation slots keep the presented cadence but carry no new
|
||||
// motion; scale them out so this reads as the rate the eye actually sees.
|
||||
result.effectiveFramesPerSecond =
|
||||
result.framesPerSecond *
|
||||
(static_cast<double>(newMotionSamples) / static_cast<double>(milliseconds.size()));
|
||||
result.framesPerSecond * (static_cast<double>(newMotionSamples) / static_cast<double>(milliseconds.size()));
|
||||
for (const double value : milliseconds) {
|
||||
const double difference = value - result.averageFrameTimeMs;
|
||||
result.jitterMs += difference * difference;
|
||||
@@ -389,8 +382,7 @@ bool wait_for_frame_worker_private_for(FrameWorkerPhase phase, std::chrono::micr
|
||||
|
||||
void wait_for_frame_worker_private(FrameWorkerPhase phase) noexcept {
|
||||
constexpr auto kWaitServiceInterval = std::chrono::milliseconds(1);
|
||||
while (!wait_for_frame_worker_private_for(phase, kWaitServiceInterval)) {
|
||||
}
|
||||
while (!wait_for_frame_worker_private_for(phase, kWaitServiceInterval)) {}
|
||||
}
|
||||
|
||||
bool wait_for_frame_worker_private_for(FrameWorkerPhase phase, std::chrono::microseconds timeout) noexcept {
|
||||
@@ -408,8 +400,7 @@ bool wait_for_frame_worker_private_for(FrameWorkerPhase phase, std::chrono::micr
|
||||
return frame_worker_phase_reached(phase);
|
||||
}
|
||||
|
||||
const bool reached =
|
||||
g_frameWorker.cv.wait_for(lock, timeout, [phase] { return frame_worker_phase_reached(phase); });
|
||||
const bool reached = g_frameWorker.cv.wait_for(lock, timeout, [phase] { return frame_worker_phase_reached(phase); });
|
||||
if (reached) {
|
||||
return true;
|
||||
}
|
||||
@@ -444,9 +435,7 @@ void stop_frame_worker() noexcept {
|
||||
g_frameWorker.contentTag = AURORA_STEREO_CONTENT_TAG_UNKNOWN;
|
||||
}
|
||||
|
||||
uint32_t align_to(uint32_t value, uint32_t alignment) noexcept {
|
||||
return (value + alignment - 1) & ~(alignment - 1);
|
||||
}
|
||||
uint32_t align_to(uint32_t value, uint32_t alignment) noexcept { return (value + alignment - 1) & ~(alignment - 1); }
|
||||
|
||||
void append_u16(std::ofstream& out, uint16_t value) {
|
||||
const char bytes[] = {static_cast<char>(value), static_cast<char>(value >> 8)};
|
||||
@@ -454,15 +443,16 @@ void append_u16(std::ofstream& out, uint16_t value) {
|
||||
}
|
||||
|
||||
void append_u32(std::ofstream& out, uint32_t value) {
|
||||
const char bytes[] = {static_cast<char>(value), static_cast<char>(value >> 8),
|
||||
static_cast<char>(value >> 16), static_cast<char>(value >> 24)};
|
||||
const char bytes[] = {static_cast<char>(value), static_cast<char>(value >> 8), static_cast<char>(value >> 16),
|
||||
static_cast<char>(value >> 24)};
|
||||
out.write(bytes, sizeof(bytes));
|
||||
}
|
||||
|
||||
bool write_bmp(const char* path, const uint8_t* pixels, uint32_t width, uint32_t height,
|
||||
uint32_t bytesPerRow, bool bgra) {
|
||||
bool write_bmp(const char* path, const uint8_t* pixels, uint32_t width, uint32_t height, uint32_t bytesPerRow,
|
||||
bool bgra) {
|
||||
std::ofstream out(path, std::ios::binary | std::ios::trunc);
|
||||
if (!out) return false;
|
||||
if (!out)
|
||||
return false;
|
||||
constexpr uint32_t pixelOffset = 14 + 40;
|
||||
const uint32_t imageSize = width * height * 4;
|
||||
out.write("BM", 2);
|
||||
@@ -523,9 +513,7 @@ struct StereoEyeTarget {
|
||||
wgpu::TextureFormat colorFormat = wgpu::TextureFormat::Undefined;
|
||||
wgpu::TextureFormat depthFormat = wgpu::TextureFormat::Undefined;
|
||||
|
||||
const webgpu::TextureWithSampler& output() const noexcept {
|
||||
return resolvedColor.texture ? resolvedColor : color;
|
||||
}
|
||||
const webgpu::TextureWithSampler& output() const noexcept { return resolvedColor.texture ? resolvedColor : color; }
|
||||
};
|
||||
std::array<StereoEyeTarget, AURORA_STEREO_EYE_COUNT> g_stereoEyeTargets;
|
||||
|
||||
@@ -541,8 +529,7 @@ void ensure_stereo_eye_target(uint32_t eyeIndex, uint32_t width, uint32_t height
|
||||
target = {};
|
||||
target.color = webgpu::create_render_texture(width, height, samples > 1);
|
||||
if (samples > 1) {
|
||||
target.resolvedColor =
|
||||
webgpu::create_render_texture(target.color.size.width, target.color.size.height, false);
|
||||
target.resolvedColor = webgpu::create_render_texture(target.color.size.width, target.color.size.height, false);
|
||||
}
|
||||
|
||||
const wgpu::TextureDescriptor depthDescriptor{
|
||||
@@ -565,8 +552,7 @@ void ensure_stereo_eye_target(uint32_t eyeIndex, uint32_t width, uint32_t height
|
||||
target.depthFormat = target.depth.format;
|
||||
}
|
||||
|
||||
std::optional<AuroraStereoFrame> request_stereo_frame(uint32_t logicalFrame,
|
||||
uint64_t contentTag) noexcept {
|
||||
std::optional<AuroraStereoFrame> request_stereo_frame(uint32_t logicalFrame, uint64_t contentTag) noexcept {
|
||||
StereoProviderRegistration registration;
|
||||
{
|
||||
std::lock_guard lock(g_stereoRegistrationMutex);
|
||||
@@ -588,8 +574,7 @@ std::optional<AuroraStereoFrame> request_stereo_frame(uint32_t logicalFrame,
|
||||
return std::nullopt;
|
||||
}
|
||||
if (frame.mode != AURORA_STEREO_FRAME_IMMERSIVE_REPLAY && frame.mode != AURORA_STEREO_FRAME_VIRTUAL_SCREEN) {
|
||||
Log.warn("Stereo frame {} has invalid mode {}; rendering in mono", logicalFrame,
|
||||
static_cast<uint32_t>(frame.mode));
|
||||
Log.warn("Stereo frame {} has invalid mode {}; rendering in mono", logicalFrame, static_cast<uint32_t>(frame.mode));
|
||||
return std::nullopt;
|
||||
}
|
||||
// A virtual-screen packet only copies the completed mono image and remains
|
||||
@@ -597,8 +582,7 @@ std::optional<AuroraStereoFrame> request_stereo_frame(uint32_t logicalFrame,
|
||||
// transforms, so it requires the exact tag latched for this sealed content.
|
||||
if (frame.mode == AURORA_STEREO_FRAME_IMMERSIVE_REPLAY &&
|
||||
(contentTag == AURORA_STEREO_CONTENT_TAG_UNKNOWN || frame.contentTag != contentTag)) {
|
||||
Log.debug("Stereo frame {} content tag does not match its sealed frame; rendering in mono",
|
||||
logicalFrame);
|
||||
Log.debug("Stereo frame {} content tag does not match its sealed frame; rendering in mono", logicalFrame);
|
||||
return std::nullopt;
|
||||
}
|
||||
|
||||
@@ -617,8 +601,7 @@ std::optional<AuroraStereoFrame> request_stereo_frame(uint32_t logicalFrame,
|
||||
const bool transformsValid = frame.mode == AURORA_STEREO_FRAME_VIRTUAL_SCREEN ||
|
||||
(finite(input.projection, 16) && finite(input.viewFromCenter, 12));
|
||||
if (input.width == 0 || input.height == 0 || !transformsValid) {
|
||||
Log.warn("Stereo frame {} has invalid eye {} dimensions or transforms; rendering in mono",
|
||||
logicalFrame, eye);
|
||||
Log.warn("Stereo frame {} has invalid eye {} dimensions or transforms; rendering in mono", logicalFrame, eye);
|
||||
return std::nullopt;
|
||||
}
|
||||
}
|
||||
@@ -648,8 +631,7 @@ gfx::StereoReplayFrame make_stereo_replay_frame(const AuroraStereoFrame& input)
|
||||
return replay;
|
||||
}
|
||||
|
||||
void encode_virtual_screen_eye(wgpu::CommandEncoder& encoder, const webgpu::PresentSource& source,
|
||||
uint32_t eyeIndex) {
|
||||
void encode_virtual_screen_eye(wgpu::CommandEncoder& encoder, const webgpu::PresentSource& source, uint32_t eyeIndex) {
|
||||
const auto& output = g_stereoEyeTargets[eyeIndex].output();
|
||||
const std::array attachments{
|
||||
wgpu::RenderPassColorAttachment{
|
||||
@@ -667,12 +649,11 @@ void encode_virtual_screen_eye(wgpu::CommandEncoder& encoder, const webgpu::Pres
|
||||
{
|
||||
const auto pass = encoder.BeginRenderPass(&descriptor);
|
||||
if (source.bindGroup && source.size.width != 0 && source.size.height != 0) {
|
||||
const auto viewport = webgpu::calculate_present_viewport(
|
||||
output.size.width, output.size.height, source.size.width, source.size.height);
|
||||
const auto viewport = webgpu::calculate_present_viewport(output.size.width, output.size.height, source.size.width,
|
||||
source.size.height);
|
||||
pass.SetPipeline(webgpu::g_CopyPipeline);
|
||||
pass.SetBindGroup(0, source.bindGroup, 0, nullptr);
|
||||
pass.SetViewport(viewport.left, viewport.top, viewport.width, viewport.height,
|
||||
viewport.znear, viewport.zfar);
|
||||
pass.SetViewport(viewport.left, viewport.top, viewport.width, viewport.height, viewport.znear, viewport.zfar);
|
||||
pass.Draw(3);
|
||||
}
|
||||
pass.End();
|
||||
@@ -720,9 +701,7 @@ std::optional<PendingStereoSink> run_stereo_sink(wgpu::CommandEncoder& encoder,
|
||||
};
|
||||
}
|
||||
|
||||
void request_surface_reconfigure() noexcept {
|
||||
g_surfaceReconfigurePending.store(true, std::memory_order_release);
|
||||
}
|
||||
void request_surface_reconfigure() noexcept { g_surfaceReconfigurePending.store(true, std::memory_order_release); }
|
||||
|
||||
void request_surface_recreate() noexcept {
|
||||
g_surfaceRecreatePending.store(true, std::memory_order_release);
|
||||
@@ -743,7 +722,8 @@ std::optional<PendingFrameCapture> encode_frame_capture(const wgpu::CommandEncod
|
||||
const webgpu::PresentSource& source) {
|
||||
const uint32_t requestedFrame = g_captureFrame.load(std::memory_order_acquire);
|
||||
const uint32_t currentFrame = gfx::current_frame();
|
||||
if (requestedFrame == UINT32_MAX || currentFrame < requestedFrame) return std::nullopt;
|
||||
if (requestedFrame == UINT32_MAX || currentFrame < requestedFrame)
|
||||
return std::nullopt;
|
||||
g_captureFrame.store(UINT32_MAX, std::memory_order_release);
|
||||
if (currentFrame != requestedFrame) {
|
||||
Log.error("Missed requested frame capture {} (current frame {})", requestedFrame, currentFrame);
|
||||
@@ -753,10 +733,10 @@ std::optional<PendingFrameCapture> encode_frame_capture(const wgpu::CommandEncod
|
||||
Log.error("Frame {} capture has no present-source texture", currentFrame);
|
||||
return std::nullopt;
|
||||
}
|
||||
const bool bgra = source.format == wgpu::TextureFormat::BGRA8Unorm ||
|
||||
source.format == wgpu::TextureFormat::BGRA8UnormSrgb;
|
||||
const bool rgba = source.format == wgpu::TextureFormat::RGBA8Unorm ||
|
||||
source.format == wgpu::TextureFormat::RGBA8UnormSrgb;
|
||||
const bool bgra =
|
||||
source.format == wgpu::TextureFormat::BGRA8Unorm || source.format == wgpu::TextureFormat::BGRA8UnormSrgb;
|
||||
const bool rgba =
|
||||
source.format == wgpu::TextureFormat::RGBA8Unorm || source.format == wgpu::TextureFormat::RGBA8UnormSrgb;
|
||||
if (!bgra && !rgba) {
|
||||
Log.error("Frame {} capture does not support texture format {}", currentFrame,
|
||||
magic_enum::enum_name(source.format));
|
||||
@@ -795,12 +775,12 @@ std::optional<PendingFrameCapture> encode_frame_capture(const wgpu::CommandEncod
|
||||
void complete_frame_capture(PendingFrameCapture& capture) {
|
||||
wgpu::MapAsyncStatus mapStatus = wgpu::MapAsyncStatus::CallbackCancelled;
|
||||
wgpu::StringView mapMessage{};
|
||||
const auto future = capture.buffer.MapAsync(
|
||||
wgpu::MapMode::Read, 0, capture.bufferSize, wgpu::CallbackMode::WaitAnyOnly,
|
||||
[&mapStatus, &mapMessage](wgpu::MapAsyncStatus status, wgpu::StringView message) {
|
||||
mapStatus = status;
|
||||
mapMessage = message;
|
||||
});
|
||||
const auto future =
|
||||
capture.buffer.MapAsync(wgpu::MapMode::Read, 0, capture.bufferSize, wgpu::CallbackMode::WaitAnyOnly,
|
||||
[&mapStatus, &mapMessage](wgpu::MapAsyncStatus status, wgpu::StringView message) {
|
||||
mapStatus = status;
|
||||
mapMessage = message;
|
||||
});
|
||||
const auto waitStatus = g_instance.WaitAny(future, 5000000000);
|
||||
if (waitStatus != wgpu::WaitStatus::Success || mapStatus != wgpu::MapAsyncStatus::Success) {
|
||||
Log.error("Frame capture readback failed wait={} map={} message={}", magic_enum::enum_name(waitStatus),
|
||||
@@ -879,7 +859,6 @@ constexpr std::array<AuroraBackend, 0> PreferredBackendOrder{};
|
||||
|
||||
bool g_initialFrame = false;
|
||||
|
||||
|
||||
AuroraInfo initialize(int argc, char* argv[], const AuroraConfig& config) noexcept {
|
||||
g_config = config;
|
||||
Log.info("Aurora initializing");
|
||||
@@ -935,15 +914,15 @@ AuroraInfo initialize(int argc, char* argv[], const AuroraConfig& config) noexce
|
||||
window::destroy_window();
|
||||
}
|
||||
} else {
|
||||
Log.error("Failed to create a window for backend {}: {}", backend_name(selectedBackend),
|
||||
SDL_GetError());
|
||||
Log.error("Failed to create a window for backend {}: {}", backend_name(selectedBackend), SDL_GetError());
|
||||
}
|
||||
if (!windowCreated) {
|
||||
/* An explicitly requested backend that cannot be brought up falls back to the BACKEND_AUTO
|
||||
* search instead of aborting, and the substitution is always reported. */
|
||||
Log.error("Requested graphics backend {} is unavailable on this system; "
|
||||
"falling back to automatic selection",
|
||||
backend_name(requestedBackend));
|
||||
Log.error(
|
||||
"Requested graphics backend {} is unavailable on this system; "
|
||||
"falling back to automatic selection",
|
||||
backend_name(requestedBackend));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -964,9 +943,10 @@ AuroraInfo initialize(int argc, char* argv[], const AuroraConfig& config) noexce
|
||||
|
||||
ASSERT(windowCreated, "Error creating window: {}", SDL_GetError());
|
||||
if (requestedBackend != BACKEND_AUTO && selectedBackend != requestedBackend) {
|
||||
Log.error("Graphics backend fallback in effect: video.graphics_api requested {}, "
|
||||
"running on {}",
|
||||
backend_name(requestedBackend), backend_name(selectedBackend));
|
||||
Log.error(
|
||||
"Graphics backend fallback in effect: video.graphics_api requested {}, "
|
||||
"running on {}",
|
||||
backend_name(requestedBackend), backend_name(selectedBackend));
|
||||
}
|
||||
|
||||
// Initialize SDL_Renderer for ImGui when we can't use a Dawn backend
|
||||
@@ -1077,11 +1057,9 @@ struct PresentationJob {
|
||||
uint32_t slidPeriods = 0;
|
||||
};
|
||||
|
||||
std::array<std::vector<std::shared_ptr<PresentationImage>>, gx::MaxInterpolatedFrames + 1>
|
||||
g_presentationImagePools;
|
||||
std::array<std::vector<std::shared_ptr<PresentationImage>>, gx::MaxInterpolatedFrames + 1> g_presentationImagePools;
|
||||
|
||||
std::shared_ptr<PresentationImage> acquire_presentation_image(size_t slot, uint32_t width,
|
||||
uint32_t height) {
|
||||
std::shared_ptr<PresentationImage> acquire_presentation_image(size_t slot, uint32_t width, uint32_t height) {
|
||||
auto& pool = g_presentationImagePools.at(slot);
|
||||
// A use count of one means only the pool holds the image, so no job can be reading it. Idle
|
||||
// images from an older surface size are dropped here instead of leaking for the run.
|
||||
@@ -1122,16 +1100,15 @@ bool present_presentation_job(const PresentationJob& job) {
|
||||
std::chrono::nanoseconds scheduleWaitDuration{};
|
||||
std::chrono::nanoseconds presentDuration{};
|
||||
bool presented = false;
|
||||
if (g_surfaceReconfigurePending.load(std::memory_order_acquire) ||
|
||||
window::native_resize_pending() || !window::is_presentable()) {
|
||||
if (g_surfaceReconfigurePending.load(std::memory_order_acquire) || window::native_resize_pending() ||
|
||||
!window::is_presentable()) {
|
||||
return false;
|
||||
}
|
||||
std::chrono::nanoseconds lateBy{};
|
||||
if (job.presentAt != PresentClock::time_point{}) {
|
||||
// How expired the deadline already is at dequeue. Positive values mean the
|
||||
// slot cannot be paced and fires immediately, which is a burst symptom.
|
||||
lateBy = std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() -
|
||||
job.presentAt);
|
||||
lateBy = std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - job.presentAt);
|
||||
}
|
||||
{
|
||||
window::SurfaceLock surfaceLock;
|
||||
@@ -1139,26 +1116,23 @@ bool present_presentation_job(const PresentationJob& job) {
|
||||
// surface lock covers all of them. The renderer mutex is deliberately not taken.
|
||||
const auto surfaceLockStarted = PresentClock::now();
|
||||
std::unique_lock surfaceOwnership(g_surfaceMutex);
|
||||
surfaceLockDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||
PresentClock::now() - surfaceLockStarted);
|
||||
surfaceLockDuration =
|
||||
std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - surfaceLockStarted);
|
||||
// Surface contention is waiting, not encoding, so keep it out of the encode timing where a
|
||||
// reconfigure would look like GPU command recording.
|
||||
const auto workStarted = PresentClock::now();
|
||||
bool surfaceSizeChanged = window::native_resize_pending();
|
||||
if (!surfaceSizeChanged &&
|
||||
!g_surfaceReconfigurePending.load(std::memory_order_acquire) &&
|
||||
if (!surfaceSizeChanged && !g_surfaceReconfigurePending.load(std::memory_order_acquire) &&
|
||||
window::is_presentable() && g_surface) {
|
||||
// native_window_size_matches compares the OS client size with the configured swapchain, which is
|
||||
// what native_fb_* reports. One window query instead of SDL's per-call ones.
|
||||
surfaceSizeChanged = !window::native_window_size_matches(
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height);
|
||||
surfaceSizeChanged = !window::native_window_size_matches(webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height);
|
||||
}
|
||||
if (!surfaceSizeChanged && window::is_presentable()) {
|
||||
const auto acquireStarted = PresentClock::now();
|
||||
auto acquired = acquire_surface_texture();
|
||||
acquireDuration =
|
||||
std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - acquireStarted);
|
||||
acquireDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - acquireStarted);
|
||||
if (acquired) {
|
||||
const wgpu::CommandEncoderDescriptor encoderDescriptor{
|
||||
.label = "Presentation encoder",
|
||||
@@ -1179,59 +1153,52 @@ bool present_presentation_job(const PresentationJob& job) {
|
||||
const auto pass = encoder.BeginRenderPass(&renderPassDescriptor);
|
||||
pass.SetPipeline(webgpu::g_CopyPipeline);
|
||||
pass.SetBindGroup(0, job.image->bindGroup, 0, nullptr);
|
||||
pass.SetViewport(0.f, 0.f,
|
||||
static_cast<float>(webgpu::g_graphicsConfig.surfaceConfiguration.width),
|
||||
static_cast<float>(webgpu::g_graphicsConfig.surfaceConfiguration.height),
|
||||
0.f, 1.f);
|
||||
pass.SetViewport(0.f, 0.f, static_cast<float>(webgpu::g_graphicsConfig.surfaceConfiguration.width),
|
||||
static_cast<float>(webgpu::g_graphicsConfig.surfaceConfiguration.height), 0.f, 1.f);
|
||||
pass.Draw(3);
|
||||
pass.End();
|
||||
const auto encodeFinished = PresentClock::now();
|
||||
encodeDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||
encodeFinished - workStarted - acquireDuration);
|
||||
encodeDuration =
|
||||
std::chrono::duration_cast<std::chrono::nanoseconds>(encodeFinished - workStarted - acquireDuration);
|
||||
const wgpu::CommandBufferDescriptor cmdBufDescriptor{
|
||||
.label = "Presentation command buffer",
|
||||
};
|
||||
const auto finishStarted = PresentClock::now();
|
||||
const auto buffer = encoder.Finish(&cmdBufDescriptor);
|
||||
finishDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||
PresentClock::now() - finishStarted);
|
||||
finishDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - finishStarted);
|
||||
const auto submitStarted = PresentClock::now();
|
||||
{
|
||||
std::lock_guard submitLock(g_queueSubmitMutex);
|
||||
g_queue.Submit(1, &buffer);
|
||||
}
|
||||
submitDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||
PresentClock::now() - submitStarted);
|
||||
submitDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - submitStarted);
|
||||
// Pace the Present() call itself, not the whole unit, so acquire/encode/submit variance stays out
|
||||
// of the cadence. Holding the image across the wait is safe while the surface lock is held.
|
||||
if (job.presentAt != PresentClock::time_point{}) {
|
||||
const auto scheduleWaitStarted = PresentClock::now();
|
||||
wait_until_precise(job.presentAt);
|
||||
scheduleWaitDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||
PresentClock::now() - scheduleWaitStarted);
|
||||
scheduleWaitDuration =
|
||||
std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - scheduleWaitStarted);
|
||||
}
|
||||
// A native resize can arrive after acquisition, so drop the obsolete image and let the render
|
||||
// worker reconfigure at its ordered frame boundary.
|
||||
if (!g_surfaceReconfigurePending.load(std::memory_order_acquire) &&
|
||||
!window::native_resize_pending() && window::is_presentable() &&
|
||||
window::native_window_size_matches(
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height)) {
|
||||
if (!g_surfaceReconfigurePending.load(std::memory_order_acquire) && !window::native_resize_pending() &&
|
||||
window::is_presentable() &&
|
||||
window::native_window_size_matches(webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height)) {
|
||||
const auto presentStarted = PresentClock::now();
|
||||
wgpu::Status presentStatus;
|
||||
{
|
||||
std::lock_guard submitLock(g_queueSubmitMutex);
|
||||
presentStatus = g_surface.Present();
|
||||
}
|
||||
presentDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(
|
||||
PresentClock::now() - presentStarted);
|
||||
presentDuration = std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - presentStarted);
|
||||
if (presentStatus == wgpu::Status::Success) {
|
||||
presented = true;
|
||||
record_successful_present(
|
||||
job.interpolated, job.logicalFrame, acquireDuration, encodeDuration,
|
||||
finishDuration, submitDuration, presentDuration,
|
||||
std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() -
|
||||
submissionStarted),
|
||||
job.interpolated, job.logicalFrame, acquireDuration, encodeDuration, finishDuration, submitDuration,
|
||||
presentDuration,
|
||||
std::chrono::duration_cast<std::chrono::nanoseconds>(PresentClock::now() - submissionStarted),
|
||||
job.duplicated, job.slidPeriods, lateBy);
|
||||
} else {
|
||||
Log.warn("Surface present failed: {}", static_cast<int>(presentStatus));
|
||||
@@ -1257,8 +1224,7 @@ bool present_presentation_job(const PresentationJob& job) {
|
||||
++s_consecutiveStalledPresents;
|
||||
const auto now = PresentClock::now();
|
||||
if (s_consecutiveStalledPresents >= kStallRebuildThreshold &&
|
||||
(s_lastStallRebuild == PresentClock::time_point{} ||
|
||||
now - s_lastStallRebuild >= kStallRebuildCooldown)) {
|
||||
(s_lastStallRebuild == PresentClock::time_point{} || now - s_lastStallRebuild >= kStallRebuildCooldown)) {
|
||||
s_lastStallRebuild = now;
|
||||
s_consecutiveStalledPresents = 0;
|
||||
rebuildRequested = true;
|
||||
@@ -1266,17 +1232,18 @@ bool present_presentation_job(const PresentationJob& job) {
|
||||
} else {
|
||||
s_consecutiveStalledPresents = 0;
|
||||
}
|
||||
Log.warn("Presentation job took {:.1f} ms (surface lock {:.1f}, acquire {:.1f}, encode {:.1f}, "
|
||||
"finish {:.1f}, submit {:.1f}, schedule wait {:.1f}, present {:.1f}){}",
|
||||
std::chrono::duration<double, std::milli>(totalDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(surfaceLockDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(acquireDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(encodeDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(finishDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(submitDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(scheduleWaitDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(presentDuration).count(),
|
||||
rebuildRequested ? "; rebuilding the surface" : "");
|
||||
Log.warn(
|
||||
"Presentation job took {:.1f} ms (surface lock {:.1f}, acquire {:.1f}, encode {:.1f}, "
|
||||
"finish {:.1f}, submit {:.1f}, schedule wait {:.1f}, present {:.1f}){}",
|
||||
std::chrono::duration<double, std::milli>(totalDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(surfaceLockDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(acquireDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(encodeDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(finishDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(submitDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(scheduleWaitDuration).count(),
|
||||
std::chrono::duration<double, std::milli>(presentDuration).count(),
|
||||
rebuildRequested ? "; rebuilding the surface" : "");
|
||||
if (rebuildRequested) {
|
||||
request_surface_reconfigure();
|
||||
}
|
||||
@@ -1345,8 +1312,7 @@ void wait_for_presenter_idle() noexcept {
|
||||
return;
|
||||
}
|
||||
std::unique_lock lock(g_presenter.mutex);
|
||||
g_presenter.cv.wait(
|
||||
lock, [] { return g_presenter.jobs.empty() && !g_presenter.presenting; });
|
||||
g_presenter.cv.wait(lock, [] { return g_presenter.jobs.empty() && !g_presenter.presenting; });
|
||||
}
|
||||
|
||||
void enqueue_presentations(std::vector<PresentationJob>&& jobs) {
|
||||
@@ -1408,18 +1374,15 @@ void stop_presenter() noexcept {
|
||||
|
||||
// `presentSource` is latched in the seal prologue: by the time this encodes, the producer's next
|
||||
// gfx::begin_frame() may already have cleared the display-copy override.
|
||||
void encode_presentation_snapshot(const wgpu::CommandEncoder& encoder,
|
||||
const webgpu::PresentSource& presentSource,
|
||||
const PresentationImage& image,
|
||||
bool includeImGui) {
|
||||
void encode_presentation_snapshot(const wgpu::CommandEncoder& encoder, const webgpu::PresentSource& presentSource,
|
||||
const PresentationImage& image, bool includeImGui) {
|
||||
ZoneScoped;
|
||||
auto viewport = webgpu::calculate_present_viewport(
|
||||
image.texture.size.width, image.texture.size.height, presentSource.size.width,
|
||||
presentSource.size.height);
|
||||
auto viewport = webgpu::calculate_present_viewport(image.texture.size.width, image.texture.size.height,
|
||||
presentSource.size.width, presentSource.size.height);
|
||||
float presentAspect = 0.f;
|
||||
if (window::get_present_aspect_ratio(presentAspect)) {
|
||||
viewport = webgpu::calculate_present_viewport_for_aspect(
|
||||
image.texture.size.width, image.texture.size.height, presentAspect);
|
||||
viewport = webgpu::calculate_present_viewport_for_aspect(image.texture.size.width, image.texture.size.height,
|
||||
presentAspect);
|
||||
}
|
||||
wgpu::BindGroup presentBindGroup = presentSource.bindGroup;
|
||||
{
|
||||
@@ -1438,8 +1401,7 @@ void encode_presentation_snapshot(const wgpu::CommandEncoder& encoder,
|
||||
const auto pass = encoder.BeginRenderPass(&renderPassDescriptor);
|
||||
pass.SetPipeline(webgpu::g_CopyPipeline);
|
||||
pass.SetBindGroup(0, presentBindGroup, 0, nullptr);
|
||||
pass.SetViewport(viewport.left, viewport.top, viewport.width, viewport.height,
|
||||
viewport.znear, viewport.zfar);
|
||||
pass.SetViewport(viewport.left, viewport.top, viewport.width, viewport.height, viewport.znear, viewport.zfar);
|
||||
pass.Draw(3);
|
||||
pass.End();
|
||||
}
|
||||
@@ -1478,6 +1440,7 @@ void shutdown() noexcept {
|
||||
{
|
||||
std::lock_guard lock(g_stereoRegistrationMutex);
|
||||
g_stereoProvider = {};
|
||||
g_stereoProviderActive.store(false, std::memory_order_release);
|
||||
#ifdef AURORA_ENABLE_GX
|
||||
g_stereoSink = {};
|
||||
#endif
|
||||
@@ -1502,14 +1465,11 @@ bool begin_frame_impl(bool pumpEvents, ImGuiFramePolicy imguiPolicy, bool* imgui
|
||||
if (pumpEvents) {
|
||||
window::pump_events();
|
||||
}
|
||||
const bool surfaceReconfigurePending =
|
||||
g_surfaceReconfigurePending.load(std::memory_order_acquire);
|
||||
const bool surfaceReconfigurePending = g_surfaceReconfigurePending.load(std::memory_order_acquire);
|
||||
const bool surfaceMutationRequired =
|
||||
surfaceReconfigurePending || !window::is_presentable() || !g_surface ||
|
||||
window::native_resize_pending() ||
|
||||
!window::native_window_size_matches(
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height);
|
||||
surfaceReconfigurePending || !window::is_presentable() || !g_surface || window::native_resize_pending() ||
|
||||
!window::native_window_size_matches(webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height);
|
||||
if (surfaceMutationRequired) {
|
||||
wait_for_presenter_idle();
|
||||
window::SurfaceLock surfaceLock;
|
||||
@@ -1526,10 +1486,8 @@ bool begin_frame_impl(bool pumpEvents, ImGuiFramePolicy imguiPolicy, bool* imgui
|
||||
if (window::is_paused()) {
|
||||
return false;
|
||||
}
|
||||
const bool consumeSurfaceReconfigure =
|
||||
g_surfaceReconfigurePending.exchange(false, std::memory_order_acq_rel);
|
||||
const bool consumeSurfaceRecreate =
|
||||
g_surfaceRecreatePending.exchange(false, std::memory_order_acq_rel);
|
||||
const bool consumeSurfaceReconfigure = g_surfaceReconfigurePending.exchange(false, std::memory_order_acq_rel);
|
||||
const bool consumeSurfaceRecreate = g_surfaceRecreatePending.exchange(false, std::memory_order_acq_rel);
|
||||
if (!g_surface || consumeSurfaceReconfigure) {
|
||||
// Reconfigure in place unless the surface was actually lost. See
|
||||
// g_surfaceRecreatePending for why destroying a live surface here is fatal
|
||||
@@ -1598,8 +1556,7 @@ struct SealedFrameContext {
|
||||
|
||||
// Phase 1: everything that touches producer-shared renderer state. Needs g_rendererGpuMutex and
|
||||
// a FIFO already drained into the recorded pass list.
|
||||
void seal_frame_locked(gfx::SealedFrame& sealedFrame, SealedFrameContext& ctx,
|
||||
uint64_t contentTag) {
|
||||
void seal_frame_locked(gfx::SealedFrame& sealedFrame, SealedFrameContext& ctx, uint64_t contentTag) {
|
||||
ZoneScopedN("Seal frame");
|
||||
const auto encoderDescriptor = wgpu::CommandEncoderDescriptor{
|
||||
.label = "Redraw encoder",
|
||||
@@ -1796,8 +1753,7 @@ std::vector<PresentationJob> encode_sealed_frame(gfx::SealedFrame& sealedFrame,
|
||||
const auto now = PresentClock::now();
|
||||
if (now > anchor) {
|
||||
const auto behind = std::chrono::duration_cast<std::chrono::nanoseconds>(now - anchor);
|
||||
const uint64_t periods =
|
||||
static_cast<uint64_t>(behind.count()) / ctx.scheduleIntervalNanos + 1u;
|
||||
const uint64_t periods = static_cast<uint64_t>(behind.count()) / ctx.scheduleIntervalNanos + 1u;
|
||||
anchor += std::chrono::nanoseconds{static_cast<int64_t>(periods * ctx.scheduleIntervalNanos)};
|
||||
}
|
||||
if (s_lastGroupAnchor != PresentClock::time_point{} && anchor <= s_lastGroupAnchor) {
|
||||
@@ -1807,8 +1763,7 @@ std::vector<PresentationJob> encode_sealed_frame(gfx::SealedFrame& sealedFrame,
|
||||
std::chrono::duration_cast<std::chrono::nanoseconds>(anchor - presentationJobs.front().presentAt);
|
||||
if (shift.count() > 0) {
|
||||
const uint32_t slidPeriods = static_cast<uint32_t>(
|
||||
(static_cast<uint64_t>(shift.count()) + ctx.scheduleIntervalNanos - 1u) /
|
||||
ctx.scheduleIntervalNanos);
|
||||
(static_cast<uint64_t>(shift.count()) + ctx.scheduleIntervalNanos - 1u) / ctx.scheduleIntervalNanos);
|
||||
for (auto& job : presentationJobs) {
|
||||
job.presentAt += shift;
|
||||
job.slidPeriods = slidPeriods;
|
||||
@@ -1843,8 +1798,7 @@ void publish_presentations(std::vector<PresentationJob>&& presentationJobs, bool
|
||||
#else
|
||||
// Keep presentation on the presenter whenever the async frame worker runs, even with
|
||||
// interpolation off, so every mode shares one surface/resize path. RenderDoc keeps the sync path.
|
||||
if (frame_worker_requested() || interpolationActive ||
|
||||
g_presenterStarted.load(std::memory_order_acquire)) {
|
||||
if (frame_worker_requested() || interpolationActive || g_presenterStarted.load(std::memory_order_acquire)) {
|
||||
enqueue_presentations(std::move(presentationJobs));
|
||||
} else {
|
||||
for (const auto& job : presentationJobs) {
|
||||
@@ -1885,14 +1839,11 @@ void record_frame_telemetry() {
|
||||
GetThreadTimes(GetCurrentThread(), &creation, &exit, &threadKernel, &threadUser) != FALSE;
|
||||
const uint64_t processCpu100ns =
|
||||
processTimesAvailable ? fileTimeValue(processKernel) + fileTimeValue(processUser) : 0;
|
||||
const uint64_t threadCpu100ns =
|
||||
threadTimesAvailable ? fileTimeValue(threadKernel) + fileTimeValue(threadUser) : 0;
|
||||
const uint64_t threadCpu100ns = threadTimesAvailable ? fileTimeValue(threadKernel) + fileTimeValue(threadUser) : 0;
|
||||
static uint64_t previousProcessCpu100ns = processCpu100ns;
|
||||
static uint64_t previousThreadCpu100ns = threadCpu100ns;
|
||||
TracyPlot("aurora: processCpuUsPerFrame",
|
||||
static_cast<int64_t>((processCpu100ns - previousProcessCpu100ns) / 10));
|
||||
TracyPlot("aurora: mainThreadCpuUsPerFrame",
|
||||
static_cast<int64_t>((threadCpu100ns - previousThreadCpu100ns) / 10));
|
||||
TracyPlot("aurora: processCpuUsPerFrame", static_cast<int64_t>((processCpu100ns - previousProcessCpu100ns) / 10));
|
||||
TracyPlot("aurora: mainThreadCpuUsPerFrame", static_cast<int64_t>((threadCpu100ns - previousThreadCpu100ns) / 10));
|
||||
previousProcessCpu100ns = processCpu100ns;
|
||||
previousThreadCpu100ns = threadCpu100ns;
|
||||
#endif
|
||||
@@ -1932,8 +1883,7 @@ bool run_frame_worker_cycle(gfx::SealedFrame& sealedFrame, uint64_t contentTag)
|
||||
// staging buffers the producer's drain has nowhere to put its commands.
|
||||
bool imguiNewFrameOwed = false;
|
||||
const bool prepared = begin_frame_impl(
|
||||
false, overlapEncode ? ImGuiFramePolicy::Deferred : ImGuiFramePolicy::Immediate,
|
||||
&imguiNewFrameOwed);
|
||||
false, overlapEncode ? ImGuiFramePolicy::Deferred : ImGuiFramePolicy::Immediate, &imguiNewFrameOwed);
|
||||
|
||||
{
|
||||
std::lock_guard lock(g_frameWorker.mutex);
|
||||
@@ -2014,12 +1964,10 @@ bool begin_frame() noexcept {
|
||||
#ifdef AURORA_ENABLE_GX
|
||||
// A surface mutation can legitimately fail preparation, and optimistic success would let GX/ImGui
|
||||
// record into a frame that was never begun, so join this path and return its real result.
|
||||
waitForSurfacePreparation =
|
||||
!window::is_presentable() || !g_surface ||
|
||||
window::native_resize_pending() || window::is_paused() ||
|
||||
!window::native_window_size_matches(
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height);
|
||||
waitForSurfacePreparation = !window::is_presentable() || !g_surface || window::native_resize_pending() ||
|
||||
window::is_paused() ||
|
||||
!window::native_window_size_matches(webgpu::g_graphicsConfig.surfaceConfiguration.width,
|
||||
webgpu::g_graphicsConfig.surfaceConfiguration.height);
|
||||
#endif
|
||||
bool workerPreparationPending = false;
|
||||
{
|
||||
@@ -2126,6 +2074,7 @@ void quiesce_frame_worker() noexcept {
|
||||
}
|
||||
}
|
||||
std::recursive_mutex& renderer_gpu_mutex() noexcept { return g_rendererGpuMutex; }
|
||||
bool stereo_frame_provider_active() noexcept { return g_stereoProviderActive.load(std::memory_order_acquire); }
|
||||
|
||||
void set_stereo_frame_provider(AuroraStereoFrameProvider provider, void* userdata) noexcept {
|
||||
std::lock_guard lock(g_stereoRegistrationMutex);
|
||||
@@ -2133,6 +2082,7 @@ void set_stereo_frame_provider(AuroraStereoFrameProvider provider, void* userdat
|
||||
.callback = provider,
|
||||
.userdata = userdata,
|
||||
};
|
||||
g_stereoProviderActive.store(provider != nullptr, std::memory_order_release);
|
||||
}
|
||||
|
||||
#ifdef AURORA_ENABLE_GX
|
||||
@@ -2250,8 +2200,7 @@ bool aurora_flush_efb_copies_to_ram() {
|
||||
}
|
||||
bool aurora_flush_efb_copy_to_ram(void* dest) {
|
||||
#ifdef AURORA_ENABLE_GX
|
||||
if (dest == nullptr || !aurora::gfx::efb_ram::has_pending(dest) ||
|
||||
!aurora::gfx::efb_ram::prepare_downloads(dest)) {
|
||||
if (dest == nullptr || !aurora::gfx::efb_ram::has_pending(dest) || !aurora::gfx::efb_ram::prepare_downloads(dest)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@@ -2297,12 +2246,14 @@ void aurora_set_log_level(AuroraLogLevel level) { aurora::g_config.logLevel = le
|
||||
void aurora_set_pause_on_focus_lost(bool value) { aurora::g_config.pauseOnFocusLost = value; }
|
||||
void aurora_set_disable_copy_filter(bool disabled) { aurora::g_config.disableCopyFilter = disabled; }
|
||||
bool aurora_get_disable_copy_filter() { return aurora::g_config.disableCopyFilter; }
|
||||
void aurora_set_stereo_stop_at_display_copy(bool enabled) {
|
||||
aurora::gfx::set_stereo_stop_at_display_copy(enabled);
|
||||
}
|
||||
void aurora_set_stereo_stop_at_display_copy(bool enabled) { aurora::gfx::set_stereo_stop_at_display_copy(enabled); }
|
||||
bool aurora_get_stereo_stop_at_display_copy() { return aurora::gfx::get_stereo_stop_at_display_copy(); }
|
||||
void aurora_set_stereo_skip_copy_clears(bool enabled) { aurora::gfx::set_stereo_skip_copy_clears(enabled); }
|
||||
bool aurora_get_stereo_skip_copy_clears() { return aurora::gfx::get_stereo_skip_copy_clears(); }
|
||||
void aurora_set_stereo_hud_screen(bool enabled, float width, float distance) {
|
||||
aurora::gfx::set_stereo_hud_screen(enabled, width, distance);
|
||||
}
|
||||
bool aurora_get_stereo_hud_screen_enabled() { return aurora::gfx::get_stereo_hud_screen_enabled(); }
|
||||
void aurora_set_background_input(bool value) {
|
||||
aurora::g_config.allowJoystickBackgroundEvents = value;
|
||||
aurora::window::set_background_input(value);
|
||||
|
||||
@@ -80,8 +80,7 @@ CopySourceRect map_texture_copy_source(const aurora::gfx::ClipRect& source, bool
|
||||
const int64_t remainder = numerator % denominator;
|
||||
const int64_t remainderMagnitude = remainder < 0 ? -remainder : remainder;
|
||||
const int64_t roundingThreshold = (denominator + 1) / 2;
|
||||
const int64_t nearestValue =
|
||||
quotient + (remainderMagnitude >= roundingThreshold ? (numerator < 0 ? -1 : 1) : 0);
|
||||
const int64_t nearestValue = quotient + (remainderMagnitude >= roundingThreshold ? (numerator < 0 ? -1 : 1) : 0);
|
||||
return MappedEdge{
|
||||
.sample = static_cast<float>(static_cast<double>(numerator) / static_cast<double>(denominator)),
|
||||
.nearest = static_cast<int32_t>(nearestValue),
|
||||
@@ -103,10 +102,6 @@ CopySourceRect map_texture_copy_source(const aurora::gfx::ClipRect& source, bool
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
u32 pack_copy_filter_samples(u8 reg, const std::array<std::array<u8, 2>, 12>& samplePattern, size_t first) {
|
||||
u32 value = static_cast<u32>(reg) << 24;
|
||||
for (size_t i = 0; i < 6; ++i) {
|
||||
@@ -118,14 +113,13 @@ u32 pack_copy_filter_samples(u8 reg, const std::array<std::array<u8, 2>, 12>& sa
|
||||
}
|
||||
|
||||
u32 pack_copy_filter0(const std::array<u8, 7>& vfilter) {
|
||||
return 0x53000000u | (static_cast<u32>(vfilter[0] & 0x3fu) << 0) |
|
||||
(static_cast<u32>(vfilter[1] & 0x3fu) << 6) | (static_cast<u32>(vfilter[2] & 0x3fu) << 12) |
|
||||
(static_cast<u32>(vfilter[3] & 0x3fu) << 18);
|
||||
return 0x53000000u | (static_cast<u32>(vfilter[0] & 0x3fu) << 0) | (static_cast<u32>(vfilter[1] & 0x3fu) << 6) |
|
||||
(static_cast<u32>(vfilter[2] & 0x3fu) << 12) | (static_cast<u32>(vfilter[3] & 0x3fu) << 18);
|
||||
}
|
||||
|
||||
u32 pack_copy_filter1(const std::array<u8, 7>& vfilter) {
|
||||
return 0x54000000u | (static_cast<u32>(vfilter[4] & 0x3fu) << 0) |
|
||||
(static_cast<u32>(vfilter[5] & 0x3fu) << 6) | (static_cast<u32>(vfilter[6] & 0x3fu) << 12);
|
||||
return 0x54000000u | (static_cast<u32>(vfilter[4] & 0x3fu) << 0) | (static_cast<u32>(vfilter[5] & 0x3fu) << 6) |
|
||||
(static_cast<u32>(vfilter[6] & 0x3fu) << 12);
|
||||
}
|
||||
|
||||
std::array<u32, 3> combined_copy_filter_coefficients(const std::array<u8, 7>& vfilter) {
|
||||
@@ -167,7 +161,7 @@ aurora::gfx::TextureHandle create_copy_texture(u32 width, u32 height, GXTexFmt t
|
||||
return aurora::gfx::new_render_texture(width, height, fmt, "Resolved Texture");
|
||||
}
|
||||
|
||||
// Reuse retired copy targets to avoid unbounded GPU texture allocation.
|
||||
// Reuse retired copy targets to avoid unbounded GPU texture allocation.
|
||||
struct CopyTexturePoolEntry {
|
||||
aurora::gx::GXState::CopyTextureKey key;
|
||||
u32 scaledWidth = 0;
|
||||
@@ -175,7 +169,7 @@ struct CopyTexturePoolEntry {
|
||||
aurora::gfx::TextureHandle handle;
|
||||
};
|
||||
std::vector<CopyTexturePoolEntry> g_copyTexturePool;
|
||||
// Keep only a few reusable copy targets per destination.
|
||||
// Keep only a few reusable copy targets per destination.
|
||||
constexpr size_t kCopyTexturePoolPerKey = 3;
|
||||
|
||||
aurora::gfx::TextureHandle acquire_copy_texture(const aurora::gx::GXState::CopyTextureKey& key, u32 width, u32 height,
|
||||
@@ -337,13 +331,9 @@ void GXSetTexCopyDst(u16 wd, u16 ht, GXTexFmt fmt, GXBool mipmap) {
|
||||
g_gxState.texCopyHalfScale = mipmap != GX_FALSE;
|
||||
}
|
||||
|
||||
void GXSetDispCopyFrame2Field(u32 mode) {
|
||||
g_gxState.dispCopyFrame2Field = mode & 3;
|
||||
}
|
||||
void GXSetDispCopyFrame2Field(u32 mode) { g_gxState.dispCopyFrame2Field = mode & 3; }
|
||||
|
||||
void GXSetCopyClamp(GXFBClamp clamp) {
|
||||
g_gxState.copyClamp = static_cast<GXFBClamp>(static_cast<u32>(clamp) & 3);
|
||||
}
|
||||
void GXSetCopyClamp(GXFBClamp clamp) { g_gxState.copyClamp = static_cast<GXFBClamp>(static_cast<u32>(clamp) & 3); }
|
||||
|
||||
u32 GXSetDispCopyYScale(f32 vscale) {
|
||||
const u32 iScale = y_scale_to_integer(vscale);
|
||||
@@ -411,8 +401,8 @@ void GXSetCopyFilter(GXBool aa, u8 sample_pattern[12][2], GXBool vf, u8 vfilter[
|
||||
|
||||
void GXSetDispCopyGamma(GXGamma gamma) {
|
||||
g_gxState.dispCopyGamma = static_cast<GXGamma>(static_cast<u32>(gamma) & 3u);
|
||||
g_gxState.bpRegCache[0x52] = (g_gxState.bpRegCache[0x52] & ~(3u << 7)) |
|
||||
((static_cast<u32>(g_gxState.dispCopyGamma) & 3u) << 7);
|
||||
g_gxState.bpRegCache[0x52] =
|
||||
(g_gxState.bpRegCache[0x52] & ~(3u << 7)) | ((static_cast<u32>(g_gxState.dispCopyGamma) & 3u) << 7);
|
||||
}
|
||||
|
||||
void GXCopyDisp(void* dest, GXBool clear) {
|
||||
@@ -422,10 +412,11 @@ void GXCopyDisp(void* dest, GXBool clear) {
|
||||
aurora::gx::fifo::drain();
|
||||
}
|
||||
const auto rect = aurora::gx::map_logical_scissor(g_gxState.dispCopySrc);
|
||||
const auto logicalDstWidth =
|
||||
std::max<u32>(g_gxState.dispCopyDstWidth != 0 ? g_gxState.dispCopyDstWidth : static_cast<u32>(g_gxState.dispCopySrc.width), 1);
|
||||
const auto logicalDstHeight =
|
||||
std::max<u32>(g_gxState.dispCopyDstHeight != 0 ? g_gxState.dispCopyDstHeight : static_cast<u32>(g_gxState.dispCopySrc.height), 1);
|
||||
const auto logicalDstWidth = std::max<u32>(
|
||||
g_gxState.dispCopyDstWidth != 0 ? g_gxState.dispCopyDstWidth : static_cast<u32>(g_gxState.dispCopySrc.width), 1);
|
||||
const auto logicalDstHeight = std::max<u32>(
|
||||
g_gxState.dispCopyDstHeight != 0 ? g_gxState.dispCopyDstHeight : static_cast<u32>(g_gxState.dispCopySrc.height),
|
||||
1);
|
||||
const auto [dstWidth, dstHeight] = scale_copy_dst(logicalDstWidth, logicalDstHeight);
|
||||
|
||||
if (!g_gxState.displayCopyTexture || g_gxState.displayCopyWidth != dstWidth ||
|
||||
@@ -444,8 +435,7 @@ void GXCopyDisp(void* dest, GXBool clear) {
|
||||
clearState.clearDepth, clearState.clearColorValue, aurora::gx::clear_depth_value(),
|
||||
GX_TF_RGBA8, nullptr, false, ©Filter, false,
|
||||
static_cast<float>(rect.height) / std::max<float>(g_gxState.dispCopySrc.height, 1.0f),
|
||||
(g_gxState.copyClamp & GX_CLAMP_TOP) != 0,
|
||||
(g_gxState.copyClamp & GX_CLAMP_BOTTOM) != 0);
|
||||
(g_gxState.copyClamp & GX_CLAMP_TOP) != 0, (g_gxState.copyClamp & GX_CLAMP_BOTTOM) != 0);
|
||||
aurora::gfx::mark_last_resolve_as_display_copy();
|
||||
aurora::gx::set_display_copy_present_source();
|
||||
}
|
||||
@@ -483,13 +473,14 @@ void GXCopyTex(void* dest, GXBool clear) {
|
||||
auto it = g_gxState.copyTextureCache.find(key);
|
||||
if (it == g_gxState.copyTextureCache.end()) {
|
||||
auto handle = acquire_copy_texture(key, scaledDstWidth, scaledDstHeight, texCopyFmt);
|
||||
it = g_gxState.copyTextureCache.emplace(key, aurora::gx::GXState::CopyTextureRef{.handle = handle, .revision = 0}).first;
|
||||
it = g_gxState.copyTextureCache.emplace(key, aurora::gx::GXState::CopyTextureRef{.handle = handle, .revision = 0})
|
||||
.first;
|
||||
}
|
||||
auto& handle = it->second;
|
||||
const u32 currentFrame = aurora::gfx::current_frame();
|
||||
const bool sampledThisFrame = handle.sampledThisFrame && handle.lastSampledFrame == currentFrame;
|
||||
const bool scaledSizeChanged = !handle.handle || handle.handle->size.width != scaledDstWidth ||
|
||||
handle.handle->size.height != scaledDstHeight;
|
||||
const bool scaledSizeChanged =
|
||||
!handle.handle || handle.handle->size.width != scaledDstWidth || handle.handle->size.height != scaledDstHeight;
|
||||
auto clearState = get_copy_clear_state(clear);
|
||||
if (sampledThisFrame || scaledSizeChanged) {
|
||||
const u32 revision = handle.revision;
|
||||
@@ -524,18 +515,26 @@ void GXCopyTex(void* dest, GXBool clear) {
|
||||
// Skip only recurring color copies so one-shot copies are never lost.
|
||||
const bool producedConsecutively = handle.revision != 0 && currentFrame - handle.lastProducedFrame <= 1;
|
||||
const bool persistentCopy = !aurora::gx::is_depth_format(texCopyFmt) && !producedConsecutively;
|
||||
if (handle.handle) {
|
||||
// Preserve both provenance and production time. Native post-processing
|
||||
// samples a copy immediately, whereas one-shot bakes such as MKW's minimap
|
||||
// are ordinary 2D game content once retained across frames.
|
||||
handle.handle->isEfbCopy = true;
|
||||
handle.handle->lastEfbCopyFrame = currentFrame;
|
||||
}
|
||||
aurora::gfx::resolve_pass(handle.handle, rect, clearState.clearColor, clearState.clearAlpha, clearState.clearDepth,
|
||||
clearState.clearColorValue, aurora::gx::clear_depth_value(), resolveFmt,
|
||||
&sourceRect.sampleRect, g_gxState.texCopyHalfScale, ©Filter, forceOpaqueAlpha,
|
||||
sourceRect.sampleRect.w() / std::max<float>(g_gxState.texCopySrc.height, 1.0f),
|
||||
(g_gxState.copyClamp & GX_CLAMP_TOP) != 0,
|
||||
(g_gxState.copyClamp & GX_CLAMP_BOTTOM) != 0, persistentCopy);
|
||||
(g_gxState.copyClamp & GX_CLAMP_TOP) != 0, (g_gxState.copyClamp & GX_CLAMP_BOTTOM) != 0,
|
||||
persistentCopy);
|
||||
++handle.revision;
|
||||
handle.lastProducedFrame = currentFrame;
|
||||
handle.width = logicalDstWidth;
|
||||
handle.height = logicalDstHeight;
|
||||
handle.format = texCopyFmt;
|
||||
handle.dataSize = GXGetTexBufferSize(static_cast<u16>(logicalDstWidth), static_cast<u16>(logicalDstHeight), texCopyFmt, GX_FALSE, 0);
|
||||
handle.dataSize = GXGetTexBufferSize(static_cast<u16>(logicalDstWidth), static_cast<u16>(logicalDstHeight),
|
||||
texCopyFmt, GX_FALSE, 0);
|
||||
aurora::gx::notify_copy_texture_created();
|
||||
g_gxState.copyTextures[dest] = handle;
|
||||
// Keep the GPU copy and download it only if guest code reads the destination.
|
||||
@@ -592,5 +591,4 @@ f32 GXGetYScaleFactor(u16 efbHeight, u16 xfbHeight) {
|
||||
|
||||
return resultScale;
|
||||
}
|
||||
|
||||
}
|
||||
+291
-160
@@ -206,15 +206,27 @@ static std::atomic_bool g_stereoSkipCopyClears{true};
|
||||
void set_stereo_stop_at_display_copy(bool value) noexcept {
|
||||
g_stereoStopAtDisplayCopy.store(value, std::memory_order_relaxed);
|
||||
}
|
||||
bool get_stereo_stop_at_display_copy() noexcept {
|
||||
return g_stereoStopAtDisplayCopy.load(std::memory_order_relaxed);
|
||||
}
|
||||
bool get_stereo_stop_at_display_copy() noexcept { return g_stereoStopAtDisplayCopy.load(std::memory_order_relaxed); }
|
||||
void set_stereo_skip_copy_clears(bool value) noexcept {
|
||||
g_stereoSkipCopyClears.store(value, std::memory_order_relaxed);
|
||||
}
|
||||
bool get_stereo_skip_copy_clears() noexcept {
|
||||
return g_stereoSkipCopyClears.load(std::memory_order_relaxed);
|
||||
bool get_stereo_skip_copy_clears() noexcept { return g_stereoSkipCopyClears.load(std::memory_order_relaxed); }
|
||||
|
||||
// The fixed virtual screen orthographic draws are placed on during immersive
|
||||
// replay, in game world units. Written from the settings overlay and read by
|
||||
// the frame worker, like the two controls above. A cleared enable, or a size or
|
||||
// distance that is not positive, leaves 2D content on its recorded GX
|
||||
// transforms, which stretches it across the whole eye.
|
||||
static std::atomic_bool g_stereoHudScreenEnabled{false};
|
||||
static std::atomic<float> g_stereoHudScreenWidth{0.f};
|
||||
static std::atomic<float> g_stereoHudScreenDistance{0.f};
|
||||
|
||||
void set_stereo_hud_screen(bool enabled, float width, float distance) noexcept {
|
||||
g_stereoHudScreenWidth.store(width, std::memory_order_relaxed);
|
||||
g_stereoHudScreenDistance.store(distance, std::memory_order_relaxed);
|
||||
g_stereoHudScreenEnabled.store(enabled, std::memory_order_relaxed);
|
||||
}
|
||||
bool get_stereo_hud_screen_enabled() noexcept { return g_stereoHudScreenEnabled.load(std::memory_order_relaxed); }
|
||||
|
||||
// Recycle command storage: discarding passes used to free their command lists too, so each frame
|
||||
// rebuilt hundreds of KB from zero capacity. The passes themselves are cheap to recreate.
|
||||
@@ -346,18 +358,15 @@ struct ResolveSamplingPlan {
|
||||
};
|
||||
|
||||
static ResolveSamplingPlan make_resolve_sampling_plan(const TextureHandle& target, GXTexFmt format,
|
||||
const ClipRect& resolveRect, const Vec4<float>& sourceRect,
|
||||
bool halfScale, bool copyFilterActive,
|
||||
bool forceOpaqueAlpha) noexcept {
|
||||
const ClipRect& resolveRect, const Vec4<float>& sourceRect,
|
||||
bool halfScale, bool copyFilterActive,
|
||||
bool forceOpaqueAlpha) noexcept {
|
||||
const auto differs = [](float lhs, float rhs) { return std::abs(lhs - rhs) > 0.01f; };
|
||||
const auto differsFromDst = [&](uint32_t dst, float src) {
|
||||
return differs(static_cast<float>(dst), src);
|
||||
};
|
||||
const bool sourceMatchesIntegerRect =
|
||||
!differs(sourceRect.x(), static_cast<float>(resolveRect.x)) &&
|
||||
!differs(sourceRect.y(), static_cast<float>(resolveRect.y)) &&
|
||||
!differs(sourceRect.z(), static_cast<float>(resolveRect.width)) &&
|
||||
!differs(sourceRect.w(), static_cast<float>(resolveRect.height));
|
||||
const auto differsFromDst = [&](uint32_t dst, float src) { return differs(static_cast<float>(dst), src); };
|
||||
const bool sourceMatchesIntegerRect = !differs(sourceRect.x(), static_cast<float>(resolveRect.x)) &&
|
||||
!differs(sourceRect.y(), static_cast<float>(resolveRect.y)) &&
|
||||
!differs(sourceRect.z(), static_cast<float>(resolveRect.width)) &&
|
||||
!differs(sourceRect.w(), static_cast<float>(resolveRect.height));
|
||||
const bool dstMatchesSource = target && !differsFromDst(target->size.width, sourceRect.z()) &&
|
||||
!differsFromDst(target->size.height, sourceRect.w());
|
||||
const bool needsShaderSampling =
|
||||
@@ -369,8 +378,7 @@ static ResolveSamplingPlan make_resolve_sampling_plan(const TextureHandle& targe
|
||||
};
|
||||
}
|
||||
|
||||
static ClipRect calculate_resolve_snapshot_rect(const wgpu::Extent3D& targetSize,
|
||||
const Vec4<float>& sourceRect,
|
||||
static ClipRect calculate_resolve_snapshot_rect(const wgpu::Extent3D& targetSize, const Vec4<float>& sourceRect,
|
||||
const ResolveSamplingPlan& samplingPlan,
|
||||
bool copyFilterActive) noexcept {
|
||||
const ClipRect fullTarget{
|
||||
@@ -381,12 +389,10 @@ static ClipRect calculate_resolve_snapshot_rect(const wgpu::Extent3D& targetSize
|
||||
};
|
||||
// Shader-sampled and converted copies can feed effects that read outside the copy rectangle
|
||||
// (MKW's DOF chain), so keep those full-target. Exact copies have a closed, crop-safe footprint.
|
||||
if (samplingPlan.needsConversion ||
|
||||
samplingPlan.needsShaderSampling ||
|
||||
targetSize.width == 0 || targetSize.height == 0 ||
|
||||
!std::isfinite(sourceRect.x()) || !std::isfinite(sourceRect.y()) ||
|
||||
!std::isfinite(sourceRect.z()) || !std::isfinite(sourceRect.w()) ||
|
||||
sourceRect.z() <= 0.0f || sourceRect.w() <= 0.0f) {
|
||||
if (samplingPlan.needsConversion || samplingPlan.needsShaderSampling || targetSize.width == 0 ||
|
||||
targetSize.height == 0 || !std::isfinite(sourceRect.x()) || !std::isfinite(sourceRect.y()) ||
|
||||
!std::isfinite(sourceRect.z()) || !std::isfinite(sourceRect.w()) || sourceRect.z() <= 0.0f ||
|
||||
sourceRect.w() <= 0.0f) {
|
||||
return fullTarget;
|
||||
}
|
||||
|
||||
@@ -398,10 +404,10 @@ static ClipRect calculate_resolve_snapshot_rect(const wgpu::Extent3D& targetSize
|
||||
const int32_t targetHeight = static_cast<int32_t>(targetSize.height);
|
||||
const int32_t left = std::clamp(static_cast<int32_t>(std::floor(sourceRect.x())) - haloX, 0, targetWidth - 1);
|
||||
const int32_t top = std::clamp(static_cast<int32_t>(std::floor(sourceRect.y())) - haloY, 0, targetHeight - 1);
|
||||
const int32_t right = std::clamp(static_cast<int32_t>(std::ceil(sourceRect.x() + sourceRect.z())) + haloX,
|
||||
left + 1, targetWidth);
|
||||
const int32_t bottom = std::clamp(static_cast<int32_t>(std::ceil(sourceRect.y() + sourceRect.w())) + haloY,
|
||||
top + 1, targetHeight);
|
||||
const int32_t right =
|
||||
std::clamp(static_cast<int32_t>(std::ceil(sourceRect.x() + sourceRect.z())) + haloX, left + 1, targetWidth);
|
||||
const int32_t bottom =
|
||||
std::clamp(static_cast<int32_t>(std::ceil(sourceRect.y() + sourceRect.w())) + haloY, top + 1, targetHeight);
|
||||
return {
|
||||
.x = left,
|
||||
.y = top,
|
||||
@@ -528,9 +534,9 @@ PipelineRef pipeline_ref(const clear::PipelineConfig& config) {
|
||||
|
||||
void resolve_pass(TextureHandle texture, ClipRect rect, bool clearColor, bool clearAlpha, bool clearDepth,
|
||||
Vec4<float> clearColorValue, float clearDepthValue, GXTexFmt resolveFormat,
|
||||
const Vec4<float>* sourceRectPixels, bool halfScale,
|
||||
const std::array<u32, 3>* copyFilterCoefficients, bool forceOpaqueAlpha,
|
||||
float copyFilterRowStride, bool clampTop, bool clampBottom, bool persistentCopy) {
|
||||
const Vec4<float>* sourceRectPixels, bool halfScale, const std::array<u32, 3>* copyFilterCoefficients,
|
||||
bool forceOpaqueAlpha, float copyFilterRowStride, bool clampTop, bool clampBottom,
|
||||
bool persistentCopy) {
|
||||
// Resolve current render pass
|
||||
if (!has_current_render_pass()) {
|
||||
Log.warn("Dropping resolve pass without an active render pass");
|
||||
@@ -579,35 +585,31 @@ void resolve_pass(TextureHandle texture, ClipRect rect, bool clearColor, bool cl
|
||||
prevPass.resolveCopyFilterActive = prevPass.resolveCopyFilterCoefficients[0] != 0 ||
|
||||
prevPass.resolveCopyFilterCoefficients[1] != 64 ||
|
||||
prevPass.resolveCopyFilterCoefficients[2] != 0;
|
||||
const auto samplingPlan = make_resolve_sampling_plan(
|
||||
prevPass.resolveTarget, resolveFormat, rect, sourceRect, halfScale,
|
||||
prevPass.resolveCopyFilterActive, forceOpaqueAlpha);
|
||||
const auto samplingPlan = make_resolve_sampling_plan(prevPass.resolveTarget, resolveFormat, rect, sourceRect,
|
||||
halfScale, prevPass.resolveCopyFilterActive, forceOpaqueAlpha);
|
||||
prevPass.resolveNeedsConversion = samplingPlan.needsConversion;
|
||||
prevPass.resolveNeedsShaderSampling = samplingPlan.needsShaderSampling;
|
||||
prevPass.resolveLinearSampling = samplingPlan.usesLinearSampling;
|
||||
prevPass.snapshotColorResolveSource =
|
||||
!gx::is_depth_format(resolveFormat) && (clearColor || clearAlpha || clearDepth);
|
||||
prevPass.snapshotColorResolveSource = !gx::is_depth_format(resolveFormat) && (clearColor || clearAlpha || clearDepth);
|
||||
Vec4<float> uniformSourceRect = sourceRect;
|
||||
float srcW = static_cast<float>(prevPass.targetSize.width);
|
||||
float srcH = static_cast<float>(prevPass.targetSize.height);
|
||||
if (prevPass.snapshotColorResolveSource) {
|
||||
prevPass.resolveSnapshotRect = calculate_resolve_snapshot_rect(
|
||||
prevPass.targetSize, sourceRect, samplingPlan, prevPass.resolveCopyFilterActive);
|
||||
prevPass.resolveSourceSnapshot = acquire_resolve_source_snapshot(
|
||||
static_cast<uint32_t>(prevPass.resolveSnapshotRect.width),
|
||||
static_cast<uint32_t>(prevPass.resolveSnapshotRect.height));
|
||||
uniformSourceRect = {
|
||||
sourceRect.x() - static_cast<float>(prevPass.resolveSnapshotRect.x),
|
||||
sourceRect.y() - static_cast<float>(prevPass.resolveSnapshotRect.y),
|
||||
sourceRect.z(), sourceRect.w()};
|
||||
prevPass.resolveSnapshotRect = calculate_resolve_snapshot_rect(prevPass.targetSize, sourceRect, samplingPlan,
|
||||
prevPass.resolveCopyFilterActive);
|
||||
prevPass.resolveSourceSnapshot =
|
||||
acquire_resolve_source_snapshot(static_cast<uint32_t>(prevPass.resolveSnapshotRect.width),
|
||||
static_cast<uint32_t>(prevPass.resolveSnapshotRect.height));
|
||||
uniformSourceRect = {sourceRect.x() - static_cast<float>(prevPass.resolveSnapshotRect.x),
|
||||
sourceRect.y() - static_cast<float>(prevPass.resolveSnapshotRect.y), sourceRect.z(),
|
||||
sourceRect.w()};
|
||||
srcW = static_cast<float>(prevPass.resolveSourceSnapshot->size.width);
|
||||
srcH = static_cast<float>(prevPass.resolveSourceSnapshot->size.height);
|
||||
}
|
||||
// GX's copy-clamp bits pin every vertical filter tap to the first or last source texel; without
|
||||
// it a filtered copy at internal resolution samples an unrelated EFB row as a visible border.
|
||||
const float clampTopPixels = clampTop ? uniformSourceRect.y() : 0.0f;
|
||||
const float clampBottomPixels =
|
||||
clampBottom ? uniformSourceRect.y() + uniformSourceRect.w() : srcH;
|
||||
const float clampBottomPixels = clampBottom ? uniformSourceRect.y() + uniformSourceRect.w() : srcH;
|
||||
const float clampTopUv = (clampTopPixels + 0.5f) / srcH;
|
||||
const float clampBottomUv = (clampBottomPixels - 0.5f) / srcH;
|
||||
// Push UV transform uniform for tex_copy_conv (crop region in UV space)
|
||||
@@ -1185,8 +1187,7 @@ static StereoDisplaySource stereo_display_source(const std::vector<RenderPass>&
|
||||
// valid display copy—not a union of every copy—is the frame shown to users.
|
||||
for (auto it = passes.rbegin(); it != passes.rend(); ++it) {
|
||||
const auto& pass = *it;
|
||||
if (!pass.efbTarget || !pass.displayCopyResolve || pass.targetSize.width == 0 ||
|
||||
pass.targetSize.height == 0) {
|
||||
if (!pass.efbTarget || !pass.displayCopyResolve || pass.targetSize.width == 0 || pass.targetSize.height == 0) {
|
||||
continue;
|
||||
}
|
||||
const int32_t passIndex = static_cast<int32_t>(std::distance(passes.begin(), it.base())) - 1;
|
||||
@@ -1216,46 +1217,82 @@ static void log_stereo_display_source_region(ClipRect region, bool foundDisplayC
|
||||
if (region.width > 0 && region.height > 0 && (!logged || region != lastLogged)) {
|
||||
logged = true;
|
||||
lastLogged = region;
|
||||
Log.info("Immersive display-copy source region: {}x{} at ({}, {}){}", region.width,
|
||||
region.height, region.x, region.y,
|
||||
foundDisplayCopy ? "" : " (full-EFB fallback)");
|
||||
Log.info("Immersive display-copy source region: {}x{} at ({}, {}){}", region.width, region.height, region.x,
|
||||
region.y, foundDisplayCopy ? "" : " (full-EFB fallback)");
|
||||
}
|
||||
}
|
||||
|
||||
// The virtual screen orthographic draws are placed on, sized from the aspect
|
||||
// ratio the game is currently presenting at: 4:3 while VILockAspectRatio holds
|
||||
// it there, otherwise the mirror window's own aspect, which is what Mario Kart
|
||||
// Wii's dynamic widescreen builds its projections from. Matching it keeps the
|
||||
// HUD unstretched on the screen.
|
||||
static stereo_replay::HudScreen stereo_hud_screen() noexcept {
|
||||
if (!g_stereoHudScreenEnabled.load(std::memory_order_relaxed)) {
|
||||
return {};
|
||||
}
|
||||
const float width = g_stereoHudScreenWidth.load(std::memory_order_relaxed);
|
||||
const float distance = g_stereoHudScreenDistance.load(std::memory_order_relaxed);
|
||||
float aspect = 0.f;
|
||||
if (!window::get_present_aspect_ratio(aspect) || !(aspect > 0.f)) {
|
||||
aspect = 16.f / 9.f;
|
||||
}
|
||||
const float halfWidth = width * 0.5f;
|
||||
return {
|
||||
.halfWidth = halfWidth,
|
||||
.halfHeight = halfWidth / aspect,
|
||||
.distance = distance,
|
||||
};
|
||||
}
|
||||
|
||||
static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame) noexcept {
|
||||
const StereoDisplaySource displaySource = stereo_display_source(g_renderPasses);
|
||||
const ClipRect displayRegion = displaySource.region;
|
||||
// This is the producer-side preparation path; eye replay can query the pure
|
||||
// helper concurrently without touching this diagnostic state.
|
||||
log_stereo_display_source_region(displayRegion, displaySource.foundDisplayCopy);
|
||||
const stereo_replay::HudScreen hudScreen = stereo_hud_screen();
|
||||
// A draw is replayed per eye when it carries the game camera (perspective) or
|
||||
// when it is 2D content the virtual screen is claiming.
|
||||
const auto replayed = [&](const gx::UniformReplayLayout& layout) noexcept {
|
||||
return layout.perspective || (hudScreen.valid() && !layout.nativeEfbEffect);
|
||||
};
|
||||
size_t requiredBytes = 0;
|
||||
size_t efbPassCount = 0;
|
||||
size_t perspectiveDrawCount = 0;
|
||||
size_t replayPerspectiveDrawCount = 0;
|
||||
size_t replayHudScreenDrawCount = 0;
|
||||
for (const auto& pass : g_renderPasses) {
|
||||
if (pass.efbTarget) {
|
||||
++efbPassCount;
|
||||
}
|
||||
for (const auto& command : pass.commands) {
|
||||
if (command.type != CommandType::Draw || command.data.draw.type != ShaderType::GX ||
|
||||
!command.data.draw.gx.uniformReplayLayout.perspective) {
|
||||
!replayed(command.data.draw.gx.uniformReplayLayout)) {
|
||||
continue;
|
||||
}
|
||||
++perspectiveDrawCount;
|
||||
const auto& draw = command.data.draw.gx;
|
||||
const auto& layout = draw.uniformReplayLayout;
|
||||
if (layout.perspective) {
|
||||
++perspectiveDrawCount;
|
||||
}
|
||||
if (!pass.efbTarget) {
|
||||
continue;
|
||||
}
|
||||
++replayPerspectiveDrawCount;
|
||||
const auto& draw = command.data.draw.gx;
|
||||
const auto& layout = draw.uniformReplayLayout;
|
||||
if (layout.perspective) {
|
||||
++replayPerspectiveDrawCount;
|
||||
} else {
|
||||
++replayHudScreenDrawCount;
|
||||
}
|
||||
const size_t projectionEnd = static_cast<size_t>(layout.projectionOffset) + sizeof(Mat4x4<float>);
|
||||
const size_t positionEnd = static_cast<size_t>(layout.positionOffset) +
|
||||
static_cast<size_t>(layout.positionMatrixCount) * sizeof(Mat3x4<float>);
|
||||
const size_t normalEnd = static_cast<size_t>(layout.normalOffset) +
|
||||
static_cast<size_t>(layout.normalMatrixCount) * sizeof(Mat3x4<float>);
|
||||
if (std::max({projectionEnd, positionEnd, normalEnd}) > draw.uniformRange.size) {
|
||||
Log.error("Stereo replay rejected an invalid GX uniform layout (range={}, projection={}, position={}, normal={})",
|
||||
draw.uniformRange.size, projectionEnd, positionEnd, normalEnd);
|
||||
Log.error(
|
||||
"Stereo replay rejected an invalid GX uniform layout (range={}, projection={}, position={}, normal={})",
|
||||
draw.uniformRange.size, projectionEnd, positionEnd, normalEnd);
|
||||
return false;
|
||||
}
|
||||
requiredBytes += static_cast<size_t>(draw.uniformRange.size) * AURORA_STEREO_EYE_COUNT;
|
||||
@@ -1264,69 +1301,118 @@ static bool prepare_stereo_replay_uniforms(const StereoReplayFrame& stereoFrame)
|
||||
static bool replayCoverageLogged = false;
|
||||
if (!replayCoverageLogged) {
|
||||
replayCoverageLogged = true;
|
||||
Log.info("Immersive replay coverage: {} of {} passes target the EFB; {} of {} perspective draws replay",
|
||||
efbPassCount, g_renderPasses.size(), replayPerspectiveDrawCount,
|
||||
perspectiveDrawCount);
|
||||
Log.info(
|
||||
"Immersive replay coverage: {} of {} passes target the EFB; {} of {} perspective draws replay; {} draws on the "
|
||||
"virtual screen",
|
||||
efbPassCount, g_renderPasses.size(), replayPerspectiveDrawCount, perspectiveDrawCount,
|
||||
replayHudScreenDrawCount);
|
||||
}
|
||||
|
||||
// end_batch_impl appends MaxUniformSize bytes after this for safe dynamic-offset reads.
|
||||
if (requiredBytes > UniformBufferSize ||
|
||||
g_uniforms.size() > UniformBufferSize - requiredBytes ||
|
||||
if (requiredBytes > UniformBufferSize || g_uniforms.size() > UniformBufferSize - requiredBytes ||
|
||||
g_uniforms.size() + requiredBytes > UniformBufferSize - gx::MaxUniformSize) {
|
||||
Log.warn("Skipping stereo replay: GX uniform buffer needs {} additional bytes ({} of {} already used)",
|
||||
requiredBytes, g_uniforms.size(), UniformBufferSize);
|
||||
return false;
|
||||
}
|
||||
|
||||
Viewport drawViewport{
|
||||
.left = static_cast<float>(displayRegion.x),
|
||||
.top = static_cast<float>(displayRegion.y),
|
||||
.width = static_cast<float>(displayRegion.width),
|
||||
.height = static_cast<float>(displayRegion.height),
|
||||
.znear = 0.0f,
|
||||
.zfar = 1.0f,
|
||||
};
|
||||
for (auto& pass : g_renderPasses) {
|
||||
if (!pass.efbTarget) {
|
||||
continue;
|
||||
}
|
||||
for (auto& command : pass.commands) {
|
||||
if (command.type == CommandType::SetViewport) {
|
||||
drawViewport = command.data.setViewport;
|
||||
continue;
|
||||
}
|
||||
if (command.type != CommandType::Draw || command.data.draw.type != ShaderType::GX ||
|
||||
!command.data.draw.gx.uniformReplayLayout.perspective) {
|
||||
!replayed(command.data.draw.gx.uniformReplayLayout)) {
|
||||
continue;
|
||||
}
|
||||
auto& draw = command.data.draw.gx;
|
||||
const auto& layout = draw.uniformReplayLayout;
|
||||
Mat4x4<float> gameProjection;
|
||||
std::memcpy(&gameProjection, g_uniforms.data() + draw.uniformRange.offset + layout.projectionOffset,
|
||||
sizeof(gameProjection));
|
||||
// Only a genuinely affine projection carries its NDC position in its clip
|
||||
// position, which is what the virtual screen reprojection consumes. GX
|
||||
// tracks the projection type separately from the matrix, so a 2D draw
|
||||
// whose matrix disagrees keeps its recorded transforms instead of being
|
||||
// folded onto the screen from a shape the composition cannot represent.
|
||||
if (!layout.perspective && !stereo_replay::is_orthographic_projection(gameProjection)) {
|
||||
continue;
|
||||
}
|
||||
for (uint32_t eyeIndex = 0; eyeIndex < AURORA_STEREO_EYE_COUNT; ++eyeIndex) {
|
||||
auto [uniform, range] = copy_uniform(draw.uniformRange);
|
||||
draw.stereoUniformRanges[eyeIndex] = range;
|
||||
const auto& eye = stereoFrame.eyes[eyeIndex];
|
||||
|
||||
Mat4x4<float> gameProjection;
|
||||
std::memcpy(&gameProjection, uniform.data() + layout.projectionOffset,
|
||||
sizeof(gameProjection));
|
||||
const auto projection =
|
||||
stereo_replay::compose_projection(eye.projection, gameProjection);
|
||||
std::memcpy(uniform.data() + layout.projectionOffset, &projection,
|
||||
sizeof(projection));
|
||||
if (layout.perspective) {
|
||||
const auto projection = stereo_replay::compose_projection(eye.projection, gameProjection);
|
||||
std::memcpy(uniform.data() + layout.projectionOffset, &projection, sizeof(projection));
|
||||
|
||||
for (uint32_t matrix = 0; matrix < layout.positionMatrixCount; ++matrix) {
|
||||
if ((layout.positionMatrixMask & (1u << matrix)) == 0) {
|
||||
continue;
|
||||
for (uint32_t matrix = 0; matrix < layout.positionMatrixCount; ++matrix) {
|
||||
if ((layout.positionMatrixMask & (1u << matrix)) == 0) {
|
||||
continue;
|
||||
}
|
||||
const size_t offset = layout.positionOffset + matrix * sizeof(Mat3x4<float>);
|
||||
Mat3x4<float> source;
|
||||
std::memcpy(&source, uniform.data() + offset, sizeof(source));
|
||||
const auto transformed = stereo_replay::compose_affine(eye.viewFromCenter, source);
|
||||
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
|
||||
}
|
||||
const size_t offset = layout.positionOffset + matrix * sizeof(Mat3x4<float>);
|
||||
Mat3x4<float> source;
|
||||
std::memcpy(&source, uniform.data() + offset, sizeof(source));
|
||||
const auto transformed = stereo_replay::compose_affine(eye.viewFromCenter, source);
|
||||
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
|
||||
}
|
||||
for (uint32_t matrix = 0; matrix < layout.normalMatrixCount; ++matrix) {
|
||||
const size_t offset = layout.normalOffset + matrix * sizeof(Mat3x4<float>);
|
||||
Mat3x4<float> source;
|
||||
std::memcpy(&source, uniform.data() + offset, sizeof(source));
|
||||
const auto transformed = stereo_replay::compose_normal(eye.viewFromCenter, source);
|
||||
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
|
||||
for (uint32_t matrix = 0; matrix < layout.normalMatrixCount; ++matrix) {
|
||||
const size_t offset = layout.normalOffset + matrix * sizeof(Mat3x4<float>);
|
||||
Mat3x4<float> source;
|
||||
std::memcpy(&source, uniform.data() + offset, sizeof(source));
|
||||
const auto transformed = stereo_replay::compose_normal(eye.viewFromCenter, source);
|
||||
std::memcpy(uniform.data() + offset, &transformed, sizeof(transformed));
|
||||
}
|
||||
} else {
|
||||
// 2D content reaches the eye entirely through its projection: the
|
||||
// draw's own position matrices lay the element out in screen space.
|
||||
// First lift viewport-local NDC into displayed-frame NDC; replay will
|
||||
// use a full-eye viewport so sub-pane elements are not transformed by
|
||||
// the recorded viewport a second time.
|
||||
const auto ndcRemap = stereo_replay::make_hud_ndc_remap(
|
||||
drawViewport.left, drawViewport.top, drawViewport.width, drawViewport.height,
|
||||
static_cast<float>(displayRegion.x), static_cast<float>(displayRegion.y),
|
||||
static_cast<float>(displayRegion.width), static_cast<float>(displayRegion.height));
|
||||
const auto projection = stereo_replay::compose_hud_screen_projection(
|
||||
eye.projection, eye.viewFromCenter, hudScreen, gameProjection, gx::UseReversedZ, ndcRemap);
|
||||
std::memcpy(uniform.data() + layout.projectionOffset, &projection, sizeof(projection));
|
||||
}
|
||||
|
||||
if (displayRegion.width > 0 && displayRegion.height > 0) {
|
||||
float renderSize[2];
|
||||
float logicalSize[2];
|
||||
std::memcpy(renderSize, uniform.data() + 8, sizeof(renderSize));
|
||||
renderSize[0] *= static_cast<float>(eye.target.size.width) /
|
||||
static_cast<float>(displayRegion.width);
|
||||
renderSize[1] *= static_cast<float>(eye.target.size.height) /
|
||||
static_cast<float>(displayRegion.height);
|
||||
std::memcpy(logicalSize, uniform.data() + 16, sizeof(logicalSize));
|
||||
if (layout.perspective) {
|
||||
renderSize[0] *= static_cast<float>(eye.target.size.width) / static_cast<float>(displayRegion.width);
|
||||
renderSize[1] *= static_cast<float>(eye.target.size.height) / static_cast<float>(displayRegion.height);
|
||||
} else {
|
||||
// Point/line expansion and GX's pixel-center correction now operate
|
||||
// in the full eye viewport. Recover the complete logical frame size
|
||||
// from this draw's logical-to-render scale.
|
||||
if (renderSize[0] != 0.0f) {
|
||||
logicalSize[0] *= static_cast<float>(displayRegion.width) / renderSize[0];
|
||||
}
|
||||
if (renderSize[1] != 0.0f) {
|
||||
logicalSize[1] *= static_cast<float>(displayRegion.height) / renderSize[1];
|
||||
}
|
||||
renderSize[0] = static_cast<float>(eye.target.size.width);
|
||||
renderSize[1] = static_cast<float>(eye.target.size.height);
|
||||
std::memcpy(uniform.data() + 16, logicalSize, sizeof(logicalSize));
|
||||
}
|
||||
std::memcpy(uniform.data() + 8, renderSize, sizeof(renderSize));
|
||||
}
|
||||
}
|
||||
@@ -1430,9 +1516,8 @@ void expire_bind_group_cache() noexcept {
|
||||
// fmt::format for a name nothing reads outside a capture.
|
||||
static const char* render_pass_label(u32 index) noexcept {
|
||||
static constexpr std::array<const char*, 16> kRenderPassLabels{
|
||||
"Render pass 0", "Render pass 1", "Render pass 2", "Render pass 3",
|
||||
"Render pass 4", "Render pass 5", "Render pass 6", "Render pass 7",
|
||||
"Render pass 8", "Render pass 9", "Render pass 10", "Render pass 11",
|
||||
"Render pass 0", "Render pass 1", "Render pass 2", "Render pass 3", "Render pass 4", "Render pass 5",
|
||||
"Render pass 6", "Render pass 7", "Render pass 8", "Render pass 9", "Render pass 10", "Render pass 11",
|
||||
"Render pass 12", "Render pass 13", "Render pass 14", "Render pass 15",
|
||||
};
|
||||
return index < kRenderPassLabels.size() ? kRenderPassLabels[index] : "Render pass";
|
||||
@@ -1536,10 +1621,11 @@ static void render_impl(std::vector<RenderPass>& renderPasses, wgpu::CommandEnco
|
||||
if (passInfo.snapshotColorResolveSource && passInfo.resolveSourceSnapshot) {
|
||||
const wgpu::TexelCopyTextureInfo src{
|
||||
.texture = passInfo.copySourceTexture,
|
||||
.origin = wgpu::Origin3D{
|
||||
.x = static_cast<uint32_t>(passInfo.resolveSnapshotRect.x),
|
||||
.y = static_cast<uint32_t>(passInfo.resolveSnapshotRect.y),
|
||||
},
|
||||
.origin =
|
||||
wgpu::Origin3D{
|
||||
.x = static_cast<uint32_t>(passInfo.resolveSnapshotRect.x),
|
||||
.y = static_cast<uint32_t>(passInfo.resolveSnapshotRect.y),
|
||||
},
|
||||
};
|
||||
const wgpu::TexelCopyTextureInfo dst{
|
||||
.texture = passInfo.resolveSourceSnapshot->texture,
|
||||
@@ -1561,9 +1647,8 @@ static void render_impl(std::vector<RenderPass>& renderPasses, wgpu::CommandEnco
|
||||
.srcView = resolveSourceView,
|
||||
.uniformRange = passInfo.resolveUniformRange,
|
||||
.dst = passInfo.resolveTarget,
|
||||
.sampleFilter = passInfo.resolveLinearSampling
|
||||
? tex_copy_conv::SampleFilter::Linear
|
||||
: tex_copy_conv::SampleFilter::Nearest,
|
||||
.sampleFilter = passInfo.resolveLinearSampling ? tex_copy_conv::SampleFilter::Linear
|
||||
: tex_copy_conv::SampleFilter::Nearest,
|
||||
.forceOpaqueAlpha = passInfo.resolveForceOpaqueAlpha,
|
||||
};
|
||||
if (passInfo.resolveNeedsConversion) {
|
||||
@@ -1575,12 +1660,12 @@ static void render_impl(std::vector<RenderPass>& renderPasses, wgpu::CommandEnco
|
||||
.texture = resolveSourceTexture,
|
||||
.origin =
|
||||
wgpu::Origin3D{
|
||||
.x = static_cast<uint32_t>(passInfo.resolveRect.x -
|
||||
(passInfo.snapshotColorResolveSource
|
||||
? passInfo.resolveSnapshotRect.x : 0)),
|
||||
.y = static_cast<uint32_t>(passInfo.resolveRect.y -
|
||||
(passInfo.snapshotColorResolveSource
|
||||
? passInfo.resolveSnapshotRect.y : 0)),
|
||||
.x = static_cast<uint32_t>(passInfo.resolveRect.x - (passInfo.snapshotColorResolveSource
|
||||
? passInfo.resolveSnapshotRect.x
|
||||
: 0)),
|
||||
.y = static_cast<uint32_t>(passInfo.resolveRect.y - (passInfo.snapshotColorResolveSource
|
||||
? passInfo.resolveSnapshotRect.y
|
||||
: 0)),
|
||||
},
|
||||
};
|
||||
const wgpu::TexelCopyTextureInfo dst{
|
||||
@@ -1635,8 +1720,8 @@ void render(SealedFrame& frame, wgpu::CommandEncoder& cmd, int32_t interpolatedF
|
||||
});
|
||||
}
|
||||
|
||||
void render_stereo_eye(SealedFrame& frame, wgpu::CommandEncoder& cmd,
|
||||
const StereoReplayFrame& stereoFrame, uint32_t eye, bool finalize) {
|
||||
void render_stereo_eye(SealedFrame& frame, wgpu::CommandEncoder& cmd, const StereoReplayFrame& stereoFrame,
|
||||
uint32_t eye, bool finalize) {
|
||||
CHECK(eye < AURORA_STEREO_EYE_COUNT, "invalid stereo eye {}", eye);
|
||||
const auto displaySource = stereo_display_source(frame.data().passes);
|
||||
// GXCopyDisp publishes the frame and then clears the EFB for the next one.
|
||||
@@ -1703,8 +1788,7 @@ static void render_pass_impl(const wgpu::RenderPassEncoder& pass, const std::vec
|
||||
int32_t sourceRegionTop = 0;
|
||||
int32_t sourceRegionRight = sourceWidth;
|
||||
int32_t sourceRegionBottom = sourceHeight;
|
||||
if (overrideTarget && invocation.replaySourceRegion.width > 0 &&
|
||||
invocation.replaySourceRegion.height > 0) {
|
||||
if (overrideTarget && invocation.replaySourceRegion.width > 0 && invocation.replaySourceRegion.height > 0) {
|
||||
const auto& region = invocation.replaySourceRegion;
|
||||
sourceRegionLeft = std::clamp(region.x, 0, sourceWidth);
|
||||
sourceRegionTop = std::clamp(region.y, 0, sourceHeight);
|
||||
@@ -1726,6 +1810,32 @@ static void render_pass_impl(const wgpu::RenderPassEncoder& pass, const std::vec
|
||||
? static_cast<float>(targetSize.height) / static_cast<float>(sourceRegionHeight)
|
||||
: 1.0f;
|
||||
|
||||
// WebGPU starts a pass scissored to the whole attachment, which is also what a
|
||||
// virtual-screen draw wants.
|
||||
const std::array<uint32_t, 4> fullTargetScissor{0, 0, targetSize.width, targetSize.height};
|
||||
std::array<uint32_t, 4> recordedScissor = fullTargetScissor;
|
||||
bool hudScreenScissor = false;
|
||||
bool scissorStateKnown = true;
|
||||
const auto apply_scissor = [&pass](const std::array<uint32_t, 4>& rect) noexcept {
|
||||
pass.SetScissorRect(rect[0], rect[1], rect[2], rect[3]);
|
||||
};
|
||||
Viewport recordedViewport{
|
||||
.left = 0.0f,
|
||||
.top = 0.0f,
|
||||
.width = static_cast<float>(targetSize.width),
|
||||
.height = static_cast<float>(targetSize.height),
|
||||
.znear = 0.0f,
|
||||
.zfar = 1.0f,
|
||||
};
|
||||
bool hudScreenViewport = false;
|
||||
bool viewportStateKnown = true;
|
||||
const auto apply_viewport = [&pass, &recordedViewport, &targetSize](bool hudScreen) noexcept {
|
||||
pass.SetViewport(hudScreen ? 0.0f : recordedViewport.left, hudScreen ? 0.0f : recordedViewport.top,
|
||||
hudScreen ? static_cast<float>(targetSize.width) : recordedViewport.width,
|
||||
hudScreen ? static_cast<float>(targetSize.height) : recordedViewport.height,
|
||||
recordedViewport.znear, recordedViewport.zfar);
|
||||
};
|
||||
|
||||
for (const auto& cmd : renderPasses[idx].commands) {
|
||||
#ifdef AURORA_GFX_DEBUG_GROUPS
|
||||
{
|
||||
@@ -1752,52 +1862,79 @@ static void render_pass_impl(const wgpu::RenderPassEncoder& pass, const std::vec
|
||||
// reproduced in clip space. Passing the raw swapped pair diverged per backend in release builds.
|
||||
const float minDepth = std::clamp(std::min(vp.znear, vp.zfar), 0.0f, 1.0f);
|
||||
const float maxDepth = std::clamp(std::max(vp.znear, vp.zfar), 0.0f, 1.0f);
|
||||
pass.SetViewport((vp.left - static_cast<float>(sourceRegionLeft)) * scaleX,
|
||||
(vp.top - static_cast<float>(sourceRegionTop)) * scaleY,
|
||||
vp.width * scaleX, vp.height * scaleY, minDepth, maxDepth);
|
||||
recordedViewport = {
|
||||
.left = (vp.left - static_cast<float>(sourceRegionLeft)) * scaleX,
|
||||
.top = (vp.top - static_cast<float>(sourceRegionTop)) * scaleY,
|
||||
.width = vp.width * scaleX,
|
||||
.height = vp.height * scaleY,
|
||||
.znear = minDepth,
|
||||
.zfar = maxDepth,
|
||||
};
|
||||
hudScreenViewport = false;
|
||||
viewportStateKnown = true;
|
||||
apply_viewport(false);
|
||||
} break;
|
||||
case CommandType::SetScissor: {
|
||||
const auto& sc = cmd.data.setScissor;
|
||||
const auto sourceLeft = std::clamp(sc.x, sourceRegionLeft, sourceRegionRight);
|
||||
const auto sourceTop = std::clamp(sc.y, sourceRegionTop, sourceRegionBottom);
|
||||
const auto sourceRight =
|
||||
std::clamp(sc.x + sc.width, sourceLeft, sourceRegionRight);
|
||||
const auto sourceBottom =
|
||||
std::clamp(sc.y + sc.height, sourceTop, sourceRegionBottom);
|
||||
const auto left = static_cast<uint32_t>(std::clamp(
|
||||
static_cast<int32_t>(std::floor(static_cast<float>(sourceLeft - sourceRegionLeft) * scaleX)), 0,
|
||||
static_cast<int32_t>(targetSize.width)));
|
||||
const auto top = static_cast<uint32_t>(std::clamp(
|
||||
static_cast<int32_t>(std::floor(static_cast<float>(sourceTop - sourceRegionTop) * scaleY)), 0,
|
||||
static_cast<int32_t>(targetSize.height)));
|
||||
const auto right = static_cast<uint32_t>(std::clamp(
|
||||
static_cast<int32_t>(std::ceil(static_cast<float>(sourceRight - sourceRegionLeft) * scaleX)),
|
||||
static_cast<int32_t>(left), static_cast<int32_t>(targetSize.width)));
|
||||
const auto bottom = static_cast<uint32_t>(std::clamp(
|
||||
static_cast<int32_t>(std::ceil(static_cast<float>(sourceBottom - sourceRegionTop) * scaleY)),
|
||||
static_cast<int32_t>(top), static_cast<int32_t>(targetSize.height)));
|
||||
pass.SetScissorRect(left, top, right - left, bottom - top);
|
||||
const auto sourceRight = std::clamp(sc.x + sc.width, sourceLeft, sourceRegionRight);
|
||||
const auto sourceBottom = std::clamp(sc.y + sc.height, sourceTop, sourceRegionBottom);
|
||||
const auto left = static_cast<uint32_t>(
|
||||
std::clamp(static_cast<int32_t>(std::floor(static_cast<float>(sourceLeft - sourceRegionLeft) * scaleX)), 0,
|
||||
static_cast<int32_t>(targetSize.width)));
|
||||
const auto top = static_cast<uint32_t>(
|
||||
std::clamp(static_cast<int32_t>(std::floor(static_cast<float>(sourceTop - sourceRegionTop) * scaleY)), 0,
|
||||
static_cast<int32_t>(targetSize.height)));
|
||||
const auto right = static_cast<uint32_t>(
|
||||
std::clamp(static_cast<int32_t>(std::ceil(static_cast<float>(sourceRight - sourceRegionLeft) * scaleX)),
|
||||
static_cast<int32_t>(left), static_cast<int32_t>(targetSize.width)));
|
||||
const auto bottom = static_cast<uint32_t>(
|
||||
std::clamp(static_cast<int32_t>(std::ceil(static_cast<float>(sourceBottom - sourceRegionTop) * scaleY)),
|
||||
static_cast<int32_t>(top), static_cast<int32_t>(targetSize.height)));
|
||||
recordedScissor = {left, top, right - left, bottom - top};
|
||||
hudScreenScissor = false;
|
||||
scissorStateKnown = true;
|
||||
apply_scissor(recordedScissor);
|
||||
} break;
|
||||
case CommandType::Draw: {
|
||||
const auto& draw = cmd.data.draw;
|
||||
switch (draw.type) {
|
||||
case ShaderType::GX: {
|
||||
const gfx::Range* uniformOverride = nullptr;
|
||||
// Only a 2D draw the virtual screen actually claimed carries a stereo
|
||||
// uniform range without being perspective.
|
||||
bool virtualScreenDraw = false;
|
||||
if (invocation.stereoEye < draw.gx.stereoUniformRanges.size() &&
|
||||
draw.gx.stereoUniformRanges[invocation.stereoEye].size != 0) {
|
||||
uniformOverride = &draw.gx.stereoUniformRanges[invocation.stereoEye];
|
||||
virtualScreenDraw = !draw.gx.uniformReplayLayout.perspective;
|
||||
} else if (invocation.interpolatedFrame >= 0 &&
|
||||
static_cast<size_t>(invocation.interpolatedFrame) <
|
||||
draw.gx.interpolatedUniformRanges.size() &&
|
||||
static_cast<size_t>(invocation.interpolatedFrame) < draw.gx.interpolatedUniformRanges.size() &&
|
||||
draw.gx.interpolatedUniformRanges[invocation.interpolatedFrame].size != 0) {
|
||||
uniformOverride = &draw.gx.interpolatedUniformRanges[invocation.interpolatedFrame];
|
||||
}
|
||||
gx::render(draw.gx, pass, encodeState, renderPasses[idx].requireReadyPipelines, uniformOverride);
|
||||
// Such a draw no longer lands where the game aimed it, while the
|
||||
// recorded scissor still describes the rectangle it occupied on the flat
|
||||
// frame (Mario Kart clips the item roulette that way). Honouring that
|
||||
// rectangle would cut the reprojected element away, so it gets the whole
|
||||
// eye and every other draw gets the game's own rectangle back.
|
||||
if (!scissorStateKnown || virtualScreenDraw != hudScreenScissor) {
|
||||
hudScreenScissor = virtualScreenDraw;
|
||||
scissorStateKnown = true;
|
||||
apply_scissor(virtualScreenDraw ? fullTargetScissor : recordedScissor);
|
||||
}
|
||||
if (!viewportStateKnown || virtualScreenDraw != hudScreenViewport) {
|
||||
hudScreenViewport = virtualScreenDraw;
|
||||
viewportStateKnown = true;
|
||||
apply_viewport(virtualScreenDraw);
|
||||
}
|
||||
gx::render(draw.gx, pass, encodeState, renderPasses[idx].requireReadyPipelines, uniformOverride,
|
||||
virtualScreenDraw ? draw.gx.exactScreenDepthPipeline : 0);
|
||||
} break;
|
||||
case ShaderType::Clear: {
|
||||
auto clearDraw = draw.clear;
|
||||
if (invocation.skipCopyClears && overrideTarget && clearDraw.copyClear &&
|
||||
renderPasses[idx].postCopyClear) {
|
||||
if (invocation.skipCopyClears && overrideTarget && clearDraw.copyClear && renderPasses[idx].postCopyClear) {
|
||||
// The scissored twin of the attachment-load-op case above: the copy's
|
||||
// EFB reset, rescaled into eye space, covers the whole eye.
|
||||
break;
|
||||
@@ -1806,30 +1943,20 @@ static void render_pass_impl(const wgpu::RenderPassEncoder& pass, const std::vec
|
||||
const auto& sc = clearDraw.scissor;
|
||||
const auto sourceLeft = std::clamp(sc.x, sourceRegionLeft, sourceRegionRight);
|
||||
const auto sourceTop = std::clamp(sc.y, sourceRegionTop, sourceRegionBottom);
|
||||
const auto sourceRight =
|
||||
std::clamp(sc.x + sc.width, sourceLeft, sourceRegionRight);
|
||||
const auto sourceBottom =
|
||||
std::clamp(sc.y + sc.height, sourceTop, sourceRegionBottom);
|
||||
const auto left = std::clamp(
|
||||
static_cast<int32_t>(
|
||||
std::floor(static_cast<float>(sourceLeft - sourceRegionLeft) * scaleX)),
|
||||
0,
|
||||
static_cast<int32_t>(targetSize.width));
|
||||
const auto top = std::clamp(
|
||||
static_cast<int32_t>(
|
||||
std::floor(static_cast<float>(sourceTop - sourceRegionTop) * scaleY)),
|
||||
0,
|
||||
static_cast<int32_t>(targetSize.height));
|
||||
const auto right = std::clamp(
|
||||
static_cast<int32_t>(
|
||||
std::ceil(static_cast<float>(sourceRight - sourceRegionLeft) * scaleX)),
|
||||
left,
|
||||
static_cast<int32_t>(targetSize.width));
|
||||
const auto bottom = std::clamp(
|
||||
static_cast<int32_t>(
|
||||
std::ceil(static_cast<float>(sourceBottom - sourceRegionTop) * scaleY)),
|
||||
top,
|
||||
static_cast<int32_t>(targetSize.height));
|
||||
const auto sourceRight = std::clamp(sc.x + sc.width, sourceLeft, sourceRegionRight);
|
||||
const auto sourceBottom = std::clamp(sc.y + sc.height, sourceTop, sourceRegionBottom);
|
||||
const auto left =
|
||||
std::clamp(static_cast<int32_t>(std::floor(static_cast<float>(sourceLeft - sourceRegionLeft) * scaleX)),
|
||||
0, static_cast<int32_t>(targetSize.width));
|
||||
const auto top =
|
||||
std::clamp(static_cast<int32_t>(std::floor(static_cast<float>(sourceTop - sourceRegionTop) * scaleY)), 0,
|
||||
static_cast<int32_t>(targetSize.height));
|
||||
const auto right =
|
||||
std::clamp(static_cast<int32_t>(std::ceil(static_cast<float>(sourceRight - sourceRegionLeft) * scaleX)),
|
||||
left, static_cast<int32_t>(targetSize.width));
|
||||
const auto bottom =
|
||||
std::clamp(static_cast<int32_t>(std::ceil(static_cast<float>(sourceBottom - sourceRegionTop) * scaleY)),
|
||||
top, static_cast<int32_t>(targetSize.height));
|
||||
clearDraw.scissor = ClipRect{
|
||||
.x = left,
|
||||
.y = top,
|
||||
@@ -1840,6 +1967,10 @@ static void render_pass_impl(const wgpu::RenderPassEncoder& pass, const std::vec
|
||||
// Clear draws set their own viewport and scissor. Stereo replay must use
|
||||
// the eye attachment extent here, not the original EFB/desktop extent.
|
||||
clear::render(clearDraw, pass, targetSize, encodeState.currentPipeline);
|
||||
// The clear helper mutates both pieces of dynamic state without a
|
||||
// matching recorded command; force the next GX draw to restore them.
|
||||
viewportStateKnown = false;
|
||||
scissorStateKnown = false;
|
||||
} break;
|
||||
}
|
||||
} break;
|
||||
|
||||
@@ -374,6 +374,12 @@ bool get_stereo_stop_at_display_copy() noexcept;
|
||||
void set_stereo_skip_copy_clears(bool value) noexcept;
|
||||
bool get_stereo_skip_copy_clears() noexcept;
|
||||
|
||||
// Places orthographic draws on a fixed virtual screen during immersive replay.
|
||||
// `width` and `distance` are in game world units; the screen's height follows
|
||||
// the game's presented aspect ratio. Also live.
|
||||
void set_stereo_hud_screen(bool enabled, float width, float distance) noexcept;
|
||||
bool get_stereo_hud_screen_enabled() noexcept;
|
||||
|
||||
void begin_offscreen(uint32_t width, uint32_t height);
|
||||
void end_offscreen();
|
||||
bool is_offscreen() noexcept;
|
||||
|
||||
@@ -10,8 +10,7 @@ namespace aurora::gfx::stereo_replay {
|
||||
// would pair an unrelated depth range with the original pipeline compare and
|
||||
// clear state, which can reject the entire eye. Replace only the four
|
||||
// perspective-frustum coefficients and preserve every depth-related element.
|
||||
inline Mat4x4<float> compose_projection(const Mat4x4<float>& eyeFrustum,
|
||||
const Mat4x4<float>& gameProjection) noexcept {
|
||||
inline Mat4x4<float> compose_projection(const Mat4x4<float>& eyeFrustum, const Mat4x4<float>& gameProjection) noexcept {
|
||||
Mat4x4<float> out = gameProjection;
|
||||
out.m0[0] = eyeFrustum.m0[0];
|
||||
out.m0[2] = eyeFrustum.m0[2];
|
||||
@@ -24,8 +23,7 @@ inline Mat4x4<float> compose_projection(const Mat4x4<float>& eyeFrustum,
|
||||
// them as vec4 * mat3x4, which is equivalent to the original column-vector
|
||||
// affine transform. Applying an eye-space delta therefore composes delta *
|
||||
// objectToCenter in the ordinary row-major notation used below.
|
||||
inline Mat3x4<float> compose_affine(const Mat3x4<float>& viewFromCenter,
|
||||
const Mat3x4<float>& objectToCenter) noexcept {
|
||||
inline Mat3x4<float> compose_affine(const Mat3x4<float>& viewFromCenter, const Mat3x4<float>& objectToCenter) noexcept {
|
||||
Mat3x4<float> out{};
|
||||
for (size_t row = 0; row < 3; ++row) {
|
||||
auto& dst = *(&out.m0 + row);
|
||||
@@ -34,16 +32,14 @@ inline Mat3x4<float> compose_affine(const Mat3x4<float>& viewFromCenter,
|
||||
dst[column] = view[0] * objectToCenter.m0[column] + view[1] * objectToCenter.m1[column] +
|
||||
view[2] * objectToCenter.m2[column];
|
||||
}
|
||||
dst[3] = view[3] + view[0] * objectToCenter.m0[3] + view[1] * objectToCenter.m1[3] +
|
||||
view[2] * objectToCenter.m2[3];
|
||||
dst[3] = view[3] + view[0] * objectToCenter.m0[3] + view[1] * objectToCenter.m1[3] + view[2] * objectToCenter.m2[3];
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// Normals receive only the eye transform's linear part. OpenXR view deltas
|
||||
// are rigid transforms, so no inverse-transpose correction is needed here.
|
||||
inline Mat3x4<float> compose_normal(const Mat3x4<float>& viewFromCenter,
|
||||
const Mat3x4<float>& objectToCenter) noexcept {
|
||||
inline Mat3x4<float> compose_normal(const Mat3x4<float>& viewFromCenter, const Mat3x4<float>& objectToCenter) noexcept {
|
||||
Mat3x4<float> out{};
|
||||
for (size_t row = 0; row < 3; ++row) {
|
||||
auto& dst = *(&out.m0 + row);
|
||||
@@ -57,4 +53,130 @@ inline Mat3x4<float> compose_normal(const Mat3x4<float>& viewFromCenter,
|
||||
return out;
|
||||
}
|
||||
|
||||
// A fixed virtual screen for the game's 2D content, sized and placed in the
|
||||
// recorded center-eye view space: a rectangle `distance` units straight ahead
|
||||
// of the game camera, `halfWidth` by `halfHeight` units across. It stays where
|
||||
// the camera puts it, so turning the head looks around it rather than dragging
|
||||
// it along.
|
||||
struct HudScreen {
|
||||
float halfWidth = 0.0f;
|
||||
float halfHeight = 0.0f;
|
||||
float distance = 0.0f;
|
||||
|
||||
[[nodiscard]] bool valid() const noexcept { return halfWidth > 0.0f && halfHeight > 0.0f && distance > 0.0f; }
|
||||
};
|
||||
|
||||
// Converts a draw's viewport-local NDC into the NDC of the complete displayed
|
||||
// frame. It is identity for a full-frame viewport. Virtual-screen replay uses a
|
||||
// full-eye host viewport, so this keeps sub-pane HUD elements in their original
|
||||
// part of the 2D screen instead of applying their viewport twice.
|
||||
struct HudNdcRemap {
|
||||
float scaleX = 1.0f;
|
||||
float scaleY = 1.0f;
|
||||
float offsetX = 0.0f;
|
||||
float offsetY = 0.0f;
|
||||
};
|
||||
|
||||
inline HudNdcRemap make_hud_ndc_remap(float viewportLeft, float viewportTop, float viewportWidth, float viewportHeight,
|
||||
float frameLeft, float frameTop, float frameWidth, float frameHeight) noexcept {
|
||||
if (!(frameWidth > 0.0f) || !(frameHeight > 0.0f)) {
|
||||
return {};
|
||||
}
|
||||
return {
|
||||
.scaleX = viewportWidth / frameWidth,
|
||||
.scaleY = viewportHeight / frameHeight,
|
||||
.offsetX = (2.0f * (viewportLeft - frameLeft) + viewportWidth) / frameWidth - 1.0f,
|
||||
.offsetY = 1.0f - (2.0f * (viewportTop - frameTop) + viewportHeight) / frameHeight,
|
||||
};
|
||||
}
|
||||
|
||||
inline Mat4x4<float> remap_hud_ndc(const Mat4x4<float>& projection, const HudNdcRemap& remap) noexcept {
|
||||
Mat4x4<float> out = projection;
|
||||
for (size_t i = 0; i < 4; ++i) {
|
||||
out.m0[i] = projection.m0[i] * remap.scaleX + projection.m3[i] * remap.offsetX;
|
||||
out.m1[i] = projection.m1[i] * remap.scaleY + projection.m3[i] * remap.offsetY;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// A GX orthographic projection is affine: apply_xf_projection writes exactly
|
||||
// (0, 0, 0, 1) into its w row, and the renderer's depth-window flip only ever
|
||||
// touches the z row. An orthographic draw's clip position is therefore already
|
||||
// its NDC position, which is what compose_hud_screen_projection relies on.
|
||||
inline bool is_orthographic_projection(const Mat4x4<float>& projection) noexcept {
|
||||
return projection.m3[0] == 0.0f && projection.m3[1] == 0.0f && projection.m3[2] == 0.0f && projection.m3[3] == 1.0f;
|
||||
}
|
||||
|
||||
// The stored GX projection has not yet passed through Aurora's final clip-depth
|
||||
// conversion. Turn its Z row into the 0..1 backend NDC value that the original
|
||||
// orthographic draw would have produced. The virtual-screen shader captures
|
||||
// this row before replacing raster depth with a stable midrange value.
|
||||
inline Vec4<float> backend_ndc_depth_row(const Mat4x4<float>& projection, bool reversedDepth) noexcept {
|
||||
Vec4<float> row{};
|
||||
for (size_t i = 0; i < 4; ++i) {
|
||||
row[i] = reversedDepth ? -projection.m2[i] : projection.m2[i] + projection.m3[i];
|
||||
}
|
||||
return row;
|
||||
}
|
||||
|
||||
// Replaces an orthographic draw's projection so its 2D output lands on the
|
||||
// fixed virtual screen instead of being stretched across the whole eye.
|
||||
//
|
||||
// The GX vertex shader computes `vec4(mv_pos, 1) * proj`, reading m0..m3 as the
|
||||
// x/y/z/w rows of that product, so for an orthographic draw m0 and m1 already
|
||||
// yield the game's NDC x/y and m2 its NDC depth. This composes three more steps
|
||||
// into the same matrix:
|
||||
//
|
||||
// 1. NDC to a point on the screen rectangle in the recorded center-eye view
|
||||
// space: (ndc.x * halfWidth, ndc.y * halfHeight, -distance).
|
||||
// 2. That space into this eye's view space, through viewFromCenter.
|
||||
// 3. Eye view space into clip space, through the OpenXR frustum's four terms.
|
||||
//
|
||||
// Each step is affine in the vertex position, so the whole chain collapses into
|
||||
// one projection matrix and the draw's own position matrices stay untouched.
|
||||
//
|
||||
// The composed Z row carries the original flat-screen NDC depth. The exact-depth
|
||||
// vertex variant captures it, then parks clip depth in the middle of the volume
|
||||
// for stable rasterization; the fragment variant exports the captured value.
|
||||
// Keeping original depth out of the VR perspective divide is what makes
|
||||
// equal-depth 2D layers deterministic under head rotation and translation.
|
||||
inline Mat4x4<float> compose_hud_screen_projection(const Mat4x4<float>& eyeFrustum, const Mat3x4<float>& viewFromCenter,
|
||||
const HudScreen& screen, const Mat4x4<float>& gameProjection,
|
||||
bool reversedDepth, const HudNdcRemap& ndcRemap = {}) noexcept {
|
||||
const Mat4x4<float> frameProjection = remap_hud_ndc(gameProjection, ndcRemap);
|
||||
// The screen point's three coordinates, each as a functional of (mv_pos, 1).
|
||||
Mat3x4<float> screenPoint{};
|
||||
for (size_t i = 0; i < 4; ++i) {
|
||||
screenPoint.m0[i] = frameProjection.m0[i] * screen.halfWidth;
|
||||
screenPoint.m1[i] = frameProjection.m1[i] * screen.halfHeight;
|
||||
screenPoint.m2[i] = 0.0f;
|
||||
}
|
||||
screenPoint.m2[3] = -screen.distance;
|
||||
|
||||
// The same functionals carried into eye view space. viewFromCenter's own
|
||||
// translation column joins the constant term, the one place the implicit 1 of
|
||||
// the homogeneous screen point contributes.
|
||||
Mat3x4<float> eyePoint{};
|
||||
for (size_t row = 0; row < 3; ++row) {
|
||||
auto& dst = *(&eyePoint.m0 + row);
|
||||
const auto& view = *(&viewFromCenter.m0 + row);
|
||||
for (size_t i = 0; i < 4; ++i) {
|
||||
dst[i] = view[0] * screenPoint.m0[i] + view[1] * screenPoint.m1[i] + view[2] * screenPoint.m2[i];
|
||||
}
|
||||
dst[3] += view[3];
|
||||
}
|
||||
|
||||
const Vec4<float> exactDepthRow = backend_ndc_depth_row(gameProjection, reversedDepth);
|
||||
Mat4x4<float> out{};
|
||||
for (size_t i = 0; i < 4; ++i) {
|
||||
out.m0[i] = eyeFrustum.m0[0] * eyePoint.m0[i] + eyeFrustum.m0[2] * eyePoint.m2[i];
|
||||
out.m1[i] = eyeFrustum.m1[1] * eyePoint.m1[i] + eyeFrustum.m1[2] * eyePoint.m2[i];
|
||||
out.m3[i] = -eyePoint.m2[i];
|
||||
// The exact-depth shader captures this original flat-screen value before
|
||||
// parking the geometry at 0.5 for rasterization.
|
||||
out.m2[i] = exactDepthRow[i];
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
} // namespace aurora::gfx::stereo_replay
|
||||
@@ -46,6 +46,15 @@ struct TextureRef {
|
||||
u32 gxFormat;
|
||||
bool hasArbitraryMips = false;
|
||||
bool isReplacement = false;
|
||||
// GXCopyTex provenance used by immersive replay. A recently produced copy can
|
||||
// be a native framebuffer effect; an older one is persistent game content
|
||||
// (Mario Kart Wii bakes its minimap this way) and belongs on the 2D screen.
|
||||
bool isEfbCopy = false;
|
||||
uint32_t lastEfbCopyFrame = UINT32_MAX;
|
||||
|
||||
[[nodiscard]] bool is_recent_efb_copy(uint32_t frame) const noexcept {
|
||||
return isEfbCopy && lastEfbCopyFrame != UINT32_MAX && frame - lastEfbCopyFrame <= 1;
|
||||
}
|
||||
|
||||
TextureRef(wgpu::Texture texture, wgpu::TextureView sampleTextureView, wgpu::TextureView attachmentTextureView,
|
||||
wgpu::Extent3D size, wgpu::TextureFormat format, uint32_t mipCount, u32 gxFormat)
|
||||
|
||||
@@ -108,13 +108,13 @@ static constexpr u8 CP_PRIMITIVE_END = 0xBF;
|
||||
|
||||
// Read helpers for big/little endian
|
||||
#if _MSC_VER
|
||||
template<typename T>
|
||||
template <typename T>
|
||||
__forceinline // Yes, this was necessary.
|
||||
inline T unaligned_load(const T* ptr) {
|
||||
inline T unaligned_load(const T* ptr) {
|
||||
return *static_cast<const __unaligned T*>(ptr);
|
||||
}
|
||||
#else
|
||||
template<typename T>
|
||||
template <typename T>
|
||||
inline T unaligned_load(const T* ptr) {
|
||||
T copy;
|
||||
memcpy(©, ptr, sizeof(T));
|
||||
@@ -138,9 +138,7 @@ static inline u32 read_u32(const u8* ptr, bool bigEndian) {
|
||||
return val;
|
||||
}
|
||||
|
||||
static bool is_draw_cmd(u8 cmd) {
|
||||
return cmd >= CP_PRIMITIVE_START && cmd <= CP_PRIMITIVE_END;
|
||||
}
|
||||
static bool is_draw_cmd(u8 cmd) { return cmd >= CP_PRIMITIVE_START && cmd <= CP_PRIMITIVE_END; }
|
||||
|
||||
static GXPrimitive primitive_from_draw_cmd(u8 cmd) {
|
||||
switch (cmd & CP_OPCODE_MASK) {
|
||||
@@ -164,9 +162,6 @@ static GXPrimitive primitive_from_draw_cmd(u8 cmd) {
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
|
||||
|
||||
static u32 bp_get(u32 reg, u32 size, u32 shift);
|
||||
static inline f32 read_f32(const u8* ptr, bool bigEndian);
|
||||
|
||||
@@ -278,11 +273,12 @@ static inline void mark_pipeline_state_dirty() noexcept {
|
||||
// Rejects a malformed XF write in release builds.
|
||||
#define XF_REQUIRE(cond, msg, ...) \
|
||||
do { \
|
||||
if (!(cond)) UNLIKELY { \
|
||||
CHECK(cond, msg, ##__VA_ARGS__); \
|
||||
Log.warn(msg, ##__VA_ARGS__); \
|
||||
return true; \
|
||||
} \
|
||||
if (!(cond)) \
|
||||
UNLIKELY { \
|
||||
CHECK(cond, msg, ##__VA_ARGS__); \
|
||||
Log.warn(msg, ##__VA_ARGS__); \
|
||||
return true; \
|
||||
} \
|
||||
} while (0)
|
||||
|
||||
static bool copy_xf_data(u32 addr, const u8* data, u32 len, bool bigEndian) {
|
||||
@@ -468,7 +464,8 @@ static bool handle_aurora(const u8* data, u32& pos, u32 size, bool bigEndian);
|
||||
|
||||
void process(const u8* data, u32 size, bool bigEndian) {
|
||||
ZoneScoped;
|
||||
// Everything decoded here mutates renderer state (GX state, the recorded command lists and the mapped staging buffers), so take the renderer GPU mutex once for the whole drain rather than once per draw command.
|
||||
// Everything decoded here mutates renderer state (GX state, the recorded command lists and the mapped staging
|
||||
// buffers), so take the renderer GPU mutex once for the whole drain rather than once per draw command.
|
||||
std::lock_guard gpuLock(aurora::renderer_gpu_mutex());
|
||||
u32 pos = 0;
|
||||
|
||||
@@ -522,9 +519,10 @@ void process(const u8* data, u32 size, bool bigEndian) {
|
||||
if (array.data == nullptr || array.stride == 0 || byteOffset + byteCount > array.size) {
|
||||
static u32 invalidIndexedXfLogCount = 0;
|
||||
if (invalidIndexedXfLogCount < 16) {
|
||||
Log.warn("Skipping indexed XF load with invalid source array: array={} idx={} stride={} offset={} bytes={} "
|
||||
"size={} dst=0x{:04X}",
|
||||
arrayType, srcArrayIdx, array.stride, byteOffset, byteCount, array.size, dstAddr);
|
||||
Log.warn(
|
||||
"Skipping indexed XF load with invalid source array: array={} idx={} stride={} offset={} bytes={} "
|
||||
"size={} dst=0x{:04X}",
|
||||
arrayType, srcArrayIdx, array.stride, byteOffset, byteCount, array.size, dstAddr);
|
||||
++invalidIndexedXfLogCount;
|
||||
}
|
||||
break;
|
||||
@@ -635,12 +633,12 @@ static void handle_bp(u32 value, bool bigEndian) {
|
||||
} else {
|
||||
const u32 ssMask = g_gxState.bpRegCache[0xFE];
|
||||
// A preceding 0xFE write is rare; the common path only has to prove the mask is already wide open.
|
||||
if (ssMask != 0x00FFFFFF) UNLIKELY {
|
||||
g_gxState.bpRegCache[0xFE] = 0x00FFFFFF;
|
||||
}
|
||||
if (ssMask != 0x00FFFFFF)
|
||||
UNLIKELY { g_gxState.bpRegCache[0xFE] = 0x00FFFFFF; }
|
||||
const u32 merged = (g_gxState.bpRegCache[regId] & ~ssMask) | (value & ssMask);
|
||||
value = (regId << 24) | (merged & 0x00FFFFFF);
|
||||
if (g_gxState.bpRegCache[regId] == value) return;
|
||||
if (g_gxState.bpRegCache[regId] == value)
|
||||
return;
|
||||
g_gxState.bpRegCache[regId] = value;
|
||||
}
|
||||
// TEV color combiner stages (0xC0, 0xC2, 0xC4, ... 0xDE)
|
||||
@@ -704,7 +702,8 @@ static void handle_bp(u32 value, bool bigEndian) {
|
||||
case 0x00: {
|
||||
g_gxState.numTexGens = bp_get(value, 4, 0);
|
||||
g_gxState.numChans = bp_get(value, 3, 4);
|
||||
// genMode owns the same numTexGens/numChans that XF 0x3F/0x09 decode, so the XF cache can no longer vouch for those two slots.
|
||||
// genMode owns the same numTexGens/numChans that XF 0x3F/0x09 decode, so the XF cache can no longer vouch for those
|
||||
// two slots.
|
||||
g_gxState.invalidateXfReg(0x3F);
|
||||
g_gxState.invalidateXfReg(0x09);
|
||||
g_gxState.numTevStages = bp_get(value, 4, 10) + 1;
|
||||
@@ -1387,7 +1386,8 @@ void reset_cp_register_cache() {
|
||||
}
|
||||
|
||||
static bool cp_register_write_unchanged(u8 addr, u32 value) {
|
||||
if (!cacheable_cp_register(addr)) return false;
|
||||
if (!cacheable_cp_register(addr))
|
||||
return false;
|
||||
|
||||
if (s_cpRegisterCacheValid[addr] && s_cpRegisterCache[addr] == value) {
|
||||
return true;
|
||||
@@ -1399,7 +1399,8 @@ static bool cp_register_write_unchanged(u8 addr, u32 value) {
|
||||
|
||||
// CP register handler - decodes CP register writes and updates g_gxState
|
||||
static void handle_cp(u8 addr, u32 value, bool bigEndian) {
|
||||
if (cp_register_write_unchanged(addr, value)) return;
|
||||
if (cp_register_write_unchanged(addr, value))
|
||||
return;
|
||||
|
||||
switch (addr) {
|
||||
// VCD low (0x50)
|
||||
@@ -1559,8 +1560,10 @@ static void handle_cp(u8 addr, u32 value, bool bigEndian) {
|
||||
|
||||
// XF register handler - decodes XF (transform unit) register writes and updates g_gxState
|
||||
static void handle_xf(const u8* data, u32& pos, u32 size, bool bigEndian) {
|
||||
// These bounds must hold in release too: CHECK() is a no-op under NDEBUG, so relying on it alone let a truncated guest display list read past `data`.
|
||||
if (pos > size || size - pos < 4) UNLIKELY {
|
||||
// These bounds must hold in release too: CHECK() is a no-op under NDEBUG, so relying on it alone let a truncated
|
||||
// guest display list read past `data`.
|
||||
if (pos > size || size - pos < 4)
|
||||
UNLIKELY {
|
||||
CHECK(false, "XF header read overrun");
|
||||
pos = size;
|
||||
return;
|
||||
@@ -1572,7 +1575,8 @@ static void handle_xf(const u8* data, u32& pos, u32 size, bool bigEndian) {
|
||||
u32 addr = header & 0xFFFF;
|
||||
u32 dataBytes = count * 4;
|
||||
// Log.warn(" xf: addr {:04x} count {} dataBytes {} pos {} -> {}", addr, count, dataBytes, pos, pos + dataBytes);
|
||||
if (size - pos < dataBytes) UNLIKELY {
|
||||
if (size - pos < dataBytes)
|
||||
UNLIKELY {
|
||||
CHECK(false, "XF data read overrun: need {} bytes at pos {}", dataBytes, pos);
|
||||
pos = size;
|
||||
return;
|
||||
@@ -1594,9 +1598,12 @@ static void handle_xf(const u8* data, u32& pos, u32 size, bool bigEndian) {
|
||||
// Skip register writes that decode to state we already hold.
|
||||
const bool cacheable = reg < g_gxState.xfRegCache.size();
|
||||
const bool unchanged = cacheable && g_gxState.xfRegMatches(reg, val);
|
||||
if (cacheable) g_gxState.storeXfReg(reg, val);
|
||||
// Viewport (0x1A-0x1F) and projection (0x20-0x26) keep their unconditional apply below; only the banks that already had skip semantics and the TexGen bank drop out here.
|
||||
if (unchanged && (reg <= 0x19 || reg >= 0x3F)) continue;
|
||||
if (cacheable)
|
||||
g_gxState.storeXfReg(reg, val);
|
||||
// Viewport (0x1A-0x1F) and projection (0x20-0x26) keep their unconditional apply below; only the banks that
|
||||
// already had skip semantics and the TexGen bank drop out here.
|
||||
if (unchanged && (reg <= 0x19 || reg >= 0x3F))
|
||||
continue;
|
||||
|
||||
switch (reg) {
|
||||
case 0x00:
|
||||
@@ -1806,20 +1813,20 @@ static void handle_draw_overrun(u8 cmd, u16 vtxCount, u32 vtxSize, u32 totalVtxB
|
||||
Log.warn(" hex dump around truncated draw cmd (pos {}-{}):{}", dumpStart, dumpEnd - 1, hex);
|
||||
const auto fmt = static_cast<GXVtxFmt>(cmd & CP_VAT_MASK);
|
||||
const auto& vtxFmt = g_gxState.vtxFmts[fmt];
|
||||
Log.warn(" truncated draw cmd=0x{:02X} fmt={} vtxCount={} vtxSize={} desc pn={} pos={} nrm={} clr0={} clr1={} "
|
||||
"tex0={} tex1={} tex2={} tex3={} tex4={} tex5={} tex6={} tex7={}",
|
||||
cmd, static_cast<u32>(fmt), vtxCount, vtxSize, g_gxState.vtxDesc[GX_VA_PNMTXIDX],
|
||||
g_gxState.vtxDesc[GX_VA_POS], g_gxState.vtxDesc[GX_VA_NRM], g_gxState.vtxDesc[GX_VA_CLR0],
|
||||
g_gxState.vtxDesc[GX_VA_CLR1], g_gxState.vtxDesc[GX_VA_TEX0], g_gxState.vtxDesc[GX_VA_TEX1],
|
||||
g_gxState.vtxDesc[GX_VA_TEX2], g_gxState.vtxDesc[GX_VA_TEX3], g_gxState.vtxDesc[GX_VA_TEX4],
|
||||
g_gxState.vtxDesc[GX_VA_TEX5], g_gxState.vtxDesc[GX_VA_TEX6], g_gxState.vtxDesc[GX_VA_TEX7]);
|
||||
Log.warn(" fmt {} attrs pos({},{}) nrm({},{}) clr0({},{}) tex0({},{}) tex1({},{})",
|
||||
static_cast<u32>(fmt), static_cast<u32>(vtxFmt.attrs[GX_VA_POS].cnt),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_POS].type), static_cast<u32>(vtxFmt.attrs[GX_VA_NRM].cnt),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_NRM].type), static_cast<u32>(vtxFmt.attrs[GX_VA_CLR0].cnt),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_CLR0].type), static_cast<u32>(vtxFmt.attrs[GX_VA_TEX0].cnt),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_TEX0].type), static_cast<u32>(vtxFmt.attrs[GX_VA_TEX1].cnt),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_TEX1].type));
|
||||
Log.warn(
|
||||
" truncated draw cmd=0x{:02X} fmt={} vtxCount={} vtxSize={} desc pn={} pos={} nrm={} clr0={} clr1={} "
|
||||
"tex0={} tex1={} tex2={} tex3={} tex4={} tex5={} tex6={} tex7={}",
|
||||
cmd, static_cast<u32>(fmt), vtxCount, vtxSize, g_gxState.vtxDesc[GX_VA_PNMTXIDX], g_gxState.vtxDesc[GX_VA_POS],
|
||||
g_gxState.vtxDesc[GX_VA_NRM], g_gxState.vtxDesc[GX_VA_CLR0], g_gxState.vtxDesc[GX_VA_CLR1],
|
||||
g_gxState.vtxDesc[GX_VA_TEX0], g_gxState.vtxDesc[GX_VA_TEX1], g_gxState.vtxDesc[GX_VA_TEX2],
|
||||
g_gxState.vtxDesc[GX_VA_TEX3], g_gxState.vtxDesc[GX_VA_TEX4], g_gxState.vtxDesc[GX_VA_TEX5],
|
||||
g_gxState.vtxDesc[GX_VA_TEX6], g_gxState.vtxDesc[GX_VA_TEX7]);
|
||||
Log.warn(" fmt {} attrs pos({},{}) nrm({},{}) clr0({},{}) tex0({},{}) tex1({},{})", static_cast<u32>(fmt),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_POS].cnt), static_cast<u32>(vtxFmt.attrs[GX_VA_POS].type),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_NRM].cnt), static_cast<u32>(vtxFmt.attrs[GX_VA_NRM].type),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_CLR0].cnt), static_cast<u32>(vtxFmt.attrs[GX_VA_CLR0].type),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_TEX0].cnt), static_cast<u32>(vtxFmt.attrs[GX_VA_TEX0].type),
|
||||
static_cast<u32>(vtxFmt.attrs[GX_VA_TEX1].cnt), static_cast<u32>(vtxFmt.attrs[GX_VA_TEX1].type));
|
||||
Log.warn("stopping FIFO decode at truncated draw: need {} bytes at pos {}, have {}", totalVtxBytes, pos, size);
|
||||
}
|
||||
|
||||
@@ -1852,12 +1859,12 @@ static u32 calculate_last_vtx_size(GXVtxFmt fmt) {
|
||||
return vtxSize;
|
||||
}
|
||||
|
||||
static void handle_draw_unmerged(GXPrimitive prim, GXVtxFmt fmt, u16 vtxCount,
|
||||
gfx::Range vertRange, uint16_t usedPnMtxMask,
|
||||
HashType matrixTopologySignature,
|
||||
HashType geometrySignature, bool interpolationIdentityActive);
|
||||
static void handle_draw_unmerged(GXPrimitive prim, GXVtxFmt fmt, u16 vtxCount, gfx::Range vertRange,
|
||||
uint16_t usedPnMtxMask, HashType matrixTopologySignature, HashType geometrySignature,
|
||||
bool interpolationIdentityActive);
|
||||
|
||||
// The per-draw geometry signature, matrix-usage mask and draw-identity hashes exist purely to feed frame interpolation (build_uniform consumes them only after its `frame_interpolation_fps() == 0` early-out).
|
||||
// The per-draw geometry signature, matrix-usage mask and draw-identity hashes exist purely to feed frame interpolation
|
||||
// (build_uniform consumes them only after its `frame_interpolation_fps() == 0` early-out).
|
||||
static inline bool frame_interpolation_identity_needed() noexcept {
|
||||
return frame_interpolation_active() && g_gxState.projType == GX_PERSPECTIVE;
|
||||
}
|
||||
@@ -1871,8 +1878,7 @@ static uint32_t matrix_index_prefix_size(GXVtxFmt fmt) noexcept {
|
||||
case GX_NONE:
|
||||
break;
|
||||
case GX_DIRECT:
|
||||
size += comp_type_size(attr, vtxFmt.attrs[i].type) *
|
||||
comp_cnt_count(attr, vtxFmt.attrs[i].cnt);
|
||||
size += comp_type_size(attr, vtxFmt.attrs[i].type) * comp_cnt_count(attr, vtxFmt.attrs[i].cnt);
|
||||
break;
|
||||
case GX_INDEX8:
|
||||
++size;
|
||||
@@ -1885,21 +1891,21 @@ static uint32_t matrix_index_prefix_size(GXVtxFmt fmt) noexcept {
|
||||
return size;
|
||||
}
|
||||
|
||||
static HashType draw_geometry_signature(GXVtxFmt fmt, const uint8_t* vertices,
|
||||
uint16_t vtxCount, uint32_t vtxStride) noexcept {
|
||||
static HashType draw_geometry_signature(GXVtxFmt fmt, const uint8_t* vertices, uint16_t vtxCount,
|
||||
uint32_t vtxStride) noexcept {
|
||||
Hasher hasher;
|
||||
hasher.update(vtxCount);
|
||||
hasher.update(vtxStride);
|
||||
|
||||
// Matrix-index bytes select an instance's current XF palette slots, so they are deliberately excluded from mesh identity.
|
||||
// Matrix-index bytes select an instance's current XF palette slots, so they are deliberately excluded from mesh
|
||||
// identity.
|
||||
const uint32_t matrixPrefix = std::min(matrix_index_prefix_size(fmt), vtxStride);
|
||||
if (matrixPrefix == 0) {
|
||||
// Most draws have no direct matrix-index prefix.
|
||||
hasher.update(vertices, static_cast<size_t>(vtxCount) * vtxStride);
|
||||
} else {
|
||||
for (uint16_t vertex = 0; vertex < vtxCount; ++vertex) {
|
||||
hasher.update(vertices + static_cast<size_t>(vertex) * vtxStride + matrixPrefix,
|
||||
vtxStride - matrixPrefix);
|
||||
hasher.update(vertices + static_cast<size_t>(vertex) * vtxStride + matrixPrefix, vtxStride - matrixPrefix);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1924,8 +1930,7 @@ struct PnMtxUsage {
|
||||
};
|
||||
|
||||
// Slot mask only.
|
||||
static uint16_t pn_mtx_mask(const uint8_t* vertices, uint16_t vtxCount,
|
||||
uint32_t vtxStride) noexcept {
|
||||
static uint16_t pn_mtx_mask(const uint8_t* vertices, uint16_t vtxCount, uint32_t vtxStride) noexcept {
|
||||
if (g_gxState.vtxDesc[GX_VA_PNMTXIDX] != GX_DIRECT) {
|
||||
return static_cast<uint16_t>(1u << std::min<uint32_t>(g_gxState.currentPnMtx, MaxPnMtx - 1));
|
||||
}
|
||||
@@ -1939,8 +1944,7 @@ static uint16_t pn_mtx_mask(const uint8_t* vertices, uint16_t vtxCount,
|
||||
return mask;
|
||||
}
|
||||
|
||||
static PnMtxUsage pn_mtx_usage(const uint8_t* vertices, uint16_t vtxCount,
|
||||
uint32_t vtxStride) noexcept {
|
||||
static PnMtxUsage pn_mtx_usage(const uint8_t* vertices, uint16_t vtxCount, uint32_t vtxStride) noexcept {
|
||||
if (g_gxState.vtxDesc[GX_VA_PNMTXIDX] != GX_DIRECT) {
|
||||
const uint32_t matrixIndex = std::min<uint32_t>(g_gxState.currentPnMtx, MaxPnMtx - 1);
|
||||
return {
|
||||
@@ -1968,7 +1972,8 @@ static PnMtxUsage pn_mtx_usage(const uint8_t* vertices, uint16_t vtxCount,
|
||||
};
|
||||
}
|
||||
|
||||
// Whether the most recent unmerged draw recorded an interpolation snapshot, so a draw merged into it knows there is a snapshot to extend.
|
||||
// Whether the most recent unmerged draw recorded an interpolation snapshot, so a draw merged into it knows there is a
|
||||
// snapshot to extend.
|
||||
static bool s_lastDrawRecordedInterpolation = false;
|
||||
|
||||
struct CachedIndexTemplate {
|
||||
@@ -1985,9 +1990,8 @@ static const CachedIndexTemplate& cached_index_template(GXPrimitive prim, u16 vt
|
||||
static std::array<CachedIndexTemplate, CacheSize> cache{};
|
||||
const u32 key = (static_cast<u32>(underlying(prim)) << 16) | vtxCount;
|
||||
auto& entry = cache[(key ^ (key >> 9)) & (CacheSize - 1)];
|
||||
if (entry.valid && entry.prim == prim && entry.vtxCount == vtxCount) LIKELY {
|
||||
return entry;
|
||||
}
|
||||
if (entry.valid && entry.prim == prim && entry.vtxCount == vtxCount)
|
||||
LIKELY { return entry; }
|
||||
|
||||
entry.valid = true;
|
||||
entry.prim = prim;
|
||||
@@ -1998,9 +2002,9 @@ static const CachedIndexTemplate& cached_index_template(GXPrimitive prim, u16 vt
|
||||
|
||||
static IndexBuffer handle_draw_idx_buf;
|
||||
|
||||
static ArrayRef<u16> offset_index_template(const CachedIndexTemplate& indexTemplate,
|
||||
u16 vtxStart) {
|
||||
// Grow-only: resizing down and back up made every merge zero-fill the buffer before the transform immediately overwrote it.
|
||||
static ArrayRef<u16> offset_index_template(const CachedIndexTemplate& indexTemplate, u16 vtxStart) {
|
||||
// Grow-only: resizing down and back up made every merge zero-fill the buffer before the transform immediately
|
||||
// overwrote it.
|
||||
const size_t count = indexTemplate.indices.size();
|
||||
if (handle_draw_idx_buf.size() < count) {
|
||||
handle_draw_idx_buf.resize(count);
|
||||
@@ -2016,7 +2020,8 @@ static ArrayRef<u16> offset_index_template(const CachedIndexTemplate& indexTempl
|
||||
struct CachedPipelineState {
|
||||
gfx::PipelineRef ref = 0;
|
||||
HashType configHash = 0;
|
||||
// Carried here so the draw can be recorded without keeping the PipelineConfig that produced it alive; it is the only field of the config the draw itself still needs.
|
||||
// Carried here so the draw can be recorded without keeping the PipelineConfig that produced it alive; it is the only
|
||||
// field of the config the draw itself still needs.
|
||||
u32 dstAlpha = UINT32_MAX;
|
||||
ShaderInfo shaderInfo{};
|
||||
};
|
||||
@@ -2032,10 +2037,8 @@ static const CachedPipelineState& cached_pipeline_state(const PipelineConfig& co
|
||||
|
||||
const HashType hash = xxh3_hash(config, static_cast<HashType>(gfx::ShaderType::GX));
|
||||
auto& entry = cache[hash & (CacheSize - 1)];
|
||||
if (entry.valid && entry.state.configHash == hash &&
|
||||
std::memcmp(&entry.config, &config, sizeof(config)) == 0) LIKELY {
|
||||
return entry.state;
|
||||
}
|
||||
if (entry.valid && entry.state.configHash == hash && std::memcmp(&entry.config, &config, sizeof(config)) == 0)
|
||||
LIKELY { return entry.state; }
|
||||
|
||||
entry.valid = true;
|
||||
entry.config = config;
|
||||
@@ -2048,7 +2051,9 @@ static const CachedPipelineState& cached_pipeline_state(const PipelineConfig& co
|
||||
return entry.state;
|
||||
}
|
||||
|
||||
// Resolving a pipeline the long way costs a ~2.7KB zero-init, a full populate_pipeline_config, an XXH3 over the whole config and a memcmp against the hash-indexed slot -- roughly 11KB of memory traffic for a result that is almost always identical to the previous draw's.
|
||||
// Resolving a pipeline the long way costs a ~2.7KB zero-init, a full populate_pipeline_config, an XXH3 over the whole
|
||||
// config and a memcmp against the hash-indexed slot -- roughly 11KB of memory traffic for a result that is almost
|
||||
// always identical to the previous draw's.
|
||||
static const CachedPipelineState& resolve_pipeline_state(GXPrimitive prim, GXVtxFmt fmt) {
|
||||
struct Memo {
|
||||
const CachedPipelineState* state = nullptr;
|
||||
@@ -2061,14 +2066,15 @@ static const CachedPipelineState& resolve_pipeline_state(GXPrimitive prim, GXVtx
|
||||
|
||||
const u32 sampleCount = gfx::get_sample_count();
|
||||
const u32 generation = g_gxState.pipelineStateGeneration;
|
||||
if (memo.state != nullptr && memo.generation == generation && memo.sampleCount == sampleCount &&
|
||||
memo.prim == prim && memo.fmt == fmt) LIKELY {
|
||||
return *memo.state;
|
||||
}
|
||||
if (memo.state != nullptr && memo.generation == generation && memo.sampleCount == sampleCount && memo.prim == prim &&
|
||||
memo.fmt == fmt)
|
||||
LIKELY { return *memo.state; }
|
||||
|
||||
PipelineConfig config{};
|
||||
populate_pipeline_config(config, prim, fmt);
|
||||
// cached_pipeline_state hands back a reference into a fixed direct-mapped table, so the address stays valid; the entry it points at can only be rewritten by another call to that function, and every such call goes through this miss path and replaces the memo in the same breath.
|
||||
// cached_pipeline_state hands back a reference into a fixed direct-mapped table, so the address stays valid; the
|
||||
// entry it points at can only be rewritten by another call to that function, and every such call goes through this
|
||||
// miss path and replaces the memo in the same breath.
|
||||
const CachedPipelineState& state = cached_pipeline_state(config);
|
||||
memo = Memo{
|
||||
.state = &state,
|
||||
@@ -2080,34 +2086,64 @@ static const CachedPipelineState& resolve_pipeline_state(GXPrimitive prim, GXVtx
|
||||
return state;
|
||||
}
|
||||
|
||||
bool submit_raw_draw(GXPrimitive prim, GXVtxFmt fmt, const uint8_t* vertices, uint16_t vtxCount,
|
||||
uint32_t vertexBytes) {
|
||||
// Exact Screen Depth changes shader outputs but no GX pipeline state. Resolve a
|
||||
// sibling pipeline only for orthographic draws that can actually reach the VR
|
||||
// screen, and memoize it with the same state epoch as the ordinary pipeline.
|
||||
static gfx::PipelineRef resolve_exact_screen_depth_pipeline(GXPrimitive prim, GXVtxFmt fmt) {
|
||||
struct Memo {
|
||||
gfx::PipelineRef ref = 0;
|
||||
u32 generation = 0;
|
||||
u32 sampleCount = 0;
|
||||
GXPrimitive prim = static_cast<GXPrimitive>(0);
|
||||
GXVtxFmt fmt = static_cast<GXVtxFmt>(0);
|
||||
};
|
||||
static Memo memo{};
|
||||
|
||||
const u32 sampleCount = gfx::get_sample_count();
|
||||
const u32 generation = g_gxState.pipelineStateGeneration;
|
||||
if (memo.ref != 0 && memo.generation == generation && memo.sampleCount == sampleCount && memo.prim == prim &&
|
||||
memo.fmt == fmt)
|
||||
LIKELY { return memo.ref; }
|
||||
|
||||
PipelineConfig config{};
|
||||
populate_pipeline_config(config, prim, fmt);
|
||||
config.shaderConfig.exactScreenDepth = 1;
|
||||
const gfx::PipelineRef ref = gfx::pipeline_ref(config);
|
||||
memo = Memo{
|
||||
.ref = ref,
|
||||
.generation = generation,
|
||||
.sampleCount = sampleCount,
|
||||
.prim = prim,
|
||||
.fmt = fmt,
|
||||
};
|
||||
return ref;
|
||||
}
|
||||
|
||||
bool submit_raw_draw(GXPrimitive prim, GXVtxFmt fmt, const uint8_t* vertices, uint16_t vtxCount, uint32_t vertexBytes) {
|
||||
ZoneScoped;
|
||||
if (vertices == nullptr || vtxCount == 0 || vertexBytes == 0) {
|
||||
return false;
|
||||
}
|
||||
|
||||
if (__gx->dirtyState != 0) UNLIKELY {
|
||||
__GXSetDirtyState();
|
||||
}
|
||||
if (__gx->dirtyState != 0)
|
||||
UNLIKELY { __GXSetDirtyState(); }
|
||||
|
||||
// Raw bridge draws consume live decoded GX state that is also maintained by the HLE producer.
|
||||
drain();
|
||||
|
||||
u32 vtxSize;
|
||||
if (g_gxState.lastVtxFmt == fmt) LIKELY {
|
||||
vtxSize = g_gxState.lastVtxSize;
|
||||
} else UNLIKELY {
|
||||
vtxSize = calculate_last_vtx_size(fmt);
|
||||
}
|
||||
if (g_gxState.lastVtxFmt == fmt)
|
||||
LIKELY { vtxSize = g_gxState.lastVtxSize; }
|
||||
else
|
||||
UNLIKELY { vtxSize = calculate_last_vtx_size(fmt); }
|
||||
|
||||
const u32 expectedVertexBytes = static_cast<u32>(vtxCount) * vtxSize;
|
||||
if (expectedVertexBytes == 0 || expectedVertexBytes != vertexBytes) {
|
||||
static uint32_t rawVertexSizeMismatchCount = 0;
|
||||
if (rawVertexSizeMismatchCount++ < 64) {
|
||||
Log.warn("raw draw vertex-size mismatch prim={} fmt={} count={} cached_stride={} expected={} supplied={}",
|
||||
static_cast<uint32_t>(prim), static_cast<uint32_t>(fmt), vtxCount, vtxSize,
|
||||
expectedVertexBytes, vertexBytes);
|
||||
static_cast<uint32_t>(prim), static_cast<uint32_t>(fmt), vtxCount, vtxSize, expectedVertexBytes,
|
||||
vertexBytes);
|
||||
}
|
||||
return false;
|
||||
}
|
||||
@@ -2116,11 +2152,8 @@ bool submit_raw_draw(GXPrimitive prim, GXVtxFmt fmt, const uint8_t* vertices, ui
|
||||
std::lock_guard gpuLock(aurora::renderer_gpu_mutex());
|
||||
const gfx::Range vertRange = gfx::push_verts(vertices, vertexBytes);
|
||||
const bool interpolationIdentityActive = frame_interpolation_identity_needed();
|
||||
const PnMtxUsage matrixUsage = interpolationIdentityActive
|
||||
? pn_mtx_usage(vertices, vtxCount, vtxSize)
|
||||
: PnMtxUsage{};
|
||||
handle_draw_unmerged(prim, fmt, vtxCount, vertRange,
|
||||
matrixUsage.mask, matrixUsage.topologySignature,
|
||||
const PnMtxUsage matrixUsage = interpolationIdentityActive ? pn_mtx_usage(vertices, vtxCount, vtxSize) : PnMtxUsage{};
|
||||
handle_draw_unmerged(prim, fmt, vtxCount, vertRange, matrixUsage.mask, matrixUsage.topologySignature,
|
||||
interpolationIdentityActive ? draw_geometry_signature(fmt, vertices, vtxCount, vtxSize) : 0,
|
||||
interpolationIdentityActive);
|
||||
return true;
|
||||
@@ -2138,18 +2171,17 @@ static bool handle_draw(u8 cmd, const u8* data, u32& pos, u32 size, bool bigEndi
|
||||
pos += 2;
|
||||
|
||||
u32 vtxSize;
|
||||
if (g_gxState.lastVtxFmt == fmt) LIKELY {
|
||||
vtxSize = g_gxState.lastVtxSize;
|
||||
} else UNLIKELY {
|
||||
vtxSize = calculate_last_vtx_size(fmt);
|
||||
}
|
||||
if (g_gxState.lastVtxFmt == fmt)
|
||||
LIKELY { vtxSize = g_gxState.lastVtxSize; }
|
||||
else
|
||||
UNLIKELY { vtxSize = calculate_last_vtx_size(fmt); }
|
||||
|
||||
u32 totalVtxBytes = vtxCount * vtxSize;
|
||||
if (pos + totalVtxBytes > size) UNLIKELY {
|
||||
handle_draw_overrun(cmd, vtxCount, vtxSize, totalVtxBytes, data, pos, size);
|
||||
return false;
|
||||
}
|
||||
|
||||
if (pos + totalVtxBytes > size)
|
||||
UNLIKELY {
|
||||
handle_draw_overrun(cmd, vtxCount, vtxSize, totalVtxBytes, data, pos, size);
|
||||
return false;
|
||||
}
|
||||
|
||||
// Push raw vertex data to buffer
|
||||
const uint8_t* vertices = data + pos;
|
||||
@@ -2157,61 +2189,61 @@ static bool handle_draw(u8 cmd, const u8* data, u32& pos, u32 size, bool bigEndi
|
||||
pos += totalVtxBytes;
|
||||
|
||||
// Try to merge with previous draw call
|
||||
if (!g_gxState.stateDirty) LIKELY {
|
||||
auto* lastDraw = gfx::get_last_draw_command<DrawData>();
|
||||
// Only if the previous draw call was a single instance draw (no lines/points handling)
|
||||
if (lastDraw != nullptr && prim != GX_LINES && prim != GX_LINESTRIP && prim != GX_POINTS &&
|
||||
lastDraw->instanceCount == 1) LIKELY {
|
||||
const auto& indexTemplate = cached_index_template(prim, vtxCount);
|
||||
const auto indices = offset_index_template(indexTemplate, lastDraw->vtxCount);
|
||||
const u32 numIndices = indexTemplate.indexCount;
|
||||
const gfx::Range idxRange = gfx::push_indices(indices);
|
||||
CHECK(lastDraw->vertRange.offset + lastDraw->vertRange.size == vertRange.offset,
|
||||
"Non-consecutive vertex ranges ({} < {})", lastDraw->vertRange.offset + lastDraw->vertRange.size,
|
||||
vertRange.offset);
|
||||
CHECK(lastDraw->idxRange.offset + lastDraw->idxRange.size == idxRange.offset,
|
||||
"Non-consecutive index ranges ({} < {})", lastDraw->idxRange.offset + lastDraw->idxRange.size,
|
||||
idxRange.offset);
|
||||
lastDraw->vertRange.size += vertRange.size;
|
||||
lastDraw->idxRange.size += idxRange.size;
|
||||
lastDraw->vtxCount += vtxCount;
|
||||
lastDraw->indexCount += numIndices;
|
||||
++gfx::g_mergedDrawCallCount;
|
||||
// This primitive now renders through the draw we merged into, so its palette slots belong to that draw's interpolation snapshot as well.
|
||||
if (s_lastDrawRecordedInterpolation) UNLIKELY {
|
||||
extend_interpolation_draw(pn_mtx_mask(vertices, vtxCount, vtxSize));
|
||||
}
|
||||
return true;
|
||||
if (!g_gxState.stateDirty)
|
||||
LIKELY {
|
||||
auto* lastDraw = gfx::get_last_draw_command<DrawData>();
|
||||
// Only if the previous draw call was a single instance draw (no lines/points handling)
|
||||
if (lastDraw != nullptr && prim != GX_LINES && prim != GX_LINESTRIP && prim != GX_POINTS &&
|
||||
lastDraw->instanceCount == 1)
|
||||
LIKELY {
|
||||
const auto& indexTemplate = cached_index_template(prim, vtxCount);
|
||||
const auto indices = offset_index_template(indexTemplate, lastDraw->vtxCount);
|
||||
const u32 numIndices = indexTemplate.indexCount;
|
||||
const gfx::Range idxRange = gfx::push_indices(indices);
|
||||
CHECK(lastDraw->vertRange.offset + lastDraw->vertRange.size == vertRange.offset,
|
||||
"Non-consecutive vertex ranges ({} < {})", lastDraw->vertRange.offset + lastDraw->vertRange.size,
|
||||
vertRange.offset);
|
||||
CHECK(lastDraw->idxRange.offset + lastDraw->idxRange.size == idxRange.offset,
|
||||
"Non-consecutive index ranges ({} < {})", lastDraw->idxRange.offset + lastDraw->idxRange.size,
|
||||
idxRange.offset);
|
||||
lastDraw->vertRange.size += vertRange.size;
|
||||
lastDraw->idxRange.size += idxRange.size;
|
||||
lastDraw->vtxCount += vtxCount;
|
||||
lastDraw->indexCount += numIndices;
|
||||
++gfx::g_mergedDrawCallCount;
|
||||
// This primitive now renders through the draw we merged into, so its palette slots belong to that draw's
|
||||
// interpolation snapshot as well.
|
||||
if (s_lastDrawRecordedInterpolation)
|
||||
UNLIKELY { extend_interpolation_draw(pn_mtx_mask(vertices, vtxCount, vtxSize)); }
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const bool interpolationIdentityActive = frame_interpolation_identity_needed();
|
||||
const PnMtxUsage matrixUsage = interpolationIdentityActive
|
||||
? pn_mtx_usage(vertices, vtxCount, vtxSize)
|
||||
: PnMtxUsage{};
|
||||
handle_draw_unmerged(prim, fmt, vtxCount, vertRange,
|
||||
matrixUsage.mask, matrixUsage.topologySignature,
|
||||
const PnMtxUsage matrixUsage = interpolationIdentityActive ? pn_mtx_usage(vertices, vtxCount, vtxSize) : PnMtxUsage{};
|
||||
handle_draw_unmerged(prim, fmt, vtxCount, vertRange, matrixUsage.mask, matrixUsage.topologySignature,
|
||||
interpolationIdentityActive ? draw_geometry_signature(fmt, vertices, vtxCount, vtxSize) : 0,
|
||||
interpolationIdentityActive);
|
||||
return true;
|
||||
}
|
||||
|
||||
static void handle_draw_unmerged(GXPrimitive prim, GXVtxFmt fmt, u16 vtxCount,
|
||||
gfx::Range vertRange, uint16_t usedPnMtxMask,
|
||||
HashType matrixTopologySignature,
|
||||
HashType geometrySignature, bool interpolationIdentityActive) {
|
||||
static void handle_draw_unmerged(GXPrimitive prim, GXVtxFmt fmt, u16 vtxCount, gfx::Range vertRange,
|
||||
uint16_t usedPnMtxMask, HashType matrixTopologySignature, HashType geometrySignature,
|
||||
bool interpolationIdentityActive) {
|
||||
ZoneScoped;
|
||||
// GX_CULL_ALL rasterizes nothing on hardware - no color, no depth.
|
||||
if (g_gxState.cullMode == GX_CULL_ALL && prim != GX_LINES && prim != GX_LINESTRIP && prim != GX_POINTS)
|
||||
UNLIKELY {
|
||||
// Leave stateDirty alone: the next draw re-resolving its pipeline is the safe direction, and nothing about this draw reached the GPU.
|
||||
return;
|
||||
}
|
||||
// Callers hold the renderer GPU mutex for the whole drain (process() and submit_raw_draw); taking it again per draw only cost a recursive re-entry.
|
||||
UNLIKELY {
|
||||
// Leave stateDirty alone: the next draw re-resolving its pipeline is the safe direction, and nothing about this
|
||||
// draw reached the GPU.
|
||||
return;
|
||||
}
|
||||
// Callers hold the renderer GPU mutex for the whole drain (process() and submit_raw_draw); taking it again per draw
|
||||
// only cost a recursive re-entry.
|
||||
const auto& indexTemplate = cached_index_template(prim, vtxCount);
|
||||
const u32 numIndices = indexTemplate.indexCount;
|
||||
const gfx::Range idxRange = gfx::push_indices(ArrayRef<u16>{
|
||||
indexTemplate.indices.data(), indexTemplate.indices.size()});
|
||||
const gfx::Range idxRange =
|
||||
gfx::push_indices(ArrayRef<u16>{indexTemplate.indices.data(), indexTemplate.indices.size()});
|
||||
|
||||
// Build pipeline, bind groups, and push draw command
|
||||
BindGroupRanges ranges{};
|
||||
@@ -2238,27 +2270,31 @@ static void handle_draw_unmerged(GXPrimitive prim, GXVtxFmt fmt, u16 vtxCount,
|
||||
|
||||
const auto pipeline = pipelineState.ref;
|
||||
|
||||
// Draw-identity hashing only feeds frame interpolation, and only for perspective draws: build_uniform reads the identity exclusively past its `!perspective || frame_interpolation_fps() == 0` early-out.
|
||||
// Draw-identity hashing only feeds frame interpolation, and only for perspective draws: build_uniform reads the
|
||||
// identity exclusively past its `!perspective || frame_interpolation_fps() == 0` early-out.
|
||||
FrameInterpolationDrawIdentity drawIdentity{};
|
||||
if (interpolationIdentityActive) UNLIKELY {
|
||||
const HashType drawShape = static_cast<HashType>(vtxCount) |
|
||||
(static_cast<HashType>(underlying(prim)) << 16) |
|
||||
(static_cast<HashType>(underlying(fmt)) << 24);
|
||||
const HashType pipelineDrawSignature = xxh3_hash(pipelineState.configHash, drawShape);
|
||||
const HashType textureSignature = xxh3_hash(bindGroups.textureBindGroup);
|
||||
const HashType materialAndTopology =
|
||||
xxh3_hash(matrixTopologySignature,
|
||||
xxh3_hash(bindGroups.textureBindGroup, pipelineDrawSignature));
|
||||
drawIdentity = FrameInterpolationDrawIdentity{
|
||||
.combined = xxh3_hash(geometrySignature, materialAndTopology),
|
||||
.pipeline = pipelineDrawSignature,
|
||||
.texture = textureSignature,
|
||||
.matrixTopology = matrixTopologySignature,
|
||||
};
|
||||
}
|
||||
if (interpolationIdentityActive)
|
||||
UNLIKELY {
|
||||
const HashType drawShape = static_cast<HashType>(vtxCount) | (static_cast<HashType>(underlying(prim)) << 16) |
|
||||
(static_cast<HashType>(underlying(fmt)) << 24);
|
||||
const HashType pipelineDrawSignature = xxh3_hash(pipelineState.configHash, drawShape);
|
||||
const HashType textureSignature = xxh3_hash(bindGroups.textureBindGroup);
|
||||
const HashType materialAndTopology =
|
||||
xxh3_hash(matrixTopologySignature, xxh3_hash(bindGroups.textureBindGroup, pipelineDrawSignature));
|
||||
drawIdentity = FrameInterpolationDrawIdentity{
|
||||
.combined = xxh3_hash(geometrySignature, materialAndTopology),
|
||||
.pipeline = pipelineDrawSignature,
|
||||
.texture = textureSignature,
|
||||
.matrixTopology = matrixTopologySignature,
|
||||
};
|
||||
}
|
||||
const bool perspective = g_gxState.projType == GX_PERSPECTIVE;
|
||||
const auto uniformRanges =
|
||||
build_uniform(info, vertRange.offset, ranges, drawIdentity, perspective, usedPnMtxMask);
|
||||
const auto uniformRanges = build_uniform(info, vertRange.offset, ranges, drawIdentity, perspective, usedPnMtxMask);
|
||||
const auto& replayLayout = uniformRanges.replayLayout;
|
||||
const gfx::PipelineRef exactScreenDepthPipeline =
|
||||
aurora::stereo_frame_provider_active() && !replayLayout.perspective && !replayLayout.nativeEfbEffect
|
||||
? resolve_exact_screen_depth_pipeline(prim, fmt)
|
||||
: 0;
|
||||
s_lastDrawRecordedInterpolation = interpolationIdentityActive;
|
||||
|
||||
uint32_t instanceCount = 1;
|
||||
@@ -2271,6 +2307,7 @@ static void handle_draw_unmerged(GXPrimitive prim, GXVtxFmt fmt, u16 vtxCount,
|
||||
}
|
||||
gfx::push_draw_command(DrawData{
|
||||
.pipeline = pipeline,
|
||||
.exactScreenDepthPipeline = exactScreenDepthPipeline,
|
||||
.vertRange = vertRange,
|
||||
.idxRange = idxRange,
|
||||
.uniformRange = uniformRanges.current,
|
||||
|
||||
+26
-26
@@ -65,7 +65,8 @@ constexpr u32 MaxIndStages = GX_MAX_INDTEXSTAGE;
|
||||
constexpr u32 MaxIndTexMtxs = 3;
|
||||
constexpr u32 MaxVtxFmt = GX_MAX_VTXFMT;
|
||||
constexpr u32 MaxPnMtx = (GX_PNMTX9 / 3) + 1;
|
||||
// Position and texture matrices share one shader array (`ubuf.postex_mtx`), mirroring XF matrix memory: rows 0..29 (slots 0..9) are position matrices and rows 30..59 (slots 10..19) are texture matrices.
|
||||
// Position and texture matrices share one shader array (`ubuf.postex_mtx`), mirroring XF matrix memory: rows 0..29
|
||||
// (slots 0..9) are position matrices and rows 30..59 (slots 10..19) are texture matrices.
|
||||
constexpr u32 MaxPostexMtx = MaxPnMtx + MaxTexMtx;
|
||||
constexpr u32 MaxIndexAttr = 12; // VA_POS -> VA_TEX7
|
||||
constexpr u32 MaxUniformSize = 3840;
|
||||
@@ -439,7 +440,8 @@ struct GXState {
|
||||
regs[0xFE] = 0x00FFFFFF;
|
||||
return regs;
|
||||
}();
|
||||
// Covers XF 0x00-0x5F: the scalar bank, viewport/projection, and the TexGen and post-transform registers that every material's display list re-emits.
|
||||
// Covers XF 0x00-0x5F: the scalar bank, viewport/projection, and the TexGen and post-transform registers that every
|
||||
// material's display list re-emits.
|
||||
std::array<u32, 0x60> xfRegCache{};
|
||||
std::array<u64, 2> xfRegCacheValid{};
|
||||
|
||||
@@ -490,13 +492,9 @@ inline float clear_depth_value() {
|
||||
|
||||
inline bool render_target_has_alpha(GXPixelFmt pixelFmt) noexcept { return pixelFmt == GX_PF_RGBA6_Z24; }
|
||||
|
||||
inline float indirect_matrix_scale_multiplier(s8 scaleExp) {
|
||||
return std::exp2f(static_cast<float>(scaleExp));
|
||||
}
|
||||
inline float indirect_matrix_scale_multiplier(s8 scaleExp) { return std::exp2f(static_cast<float>(scaleExp)); }
|
||||
|
||||
inline s32 indirect_matrix_mantissa(float value) noexcept {
|
||||
return static_cast<s32>(std::lround(value * 1024.0f));
|
||||
}
|
||||
inline s32 indirect_matrix_mantissa(float value) noexcept { return static_cast<s32>(std::lround(value * 1024.0f)); }
|
||||
|
||||
inline s32 indirect_matrix_shift(s8 scaleExp) noexcept { return -static_cast<s32>(scaleExp); }
|
||||
|
||||
@@ -542,7 +540,7 @@ struct AttrConfig {
|
||||
u8 stride = 0; // Array stride
|
||||
u8 frac = 0;
|
||||
bool le = true;
|
||||
u8 nrmIndexCount = 0; // GX_NRM_NBT3 stores three separate normal/tangent/binormal indices.
|
||||
u8 nrmIndexCount = 0; // GX_NRM_NBT3 stores three separate normal/tangent/binormal indices.
|
||||
};
|
||||
struct ShaderConfig {
|
||||
u8 fogType = GX_FOG_NONE;
|
||||
@@ -550,7 +548,10 @@ struct ShaderConfig {
|
||||
u8 lineMode : 2 = 0; // 1 = GX_LINES, 2 = GX_LINESTRIP, 3 = GX_POINTS
|
||||
u8 dualTexEnabled : 1 = 0;
|
||||
u8 fogRangeAdjust : 1 = 0;
|
||||
u8 pad1 : 4 = 0;
|
||||
// Alternate replay-only shader variant that exports the original flat-screen
|
||||
// depth for draws reprojected onto the VR virtual screen.
|
||||
u8 exactScreenDepth : 1 = 0;
|
||||
u8 pad1 : 3 = 0;
|
||||
u8 numTexGens = 0;
|
||||
u32 zTexture = 0; // bias[0:23], format[24:25], op[26:27]; 0 disables shader depth output.
|
||||
std::array<AttrConfig, MaxVtxAttr> attrs;
|
||||
@@ -593,8 +594,7 @@ inline bool tev_stage_needs_fixed_texcoord_state(const ShaderConfig& config, u32
|
||||
const auto& stage = config.tevStages[stageIdx];
|
||||
const bool hasCoordOperation = stage.indTexMtxId != GX_ITM_OFF || stage.indTexWrapS != GX_ITW_OFF ||
|
||||
stage.indTexWrapT != GX_ITW_OFF || stage.indTexAddPrev;
|
||||
const bool feedsNextAddPrev =
|
||||
stageIdx + 1 < config.tevStageCount && config.tevStages[stageIdx + 1].indTexAddPrev;
|
||||
const bool feedsNextAddPrev = stageIdx + 1 < config.tevStageCount && config.tevStages[stageIdx + 1].indTexAddPrev;
|
||||
return hasCoordOperation || feedsNextAddPrev;
|
||||
}
|
||||
|
||||
@@ -626,11 +626,11 @@ inline int tev_indirect_wrap_mask(GXIndTexWrap wrap) noexcept {
|
||||
|
||||
inline constexpr s32 tev_s24_wrap(s32 value) noexcept {
|
||||
const u32 wrapped = static_cast<u32>(value) & 0x00ffffffu;
|
||||
return (wrapped & 0x00800000u) != 0 ? static_cast<s32>(wrapped) - 0x01000000
|
||||
: static_cast<s32>(wrapped);
|
||||
return (wrapped & 0x00800000u) != 0 ? static_cast<s32>(wrapped) - 0x01000000 : static_cast<s32>(wrapped);
|
||||
}
|
||||
|
||||
// GX falls back to texcoord 0 when a TEV order names GX_TEXCOORD_NULL or a coordinate beyond the configured texgen count.
|
||||
// GX falls back to texcoord 0 when a TEV order names GX_TEXCOORD_NULL or a coordinate beyond the configured texgen
|
||||
// count.
|
||||
inline int tev_effective_texcoord(const ShaderConfig& config, GXTexCoordID texCoordId) noexcept {
|
||||
if (config.numTexGens == 0) {
|
||||
return -1;
|
||||
@@ -643,10 +643,9 @@ inline int tev_effective_texcoord(const ShaderConfig& config, GXTexCoordID texCo
|
||||
inline bool tev_stage_combiner_uses_texture(const TevStage& stage) noexcept {
|
||||
const auto& color = stage.colorPass;
|
||||
const auto& alpha = stage.alphaPass;
|
||||
return color.a == GX_CC_TEXC || color.a == GX_CC_TEXA || color.b == GX_CC_TEXC ||
|
||||
color.b == GX_CC_TEXA || color.c == GX_CC_TEXC || color.c == GX_CC_TEXA ||
|
||||
color.d == GX_CC_TEXC || color.d == GX_CC_TEXA || alpha.a == GX_CA_TEXA ||
|
||||
alpha.b == GX_CA_TEXA || alpha.c == GX_CA_TEXA || alpha.d == GX_CA_TEXA;
|
||||
return color.a == GX_CC_TEXC || color.a == GX_CC_TEXA || color.b == GX_CC_TEXC || color.b == GX_CC_TEXA ||
|
||||
color.c == GX_CC_TEXC || color.c == GX_CC_TEXA || color.d == GX_CC_TEXC || color.d == GX_CC_TEXA ||
|
||||
alpha.a == GX_CA_TEXA || alpha.b == GX_CA_TEXA || alpha.c == GX_CA_TEXA || alpha.d == GX_CA_TEXA;
|
||||
}
|
||||
|
||||
struct TevStageTextureDependency {
|
||||
@@ -657,14 +656,12 @@ struct TevStageTextureDependency {
|
||||
bool canSampleTexture = false;
|
||||
};
|
||||
|
||||
inline bool tev_texture_sample_enabled(const TevStageTextureDependency& dependency,
|
||||
bool sampleRequested) noexcept {
|
||||
inline bool tev_texture_sample_enabled(const TevStageTextureDependency& dependency, bool sampleRequested) noexcept {
|
||||
return sampleRequested && dependency.canSampleTexture;
|
||||
}
|
||||
|
||||
// Pure TEV order/dependency analysis used by both ShaderInfo and WGSL generation.
|
||||
inline TevStageTextureDependency tev_stage_texture_dependency(const ShaderConfig& config,
|
||||
u32 stageIdx) noexcept {
|
||||
inline TevStageTextureDependency tev_stage_texture_dependency(const ShaderConfig& config, u32 stageIdx) noexcept {
|
||||
if (stageIdx >= config.tevStageCount) {
|
||||
return {};
|
||||
}
|
||||
@@ -712,9 +709,11 @@ struct GXBindGroups {
|
||||
// Bind group resolved at draw-build time so that gx::render does not have to hash-map the ref again for every draw.
|
||||
WGPUBindGroup resolvedTextureBindGroup = nullptr;
|
||||
};
|
||||
// Which matrix-memory slots a generated shader can actually read, and where each one lives in the compacted uniform arrays.
|
||||
// Which matrix-memory slots a generated shader can actually read, and where each one lives in the compacted uniform
|
||||
// arrays.
|
||||
struct UniformMatrixLayout {
|
||||
// `postexSlots`/`nrmSlots` entry meaning "the matrix selected by the current matrix index", which is only known per draw.
|
||||
// `postexSlots`/`nrmSlots` entry meaning "the matrix selected by the current matrix index", which is only known per
|
||||
// draw.
|
||||
static constexpr u8 kCurrentPnMtx = 0xFE;
|
||||
// `postexRemap` entry for a slot that this shader never reads.
|
||||
static constexpr u8 kAbsent = 0xFF;
|
||||
@@ -727,7 +726,8 @@ struct UniformMatrixLayout {
|
||||
std::array<u8, MaxPnMtx> nrmSlots{};
|
||||
u8 postexCount = 0;
|
||||
u8 nrmCount = 0;
|
||||
// Position slots 0..MaxPnMtx-1 are uploaded 1:1 at compact indices 0..9, so `current_pnmtx` and any per-vertex index need no remapping.
|
||||
// Position slots 0..MaxPnMtx-1 are uploaded 1:1 at compact indices 0..9, so `current_pnmtx` and any per-vertex index
|
||||
// need no remapping.
|
||||
bool absolutePosRegion = false;
|
||||
};
|
||||
// Output info from shader generation
|
||||
|
||||
@@ -66,12 +66,14 @@ void clear_shader_module_cache() {
|
||||
}
|
||||
|
||||
void render(const DrawData& data, const wgpu::RenderPassEncoder& pass, DrawEncodeState& state,
|
||||
bool requireReadyPipeline, const gfx::Range* uniformRangeOverride) {
|
||||
if (!gfx::bind_pipeline(data.pipeline, pass, state.currentPipeline, requireReadyPipeline)) {
|
||||
bool requireReadyPipeline, const gfx::Range* uniformRangeOverride, gfx::PipelineRef pipelineOverride) {
|
||||
const gfx::PipelineRef pipeline = pipelineOverride != 0 ? pipelineOverride : data.pipeline;
|
||||
if (!gfx::bind_pipeline(pipeline, pass, state.currentPipeline, requireReadyPipeline)) {
|
||||
return;
|
||||
}
|
||||
|
||||
// An interpolated presentation slot re-encodes the identical draw with only this range replaced; overriding here avoids copying the whole DrawData per draw per slot.
|
||||
// An interpolated presentation slot re-encodes the identical draw with only this range replaced; overriding here
|
||||
// avoids copying the whole DrawData per draw per slot.
|
||||
const gfx::Range& uniformRange = uniformRangeOverride != nullptr ? *uniformRangeOverride : data.uniformRange;
|
||||
const std::array offsets{uniformRange.offset};
|
||||
pass.SetBindGroup(1, gfx::g_uniformBindGroup, offsets.size(), offsets.data());
|
||||
@@ -90,7 +92,6 @@ void render(const DrawData& data, const wgpu::RenderPassEncoder& pass, DrawEncod
|
||||
pass.SetIndexBuffer(gfx::g_indexBuffer, wgpu::IndexFormat::Uint16, 0, wgpu::kWholeSize);
|
||||
state.indexBufferBound = true;
|
||||
}
|
||||
pass.DrawIndexed(data.indexCount, data.instanceCount,
|
||||
static_cast<uint32_t>(data.idxRange.offset / sizeof(uint16_t)));
|
||||
pass.DrawIndexed(data.indexCount, data.instanceCount, static_cast<uint32_t>(data.idxRange.offset / sizeof(uint16_t)));
|
||||
}
|
||||
} // namespace aurora::gx
|
||||
@@ -6,6 +6,9 @@
|
||||
namespace aurora::gx {
|
||||
struct DrawData {
|
||||
gfx::PipelineRef pipeline;
|
||||
// Same GX state with exact fragment-depth export enabled. Bound only when
|
||||
// this draw is actually reprojected onto the VR virtual screen.
|
||||
gfx::PipelineRef exactScreenDepthPipeline;
|
||||
gfx::Range vertRange;
|
||||
gfx::Range idxRange;
|
||||
gfx::Range uniformRange;
|
||||
@@ -19,14 +22,12 @@ struct DrawData {
|
||||
uint32_t dstAlpha;
|
||||
};
|
||||
|
||||
constexpr uint32_t GXPipelineConfigVersion = 19;
|
||||
constexpr uint32_t GXPipelineConfigVersion = 20;
|
||||
|
||||
constexpr GXFogType effective_pipeline_fog_type(GXFogType fogType, GXZTexOp zTextureOp,
|
||||
bool zCompLocBeforeTex, GXBlendMode blendMode,
|
||||
GXLogicOp logicOp) noexcept {
|
||||
constexpr GXFogType effective_pipeline_fog_type(GXFogType fogType, GXZTexOp zTextureOp, bool zCompLocBeforeTex,
|
||||
GXBlendMode blendMode, GXLogicOp logicOp) noexcept {
|
||||
const bool usesLateZTexture = zTextureOp != GX_ZT_DISABLE && !zCompLocBeforeTex;
|
||||
const bool usesUnsupportedLogicFog =
|
||||
blendMode == GX_BM_LOGIC && logicOp != GX_LO_OR;
|
||||
const bool usesUnsupportedLogicFog = blendMode == GX_BM_LOGIC && logicOp != GX_LO_OR;
|
||||
return usesLateZTexture && usesUnsupportedLogicFog ? GX_FOG_NONE : fogType;
|
||||
}
|
||||
|
||||
@@ -49,33 +50,33 @@ inline bool valid_pipeline_config(const PipelineConfig& config) noexcept {
|
||||
const auto in_range = [](auto value, auto maximum) {
|
||||
using Value = decltype(value);
|
||||
return static_cast<std::underlying_type_t<Value>>(value) >= 0 &&
|
||||
static_cast<std::underlying_type_t<Value>>(value) <=
|
||||
static_cast<std::underlying_type_t<Value>>(maximum);
|
||||
static_cast<std::underlying_type_t<Value>>(value) <= static_cast<std::underlying_type_t<Value>>(maximum);
|
||||
};
|
||||
const bool validSamples = config.msaaSamples == 1 || config.msaaSamples == 2 ||
|
||||
config.msaaSamples == 4 || config.msaaSamples == 8;
|
||||
return config.version == GXPipelineConfigVersion && validSamples &&
|
||||
in_range(config.depthFunc, GX_ALWAYS) && in_range(config.cullMode, GX_CULL_ALL) &&
|
||||
in_range(config.blendMode, GX_BM_SUBTRACT) &&
|
||||
in_range(config.blendFacSrc, GX_BL_INVDSTALPHA) &&
|
||||
in_range(config.blendFacDst, GX_BL_INVDSTALPHA) &&
|
||||
const bool validSamples =
|
||||
config.msaaSamples == 1 || config.msaaSamples == 2 || config.msaaSamples == 4 || config.msaaSamples == 8;
|
||||
return config.version == GXPipelineConfigVersion && validSamples && in_range(config.depthFunc, GX_ALWAYS) &&
|
||||
in_range(config.cullMode, GX_CULL_ALL) && in_range(config.blendMode, GX_BM_SUBTRACT) &&
|
||||
in_range(config.blendFacSrc, GX_BL_INVDSTALPHA) && in_range(config.blendFacDst, GX_BL_INVDSTALPHA) &&
|
||||
in_range(config.blendOp, GX_LO_SET) && in_range(config.pixelFmt, GX_PF_YUV420);
|
||||
}
|
||||
|
||||
wgpu::RenderPipeline create_pipeline([[maybe_unused]] const PipelineConfig& config);
|
||||
void clear_shader_module_cache();
|
||||
|
||||
// Per-pass encoder state carried across the draws of one render pass, so the replay loop can elide Dawn calls that would re-bind what is already bound.
|
||||
// Per-pass encoder state carried across the draws of one render pass, so the replay loop can elide Dawn calls that
|
||||
// would re-bind what is already bound.
|
||||
struct DrawEncodeState {
|
||||
gfx::PipelineRef currentPipeline = UINTPTR_MAX;
|
||||
// The bind group currently occupying slot 2.
|
||||
WGPUBindGroup boundTextureBindGroup = nullptr;
|
||||
// The pass-wide index buffer binding is established lazily by the first indexed draw; every draw then addresses it with firstIndex instead of a per-draw SetIndexBuffer.
|
||||
// The pass-wide index buffer binding is established lazily by the first indexed draw; every draw then addresses it
|
||||
// with firstIndex instead of a per-draw SetIndexBuffer.
|
||||
bool indexBufferBound = false;
|
||||
};
|
||||
|
||||
void render(const DrawData& data, const wgpu::RenderPassEncoder& pass, DrawEncodeState& state,
|
||||
bool requireReadyPipeline, const gfx::Range* uniformRangeOverride = nullptr);
|
||||
bool requireReadyPipeline, const gfx::Range* uniformRangeOverride = nullptr,
|
||||
gfx::PipelineRef pipelineOverride = 0);
|
||||
|
||||
void queue_surface(const u8* dlStart, uint32_t dlSize, bool bigEndian) noexcept;
|
||||
} // namespace aurora::gx
|
||||
@@ -22,7 +22,6 @@ using namespace std::string_view_literals;
|
||||
|
||||
static Module Log("aurora::gfx::gx");
|
||||
|
||||
|
||||
static inline std::string_view chan_comp(GXTevColorChan chan) noexcept {
|
||||
switch (chan) {
|
||||
case GX_CH_RED:
|
||||
@@ -356,17 +355,15 @@ static std::string tev_scale_index(GXTevScale scale) {
|
||||
}
|
||||
}
|
||||
|
||||
static std::string tev_regular_op(GXTevOp op, GXTevBias bias, GXTevScale scale, std::string_view a,
|
||||
std::string_view b, std::string_view c, std::string_view d,
|
||||
std::string_view suffix) {
|
||||
static std::string tev_regular_op(GXTevOp op, GXTevBias bias, GXTevScale scale, std::string_view a, std::string_view b,
|
||||
std::string_view c, std::string_view d, std::string_view suffix) {
|
||||
CHECK(op == GX_TEV_ADD || op == GX_TEV_SUB, "invalid regular tev op {}", underlying(op));
|
||||
return fmt::format("tev_regular_{4}({0}, {1}, {2}, {3}, {5}, {6}, {7})", a, b, c, d, suffix,
|
||||
tev_bias_i32(bias), tev_scale_index(scale), op == GX_TEV_SUB ? "true" : "false");
|
||||
return fmt::format("tev_regular_{4}({0}, {1}, {2}, {3}, {5}, {6}, {7})", a, b, c, d, suffix, tev_bias_i32(bias),
|
||||
tev_scale_index(scale), op == GX_TEV_SUB ? "true" : "false");
|
||||
}
|
||||
|
||||
static std::string tev_op(GXTevOp op, GXTevBias bias, GXTevScale scale, std::string_view a, std::string_view b,
|
||||
std::string_view c, std::string_view d, std::string_view zero,
|
||||
std::string_view suffix) {
|
||||
std::string_view c, std::string_view d, std::string_view zero, std::string_view suffix) {
|
||||
switch (op) {
|
||||
DEFAULT_FATAL("unimplemented tev op {}", underlying(op));
|
||||
case GX_TEV_ADD:
|
||||
@@ -403,8 +400,8 @@ static std::string tev_op(GXTevOp op, GXTevBias bias, GXTevScale scale, std::str
|
||||
}
|
||||
}
|
||||
|
||||
static std::string tev_color_op(GXTevOp op, GXTevBias bias, GXTevScale scale, bool clamp,
|
||||
std::string_view a, std::string_view b, std::string_view c, std::string_view d) {
|
||||
static std::string tev_color_op(GXTevOp op, GXTevBias bias, GXTevScale scale, bool clamp, std::string_view a,
|
||||
std::string_view b, std::string_view c, std::string_view d) {
|
||||
const auto overflow = [](std::string_view reg) { return fmt::format("tev_overflow_vec3f({})", reg); };
|
||||
std::string expr = tev_op(op, bias, scale, overflow(a), overflow(b), overflow(c), d, "vec3(0)"sv, "vec3f"sv);
|
||||
return clamp ? fmt::format("clamp({}, vec3f(0.0), vec3f(1.0))", expr)
|
||||
@@ -425,21 +422,24 @@ static bool alpha_compare_uses_color_inputs(GXTevOp op) noexcept {
|
||||
}
|
||||
}
|
||||
|
||||
static std::string tev_alpha_op(GXTevOp op, GXTevBias bias, GXTevScale scale, bool clamp,
|
||||
std::string_view a, std::string_view b, std::string_view c, std::string_view d,
|
||||
std::string_view colorA, std::string_view colorB) {
|
||||
static std::string tev_alpha_op(GXTevOp op, GXTevBias bias, GXTevScale scale, bool clamp, std::string_view a,
|
||||
std::string_view b, std::string_view c, std::string_view d, std::string_view colorA,
|
||||
std::string_view colorB) {
|
||||
const auto scalarOverflow = [](std::string_view reg) { return fmt::format("tev_overflow_f32({})", reg); };
|
||||
std::string expr;
|
||||
if (alpha_compare_uses_color_inputs(op)) {
|
||||
const auto colorOverflow = [](std::string_view reg) { return fmt::format("tev_overflow_vec3f({})", reg); };
|
||||
// GX alpha compares R8/GR16/BGR24 using the color combiner's A/B inputs, while C and D remain the scalar alpha-combiner inputs.
|
||||
expr = tev_op(op, bias, scale, colorOverflow(colorA), colorOverflow(colorB), scalarOverflow(c), d, "0.0"sv,
|
||||
"f32"sv);
|
||||
// GX alpha compares R8/GR16/BGR24 using the color combiner's A/B inputs, while C and D remain the scalar
|
||||
// alpha-combiner inputs.
|
||||
expr =
|
||||
tev_op(op, bias, scale, colorOverflow(colorA), colorOverflow(colorB), scalarOverflow(c), d, "0.0"sv, "f32"sv);
|
||||
} else {
|
||||
// The numeric RGB8 op aliases are the alpha combiner's A8 compares and therefore continue to compare scalar alpha A/B inputs here.
|
||||
// The numeric RGB8 op aliases are the alpha combiner's A8 compares and therefore continue to compare scalar alpha
|
||||
// A/B inputs here.
|
||||
expr = tev_op(op, bias, scale, scalarOverflow(a), scalarOverflow(b), scalarOverflow(c), d, "0.0"sv, "f32"sv);
|
||||
}
|
||||
return clamp ? fmt::format("clamp({}, 0.0, 1.0)", expr) : fmt::format("clamp({}, -1024.0 / 255.0, 1023.0 / 255.0)", expr);
|
||||
return clamp ? fmt::format("clamp({}, 0.0, 1.0)", expr)
|
||||
: fmt::format("clamp({}, -1024.0 / 255.0, 1023.0 / 255.0)", expr);
|
||||
}
|
||||
|
||||
struct AlphaCompareExpr {
|
||||
@@ -611,7 +611,8 @@ constexpr std::array<std::string_view, GX_CA_ZERO + 1> TevAlphaArgNames{
|
||||
"APREV"sv, "A0"sv, "A1"sv, "A2"sv, "TEXA"sv, "RASA"sv, "KONST"sv, "ZERO"sv,
|
||||
};
|
||||
|
||||
auto fetch_attr(const AttrConfig& mapping, std::string_view buf, std::string_view offs, bool le, u8 cntOverride = 0) -> std::string {
|
||||
auto fetch_attr(const AttrConfig& mapping, std::string_view buf, std::string_view offs, bool le, u8 cntOverride = 0)
|
||||
-> std::string {
|
||||
const u8 cnt = cntOverride != 0 ? cntOverride : mapping.cnt;
|
||||
switch (mapping.compType) {
|
||||
case GX_U8:
|
||||
@@ -695,8 +696,8 @@ auto normal_group_load(const ShaderConfig& config, u8 group, std::string_view vi
|
||||
le = mapping.le;
|
||||
} else if (mapping.attrType == GX_INDEX16) {
|
||||
const std::string indexOffs = offset_plus(offs, mapping.nrmIndexCount == 3 ? group * 2u : 0);
|
||||
offs = fmt::format("ubuf.array_start[{}] + raw_fetch_u16_1(&{}, {}, {}) * {}u + {}u", GX_VA_NRM - GX_VA_POS,
|
||||
buf, indexOffs, le, mapping.stride, groupOffset);
|
||||
offs = fmt::format("ubuf.array_start[{}] + raw_fetch_u16_1(&{}, {}, {}) * {}u + {}u", GX_VA_NRM - GX_VA_POS, buf,
|
||||
indexOffs, le, mapping.stride, groupOffset);
|
||||
buf = "abuf"sv;
|
||||
le = mapping.le;
|
||||
} else {
|
||||
@@ -809,8 +810,7 @@ auto lighting_func(const ShaderConfig& config, const ColorChannelConfig& cc, u8
|
||||
if (alpha) {
|
||||
return fmt::format("\n {0}.a = f32(i32(round({1}.a * 255.0))) / 255.0;", outVar, matSrc);
|
||||
}
|
||||
return fmt::format("\n {0} = vec4f(vec3f(vec3i(round({1}.rgb * 255.0))) / 255.0, 1.0);", outVar,
|
||||
matSrc);
|
||||
return fmt::format("\n {0} = vec4f(vec3f(vec3i(round({1}.rgb * 255.0))) / 255.0, 1.0);", outVar, matSrc);
|
||||
}
|
||||
GXDiffuseFn diffFn = cc.diffFn;
|
||||
std::string lightAttnFn;
|
||||
@@ -824,9 +824,8 @@ auto lighting_func(const ShaderConfig& config, const ColorChannelConfig& cc, u8
|
||||
attn = max(0.0, cos_attn / dist_attn);)""");
|
||||
} else if (cc.attnFn == GX_AF_SPEC) {
|
||||
std::string_view normal = UsePerPixelLighting ? "in.mv_nrm"sv : "mv_nrm"sv;
|
||||
std::string dist_attn = diffFn != GX_DF_NONE
|
||||
? "dot(normalize(light.dist_att), vec3f(1.0, attn, attn * attn))"
|
||||
: "dot(light.dist_att, vec3f(1.0, attn, attn * attn))";
|
||||
std::string dist_attn = diffFn != GX_DF_NONE ? "dot(normalize(light.dist_att), vec3f(1.0, attn, attn * attn))"
|
||||
: "dot(light.dist_att, vec3f(1.0, attn, attn * attn))";
|
||||
lightAttnFn = fmt::format(R"""(
|
||||
attn = select(0.0, max(0.0, dot({0}, light.dir)), dot({0}, ldir) >= 0.0);
|
||||
var cos_attn = dot(light.cos_att, vec3f(1.0, attn, attn * attn));
|
||||
@@ -842,14 +841,14 @@ auto lighting_func(const ShaderConfig& config, const ColorChannelConfig& cc, u8
|
||||
var dist2 = dot(ldir, ldir);
|
||||
var dist = sqrt(dist2);
|
||||
ldir = select({1}, ldir / max(dist, 1e-20), dist > 0.0);)""",
|
||||
posVar, normal);
|
||||
posVar, normal);
|
||||
} else {
|
||||
lightDirSetup = fmt::format(R"""(
|
||||
var ldir = light.pos - {0};
|
||||
var dist2 = dot(ldir, ldir);
|
||||
var dist = sqrt(dist2);
|
||||
ldir = ldir / dist;)""",
|
||||
posVar);
|
||||
posVar);
|
||||
}
|
||||
std::string_view lightDiffFn;
|
||||
if (diffFn == GX_DF_NONE) {
|
||||
@@ -869,10 +868,8 @@ auto lighting_func(const ShaderConfig& config, const ColorChannelConfig& cc, u8
|
||||
}
|
||||
const auto outputTarget = alpha ? fmt::format("{}.a", outVar) : outVar;
|
||||
const auto outputValue =
|
||||
alpha
|
||||
? fmt::format("f32((material * (lacc + (lacc >> {}))) >> {}) / 255.0", shift7, shift8)
|
||||
: fmt::format("vec4f(vec3f((material * (lacc + (lacc >> {}))) >> {}) / 255.0, 1.0)", shift7,
|
||||
shift8);
|
||||
alpha ? fmt::format("f32((material * (lacc + (lacc >> {}))) >> {}) / 255.0", shift7, shift8)
|
||||
: fmt::format("vec4f(vec3f((material * (lacc + (lacc >> {}))) >> {}) / 255.0, 1.0)", shift7, shift8);
|
||||
return fmt::format(R"""(
|
||||
{{
|
||||
var lighting = {11}(round({5}{8} * 255.0));
|
||||
@@ -993,6 +990,15 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
"\n let clip_base = select(clip_a, clip_b, use_b);"
|
||||
"\n out.pos = vec4f(clip_base.xy + offset_ndc * clip_base.w, clip_base.zw);";
|
||||
}
|
||||
if (config.exactScreenDepth) {
|
||||
vtxOutAttrs += fmt::format("\n @location({}) @interpolate(flat) exact_screen_depth: f32,", vtxOutIdx++);
|
||||
// Virtual-screen composition stores the original backend-convention NDC
|
||||
// depth in the otherwise replaceable projection Z row. Capture it before
|
||||
// parking raster depth at midrange; this avoids growing every GX uniform.
|
||||
vtxXfrAttrsPre +=
|
||||
"\n out.exact_screen_depth = clamp(out.pos.z, 0.0, 1.0);"
|
||||
"\n out.pos.z = -0.5 * out.pos.w;";
|
||||
}
|
||||
if constexpr (UseReversedZ) {
|
||||
vtxXfrAttrsPre += "\n out.pos.z = -out.pos.z;";
|
||||
} else {
|
||||
@@ -1019,8 +1025,8 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
std::array<bool, MaxTexCoord> emittedFixedTexCoord{};
|
||||
const auto fixed_texcoord_name = [&](u32 texCoordId) {
|
||||
if (!emittedFixedTexCoord[texCoordId]) {
|
||||
fragmentFnPre += fmt::format(
|
||||
"\n let tex{0}_fixed_uv = vec2i(tex{0}_uv * ubuf.texcoord{0}_scale.xy * 128.0);", texCoordId);
|
||||
fragmentFnPre +=
|
||||
fmt::format("\n let tex{0}_fixed_uv = vec2i(tex{0}_uv * ubuf.texcoord{0}_scale.xy * 128.0);", texCoordId);
|
||||
emittedFixedTexCoord[texCoordId] = true;
|
||||
}
|
||||
return fmt::format("tex{}_fixed_uv", texCoordId);
|
||||
@@ -1042,18 +1048,18 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
}
|
||||
{
|
||||
std::string_view outReg = regName[stage.colorOp.outReg];
|
||||
std::string op = tev_color_op(
|
||||
stage.colorOp.op, stage.colorOp.bias, stage.colorOp.scale, stage.colorOp.clamp, colorA, colorB,
|
||||
color_arg_reg(stage.colorPass.c, idx, config, stage), color_arg_reg(stage.colorPass.d, idx, config, stage));
|
||||
std::string op = tev_color_op(stage.colorOp.op, stage.colorOp.bias, stage.colorOp.scale, stage.colorOp.clamp,
|
||||
colorA, colorB, color_arg_reg(stage.colorPass.c, idx, config, stage),
|
||||
color_arg_reg(stage.colorPass.d, idx, config, stage));
|
||||
fragmentFn += fmt::format("\n {0} = vec4f({1}, {0}.a);", outReg, op);
|
||||
}
|
||||
{
|
||||
std::string_view outReg = regName[stage.alphaOp.outReg];
|
||||
std::string op = tev_alpha_op(
|
||||
stage.alphaOp.op, stage.alphaOp.bias, stage.alphaOp.scale, stage.alphaOp.clamp,
|
||||
alpha_arg_reg(stage.alphaPass.a, idx, config, stage), alpha_arg_reg(stage.alphaPass.b, idx, config, stage),
|
||||
alpha_arg_reg(stage.alphaPass.c, idx, config, stage), alpha_arg_reg(stage.alphaPass.d, idx, config, stage),
|
||||
colorA, colorB);
|
||||
std::string op = tev_alpha_op(stage.alphaOp.op, stage.alphaOp.bias, stage.alphaOp.scale, stage.alphaOp.clamp,
|
||||
alpha_arg_reg(stage.alphaPass.a, idx, config, stage),
|
||||
alpha_arg_reg(stage.alphaPass.b, idx, config, stage),
|
||||
alpha_arg_reg(stage.alphaPass.c, idx, config, stage),
|
||||
alpha_arg_reg(stage.alphaPass.d, idx, config, stage), colorA, colorB);
|
||||
fragmentFn += fmt::format("\n {0}.a = {1};", outReg, op);
|
||||
}
|
||||
}
|
||||
@@ -1170,8 +1176,7 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
vtxXfrAttrs += fmt::format("\n var tc{} = vec4f({}, 1.0, 1.0);", i,
|
||||
vtx_attr(config, GXAttr(GX_VA_TEX0 + (tcg.src - GX_TG_TEX0))));
|
||||
} else if (tcg.src == GX_MAX_TEXGENSRC) {
|
||||
vtxXfrAttrs += fmt::format("\n var tc{} = vec4f({}, 1.0, 1.0);", i,
|
||||
vtx_attr(config, GXAttr(GX_VA_TEX0 + i)));
|
||||
vtxXfrAttrs += fmt::format("\n var tc{} = vec4f({}, 1.0, 1.0);", i, vtx_attr(config, GXAttr(GX_VA_TEX0 + i)));
|
||||
} else if (tcg.src == GX_TG_POS) {
|
||||
vtxXfrAttrs += fmt::format("\n var tc{} = vec4f({}, 1.0);", i, vtx_attr(config, GX_VA_POS));
|
||||
} else if (tcg.src == GX_TG_NRM) {
|
||||
@@ -1198,10 +1203,9 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
vtxXfrAttrs += fmt::format("\n var tc{0}_tmp = tc{0}.xyz;", i);
|
||||
} else {
|
||||
const u32 texMtxSlot = (tcg.mtx) / 3;
|
||||
const u32 texMtxIdx = texMtxSlot < MaxPostexMtx ? info.matrixLayout.postexRemap[texMtxSlot]
|
||||
: texMtxSlot;
|
||||
CHECK(texMtxIdx != UniformMatrixLayout::kAbsent,
|
||||
"texgen {} matrix slot {} missing from the uniform layout", i, texMtxSlot);
|
||||
const u32 texMtxIdx = texMtxSlot < MaxPostexMtx ? info.matrixLayout.postexRemap[texMtxSlot] : texMtxSlot;
|
||||
CHECK(texMtxIdx != UniformMatrixLayout::kAbsent, "texgen {} matrix slot {} missing from the uniform layout", i,
|
||||
texMtxSlot);
|
||||
vtxXfrAttrs += fmt::format("\n var tc{0}_tmp = tc{0} * ubuf.postex_mtx[{1}];", i, texMtxIdx);
|
||||
}
|
||||
if (tcg.type == GX_TG_MTX2x4) {
|
||||
@@ -1395,16 +1399,16 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
u32 mtxIdx = stage.indTexMtxId - GX_ITM_S0;
|
||||
u32 regTexCoord = static_cast<u32>(textureDependency.texCoordId);
|
||||
const auto fixedUv = fixed_texcoord_name(regTexCoord);
|
||||
fragmentFnPre += fmt::format(
|
||||
"\n var ind{0}_offset_fixed = ({1} * vec2i(ind{0}_coord.x)) >> vec2u(8u);", i, fixedUv);
|
||||
fragmentFnPre +=
|
||||
fmt::format("\n var ind{0}_offset_fixed = ({1} * vec2i(ind{0}_coord.x)) >> vec2u(8u);", i, fixedUv);
|
||||
appendScaleShift(mtxIdx);
|
||||
} else if (stage.indTexMtxId >= GX_ITM_T0 && stage.indTexMtxId <= GX_ITM_T2 && hasBaseCoord) {
|
||||
// Dynamic T: (fixed UV * indirect T) >> 8, then exponent shift.
|
||||
u32 mtxIdx = stage.indTexMtxId - GX_ITM_T0;
|
||||
u32 regTexCoord = static_cast<u32>(textureDependency.texCoordId);
|
||||
const auto fixedUv = fixed_texcoord_name(regTexCoord);
|
||||
fragmentFnPre += fmt::format(
|
||||
"\n var ind{0}_offset_fixed = ({1} * vec2i(ind{0}_coord.y)) >> vec2u(8u);", i, fixedUv);
|
||||
fragmentFnPre +=
|
||||
fmt::format("\n var ind{0}_offset_fixed = ({1} * vec2i(ind{0}_coord.y)) >> vec2u(8u);", i, fixedUv);
|
||||
appendScaleShift(mtxIdx);
|
||||
}
|
||||
}
|
||||
@@ -1445,8 +1449,8 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
// ShaderInfo only emits texN_size_bias and texture bindings for a texture that can actually be sampled.
|
||||
if (willSampleTexture) {
|
||||
u32 texMapId = static_cast<u32>(textureDependency.texMapId);
|
||||
fragmentFnPre += fmt::format(
|
||||
"\n var ind{0}_uv = vec2f(t_TexCoord) / (ubuf.tex{1}_size_bias.xy * 128.0);", i, texMapId);
|
||||
fragmentFnPre +=
|
||||
fmt::format("\n var ind{0}_uv = vec2f(t_TexCoord) / (ubuf.tex{1}_size_bias.xy * 128.0);", i, texMapId);
|
||||
uvIn = fmt::format("ind{0}_uv", i);
|
||||
}
|
||||
}
|
||||
@@ -1466,8 +1470,7 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
}
|
||||
|
||||
std::string fogDepthExpr = UseReversedZ ? "in.pos.z" : "(1.0 - in.pos.z)";
|
||||
std::string fogZCoordExpr =
|
||||
fmt::format("u32(round(clamp({}, 0.0, 1.0) * 16777216.0))", fogDepthExpr);
|
||||
std::string fogZCoordExpr = fmt::format("u32(round(clamp({}, 0.0, 1.0) * 16777216.0))", fogDepthExpr);
|
||||
if (usesZTextureDepth) {
|
||||
const u32 zTexBias = config.zTexture & 0x00FFFFFFu;
|
||||
const u32 zTexFmt = (config.zTexture >> 24) & 0x3u;
|
||||
@@ -1623,15 +1626,16 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
if (discard.constant == 1) {
|
||||
fragmentFn += "\n // Alpha compare\n discard;";
|
||||
} else if (discard.constant != 0) {
|
||||
fragmentFn += "\n // Alpha compare"
|
||||
"\n let alphaCompare = u32(round(clamp(prev.a, 0.0, 1.0) * 255.0));";
|
||||
fragmentFn +=
|
||||
"\n // Alpha compare"
|
||||
"\n let alphaCompare = u32(round(clamp(prev.a, 0.0, 1.0) * 255.0));";
|
||||
fragmentFn += fmt::format("\n if ({}) {{ discard; }}", discard.expr);
|
||||
}
|
||||
}
|
||||
|
||||
std::string fragmentReturnType = "@location(0) vec4f";
|
||||
std::string fragmentReturn = " return prev;";
|
||||
if (usesZTextureDepth) {
|
||||
if (usesZTextureDepth || config.exactScreenDepth) {
|
||||
uniformPre +=
|
||||
"\n"
|
||||
"struct FragmentOutput {\n"
|
||||
@@ -1639,7 +1643,15 @@ wgpu::ShaderModule build_shader(const ShaderConfig& config) noexcept {
|
||||
" @builtin(frag_depth) depth: f32,\n"
|
||||
"};";
|
||||
|
||||
fragmentFn += fmt::format("\n let fragDepth = {}ztexDepth;", UseReversedZ ? "" : "1.0 - ");
|
||||
if (config.exactScreenDepth) {
|
||||
// Fragment depth is already in window coordinates. Reapply the recorded
|
||||
// viewport's clamped depth window exactly as the fixed pipeline did.
|
||||
fragmentFn +=
|
||||
"\n let fragDepth = ubuf.exact_screen_depth_range.x + "
|
||||
"in.exact_screen_depth * ubuf.exact_screen_depth_range.y;";
|
||||
} else {
|
||||
fragmentFn += fmt::format("\n let fragDepth = {}ztexDepth;", UseReversedZ ? "" : "1.0 - ");
|
||||
}
|
||||
fragmentReturnType = "FragmentOutput";
|
||||
fragmentReturn =
|
||||
" var out: FragmentOutput;\n"
|
||||
@@ -2022,7 +2034,7 @@ struct Uniform {{
|
||||
current_pnmtx: u32,
|
||||
render_viewport_size: vec2f,
|
||||
logical_viewport_size: vec2f,
|
||||
pad: vec2u,
|
||||
exact_screen_depth_range: vec2f,
|
||||
array_start: array<u32, 12>,{0}
|
||||
}};
|
||||
@group(0) @binding(0)
|
||||
@@ -2050,8 +2062,7 @@ fn fs_main(in: VertexOutput) -> {9} {{{6}{5}
|
||||
}}
|
||||
)""",
|
||||
uniBufAttrs, texBindings, vtxOutAttrs, vtxInAttrs, vtxXfrAttrs, fragmentFn,
|
||||
fragmentFnPre, vtxXfrAttrsPre, uniformPre, fragmentReturnType,
|
||||
fragmentReturn);
|
||||
fragmentFnPre, vtxXfrAttrsPre, uniformPre, fragmentReturnType, fragmentReturn);
|
||||
wgpu::ShaderSourceWGSL wgslDescriptor{};
|
||||
wgslDescriptor.code = shaderSource.c_str();
|
||||
const auto label = fmt::format("GX Shader {:x}", hash);
|
||||
|
||||
@@ -61,8 +61,7 @@ Fog fog_uniform() {
|
||||
rangeWidth = 1.0f;
|
||||
}
|
||||
const int rangeCenter = static_cast<int>(rangeBase & 0x3ffu) - 342;
|
||||
const float screenSpaceCenter = rangeEnabled ? ((static_cast<float>(rangeCenter) / rangeWidth) * 2.0f) - 1.0f
|
||||
: 0.0f;
|
||||
const float screenSpaceCenter = rangeEnabled ? ((static_cast<float>(rangeCenter) / rangeWidth) * 2.0f) - 1.0f : 0.0f;
|
||||
fog.rangeBase = {screenSpaceCenter, rangeEnabled ? rangeWidth : 1.0f, 0.0f, 0.0f};
|
||||
|
||||
std::array<float, 12> rangeK{};
|
||||
@@ -129,8 +128,8 @@ Vec4<float> texture_size_bias(const gfx::TextureBind& tex) {
|
||||
}
|
||||
|
||||
Vec4<float> texcoord_scale(const TexCoordScale& scale) {
|
||||
return {static_cast<float>(scale.scaleS) + 1.0f, static_cast<float>(scale.scaleT) + 1.0f,
|
||||
scale.biasS ? 1.0f : 0.0f, scale.biasT ? 1.0f : 0.0f};
|
||||
return {static_cast<float>(scale.scaleS) + 1.0f, static_cast<float>(scale.scaleT) + 1.0f, scale.biasS ? 1.0f : 0.0f,
|
||||
scale.biasT ? 1.0f : 0.0f};
|
||||
}
|
||||
|
||||
void mark_texture_sample(const TevStageTextureDependency& dependency, ShaderInfo& info) {
|
||||
@@ -149,8 +148,8 @@ void mark_tev_reg_read(GXTevRegID reg, bool alpha, ShaderInfo& info) {
|
||||
}
|
||||
}
|
||||
|
||||
void color_arg_reg_info(GXTevColorArg arg, const TevStage& stage,
|
||||
const TevStageTextureDependency& textureDependency, ShaderInfo& info) {
|
||||
void color_arg_reg_info(GXTevColorArg arg, const TevStage& stage, const TevStageTextureDependency& textureDependency,
|
||||
ShaderInfo& info) {
|
||||
switch (arg) {
|
||||
case GX_CC_CPREV:
|
||||
mark_tev_reg_read(GX_TEVPREV, false, info);
|
||||
@@ -229,8 +228,8 @@ void color_arg_reg_info(GXTevColorArg arg, const TevStage& stage,
|
||||
}
|
||||
}
|
||||
|
||||
void alpha_arg_reg_info(GXTevAlphaArg arg, const TevStage& stage,
|
||||
const TevStageTextureDependency& textureDependency, ShaderInfo& info) {
|
||||
void alpha_arg_reg_info(GXTevAlphaArg arg, const TevStage& stage, const TevStageTextureDependency& textureDependency,
|
||||
ShaderInfo& info) {
|
||||
switch (arg) {
|
||||
case GX_CA_APREV:
|
||||
mark_tev_reg_read(GX_TEVPREV, true, info);
|
||||
@@ -296,7 +295,8 @@ ShaderInfo build_shader_info(const ShaderConfig& config) noexcept {
|
||||
ZoneScoped;
|
||||
|
||||
ShaderInfo info{
|
||||
// vtx_start, current_pnmtx, render/logical viewport size, array_start, pad, proj
|
||||
// vtx_start, current_pnmtx, render/logical viewport size, depth range,
|
||||
// array_start, proj
|
||||
.uniformSize = 4 + 4 + 8 + 8 + 8 + 48 + 64,
|
||||
};
|
||||
|
||||
@@ -435,7 +435,8 @@ ShaderInfo build_shader_info(const ShaderConfig& config) noexcept {
|
||||
continue;
|
||||
}
|
||||
if (info.indexAttr.test(GX_VA_TEX0MTXIDX + i)) {
|
||||
// A per-vertex texture matrix index addresses raw matrix memory and can name any row, including the position rows.
|
||||
// A per-vertex texture matrix index addresses raw matrix memory and can name any row, including the position
|
||||
// rows.
|
||||
dynamicTexMtx = true;
|
||||
continue;
|
||||
}
|
||||
@@ -444,14 +445,16 @@ ShaderInfo build_shader_info(const ShaderConfig& config) noexcept {
|
||||
}
|
||||
const u32 slot = static_cast<u32>(tcg.mtx) / 3;
|
||||
if (slot >= MaxPostexMtx) {
|
||||
// Out of range for the shader array; preserve the uncompacted layout rather than inventing a different (still invalid) index.
|
||||
// Out of range for the shader array; preserve the uncompacted layout rather than inventing a different (still
|
||||
// invalid) index.
|
||||
dynamicTexMtx = true;
|
||||
continue;
|
||||
}
|
||||
literalSlots.set(slot);
|
||||
}
|
||||
|
||||
// At most MaxPostexMtx entries can ever be pushed: the absolute layout pushes the ten position slots plus at most ten texture slots, and the compacted layout pushes the current matrix plus at most one literal slot per texgen.
|
||||
// At most MaxPostexMtx entries can ever be pushed: the absolute layout pushes the ten position slots plus at most
|
||||
// ten texture slots, and the compacted layout pushes the current matrix plus at most one literal slot per texgen.
|
||||
const auto pushPostex = [&](u8 slot) {
|
||||
CHECK(layout.postexCount < layout.postexSlots.size(), "postex matrix layout overflow");
|
||||
layout.postexSlots[layout.postexCount++] = slot;
|
||||
@@ -463,7 +466,8 @@ ShaderInfo build_shader_info(const ShaderConfig& config) noexcept {
|
||||
|
||||
layout.absolutePosRegion = dynamicPnMtx || dynamicTexMtx;
|
||||
if (!layout.absolutePosRegion) {
|
||||
// Compact slot 0 always holds the matrix selected by the current matrix index; build_uniform writes 0 into `current_pnmtx` to match.
|
||||
// Compact slot 0 always holds the matrix selected by the current matrix index; build_uniform writes 0 into
|
||||
// `current_pnmtx` to match.
|
||||
pushPostex(UniformMatrixLayout::kCurrentPnMtx);
|
||||
pushNrm(UniformMatrixLayout::kCurrentPnMtx);
|
||||
}
|
||||
@@ -474,7 +478,8 @@ ShaderInfo build_shader_info(const ShaderConfig& config) noexcept {
|
||||
layout.postexRemap[slot] = layout.postexCount;
|
||||
pushPostex(static_cast<u8>(slot));
|
||||
if (layout.absolutePosRegion) {
|
||||
// The normal array is indexed by the same expression as the position array, so its compact indices have to line up with the position slots.
|
||||
// The normal array is indexed by the same expression as the position array, so its compact indices have to line
|
||||
// up with the position slots.
|
||||
pushNrm(static_cast<u8>(slot));
|
||||
}
|
||||
}
|
||||
@@ -485,7 +490,8 @@ ShaderInfo build_shader_info(const ShaderConfig& config) noexcept {
|
||||
layout.postexRemap[slot] = layout.postexCount;
|
||||
pushPostex(static_cast<u8>(slot));
|
||||
}
|
||||
// `postex_mtx[in_pnmtxidx]` and `nrm_mtx[in_pnmtxidx]` are emitted unconditionally, so neither array can ever be empty.
|
||||
// `postex_mtx[in_pnmtxidx]` and `nrm_mtx[in_pnmtxidx]` are emitted unconditionally, so neither array can ever be
|
||||
// empty.
|
||||
info.uniformSize += sizeof(Mat3x4<float>) * (layout.postexCount + layout.nrmCount);
|
||||
}
|
||||
if (config.fogType != GX_FOG_NONE) {
|
||||
@@ -543,9 +549,14 @@ static u32 line_texcoord_mask() noexcept {
|
||||
}
|
||||
|
||||
namespace {
|
||||
// Largest possible staged prefix: scalar head (80) + line/point block (16) + projection matrix + every matrix the uncompacted layout can hold.
|
||||
constexpr size_t kStagedUniformBytes =
|
||||
96 + sizeof(Mat4x4<float>) + sizeof(Mat3x4<float>) * (MaxPostexMtx + MaxPnMtx);
|
||||
// GX's fixed EFB workspace, the reference the effect-buffer test below reduces
|
||||
// against.
|
||||
constexpr u16 kEfbWidth = 640;
|
||||
constexpr u16 kEfbHeight = 528;
|
||||
|
||||
// Largest possible staged prefix: scalar head (80) + line/point block (16) + projection matrix + every matrix the
|
||||
// uncompacted layout can hold.
|
||||
constexpr size_t kStagedUniformBytes = 96 + sizeof(Mat4x4<float>) + sizeof(Mat3x4<float>) * (MaxPostexMtx + MaxPnMtx);
|
||||
|
||||
// The host viewport always receives the normalized GX depth window (render_pass_impl clamps to minDepth <= maxDepth).
|
||||
static Mat4x4<float> effective_projection() noexcept {
|
||||
@@ -569,7 +580,8 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
|
||||
auto [buf, range] = gfx::map_uniform(info.uniformSize);
|
||||
|
||||
const auto& layout = info.matrixLayout;
|
||||
// `postex_mtx[current_pnmtx]` indexes raw matrix memory, so a current-matrix index of 10 or more selects a texture matrix; the normal array only covers the position rows and clamps.
|
||||
// `postex_mtx[current_pnmtx]` indexes raw matrix memory, so a current-matrix index of 10 or more selects a texture
|
||||
// matrix; the normal array only covers the position rows and clamps.
|
||||
const u32 currentPostexSlot = std::min<u32>(g_gxState.currentPnMtx, MaxPostexMtx - 1);
|
||||
const u32 currentNrmSlot = std::min<u32>(g_gxState.currentPnMtx, MaxPnMtx - 1);
|
||||
|
||||
@@ -583,15 +595,23 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
|
||||
const auto stage_u32 = [&](u32 value) noexcept { stage(&value, sizeof(value)); };
|
||||
const auto stage_f32 = [&](f32 value) noexcept { stage(&value, sizeof(value)); };
|
||||
|
||||
const Mat4x4<float> effectiveProj = effective_projection();
|
||||
const auto& viewport = g_gxState.renderViewport;
|
||||
const float depthNear = std::clamp(std::min(viewport.znear, viewport.zfar), 0.0f, 1.0f);
|
||||
const float depthFar = std::clamp(std::max(viewport.znear, viewport.zfar), 0.0f, 1.0f);
|
||||
|
||||
stage_u32(vtxStart);
|
||||
// With a compacted position region the live matrix is uploaded to slot 0, so every `postex_mtx[in_pnmtxidx]` / `nrm_mtx[in_pnmtxidx]` read has to resolve to 0 as well.
|
||||
// With a compacted position region the live matrix is uploaded to slot 0, so every `postex_mtx[in_pnmtxidx]` /
|
||||
// `nrm_mtx[in_pnmtxidx]` read has to resolve to 0 as well.
|
||||
stage_u32(layout.absolutePosRegion ? g_gxState.currentPnMtx : 0u);
|
||||
stage_f32(g_gxState.renderViewport.width);
|
||||
stage_f32(g_gxState.renderViewport.height);
|
||||
stage_f32(g_gxState.logicalViewport.width);
|
||||
stage_f32(g_gxState.logicalViewport.height);
|
||||
std::memset(staged.data() + stagedSize, 0, 8); // pad
|
||||
stagedSize += 8;
|
||||
// Fragment-depth output bypasses the fixed viewport transform, so exact
|
||||
// screen depth applies that same clamped window explicitly.
|
||||
stage_f32(depthNear);
|
||||
stage_f32(depthFar - depthNear);
|
||||
for (const auto& vaRange : ranges.vaRanges) {
|
||||
stage_u32(vaRange.offset);
|
||||
}
|
||||
@@ -609,26 +629,22 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
|
||||
}
|
||||
}
|
||||
const size_t projectionOffset = stagedSize;
|
||||
const Mat4x4<float> effectiveProj = effective_projection();
|
||||
stage(&effectiveProj, sizeof(effectiveProj));
|
||||
|
||||
const size_t positionOffset = stagedSize;
|
||||
uint32_t positionMatrixMask = 0;
|
||||
for (u32 i = 0; i < layout.postexCount; ++i) {
|
||||
const u32 slot =
|
||||
layout.postexSlots[i] == UniformMatrixLayout::kCurrentPnMtx ? currentPostexSlot
|
||||
: layout.postexSlots[i];
|
||||
layout.postexSlots[i] == UniformMatrixLayout::kCurrentPnMtx ? currentPostexSlot : layout.postexSlots[i];
|
||||
if (slot < MaxPnMtx) {
|
||||
positionMatrixMask |= 1u << i;
|
||||
}
|
||||
stage(slot < MaxPnMtx ? &g_gxState.pnMtx[slot].pos : &g_gxState.texMtxs[slot - MaxPnMtx],
|
||||
sizeof(Mat3x4<float>));
|
||||
stage(slot < MaxPnMtx ? &g_gxState.pnMtx[slot].pos : &g_gxState.texMtxs[slot - MaxPnMtx], sizeof(Mat3x4<float>));
|
||||
}
|
||||
|
||||
const size_t normalOffset = stagedSize;
|
||||
for (u32 i = 0; i < layout.nrmCount; ++i) {
|
||||
const u32 slot = layout.nrmSlots[i] == UniformMatrixLayout::kCurrentPnMtx ? currentNrmSlot
|
||||
: layout.nrmSlots[i];
|
||||
const u32 slot = layout.nrmSlots[i] == UniformMatrixLayout::kCurrentPnMtx ? currentNrmSlot : layout.nrmSlots[i];
|
||||
stage(&g_gxState.pnMtx[slot].nrm, sizeof(Mat3x4<float>));
|
||||
}
|
||||
buf.append(staged.data(), stagedSize);
|
||||
@@ -640,7 +656,8 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
|
||||
}
|
||||
}
|
||||
if (info.lightingEnabled) {
|
||||
// Sanitizing and normalizing light directions is substantially more expensive than copying the uniform data, while lights normally remain unchanged across many draws.
|
||||
// Sanitizing and normalizing light directions is substantially more expensive than copying the uniform data, while
|
||||
// lights normally remain unchanged across many draws.
|
||||
static_assert(sizeof(g_gxState.lights) == 80 * GX::MaxLights);
|
||||
if (g_gxState.preparedLightsDirty) {
|
||||
for (size_t i = 0; i < g_gxState.preparedLights.size(); ++i) {
|
||||
@@ -702,14 +719,28 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
|
||||
buf.append(texcoord_scale(g_gxState.texCoordScales[i]));
|
||||
}
|
||||
}
|
||||
// A freshly produced, downscaled EFB copy is an effect buffer by construction
|
||||
// (Mario Kart Wii's bloom chain reduces to 128x128 and 64x64), and a fresh
|
||||
// full-resolution one blended back in is a blur or haze pass. Older one-shot
|
||||
// copies are persistent game art (notably MKW's baked minimap), so they stay
|
||||
// eligible for the virtual screen even when small and alpha blended.
|
||||
bool samplesRecentEfbCopy = false;
|
||||
bool samplesReducedEfbCopy = false;
|
||||
const u32 currentFrame = gfx::current_frame();
|
||||
for (int i = 0; i < info.sampledTextures.size(); ++i) {
|
||||
if (!info.sampledTextures.test(i)) {
|
||||
continue;
|
||||
}
|
||||
const auto& tex = get_texture(static_cast<GXTexMapID>(i));
|
||||
// CHECK(tex, "unbound texture {}", i);
|
||||
if (tex.ref && tex.ref->is_recent_efb_copy(currentFrame)) {
|
||||
samplesRecentEfbCopy = true;
|
||||
samplesReducedEfbCopy =
|
||||
samplesReducedEfbCopy || tex.texObj.width() * 2 <= kEfbWidth || tex.texObj.height() * 2 <= kEfbHeight;
|
||||
}
|
||||
buf.append(texture_size_bias(tex));
|
||||
}
|
||||
const bool blends = g_gxState.blendMode == GX_BM_BLEND || g_gxState.blendMode == GX_BM_SUBTRACT;
|
||||
|
||||
const UniformReplayLayout replayLayout{
|
||||
.projectionOffset = static_cast<uint32_t>(projectionOffset),
|
||||
@@ -719,6 +750,7 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
|
||||
.positionMatrixCount = layout.postexCount,
|
||||
.normalMatrixCount = layout.nrmCount,
|
||||
.perspective = perspective,
|
||||
.nativeEfbEffect = !perspective && samplesRecentEfbCopy && (samplesReducedEfbCopy || blends),
|
||||
};
|
||||
|
||||
if (!perspective || frame_interpolation_fps() == 0) {
|
||||
@@ -739,9 +771,7 @@ UniformRanges build_uniform(const ShaderInfo& info, u32 vtxStart, const BindGrou
|
||||
.positionOffset = positionOffset,
|
||||
.normalOffset = normalOffset,
|
||||
// A compacted position region holds the current matrix at slot 0.
|
||||
.currentMatrix = layout.absolutePosRegion
|
||||
? std::min<size_t>(g_gxState.currentPnMtx, MaxPnMtx - 1)
|
||||
: 0,
|
||||
.currentMatrix = layout.absolutePosRegion ? std::min<size_t>(g_gxState.currentPnMtx, MaxPnMtx - 1) : 0,
|
||||
.indexedMatrices = info.indexAttr.test(GX_VA_PNMTXIDX),
|
||||
});
|
||||
g_gxState.stateDirty = false;
|
||||
|
||||
@@ -13,6 +13,10 @@ struct UniformReplayLayout {
|
||||
uint8_t positionMatrixCount = 0;
|
||||
uint8_t normalMatrixCount = 0;
|
||||
bool perspective = false;
|
||||
// A 2D draw compositing the framebuffer back over itself: bloom, blur and the
|
||||
// rest of the native post-processing chain. It belongs to the rendered image,
|
||||
// not to the game's 2D layer, so it must stay where the game aimed it.
|
||||
bool nativeEfbEffect = false;
|
||||
};
|
||||
|
||||
struct UniformRanges {
|
||||
|
||||
@@ -134,6 +134,7 @@ std::chrono::nanoseconds wait_for_frame_worker_sealed() noexcept;
|
||||
bool wait_for_frame_worker_for(std::chrono::microseconds timeout) noexcept;
|
||||
void quiesce_frame_worker() noexcept;
|
||||
std::recursive_mutex& renderer_gpu_mutex() noexcept;
|
||||
bool stereo_frame_provider_active() noexcept;
|
||||
|
||||
template <typename T>
|
||||
class ArrayRef {
|
||||
|
||||
@@ -4,6 +4,7 @@
|
||||
#include "gx_test_common.hpp"
|
||||
#include "gfx/efb_ram_encoder.hpp"
|
||||
#include "gfx/tex_copy_format_contract.hpp"
|
||||
#include "gfx/texture.hpp"
|
||||
#include "gx/shader_info.hpp"
|
||||
#include "gx/pipeline.hpp"
|
||||
#include "__gx.h"
|
||||
@@ -24,19 +25,16 @@ TEST(GXPipelineConfig, RejectsInvalidDeserializedEnums) {
|
||||
}
|
||||
|
||||
TEST(GXPipelineConfig, FoggedLateZLogicOrPreservesEggDofMask) {
|
||||
EXPECT_EQ(aurora::gx::effective_pipeline_fog_type(GX_FOG_PERSP_LIN, GX_ZT_REPLACE,
|
||||
false, GX_BM_LOGIC, GX_LO_OR),
|
||||
EXPECT_EQ(aurora::gx::effective_pipeline_fog_type(GX_FOG_PERSP_LIN, GX_ZT_REPLACE, false, GX_BM_LOGIC, GX_LO_OR),
|
||||
GX_FOG_PERSP_LIN);
|
||||
|
||||
// Keep the existing defensive suppression for late-Z mask passes whose
|
||||
// integer logic operation still has only an inexact WebGPU fallback.
|
||||
EXPECT_EQ(aurora::gx::effective_pipeline_fog_type(GX_FOG_PERSP_LIN, GX_ZT_REPLACE,
|
||||
false, GX_BM_LOGIC, GX_LO_COPY),
|
||||
EXPECT_EQ(aurora::gx::effective_pipeline_fog_type(GX_FOG_PERSP_LIN, GX_ZT_REPLACE, false, GX_BM_LOGIC, GX_LO_COPY),
|
||||
GX_FOG_NONE);
|
||||
|
||||
// Logic fallbacks do not require this workaround when Z-texturing is early.
|
||||
EXPECT_EQ(aurora::gx::effective_pipeline_fog_type(GX_FOG_PERSP_LIN, GX_ZT_REPLACE,
|
||||
true, GX_BM_LOGIC, GX_LO_COPY),
|
||||
EXPECT_EQ(aurora::gx::effective_pipeline_fog_type(GX_FOG_PERSP_LIN, GX_ZT_REPLACE, true, GX_BM_LOGIC, GX_LO_COPY),
|
||||
GX_FOG_PERSP_LIN);
|
||||
}
|
||||
|
||||
@@ -66,6 +64,31 @@ TEST(GXShaderInfo, ShaderLightSanitizesNonFiniteDirectionForUniforms) {
|
||||
EXPECT_FLOAT_EQ(zero.dir[2], 0.0f);
|
||||
}
|
||||
|
||||
TEST_F(GXFifoTest, UniformRetainsViewportWindowForVrReplay) {
|
||||
gxState().proj = {};
|
||||
gxState().proj.m0 = {1.0f, 0.0f, 0.0f, 0.0f};
|
||||
gxState().proj.m1 = {0.0f, 1.0f, 0.0f, 0.0f};
|
||||
gxState().proj.m2 = {0.0f, 0.0f, -0.001f, -0.5f};
|
||||
gxState().proj.m3 = {0.0f, 0.0f, 0.0f, 1.0f};
|
||||
gxState().renderViewport.znear = 0.2f;
|
||||
gxState().renderViewport.zfar = 0.8f;
|
||||
|
||||
aurora::gfx::testing::reset_uniform_allocations();
|
||||
const auto info = aurora::gx::build_shader_info({});
|
||||
const auto ranges = aurora::gx::build_uniform(info, 0, aurora::gx::BindGroupRanges{},
|
||||
aurora::gx::FrameInterpolationDrawIdentity{}, false);
|
||||
const auto& bytes = aurora::gfx::testing::uniform_allocation(ranges.current.offset);
|
||||
|
||||
const auto read_float = [&](size_t offset) {
|
||||
float value = 0.0f;
|
||||
std::memcpy(&value, bytes.data() + offset, sizeof(value));
|
||||
return value;
|
||||
};
|
||||
EXPECT_FLOAT_EQ(read_float(24), 0.2f);
|
||||
EXPECT_FLOAT_EQ(read_float(28), 0.6f);
|
||||
EXPECT_EQ(ranges.replayLayout.projectionOffset, 80u);
|
||||
}
|
||||
|
||||
TEST(GXLighting, SpotCoefficientsAndPositionGetterMatchRevolutionSdk) {
|
||||
GXLightObj light{};
|
||||
GXInitLightSpot(&light, 60.0f, GX_SP_SHARP);
|
||||
@@ -104,9 +127,8 @@ TEST(EfbRamEncoderContract, Z24X8WritesNativeFourByFourArGbPlanes) {
|
||||
}
|
||||
|
||||
std::array<u8, 64> encoded{};
|
||||
ASSERT_TRUE(aurora::gfx::efb_ram::encode(
|
||||
encoded.data(), encoded.size(), GX_TF_Z24X8, 4, 4, rgba.data(), 4, 4, 16,
|
||||
aurora::gfx::efb_ram::HostPixelOrder::RGBA));
|
||||
ASSERT_TRUE(aurora::gfx::efb_ram::encode(encoded.data(), encoded.size(), GX_TF_Z24X8, 4, 4, rgba.data(), 4, 4, 16,
|
||||
aurora::gfx::efb_ram::HostPixelOrder::RGBA));
|
||||
EXPECT_EQ(aurora::gfx::efb_ram::encoded_size(GX_TF_Z24X8, 4, 4), encoded.size());
|
||||
|
||||
for (u32 i = 0; i < 16; ++i) {
|
||||
@@ -115,8 +137,7 @@ TEST(EfbRamEncoderContract, Z24X8WritesNativeFourByFourArGbPlanes) {
|
||||
EXPECT_EQ(encoded[32 + i * 2 + 0], 0x40 + i) << "GB plane texel " << i;
|
||||
EXPECT_EQ(encoded[32 + i * 2 + 1], 0x80 + i) << "GB plane texel " << i;
|
||||
|
||||
const u32 depth = (static_cast<u32>(encoded[i * 2 + 1]) << 16) |
|
||||
(static_cast<u32>(encoded[32 + i * 2 + 0]) << 8) |
|
||||
const u32 depth = (static_cast<u32>(encoded[i * 2 + 1]) << 16) | (static_cast<u32>(encoded[32 + i * 2 + 0]) << 8) |
|
||||
encoded[32 + i * 2 + 1];
|
||||
EXPECT_EQ(depth, ((0x10u + i) << 16) | ((0x40u + i) << 8) | (0x80u + i));
|
||||
}
|
||||
@@ -146,8 +167,12 @@ static std::vector<u8> bp_cmd(u8 reg, u32 value) {
|
||||
}
|
||||
|
||||
static std::vector<u8> cp_cmd(u8 reg, u32 value) {
|
||||
return {0x08, reg, static_cast<u8>((value >> 24) & 0xFF), static_cast<u8>((value >> 16) & 0xFF),
|
||||
static_cast<u8>((value >> 8) & 0xFF), static_cast<u8>(value & 0xFF)};
|
||||
return {0x08,
|
||||
reg,
|
||||
static_cast<u8>((value >> 24) & 0xFF),
|
||||
static_cast<u8>((value >> 16) & 0xFF),
|
||||
static_cast<u8>((value >> 8) & 0xFF),
|
||||
static_cast<u8>(value & 0xFF)};
|
||||
}
|
||||
|
||||
static std::vector<u8> xf_cmd(u16 addr, std::initializer_list<u32> values) {
|
||||
@@ -180,8 +205,8 @@ TEST(FrameInterpolationContract, RequiresStablePerspectiveDrawSequence) {
|
||||
aurora::gx::set_frame_interpolation_fps(0);
|
||||
aurora::gx::begin_frame_interpolation();
|
||||
};
|
||||
const auto buildFrame = [](const aurora::gx::ShaderInfo& info, aurora::HashType signatureBase,
|
||||
bool perspective, bool reverse = false) {
|
||||
const auto buildFrame = [](const aurora::gx::ShaderInfo& info, aurora::HashType signatureBase, bool perspective,
|
||||
bool reverse = false) {
|
||||
aurora::gx::begin_frame_interpolation();
|
||||
aurora::gx::BindGroupRanges ranges{};
|
||||
uint32_t interpolatedUniforms = 0;
|
||||
@@ -193,9 +218,9 @@ TEST(FrameInterpolationContract, RequiresStablePerspectiveDrawSequence) {
|
||||
.texture = 1,
|
||||
};
|
||||
const auto uniforms = aurora::gx::build_uniform(info, 0, ranges, identity, perspective);
|
||||
interpolatedUniforms += static_cast<uint32_t>(std::count_if(
|
||||
uniforms.interpolated.begin(), uniforms.interpolated.end(),
|
||||
[](const aurora::gfx::Range& range) { return range.size != 0; }));
|
||||
interpolatedUniforms +=
|
||||
static_cast<uint32_t>(std::count_if(uniforms.interpolated.begin(), uniforms.interpolated.end(),
|
||||
[](const aurora::gfx::Range& range) { return range.size != 0; }));
|
||||
}
|
||||
aurora::gx::finalize_frame_interpolation();
|
||||
return interpolatedUniforms;
|
||||
@@ -258,9 +283,9 @@ TEST(FrameInterpolationContract, RequiresStablePerspectiveDrawSequence) {
|
||||
.texture = 7,
|
||||
};
|
||||
const auto uniforms = aurora::gx::build_uniform(info, 0, ranges, identity, true);
|
||||
interpolatedUniforms += static_cast<uint32_t>(std::count_if(
|
||||
uniforms.interpolated.begin(), uniforms.interpolated.end(),
|
||||
[](const aurora::gfx::Range& range) { return range.size != 0; }));
|
||||
interpolatedUniforms +=
|
||||
static_cast<uint32_t>(std::count_if(uniforms.interpolated.begin(), uniforms.interpolated.end(),
|
||||
[](const aurora::gfx::Range& range) { return range.size != 0; }));
|
||||
}
|
||||
aurora::gx::finalize_frame_interpolation();
|
||||
return interpolatedUniforms;
|
||||
@@ -332,20 +357,15 @@ TEST(FrameInterpolationContract, IndexedPaletteInterpolationPreservesSharedSeams
|
||||
|
||||
aurora::Mat3x4<float> rotatedMidpoint{};
|
||||
aurora::Mat3x4<float> translatedMidpoint{};
|
||||
ASSERT_TRUE(aurora::gx::interpolate_indexed_transform(
|
||||
identity, quarterTurn, 0.5f, rotatedMidpoint));
|
||||
ASSERT_TRUE(aurora::gx::interpolate_indexed_transform(
|
||||
translatedPrevious, translatedCurrent, 0.5f, translatedMidpoint));
|
||||
ASSERT_TRUE(aurora::gx::interpolate_indexed_transform(identity, quarterTurn, 0.5f, rotatedMidpoint));
|
||||
ASSERT_TRUE(
|
||||
aurora::gx::interpolate_indexed_transform(translatedPrevious, translatedCurrent, 0.5f, translatedMidpoint));
|
||||
|
||||
const auto transformPoint = [](const aurora::Mat3x4<float>& matrix,
|
||||
const std::array<float, 3>& point) {
|
||||
const auto transformPoint = [](const aurora::Mat3x4<float>& matrix, const std::array<float, 3>& point) {
|
||||
return std::array<float, 3>{
|
||||
matrix.m0.x() * point[0] + matrix.m0.y() * point[1] +
|
||||
matrix.m0.z() * point[2] + matrix.m0.w(),
|
||||
matrix.m1.x() * point[0] + matrix.m1.y() * point[1] +
|
||||
matrix.m1.z() * point[2] + matrix.m1.w(),
|
||||
matrix.m2.x() * point[0] + matrix.m2.y() * point[1] +
|
||||
matrix.m2.z() * point[2] + matrix.m2.w(),
|
||||
matrix.m0.x() * point[0] + matrix.m0.y() * point[1] + matrix.m0.z() * point[2] + matrix.m0.w(),
|
||||
matrix.m1.x() * point[0] + matrix.m1.y() * point[1] + matrix.m1.z() * point[2] + matrix.m1.w(),
|
||||
matrix.m2.x() * point[0] + matrix.m2.y() * point[1] + matrix.m2.z() * point[2] + matrix.m2.w(),
|
||||
};
|
||||
};
|
||||
|
||||
@@ -367,8 +387,7 @@ TEST(FrameInterpolationContract, IndexedPaletteInterpolationPreservesSharedSeams
|
||||
{0.0f, 0.0f, 1.0f, 0.0f},
|
||||
};
|
||||
aurora::Mat3x4<float> shearedMidpoint{};
|
||||
ASSERT_TRUE(aurora::gx::interpolate_indexed_transform(
|
||||
identity, sheared, 0.5f, shearedMidpoint));
|
||||
ASSERT_TRUE(aurora::gx::interpolate_indexed_transform(identity, sheared, 0.5f, shearedMidpoint));
|
||||
EXPECT_FLOAT_EQ(shearedMidpoint.m0.y(), 0.25f);
|
||||
EXPECT_FLOAT_EQ(shearedMidpoint.m0.w(), 2.0f);
|
||||
}
|
||||
@@ -377,10 +396,8 @@ TEST(FrameInterpolationContract, IndexedPaletteHistoryKeepsAbsoluteVertexSlots)
|
||||
constexpr uint16_t usedMask = (1u << 0) | (1u << 1);
|
||||
constexpr size_t projectionOffset = 0;
|
||||
constexpr size_t positionOffset = sizeof(aurora::Mat4x4<float>);
|
||||
constexpr size_t normalOffset =
|
||||
positionOffset + aurora::gx::MaxPnMtx * sizeof(aurora::Mat3x4<float>);
|
||||
constexpr size_t uniformSize =
|
||||
normalOffset + aurora::gx::MaxPnMtx * sizeof(aurora::Mat3x4<float>);
|
||||
constexpr size_t normalOffset = positionOffset + aurora::gx::MaxPnMtx * sizeof(aurora::Mat3x4<float>);
|
||||
constexpr size_t uniformSize = normalOffset + aurora::gx::MaxPnMtx * sizeof(aurora::Mat3x4<float>);
|
||||
const aurora::gx::FrameInterpolationDrawIdentity identity{
|
||||
.combined = 0x1234,
|
||||
.pipeline = 0x5678,
|
||||
@@ -395,32 +412,28 @@ TEST(FrameInterpolationContract, IndexedPaletteHistoryKeepsAbsoluteVertexSlots)
|
||||
{0.0f, 0.0f, 1.0f, 0.0f},
|
||||
};
|
||||
};
|
||||
const auto recordFrame = [&](const aurora::gx::FrameInterpolationDrawIdentity& drawIdentity,
|
||||
float slot0X, float slot1X,
|
||||
std::array<uint8_t, uniformSize>& source) {
|
||||
const auto recordFrame = [&](const aurora::gx::FrameInterpolationDrawIdentity& drawIdentity, float slot0X,
|
||||
float slot1X, std::array<uint8_t, uniformSize>& source) {
|
||||
g_gxState.pnMtx[0].pos = matrixAt(slot0X);
|
||||
g_gxState.pnMtx[1].pos = matrixAt(slot1X);
|
||||
g_gxState.pnMtx[0].nrm = matrixAt(0.0f);
|
||||
g_gxState.pnMtx[1].nrm = matrixAt(0.0f);
|
||||
std::memcpy(source.data() + positionOffset, &g_gxState.pnMtx[0].pos,
|
||||
std::memcpy(source.data() + positionOffset, &g_gxState.pnMtx[0].pos, sizeof(aurora::Mat3x4<float>));
|
||||
std::memcpy(source.data() + positionOffset + sizeof(aurora::Mat3x4<float>), &g_gxState.pnMtx[1].pos,
|
||||
sizeof(aurora::Mat3x4<float>));
|
||||
std::memcpy(source.data() + positionOffset + sizeof(aurora::Mat3x4<float>),
|
||||
&g_gxState.pnMtx[1].pos, sizeof(aurora::Mat3x4<float>));
|
||||
std::memcpy(source.data() + normalOffset, &g_gxState.pnMtx[0].nrm,
|
||||
std::memcpy(source.data() + normalOffset, &g_gxState.pnMtx[0].nrm, sizeof(aurora::Mat3x4<float>));
|
||||
std::memcpy(source.data() + normalOffset + sizeof(aurora::Mat3x4<float>), &g_gxState.pnMtx[1].nrm,
|
||||
sizeof(aurora::Mat3x4<float>));
|
||||
std::memcpy(source.data() + normalOffset + sizeof(aurora::Mat3x4<float>),
|
||||
&g_gxState.pnMtx[1].nrm, sizeof(aurora::Mat3x4<float>));
|
||||
return aurora::gx::record_interpolation_draw(
|
||||
drawIdentity, projection, usedMask,
|
||||
aurora::gx::InterpolatedUniformLayout{
|
||||
.sourceUniformData = source.data(),
|
||||
.uniformSize = source.size(),
|
||||
.projectionOffset = projectionOffset,
|
||||
.positionOffset = positionOffset,
|
||||
.normalOffset = normalOffset,
|
||||
.currentMatrix = 0,
|
||||
.indexedMatrices = true,
|
||||
});
|
||||
return aurora::gx::record_interpolation_draw(drawIdentity, projection, usedMask,
|
||||
aurora::gx::InterpolatedUniformLayout{
|
||||
.sourceUniformData = source.data(),
|
||||
.uniformSize = source.size(),
|
||||
.projectionOffset = projectionOffset,
|
||||
.positionOffset = positionOffset,
|
||||
.normalOffset = normalOffset,
|
||||
.currentMatrix = 0,
|
||||
.indexedMatrices = true,
|
||||
});
|
||||
};
|
||||
|
||||
aurora::gx::set_frame_interpolation_fps(0);
|
||||
@@ -445,10 +458,8 @@ TEST(FrameInterpolationContract, IndexedPaletteHistoryKeepsAbsoluteVertexSlots)
|
||||
ASSERT_EQ(interpolated.size(), uniformSize);
|
||||
aurora::Mat3x4<float> slot0Midpoint{};
|
||||
aurora::Mat3x4<float> slot1Midpoint{};
|
||||
std::memcpy(static_cast<void*>(&slot0Midpoint), interpolated.data() + positionOffset,
|
||||
sizeof(slot0Midpoint));
|
||||
std::memcpy(static_cast<void*>(&slot1Midpoint),
|
||||
interpolated.data() + positionOffset + sizeof(slot0Midpoint),
|
||||
std::memcpy(static_cast<void*>(&slot0Midpoint), interpolated.data() + positionOffset, sizeof(slot0Midpoint));
|
||||
std::memcpy(static_cast<void*>(&slot1Midpoint), interpolated.data() + positionOffset + sizeof(slot0Midpoint),
|
||||
sizeof(slot1Midpoint));
|
||||
EXPECT_FLOAT_EQ(slot0Midpoint.m0.w(), 45.0f);
|
||||
EXPECT_FLOAT_EQ(slot1Midpoint.m0.w(), 55.0f);
|
||||
@@ -584,8 +595,7 @@ TEST(TevTexcoordStateContract, DirectStageFeedsFollowingAddPrevInFixedPoint) {
|
||||
// WGSL must likewise avoid producing a regular-texture UV expression for
|
||||
// stage 1: tex1_size_bias is intentionally absent from this uniform layout.
|
||||
const auto zeroTexgenDependency = aurora::gx::tev_stage_texture_dependency(config, 1);
|
||||
EXPECT_FALSE(aurora::gx::tev_texture_sample_enabled(zeroTexgenDependency,
|
||||
zeroTexgenDependency.combinerUsesTexture));
|
||||
EXPECT_FALSE(aurora::gx::tev_texture_sample_enabled(zeroTexgenDependency, zeroTexgenDependency.combinerUsesTexture));
|
||||
|
||||
// A standalone direct stage still uses the normalized fast path.
|
||||
config.numTexGens = 2;
|
||||
@@ -608,7 +618,6 @@ TEST(TevTexcoordStateContract, DirectStageFeedsFollowingAddPrevInFixedPoint) {
|
||||
config.zTexture = static_cast<u32>(GX_ZT_REPLACE) << 26;
|
||||
EXPECT_TRUE(aurora::gx::tev_z_texture_enabled(config));
|
||||
EXPECT_EQ(aurora::gx::tev_z_texture_stage(config), -1);
|
||||
|
||||
}
|
||||
|
||||
TEST(TevRegisterLivenessContract, RgbWriteDoesNotHideSameStageOldAlphaRead) {
|
||||
@@ -2157,8 +2166,8 @@ TEST_F(GXFifoTest, RawDrawPreservesHorizontalXzQuadVertices) {
|
||||
append_vertex(53.659431f, 0.27f, -53.659435f, 1, 1);
|
||||
const auto expected = vertices;
|
||||
|
||||
ASSERT_TRUE(aurora::gx::fifo::submit_raw_draw(
|
||||
GX_QUADS, GX_VTXFMT0, vertices.data(), 4, static_cast<uint32_t>(vertices.size())));
|
||||
ASSERT_TRUE(aurora::gx::fifo::submit_raw_draw(GX_QUADS, GX_VTXFMT0, vertices.data(), 4,
|
||||
static_cast<uint32_t>(vertices.size())));
|
||||
EXPECT_EQ(aurora::gfx::testing::last_pushed_vertices(), expected);
|
||||
}
|
||||
|
||||
@@ -2183,25 +2192,19 @@ TEST_F(GXFifoTest, DrawTopologyTemplatesPreserveExactGxIndexOrder) {
|
||||
return aurora::gfx::testing::last_pushed_indices();
|
||||
};
|
||||
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_QUADS, 8),
|
||||
(std::vector<u16>{0, 1, 2, 2, 3, 0, 4, 5, 6, 6, 7, 4}));
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_QUADS, 8), (std::vector<u16>{0, 1, 2, 2, 3, 0, 4, 5, 6, 6, 7, 4}));
|
||||
g_gxState.stateDirty = true;
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLES, 6),
|
||||
(std::vector<u16>{0, 1, 2, 3, 4, 5}));
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLES, 6), (std::vector<u16>{0, 1, 2, 3, 4, 5}));
|
||||
g_gxState.stateDirty = true;
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLEFAN, 5),
|
||||
(std::vector<u16>{0, 1, 2, 0, 2, 3, 0, 3, 4}));
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLEFAN, 5), (std::vector<u16>{0, 1, 2, 0, 2, 3, 0, 3, 4}));
|
||||
g_gxState.stateDirty = true;
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLEFAN, 2),
|
||||
(std::vector<u16>{0, 1}));
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLEFAN, 2), (std::vector<u16>{0, 1}));
|
||||
g_gxState.stateDirty = true;
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLESTRIP, 6),
|
||||
(std::vector<u16>{0, 1, 2, 2, 1, 3, 2, 3, 4, 4, 3, 5}));
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_TRIANGLESTRIP, 6), (std::vector<u16>{0, 1, 2, 2, 1, 3, 2, 3, 4, 4, 3, 5}));
|
||||
g_gxState.stateDirty = true;
|
||||
EXPECT_TRUE(decodeAndReadIndices(GX_TRIANGLESTRIP, 0).empty());
|
||||
g_gxState.stateDirty = true;
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_LINES, 2),
|
||||
(std::vector<u16>{0, 1, 3, 3, 2, 0}));
|
||||
EXPECT_EQ(decodeAndReadIndices(GX_LINES, 2), (std::vector<u16>{0, 1, 3, 3, 2, 0}));
|
||||
}
|
||||
|
||||
TEST_F(GXFifoTest, MergedDrawOffsetsCachedTopologyWithoutJoiningPrimitives) {
|
||||
@@ -2216,8 +2219,7 @@ TEST_F(GXFifoTest, MergedDrawOffsetsCachedTopologyWithoutJoiningPrimitives) {
|
||||
decode_fifo(fifo);
|
||||
|
||||
EXPECT_EQ(aurora::gfx::g_mergedDrawCallCount, 1u);
|
||||
EXPECT_EQ(aurora::gfx::testing::last_pushed_indices(),
|
||||
(std::vector<u16>{3, 4, 5}));
|
||||
EXPECT_EQ(aurora::gfx::testing::last_pushed_indices(), (std::vector<u16>{3, 4, 5}));
|
||||
}
|
||||
|
||||
TEST_F(GXFifoTest, TexBufferSize_UsesExactLinearPcFormatSizes) {
|
||||
@@ -2637,9 +2639,18 @@ TEST_F(GXFifoTest, LoadNrmMtxImm_Identity) {
|
||||
|
||||
TEST_F(GXFifoTest, LoadNrmMtxImm_ArbitraryValues) {
|
||||
aurora::Mat3x4<float> mtx{};
|
||||
mtx.m0[0] = 0.5f; mtx.m0[1] = -0.5f; mtx.m0[2] = 0.7f; mtx.m0[3] = 999.0f;
|
||||
mtx.m1[0] = 0.3f; mtx.m1[1] = 0.8f; mtx.m1[2] = -0.1f; mtx.m1[3] = 888.0f;
|
||||
mtx.m2[0] = -0.6f; mtx.m2[1] = 0.2f; mtx.m2[2] = 0.9f; mtx.m2[3] = 777.0f;
|
||||
mtx.m0[0] = 0.5f;
|
||||
mtx.m0[1] = -0.5f;
|
||||
mtx.m0[2] = 0.7f;
|
||||
mtx.m0[3] = 999.0f;
|
||||
mtx.m1[0] = 0.3f;
|
||||
mtx.m1[1] = 0.8f;
|
||||
mtx.m1[2] = -0.1f;
|
||||
mtx.m1[3] = 888.0f;
|
||||
mtx.m2[0] = -0.6f;
|
||||
mtx.m2[1] = 0.2f;
|
||||
mtx.m2[2] = 0.9f;
|
||||
mtx.m2[3] = 777.0f;
|
||||
|
||||
GXLoadNrmMtxImm(&mtx, GX_PNMTX0);
|
||||
auto bytes = capture_fifo();
|
||||
@@ -2770,9 +2781,18 @@ TEST_F(GXFifoTest, LoadTexMtx3x4_Identity) {
|
||||
|
||||
TEST_F(GXFifoTest, LoadTexMtx3x4_ArbitraryValues) {
|
||||
aurora::Mat3x4<float> mtx{};
|
||||
mtx.m0[0] = 2.0f; mtx.m0[1] = 0.5f; mtx.m0[2] = 0.0f; mtx.m0[3] = 10.0f;
|
||||
mtx.m1[0] = -0.5f; mtx.m1[1] = 3.0f; mtx.m1[2] = 0.0f; mtx.m1[3] = 20.0f;
|
||||
mtx.m2[0] = 0.0f; mtx.m2[1] = 0.0f; mtx.m2[2] = 1.5f; mtx.m2[3] = -5.0f;
|
||||
mtx.m0[0] = 2.0f;
|
||||
mtx.m0[1] = 0.5f;
|
||||
mtx.m0[2] = 0.0f;
|
||||
mtx.m0[3] = 10.0f;
|
||||
mtx.m1[0] = -0.5f;
|
||||
mtx.m1[1] = 3.0f;
|
||||
mtx.m1[2] = 0.0f;
|
||||
mtx.m1[3] = 20.0f;
|
||||
mtx.m2[0] = 0.0f;
|
||||
mtx.m2[1] = 0.0f;
|
||||
mtx.m2[2] = 1.5f;
|
||||
mtx.m2[3] = -5.0f;
|
||||
|
||||
GXLoadTexMtxImm(&mtx, GX_TEXMTX0, GX_MTX3x4);
|
||||
auto bytes = capture_fifo();
|
||||
@@ -2888,8 +2908,14 @@ TEST_F(GXFifoTest, LoadTexMtx2x4_Identity) {
|
||||
|
||||
TEST_F(GXFifoTest, LoadTexMtx2x4_ArbitraryValues) {
|
||||
aurora::Mat3x4<float> mtx{};
|
||||
mtx.m0[0] = 0.5f; mtx.m0[1] = -1.0f; mtx.m0[2] = 0.25f; mtx.m0[3] = 100.0f;
|
||||
mtx.m1[0] = 3.0f; mtx.m1[1] = 0.0f; mtx.m1[2] = -2.5f; mtx.m1[3] = -50.0f;
|
||||
mtx.m0[0] = 0.5f;
|
||||
mtx.m0[1] = -1.0f;
|
||||
mtx.m0[2] = 0.25f;
|
||||
mtx.m0[3] = 100.0f;
|
||||
mtx.m1[0] = 3.0f;
|
||||
mtx.m1[1] = 0.0f;
|
||||
mtx.m1[2] = -2.5f;
|
||||
mtx.m1[3] = -50.0f;
|
||||
// Row 2 values should be ignored by the encoder
|
||||
mtx.m2[0] = 999.0f;
|
||||
|
||||
@@ -3497,8 +3523,8 @@ TEST_F(GXFifoTest, TexCoordGen_Identity) {
|
||||
}
|
||||
|
||||
TEST_F(GXFifoTest, MatrixIndexA_DecodesTexMatricesFromCpPacket) {
|
||||
const u32 value = (GX_PNMTX3 << 0) | (GX_TEXMTX3 << 6) | (GX_TEXMTX4 << 12) | (GX_IDENTITY << 18) |
|
||||
(GX_TEXMTX7 << 24);
|
||||
const u32 value =
|
||||
(GX_PNMTX3 << 0) | (GX_TEXMTX3 << 6) | (GX_TEXMTX4 << 12) | (GX_IDENTITY << 18) | (GX_TEXMTX7 << 24);
|
||||
auto bytes = cp_cmd(0x30, value);
|
||||
|
||||
reset_gx_state();
|
||||
@@ -4179,11 +4205,44 @@ TEST_F(GXFifoTest, CopyTexColorFormatMarksResolvePersistent) {
|
||||
|
||||
GXSetTexCopySrc(336, 300, 152, 114);
|
||||
GXSetTexCopyDst(152, 114, GX_TF_RGBA8, GX_FALSE);
|
||||
aurora::gfx::testing::set_current_frame(42);
|
||||
GXCopyTex(image.data(), GX_FALSE);
|
||||
|
||||
const auto& records = aurora::gfx::testing::resolve_pass_records();
|
||||
ASSERT_EQ(records.size(), 1u);
|
||||
EXPECT_TRUE(records.front().persistentCopy);
|
||||
ASSERT_TRUE(records.front().texture);
|
||||
EXPECT_TRUE(records.front().texture->isEfbCopy);
|
||||
EXPECT_EQ(records.front().texture->lastEfbCopyFrame, 42u);
|
||||
EXPECT_TRUE(records.front().texture->is_recent_efb_copy(42));
|
||||
EXPECT_TRUE(records.front().texture->is_recent_efb_copy(43));
|
||||
EXPECT_FALSE(records.front().texture->is_recent_efb_copy(44));
|
||||
|
||||
GXTexObj_ texObj{};
|
||||
texObj.mWidth = 152;
|
||||
texObj.mHeight = 114;
|
||||
texObj.mFormat = GX_TF_RGBA8;
|
||||
gxState().textures[GX_TEXMAP0] = aurora::gfx::TextureBind{texObj, records.front().texture};
|
||||
aurora::gx::ShaderConfig shader{};
|
||||
shader.numTexGens = 1;
|
||||
shader.tevStageCount = 1;
|
||||
shader.tevStages[0].texCoordId = GX_TEXCOORD0;
|
||||
shader.tevStages[0].texMapId = GX_TEXMAP0;
|
||||
shader.tevStages[0].colorPass.d = GX_CC_TEXC;
|
||||
shader.tevStages[0].alphaPass.d = GX_CA_TEXA;
|
||||
const auto info = aurora::gx::build_shader_info(shader);
|
||||
|
||||
aurora::gfx::testing::reset_uniform_allocations();
|
||||
const auto freshLayout = aurora::gx::build_uniform(info, 0, aurora::gx::BindGroupRanges{},
|
||||
aurora::gx::FrameInterpolationDrawIdentity{}, false);
|
||||
EXPECT_TRUE(freshLayout.replayLayout.nativeEfbEffect);
|
||||
|
||||
// Once retained instead of regenerated, the same reduced alpha texture is a
|
||||
// one-shot 2D bake (the path used by MKW's minimap), not live post-processing.
|
||||
aurora::gfx::testing::set_current_frame(44);
|
||||
const auto retainedLayout = aurora::gx::build_uniform(info, 0, aurora::gx::BindGroupRanges{},
|
||||
aurora::gx::FrameInterpolationDrawIdentity{}, false);
|
||||
EXPECT_FALSE(retainedLayout.replayLayout.nativeEfbEffect);
|
||||
}
|
||||
|
||||
TEST_F(GXFifoTest, RecurringColorCopyKeepsLaterResolveSkippable) {
|
||||
@@ -4826,9 +4885,18 @@ TEST_F(GXFifoTest, LoadPTTexMtx_Identity) {
|
||||
|
||||
TEST_F(GXFifoTest, LoadPTTexMtx_ArbitraryValues) {
|
||||
aurora::Mat3x4<float> mtx{};
|
||||
mtx.m0[0] = 2.0f; mtx.m0[1] = 0.5f; mtx.m0[2] = 0.0f; mtx.m0[3] = 10.0f;
|
||||
mtx.m1[0] = -0.5f; mtx.m1[1] = 3.0f; mtx.m1[2] = 0.0f; mtx.m1[3] = 20.0f;
|
||||
mtx.m2[0] = 0.0f; mtx.m2[1] = 0.0f; mtx.m2[2] = 1.5f; mtx.m2[3] = -5.0f;
|
||||
mtx.m0[0] = 2.0f;
|
||||
mtx.m0[1] = 0.5f;
|
||||
mtx.m0[2] = 0.0f;
|
||||
mtx.m0[3] = 10.0f;
|
||||
mtx.m1[0] = -0.5f;
|
||||
mtx.m1[1] = 3.0f;
|
||||
mtx.m1[2] = 0.0f;
|
||||
mtx.m1[3] = 20.0f;
|
||||
mtx.m2[0] = 0.0f;
|
||||
mtx.m2[1] = 0.0f;
|
||||
mtx.m2[2] = 1.5f;
|
||||
mtx.m2[3] = -5.0f;
|
||||
|
||||
GXLoadTexMtxImm(&mtx, GX_PTTEXMTX0, GX_MTX3x4);
|
||||
auto bytes = capture_fifo();
|
||||
|
||||
@@ -46,6 +46,7 @@ std::recursive_mutex& renderer_gpu_mutex() noexcept {
|
||||
static std::recursive_mutex mutex;
|
||||
return mutex;
|
||||
}
|
||||
bool stereo_frame_provider_active() noexcept { return true; }
|
||||
} // namespace aurora
|
||||
|
||||
extern "C" bool aurora_wait_for_frame_worker_for(uint32_t) { return true; }
|
||||
@@ -293,12 +294,9 @@ std::pair<ByteBuffer, Range> map_uniform(size_t length) {
|
||||
s_uniformAllocations.emplace_back(length, 0);
|
||||
auto& uniformBytes = s_uniformAllocations.back();
|
||||
return {ByteBuffer{uniformBytes.data(), uniformBytes.size()},
|
||||
Range{static_cast<uint32_t>(s_uniformAllocations.size() - 1),
|
||||
static_cast<uint32_t>(length)}};
|
||||
}
|
||||
std::pair<ByteBuffer, Range> copy_uniform(Range source) {
|
||||
return map_uniform(source.size);
|
||||
Range{static_cast<uint32_t>(s_uniformAllocations.size() - 1), static_cast<uint32_t>(length)}};
|
||||
}
|
||||
std::pair<ByteBuffer, Range> copy_uniform(Range source) { return map_uniform(source.size); }
|
||||
uint32_t align_uniform(uint32_t value) { return (value + 255u) & ~255u; }
|
||||
|
||||
Vec2<uint32_t> get_render_target_size() noexcept { return s_renderTargetSize; }
|
||||
@@ -321,15 +319,9 @@ void reset_vertex_push_record() noexcept {
|
||||
s_trackDrawCommands = false;
|
||||
g_mergedDrawCallCount = 0;
|
||||
}
|
||||
const std::vector<uint8_t>& last_pushed_vertices() noexcept {
|
||||
return s_lastPushedVertices;
|
||||
}
|
||||
const std::vector<uint16_t>& last_pushed_indices() noexcept {
|
||||
return s_lastPushedIndices;
|
||||
}
|
||||
void reset_uniform_allocations() noexcept {
|
||||
s_uniformAllocations.clear();
|
||||
}
|
||||
const std::vector<uint8_t>& last_pushed_vertices() noexcept { return s_lastPushedVertices; }
|
||||
const std::vector<uint16_t>& last_pushed_indices() noexcept { return s_lastPushedIndices; }
|
||||
void reset_uniform_allocations() noexcept { s_uniformAllocations.clear(); }
|
||||
const std::vector<uint8_t>& uniform_allocation(size_t index) noexcept {
|
||||
CHECK(index < s_uniformAllocations.size(), "uniform test allocation {} out of range {}", index,
|
||||
s_uniformAllocations.size());
|
||||
@@ -339,9 +331,7 @@ void use_draw_command_tracking(bool enabled) noexcept {
|
||||
s_trackDrawCommands = enabled;
|
||||
s_lastGxDraw.reset();
|
||||
}
|
||||
void use_real_vertex_format_helpers(bool enabled) noexcept {
|
||||
s_useRealVertexFormatHelpers = enabled;
|
||||
}
|
||||
void use_real_vertex_format_helpers(bool enabled) noexcept { s_useRealVertexFormatHelpers = enabled; }
|
||||
} // namespace aurora::gfx::testing
|
||||
|
||||
// --- Pipeline/draw command stubs ---
|
||||
@@ -405,8 +395,8 @@ void reset_resolve_pass_records() noexcept { s_resolvePassRecords.clear(); }
|
||||
|
||||
const std::vector<ResolvePassRecord>& resolve_pass_records() noexcept { return s_resolvePassRecords; }
|
||||
|
||||
void set_framebuffer_sizes(uint32_t logicalWidth, uint32_t logicalHeight,
|
||||
uint32_t targetWidth, uint32_t targetHeight) noexcept {
|
||||
void set_framebuffer_sizes(uint32_t logicalWidth, uint32_t logicalHeight, uint32_t targetWidth,
|
||||
uint32_t targetHeight) noexcept {
|
||||
s_logicalFbSize = {logicalWidth, logicalHeight};
|
||||
s_renderTargetSize = {targetWidth, targetHeight};
|
||||
}
|
||||
@@ -431,9 +421,9 @@ TextureHandle new_conv_texture(uint32_t width, uint32_t height, u32 gxFormat, co
|
||||
void write_texture(const TextureRef& ref, ArrayRef<uint8_t> data) noexcept {}
|
||||
void resolve_pass(TextureHandle texture, ClipRect rect, bool clearColor, bool clearAlpha, bool clearDepth,
|
||||
Vec4<float> clearColorValue, float clearDepthValue, GXTexFmt resolveFormat,
|
||||
const Vec4<float>* sourceRectPixels, bool halfScale,
|
||||
const std::array<u32, 3>* copyFilterCoefficients, bool forceOpaqueAlpha,
|
||||
float copyFilterRowStride, bool clampTop, bool clampBottom, bool persistentCopy) {
|
||||
const Vec4<float>* sourceRectPixels, bool halfScale, const std::array<u32, 3>* copyFilterCoefficients,
|
||||
bool forceOpaqueAlpha, float copyFilterRowStride, bool clampTop, bool clampBottom,
|
||||
bool persistentCopy) {
|
||||
testing::ResolvePassRecord record;
|
||||
record.texture = std::move(texture);
|
||||
record.rect = rect;
|
||||
|
||||
@@ -2,6 +2,9 @@
|
||||
|
||||
#include <gtest/gtest.h>
|
||||
|
||||
#include <array>
|
||||
#include <cmath>
|
||||
|
||||
namespace aurora::gfx::stereo_replay {
|
||||
namespace {
|
||||
|
||||
@@ -27,9 +30,7 @@ TEST(StereoReplayTest, EyeFrustumPreservesGameDepthMapping) {
|
||||
EXPECT_FLOAT_EQ(result.m1[2], eye.m1[2]);
|
||||
for (size_t row = 0; row < 4; ++row) {
|
||||
for (size_t column = 0; column < 4; ++column) {
|
||||
const bool frustumTerm =
|
||||
(row == 0 && (column == 0 || column == 2)) ||
|
||||
(row == 1 && (column == 1 || column == 2));
|
||||
const bool frustumTerm = (row == 0 && (column == 0 || column == 2)) || (row == 1 && (column == 1 || column == 2));
|
||||
if (!frustumTerm) {
|
||||
EXPECT_FLOAT_EQ(result[row][column], game[row][column]);
|
||||
}
|
||||
@@ -37,5 +38,122 @@ TEST(StereoReplayTest, EyeFrustumPreservesGameDepthMapping) {
|
||||
}
|
||||
}
|
||||
|
||||
Mat4x4<float> game_orthographic_projection() {
|
||||
// x over [0, 640) and y over [0, 456) mapped to NDC, with a shallow depth
|
||||
// window, as GX builds an orthographic projection for a 2D layer.
|
||||
Mat4x4<float> game{};
|
||||
game.m0 = {2.0f / 640.0f, 0.0f, 0.0f, -1.0f};
|
||||
game.m1 = {0.0f, -2.0f / 456.0f, 0.0f, 1.0f};
|
||||
game.m2 = {0.0f, 0.0f, -1.0f / 1000.0f, -0.5f};
|
||||
game.m3 = {0.0f, 0.0f, 0.0f, 1.0f};
|
||||
return game;
|
||||
}
|
||||
|
||||
float dot4(const Vec4<float>& row, const Vec4<float>& v) {
|
||||
return row[0] * v[0] + row[1] * v[1] + row[2] * v[2] + row[3] * v[3];
|
||||
}
|
||||
|
||||
const std::array<Vec4<float>, 5> kVertices{{
|
||||
{0.0f, 0.0f, 0.0f, 1.0f},
|
||||
{640.0f, 456.0f, 0.0f, 1.0f},
|
||||
{320.0f, 228.0f, -250.0f, 1.0f},
|
||||
{97.0f, 401.0f, 640.0f, 1.0f},
|
||||
{-30.0f, 12.5f, 33.0f, 1.0f},
|
||||
}};
|
||||
|
||||
TEST(StereoReplayTest, OrthographicProjectionIsRecognizedByItsWRow) {
|
||||
const auto game = game_orthographic_projection();
|
||||
EXPECT_TRUE(is_orthographic_projection(game));
|
||||
|
||||
Mat4x4<float> perspective = game;
|
||||
perspective.m3 = {0.0f, 0.0f, -1.0f, 0.0f};
|
||||
EXPECT_FALSE(is_orthographic_projection(perspective));
|
||||
}
|
||||
|
||||
TEST(StereoReplayTest, HudViewportNdcIsLiftedIntoTheDisplayedFrame) {
|
||||
// Bottom-right quarter of a 608x456 displayed frame.
|
||||
const auto remap = make_hud_ndc_remap(304.0f, 228.0f, 304.0f, 228.0f, 0.0f, 0.0f, 608.0f, 456.0f);
|
||||
EXPECT_FLOAT_EQ(remap.scaleX, 0.5f);
|
||||
EXPECT_FLOAT_EQ(remap.scaleY, 0.5f);
|
||||
EXPECT_FLOAT_EQ(remap.offsetX, 0.5f);
|
||||
EXPECT_FLOAT_EQ(remap.offsetY, -0.5f);
|
||||
|
||||
Mat4x4<float> local{};
|
||||
local.m0 = {1.0f, 0.0f, 0.0f, 0.0f};
|
||||
local.m1 = {0.0f, 1.0f, 0.0f, 0.0f};
|
||||
local.m3 = {0.0f, 0.0f, 0.0f, 1.0f};
|
||||
const auto frame = remap_hud_ndc(local, remap);
|
||||
const Vec4<float> topLeft{-1.0f, 1.0f, 0.0f, 1.0f};
|
||||
const Vec4<float> bottomRight{1.0f, -1.0f, 0.0f, 1.0f};
|
||||
EXPECT_FLOAT_EQ(dot4(frame.m0, topLeft), 0.0f);
|
||||
EXPECT_FLOAT_EQ(dot4(frame.m1, topLeft), 0.0f);
|
||||
EXPECT_FLOAT_EQ(dot4(frame.m0, bottomRight), 1.0f);
|
||||
EXPECT_FLOAT_EQ(dot4(frame.m1, bottomRight), -1.0f);
|
||||
}
|
||||
|
||||
TEST(StereoReplayTest, HudScreenProjectionMatchesTheChainItComposes) {
|
||||
const auto game = game_orthographic_projection();
|
||||
Mat4x4<float> eyeFrustum{};
|
||||
eyeFrustum.m0 = {1.15f, 0.0f, 0.08f, 0.0f};
|
||||
eyeFrustum.m1 = {0.0f, 1.02f, -0.03f, 0.0f};
|
||||
|
||||
// A head turned a little and offset from the recorded center eye.
|
||||
const float angle = 0.3f;
|
||||
const float c = std::cos(angle);
|
||||
const float s = std::sin(angle);
|
||||
Mat3x4<float> viewFromCenter{};
|
||||
viewFromCenter.m0 = {c, 0.0f, s, 15.0f};
|
||||
viewFromCenter.m1 = {0.0f, 1.0f, 0.0f, -4.0f};
|
||||
viewFromCenter.m2 = {-s, 0.0f, c, 7.0f};
|
||||
|
||||
const HudScreen screen{.halfWidth = 600.0f, .halfHeight = 337.5f, .distance = 1000.0f};
|
||||
const auto composed = compose_hud_screen_projection(eyeFrustum, viewFromCenter, screen, game, true);
|
||||
const auto exactDepth = backend_ndc_depth_row(game, true);
|
||||
|
||||
for (const auto& v : kVertices) {
|
||||
// The same chain, one step at a time: game NDC, a point on the screen
|
||||
// rectangle, that point in eye view space, then the eye's clip space.
|
||||
const float ndcX = dot4(game.m0, v);
|
||||
const float ndcY = dot4(game.m1, v);
|
||||
const Vec4<float> screenPoint{ndcX * screen.halfWidth, ndcY * screen.halfHeight, -screen.distance, 1.0f};
|
||||
const float eyeX = dot4(viewFromCenter.m0, screenPoint);
|
||||
const float eyeY = dot4(viewFromCenter.m1, screenPoint);
|
||||
const float eyeZ = dot4(viewFromCenter.m2, screenPoint);
|
||||
|
||||
EXPECT_NEAR(dot4(composed.m0, v), eyeFrustum.m0[0] * eyeX + eyeFrustum.m0[2] * eyeZ, 1e-2f);
|
||||
EXPECT_NEAR(dot4(composed.m1, v), eyeFrustum.m1[1] * eyeY + eyeFrustum.m1[2] * eyeZ, 1e-2f);
|
||||
const float clipW = -eyeZ;
|
||||
EXPECT_NEAR(dot4(composed.m3, v), clipW, 1e-2f);
|
||||
EXPECT_NEAR(dot4(composed.m2, v), dot4(exactDepth, v), 1e-6f);
|
||||
}
|
||||
}
|
||||
|
||||
TEST(StereoReplayTest, HudScreenParksRasterDepthAtMidrangeUnderHeadMotion) {
|
||||
const auto game = game_orthographic_projection();
|
||||
Mat4x4<float> eyeFrustum{};
|
||||
eyeFrustum.m0 = {1.15f, 0.0f, 0.08f, 0.0f};
|
||||
eyeFrustum.m1 = {0.0f, 1.02f, -0.03f, 0.0f};
|
||||
const float angle = 0.35f;
|
||||
const float c = std::cos(angle);
|
||||
const float s = std::sin(angle);
|
||||
Mat3x4<float> moved{};
|
||||
moved.m0 = {c, 0.0f, s, 21.0f};
|
||||
moved.m1 = {0.0f, 1.0f, 0.0f, -9.0f};
|
||||
moved.m2 = {-s, 0.0f, c, 13.0f};
|
||||
|
||||
const HudScreen screen{.halfWidth = 600.0f, .halfHeight = 337.5f, .distance = 1000.0f};
|
||||
const auto composed = compose_hud_screen_projection(eyeFrustum, moved, screen, game, true);
|
||||
|
||||
// The exact-depth shader captures composed Z, then parks clip Z at -0.5W.
|
||||
// Aurora's following reversed-depth conversion negates that to +0.5W, so
|
||||
// rasterization stays stable even though W varies across the rotated screen.
|
||||
for (const auto& v : kVertices) {
|
||||
const float w = dot4(composed.m3, v);
|
||||
ASSERT_GT(w, 0.0f);
|
||||
const float parkedClipZ = -0.5f * w;
|
||||
EXPECT_NEAR(-parkedClipZ / w, 0.5f, 1e-5f);
|
||||
}
|
||||
}
|
||||
|
||||
} // namespace
|
||||
} // namespace aurora::gfx::stereo_replay
|
||||
@@ -51,6 +51,8 @@ struct RuntimeUserConfig {
|
||||
std::optional<float> vrRenderScale;
|
||||
std::optional<float> vrWorldUnitsPerMeter;
|
||||
std::optional<float> vrHudDistanceMeters;
|
||||
std::optional<float> vrHudWidthMeters;
|
||||
std::optional<bool> vrHudVirtualScreen;
|
||||
std::optional<bool> vrStopAtDisplayCopy;
|
||||
std::optional<bool> vrSkipCopyClears;
|
||||
std::optional<float> audioVolume;
|
||||
@@ -324,6 +326,13 @@ inline void EnsureConfigFile() {
|
||||
"render_scale = 1.0\n"
|
||||
"world_units_per_meter = 500.0\n"
|
||||
"hud_distance_meters = 2.0\n"
|
||||
"hud_width_meters = 2.4\n"
|
||||
"# During a race, put the game's 2D layer (minimap, position,\n"
|
||||
"# item roulette, lap times) on a virtual screen fixed in front of\n"
|
||||
"# the kart camera instead of stretching it across the whole view.\n"
|
||||
"# Changeable live from the F10 menu; the two sizes above place\n"
|
||||
"# that screen and the menu screen alike and are read at launch.\n"
|
||||
"hud_virtual_screen = true\n"
|
||||
"# EFB replay controls for the per-eye views, changeable live\n"
|
||||
"# from the F10 menu. stop_at_display_copy ends each eye at the\n"
|
||||
"# frame's final GXCopyDisp; skip_copy_clears drops the EFB\n"
|
||||
@@ -484,6 +493,11 @@ inline RuntimeUserConfig ParseConfigDocument(const toml::value& document) {
|
||||
value && *value >= 0.25f && *value <= 10.0f) {
|
||||
config.vrHudDistanceMeters = *value;
|
||||
}
|
||||
if (auto value = FindConfigFloat(document, "vr", "hud_width_meters");
|
||||
value && *value >= 0.25f && *value <= 20.0f) {
|
||||
config.vrHudWidthMeters = *value;
|
||||
}
|
||||
config.vrHudVirtualScreen = FindConfigValue<bool>(document, "vr", "hud_virtual_screen");
|
||||
config.vrStopAtDisplayCopy = FindConfigValue<bool>(document, "vr", "stop_at_display_copy");
|
||||
config.vrSkipCopyClears = FindConfigValue<bool>(document, "vr", "skip_copy_clears");
|
||||
|
||||
@@ -713,6 +727,11 @@ inline bool SetVrEnabled(bool value) {
|
||||
return WriteSetting("vr", "enabled", value ? "true" : "false");
|
||||
}
|
||||
|
||||
inline bool SetVrHudVirtualScreen(bool value) {
|
||||
Mutable().vrHudVirtualScreen = value;
|
||||
return WriteSetting("vr", "hud_virtual_screen", value ? "true" : "false");
|
||||
}
|
||||
|
||||
inline bool SetVrStopAtDisplayCopy(bool value) {
|
||||
Mutable().vrStopAtDisplayCopy = value;
|
||||
return WriteSetting("vr", "stop_at_display_copy", value ? "true" : "false");
|
||||
@@ -955,6 +974,14 @@ inline float VrHudDistanceMeters(float fallback = 2.0f) {
|
||||
return std::clamp(Get().vrHudDistanceMeters.value_or(fallback), 0.25f, 10.0f);
|
||||
}
|
||||
|
||||
inline float VrHudWidthMeters(float fallback = 2.4f) {
|
||||
return std::clamp(Get().vrHudWidthMeters.value_or(fallback), 0.25f, 20.0f);
|
||||
}
|
||||
|
||||
inline bool VrHudVirtualScreen(bool fallback = true) {
|
||||
return Get().vrHudVirtualScreen.value_or(fallback);
|
||||
}
|
||||
|
||||
inline bool VrStopAtDisplayCopy(bool fallback = true) {
|
||||
return Get().vrStopAtDisplayCopy.value_or(fallback);
|
||||
}
|
||||
|
||||
@@ -74,6 +74,9 @@ struct MkwVRPolicyConfig {
|
||||
bool immersive_races = true;
|
||||
float world_units_per_meter = 500.0f;
|
||||
float hud_distance_meters = 2.0f;
|
||||
// Width of the virtual screen, shared by the menu quad layer and the
|
||||
// immersive race HUD so 2D content keeps its place across the transition.
|
||||
float hud_width_meters = 2.4f;
|
||||
float hud_scale = 1.0f;
|
||||
};
|
||||
|
||||
|
||||
@@ -102,6 +102,7 @@ bool g_showFps = RuntimeConfigFile::ShowFps(true);
|
||||
bool g_vrEnabled = RuntimeConfigFile::VrEnabled(false);
|
||||
bool g_vrStopAtDisplayCopy = RuntimeConfigFile::VrStopAtDisplayCopy(true);
|
||||
bool g_vrSkipCopyClears = RuntimeConfigFile::VrSkipCopyClears(true);
|
||||
bool g_vrHudVirtualScreen = RuntimeConfigFile::VrHudVirtualScreen(true);
|
||||
uint32_t g_disabledPostProcessingPaths = RuntimeConfigFile::DisabledPostProcessingPaths(0);
|
||||
std::array<int32_t, PAD_MAX_CONTROLLERS> g_configuredControllerIndices = [] {
|
||||
std::array<int32_t, PAD_MAX_CONTROLLERS> indices{};
|
||||
@@ -705,6 +706,16 @@ void DrawAudioSettings() {
|
||||
}
|
||||
}
|
||||
|
||||
// The virtual screen's placement comes from the launch-time [vr] geometry, the
|
||||
// same metres the menu quad is built from, converted into the world units the
|
||||
// eye replay works in.
|
||||
void ApplyVrHudVirtualScreen() {
|
||||
const float unitsPerMeter = RuntimeConfigFile::VrWorldUnitsPerMeter(500.0f);
|
||||
aurora_set_stereo_hud_screen(g_vrHudVirtualScreen,
|
||||
RuntimeConfigFile::VrHudWidthMeters(2.4f) * unitsPerMeter,
|
||||
RuntimeConfigFile::VrHudDistanceMeters(2.0f) * unitsPerMeter);
|
||||
}
|
||||
|
||||
void DrawGraphicsSettings() {
|
||||
g_displayMode = static_cast<int>(aurora_get_display_mode());
|
||||
struct EffectFlag {
|
||||
@@ -827,6 +838,19 @@ void DrawGraphicsSettings() {
|
||||
"Both apply on the next frame. Turning either off restores the raw replay and is "
|
||||
"expected to black out the eyes.");
|
||||
ImGui::PopTextWrapPos();
|
||||
ImGui::Separator();
|
||||
ImGui::Text("VR 2D layer");
|
||||
if (ImGui::Checkbox("2D layer on a virtual screen", &g_vrHudVirtualScreen)) {
|
||||
ApplyVrHudVirtualScreen();
|
||||
RuntimeConfigFile::SetVrHudVirtualScreen(g_vrHudVirtualScreen);
|
||||
}
|
||||
if (ImGui::IsItemHovered()) {
|
||||
ImGui::SetTooltip(
|
||||
"Puts the minimap, race position, item roulette and the rest of the race HUD on a "
|
||||
"screen fixed ahead of the kart camera. Turn off to leave them stretched across "
|
||||
"the whole view. Its size and distance are the [vr] hud_width_meters and "
|
||||
"hud_distance_meters read at launch.");
|
||||
}
|
||||
}
|
||||
|
||||
void DrawFpsOverlay() {
|
||||
@@ -1046,6 +1070,7 @@ void InitializeRuntimeSettings() noexcept {
|
||||
aurora_set_disable_copy_filter(g_disableCopyFilter);
|
||||
aurora_set_stereo_stop_at_display_copy(g_vrStopAtDisplayCopy);
|
||||
aurora_set_stereo_skip_copy_clears(g_vrSkipCopyClears);
|
||||
ApplyVrHudVirtualScreen();
|
||||
aurora_set_skip_unready_pipelines(g_skipUnreadyPipelines);
|
||||
g_strapInputAccepted.store(false, std::memory_order_relaxed);
|
||||
g_startupDismissFrame.store(UINT64_MAX, std::memory_order_relaxed);
|
||||
|
||||
@@ -50,6 +50,9 @@ MkwVRPolicyConfig SanitizeConfig(const MkwVRPolicyConfig& config) noexcept {
|
||||
if (!IsFinitePositive(&sanitized.hud_distance_meters)) {
|
||||
sanitized.hud_distance_meters = kDefaultConfig.hud_distance_meters;
|
||||
}
|
||||
if (!IsFinitePositive(&sanitized.hud_width_meters)) {
|
||||
sanitized.hud_width_meters = kDefaultConfig.hud_width_meters;
|
||||
}
|
||||
if (!IsFinitePositive(&sanitized.hud_scale)) {
|
||||
sanitized.hud_scale = kDefaultConfig.hud_scale;
|
||||
}
|
||||
|
||||
@@ -43,6 +43,7 @@ void ConfigurePolicy(bool enabled) noexcept {
|
||||
config.immersive_races = true;
|
||||
config.world_units_per_meter = RuntimeConfigFile::VrWorldUnitsPerMeter(500.0f);
|
||||
config.hud_distance_meters = RuntimeConfigFile::VrHudDistanceMeters(2.0f);
|
||||
config.hud_width_meters = RuntimeConfigFile::VrHudWidthMeters(2.4f);
|
||||
MkwVRPolicyConfigure(config);
|
||||
MkwVRInstrumentationInitialize();
|
||||
}
|
||||
@@ -458,6 +459,7 @@ private:
|
||||
presentation.mode = immersive ? OpenXRD3D12FrameMode::ImmersiveProjection
|
||||
: OpenXRD3D12FrameMode::VirtualScreen;
|
||||
presentation.quad_distance_meters = policy.config.hud_distance_meters;
|
||||
presentation.quad_width_meters = policy.config.hud_width_meters;
|
||||
|
||||
OpenXRD3D12Frame frame{};
|
||||
const OpenXRD3D12BeginStatus begin = backend_->BeginFrame(presentation, frame);
|
||||
|
||||
Reference in new issue
Block a user