Commit Graph
12279 Commits
Author SHA1 Message Date
Ryan Houdek ea20429351 Docs: Update for release FEX-2507.1 FEX-2507.1 2025-07-11 11:37:44 -07:00
Alyssa Rosenzweig 91828efa7a JIT: fix divisor masking
oversight. should fix Steam.

Fixes: de4becc26 ("OpcodeDispatcher: mask certain divisors")
Closes: #4652
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-11 11:34:45 -07:00
Billy Laws cce605d5e0 PoolBufferWithTimedRetirement: Unclaim in dtor
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.

Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
2025-07-11 11:34:07 -07:00
Ryan Houdek 3ba84ad06a Docs: Update for release FEX-2507 FEX-2507 2025-07-07 23:49:56 -07:00
Ryan Houdek c6aae9e05a Merge pull request #4634 from Sonicadvance1/fix_horizon
EmulatedFiles: Emulate `current_clocksource`
2025-07-07 21:16:54 -07:00
Ryan Houdek 95b4618833 Merge pull request #4644 from ChanthMiao/fix/sigframe_mistake
Fix: wrong magic value in fpstate.
2025-07-06 17:07:31 -07:00
Changwei Miao 1fe17d55d9 Fix: wrong magic value in fpstate.
FEX should only set fpx_sw_bytes.magic1 with FP_XSTATE_MAGIC when
avx is enabled. Otherwise it may cause segfault in ntdll::save_context,
which requires access to extended xstate info if magic1 equals FP_XSTATE_MAGIC.

Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2025-07-06 18:44:37 +08:00
Ryan Houdek ead73371d9 EmulatedFiles: Emulate current_clocksource
WINE uses this to determine TSC frequency and because it doesn't say
`tsc` on ARM devices, it was ignoring TSC and instead using CPU maximum
frequency.

This was causing Horizon to think the TSC ran at whatever the max
frequency of a core was  (1.8Ghz to 2.6Ghz depending?) This was causing
all of Horizon Zero Dawn's physics to run at slower than real time
speeds because our 1Ghz (on Orion) TSC is significantly lower than the
max clock speeds of the cores.

This is still a bug in Wine that it is using the maximum CPU clock speed
in the case of current_clocksource not being TSC, but that's a battle
for a different time.
2025-07-05 21:10:53 -07:00
Ryan Houdek 6f089a4323 Merge pull request #4643 from tstellar/llvm-21
Fix build with LLVM >= 21
2025-07-05 16:51:38 -07:00
Ryan Houdek 640f024551 Merge pull request #4641 from Sonicadvance1/optimize_sincos
JIT: Optimize x87 FSINCOS
2025-07-05 16:26:37 -07:00
Tom Stellard 99920f89dd Fix build with LLVM >= 21 2025-07-05 16:50:34 +00:00
Ryan Houdek 8ea276267f InstcountCI: Update 2025-07-03 18:05:52 -07:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Ryan Houdek fa0a54deb9 Merge pull request #4640 from alyssarosenzweig/bug/fix-hades
Fix Hades
2025-07-03 14:53:01 -07:00
Alyssa Rosenzweig 046043090f unittests: add move merging test
this hits a nasty case with post-RA merging. fails on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:56 -04:00
Alyssa Rosenzweig 61150a18cc RegisterAllocationPass: fix bookkeeping with merging
this fixes Hades.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:53 -04:00
Ryan Houdek 95ca20cfee Merge pull request #4636 from alyssarosenzweig/bug/ra-invariant
RegisterAllocationPass: assert an invariant in post-RA prop
2025-07-02 18:17:08 -07:00
Alyssa Rosenzweig 5a536d47fd RegisterAllocationPass: assert an invariant in post-RA prop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:53:47 -04:00
Ryan Houdek afbc7da027 Merge pull request #4629 from alyssarosenzweig/opt/cpuid-basic
Optimize some constant cpuid/xgetbv cases
2025-07-02 10:25:01 -07:00
Alyssa Rosenzweig c093c08c40 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 360d8c629e RegisterAllocationPass: optimize cpuid
for constant function where we don't have a leaf. this isn't fully general but
we can't do better without a more general post-RA optimizer. i'm not inclined to
do that unless/until we get hot blocks demonstrating its value (that we can
compare against the JIT time hit of the heavier-duty optimizer.)

however this special case we can (and should) optimize for now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig abb41d39e4 RegisterAllocationPass: optimize xgetbv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig d4eb4ef594 IR: plumb CPUID into RA pass
for cpuid folding.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 02f45854e8 IR: include a fence in CPUID
easier for post-RA to chew thru.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Ryan Houdek 62de1004df Merge pull request #4635 from neobrain/fix_libfwd_wl_more
LibraryForwarding/wayland: Add new interface objects
2025-07-02 09:01:11 -07:00
Tony Wasserka 7f216ca02f LibraryForwarding/wayland: Add new interface objects 2025-07-02 16:51:03 +02:00
LC 492b0fdda8 Merge pull request #4632 from neobrain/feature_nix
Build: Add nix-based helpers to facilitate cross-compilation
2025-07-01 16:23:58 -04:00
LC bb072c0112 Merge pull request #4633 from Sonicadvance1/noexec_test
unittests: Adds unittest for no-exec testing
2025-07-01 16:20:44 -04:00
Ryan Houdek afabe7cb47 Merge pull request #4627 from alyssarosenzweig/opt/long-div-peephole-ready
Optimize long division
2025-06-30 13:59:48 -07:00
Ryan Houdek 38e0fc2434 unittests: Adds unittest for no-exec testing
In preparation for #4474
2025-06-30 13:19:33 -07:00
Alyssa Rosenzweig 16a70eafc6 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:59:27 -04:00
Alyssa Rosenzweig de4becc26e OpcodeDispatcher: mask certain divisors
needed for fusing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig af23f4325f OpcodeDispatcher: reorder xor-with-self sequence
this lets us peephole fuse things even when there are flags calculated in the
way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Ryan Houdek 94af96df8f Merge pull request #4631 from neobrain/fix_libfwd_wl_cutter
LibraryForwarding: Fix various Wayland issues
2025-06-30 11:38:13 -07:00
Ryan Houdek e685ab818e Merge pull request #4619 from neobrain/refactor_drop_config_h_in
Remove code generation build step for install prefix
2025-06-30 11:35:26 -07:00
Ryan Houdek 5f2a72b65b Merge pull request #4624 from Sonicadvance1/static_analysis_fixes
Some static code analysis fixes
2025-06-30 11:28:06 -07:00
Tony Wasserka c1842a6167 Build: Add nix-based helpers to manage toolchains for ARM64EC/WOW64
The cmake_configure_woa*.sh scripts will automatically install any required
cross-toolchains required to enable either ARM64EC or WOW64 builds of FEX,
and it will initialize the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell WineOnArm/shell.nix`,
which will make the toolchain available via environment variables. This also
generates a meson crossfile for building VKD3D or vkd3d-proton.
2025-06-30 16:50:13 +02:00
Tony Wasserka 3f3907b5d1 Build: Add nix-based helpers to manage toolchains for FEXLinuxTests
The cmake_enable_flt.sh script will automatically install the required
cross-toolchains required to build FEXLinuxTests, and it will reconfigure
the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell FEXLinuxTests/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 1899465390 Build: Add nix-based helpers to manage toolchains for library forwarding
The cmake_enable_libfwd.sh script will automatically install any required
cross- toolchains and development headers required to enable library
forwarding, and it will reconfigure the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell LibraryForwarding/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 18360d4ccb LibraryForwarding/wayland: Add more method signatures 2025-06-27 10:57:10 +02:00
Tony Wasserka a691c3cd99 LibraryForwarding/wayland: Fix mprotect call when crossing page boundaries 2025-06-27 10:57:10 +02:00
Tony Wasserka 70bc561bbf LibraryForwarding/wayland: Fix wl_proxy_marshal_array_constructor 2025-06-27 10:57:10 +02:00
Tony Wasserka b9222d8431 LibraryForwarding/unittests: Fix build with clang 20 2025-06-27 10:57:10 +02:00
Tony Wasserka a5fad89e57 Merge pull request #4621 from neobrain/feature_update_readme
Update Readme.md
2025-06-26 21:54:53 +02:00
Tony Wasserka a14360b89d Update Readme.md 2025-06-26 21:41:23 +02:00
Tony Wasserka c9aaedd217 CPack: Update package description 2025-06-26 21:41:23 +02:00
Alyssa Rosenzweig cda15ce9ea RegisterAllocationPass: optimize long divsion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 62410c4381 InstCountCI: add another udiv case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:22 -04:00
Ryan Houdek cf82b56dd8 CPUBackend: Remove unused variable
CID 482006
2025-06-19 16:51:30 -07:00