Commit Graph
14214 Commits
Author SHA1 Message Date
LC df73e84725 Merge pull request #5507 from Sonicadvance1/155
CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588
2026-05-21 22:39:14 -04:00
Ryan Houdek 5bf07c2e77 CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588 2026-05-21 18:04:29 -07:00
Ryan Houdek bb0d142a65 Merge pull request #5441 from neobrain/feature_mmap_code_cache
CodeCache: Implement lazy code loading
2026-05-20 17:19:11 -07:00
LC c98cef0da1 Merge pull request #5506 from Sonicadvance1/154
HostFeatures: Only enable `dc zva` optimization on Ampere CPUs
2026-05-19 22:29:11 -04:00
Ryan Houdek 1d9c52be02 InstcountCI: Update 2026-05-19 17:38:10 -07:00
Ryan Houdek a6c9df1a64 HostFeatures: Only enable dc zva optimization on Ampere CPUs
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.

So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.

microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```

microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```
2026-05-19 17:28:39 -07:00
Ryan Houdek f66368b191 Merge pull request #5505 from fixedcat/main
Arm64: Fix byte-size handling in unaligned STLXR emulation
2026-05-19 12:28:23 -07:00
LC e4daea406e Merge pull request #5503 from Sonicadvance1/154
FEXCore: Allow InterruptFaultPage to be significantly further away
2026-05-19 12:22:18 -04:00
fixedcat 6216f22cb9 Arm64: Fix byte-size handling in unaligned STLXR emulation 2026-05-19 18:53:44 +08:00
LC d4c80d9094 Merge pull request #5422 from Sonicadvance1/138
FEXCore: Add support for developer single stepping, read/write watching.
2026-05-19 01:11:04 -04:00
Ryan Houdek 27324ded87 Merge pull request #5499 from bylaws/winstuff
Windows additions for code caching
2026-05-18 16:05:59 -07:00
Ryan Houdek 7d1c625e32 FEXCore: Allow InterruptFaultPage to be significantly further away
We are actually quite close to a single page of CPU state per thread and
any additional changes are likely to cause it to overflow which would
hit these asserts. As we saw with the libc++ implementation of mutexes,
just one object type changing size could push it over the edge.

Future proof this by ensuring we can have this be sixteen pages per
thread before needing to hit more complex implementations. Which I don't
see us getting that large of CPU context tracking.
2026-05-18 15:57:25 -07:00
Ryan Houdek b4fe65f2c0 Merge pull request #5496 from neobrain/fix_libfwd_findpkg
Library Forwarding: Various build system improvements
2026-05-18 14:55:53 -07:00
Ryan Houdek f5efdac2e5 Merge pull request #5423 from bylaws/depenencey
WOW64: Support disabling DEP
2026-05-18 14:39:44 -07:00
Billy Laws 23de875516 Windows: Add NtUnmapViewOfSection prototype 2026-05-17 23:06:02 +01:00
Billy Laws 82030b8286 Windows: Declare winternl relocation APIs 2026-05-17 23:03:11 +01:00
Ryan Houdek af4da43bb8 Merge pull request #5494 from neobrain/fix_determine_va
Allocator: Fix and optimize VA range detection
2026-05-17 14:51:59 -07:00
Ryan Houdek 33f3b8659c Merge pull request #5495 from neobrain/fix_portable_config
Config: Use more sensible default for portable config location
2026-05-17 14:51:04 -07:00
Ryan Houdek ed724a61a7 Merge pull request #5498 from bylaws/evmd
ImageTracker: Support using image IDs as an extended volatile metadata key
2026-05-17 14:49:16 -07:00
Billy Laws 053bd74aa3 Windows/Common: Add ScopedHandle::reset() 2026-05-17 22:35:06 +01:00
Billy Laws 3c0410e59f ImageTracker: Support using image IDs as an extended volatile metadata key 2026-05-17 19:34:41 +01:00
Tony Wasserka e621f6c753 LibraryForwarding/Build: Explicitly look up LLVM headers
This could previously set up incorrect header paths when clang and LLVM were
installed in different directories (such as when using nix).
2026-05-15 16:07:13 +02:00
Tony Wasserka b05f000f42 LibraryForwarding/Build: Try harder to properly discover header locations 2026-05-15 16:07:13 +02:00
Tony Wasserka e60bfc6d23 LibraryForwarding/Build: Allow specifying system header location externally 2026-05-15 16:07:09 +02:00
Tony Wasserka 9a3d3201f9 Config: Use more sensible default for portable config location 2026-05-15 15:59:22 +02:00
Tony Wasserka abf9724424 Allocator: Fix and optimize VA range detection 2026-05-15 15:47:38 +02:00
LC ab9a8c62ab Merge pull request #5493 from Sonicadvance1/153
FEXGetConfig: Even more correctness changes for X2E
2026-05-14 09:39:31 -04:00
Ryan Houdek 1d71650379 FEXGetConfig: Even more correctness changes for X2E
Some of the information was incorrect, so make sure it shows the
hardware correctly.
2026-05-13 20:00:57 -07:00
Tony Wasserka a040740974 CodeCache: Ensure atomicity of code page finalization 2026-05-13 22:53:09 +02:00
Tony Wasserka 5be0dc9fc5 CodeCache: Implement lazy code loading 2026-05-13 21:25:47 +02:00
Tony Wasserka 52ad434d24 LinuxSyscalls: Defer MappedResource deletion until after code invalidation
This ensures that any code buffer memory owned by the MappedResource is
invalidated before being deallocated.
2026-05-13 21:25:47 +02:00
Tony Wasserka d69d111bb6 CodeCache: Align code section within cache files
This allows mapping the code directly into memory for execution.
2026-05-13 21:25:47 +02:00
LC 50f4494875 Merge pull request #5490 from Sonicadvance1/152
FEXGetConfig: Showcase RMW versus loadstore atomic differences
2026-05-12 16:32:04 -04:00
Ryan Houdek 8a4982383a FEXGetConfig: Showcase RMW versus loadstore atomic differences
This differs on X2E, so it's good to showcase it.
2026-05-12 12:38:44 -07:00
Ryan Houdek 0d72890482 Merge pull request #5432 from pmatos/f64-fprem
JIT-inline FPREM/FPREM1 for reduced precision x87 path
2026-05-11 14:31:31 -07:00
Ryan Houdek 9be7d6d112 Merge pull request #5489 from neobrain/feature_better_fexbash
FEXBash: Drop implicit -c and add colored PS1
2026-05-11 12:22:00 -07:00
Tony Wasserka 3c4121ba07 FEXBash: Use a shiny rainbow for PS1 2026-05-11 20:30:21 +02:00
Tony Wasserka b6e44b04d6 FEXBash: Clean up path handling 2026-05-11 20:30:21 +02:00
Tony Wasserka 02c11afbc6 FEXBash: Don't imply "-c" to behave more closely like bash
Implicitly adding "-c" breaks argument passing for scripts. For example, the
command "FEXBash ./steam.sh -silent" will process steam.sh but the script
wouldn't see the "-silent" argument previously.
2026-05-11 19:49:15 +02:00
Paulo Matos 2b8f5b57eb instcountci: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 0519c9467c asm_tests: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 84d968c7d2 JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Ryan Houdek a04b0241c2 Docs: Update for release FEX-2605 FEX-2605 2026-05-08 19:28:30 -07:00
Ryan Houdek 670fd19d33 Merge pull request #5486 from Sonicadvance1/151
Allocator: Mark large unmapped regions as DONTDUMP
2026-05-08 19:28:13 -07:00
Ryan Houdek a66544f3f4 Allocator: Mark large unmapped regions as DONTDUMP
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.

This will speed up coredumps.
2026-05-08 16:27:04 -07:00
Ryan Houdek 1bfb3aefcc Merge pull request #5485 from Sonicadvance1/150
Windows: Setup `tu_override_uncached_as_cache_coherent` inside of dlls
2026-05-08 15:58:43 -07:00
Ryan Houdek b7bfbc3fcd Windows: Setup tu_override_uncached_as_cache_coherent inside of dlls
To not have this environment variable accidently be enabled on arm64
native Wine games, we need to set it from inside of FEX.

Requires the FEX dlls to set them directly rather than launch scripts.
2026-05-08 13:39:51 -07:00
Ryan Houdek e517f3259c Merge pull request #5484 from neobrain/fix_code_cache_no_guest_wrappers
CodeCache: Fix crash when guest library wrappers aren't installed
2026-05-07 10:59:32 -07:00
Tony Wasserka 8afda92a64 CodeCache: Fix crash when guest library wrappers aren't installed 2026-05-07 18:56:35 +02:00
Ryan Houdek ed216c8d4d Merge pull request #5449 from neobrain/opt_code_cache_writing
CodeCache: Slightly optimize cache file writing
2026-05-06 18:06:51 -07:00