Commit Graph
14517 Commits
Author SHA1 Message Date
Ryan Houdek db9414a756 Merge pull request #5655 from neobrain/feature_woa_cache_loading
Windows/ImageTracker: Adapt code cache loading logic to FEXOfflineCompiler
2026-07-06 12:40:40 -07:00
Ryan Houdek 5a0bf1bb5f Merge pull request #5657 from neobrain/fix_foc_syscall_abi_woa
FEXOfflineCompiler: Fix improper syscall ABI on WoA
2026-07-06 12:37:52 -07:00
LC 444c37fe2e Merge pull request #5656 from neobrain/fix_invalid_iterator_deref
Core: Fix dereference of invalid iterator
2026-07-06 10:13:04 -04:00
Tony Wasserka d3d735370f FEXOfflineCompiler: Fix improper syscall ABI on WoA 2026-07-06 15:16:30 +02:00
Tony Wasserka 852e93aa74 Core: Fix dereference of invalid iterator 2026-07-06 15:07:37 +02:00
Tony Wasserka dc1be2efe8 Windows/ImageTracker: Unindent refactored code 2026-07-06 14:58:16 +02:00
Tony Wasserka f2734ac608 Windows/ImageTracker: Adapt code cache loading logic to FEXOfflineCompiler 2026-07-06 14:58:16 +02:00
Tony Wasserka 57ca49dc5c Merge pull request #5501 from bylaws/finishloadwin
Windows/ImageTracker: Wire up LoadCache and EnableLoadedSection
2026-07-06 14:58:01 +02:00
Billy Laws 5ef3134a4d Windows/ImageTracker: Wire up LoadCache and EnableLoadedSection
The lazy code loading refactor replaced LoadData with the new
LoadCache/EnableLoadedSection API but left the Windows path as TODOs.
Implement the wiring: LoadAOTImages now calls LoadCache +
RegisterMappedCodeBuffer for each mapped cache file, and HandleImageMap
calls EnableLoadedSection (with nullptr thread since lazy mapping is not
yet implemented on Windows).
2026-07-06 14:44:59 +02:00
Ryan Houdek 5b91642883 Merge pull request #5654 from ShadowCurse/cache_va_size
Allocator: fix the caching of host va size
2026-07-05 18:32:10 -07:00
Egor Lazarchuk 1eb5abb8db Allocator: rename DetermineVASize to GetHostVABits
`DetermineVASize` does not return the size of VA, but the number of bits
it can use. Change the naming to make it more self explanatory.
In the mean time also move `HostVASize` global into `GetHostVABits`
since it is not and should not be used directly.
2026-07-05 12:46:41 +01:00
Egor Lazarchuk a65e1bf7e5 Allocator: fix the caching of host va size
Commit abf9724 ("Allocator: Fix and optimize VA range detection")
removed assignment to the `HostVASize` global thus making each call to
`DetermineVASize` redo all the work with potential to produce incorrect
results. Set the global again to fix this.
2026-07-05 12:46:29 +01:00
Ryan Houdek 3f98202b0f Merge pull request #5651 from lioncash/perm
unittests: Add stress tests for VPERMIL{PD, PS}
2026-07-04 12:13:24 -07:00
LC e7727c39f3 unittests: Add stress tests for VPERMIL{PD, PS}
While unlikely to be used in practice over other kind of
shuffling and blending, these should also have stress tests
to make sure they do the right thing.
2026-07-04 15:00:30 -04:00
Ryan Houdek 17e5637664 Merge pull request #5650 from lioncash/aes256
[SVE256] Handle 256-bit AES operations
2026-07-03 19:06:36 -07:00
LC 7fd9b897c2 [SVE256] Handle 256-bit AES operations
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.

Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
Ryan Houdek f8967aa207 Merge pull request #5649 from lioncash/pclmul256
[SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
2026-07-03 18:33:53 -07:00
LC 5145324806 [SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
Since vixl now handles this, we can drop this support right in.
2026-07-03 21:22:05 -04:00
Ryan Houdek 4233fb6270 Merge pull request #5648 from lioncash/aes
[SVE256] Ensure SSE insertion behavior for AES/SHA/PCLMUL operations
2026-07-03 17:33:19 -07:00
LC d11b19fd2b [SVE256] Ensure insertion behavior for PCLMUL SSE operations
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC 684c568033 [SVE256] Ensure insertion behavior for SHA SSE operations
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC ab4fb7b3ad [SVE256] Ensure insertion behavior for AES operations on SSE
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
Ryan Houdek 91017dbedb Merge pull request #5647 from lioncash/perm128
unittests: Add stress test for vperm2f128/vperm2i128
2026-07-03 15:32:34 -07:00
LC f5e8a051a8 unittests: Add stress test for vperm2f128/vperm2i128
Lets us test all possible immediate encodings for proper behavior.
2026-07-03 16:17:12 -04:00
LC 1d695f6db4 Merge pull request #5646 from Sonicadvance1/181
Github: More dependabot things
2026-07-02 22:22:45 -04:00
Ryan Houdek d4bdfd0592 Github: More dependabot things
They just never stop.
2026-07-02 19:05:12 -07:00
Ryan Houdek 1cc4b93e7a Docs: Update for release FEX-2607 FEX-2607 2026-07-02 17:47:31 -07:00
LC b4e2f5118a Merge pull request #5645 from Sonicadvance1/180
Windows: Fixes SHM stats reallocation
2026-07-02 19:33:37 -04:00
Ryan Houdek 812b6398e5 Windows: Fixes SHM stats reallocation
This was accidentally setting `CurrentSize` instead of just returning
the newly allocated size to the frontend. This was causing the frontend
to then fail to detect the reallocation actually occured and no longer
get stats for new threads.

Also happened to not use `NewSize` but instead `CurrentSize * 2` which
didn't matter as it matched the growth pattern, but was technically
incorrect.

Fixes SHM stats since the introduction of the unixlib, ezpz.
2026-07-02 13:38:14 -07:00
LC 1db45e2a70 Merge pull request #5639 from Sonicadvance1/179
FEXServerClient: Workaround sun_path 108 byte limit
2026-07-02 04:03:48 -04:00
Ryan Houdek 6bcadde658 Merge pull request #5644 from lioncash/vmov
VectorOps: Eliminate unnecessary moves in VMov if applicable
2026-07-01 17:28:42 -07:00
Ryan Houdek 5f1c8efe0e Merge pull request #5643 from lioncash/rec
VectorOps: Avoid temporary if able in 256-bit VFRecp
2026-07-01 17:25:33 -07:00
Ryan Houdek b9aeccf13b Merge pull request #5642 from lioncash/feature
HostFeatures: Put SVE support querying into single function
2026-07-01 15:29:02 -07:00
Ryan Houdek 16f90b33f3 Merge pull request #5641 from lioncash/minmax
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
2026-07-01 14:17:18 -07:00
Ryan Houdek 5d8d052a77 Merge pull request #5640 from lioncash/move
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
2026-07-01 12:41:03 -07:00
Tony Wasserka 7fa4d78269 Merge pull request #5635 from Sonicadvance1/178
FEXOfflineCompiler: Fixes HostFeature detection under Win32
2026-07-01 17:15:13 +02:00
Ryan Houdek 6bf0db7df6 FEXServerClient: Workaround sun_path 108 byte limit
We really don't want to do this, but in the case that the AF_UNIX path
is longer than the 108-byte limit that sun_path provides we don't really
have a choice. The alternative choice would be to switch /entirely/ away
from AF_UNIX and instead use pipes. We need a bandage fix for now, so
throw the socket in to a temp folder if the path is too long.
2026-06-30 17:31:23 -07:00
Ryan Houdek 110313e7de Merge pull request #5638 from lioncash/swap
IR: Add constant for swapping midsections of 256-bit vectors around
2026-06-30 16:38:56 -07:00
Ryan Houdek 44e24c9e6b Merge pull request #5637 from mrpippy/unicodestring
Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism
2026-06-30 16:23:08 -07:00
Brendan Shanks 6c58fef220 Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism. 2026-06-30 15:37:55 -07:00
LC 3d593ce87d Merge pull request #5626 from Sonicadvance1/177
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 18:15:17 -04:00
Ryan Houdek 36e1b5107e InstcountCI: Update 2026-06-30 15:01:46 -07:00
Ryan Houdek f89123f489 unittests/ASM: Allow up to 3-bits of precision loss 2026-06-30 15:00:16 -07:00
Paulo Matos b7280a765d asm_tests: Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Ryan Houdek a4f89b79a3 Merge pull request #5636 from lioncash/move
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
2026-06-30 12:35:13 -07:00
Ryan Houdek 5bde4d875a FEXOfflineCompiler: Fixes HostFeature detection under Win32
CPUFeature detection is marginally different between Linux and Windows.
FOC was only using the Linux path which had two broken things happening
to it.
- Feature detection was incorrect and enabling/disabling features
  differently from wow64/arm64ec .dll files
- HostType was being set as Linux even though it was generating code for
  WINE

Ensure that when built for Win32 that it uses the correct feature
fetching.

One thing that is still incorrect is that 64-bit or 32-bit is determined
at compile time on win32, whereas the Linux side parses an ELF and
determines bitness at runtime. This doesn't fix that remaining problem
there.
2026-06-30 11:39:02 -07:00
Ryan Houdek a79c471c31 Merge pull request #5634 from lioncash/shuffle
unittests: Add selector tests for VPSHUF{D, HW, LW}
2026-06-30 11:26:09 -07:00
LC c1e29f9013 VectorOps: Eliminate unnecessary moves in VMov if applicable
If the destination and source don't match, then we can just zero
and insert directly into the destination instead of a temporary.
2026-06-30 05:06:26 -04:00
LC 4c27dfd5eb VectorOps: Avoid temporary if able in 256-bit VFRecp
If we're non-aliasing, we can make use of the destination reg directly.
Makes the non-RPRES path a little nicer.
2026-06-30 04:36:20 -04:00