Commit Graph
1968 Commits
Author SHA1 Message Date
LC 4f2d2e646e Merge pull request #4542 from Sonicadvance1/fexcore_reconstructions_getting_saved_today
FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
2025-05-01 13:28:16 -04:00
Ryan Houdek fc052efb91 FEXCore: Fixes x87 reduced precision
With the change from #4538 I had accidentally broken x87 reduced
precision.

This is due to the fact that we accidentally lost ABI information about
interpreter fallbacks supporting `preserve_all` or not. So now instead
of having some ABI callbacks supporting it and some not, just force
usage of `preserve_all` if it is supported by the compiler entirely.

Fixes Steam when x87 reduced precision is enabled.
2025-04-30 17:37:10 -07:00
Ryan Houdek c5754145c5 FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
As said in the implementation of this struct commit message. This new
pair struct optimizes specific cases of small forward only increments
that can fit in to 8-bit space, and small forward or backward jump cases
that fit in to 16-bit space.

Some stats of this change:
- Steam: 5.88MB down to 4.34MB. 73.8% the space consumed
- Steamwebhelper: 15.8MB down to 13.24MB. 83.8% space consumed
- Sonic Mania: 3.6MB down to 2.58MB. 71.6% space consumed

As for absolute stats when compared to all code buffer size:
- Steam: 86MB of code buffer to 5.88MB -> 4.34MB of RIP reconstruction.
  - 6.8% -> 5% code buffer space used for RIP reconstruction
- Steamwebhelper: 285MB of code buffer to 17MB -> 14.26MB of RIP reconstruction.
  - 5.9% -> 4.9% code buffer space used for RIP reconstruction
- Sonic Mania: 48.53MB of code buffer to 3.55MB -> 2.53MB of RIP reconstruction.
  - 7.3% -> 5.2% code buffer space used for RIP reconstruction
2025-04-30 15:41:43 -07:00
Ryan Houdek 5b9354cc74 JIT: Move interpreter ABI handlers in to the dispatcher
Fixes #4535

A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher

This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.

This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
2025-04-29 22:35:39 -07:00
Ryan Houdek f7049a6478 Arm64Emitter: On spill return stack used and stop clobbering TMP4
TMP4 was used before we passed in a tmp register. Now use that temp
register.

Also return the amount of stack used on the push function. This will be
used in a bit.
2025-04-29 22:31:35 -07:00
Ryan Houdek e1d032b5a6 Merge pull request #4540 from pmatos/MProtectLastPage
mprotect last page of CodeBuffer
2025-04-27 11:07:14 -07:00
Paulo Matos 791502afef Protect last page of CodeBuffer
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.

In addition, there a couple of clang-tidy fixes which should be NFC.
2025-04-25 20:41:18 +02:00
Tony Wasserka e89a913b69 JIT: Make memory write visible to other threads reading the same location 2025-04-25 09:47:13 +02:00
Tony Wasserka 293d77d412 JIT: Make code patching during (un-/)linking thread-safe 2025-04-25 09:21:09 +02:00
Ryan Houdek 6beb4b0f8b Merge pull request #4533 from Sonicadvance1/align_tail_in_the_pale_moonlight
JIT: Align JITCodeTail to native alignment
2025-04-24 16:18:45 -07:00
Ryan Houdek 32764ddf81 JIT: Moves VPCMPESTRX handler to use vectors
I pushed this off from the previous changes that were converting things
to vector as less important. It has now become more important to keep
these in vector registers until beyond the ABI boundary.

This will reduce burden on our JIT backend and just changes where the
movement in to GPRs occurs. Necessary for #4535
2025-04-24 10:09:26 -07:00
Ryan Houdek 735f537846 JIT: Align JITCodeTail to native alignment
Removes UB
2025-04-24 08:40:43 -07:00
Ryan Houdek b7790e10e9 Merge pull request #4532 from pmatos/X87StateBlockReset
X87 state block reset
2025-04-24 08:37:32 -07:00
Paulo Matos d377e26106 Reset MMXState to X87 at the start of each block
Ensures that blocks always start with the same state independently of predecessors
which allows independent compilation of blocks.
Starting in the X87 state is better than starting in MMX state because
MMX state is more work to initialize.
2025-04-24 15:02:41 +02:00
Ryan Houdek adcefe62cf InstcountCI: Changes how code size is calculated
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.

Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.

So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.

This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.

This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.

This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
2025-04-23 12:56:46 -07:00
Billy Laws 30fefbd57d X86Tables: Set FLAGS_BLOCK_END for more faulting ops 2025-04-19 14:19:19 +01:00
Tony Wasserka 0fe28129a0 JIT: Fix warning about unused variable 2025-04-18 11:16:46 +02:00
LC 572d6e0395 Merge pull request #4509 from Sonicadvance1/in_the_twilight_of_the_pale_blue_moon
A couple of barrier and timing fixes.
2025-04-17 17:30:03 -04:00
Billy Laws 416267a238 OpcodeDispatcher: Safely clobber NZCV in FCOMIF64
Also fix a small typo that broke the !flagm2 path.
2025-04-16 13:06:37 +01:00
Ryan Houdek 00f8181c3a OpcodeDispatcher: Fix CPUID being a instruction fence
It is common practice for games to use CPUID as an instruction barrier
for various reasons. Ensure that we respect this by adding support for
an instruction barrier.
2025-04-15 15:42:43 -07:00
Ryan Houdek 60565cc2ef OpcodeDispatcher: Implement support for self-synchronizing cycle counter
FEAT_ECV added a new synchronizing cycle counter instruction that
restrict speculation across the cycle counter access. Because it
restricts speculation, it effectively acts like an isb and load dsb.

Luckily for us, this actually matches behaviour for what rdtscp does, so
we can take advantage of it if the host supports FEAT_ECV.
2025-04-15 15:42:43 -07:00
Paulo Matos cb972e165b Fix cast in ValidateCode impl 2025-04-15 09:35:46 +02:00
Ryan Houdek 211bec65d2 Merge pull request #4492 from bylaws/badencodings
Frontend: Be more tolerant of bad instruction encodings
2025-04-13 19:09:07 -07:00
Billy Laws 039aaae041 FEXCore: Add a pre-compilation frontend callback to SyscallHandler 2025-04-11 12:07:16 +01:00
Billy Laws 9625201cbf Frontend: Remove asserts on invalid instruction encodings 2025-04-11 12:06:01 +01:00
Billy Laws a1c6378317 Frontend: Only accept POP opcodes with a 0 ModRM.reg field 2025-04-09 23:04:11 +01:00
Billy Laws f064013b9a Frontend: Enforce FLAGS_SF_MOD_MEM_ONLY/FLAGS_SF_MOD_REG_ONLY 2025-04-09 23:02:09 +01:00
Billy Laws d4480c3566 X86Tables: Mark VMOVNTDQ as FLAGS_SF_MOD_MEM_ONLY 2025-04-09 23:01:26 +01:00
Billy Laws b3a4de7aa8 X86Tables: Mark MASKMOVQ as FLAGS_SF_MOD_REG_ONLY 2025-04-09 23:01:26 +01:00
Billy Laws 60d8131a4d X86Tables: Drop FLAGS_SF_MOD_MEM_ONLY from (V)MOV(L/H)PS
These encodings are shared with MOVLHPS/MOVHLPS and FEX handles both variants.
2025-04-09 23:01:26 +01:00
Billy Laws 6d78aefa13 Frontend: Ignore REX register extension for MMX registers 2025-04-09 23:01:26 +01:00
Ryan Houdek 7ca757bb6d FEXCore: Move CPUInfo to FEX
This is only ever used in the frontend now.
2025-04-08 22:54:43 -07:00
Ryan Houdek 5767c61a91 FEXCore/Emitter: Stop creating a vector on the heap
In the Push/Pop CalleeSavedRegisters these vectors were getting created
on the heap, allocating memory and then just iterating them.

Just use a std::array which makes it stop allocating memory and saves
the number of instructions.
2025-04-08 18:36:47 -07:00
Ryan Houdek ec1c7797f4 OpcodeDispatcher: Implement support for sha256rnds2 using ARM instructions
The big one.
2025-04-02 12:51:03 -07:00
Tony Wasserka 0bd924eb7e Merge pull request #4460 from Sonicadvance1/dead_code
FEXCore/Frontend: Remove logically dead code
2025-04-02 09:20:41 +02:00
Ryan Houdek 9b0bb29d78 IR: Implement support for sha256h{2,}
I keep carrying this patch around. Not yet wired up to the instruction
implementation yet, but I don't want to forget about it.
2025-04-01 17:53:48 -07:00
Ryan Houdek c3e71de1d7 FEXCore/Frontend: Remove logically dead code
This code can't get hit.
2025-04-01 16:09:07 -07:00
Ryan Houdek 1b18bfaff5 Merge pull request #4473 from bylaws/win32-f
Windows: Small fixups
2025-04-01 10:53:28 -07:00
Ryan Houdek 8aecdc536c Merge pull request #4471 from alyssarosenzweig/opt/cvtss2si
Optimize float->integer conversions with Feat_FRINTTS
2025-04-01 08:52:50 -07:00
Billy Laws 8f50106187 Context: Fix incorrect ifdef on ARM64EC 2025-03-31 23:57:58 +01:00
Billy Laws 02d3a319f9 OpTables: Disable thunk opcodes on win32
They are of no use here, and are quite frequent in never-taken blocks in Denuvo games
so treating them as invalid avoids wasting some time.
2025-03-31 23:57:58 +01:00
Ryan Houdek cdaf1c5262 Arm64Emitter: Removes warning 2025-03-31 13:51:50 -07:00
Alyssa Rosenzweig c4f7b27459 OpcodeDispatcher: accelerate F->I conversions with FRINTTS
this should significantly help perf on supported platforms.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:09:54 -04:00
Alyssa Rosenzweig 1f08f8df0d IR: allow VUShrNI with bitshift=0
encodes to Xtn, we need this to narrow.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Alyssa Rosenzweig 0038a0b19c IR: plumb Vector_FToISized op
this exposes the frint* opcodes in a new ir op

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-31 16:03:18 -04:00
Ryan Houdek f9b369c550 Telemetry: Removes unnecessary indirection
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.

Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).

This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
2025-03-29 15:10:37 -07:00
Ryan Houdek 0599d80b13 FEXCore/Frontend: Changes how VEX operand encoding flags are encoded
These three options are mutually exclusive with each other and could
potentially result in invalid encodings of the table on accident.

Change over to a 2-bit bitfield to encode if the operand that consumes
the VEX option is none, destination, 1st src, or 2nd src.

This ensures the table can't ever be incorrectly encoded.
2025-03-29 13:51:02 -07:00
Lioncache 88e6c48db7 X86Tables: Replace deprecated std::is_trivial template
This is deprecated in C++26, so we can just use a more specific type trait.
2025-03-29 00:55:15 -04:00
Ryan Houdek 0ca34d11ad Merge pull request #4454 from alyssarosenzweig/silly-nop
Fix 66 90 decoding to a nop
2025-03-28 14:50:02 -07:00
Alyssa Rosenzweig 3a2ca41724 OpcodeDispatcher: handle 66 90 as a NOP
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-03-28 13:03:34 -04:00