Commit Graph
1995 Commits
Author SHA1 Message Date
Ryan Houdek cdea8d7f74 Merge pull request #4593 from alyssarosenzweig/ici/zeroing-sub-regs
InstructionCountCI: add more cases for mov 0/~0
2025-05-29 12:26:15 -07:00
Ryan Houdek ef6dc3d802 Merge pull request #4592 from alyssarosenzweig/opt/x87-tag
OpcodeDispatcher: optimize X87FTWTag
2025-05-29 12:26:04 -07:00
Alyssa Rosenzweig 0fbe69ebcf OpcodeDispatcher: optimize xor-with-self flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:53:38 -04:00
Alyssa Rosenzweig ece817c691 OpcodeDispatcher: optimize logical flags
seems to be strictly better.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:34:50 -04:00
Alyssa Rosenzweig 3fbc8204b7 OpcodeDispatcher: clean up logical flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:28:31 -04:00
Alyssa Rosenzweig b8dd5d95b0 OpcodeDispatcher: optimize X87FTWTag
using bit twiddling tricks :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:13:16 -04:00
Ryan Houdek 7cd52febc2 FEXCore/CPUID: Remove warning
This is currently only used on win32 builds because Linux doesn't
understand TPIDRRO.
2025-05-28 09:28:41 -07:00
Alyssa Rosenzweig 6927c7577a RegisterAllocationPass: delete trivial instructions
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 61ae53cc03 IR: add paired PushTwo/PopTwo helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 0e67f30103 OpcodeDispatcher: make Push do the right thing and use it
this both optimizes and bug-fixes pusha while deleting a snotton of code.

Closes: https://github.com/FEX-Emu/FEX/issues/4589
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Ryan Houdek dc9f8aa855 Merge pull request #4580 from alyssarosenzweig/ir/inline-ra
IR: Inline registers into the IR
2025-05-26 09:44:51 -07:00
StanfordZhang 23cda2c961 Update Arm64Emitter.cpp
fix callee saved floating-point arguments issue
2025-05-23 22:14:37 +08:00
Alyssa Rosenzweig 65ee1fafa8 IR: drop RegisterAllocationData
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 49f8332c5b JIT: use registers directly from the IR
This is the flag day change from the series, using all the new shiny
infrastructre we added.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig afce108ed7 JIT: make almost all the DEF_OPs common
this deduplicates a bunch of #defines, letting us change the signature easier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 30f3b545af RegisterAllocationPass: use Header spill slots instead
Removes even more RAData dependence.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig bf1597920c IR: use post-RA flag
rather than implicitly depending on the RA data.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig a01e29ac99 IR: extend the IR header with RA info
Beyond the actual registers allocated, there are two pieces of sideband data we
store in the RAData object:

* # of spill slots (explicitly)
* whether RA has run (implicitly by the existence of RAData)

We want to get rid of RAData, so we'll move these to the header.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig c7797d56c9 IR: don't use GetOpSize in ExitFunction
nothing else does this, and it complicates upcoming refactor to move away from
IR builder helpers doing IR dereferencing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:51:11 -04:00
Alyssa Rosenzweig dfd1aedae5 JIT: use .ID() even less
oops, missed a "/g" with th sed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 16:23:22 -04:00
Alyssa Rosenzweig ecc6fea54e JIT: stop using .ID() pattern
sed -ie 's/.ID()//' *

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:46:32 -04:00
Alyssa Rosenzweig 99446da7c1 JIT: add Reg helpers taking OrderedNodeWrappers
more ergonomic and will give us freedom to migrate things easier soon.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:43:20 -04:00
Tony Wasserka 39a5c2021e ThreadPoolAllocator: Rename FixedSizePoolAllocation to PoolBufferWithTimedRetirement
This more accurately reflects that the core feature of the helper is the
timer-based unclaiming of buffers instead of the allocation size.
2025-05-14 13:29:44 +02:00
LC 1dff7073de Merge pull request #4554 from Sonicadvance1/remove_unused_argument
NFC: FEXCore: Removes unused argument on CreateThread
2025-05-05 14:30:24 -04:00
Ryan Houdek 92a82c3134 FEXCore: Removes unused argument on CreateThread
ParentTID is purely a Linux construct and has been moved entirely to the
frontend at this point. Remove this argument which is now unused.
2025-05-05 11:17:02 -07:00
Tony Wasserka cdaa65f6fd Arm64Emitter: Fix overalignment in Align16B
Previously, 16 additional bytes were emitted if the buffer was already
aligned.
2025-05-05 16:02:42 +02:00
Ryan Houdek 7ed17f68c5 FEXCore: Remove unused InvalidateGuestCodeRange with callback 2025-05-02 01:15:50 -07:00
LC 4f2d2e646e Merge pull request #4542 from Sonicadvance1/fexcore_reconstructions_getting_saved_today
FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
2025-05-01 13:28:16 -04:00
Ryan Houdek fc052efb91 FEXCore: Fixes x87 reduced precision
With the change from #4538 I had accidentally broken x87 reduced
precision.

This is due to the fact that we accidentally lost ABI information about
interpreter fallbacks supporting `preserve_all` or not. So now instead
of having some ABI callbacks supporting it and some not, just force
usage of `preserve_all` if it is supported by the compiler entirely.

Fixes Steam when x87 reduced precision is enabled.
2025-04-30 17:37:10 -07:00
Ryan Houdek c5754145c5 FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
As said in the implementation of this struct commit message. This new
pair struct optimizes specific cases of small forward only increments
that can fit in to 8-bit space, and small forward or backward jump cases
that fit in to 16-bit space.

Some stats of this change:
- Steam: 5.88MB down to 4.34MB. 73.8% the space consumed
- Steamwebhelper: 15.8MB down to 13.24MB. 83.8% space consumed
- Sonic Mania: 3.6MB down to 2.58MB. 71.6% space consumed

As for absolute stats when compared to all code buffer size:
- Steam: 86MB of code buffer to 5.88MB -> 4.34MB of RIP reconstruction.
  - 6.8% -> 5% code buffer space used for RIP reconstruction
- Steamwebhelper: 285MB of code buffer to 17MB -> 14.26MB of RIP reconstruction.
  - 5.9% -> 4.9% code buffer space used for RIP reconstruction
- Sonic Mania: 48.53MB of code buffer to 3.55MB -> 2.53MB of RIP reconstruction.
  - 7.3% -> 5.2% code buffer space used for RIP reconstruction
2025-04-30 15:41:43 -07:00
Ryan Houdek 5b9354cc74 JIT: Move interpreter ABI handlers in to the dispatcher
Fixes #4535

A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher

This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.

This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
2025-04-29 22:35:39 -07:00
Ryan Houdek f7049a6478 Arm64Emitter: On spill return stack used and stop clobbering TMP4
TMP4 was used before we passed in a tmp register. Now use that temp
register.

Also return the amount of stack used on the push function. This will be
used in a bit.
2025-04-29 22:31:35 -07:00
Ryan Houdek e1d032b5a6 Merge pull request #4540 from pmatos/MProtectLastPage
mprotect last page of CodeBuffer
2025-04-27 11:07:14 -07:00
Paulo Matos 791502afef Protect last page of CodeBuffer
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.

In addition, there a couple of clang-tidy fixes which should be NFC.
2025-04-25 20:41:18 +02:00
Tony Wasserka e89a913b69 JIT: Make memory write visible to other threads reading the same location 2025-04-25 09:47:13 +02:00
Tony Wasserka 293d77d412 JIT: Make code patching during (un-/)linking thread-safe 2025-04-25 09:21:09 +02:00
Ryan Houdek 6beb4b0f8b Merge pull request #4533 from Sonicadvance1/align_tail_in_the_pale_moonlight
JIT: Align JITCodeTail to native alignment
2025-04-24 16:18:45 -07:00
Ryan Houdek 32764ddf81 JIT: Moves VPCMPESTRX handler to use vectors
I pushed this off from the previous changes that were converting things
to vector as less important. It has now become more important to keep
these in vector registers until beyond the ABI boundary.

This will reduce burden on our JIT backend and just changes where the
movement in to GPRs occurs. Necessary for #4535
2025-04-24 10:09:26 -07:00
Ryan Houdek 735f537846 JIT: Align JITCodeTail to native alignment
Removes UB
2025-04-24 08:40:43 -07:00
Ryan Houdek b7790e10e9 Merge pull request #4532 from pmatos/X87StateBlockReset
X87 state block reset
2025-04-24 08:37:32 -07:00
Paulo Matos d377e26106 Reset MMXState to X87 at the start of each block
Ensures that blocks always start with the same state independently of predecessors
which allows independent compilation of blocks.
Starting in the X87 state is better than starting in MMX state because
MMX state is more work to initialize.
2025-04-24 15:02:41 +02:00
Ryan Houdek adcefe62cf InstcountCI: Changes how code size is calculated
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.

Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.

So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.

This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.

This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.

This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
2025-04-23 12:56:46 -07:00
Billy Laws 30fefbd57d X86Tables: Set FLAGS_BLOCK_END for more faulting ops 2025-04-19 14:19:19 +01:00
Tony Wasserka 0fe28129a0 JIT: Fix warning about unused variable 2025-04-18 11:16:46 +02:00
LC 572d6e0395 Merge pull request #4509 from Sonicadvance1/in_the_twilight_of_the_pale_blue_moon
A couple of barrier and timing fixes.
2025-04-17 17:30:03 -04:00
Billy Laws 416267a238 OpcodeDispatcher: Safely clobber NZCV in FCOMIF64
Also fix a small typo that broke the !flagm2 path.
2025-04-16 13:06:37 +01:00
Ryan Houdek 00f8181c3a OpcodeDispatcher: Fix CPUID being a instruction fence
It is common practice for games to use CPUID as an instruction barrier
for various reasons. Ensure that we respect this by adding support for
an instruction barrier.
2025-04-15 15:42:43 -07:00
Ryan Houdek 60565cc2ef OpcodeDispatcher: Implement support for self-synchronizing cycle counter
FEAT_ECV added a new synchronizing cycle counter instruction that
restrict speculation across the cycle counter access. Because it
restricts speculation, it effectively acts like an isb and load dsb.

Luckily for us, this actually matches behaviour for what rdtscp does, so
we can take advantage of it if the host supports FEAT_ECV.
2025-04-15 15:42:43 -07:00
Paulo Matos cb972e165b Fix cast in ValidateCode impl 2025-04-15 09:35:46 +02:00
Ryan Houdek 211bec65d2 Merge pull request #4492 from bylaws/badencodings
Frontend: Be more tolerant of bad instruction encodings
2025-04-13 19:09:07 -07:00