Commit Graph
433 Commits
Author SHA1 Message Date
Stefanos Kornilios Misis Poiitidis 3a423fbc41 Synchronized Block Linking 2022-08-08 03:33:23 +03:00
lioncash 9e524a30d6 General: Resolve fmt deprecation warnings
fmt 9.0.0 deprecates implicit conversions of unscoped enum values to
integers to be consistent with scoped enum behavior.

Fairly trivial to resolve
2022-08-04 11:39:51 -04:00
lioncash 5f9052a675 CoreState: Use a union for representing xmm data
In the event that the host doesn't support the requirements for running
AVX-enabled applications (SVE2 with at least 256-bit wide vectors) but
still wanted to run regular SSE-enabled applications, they would be
taking a performance hit due to a load/store pessimization (necessary in
order for 256-bit loads/stores to work)

However, we can add an alternate view into the xmm data that would allow
those hosts to use the previous optimization, while still supporting
AVX-capable hosts.
2022-07-26 16:56:54 -04:00
Ryan Houdek 8037231376 Merge pull request #1850 from Sonicadvance1/fix_faulting_instructions
FEXCore: Fix-up edge case behaviour on faulting instructions
2022-07-19 10:23:12 -07:00
Ryan Houdek ff1d51c7bd FEXCore: Fix-up edge case behaviour on faulting instructions
x86 has six instructions that will fault on us that we mostly handled. A
few of these weren't being handled correctly.

One problem is that RIP needs to synchronize differently depending on
which fault instruction it is. Some instructions fault at the
instruction RIP, some at the instruction afterwards.

Additionally some of the metadata generated around the signal delegation
wasn't correct.

With behaviour of all of these instructions changed, it will now be
easier to switch over to non-faulting guest synchronous signals off of
these instructions. This doesn't go far enough to change that behaviour
yet.

I noticed this when looking at Elden Ring's weird faulting behaviour
with ud2 and `int 0x2d`. This made me investigate since I had a
suspicion that we weren't handling both cases correctly.

With this change, Elden Ring is now stabilized and works under FEX.
2022-07-17 20:56:31 -07:00
Ryan Houdek ba0887defa Misc: Convert assert logs to assume+assert that can be
Most of these won't make a performance difference. But we should be
using the assume version everywhere we can.
2022-07-17 12:51:43 -07:00
lioncash 34aa6bf6c7 CoreState: Move flag variables into padding
Allows for a smaller struct size.

This movement of the struct members also requires us to modify the
DeadContextStorePass to take the new locations into account.

On the plus side, we get to remove all of the padding bits in the
CPUState struct, so we can remove the handling for them.
2022-07-13 16:57:16 -04:00
lioncash 9d437b8863 CoreState: Expand xmm registers
Expands them to add the high lanes added in AVX
2022-07-13 14:29:48 -04:00
Ryan Houdek fb41ba172d Merge pull request #1835 from wannacu/main
AOTIR: Fix IRList delete
2022-07-07 10:03:33 -07:00
Tony Wasserka 4f8525da4d ValueDominanceValidation: Avoid stack exhaustion when aggregating predecessors
The recursive algorithm used here previously led to deeply nested function
calls, which eventually exhausted the available stack space. The simple
non-recursive algorithm used now avoids this problem at the expense of
small overhead.
2022-07-07 16:46:38 +02:00
wannacu 227462e4d8 AOTIR: Fix IRList delete 2022-07-07 14:39:00 +08:00
Stefanos Kornilios Misis Poiitidis 4139332ad9 GDBSymbols: Cleanups 2022-06-27 14:44:07 +03:00
Stefanos Kornilios Misis Poiitidis ae5cfcc249 Symbols: Add SymName function 2022-06-24 19:38:40 +03:00
Stefanos Kornilios Misis Poiitidis 8f578b57f2 GDBSymbols: Cleanups 2022-06-24 17:02:39 +03:00
Stefanos Kornilios Misis Poiitidis 3871646611 GDBSymbols: Add gdb reader for fex, GDBSymbols option to enable 2022-06-24 16:27:16 +03:00
Stefanos Kornilios Misis Poiitidis 9ce94266e1 IR: Remove GuestCallDirect, GuestCallIndirect 2022-06-23 19:46:25 +03:00
Tony Wasserka f27830bf41 Fix inconsistent allocation schemes used for RegisterAllocationData 2022-06-23 17:16:44 +02:00
lioncash 58ae49372c CoreState: Add register size constants
Gets rid of some magic numbers and reduces the number of things that
need to manually change (e.g. when supporting AVX and needing to
increase the xmm size).
2022-06-17 10:55:07 -04:00
Ryan Houdek 05d15fa052 AOTIR: Fix RAData free
This thing is allocated with malloc so it needs to be free'd with free.
This was poisoning asan runs.
2022-06-16 17:01:06 -07:00
Stefanos Kornilios Misis Poiitidis 2c3baaad1c Context: Split LocalIR to PrecompiledIR and DebugStore 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Misis Poiitidis 01bd43aa2b IPR: Store IR in the code bufffer, use executer function, cleanups 2022-06-10 10:39:42 +03:00
Stefanos Kornilios Mitsis Poiitidis 2d3c6efae2 Merge pull request #1761 from FEX-Emu/skmp/ir-meta-fns
IR: add IsFragmentExit, IsBlockExit
2022-06-10 08:15:03 +03:00
Ryan Houdek 9e9ceb3894 Merge pull request #1762 from lioncash/pcl
OpcodeDispatcher: Handle CLMUL opcode extension
2022-06-09 12:55:54 -07:00
Stefanos Kornilios Misis Poiitidis bd70088877 IR: add IsFragmentExit, IsBlockExit 2022-06-09 20:44:24 +03:00
lioncash 0b24758f41 IR: Add PCLMUL IR opcode 2022-06-09 11:22:29 -04:00
lioncash a4e2e36243 IR.json: Correct 'Dest' key to 'Desc'
In a few places the description key was accidentally written as Dest.
2022-06-08 10:58:49 -04:00
Ryan Houdek 4bfd1dde1f IR: Implements support for Yield IR op
This is an IR op that produces nor consumes any SSA values, but has side
effects.

Turns in to the pause instruction on x86 and yield instruction on
AArch64.
2022-05-26 08:46:14 -07:00
Stefanos Kornilios Misis Poiitidis a284adcd19 SMC: Add mprotect based tracking, --smc=mtrack, make default 2022-05-16 15:51:09 +03:00
Ryan Houdek 2feae06209 Arm64: Adds support for RCPC2 extension
This allows us to have RCPC loadstore operations with a 9-bit signed
offset.

This gives us a small range of [-256,256) of immediate encoding range on
our TSO loadstore operations.
Updates the inline constant pass in ConstProp to support this range on
TSO IR ops if the host supports RCPC2.

Apple M1 supports this extension, didn't test with Cortex-X2/A710.
2022-05-12 07:13:47 -07:00
Stefanos Kornilios Misis Poiitidis 94d2ed85a7 Core: Add support for process-wide code invalidation, rename IR invalidate op to do thread specific invalidation 2022-05-02 01:49:49 +03:00
Ryan Houdek 8a7f39559c Merge pull request #1627 from Sonicadvance1/wip_reclaimable_pool_allocator
FEXCore: Reclaimable thread pool allocator
2022-04-26 10:21:36 -07:00
Ryan Houdek 753d0ede6c FEXCore: Reclaimable thread pool allocator
Creates a pool allocator for OpcodeDispatcher and IRCompaction that
shares memory allocations between threads in a pool and supports
reclaiming stale allocations from participating threads.

A thread will use a heuristic to keep its claimed memory allocation
around if it is allocating a lot of code. If it slows down then it will
start putting the memory allocation back in to the thread pool.

Additionally if the allocation has been "disowned" and gone to sleep
while still retaining the allocation, then another thread can inspect
 these stale allocations and reclaim it from the idling thread. Saving
further memory.

This needs some more work and cleanup but this is an interesting concept
that saves a decent amount of memory even in a basic test.

Causes teeworlds' title screen to go from 754MB to 599MB in my simple
test. 79.4% the memory usage is a good start.
2022-04-26 10:01:56 -07:00
wannacu 7b379fc3cf AOTIR: copy RAData and IRList, make sure data is accessible 2022-04-26 09:27:18 +08:00
Ryan Houdek 42a6320935 Merge pull request #1585 from CallumDev/x87f64
Emulate reduced-precision X87 with 64-bit host FPU ops
2022-04-21 09:22:03 -07:00
CallumDev d4d5f4d1dd Introduce F64 codegen for reduced precision X87 2022-04-22 01:36:44 +09:30
Ryan Houdek 253333a4cf CompileService: Removes no longer necessary service thread
Since we are masking signals before compiling code, we no longer will
receive a signal in the middle of compiling code.

This makes the compile service never be invoked so we can just remove
it.

We still have some locations in the syscall handling that isn't signal
safe, but compileservice wouldn't have fixed those anyway.
2022-04-19 18:52:01 -07:00
Ryan Houdek fb69300397 FEXCore: Delete IR after it is used
For the JIT cores we don't need to keep IR around, it's only necessary
for the Interpreter. So once the AOT IR service is done dealing with the
IR, check to see if we can delete it.

This causes teeworld's title screen memory usage to go from 730MB to
566MB. 77.5% the memory usage there.

This is effectively an infinite memory leak if the codespace wasn't ever
overwritten or invalidated. So larger memory usage programs would end up
having a larger impact.
2022-03-13 19:01:40 -07:00
Ryan Houdek b190150281 FEXCore: Merges redundant string trimming implementations 2022-03-13 12:50:01 -07:00
Ryan Houdek 4cb6918506 Change page define usages over to self-defined
In the case of an AArch64 builder is using 16kb or 64kb pages like is
common on servers then it would fail to compile, even if the resulting
application would only ever run on 4k page hosts.

Resolve this by removing the build check and hardcoding 4kb pages for
each of our uses. We still require 4kb pages to run, so this mostly just
removes the weirdness where it is 16kb builder + 4k runner. Would have
broken some of our assumptions when running.
2022-03-06 07:33:10 -08:00
Ryan Houdek 5fbd01536f IR: Fixes some GPRPair IR op definitions
These were always wrong but how it the operations were handled meant
that it happened to work even though the IR representation was broken
2022-03-06 03:54:45 -08:00
Ryan Houdek 11a07eb3f2 FEXCore: Adds support for RDRAND/RDSEED
This matches the AArch64 implementation fairly well.
Bundles RDRAND and RDSEED together for simplification, both instructions
are a single flag on AArch64.
2022-03-06 03:40:25 -08:00
Ryan Houdek 7a9492ceca IRPasses: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek 73abf9b9dd IRParser: Resolves fallout from recent JSON changes 2022-03-04 04:07:25 -08:00
Ryan Houdek 0387c24e21 IREmitter: Resolves the fallout from recent JSON changes 2022-03-04 04:07:25 -08:00
Ryan Houdek 9d14bc7846 IR: Updates generators and json to new format
This greatly simplifies the IR format by using string parsing for
gathering the information.
Tons of redundant information removed.
Significantly more difficult to mess up adding a new IR op.

Significantly improves the generator functions in IREmitter
2022-03-04 03:51:44 -08:00
Ryan Houdek ac32ecadbe Merge pull request #1597 from Sonicadvance1/3dnow_and_back_again
OpcodeDispatcher: Implements all the 3DNow! instructions
2022-03-04 03:36:40 -08:00
Ryan Houdek 91f780223b Stop self-defining PAGE_SIZE
We only work on targets with 4096 byte page sizes.
Adds a cmake compile test to ensure this is adhered to.
2022-03-01 03:30:43 -08:00
Ryan Houdek a0efd2b01f Resolve comments. 2022-02-28 21:04:03 -08:00
Ryan Houdek 6501715a3c Allow classifying syscalls with flags
In some cases we can generate more optimal code if we have more
information about a syscall which number gets const-propagated.

In particular optimizing through syscalls, not synchronizing state, and
never returning.

- Noreturn is used by a syscall that never returns, like exit.

This means that it never needs to try and synchronize state coming back

- Not synchronizing state and optimizing through syscalls

Useful for syscalls that don't read the state past arguments and only
returns a value.
2022-02-28 21:03:54 -08:00
Ryan Houdek c7dd176799 IR: Adds a VRev64 op
This directly matches the AArch64 instruction and will be used shortly
2022-02-28 04:00:13 -08:00