Alyssa Rosenzweig
b0b4ad2083
OpcodeDispatcher: fuse xlat address
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig
ee4bee4fef
OpcodeDispatcher: fuse BT address
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig
c3a0f5a2f6
OpcodeDispatcher: fuse sgdt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig
0413a6bf68
OpcodeDispatcher: improve bmi2 shift
...
allow upper garbage, use simpler clean.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig
7bd036d1ae
OpcodeDispatcher: refactor address modes
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-06-01 09:42:32 -04:00
Alyssa Rosenzweig
665491adf8
OpcodeDispatcher: drop weird !flagm special case
...
now that bfi is coalesced, this is a win.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig
24cb02f4ff
FEXCore: remove IRCompaction
...
New RA does not need it for correctness, and the slight slow down to new RA from
not compacting first is much smaller than the cost of compaction. Overall speeds
up node.js start time by ~6% on top of new RA.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig
2bbcf72e27
OpcodeDispatcher: optimize asr masking
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-23 08:27:13 -04:00
Alyssa Rosenzweig
aa3a92aa60
OpcodeDispatcher: merge asr impls
...
so we can DRY the next patch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-23 08:27:13 -04:00
Alyssa Rosenzweig
53567a6526
OpcodeDispatcher: optimize movsxd
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-22 21:00:25 -04:00
Alyssa Rosenzweig
90fb5f038b
OpcodeDispatcher: optimize movsx
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-22 21:00:25 -04:00
Ryan Houdek
ca70e387ec
Merge pull request #3648 from alyssarosenzweig/ra/pair-extract
...
Slightly improve pair coalescing + memcpy fix from RA branch
2024-05-21 16:19:35 -07:00
Ryan Houdek
7b4e48480b
Merge pull request #3646 from alyssarosenzweig/opt/minor-disp
...
OpcodeDispatcher: eliminate some Bfe's
2024-05-21 15:57:52 -07:00
Alyssa Rosenzweig
a31c3c1c15
OpcodeDispatcher: use ExtractPair
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig
1a467f0ebd
IR: reduce memcpy worstcase reg pressure
...
This avoids a bunch of sharp edges for RA at a small cost when obscure
segment registers are used.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-21 16:54:06 -04:00
Alyssa Rosenzweig
50e56358c3
OpcodeDispatcher: eliminate Bfe's with cmpxchg
...
ConstProp was catching these but they're pointless.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-21 16:39:00 -04:00
Alyssa Rosenzweig
465dbc260f
OpcodeDispatcher: eliminate Bfe's with lea
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-21 16:39:00 -04:00
Alyssa Rosenzweig
769b2c2a46
OpcodeDispatcher: allow garbage on shift dests
...
doesn't matter for left shifts (we mask off the garbage), or 32-bit shifts, or
shifts where we explicitly sbfe after.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-20 10:33:15 -04:00
Alyssa Rosenzweig
3b2100307e
OpcodeDispatcher: allow garbage on more shifts
...
we're masking anyway
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-20 10:33:15 -04:00
Ryan Houdek
926eefc86c
Merge pull request #3635 from alyssarosenzweig/opt/flag-store
...
OpcodeDispatcher: reorder some moves
2024-05-16 10:58:40 -07:00
Alyssa Rosenzweig
7b39e57e72
OpcodeDispatcher: defer overwritten store
...
this can save moves, as it's a bit easier to reason about the live ranges.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-16 08:31:58 -04:00
Ryan Houdek
010028e381
FEXCore: Fixes the difference between CPL-0 and undefined instructions
...
undefined instructions are expected to return SIGILL, while implemented
instructions that aren't available in CPL-3 are expected to SIGSEGV.
Noticed this while testing out CPU-Z, it installs a kernel module and
does a bunch of `RDMSR` and `OUTS` instructions. Decided to walk through
the rest of the instructions in the `System Instruction Reference`
section.
Turns out there's a bunch of oddities in there that we don't support.
First step is to go through all the explicitl SIGILL and SIGSEGV and
implement a test for them.
Next step will be implementing the remaining operations that are
considered "System" operations but are still available in CPL-3.
This list includes:
- lar
- lgdt
- lsl
- sidt
- sldt
- stac
- clac
- verr
- verw
2024-05-13 11:12:26 -07:00
Alyssa Rosenzweig
a2fc51fc7b
IR: specify registers, not offsets for SRA
...
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig
b91b0e9d65
IR: infer SRA static class
...
no need to stick it in the IR.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig
74489a4177
IR: remove dead SRA flags
...
I don't know what these were meant for, and I don't care (-:
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-05-08 14:01:42 -04:00
Ryan Houdek
c8704a7f71
OpcodeDispatcher: Implement support for SMSW
...
Found out that Far Cry uses this instruction and it is viable to use in
CPL-3. This only returns constant data but its behaviour is a little
quirky.
This instruction has a weird behaviour that the 32-bit operation does an
insert in to the 64-bit destination, which might be an Intel versus AMD
behaviour. I don't have an Intel machine available to test if that
theory is true although. This assumption would match similar behaviour
where segment registers are inserted instead of zext.
Gets the game farther but then it crashes in a `___ascii_strnicmp`
function where the arguments end up being `___ascii_strnicmp(nullptr, "Color", 5);`.
2024-04-18 07:41:39 -07:00
Paulo Matos
2b4ec88dae
Whole-tree reformat
...
This follows discussions from #3413 .
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Ryan Houdek
271700e9f6
Merge pull request #3568 from lioncash/const
...
X87: Simplify constant loading for FLD family
2024-04-11 00:32:26 -07:00
Ryan Houdek
1ba678f631
Merge pull request #3562 from Sonicadvance1/fix_rsp_store_tso
...
OpcodeDispatcher: Fixes disabling TSO access on RSP SIB stores
2024-04-11 00:32:14 -07:00
Lioncache
4cb2432b5c
OpcodeDispatcher: Make use of new x87 constants
...
Now we can load these directly instead of needing to manually materialize them.
2024-04-09 10:17:15 -04:00
Lioncache
b0aeb501f4
OpcodeDispatcher: Add helper for making segment offset addresses
...
There's quite a few places where the segment offset appending is open-coded
throughout the opcode dispatcher, but we can pull these out into a few
helpers to make the sites a little more compact and declarative.
2024-04-08 17:50:58 -04:00
Ryan Houdek
0e93fd0f3e
OpcodeDispatcher: Fixes disabling TSO access on RSP SIB stores
...
GPR Direct/Indirect already had this and SIB version also already
supported on the load side. Fixes this missed behaviour.
2024-04-08 11:32:01 -07:00
Alyssa Rosenzweig
098859caf7
OpcodeDispatcher: use _ShiftFlags for ASHR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig
c632543451
OpcodeDispatcher: use _ShiftFlags for SHRD
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig
801cf72f95
OpcodeDispatcher: use _ShiftFlags for SHLD
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig
650cd2c46e
OpcodeDispatcher: use _ShiftFlags for SHR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig
5b48ce2228
OpcodeDispatcher: use _ShiftFlags for SHL
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig
a05cc06ab4
OpcodeDispatcher: unify imm/1-bit ASHR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
031e756a78
OpcodeDispatcher: unify imm/1-bit SHR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
2a9f1ce8cb
OpcodeDispatcher: unify imm/1-bit SHL
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
8c53a9f051
OpcodeDispatcher: use LoadConstantShift for rotates
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
cf26ec7898
OpcodeDispatcher: use LoadConstantShift for SHRD
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
582c3dae6e
OpcodeDispatcher: use LoadConstantShift for SHLD
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
2abac03ab0
OpcodeDispatcher: add LoadConstantShift helper
...
shows up a bunch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
8cc684fa12
OpcodeDispatcher: drop misinformed comment
...
tbnz only tests a single bit, not a mask.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
d92de1d947
OpcodeDispatcher: drop result masking for shifts
...
flag calcs are fine with upper garbage.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig
bd0b5eceb8
Merge pull request #3545 from alyssarosenzweig/opt/pf-scalar
...
Use scalar integer code to calculate PF
2024-04-02 11:28:53 -04:00
Ryan Houdek
e8abc88702
Merge pull request #3542 from alyssarosenzweig/ra/rep
...
Eliminate xblock liveness with rep cmp/lod/scas
2024-04-02 04:24:24 -07:00
Ryan Houdek
29c6281e11
Merge pull request #3539 from alyssarosenzweig/ra/rol-ror2
...
rewrite ROL/ROR
2024-04-02 00:17:08 -07:00
Ryan Houdek
cd9ffd2045
Merge pull request #3536 from alyssarosenzweig/ra/rcl-rcr
...
OpcodeDispatcher: eliminate xblock liveness for rcl/rcr
2024-04-01 11:44:37 -07:00