This is a new procfs symlink path that changes behaviour of binfmt_misc
when exposed. We need to check both procfs/exe and procfs/interpreter
and see if they exist AND also differ.
Once/if they do then we can disable a bunch of checking of paths once
they do. The fallback when none of this is supported has the same
behaviour has previously where it still does all the regular checking.
During binfmt_misc install cmake will check the kernel version for the
raw binfmt_misc writing. Which will never pass until we have a real
kernel version that it is upstreamed in.
For update-binfmts we add a new optional argument where the tool will
drop the flag if the host kernel version isn't new enough to handle the
option.
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.
My concrete IR proposals are that:
* SSA values must be killed in the same block that they are defined.
* Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
* LoadGPR/StoreGPR are eliminated in favour of SSA within a block.
This has a lot of nice properties for our setting:
* Except for some internal REP instruction emulation (etc), we already have
registers for everything that escapes block boundaries, so this form is very
easy to go into -- straightforward local value numbering, not a full into
SSA pass.
* Spilling is entirely local (if it happens at all), since everything is in
registers at block boundaries. This is excellent, because Belady's algorithm
lets us spill nearly optimally in linear-time for individual blocks. (And
the global version of Belady's algorithm is massively more complicated...)
A nice fit for a JIT.
Relatedly, it turns out allowing spilling is probably a decent decision,
since the same spiller code can be used to rematerialize constants in a
straightforward way. This is an issue with the current RA.
* Register assignment is entirely local. For the same reason, we can assign
registers "optimally" in linear time & memory (e.g. with linear scan). And
the impl is massively simpler than a full blown SSA-based tree scan RA. For
example, we don't have to worry about parallel copies or coalescing phis or
anything. Massively nicer algorithm to deal with.
* SSA value names can be block local which makes the validation implicit :~)
It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.
Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now the calculation of PF is entirely deferred, by inverting our internal
representation of PF. All the (e.g.) logical op needs to do is store the low
8-bits of the result.
This is a bit of a mixed bag. Primary ALU ops all save an instruction, by
skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky
cases are:
* Zeroing PF. This now requires writing 1 instead of 0, which may require an
extra move for the constant. Some of this will go away when we merge PF+AF
into a single register, which is next up on the list. In that case, the
two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing
x86 view of PF), that will get absorbed into the or, if we also write AF. If
we leave AF undefined and need to write a 1, that's a single mov instruction
and we couldn't do better anyway if not inverted (since we'd still have a mov
wzr even then). So in view of the future work, this isn't something I'm
concerned about.
* Float comparisons that put Unordered into PF. These require an extra invert to
match the new convention. These are already so unnecessarily bloated that I'm
not convinced I'm making things materially worse here. But we realistically
need multiple destination support in the IR to fix this particular mess.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
AF is calculated as:
((Src1 ^ Src2) ^ Res)[4]
Due to the extract, this is equivalent to
((Src1 ^ Src2) ^ (Res ^ 1))[4]
We already store (Res ^ 1) as the PF byte. So, it suffices to store
AF Byte = Src1 ^ Src2
and then we can recover the flag value
AF = (AF Byte ^ PF Byte)[4]
This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:
* Both PF and AF written together, the coupling is correct.
* PF written but AF invalidated, irrelevant.
* Both invalidated, irrelevant.
None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Logical ops leave AF undefined so we can't expect it to be zero after. Mask the
result of lahf to avoid testing UB. These unit tests would regress from the work
in this MR otherwise.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Use an explicit invalidate, so we can zero easily enough if we need to for
debugging later but we can save the instrs ordinarily.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
FEX has a problem with large blocks that uses a ton of constants spread
throughout the block. Once a block gets large enough with enough
constants that have large live ranges, FEX slows down to unusable speeds
due to the register allocator spending more time calculating node
interferences than anything else in the program.
This adds a little heuristic to ensure that constants aren't reused if
the previous value is past a certain distance threshold. This threshold
works well enough that XeSS's pedantic initialization code doesn't have
issues now. See https://github.com/FEX-Emu/FEX/issues/2688 for more
information about that.
FEX itself should work to remove bad constant usages to make this pass
less necessary anyway. In most cases we are materializing duplicated 0,
1, and masks which could be done without a constant entirely.
Maybe once we've improve that enough we could remove this constant
pooling entirely.
To note, this doesn't fix the issue that XeSS causes our register
allocator, this is purely a heuristic workaround.
This class is very expensive to initialize so if you happen to have the
disassembler configuration enabled you were eating a very bad
initialization cost for no reason.
Only initialize the data member if any disassembler runtime option is
enabled, this completely removes the overhead.
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.
This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.
Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.
Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
Currently in main today, FEX fails to compact OF/CF/ZF/SF and PF.
This is due to recent optimizations with flag calculations on each of
these. Now that we have a centralized location where we compact and set
our internal representation of flags we can do this in one location.
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
Currently FEX's internal EFLAGS representation is a perfect 1:1 mapping
between bit offset and byte offset. This is going to change with #3038.
There should be no reason that the frontend needs to understand how to
reconstruct the compacted flags from the internal representation.
Adds context helpers and moves all the logic to FEXCore. The locations
that previously needed to handle this have been converted over to use
this.
32-bit or 64-bit addition without carry-in. This matches the baseline hardware
semantic. Generalizing to support other cases can come later, this should be a
win already.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For correct carry/overflow behaviour, we need to use a compare of the right size. The existing logic
to look at the source sizes doesn't work for this, since a 32-bit NEG instruction will compare a
32-bit source with a 64-bit _Constant(0) .. which needs a 32-bit compare but the existing logic
would use a 64-bit compare. This is not yet a bug fix, since the overflow code is currently in
software for 32-bit negates so it's irrelevant. But it should prevent regressions from using native
compares later in this series. Presumably this was intended all along but left as-is to avoid
disturbing instcountci once noticed. Time to disturb CI!
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
AddNZCV is a new op to return the NZCV for an addition directly, which lets us
skip software flag calculation in some cases. In the future it would be nice to
fuse this into the Add itself as a second destination to avoid repeating the
addition, but that's a very involved change and right now I'm building FEX on an
old Chromebook because my M1 kernel is FUBAR.
Similarly, SubNZCV returns flags for Sub. This has the extra twist of needing to
invert the carry bit due to the inverted definition between arm64 and x86_64.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This will let us reuse it in some cases. In some of these
implementations there is a bad code smell around using the zero register
but that isn't going to get solved in this commit.
A bunch of the AES operations take a zero register upfront and we
currently materialize it for each instruction.
Considering that most AES operations are used back to back, we can
eliminate these materializations by caching it between instructions.
Additionally removes a move in the optimal case when destination matches
the state register, which is exactly what the SSE operation ends up
doing.
AESKeyGenAssist has an edge case that if the destination RA overlaps the
zero register then we still need to eat a move, hopefully doesn't happen
too frequently in practice. This is also the lesser used instruction so
it isn't a big deal. RA constraints could solve that still.
This instruction has xmm0 be one of the implicit sources. We were
loading xmm0 twice. #2700 would also fix this but that breaks other
things for some reason.
We can load the swizzle table from our constant pool now. This removes
the only usage of VTMP3 from our Arm64 JIT.
I would say the this is now optimal for the version without RCON set.
With RCON we could technically make some of the move of the constant
more optimal.
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.
This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
These functions are called a lot....a lot a lot.
Optimize these in to a couple of ALU operations instead of a whole
table lookup. Confirming with output assembly that this becomes more
optimal.
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.
I would say this is now optimal.
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.
Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.
When consuming our own controlled data, we don't want the range clamping
to be enabled.
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.
This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.
Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
Noticed that we hadn't ever enabled this, which was a concern when our
GPR operations weren't as strict about leaving garbage in the upper bits
when operating as a 32-bit operation.
Now that our ALU operations are more strict about enforcing upper bit
zeroing we can enable this.
This causes Half-Life: Source FPS to get to > 200FPS finally. Causes
significant performance improvements for 32-bit games because we're no
longer redundantly moving registers before and after every operation.
Causing a bunch of 3-4 instruction sequences to convert to 1.
RAValidation was making an assumption that GPR register class would only
have up to 16 registers for either SRA or dynamic registers.
When running a 32-bit application we allow 17 GPRs to be dynamically
allocated, since we can take 8 back from SRA in that case.
Just split the two classes in the RAValidation pass since they will
never overlap their allocation.
Fixes validation in `32Bit_Secondary/15_XX_0.asm` locally that changed
behaviour due to tinkering.
When the application calls exit_group we no longer need to care about
cleanup because the entire process group is leaving.
Just immediately call exit group and get out. Might revisit this in the
future.
Fixes#2752
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.
This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.
In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.
[1] e89321dc60 for scalar.
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
FPR and converted it in-place.
These are now optimal in the case of AFP is unsupported.
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.
Needs #2994 merged first.
Huge thanks to @dougallj for the optimization idea!
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.
Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.
Needs #2993 merged first.
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.
Thanks to @rygorous for the idea!
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
Given the operation is:
Dst = Vector1 / Vector2
If Dst and Vector1 alias one another, then we can just perform
the division as is without any moving of data around.
Zero-extension will already occur if necessary upon storing.
Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
This now uses the new load element IR operation and makes these
instructions optimal.
LRPCPC3 will introduce instructions in the future for TSO emulation to
help these operations, but that doesn't exist today.
This can be done without storing any data to memory and also
compressing the amount of instructions being used.
Thanks to @dougallj for the optimization suggestions.
VRev32 matches Arm64 semantics directly.
LoadNamedVectorConstant allows FEX to quickly load "named constants".
This will allow us to have specific hardcoded vector constant values
that we can load with a ldr(State)+ldr(Value) and will be more abused in
the future.
This also allows us to do a very simple optimization in the future where
we can optimize away redundant loads of these loads if they are used
multiple times in the same block. (Not implemented here).
Two optimizations here:
1) The final VInsElement was generating three instructions
- This itself could have been change to vzip, which would have
removed two instructions.
2) Optimize how the pairwise elements are calculated to shave one
instruction off the calculation.
- addp odd elements and even elemnts first
- Then transpose those elements
- Then use one final addp to generate the result in the correct
order.
The ext+uabdl+addp pairs of operations could be reordered to shave off
one temporary register usage if we really care later.
Can be slightly more optimal with a slightly change algorithm but will
require implementing some IR ops which can be put off. It's only about
an instruction savings.
Even on platforms without SVE these have improved slightly but the real
improvement comes from SVE.
Adds some new InstCountCI files for SVE128 enabled testing.
Also enables SVE128 in the VEX maps. Host features should probably
enable SVE128 when SVE256/AVX is enabled, but that isn't the case today.
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.
With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).
This is a significant improvement even on platforms that only support
128-bit SVE.
With ASIMD FEX would never optimize BSL out of fear if some registers
overlapped it would break things. So it had previously always moved to a
temporary first and then moved the result back out when done.
Now instead check upfront if any of the source registers overlap the
destination. If the destination register overlaps any of the three
sources we can bsl, bit, or bif depending on which register gets
overlapped.
Worst case the destination doesn't overlap any of the source registers
and still needs these moves.
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.
Back to back bfi is actually less optimal due to dependency tracking.
With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
This paves the way to optimizing pushes in to both push operations and
push pair operations to more optimally match Arm64 push support.
While this does the first step for supporting the base push, we'll leave
optimizing push pairs to future work.
This is a bit of tricky operation where due to our our usage of SSA, the
incoming source isn't guaranteed to end its live-range at this
instruction.
This gives us a behaviour where to be optimal we need to take different
paths depending on if the incoming address register is the same as the
destination node.
Once we have form of RA constraints or non-SSA IR form that can
guarantee this restriction then this will go away.
FEX doesn't use the platform register on wine platforms so there is no
reason to save and restore it.
On Linux we can still use it at some point but for now it isn't part of
our RA.
Changes the idiom used for constant mask generation to a ternary.
This pattern is definitely used elsewhere in code but we can get rid of
all instances here.
This takes a similar approach to deferred signal handling and allows any given
thread to be interrupted while running JIT code by protecting the appropriate
page as RO. When the thread then enters a new block, it will try to acccess
that page and segfault. This is safer than just sending a signal to the thread
as that could stop in a place where JIT context couldn't be recovered correctly.
With WOW, all allocations from 64-bit code use the full address space
and limiting is handled on the syscall thunk side so theres need to
worry about STL allocations stealing AS.
This relies on wine's behaviour passing through linux paths and env vars,
so that the config in the user's home directory can be accessed outside
of the wine prefix.
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.
Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.
It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.
Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.
A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.
[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
- 3.6 - Restricted Subset of Segmentation
- `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
registers; the base for CS, DS, ES, and SS is ignored for 32-bit
mode, same as 64-bit mode (treated as zero).`
- 4.2.17 - MOV to Segment Register
- Will fault if SS is written (Breaking anything that writes to
SS).
- Will not fault if CS, DS, ES are written (Thus it sets the
segment but gets ignored due to 3.6).
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now all vbroadcast implementations go down the more optimal path.
For non-SVE 128-bit cases where we only have 128-bit wide registers,
we behave like ld1rqb and just act as a normal 128-bit load for
interface convenience.
In the case of running on a 128-bit SVE system this predicate wasn't
setup. Since we never had any predicate usage before this wasn't an
issue. Now that #2914 is using the 128-bit predicate we need to make
sure that we are generating it.
Allows the implementations of the vbroadcast instructions to perform the
load and broadcast in one operation as opposed to doing the load and then
broadcast separately.
Notably, the broadcasting loads can also be used on systems that have SVE 128-bit
support as well, not only 256-bit.
On non-SVE systems, we use the equivalent AdvSIMD instructions.
Instead of allocating a temporary copy of the string, return a view of
it instead. Should improve the performance of system calls that take
file paths. Since it was allocating a string for every single syscall
that uses them in this case.
For config values that were string objects we were unnecessary creating
copies each time the string was accessed.
Convert the () operator over to returning a reference.
The current implementation uses orr excessively. This has FEX missing
hardware optimization opportunities where some CPU cores will zero-cycle
move constants that fit in to the 16-bits of movz/movk.
First evaluate up front if the number of 16-bit segments is > 1, in
those cases we should check if it is a bitfield that can be moved in one
instruction with orr.
After that point we will use movz for 16-bit constant moves.
Additionally this optimizes the case where a constant of zero is loaded
to be a `mov <reg>, zr` which gets renamed in most hardware.
Commonly we are doing a BFI into a 32-bit register, which is hitting the
ubfx (lsr alias) path.
In the case of 32-bit destination we can also do a regular move, which
will take advantage of CPU's rename functionality and give a minor speed
boost.
Didn't notice this in the previous PR, When DUMPIR=stderr without and
selection of where to place it in PASSMANAGERDUMPIR it was supposed to
put the dumper at the end of the passes.
We need to make sure that it it placed at the end of the passes rather
than current `it`.
This will allow investigating the Arm64 directly next to the test, plus
publicly linking directly to badly behaving tests.
Perfect for nerdsniping implementations.
We can perform the SQRT first and then broadcast 1.0 into the destination
since all the intermediary work is done, meaning we don't have to worry
about Dst and Vector aliasing one another.
Pretty sure this is why CI is unhappy. If a test in a different file has
the same name then it is highly likely to conflict when nasm is
generating files and will overwrite and erase, causing CI to break.
Include the incoming json filename as part of the asm keys so it can't
conflict here.
If DumpIR is enabled but the PassManagerDumpIR option isn't enabled then
this currently does nothing.
As a convenience, enable dumping the final optimized IR if an option
hasn't been specified.
None of the instructions here are optimal, a couple are close.
Will be splitting each of the map tables in to their own json files
since each one can get fairly large.
We need to move the modifier enum out of the SVEMemOperand class
since it's also used with adr. Plus, this can also be convenient
not being tied down to the class itself.
This also makes accessing modifiers less noisy, since the class
When a 32-bit adcx instruction was encountered, it was getting treated
as a 16-bit adcx instruction instead. This is because of the 0x66 prefix
required to handle this instruction.
Adds a unit test to ensure it doesn't break again.
Implements CI for tracking instruction counts for generate blocks of
code when transforming from x86 to ARM64 assembly.
This will end up encompassing every instruction in our instruction
tables similarly to how our assembly tests try to test everything in our
instruction tables.
Incidentally, the data for this CI is generated using our assembly
tests. By enabling disassembly and instruction stats when executing a
suite of instructions, this gives the stats that can be added to a json
file.
The current implementation only implements the SecondGroup table of
instructions because it is a relatively small table and has known
inefficiencies in the instruction implementations. As this gets merged I
will be adding more tables of instructions to additional json files for
testing.
These JSON files will support adjusting CPU features regardless of the
host features so it can test implementations depending on different CPU
features. This will let us test things like one instruction having
different "optimal" implementations depending on if it supports SVE128,
SVE256, SVEI8MM, etc.
This initial instruction auditing is what found the bug in our vector
shift instructions by size of zero. If inspecting the result of the CI
run, you can tell that these instructions still aren't "optimal" because
they are doing loads and stores that can be eliminated.
The "Optimal" in the JSON is purely for human readable and grepping
ability to see what is optimal versus not. Same with the "Comment"
section.
According to my auditing spreadsheet, the total number of instructions
that will end up in these json files will be about 1000, but we will
likely end up with more since there will be edge cases that can be more
optimal depending on arguments.
This was confusingly split between Arm64Emitter, Arm64Dispatcher, and
Arm64JIT.
- Arm64JIT objects were unnecessary and free to be deleted.
- Arm64Dispatcher simulator and decoder moved to Arm64Emitter
- Arm64Emitter disassembler and decoder renamed
- Dropped usage of the PrintDisassembler since it is hardcoded to go
through a FILE* type
- We instead want its output to go through LogMan, which means using a
split Decoder+Disassembler object pair.
- Can't reuse the object from the vixl simulator since the simulator
registers the decoder as a visitor, causing the simulator to execute
while disassembling instructions if reused.
- Disassembly output for blocks and dispatcher now output through Logman
- Blocks wrapped in Begin/End text for tracking purposes for CI.
We don't currently have a device in CI that can run SVE with 128-bit
width registers. Until we have a device with this, make sure the vixl
simulator is also running the ASM tests in this width.
This was causing test failure locally where some values were set to
uninitialized data. Ensure that gregs, YMM, and MMX registers are all
zero initialized.
Requires the IR headerop to house the number of host instructions this
code is translating for the stats.
Fixes compiling with disassembly enabled, will be used with the
instruction count CI.
This is incredibly useful and I find myself hacking this feature in
every time I am optimizing IR. Adds a new configuration option which
allows dumping IR at various times.
Before any optimization passes has happened
After all optimizations passes have happened
Before and After each IRPass to see what is breaking something.
Needs #2864 merged first
This is a /very/ simple optimization purely because of a choice that ARM
made with SVE in latest Cortex.
Cortex-A715:
- sxtl/sxtl2/uxtl/uxtl2 can execute 1 instruction per cycle.
- sunpklo/sunpkhi/uunpklo/uunpkhi can execute 2 instructions per cycle.
Cortex-X3:
- sxtl/sxtl2/uxtl/uxtl2 can execute 2 instruction per cycle.
- sunpklo/sunpkhi/uunpklo/uunpkhi can execute 4 instructions per cycle.
This is fairly quirky since this optimization only works on SVE systems
with 128-bit Vector length. Which since it is all of the current
consumer platforms, it will work.
We need to know the difference between the host supporting SVE with
128-bit registers versus 256-bit registers. Ensure we know the
difference.
No functional change here.
This allows use to both enable and disable regardless of what the host
supports. This replaces the old `EnableAVX` option.
Unlike the old EnableAVX option which was a binary option which could
only disable, each of these options are technically trinary states.
Not setting an option gives you the default detection, while explicitly
enabling or disabling will toggle the option regardless of what the host
supports.
This will be used by the instruction count CI in the future.
Moves the dummy handlers over to this library. This will end up getting
used for more than the mingw test harness runner once the instruction
count CI is operational.
This was a debug LoadConstant that would load the entry in to a temprary
register to make it easier to see what RIP a block was in.
This was implemented when FEX stopped storing the RIP in the CPU state
for every block. This is now no longer necessary since FEX stores the
in the tail data of the block.
This was affecting instructioncountci when in a debug build.
I use this locally when looking for optimization opportunities in the
JIT.
The instruction count CI in the future will use this as well.
Just get it upstreamed right away.
`eor <reg>, <reg>, <reg>` is not the optimal way to zero a vector
register on ARM CPUs. Instead we should move by constant or zero
register to take advantage of zero-latency moves.
While the ENABLE_LLD and ENABLE_MOLD options are nice, they don't handle
the case when the linker of `lld` or `mold` doesn't match the compiler.
This particularly crops up when overriding the C compiler to a new
version of clang but the globally installed `ld.lld` is still the old
clang version.
This then causes clang to fail with unusual errors when upstream breaks
compatibility with itself.
Easy enough to use by passing the linker to cmake:
`-DUSE_LINKER=/usr/bin/ld.lld-15`
This also removes the ENABLE_LLD and ENABLE_MOLD options to use
USE_LINKER directly.
- ldd: `-DUSE_LINKER=lld`
- mold: `-DUSE_LINKER=mold`
Example of compiler failure when built with clang-15 but attempting to
link with ld.lld 14:
```bash
ld.lld-14: error: unittests/APITests/CMakeFiles/Filesystem.dir/Filesystem.cpp.o: Opaque pointers are only supported in -opaque-pointers mode (Producer: 'LLVM15.0.7' Reader: 'LLVM 14.0.6')
```
This needs to default to 64-bit addresses, this was previously
defaulting to 32-bit which was meaning the destination address was
getting truncated. In a 32-bit process the address is still 32-bit.
I'm actually surprised this hasn't caused spurious SIGSEGV before this
point.
Adds a 32-bit test to ensure that side is tested as well.
This is more obvious. llvm-mca says TST is half the cycle count of CMN
for whatever it's defaulting to. dougallj's reference shows both as the
same performance.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
In the non-immediate cases, we can amortize some work between the two
flags to come out 1 instruction ahead.
In the immediate case, costs us an extra 2 instructions compared to
before we packed NZCV flags, but this mitigates a bigger instr count
regression that this PR would otherwise have. Coming out ahead will
require FlagM and smarter RA, but is doable.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Same technique as the left shifts. Gets rid of all our COND_FLAG_SET
use, which is good because it's a performance footgun.
Overall saves 17 instructions (!!!!) from the flag calculation code for
`sar eax, cl`.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Similar to the immediate case, but now we select between the entire old
and new NZCV registers. This is faster than selecting each bit
independently. Saves 11 instructions for calculating flags for "shl eax,
cl".
It is undefined in this case. We prefer to zero (rather than preserve
the existing value) as it avoids a costly RMW of the NZCV register.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We need to be careful to preserve V if needed. For `shl 1` and `shr 1`,
saves 2 instruction overall compared to before the PR. For `sar 1`,
saves 3 overall.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We can just return zero, no need to do a pointless Bfe. Saves yet
another instruction for GetPackedRLAG in a test I'm looking at.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We know that bit 0 is CF, so we can do CF first and then avoid setting
Original to 2 (for reserved) with a silly `or xzr, #2` instruction.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Faster sign/negate testing for 32-bit/64-bit inputs. This could maybe be
extended to 8/16-bit if we have FlagM but that's for later.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If we can prove that a flag bit could not possibly be set, we can use
orlshl rather than bfi, which can be more efficient.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
In some cases we just want to insert in one bit at a time, add a helper
to zero the 4 flags together so we can avoid the extra RMW cycle at the
beginning.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We can set N more efficiently with some bit math, and zero ZCV at the
same time. In the future we'll be able to use TST for this to make it
even faster.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Later, this will let us take advantage of the arm64 flags. For
now, this just turns some strb's into bfi's for dubious benefit.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If we read or write NZCV flags we need to call CalculateDeferredFlags on
block boundaries, if only to flush out the cached copy.
Also, when leaving a block we call it to flush out. This is annoyingly
invasive but I don't know of a better way to do this that doesn't
involve rearchitecting the dispatcher.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We'll add an extra caching layer in a moment so can't call _LoadFlag
directly and expect correct results.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Maps to arm64 tst, except properly SSA. This will need some RA support
to avoid redundant mrs/msr sequences.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
On arm64, orr (with a shifted register) is maybe fewer cycles and
definitely easier on the RA than bfi (=> fewer moves generated). So,
it's preferred when we know the corresponding bit is 0 in the
destination.
It's not useful on other targets, so it's gated behind a backend feature
bit.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Reserve 4 bytes of "flags" to model the 32-bit arm64 NZCV register, so
we can start porting FEX's flag handling code over to using NZCV without
needing the whole compiler to be aware of instructions that might
clobber host flags.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This IR operation was limited to GPR only previously for the values
getting compared.
This adds support for GPRPair (and technically FPR) so that it can be
used directly with GPR pairs. I say technically FPR because the
IREmitter disallowed FPRs for the comparison, but this was already
supported in all of the backends, we just didn't ever use it.
Some minor changes to the constant prop pass to ensure that we don't try
to propagate a select in to a CondJump, otherwise pair comparisons would
be duplicated. This code is expecting to be able to merge a simple
comparison in to a `cbnz`, which doesn't happen with GPR pairs.
When cmake's `install` function is invoked with a relative path, then it
is interpreted as being relative to the `CMAKE_INSTALL_PREFIX` variable.
This variable follows both `DESTDIR` and `CMAKE_INSTALL_PREFIX` so it is
best to use relative addresses in the install path.
Thanks for the report Mike!
Fixes#2849
Currently only does a build, doing a CI run means figuring out why
TestHarnessRunner doesn't find libraries correctly.
I want to ensure we don't break building at least while sorting out the
rest of this.
Inspired by #2832 by going through the source and removing uses of
`assert`.
assert doesn't work for us in debug builds because some games will
capture SIGABRT and continue running. So we need to use FEX's built in
assert handlers which call `FEX_TRAP_EXECUTION` which will take down the
FEX process as expected.
There wasn't too much usage of this in the source, so this is relatively
straightforward.
We can move asserts and the base opcode into the implementing function.
While we're at it, we can add an assert to ensure predicate registers are in range.
We can centralize the base opcode and some of the asserts in the implementing function.
We can also add an assert that validates the predicate register range
and also make the offset assertions much more informative.
Since everything we need is already in the ARMEmitter namespace, we
don't need to qualify all type usages, which reduces the verbosity
a little during reading.
This is actually already implemented but was grouped with the
SVE floating-point convert to integer group. We can extract this
out to be organized a little nicer.
We can move the op into the implementing function and get rid of
an unnecessary function.
We can also add an assert to ensure the predicates are valid as well.
This test was written to test SMC where one thread is doing execution
while the other thread is modifying.
According to the printf documentation it is supposed to "wait for code
to be modified" but actually it was testing a race between a printf on
one thread and the primary thread modifying the code.
Fix this test so it is actually waiting for code modification to happen
rather than testing a race condition. This is likely what the original
author intended.
CI is hitting this flake more frequently now because it is even faster
it seems, so fixing this test is necessary to resolve these flakes.
insert() will only ever perform an insertion if the relevant element
doesn't exist within the set, so we were doing an unnecessary lookup
in two spots.
We don't use emplace() here because it will need to construct the
key for the element inside the allocated node (since the key may
be non-copyable/non-movable), causing an allocation even if an insert
doesn't actually occur.
Conversely, insert() will not need to allocate and construct a node
ahead of time if an element already exists in the set.
Every time I look at this I can never remember which node is the value
node and which one is the store node.
Rename the opaque `Node` to `ValueNode` and have documentation comments
to explain which one it is. This way I won't forget again.
Tests for the regression from 7e6bb04db ("OpcodeDispatcher: Extract
CalculatePF"). This fails on main but passes with this PR.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Inspired from: https://github.com/dotnet/runtime/issues/8072
Currently FEX is /very/ heavy handed with our backpatching where we wrap
every backpatched loadstore with `dmb ish`.
This can be relaxed slightly according to the linked issue.
For TSO load instructions the instruction sequence changes to:
ldr <args>;
dmb ld; <-- Slightly less strict dmb
For TSO store instructions the instruction sequence changes to:
dmb ish; <-- Still the all encompassing dmb
str <args>;
For backpatching loadstores this does the same thing where only one side
needs the nop and it uses the same instruction sequence when
backpatched.
The minor change is that on load backpatching, we are no longer backing
up a single instruction, instead just re-executing the instruction we
patched directly.
Took a long time to come back to this (Last looked in August 2020).
Previously when I was implementing this idea it didn't work, but that
was because our CompareExchange operation was broken back then. With the
CAS now, it should just work.
We expect that PF is written more often than it's read, so we want to
get the expensive popcount out of the hot path. (Thank you to Dougall
for suggesting that.)
There are two cases:
1. PF is written by an integer instruction. In this case, we calculate
with the formula `popcount(x ^ 1) & 1`.
2. PF is written by a float instruction, copying a host flag.
What we really want is to defer the relatively expensive popcount. So,
to unify these cases, we have integer instructions write `x ^ 1` and
(unchanged) float instructions write the host flag. Then, when reading
PF, we do `popcount(value) & 1` on the byte read in.
If PF is written but not read, this saves the expensive popcount and
leaves only the cheap xor.
If PF is written by an integer op and read, this maybe shuffles some
code but does not materially change anything.
If PF is written by a float op and read, this is worse because now we're
doing an extra pointless popcount. This is a tradeoff... However, this
is only relevant to unordered float comparisons, which I expect to be
obscure for games. So this should be worth it over all (for games, if
not weird numerical computing workloads).
How does this connect to my register zeroing quest? The constant folding
code doesn't currently deal with FPRs and I'm not in a mood to change
this. So before, a block ending with `xor eax, eax` would still do a
popcount for PF. Now it just writes a constant 1 since the xor constant
folds and the popcount never happens at all.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For flag calculation after moving a constant. This cleans up the code
generated for zeroing at the end of a block.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Lets us deduplicate the asserts and put them in one spot. We can also
improve the assert message to also indicate the valid range.
While we're in the area, we can collapse a few case paths
as a result of this assert movement.
Usually uploading of results takes about two seconds.
Sometimes github's connection to the runner flakes and it stalls out the
upload action for some reason.
Github's default timeout is SIX HOURS.
Change this to a one-minute timeout on the upload step so it quickly
goes away when Github's internet flakes out.
For #2804 so it can compare if GPRs match more easily.
Only adding to GPR and Literal types, since its ambiguous what this
would mean for the memory accessing types. Going to leave those other
ones alone for now.
Two bugs here that caused thunking X11 thunking in Wine/Proton to not
work.
The easier of the two. The various variadic functions that we thunk
actually take key:value pairs where the first is a string pointer, and
the value can be various things.
We need to handle these as true key:value pairs rather than finding the
first nullptr and dropping the remainder.
Additionally, there are 12 keys that specify a callback that FEX needs
to catch and convert to host callable. Wine is the first application
that I have seen that actually uses this. If these callbacks aren't
wired up then it it can miss events.
The harder of the two problems is the `libX11_Variadic_u64` function was
subtly incorrect. Nothing had previously truly exercised this and my
test program didn't notice anything wrong while writing it.
The first incorrect thing was that it was subtracting the nullptr ender
variable before the stack size calculation, causing the value to
overwrite the stack if the number of remaining elements was event.
Secondly the assembly that was storing two elements per step was
decrementing the counter by 8 instead of two. Didn't pick this up before
since I believe the code was only hitting the non-pair path before.
This gets Proton thunking working under FEX now.
I was hitting an issue where thunking in Wine+Vulkan applications was breaking
unless both OpenGL and Vulkan was enabled.
Turns out this was because Vulkan enabled only XCB, which didn't enable
X11. So when XCB is thunked but X11 isn't, this causes weird issues
where X11 calls in to XCB functions and gets in desync'ed state. Causing
hangs to appear in xcb_take_socket.
Now we can enabled just `Vulkan` as a thunk and it'll work fine.
```
(gdb) bt
from target:/usr/lib/fex-emu/HostThunks//libvulkan-host.so
```
Now that we have the helper for encoding immediate shifts,
we can trivially implement the remaining missing instruction
in the bitwise logical unpredicated group.
Prevents the code invalidation mutex from being locked as shared recursively,
since it is locked before entering ThreadExitFunctionLink and would end up
being locked again by ThreadAddBlockLink.
This fixes a deadlock on Windows.
This lets us deduplicate the behavior rather than open-coding it everywhere
and also makes it nicer to implement instructions that make use of this
encoding pattern.
Fixes#2754
These panicking fallbacks are at times not ending up in as plt calls
for some reason that I haven't been able to reproduce locally.
So far the only way I can reproduce is building with Canonical's PPA
build system, since rebuilding locally didn't resolve the issue.
This will change the failure mode from these panicking asserts happening
at call time, to dlopen failing during relocation when loading the
thunk. Which LD_DEBUG=all can be used for debugging relocation failure
in that case.
Allows us to consume an array of strings and convert it to an mask of
enum values. This is a quality of life change that allows us to specify
a mask of options.
The first configuration option added to support this is to control the
vixl disassembler. Now by default the vixl disassembler doesn't
disassemble any blocks and needs to be enabled individually.
eg:
```
FEXLoader --disassemble=blocks <args>
FEXLoader --disassemble=dispatcher <args>
FEXLoader --disassemble=blocks,dispatcher <args>
```
Has the additional convenience option of just passing in numbers as
well.
```
FEXLoader --disassemble=2 <args>
FEXLoader --disassemble=1 <args>
FEXLoader --disassemble=3 <args>
```
Also of course all of this works through environment variables.
```
FEX_DISASSEMBLE=blocks FEXInterpreter <args>
FEX_DISASSEMBLE=dispatcher FEXInterpreter <args>
FEX_DISASSEMBLE=blocks,dispatcher FEXInterpreter <args>
```
While only used fairly sparingly now, this is likely to have some
additional configurations using this in the future. Since we already
have some configs that are basically using enums, but just by doing
string comparisons.
This was asked for by a developer, so I figured I would throw it
together quick.
Ensures that we handle the AVX2 VSIB byte in a decent way.
As is, we can't compute the [index * scale] variant portion
of the entire address operand, since the scale needs to act
on every element of the vector after sign extension.
What we can do though, is compute the base address and add
the displacement to it ahead of time though.
All of these IR operations were being fairly inefficient in their
address calculation. All of these are known using power of 2 stride
indexing. So all of these can be converted from three instructions to
one.
These are always used for x87 stack accesses so each one gets an
improvement.
Before:
```asm
0x0000ffff6a800248 d2800200 mov x0, #0x10
0x0000ffff6a80024c 9b007e80 mul x0, x20, x0
0x0000ffff6a800250 8b000380 add x0, x28, x0
0x0000ffff6a800254 fd417805 ldr d5, [x0, #752]
```
After:
```asm
0x0000ffff91e80240 8b141380 add x0, x28, x20, lsl #4
0x0000ffff91e80244 fd417805 ldr d5, [x0, #752]
```
Currently we're clearing icache including the data that lives on the
tail of the block. Instead only clear the code that the was emitted and
not tail data.
Additionally only disasm the code rather than all the tail data as well,
as it gets unwieldy if viewing.
If we are loading exactly the flags we need from the RFLAGS (ensuring we
don't load the reserved flag in bit 1) then we don't need to do a mask
on the result.
Additionally there is some bad code-motion around selects that was
causing SBFE operations to occur on constants. Ensure that we const-prop
any SBFE operations to clean this up.
This PR along with #2783 causes FMOV blow-up to go from 41 instruction
to 31 instructions.
Noticed this when inspecting some code that was moving constant
`0x80808080` in to a register. Was using two move instructions when it
could have used a single bitmask move.
This now checks to see if a constant can be 32-bit encoded in a logical
bitmask move and uses that.
This instruction doesn't match ARM semantics very well since it returns
the position of the minimum element.
But at the very least the insert in to the final instruction can be a
bit more optimal, Converts an 5 inst eor+mov+mov+mov+mov in to 2 inst
mov+mov.
This works because `VUMinV` already zero extends the vector so the
position only needs to be inserted at the end.
32-bit and 64-bit SH{L,R}D matches behaviour of EXTR. Optimize to using
this op in that case.
This converts the lsl+lsr+orr sequence in to a single extr instruction.
16-bit still goes down the old path.
Weirdly this code manages to have a bad insert for no reason? But
unrelated since this happens in the old code as well.
```
%4(GPRFixed3) i64 = LoadRegister #0x0, #0x20, GPR, GPRFixed, u8:Tmp:Size
%5(GPR0) i64 = LoadRegister #0x0, #0x8, GPR, GPRFixed, u8:Tmp:Size
%6(GPRFixed0) i64 = Extr %5(GPR0) i64, %4(GPRFixed3) i64, #0x3e
```
Not sure why the SRA fails on that second LoadRegister.
There was one holdout variable that was in a TLS object in FEXCore. Move
it to the frontend with the rest of the TLS variables.
Allows us to remove "Frontend" TLS management to be the only TLS
management.
This pass is currently doing nothing in main.
Ever since we have enforced that LoadContext/StoreContext doesn't touch
GPRs and FPRs, this has only been eliminating flags.
Remove that usage of LoadContext/StoreContext and replace with their
their replacement of LoadRegister/StoreRegister for tracking GPR and FPR
accesses.
Stripped from #2700 since this is safe to merge.
Noticed while looking at #2700.
Testing doesn't currently see this as a bug but will once #2700 starts
optimizing StoreRegister+LoadRegister pairs.
Doesn't fix the issues in that PR, but this is one.
While we're in the area implementing the Scalar + Vector variants,
we may as well cross off the Vector + Immediate variants and
complete all of the load variants for the regular LD1{*} loads
Our regex would only ever capture a single digit, so versions that had
more than one digit per section would lose additional digits.
Fixes and moves the helper to a cmake file to be shared between
GuestLibs and HostLibs.
Uses the fix in xcb because Fedora ships an older version that doesn't
have some of FEX's newer symbols.
Removes the @PREFIX_ARCH@ replacement string in the thunks path.
The library prefix paths now get generated upfront and everything gets
replaced to handle the differences between multiarch distros.
Fixes Thunks on Arch and Fedora.
This is less noisy with no loss of clarity, and follows the notation
used by both LLVM IR and NIR. (So, it should be familiar.)
Change done with:
sed -i -e 's/%ssa/%/g' $(git grep -l '%ssa')
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This does duplicate the _Constant(1) but it doesn't matter because it
gets inlined into the eor anyway. There is no functional change here.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We store garbage in the upper bits. That's ok, but it means we need to
mask on read for correct behaviour.
Closes#2767
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
We can fold the Not into the And. This requires flipping the arguments
to Andn, but we do not flip the order of the assignments since that
requires an extra register in a test I'm looking at.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
WIN32 has a define already called `GetObject` and will cause our
symbol to have an A appended to it and break linking.
Just rename it to `GetTelemetryValue`
Noticed during introspection that we were generating zero constants
redundantly. Bunch of single cycle hits or zero-register renames.
Every time a `SetRFLAG` helper was called, it was /always/ doing a BFE
on everything passed in to extract the lowest bit. In nearly all cases
the data getting passed in is already only the lowest bit.
Instead, stop the helper from doing this BFE, and ensure the
OpcodeDispatcher does BFE in the couple of cases it still needs to do.
As I was skimming through all these to ensure BFE isn't necessary, I did
notice that some of the BCD instructions are wrong or questionable. So I
left a comment on those so we can come back to it.
These address calculations were failing to understand that they can be
optimized. When TSO emulation is disabled these were fine, but with TSO
we were eating one more instruction.
Before:
```
add x20, x12, #0x4 (4)
dmb ish
ldr s16, [x20]
dmb ish
```
After:
```
dmb ish
ldr s16, [x12, #4]
dmb ish
```
Also left a note that once LRCPC3 is supported in hardware that we can do a similar optimization there.
When this instruction returns the index in to the ecx register, this is
defined as a 32-bit result. This means it actually gets zero-extended to
the full 64-bit GPR size on 64-bit processes.
Previously FEX was doing a 32-bit insert which leaves garbage data in
the upper 32-bits of the RCX register.
Adds a unit test to ensure the result is zero extended.
Fixes running Java games under FEX now that SSE4.2 is exposed.
ARM64 BFI doesn't allow you to encode two source registers here to match
our SSA semantics. Also since we don't support RA constraints to ensure
that these match, just do the optimal case in the backend.
Leave a comment for future RA contraint excavators to make this more
optimal
Loaded 100 of 444 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.