This is a new procfs symlink path that changes behaviour of binfmt_misc
when exposed. We need to check both procfs/exe and procfs/interpreter
and see if they exist AND also differ.
Once/if they do then we can disable a bunch of checking of paths once
they do. The fallback when none of this is supported has the same
behaviour has previously where it still does all the regular checking.
During binfmt_misc install cmake will check the kernel version for the
raw binfmt_misc writing. Which will never pass until we have a real
kernel version that it is upstreamed in.
For update-binfmts we add a new optional argument where the tool will
drop the flag if the host kernel version isn't new enough to handle the
option.
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.
My concrete IR proposals are that:
* SSA values must be killed in the same block that they are defined.
* Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
* LoadGPR/StoreGPR are eliminated in favour of SSA within a block.
This has a lot of nice properties for our setting:
* Except for some internal REP instruction emulation (etc), we already have
registers for everything that escapes block boundaries, so this form is very
easy to go into -- straightforward local value numbering, not a full into
SSA pass.
* Spilling is entirely local (if it happens at all), since everything is in
registers at block boundaries. This is excellent, because Belady's algorithm
lets us spill nearly optimally in linear-time for individual blocks. (And
the global version of Belady's algorithm is massively more complicated...)
A nice fit for a JIT.
Relatedly, it turns out allowing spilling is probably a decent decision,
since the same spiller code can be used to rematerialize constants in a
straightforward way. This is an issue with the current RA.
* Register assignment is entirely local. For the same reason, we can assign
registers "optimally" in linear time & memory (e.g. with linear scan). And
the impl is massively simpler than a full blown SSA-based tree scan RA. For
example, we don't have to worry about parallel copies or coalescing phis or
anything. Massively nicer algorithm to deal with.
* SSA value names can be block local which makes the validation implicit :~)
It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.
Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now the calculation of PF is entirely deferred, by inverting our internal
representation of PF. All the (e.g.) logical op needs to do is store the low
8-bits of the result.
This is a bit of a mixed bag. Primary ALU ops all save an instruction, by
skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky
cases are:
* Zeroing PF. This now requires writing 1 instead of 0, which may require an
extra move for the constant. Some of this will go away when we merge PF+AF
into a single register, which is next up on the list. In that case, the
two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing
x86 view of PF), that will get absorbed into the or, if we also write AF. If
we leave AF undefined and need to write a 1, that's a single mov instruction
and we couldn't do better anyway if not inverted (since we'd still have a mov
wzr even then). So in view of the future work, this isn't something I'm
concerned about.
* Float comparisons that put Unordered into PF. These require an extra invert to
match the new convention. These are already so unnecessarily bloated that I'm
not convinced I'm making things materially worse here. But we realistically
need multiple destination support in the IR to fix this particular mess.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
AF is calculated as:
((Src1 ^ Src2) ^ Res)[4]
Due to the extract, this is equivalent to
((Src1 ^ Src2) ^ (Res ^ 1))[4]
We already store (Res ^ 1) as the PF byte. So, it suffices to store
AF Byte = Src1 ^ Src2
and then we can recover the flag value
AF = (AF Byte ^ PF Byte)[4]
This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:
* Both PF and AF written together, the coupling is correct.
* PF written but AF invalidated, irrelevant.
* Both invalidated, irrelevant.
None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Logical ops leave AF undefined so we can't expect it to be zero after. Mask the
result of lahf to avoid testing UB. These unit tests would regress from the work
in this MR otherwise.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Use an explicit invalidate, so we can zero easily enough if we need to for
debugging later but we can save the instrs ordinarily.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
FEX has a problem with large blocks that uses a ton of constants spread
throughout the block. Once a block gets large enough with enough
constants that have large live ranges, FEX slows down to unusable speeds
due to the register allocator spending more time calculating node
interferences than anything else in the program.
This adds a little heuristic to ensure that constants aren't reused if
the previous value is past a certain distance threshold. This threshold
works well enough that XeSS's pedantic initialization code doesn't have
issues now. See https://github.com/FEX-Emu/FEX/issues/2688 for more
information about that.
FEX itself should work to remove bad constant usages to make this pass
less necessary anyway. In most cases we are materializing duplicated 0,
1, and masks which could be done without a constant entirely.
Maybe once we've improve that enough we could remove this constant
pooling entirely.
To note, this doesn't fix the issue that XeSS causes our register
allocator, this is purely a heuristic workaround.
This class is very expensive to initialize so if you happen to have the
disassembler configuration enabled you were eating a very bad
initialization cost for no reason.
Only initialize the data member if any disassembler runtime option is
enabled, this completely removes the overhead.
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.
This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.
Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.
Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
Currently in main today, FEX fails to compact OF/CF/ZF/SF and PF.
This is due to recent optimizations with flag calculations on each of
these. Now that we have a centralized location where we compact and set
our internal representation of flags we can do this in one location.
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
Currently FEX's internal EFLAGS representation is a perfect 1:1 mapping
between bit offset and byte offset. This is going to change with #3038.
There should be no reason that the frontend needs to understand how to
reconstruct the compacted flags from the internal representation.
Adds context helpers and moves all the logic to FEXCore. The locations
that previously needed to handle this have been converted over to use
this.
32-bit or 64-bit addition without carry-in. This matches the baseline hardware
semantic. Generalizing to support other cases can come later, this should be a
win already.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For correct carry/overflow behaviour, we need to use a compare of the right size. The existing logic
to look at the source sizes doesn't work for this, since a 32-bit NEG instruction will compare a
32-bit source with a 64-bit _Constant(0) .. which needs a 32-bit compare but the existing logic
would use a 64-bit compare. This is not yet a bug fix, since the overflow code is currently in
software for 32-bit negates so it's irrelevant. But it should prevent regressions from using native
compares later in this series. Presumably this was intended all along but left as-is to avoid
disturbing instcountci once noticed. Time to disturb CI!
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
AddNZCV is a new op to return the NZCV for an addition directly, which lets us
skip software flag calculation in some cases. In the future it would be nice to
fuse this into the Add itself as a second destination to avoid repeating the
addition, but that's a very involved change and right now I'm building FEX on an
old Chromebook because my M1 kernel is FUBAR.
Similarly, SubNZCV returns flags for Sub. This has the extra twist of needing to
invert the carry bit due to the inverted definition between arm64 and x86_64.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This will let us reuse it in some cases. In some of these
implementations there is a bad code smell around using the zero register
but that isn't going to get solved in this commit.
A bunch of the AES operations take a zero register upfront and we
currently materialize it for each instruction.
Considering that most AES operations are used back to back, we can
eliminate these materializations by caching it between instructions.
Additionally removes a move in the optimal case when destination matches
the state register, which is exactly what the SSE operation ends up
doing.
AESKeyGenAssist has an edge case that if the destination RA overlaps the
zero register then we still need to eat a move, hopefully doesn't happen
too frequently in practice. This is also the lesser used instruction so
it isn't a big deal. RA constraints could solve that still.
This instruction has xmm0 be one of the implicit sources. We were
loading xmm0 twice. #2700 would also fix this but that breaks other
things for some reason.
We can load the swizzle table from our constant pool now. This removes
the only usage of VTMP3 from our Arm64 JIT.
I would say the this is now optimal for the version without RCON set.
With RCON we could technically make some of the move of the constant
more optimal.
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.
This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
These functions are called a lot....a lot a lot.
Optimize these in to a couple of ALU operations instead of a whole
table lookup. Confirming with output assembly that this becomes more
optimal.
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.
I would say this is now optimal.
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.
Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.
When consuming our own controlled data, we don't want the range clamping
to be enabled.
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.
This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.
Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
Noticed that we hadn't ever enabled this, which was a concern when our
GPR operations weren't as strict about leaving garbage in the upper bits
when operating as a 32-bit operation.
Now that our ALU operations are more strict about enforcing upper bit
zeroing we can enable this.
This causes Half-Life: Source FPS to get to > 200FPS finally. Causes
significant performance improvements for 32-bit games because we're no
longer redundantly moving registers before and after every operation.
Causing a bunch of 3-4 instruction sequences to convert to 1.
RAValidation was making an assumption that GPR register class would only
have up to 16 registers for either SRA or dynamic registers.
When running a 32-bit application we allow 17 GPRs to be dynamically
allocated, since we can take 8 back from SRA in that case.
Just split the two classes in the RAValidation pass since they will
never overlap their allocation.
Fixes validation in `32Bit_Secondary/15_XX_0.asm` locally that changed
behaviour due to tinkering.
When the application calls exit_group we no longer need to care about
cleanup because the entire process group is leaving.
Just immediately call exit group and get out. Might revisit this in the
future.
Fixes#2752
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.
This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.
In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.
[1] e89321dc60 for scalar.
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
FPR and converted it in-place.
These are now optimal in the case of AFP is unsupported.
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.
Needs #2994 merged first.
Huge thanks to @dougallj for the optimization idea!
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.
Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.
Needs #2993 merged first.
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.
Thanks to @rygorous for the idea!
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
Given the operation is:
Dst = Vector1 / Vector2
If Dst and Vector1 alias one another, then we can just perform
the division as is without any moving of data around.
Zero-extension will already occur if necessary upon storing.
Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
This now uses the new load element IR operation and makes these
instructions optimal.
LRPCPC3 will introduce instructions in the future for TSO emulation to
help these operations, but that doesn't exist today.
This can be done without storing any data to memory and also
compressing the amount of instructions being used.
Thanks to @dougallj for the optimization suggestions.
VRev32 matches Arm64 semantics directly.
LoadNamedVectorConstant allows FEX to quickly load "named constants".
This will allow us to have specific hardcoded vector constant values
that we can load with a ldr(State)+ldr(Value) and will be more abused in
the future.
This also allows us to do a very simple optimization in the future where
we can optimize away redundant loads of these loads if they are used
multiple times in the same block. (Not implemented here).
Two optimizations here:
1) The final VInsElement was generating three instructions
- This itself could have been change to vzip, which would have
removed two instructions.
2) Optimize how the pairwise elements are calculated to shave one
instruction off the calculation.
- addp odd elements and even elemnts first
- Then transpose those elements
- Then use one final addp to generate the result in the correct
order.
The ext+uabdl+addp pairs of operations could be reordered to shave off
one temporary register usage if we really care later.
Can be slightly more optimal with a slightly change algorithm but will
require implementing some IR ops which can be put off. It's only about
an instruction savings.
Even on platforms without SVE these have improved slightly but the real
improvement comes from SVE.
Adds some new InstCountCI files for SVE128 enabled testing.
Also enables SVE128 in the VEX maps. Host features should probably
enable SVE128 when SVE256/AVX is enabled, but that isn't the case today.
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.
With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).
This is a significant improvement even on platforms that only support
128-bit SVE.
With ASIMD FEX would never optimize BSL out of fear if some registers
overlapped it would break things. So it had previously always moved to a
temporary first and then moved the result back out when done.
Now instead check upfront if any of the source registers overlap the
destination. If the destination register overlaps any of the three
sources we can bsl, bit, or bif depending on which register gets
overlapped.
Worst case the destination doesn't overlap any of the source registers
and still needs these moves.
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.
Back to back bfi is actually less optimal due to dependency tracking.
With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
This paves the way to optimizing pushes in to both push operations and
push pair operations to more optimally match Arm64 push support.
While this does the first step for supporting the base push, we'll leave
optimizing push pairs to future work.
This is a bit of tricky operation where due to our our usage of SSA, the
incoming source isn't guaranteed to end its live-range at this
instruction.
This gives us a behaviour where to be optimal we need to take different
paths depending on if the incoming address register is the same as the
destination node.
Once we have form of RA constraints or non-SSA IR form that can
guarantee this restriction then this will go away.
FEX doesn't use the platform register on wine platforms so there is no
reason to save and restore it.
On Linux we can still use it at some point but for now it isn't part of
our RA.
Changes the idiom used for constant mask generation to a ternary.
This pattern is definitely used elsewhere in code but we can get rid of
all instances here.
This takes a similar approach to deferred signal handling and allows any given
thread to be interrupted while running JIT code by protecting the appropriate
page as RO. When the thread then enters a new block, it will try to acccess
that page and segfault. This is safer than just sending a signal to the thread
as that could stop in a place where JIT context couldn't be recovered correctly.
With WOW, all allocations from 64-bit code use the full address space
and limiting is handled on the syscall thunk side so theres need to
worry about STL allocations stealing AS.
This relies on wine's behaviour passing through linux paths and env vars,
so that the config in the user's home directory can be accessed outside
of the wine prefix.
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.
Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.
It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.
Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.
A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.
[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
- 3.6 - Restricted Subset of Segmentation
- `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
registers; the base for CS, DS, ES, and SS is ignored for 32-bit
mode, same as 64-bit mode (treated as zero).`
- 4.2.17 - MOV to Segment Register
- Will fault if SS is written (Breaking anything that writes to
SS).
- Will not fault if CS, DS, ES are written (Thus it sets the
segment but gets ignored due to 3.6).
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now all vbroadcast implementations go down the more optimal path.
For non-SVE 128-bit cases where we only have 128-bit wide registers,
we behave like ld1rqb and just act as a normal 128-bit load for
interface convenience.
In the case of running on a 128-bit SVE system this predicate wasn't
setup. Since we never had any predicate usage before this wasn't an
issue. Now that #2914 is using the 128-bit predicate we need to make
sure that we are generating it.
Allows the implementations of the vbroadcast instructions to perform the
load and broadcast in one operation as opposed to doing the load and then
broadcast separately.
Notably, the broadcasting loads can also be used on systems that have SVE 128-bit
support as well, not only 256-bit.
On non-SVE systems, we use the equivalent AdvSIMD instructions.
Instead of allocating a temporary copy of the string, return a view of
it instead. Should improve the performance of system calls that take
file paths. Since it was allocating a string for every single syscall
that uses them in this case.
For config values that were string objects we were unnecessary creating
copies each time the string was accessed.
Convert the () operator over to returning a reference.
The current implementation uses orr excessively. This has FEX missing
hardware optimization opportunities where some CPU cores will zero-cycle
move constants that fit in to the 16-bits of movz/movk.
First evaluate up front if the number of 16-bit segments is > 1, in
those cases we should check if it is a bitfield that can be moved in one
instruction with orr.
After that point we will use movz for 16-bit constant moves.
Additionally this optimizes the case where a constant of zero is loaded
to be a `mov <reg>, zr` which gets renamed in most hardware.
Commonly we are doing a BFI into a 32-bit register, which is hitting the
ubfx (lsr alias) path.
In the case of 32-bit destination we can also do a regular move, which
will take advantage of CPU's rename functionality and give a minor speed
boost.
Didn't notice this in the previous PR, When DUMPIR=stderr without and
selection of where to place it in PASSMANAGERDUMPIR it was supposed to
put the dumper at the end of the passes.
We need to make sure that it it placed at the end of the passes rather
than current `it`.
This will allow investigating the Arm64 directly next to the test, plus
publicly linking directly to badly behaving tests.
Perfect for nerdsniping implementations.
We can perform the SQRT first and then broadcast 1.0 into the destination
since all the intermediary work is done, meaning we don't have to worry
about Dst and Vector aliasing one another.
Pretty sure this is why CI is unhappy. If a test in a different file has
the same name then it is highly likely to conflict when nasm is
generating files and will overwrite and erase, causing CI to break.
Include the incoming json filename as part of the asm keys so it can't
conflict here.
If DumpIR is enabled but the PassManagerDumpIR option isn't enabled then
this currently does nothing.
As a convenience, enable dumping the final optimized IR if an option
hasn't been specified.
None of the instructions here are optimal, a couple are close.
Will be splitting each of the map tables in to their own json files
since each one can get fairly large.
We need to move the modifier enum out of the SVEMemOperand class
since it's also used with adr. Plus, this can also be convenient
not being tied down to the class itself.
This also makes accessing modifiers less noisy, since the class
When a 32-bit adcx instruction was encountered, it was getting treated
as a 16-bit adcx instruction instead. This is because of the 0x66 prefix
required to handle this instruction.
Adds a unit test to ensure it doesn't break again.
Implements CI for tracking instruction counts for generate blocks of
code when transforming from x86 to ARM64 assembly.
This will end up encompassing every instruction in our instruction
tables similarly to how our assembly tests try to test everything in our
instruction tables.
Incidentally, the data for this CI is generated using our assembly
tests. By enabling disassembly and instruction stats when executing a
suite of instructions, this gives the stats that can be added to a json
file.
The current implementation only implements the SecondGroup table of
instructions because it is a relatively small table and has known
inefficiencies in the instruction implementations. As this gets merged I
will be adding more tables of instructions to additional json files for
testing.
These JSON files will support adjusting CPU features regardless of the
host features so it can test implementations depending on different CPU
features. This will let us test things like one instruction having
different "optimal" implementations depending on if it supports SVE128,
SVE256, SVEI8MM, etc.
This initial instruction auditing is what found the bug in our vector
shift instructions by size of zero. If inspecting the result of the CI
run, you can tell that these instructions still aren't "optimal" because
they are doing loads and stores that can be eliminated.
The "Optimal" in the JSON is purely for human readable and grepping
ability to see what is optimal versus not. Same with the "Comment"
section.
According to my auditing spreadsheet, the total number of instructions
that will end up in these json files will be about 1000, but we will
likely end up with more since there will be edge cases that can be more
optimal depending on arguments.
This was confusingly split between Arm64Emitter, Arm64Dispatcher, and
Arm64JIT.
- Arm64JIT objects were unnecessary and free to be deleted.
- Arm64Dispatcher simulator and decoder moved to Arm64Emitter
- Arm64Emitter disassembler and decoder renamed
- Dropped usage of the PrintDisassembler since it is hardcoded to go
through a FILE* type
- We instead want its output to go through LogMan, which means using a
split Decoder+Disassembler object pair.
- Can't reuse the object from the vixl simulator since the simulator
registers the decoder as a visitor, causing the simulator to execute
while disassembling instructions if reused.
- Disassembly output for blocks and dispatcher now output through Logman
- Blocks wrapped in Begin/End text for tracking purposes for CI.
We don't currently have a device in CI that can run SVE with 128-bit
width registers. Until we have a device with this, make sure the vixl
simulator is also running the ASM tests in this width.
This was causing test failure locally where some values were set to
uninitialized data. Ensure that gregs, YMM, and MMX registers are all
zero initialized.
Requires the IR headerop to house the number of host instructions this
code is translating for the stats.
Fixes compiling with disassembly enabled, will be used with the
instruction count CI.
This is incredibly useful and I find myself hacking this feature in
every time I am optimizing IR. Adds a new configuration option which
allows dumping IR at various times.
Before any optimization passes has happened
After all optimizations passes have happened
Before and After each IRPass to see what is breaking something.
Needs #2864 merged first
This is a /very/ simple optimization purely because of a choice that ARM
made with SVE in latest Cortex.
Cortex-A715:
- sxtl/sxtl2/uxtl/uxtl2 can execute 1 instruction per cycle.
- sunpklo/sunpkhi/uunpklo/uunpkhi can execute 2 instructions per cycle.
Cortex-X3:
- sxtl/sxtl2/uxtl/uxtl2 can execute 2 instruction per cycle.
- sunpklo/sunpkhi/uunpklo/uunpkhi can execute 4 instructions per cycle.
This is fairly quirky since this optimization only works on SVE systems
with 128-bit Vector length. Which since it is all of the current
consumer platforms, it will work.
We need to know the difference between the host supporting SVE with
128-bit registers versus 256-bit registers. Ensure we know the
difference.
No functional change here.
This allows use to both enable and disable regardless of what the host
supports. This replaces the old `EnableAVX` option.
Unlike the old EnableAVX option which was a binary option which could
only disable, each of these options are technically trinary states.
Not setting an option gives you the default detection, while explicitly
enabling or disabling will toggle the option regardless of what the host
supports.
This will be used by the instruction count CI in the future.
Moves the dummy handlers over to this library. This will end up getting
used for more than the mingw test harness runner once the instruction
count CI is operational.
This was a debug LoadConstant that would load the entry in to a temprary
register to make it easier to see what RIP a block was in.
This was implemented when FEX stopped storing the RIP in the CPU state
for every block. This is now no longer necessary since FEX stores the
in the tail data of the block.
This was affecting instructioncountci when in a debug build.
I use this locally when looking for optimization opportunities in the
JIT.
The instruction count CI in the future will use this as well.
Just get it upstreamed right away.
`eor <reg>, <reg>, <reg>` is not the optimal way to zero a vector
register on ARM CPUs. Instead we should move by constant or zero
register to take advantage of zero-latency moves.
2023-08-08 22:24:11 -07:00
412 changed files with 72100 additions and 4493 deletions
Loaded 100 of 412 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.