This flag breaks FEX heavily for now.
glibc 2.38 started using this flag as an optimization for posix_spawn.
It will fall back to a "non-optimized" implementation if the clone
syscall returns EINVAL. For now do this while we investigate a more
proper implementation.
Should be backported to 2312.1.
In some situations TestNZ is generated with a constant that is using a
constant that can't fit inside of the tst instruction.
This was found in libGLX with virgl, crashing invalid instruction
generation and crashing steamwebhelper
These functions only want the GPRs returned for SRA. This is because the
signal handler needs this map to relation between x86 GPRs and AArch64
GPRs.
When we added AF and PF to the SRA array we accidentally started
returning two more GPRs to the frontend. This caused the signal
delegator to start corrupting the members after GPRs in FEX's CoreState.
Corrupting 16-bytes after the gregs[] array.
This included corrupting:
- es_idx, cs_idx, ss_idx, ds_idx, gs_idx, fs_idx, _pad[]
they're all copypastes of each other, unify into one general "bit test & perform
action" template. this means most of the wins from the previous commits now
apply for bt* without more copypaste.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
if the shift is < N, and we grab bit 0 after, we only need to consider <=N
bits of the source. this lets us use 32-bit lsr for 32-bit bt, which will
reduce masking in the next commit.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Split from #3284 without changing ownership semantics while I reduce the
debugging surface here.
Removes one usage of ParentThread from FEXCore. Which can be done since
it is no longer an opaque structure, we can read the StatusCode
directly.
No functional change.
When COND_AL was added it wasn't added to this helper. Since there is a
gap between the last condition, just early check the value.
Fixes reading beyond the end of the array
Since we're invoking curl directly, we don't need to wrap it in `sh -c`
with this function.
Fixes an issue where curl downloads to a non-escaped path weren't
working. Now they do.
As I was poking around erofs-utils documentation, I found out that
fsck.erofs actually provides an option for extracting erofs images
without using fuse.
This finally puts the erofs handling on feature parity with squashfs.
O_TMPFILE has a few minor problems that I have been thinking about for a
while. I just recently got reminded about this and remembered that most
problems get resolved by using memfd_create.
- O_TMPFILE is only supported on some filesystems.
- Supported filesystem must be the one mounted to the pathname being
opened.
- Only a minor inconvenience as tmpfs and all related filesystems
support this.
- An inode is actually created on whatever filesystem is backing the
folder.
- `/tmp/` must exist as a directory
- If this folder happened to not be mounted then these temporary
files wouldn't have been created.
- memfd_create doesn't have a folder that needs to exist.
- We were leaving the files open as read/write
- While we were rewinding the file offset, an misbehaving application
could have wrote garbage to the temp file.
- memfd sealing allows us to open the FD as RW and then seal its
capabilities, making it a read-only FD.
- We were leaking FDs opened with O_CLOEXEC
- We could have just opened the O_TMPFILE with O_CLOEXEC
- memfd also just supports this flag, so use it.
- No real issues, just nice to be sanitary here.
Overall this doesn't really change any behaviour, but it is nice to
cleanup some of the edges there.
These are only used by gdbserver for filling out its XML data structures
so just remove them from FEXCore.
Also fixes the ordering on RegNames to match the definition of the enum
class definition in CoreState. This has been out of correct order since
we reordered registers months ago.
As we are moving more and more OS specific code to the frontend, this is
another set of functions that can be moved to FEXLoader from FEXCore.
No functional change here, only code moved from protected to private and
to FEXLoader's SignalDelegator.
Once more thread handling is moved to the frontend we can move even more
out of FEXCore. As follows:
- CheckXIDHandler can get moved.
- First pthread FEX makes would just call this.
- Register/UnregisterTLSState
- This can happen in the clone/thread handler once the frontend
handles it.
This leaves very little in the backend and is mostly an interface for
passing signal data to the frontend that it needs once a signal has
occured.
It additionally also is used for `SignalThread`.
The frontend needs to be in control of how threads are created. This is
inherent to the fact that OS threads are OS specific. We currently have
this weird split that when initializing the FEXCore context, we create a
parent thread at all times.
This does some initial cleanup that gets the core initialization nearly
decoupled.
This has long since been unused. Originally implemented for some fuzzing
tests but has been abandoned and that should likely be implemented some
other way.
When a 32-bit imul was being executed it had a chance of returning
garbage data in the upper 32-bits of the 64-bit result.
While this didn't typically cause problems, this gets exacerbated from
32-bit applications executing multiplies for address calculations.
A combination of commits 7146691360 and
d01b457727 exposed this problem where
previously there would be multiple moves between the calculation and
data use which would have zero'd the upper bits for us previously.
Now that we are no longer doing that, we need to make sure the opcode
dispatcher doesn't generate broken code instead.
Fixes Dungeon Defenders, which hasn't worked since FEX-2308.
Adds an ASM test that ensures we don't break it again.
This is an option that has been long overdue for removal. It's original
intention was primarily to lie to the guest application about the number
of cores in the system. This allowed us to say that the system was only
single threaded which worked around some threading bugs that we had
early on.
This is no longer the case and now it is a confusing remnant of the past
that people think they need to set. Incorrectly assuming that "0" by
default means that FEX is doing some sort of disabling of threading and
forcing all emulation down one CPU core. This is not the case and has
never been the case, so removing the option makes that idea go away.
Stop lying to the application about getcpu, sched_getaffinity, and
sched_setaffinity.
- getcpu would wrap the cpu result modulo the count of cores
- sched_setaffinity wouldn't work at all
- sched_getaffinity lied and always reported full affinity of config
option
Split off from #3282 to reduce burden.
We can read the data member directly now since it isn't opaque. In fact
we already do in the signal handlers. Removes these redundant helpers.
Removes one usage of ParentThread in FEXCore.
This usually happens on backwards memcpy where we know the direction of
the copy because the code will typically do as follows:
```
std
rep movsb
cld
```
This is because the direction flag is part of the ABI and needs to be
set back to the forward direction if it was modified.
This typically doesn't get picked up on forward copies because we won't
have visibility of a cld instruction in the block.
This optimization allows us to only emit half of the code for the memcpy
if it is a compile time constant.
There's definitely some future task that could assume forward direction
if unknown and recompile the code if the assumption has failed, but not
doing that here.
While auditing our JIT to see if we have any AFP issues I noticed this.
RPRES has a hard dependency that AFP exists in order to be used, we were
hitting a case where RPRES would be enabled, but AFP is disabled by
default.
RPRES only changes behaviour when FPCR.AH is set which requires AFP.
While this doesn't affect any hardware today, it likely will in the
future.
Removes new warning from clang:
```
FEXCore/include/FEXCore/Core/CoreState.h:119:35: warning: not packing field 'DeferredSignalRefCount' as it is non-POD for the purposes of layout [-Wpacked-non-pod]
119 | NonAtomicRefCounter<uint64_t> DeferredSignalRefCount;
| ^
```
We already ensure sane data layout that this is unnecessary anyway.
Suspend may be called on a thread before it has finished WOW64 initialisation,
keep track of all initialized threads and fallback to direct
NtSuspendThread when this is the case.
The frontend shouldn't need to know any information about how to
reconstruct eflags. Just give us the information we need and it'll work
out.
There are still some inherit limitations of this and some edge cases
that might give invalid data, but it is roughly as close as it was
before.
Just provide if the PC was in the JIT, the host GPRs, and the PState object from the signal
information and FEXCore does the rest.
We don't need to change the signature for `SetFlagsFromCompactedEFLAGS`
because during reloading of register state automatically does this for
us.
Goal is to improve FEX performance with real world games. Let's start with
Sonic Mania. Two hot blocks from #2681.
I'm only bothering to track flagm for this because I'm super biased.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Many flag-generating instructions like cmp need to save calculations for
deferred PF and AF flag calculation. Currently, they require a store per flag,
which is prohibitively expensive for hot instructions like cmp. By instead
pinning PF/AF temporary results to registers (x26/x27 by convention here), we
eliminate many stores altogether and turn the rest into zero-cycle moves (on
64-bit at least, this isn't optimal for 32-bit emulation due to CTX->GetGPRSize
shenanigans, need to check if this requirement can be lifted..).
To implement, we model as SRA and then the existing SRA code is able to generate
good code with little manual tuning. (Future work will get us to excellent code
with more tuning ;) ).
The tradeoff is reducing the working dynamic GPR set by 2 registers, which might
increase spilling in some cases. I think it's worth it in practice, though.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
- sha1nexte
- Takes advantage of sha1h if supported
- Does the operation in a vector otherwise
- sha1msg2
- Instead of dumping everything to GPRs, we can do this with vectors
- Mostly matches ARM's sha1su1 instruction, but it is /just/
different enough to be annoying.
- sha256msg1
- Directly matches sha256u0
- Leaves the previous implementation alone
Not supposed to touch flags at all, so don't! instead of making a terrible mess
of csels. a lot less instructions, and probably faster because the branch should
be predicted correctly in practice in hot loops.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Instruction count CI has transformed the way we work on FEX… I love the system
and want to make it better. there’s one part of instruction count CI that isn’t
so lovable: the problematic “optimal” flag on instructions.
There are several issues with this flag, both philosophical and practical.
– it is tedious to update the optimal flag when making an implementation
optimal. The effect of that is discouraging people from making instructions,
optimal, or encouraging people to fail to update the flag, and dilute the value
of it. Either way, since we care far more about optimal implementations, then we
do about updating the flag, clearly we should prioritize the implementation and
not the flag. This issue was not obvious at the outset, when instruction count,
CI was introduced, and still quite small. The problem magnified when we started
duplicating instructions in bulk for different combinations of CPU features
(flagm, AFP, etc.) that intern multiplies the manual work required to update the
flags by the corresponding constant factor. if it comes down to a choice between
removing this extra coverage and removing the flag, I think we all agree that
removing the flag is the lesser evil.
– The definition of “optimal” is fundamentally problematic. I have often
improved the instruction count of an instruction that was already “optimal”.
This is all kinds of silly, and calls into question whether there’s any value
whatsoever in the existing classifications of the flag. Furthermore, it is often
unknowable, whether an implementation really is optimal. Is it possible to
implement BZHI (with flag calculations) in fewer than eight instructions? We
don’t know, and it’s silly to pretend that we do.
– as a consequence of the problematic definitions , there are so many errors in
both directions that I don’t think there’s much value in preserving the existing
classification at the expense of +progress. Being able to say “32% of
instructions are translated optimally” is neat, but it really doesn’t tell us
anything whatsoever when you dig a little deeper.
So, as the flag is misleading at best and perhaps harmful at worst, let’s remove
it and make the instruction count CI, more useful overall. let’s let the
expected count and the assembly speak for themselves, and cut away the chaff. if
we want a meaningless number to report to management, we can instead calculate
the average blowup factor ;-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If a bind fails then this is usually because PID reuse happened and the
unix domain socket path still existed on the filesystem. Alternatively
the process did an execve which left a dangling unix domain file.
Unlink the file and try again in these cases. It will almost certainly
work the second time.
Instead of relying on sleeping in accept, use poll. This allows us to
early shutdown instead of leaving a thread hanging stuck in accept.
Doesn't change behaviour with FEX gdbserver waiting for attach, but in
the future when we all full process trees to have gdbserver running this
will be more important.
Using the server mount folder works most of the time, but when running
under pressure-vessel this stacks directories in a weird way because the
mount folder has some tricks applied to it.
Expose the temp folder being used directly instead.
These are setup to be nullptr by default. Instead of providing no-op
lock instructions just have them be nullptr.
It's already part of the API that they need to be nullptr checked before
calling and this matches behaviour of the real libX11 library. Removes
some spam that is unnecessary.
Unused since TestNZ+NZCVSelect accomplishes the same and good riddance. Might
bring them back later for tbz/tbnz, but certainly not in this Selectful form.
(I added them when I thought we were going to RA the flags. With the more
effective static approach, we don't need this for that.)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now unused. If we bring it back, it should be brought back as CSSC only. On
non-CSSC platforms, an explicit cmp + predicated neg can be better.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Separate out the NZCV bits from the more complex stuff so we can specially
optimize the branches.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
In this mode, rather than the branch comparing its arguments and then jumping
based on the result, the branch simply jumps by the native comparison based on
the NZCV value. This allows us to map x86 branches to arm64 branches 1:1.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Resolving the issue that we can only ever have one gdbserver process
running consuming port 8086 (even if the port number is cute).
Doesn't give us anything yet but in the future will allow us to have
whole process trees running gdbservers that we can attach to.
Usually better in practice... some rotates are slightly regressed by this but
they were already terrible.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
via the 0 reg. We really need a more generalized approach to taking advantage of
wzr, but this optimizes the special case I care about for seta (saving a move to
make the impl optimal).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This will replace Select soon, as it lets us take advantage of
NZCV-generating instructions and it doesn't clobber NZCV.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Do it explicitly for sve-256 and punt on optimizing, so we avoid regressing code
gen otherwise.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Rather than the context. Effectively a static register allocation scheme for
flags. This will let us optimize out a LOT of flag handling code, keeping things
in NZCV rather than needing to copy between NZCV and memory all the time.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Some opcodes only clobber NZCV under certain circumstances, we don't yet have
a good way of encoding that. In the mean time this hot fixes some would-be
instcountci regressions.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Again we need to handle this one specially because the dispatcher can't insert
restore code after the branch. It should be optimized in the near future, don't
worry.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Semantics differ markedly from the non-NZCV flags, splitting this out makes it a
lot easier to do things correctly imho. Gets the dest/src size correct
(important for spilling), as well as makes our existing opt passes skip this
which is needed for correctness at the moment anyway.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This fixes an issue where CPU tunables were ending up in the thunk
generator which means if your CPU doesn't support all the features on
the *Builder* then it would crash with SIGILL. This was happening with
Canonical's runners because they typically only support ARMv8.2 but we
are compiling packages to run on ARMv8.4 devices.
cc: FEX-2311.1
SHA instructions are very large right now and cause register spilling
due to their codegen. Ender Lilies has a really large block in a
function called `sha1_block_data_order` that was causing FEX to spill
NZCV flags incorrectly. The assumption which held true before NZCV
optimizations were a thing was that all flags were either 1-bit in an
8-bit container, or just 8-bit (x87 TOP flag).
NZCV host flags broke this assumption by making its flags 32-bit which
ended up breaking when encounting spilling situations.
Replace every instance of the Op overwrite pattern, and ban that anti-pattern
from the codebase in the future. This will prevent piles of NZCV related
regressions.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The "create op with wrong opcode, then change the opcode" pattern is REALLY
dangerous. This does not address that. But when we start doing NZCV trickery, it
will get /more/ dangerous, and so it's time to add a helper and make the
convenient thing the safe(r) thing. This helper correctly saves NZCV /before/
the instruction like the real builders would. It also provides a spot for future
safety asserts if someone is motivated.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
It looks like currently FEX has a bug around implicit flag clobbering
with this pull request where an IR operation that implicitly clobbers
flags isn't correctly saving the NZCV flags before doing the operation.
Adds a unit test that specifically captures this issue. RAX will be 1 or
0 depending on if the flags are clobbered incorrectly or not.
When attempting to debug #3162 I had noticed spurious behaviour around
what I assumed to be eflags getting corrupt around inlined syscalls.
This turned out to be a red herring but to ensure we are still testing
this, create a fully fleshed out unit test.
This test ensures a couple of things.
1) A flag that is set or unset before a syscall doesn't have its data
corrupt
2) An inline syscall doesn't corrupt the eflags, checking the eflag
result after returning from the syscall.
3) A signal occuring while in an inline syscall returns the correct
eflags information in the signal handler information
This test gets accomplished by setting or unsetting a particular flag
and then calling the futex syscall in a way that is guaranteed to be
inlined and also wait forever. Then the parent thread will signal with a
SIGTERM and read back the signal information. It does this multiple
times for each flag we care about.
While tracking issues in #3162, I had encountered a random crash that I
started hunting. It was very quickly apparent that this crash was
unrelated to that PR. I just happened to be running a unittest that was
creating and tearing down a bunch of threads that exacerbated the
problem.
See as follows with the strace output:
```
[pid 269497] munmap(0x7fffde1ff000, 16777216) = 0
[pid 269497] munmap(0x7fffde1ff000, 16777216 <unfinished ...>
[pid 268982] mmap(NULL, 16777216, PROT_READ|PROT_WRITE|PROT_EXEC, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7fffde1ff000
[pid 269497] <... munmap resumed>) = 0
```
One thread is freeing some memory with munmap, another one then does a mmap and gets the same address back.
Nothing too crazy at initial glance, but taking a closelier look, we can
see that there are two strange oddities:
1) We are double unmapping the same address range through munmap
2) The second munmap is interrupted and returns AFTER the mmap.
This has the unfortunate side-effect that the mmap that just returned
the same address has actually just been unmapped! This was resulting in
spurious crashes around thread creation that was SUPER hard to nail
down.
The problem comes down to how code buffer objects are managed, in
particular how the Arm64Emitter and Dispatcher handled its buffers.
Arm64Emitter is inherited by two classes; Dispatcher, and Arm64JITCore.
On class destruction the emitter would free its internal tracking
buffer. Additionally on destruction, the Arm64JITCore would walk through
all of its CodeBuffers and free them. The problem ends up being that in
the Arm64JITCore, it would free its code buffers which also ended up
being the current active buffer bound to the Arm64Emitter. Thus causing
the Arm64Emitter to come back around and try to free the same buffer
again.
This is a double-free problem! and was only visible on thread exiting!
Can't track double frees with mmap and munmap with current tooling!
This problem typically didn't occur because of how fast the destruction
usually takes and jemalloc inbetween also typically means the problem
doesn't occur. Initially thinking this was a threaded pool allocator bug
because typically the new allocation would end up in there once a new
thread was spinning up.
Now we change behaviour, Arm64Emitter doesn't do any buffer management
itself, instead just passing an initial buffer on to its internal buffer
tracking if given one up front.
This leaves the Dispatcher and the Arm64JITCore to do their buffer
management and ensuring there is no double free.
The day is saved!
This mask was being used incorrectly, it's a GPR spill mask for host
GPRs not an index in to the SRA array. Search the array of SRA registers
for the first one in the mask first to use as a temporary.
Fixes an issue with 32-bit inline syscalls where the first register
being spilled was r8, which was beyond the size of SRA registers on
32-bit processes. This would cause FEX to read the value just after
x32::SRA which is x32::RA. This would mean it would use r20 as a
temporary, corrupting the register in the process.
I noticed this while poking at #3162, but also when I was looking at a
memory buffer ownership problem.
Requires #3249 to be merged first
Library alerting has been disabled for now, and storing IR while
gdbserver is running is removed.
Otherwise no functional change.
When attempting to read files that aren't backed by a filesystem then
our current read file helpers fail since they query the file size
upfront.
Change the helper so that it doesn't query the size and just reads the file if it
can be opened. This lets us read `/proc/self/maps` using helpers.
GDBServer is inherently OS specific which is why all this code is
removed when compiling for mingw/win32. This should get moved to the
frontend before we start landing more work to clean this interface up.
Not really any functional change.
Changes:
FEXCore/Context: Adds new public interfaces, these were previously
private.
- WaitForIdle
- If `Pause` was called or the process is shutting down then this
will wait until all threads have paused or exited.
- WaitForThreadsToRun
- If `Pause` was previously called and then `Run` was called to get
them running again, this waits until all the threads have come out
of idle to avoid races.
- GetThreads
- Returns the `InternalThreadData` for all the current threads.
- GDBServer needs to know all the internal thread data state when the
threads are paused which is what this gives it.
GDBServer:
- Removes usages of internal data structures where possible.
- This gets it clean enough that moving it out of FEXCore is now
possible.
These should always be used in the dispatcher rather than the raw jumps they
translate to, as they ensure that flags are flushed. Eliminates a class of bugs
that will become a lot easier to hit with the new nzcv work.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Requires #3238 to be merged first since this uses the tbx IR operation.
Worst case is now a three instruction sequence of ldr+ldr+tbx.
Some operations are special-cased, which definitely doesn't cover all
possible cases we could use without tbx, but as a worst case improvement
this is a significant improvement.
We can't sanely test all 256 swizzle masks, so walk the array of random
data for each swizzle and CRC the results each step along the way.
This will give us a final result rather than each individual step.
A bunch of blendps swizzles weren't optimal. This optimizes all swizzles
to be optimal.
Two instructions can be more optimal without a tbx but the rest required
tbx to be optimal since they don't match ARM's swizzle mechanics.
gdbserver is currently entirely broken so this doesn't change behaviour.
The gdb pause check that we originally had an excessive amount of
overhead.
Instead use the pending interrupt fault check that was wired
up for wine.
This makes the check very lightweight and makes it more reasonable to
implement a way to have gdbserver support attaching to a process.
gdb gets angry if we return text with `<No Name>` in an xml text field.
Instead only return a name if we have one and gdb will take care of the
rest.
Additionally change the formatting of the return packet, it doesn't need
the xml version header.
Messed up when originally implementing this, substr's second argument is
requested substring length, not the ending position.
Noticed this while trying to parse multiple FEX_HOSTFEATURES options.
Everytime I want to quickly output a value for testing I tend to use
Print which didn't work under the simulator.
Give this a quick fix to wire up the jump to the vixl sim.
Currently we don't get why an IR emit failed in the assert message. Put
the code in to the message so it is easier to see.
This also resolved the issue that when in RelWithDebInfo the assert line
would typically be the end of the IR emission function, so you couldn't
see which assert actually triggered. Now since the message is printed
this is easier
Before:
```
[ASSERT]
```
After:
```
[ASSERT] Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit
```
This is a lot easier and better data than what #3227 proposed.
If the fetch and nonfetch versions mismatched then the DCE optimization
which changes the IR operation would break handily.
Add a comment as a reminder so if anyone touches this they will
understand.
When the result of an atomic fetch operation is unused then we can
safely convert it to a non-fetching version of the operation.
This happens hundreds of times per process as far as I can tell. No idea
if this actually helps any hardware, but theoretically it can allow CPUs
to not stall waiting for writeback of the atomic operation.
Couldn't detect any performance improvements in the various little
things I was poking at least. Very trivial to support so add it.
Unaligned variants are already handled in our unaligned fault handler
since the only difference is the acquire semantic is dropped and the
destination register is the zero register.
Secondary ALU operations were missed and when the operation is 4-bytes
in size we can also allow garbage upper bits since the JIT will emit a
32-bit operation for this instruction which is safe.
Optimizes some bad codegen around 32-bit ALU operations.
Instead of using VInsElement in pi2fw and pf2iw, just use uzp1 to ensure
we don't unintentionally add to RA pressure.
Additionally we can generate the constant needed for pmulhrw directly
using the movi instruction. Converts two instructions in to one.
Under FEX's current constraints this makes all 3DNow! instructions
optimal.
This PR has a bug around flags calculation and REP LODS{B,W,D,Q}.
This currently passes on main but fails on #3162.
Bug only occurs in 32-bit instead of 64-bit with the same test. Should
help diagnose the bugs in #3162.
When SubShift (LSL) occurs with both sources constant then optimize away
the calculation.
Additionally if add is found to have one immediate constant where the
inverse of the constant fits in to ImmAddSub range, then invert the
constant and change it in to a sub.
This optimizes the cases when direction flag is known upfront in an
instruction.
Previously this moved two constant, did a compare and a csel. Four
instructions in total. It also corrupts NZCV which we want to use for
other things.
This new codegen emits one constant and one subtract instruction, two
instructions total and doesn't touch NZCV.
More optimal!
Audit the code base and mark any instruction that implicitly clobbers flags so
it can get special handling in the dispatcher to spill NZCV ahead of emitting.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Lots of instructions clobber NZCV inadvertently but are not intended to write to
the host flags from the IR point-of-view. As an example, Abs logically has no
side effects but physically clobbers NZCV due to its cmp/csneg impl on non-CSSC
hw. Add infrastructure to model this in the IR so we can deal with it when we
start using NZCV for things.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
So we don't need to mark VInsertElement as implicit clobber in the common case.
Only afects sve256 which doesn't exist yet.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This lets most of the ASM tests run on 16K Linux hosts which is good because I
have a Mac and I'm bad at computer.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
With the previous RCLSE pass optimization that fixes store->load
forwarding, this pass started optimizing harder.
This hit a bug with this vmov removal that previously didn't get hit.
In particular this would eliminate vmov IR operations even if they were
zero extending a vector.
Since we have dramatically cleaned up the amount of vmov IR operations
we are generating, remove this optimization entirely. In the games I
tested, the only game that hit this "optimization" was Ender Lilies and
it started generating broken code for the single block of instructions
that did.
Adds a unit test for this case just in-case it comes back in the future
for some reason.
Fixes an issue where Ender Lilies would flash the screen to black every
time an enemy hit the player character.
These instructions aren't super amazing due to the fact that they have
both a source mask and a destination duplication mask.
Setup a case where we can generate more optimal code in /most/ cases.
There are a few that still fall down a "bad" path for the result
broadcast but in most cases they are optimal. Still to be seen what
games typically use the broadcast mask as.
AVX in its infinite wisdom expanded DPPS to 256-bit, while leaving DPPD
to only support 128-bit still. This leaves the original implementation
alone for 256-bit DPPS since I don't want to break it.
This is another instruction that gets a free optimization when
SVE-128bit is supported!
If no registers alias, then we can move the first source directly into the
destination and then perform the FCADD operation as opposed to using a
temporary.
These annotations allow for a given type or parameter to be treated as
"compatible" even if data layout analysis can't infer this automatically.
assume_compatible_data_layout is more powerful than is_opaque, since it
allows for structs containing members of a certain type to be automatically
inferred as "compatible".
Conversely however, is_opaque enforces that the underlying data is never
accessed directly, since non-pointer uses of the type would still be
detected as "incompatible".
This annotation can be used for data types that can't be repacked
automatically even with custom repack annotations. With ptr_passthrough,
the types are wrapped in guest_layout and passed to the host like that.
Previously, two functions with the same signature would always be wrapped
in the same logic. This change allows customizing one function with
annotations while leaving the other one unchanged.
We can perform less moves by checking for scenarios where aliasing
occurs. Since addition is commutative (usually, general-case anyway),
order of inputs doesn't strictly matter here.
In the event no source vectors alias the destination,
we can just move the first source vector into it and
then perform the divide without needing to move afterword.
When a syscall from the *at series is provided an FD but the path is
absolute then dirfd should be ignored. We weren't correctly doing this.
Now if the path is absolute, but set the argument to the special
AT_FDCWD..
Fixes#3204
We can avoid needing to use movprfx here by moving
directly into the destination when possible and just
doing the UMAX directly.
Also expands the unsigned max tests to test values with
the sign bit set to ensure all behavior is caught.
Since SMAX performs a comparison and returns the max value regardless
of how the operands are provided, we can check for when the second
input aliases the destination.
Removes the truncating move that we perform inside the StoreResult
function and instead delegates the responsibility to the instruction
implementations themselves.
This removes a lot of redundant moves that occur on 128-bit variants
of AVX instructions.
Also fixes a weird case where we were handling 128-bit SVE
in VBroadcastFromMem when we already have AdvSIMD instructions
that will perfom the zero-extension behavior for us.
Allows for easier expansion without needing to expand the function definitons.
Also makes a few usages significantly less verbose and makes specifying
options a little more declarative, rather than having to memorize what
each argument is specifying.
Notably this bugfix version also introduces support for formatting
std::atomic types and std::atomic_flag.
Also, of course keeps our tracked external up to date.
A couple of games were hitting these. Not sure how they were missed in
PR #3159 but adds the missing one.
Small rearrangement to make this easier as well. Hopefully thunk stuff
lands sooner rather than later to automate this for Vulkan.
Maybe `-isystem` instead of `-I` needs to be used unlike what #2076,
might depend on what is installed on the host system.
Some simple tests to showcase instructions that we can optimize.
- Back to back pushes could be optimized
- Back to back scalar vector operations can be optimized
- Show with AFP that back to back scalar is already optimal
- Also ensures we don't break this stage.
We don't need zero in the upper bits for a push.
Makes a couple variants optimal.
Adds missing tests to the 32-bit file, since only 32-bit can push a
32-bit register.
These IR operations are required to support AFP's NEP mode which does
vector insert in to the destination register. Additionally it gives us
tracking information to allow optimizing out redundant inserts on
devices that don't support AFP natively.
In order to match x86 semantics we need to support binary and unary
scalar operations that do a final insert in to a vector. With optional
zeroing of the top 128-bits for AVX variants.
A tricky thing is that in binary operations this means that the
destination and first source have an intrinsically linked property
depending on if it is SSE or AVX.
SSE example:
- addss xmm0, xmm1
- xmm0 is both the destination and the first source.
- This means xmm0[31:0] = xmm0[31:0] + xmm1[31:0]
- Bits [127:32] are UNMODIFIED.
FEX's JIT jumps through some hoops so that if the destination register
equals the first source register, then it hits the optimal path the
AFP.NEP will insert in to the result. AVX throws a small wrench in to
this due to changed behaviour
AVX example:
- vaddss xmm0, xmm1, xmm2
- xmm0 is ONLY the destination, xmm1 and xmm2 are the sources
- This operation copies the bits above the scalar result from the
first source (xmm1).
- Additionally this will zero bits above the original 128-bit xmm
register.
- xmm0[31:0] = xmm1[31:0] + xmm2[31:0]
- xmm0[127:32] = xmm1[127:32]
- ymm0[255:127] = 0
This causes these instructions to support a fairly large table depending
on if the instruction is an SSE or AVX instruction, plus if the host CPU
supports AFP or not.
So while fairly complex, it's handling all the edge cases and gives us
optimization opportunities as we move forward. Currently on non-AFP
supporting devices this has a minor benefit that these IR operations
remove one temporary register, lowering the Register Allocation
overhead.
In the coming weeks I am likely to introduce an optimization pass that
removes redundant inserts because FEX currently does /really/ badly with
scalar code loops.
Needs #3184 merged first.
When FEX is in the JIT we need to make sure to enable NEP and AH and
then disable when leaving.
Explicitly disabled when the vixl simulator is used since even
attempting to set the bits will cause it to fault out. Ensures
InstCountCI keeps working.
There are some cases where we want to test multiple instructions where
we can do optimizations that would overwise be hard to see.
eg:
```asm
; Can be optimized to a single stp
push eax
push ebx
; Can remove half of the copy since we know the direction
cld
rep movsb
; Can remove a redundant insert
addss xmm0, xmm1
addss xmm0, xmm2
```
This lets us have arbitrary sized code in instruction count CI, with the
original json key becoming only a label if the instruction array is
provided.
There are still some major limitations to this, instructions that
generate side-effects might have "garbage" after the end of the block
that isn't correctly accounted for. So care must be taken.
Example in the json
```json
"push ax, bx": {
"ExpectedInstructionCount": 4,
"Optimal": "No",
"Comment": "0x50",
"x86Insts": [
"push ax",
"push bx"
],
"ExpectedArm64ASM": [
"uxth w20, w4",
"strh w20, [x8, #-2]!",
"uxth w20, w7",
"strh w20, [x8, #-2]!"
]
}
```
This adds all the missing atomic tests in to their own tests files.
This includes all of them except a few choice ones that are in their
original files.
- BTC, BTR, BTS are in their Secondary/SecondaryGroup files
- CMPXCHG, CMPXCHG8B, CMPXCHG16B are in their Secondary/SecondaryGroup
files
- These always imply lock semantics even without the prefix.
Six of the EFLAGS can't be used directly in a bitmask because they are
either contained in a different flags location or has multiple bits
stored in it.
SF, ZF, CF, OF are stored in ARM's NZCV format in offset 24.
PF calculation is deferred but stored in the regular offset.
AF is also deferred in relation to the PF but stored in the regular
offset.
These /need/ to be reconstructed using the `ReconstructCompactedEFLAGS`
function when wanting to read the EFLAGS.
When setting these flags they /need/ to be set using
`SetFlagsFromCompactedEFLAGS`.
If either of these functions are not used when managing EFLAGs then the
internal representation will get mangled and the state will be
corrupted.
Having a little `_RAW` on these to signify that these aren't just
regular single bit representations like the other flags in EFLAGS should
make us puzzle about this issue before writing more broken code that
tries accessing it directly.
This allows us to use reciprocal instructions which matches precision of
what x86 expects rather than converting everything to float divides.
Currently no hardware supports this, and even the upcoming X4/A720/A520
won't support it, but it was trivial to implement so wire it up.
Suggested by Alyssa. Adding an IR operation can be a little tedious
since you need to add the definition to JIT.cpp for the dispatch switch,
JITClass.h for the function declared, and then actually defining the
implementation in the correct file.
Instead support the common case where an IR operation just gets
dispatched through to the regular handler. This lets the developer just
put the function definition in to the json and the relevent cpp file and
it just gets picked up.
Some minor things:
- Needs to support dynamic dispatch for {Load,Store}Register and
{Load,Store}Mem
- This is just a bool in the json
- It needs to not output JIT dispatch for some IR operations
- SSE4.2 string instructions and x87 operations
- These go down the "Unhandled" path
- Needs to support a Dispatcher function override
- This is just for handling NoOp IR operations that get used for
other reasons.
- Finally removes VSMul and VUMul, consolidating to VMul
- Unlike V{U,S}Mull, signed or unsigned doesn't change behaviour here
- Fixed a couple random handler names not matching the IR operation
name.
This syscall requires a valid pointer otherwise it returns EFAULT.
When going through the glibc helper it can crash before reaching the raw
syscall even.
Enables in InstCountCI so Pi users can run InstCountCI can run the tests
without breaking on crypto operations.
When crypto is enabled or disabled just wholesale change AES, CRC32, and
PMULL 128-bit in one step. We don't really care about partial support
here.
The motivation towards just having a pointer array in CpuState was that
initialization was fairly cheap and that we have limited space inside
the encoding depending on what we want to do.
Initialization cost is still a concern but doing a memcpy of 128-bytes
isn't that big of a deal.
Limited space in CpuState, while a concern isn't a significant one.
- Needs to currently be less than 1 page in size
- Needs to be under the architectural offset limitations of loadstore
scaled offsets. Which is 65KB for 128-bit vectors
Still keeps the pointer array around for cases when we would need
synthesize an address offset and it's just easier to load the
process-wide table.
The performance improvement here is removing the dependency in the
ldr+ldr chain. In microbenchmarks this has shown to have an improvement
of ~4% by removing this dependency chain on Cortex-X1C.
This is the cause of a bunch of redundant moves that shows up in
InstCountCI. Fixing this aliasing and pre-colouring issue causes a ton
of 256-bit operations to become optimal.
Using the cached zero value is less efficient than loading it in to the
register for all these cases.
Lets us use rename hardware more efficiently and removes a dependency
chain on a single register.
Original:
```
movi v2.2d, #0x0
mov z16.d, p7/m, z2.d
<... 16 more times>
mov z31.d, p7/m, z2.d
```
Result:
```
movi v16.2d, #0x0
<... 16 more times>
movi v31.2d, #0x0
```
This is a quality of life improvement for people that want to tinker
with the InstCountCI but they may not necessarily have an Arm64 device
available immediately for poking.
As long as the vixl disassembler is enabled then the InstCountCI tests
can run and get bit-accurate encodings just like on an Arm64 device.
This also ensures that behaviour is consistent with or without the vixl
simulator enabled which is very important when running on x86 hosts.
This runs the data layout analysis pass added in the previous change twice:
Once for the host architecture and once for the guest architecture. This
allows the new DataLayoutCompareAction to query architecture differences for
each type, which can then be used to instruct code generation accordingly.
Currently, type compatibility is classified into 3 categories:
* Fully compatible (same size/alignment for the type itself and any members)
* Repackable (incompatibility can be resolved with emission of automatable
repacking code, e.g. when struct members are located at differing offsets
due to padding bytes)
* Incompatible
The set of these types is tracked in AnalysisAction, to which extensive
verification logic is added to detect potential incompatibilities and to
enforce use of annotatations where needed.
This was only required on x86 devices trying to escape the emulation.
Since x86 is now remove, this is entirely unnecessary.
When Steam launches applications with `/bin/sh`, this will remain under
the emulation and not escape these days.
With the removal of the x86 JIT, there is no need to have these be
independent classes.
Merges the Arm64Dispatcher in to the base Dispatcher class.
No functional change, just moving code.
Similar to previous tests, vpgatherqq and vgatherqpd are equivalent
instructions. So the tests are the same with the mnemonic changed.
This adds tests for an additional two sets of instructions. Getting us
full coverage of all eight instructions if we include the tests from
PR #3167 and #3166
Tests the same things as described in #3165
In addition, since these tests use 64-bit indices for address
calculation, we can easily generate and indice vector that tests
overflow. So every test at every displacement ALSO gains an additional
overflow test to ensure correct behaviour around pointer overflow
calculation.
Similar to previous tests, vgatherqd and vgatherqps are equivalent
instructions. So the tests are the same with the mnemonic changed.
This adds tests for an additional two sets of instructions, Getting us
up to six total over the eight if we include the tests from #3166.
Tests the same things as described in #3165
In addition, since these tests use 64-bit indices for address
calculation, we can easily generate and indice vector that tests
overflow. So every test at every displacement ALSO gains and additional
overflow test to ensure correct behaviour around pointer overflow
calculation.
Just like the previous tests, vpgatherdq and vgatherpq are equivalent
instructions. So the tests are the same except for the instruction
mnemonic again.
This adds unittests for two more of the eight gather instructions.
Getting us up to testing four in total.
Specifically this adds tests for 32-bit indices while loading 64-bit
element instructions.
Same thing as PR #3165 for what it tests versus doesn't.
vpgatherdd and vgatherps are effectively the same instructions, so the
tests are the same except for the instruction mnemonic.
This adds unit tests for two of the eight gather instructions.
Specifically this adds tests for the 32-bit indices loading 32-bit
elements instructions.
What it tests:
- Tests all displacement scales
- Tests multiple mask arrangements
- Ensures the mask register is zero'd after the instruction
What it doesn't test:
- Doesn't test address size calculation overflow
- Only would happen on 32-bit with 32-bit indices, or /really/ high
base addresses
- The instruction should behave as a mask to the address size
- Effectively behaves like `(uint64_t)(base + index << ilog2(scale))`
- Better idea is to just not expose AVX to 32-bit applications
- Doesn't test VSIB immediate displacement
- This just ends up being base_addr + imm so it isn't too interesting
- We can add more tests in the future if we think we messed that up
- Doesn't test partial fault behaviour
- Because that's a nightmare.
Specifically keeps each instruction test small and isolated so if a
single register fails it is very easily to nail down which operation did
it.
I know some of our ASM tests do a chunk of work and spit out a result at
the end which can be difficult to debug in some cases. Didn't want to do
that which is why the tests are spread out across 16 files for these
single class of instructions.
If we ever get around to fusing ops with shifts in the ConstProp optimizer (may
or may not be worthwhile), this will delete an instruction from things like "or
al, bh".
Even though lsr is the same speed as bfe on Firestorm, I feel if you ask for
garbage you should get garbage C:
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Pointless, upper bits ignored anyway. Deletes piles of uxt and even some 32-bit
instruction moves.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
To load 8-bit sources without bfe'ing for al/bl/cl if the caller knows it
doesn't need masking behaviour, but without lying about the size so the extract
for ah/bh/ch will still work properly.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.
Deletes a pile of code both in FEX and the output :-)
ADC/SBC could probably get similar treatment later.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Use the raw popcount rather than the final PF and use some sneaky bit math to
come out 1 instruction ahead.
Closes#3117
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Mostly copypaste of Orlshl... we really should deduplicate this mess somehow.
Maybe a shift enum on the core Or op?
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
If we const-prop the required functions and leafs then we can directly
encode the CPUID information rather than jumping out of the JIT.
In testing almost all CPUID executions const-prop which function is
getting called. Worst case that I found was only 85% const-prop rate.
This isn't quite 100% optimal since we need to call the RCLSE and
Constprop passes after we optimize these, which would remove some
redundant moves.
Sadly there seems to be a bug in the constprop pass that starts crashing
applications if that is done.
Easily enough tested by running Half-Life 2 and it immediately hitting
SIGILL.
Even without this optimization, this is stil a significant savings since
we aren't jumping out of the JIT anymore for these optimized CPUIDs.
Most CPUID routines return constant data, there are four that don't.
Some CPUID functions also need the leaf descriptor, so we need to
describe that as well.
Functions that don't return constant data:
- function 1Ah - Returns different data depending on current CPU core
- function 8000_000{2,3,4} - Different data based on CPU core
Functions that need leaf constprop:
- 4h, 7h, Dh, 4000_0001h, 8000_001Dh
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This option was disabled a few months ago when we switched the server
socket from a filesystem unix socket to an abstract socket.
This partially broke our chroot scripts which relied on this option
existing.
Readds support for an explicitly named abstract socket named from
config.
This is a workaround for dealing with chroots that change users.
They end up changing a user while doing operations and then can't
connect to the FEXServer anymore because environment variables have been
wiped away.
movprfx is invalid to use when the source register matches the movprfx
destination.
This was getting picked up on by `TwoByte/0F_D1.asm` now that RCLSE is
working better now.
The bug that was causing crashes with this was due to inline syscalls.
Now that this is fixed we can re-enable store->load operations.
This allows constant propagation to work significantly better, which
means inline syscalls start working again. This can significantly
improve syscall performance in some cases.
This is most likely to improve performance in dxsetup and vc_redist but
hard to get a real profile.
Additionally this will let us inline cpuid results in the future which
is pretty nice.
Ever since we reordered registers in `X86Enums.h` this has silently been
broken. This wasn't hit because RCLSE has been broken ever since SRA was
added, so inlinesyscalls just weren't ever happening.
Quick fix while I think of a way to more strictly correlate these
registers so it doesn't happen again.
The range was slightly incorrect which mostly wouldn't have caused
issues.
The lowest byte would have just generated slightly less optimal code.
The upper byte could have generated broken code, which our CI couldn't
catch since TSO instructions only get enabled when multiple threads are
in-flight.
Easy enough to fix.
This would have caused core to try and initialize a custom core on
Arm64, which causes a std::function assert because it doesn't support
that.
Users would likely get hit by this immediately since we deleted the
interpreter and shifted all the core numbers.
Originally this was going to use setf8/setf16, but it looks like the approach of
shift-and-test turns out to be faster. As a bonus this is a nice delete-the-code
win :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The only reason we need to XOR arguments for AF is to get bit 4 correct. But if
the operand in question is known to have bit 4 clear, the XOR will be an
effective no-op and can be skipped. This saves an instruction in a bunch of
common cases, like inc/dec. If we dedicated a register to AF to eliminate the
store, we would not save an instruction from this but would still come out ahead
due to an eor turning into a (zero cycle?) mov that can be handled by the
renamer.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Add new synthetic condition codes that do an AND as their relational operator,
testing the result. This is 1 IR op for things like
(A & B) == 0 ? C : D
This can translate to
tst A, B
csel A, B, eq
In the future, if A is the NZCV register and B is a supported immediate, eg
(NZCV & 0x80000000) == 0 ? C : D
this will be able to translate to a single instruction with the appropriate
condition
csel A, B, pl
but that needs RA support.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This provides more robust handling than a signal based approach, as the
suspender is able to wait for the suspendee to reach a suitable position and
flush its context to memory before returning.
This should support most simple cases of SMC, however programs which make use
of separate shared memory mappings for writing and execution are not handled.
The overall approach is the same as is done for linux, where RWX mappings are
protected to RX and then when a write occurs the signal handler invalidates the
faulting page and reprotects it to RWX until code in that page is jitted again.
When an exception occurs, pretend that we were just at the point of JIT entry
so the stack can be unwound to the wow64 SEH handler, which then handles
dispatching the exception to the x86 guest with the restored context.
This allows for running x86 applications under wine without having to run all
of wine under FEX. The JIT is invoked when running application code and then
left when handling NT syscalls or unix calls to e.g. the Vulkan driver.
The MinGW supplied import libraries are incomplete and miss a lot of
functions necessary to implement lower level windows code. To avoid
needing to many resolve every function, pull in .def files from wine
that detail the entire ntdll and wow64 APIs.
These are cut down versions of wine headers containing only what is necessary
for WOW. This shouldn't carry any license implications for FEX, as per the
LGPLv3 license:
```
The object code form of an Application may incorporate material from a header
file that is part of the Library. You may convey such object code under terms
of your choice, provided that, if the incorporated material is not limited to
numerical parameters, data structure layouts and accessors, or small macros,
inline functions and templates (ten or fewer lines in length), you do both of
the following:
a) Give prominent notice with each copy of the object code that the Library is
used in it and that the Library and its use are covered by this License.
b) Accompany the object code with a copy of the GNU GPL and this license
document.
```
This is blocking performance improvements. This backend is almost
unilaterally unused except for when I'm testing if games run on Radeon
video drivers.
Hopefully AmpereOne and Orin/Grace can fulfill this role when they
launch next year.
This is an atomicFetchCLR, removes two mvn instructions that are back to
back negating the source.
We didn't have this instruction combination in InstCountCI so will be a
bit hard to see.
It is scarcely used today, and like the x86 jit, it is a significant
maintainence burden complicating work on FEXCore and arm64 optimization. Remove
it, bringing us down to 2 backends.
1 down, 1 to go.
Some interpreter scaffolding remains for x87 fallbacks. That is not a problem
here.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.
The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.
To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
clang benchmark and seeing how frequently it was still writing.
- 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
symbols.
- In most cases 100ms is enough that you won't notice the blip in
perf.
One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
We can cut down on a few of the generated moves. For
the case where both destinations alias one another,
we can just calculate the high part instead of both of them.
Allows viewing the codegen for cases where conditional zeroing is performed.
Also fixes up the vperm2f variants shorthanding one of the registers
to make everything a little more explicit.
cmov was quite terrible in its implementation. Some things of note:
- NZCV cache would cause store for no reason
- {16,32}-bit would zero extend sources for no reason
- 16-bit would zero extend result for no reason
A bunch of flag testing is still doing a ubfx plus compare against zero
when it could end up being a tst instead, but this is a step in the
right direction and switches over to explicit sized selects.
If we are going to throw away the updated value of CF anyway there is no point
wasting an instruction to invert CF. Add an IR toggle for that so the arm64 JIT
can make better choices.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Changes the helper which all the source uses to still calculate the size
implicitly. This is going to take a while to convert all implicit uses
over to the explicit operation.
Get us started by at least having the IR operation itself be explicit.
Wide shifts under SVE use 64-bit source elements. If a smaller element
overlaps the 64-bit shift element then it uses that shift
eg:
- Src1[15:0] >> Shift[63:0]
- Src1[31:16] >> Shift[63:0]
- Src1[47:32] >> Shift[63:0]
- Src1[63:48] >> Shift[63:0]
- After this point it will switch to the next 64-bit shift element
- Src1[79:64] >> Shift[127:64]
- Src1[95:80] >> Shift[127:64]
- Src1[111:96] >> Shift[127:64]
- Src1[127:112] >> Shift[127:64]
As seen, we can skip the duplication of the scalar element if the OpSize
is 64-bit, this makes MMX emulation slightly more optimal here.
This also means that a few instructions that weren't claimed to be
optimal actually are since they need the duplication operation (which
vixl always labels as a mov).
Hades and the vcruntime hits this very hard in memmove.
`86.56% [JIT] tid 458574 [.] JIT_0x18000c375_0x7fffc94790c8`
```asm
0x00007fffc94790f8: ldaprb w3, [x2]
0x00007fffc94790fc: stlrb w3, [x1]
0x00007fffc9479100: add x1, x1, #0x1
0x00007fffc9479104: add x2, x2, #0x1
0x00007fffc9479108: sub x0, x0, #0x1
0x00007fffc947910c: cbnz x0, 0x7fffc94790f8
```
This performance is terrible because Cortex's LRCPC performance is bottom-tier.
Work around the performance issue by forcing things to do larger moves with vector moves instead.
The only version of this instruction that was generating optimal code
was the one with 64-bit destination and source.
Optimizes the rest of the operating sizes so that they are all optimal
at one instruction translations
16-bit MOVBE is a bit of a special case where it loads 16-bits in to the
bottom of the GPR without clearing the upper bits of the register.
Which means 32-bits or 64-bits depending on operating mode.
Arm64 doesn't support a 16-bit bswap so it needs to operate at 32-bits
instead. We then can insert the resulting bits of the 32-bit rev with a
bfxil in to the lower bits of the resulting destination register.
This allows 16-bit movbe to be optimal now.
Optimal blendps is worst case 2 instructions.
FEX's RA doesn't quite get there since it can't see through multiple
instructions with SRA destinations. That'll be fixed in the future.
Optimal blendps is always one instruction, one is a no-op.
We always hit this.
Cleans up the code which had special cased some 32-bit optimization
which is unnecessary now that both 8-bit and 16-bit are also optimized.
When FEX does a VExtractToGPR, the result is zero extended to the full
GPR register size. This means we don't need to do a zero extend when
storing to a guest GPR.
Makes pextr{b,w} optimal now.
Needs #3088 merged first.
If CI faults out due to a bug then we would have no log as to which
instruction caused the issue.
I find myself adding this each time an assert fires to see what
instruction it was working on. Just add it directly.
I managed to miss a whole section of instructions from the secondary
opsize tables. This resulted in four instructions missing from the
database.
Adds cmppd, pinsrw, pextrw, and shufpd which are all non-optimal
instruction implementations.
In the case that source registers are sequential then this turns in to a
load of the vector constant (2 instructions) and the single tbl
instruction.
If the registers aren't sequential then the tbl turns in to 2 moves and
then the single tbl, which with zero-cycle rename isn't too bad.
Since this is a worst case option this is significantly better than the
previous implementation doing a bunch of inserts which was always 9
instructions.
We should still strive to implement faster versions without the use of
TBL2 if possible but this makes it less of a concern.
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.
The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
Hits a whole bunch of common cases, most of which then emit optimal code
generation.
Two cases that use VInsElement hit the RA quirk where the SRA
destination is dead but RA doesn't see it, so it ends up doing a couple
moves. If RA gets fixed then those two moves will go away.
There are definitely still cases that we could emit more optimal code.
Additionally we could implement a TBL2 IR operation to do a LUT approach
for ones we don't cover.
Problem with implementing a TBL2 ir operation is that we have no way to
ensure registers are sequential so we would need to always do moves
```asm
ldr v2, <LUT Table>
mov v0, v16
mov v1, v18
tbl v16.16b, { v0.16b, v1.16b }, v2.16b
```
Which to be fair isn't terrible, and if we're lucky that the guest uses
sequential registers we can naturally get the more optimal code path.
Ideally our RA could push some operations in to sequential registers but
that's not possible currently.
I'll do a follow-up PR that implements TBL2.
Move instruction to itself here is a nop.
Need to be careful about AVX operations which use a different handler
since those might actually zero the upper bits on 128-bit move
{Load,Store}RegisterSRA always loads or stores GPRSize. 8-bit and 16-bit
are vestigial and all OpcodeDispatcher usage will load the full GPR size
(32-bit or 64-bit) and then extract or insert as necessary.
This cleans up a few bits of codegen in InstCountCI.
Currently FEX will always jump out of the JIT any time FCW was getting
written to, ensuring that the softfloat state is setup to rounding at
the time of FCW getting written.
This has the unintended side-effect that even in "x87 reduced precision"
mode we were jumping out of the JIT.
This hit a real world use case of an installer reloading FCW after every
x87 operation and generating a block with 2297 instructions.
Instead when jumping out of the JIT for handling x87 operations, load
FCW and pass it as the first argument of the handler. Setting the
softfloat state at that point.
This helps the installer's hottest block by cutting it down to 1477
instructions. 64.3% of the original size. The code block is still
burning 90% of the CPU time of the installer but the performance is
significantly better while it is doing its decompression.
In order to optimize this installer's block of code more then we will
likely need to optimize out x87 stack usage.
This leaves us with two temporary vectors that the JIT can use.
As of last month we stopped using v2 and v3 as temporaries and these can
now be given back to the JIT.
Ensures that the registers are still sequentially ordered and adds
support for spilling the FPR counts that are aligned by 2 instead of 4.
Adds a couple of instructions to filling and spilling but isn't that big
of an issue.
InstcountCI has some ridiculously large changes just because RA is
starting at a new register number.
This wasn't implemented initially for the interpreter and x86 JIT.
This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.
The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
These instructions set all the flags to undefined and moves the
resulting bit in to CF. No need to calculate the deferred flags when
we are about to write over them.
We have supported this since #163 but we haven't been exposing the
feature in hwcap2.
We have exposed it in CPUID this entire time, just not in hwcap2.
Arm64 store with writeback when source register is the same register as
the address is undefined behaviour.
Depending on hardware details this can do a whole bunch of things.
This situation happens when the x86 code does `push rsp` which is quite
common for applications to do. We would then convert this to a `str x8, [x8, #-8]!`
Which results in undefined behaviour.
Now that redundant loads are optimized this showed up as an issue. Adds
a unit test to ensure we don't hit this again.
When the destination overlaps one of the sources we must be careful to
follow a movprfx rule.
```
The destination register must not refer to architectural register state
referenced by any other source operand register of this instruction.
```
We ended up in a situation in the vpmulh{u,}w AVX tests where zm was
overlapping the destination which violated that rule. This also
generated invalid code for this instruction.
```
[INFO] movprfx z6, z4
[INFO] umulh z6.h, p6/m, z6.h, z6.h
```
As seen, we were overwriting one of the sources because the destination
overlapped it. Now instead check if each individual overlap so invalid
code isn't generated.
InstCountCI results aren't affected since this only happens in
situations with multiple instructions.
Now that the RCLSE pass finally optimizes redundant loads again this
optimization that lives in the OpcodeDispatcher can be removed.
With InstCountCI reran, the pblendvb results don't change at all, as
expected.
This is taking steps to start fixing RCLSE which was started by #2700.
Same situation as that PR, since #2170 when we converted
{Load,Store}Context in to {Load,Store}Register we broke this pass
entirely. It hasn't been doing anything for redundant GPRs and FPRs
since at least November of last year.
Technically it was potentially still optimizing redundant MMX
accesses, but it is so broken that it doesn't matter.
Instead of going all in like #2700 did, tear down the pass and start
again. We are now /only/ optimizing redundant context/register loads.
This fixes an issue that comes up commonly where the same register used
as sources was getting loaded twice, causing redundant moves.
`packsswb xmm0, xmm0` for example was generating a four instruction
sequence instead of three instructions because we weren't eliminating
the redundant load.
Going to take reimplementing all the optimizations that this pass does
in steps. This way we can track any regression in the independent steps
unlike what happened in #2700.
Confirmed that Proton/Sonic Mania still works after this.
LogMan::Throw::AFmt(Thread->ThreadManager.GetTID()==FHU::Syscalls::gettid(),"Must be called from owning thread {}, not {}",Thread->ThreadManager.GetTID(),FHU::Syscalls::gettid());
Loaded 100 of 747 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.