Noticed while taking a look at the relocations that we were technically
not doing alignment before writing down code size.
- Make sure Align16B isn't used with unaligned code with assert
- Switch an `Align` over to `Align(16)` to force 16-byte alignment
- Without NOP insertion, as this is data at this point, so just zeros.
- Record data size after that alignment
- Remove the `Align16B` that occurred afterwards
- Previous query between alignments would leave us with up to 12 bytes
unaccounted for.
- Ensure everything is using the correct sizes by not querying again
- Ensure that emission buffer abuse can't happen by zeroing the buffer.
This needs to be enabled when *generating* caches, not at runtime when we're
loading them (unless we're compiling for validation).
Previous code would incorrectly disable NOP padding in FEXOfflineCompiler and
instead enable it at runtime when it wasn't needed.
When thunks are jumping out, games are jumping /entirely/ out of their
controlled code, which means we don't need to save and restore
NZCV,PF,AF.
Some CPUs don't fully rename direct accesses to this register which adds
up during thunking. FPCR is also in the same situation where it'll force
pipeline flushes and isn't renamed away, but we can't really avoid that.
Improves performance at least in Detroit: Become human where the game
spends ~48% CPU time inside of the thunk trampoline for
`vkUpdateDescriptorSets`.
I plan on a follow-up PR where I converge all these options into a
struct argument instead, but that's a follow-up since I don't want to
burn a bunch of time right now.
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.
The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
Recent CPUs do nop fusion with the following instruction, this gives the
CPU the best chance to do fusion with something that actually does work.
Very trivial, doesn't do this for the more complex handling below these
as counting the number of moves before nop emitting is messy.
Instead of just a trivial pad being on or off, support a tri-state
on/off/auto where on will always pad, off will never pad, and auto will
pad only if code caching is enabled.
Further augment this by allowing a byte-width to be passed in, which can
be used with pointers to force a 48-bit VA width to only ever pad to
three instructions, reducing the common worst-case situation from 4
instructions to 3. This works because we're not going to expose a VA
width larger than 47-bit to the guest.
Fixes the handful of use-cases that explicitly chose their NOP padding,
and a bug in Arm64Relocations.cpp where it was incorrectly asking to not
receive padding even though it requires it.
When loading code caches, these constants get patched up for the new guest
address. The new value may be larger than the original, so the padding bytes
ensure the maximum of 16 bytes of encoding space is always available.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
The host stack pointer can't be reused since explicit bounds checks
would be far too expensive, and on-stack signals prevent implicit ones
using guard pages from working.
TMP4 was used before we passed in a tmp register. Now use that temp
register.
Also return the amount of stack used on the push function. This will be
used in a bit.
In the Push/Pop CalleeSavedRegisters these vectors were getting created
on the heap, allocating memory and then just iterating them.
Just use a std::array which makes it stop allocating memory and saves
the number of instructions.
stop doing weird special cases. just dump all the regs except what aapcs64 says
we don't have to.
this fixes saving x18 across thunks and things. so probably fixes things *cry*
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
* x18 wasn't getting spilled even though it was supposed to be.
* arm64ec preserve_all definitions were all messed up, specialize these to fix.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>