While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.
This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.
Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.
We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
Fixes#5230
Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.
Note: I have no idea how fmt will handle that format string with an
array of uint8_t.
Signed-off-by: crueter <crueter@eden-emu.dev>
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.
8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.
Both scalar and vector wide.
Fixes#5326
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
The current CodeBuffer regrowth code discards any existing contents. This is
undesirable with code caching, since those contents can't be re-fetched from
the disk cache and instead need to be recompiled at runtime.
Additionally, FEXOfflineCompiler obviously should never discard compiled code.
Using a large enough CodeBuffer right away reduces the likelihood that FEX
runs into such scenarios.
When thunks are jumping out, games are jumping /entirely/ out of their
controlled code, which means we don't need to save and restore
NZCV,PF,AF.
Some CPUs don't fully rename direct accesses to this register which adds
up during thunking. FPCR is also in the same situation where it'll force
pipeline flushes and isn't renamed away, but we can't really avoid that.
Improves performance at least in Detroit: Become human where the game
spends ~48% CPU time inside of the thunk trampoline for
`vkUpdateDescriptorSets`.
I plan on a follow-up PR where I converge all these options into a
struct argument instead, but that's a follow-up since I don't want to
burn a bunch of time right now.
Otherwise we hit an assert in FEXCore backend with code discovery
hitting things that look like AVX.
Fixes a crash in Uplay.
Also adds a test to just ensure that the instruction faults out and is
captured instead of crashing in FEX itself.
This breaks relocations currently due to not handling negatives and also
an interesting overwriting problem.
Not that big of a deal, it's only a minor optimization anyway.
Fixes#5227
Integrate Zydis as an optional dependency to enable x86/x86-64 guest
instruction disassembly during JIT compilation.
Build with -DENABLE_ZYDIS=TRUE.
Use FEX_X86DISASSEMBLE=1 at runtime to output guest x86 instructions
for each compiled block.
There's no longer a distinction between AArch64 and x86 and everything
effectively falls under "Common" now. This means flattening the entire
structure just cleans it up.
NFC. (Although instcountCI will update because of a couple pointer
offsets changing)
Fixes crash in thunks that use callbacks, introduced in #5148.
The dispatcher would call the syscallhandler to get the VDSO thunk
callback. But due to reordering initialization, the VDSO thunk would
have not been loaded at that point. This would cause thunks that use
callbacks to crash with a nullptr exception.
Instead, defer the thunk callback pointer loading until the thread
starts executing, and load the pointer in to our thread state's pointer
struct instead.
Didn't get caught in my initial test sweep since I didn't run a Wine
game with thunks.
This wasn't handling negatives correctly which was causing xalia.exe to
assert. Disable for now rather than further changing logic, with a TODO that it
should get fixed in the future.
Opcode handlers are written with the assumption that LoadSource will not
touch flags and this would be an annoying assumption to change. As this
is such an edge case anyway just don't defer flags and force a load of
the saved value before _TelemetrySetValue (which are implicitly saved
before it).
Fixes the following snippet in upc.exe:
AND word ptr [ESP + ECX*0x1 + 0x80000000],DX
BTR CX,DX
ADC CX,word ptr SS:[EAX + ECX*0x1 + 0x80000000]