Dramatically reduces memory consumption of FEX's per-thread lookup
structures. Primarily because L2 cache entirely goes away which can end
up reaching hundreds of megabytes or over a gigabyte of memory in some
cases, but also because L1 cache dynamically scales based on load.
Useful for conserving memory on systems with less than 16GB of RAM and
are UMA, like Asahi users inside of muvm.
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.
This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.
Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.
We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
Fixes#5230
Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.
Note: I have no idea how fmt will handle that format string with an
array of uint8_t.
Signed-off-by: crueter <crueter@eden-emu.dev>
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.
8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.
Both scalar and vector wide.
Fixes#5326
Presumably this was done as a hack to highlight the line in red when rendering
documentation to markdown. Since it's only used for two instructions, drop this
use to ease generation of C++ docstrings.
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
The current CodeBuffer regrowth code discards any existing contents. This is
undesirable with code caching, since those contents can't be re-fetched from
the disk cache and instead need to be recompiled at runtime.
Additionally, FEXOfflineCompiler obviously should never discard compiled code.
Using a large enough CodeBuffer right away reduces the likelihood that FEX
runs into such scenarios.
When thunks are jumping out, games are jumping /entirely/ out of their
controlled code, which means we don't need to save and restore
NZCV,PF,AF.
Some CPUs don't fully rename direct accesses to this register which adds
up during thunking. FPCR is also in the same situation where it'll force
pipeline flushes and isn't renamed away, but we can't really avoid that.
Improves performance at least in Detroit: Become human where the game
spends ~48% CPU time inside of the thunk trampoline for
`vkUpdateDescriptorSets`.
I plan on a follow-up PR where I converge all these options into a
struct argument instead, but that's a follow-up since I don't want to
burn a bunch of time right now.
Turns out casting a vector of data to a string is a bad idea, who knew?!
This bug was exposed by the rpmalloc PR because the
`/run/host/container-manager` file isn't null terminated. So the cast
fextl::string through std::vector::data had the /potential/ to not be
null terminated.
This was just a bug hiding in the code, convert it over to use
fextl::string directly and avoid the entire dance.
Fixes#5236
Otherwise we hit an assert in FEXCore backend with code discovery
hitting things that look like AVX.
Fixes a crash in Uplay.
Also adds a test to just ensure that the instruction faults out and is
captured instead of crashing in FEX itself.
This breaks relocations currently due to not handling negatives and also
an interesting overwriting problem.
Not that big of a deal, it's only a minor optimization anyway.
Fixes#5227
Previously, FEX would use `$HOME/.fex-emu` for its data and config if
`XDG_CONFIG_HOME` and/or `XDG_DATA_HOME` were unset. This doesn't follow
XDG, so instead we do a fallback to `$HOME/.config` and
`$HOME/.local/share` respectively if XDG env vars are unset. Also,
pre-emptively creates those directories since `~/.local/share` and
`~/.config` existing is technically not a guarantee if XDG dirs are unset.
Signed-off-by: crueter <crueter@eden-emu.dev>