Currently FEX doesn't properly support partial decoded instructions,
which behave slightly differently than full noexec or invalid
instruction decodings. Before this commit we didn't even have a way to
detect the difference.
Primary difference is that the faulting RIP is the beginning of
instruction decode, while the fault address is the first byte that
couldn't be fetched due to memory permissions. This shows up as a
difference between the RIP in mcontext and si_addr in siginfo in the
Linux signal handler.
Right now just change the log so we can determine if we need to support
this edge case.
This saves a whole bunch of memory. Cutting `Just Cause 2`'s title
screen from 1132MB anonymous FEX memory down to 438MB. 629MB in L2
alone.
L2 is primarily a means to reduce overhead in map queries, so it's all
about performance. But because it consumes a lot of people it's kind of
hard.
One idea is that the L2 lookups can be moved to shared data structures,
since we already pull the shared lock when doing an L2 lookup this is
already halfway there.
Side note, we're using unique locks even with read-only code paths
which we can't use the shared lock because this terrible recursive
mutex!
Instead of outright changing L2 behaviour and potentially wrecking
havoc, add a config option for now so testing can happen over time.
before:
```
Total FEX Anon memory resident: 1132 mB
JIT resident: 60 mB
OpDispatcher resident: 97 mB
Frontend resident: 37 mB
CPUBackend resident: 500 kB
Lookup cache resident: 629 mB
Lookup L1 cache resident: 108 mB
ThreadStates resident: 436 kB
```
after:
```
Total FEX Anon memory resident: 438 mB
JIT resident: 62 mB
OpDispatcher resident: 56 mB
Frontend resident: 22 mB
CPUBackend resident: 496 kB
Lookup cache resident: 0 (null)
Lookup L1 cache resident: 109 mB
ThreadStates resident: 436 kB
```
Trivial fix, if a symlink gets deleted between checking if it is a
symlink versus getting a path of it then it resulted in a crash.
Happened periodically for me.
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.
The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.
Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.
I didn't do that exercise since that can be followed up in a subsequent
PR.
Use getuid() instead of geteuid() when determining FEXServer socket names.
Ensures setuid binaries (like chrome-sandbox from Discord) connect to their parent
user's FEXServer instance instead of trying to spawn a separate server.
Fixes connection errors when running applications that spawn setuid children.
Fix a typo in unittests/ASM/Secondary/09_XX_07.asm which caused it to
jump back to the test_32bit label instead of test_64bit.
On Arm systems with FEAT_RNG support, a significant amount of time may
be required before successive uses of RNDRRS. This is because it returns
a random number with fresh full entropy, and it can take a while to
collect the new entropy. That time may be hundreds or thousands of
instructions, so by jumping back to test_32bit the 64-bit test will
alway fail because an RNDRRS has been executed too recently.
Signed-off-by: Rebecca Cran <rebecca@bsdio.com>
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.
With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
Misc resident: 54 MiB
JEMalloc resident: 208 MiB
```
So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
These don't require being bound to class state directly, and so the list
management can be completely opaque to the outside (also means less
rebuilding if these change)
With the gather overflow fixes in place, I've been having this running
for a while. Now that we just kicked out a release, enable AVX even on
32-bit.
We'll need to eventually create a list of games that explode with AVX
enabled, but that same list would match what happens on real x86 hosts,
so there can be some collaboration there.
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.
Correct the implementation to use the correct register size for FMA.