Currently FEX doesn't properly support partial decoded instructions,
which behave slightly differently than full noexec or invalid
instruction decodings. Before this commit we didn't even have a way to
detect the difference.
Primary difference is that the faulting RIP is the beginning of
instruction decode, while the fault address is the first byte that
couldn't be fetched due to memory permissions. This shows up as a
difference between the RIP in mcontext and si_addr in siginfo in the
Linux signal handler.
Right now just change the log so we can determine if we need to support
this edge case.
This saves a whole bunch of memory. Cutting `Just Cause 2`'s title
screen from 1132MB anonymous FEX memory down to 438MB. 629MB in L2
alone.
L2 is primarily a means to reduce overhead in map queries, so it's all
about performance. But because it consumes a lot of people it's kind of
hard.
One idea is that the L2 lookups can be moved to shared data structures,
since we already pull the shared lock when doing an L2 lookup this is
already halfway there.
Side note, we're using unique locks even with read-only code paths
which we can't use the shared lock because this terrible recursive
mutex!
Instead of outright changing L2 behaviour and potentially wrecking
havoc, add a config option for now so testing can happen over time.
before:
```
Total FEX Anon memory resident: 1132 mB
JIT resident: 60 mB
OpDispatcher resident: 97 mB
Frontend resident: 37 mB
CPUBackend resident: 500 kB
Lookup cache resident: 629 mB
Lookup L1 cache resident: 108 mB
ThreadStates resident: 436 kB
```
after:
```
Total FEX Anon memory resident: 438 mB
JIT resident: 62 mB
OpDispatcher resident: 56 mB
Frontend resident: 22 mB
CPUBackend resident: 496 kB
Lookup cache resident: 0 (null)
Lookup L1 cache resident: 109 mB
ThreadStates resident: 436 kB
```
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.
The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.
Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.
I didn't do that exercise since that can be followed up in a subsequent
PR.
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.
With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
Misc resident: 54 MiB
JEMalloc resident: 208 MiB
```
So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.
Correct the implementation to use the correct register size for FMA.
Turns out a simpler way was added to the docs at some point and I never
noticed.
Before:
text data bss dec hex filename
4159895 1471360 4336824 9968079 9819cf Bin/FEX
After:
text data bss dec hex filename
4157159 1471360 4336824 9965343 980f1f Bin/FEX
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.
Before:
text data bss dec hex filename
4160559 1471360 4336824 9968743 981c67 Bin/FEX
After:
text data bss dec hex filename
4159927 1471360 4336824 9968111 9819ef Bin/FEX
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.
Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.
Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.
This tracking will be the first step towards seeing where our memory
usage is going.