This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.
With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
Misc resident: 54 MiB
JEMalloc resident: 208 MiB
```
So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.
Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.
This tracking will be the first step towards seeing where our memory
usage is going.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
This new variable length pair of integer implementation is taking direct
advantage of the most common aspects of FEX's JIT in that most x86
instructions are <= 8-bytes in length, and the ARM implementations of
those are /usually/ 16 instructions in length or less. Also only
unsigned offsets in this implementation since the common case is forward
incrementing.
This converts a majority of 16-bit vl encodings in to an 8-bit encoding
instead, shaving space off the RIP reconstruction data.
An additional optimization is for the 16-bit pair of integers, we
continue this optimization through but with more bits and changing over
to signed. This captures the second most common cases of /slightly/
larger increments and small loops.
Pairs of 32-bit and 64-bit integers are unoptimized since they are
uncommon.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
This would become an issue when multiple threads are contending with the SIGBUS handler on the same code.
We were failing to mask the VR, OPC, Rm, and Option bits, resulting in a
comparison below always resulting in a false result if another thread
managed to backpatch.
This was just unlikely to be seen on LRCPC2 supporting hardware and
since we fixed `LDSTUNSCALED_MASK` before, this wasn't really getting
seen.
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only
To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.
Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
When #2722 implemented this initially and #4271 switched over to signed
int16_t there was assumptions made that int16_t was a reasonable
trade-off in encoding size versus needing to deal with 8-bit values
being too small in some cases.
In the common case we are almost always encoding 8-bit values because
instructions are typically linear (and less than 15-bytes in size), but
16-bit was chosen because optimizing JIT and multiple instructions that
don't cause exceptions can add up to larger than 8-bit.
Instead of hardcoding 16-bit values, implement a variable length integer
class where ~96.8% of values are 8-bit encoded, and the remaining 3.19% are encoded using 16-bit.
Due to some constraints that #4271 put in place, we can basically
guarantee currently that branch targets are within 16-bit. The VL class
does support 32-bit and 64-bit as well so if we change behaviour then
nothing needs to change.
Some stats when running Sonic Mania with multiblock enabled.
Encoded integers: 3,504,907
Encoded 8-bit: 3,393,095 (96.8%)
Encoded 16-bit: 111,812 (3.19%)
Encoded 32/64-bit: 0
Encoded Size: 3,615,181 bytes (3.44MiB)
Fixed encoded size: 7,007,604 bytes (6.68MiB)
Definitely worth using and saves the headache of large RIP/PC offsets
causing problems.
When multiple threads simultaneously SIGBUS on the same address, one of them
will perform the backpatching while the other will detect the backpatched
instruction sequence and hence report the SIGBUS as "handled".
This typo broke the instruction detection logic: The second thread would
assume the source of the SIGBUS was unrelated to TSO emulation and hence
report the signal as unhandled (generally triggering program abortion).
In practice, this problem did not manifest as FEX does not currently share
CodeBuffers between threads.
Anything less than three pages can't be used for FEX allocations due to
VMA implementation details. Plus we may have reduced a single page
reservation to zero with the prior ObjectAlloc size reservation.
When small regions were being used for VMA allocations (less than 64
pages), this function was truncating the result to zero. Resulting in
incorrect `LiveVMARegion` size calculations. It would calculate that the
FlexBitSet consumes zero bits of space, even though it needs to use at
least 2, or 3 if we actually want to allocate anything from that
LiveVMARegion.
This was noticed in this PR because our VMA region tracking is being
used more, which has a more likely chance to have small VMA regions for
allocating from. Cause a 1page allocation to try and use a 3 page VMA
region for allocation, but failing because the FlexBitSet size wasn't
calculated correctly.
- Page layout:
- [0x0, 0x1000): struct LiveVMARegion
- [0x1000, 0x2000): FlexBitSet<uint64_t> UsedPages
- ^ This space wasn't allocated/mprotected due to the size not
calculating correctly.
- [0x2000, 0x3000): Memory for allocation