This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.
With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
Misc resident: 54 MiB
JEMalloc resident: 208 MiB
```
So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.
Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.
This tracking will be the first step towards seeing where our memory
usage is going.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
This new variable length pair of integer implementation is taking direct
advantage of the most common aspects of FEX's JIT in that most x86
instructions are <= 8-bytes in length, and the ARM implementations of
those are /usually/ 16 instructions in length or less. Also only
unsigned offsets in this implementation since the common case is forward
incrementing.
This converts a majority of 16-bit vl encodings in to an 8-bit encoding
instead, shaving space off the RIP reconstruction data.
An additional optimization is for the 16-bit pair of integers, we
continue this optimization through but with more bits and changing over
to signed. This captures the second most common cases of /slightly/
larger increments and small loops.
Pairs of 32-bit and 64-bit integers are unoptimized since they are
uncommon.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
This would become an issue when multiple threads are contending with the SIGBUS handler on the same code.
We were failing to mask the VR, OPC, Rm, and Option bits, resulting in a
comparison below always resulting in a false result if another thread
managed to backpatch.
This was just unlikely to be seen on LRCPC2 supporting hardware and
since we fixed `LDSTUNSCALED_MASK` before, this wasn't really getting
seen.
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only
To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.
Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
When #2722 implemented this initially and #4271 switched over to signed
int16_t there was assumptions made that int16_t was a reasonable
trade-off in encoding size versus needing to deal with 8-bit values
being too small in some cases.
In the common case we are almost always encoding 8-bit values because
instructions are typically linear (and less than 15-bytes in size), but
16-bit was chosen because optimizing JIT and multiple instructions that
don't cause exceptions can add up to larger than 8-bit.
Instead of hardcoding 16-bit values, implement a variable length integer
class where ~96.8% of values are 8-bit encoded, and the remaining 3.19% are encoded using 16-bit.
Due to some constraints that #4271 put in place, we can basically
guarantee currently that branch targets are within 16-bit. The VL class
does support 32-bit and 64-bit as well so if we change behaviour then
nothing needs to change.
Some stats when running Sonic Mania with multiblock enabled.
Encoded integers: 3,504,907
Encoded 8-bit: 3,393,095 (96.8%)
Encoded 16-bit: 111,812 (3.19%)
Encoded 32/64-bit: 0
Encoded Size: 3,615,181 bytes (3.44MiB)
Fixed encoded size: 7,007,604 bytes (6.68MiB)
Definitely worth using and saves the headache of large RIP/PC offsets
causing problems.
When multiple threads simultaneously SIGBUS on the same address, one of them
will perform the backpatching while the other will detect the backpatched
instruction sequence and hence report the SIGBUS as "handled".
This typo broke the instruction detection logic: The second thread would
assume the source of the SIGBUS was unrelated to TSO emulation and hence
report the signal as unhandled (generally triggering program abortion).
In practice, this problem did not manifest as FEX does not currently share
CodeBuffers between threads.
Anything less than three pages can't be used for FEX allocations due to
VMA implementation details. Plus we may have reduced a single page
reservation to zero with the prior ObjectAlloc size reservation.
When small regions were being used for VMA allocations (less than 64
pages), this function was truncating the result to zero. Resulting in
incorrect `LiveVMARegion` size calculations. It would calculate that the
FlexBitSet consumes zero bits of space, even though it needs to use at
least 2, or 3 if we actually want to allocate anything from that
LiveVMARegion.
This was noticed in this PR because our VMA region tracking is being
used more, which has a more likely chance to have small VMA regions for
allocating from. Cause a 1page allocation to try and use a 3 page VMA
region for allocation, but failing because the FlexBitSet size wasn't
calculated correctly.
- Page layout:
- [0x0, 0x1000): struct LiveVMARegion
- [0x1000, 0x2000): FlexBitSet<uint64_t> UsedPages
- ^ This space wasn't allocated/mprotected due to the size not
calculating correctly.
- [0x2000, 0x3000): Memory for allocation
When running on a system with a 48-bit VA, if FEX does any allocations
between us reserving the upper 128TB and the application running, then
/technically/ we are intersecting with the application's memory region
in the lower 47-bits.
This didn't typically result in any problems due to how ASLR works, but
if we did any large allocations (like #4291 wants with 128MB VMA region)
then these typically get pushed higher in the VA space.
Again not usually a problem, but if you happen to be running an
application that is using MAP_FIXED with hardcoded addresses then this
can stomp over FEX-Emu memory causing problems.
This is what happens with Wine, it reserves the upper-32MB of its 47-bit
VA space, which is /highly/ likely to stomp on FEX memory. In-fact it
likely occurs all the time, we just got lucky with whatever it was
clobbering wasn't used at the time.
On 39-bit VA systems this isn't a problem because the mmap fails
outright with a warning message from WINE.
Because we are already reserving the upper 128TB of VA space, instead
just always enable our allocator and use the regions that were reserved.
We need to be a little bit careful to ensure we don't accidentally
allocate more memory post-reservation but that just requires a small
adjustment to our unique_ptr and constructor for the 64BitAllocator.
This means /all/ FEX-Emu allocations will be in the upper 128TB VA space
when running 64-bit applications on a 48-bit VA system. Which is kind of
nice.
Fixes WINE in #4291 when the allocator stats are bumped to 128MB per
process.
These parsing strings are tiny, less than 64 bytes all the time. Just
stack allocate the buffer. Makes it safer to use during extenuating
circumstances as well, like SIGBUS and SIGSEGV.
String parsing was adding a newline but then we failed to use it. I
don't think it caused any issues considering how infrequent instant
profiler objects are used.
The immediate offset masking was at the completely wrong offset when I
wrote these handlers. No idea how I managed to mess those up so badly.
Should fix at least some of the issues with #4216
It is not an error that pread returns /less/ than what was requested. In
fact it's very common for the Linux kernel to return less than the data
requested from procfs.
procfs keeps coming back to bite this function, previously it was fstat
returning size of 0 which it hit. Now it only feeds data as much as it
wants per loop. In particular /proc/self/maps would only read ~3k bytes
on my system, but not be complete.
To fully fix the issue, always make sure to keep reading until there is
either an error OR zero is reached!