Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.
This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.
Based on top of #5362 so the THP disable controls are in.
Add support for question mark and plus mark in regex, supply testing for star
Added more characters to the regex alphabets, add more test case
Added support for regex matching of configs, awaiting reviews
Rename variable to CamelCase
Addresses PR reviews
Remove unnecessary features and test cases
Rewrite to naive regex with dp
Addresses PR reviews
Build fixes
Code Review
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.
With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
Disables THP on some key locations that are fairly sparse
- rpmalloc
- This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
- These get in the hundreds of megabytes, while not being sparse they
trend towards only using a handful of pages and ballooning to 2MB
per thread is quite heavy.
- Lookup cache
- L1 specifically gets hit here which adds a decent chunk of overhead
due to sparsity.
Win32 for all of these also aren't handled, but that will need to be a
followup.
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.
We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
Fixes#5230
Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.
Note: I have no idea how fmt will handle that format string with an
array of uint8_t.
Signed-off-by: crueter <crueter@eden-emu.dev>
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.
In Bayonetta's title screen it went from 963MB down to 834MB resident.
ankerl::unordered_dense is faster on average and has less memory usage
than tsl::robin_map. It is pretty significantly faster than std but
we'll keep that as is for now.
Obviously, this will need a lot of testing.
Signed-off-by: crueter <crueter@eden-emu.dev>
This is fundamentally a frontend only problem, and also Linux only.
Moves it to the frontend where it belongs.
There's likely more things in Allocator.cpp that can be moved to the
frontend but this is the first thing.
NFC
There's no longer a distinction between AArch64 and x86 and everything
effectively falls under "Common" now. This means flattening the entire
structure just cleans it up.
NFC. (Although instcountCI will update because of a couple pointer
offsets changing)
Fixes crash in thunks that use callbacks, introduced in #5148.
The dispatcher would call the syscallhandler to get the VDSO thunk
callback. But due to reordering initialization, the VDSO thunk would
have not been loaded at that point. This would cause thunks that use
callbacks to crash with a nullptr exception.
Instead, defer the thunk callback pointer loading until the thread
starts executing, and load the pointer in to our thread state's pointer
struct instead.
Didn't get caught in my initial test sweep since I didn't run a Wine
game with thunks.
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>