I forgot on Linux by default we didn't have the syscall instructions
count as block end. Change this so that it counts as block end now.
This has the additional benefit that now the frontend needs to modify
the RIP manually as well which is fine as it's what arm64ec and wow64
does.
Also add back the UnimplementedOp in RDPID that accidentally got caught
up. Also increment DiskCache version as both changes will change
codegen.
Fixes#5942
CHPE_V2_CPU_AREA_INFO is the mechanism used on ARM64EC to coordinate
thread state, suspension and exception handling with the kernel.
It's typically set up by ntdll for ARM64EC processes, but Wine uses
the TEB pointer without restricting where it's coming from. Set
ChpeV2CpuAreaInfo in FEX WOW64 thread initialization.
Fixes 32-bit debug events and a number of CPU context handling tests.
Reduces the number of wineserver round-trips required for suspend and
avoids the need for remote thread creation for inter-process suspends.
We would check the FileID of mapped sections, but BSS is an anonymous mapping.
Grab the ELF image extents when we parse the file, and add an additional check
to the relocation filter to bail out additional relocs if we know that size.
With PR #5863, the bug that was breaking rendering in d3dx9 and Fallout
New Vegas is now resolved. This means we can now enable the feature by
default again as all known bugs are resolved.
With 3DNow! and full x87 softfloat enabled in FO:NV inside of the starting house. The game is
running at around 40-43FPS, doing ~22.6 to 34 million soft float
operations per second.
With 3DNow! disabled the game is running 30-31FPS, doing ~24-50 million
softfloat operations per second.
In both cases if x87 reduced precision is enabled then the game is
running at around 108-113FPS with zero softfloat instances getting
counted. Can't see the performance difference at that speed.
The LookupKey based on data available at Lookup time gives a list of possible
candidate entries, which may have different guest code sizes/footprints.
If only one candidate, it's stored in-line in the map like before - if we grow
past that, an additional (multi)map is allocated, sorted by footprint.
The footprint sorting lets us reduce the amount of hashes performed at Lookup
to the strict minimum.
Anon entries keys are based on a limited size decoded prefix, so this is an
important part of geting the best hit rate possible out of all the other work
with anon keys, as there's often multiple candidates with the same prefix.
Cap the max amount of entries per bucket to limit growth.
This allows WTF to catch the allocations just like on Linux. Punch our
unixlib path all the way through to rpmalloc so it gets named and
tracked properly.
Otherwise GuestSize ends up different across compiles of the same valid code
extents, including the copy of it baked in JITCodeTail of the host code itself.
JIT invalidation around syscalls and kernel calls is mighty fickle.
This is why the previous code path was born. In order to properly handle
this case we need strong coordination between Wine and FEX around
invalidating code and memory protections. This doesn't quite exist today
so we're kind of stuck with a kludge solution. Instead of forcing the
Persona 5 code invalidation on to every process, only do it on P5R.
This fixes a hang in msiexec with PhysX trying to do a blocking read and
jitting code or creating threads, while also maintaining the P5R
approach. The full comment is in the file about the reasoning.