The instruction decode tables for crc32 introduced some dumb.
`F2h` and `F2h && 66h` prefixes both work for crc32.
This is a failure on Intel's part for sticking crc32 in to the vector
table.
MOVBE without any prefixes also does the same garbage where prefix `66h`
acts as an operand prefix size ONLY.
This table is particularly terrible. CRC32 is the first instruction in
this table that needs either prefix `72h` OR `66h && F2h`
For 8bit CRC32, this ignores the 66h operand size override prefix.
- But our table decoding didn't handle this
For 16bit/32bit/64bit CRC32 this behaviour changes depending on 66h
prefix AND REX.W
- 66h prefix is ignored when REX.W is set, always 64bit but it falls
down the other table path
This is an absolutely weird edge case that nobody should hit, but here
we are.
This was mainly an optimization around memory usage. ALU ops tend to
bloat the IR quite heavily, but I also noticed a 2-4% uplift in
performance of some applications. So a nice side effect.
Should let us more aggressively target reducing our IR intrusive
allocator size since this is quite reduced.
In a pedantic heavy ALU op code block this reduces the number of IR ops
from 14,756 IR ops to 2,016 prior to optimization.
After optimization both had reduced down to 50 IR ops, proving the
output IR was the same.
Dynamically linking xxhash is causing problems with pressure-vessel.
With this in place we only have the typical C++ dependencies
```
$ ldd ./Bin/FEXLoader
linux-vdso.so.1 (0x00007fff44d9d000)
libstdc++.so.6 => /lib/x86_64-linux-gnu/libstdc++.so.6 (0x00007f4c4d884000)
libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6 (0x00007f4c4d7a0000)
libgcc_s.so.1 => /lib/x86_64-linux-gnu/libgcc_s.so.1 (0x00007f4c4d786000)
libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007f4c4d55e000)
/lib64/ld-linux-x86-64.so.2 (0x00007f4c4e0fa000)
```
By default we won't build with the interpeter to reduce user confusion.
The interpreter isn't really useful to end users so remove it.
Completely removes it from building except for the fallback operations.
This also removes the selection from FEXConfig to remove selection
confusion there.
File Stats:
FEXLoader Size with Interpreter: 3422768 bytes
FEXLoader Size without Interpreter: 3301944 bytes
Size difference: 96.4699915%
Bytes removed: 120824 bytes
4k pages removed: 29.498046875 -> 30 rounded up
VM Stats (Reported from bloaty):
Memory Size with Interpreter: 6.50Mi
Memory Size without Interpreter: 6.38Mi
Size difference: 98.1538462%
If an application is forking heavily with threaded file accesses
happening then the mutex can end up in an unknown state.
On fork make sure to lock the mutex then immediately unlock after fork
occurs.
This final step resolves hanging that pressure-vessel hits on startup.
Since it is doing a ton of file opening and forking during
initialization.
Instead of just a basic mutex, also mask the signals.
This fixes the problem where we can end up receiving a signal in the
middle of memory allocation. Thus leaving the locked mutex in a broken
state.
This more closely matches the Linux kernel behaviour.
Since if you're in the middle of a memory allocating syscall, you won't
get signaled.
Currently force disabled until the rest of SSE 4.2 is enabled
This is to remind us in the future that SSE4.2 can only be enabled in
CPUID with CRC32 instruction support.
This isn't correct and breaks games.
This makes the FREM and REM1 implementation the same.
While not 100% correct, it is still better than before.
New issues will be created to handle the differences in the future.
Fixes#1374.
Also fixes most of the HL2 issues, just not the seam issue.
Some of these behaviours have changed now, particularly around signal
handling.
Some things still fail now of course. But most everything is now
documented as to why it is failing or disabled.
Fixes#955
On Set, we have four options that need to be converted.
On Get, we have two options that need to be converted.
This fixes a crash that Tomb Raider 2013 was having on launch.
Makes curl do its continue feature to give the users the best chance of
downloading a rootfs. We don't need to restart the full file transfer on
failure. Helps people with slower connections.
On failure to download, asks the user if they want to retry the download
rather than just exiting with a weird error about hash failure.
Once the image is downloaded, now changes options depending on if
squashfuse or unsquashfs works.
Prevents the user from selecting a bad option and getting unexpected
behaviour. Ideally we would do a squashfs mount test as well for
platforms that don't have working FUSE, like termux. This is harder to
get right and its for an unsupported platform, so I'm not going to
invest more time with it.
Fixes#1525Fixes#1526Fixes#1527
Location to check if curl, squashfuse, and unsquashfs are working.
unsquashfs is a bit more complex where it needs to parse the help output
to see if zstd is supported
In the case of launching without stdout/stderr then redirection could
have these constants be a redirected FD that sits in the same fd number.
Use -2 to indicate no redirection.
Use -1 to indicate closing traditional stderr/stdout
The rest will indicate if stdout and stderr should be replaced as
normal.
Making sure not to close the incoming fds if they matched the
stdout/stderr FD numbers.