This was originally set to zero out of concern for any application that
is doing anti-emulation, anti-cheat, anti-VM checks.
This concern is likely unwarranted, and if any application/game starts
hitting this as a problem then we can throw an application profile at it
instead.
This makes pressure-vessel emulation checking more optimal by it not
having to do a uname dance.
The latest ARMv8.0 toolchain is implementing fetch_add with a bic rather
than an and. Not sure why they started doing this but support the
remaining logical operations in our unaligned atomics handler.
Fixes the Interpreter ARMv8.0 atomic ops.
Fixes#1742
The pause instruction is architecturally defined to be the `REP NOP`
instruction.
This allows people to use this instruction as a backwards compatible
pause without checking for CPUID support. In fact there is no way to
check if the hardware implements this as a `REP NOP` or a `PAUSE`.
If you have new enough hardware then this just ends up being a PAUSE.
Pass this PAUSE over to our host to help out applications that are
writing spin loops with a PAUSE in it, which we were deleting
previously.
This is an IR op that produces nor consumes any SSA values, but has side
effects.
Turns in to the pause instruction on x86 and yield instruction on
AArch64.
Previously if the instruction was encoded to use rsp, rbp, rsi, or rdi
then due to how these were encoded in modrm this would hit the frontend
path for writing to the high 8 bits of a 16bit register.
This is because it's instruction specific if an 8-bit modrm instruction
chooses to use the high 8-bit region or the upper 4 registers.
See the ModRM.reg section of `ModRM.reg and .r/m Field Encodings`
specifically to see what each encoding stands for. Has four different
meanings per encoding depending on instruction.
This instruction doesn't actually write to registers at 8-bit size, it
extracts an element at 8-bit size and then zero extends it to the full
GPR.
When storing to memory it always stores to memory at the size of the
element extracted.
I grepped around the instruction tables to see if there were any other
instances of this mistake. This was the only one.
Fixes#1472 and also gets Psychonauts 2 running.
This does the setup for handling the named region object loading and
closing using the async interface.
This exercises the async interface while the async thread itself only
does the minimum no-op steps required to fake loading and saving.
The no-op interface is hooked up to the point of exercising it in the
most minimal of sense.
If the configuration is set to enable read-only or read/write object
code then it will spin up the async worker thread as well, but it
doesn't do anything yet.
x86 has eight instructions that are non-temporal.
Only one of which is a load-NT.
Adds a memory access type classification to our LoadSource/StoreResult
helpers.
This lets us explicitly choose Default, TSO, NonTSO, and Stream.
Stream currently just behaves like NonTSO so at some point in the future
we can add non-temporal loadstores to the IR.
Main thing is to move these NT accesses to non-TSO.
Documentation claims that these insert the lower 32-bits leaving the
upper bits unaffected.
Hardware testing proves that the upper 32-bits of the base registers are
zero'd.
Additional documentation also concurs that this is the case.
This allows us to have RCPC loadstore operations with a 9-bit signed
offset.
This gives us a small range of [-256,256) of immediate encoding range on
our TSO loadstore operations.
Updates the inline constant pass in ConstProp to support this range on
TSO IR ops if the host supports RCPC2.
Apple M1 supports this extension, didn't test with Cortex-X2/A710.
The JIT currently doesn't use this. This is just the handling code
itself.
One line disabled handling Guest RIP move relocations until the JIT
object cache is enabled.
The JIT currently doesn't use this. This is just the handling code
itself.
One line disabled handling Guest RIP move relocations until the JIT
Object cache is enabled.
It wasn't using acquire semantics, only release semantics.
This was causing the swap to load data from a stale cacheline, causing
the futex system in glibc to break.
This break only occured if you tried going down the RCPC codepath
because of edge case memory ordering problems.
This then enables the RCPC code path now since it works.
In a large number of cases we are moving pointers within a 4GB region
and some marginal pointers that are within 1MB.
This is only used in the case that MOVZ can't be used.
NOP padding still occurs after these instructions to ensure that if they
are being used with relocations it will still get padded to a full 4
instruction length.
Not all hardware fuses these and LLVM claims that Cortex beyond A72 even
doesn't, but it'll still be faster.