Commit Graph
1677 Commits
Author SHA1 Message Date
Ryan Houdek c691d70919 Merge pull request #1467 from Sonicadvance1/fix_really_old_ubuntu
Improves compile ability for older libraries
2021-12-24 18:50:39 -08:00
Ryan Houdek edce981824 Merge pull request #1463 from Sonicadvance1/enable_threads_default
Config: Enables all host threads by default
2021-12-24 18:50:20 -08:00
Ryan Houdek 9bffaeea40 Merge pull request #1462 from Sonicadvance1/sanitize_core_option
Config: Sanitize Core option
2021-12-24 18:50:11 -08:00
Ryan Houdek d708cbad5e Config: Always fixup Threads option
In the case of nothing being set then with the default being zero now we
will not have fixed it up to calculate the number of threads based on
host core count.
2021-12-24 16:32:58 -08:00
Ryan Houdek 2079f6b3c7 Improves compile ability for older libraries
Adds a header only include utility folder that can be included from
everywhere.

Contains syscall helpers for older glibc and defines for older Linux
uapi headers missing some defines.
2021-12-24 15:43:25 -08:00
Ryan Houdek ef7b77dff7 Config: Enables all host threads by default
We've solved the few threading problems that we had before. So now allow
FEX to query the host for the number of threads.
2021-12-24 13:22:01 -08:00
Ryan Houdek 777aadb73e Config: Sanitize Core option
If set to an invalid Core option from the json file then sanitize it
back to JIT.

Otherwise FEX has a chance of just crashing.
2021-12-24 13:21:47 -08:00
Ryan Houdek ccd06e2097 GDBServer: Fixes long string packet encodings
Fixes thread, memory-map, and OS data packet types.
These were attempting to substr when the encode function already handles
that.
Was making it so gdb was only ever receiving the first 1000 bytes of the
data and then decoding incorrectly.
2021-12-24 13:21:35 -08:00
Ryan Houdek 17480c0e2d HostFeatures: Detect if the host CPU suports float exceptions
On x86 this is always supported.
On ARM this is only supported if FPCR writes actually enable the things.
Also detects the AFP feature for flushing input denormals to zero.

These are all part of the x86 MXCSR.
No Cortex supports FPCR exceptions, while Apple M1 CPUs support
Exceptions but not the true "AFP" extension
Apple instead supports some additional flags in their
`SYS_APL_AFPCR_EL0` register for enabling this.
2021-12-17 16:12:31 -08:00
Ryan Houdek 2cee9e5d4b CPUID: Fixes crash on unknown CPU
If the CPU is unknown inside of the ARM CPU detection then the
MIDROption selected could have fallen down a path where it is set to
nullptr.

Resolve this crash by doing a nullptr check.
2021-12-16 21:50:42 -08:00
Ryan Houdek fa6f1b1d90 Merge pull request #1446 from Sonicadvance1/more_fault_reconstruction
Dispatcher: Adds more state reconstruction to state restore
2021-12-15 03:04:09 -08:00
Ryan Houdek e24eb7a72d Merge pull request #1447 from Sonicadvance1/consolidate_host_features
HostFeatures: Consolidates HostFeatures flags
2021-12-15 02:55:42 -08:00
Ryan Houdek fd59fb1a7a OpcodeDispatcher: Implements support for RDTSCP
Have fun
2021-12-15 01:18:34 -08:00
Ryan Houdek 04deeb3911 IR: Adds Processor ID IR op
For the x86-64 JIT this is implemented with pulling rdtscp's result for
this value.
For Interpreter and AArch64 JIT this is implemented with the getcpu
syscall.

Theoretically AArch64 could implement this with MPIDR_EL1 but because
SoC vendors hecked this up, we can't. Thanks.
Kernel just returns zero + reserved bits if you try reading it.
2021-12-15 01:17:24 -08:00
Ryan Houdek 23cb0dea00 HostFeatures: Consolidates HostFeatures flags
Some of these were in the Emitter class and some were in the
HostFeatures.

Merge these together since in the future I'm going to be using all of
this data as a key for our AOT code cache.
2021-12-14 02:41:42 -08:00
Ryan Houdek bcad7a9eea Dispatcher: Adds more state reconstruction to state restore
Still only setting the new state if RIP is affected for now.
Noticed a bug where we weren't setting our frame RIP to the new RIP on
32-bit.
Decided to walk through more of the state setting while fixing that.

This gets #1214 further but then it eventually crashes with a read to
0x11.
2021-12-14 02:36:24 -08:00
lioncash 06e4a5a5b7 CPUID: Signify full support for BMI2
With PDEP support dropped in, we now support all of BMI2, so we can
signify that we support it in our emulated CPUID.
2021-12-13 14:07:41 -05:00
lioncash 6ff80670b3 OpcodeDispatcher: Handle PDEP
Now all of BMI2 is handled.
2021-12-13 14:06:55 -05:00
lioncash 38eea80b8d IR: Add PDep IR opcode 2021-12-13 13:54:15 -05:00
Ryan Houdek 9394e49c95 Merge pull request #1431 from Sonicadvance1/expose_arm_names
CPUID Expose Hybrid flag and CPU names
2021-12-12 18:07:53 -08:00
Ryan Houdek 91984003b6 Merge pull request #1440 from lioncash/pext
OpcodeDispatcher: Handle PEXT
2021-12-10 16:27:49 -08:00
lioncash dbf571fdfb OpcodeDispatcher: Handle PEXT 2021-12-10 19:03:58 -05:00
lioncash b6abcc5e3c IR: Add PExt opcode 2021-12-10 19:03:54 -05:00
Tony Wasserka a7c0997daf OpcodeDispatcher: Mark const-initialized instruction tables as constexpr
This allows the compiler back these tables into the executable, which
reduces the amount of work the function has to do at runtime.

Reduces the runtime of this function by 30% relative to the previous commit.
2021-12-10 13:25:24 +01:00
Tony Wasserka 6eba3f331e OpcodeDispatcher: Store instruction tables as arrays instead of std::vector
Reduces the runtime of this function by 90%.
2021-12-10 13:25:23 +01:00
Tony Wasserka bcc75e3312 X86Tables: Use stack-allocated arrays instead of std::vector 2021-12-10 13:25:23 +01:00
Tony Wasserka 30672e1517 X86Tables: Remove unneeded initialization code from the Debug mode path
The tables have recently been changed to be zeroed out as a whole on startup.

Speeds up InstallDebugInfo by about two orders of magnitude and reduces Debug
executable size by 1.8%.
2021-12-10 13:25:23 +01:00
Ryan Houdek d4655fbb17 Merge pull request #1424 from Sonicadvance1/aot_code_movement
FEXCore: Reorganizes some AOT related code
2021-12-09 13:55:50 -08:00
Ryan Houdek 5f0dfcd715 Merge pull request #1422 from Sonicadvance1/implement_clzero
Implements CLZero instruction
2021-12-09 13:55:31 -08:00
Ryan Houdek 9d43904792 Merge pull request #1436 from lioncash/context-const
Context: Take some arguments as pointer-to-const
2021-12-09 12:23:34 -08:00
Ryan Houdek d74cf6d8d8 CPUID Expose Hybrid flag and CPU names
Had some idle time so I implemented this logic.

We do some tricky logic to have a big.little configuration even with
unknown CPU core types. Promoting or demoting a single MIDR depending on
if we have a mixed configuration or not.

In a non-hybrid design we only claim product names inside the CPUID
product string.

This will appear if you `/proc/cpuinfo` or read the CPUID registers
directly

eg on Snapdragon 888:
processor       : 0
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 1
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 2
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 3
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 4
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 5
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 6
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 7
model name      : FEX-2112-1-g13b14b85            Cortex-X1

eg on Macbook Pro VM which can't see the CPU type:
processor       : 0
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 1
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 2
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 3
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 4
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 5
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 6
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 7
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
2021-12-09 11:39:52 -08:00
Ryan Houdek 92cb9d477e Arm64: Fixes vixl assertions around ubfm usage
vixl has an assert check to ensure the register sizes are the same for
ubfm.

32-bit Inline syscalls hit this for both arguments and return value.
VCastFromGPR would have hit this but there aren't any x86 instructions
that move 8-bit and 16-bit values in to a vector register.
2021-12-09 11:34:20 -08:00
lioncash 7ed6007252 Context: std::move functions in signal registration functions
Prevents potential reallocations, given we're using std::function here.
2021-12-09 14:31:47 -05:00
lioncash c73991f467 Context: Take some arguments as pointer-to-const
Several API functions act as state querying functions. These can take
some parameters by const to communicate that we don't intend to modify
the respective passed in instance.
2021-12-09 14:26:35 -05:00
Ryan Houdek 13087f8425 Merge pull request #1426 from Sonicadvance1/fcmp_unittests
Arm64: Fixes MapSelectCC FGT flag
2021-12-08 05:26:15 -08:00
Ryan Houdek f2de640395 Arm64: Fixes MapSelectCC FGT flag
As a continuation to #1404, this was mapped to the incorrect Arm64 flags
2021-12-07 18:43:50 -08:00
Ryan Houdek 9a642158e0 Workaround fmt not handling nullptr strings
Simple ternary check each time a name is used
2021-12-07 12:23:56 -08:00
Ryan Houdek 6404aba6e2 X86Tables: Build Unknown op definition tables at compile time
It doesn't make any sense anymore to have specific instruction names
set for UND versus a nullptr string anymore.

Was useful when we could use it to determine the difference between
undefined from the start versus set in the tables but with unknown
decoding. Which is an edge case.

Now instead just zero initialize the data, which means it is an unknown
type and nullptr name. Which works for use.

Improves initialization time of the InitializeInfoTables function from
423 microseconds to 37 microseconds.
2021-12-06 18:57:02 -08:00
Ryan Houdek 45f919683a FEXCore: Reorganizes some AOT related code
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.

This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.

Two minor behaviour changes that got mixed up with this change.

The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.

Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
2021-12-06 18:42:21 -08:00
Ryan Houdek d6e4da7e77 CPUID: Exposes support for CLZero
Only exposed if the host if the HostFeatures for ARM claim to support
it.
2021-12-03 21:20:40 -08:00
Ryan Houdek 18b223811a OpcodeDispatcher: Implements CLZero instruction
This instruction zeroes a cacheline in memory that is weakly ordered and
non-temporal.

It uses the RAX register for where in memory to clear and aligns the
address on cacheline regardless of actual alignment.
2021-12-03 21:15:51 -08:00
Ryan Houdek e1d21f0bff IR: Implements new CacheLineZero op
This zeroes out an emulated 64byte cacheline. Writing zeros to memory.
This very specifically is only 64bytes to match x86 behaviour.
Also specifically non-temporal and weakly ordered. Which matches x86
CLZero behaviour.
2021-12-03 21:14:02 -08:00
Ryan Houdek d2783f2edd HostFeatures: Adds new host feature flag for CLZero
If the DCZID block size matches the CLZero cacheline size. Then claim we
support CLZero here.
Otherwise we don't want to expose support for it.
2021-12-03 21:12:52 -08:00
lioncash cdab71e612 RegisterAllocationPass: Resolve sign comparison mismatch in CalculateNodeInterference()
GetSSACount() returns a size_t, rather than a signed value.
2021-12-02 18:13:01 -05:00
Ryan Houdek 8ef0278338 Merge pull request #1418 from lioncash/branch
JIT: Eliminate redundant jump target map lookups
2021-12-02 14:45:44 -08:00
lioncash 98f54eab9c Arm64/BranchOps: Mark MapBranchCC as static
This isn't used outside of this translation unit, so we can make that
explicit.
2021-12-02 17:18:44 -05:00
Ryan Houdek ffee22d3fd Merge pull request #1417 from lioncash/cpuid
CPUID: Call handler functions directly
2021-12-02 14:14:20 -08:00
lioncash e648e90c2e JIT: Eliminate redundant jump target map lookups
try_emplace() will return the existing entry in the map if it
exists already, so we don't need to perform a find() and then
try_emplace(). We can just use try_emplace() by itself.
2021-12-02 17:03:04 -05:00
Ryan Houdek b93871ff55 Merge pull request #1416 from lioncash/strong
IR: Convert NodeID into a strong type
2021-12-02 13:39:24 -08:00
lioncash 352f0a1133 CPUID: Call handler functions directly
Given the calls are all internal, we can store the function pointers
directly and call them with the this pointer instead of indirecting
through std::bind and std::function.
2021-12-02 16:34:43 -05:00