Compare commits

...
200 Commits
Author SHA1 Message Date
Ryan Houdek affbb40cc1 Docs: Update for release FEX-2201 2022-01-03 19:08:53 -08:00
Ryan Houdek b276750c8a Merge pull request #1488 from Sonicadvance1/python_auto_install
Scripts: Adds a python script that can hand hold a user through FEX install
2022-01-03 19:07:41 -08:00
Ryan Houdek 6d5322b349 Readme: Updates with links to the wiki's quick start guide and python command
Quick start guide on the wiki is currently not populated but will be
soon.
2022-01-03 16:05:05 -08:00
Ryan Houdek f836ff0897 Scripts: Adds a python script that can hand hold a user through FEX install
This is definitely a bit divisive but overall this is a win.
This is a user pattern that is emerging in a bunch of projects.
Allow an officially sourced script that lets you pipe a script directly
in to python/bash and setup the environment entirely.

This only supports Ubuntu {20.04, 21.04, 21.10, 22.04} which matches
exactly what we expose in the PPA.

Once this is in the repo, and our PPA is updated to the latest release
tag you can run this script like:
`curl --silent <Direct Github raw link> | python3`

Once the PPA is updated, the README and Wiki will be updated with Quick
Start guides to use this path.

The script steps
1) Checks if ARMv8, or fail
2) Checks if supported Ubuntu, or fail
3) Checks if PPA installed
3a) Install PPA if not, or fail
4) Check if packages are installed
4a) Install non-installed packages, or fail
5) Check if RootFS is configured/exists
5a) Run through FEXRootFSFetcher to get/setup rootfs, or fail
6) Attempt to run emulated uname -a through FEX, or fail
7) Provide some examples for how to use FEX
8) Exit with success!
2022-01-03 14:20:09 -08:00
Mai M 0bf577654d Merge pull request #1491 from Sonicadvance1/fix_signal_pause
FEXCore: Don't use Run() inside RunUntilExit()
2022-01-03 06:56:43 -05:00
Mai M ac63a2bac4 Merge pull request #1493 from Sonicadvance1/fix_pidfd_send_signal32
Linux: Fixes pidfd_send_signal for 32-bit
2022-01-03 06:37:13 -05:00
Ryan Houdek 6b8e6c18ee SignalDelegator: Increase signal stack size
Just to help in the case that a program will dequeue a bucketload of
signals in one go. Basic signal stack size wasn't enough.
2022-01-03 03:29:35 -08:00
Ryan Houdek 93dccffc81 unittests: Updates posix tests with new passing tests 2022-01-03 03:29:35 -08:00
Ryan Houdek 22810208b4 FEXCore: Don't use Run() inside RunUntilExit()
Run() is expected to be used after a pause at this point.
Make RunUntilExit() notify using the StartRunning variable entirely.

This is because if we used Run() then the initial thread hasn't yet set
up its core or signal handlers. So it'll ignore SIG63. Then its reason
for pause is still set to Resume.
Once an application tries sending SIG63, we try restoring the context
from a paused state that does exist and crash
2022-01-03 03:29:35 -08:00
Mai M 154468c172 Merge pull request #1492 from Sonicadvance1/disable_unsync_log_message
Dispatcher: Disables unsync context message
2022-01-03 06:26:08 -05:00
Mai M 0e45b98fb2 Merge pull request #1490 from Sonicadvance1/fix_cpuid_memcpy
CPUID: Fixes another memory overflow issue
2022-01-03 06:24:47 -05:00
Ryan Houdek 2d573f7bf9 Linux: Fixes pidfd_send_signal for 32-bit
Due to the way this syscall handles siginfo_t, we only need to shift
data through the padding.

This is due to the syscall not allowing you to send siginfo_t with
si_code being positive. Which is what the kernel uses to interpret the
data if it is reinterpreting it.
2022-01-03 03:19:12 -08:00
Ryan Houdek 01587270d9 Dispatcher: Disables unsync context message
This message can actually cause a perfectly running application to crash
due to signaling at the wrong time.

Instead constexpr disable it so if you want to see it. You can still
enable it
2022-01-03 03:02:33 -08:00
Ryan Houdek 0a63a15287 CPUID: Fixes another memory overflow issue
Just like the previous one, but this time when we are NOT on a release
2022-01-03 02:56:48 -08:00
Mai M a65e7a09e7 Merge pull request #1489 from Sonicadvance1/release_process_documentation
Docs: Adds documentation about the FEX monthly release process
2022-01-03 01:01:38 -05:00
Ryan Houdek 7e5c561849 Docs: Adds documentation about the FEX monthly release process
Walks through all the steps required for a monthly release.
2022-01-02 21:13:54 -08:00
Mai M a45bc7598c Merge pull request #1487 from Sonicadvance1/rootfs_stop_trying_so_hard
RootFS: Stop trying to retry rootfs after five times
2022-01-01 03:40:40 -05:00
Ryan Houdek 9f8071ce8f RootFS: Stop trying to retry rootfs after five times
In most cases this will be fixed on the first retry, but give it a
chance at five.
Otherwise you have a chance to infinite loop.
2021-12-31 20:04:57 -08:00
Mai M 0808599813 Merge pull request #1486 from Sonicadvance1/add_version_override
CMake: Adds an OVERRIDE_VERSION option
2021-12-31 02:51:15 -05:00
Mai M 340699ca33 Merge pull request #1485 from Sonicadvance1/fix_fortify
FEXCore: Fixes a memcpy overflow in processor brand
2021-12-31 02:51:06 -05:00
Ryan Houdek 00a572c00c CMake: Adds an OVERRIDE_VERSION option
This is intended to be used by a package maintainer to override the
version when the git repo or git executable isn't available.

Not expected to be used by normal users. Instead by our automated ppa
tooling.
2021-12-30 22:55:46 -08:00
Ryan Houdek d300181ab9 FEXCore: Fixes a memcpy overflow in processor brand
When our FEX version string has less than 16 characters this would
result in size_t underflow, resulting in a crash.

This only ever occurs on a release tag so it went uncaught at first.
The PPA builder managed to uncover this problem since it only deals with
release tags.
2021-12-30 22:54:44 -08:00
Ryan Houdek 5146651908 Merge pull request #1482 from Sonicadvance1/override_tune_cpu
CMake: Adds a TUNE_CPU option
2021-12-29 20:36:04 -08:00
Ryan Houdek 875bae41a3 CMake: Adds a TUNE_CPU option
By default we tune to the native CPU, using some heuristics to determine
the true native CPU since Apple doesn't expose an MIDR.

Adds an option that allows the user to pass in the CPU to tune for.
This will be useful for debugging and also for package building.
2021-12-29 17:27:19 -08:00
Mai M 95bd309db9 Merge pull request #1481 from Sonicadvance1/fix_config_dependency
FEXCore: Fixes FEXCore_Base config dependency
2021-12-29 20:07:22 -05:00
Ryan Houdek 6f1e7e7a4c FEXCore: Fixes FEXCore_Base config dependency
This depends on Config. Then everything else links to this, pulling in
the dependency.
2021-12-29 16:52:26 -08:00
Stefanos Kornilios Mitsis Poiitidis b078f51199 Merge pull request #1480 from Sonicadvance1/fix_asmtest_dependencies
unittests: Fixes a missing dependency on ASM tests
2021-12-29 15:37:32 +02:00
Ryan Houdek d03856f0b7 unittests: Fixes a missing dependency on ASM tests
The output asm folder needs to be created before the config file can be
generated. Otherwise the python script will fail.
2021-12-29 00:27:13 -08:00
Ryan Houdek 98210fdf47 Merge pull request #1479 from Sonicadvance1/fix_msqid_ds
Linux: Finish 32-bit msqid_ds usage
2021-12-28 23:37:20 -08:00
Ryan Houdek 3d1cef4afe Linux: Documents msgctl syscall only supporting IPC_64
I thought this was incorrect but since it only supports IPC_64 it is
actually correct.
If a 32-bit application wants to use the old encoding then it needs to
use the ipc syscall instead.
Fixes #1253
2021-12-28 23:04:53 -08:00
Ryan Houdek 060ab98378 Linux: Implements 32-bit msgctl IPC_SET
Needed struct version conversion support.
Removes the remaining UNHANDLED defines since we support them all now.
2021-12-28 23:02:18 -08:00
Ryan Houdek 1d5bd1520a Merge pull request #1478 from Sonicadvance1/stop_hardcode
Stop hardcoding /usr/bin paths
2021-12-28 22:35:15 -08:00
Ryan Houdek 9eca823e84 Stop hardcoding /usr/bin paths
Instead of using execve, use execvpe.
This glibc helper will search PATH or `confstr(_CS_PATH)`

Fixes #1475
2021-12-28 21:40:22 -08:00
Ryan Houdek cd4269f4e8 Merge pull request #1477 from Sonicadvance1/implement_into
FEXCore: Implements recoverable INTO instruction
2021-12-28 20:38:32 -08:00
Ryan Houdek 2f1d44f838 unittests: Basic INTO test
This only tests for the non-faulting INTO instruction.

This is because our ASM tests can not test for signals. Nor can it
recover.
2021-12-28 19:58:35 -08:00
Ryan Houdek 95eb456065 FEXCore: Implements recoverable INTO instruction
This instruction is only available on 32-bit x86. On overflow flag set,
this instruction will raise a SIGSEGV with overflow exception set in
siginfo_t's si_code.

This implements the everything required to raise the signal and modify
our signal results that we are giving to the guest.

This is effectively the first step in generating signal frames with
custom data in it but only enough for INTO today.
2021-12-28 19:58:35 -08:00
Ryan Houdek d4b31cd4c6 Merge pull request #1476 from Sonicadvance1/squashfs_robust_lock
RootFS: Be more robust against stale lock files
2021-12-28 17:26:54 -08:00
Ryan Houdek c611eb228f RootFS: Be more robust against stale lock files
In the case of having a stale lock file on the system, usually due to an
unclean shutdown, have FEX retry if the socket for that lock file is
also not available.

This resolves the exact case of:
```
FEXBash xeyes &
kill -9 `pidof FEXMountDaemon`
sudo shutdown -r now
FEXBash xeyes & <--- This command would now fail
```

The first command would do:
 - Check lock file and socket
 - Spin up FEXMountDaemon
 - Create Lock file and Socket file
 - Start executing XEyes in FEXInterpreter

The second command would do:
 - Kills the FEXMountDaemon process
 - This leaves squashfuse running
 - The lock file and socket file are now unmanaged and just files on the filesystem

The third command would do:
 - unmounts the squashfuse mount
 - Leaving a dangling folder as a mount point

The fourth command would do:
 - Check for lock file and socket
 - Attempt to use lock file, socket, and mount that is dangling
 - Fail with error

Now instead fourth command will do:
 - Check for lock file and socket
 - Check if filesystem is valid
 - Knows that lock file exists but socket isn't active at all
 - Deletes lock file
 - Spin up FEXMountDaemon
 - Create Lock File and Socket File
 - Start executing XEyes in FEXInterpreter
2021-12-28 16:48:46 -08:00
Ryan Houdek bc120b9f97 Merge pull request #1471 from Sonicadvance1/remaining_syscalls
Linux: Implements the remaining 32-bit syscalls
2021-12-27 19:25:35 -08:00
Ryan Houdek 995672921f Linux: Implements the remaining 32-bit syscalls
The main one that we can't implement is readdir. Falls in to the same
problem space as getdents.

With this, we have the full entry tables filled out for both 64-bit and
32-bit. With some holes in the implementation, we have almost all
coverage now.
2021-12-27 02:19:30 -08:00
Ryan Houdek 514b6a822f Linux: Oops, forgot to pass usig to io_pggetevents
Easy enough to fix
2021-12-27 02:17:52 -08:00
Ryan Houdek 0945c728e9 Linux: Oops, missed fsmount syscall implementation
Simple fix
2021-12-27 02:17:20 -08:00
Ryan Houdek 3a3e2776ba Adds Game bug issue template 2021-12-25 12:29:59 -08:00
Ryan Houdek b72245a2d5 Merge pull request #1469 from Sonicadvance1/remove_0x
FEXRootFSFetcher: Remove 0x prefix on hash
2021-12-24 20:46:43 -08:00
Ryan Houdek 9c49bd3c8e FEXRootFSFetcher: Remove 0x prefix on hash
So as to not confuse users who just copy and paste the hash without
thinking
2021-12-24 20:33:23 -08:00
Ryan Houdek c691d70919 Merge pull request #1467 from Sonicadvance1/fix_really_old_ubuntu
Improves compile ability for older libraries
2021-12-24 18:50:39 -08:00
Ryan Houdek f3a27a57f1 Merge pull request #1468 from Sonicadvance1/experimental_libcxx
CMake: Adds experimental libc++ option
2021-12-24 18:50:31 -08:00
Ryan Houdek edce981824 Merge pull request #1463 from Sonicadvance1/enable_threads_default
Config: Enables all host threads by default
2021-12-24 18:50:20 -08:00
Ryan Houdek 9bffaeea40 Merge pull request #1462 from Sonicadvance1/sanitize_core_option
Config: Sanitize Core option
2021-12-24 18:50:11 -08:00
Ryan Houdek 88ce9b5cd2 Merge pull request #1461 from Sonicadvance1/fix_gdb_map
GDBServer: Fixes long string packet encodings
2021-12-24 18:50:01 -08:00
Ryan Houdek 72e8a997f6 Merge pull request #1460 from Sonicadvance1/FEXRootFSFetcher
FEXRootFSFetcher: Adds a new tool to help set up a new RootFS
2021-12-24 18:49:50 -08:00
Ryan Houdek d708cbad5e Config: Always fixup Threads option
In the case of nothing being set then with the default being zero now we
will not have fixed it up to calculate the number of threads based on
host core count.
2021-12-24 16:32:58 -08:00
Ryan Houdek 152eaff00b CMake: Adds experimental libc++ option
FEX currently doesn't compile with it mainly because libc++ doesn't
support c++ pmr at all.
2021-12-24 15:46:58 -08:00
Ryan Houdek 2079f6b3c7 Improves compile ability for older libraries
Adds a header only include utility folder that can be included from
everywhere.

Contains syscall helpers for older glibc and defines for older Linux
uapi headers missing some defines.
2021-12-24 15:43:25 -08:00
Ryan Houdek 2c31080fd4 FEXRootFSFetcher: Adds a new tool to help set up a new RootFS
This tool supports both a zenity and tty interface.
TTY will be presented if available while Zenity will be used otherwise.

This tool pulls a rootfs list from https://rootfs.fex-emu.org/
It then allows you to select a rootfs from the list, download it, place
it in to the correct working folder, extract it if desired, and set it
as the current default RootFS.

This requires curl and potentially zenity installed to use.
Maybe also unsquashfs if the user chooses to extract the image.

This is an all in one tool to quickly get a new user up and running.

Additionally if you pass in a file path in to the tool, it will generate
an xxhash of the file and exit out. This is the hash used to ensure the
files are valid.
2021-12-24 13:38:20 -08:00
Ryan Houdek ef7b77dff7 Config: Enables all host threads by default
We've solved the few threading problems that we had before. So now allow
FEX to query the host for the number of threads.
2021-12-24 13:22:01 -08:00
Ryan Houdek 777aadb73e Config: Sanitize Core option
If set to an invalid Core option from the json file then sanitize it
back to JIT.

Otherwise FEX has a chance of just crashing.
2021-12-24 13:21:47 -08:00
Ryan Houdek ccd06e2097 GDBServer: Fixes long string packet encodings
Fixes thread, memory-map, and OS data packet types.
These were attempting to substr when the encode function already handles
that.
Was making it so gdb was only ever receiving the first 1000 bytes of the
data and then decoding incorrectly.
2021-12-24 13:21:35 -08:00
Ryan Houdek 09a5f8c6b5 FEX: Move common config loading to a helper 2021-12-24 13:21:24 -08:00
Ryan Houdek 8f835678f5 Merge pull request #1466 from Sonicadvance1/fix_old_ubuntu3
Linux: Fixes build on Ubuntu 20.04 take 3
2021-12-24 13:20:45 -08:00
Ryan Houdek 8f2cd39802 Linux: Fixes build on Ubuntu 20.04 take 3
Fixes #1457
Header is necessary for x64 types as well for struct verifier
2021-12-24 13:05:32 -08:00
Ryan Houdek 059e8a9099 Merge pull request #1465 from Sonicadvance1/fix_old_ubuntu2
Linux: Fixes build on Ubuntu 20.04 take 2
2021-12-24 12:40:34 -08:00
Ryan Houdek 5fae07f6ee Linux: Fixes build on Ubuntu 20.04 take 2
Fixes #1457

Header include order matters here. Rude.
2021-12-24 12:23:52 -08:00
Ryan Houdek 02005e71ca Merge pull request #1464 from Sonicadvance1/fix_old_ubuntu
Linux: Fixes build on Ubuntu 20.04
2021-12-24 12:11:49 -08:00
Ryan Houdek 59cc5228b8 Linux: Fixes build on Ubuntu 20.04
Fixes #1457
2021-12-24 02:46:29 -08:00
Ryan Houdek 117cbde226 Merge pull request #1455 from Sonicadvance1/detect_fpcr_exceptions
HostFeatures: Detect if the host CPU suports float exceptions
2021-12-23 13:42:08 -08:00
Ryan Houdek 565d1e27d7 Merge pull request #1456 from Sonicadvance1/fix_sized_constant
IR: Fixes sized constant mask
2021-12-18 23:31:33 -08:00
Ryan Houdek 22ecff05c9 IR: Fixes sized constant mask
This mask was subtracting backwards, which made anything smaller than 8
byte constants not mask correctly
2021-12-18 23:14:55 -08:00
Ryan Houdek 17480c0e2d HostFeatures: Detect if the host CPU suports float exceptions
On x86 this is always supported.
On ARM this is only supported if FPCR writes actually enable the things.
Also detects the AFP feature for flushing input denormals to zero.

These are all part of the x86 MXCSR.
No Cortex supports FPCR exceptions, while Apple M1 CPUs support
Exceptions but not the true "AFP" extension
Apple instead supports some additional flags in their
`SYS_APL_AFPCR_EL0` register for enabling this.
2021-12-17 16:12:31 -08:00
Ryan Houdek 0608d95322 Updates externals vixl 2021-12-17 16:11:40 -08:00
Ryan Houdek 47bd47950b Merge pull request #1454 from Sonicadvance1/fix_aarch64_ubuntu_20.04
Linux: Fixes Ubuntu 20.04 compilation on AArch64
2021-12-16 22:41:05 -08:00
Ryan Houdek 9aab7e07f4 Linux: Fixes Ubuntu 20.04 compilation on AArch64
This header was missing
2021-12-16 22:23:32 -08:00
Mai M f2fd9d9e3f Merge pull request #1451 from Sonicadvance1/fix_semid_ds
Linux: Fixes semid_ds_64 definition
2021-12-17 01:19:41 -05:00
Mai M 45430e9569 Merge pull request #1452 from Sonicadvance1/fix_32bit_shmctl
Linux: Fixes 32-bit shmctl
2021-12-17 01:19:30 -05:00
Mai M be3e3a351a Merge pull request #1453 from Sonicadvance1/fix_cpuid_crash
CPUID: Fixes crash on unknown CPU
2021-12-17 01:19:19 -05:00
Ryan Houdek 2cee9e5d4b CPUID: Fixes crash on unknown CPU
If the CPU is unknown inside of the ARM CPU detection then the
MIDROption selected could have fallen down a path where it is set to
nullptr.

Resolve this crash by doing a nullptr check.
2021-12-16 21:50:42 -08:00
Ryan Houdek 03cc35f341 Linux: Fixes 32-bit shmctl
We were dereferencing the shmun ptr instead of using it directly.
Resulting in an almost immediate crash

Additionally IPC_SET doesn't write back to the shmid_ds provided.

Additional still SHM_STATE/SHM_STAT_ANY/IPC_STAT wasn't writing its
result back to the guest. Resulting in invalid stat information for the
guest.

Additionally SHM_INFO doesn't follow IPC64 behaviour since there is no
shm_info 64-bit type for a 32-bit OS.
The kernel just truncates the results in this case. Could be an
oversight on the kernel dev's side?

Fixes #1252
2021-12-15 23:34:15 -08:00
Ryan Houdek 3aead5ef65 Linux: Move some 32-bit ipc types to Types.h
This way we can run the struct verifier over it
2021-12-15 23:29:08 -08:00
Ryan Houdek b9abd091d5 Linux: Fixes semid_ds_64 definition
Wasn't quite the expected definition.

Fixes #1450
2021-12-15 22:02:37 -08:00
Ryan Houdek 43dc232e5a Merge pull request #1442 from Sonicadvance1/more_aot_changes
More AOT code movement
2021-12-15 03:27:01 -08:00
Ryan Houdek ea73c9d7ea Merge pull request #1443 from Sonicadvance1/squashfs_fixes
FEXMountDaemon squashfs fixes.
2021-12-15 03:04:15 -08:00
Ryan Houdek fa6f1b1d90 Merge pull request #1446 from Sonicadvance1/more_fault_reconstruction
Dispatcher: Adds more state reconstruction to state restore
2021-12-15 03:04:09 -08:00
Ryan Houdek e24eb7a72d Merge pull request #1447 from Sonicadvance1/consolidate_host_features
HostFeatures: Consolidates HostFeatures flags
2021-12-15 02:55:42 -08:00
Ryan Houdek e05d116b05 Merge pull request #1448 from Sonicadvance1/fix_proton_experimental
Linux: Adds the safe syscall unimplemented gap for x86-64
2021-12-15 02:55:35 -08:00
Ryan Houdek ee821b9cbf Merge pull request #1449 from Sonicadvance1/implement_rdtscp
Implements RDTSCP
2021-12-15 02:52:47 -08:00
Ryan Houdek c29e563836 unittests: Adds tests for RDTSCP 2021-12-15 01:21:16 -08:00
Ryan Houdek fd59fb1a7a OpcodeDispatcher: Implements support for RDTSCP
Have fun
2021-12-15 01:18:34 -08:00
Ryan Houdek 04deeb3911 IR: Adds Processor ID IR op
For the x86-64 JIT this is implemented with pulling rdtscp's result for
this value.
For Interpreter and AArch64 JIT this is implemented with the getcpu
syscall.

Theoretically AArch64 could implement this with MPIDR_EL1 but because
SoC vendors hecked this up, we can't. Thanks.
Kernel just returns zero + reserved bits if you try reading it.
2021-12-15 01:17:24 -08:00
Ryan Houdek 4a7c65b20d Linux: Adds the safe syscall unimplemented gap for x86-64
Syscalls on x86-64 have a gap in the range of [335, 424) where these are
defined as returning ENOSYS.
This was a decision that the Linux developers decided on so x86-64 can
realign its syscall numbers to a common infrastructure.

Turns out that Proton Experimental was using a 32-bit only syscall to
check if the feature existed and was listening for ENOSYS.
They did this unconditionally on both x86 and x86-64 and it hits this
gap section.

This was introduced in Proton Experimental with commit
563fb0fbe2a49734aede87bedccb20daefc553b0
Message: "ntdll: Use clock_gettime64 if supported."

This is technically valid since it falls within this gap section but
leaves a bad taste.

We can safely ignore this section so set it up with the
UnimplementedSyscallSafe handler and also assign it to Syscall MAX so it
can go down our fast handler.

Not that this is likely to be a performance critical path.

Fixes Proton Experimental crashing when run under FEX.
2021-12-14 03:45:55 -08:00
Ryan Houdek 23cb0dea00 HostFeatures: Consolidates HostFeatures flags
Some of these were in the Emitter class and some were in the
HostFeatures.

Merge these together since in the future I'm going to be using all of
this data as a key for our AOT code cache.
2021-12-14 02:41:42 -08:00
Ryan Houdek bcad7a9eea Dispatcher: Adds more state reconstruction to state restore
Still only setting the new state if RIP is affected for now.
Noticed a bug where we weren't setting our frame RIP to the new RIP on
32-bit.
Decided to walk through more of the state setting while fixing that.

This gets #1214 further but then it eventually crashes with a read to
0x11.
2021-12-14 02:36:24 -08:00
Ryan Houdek 0a3a270a63 Merge pull request #1445 from lioncash/pdep
OpcodeDispatcher: Handle PDEP
2021-12-13 16:43:35 -08:00
lioncash 06e4a5a5b7 CPUID: Signify full support for BMI2
With PDEP support dropped in, we now support all of BMI2, so we can
signify that we support it in our emulated CPUID.
2021-12-13 14:07:41 -05:00
lioncash 6ff80670b3 OpcodeDispatcher: Handle PDEP
Now all of BMI2 is handled.
2021-12-13 14:06:55 -05:00
lioncash 38eea80b8d IR: Add PDep IR opcode 2021-12-13 13:54:15 -05:00
Ryan Houdek c43af0e10f Merge pull request #1444 from neobrain/fix_jemalloc_overrides
Update jemalloc for fixed libc overrides
2021-12-13 04:48:05 -08:00
Tony Wasserka 2fa8cd7e97 Update jemalloc for fixed libc overrides 2021-12-13 13:38:39 +01:00
Ryan Houdek 5d0734a7f2 FEXMountDaemon: Attempt to increase FD max when close to limit
On FEXMountDaemon startup, fetch the FD limit and current number of open
files. This allows us to track the number of open FDs we have.

Once we get close the the safe threshold we then try increasing the soft
limit towards the hard limit. If we bump up to the max (Say someone
setting the hard limit to 256) then print some fairly big warning
messages.

With the previously fixed FD leak, we could very quickly hit the pipe
limit because Steam is constantly opening new processes all the time.
With the leak fixed, running Steam and idling, FEXMountDaemon only has
around 16 pipes being watched.
2021-12-12 22:51:53 -08:00
Ryan Houdek 3e0e922fef FEXMountDaemon: Fixes EPoll thread not shutting down on SIGTERM
This is another edge case where if the process was sent a SIGTERM then
it can get in to a weird state where pipe tracking is broken.

We already check for zero pipes being tracked at the end of this loop,
just check for the shutdown variable to be set instead
2021-12-12 22:51:15 -08:00
Ryan Houdek e2004f4999 FEXMountDaemon: Fix leaked pipe FDs
Even though we were removing the pipe from the epoll interest list,
we were failing to close the pipe fd. This was leaking the FD and
running in to the Linux process FD limitation over time.

This fixes a class of issues where if you were running FEXMountDaemon
that hit the max, then there were FEX processes not being watched. Which
means that FEXMountDaemon could exit while FEX was still trying to run.

Usually a non-issue since squashfuse would keep running in the
background since open files in the mount were still in use.

Could hit fun race conditions because of it though.
2021-12-12 22:45:40 -08:00
Ryan Houdek 13f37efe12 SquashFS: Work around a race condition on FEXMountDaemon ack
On an overburdened system then FEXMountDaemon may notrespond within the two
seconds of FEX waiting for an ack.
In this case then we hit a bad edge case where we overwrite the lock
file with a new FEXMountDaemon instance then spin up a new
FEXMountDaemon.
This new FEXMountDaemon will recognize that one is still running and
early exit, but now that we've overwritten the lock file with a new
location, all future FEX instances won't be able to find the squashfs
rootfs.

Instead, ignore that the ack wasn't received and attempt running
instead.

FEXMountDaemon should still pick up the pipe message and monitor FEX
regardless.

This fixes the `/usr not found` messages and crashes resulting from it.
2021-12-12 22:44:48 -08:00
Ryan Houdek 80e66a36eb FEXConfig: Adds AOT options to the UI 2021-12-12 18:22:02 -08:00
Ryan Houdek 4998d35ec5 FEXLoader: Moves AOT handling to its own independent files
This is only moving it outside of FEXLoader to make it easier to work
on.

I'll be hammering on this more soon.
2021-12-12 18:19:12 -08:00
Ryan Houdek f0655874e3 ELFCodeLoader: Fixes bug in ELF section tracking
Each section map invocation was clearing the sections pointer, meaning
we only ever had one section in the vector.

We want to keep all of the sections around from the ELFCodeLoader.
Fixes a bug where we were effectively never catching any code segments
from the executable passed in.
2021-12-12 18:15:06 -08:00
Ryan Houdek 9394e49c95 Merge pull request #1431 from Sonicadvance1/expose_arm_names
CPUID Expose Hybrid flag and CPU names
2021-12-12 18:07:53 -08:00
Ryan Houdek 91984003b6 Merge pull request #1440 from lioncash/pext
OpcodeDispatcher: Handle PEXT
2021-12-10 16:27:49 -08:00
lioncash dbf571fdfb OpcodeDispatcher: Handle PEXT 2021-12-10 19:03:58 -05:00
lioncash b6abcc5e3c IR: Add PExt opcode 2021-12-10 19:03:54 -05:00
Ryan Houdek e99a23cfc6 Merge pull request #1439 from neobrain/opt_x86tables_init
Speed up initialization of X86Tables
2021-12-10 09:11:47 -08:00
Ryan Houdek 425d9323d5 Merge pull request #1295 from neobrain/feature_new_thunk_gen
Implement thunk library generation using libclang
2021-12-10 09:10:00 -08:00
Tony Wasserka a7c0997daf OpcodeDispatcher: Mark const-initialized instruction tables as constexpr
This allows the compiler back these tables into the executable, which
reduces the amount of work the function has to do at runtime.

Reduces the runtime of this function by 30% relative to the previous commit.
2021-12-10 13:25:24 +01:00
Tony Wasserka 6eba3f331e OpcodeDispatcher: Store instruction tables as arrays instead of std::vector
Reduces the runtime of this function by 90%.
2021-12-10 13:25:23 +01:00
Tony Wasserka bcc75e3312 X86Tables: Use stack-allocated arrays instead of std::vector 2021-12-10 13:25:23 +01:00
Tony Wasserka 30672e1517 X86Tables: Remove unneeded initialization code from the Debug mode path
The tables have recently been changed to be zeroed out as a whole on startup.

Speeds up InstallDebugInfo by about two orders of magnitude and reduces Debug
executable size by 1.8%.
2021-12-10 13:25:23 +01:00
Tony Wasserka b6e46fd44d Thunks: Update documentation to reflect generator changes 2021-12-10 11:25:01 +01:00
Tony Wasserka 32a0e37569 Thunks: Remove now unneeded generator files 2021-12-10 11:25:01 +01:00
Tony Wasserka 5b6175b702 Remove now unused Vulkan-Docs submodule 2021-12-10 11:25:01 +01:00
Tony Wasserka 0285e35c87 Thunks: Use libclang-based code generation for libvulkan 2021-12-10 11:25:01 +01:00
Tony Wasserka bbfb8713a7 Thunks: Drop struct verifier tests for Vulkan
This is currently not supported with the new generator.
2021-12-10 11:25:00 +01:00
Tony Wasserka c1d6967cb6 Thunks/gen: Add support for expanding symbol lists using a macro 2021-12-10 11:25:00 +01:00
Tony Wasserka 20e24eae9c Thunks: Use libclang-based code generation for libdrm 2021-12-10 11:25:00 +01:00
Tony Wasserka b8d9027680 Thunks: Use libclang-based code generation for libxshmfence 2021-12-10 11:25:00 +01:00
Tony Wasserka 688ef9f5af Thunks: Use libclang-based code generation for libxcb-xfixes 2021-12-10 11:25:00 +01:00
Tony Wasserka 181c9074ab Thunks: Use libclang-based code generation for libxcb-sync 2021-12-10 11:25:00 +01:00
Tony Wasserka 6f85f64bfd Thunks: Use libclang-based code generation for libxcb-shm 2021-12-10 11:25:00 +01:00
Tony Wasserka ba6ee61db2 Thunks: Use libclang-based code generation for libxcb-randr 2021-12-10 11:25:00 +01:00
Tony Wasserka e735281b2c Thunks: Use libclang-based code generation for libxcb-present 2021-12-10 11:24:59 +01:00
Tony Wasserka c8b2d714a8 Thunks: Use libclang-based code generation for libxcb-glx 2021-12-10 11:24:59 +01:00
Tony Wasserka bba0625a57 Thunks: Use libclang-based code generation for libxcb-dri3 2021-12-10 11:24:59 +01:00
Tony Wasserka c9a7b36210 Thunks: Use libclang-based code generation for libxcb-dri2 2021-12-10 11:24:59 +01:00
Tony Wasserka 99bf02f27f Thunks/gen: Add support for libraries with "-" in the filename 2021-12-10 11:24:59 +01:00
Tony Wasserka a2c02ae51e Thunks: Use libclang-based code generation for libxcb
Notably, xcb_take_socket uses the callback_guest annotation, since the
given callback is never called on the host (instead it's manually forwarded
back to a helper guest thread for calling).
2021-12-10 11:24:59 +01:00
Tony Wasserka c7b59143c6 Thunks/gen: Add custom guest entrypoint annotation 2021-12-10 11:24:59 +01:00
Tony Wasserka 1e23e61572 Thunks: Use libclang-based code generation for libEGL 2021-12-10 11:24:59 +01:00
Tony Wasserka f4dd1a895e Thunks: Use libclang-based code generation for libGL 2021-12-10 11:24:58 +01:00
Tony Wasserka 3a2f9a4d46 Thunks/gen: Support guest-side symbol tables and custom host symbol loaders 2021-12-10 11:24:58 +01:00
Tony Wasserka 4a848d7202 Thunks: Use libclang-based code generation for libX11 2021-12-10 11:24:58 +01:00
Tony Wasserka 1b58ed9f57 Thunks/X11: Fix incorrect library version 2021-12-10 11:24:58 +01:00
Tony Wasserka 1c86f7ed36 Thunks/gen: Support annotating guest-exclusive function pointers
Function pointer arguments given to the guest thunk library aren't
callable on the host, so they need special handling on a case-by-case
basis. Similar problems arise when the guest-side tries to consume a
pointer returned from a native host library. To automate some common
scenarios, this change adds two new annotations:

"callback_guest" indicates the callback parameter is never called on the
host and hence can be marshalled like any other argument. Accidental host
calls to the function pointer are prevented by wrapping it in an opaque
type alias.

"returns_guest_pointer" indicates that the host function returns a pointer
usable in the guest context. This applies e.g. to functions that derive
the returned pointer from input guest pointer arguments.
2021-12-10 11:24:58 +01:00
Tony Wasserka 2733b2ee1e Thunks/gen: Add annotation for custom host implementations 2021-12-10 11:24:58 +01:00
Tony Wasserka 4ceb2dfdf2 Thunks: Use libclang-based code generation for libXext 2021-12-10 11:24:58 +01:00
Tony Wasserka bd380e0f15 Thunks: Use libclang-based code generation for libXrender 2021-12-10 11:24:58 +01:00
Tony Wasserka 73ec786c60 Thunks: Use libclang-based code generation for libasound 2021-12-10 11:24:57 +01:00
Tony Wasserka e3a2c8dc80 Thunks: Use libclang-based code generation for libXfixes 2021-12-10 11:24:57 +01:00
Tony Wasserka d3b14df840 Thunks/gen: Add versioning support 2021-12-10 11:24:57 +01:00
Tony Wasserka a12ab8f98a Thunks/gen: Add support for "callback_stub" annotations
Some applications set callbacks that never get called in practice (such as
error handlers). It's sensible to just not implement these instead of
cluttering the code with effectively unused callback wrappers.
2021-12-10 11:24:57 +01:00
Tony Wasserka 966b9a69d8 Thunks/gen: Add support for function pointer parameters ("callbacks") 2021-12-10 11:24:57 +01:00
Tony Wasserka 927d3d00e2 Thunks/gen: Add support for thunking variadic functions 2021-12-10 11:24:57 +01:00
Tony Wasserka bfb9cabeb8 Thunks/gen: Add support for generation of host library functions 2021-12-10 11:24:57 +01:00
Tony Wasserka 22466a973c Thunks/gen: Add unit tests 2021-12-10 11:24:57 +01:00
Tony Wasserka c05e1c9797 Thunks: Add a new code generator based on libclang 2021-12-10 11:24:56 +01:00
Ryan Houdek c65be9f55d Merge pull request #1438 from Sonicadvance1/remove_jemalloc_from_fexconfig
FEXConfig: Removes jemalloc usage
2021-12-09 18:10:32 -08:00
Ryan Houdek edc31fe0a3 FEXConfig: Removes jemalloc usage
Splits out the few required dependencies to a FEXCore_Base static
library.

FEXCore then links to this directly.

Then make it so the FEX Common code links to FEXCore_Base so the
jemalloc dependency doesn't get pulled in.
2021-12-09 17:54:53 -08:00
Ryan Houdek d4655fbb17 Merge pull request #1424 from Sonicadvance1/aot_code_movement
FEXCore: Reorganizes some AOT related code
2021-12-09 13:55:50 -08:00
Ryan Houdek 5f0dfcd715 Merge pull request #1422 from Sonicadvance1/implement_clzero
Implements CLZero instruction
2021-12-09 13:55:31 -08:00
Ryan Houdek e62cb417c2 Merge pull request #1437 from Sonicadvance1/disable_flake_gvisor
gvisor: Disables flaky test
2021-12-09 12:58:29 -08:00
Ryan Houdek df486c0786 gvisor: Disables flaky test 2021-12-09 12:37:58 -08:00
Ryan Houdek 9d43904792 Merge pull request #1436 from lioncash/context-const
Context: Take some arguments as pointer-to-const
2021-12-09 12:23:34 -08:00
Ryan Houdek 5654f9a030 Merge pull request #1435 from Sonicadvance1/fix_vixl_assertions
Arm64: Fixes vixl assertions around ubfm usage
2021-12-09 11:54:41 -08:00
Ryan Houdek d74cf6d8d8 CPUID Expose Hybrid flag and CPU names
Had some idle time so I implemented this logic.

We do some tricky logic to have a big.little configuration even with
unknown CPU core types. Promoting or demoting a single MIDR depending on
if we have a mixed configuration or not.

In a non-hybrid design we only claim product names inside the CPUID
product string.

This will appear if you `/proc/cpuinfo` or read the CPUID registers
directly

eg on Snapdragon 888:
processor       : 0
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 1
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 2
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 3
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 4
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 5
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 6
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 7
model name      : FEX-2112-1-g13b14b85            Cortex-X1

eg on Macbook Pro VM which can't see the CPU type:
processor       : 0
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 1
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 2
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 3
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 4
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 5
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 6
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 7
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
2021-12-09 11:39:52 -08:00
Ryan Houdek 84905d2856 unittests: Adds 32-bit inline syscall test
Triggers the ubfm bug on AArch64.
Uses a fairly benign syscall that gets inlined
2021-12-09 11:34:35 -08:00
Ryan Houdek 92cb9d477e Arm64: Fixes vixl assertions around ubfm usage
vixl has an assert check to ensure the register sizes are the same for
ubfm.

32-bit Inline syscalls hit this for both arguments and return value.
VCastFromGPR would have hit this but there aren't any x86 instructions
that move 8-bit and 16-bit values in to a vector register.
2021-12-09 11:34:20 -08:00
lioncash 7ed6007252 Context: std::move functions in signal registration functions
Prevents potential reallocations, given we're using std::function here.
2021-12-09 14:31:47 -05:00
Ryan Houdek b6499ac724 Merge pull request #1434 from lioncash/syscallthread
{x32, x64}/Thread: Make use of .data() instead of .at(0)
2021-12-09 11:28:39 -08:00
lioncash c73991f467 Context: Take some arguments as pointer-to-const
Several API functions act as state querying functions. These can take
some parameters by const to communicate that we don't intend to modify
the respective passed in instance.
2021-12-09 14:26:35 -05:00
lioncash 0bf0fe4779 {x32, x64}/Thread: Make use of .data() instead of .at(0)
Follows the previous PRs of using .data() instead of .at(0) to retrieve
the buffer pointers.
2021-12-09 14:14:40 -05:00
Ryan Houdek 097f48f3ff Merge pull request #1433 from lioncash/syscallmsg
x32/Socket: Resolve sign comparison mismatch cases
2021-12-09 11:09:07 -08:00
Ryan Houdek 75e5df545a Merge pull request #1432 from lioncash/syscall-warn
LinuxSyscalls: Enable warnings
2021-12-09 10:51:21 -08:00
lioncash 171d5f7263 x32/Socket: Use .data() instead of .at(0) for pointer retrieval
.at(0) can blow up in the case we receive a message from user code with msg_iovlen=0
.data() will simply return a pointer to the empty buffer and we can
proceed with forwarding the call to the syscall.
2021-12-09 13:49:15 -05:00
lioncash 970067d19b x32/Socket: Resolve sign comparison mismatch cases 2021-12-09 13:43:26 -05:00
lioncash c26ff60949 LinuxSyscalls: Enable warnings
Enables some basic warnings that we have enabled for FEXCore in order to
prevent some small things from slipping through.
2021-12-09 12:58:54 -05:00
lioncash 04d830ed39 LinuxSyscalls: Minor tidy up of CMakeLists 2021-12-09 12:46:39 -05:00
Ryan Houdek b9c49027c7 Merge pull request #1409 from Sonicadvance1/reentrantmutex
Adds a new ReentrantMutex to use for FEXCore
2021-12-08 14:00:12 -08:00
Ryan Houdek 9860e8b71b Merge pull request #1430 from lioncash/net
NetStream: Move NetBuf definition into cpp file
2021-12-08 13:10:06 -08:00
lioncash a1e94a9863 NetStream: Move NetBuf definition into cpp file
Keeps the NetBuf class completely internal and also lessens the header
dependencies for NetStream.
2021-12-08 10:41:55 -05:00
lioncash 2945c13dcb NetStream: Mark virtual functions as override
Makes it explicit that we're overriding the interface of the class being
derived from.
2021-12-08 10:33:18 -05:00
lioncash 4f68821aef NetStream: Mark constructors as explicit
Prevents implicit conversion of ints to NetStreams.
2021-12-08 10:30:38 -05:00
Ryan Houdek 13087f8425 Merge pull request #1426 from Sonicadvance1/fcmp_unittests
Arm64: Fixes MapSelectCC FGT flag
2021-12-08 05:26:15 -08:00
Ryan Houdek fd0424768b Merge pull request #1427 from Sonicadvance1/enable_unaligned_atomic_tests
unittests: Enables remaining unaligned atomic tests on ARMv8.0
2021-12-08 05:26:03 -08:00
Ryan Houdek 5c4112f103 Merge pull request #1428 from Sonicadvance1/fix_known_failures
unittests: Fixes known failures
2021-12-08 05:25:50 -08:00
Ryan Houdek e8670bbab2 Merge pull request #1425 from Sonicadvance1/x86tables_unknown
X86Tables: Build Unknown op definition tables at compile time
2021-12-08 00:22:57 -08:00
Ryan Houdek 13b14b857b unittests: Fixes known failures
These were just incorrect results and I had failed to fix them.
2021-12-07 22:06:36 -08:00
Ryan Houdek 48c7ff2a23 unittests: Enables remaining unaligned atomic tests on ARMv8.0 2021-12-07 21:54:03 -08:00
Ryan Houdek 77a032a286 unittests: Adds tests for handling cmp merging
As a continuation to #1404 these unit tests were needing to be added.
This tests both select and conditional branch merging for integer ops
and float ops.

This is what discovered the issue fixed in f2de640395
2021-12-07 18:45:54 -08:00
Ryan Houdek f2de640395 Arm64: Fixes MapSelectCC FGT flag
As a continuation to #1404, this was mapped to the incorrect Arm64 flags
2021-12-07 18:43:50 -08:00
Ryan Houdek 9a642158e0 Workaround fmt not handling nullptr strings
Simple ternary check each time a name is used
2021-12-07 12:23:56 -08:00
Ryan Houdek dc44caa178 Github: Adds the new APITests to the github workflow 2021-12-07 11:55:26 -08:00
Ryan Houdek 081a003c6c unittests: Adds a new catch2 test for testing the InterruptableConditionVariable 2021-12-07 11:55:26 -08:00
Ryan Houdek 4c5fc6e813 FEXCore: Uses the InterruptableConditionVariable for StartRunning event
This can be used with the gdbstub and on pause could cause longjumps
out of the event.
Resolves the hang on InternalThreadState destruction.
2021-12-07 11:55:26 -08:00
Ryan Houdek 253cdb552d FEXCore/Utils: Adds InterruptableConditionVariable class
Mutex destruction is not safe in the face of longjump and signals.
Adds a new class that is very simple and will still work in this
instance.
2021-12-07 11:55:26 -08:00
Ryan Houdek 6404aba6e2 X86Tables: Build Unknown op definition tables at compile time
It doesn't make any sense anymore to have specific instruction names
set for UND versus a nullptr string anymore.

Was useful when we could use it to determine the difference between
undefined from the start versus set in the tables but with unknown
decoding. Which is an edge case.

Now instead just zero initialize the data, which means it is an unknown
type and nullptr name. Which works for use.

Improves initialization time of the InitializeInfoTables function from
423 microseconds to 37 microseconds.
2021-12-06 18:57:02 -08:00
Ryan Houdek 45f919683a FEXCore: Reorganizes some AOT related code
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.

This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.

Two minor behaviour changes that got mixed up with this change.

The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.

Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
2021-12-06 18:42:21 -08:00
Ryan Houdek c3a6890a40 CMake: Adds Catch build when tests are enabled 2021-12-06 18:27:12 -08:00
Tony Wasserka 07be5a0bae Add Catch2 submodule 2021-12-06 18:27:11 -08:00
Ryan Houdek d6e4da7e77 CPUID: Exposes support for CLZero
Only exposed if the host if the HostFeatures for ARM claim to support
it.
2021-12-03 21:20:40 -08:00
Ryan Houdek ed5d7f6a62 unittests: Adds CLZero unit tests
First test just ensures that an aligned cacheline clear will clear out
the data.
Second test ensures an unaligned address will still only clear the
cacheline the address lies in without touching the other surrounding
cachelines.

Disabled on the host runner since the CI host runner doesn't support the
instruction
2021-12-03 21:19:14 -08:00
Ryan Houdek 18b223811a OpcodeDispatcher: Implements CLZero instruction
This instruction zeroes a cacheline in memory that is weakly ordered and
non-temporal.

It uses the RAX register for where in memory to clear and aligns the
address on cacheline regardless of actual alignment.
2021-12-03 21:15:51 -08:00
Ryan Houdek e1d21f0bff IR: Implements new CacheLineZero op
This zeroes out an emulated 64byte cacheline. Writing zeros to memory.
This very specifically is only 64bytes to match x86 behaviour.
Also specifically non-temporal and weakly ordered. Which matches x86
CLZero behaviour.
2021-12-03 21:14:02 -08:00
Ryan Houdek d2783f2edd HostFeatures: Adds new host feature flag for CLZero
If the DCZID block size matches the CLZero cacheline size. Then claim we
support CLZero here.
Otherwise we don't want to expose support for it.
2021-12-03 21:12:52 -08:00
Ryan Houdek 3377e5a50e HarnessHelpers: Fixes accidental fmt::format
Was meant to be fmt::print. Happened with 62dfccc989

oops
2021-12-03 21:01:59 -08:00
247 changed files with 20655 additions and 16615 deletions

No files matched your search

@@ -0,0 +1,45 @@
---
name: Potential Game Bug
about: A bug in FEX-Emu that causes a problem in a game
title: "[Game]: [Short Problem Description]"
labels: Game related
assignees: ''
---
**What Game**
The game name.
A link to the storefront where to get the game. GOG, Steam, Itch.io, etc
**Describe the bug**
A clear and concise description of what the bug is.
**To Reproduce**
Steps to reproduce the behavior:
1. Go to '...'
2. Click on '....'
3. Scroll down to '....'
4. See error
**Expected behavior**
A clear and concise description of what you expected to happen.
**Screenshots and Video**
If applicable, add screenshots and video to help explain your problem.
**System information:**
- OS: [eg: Ubuntu 21.10]
- CPU/SoC: [eg: Snapdragon 888, Intel Core i8-12900k]
- Video driver version: [eg: OpenGL ES 3.2 Mesa 22.0.0-devel (git-9ff086052a)]
- RootFS used: [eg: Ubuntu 21.10 Official Rootfs]
- FEX version: (FEXGetConfig --version) [eg: FEX-2112-155-gc691d709]
- Thunks Enabled: [Yes/No]
**Additional context**
- Is this an x86 or x86-64 game: [x86/x86-64/Both]
- Does this reproduce on x86-64 host with FEX: [Yes/No/Untested]
- Does this reproduce on AArch64 with Radeon/Intel/Nvidia: [Yes/No/Untested]
- Is this a Vulkan game: [Yes/No/Unknown]
- If Yes, What is your Vulkan driver:
Add any other context about the problem here.
+11
View File
@@ -140,6 +140,17 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target api_tests
- name: APITest Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+3 -4
View File
@@ -42,7 +42,6 @@
[submodule "External/xxhash"]
path = External/xxhash
url = https://github.com/FEX-Emu/xxHash.git
[submodule "External/Vulkan-Docs"]
shallow = true
path = External/Vulkan-Docs
url = https://github.com/KhronosGroup/Vulkan-Docs.git
[submodule "External/Catch2"]
path = External/Catch2
url = https://github.com/catchorg/Catch2.git
+96 -53
View File
@@ -18,11 +18,16 @@ option(ENABLE_STATIC_PIE "Enables static-pie build" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
# These options are meant for package management
set (TUNE_CPU "native" CACHE STRING "Override the CPU the build is tuned for")
set (OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version in the format of <MMYY>{.<REV>}")
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
if (CMAKE_BUILD_TYPE MATCHES "DEBUG")
set(ENABLE_ASSERTIONS TRUE)
@@ -89,6 +94,12 @@ if (ENABLE_LLD)
link_libraries(${LD_OVERRIDE})
endif()
if (ENABLE_LIBCXX)
message(WARNING "This is an unsupported configuration and should only be used for testing")
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -std=c++11 -stdlib=libc++")
set(CMAKE_EXE_LINKER_FLAGS "${CMAKE_EXE_LINKER_FLAGS} -stdlib=libc++ -lc++abi")
endif()
if (NOT ENABLE_OFFLINE_TELEMETRY)
# Disable FEX offline telemetry entirely if asked
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
@@ -268,6 +279,16 @@ endif()
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
if (BUILD_TESTS)
option(CATCH_BUILD_STATIC_LIBRARY "" ON)
set(CATCH_BUILD_STATIC_LIBRARY ON)
add_subdirectory(External/Catch2/)
# Pull in catch_discover_tests definition
list(APPEND CMAKE_MODULE_PATH "${CMAKE_CURRENT_SOURCE_DIR}/External/Catch2/contrib/")
include(Catch)
endif()
add_subdirectory(External/cpp-optparse/)
include_directories(External/cpp-optparse/)
@@ -314,33 +335,42 @@ if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
endif()
endif()
if(_M_ARM_64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
# Clang 12.0 fixed the -mcpu=native bug with mixed big.little implementers
# Clang can not currently check for native Apple M1 type in hypervisor. Currently disabled
check_cxx_compiler_flag("-mcpu=native" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=native")
if (TUNE_CPU STREQUAL "native")
if(_M_ARM_64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
# Clang 12.0 fixed the -mcpu=native bug with mixed big.little implementers
# Clang can not currently check for native Apple M1 type in hypervisor. Currently disabled
check_cxx_compiler_flag("-mcpu=native" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=native")
endif()
else()
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo" "${CMAKE_CXX_COMPILER_VERSION}"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=${AARCH64_CPU}")
endif()
endif()
else()
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo" "${CMAKE_CXX_COMPILER_VERSION}"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=${AARCH64_CPU}")
check_cxx_compiler_flag("-march=native" COMPILER_SUPPORTS_MARCH_NATIVE)
if(COMPILER_SUPPORTS_MARCH_NATIVE)
add_compile_options("-march=native")
endif()
endif()
else()
check_cxx_compiler_flag("-march=native" COMPILER_SUPPORTS_MARCH_NATIVE)
if(COMPILER_SUPPORTS_MARCH_NATIVE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -march=native")
check_cxx_compiler_flag("-mcpu=${TUNE_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=${TUNE_CPU}")
else()
message(FATAL_ERROR "Trying to compile cpu type '${TUNE_CPU}' but the compiler doesn't support this")
endif()
endif()
@@ -414,6 +444,10 @@ if (BUILD_TESTS)
enable_testing()
message(STATUS "Unit tests are enabled")
endif()
add_subdirectory(FEXHeaderUtils/)
include_directories(FEXHeaderUtils/)
add_subdirectory(External/FEXCore)
# Binfmt_misc files must be installed prior to Source/ installs
@@ -432,6 +466,8 @@ if (BUILD_TESTS)
endif()
if (BUILD_THUNKS)
add_subdirectory(ThunkLibs/Generator)
include(ExternalProject)
ExternalProject_Add(host-libs
@@ -441,9 +477,10 @@ if (BUILD_THUNKS)
CMAKE_ARGS
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DVULKAN_XML=${CMAKE_SOURCE_DIR}/External/Vulkan-Docs/xml/vk.xml"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
)
install(
@@ -465,9 +502,10 @@ if (BUILD_THUNKS)
"-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DVULKAN_XML=${CMAKE_SOURCE_DIR}/External/Vulkan-Docs/xml/vk.xml"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
)
install(
@@ -484,43 +522,48 @@ set(FEX_VERSION_MAJOR "0")
set(FEX_VERSION_MINOR "0")
set(FEX_VERSION_PATCH "0")
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe --abbrev=0
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
RESULT_VARIABLE GIT_ERROR
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
if (OVERRIDE_VERSION STREQUAL "detect")
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe --abbrev=0
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
RESULT_VARIABLE GIT_ERROR
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
if (NOT ${GIT_ERROR} EQUAL 0)
# Likely built in a way that doesn't have tags
# Setup a version tag that is unknown
set(GIT_DESCRIBE_STRING "FEX-0000")
if (NOT ${GIT_ERROR} EQUAL 0)
# Likely built in a way that doesn't have tags
# Setup a version tag that is unknown
set(GIT_DESCRIBE_STRING "FEX-0000")
endif()
endif()
else()
set(GIT_DESCRIBE_STRING "FEX-${OVERRIDE_VERSION}")
endif()
# Change something like `FEX-2106.1-76-<hash>` in to a list
string(REPLACE "-" ";" DESCRIBE_LIST ${GIT_DESCRIBE_STRING})
# Parse the version here
# Change something like `FEX-2106.1-76-<hash>` in to a list
string(REPLACE "-" ";" DESCRIBE_LIST ${GIT_DESCRIBE_STRING})
# Extract the `2106.1` element
list(GET DESCRIBE_LIST 1 DESCRIBE_LIST)
# Extract the `2106.1` element
list(GET DESCRIBE_LIST 1 DESCRIBE_LIST)
# Change `2106.1` in to a list
string(REPLACE "." ";" DESCRIBE_LIST ${DESCRIBE_LIST})
# Change `2106.1` in to a list
string(REPLACE "." ";" DESCRIBE_LIST ${DESCRIBE_LIST})
# Calculate list size
list(LENGTH DESCRIBE_LIST LIST_SIZE)
# Calculate list size
list(LENGTH DESCRIBE_LIST LIST_SIZE)
# Pull out the major version
list(GET DESCRIBE_LIST 0 FEX_VERSION_MAJOR)
# Pull out the major version
list(GET DESCRIBE_LIST 0 FEX_VERSION_MAJOR)
# Minor version only exists if there is a .1 at the end
# eg: 2106 versus 2106.1
if (LIST_SIZE GREATER 1)
list(GET DESCRIBE_LIST 1 FEX_VERSION_MINOR)
endif()
# Minor version only exists if there is a .1 at the end
# eg: 2106 versus 2106.1
if (LIST_SIZE GREATER 1)
list(GET DESCRIBE_LIST 1 FEX_VERSION_MINOR)
endif()
# Package creation
Vendored Submodule
+1
Submodule External/Catch2 added at c4e3767e26.
+23 -18
View File
@@ -37,27 +37,32 @@ endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
# Find our git hash
find_package(Git)
set(GIT_SHORT_HASH "Unknown")
set(GIT_DESCRIBE_STRING "FEX-Unknown")
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} rev-parse --short HEAD
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_SHORT_HASH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
if (OVERRIDE_VERSION STREQUAL "detect")
# Find our git hash
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} rev-parse --short HEAD
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_SHORT_HASH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
endif()
else()
set(GIT_SHORT_HASH "${OVERRIDE_VERSION}")
set(GIT_DESCRIBE_STRING "FEX-${OVERRIDE_VERSION}")
endif()
configure_file(
+35 -28
View File
@@ -1,7 +1,14 @@
set (MAN_DIR ${CMAKE_INSTALL_PREFIX}/share/man CACHE PATH "MAN_DIR")
set (SRCS
set (FEXCORE_BASE_SRCS
Common/Paths.cpp
Interface/Config/Config.cpp
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
)
set (SRCS
Common/JitSymbols.cpp
Common/SoftFloat-3e/extF80_add.c
Common/SoftFloat-3e/extF80_div.c
@@ -70,7 +77,6 @@ set (SRCS
Common/SoftFloat-3e/f32_to_extF80.c
Common/SoftFloat-3e/s_normSubnormalF32Sig.c
Common/SoftFloat-3e/s_f32UIToCommonNaN.c
Interface/Config/Config.cpp
Interface/Context/Context.cpp
Interface/Core/LookupCache.cpp
Interface/Core/BlockSamplingData.cpp
@@ -120,6 +126,7 @@ set (SRCS
Interface/Core/X86Tables/X87Tables.cpp
Interface/Core/X86Tables/XOPTables.cpp
Interface/HLE/Thunks/Thunks.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IRParser.cpp
Interface/IR/IREmitter.cpp
@@ -140,9 +147,6 @@ set (SRCS
Interface/IR/Passes/SyscallOptimization.cpp
Utils/Allocator.cpp
Utils/Allocator/64BitAllocator.cpp
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
@@ -195,7 +199,7 @@ if (ENABLE_JIT_ARM64)
Interface/Core/JIT/Arm64/VectorOps.cpp)
endif()
set (LIBS vixl dl fmt::fmt xxhash tiny-json)
set (LIBS vixl dl xxhash tiny-json)
if (ENABLE_JEMALLOC)
list (APPEND LIBS FEX_jemalloc)
endif()
@@ -296,17 +300,10 @@ install(FILES ${OUTPUT_MAN_NAME_COMPRESS} DESTINATION ${MAN_DIR}/man1)
check_cxx_compiler_flag(-fdiagnostics-color=always GCC_COLOR)
check_cxx_compiler_flag(-fcolor-diagnostics CLANG_COLOR)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
add_dependencies(${Name} CONFIG_INC)
target_link_libraries(${Name} ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
function(AddDefaultOptionsToTarget Name)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES VISIBILITY_INLINES_HIDDEN TRUE)
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
target_include_directories(${Name} PRIVATE IncludePrivate/)
@@ -316,6 +313,7 @@ function(AddObject Name Type)
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
target_compile_definitions(${Name} PRIVATE ${DEFINES})
add_dependencies(${Name} CONFIG_INC)
target_compile_options(${Name}
PRIVATE
@@ -338,20 +336,6 @@ function(AddObject Name Type)
PRIVATE
"-fcolor-diagnostics")
endif()
endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES VISIBILITY_INLINES_HIDDEN TRUE)
set_target_properties(${Name} PROPERTIES VERSION ${FEXCore_VERSION} SOVERSION ${FEXCore_VERSION})
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
target_include_directories(${Name} PUBLIC "${PROJECT_SOURCE_DIR}/include/")
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(${Name}
@@ -363,6 +347,29 @@ function(AddLibrary Name Type)
endif()
endfunction()
# Build FEXCore_Config static library
add_library(FEXCore_Base STATIC ${FEXCORE_BASE_SRCS})
target_link_libraries(FEXCore_Base fmt::fmt tiny-json)
AddDefaultOptionsToTarget(FEXCore_Base)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
target_link_libraries(${Name} FEXCore_Base ${LIBS})
AddDefaultOptionsToTarget(${Name})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} FEXCore_Base ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
AddDefaultOptionsToTarget(${Name})
endfunction()
AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
+17 -1
View File
@@ -416,7 +416,9 @@ namespace JSON {
Meta->Load();
// Do configuration option fix ups after everything is reloaded
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THREADS)) {
{
// Always fix up the number of threads and create the configuration
// Otherwise the application could receive zero as the number of threads
FEX_CONFIG_OPT(Cores, THREADS);
if (Cores == 0) {
// When the number of emulated CPU cores is zero then auto detect
@@ -424,6 +426,20 @@ namespace JSON {
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CORE)) {
// Sanitize Core option
FEX_CONFIG_OPT(Core, CORE);
#if (_M_X86_64)
constexpr uint32_t MaxCoreNumber = 2;
#else
constexpr uint32_t MaxCoreNumber = 1;
#endif
if (Core > MaxCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, std::to_string(FEXCore::Config::CONFIG_IRJIT));
}
}
std::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
+1 -1
View File
@@ -32,7 +32,7 @@
},
"Threads": {
"Type": "uint32",
"Default": "1",
"Default": "0",
"ShortArg": "T",
"Desc": [
"Number of physical hardware threads to tell the process we have.",
+20 -12
View File
@@ -51,7 +51,7 @@ namespace FEXCore::Context {
CTX->CustomExitHandler = std::move(handler);
}
ExitHandler GetExitHandler(FEXCore::Context::Context *CTX) {
ExitHandler GetExitHandler(const FEXCore::Context::Context *CTX) {
return CTX->CustomExitHandler;
}
@@ -71,23 +71,23 @@ namespace FEXCore::Context {
return CTX->RunUntilExit();
}
int GetProgramStatus(FEXCore::Context::Context *CTX) {
int GetProgramStatus(const FEXCore::Context::Context *CTX) {
return CTX->GetProgramStatus();
}
FEXCore::Context::ExitReason GetExitReason(FEXCore::Context::Context *CTX) {
FEXCore::Context::ExitReason GetExitReason(const FEXCore::Context::Context *CTX) {
return CTX->ParentThread->ExitReason;
}
bool IsDone(FEXCore::Context::Context *CTX) {
bool IsDone(const FEXCore::Context::Context *CTX) {
return CTX->IsPaused();
}
void GetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
void GetCPUState(const FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(State, CTX->ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
}
void SetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
void SetCPUState(FEXCore::Context::Context *CTX, const FEXCore::Core::CPUState *State) {
memcpy(CTX->ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
@@ -115,11 +115,11 @@ namespace FEXCore::Context {
}
void RegisterHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterHostSignalHandler(Signal, Func, Required);
CTX->RegisterHostSignalHandler(Signal, std::move(Func), Required);
}
void RegisterFrontendHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterFrontendHostSignalHandler(Signal, Func, Required);
CTX->RegisterFrontendHostSignalHandler(Signal, std::move(Func), Required);
}
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
@@ -162,12 +162,20 @@ namespace FEXCore::Context {
return CTX->CPUID.RunFunction(Function, Leaf);
}
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader) {
CTX->AOTIRLoader = CacheReader;
FEX_DEFAULT_VISIBILITY FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CTX->CPUID.RunFunctionName(Function, Leaf, CPU);
}
void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter) {
CTX->AOTIRWriter = CacheWriter;
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader) {
CTX->SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
CTX->SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(FEXCore::Context::Context *CTX, std::function<void(const std::string&)> CacheRenamer) {
CTX->SetAOTIRRenamer(CacheRenamer);
}
void FinalizeAOTIRCache(FEXCore::Context::Context *CTX) {
+22 -68
View File
@@ -4,6 +4,7 @@
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
@@ -57,38 +58,6 @@ namespace FEXCore::Context {
MODE_SINGLESTEP = 1,
};
struct AOTIRInlineEntry {
uint64_t GuestHash;
uint64_t GuestLength;
/* RAData followed by IRData */
uint8_t InlineData[0];
IR::RegisterAllocationData *GetRAData();
IR::IRListView *GetIRData();
};
struct AOTIRInlineIndexEntry {
uint64_t GuestStart;
uint64_t DataOffset;
};
struct AOTIRInlineIndex {
uint64_t Count;
uint64_t DataBase;
AOTIRInlineIndexEntry Entries[0];
AOTIRInlineEntry *Find(uint64_t GuestStart);
AOTIRInlineEntry *GetInlineEntry(uint64_t DataOffset);
};
struct AOTIRCaptureCacheEntry {
std::unique_ptr<std::ostream> Stream;
std::map<uint64_t, uint64_t> Index;
void AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView *IRList, FEXCore::IR::RegisterAllocationData *RAData);
};
struct Context {
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
@@ -156,32 +125,6 @@ namespace FEXCore::Context {
CustomCPUFactoryType CustomCPUFactory;
FEXCore::Context::ExitHandler CustomExitHandler;
struct AOTIRCacheEntry {
AOTIRInlineIndex *Array;
void *mapping;
size_t size;
};
std::unordered_map<std::string, AOTIRCacheEntry> AOTIRCache;
std::function<int(const std::string&)> AOTIRLoader;
std::function<std::unique_ptr<std::ostream>(const std::string&)> AOTIRWriter;
std::unordered_map<std::string, AOTIRCaptureCacheEntry> AOTIRCaptureCache;
struct AddrToFileEntry {
uint64_t Start;
uint64_t Len;
uint64_t Offset;
std::string fileid;
std::string filename;
void *CachedFileEntry;
bool ContainsCode;
};
using AddrToFileMapType = std::map<uint64_t, AddrToFileEntry>;
AddrToFileMapType AddrToFile;
std::map<std::string, std::string> FilesWithCode;
AddrToFileMapType::iterator FindAddrForFile(uint64_t Entry, uint64_t Length);
#ifdef BLOCKSTATS
std::unique_ptr<FEXCore::BlockSamplingData> BlockData;
#endif
@@ -253,9 +196,6 @@ namespace FEXCore::Context {
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
bool LoadAOTIRCache(int streamfd);
void FinalizeAOTIRCache();
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer);
/**
* @brief Initializes the JIT compilers for the thread
*
@@ -337,6 +277,26 @@ namespace FEXCore::Context {
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
void FinalizeAOTIRCache() {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) {
IRCaptureCache.SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
IRCaptureCache.SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(std::function<void(const std::string&)> CacheRenamer) {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
@@ -362,13 +322,7 @@ namespace FEXCore::Context {
std::mutex ExitMutex;
std::unique_ptr<GdbServer> DebugServer;
std::shared_mutex AOTIRCacheLock;
std::shared_mutex AOTIRCaptureCacheWriteoutLock;
std::atomic<bool> AOTIRCaptureCacheWriteoutFlusing;
std::queue<std::function<void()>> AOTIRCaptureCacheWriteoutQueue;
void AOTIRCaptureCacheWriteoutQueue_Flush();
void AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn);
IR::AOTIRCaptureCache IRCaptureCache;
bool StartPaused = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
@@ -1,4 +1,5 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
@@ -15,37 +16,16 @@ namespace FEXCore::CPU {
#define STATE x28
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
CPU.SetUp();
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
if (SupportsAtomics) {
if (ctx->HostFeatures.SupportsAtomics) {
// Hypervisor can hide this on the c630?
Features.Combine(vixl::CPUFeatures::Feature::kLORegions);
}
SetCPUFeatures(Features);
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _M_ARM_64
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
__asm volatile ("mrs %[ctr], ctr_el0"
: [ctr] "=r"(CTR));
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
#endif
}
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
@@ -58,12 +58,9 @@ const std::array<aarch64::VRegister, 12> RAFPR = {
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(size_t size);
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
vixl::aarch64::CPU CPU;
bool SupportsAtomics{};
bool SupportsRCPC{};
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void SpillStaticRegs(bool FPRs = true, uint32_t SpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t FillMask = ~0U);
@@ -79,9 +76,6 @@ protected:
uint32_t SpillSlots{};
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
@@ -24,6 +24,7 @@ struct X86ContextBackup {
uint64_t GPRs[23];
FEXCore::x86_64::_libc_fpstate FPRState;
uint64_t sa_mask;
bool FaultToTopAndGeneratedException;
// Guest state
int Signal;
@@ -46,6 +47,7 @@ struct ArmContextBackup {
uint32_t FPCR;
__uint128_t FPRs[32];
uint64_t sa_mask;
bool FaultToTopAndGeneratedException;
// Guest state
int Signal;
+345 -64
View File
@@ -5,14 +5,16 @@ desc: Handles presented capability bits for guest cpu
$end_info$
*/
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include "Common/StringConv.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Utils/FileLoading.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include <FEXHeaderUtils/Syscalls.h>
#include "git_version.h"
#include <cstring>
@@ -21,6 +23,66 @@ $end_info$
#endif
namespace FEXCore {
namespace ProductNames {
#ifdef _M_ARM_64
static const char ARM_UNKNOWN[] = "Unknown ARM CPU";
static const char ARM_A57[] = "Cortex-A57";
static const char ARM_A72[] = "Cortex-A72";
static const char ARM_A73[] = "Cortex-A73";
static const char ARM_A75[] = "Cortex-A75";
static const char ARM_A76[] = "Cortex-A76";
static const char ARM_A76AE[] = "Cortex-A76AE";
static const char ARM_V1[] = "Neoverse V1";
static const char ARM_A77[] = "Cortex-A77";
static const char ARM_A78[] = "Cortex-A78";
static const char ARM_A78AE[] = "Cortex-A78AE";
static const char ARM_A78C[] = "Cortex-A78C";
static const char ARM_A710[] = "Cortex-A710";
static const char ARM_X1[] = "Cortex-X1";
static const char ARM_X2[] = "Cortex-X2";
static const char ARM_N1[] = "Neoverse N1";
static const char ARM_N2[] = "Neoverse N2";
static const char ARM_E1[] = "Neoverse E1";
static const char ARM_A35[] = "Cortex-A35";
static const char ARM_A53[] = "Cortex-A53";
static const char ARM_A55[] = "Cortex-A55";
static const char ARM_A65[] = "Cortex-A65";
static const char ARM_A510[] = "Cortex-A510";
static const char ARM_Kryo200[] = "Kryo 2xx";
static const char ARM_Kryo300[] = "Kryo 3xx";
static const char ARM_Kryo400[] = "Kryo 4xx/5xx";
static const char ARM_Kryo200S[] = "Kryo 2xx S";
static const char ARM_Kryo300S[] = "Kryo 3xx S";
static const char ARM_Kryo400S[] = "Kryo 4xx/5xx S";
static const char ARM_Denver[] = "Nvidia Denver";
static const char ARM_Carmel[] = "Nvidia Carmel";
static const char ARM_Firestorm[] = "Apple Firestorm";
static const char ARM_Icestorm[] = "Apple Icestorm";
#else
static const char UNKNOWN[] = "Unknown CPU";
#endif
}
static uint32_t GetCPUID() {
uint32_t CPU{};
FHU::Syscalls::getcpu(&CPU, nullptr);
return CPU;
}
static uint32_t CalculateNumberOfCPUs() {
size_t CPUs = 1;
while(std::filesystem::exists("/sys/devices/system/cpu/cpu" + std::to_string(CPUs))) {
CPUs++;
}
return CPUs;
}
constexpr uint32_t SUPPORTS_AVX = 0;
// #define CPUID_AMD
#ifdef CPUID_AMD
@@ -49,60 +111,241 @@ static uint32_t GetCycleCounterFrequency() {
return Result;
}
static bool GetHostHybridFlag() {
int MaxCPUs = 64;
size_t AllocSize = CPU_ALLOC_SIZE(MaxCPUs);
cpu_set_t *Set = CPU_ALLOC(MaxCPUs);
CPU_ZERO_S(AllocSize, Set);
void CPUIDEmu::SetupHostHybridFlag() {
size_t CPUs = CalculateNumberOfCPUs();
PerCPUData.resize(CPUs);
int Result{};
for (;;) {
Result = sched_getaffinity(0, AllocSize, Set);
if (Result == 0 ||
(Result == -1 && errno != EINVAL)) {
break;
}
MaxCPUs <<= 1;
CPU_FREE(Set);
Set = CPU_ALLOC(MaxCPUs);
AllocSize = CPU_ALLOC_SIZE(MaxCPUs);
CPU_ZERO_S(AllocSize, Set);
}
if (Result != 0) {
return false;
}
int CPUs = CPU_COUNT_S(AllocSize, Set);
bool Hybrid = false;
uint64_t MIDR{};
for (int i = 0; i < CPUs; ++i) {
if (CPU_ISSET_S(i, AllocSize, Set)) {
std::error_code ec{};
std::string MIDRPath = "/sys/devices/system/cpu/cpu" + std::to_string(i) + "/regs/identification/midr_el1";
if (std::filesystem::exists(MIDRPath, ec)) {
std::vector<char> Data{};
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
if (FEXCore::StrConv::Conv(&Data.at(0), &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
break;
}
MIDR = NewMIDR;
for (size_t i = 0; i < CPUs; ++i) {
std::error_code ec{};
std::string MIDRPath = "/sys/devices/system/cpu/cpu" + std::to_string(i) + "/regs/identification/midr_el1";
if (std::filesystem::exists(MIDRPath, ec)) {
std::vector<char> Data{};
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
if (FEXCore::StrConv::Conv(&Data.at(0), &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
}
// Truncate to 32-bits, top 32-bits are all reserved in MIDR
PerCPUData[i].ProductName = ProductNames::ARM_UNKNOWN;
PerCPUData[i].MIDR = NewMIDR;
MIDR = NewMIDR;
}
}
}
}
CPU_FREE(Set);
return Hybrid;
struct CPUMIDR {
uint8_t Implementer;
uint16_t Part;
bool DefaultBig; // Defaults to a big core
const char *ProductName{};
};
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 35> CPUMIDRs = {{
// Typically big CPU cores
{0x61, 0x023, 1, ProductNames::ARM_Firestorm}, // Apple M1 Firestorm
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd4b, 1, ProductNames::ARM_A78C}, // A78C
{0x41, 0xd4a, 1, ProductNames::ARM_E1}, // E1
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd48, 1, ProductNames::ARM_X2}, // X2
{0x41, 0xd47, 1, ProductNames::ARM_A710}, // A710
{0x41, 0xd44, 1, ProductNames::ARM_X1}, // X1
{0x41, 0xd42, 1, ProductNames::ARM_A78AE}, // A78AE
{0x41, 0xd41, 1, ProductNames::ARM_A78}, // A78
{0x41, 0xd40, 1, ProductNames::ARM_V1}, // V1
{0x41, 0xd0e, 1, ProductNames::ARM_A76AE}, // A76AE
{0x41, 0xd0d, 1, ProductNames::ARM_A77}, // A77
{0x41, 0xd0c, 1, ProductNames::ARM_N1}, // N1
{0x41, 0xd0b, 1, ProductNames::ARM_A76}, // A76
{0x51, 0x804, 1, ProductNames::ARM_Kryo400}, // Kryo 4xx Gold (A76 based)
{0x41, 0xd0a, 1, ProductNames::ARM_A75}, // A75
{0x51, 0x802, 1, ProductNames::ARM_Kryo300}, // Kryo 3xx Gold (A75 based)
{0x41, 0xd09, 1, ProductNames::ARM_A73}, // A73
{0x51, 0x800, 1, ProductNames::ARM_Kryo200}, // Kryo 2xx Gold (A73 based)
{0x41, 0xd08, 1, ProductNames::ARM_A72}, // A72
{0x4e, 0x004, 1, ProductNames::ARM_Carmel}, // Carmel
// Denver rated above A57 to match TX2 weirdness
{0x4e, 0x003, 1, ProductNames::ARM_Denver}, // Denver
{0x41, 0xd07, 1, ProductNames::ARM_A57}, // A57
// Typically Little CPU cores
{0x61, 0x022, 0, ProductNames::ARM_Icestorm}, // Apple M1 Icestorm
{0x41, 0xd46, 0, ProductNames::ARM_A510}, // A510
{0x41, 0xd06, 0, ProductNames::ARM_A65}, // A65
{0x41, 0xd05, 0, ProductNames::ARM_A55}, // A55
{0x51, 0x805, 0, ProductNames::ARM_Kryo400S}, // Kryo 4xx/5xx Silver (A55 based)
{0x51, 0x803, 0, ProductNames::ARM_Kryo300S}, // Kryo 3xx Silver (A55 based)
{0x41, 0xd03, 0, ProductNames::ARM_A53}, // A53
{0x51, 0x801, 0, ProductNames::ARM_Kryo200S}, // Kryo 2xx Silver (A53 based)
{0x41, 0xd04, 0, ProductNames::ARM_A35}, // A35
{0x41, 0, 0, ProductNames::ARM_UNKNOWN}, // Invalid CPU or Apple CPU inside Parallels VM
{0x0, 0, 0, ProductNames::ARM_UNKNOWN}, // Invalid starting point is lowest ranked
}};
auto FindDefinedMIDR = [](uint32_t MIDR) -> const CPUMIDR* {
uint8_t Implementer = MIDR >> 24;
uint16_t Part = (MIDR >> 4) & 0xFFF;
for (auto &MIDROption : CPUMIDRs) {
if (MIDROption.Implementer == Implementer &&
MIDROption.Part == Part) {
return &MIDROption;
}
}
return nullptr;
};
if (Hybrid) {
// Walk the MIDRs and calculate big little designs
std::vector<const CPUMIDR*> BigCores;
std::vector<const CPUMIDR*> LittleCores;
// Separate CPU cores out to big or little selected
for (size_t i = 0; i < CPUs; ++i) {
uint32_t MIDR = PerCPUData[i].MIDR;
auto MIDROption = FindDefinedMIDR(MIDR);
if (MIDROption) {
// Found one
if (MIDROption->DefaultBig) {
BigCores.emplace_back(MIDROption);
}
else {
LittleCores.emplace_back(MIDROption);
}
}
else {
// If we didn't insert this MIDR then claim it is a little core.
LittleCores.emplace_back(&CPUMIDRs.back());
}
}
if (LittleCores.empty()) {
// If we only ended up with big cores then we need to move some to be little cores
uint32_t LowestMIDR = ~0U;
uint32_t LowestMIDRIdx = 0;
// Walk all the big cores
for (size_t i = 0; i < BigCores.size(); ++i) {
uint8_t Implementer = BigCores[i]->Implementer;
uint16_t Part = BigCores[i]->Part;
// Walk our list of CPUMIDRs to find the most little core
for (size_t j = LowestMIDRIdx; j < CPUMIDRs.size(); ++j) {
auto &MIDROption = CPUMIDRs[i];
if ((MIDROption.Implementer == Implementer &&
MIDROption.Part == Part) ||
(MIDROption.Implementer == 0 &&
MIDROption.Part == 0)) {
LowestMIDRIdx = j;
LowestMIDR = MIDR;
break;
}
}
}
// Now we WILL have found a big core to demote to little status
// Demote them
std::erase_if(BigCores, [&LittleCores, LowestMIDR](auto *Entry) {
// Demote by erase copy to little array
uint8_t Implementer = LowestMIDR >> 24;
uint16_t Part = (LowestMIDR >> 4) & 0xFFF;
if (Entry->Implementer == Implementer &&
Entry->Part == Part) {
// Add it to the BigCore list
LittleCores.emplace_back(Entry);
return true;
}
return false;
});
}
if (BigCores.empty()) {
// We never found a CPU core we understand
// Grab the first core, consider it as little, move everything else to Big
uint32_t LittleMIDR = PerCPUData[0].MIDR;
// Now walk the little cores and move them to Big if they don't match
std::erase_if(LittleCores, [&BigCores, LittleMIDR](auto *Entry) {
// You're promoted now
uint8_t Implementer = LittleMIDR >> 24;
uint16_t Part = (LittleMIDR >> 4) & 0xFFF;
if (Entry->Implementer != Implementer ||
Entry->Part != Part) {
// Add it to the BigCore list
BigCores.emplace_back(Entry);
return true;
}
return false;
});
}
// Now walk the per CPU data one more time and set if it is big or little
for (auto &Data : PerCPUData) {
uint8_t Implementer = Data.MIDR >> 24;
uint16_t Part = (Data.MIDR >> 4) & 0xFFF;
bool FoundBig{};
const CPUMIDR *MIDR{};
for (auto Big : BigCores) {
if (Big->Implementer == Implementer &&
Big->Part == Part) {
FoundBig = true;
MIDR = Big;
break;
}
}
if (!FoundBig) {
for (auto Little : LittleCores) {
if (Little->Implementer == Implementer &&
Little->Part == Part) {
MIDR = Little;
break;
}
}
}
Data.IsBig = FoundBig;
if (MIDR) {
Data.ProductName = MIDR->ProductName ?: ProductNames::ARM_UNKNOWN;
}
else {
Data.ProductName = ProductNames::ARM_UNKNOWN;
}
}
}
else {
// If we aren't hybrid then just claim everything is big
for (size_t i = 0; i < CPUs; ++i) {
uint32_t MIDR = PerCPUData[i].MIDR;
auto MIDROption = FindDefinedMIDR(MIDR);
PerCPUData[i].IsBig = true;
if (MIDROption) {
PerCPUData[i].ProductName = MIDROption->ProductName ?: ProductNames::ARM_UNKNOWN;
}
else {
PerCPUData[i].ProductName = ProductNames::ARM_UNKNOWN;
}
}
}
}
#else
@@ -119,16 +362,21 @@ static uint32_t GetCycleCounterFrequency() {
return 0;
}
static bool GetHostHybridFlag() {
void CPUIDEmu::SetupHostHybridFlag() {
uint32_t eax, ebx, ecx, edx;
__cpuid(0, eax, ebx, ecx, edx);
if (eax >= 0x7) {
__cpuid(0x7, eax, ebx, ecx, edx);
// Bit 15 of edx claims hybrid CPU
return (edx & (1U << 15)) != 0;
Hybrid = (edx & (1U << 15)) != 0;
}
return false;
size_t CPUs = CalculateNumberOfCPUs();
PerCPUData.resize(CPUs);
for (size_t i = 0; i < CPUs; ++i) {
PerCPUData[i].IsBig = true;
PerCPUData[i].ProductName = ProductNames::UNKNOWN;
}
}
#endif
@@ -386,7 +634,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 5) | // AVX2 support
(1 << 6) | // FPU data pointer updated only on exception
(1 << 7) | // SMEP support
(0 << 8) | // BMI2
(1 << 8) | // BMI2
(0 << 9) | // Enhanced REP MOVSB/STOSB
(1 << 10) | // INVPCID for system software control of process-context
(0 << 11) | // Restricted transactional memory
@@ -549,6 +797,18 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
if (Hybrid) {
uint32_t CPU = GetCPUID();
auto &Data = PerCPUData[CPU];
// 0x40 is a big CPU
// 0x20 is a little CPU
Res.eax |= (Data.IsBig ? 0x40 : 0x20) << 24;
}
return Res;
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
@@ -636,7 +896,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) {
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(0 << 27) | // RDTSCP
(1 << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(0 << 30) | // 3DNow! Extensions
@@ -644,27 +904,44 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) {
return Res;
}
constexpr char ProcessorBrand[48] = {
constexpr char ProcessorBrand[32] = {
GIT_DESCRIBE_STRING
"\0"
};
constexpr ssize_t DESCRIBE_STR_SIZE = std::char_traits<char>::length(GIT_DESCRIBE_STRING);
static_assert(DESCRIBE_STR_SIZE < 32);
//Processor brand string
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
memcpy(&Res, &ProcessorBrand[0], sizeof(FEXCore::CPUID::FunctionResults));
return Res;
return Function_8000_0002h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
memcpy(&Res, &ProcessorBrand[16], sizeof(FEXCore::CPUID::FunctionResults));
return Res;
return Function_8000_0003h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf) {
return Function_8000_0004h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
memcpy(&Res, &ProcessorBrand[32], sizeof(FEXCore::CPUID::FunctionResults));
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[0], std::min(16L, DESCRIBE_STR_SIZE));
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[16], std::max(0L, DESCRIBE_STR_SIZE - 16));
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
auto &Data = PerCPUData[CPU];
memcpy(&Res, Data.ProductName, std::min(strlen(Data.ProductName), sizeof(FEXCore::CPUID::FunctionResults)));
return Res;
}
@@ -757,7 +1034,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) {
Res.ebx =
(0 << 2) | // XSaveErPtr: Saving and restoring error pointers
(0 << 1) | // IRPerf: Instructions retired count support
(0 << 0); // CLZERO support
(CTX->HostFeatures.SupportsCLZERO << 0); // CLZERO support
uint32_t CoreCount = Cores() - 1;
Res.ecx =
@@ -920,6 +1197,10 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// 0x16: Processor frequency information
// 0x17: SoC vendor attribute enumeration
// 0x1A: Hybrid Information Sub-leaf
#ifndef CPUID_AMD
RegisterFunction(0x1A, &CPUIDEmu::Function_1Ah);
#endif
// Largest extended function number
RegisterFunction(0x8000'0000, &CPUIDEmu::Function_8000_0000h);
// Processor vendor
@@ -960,7 +1241,7 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// 0x8000'001F: AMD Secure Encryption
// Setup some state tracking
Hybrid = GetHostHybridFlag();
SetupHostHybridFlag();
}
}
+31
View File
@@ -23,6 +23,10 @@ private:
constexpr static uint32_t CPUID_VENDOR_AMD3 = 0x444D4163; // "cAMD"
public:
// X86 cacheline size effectively has to be hardcoded to 64
// if we report anything differently then applications are likely to break
constexpr static uint64_t CACHELINE_SIZE = 64;
void Init(FEXCore::Context::Context *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, uint32_t Leaf) {
@@ -34,6 +38,16 @@ public:
return (this->*Handler->second)(Leaf);
}
FEXCore::CPUID::FunctionResults RunFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
if (Function == 0x8000'0002U)
return Function_8000_0002h(Leaf, CPU % PerCPUData.size());
else if (Function == 0x8000'0003U)
return Function_8000_0003h(Leaf, CPU % PerCPUData.size());
else
return Function_8000_0004h(Leaf, CPU % PerCPUData.size());
}
private:
FEXCore::Context::Context *CTX;
bool Hybrid{};
@@ -45,6 +59,14 @@ private:
}
std::unordered_map<uint32_t, FunctionHandler> FunctionHandlers;
struct CPUData {
const char *ProductName{};
#ifdef _M_ARM_64
uint32_t MIDR{};
#endif
bool IsBig{};
};
std::vector<CPUData> PerCPUData{};
// Functions
FEXCore::CPUID::FunctionResults Function_0h(uint32_t Leaf);
@@ -55,11 +77,17 @@ private:
FEXCore::CPUID::FunctionResults Function_07h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_0Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0002h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0003h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0004h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0002h(uint32_t Leaf, uint32_t CPU);
FEXCore::CPUID::FunctionResults Function_8000_0003h(uint32_t Leaf, uint32_t CPU);
FEXCore::CPUID::FunctionResults Function_8000_0004h(uint32_t Leaf, uint32_t CPU);
FEXCore::CPUID::FunctionResults Function_8000_0005h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0006h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0007h(uint32_t Leaf);
@@ -68,5 +96,8 @@ private:
FEXCore::CPUID::FunctionResults Function_8000_0019h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_001Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_Reserved(uint32_t Leaf);
void SetupHostHybridFlag();
};
}
+112 -426
View File
@@ -41,6 +41,7 @@ $end_info$
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <algorithm>
#include <array>
@@ -50,6 +51,7 @@ $end_info$
#include <cstdint>
#include <filesystem>
#include <functional>
#include <fstream>
#include <map>
#include <memory>
#include <mutex>
@@ -61,7 +63,6 @@ $end_info$
#include <string.h>
#include <string>
#include <string_view>
#include <sstream>
#include <sys/mman.h>
#include <sys/stat.h>
#include <sys/syscall.h>
@@ -141,54 +142,8 @@ std::string_view const& GetGRegName(unsigned Reg) {
} // namespace FEXCore::Core
namespace FEXCore::Context {
void Context::AOTIRCaptureCacheWriteoutQueue_Flush() {
{
std::shared_lock lk{AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
for (;;) {
AOTIRCaptureCacheWriteoutLock.lock();
std::function<void()> fn = std::move(AOTIRCaptureCacheWriteoutQueue.front());
bool MaybeEmpty = false;
AOTIRCaptureCacheWriteoutQueue.pop();
MaybeEmpty = AOTIRCaptureCacheWriteoutQueue.size() == 0;
AOTIRCaptureCacheWriteoutLock.unlock();
fn();
if (MaybeEmpty) {
std::shared_lock lk{AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
}
LOGMAN_MSG_A_FMT("Must never get here");
}
void Context::AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn) {
bool Flush = false;
{
std::unique_lock lk{AOTIRCaptureCacheWriteoutLock};
AOTIRCaptureCacheWriteoutQueue.push(fn);
if (AOTIRCaptureCacheWriteoutQueue.size() > 10000) {
Flush = true;
}
}
bool test_val = false;
if (Flush && AOTIRCaptureCacheWriteoutFlusing.compare_exchange_strong(test_val, true)) {
AOTIRCaptureCacheWriteoutQueue_Flush();
}
}
Context::Context() {
Context::Context()
: IRCaptureCache {this} {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
#endif
@@ -211,10 +166,6 @@ namespace FEXCore::Context {
}
Threads.clear();
}
for (auto &Mod: AOTIRCache) {
FEXCore::Allocator::munmap(Mod.second.mapping, Mod.second.size);
}
}
static FEXCore::Core::CPUState CreateDefaultCPUState() {
@@ -323,7 +274,7 @@ namespace FEXCore::Context {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Pause);
if (Thread->RunningEvents.Running.load()) {
// Only attempt to stop this thread if it is running
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
}
@@ -386,7 +337,7 @@ namespace FEXCore::Context {
}
void Context::Stop(bool IgnoreCurrentThread) {
pid_t tid = gettid();
pid_t tid = FHU::Syscalls::gettid();
FEXCore::Core::InternalThreadState* CurrentThread{};
// Tell all the threads that they should stop
@@ -426,20 +377,25 @@ namespace FEXCore::Context {
void Context::StopThread(FEXCore::Core::InternalThreadState *Thread) {
if (Thread->RunningEvents.Running.exchange(false)) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Stop);
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
void Context::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
if (Thread->RunningEvents.Running.load()) {
Thread->SignalReason.store(Event);
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
FEXCore::Context::ExitReason Context::RunUntilExit() {
if(!StartPaused)
Run();
if(!StartPaused) {
// We will only have one thread at this point, but just in case run notify everything
std::lock_guard lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->StartRunning.NotifyAll();
}
}
ExecutionThread(ParentThread);
while(true) {
@@ -474,7 +430,6 @@ namespace FEXCore::Context {
};
LocalLoader->AddIR(IRHandler);
}
struct ExecutionThreadHandler {
@@ -502,7 +457,7 @@ namespace FEXCore::Context {
void Context::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = ::gettid();
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
SignalDelegation->RegisterTLSState(Thread);
ThunkHandler->RegisterTLSState(Thread);
@@ -665,6 +620,58 @@ namespace FEXCore::Context {
}
}
static void IRDumper(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
const auto DumpIRStr = Thread->CTX->Config.DumpIR();
if (DumpIRStr =="stderr") {
f = stderr;
}
else if (DumpIRStr =="stdout") {
f = stdout;
}
else {
const auto fileName = fmt::format("{}/{:x}{}", DumpIRStr, GuestRIP, RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.c_str(), "w");
CloseAfter = true;
}
if (f) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fmt::print(f,"IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
if (CloseAfter) {
fclose(f);
}
}
};
static void ValidateIR(FEXCore::Core::InternalThreadState *Thread) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
static auto compaction = IR::CreateIRCompaction();
compaction->Run(Thread->OpDispatcher.get());
auto NewIR = Thread->OpDispatcher->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
auto reparsed = IR::Parse(&out);
if (reparsed == nullptr) {
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
} else {
std::stringstream out2;
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
LogMan::Msg::IFmt("one:\n {}", out.str());
LogMan::Msg::IFmt("two:\n {}", out2.str());
LOGMAN_MSG_A_FMT("Parsed IR doesn't match\n");
}
}
}
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
uint8_t const *GuestCode{};
GuestCode = reinterpret_cast<uint8_t const*>(GuestRIP);
@@ -734,7 +741,7 @@ namespace FEXCore::Context {
else {
if (Thread->OpDispatcher->HandledLock != IsLocked) {
HadDispatchError = true;
LogMan::Msg::EFmt("Missing LOCK HANDLER at 0x{:x}{{'{}'}}", Block.Entry + BlockInstructionsLength, TableInfo->Name);
LogMan::Msg::EFmt("Missing LOCK HANDLER at 0x{:x}{{'{}'}}", Block.Entry + BlockInstructionsLength, TableInfo->Name ?: "UND");
}
BlockInstructionsLength += DecodedInfo->InstSize;
TotalInstructionsLength += DecodedInfo->InstSize;
@@ -771,73 +778,32 @@ namespace FEXCore::Context {
Thread->OpDispatcher->Finalize();
const auto IRDumper = [Thread, GuestRIP](IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
const auto DumpIRStr = Thread->CTX->Config.DumpIR();
if (DumpIRStr =="stderr") {
f = stderr;
}
else if (DumpIRStr =="stdout") {
f = stdout;
}
else {
const auto fileName = fmt::format("{}/{:x}{}", DumpIRStr, GuestRIP, RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.c_str(), "w");
CloseAfter = true;
// Debug
{
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread, GuestRIP, nullptr);
}
if (f) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fmt::print(f,"IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
if (CloseAfter) {
fclose(f);
}
}
};
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(nullptr);
}
if (Thread->CTX->Config.ValidateIRarser) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
static auto compaction = IR::CreateIRCompaction();
compaction->Run(Thread->OpDispatcher.get());
auto NewIR = Thread->OpDispatcher->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
auto reparsed = IR::Parse(&out);
if (reparsed == nullptr) {
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
} else {
std::stringstream out2;
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
LogMan::Msg::IFmt("one:\n {}", out.str());
LogMan::Msg::IFmt("two:\n {}", out2.str());
LOGMAN_MSG_A_FMT("Parsed IR doesn't match\n");
}
if (Thread->CTX->Config.ValidateIRarser) {
ValidateIR(Thread);
}
}
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(Thread->OpDispatcher.get());
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
// Debug
{
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread, GuestRIP, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
if (Thread->OpDispatcher->ShouldDump) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
LogMan::Msg::IFmt("IR 0x{:x}:\n{}\n@@@@@\n", GuestRIP, out.str());
if (Thread->OpDispatcher->ShouldDump) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
LogMan::Msg::IFmt("IR 0x{:x}:\n{}\n@@@@@\n", GuestRIP, out.str());
}
}
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
@@ -855,63 +821,6 @@ namespace FEXCore::Context {
};
}
AOTIRInlineEntry *AOTIRInlineIndex::GetInlineEntry(uint64_t DataOffset) {
uintptr_t This = (uintptr_t)this;
return (AOTIRInlineEntry*)(This + DataBase + DataOffset);
}
AOTIRInlineEntry *AOTIRInlineIndex::Find(uint64_t GuestStart) {
ssize_t l = 0;
ssize_t r = Count - 1;
while (l <= r) {
size_t m = l + (r - l) / 2;
if (Entries[m].GuestStart == GuestStart)
return GetInlineEntry(Entries[m].DataOffset);
else if (Entries[m].GuestStart < GuestStart)
l = m + 1;
else
r = m - 1;
}
return nullptr;
}
IR::RegisterAllocationData *AOTIRInlineEntry::GetRAData() {
return (IR::RegisterAllocationData *)InlineData;
}
IR::IRListView *AOTIRInlineEntry::GetIRData() {
auto RAData = GetRAData();
auto Offset = RAData->Size(RAData->MapCount);
return (IR::IRListView *)&InlineData[Offset];
}
void AOTIRCaptureCacheEntry::AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView *IRList, FEXCore::IR::RegisterAllocationData *RAData) {
auto Inserted = Index.emplace(GuestRIP, Stream->tellp());
if (Inserted.second) {
//GuestHash
Stream->write((const char*)&Hash, sizeof(Hash));
//GuestLength
Stream->write((const char*)&Length, sizeof(Length));
// RAData (inline)
// In file, IsShared is always set
auto Shared = RAData->IsShared;
RAData->IsShared = true;
Stream->write((const char*)RAData, RAData->Size(RAData->MapCount));
RAData->IsShared = Shared;
// IRData (inline)
IRList->Serialize(*Stream);
}
}
Context::CompileCodeResult Context::CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
FEXCore::IR::IRListView *IRList {};
FEXCore::Core::DebugData *DebugData {};
@@ -935,55 +844,17 @@ namespace FEXCore::Context {
GeneratedIR = false;
}
// AOT IR bookkeeping and cache
{
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
if (!file->second.ContainsCode) {
file->second.ContainsCode = true;
FilesWithCode[file->second.fileid] = file->second.filename;
}
}
}
if (IRList == nullptr && Config.AOTIRLoad) {
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
auto Mod = (AOTIRInlineIndex*)file->second.CachedFileEntry;
if (Mod == nullptr) {
file->second.CachedFileEntry = Mod = AOTIRCache[file->second.fileid].Array;
}
if (Mod != nullptr)
{
auto AOTEntry = Mod->Find(GuestRIP - file->second.Start + file->second.Offset);
if (AOTEntry) {
// verify hash
auto MappedStart = GuestRIP;
auto hash = XXH3_64bits((void*)MappedStart, AOTEntry->GuestLength);
if (hash == AOTEntry->GuestHash) {
IRList = AOTEntry->GetIRData();
//LogMan::Msg::DFmt("using {} + {:x} -> {:x}\n", file->second.fileid, AOTEntry->first, GuestRIP);
RAData = AOTEntry->GetRAData();;
DebugData = new FEXCore::Core::DebugData();
StartAddr = MappedStart;
Length = AOTEntry->GuestLength;
GeneratedIR = true;
} else {
LogMan::Msg::IFmt("AOTIR: hash check failed {:x}\n", MappedStart);
}
} else {
//LogMan::Msg::IFmt("AOTIR: Failed to find {:x}, {:x}, {}\n", GuestRIP, GuestRIP - file->second.Start + file->second.Offset, file->second.fileid);
}
}
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(GuestRIP, IRList);
if (_GeneratedIR) {
// Setup pointers to internal structures
IRList = IRCopy;
RAData = RACopy;
DebugData = DebugDataCopy;
StartAddr = _StartAddr;
Length = _Length;
GeneratedIR = _GeneratedIR;
}
}
@@ -1024,114 +895,6 @@ namespace FEXCore::Context {
};
}
static bool readAll(int fd, void *data, size_t size) {
int rv = read(fd, data, size);
if (rv != size)
return false;
else
return true;
}
bool Context::LoadAOTIRCache(int streamfd) {
uint64_t tag;
if (!readAll(streamfd, (char*)&tag, sizeof(tag)) || tag != 0xDEADBEEFC0D30004)
return false;
std::string Module;
uint64_t ModSize;
uint64_t IndexSize;
lseek(streamfd, -sizeof(ModSize), SEEK_END);
if (!readAll(streamfd, (char*)&ModSize, sizeof(ModSize)))
return false;
Module.resize(ModSize);
lseek(streamfd, -sizeof(ModSize) - ModSize, SEEK_END);
if (!readAll(streamfd, (char*)&Module[0], Module.size()))
return false;
lseek(streamfd, -sizeof(ModSize) - ModSize - sizeof(IndexSize), SEEK_END);
if (!readAll(streamfd, (char*)&IndexSize, sizeof(IndexSize)))
return false;
struct stat fileinfo;
if (fstat(streamfd, &fileinfo) < 0)
return false;
size_t Size = (fileinfo.st_size + 4095) & ~4095;
size_t IndexOffset = fileinfo.st_size - IndexSize -sizeof(ModSize) - ModSize - sizeof(IndexSize);
void *FilePtr = FEXCore::Allocator::mmap(nullptr, Size, PROT_READ, MAP_SHARED, streamfd, 0);
if (FilePtr == MAP_FAILED)
return false;
auto Array = (AOTIRInlineIndex *)((char*)FilePtr + IndexOffset);
AOTIRCache.insert({Module, {Array, FilePtr, Size}});
LogMan::Msg::DFmt("AOTIR: Module {} has {} functions", Module, Array->Count);
return true;
}
void Context::WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
std::shared_lock lk(AOTIRCacheLock);
for( const auto &File: FilesWithCode) {
Writer(File.first, File.second);
}
}
void Context::FinalizeAOTIRCache() {
AOTIRCaptureCacheWriteoutQueue_Flush();
std::unique_lock lk(AOTIRCacheLock);
for (auto& [String, Entry] : AOTIRCaptureCache) {
if (!Entry.Stream) {
continue;
}
const auto ModSize = String.size();
auto &stream = Entry.Stream;
// pad to 32 bytes
constexpr char Zero = 0;
while(stream->tellp() & 31)
stream->write(&Zero, 1);
// AOTIRInlineIndex
const auto FnCount = Entry.Index.size();
const size_t DataBase = -stream->tellp();
stream->write((const char*)&FnCount, sizeof(FnCount));
stream->write((const char*)&DataBase, sizeof(DataBase));
for (const auto& [GuestStart, DataOffset] : Entry.Index) {
//AOTIRInlineIndexEntry
// GuestStart
stream->write((const char*)&GuestStart, sizeof(GuestStart));
// DataOffset
stream->write((const char*)&DataOffset, sizeof(DataOffset));
}
// End of file header
const auto IndexSize = FnCount * sizeof(AOTIRInlineIndexEntry) + sizeof(DataBase) + sizeof(FnCount);
stream->write((const char*)&IndexSize, sizeof(IndexSize));
stream->write(String.c_str(), ModSize);
stream->write((const char*)&ModSize, sizeof(ModSize));
}
}
void Context::CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto NewBlock = CompileBlock(Frame, GuestRIP);
@@ -1143,19 +906,6 @@ namespace FEXCore::Context {
}
}
Context::AddrToFileMapType::iterator Context::FindAddrForFile(uint64_t Entry, uint64_t Length) {
// Thread safety here! We are returning an iterator to the map object
// This needs the AOTIRCacheLock locked prior to coming in to the function
auto file = AddrToFile.lower_bound(Entry);
if (file != AddrToFile.begin()) {
--file;
if (file->second.Start <= Entry && (file->second.Start + file->second.Len) >= (Entry + Length)) {
return file;
}
}
return AddrToFile.end();
}
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
@@ -1227,56 +977,19 @@ namespace FEXCore::Context {
}
}
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || Config.LibraryJITNaming()) {
std::shared_lock lk(AOTIRCacheLock);
auto file = FindAddrForFile(StartAddr, Length);
// Only go down this path if we actually found a library region
if (file != AddrToFile.end()) {
if (DebugData && Config.LibraryJITNaming()) {
Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, file->second.filename);
}
// Add to AOT cache if aot generation is enabled
if (GeneratedIR && RAData &&
(Config.AOTIRCapture() || Config.AOTIRGenerate())) {
auto hash = XXH3_64bits((void*)StartAddr, Length);
auto LocalRIP = GuestRIP - file->second.Start + file->second.Offset;
auto LocalStartAddr = StartAddr - file->second.Start + file->second.Offset;
auto fileid = file->second.fileid;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRList, RAData, fileid]() {
auto *AotFile = &AOTIRCaptureCache[fileid];
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(fileid);
uint64_t tag = 0xDEADBEEFC0D30004;
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
});
if (Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
if (DecrementRefCount)
--Thread->CompileBlockReentrantRefCount;
Thread->CPUBackend->ClearCache();
return (uintptr_t)CodePtr;
}
}
}
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
}
if (IRCaptureCache.PostCompileCode(
Thread,
CodePtr,
GuestRIP,
StartAddr,
Length,
RAData,
IRList,
DebugData,
GeneratedIR,
DecrementRefCount)) {
// Early exit
return (uintptr_t)CodePtr;
}
if (DecrementRefCount)
@@ -1406,38 +1119,11 @@ namespace FEXCore::Context {
}
void Context::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// TODO: Support overlapping maps and region splitting
auto base_filename = std::filesystem::path(filename).filename().string();
if (!base_filename.empty()) {
auto filename_hash = XXH3_64bits(filename.c_str(), filename.size());
auto fileid = base_filename + "-" + std::to_string(filename_hash) + "-";
// append optimization flags to the fileid
fileid += (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) ? "S" : "s";
fileid += Config.TSOEnabled ? "T" : "t";
fileid += Config.ABILocalFlags ? "L" : "l";
fileid += Config.ABINoPF ? "p" : "P";
std::unique_lock lk(AOTIRCacheLock);
AddrToFile.insert({ Base, { Base, Size, Offset, fileid, filename, nullptr, false} });
if (Config.AOTIRLoad && !AOTIRCache.contains(fileid) && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
if (streamfd != -1) {
LoadAOTIRCache(streamfd);
close(streamfd);
}
}
}
IRCaptureCache.AddNamedRegion(Base, Size, Offset, filename);
}
void Context::RemoveNamedRegion(uintptr_t Base, uintptr_t Size) {
std::unique_lock lk(AOTIRCacheLock);
// TODO: Support partial removing
AddrToFile.erase(Base);
IRCaptureCache.RemoveNamedRegion(Base, Size);
}
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
@@ -11,6 +11,8 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <array>
#include <bit>
#include <cmath>
@@ -37,7 +39,7 @@ static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread), Arm64Emitter(MAX_DISPATCHER_CODE_SIZE) {
: Dispatcher(ctx, Thread), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE) {
SRAEnabled = config.StaticRegisterAssignment;
SetAllowAssembler(true);
@@ -337,6 +339,31 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
hlt(0);
}
{
// Guest Overflow handler
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
LoadConstant(x0, reinterpret_cast<uint64_t>(&SynchronousFaultData));
LoadConstant(w1, 1);
strb(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)));
LoadConstant(w1, X86State::X86_TRAPNO_OF);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)));
LoadConstant(w1, 0x80);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)));
LoadConstant(x1, 0);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)));
// hlt/udf = SIGILL
// brk = SIGTRAP
// ??? = SIGSEGV
// Force a SIGSEGV by loading zero
ldr(x1, MemOperand(x1));
}
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
@@ -428,7 +455,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
GetBuffer()->SetExecutable();
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(::gettid());
std::string Name = "Dispatch_" + std::to_string(FHU::Syscalls::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr), Name);
}
if (CTX->Config.GlobalJITNaming()) {
@@ -81,6 +81,10 @@ ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, vo
Context->UContextLocation = 0;
Context->SigInfoLocation = 0;
// Store fault to top status and then reset it
Context->FaultToTopAndGeneratedException = SynchronousFaultData.FaultToTopAndGeneratedException;
SynchronousFaultData.FaultToTopAndGeneratedException = false;
return Context;
}
@@ -118,7 +122,8 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
[[maybe_unused]] auto *guest_siginfo = reinterpret_cast<siginfo_t*>(Context->SigInfoLocation);
// If the guest modified the RIP then we need to take special precautions here
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP]) {
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP] ||
Context->FaultToTopAndGeneratedException) {
// Hack! Go back to the top of the dispatcher top
// This is only safe inside the JIT rather than anything outside of it
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
@@ -127,6 +132,14 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP];
// XXX: Full context setting
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
Frame->State.flags[1] = 1;
Frame->State.flags[9] = 1;
#define COPY_REG(x) \
Frame->State.gregs[X86State::REG_##x] = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_##x];
COPY_REG(R8);
@@ -146,13 +159,29 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
FEXCore::x86_64::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86_64::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
// Copy float registers
memcpy(Frame->State.mm, fpstate->_st, sizeof(Frame->State.mm));
memcpy(Frame->State.xmm, fpstate->_xmm, sizeof(Frame->State.xmm));
// FCW store default
Frame->State.FCW = fpstate->fcw;
Frame->State.FTW = fpstate->ftw;
// Deconstruct FSW
Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] = (fpstate->fsw >> 8) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] = (fpstate->fsw >> 9) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C2_LOC] = (fpstate->fsw >> 10) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C3_LOC] = (fpstate->fsw >> 14) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] = (fpstate->fsw >> 11) & 0b111;
}
}
else {
auto *guest_uctx = reinterpret_cast<FEXCore::x86::ucontext_t*>(Context->UContextLocation);
[[maybe_unused]] auto *guest_siginfo = reinterpret_cast<FEXCore::x86::siginfo_t*>(Context->SigInfoLocation);
// If the guest modified the RIP then we need to take special precautions here
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP]) {
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP] ||
Context->FaultToTopAndGeneratedException) {
// Hack! Go back to the top of the dispatcher top
// This is only safe inside the JIT rather than anything outside of it
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
@@ -160,17 +189,54 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
// XXX: Full context setting
// First 32-bytes of flags is EFLAGS broken out
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
Frame->State.flags[1] = 1;
Frame->State.flags[9] = 1;
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP];
Frame->State.cs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS];
Frame->State.ds = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS];
Frame->State.es = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES];
Frame->State.fs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS];
Frame->State.gs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_GS];
Frame->State.ss = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_SS];
#define COPY_REG(x) \
Frame->State.gregs[X86State::REG_##x] = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_##x];
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
FEXCore::x86::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
// Copy float registers
for (size_t i = 0; i < 8; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
memcpy(&Frame->State.mm[i], &fpstate->_st[i], 10);
}
// Extended XMM state
memcpy(fpstate->_xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
// FCW store default
Frame->State.FCW = fpstate->fcw;
Frame->State.FTW = fpstate->ftw;
// Deconstruct FSW
Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] = (fpstate->fsw >> 8) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] = (fpstate->fsw >> 9) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C2_LOC] = (fpstate->fsw >> 10) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C3_LOC] = (fpstate->fsw >> 14) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] = (fpstate->fsw >> 11) & 0b111;
}
}
}
@@ -267,7 +333,10 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// This can also happen if we have put a signal on hold, then we just reenabled the signal
// So we are in the syscall handler
// Only throw a log message in this case
LogMan::Msg::EFmt("Signals in dispatcher have unsynchronized context");
if constexpr (false) {
// XXX: Messages in the signal handler can cause us to crash
LogMan::Msg::EFmt("Signals in dispatcher have unsynchronized context");
}
}
}
}
@@ -334,8 +403,23 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_CSGSFS] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
// aarch64 and x86_64 siginfo_t matches. We can just copy this over
// SI_USER could also potentially have random data in it, needs to be bit perfect
// For guest faults we don't have a real way to reconstruct state to a real guest RIP
*guest_siginfo = *HostSigInfo;
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = SynchronousFaultData.err_code;
// Overwrite si_code
guest_siginfo->si_code = SynchronousFaultData.si_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
}
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_OLDMASK] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_CR2] = 0;
@@ -380,11 +464,6 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_stack.ss_sp = GuestStack->ss_sp;
guest_uctx->uc_stack.ss_size = GuestStack->ss_size;
// aarch64 and x86_64 siginfo_t matches. We can just copy this over
// SI_USER could also potentially have random data in it, needs to be bit perfect
// For guest faults we don't have a real way to reconstruct state to a real guest RIP
*guest_siginfo = *HostSigInfo;
Frame->State.gregs[X86State::REG_RSI] = SigInfoLocation;
Frame->State.gregs[X86State::REG_RDX] = UContextLocation;
}
@@ -421,8 +500,16 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS] = Frame->State.fs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES] = Frame->State.es;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS] = Frame->State.ds;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_siginfo->si_code = SynchronousFaultData.si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = SynchronousFaultData.err_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
guest_siginfo->si_code = HostSigInfo->si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
}
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS] = Frame->State.cs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL] = 0;
@@ -470,7 +557,6 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// These three elements are in every siginfo
guest_siginfo->si_signo = HostSigInfo->si_signo;
guest_siginfo->si_errno = HostSigInfo->si_errno;
guest_siginfo->si_code = HostSigInfo->si_code;
switch (Signal) {
case SIGSEGV:
@@ -49,12 +49,19 @@ public:
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint64_t PauseReturnInstruction{};
/** @} */
uint32_t SignalHandlerRefCounter{};
struct SynchronousFaultDataStruct {
bool FaultToTopAndGeneratedException{};
uint32_t TrapNo;
uint32_t err_code;
uint32_t si_code;
} SynchronousFaultData;
uint64_t Start{};
uint64_t End{};
@@ -11,6 +11,7 @@
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cmath>
#include <memory>
@@ -284,6 +285,24 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
ud2();
}
{
// Guest Overflow handler
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = getCurr<uint64_t>();
// ud2 = SIGILL
// int3 = SIGTRAP
// hlt = SIGSEGV
mov(rax, reinterpret_cast<uint64_t>(&SynchronousFaultData));
add(byte [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)], 1);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)], X86State::X86_TRAPNO_OF);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)], 0);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)], 0x80);
hlt();
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
// using CallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
@@ -312,7 +331,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
End = Start + getSize();
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(::gettid());
std::string Name = "Dispatch_" + std::to_string(FHU::Syscalls::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(Start), End-Start, Name);
}
if (CTX->Config.GlobalJITNaming()) {
+5 -5
View File
@@ -399,12 +399,12 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
@@ -664,7 +664,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining",
DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name, InstructionSize, Bytes);
DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
}
@@ -680,12 +680,12 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
+5 -5
View File
@@ -480,7 +480,7 @@ std::string buildMemoryMap() {
xml << "<?xml version='1.0'?>\n";
xml << "<!DOCTYPE memory-map>\n";
xml << "<memory-map>";
xml << "<memory-map>\n";
std::fstream fs("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::string Line;
@@ -494,7 +494,7 @@ std::string buildMemoryMap() {
}
}
xml << "</memory-map>";
xml << "</memory-map>\n";
xml << std::flush;
@@ -595,20 +595,20 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
ThreadString = ss.str();
}
return {encode(ThreadString.substr(offset, length)), HandledPacketType::TYPE_ACK};
return {encode(ThreadString), HandledPacketType::TYPE_ACK};
}
if (object == "memory-map") {
if (offset == 0) {
MemoryMapString = buildMemoryMap();
}
return {encode(MemoryMapString.substr(offset, length)), HandledPacketType::TYPE_ACK};
return {encode(MemoryMapString), HandledPacketType::TYPE_ACK};
}
if (object == "osdata") {
if (offset == 0) {
OSDataString = buildOSData();
}
return {encode(OSDataString.substr(offset, length)), HandledPacketType::TYPE_ACK};
return {encode(OSDataString), HandledPacketType::TYPE_ACK};
}
return {"", HandledPacketType::TYPE_UNKNOWN};
+88
View File
@@ -1,3 +1,4 @@
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#ifdef _M_ARM_64
@@ -13,14 +14,101 @@
namespace FEXCore {
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
static uint32_t GetDCZID() {
uint64_t Result{};
__asm("mrs %[Res], DCZID_EL0"
: [Res] "=r" (Result));
return Result;
}
static uint32_t GetFPCR() {
uint64_t Result{};
__asm ("mrs %[Res], FPCR"
: [Res] "=r" (Result));
return Result;
}
static void SetFPCR(uint64_t Value) {
__asm ("msr FPCR, %[Value]"
:: [Value] "r" (Value));
}
#else
static uint32_t GetDCZID() {
// Return unsupported
return DCZID_DZP_MASK;
}
#endif
HostFeatures::HostFeatures() {
#ifdef _M_ARM_64
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
// Only supported when FEAT_AFP is supported
SupportsFlushInputsToZero = Features.Has(vixl::CPUFeatures::Feature::kAFP);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
__asm volatile ("mrs %[ctr], ctr_el0"
: [ctr] "=r"(CTR));
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#endif
#ifdef _M_X86_64
Xbyak::util::Cpu Features{};
SupportsAES = Features.has(Xbyak::util::Cpu::tAESNI);
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
#else
// Test if this CPU supports float exception trapping by attempting to enable
// On unsupported these bits are architecturally defined as RAZ/WI
constexpr uint32_t ExceptionEnableTraps =
(1U << 8) | // Invalid Operation float exception trap enable
(1U << 9) | // Divide by zero float exception trap enable
(1U << 10) | // Overflow float exception trap enable
(1U << 11) | // Underflow float exception trap enable
(1U << 12) | // Inexact float exception trap enable
(1U << 15); // Input Denormal float exception trap enable
uint32_t OriginalFPCR = GetFPCR();
uint32_t FPCR = OriginalFPCR | ExceptionEnableTraps;
SetFPCR(FPCR);
FPCR = GetFPCR();
SupportsFloatExceptions = (FPCR & ExceptionEnableTraps) == ExceptionEnableTraps;
// Set FPCR back to original just in case anything changed
SetFPCR(OriginalFPCR);
#endif
// Check if we can support cacheline clears
uint32_t DCZID = GetDCZID();
if ((DCZID & DCZID_DZP_MASK) == 0) {
uint32_t DCZID_Log2 = DCZID & DCZID_BS_MASK;
uint32_t DCZID_Bytes = (1 << DCZID_Log2) * sizeof(uint32_t);
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
SupportsCLZERO = DCZID_Bytes == CPUIDEmu::CACHELINE_SIZE;
}
}
}
+17
View File
@@ -1,9 +1,26 @@
#pragma once
#include <cstdint>
namespace FEXCore {
class HostFeatures final {
public:
HostFeatures();
/**
* @brief Backend features that change how codegen is generated from IR
*
* Specifically things that affect the IR->Codegen process
* Not the x86->IR process
*/
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
bool SupportsAES{};
bool SupportsCLZERO{};
bool SupportsAtomics{};
bool SupportsRCPC{};
// Float exception behaviour
bool SupportsFlushInputsToZero{};
bool SupportsFloatExceptions{};
};
}
@@ -493,6 +493,54 @@ DEF_OP(Extr) {
}
}
DEF_OP(PDep) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
if (OpSize != 4 && OpSize != 8) {
LOGMAN_MSG_A_FMT("Unknown PDep Size: {}\n", OpSize);
return;
}
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(0))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(0));
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(1))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(1));
uint64_t Result = 0;
for (uint64_t Index = 0; Mask > 0; Index++) {
const uint64_t Offset = std::countr_zero(Mask);
Mask &= Mask - 1;
Result |= ((Input >> Index) & 1) << Offset;
}
GD = Result;
}
DEF_OP(PExt) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
if (OpSize != 4 && OpSize != 8) {
LOGMAN_MSG_A_FMT("Unknown PExt Size: {}\n", OpSize);
return;
}
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(0))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(0));
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(1))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(1));
uint64_t Result = 0;
for (uint64_t Offset = 0; Mask > 0; Offset++) {
const uint64_t Index = std::countr_zero(Mask);
Mask &= Mask - 1;
Result |= ((Input >> Index) & 1) << Offset;
}
GD = Result;
}
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
uint8_t OpSize = IROp->Size;
@@ -73,6 +73,8 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(PDEP, PDep);
REGISTER_OP(PEXT, PExt);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
@@ -157,6 +159,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
// Misc ops
REGISTER_OP(DUMMY, NoOp);
@@ -172,6 +175,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
// Move ops
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
@@ -96,6 +96,8 @@ namespace FEXCore::CPU {
DEF_OP(Rol);
DEF_OP(Ror);
DEF_OP(Extr);
DEF_OP(PDep);
DEF_OP(PExt);
DEF_OP(LDiv);
DEF_OP(LUDiv);
DEF_OP(LRem);
@@ -180,6 +182,7 @@ namespace FEXCore::CPU {
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
@@ -190,6 +193,7 @@ namespace FEXCore::CPU {
DEF_OP(Print);
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -264,5 +264,22 @@ DEF_OP(CacheLineClear) {
CacheLineFlush(MemData);
}
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
uintptr_t MemData = *GetSrc<uintptr_t*>(Data->SSAData, Op->Addr);
// Force cacheline alignment
MemData = MemData & ~(CPUIDEmu::CACHELINE_SIZE - 1);
using DataType = uint64_t;
DataType *MemData64 = reinterpret_cast<DataType*>(MemData);
// 64-byte cache line zero
for (size_t i = 0; i < (CPUIDEmu::CACHELINE_SIZE / sizeof(DataType)); ++i) {
MemData64[i] = 0;
}
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -8,6 +8,8 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#ifdef _M_X86_64
#include <xmmintrin.h>
@@ -46,7 +48,7 @@ DEF_OP(Break) {
StopThread(Data->State);
break;
case FEXCore::IR::Break_InvalidInstruction:
tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGILL);
FHU::Syscalls::tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGILL);
break;
default: LOGMAN_MSG_A_FMT("Unknown Break Reason: {}", Op->Reason); break;
}
@@ -140,6 +142,12 @@ DEF_OP(Print) {
LOGMAN_MSG_A_FMT("Unknown value size: {}", OpSize);
}
DEF_OP(ProcessorID) {
uint32_t CPU, CPUNode;
FHU::Syscalls::getcpu(&CPU, &CPUNode);
GD = (CPUNode << 12) | CPU;
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -8,6 +8,7 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <bit>
#include <cstdint>
namespace FEXCore::CPU {
+126 -3
View File
@@ -8,6 +8,10 @@ $end_info$
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source / Divisor;
@@ -80,8 +84,6 @@ DEF_OP(CycleCounter) {
#endif
}
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
DEF_OP(Add) {
auto Op = IROp->C<IR::IROp_Add>();
uint8_t OpSize = IROp->Size;
@@ -511,6 +513,125 @@ DEF_OP(Extr) {
}
}
DEF_OP(PDep) {
auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const Register Input = GRS(Op->Args(0).ID());
const Register Mask = GRS(Op->Args(1).ID());
const Register Dest = GRS(Node);
const Register ShiftedBitReg = OpSize <= 4 ? TMP1.W() : TMP1;
const Register BitReg = OpSize <= 4 ? TMP2.W() : TMP2;
const Register SubMaskReg = OpSize <= 4 ? TMP3.W() : TMP3;
const Register IndexReg = OpSize <= 4 ? TMP4.W() : TMP4;
const Register SizedZero = OpSize <= 4 ? Register{wzr} : Register{xzr};
const Register InputReg = OpSize <= 4 ? SRA64[0].W() : SRA64[0];
const Register MaskReg = OpSize <= 4 ? SRA64[1].W() : SRA64[1];
const Register DestReg = OpSize <= 4 ? SRA64[2].W() : SRA64[2];
const auto SpillCode = 1U << InputReg.GetCode() |
1U << MaskReg.GetCode() |
1U << DestReg.GetCode();
aarch64::Label EarlyExit;
aarch64::Label NextBit;
aarch64::Label Done;
cbz(Mask, &EarlyExit);
mov(IndexReg, SizedZero);
// We sadly need to spill regs for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, SpillCode);
mov(InputReg, Input);
mov(MaskReg, Mask);
mov(DestReg, SizedZero);
// Main loop
bind(&NextBit);
rbit(ShiftedBitReg, MaskReg);
clz(ShiftedBitReg, ShiftedBitReg);
lsrv(BitReg, InputReg, IndexReg);
and_(BitReg, BitReg, 1);
sub(SubMaskReg, MaskReg, 1);
add(IndexReg, IndexReg, 1);
ands(MaskReg, MaskReg, SubMaskReg);
lslv(ShiftedBitReg, BitReg, ShiftedBitReg);
orr(DestReg, DestReg, ShiftedBitReg);
b(&NextBit, Condition::ne);
// Store result in a temp so it doesn't get clobbered.
// and restore it after the re-fill below.
mov(IndexReg, DestReg);
// Restore our registers before leaving
// TODO: Also remove along with above TODO.
FillStaticRegs(false, SpillCode);
mov(Dest, IndexReg);
b(&Done);
// Early exit
bind(&EarlyExit);
mov(Dest, SizedZero);
// All done with nothing to do.
bind(&Done);
}
DEF_OP(PExt) {
auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const Register Input = GRS(Op->Args(0).ID());
const Register Mask = GRS(Op->Args(1).ID());
const Register Dest = GRS(Node);
const Register MaskReg = OpSize <= 4 ? TMP1.W() : TMP1;
const Register BitReg = OpSize <= 4 ? TMP2.W() : TMP2;
const Register SubMaskReg = OpSize <= 4 ? TMP3.W() : TMP3;
const Register Offset = OpSize <= 4 ? TMP4.W() : TMP4;
const Register SizedZero = OpSize <= 4 ? Register{wzr} : Register{xzr};
aarch64::Label EarlyExit;
aarch64::Label NextBit;
aarch64::Label Done;
cbz(Mask, &EarlyExit);
mov(MaskReg, Mask);
mov(Offset, SizedZero);
// We sadly need to spill a reg for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, 1U << Mask.GetCode());
mov(Mask, SizedZero);
// Main loop
bind(&NextBit);
rbit(BitReg, MaskReg);
clz(BitReg, BitReg);
sub(SubMaskReg, MaskReg, 1);
ands(MaskReg, SubMaskReg, MaskReg);
lsrv(BitReg, Input, BitReg);
and_(BitReg, BitReg, 1);
lslv(BitReg, BitReg, Offset);
add(Offset, Offset, 1);
orr(Mask, BitReg, Mask);
b(&NextBit, Condition::ne);
mov(Dest, Mask);
// Restore our mask register before leaving
// TODO: Also remove along with above TODO.
FillStaticRegs(false, 1U << Mask.GetCode());
b(&Done);
// Early exit
bind(&EarlyExit);
mov(Dest, SizedZero);
// All done with nothing to do.
bind(&Done);
}
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
uint8_t OpSize = IROp->Size;
@@ -916,7 +1037,7 @@ Condition MapSelectCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FLU: return Condition::lt;
case FEXCore::IR::COND_FGE: return Condition::ge;
case FEXCore::IR::COND_FLEU:return Condition::le;
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FGT: return Condition::gt;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:
@@ -1099,6 +1220,8 @@ void Arm64JITCore::RegisterALUHandlers() {
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(PDEP, PDep);
REGISTER_OP(PEXT, PExt);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
@@ -19,7 +19,7 @@ DEF_OP(CASPair) {
auto Desired = GetSrcPair<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetReg<RA_64>(Op->Header.Args[2].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP3, Expected.first);
mov(TMP4, Expected.second);
@@ -110,7 +110,7 @@ DEF_OP(CAS) {
auto Desired = GetReg<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetReg<RA_64>(Op->Header.Args[2].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP2, Expected);
switch (OpSize) {
case 1: casalb(TMP2.W(), Desired.W(), MemOperand(MemSrc)); break;
@@ -218,7 +218,7 @@ DEF_OP(AtomicAdd) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: staddlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: staddlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
@@ -276,7 +276,7 @@ DEF_OP(AtomicSub) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (IROp->Size) {
case 1: staddlb(TMP2.W(), MemOperand(MemSrc)); break;
@@ -335,7 +335,7 @@ DEF_OP(AtomicAnd) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (IROp->Size) {
case 1: stclrlb(TMP2.W(), MemOperand(MemSrc)); break;
@@ -394,7 +394,7 @@ DEF_OP(AtomicOr) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: stsetlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: stsetlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
@@ -452,7 +452,7 @@ DEF_OP(AtomicXor) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: steorlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: steorlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
@@ -510,7 +510,7 @@ DEF_OP(AtomicSwap) {
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (IROp->Size) {
case 1: swplb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -568,7 +568,7 @@ DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldaddalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -629,7 +629,7 @@ DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (IROp->Size) {
case 1: ldaddalb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -691,7 +691,7 @@ DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
switch (IROp->Size) {
case 1: ldclralb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -753,7 +753,7 @@ DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldsetalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldsetalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -814,7 +814,7 @@ DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldeoralb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldeoralh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -318,7 +318,7 @@ DEF_OP(InlineSyscall) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
else {
uxtw(RegArgs[i], Reg);
uxtw(RegArgs[i].W(), Reg);
}
}
}
@@ -331,7 +331,7 @@ DEF_OP(InlineSyscall) {
mov(RegArgs[i], GetReg<RA_64>(Op->Header.Args[i].ID()));
}
else {
uxtw(RegArgs[i], GetReg<RA_32>(Op->Header.Args[i].ID()));
uxtw(RegArgs[i], GetReg<RA_64>(Op->Header.Args[i].ID()));
}
}
}
@@ -354,7 +354,7 @@ DEF_OP(InlineSyscall) {
mov(GetReg<RA_64>(Node), x0);
}
else {
uxtw(GetReg<RA_64>(Node), w0);
uxtw(GetReg<RA_64>(Node), x0);
}
}
@@ -39,11 +39,11 @@ DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
switch (Op->Header.ElementSize) {
case 1:
uxtb(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()).W());
uxtb(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 2:
uxth(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()).W());
uxth(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 4:
+1 -1
View File
@@ -333,7 +333,7 @@ void Arm64JITCore::FreeCodeBuffer(CodeBuffer Buffer) {
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: Arm64Emitter(0)
: Arm64Emitter(ctx, 0)
, CTX {ctx}
, ThreadState {Thread} {
{
@@ -176,6 +176,7 @@ private:
struct CompilerSharedData {
uint64_t SignalReturnInstruction{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
@@ -233,6 +234,8 @@ private:
DEF_OP(Rol);
DEF_OP(Ror);
DEF_OP(Extr);
DEF_OP(PDep);
DEF_OP(PExt);
DEF_OP(LDiv);
DEF_OP(LUDiv);
DEF_OP(LRem);
@@ -321,6 +324,7 @@ private:
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
@@ -331,6 +335,7 @@ private:
DEF_OP(Print);
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -623,7 +623,7 @@ DEF_OP(LoadMemTSO) {
LOGMAN_MSG_A_FMT("LoadMemTSO: No offset allowed");
}
if (SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
@@ -939,13 +939,35 @@ DEF_OP(CacheLineClear) {
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
mov(TMP1, MemReg);
for (size_t i = 0; i < std::max(1U, DCacheLineSize / 64U); ++i) {
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
dc(DataCacheOp::CVAU, TMP1);
add(TMP1, TMP1, DCacheLineSize);
add(TMP1, TMP1, CTX->HostFeatures.DCacheLineSize);
}
dsb(InnerShareable, BarrierAll);
}
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
if (CTX->HostFeatures.SupportsCLZERO) {
// We can use this instruction directly
dc(DataCacheOp::ZVA, MemReg);
}
else {
// We must walk the cacheline ourselves
// Force cacheline alignment
and_(TMP1, MemReg, ~(CPUIDEmu::CACHELINE_SIZE - 1));
// This will end up being four STPs
// Depending on uarch it could be slightly more efficient in instructions emitted
// and uops to use vector pair STP, but we want the non-temporal bit specifically here
for (size_t i = 0; i < CPUIDEmu::CACHELINE_SIZE; i += 16) {
stnp(xzr, xzr, MemOperand(TMP1, i, Offset));
}
}
}
#undef DEF_OP
void Arm64JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -972,6 +994,7 @@ void Arm64JITCore::RegisterMemoryHandlers() {
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
}
}
@@ -40,9 +40,13 @@ DEF_OP(Break) {
switch (Op->Reason) {
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
case FEXCore::IR::Break_Overflow: // overflow
hlt(4);
break;
case FEXCore::IR::Break_Overflow: // overflow
ResetStack();
LoadConstant(TMP1, ThreadSharedData.Dispatcher->OverflowExceptionInstructionAddress);
br(TMP1);
break;
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
@@ -154,6 +158,58 @@ DEF_OP(Print) {
PopDynamicRegsAndLR();
}
DEF_OP(ProcessorID) {
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
// 16bit LoadConstant to be a single instruction
// We must always spill at least one register (x8) so this value always has a bit set
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(x0, SpillMask & 0xFFFF);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Allocate some temporary space for storing the uint32_t CPU and Node IDs
sub(sp, sp, 16);
// Load the getcpu syscall number
LoadConstant(x8, SYS_getcpu);
// CPU pointer in x0
add(x0, sp, 0);
// Node in x1
add(x1, sp, 4);
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
// Load the values returned by the kernel
ldp(w0, w1, MemOperand(sp));
// Deallocate stack space
sub(sp, sp, 16);
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Now store the result in the destination in the expected format
// uint32_t Res = (node << 12) | cpu;
// CPU is in w0
// Node is in w1
orr(GetReg<RA_64>(Node), x0, Operand(x1, LSL, 12));
}
#undef DEF_OP
void Arm64JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -170,6 +226,7 @@ void Arm64JITCore::RegisterMiscHandlers() {
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
#undef REGISTER_OP
}
}
@@ -671,6 +671,36 @@ DEF_OP(Extr) {
}
}
DEF_OP(PDep) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const auto Input = GRS(Op->Args(0).ID());
const auto Mask = GRS(Op->Args(1).ID());
const auto Dest = GRD(Node);
if (OpSize == 4) {
pdep(Dest.cvt32(), Input.cvt32(), Mask.cvt32());
} else {
pdep(Dest.cvt64(), Input.cvt64(), Mask.cvt64());
}
}
DEF_OP(PExt) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const auto Input = GRS(Op->Args(0).ID());
const auto Mask = GRS(Op->Args(1).ID());
const auto Dest = GRD(Node);
if (OpSize == 4) {
pext(Dest.cvt32(), Input.cvt32(), Mask.cvt32());
} else {
pext(Dest.cvt64(), Input.cvt64(), Mask.cvt64());
}
}
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
uint8_t OpSize = IROp->Size;
@@ -1248,6 +1278,8 @@ void X86JITCore::RegisterALUHandlers() {
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(PDEP, PDep);
REGISTER_OP(PEXT, PExt);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
@@ -358,6 +358,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.OverflowExceptionInstructionAddress = Dispatcher->OverflowExceptionInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
@@ -167,7 +167,7 @@ private:
struct CompilerSharedData {
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
@@ -233,6 +233,8 @@ private:
DEF_OP(Rol);
DEF_OP(Ror);
DEF_OP(Extr);
DEF_OP(PDep);
DEF_OP(PExt);
DEF_OP(LDiv);
DEF_OP(LUDiv);
DEF_OP(LRem);
@@ -315,6 +317,7 @@ private:
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
@@ -325,6 +328,7 @@ private:
DEF_OP(Print);
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -595,6 +595,23 @@ DEF_OP(CacheLineClear) {
clflush(ptr [MemReg]);
}
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
// Align by cacheline
mov (TMP1, CPUIDEmu::CACHELINE_SIZE - 1);
andn(TMP1, TMP1, MemReg.cvt64());
xor_(TMP2, TMP2);
using DataType = uint64_t;
// 64-byte cache line zero
for (size_t i = 0; i < CPUIDEmu::CACHELINE_SIZE; i += sizeof(DataType)) {
mov (qword [TMP1 + i], TMP2);
}
}
#undef DEF_OP
void X86JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -615,6 +632,7 @@ void X86JITCore::RegisterMemoryHandlers() {
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
}
}
@@ -49,9 +49,13 @@ DEF_OP(Break) {
switch (Op->Reason) {
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
case FEXCore::IR::Break_Overflow: // overflow
ud2();
break;
case FEXCore::IR::Break_Overflow: // overflow
// Need to be outside of JIT cache space to ensure cache clearing correctness
mov(TMP1, ThreadSharedData.Dispatcher->OverflowExceptionInstructionAddress);
jmp(TMP1);
break;
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
@@ -155,6 +159,13 @@ DEF_OP(Print) {
PopRegs();
}
DEF_OP(ProcessorID) {
// Cyclecounter in EDX:EAX
// IA32_TSC_AUX in ECX
rdtscp();
mov (GetDst<RA_32>(Node), ecx);
}
#undef DEF_OP
void X86JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -171,6 +182,7 @@ void X86JITCore::RegisterMiscHandlers() {
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
#undef REGISTER_OP
}
}
+72 -19
View File
@@ -2449,6 +2449,22 @@ void OpDispatchBuilder::MULX(OpcodeArgs) {
StoreResult(GPRClass, Op, Op->Dest, ResultHi, -1);
}
void OpDispatchBuilder::PDEP(OpcodeArgs) {
auto* Input = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto* Mask = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
auto Result = _PDep(Input, Mask);
StoreResult(GPRClass, Op, Op->Dest, Result, -1);
}
void OpDispatchBuilder::PEXT(OpcodeArgs) {
auto* Input = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
auto* Mask = LoadSource(GPRClass, Op, Op->Src[1], Op->Flags, -1);
auto Result = _PExt(Input, Mask);
StoreResult(GPRClass, Op, Op->Dest, Result, -1);
}
void OpDispatchBuilder::ADXOp(OpcodeArgs) {
// Handles ADCX and ADOX
@@ -5058,8 +5074,9 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
break;
}
const uint8_t GPRSize = CTX->GetGPRSize();
if (setRIP) {
const uint8_t GPRSize = CTX->GetGPRSize();
BlockSetRIP = setRIP;
@@ -5077,6 +5094,8 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
SetFalseJumpTarget(CondJump, FalseBlock);
SetCurrentCodeBlock(FalseBlock);
auto NewRIP = GetDynamicPC(Op);
_StoreContext(GPRClass, GPRSize, offsetof(FEXCore::Core::CPUState, rip), NewRIP);
_Break(Reason, Literal);
// Make sure to start a new block after ending this one
@@ -5145,6 +5164,32 @@ void OpDispatchBuilder::StoreFenceOrCLFlush(OpcodeArgs) {
}
}
void OpDispatchBuilder::CLZeroOp(OpcodeArgs) {
OrderedNode *DestMem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1, false);
_CacheLineZero(DestMem);
}
void OpDispatchBuilder::RDTSCPOp(OpcodeArgs) {
const uint8_t GPRSize = CTX->GetGPRSize();
// RDTSCP is slightly different than RDTSC
// IA32_TSC_AUX is returned in RCX
// All previous loads are globally visible
// - Explicitly does not wait for stores to be globally visible
// - Explicitly use an MFENCE before this instruction if you want this behaviour
// This instruction is not an execution fence, so subsequent instructions can execute after this
// - Explicitly use an LFENCE after RDTSCP if you want to block this behaviour
_Fence({FEXCore::IR::Fence_Load});
auto Counter = _CycleCounter();
auto CounterLow = _Bfe(32, 0, Counter);
auto CounterHigh = _Bfe(32, 32, Counter);
auto ID = _ProcessorID();
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RAX), CounterLow);
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RCX), ID);
_StoreContext(GPRClass, GPRSize, GPROffset(X86State::REG_RDX), CounterHigh);
}
void OpDispatchBuilder::UnimplementedOp(OpcodeArgs) {
const uint8_t GPRSize = CTX->GetGPRSize();
@@ -5171,7 +5216,7 @@ void OpDispatchBuilder::InvalidOp(OpcodeArgs) {
#undef OpcodeArgs
void InstallOpcodeHandlers(Context::OperatingMode Mode) {
const std::vector<std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr>> BaseOpTable = {
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> BaseOpTable[] = {
// Instructions
{0x00, 6, &OpDispatchBuilder::ALUOp},
@@ -5241,7 +5286,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xFC, 2, &OpDispatchBuilder::FLAGControlOp},
};
const std::vector<std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr>> BaseOpTable_32 = {
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> BaseOpTable_32[] = {
{0x06, 1, &OpDispatchBuilder::PUSHSegmentOp<FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x07, 1, &OpDispatchBuilder::POPSegmentOp<FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x0E, 1, &OpDispatchBuilder::PUSHSegmentOp<FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX>},
@@ -5254,13 +5299,14 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0x60, 1, &OpDispatchBuilder::PUSHAOp},
{0x61, 1, &OpDispatchBuilder::POPAOp},
{0xCE, 1, &OpDispatchBuilder::INTOp},
};
const std::vector<std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr>> BaseOpTable_64 = {
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> BaseOpTable_64[] = {
{0x63, 1, &OpDispatchBuilder::MOVSXDOp},
};
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> TwoByteOpTable = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> TwoByteOpTable[] = {
// Instructions
{0x0B, 1, &OpDispatchBuilder::INTOp},
{0x0E, 1, &OpDispatchBuilder::X87EMMS},
@@ -5403,16 +5449,16 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0x37, 1, &OpDispatchBuilder::CallbackReturnOp},
};
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> TwoByteOpTable_32 = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> TwoByteOpTable_32[] = {
{0x05, 1, &OpDispatchBuilder::NOPOp},
};
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> TwoByteOpTable_64 = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> TwoByteOpTable_64[] = {
{0x05, 1, &OpDispatchBuilder::SyscallOp},
};
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
const std::vector<std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> PrimaryGroupOpTable = {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> PrimaryGroupOpTable[] = {
// GROUP 1
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 0), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 1), 1, &OpDispatchBuilder::SecondaryALUOp},
@@ -5533,7 +5579,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
};
#undef OPD
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> RepModOpTable = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> RepModOpTable[] = {
{0x10, 2, &OpDispatchBuilder::MOVSSOp},
{0x12, 1, &OpDispatchBuilder::MOVSLDUPOp},
{0x16, 1, &OpDispatchBuilder::MOVSHDUPOp},
@@ -5565,7 +5611,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<4, true>},
};
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> RepNEModOpTable = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> RepNEModOpTable[] = {
{0x10, 2, &OpDispatchBuilder::MOVSDOp},
{0x12, 1, &OpDispatchBuilder::MOVDDUPOp},
{0x19, 7, &OpDispatchBuilder::NOPOp},
@@ -5592,7 +5638,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xF0, 1, &OpDispatchBuilder::MOVVectorOp},
};
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> OpSizeModOpTable = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> OpSizeModOpTable[] = {
{0x10, 2, &OpDispatchBuilder::MOVVectorOp},
{0x12, 2, &OpDispatchBuilder::MOVLPOp},
{0x14, 1, &OpDispatchBuilder::PUNPCKLOp<8>},
@@ -5706,7 +5752,7 @@ constexpr uint16_t PF_F3 = 1;
constexpr uint16_t PF_66 = 2;
constexpr uint16_t PF_F2 = 3;
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_6) << 5) | (prefix) << 3 | (Reg))
const std::vector<std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> SecondaryExtensionOpTable = {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryExtensionOpTable[] = {
// GROUP 8
{OPD(FEXCore::X86Tables::TYPE_GROUP_8, PF_NONE, 4), 1, &OpDispatchBuilder::BTOp<1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_8, PF_F3, 4), 1, &OpDispatchBuilder::BTOp<1>},
@@ -5788,15 +5834,20 @@ constexpr uint16_t PF_F2 = 3;
};
#undef OPD
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> SecondaryModRMExtensionOpTable = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryModRMExtensionOpTable[] = {
// REG /2
{((1 << 3) | 0), 1, &OpDispatchBuilder::UnimplementedOp},
// REG /7
{((3 << 3) | 1), 1, &OpDispatchBuilder::RDTSCPOp},
{((3 << 3) | 4), 1, &OpDispatchBuilder::CLZeroOp},
};
// Top bit indicating if it needs to be repeated with {0x40, 0x80} or'd in
// All OPDReg versions need it
#define OPDReg(op, reg) ((1 << 15) | ((op - 0xD8) << 8) | (reg << 3))
#define OPD(op, modrmop) (((op - 0xD8) << 8) | modrmop)
const std::vector<std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> X87OpTable = {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> X87OpTable[] = {
{OPDReg(0xD8, 0) | 0x00, 8, &OpDispatchBuilder::FADD<32, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 1) | 0x00, 8, &OpDispatchBuilder::FMUL<32, false, OpDispatchBuilder::OpResult::RES_ST0>},
@@ -6034,7 +6085,7 @@ constexpr uint16_t PF_F2 = 3;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
const std::vector<std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> H0F38Table = {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_66, 0x00), 1, &OpDispatchBuilder::PSHUFBOp},
{OPD(PF_38_NONE, 0x01), 1, &OpDispatchBuilder::PHADD<2>},
@@ -6116,7 +6167,7 @@ constexpr uint16_t PF_F2 = 3;
#define OPD(REX, prefix, opcode) ((REX << 9) | (prefix << 8) | opcode)
#define PF_3A_NONE 0
#define PF_3A_66 1
const std::vector<std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> H0F3ATable = {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F3ATable[] = {
{OPD(0, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<4, false>},
{OPD(0, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<8, false>},
{OPD(0, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::VectorRound<4, true>},
@@ -6151,7 +6202,7 @@ constexpr uint16_t PF_F2 = 3;
#undef OPD
#define OPD(map_select, pp, opcode) (((map_select - 1) << 10) | (pp << 8) | (opcode))
const std::vector<std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> VEXTable = {
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> VEXTable[] = {
{OPD(1, 0b01, 0x6E), 2, &OpDispatchBuilder::UnimplementedOp},
{OPD(1, 0b10, 0x6F), 1, &OpDispatchBuilder::UnimplementedOp},
@@ -6178,6 +6229,8 @@ constexpr uint16_t PF_F2 = 3;
{OPD(2, 0b00, 0xF2), 1, &OpDispatchBuilder::ANDNBMIOp},
{OPD(2, 0b00, 0xF5), 1, &OpDispatchBuilder::BZHI},
{OPD(2, 0b10, 0xF5), 1, &OpDispatchBuilder::PEXT},
{OPD(2, 0b11, 0xF5), 1, &OpDispatchBuilder::PDEP},
{OPD(2, 0b11, 0xF6), 1, &OpDispatchBuilder::MULX},
{OPD(2, 0b00, 0xF7), 1, &OpDispatchBuilder::BEXTRBMIOp},
{OPD(2, 0b01, 0xF7), 1, &OpDispatchBuilder::BMI2Shift},
@@ -6189,14 +6242,14 @@ constexpr uint16_t PF_F2 = 3;
#undef OPD
#define OPD(group, pp, opcode) (((group - X86Tables::InstType::TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
const std::vector<std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr>> VEXGroupTable = {
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> VEXGroupTable[] = {
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b001), 1, &OpDispatchBuilder::BLSRBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b010), 1, &OpDispatchBuilder::BLSMSKBMIOp},
{OPD(X86Tables::InstType::TYPE_VEX_GROUP_17, 0, 0b011), 1, &OpDispatchBuilder::BLSIBMIOp},
};
#undef OPD
const std::vector<std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr>> EVEXTable = {
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> EVEXTable[] = {
{0x10, 2, &OpDispatchBuilder::UnimplementedOp},
{0x59, 1, &OpDispatchBuilder::UnimplementedOp},
{0x7F, 1, &OpDispatchBuilder::UnimplementedOp},
@@ -338,6 +338,8 @@ public:
void BMI2Shift(OpcodeArgs);
void BZHI(OpcodeArgs);
void MULX(OpcodeArgs);
void PDEP(OpcodeArgs);
void PEXT(OpcodeArgs);
void RORX(OpcodeArgs);
// ADX Ops
@@ -469,6 +471,8 @@ public:
void FenceOp(OpcodeArgs);
void StoreFenceOrCLFlush(OpcodeArgs);
void CLZeroOp(OpcodeArgs);
void RDTSCPOp(OpcodeArgs);
void PSADBW(OpcodeArgs);
+2 -1
View File
@@ -1,5 +1,6 @@
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <unistd.h>
#include <signal.h>
@@ -91,7 +92,7 @@ namespace FEXCore {
HostSignalHandler &Handler = HostHandlers[Signal];
if (!Thread) {
LogMan::Msg::EFmt("[{}] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", ::gettid());
LogMan::Msg::EFmt("[{}] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", FHU::Syscalls::gettid());
}
else {
if (Handler.Handler &&
+4 -41
View File
@@ -6,44 +6,7 @@
namespace FEXCore::X86Tables::X86InstDebugInfo {
void InstallDebugInfo() {
using namespace FEXCore::X86Tables;
auto NoFlags = Flags {0};
for (auto &BaseOp : BaseOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : SecondBaseOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : RepModOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : RepNEModOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : OpSizeModOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : PrimaryInstGroupOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : SecondInstGroupOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : SecondModRMTableOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : X87Ops)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : DDDNowOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : H0F38TableOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : H0F3ATableOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : VEXTableOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : VEXTableGroupOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : XOPTableOps)
BaseOp.DebugInfo = NoFlags;
for (auto &BaseOp : XOPTableGroupOps)
BaseOp.DebugInfo = NoFlags;
const std::vector<std::tuple<uint8_t, uint8_t, Flags>> BaseOpTable = {
const std::tuple<uint8_t, uint8_t, Flags> BaseOpTable[] = {
{0x50, 8, {FLAGS_MEM_ACCESS}},
{0x58, 8, {FLAGS_MEM_ACCESS}},
@@ -62,7 +25,7 @@ void InstallDebugInfo() {
{0xF4, 1, {FLAGS_DEBUG}},
};
const std::vector<std::tuple<uint8_t, uint8_t, Flags>> TwoByteOpTable = {
const std::tuple<uint8_t, uint8_t, Flags> TwoByteOpTable[] = {
{0x0B, 1, {FLAGS_DEBUG}},
{0x19, 7, {FLAGS_DEBUG}},
{0x28, 2, {FLAGS_MEM_ALIGN_16}},
@@ -78,14 +41,14 @@ void InstallDebugInfo() {
{0xFF, 1, {FLAGS_DEBUG}},
};
const std::vector<std::tuple<uint8_t, uint8_t, Flags>> PrimaryGroupOpTable = {
const std::tuple<uint8_t, uint8_t, Flags> PrimaryGroupOpTable[] = {
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 6), 2, {FLAGS_DIVIDE}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 6), 2, {FLAGS_DIVIDE}},
#undef OPD
};
const std::vector<std::tuple<uint16_t, uint8_t, Flags>> SecondaryExtensionOpTable = {
const std::tuple<uint16_t, uint8_t, Flags> SecondaryExtensionOpTable[] = {
#define PF_NONE 0
#define PF_F3 1
#define PF_66 2
+1 -1
View File
@@ -9,12 +9,12 @@ $end_info$
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#include <cstring>
#include <vector>
#include <sys/mman.h>
#include <bits/mman-map-flags-generic.h>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
+17 -56
View File
@@ -10,25 +10,24 @@ $end_info$
namespace FEXCore::X86Tables {
X86InstInfo BaseOps[MAX_PRIMARY_TABLE_SIZE];
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps{};
std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps{};
std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> RepModOps{};
std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps{};
std::array<X86InstInfo, MAX_OPSIZE_MOD_TABLE_SIZE> OpSizeModOps{};
X86InstInfo SecondBaseOps[MAX_SECOND_TABLE_SIZE];
X86InstInfo RepModOps[MAX_REP_MOD_TABLE_SIZE];
X86InstInfo RepNEModOps[MAX_REPNE_MOD_TABLE_SIZE];
X86InstInfo OpSizeModOps[MAX_OPSIZE_MOD_TABLE_SIZE];
X86InstInfo PrimaryInstGroupOps[MAX_INST_GROUP_TABLE_SIZE];
X86InstInfo SecondInstGroupOps[MAX_INST_SECOND_GROUP_TABLE_SIZE];
X86InstInfo SecondModRMTableOps[MAX_SECOND_MODRM_TABLE_SIZE];
X86InstInfo X87Ops[MAX_X87_TABLE_SIZE];
X86InstInfo DDDNowOps[MAX_3DNOW_TABLE_SIZE];
X86InstInfo H0F38TableOps[MAX_0F_38_TABLE_SIZE];
X86InstInfo H0F3ATableOps[MAX_0F_3A_TABLE_SIZE];
X86InstInfo VEXTableOps[MAX_VEX_TABLE_SIZE];
X86InstInfo VEXTableGroupOps[MAX_VEX_GROUP_TABLE_SIZE];
X86InstInfo XOPTableOps[MAX_XOP_TABLE_SIZE];
X86InstInfo XOPTableGroupOps[MAX_XOP_GROUP_TABLE_SIZE];
X86InstInfo EVEXTableOps[MAX_EVEX_TABLE_SIZE];
std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps{};
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps{};
std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps{};
std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87Ops{};
std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps{};
std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps{};
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps{};
std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps{};
std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps{};
std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> XOPTableOps{};
std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> XOPTableGroupOps{};
std::array<X86InstInfo, MAX_EVEX_TABLE_SIZE> EVEXTableOps{};
void InitializeBaseTables(Context::OperatingMode Mode);
void InitializeSecondaryTables(Context::OperatingMode Mode);
@@ -49,44 +48,6 @@ uint64_t NumInsts{};
#endif
void InitializeInfoTables(Context::OperatingMode Mode) {
using namespace FEXCore::X86Tables::InstFlags;
auto UnknownOp = X86InstInfo{"UND", TYPE_UNKNOWN, FLAGS_NONE, 0, nullptr};
for (auto &BaseOp : BaseOps)
BaseOp = UnknownOp;
for (auto &BaseOp : SecondBaseOps)
BaseOp = UnknownOp;
for (auto &BaseOp : RepModOps)
BaseOp = UnknownOp;
for (auto &BaseOp : RepNEModOps)
BaseOp = UnknownOp;
for (auto &BaseOp : OpSizeModOps)
BaseOp = UnknownOp;
for (auto &BaseOp : PrimaryInstGroupOps)
BaseOp = UnknownOp;
for (auto &BaseOp : SecondInstGroupOps)
BaseOp = UnknownOp;
for (auto &BaseOp : SecondModRMTableOps)
BaseOp = UnknownOp;
for (auto &BaseOp : X87Ops)
BaseOp = UnknownOp;
for (auto &BaseOp : DDDNowOps)
BaseOp = UnknownOp;
for (auto &BaseOp : H0F38TableOps)
BaseOp = UnknownOp;
for (auto &BaseOp : H0F3ATableOps)
BaseOp = UnknownOp;
for (auto &BaseOp : VEXTableOps)
BaseOp = UnknownOp;
for (auto &BaseOp : VEXTableGroupOps)
BaseOp = UnknownOp;
for (auto &BaseOp : XOPTableOps)
BaseOp = UnknownOp;
for (auto &BaseOp : XOPTableGroupOps)
BaseOp = UnknownOp;
for (auto &BaseOp : EVEXTableOps)
BaseOp = UnknownOp;
InitializeBaseTables(Mode);
InitializeSecondaryTables(Mode);
InitializePrimaryGroupTables(Mode);
@@ -288,13 +288,13 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xEA, 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(BaseOps, BaseOpTable, std::size(BaseOpTable));
GenerateTable(&BaseOps.at(0), BaseOpTable, std::size(BaseOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(BaseOps, BaseOpTable_64, std::size(BaseOpTable_64));
GenerateTable(&BaseOps.at(0), BaseOpTable_64, std::size(BaseOpTable_64));
}
else {
GenerateTable(BaseOps, BaseOpTable_32, std::size(BaseOpTable_32));
GenerateTable(&BaseOps.at(0), BaseOpTable_32, std::size(BaseOpTable_32));
}
}
}
@@ -48,6 +48,6 @@ void InitializeDDDTables() {
{0xB7, 1, X86InstInfo{"PMULHRW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
};
GenerateTable(DDDNowOps, DDDNowOpTable, std::size(DDDNowOpTable));
GenerateTable(&DDDNowOps.at(0), DDDNowOpTable, std::size(DDDNowOpTable));
}
}
@@ -30,6 +30,6 @@ void InitializeEVEXTables() {
{0xE7, 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
};
GenerateTable(EVEXTableOps, EVEXTable, std::size(EVEXTable));
GenerateTable(&EVEXTableOps.at(0), EVEXTable, std::size(EVEXTable));
}
}
@@ -106,6 +106,6 @@ void InitializeH0F38Tables() {
};
#undef OPD
GenerateTable(H0F38TableOps, H0F38Table, std::size(H0F38Table));
GenerateTable(&H0F38TableOps.at(0), H0F38Table, std::size(H0F38Table));
}
}
@@ -60,10 +60,10 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(H0F3ATableOps, H0F3ATable, std::size(H0F3ATable));
GenerateTable(&H0F3ATableOps.at(0), H0F3ATable, std::size(H0F3ATable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(H0F3ATableOps, H0F3ATable_64, std::size(H0F3ATable_64));
GenerateTable(&H0F3ATableOps.at(0), H0F3ATable_64, std::size(H0F3ATable_64));
}
}
}
@@ -163,14 +163,13 @@ void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
GenerateTable(&PrimaryInstGroupOps.at(0), PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_64, std::size(PrimaryGroupOpTable_64));
GenerateTable(&PrimaryInstGroupOps.at(0), PrimaryGroupOpTable_64, std::size(PrimaryGroupOpTable_64));
}
else {
GenerateTable(PrimaryInstGroupOps, PrimaryGroupOpTable_32, std::size(PrimaryGroupOpTable_32));
GenerateTable(&PrimaryInstGroupOps.at(0), PrimaryGroupOpTable_32, std::size(PrimaryGroupOpTable_32));
}
}
}
@@ -488,7 +488,7 @@ void InitializeSecondaryGroupTables() {
};
#undef OPD
GenerateTable(SecondInstGroupOps, SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
GenerateTable(&SecondInstGroupOps.at(0), SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
}
}
@@ -47,15 +47,15 @@ void InitializeSecondaryModRMTables() {
// REG /7
{((3 << 3) | 0), 1, X86InstInfo{"SWAPGS", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 1), 1, X86InstInfo{"RDTSCP", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 1), 1, X86InstInfo{"RDTSCP", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 2), 1, X86InstInfo{"MONITORX", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 3), 1, X86InstInfo{"MWAITX", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 4), 1, X86InstInfo{"CLZERO", TYPE_INST, FLAGS_SF_SRC_RAX, 0, nullptr}},
{((3 << 3) | 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((3 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(SecondModRMTableOps, SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
GenerateTable(&SecondModRMTableOps.at(0), SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
}
}
@@ -582,18 +582,18 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xFF, 1, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(SecondBaseOps, TwoByteOpTable, std::size(TwoByteOpTable));
GenerateTable(&SecondBaseOps.at(0), TwoByteOpTable, std::size(TwoByteOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(SecondBaseOps, TwoByteOpTable_64, std::size(TwoByteOpTable_64));
if (Mode == Context::MODE_64BIT) {
GenerateTable(&SecondBaseOps.at(0), TwoByteOpTable_64, std::size(TwoByteOpTable_64));
}
else {
GenerateTable(SecondBaseOps, TwoByteOpTable_32, std::size(TwoByteOpTable_32));
GenerateTable(&SecondBaseOps.at(0), TwoByteOpTable_32, std::size(TwoByteOpTable_32));
}
GenerateTableWithCopy(RepModOps, RepModOpTable, std::size(RepModOpTable), SecondBaseOps);
GenerateTableWithCopy(RepNEModOps, RepNEModOpTable, std::size(RepNEModOpTable), SecondBaseOps);
GenerateTableWithCopy(OpSizeModOps, OpSizeModOpTable, std::size(OpSizeModOpTable), SecondBaseOps);
GenerateTableWithCopy(&RepModOps.at(0), RepModOpTable, std::size(RepModOpTable), &SecondBaseOps.at(0));
GenerateTableWithCopy(&RepNEModOps.at(0), RepNEModOpTable, std::size(RepNEModOpTable), &SecondBaseOps.at(0));
GenerateTableWithCopy(&OpSizeModOps.at(0), OpSizeModOpTable, std::size(OpSizeModOpTable), &SecondBaseOps.at(0));
}
@@ -394,8 +394,9 @@ void InitializeVEXTables() {
{OPD(2, 0b11, 0xF3), 1, X86InstInfo{"", TYPE_VEX_GROUP_17, FLAGS_NONE, 0, nullptr}}, // VEX Group 17
{OPD(2, 0b00, 0xF5), 1, X86InstInfo{"BZHI", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
{OPD(2, 0b01, 0xF5), 1, X86InstInfo{"PEXT", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b11, 0xF5), 1, X86InstInfo{"PDEP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
// AMD reference manual is incorrect. PEXT actually maps to 0b10, not 0b01.
{OPD(2, 0b10, 0xF5), 1, X86InstInfo{"PEXT", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF5), 1, X86InstInfo{"PDEP", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(2, 0b11, 0xF6), 1, X86InstInfo{"MULX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
@@ -509,7 +510,7 @@ void InitializeVEXTables() {
};
#undef OPD
GenerateTable(VEXTableOps, VEXTable, std::size(VEXTable));
GenerateTable(VEXTableGroupOps, VEXGroupTable, std::size(VEXGroupTable));
GenerateTable(&VEXTableOps.at(0), VEXTable, std::size(VEXTable));
GenerateTable(&VEXTableGroupOps.at(0), VEXGroupTable, std::size(VEXGroupTable));
}
}
+14 -30
View File
@@ -17,20 +17,19 @@ extern uint64_t Total;
extern uint64_t NumInsts;
#endif
struct U8U8InfoStruct {
uint8_t first, second;
X86InstInfo Info;
};
struct U16U8InfoStruct {
uint16_t first;
template <typename OpcodeType>
struct X86TablesInfoStruct {
OpcodeType first;
uint8_t second;
X86InstInfo Info;
};
using U8U8InfoStruct = X86TablesInfoStruct<uint8_t>;
using U16U8InfoStruct = X86TablesInfoStruct<uint16_t>;
static inline void GenerateTable(X86InstInfo *FinalTable, U8U8InfoStruct const *LocalTable, size_t TableSize) {
template<typename OpcodeType>
static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
for (size_t j = 0; j < TableSize; ++j) {
U8U8InfoStruct const &Op = LocalTable[j];
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
@@ -45,26 +44,10 @@ static inline void GenerateTable(X86InstInfo *FinalTable, U8U8InfoStruct const *
}
};
static inline void GenerateTable(X86InstInfo *FinalTable, U16U8InfoStruct const *LocalTable, size_t TableSize) {
template<typename OpcodeType>
static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize, X86InstInfo *OtherLocal) {
for (size_t j = 0; j < TableSize; ++j) {
U16U8InfoStruct const &Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
FinalTable[OpNum + i] = Info;
#ifndef NDEBUG
++Total;
if (Info.Type == TYPE_INST)
NumInsts++;
#endif
}
}
};
static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, U8U8InfoStruct const *LocalTable, size_t TableSize, X86InstInfo *OtherLocal) {
for (size_t j = 0; j < TableSize; ++j) {
U8U8InfoStruct const &Op = LocalTable[j];
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
@@ -84,9 +67,10 @@ static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, U8U8InfoStruct
}
};
static inline void GenerateX87Table(X86InstInfo *FinalTable, U16U8InfoStruct const *LocalTable, size_t TableSize) {
template<typename OpcodeType>
static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
for (size_t j = 0; j < TableSize; ++j) {
U16U8InfoStruct const &Op = LocalTable[j];
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
@@ -263,6 +263,6 @@ void InitializeX87Tables() {
#undef OPD
#undef OPDReg
GenerateX87Table(X87Ops, X87OpTable, std::size(X87OpTable));
GenerateX87Table(&X87Ops.at(0), X87OpTable, std::size(X87OpTable));
}
}
@@ -130,7 +130,7 @@ void InitializeXOPTables() {
};
#undef OPD
GenerateTable(XOPTableOps, XOPTable, std::size(XOPTable));
GenerateTable(XOPTableGroupOps, XOPGroupTable, std::size(XOPGroupTable));
GenerateTable(&XOPTableOps.at(0), XOPTable, std::size(XOPTable));
GenerateTable(&XOPTableGroupOps.at(0), XOPGroupTable, std::size(XOPGroupTable));
}
}
+408
View File
@@ -0,0 +1,408 @@
#include "Interface/Context/Context.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <cstddef>
#include <cstdint>
#include <filesystem>
#include <fstream>
#include <mutex>
#include <sys/mman.h>
#include <sys/stat.h>
#include <unistd.h>
#include <xxhash.h>
namespace FEXCore::IR {
AOTIRInlineEntry *AOTIRInlineIndex::GetInlineEntry(uint64_t DataOffset) {
uintptr_t This = (uintptr_t)this;
return (AOTIRInlineEntry*)(This + DataBase + DataOffset);
}
AOTIRInlineEntry *AOTIRInlineIndex::Find(uint64_t GuestStart) {
ssize_t l = 0;
ssize_t r = Count - 1;
while (l <= r) {
size_t m = l + (r - l) / 2;
if (Entries[m].GuestStart == GuestStart)
return GetInlineEntry(Entries[m].DataOffset);
else if (Entries[m].GuestStart < GuestStart)
l = m + 1;
else
r = m - 1;
}
return nullptr;
}
IR::RegisterAllocationData *AOTIRInlineEntry::GetRAData() {
return (IR::RegisterAllocationData *)InlineData;
}
IR::IRListView *AOTIRInlineEntry::GetIRData() {
auto RAData = GetRAData();
auto Offset = RAData->Size(RAData->MapCount);
return (IR::IRListView *)&InlineData[Offset];
}
void AOTIRCaptureCacheEntry::AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView *IRList, FEXCore::IR::RegisterAllocationData *RAData) {
auto Inserted = Index.emplace(GuestRIP, Stream->tellp());
if (Inserted.second) {
//GuestHash
Stream->write((const char*)&Hash, sizeof(Hash));
//GuestLength
Stream->write((const char*)&Length, sizeof(Length));
RAData->Serialize(*Stream);
// IRData (inline)
IRList->Serialize(*Stream);
}
}
static bool readAll(int fd, void *data, size_t size) {
int rv = read(fd, data, size);
if (rv != size)
return false;
else
return true;
}
bool LoadAOTIRCache(AOTCacheType *AOTIRCache, int streamfd) {
uint64_t tag;
if (!readAll(streamfd, (char*)&tag, sizeof(tag)) || tag != FEXCore::IR::AOTIR_COOKIE)
return false;
std::string Module;
uint64_t ModSize;
uint64_t IndexSize;
lseek(streamfd, -sizeof(ModSize), SEEK_END);
if (!readAll(streamfd, (char*)&ModSize, sizeof(ModSize)))
return false;
Module.resize(ModSize);
lseek(streamfd, -sizeof(ModSize) - ModSize, SEEK_END);
if (!readAll(streamfd, (char*)&Module[0], Module.size()))
return false;
lseek(streamfd, -sizeof(ModSize) - ModSize - sizeof(IndexSize), SEEK_END);
if (!readAll(streamfd, (char*)&IndexSize, sizeof(IndexSize)))
return false;
struct stat fileinfo;
if (fstat(streamfd, &fileinfo) < 0)
return false;
size_t Size = (fileinfo.st_size + 4095) & ~4095;
size_t IndexOffset = fileinfo.st_size - IndexSize -sizeof(ModSize) - ModSize - sizeof(IndexSize);
void *FilePtr = FEXCore::Allocator::mmap(nullptr, Size, PROT_READ, MAP_SHARED, streamfd, 0);
if (FilePtr == MAP_FAILED) {
return false;
}
auto Array = (AOTIRInlineIndex *)((char*)FilePtr + IndexOffset);
AOTIRCache->insert({Module, {Array, FilePtr, Size}});
LogMan::Msg::DFmt("AOTIR: Module {} has {} functions", Module, Array->Count);
return true;
}
AOTIRCaptureCache::~AOTIRCaptureCache() {
for (auto &Mod: AOTIRCache) {
FEXCore::Allocator::munmap(Mod.second.mapping, Mod.second.size);
}
}
void AOTIRCaptureCache::FinalizeAOTIRCache() {
AOTIRCaptureCacheWriteoutQueue_Flush();
std::unique_lock lk(AOTIRCacheLock);
for (auto& [String, Entry] : AOTIRCaptureCacheMap) {
if (!Entry.Stream) {
continue;
}
const auto ModSize = String.size();
auto &stream = Entry.Stream;
// pad to 32 bytes
constexpr char Zero = 0;
while(stream->tellp() & 31)
stream->write(&Zero, 1);
// AOTIRInlineIndex
const auto FnCount = Entry.Index.size();
const size_t DataBase = -stream->tellp();
stream->write((const char*)&FnCount, sizeof(FnCount));
stream->write((const char*)&DataBase, sizeof(DataBase));
for (const auto& [GuestStart, DataOffset] : Entry.Index) {
//AOTIRInlineIndexEntry
// GuestStart
stream->write((const char*)&GuestStart, sizeof(GuestStart));
// DataOffset
stream->write((const char*)&DataOffset, sizeof(DataOffset));
}
// End of file header
const auto IndexSize = FnCount * sizeof(FEXCore::IR::AOTIRInlineIndexEntry) + sizeof(DataBase) + sizeof(FnCount);
stream->write((const char*)&IndexSize, sizeof(IndexSize));
stream->write(String.c_str(), ModSize);
stream->write((const char*)&ModSize, sizeof(ModSize));
// Close the stream
stream->close();
// Rename the file to atomically update the cache with the temporary file
AOTIRRenamer(String);
}
}
void AOTIRCaptureCache::AOTIRCaptureCacheWriteoutQueue_Flush() {
{
std::shared_lock lk{AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
for (;;) {
AOTIRCaptureCacheWriteoutLock.lock();
std::function<void()> fn = std::move(AOTIRCaptureCacheWriteoutQueue.front());
bool MaybeEmpty = false;
AOTIRCaptureCacheWriteoutQueue.pop();
MaybeEmpty = AOTIRCaptureCacheWriteoutQueue.size() == 0;
AOTIRCaptureCacheWriteoutLock.unlock();
fn();
if (MaybeEmpty) {
std::shared_lock lk{AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
}
LOGMAN_MSG_A_FMT("Must never get here");
}
void AOTIRCaptureCache::AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn) {
bool Flush = false;
{
std::unique_lock lk{AOTIRCaptureCacheWriteoutLock};
AOTIRCaptureCacheWriteoutQueue.push(fn);
if (AOTIRCaptureCacheWriteoutQueue.size() > 10000) {
Flush = true;
}
}
bool test_val = false;
if (Flush && AOTIRCaptureCacheWriteoutFlusing.compare_exchange_strong(test_val, true)) {
AOTIRCaptureCacheWriteoutQueue_Flush();
}
}
void AOTIRCaptureCache::WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
std::shared_lock lk(AOTIRCacheLock);
for( const auto &File: FilesWithCode) {
Writer(File.first, File.second);
}
}
AOTIRCaptureCache::PreGenerateIRFetchResult AOTIRCaptureCache::PreGenerateIRFetch(uint64_t GuestRIP, FEXCore::IR::IRListView *IRList) {
{
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
if (!file->second.ContainsCode) {
file->second.ContainsCode = true;
FilesWithCode[file->second.fileid] = file->second.filename;
}
}
}
PreGenerateIRFetchResult Result{};
if (IRList == nullptr && CTX->Config.AOTIRLoad()) {
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
auto Mod = (FEXCore::IR::AOTIRInlineIndex*)file->second.CachedFileEntry;
if (Mod == nullptr) {
file->second.CachedFileEntry = Mod = AOTIRCache[file->second.fileid].Array;
}
if (Mod != nullptr)
{
auto AOTEntry = Mod->Find(GuestRIP - file->second.Start + file->second.Offset);
if (AOTEntry) {
// verify hash
auto MappedStart = GuestRIP;
auto hash = XXH3_64bits((void*)MappedStart, AOTEntry->GuestLength);
if (hash == AOTEntry->GuestHash) {
Result.IRList = AOTEntry->GetIRData();
//LogMan::Msg::DFmt("using {} + {:x} -> {:x}\n", file->second.fileid, AOTEntry->first, GuestRIP);
Result.RAData = AOTEntry->GetRAData();;
Result.DebugData = new FEXCore::Core::DebugData();
Result.StartAddr = MappedStart;
Result.Length = AOTEntry->GuestLength;
Result.GeneratedIR = true;
} else {
LogMan::Msg::IFmt("AOTIR: hash check failed {:x}\n", MappedStart);
}
} else {
//LogMan::Msg::IFmt("AOTIR: Failed to find {:x}, {:x}, {}\n", GuestRIP, GuestRIP - file->second.Start + file->second.Offset, file->second.fileid);
}
}
}
}
return Result;
}
bool AOTIRCaptureCache::PostCompileCode(
FEXCore::Core::InternalThreadState *Thread,
void* CodePtr,
uint64_t GuestRIP,
uint64_t StartAddr,
uint64_t Length,
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR,
bool DecrementRefCount) {
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || CTX->Config.LibraryJITNaming()) {
std::shared_lock lk(AOTIRCacheLock);
auto file = FindAddrForFile(StartAddr, Length);
// Only go down this path if we actually found a library region
if (file != AddrToFile.end()) {
if (DebugData && CTX->Config.LibraryJITNaming()) {
CTX->Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, file->second.filename);
}
// Add to AOT cache if aot generation is enabled
if (GeneratedIR && RAData &&
(CTX->Config.AOTIRCapture() || CTX->Config.AOTIRGenerate())) {
auto hash = XXH3_64bits((void*)StartAddr, Length);
auto LocalRIP = GuestRIP - file->second.Start + file->second.Offset;
auto LocalStartAddr = StartAddr - file->second.Start + file->second.Offset;
auto fileid = file->second.fileid;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRList, RAData, fileid]() {
auto *AotFile = &AOTIRCaptureCacheMap[fileid];
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(fileid);
uint64_t tag = FEXCore::IR::AOTIR_COOKIE;
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
});
if (CTX->Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
if (DecrementRefCount) {
--Thread->CompileBlockReentrantRefCount;
}
Thread->CPUBackend->ClearCache();
return true;
}
}
}
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
}
}
return false;
}
AOTIRCaptureCache::AddrToFileMapType::iterator AOTIRCaptureCache::FindAddrForFile(uint64_t Entry, uint64_t Length) {
// Thread safety here! We are returning an iterator to the map object
// This needs the AOTIRCacheLock locked prior to coming in to the function
auto file = AddrToFile.lower_bound(Entry);
if (file != AddrToFile.begin()) {
--file;
if (file->second.Start <= Entry && (file->second.Start + file->second.Len) >= (Entry + Length)) {
return file;
}
}
return AddrToFile.end();
}
void AOTIRCaptureCache::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// TODO: Support overlapping maps and region splitting
auto base_filename = std::filesystem::path(filename).filename().string();
if (!base_filename.empty()) {
auto filename_hash = XXH3_64bits(filename.c_str(), filename.size());
auto fileid = base_filename + "-" + std::to_string(filename_hash) + "-";
// append optimization flags to the fileid
fileid += (CTX->Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) ? "S" : "s";
fileid += CTX->Config.TSOEnabled ? "T" : "t";
fileid += CTX->Config.ABILocalFlags ? "L" : "l";
fileid += CTX->Config.ABINoPF ? "p" : "P";
std::unique_lock lk(AOTIRCacheLock);
AddrToFile.insert({ Base, { Base, Size, Offset, fileid, filename, nullptr, false} });
if (CTX->Config.AOTIRLoad && !AOTIRCache.contains(fileid) && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
if (streamfd != -1) {
FEXCore::IR::LoadAOTIRCache(&AOTIRCache, streamfd);
close(streamfd);
}
}
}
}
void AOTIRCaptureCache::RemoveNamedRegion(uintptr_t Base, uintptr_t Size) {
std::unique_lock lk(AOTIRCacheLock);
// TODO: Support partial removing
AddrToFile.erase(Base);
}
}
+162
View File
@@ -0,0 +1,162 @@
#pragma once
#include <FEXCore/Config/Config.h>
#include <atomic>
#include <cstdint>
#include <functional>
#include <fstream>
#include <memory>
#include <map>
#include <unordered_map>
#include <shared_mutex>
#include <queue>
namespace FEXCore::Core {
struct DebugData;
}
namespace FEXCore::IR {
class RegisterAllocationData;
class IRListView;
constexpr auto COOKIE_VERSION = [](const char CookieText[4], uint32_t Version) {
uint64_t Cookie = Version;
Cookie <<= 32;
// Make the cookie text be the lower bits
Cookie |= CookieText[3];
Cookie <<= 8;
Cookie |= CookieText[2];
Cookie <<= 8;
Cookie |= CookieText[1];
Cookie <<= 8;
Cookie |= CookieText[0];
return Cookie;
};
constexpr static uint32_t AOTIR_VERSION = 0x0000'00004;
constexpr static uint64_t AOTIR_COOKIE = COOKIE_VERSION("FEXI", AOTIR_VERSION);
struct AOTIRInlineEntry {
uint64_t GuestHash;
uint64_t GuestLength;
/* RAData followed by IRData */
uint8_t InlineData[0];
IR::RegisterAllocationData *GetRAData();
IR::IRListView *GetIRData();
};
struct AOTIRInlineIndexEntry {
uint64_t GuestStart;
uint64_t DataOffset;
};
struct AOTIRInlineIndex {
uint64_t Count;
uint64_t DataBase;
AOTIRInlineIndexEntry Entries[0];
AOTIRInlineEntry *Find(uint64_t GuestStart);
AOTIRInlineEntry *GetInlineEntry(uint64_t DataOffset);
};
struct AOTIRCaptureCacheEntry {
std::unique_ptr<std::ofstream> Stream;
std::map<uint64_t, uint64_t> Index;
void AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView *IRList, FEXCore::IR::RegisterAllocationData *RAData);
};
struct AOTIRCacheEntry {
AOTIRInlineIndex *Array;
void *mapping;
size_t size;
};
using AOTCacheType = std::unordered_map<std::string, FEXCore::IR::AOTIRCacheEntry>;
bool LoadAOTIRCache(AOTCacheType *AOTIRCache, int streamfd);
class AOTIRCaptureCache final {
public:
AOTIRCaptureCache(FEXCore::Context::Context *ctx) : CTX {ctx} {}
~AOTIRCaptureCache();
void FinalizeAOTIRCache();
void AOTIRCaptureCacheWriteoutQueue_Flush();
void AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn);
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer);
struct PreGenerateIRFetchResult {
FEXCore::IR::IRListView *IRList {};
FEXCore::IR::RegisterAllocationData *RAData {};
FEXCore::Core::DebugData *DebugData {};
uint64_t StartAddr {};
uint64_t Length {};
bool GeneratedIR {};
};
[[nodiscard]] PreGenerateIRFetchResult PreGenerateIRFetch(uint64_t GuestRIP, FEXCore::IR::IRListView *IRList);
bool PostCompileCode(FEXCore::Core::InternalThreadState *Thread,
void* CodePtr,
uint64_t GuestRIP,
uint64_t StartAddr,
uint64_t Length,
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR,
bool DecrementRefCount);
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
// Callbacks
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) {
AOTIRLoader = CacheReader;
}
void SetAOTIRWriter(std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
AOTIRWriter = CacheWriter;
}
void SetAOTIRRenamer(std::function<void(const std::string&)> CacheRenamer) {
AOTIRRenamer = CacheRenamer;
}
private:
FEXCore::Context::Context *CTX;
std::shared_mutex AOTIRCacheLock;
std::shared_mutex AOTIRCaptureCacheWriteoutLock;
std::atomic<bool> AOTIRCaptureCacheWriteoutFlusing;
std::queue<std::function<void()>> AOTIRCaptureCacheWriteoutQueue;
std::map<std::string, std::string> FilesWithCode;
struct AddrToFileEntry {
uint64_t Start;
uint64_t Len;
uint64_t Offset;
std::string fileid;
std::string filename;
void *CachedFileEntry;
bool ContainsCode;
};
using AddrToFileMapType = std::map<uint64_t, AddrToFileEntry>;
AddrToFileMapType AddrToFile;
FEXCore::IR::AOTCacheType AOTIRCache;
std::function<int(const std::string&)> AOTIRLoader;
std::function<std::unique_ptr<std::ofstream>(const std::string&)> AOTIRWriter;
std::function<void(const std::string&)> AOTIRRenamer;
std::unordered_map<std::string, FEXCore::IR::AOTIRCaptureCacheEntry> AOTIRCaptureCacheMap;
AddrToFileMapType::iterator FindAddrForFile(uint64_t Entry, uint64_t Length);
};
}
+61 -1
View File
@@ -191,6 +191,19 @@
]
},
"ProcessorID": {
"Desc": ["Returns the processor ID correlating to the current running CPU",
"This may be out of date by time this instruction is executed so care must be taken",
"This same information can be gotten from syscall getcpu(&cpu, &node)",
"uint32_t Res = (node << 12) | cpu;",
"This means it has a limitation of 4096 CPU cores. Which is fine and matches x86 behaviour"
],
"OpClass": "Misc",
"HasDest": true,
"DestClass": "GPR",
"DestSize": 8
},
"SignalReturn": {
"HasSideEffects": true,
"OpClass": "Branch"
@@ -858,7 +871,22 @@
},
"CacheLineClear": {
"Desc": ["Does a 64 byte cacheline clear at the address specified"
"Desc": ["Does a 64 byte cacheline clear at the address specified",
"Only clears the data cachelines. Doesn't do any zeroing"
],
"HasSideEffects": true,
"OpClass": "Memory",
"SSAArgs": "1",
"SSANames": [
"Addr"
]
},
"CacheLineZero": {
"Desc": ["Does a 64 byte zero at the address specified",
"Writing zeroes to memory",
"It is specifically non-temporal and weakly ordered",
"This matches CLZero behaviour"
],
"HasSideEffects": true,
"OpClass": "Memory",
@@ -1074,6 +1102,38 @@
]
},
"PDep": {
"Desc": [
"Performs a parallel bit deposit.",
"Takes the contiguous low-order bits and deposits them into",
"the destination at the locations specified by the Mask."
],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"SSAArgs": "2",
"SSANames": [
"Input",
"Mask"
]
},
"PExt": {
"Desc": [
"Performs a parallel bit extract.",
"Each bit set in the mask will select the corresponding bit in the Input",
"and transfers them to the lower contiguous bits in the destination."
],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"SSAArgs": "2",
"SSANames": [
"Input",
"Mask"
]
},
"LDiv": {
"Desc": ["Integer long signed division returning lower bits",
"The Lower and Upper registers will be concated together to generate a dividend twice the size",
+3
View File
@@ -2,6 +2,9 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <array>
#include <sys/mman.h>
#ifdef ENABLE_JEMALLOC
#include <jemalloc/jemalloc.h>
+1 -1
View File
@@ -4,6 +4,7 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <algorithm>
#include <array>
@@ -17,7 +18,6 @@
#include <new>
#include <sstream>
#include <sys/mman.h>
#include <bits/mman-map-flags-generic.h>
#include <sys/utsname.h>
#include <type_traits>
#include <utility>
+42 -11
View File
@@ -1,11 +1,42 @@
#include <FEXCore/Utils/NetStream.h>
#include <array>
#include <cstring>
#include <iterator>
#include <sys/socket.h>
#include <unistd.h>
namespace FEXCore::Utils {
int NetStream::NetBuf::flushBuffer(const char *buffer, size_t size) {
namespace {
class NetBuf final : public std::streambuf {
public:
explicit NetBuf(int socketfd) : socket{socketfd} {
reset_output_buffer();
}
~NetBuf() override {
close(socket);
}
private:
std::streamsize xsputn(const char* buffer, std::streamsize size) override;
std::streambuf::int_type underflow() override;
std::streambuf::int_type overflow(std::streambuf::int_type ch) override;
int sync() override;
void reset_output_buffer() {
// we always leave room for one extra char
setp(std::begin(output_buffer), std::end(output_buffer) -1);
}
int flushBuffer(const char *buffer, size_t size);
int socket;
std::array<char, 1400> output_buffer;
std::array<char, 1500> input_buffer; // enough for a typical packet
};
int NetBuf::flushBuffer(const char *buffer, size_t size) {
size_t total = 0;
// Send data
@@ -21,7 +52,7 @@ int NetStream::NetBuf::flushBuffer(const char *buffer, size_t size) {
return 0;
}
std::streamsize NetStream::NetBuf::xsputn(const char* buffer, std::streamsize size) {
std::streamsize NetBuf::xsputn(const char* buffer, std::streamsize size) {
size_t buf_remaining = epptr() - pptr();
// Check if the string fits neatly in our buffer
@@ -45,23 +76,23 @@ std::streamsize NetStream::NetBuf::xsputn(const char* buffer, std::streamsize si
}
}
std::streambuf::int_type NetStream::NetBuf::overflow(std::streambuf::int_type ch) {
std::streambuf::int_type NetBuf::overflow(std::streambuf::int_type ch) {
// we always leave room for one extra char
*pptr() = (char) ch;
pbump(1);
return sync();
}
int NetStream::NetBuf::sync() {
int NetBuf::sync() {
// Flush and reset output buffer to zero
if(flushBuffer(pbase(), pptr() - pbase()) < 0) {
if (flushBuffer(pbase(), pptr() - pbase()) < 0) {
return -1;
}
reset_output_buffer();
return 0;
}
std::streambuf::int_type NetStream::NetBuf::underflow() {
std::streambuf::int_type NetBuf::underflow() {
ssize_t size = recv(socket, (void *)std::begin(input_buffer), sizeof(input_buffer), 0);
if (size <= 0) {
@@ -73,12 +104,12 @@ std::streambuf::int_type NetStream::NetBuf::underflow() {
return traits_type::to_int_type(*gptr());
}
} // Anonymous namespace
NetStream::NetStream(int socketfd) : std::iostream(new NetBuf(socketfd)) {}
NetStream::~NetStream() {
delete rdbuf();
}
NetStream::NetBuf::~NetBuf() {
close(socket);
}
}
} // namespace FEXCore::Utils
-1
View File
@@ -12,7 +12,6 @@
#include <sys/mman.h>
#include <sys/signal.h>
#include <sys/syscall.h>
#include <bits/mman-map-flags-generic.h>
#include <deque>
#include <unistd.h>
+12 -7
View File
@@ -94,7 +94,7 @@ namespace FEXCore::Context {
FEX_DEFAULT_VISIBILITY FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader);
FEX_DEFAULT_VISIBILITY void SetExitHandler(FEXCore::Context::Context *CTX, ExitHandler handler);
FEX_DEFAULT_VISIBILITY ExitHandler GetExitHandler(FEXCore::Context::Context *CTX);
FEX_DEFAULT_VISIBILITY ExitHandler GetExitHandler(const FEXCore::Context::Context *CTX);
/**
* @brief Pauses execution on the CPU core
@@ -134,7 +134,7 @@ namespace FEXCore::Context {
*
* @return The program exit status
*/
FEX_DEFAULT_VISIBILITY int GetProgramStatus(FEXCore::Context::Context *CTX);
FEX_DEFAULT_VISIBILITY int GetProgramStatus(const FEXCore::Context::Context *CTX);
/**
* @brief Tells the core to shutdown
@@ -157,7 +157,7 @@ namespace FEXCore::Context {
*
* @return The ExitReason for the parentthread
*/
FEX_DEFAULT_VISIBILITY ExitReason GetExitReason(FEXCore::Context::Context *CTX);
FEX_DEFAULT_VISIBILITY ExitReason GetExitReason(const FEXCore::Context::Context *CTX);
/**
* @brief [[theadsafe]] Checks if the Context is either done working or paused(in the case of single stepping)
@@ -168,7 +168,7 @@ namespace FEXCore::Context {
*
* @return true if the core is done or paused
*/
FEX_DEFAULT_VISIBILITY bool IsDone(FEXCore::Context::Context *CTX);
FEX_DEFAULT_VISIBILITY bool IsDone(const FEXCore::Context::Context *CTX);
/**
* @brief Gets a copy the CPUState of the parent thread
@@ -176,7 +176,8 @@ namespace FEXCore::Context {
* @param CTX The context that we created
* @param State The state object to populate
*/
FEX_DEFAULT_VISIBILITY void GetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State);
FEX_DEFAULT_VISIBILITY void GetCPUState(const FEXCore::Context::Context *CTX,
FEXCore::Core::CPUState *State);
/**
* @brief Copies the CPUState provided to the parent thread
@@ -184,7 +185,8 @@ namespace FEXCore::Context {
* @param CTX The context that we created
* @param State The satate object to copy from
*/
FEX_DEFAULT_VISIBILITY void SetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State);
FEX_DEFAULT_VISIBILITY void SetCPUState(FEXCore::Context::Context *CTX,
const FEXCore::Core::CPUState *State);
/**
* @brief Allows the frontend to pass in a custom CPUBackend creation factory
@@ -232,11 +234,14 @@ namespace FEXCore::Context {
FEX_DEFAULT_VISIBILITY void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation);
FEX_DEFAULT_VISIBILITY void SetSyscallHandler(FEXCore::Context::Context *CTX, FEXCore::HLE::SyscallHandler *Handler);
FEX_DEFAULT_VISIBILITY FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf);
FEX_DEFAULT_VISIBILITY FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf, uint32_t CPU);
FEX_DEFAULT_VISIBILITY void AddNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length, uintptr_t Offset, const std::string& Name);
FEX_DEFAULT_VISIBILITY void RemoveNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length);
FEX_DEFAULT_VISIBILITY void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader);
FEX_DEFAULT_VISIBILITY void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter);
FEX_DEFAULT_VISIBILITY void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter);
FEX_DEFAULT_VISIBILITY void SetAOTIRRenamer(FEXCore::Context::Context *CTX, std::function<void(const std::string&)> CacheRenamer);
FEX_DEFAULT_VISIBILITY void FinalizeAOTIRCache(FEXCore::Context::Context *CTX);
FEX_DEFAULT_VISIBILITY void WriteFilesWithCode(FEXCore::Context::Context *CTX, std::function<void(const std::string& fileid, const std::string& filename)> Writer);
FEX_DEFAULT_VISIBILITY void FlushCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length);
+21
View File
@@ -2,8 +2,11 @@
#include <FEXCore/Utils/CompilerDefs.h>
#include <algorithm>
#include <cstddef>
#include <cstdint>
#include <cstring>
#include <signal.h>
namespace FEXCore {
namespace x86_64 {
@@ -151,6 +154,24 @@ namespace FEXCore {
FEXCore::x86::sigval_t sigval;
} _timer;
} _sifields;
siginfo_t() = delete;
operator ::siginfo_t() const {
::siginfo_t val{};
val.si_signo = si_signo;
val.si_errno = si_errno;
val.si_code = si_code;
// Host siginfo has a pad member that is set to zeros
val.__pad0 = 0;
// Copy over the union
// The union is different sizes on 64-bit versus 32-bit
memcpy(val._sifields._pad, _sifields.pad, std::min(sizeof(val._sifields._pad), sizeof(_sifields.pad)));
return val;
}
};
static_assert(sizeof(FEXCore::x86::siginfo_t) == 128, "This needs to be the right size");
@@ -5,6 +5,7 @@
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/InterruptableConditionVariable.h>
#include <FEXCore/Utils/Threads.h>
#include <unordered_map>
@@ -83,7 +84,7 @@ namespace FEXCore::Core {
std::atomic<SignalEvent> SignalReason{SignalEvent::Nothing};
std::unique_ptr<FEXCore::Threads::Thread> ExecutionThread;
Event StartRunning;
InterruptableConditionVariable StartRunning;
Event ThreadWaiting;
std::unique_ptr<FEXCore::IR::OpDispatchBuilder> OpDispatcher;
+19 -17
View File
@@ -2,6 +2,7 @@
#include <FEXCore/Utils/CompilerDefs.h>
#include <array>
#include <cstdint>
#include <cstring>
#include <type_traits>
@@ -487,27 +488,28 @@ constexpr size_t MAX_XOP_GROUP_TABLE_SIZE = (1 << 6);
constexpr size_t MAX_EVEX_TABLE_SIZE = 256;
extern FEX_DEFAULT_VISIBILITY X86InstInfo BaseOps[MAX_PRIMARY_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo SecondBaseOps[MAX_SECOND_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo RepModOps[MAX_REP_MOD_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo RepNEModOps[MAX_REPNE_MOD_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo OpSizeModOps[MAX_OPSIZE_MOD_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo PrimaryInstGroupOps[MAX_INST_GROUP_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo SecondInstGroupOps[MAX_INST_SECOND_GROUP_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo SecondModRMTableOps[MAX_SECOND_MODRM_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo X87Ops[MAX_X87_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo DDDNowOps[MAX_3DNOW_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo H0F38TableOps[MAX_0F_38_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo H0F3ATableOps[MAX_0F_3A_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> RepModOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_OPSIZE_MOD_TABLE_SIZE> OpSizeModOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87Ops;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps;
// VEX
extern FEX_DEFAULT_VISIBILITY X86InstInfo VEXTableOps[MAX_VEX_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo VEXTableGroupOps[MAX_VEX_GROUP_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps;
// XOP
extern FEX_DEFAULT_VISIBILITY X86InstInfo XOPTableOps[MAX_XOP_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY X86InstInfo XOPTableGroupOps[MAX_XOP_GROUP_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> XOPTableOps;
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> XOPTableGroupOps;
// EVEX
extern FEX_DEFAULT_VISIBILITY X86InstInfo EVEXTableOps[MAX_EVEX_TABLE_SIZE];
extern FEX_DEFAULT_VISIBILITY std::array<X86InstInfo, MAX_EVEX_TABLE_SIZE> EVEXTableOps;
}
+1 -1
View File
@@ -43,7 +43,7 @@ friend class FEXCore::IR::PassManager;
IRPair<IROp_Constant> _Constant(uint8_t Size, uint64_t Constant) {
auto Op = AllocateOp<IROp_Constant, IROps::OP_CONSTANT>();
uint64_t Mask = ~0ULL >> (Size - 64);
uint64_t Mask = ~0ULL >> (64 - Size);
Op.first->Constant = (Constant & Mask);
Op.first->Header.Size = Size / 8;
Op.first->Header.ElementSize = Size / 8;
@@ -28,7 +28,9 @@ union PhysicalRegister {
static_assert(sizeof(PhysicalRegister) == 1);
class RegisterAllocationData {
// This class is serialized, can't have any holes in the structure
// otherwise ASAN complains about reading uninitialized memory
class FEX_PACKED RegisterAllocationData {
public:
uint32_t SpillSlotCount {};
uint32_t MapCount {};
@@ -43,6 +45,16 @@ class RegisterAllocationData {
static size_t Size(uint32_t NodeCount) {
return sizeof(RegisterAllocationData) + NodeCount * sizeof(Map[0]);
}
void Serialize(std::ostream& stream) const {
stream.write((const char*)&SpillSlotCount, sizeof(SpillSlotCount));
stream.write((const char*)&MapCount, sizeof(MapCount));
// RAData (inline)
// In file, IsShared is always set
bool _IsShared = true;
stream.write((const char*)&_IsShared, sizeof(IsShared));
stream.write((const char*)&Map[0], sizeof(Map[0]) * MapCount);
}
};
struct RegisterAllocationDataDeleter {
@@ -0,0 +1,89 @@
#pragma once
#include <atomic>
#include <chrono>
#include <climits>
#include <cstdint>
#include <linux/futex.h>
#include <sys/syscall.h>
#include <unistd.h>
namespace FEXCore {
/**
* @brief A condition variable that is robust against use of longjmp in signal handlers.
*
* This is opposed to common `std::condition_variable` implementations:
* Longjmp'ing in a signal handler while interrupting a pending `wait_for()`
* call can leave the condition variable in an invalid state that breaks later
* uses of that object and may cause hangs as a consequence.
*/
class InterruptableConditionVariable final {
public:
bool Wait(struct timespec *Timeout = nullptr) {
while (true) {
uint32_t Expected = SIGNALED;
uint32_t Desired = UNSIGNALED;
// If the mutex was already signaled then we can early exit
if (Mutex.compare_exchange_strong(Expected, Desired)) {
return true;
}
constexpr int Op = FUTEX_WAIT | FUTEX_PRIVATE_FLAG;
// WAIT will keep sleeping on the futex word while it is `val`
int Result = ::syscall(SYS_futex,
&Mutex,
Op,
Desired, // val
Timeout, // Timeout/val2
nullptr, // Addr2
0); // val3
if (Timeout && Result == -1 && errno == ETIMEDOUT) {
return false;
}
}
}
template<class Rep, class Period>
bool WaitFor(std::chrono::duration<Rep, Period> const& time) {
struct timespec Timeout{};
auto SecondsDuration = std::chrono::duration_cast<std::chrono::seconds>(time);
Timeout.tv_sec = SecondsDuration.count();
Timeout.tv_nsec = std::chrono::duration_cast<std::chrono::nanoseconds>(time - SecondsDuration).count();
return Wait(&Timeout);
}
void NotifyOne() {
DoNotify(1);
}
void NotifyAll() {
// Maximum number of waiters
DoNotify(INT_MAX);
}
private:
std::atomic<uint32_t> Mutex{};
constexpr static uint32_t SIGNALED = 1;
constexpr static uint32_t UNSIGNALED = 0;
void DoNotify(int Waiters) {
uint32_t Expected = UNSIGNALED;
uint32_t Desired = SIGNALED;
// If the mutex was in an unsignaled state then signal
if (Mutex.compare_exchange_strong(Expected, Desired)) {
constexpr int Op = FUTEX_WAKE | FUTEX_PRIVATE_FLAG;
::syscall(SYS_futex,
&Mutex,
Op,
Waiters, // val - Number of waiters to wake
0, // val2
&Mutex, // Addr2 - Mutex to do the operation on
0); // val3
}
}
};
}
+3 -36
View File
@@ -2,45 +2,12 @@
#include <FEXCore/Utils/CompilerDefs.h>
#include <array>
#include <iostream>
#include <iterator>
#include <string.h>
namespace FEXCore::Utils {
class FEX_DEFAULT_VISIBILITY NetStream : public std::iostream {
public:
NetStream(int socketfd) : std::iostream(new NetBuf(socketfd)) {}
virtual ~NetStream();
private:
class NetBuf : public std::streambuf {
public:
NetBuf(int socketfd) {
socket = socketfd;
reset_output_buffer();
}
virtual ~NetBuf();
protected:
virtual std::streamsize xsputn(const char* buffer, std::streamsize size);
virtual std::streambuf::int_type underflow();
virtual std::streambuf::int_type overflow(std::streambuf::int_type ch);
virtual int sync();
private:
void reset_output_buffer() {
// we always leave room for one extra char
setp(std::begin(output_buffer), std::end(output_buffer) -1);
}
int flushBuffer(const char *buffer, size_t size);
int socket;
std::array<char, 1400> output_buffer;
std::array<char, 1500> input_buffer; // enough for a typical packet
};
explicit NetStream(int socketfd);
~NetStream() override;
};
}
} // namespace FEXCore::Utils
Submodule External/Vulkan-Docs deleted from a0960966d5.
+1 -1
+61
View File
@@ -0,0 +1,61 @@
# Check for syscall support here
check_cxx_source_compiles(
"
#include <sched.h>
int main() {
return ::getcpu(nullptr, nullptr);
}"
compiles)
if (compiles)
message(STATUS "Has getcpu helper")
add_definitions(-DHAS_SYSCALL_GETCPU=1)
endif ()
check_cxx_source_compiles(
"
#include <unistd.h>
int main() {
return ::gettid();
}"
compiles)
if (compiles)
message(STATUS "Has gettid helper")
add_definitions(-DHAS_SYSCALL_GETTID=1)
endif ()
check_cxx_source_compiles(
"
#include <signal.h>
int main() {
return ::tgkill(0, 0, 0);
}"
compiles)
if (compiles)
message(STATUS "Has tgkill helper")
add_definitions(-DHAS_SYSCALL_TGKILL=1)
endif ()
check_cxx_source_compiles(
"
#include <sys/stat.h>
int main() {
return ::statx(0, nullptr, 0, 0, nullptr);
}"
compiles)
if (compiles)
message(STATUS "Has statx helper")
add_definitions(-DHAS_SYSCALL_STATX=1)
endif ()
check_cxx_source_compiles(
"
#include <stdio.h>
int main() {
return ::renameat2(0, nullptr, 0, nullptr, 0);
}"
compiles)
if (compiles)
message(STATUS "Has renameat2 helper")
add_definitions(-DHAS_SYSCALL_RENAMEAT2=1)
endif ()
+73
View File
@@ -0,0 +1,73 @@
#pragma once
#include <cstdint>
#include <sched.h>
#include <signal.h>
#include <stdio.h>
#include <syscall.h>
#include <sys/stat.h>
#include <unistd.h>
namespace FHU::Syscalls {
#ifndef MAP_FIXED_NOREPLACE
#define MAP_FIXED_NOREPLACE 0x100000
#endif
#ifndef SEM_STAT_ANY
#define SEM_STAT_ANY 20
#endif
#ifndef SHM_STAT_ANY
#define SHM_STAT_ANY 15
#endif
#ifndef MSG_STAT_ANY
#define MSG_STAT_ANY 13
#endif
#ifndef CLONE_PIDFD
#define CLONE_PIDFD 0x00001000
#endif
inline int32_t getcpu(uint32_t *cpu, uint32_t *node) {
// Third argument is unused
#if defined(HAS_SYSCALL_GETCPU) && HAS_SYSCALL_GETCPU
return ::getcpu(cpu, node, nullptr);
#else
return ::syscall(SYS_getcpu, cpu, node, nullptr);
#endif
}
inline int32_t gettid() {
#if defined(HAS_SYSCALL_GETTID) && HAS_SYSCALL_GETTID
return ::gettid();
#else
return ::syscall(SYS_gettid);
#endif
}
inline int32_t tgkill(pid_t tgid, pid_t tid, int sig) {
#if defined(HAS_SYSCALL_GETTID) && HAS_SYSCALL_GETTID
return ::tgkill(tggid, tid, sig);
#else
return ::syscall(SYS_tgkill, tgid, tid, sig);
#endif
}
inline int32_t statx(int dirfd, const char *pathname, int32_t flags, uint32_t mask, void *statxbuf) {
#if defined(HAS_SYSCALL_STATX) && HAS_SYSCALL_STATX
return ::statx(dirfd, pathname, flags, mask, statxbuf);
#else
return ::syscall(SYS_statx, dirfd, pathname, flags, mask, statxbuf);
#endif
}
inline int32_t renameat2(int olddirfd, const char *oldpath, int newdirfd, const char *newpath, unsigned int flags) {
#if defined(HAS_SYSCALL_STATX) && HAS_SYSCALL_STATX
return ::renameat2(olddirfd, oldpath, newdirfd, newpath, flags);
#else
return ::syscall(SYS_renameat2, olddirfd, oldpath, newdirfd, newpath, flags);
#endif
}
}
+15
View File
@@ -4,10 +4,25 @@ It has native support for a rootfs overlay, so you don't need to chroot, as well
FEX presents a Linux 5.0 interface to the guest, and supports both AArch64 and x86-64 as hosts.
FEX is very much work in progress, so expect things to change.
## Quick start guide
### For Ubuntu 20.04, 21.04, 21.10, 22.04
Execute the following command in the terminal to install FEX through a PPA.
`curl --silent https://raw.githubusercontent.com/FEX-Emu/FEX/main/Scripts/InstallFEX.py | python3`
This command will walk you through installing FEX through a PPA, and downloading a RootFS for use with FEX.
Ubuntu PPA is updated with our monthly releases.
### For everyone else
Follow the guide on the official FEX-Emu Wiki [Here](https://wiki.fex-emu.org/index.php/QuickStartGuide)
## Getting Started
FEX has been tested to build and run on ARMv8.0, ARMv8.1+, and x86-64(AVX or newer) hardware.
ARMv7 and older x86 hardware will not work.
Expected operating system usage is Linux. FEX has been tested with Ubuntu 20.04, 20.10, and 21.04. Also Arch Linux.
On AArch64 hosts the user **MUST** have an x86-64 RootFS [Creating a RootFS](#RootFS-Generation).
### Navigating the Source
+399
View File
@@ -0,0 +1,399 @@
#!/usr/bin/python3
import os
import subprocess
import sys
import tempfile
_Arch = None
def GetArch():
global _Arch
if _Arch == None:
_Arch = subprocess.check_output(['uname', '-m']).decode("utf-8").strip()
return _Arch
_Distro = None
def GetDistro():
global _Distro
# Query files in order
# /etc/lsb-release
# /etc/os-release
if _Distro == None:
if os.path.exists("/etc/lsb-release"):
File = open("/etc/lsb-release", "r")
Lines = File.readlines()
File.close()
Found = 0
Distro = ""
Version = ""
for Line in Lines:
Key, Val = Line.split("=", 1)
if Key == "DISTRIB_ID":
Distro = Val.strip().lower()
Found+=1
if Key == "DISTRIB_RELEASE":
Version = Val.strip()
Found+=1
if Found == 2:
_Distro = [Distro, Version]
return _Distro
if os.path.exists("/etc/os-release"):
File = open("/etc/os-release", "r")
Lines = File.readlines()
File.close()
Found = 0
Distro = ""
Version = ""
for Line in Lines:
Key, Val = Line.split("=", 1)
if Key == "ID":
Distro = Val.strip()
Found+=1
if Key == "VERSION_ID":
# Strip the double quotes from the version id
Version = Val.strip()[1:-1]
Found+=1
if Found == 2:
_Distro = [Distro, Version]
return _Distro
# Unknown
_Distro = ["Unknown", "0.0"]
return _Distro
def IsSupportedArch():
Arch = GetArch()
return Arch == "aarch64"
def IsSupportedDistro():
Distro = GetDistro()
# We only support Ubuntu
if Distro[0] == "ubuntu":
# We only support what is available in ppa:fex-emu/fex
return Distro[1] == "20.04" or \
Distro[1] == "21.04" or \
Distro[1] == "21.10" or \
Distro[1] == "22.04"
return False
_ArchVersion = None
def ListContainsRequired(Features, RequiredFeatures):
for Req in RequiredFeatures:
if not Req in Features:
return False
return True
def GetCPUFeaturesVersion():
global _ArchVersion
# Also LOR but kernel doesn't expose this
v8_1Mandatory = ["atomics", "asimdrdm", "crc32"]
v8_2Mandatory = v8_1Mandatory + ["dcpop"]
v8_3Mandatory = v8_2Mandatory + ["fcma", "jscvt", "lrcpc", "paca", "pacg"]
v8_4Mandatory = v8_3Mandatory + ["asimddp", "flagm", "ilrcpc", "uscat"]
# fphp asimdhp asimddp
if _ArchVersion == None:
File = open("/proc/cpuinfo", "r")
Lines = File.readlines()
File.close()
# Minimum spec is ARMv8.0
_ArchVersion = "8.0"
for Line in Lines:
if "Features" in Line:
Features = Line.split(":")[1].strip().split(" ")
# We don't care beyond 8.4 right now
if ListContainsRequired(Features, v8_4Mandatory):
_ArchVersion = "8.4"
elif ListContainsRequired(Features, v8_3Mandatory):
_ArchVersion = "8.3"
elif ListContainsRequired(Features, v8_2Mandatory):
_ArchVersion = "8.2"
elif ListContainsRequired(Features, v8_1Mandatory):
_ArchVersion = "8.1"
break;
return _ArchVersion
_PPAInstalled = None
FEXPPA = "http://ppa.launchpad.net/fex-emu/fex/ubuntu"
def GetPPAStatus():
global _PPAInstalled
if _PPAInstalled == None:
_PPAInstalled = False
CacheResults = subprocess.check_output(['apt-cache', 'policy']).decode("utf-8")
for Line in CacheResults.split("\n"):
if "http" in Line:
Line = Line.strip()
LineSplit = Line.split(" ")
# 'status' 'URL' 'series' 'arch' 'type'
if LineSplit[1] == FEXPPA:
_PPAInstalled = True
break
return _PPAInstalled
def InstallPPA():
print ("Installing PPA: ppa:fex-emu/fex")
print ("This bit will ask for your password")
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "add-apt-repository", "-y", "ppa:fex-emu/fex"])
DidInstall = CmdResult == 0
except KeyboardInterrupt:
DidInstall = False
pass
if DidInstall:
print("PPA installed")
else:
print("PPA failed to install")
return DidInstall
ARMVersionToPackage = {
"8.0": "fex-emu-armv8.0",
"8.1": "fex-emu-armv8.0",
"8.2": "fex-emu-armv8.2",
"8.3": "fex-emu-armv8.2",
"8.4": "fex-emu-armv8.4",
}
def GetPackagesToInstall():
return [
ARMVersionToPackage[GetCPUFeaturesVersion()],
"fex-emu-binfmt32",
"fex-emu-binfmt64",
]
def UpdatePPA():
print ("Updating apt sources")
print ("This bit will ask for your password")
DidUpdate = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "update"])
DidUpdate = CmdResult == 0
except KeyboardInterrupt:
DidUpdate = False
pass
if DidUpdate:
print("PPA installed")
else:
print("PPA failed to install")
return DidUpdate
def CheckAndInstallPackageUpdates():
PackagesToInstall = GetPackagesToInstall()
for Package in PackagesToInstall[:]:
UpgradableStatus = subprocess.check_output(["apt", "list", "--upgradable", Package]).decode("utf-8")
Found = False
for Line in UpgradableStatus.split("\n"):
# If the package exists to be upgraded then it will appear in this list
# We need to check multiple lines
# $ apt list --upgradable <Package>
# With upgrade available
# Listing... Done
# <Package>/<Repo> <NewVersion> <arch> [upgradable from: <Installed version>]
# Without upgrade available
# Listing... Done
# <EOF>
if Package in Line and "upgradable" in Line:
Found = True
if Found == False:
PackagesToInstall.remove(Package)
if len(PackagesToInstall) > 0:
print ("Found updates for packages: {}".format(PackagesToInstall))
print ("This bit may ask for your password")
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + PackagesToInstall)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages updated")
else:
print("Packages failed to update")
return DidInstall
return True
def CheckPackageInstallStatus():
PackagesToInstall = GetPackagesToInstall()
for Package in PackagesToInstall[:]:
CmdResult = subprocess.call(["dpkg", "-s", Package], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
if CmdResult == 0:
PackagesToInstall.remove(Package)
return PackagesToInstall
def InstallPackages(Packages):
print("Installing packages: {}".format(Packages))
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + Packages)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages installed")
else:
print("Packages failed to install")
return DidInstall
_RootFSPath = None
def GetRootFSPath():
global _RootFSPath
if _RootFSPath == None:
# Follows the same logic as FEXCore::Config::GetDataDirectory()
HomeDir = os.getenv("HOME")
if HomeDir == None:
HomeDir = os.getenv("PWD")
if HomeDir == None:
HomeDir = "."
Path = HomeDir
DataXDG = os.getenv("XDG_DATA_HOME")
if DataXDG != None:
Path = DataXDG
Path = Path + "/.fex-emu"
DataOverride = os.getenv("FEX_APP_DATA_LOCATION")
if DataOverride != None:
Path = DataOverride
_RootFSPath = Path + "/RootFS/"
return _RootFSPath
def CheckRootFSInstallStatus():
# Matches what is available on https://rootfs.fex-emu.org/file/fex-rootfs/RootFS_links.txt
UbuntuVersionToRootFS = {
"20.04": "Ubuntu_21_04.sqsh",
"21.04": "Ubuntu_21_04.sqsh",
"21.10": "Ubuntu_21_10.sqsh",
"22.04": "Ubuntu_21_10.sqsh",
}
return os.path.exists(GetRootFSPath() + UbuntuVersionToRootFS[GetDistro()[1]])
def TryInstallRootFS():
DidInstall = False
try:
CmdResult = subprocess.call(["FEXRootFSFetcher"])
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
return DidInstall
def TryBasicProgramExecution():
return subprocess.call(["FEXInterpreter", "/usr/bin/uname", "-a"]) == 0
def ExitWithStatus(Status):
# Remove the cached credentials
subprocess.call(["sudo", "-K"])
sys.exit(Status)
def main():
# Only run on supported arch
if not IsSupportedArch():
print ( "{} is not a supported architecture".format(GetArch()))
ExitWithStatus(-1)
if not IsSupportedDistro():
Distro = GetDistro()
print ( "'{} {}' is not a supported distro".format(Distro[0], Distro[1]))
ExitWithStatus(-1)
if GetDistro()[0] == "ubuntu":
print ("Getting PPA status: {}".format(("NotInstalled", "Installed")[GetPPAStatus()]))
if GetPPAStatus():
if not UpdatePPA():
print ("apt sources failed to update. Not continuing")
ExitWithStatus(-1)
if not CheckAndInstallPackageUpdates():
print ("apt packages failed to update. Not continuing")
ExitWithStatus(-1)
else:
if not InstallPPA():
print ("PPA failed to install. Not continuing")
ExitWithStatus(-1)
Packages = CheckPackageInstallStatus()
if len(Packages) > 0:
if not InstallPackages(Packages):
print ("Failed to install packages. Not continuing")
ExitWithStatus(-1)
if not CheckRootFSInstallStatus():
print ("RootFS not found. Running FEXRootFSFetcher to get rootfs")
if not TryInstallRootFS():
print ("Failed to install RootFS. Not continuing")
ExitWithStatus(-1)
print ("FEX is now installed. Trying basic program run")
if not TryBasicProgramExecution():
print ("FEXInterpreter failed to run. Not continuing")
ExitWithStatus(-1)
print ("")
print ("===================================================")
print ("FEX test run executed. You should be set to run FEX")
print ("===================================================")
print ("Usage examples:")
print ("# steam is a bash script. Wrap with FEXBash")
print ("\tFEXBash steam")
print ("# Full path execution execution will wrap the application if it exists in the rootfs")
print ("\tFEXInterpreter /usr/bin/uname")
print ("# Freestanding x86/x86-64 programs can be executed directly. binfmt_misc will redirect to FEX")
print ("\t$HOME/PetalCrashOnline.AppImage")
print ("# If you need a terminal that emulates everything.")
print ("# Run FEXBash without arguments. Double check uname to see if running under FEX")
print ("\tFEXBash")
ExitWithStatus(0)
if __name__ == "__main__":
sys.exit(main())
+1 -1
View File
@@ -9,6 +9,6 @@ set(SRCS
SocketLogging.cpp)
add_library(${NAME} STATIC ${SRCS})
target_link_libraries(${NAME} cpp-optparse json-maker FEXCore)
target_link_libraries(${NAME} FEXCore_Base cpp-optparse json-maker)
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/External/cpp-optparse/)
target_include_directories(${NAME} PRIVATE ${CMAKE_BINARY_DIR}/generated)
+39
View File
@@ -1,8 +1,10 @@
#include "Common/ArgumentLoader.h"
#include "Common/Config.h"
#include <FEXCore/Config/Config.h>
#include <cstring>
#include <filesystem>
#include <fstream>
#include <map>
#include <list>
@@ -38,4 +40,41 @@ namespace FEX::Config {
}
}
std::string LoadConfig(
bool NoFEXArguments,
bool LoadProgramConfig,
int argc,
char **argv,
char **const envp) {
FEXCore::Config::Initialize();
FEXCore::Config::AddLayer(FEXCore::Config::CreateMainLayer());
if (NoFEXArguments) {
FEX::ArgLoader::LoadWithoutArguments(argc, argv);
}
else {
FEXCore::Config::AddLayer(std::make_unique<FEX::ArgLoader::ArgLoader>(argc, argv));
}
FEXCore::Config::AddLayer(FEXCore::Config::CreateEnvironmentLayer(envp));
FEXCore::Config::Load();
auto Args = FEX::ArgLoader::Get();
if (LoadProgramConfig) {
if (Args.empty()) {
// Early exit if we weren't passed an argument
return {};
}
std::string Program = Args[0];
// These layers load on initialization
auto ProgramName = std::filesystem::path(Program).filename();
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, true));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, false));
return Program;
}
return {};
}
}
+8
View File
@@ -18,4 +18,12 @@ namespace FEX::Config {
};
void SaveLayerToJSON(const std::string& Filename, FEXCore::Config::Layer *const Layer);
std::string LoadConfig(
bool NoFEXArguments,
bool LoadProgramConfig,
int argc,
char **argv,
char **const envp
);
}
+57 -25
View File
@@ -132,7 +132,6 @@ bool SendSocketPipe(std::string const &MountPath) {
strncpy(addr.sun_path, SocketPath.data(), sizeof(addr.sun_path));
if (connect(socket_fd, reinterpret_cast<struct sockaddr*>(&addr), sizeof(addr)) == -1) {
LogMan::Msg::DFmt("Couldn't connect to AF_UNIX socket: {} {}", errno, strerror(errno));
close(socket_fd);
return false;
}
@@ -155,14 +154,13 @@ bool SendSocketPipe(std::string const &MountPath) {
int Result = ppoll(&pfd, 1, &ts, nullptr);
if (Result == -1 || Result == 0) {
// didn't get ack back in time
// Close our read pipe
close(fds[0]);
// close our write pipe
close(fds[1]);
// close socket
close(socket_fd);
return false;
// Assume an overburdened system at this point
// If the FEXMountDaemon is alive but slept for more than our timeout
// then we can spuriously throw errors
//
// Returning false here would cause FEX to try and spin up a new FEXMountDaemon
// and then the FEXMountDaemon would check to see if the lock exists
// Then would early exit and not mount a new path
}
// We've sent the message which means we're done with the socket
@@ -215,11 +213,13 @@ std::string GetRootFSLockFile() {
return LockPath;
}
bool Setup(char **const envp) {
// We need to setup the rootfs here
// If the configuration is set to use a folder then there is nothing to do
// If it is setup to use a squashfs then we need to do something more complex
enum class ErrorResult {
ERROR_SUCCESS,
ERROR_FAIL,
ERROR_TRYAGAIN,
};
ErrorResult SetupSquashFS(char **const envp) {
FEX_CONFIG_OPT(LDPath, ROOTFS);
// XXX: Disabled for now. Causing problems due to weird filesystem problems
// Can reproduce by attempting to run an application under pressure-vessel
@@ -231,7 +231,7 @@ bool Setup(char **const envp) {
// If we are inside of a rootfs/container then drop the rootfs path
// Root is already our rootfs
FEXCore::Config::Erase(FEXCore::Config::CONFIG_ROOTFS);
return true;
return ErrorResult::ERROR_SUCCESS;
}
}
@@ -244,8 +244,10 @@ bool Setup(char **const envp) {
// If the lock file exists and we can send the process a pipe then nothing to do
// Otherwise we need to spin up a new mount daemon
std::string MountPath{};
if (CheckLockExists(LockPath, &MountPath) && SendSocketPipe(MountPath)) {
return true;
bool LockExists = CheckLockExists(LockPath, &MountPath);
bool SentSocketPipe = SendSocketPipe(MountPath);
if (LockExists && SentSocketPipe) {
return ErrorResult::ERROR_SUCCESS;
}
pid_t ParentTID = ::getpid();
@@ -255,21 +257,21 @@ bool Setup(char **const envp) {
// Make the temporary mount folder
if (mkdtemp(TempFolder) == nullptr) {
LogMan::Msg::EFmt("Couldn't create temporary mount name: {}", TempFolder);
return false;
return ErrorResult::ERROR_FAIL;
}
// Change the permissions
if (chmod(TempFolder, 0777) != 0) {
LogMan::Msg::EFmt("Couldn't change permissions on temporary mount: {}", TempFolder);
rmdir(TempFolder);
return false;
return ErrorResult::ERROR_FAIL;
}
// Open some pipes for communicating with the new processes
int fds[2]{};
if (pipe2(fds, 0) != 0) {
LogMan::Msg::EFmt("Couldn't open pipe");
return false;
return ErrorResult::ERROR_FAIL;
}
// Convert the write pipe to a string to pass to the child process
@@ -281,13 +283,13 @@ bool Setup(char **const envp) {
// Child
close(fds[0]); // Close read end of pipe
const char *argv[5];
argv[0] = FEX_INSTALL_PREFIX "/bin/FEXMountDaemon";
argv[0] = "FEXMountDaemon";
argv[1] = LDPath().c_str();
argv[2] = TempFolder;
argv[3] = PipeString.c_str();
argv[4] = nullptr;
if (execve(argv[0], (char * const*)argv, envp) == -1) {
if (execvpe(argv[0], (char * const*)argv, envp) == -1) {
// Let the parent know that we couldn't execute for some reason
uint64_t error{1};
write(fds[1], &error, sizeof(error));
@@ -319,13 +321,13 @@ bool Setup(char **const envp) {
if (Result != sizeof(ChildResult)) {
LogMan::Msg::DFmt("Spurious read error");
return false;
return ErrorResult::ERROR_FAIL;
}
if (ChildResult == 1) {
// Error
LogMan::Msg::DFmt("FEXMountDaemon couldn't mount child for some reason");
return false;
return ErrorResult::ERROR_FAIL;
}
// Open the lock to let the daemon know that it has an active user
@@ -333,7 +335,16 @@ bool Setup(char **const envp) {
// Check if we have an directory inside our temp folder
if (!SanityCheckPath(TempFolder)) {
return false;
if (!SentSocketPipe) {
// If the lock exists but we weren't able to send the socket pipe
// We can remove the lock and try again. This is likely due to a stale
// lock file due to unclean shutdown
unlink(LockPath.c_str());
return ErrorResult::ERROR_TRYAGAIN;
}
LogMan::Msg::DFmt("\tOpened RootFS lock file {} but RootFS doesn't exist", LockPath);
LogMan::Msg::DFmt("\tStale lock file or race on startup? Try `rm {}` and run again?", LockPath);
return ErrorResult::ERROR_FAIL;
}
// Send the new FEXMountDaemon a pipe to listen to
@@ -341,11 +352,32 @@ bool Setup(char **const envp) {
// If everything has passed then we can now update the rootfs path
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, TempFolder);
return true;
return ErrorResult::ERROR_SUCCESS;
}
}
return ErrorResult::ERROR_SUCCESS;
}
bool Setup(char **const envp, uint32_t TryCount) {
// We need to setup the rootfs here
// If the configuration is set to use a folder then there is nothing to do
// If it is setup to use a squashfs then we need to do something more complex
if (TryCount > 5) {
// Fail if we have retried too many times
return false;
}
// Nothing to do
auto Result = SetupSquashFS(envp);
if (Result == ErrorResult::ERROR_FAIL) {
return false;
}
else if (Result == ErrorResult::ERROR_TRYAGAIN) {
return Setup(envp, TryCount + 1);
}
// ERROR_SUCCESS
return true;
}
+1 -1
View File
@@ -12,6 +12,6 @@ namespace FEX::RootFS {
std::string GetRootFSSocketFile(std::string const &MountPath);
// Checks if the rootfs lock exists
bool CheckLockExists(std::string const &LockPath, std::string *MountPath = nullptr);
bool Setup(char **const envp);
bool Setup(char **const envp, uint32_t TryCount = 0);
void Shutdown();
}
+2 -1
View File
@@ -2,6 +2,7 @@
#include <FEXCore/Utils/NetStream.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <atomic>
#include <fcntl.h>
@@ -48,7 +49,7 @@ namespace FEX::SocketLogging {
.Timestamp = Timestamp,
.PacketType = Type,
.PID = ::getpid(),
.TID = ::gettid(),
.TID = FHU::Syscalls::gettid(),
};
return Msg;
+162
View File
@@ -0,0 +1,162 @@
#include "ELFCodeLoader2.h"
#include "Linux/Utils/ELFContainer.h"
#include <FEXCore/Core/Context.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cstddef>
#include <set>
#include <sys/resource.h>
#include <sys/sysinfo.h>
#include <thread>
#include <queue>
namespace FEX::AOT {
void AOTGenSection(FEXCore::Context::Context *CTX, ELFCodeLoader2::LoadedSection &Section) {
// Make sure this section is executable and big enough
if (!Section.Executable || Section.Size < 16)
return;
std::set<uintptr_t> InitialBranchTargets;
// Load the ELF again with symbol parsing this time
ELFLoader::ELFContainer container{Section.Filename, "", true};
// Add symbols to the branch targets list
container.AddSymbols([&](ELFLoader::ELFSymbol* sym) {
auto Destination = sym->Address + Section.ElfBase;
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) ) {
return; // outside of current section, unlikely to be real code
}
InitialBranchTargets.insert(Destination);
});
LogMan::Msg::IFmt("Symbol seed: {}", InitialBranchTargets.size());
// Add unwind entries to the branch target list
container.AddUnwindEntries([&](uintptr_t Entry) {
auto Destination = Entry + Section.ElfBase;
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) ) {
return; // outside of current section, unlikely to be real code
}
InitialBranchTargets.insert(Destination);
});
LogMan::Msg::IFmt("Symbol + Unwind seed: {}", InitialBranchTargets.size());
// Scan the executable section and try to find function entries
for (size_t Offset = 0; Offset < (Section.Size - 16); Offset++) {
uint8_t *pCode = (uint8_t *)(Section.Base + Offset);
// Possible CALL <disp32>
if (*pCode == 0xE8) {
uintptr_t Destination = (int)(pCode[1] | (pCode[2] << 8) | (pCode[3] << 16) | (pCode[4] << 24));
Destination += (uintptr_t)pCode + 5;
auto DestinationPtr = (uint8_t*)Destination;
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) )
continue; // outside of current section, unlikely to be real code
if (DestinationPtr[0] == 0 && DestinationPtr[1] == 0)
continue; // add al, [rax], unlikely to be real code
InitialBranchTargets.insert(Destination);
}
// endbr64 marker marks an indirect branch destination
if (pCode[0] == 0xf3 && pCode[1] == 0x0f && pCode[2] == 0x1e && pCode[3] == 0xfa) {
InitialBranchTargets.insert((uintptr_t)pCode);
}
}
uint64_t SectionMaxAddress = Section.Base + Section.Size;
std::set<uint64_t> Compiled;
std::atomic<int> counter = 0;
std::queue<uint64_t> BranchTargets;
// Setup BranchTargets, Compiled sets from InitiaBranchTargets
Compiled.insert(InitialBranchTargets.begin(), InitialBranchTargets.end());
for (auto BranchTarget: InitialBranchTargets) {
BranchTargets.push(BranchTarget);
}
InitialBranchTargets.clear();
std::mutex QueueMutex;
std::vector<std::thread> ThreadPool;
for (int i = 0; i < get_nprocs_conf(); i++) {
std::thread thd([&BranchTargets, CTX, &counter, &Compiled, &Section, &QueueMutex, SectionMaxAddress]() {
// Set the priority of the thread so it doesn't overwhelm the system when running in the background
setpriority(PRIO_PROCESS, FHU::Syscalls::gettid(), 19);
// Setup thread - Each compilation thread uses its own backing FEX thread
FEXCore::Core::CPUState state;
auto Thread = FEXCore::Context::CreateThread(CTX, &state, FHU::Syscalls::gettid());
std::set<uint64_t> ExternalBranchesLocal;
FEXCore::Context::ConfigureAOTGen(Thread, &ExternalBranchesLocal, SectionMaxAddress);
for (;;) {
uint64_t BranchTarget;
// Get a entrypoint to process from the queue
QueueMutex.lock();
if (BranchTargets.empty()) {
QueueMutex.unlock();
break; // no entrypoint to process - exit
}
BranchTarget = BranchTargets.front();
BranchTargets.pop();
QueueMutex.unlock();
// Compile entrypoint
counter++;
FEXCore::Context::CompileRIP(Thread, BranchTarget);
// Are there more branches?
if (ExternalBranchesLocal.size() > 0) {
// Add them to the "to process" list
QueueMutex.lock();
for(auto Destination: ExternalBranchesLocal) {
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) )
continue;
if (Compiled.contains(Destination))
continue;
Compiled.insert(Destination);
BranchTargets.push(Destination);
}
QueueMutex.unlock();
ExternalBranchesLocal.clear();
}
}
// All entryproints processed, cleanup this thread
FEXCore::Context::DestroyThread(CTX, Thread);
});
// Add to the thread pool
ThreadPool.push_back(std::move(thd));
}
// Make sure all threads are finished
for (auto & Thread: ThreadPool) {
Thread.join();
}
ThreadPool.clear();
LogMan::Msg::IFmt("\nAll Done: {}", counter.load());
}
}
+7
View File
@@ -0,0 +1,7 @@
#pragma once
#include "ELFCodeLoader2.h"
namespace FEX::AOT {
void AOTGenSection(FEXCore::Context::Context *CTX, ELFCodeLoader2::LoadedSection &Section);
}
+4 -1
View File
@@ -2,7 +2,10 @@ add_subdirectory(LinuxSyscalls)
list(APPEND LIBS FEXCore Common CommonCore)
add_executable(FEXLoader FEXLoader.cpp)
add_executable(FEXLoader
FEXLoader.cpp
AOT/AOTGenerator.cpp)
target_include_directories(FEXLoader
PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/Source/
+5 -5
View File
@@ -19,6 +19,7 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <elf.h>
#include <fcntl.h>
@@ -84,8 +85,7 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
return false;
} else {
auto Filename = get_fdpath(file.fd);
Sections = std::make_unique<std::vector<LoadedSection>>();
Sections->push_back({Base, (uintptr_t)rv, size, (off_t)off, Filename, (prot & PROT_EXEC) != 0});
Sections.push_back({Base, (uintptr_t)rv, size, (off_t)off, Filename, (prot & PROT_EXEC) != 0});
return true;
}
@@ -117,7 +117,7 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
if (Elf.ehdr.e_type == ET_DYN) {
// needs base address
auto TotalSize = CalculateTotalElfSize(Elf.phdrs) + (BrkBase ? BRK_SIZE : 0);
auto TotalSize = CalculateTotalElfSize(Elf.phdrs) + (BrkBase ? BRK_SIZE : 0);
LoadBase = (uintptr_t)Mapper(0, TotalSize, PROT_NONE, MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);
if ((void*)LoadBase == MAP_FAILED) {
return {};
@@ -211,7 +211,7 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
bool Executable;
};
std::unique_ptr<std::vector<LoadedSection>> Sections;
std::vector<LoadedSection> Sections;
ELFCodeLoader2(std::string const &Filename, std::string const &RootFS, [[maybe_unused]] std::vector<std::string> const &args, std::vector<std::string> const &ParsedArgs, char **const envp = nullptr, FEXCore::Config::Value<std::string> *AdditionalEnvp = nullptr) :
Args {args} {
@@ -298,7 +298,7 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
}
void FreeSections() {
Sections.reset();
Sections.clear();
}
virtual uint64_t StackSize() const override { return STACK_SIZE; }
+25 -176
View File
@@ -5,6 +5,7 @@ desc: Glues the ELF loader, FEXCore and LinuxSyscalls to launch an elf under fex
$end_info$
*/
#include "AOT/AOTGenerator.h"
#include "Common/ArgumentLoader.h"
#include "Common/RootFSSetup.h"
#include "Common/SocketLogging.h"
@@ -38,6 +39,7 @@ $end_info$
#include <mutex>
#include <queue>
#include <set>
#include <sstream>
#include <string>
#include <sys/auxv.h>
#include <sys/resource.h>
@@ -215,154 +217,6 @@ bool IsInterpreterInstalled() {
std::filesystem::exists("/proc/sys/fs/binfmt_misc/FEX-x86_64", ec));
}
void AOTGenSection(FEXCore::Context::Context *CTX, ELFCodeLoader2::LoadedSection &Section) {
// Make sure this section is executable and big enough
if (!Section.Executable || Section.Size < 16)
return;
std::set<uintptr_t> InitialBranchTargets;
// Load the ELF again with symbol parsing this time
ELFLoader::ELFContainer container{Section.Filename, "", false};
// Add symbols to the branch targets list
container.AddSymbols([&](ELFLoader::ELFSymbol* sym) {
auto Destination = sym->Address + Section.ElfBase;
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) ) {
return; // outside of current section, unlikely to be real code
}
InitialBranchTargets.insert(Destination);
});
LogMan::Msg::IFmt("Symbol seed: {}", InitialBranchTargets.size());
// Add unwind entries to the branch target list
container.AddUnwindEntries([&](uintptr_t Entry) {
auto Destination = Entry + Section.ElfBase;
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) ) {
return; // outside of current section, unlikely to be real code
}
InitialBranchTargets.insert(Destination);
});
LogMan::Msg::IFmt("Symbol + Unwind seed: {}", InitialBranchTargets.size());
// Scan the executable section and try to find function entries
for (size_t Offset = 0; Offset < (Section.Size - 16); Offset++) {
uint8_t *pCode = (uint8_t *)(Section.Base + Offset);
// Possible CALL <disp32>
if (*pCode == 0xE8) {
uintptr_t Destination = (int)(pCode[1] | (pCode[2] << 8) | (pCode[3] << 16) | (pCode[4] << 24));
Destination += (uintptr_t)pCode + 5;
auto DestinationPtr = (uint8_t*)Destination;
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) )
continue; // outside of current section, unlikely to be real code
if (DestinationPtr[0] == 0 && DestinationPtr[1] == 0)
continue; // add al, [rax], unlikely to be real code
InitialBranchTargets.insert(Destination);
}
// endbr64 marker marks an indirect branch destination
if (pCode[0] == 0xf3 && pCode[1] == 0x0f && pCode[2] == 0x1e && pCode[3] == 0xfa) {
InitialBranchTargets.insert((uintptr_t)pCode);
}
}
uint64_t SectionMaxAddress = Section.Base + Section.Size;
std::set<uint64_t> Compiled;
std::atomic<int> counter = 0;
std::queue<uint64_t> BranchTargets;
// Setup BranchTargets, Compiled sets from InitiaBranchTargets
Compiled.insert(InitialBranchTargets.begin(), InitialBranchTargets.end());
for (auto BranchTarget: InitialBranchTargets) {
BranchTargets.push(BranchTarget);
}
InitialBranchTargets.clear();
std::mutex QueueMutex;
std::vector<std::thread> ThreadPool;
for (int i = 0; i < get_nprocs_conf(); i++) {
std::thread thd([&BranchTargets, CTX, &counter, &Compiled, &Section, &QueueMutex, SectionMaxAddress]() {
// Set the priority of the thread so it doesn't overwhelm the system when running in the background
setpriority(PRIO_PROCESS, ::gettid(), 19);
// Setup thread - Each compilation thread uses its own backing FEX thread
FEXCore::Core::CPUState state;
auto Thread = FEXCore::Context::CreateThread(CTX, &state, gettid());
std::set<uint64_t> ExternalBranchesLocal;
FEXCore::Context::ConfigureAOTGen(Thread, &ExternalBranchesLocal, SectionMaxAddress);
for (;;) {
uint64_t BranchTarget;
// Get a entrypoint to process from the queue
QueueMutex.lock();
if (BranchTargets.empty()) {
QueueMutex.unlock();
break; // no entrypoint to process - exit
}
BranchTarget = BranchTargets.front();
BranchTargets.pop();
QueueMutex.unlock();
// Compile entrypoint
counter++;
FEXCore::Context::CompileRIP(Thread, BranchTarget);
// Are there more branches?
if (ExternalBranchesLocal.size() > 0) {
// Add them to the "to process" list
QueueMutex.lock();
for(auto Destination: ExternalBranchesLocal) {
if (! (Destination >= Section.Base && Destination <= (Section.Base + Section.Size)) )
continue;
if (Compiled.contains(Destination))
continue;
Compiled.insert(Destination);
BranchTargets.push(Destination);
}
QueueMutex.unlock();
ExternalBranchesLocal.clear();
}
}
// All entryproints processed, cleanup this thread
FEXCore::Context::DestroyThread(CTX, Thread);
});
// Add to the thread pool
ThreadPool.push_back(std::move(thd));
}
// Make sure all threads are finished
for (auto & Thread: ThreadPool) {
Thread.join();
}
ThreadPool.clear();
LogMan::Msg::IFmt("\nAll Done: {}", counter.load());
}
int main(int argc, char **argv, char **const envp) {
const bool IsInterpreter = RanAsInterpreter(argv[0]);
@@ -381,33 +235,19 @@ int main(int argc, char **argv, char **const envp) {
}
#endif
FEXCore::Config::Initialize();
FEXCore::Config::AddLayer(FEXCore::Config::CreateMainLayer());
auto Program = FEX::Config::LoadConfig(
IsInterpreter,
true,
argc, argv, envp
);
if (IsInterpreter) {
FEX::ArgLoader::LoadWithoutArguments(argc, argv);
}
else {
FEXCore::Config::AddLayer(std::make_unique<FEX::ArgLoader::ArgLoader>(argc, argv));
}
FEXCore::Config::AddLayer(FEXCore::Config::CreateEnvironmentLayer(envp));
FEXCore::Config::Load();
auto Args = FEX::ArgLoader::Get();
auto ParsedArgs = FEX::ArgLoader::GetParsedArgs();
if (Args.empty()) {
if (Program.empty()) {
// Early exit if we weren't passed an argument
return 0;
}
std::string Program = Args[0];
// These layers load on initialization
auto ProgramName = std::filesystem::path(Program).filename();
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, true));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, false));
auto Args = FEX::ArgLoader::Get();
auto ParsedArgs = FEX::ArgLoader::GetParsedArgs();
// Reload the meta layer
FEXCore::Config::ReloadMetaLayer();
@@ -514,6 +354,7 @@ int main(int argc, char **argv, char **const envp) {
if (LDPath().empty() ||
std::filesystem::exists(LDPath(), ec) == false) {
fmt::print(stderr, "RootFS path doesn't exist. This is required on AArch64 hosts\n");
fmt::print(stderr, "Use FEXRootFSFetcher to download a RootFS\n");
}
#endif
return -ENOEXEC;
@@ -629,8 +470,8 @@ int main(int argc, char **argv, char **const envp) {
return open(filepath.c_str(), O_RDONLY);
});
FEXCore::Context::SetAOTIRWriter(CTX, [](const std::string& fileid) -> std::unique_ptr<std::ostream> {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir");
FEXCore::Context::SetAOTIRWriter(CTX, [](const std::string& fileid) -> std::unique_ptr<std::ofstream> {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir.tmp");
auto AOTWrite = std::make_unique<std::ofstream>(filepath, std::ios::out | std::ios::binary);
if (*AOTWrite) {
std::filesystem::resize_file(filepath, 0);
@@ -642,13 +483,21 @@ int main(int argc, char **argv, char **const envp) {
return AOTWrite;
});
for(const auto &Section: *Loader.Sections) {
FEXCore::Context::SetAOTIRRenamer(CTX, [](const std::string& fileid) -> void {
auto TmpFilepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir.tmp");
auto NewFilepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir");
// Rename the temporary file to atomically update the file
std::filesystem::rename(TmpFilepath, NewFilepath);
});
for(const auto &Section: Loader.Sections) {
FEXCore::Context::AddNamedRegion(CTX, Section.Base, Section.Size, Section.Offs, Section.Filename);
}
if (AOTIRGenerate()) {
for(auto &Section: *Loader.Sections) {
AOTGenSection(CTX, Section);
for(auto &Section: Loader.Sections) {
FEX::AOT::AOTGenSection(CTX, Section);
}
} else {
FEXCore::Context::RunUntilExit(CTX);
@@ -691,7 +540,7 @@ int main(int argc, char **argv, char **const envp) {
FEXCore::Allocator::ClearHooks();
FEXCore::Allocator::ReclaimMemoryRegion(Base48Bit);
// Allocator is now original system allocator
auto ProgramName = std::filesystem::path(Program).filename();
FEXCore::Telemetry::Shutdown(ProgramName);
if (ShutdownReason == FEXCore::Context::ExitReason::EXIT_SHUTDOWN) {
return ProgramStatus;
+2 -1
View File
@@ -20,6 +20,7 @@
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <fmt/format.h>
@@ -234,7 +235,7 @@ namespace FEX::HarnessHelper {
return;
}
fmt::format("{}: 0x{:016x} {} 0x{:016x} (Expected)\n", Name, A, A==B ? "==" : "!=", B);
fmt::print("{}: 0x{:016x} {} 0x{:016x} (Expected)\n", Name, A, A==B ? "==" : "!=", B);
};
const auto CheckGPRs = [&Matches, DumpGPRs](const std::string& Name, uint64_t A, uint64_t B) {
+31 -11
View File
@@ -11,6 +11,7 @@ add_library(LinuxEmulation STATIC
x32/FD.cpp
x32/FS.cpp
x32/Info.cpp
x32/IO.cpp
x32/Memory.cpp
x32/Msg.cpp
x32/NotImplemented.cpp
@@ -18,6 +19,7 @@ add_library(LinuxEmulation STATIC
x32/Sched.cpp
x32/Signals.cpp
x32/Socket.cpp
x32/Stubs.cpp
x32/Thread.cpp
x32/Time.cpp
x32/Timer.cpp
@@ -59,20 +61,38 @@ add_library(LinuxEmulation STATIC
Syscalls/Stubs.cpp
)
target_link_libraries(LinuxEmulation FEXCore FEX_Utils)
target_include_directories(LinuxEmulation PRIVATE ${CMAKE_BINARY_DIR}/generated)
target_include_directories(LinuxEmulation PRIVATE ${PROJECT_SOURCE_DIR}/External/drm-headers/include/)
target_compile_options(LinuxEmulation
PRIVATE
-Wall
-Werror=cast-qual
-Werror=ignored-qualifiers
-Werror=implicit-fallthrough
-Wno-trigraphs
)
target_include_directories(LinuxEmulation
PRIVATE
${CMAKE_BINARY_DIR}/generated
${PROJECT_SOURCE_DIR}/External/drm-headers/include/
)
target_link_libraries(LinuxEmulation
PRIVATE
FEXCore
FEX_Utils
)
set(HEADERS_TO_VERIFY
x32/Types.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/asound.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/drm.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/streams.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/usbdev.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/input.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/sockios.h x86_32 # This needs to match structs to 32bit structs
x32/Types.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/asound.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/drm.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/streams.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/usbdev.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/input.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/sockios.h x86_32 # This needs to match structs to 32bit structs
x32/Ioctl/joystick.h x86_32 # This needs to match structs to 32bit structs
x64/Types.h x86_64 # This needs to match structs to 64bit structs
x64/Types.h x86_64 # This needs to match structs to 64bit structs
)
list(LENGTH HEADERS_TO_VERIFY ARG_COUNT)
@@ -20,6 +20,7 @@ $end_info$
#include <fcntl.h>
#include <filesystem>
#include <ostream>
#include <sstream>
#include <stdio.h>
#include <system_error>
#include <unistd.h>
@@ -57,9 +58,6 @@ namespace FEX::EmulatedFile {
auto res_10 = FEXCore::Context::RunCPUIDFunction(ctx, 0x10, 0);
auto res_8000_0001 = FEXCore::Context::RunCPUIDFunction(ctx, 0x8000'0001, 0);
auto res_8000_0002 = FEXCore::Context::RunCPUIDFunction(ctx, 0x8000'0002, 0);
auto res_8000_0003 = FEXCore::Context::RunCPUIDFunction(ctx, 0x8000'0003, 0);
auto res_8000_0004 = FEXCore::Context::RunCPUIDFunction(ctx, 0x8000'0004, 0);
auto res_8000_0007 = FEXCore::Context::RunCPUIDFunction(ctx, 0x8000'0007, 0);
auto res_8000_0008 = FEXCore::Context::RunCPUIDFunction(ctx, 0x8000'0008, 0);
auto res_8000_000a = FEXCore::Context::RunCPUIDFunction(ctx, 0x8000'000a, 0);
@@ -105,12 +103,6 @@ namespace FEX::EmulatedFile {
vendorid.cpuid = {res_0.eax, res_0.ebx, res_0.edx, res_0.ecx};
vendorid.null = 0;
ModelName modelname {};
modelname.cpuid_2 = res_8000_0002;
modelname.cpuid_3 = res_8000_0003;
modelname.cpuid_4 = res_8000_0004;
modelname.null = 0;
Info info {res_1};
uint32_t Family = info.FamilyID + (info.FamilyID == 0xF ? info.ExFamilyID : 0);
@@ -579,6 +571,15 @@ namespace FEX::EmulatedFile {
cpu_stream << "vendor_id\t: " << vendorid.Str << std::endl;
cpu_stream << "cpu family\t: " << Family << std::endl;
cpu_stream << "model\t\t: " << (info.Model + (info.FamilyID >= 6 ? (info.ExModelID << 4) : 0)) << std::endl;
ModelName modelname {};
auto res_8000_0002 = FEXCore::Context::RunCPUIDFunctionName(ctx, 0x8000'0002, 0, i);
auto res_8000_0003 = FEXCore::Context::RunCPUIDFunctionName(ctx, 0x8000'0003, 0, i);
auto res_8000_0004 = FEXCore::Context::RunCPUIDFunctionName(ctx, 0x8000'0004, 0, i);
modelname.cpuid_2 = res_8000_0002;
modelname.cpuid_3 = res_8000_0003;
modelname.cpuid_4 = res_8000_0004;
modelname.null = 0;
cpu_stream << "model name\t: " << modelname.Str << std::endl;
cpu_stream << "stepping\t: " << info.Stepping << std::endl;
cpu_stream << "microcode\t: 0x0" << std::endl;
@@ -11,9 +11,9 @@ $end_info$
#include "Tests/LinuxSyscalls/x64/Syscalls.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <algorithm>
#include <bits/statx-generic.h>
#include <errno.h>
#include <cstring>
#include <fcntl.h>
@@ -596,11 +596,11 @@ uint64_t FileManager::Statx(int dirfd, const char *pathname, int flags, uint32_t
auto Path = GetEmulatedPath(SelfPath);
if (!Path.empty()) {
uint64_t Result = ::statx(dirfd, Path.c_str(), flags, mask, statxbuf);
uint64_t Result = FHU::Syscalls::statx(dirfd, Path.c_str(), flags, mask, statxbuf);
if (Result != -1)
return Result;
}
return ::statx(dirfd, SelfPath, flags, mask, statxbuf);
return FHU::Syscalls::statx(dirfd, SelfPath, flags, mask, statxbuf);
}
uint64_t FileManager::Mknod(const char *pathname, mode_t mode, dev_t dev) {
Loaded 100 of 247 files, more files were not shown because too many files have changed in this diff. Show more