There were two bugs in here.
The first bug here is with with the CMPXCHG emulation.
1) If the *Expected* value did *NOT* match what was in memory
2) *AND* The *Desired* value matched the memory value
3) The CAS would incorrectly return success for this CMPXCHG
4) Thus setting ZF incorrectly
The second bug comes from atomic memory operations (Add, CLR, EOR, SET, SWAP).
This operation is a Load + <Op> + CAS
1) If the memory backing between the Load and CAS changes
2) The CAS would then fail
3) On Atomic memory operations this should then retry to ensure it completes successfully
4) We were not retrying on failure in this case
5) Thus something like `LOCK INC` would have never atomically incremented correctly
6) We can't use our ASM unit tests to test this, since it needs thread contention.
This makes it a little more straightforward to see what these return
values mean without needing to look at the implementation.
These tuples were also getting a little bit large.
Same behavior, but allows clang to elide pushing all of these values on
and off the stack, particularly given these are called quite frequently
throughout the opcode dispatcher.
Don't print how many instructions are installed in the tables.
This isn't useful anymore
Not installing signal 32 and 33 are something we don't support right now. Stop complaining in that case.
Stop printing when a thread is starting up and shutting down. If you want to see this then gdb shows it well.
Don't print clone flags unless we are hitting a case where we are printing another log message.
Only the frontends need to deal with ELF files specifically.
The backend doesn't need to be aware of them at all.
Since the ELF handling is the frontend's responsibility, move all the code to the frontend.
si_addr will still be incorrect. What matters more here is that SIGCHLD gets correct information.
The guest needs SIGCHLD ifnromation to be filled out correctly, otherwise TTY handoff hangs
with the child process stopped.
I saw a red herring that I thought the high cpu usage in steamwebhelper could come from signal handlers.
This turned out to not be the case, but now I've got this implemented.
Installs the few signal handlers that we need upfront but for everything that isn't a mandatory signal
we instead now wait until the guest also installs that signal handler.
This fixes#1107
Places them in RO where they can't be modified.
While we're in the area, we can use an alias to prevent duplicated array
types, and also add a static assert to ensure the arrays are always the
same size.
This allows us to avoid needing to bounds check several accesses in a
row that we know will always be successful.
Transparent huge pages is a feature that the linux kernel opportunistically uses.
Depending on kernel configuration this feature is either enabled always, or when you madvise the region.
To ensure we hit both cases, madvise the regions we allocate in the 64-bit VMA allocator always
Can reduce kernel bookkeeping memory usage for our abusive allocator
When running the cpuid application (http://www.etallen.com/cpuid.html) I noticed
that we were returning some garbage data here.
After initially implementing support for leaf functions, it still didn't resolve the issue.
So I had to fix those in the x86-64 JIT.
I then went through and solved more issues with the function results.
- We now return a more sane CPU family that is near the feature set we support
- APICID now understands how to fill out the data correctly depending on emulated core counts
- Disabled some CPU features that we don't actually support
- Found two more cache functions that we weren't populating
- Filled with generic data cache size data
- Only thing that matters is that we ensure that cacheline size is reported as 64bytes
- L1D: 32KB, L1I: 32KB, L2: 512KB, L3: 8MB claimed for caches
- Implemented Leafs for functions
- 7h - Only has leaf 0
- Dh - Extended CPU features support
- Another register that lets you claim support for x87, SSE, and AVX
- Leaf 1 & 2 has some additional data
- 4h & 8000'0001Dh - Extended cache properties
- Almost the same as each other. One reports slightly less data though
After FEX has forked, there aren't any other threads in the process but their stacks remain.
We need to have some book keeping in place to have the stack ranges available to clean up
after fork.
We now keep both live stacks and dead stacks in a dequeue and on fork we will walk both to
clean up all stack objects that aren't our current thread.
rsi is a SSA argument, so we need to make sure to move the leaf argument first.
the leaf argument was getting corrupted when moving to the ABI.
Will be necessary once CPUID supports leafs
Passing in the string directly through the format string input can
unintentionally cause the output string to be interpreted as a format
string.
We can specify a separate format string to ensure it always prints
without any potential mangling.
<iostream> injects a static constructor in translation units that
include it, even if its facilities aren't used.
We can make use of <sstream> to avoid needing to execute those on
startup.