We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.
In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.
This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.
Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
Presumably this was done as a hack to highlight the line in red when rendering
documentation to markdown. Since it's only used for two instructions, drop this
use to ease generation of C++ docstrings.
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.
The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
Enables memcpy optimization of 80bit floats on reduced precision.
Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.
If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
Turns out a simpler way was added to the docs at some point and I never
noticed.
Before:
text data bss dec hex filename
4159895 1471360 4336824 9968079 9819cf Bin/FEX
After:
text data bss dec hex filename
4157159 1471360 4336824 9965343 980f1f Bin/FEX
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.
Before:
text data bss dec hex filename
4160559 1471360 4336824 9968743 981c67 Bin/FEX
After:
text data bss dec hex filename
4159927 1471360 4336824 9968111 9819ef Bin/FEX