In the event that the host doesn't support the requirements for running
AVX-enabled applications (SVE2 with at least 256-bit wide vectors) but
still wanted to run regular SSE-enabled applications, they would be
taking a performance hit due to a load/store pessimization (necessary in
order for 256-bit loads/stores to work)
However, we can add an alternate view into the xmm data that would allow
those hosts to use the previous optimization, while still supporting
AVX-capable hosts.
This allows us to have RCPC loadstore operations with a 9-bit signed
offset.
This gives us a small range of [-256,256) of immediate encoding range on
our TSO loadstore operations.
Updates the inline constant pass in ConstProp to support this range on
TSO IR ops if the host supports RCPC2.
Apple M1 supports this extension, didn't test with Cortex-X2/A710.
Creates a pool allocator for OpcodeDispatcher and IRCompaction that
shares memory allocations between threads in a pool and supports
reclaiming stale allocations from participating threads.
A thread will use a heuristic to keep its claimed memory allocation
around if it is allocating a lot of code. If it slows down then it will
start putting the memory allocation back in to the thread pool.
Additionally if the allocation has been "disowned" and gone to sleep
while still retaining the allocation, then another thread can inspect
these stale allocations and reclaim it from the idling thread. Saving
further memory.
This needs some more work and cleanup but this is an interesting concept
that saves a decent amount of memory even in a basic test.
Causes teeworlds' title screen to go from 754MB to 599MB in my simple
test. 79.4% the memory usage is a good start.
This is a validation pass that attempts to prove the output from the RA pass is valid.
The current design should be able to prove that the RA result is internally
consistent. That no-matter what control flow path you take thought the control
flow graph, the physical registers and spill slots will always contain a single
possible SSA value.
It also checks that the SSA values in the IR actually line up with the SSA value
in the physical register.
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
If the long divides don't have anything in the upper bits then they can be optimized away.
Teeworlds, SuperTuxKart, and RRootage were between 50%-55% getting removed
Configuration mapping was duplicated between three different tables.
Additionally default configuration values were strewn about. Making it confusing as to what the default value would end up being
Adds a new ConfigValues.inl header that defines a few things right next to each other.
Defines the enum name as usual.
Defines the JSON config option name.
Defines the Environment config option name
Defines the default value that the configuration should be
RA Pass and passmanager compaction pass don't need to be independent.
This saves some memory.
In the future we should allow passes to fetch passes by name, that will
be part of the PassManager improvement in the future
Based on @phire's idea, Inline constants are constants that can be used as imms in instructions.
This
- Adds OP_INLINECONSTANT to represent them in the IR
- Modifies the ConstProp pass to transform OP_CONSTANT to OP_INLINECONSTANT based on per-platform heuristics
- Modifies the backends to use imms
- Supported opcodes: Add, Sub, Lsl, Lsr, Asr, Ror, Rol
Mutlblock pass that eliminates flag stores that are never read.
This is a sibling to RCLSE and will likely be absorbed by it when we add MB support there.
The Passmanager passes need very little or zero x86 knowledge to do
their work. They don't need the x86 OpDispatcher handling code at all.
Convert this entirely over to using the IREmitter class so it can
actually optimize IR code from the IRLoader frontend