This is embarassing. We were throwing away performance by failing to
use the softfloat library's inline helpers.
Turns out we needed to define `INLINE` to something in order for them to
work.
Feels bad.
Somehow I had completely missed this and recent reciprocal tests have
exposed it as a problem. When AFP is supported but not RPRES then we
were hitting this code path.
We were failing to insert in to the destination correctly, which because
the reciprocal is calculated using fdiv using a synthesized constant,
this would just zero the remaining portion of the register.
When the shift size is exactly 16bytes, then it turns in to a move.
If the shift size is above 16-bytes then synthesize the zero register in
the OpcodeDispatcher, so the backend doesn't synthesize and not cache.
This was only wired up for 256-bit SVE and wasn't ever hit for 128-bit
SVE. Ensure it works with 128-bit SVE, so mulvl needs to know when
128-bit is used. Then wire it up for vmaskmovps/pd. This saves one
instruction per operation.
Fixes#3791.
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.
* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.
Fixes#4319.
If AFP.AH is supported then fmin/fmax behaves like the x86 min/max
instruction so we don't need to jump through any additional hoops.
Support this use case to save a few instructions when AFP is supported.
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.
OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).
Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
```diff
"sha1rnds4 xmm0, xmm1, 10b": {
- "ExpectedInstructionCount": 55,
+ "ExpectedInstructionCount": 10,
```
So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
Saw these while scanning around. Funnily it makes it look like libnss is
worse off because there are multiple instructions using the same table
lookup to swizzle. So one instruction turns in to two.
We don't have a way to choose one path or the other, so it's usually
better to go the route that the instruction in a vacuum is improved, so
on average it is also improved.
Only saves a handful of instructions, but still an improvement.
```
"sha1msg2 xmm0, xmm1": {
- "ExpectedInstructionCount": 11,
+ "ExpectedInstructionCount": 7,
```
This differs from the existing GPUVis backend in a number of ways:
* Tracy is optimized for minimal overhead and nanosecond-resolution profiling
* Tracy supports live tracing (in addition to capture-based operation)
* Tracy has a richer feature set and a more polished UI (notably, statistics and histograms are generated out-of-the-box)
* GPUVis supports tracing multiple processes, whereas Tracy is single-process only
To use this backend, one of the environment variables FEX_PROFILE_TARGET_NAME
or FEX_PROFILE_TARGET_PATH must be defined to select the application under
profile by name or by path suffix.
Additionally, FEX_PROFILE_WAIT_FOR_FORK=1 may be needed for games that fork on startup.
Based on #4291 and #4324. Ideally this gets merged at the same time so
we can have Mangohud be on version 2 before giving them an upstream
patch.
Performance-wise this change falls within noise of my x87 microbench.
This just lets us track the number of float fallbacks FEX does, letting
us detect things like x87 fallbacks and how frequent they are, so we can
detect if a game might be slow or stuttering because of these fallbacks.
This reduces our codegen size and removes a few umov instructions.
Performance falls within noise but this small change will allow us to do
more vector optimizations in C code in the future.
With ASIMD this can be decently faster. With my microbenchmark this
makes pcmpistri ~6% faster.
With #4324 this can be made even faster since the incoming data can stay
in vector registers; Removing some overhead of umov.
This is preparation work to allow passing the corestate to the x87 soft
float handlers directly for some profile stats.
Performance-wise, this change falls within noise because it basically
moves the GPR->Vector moves from the JIT in to C code, my microbench saw
the largest excursion of 5% but that's still within noise in the current
design of my bench.
A more tangible win from this change alone is less codegen on the JIT
side.
With the prior approach, backwards jumps into existing blocks would
explore the overlapping part rather than splitting the block, generating
needless code and wasting time decoding. Similarly, the current block
wouldn't be split when it is extended to overlap with a pending jump target.
Solve this by tracking blocks in a sorted vector and splitting existing blocks
on jumps when appropriate, in order to avoid any possibility of overlapping
blocks, which would break the lookup, misaligned and zero instruction blocks
are disallowed.