Commit Graph
27 Commits
Author SHA1 Message Date
Simon Scherer fd64ea54fe InstcountCI: Update 2026-09-04 07:32:09 +02:00
Simon Scherer 30e4650a8b InstcountCI: Update 2026-09-03 08:32:59 +02:00
Ryan Houdek fbf5187794 InstcountCI: Update 2026-09-02 18:35:36 -07:00
Ryan Houdek 20e9e2044e InstcountCI: Update 2026-09-02 17:15:36 -07:00
Justin Becker 3471ca67c4 Bump DiskCache version number 2026-09-02 16:40:17 -07:00
Simon Scherer bce8b13366 InstcountCI: Update 2026-09-02 07:47:24 +02:00
Pierre-Loup A. Griffais c01bfcd52f InstcountCI: Update 2026-09-01 03:51:12 -07:00
Ryan Houdek 1558a2c65d InstcountCI: Update 2026-08-31 19:23:20 -07:00
Pierre-Loup A. Griffais 0d3a5f96f0 InstcountCI: Update 2026-08-30 17:17:54 -07:00
Ryan Houdek b695c5ae11 InstCountCI: Update
Only adds DiskCache version to json to ensure it's tracked. No asm
changes here.
2026-08-30 14:26:45 -07:00
LC d78963c021 VectorOps: Avoid dup if able in VInsElement 128-bit element path
We don't need to broadcast if we're inserting across registers into the
equivalent position, since we already have a predicate around that can
satisfy that.
2026-06-21 20:32:37 -04:00
LC b40f9db735 [SVE256] Remove heavy handed moves from scalar compares
(See #3799)

I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)

This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.
2026-06-21 02:27:53 -04:00
LC 99baa4f3d9 [SVE256] Handle SSE insertions for CMPSD/CMPSS
See #3799
2026-06-17 21:05:31 -04:00
Ryan Houdek db14975828 InstcountCI: Update 2025-12-04 01:45:32 -08:00
Ryan Houdek 11946ffc4c InstcountCI: Update 2025-10-24 11:11:54 -07:00
Billy Laws c45abaaa6b Update InstCountCI 2025-09-09 21:34:59 +01:00
Billy Laws 203ac51681 Update InstCountCI 2025-08-07 14:36:12 +01:00
Ryan Houdek b76f819759 InstcountCI: Update 2025-03-06 18:48:43 -08:00
Ryan Houdek 69cfc78ee1 InstcountCI: Update 2025-02-24 14:54:37 -08:00
Paulo Matos dc93e30451 Disable RPRES in instcounci files
This was giving false changes on RPRES enabled HW.
2024-10-23 15:09:17 +02:00
Paulo Matos 9acd325aa4 instcountci: Fixes AFP.NEP handling on scalar insertions 2024-06-19 10:02:54 +02:00
Alyssa Rosenzweig 32150cf7b5 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 12:01:23 -04:00
Alyssa Rosenzweig cea551c2ac InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig bd4464bd5e InstructionCountCI: Remove Optimal flags
Instruction count CI has transformed the way we work on FEX… I love the system
and want to make it better. there’s one part of instruction count CI that isn’t
so lovable: the problematic “optimal” flag on instructions.

There are several issues with this flag, both philosophical and practical.

– it is tedious to update the optimal flag when making an implementation
optimal. The effect of that is discouraging people from making instructions,
optimal, or encouraging people to fail to update the flag, and dilute the value
of it. Either way, since we care far more about optimal implementations, then we
do about updating the flag, clearly we should prioritize the implementation and
not the flag. This issue was not obvious at the outset, when instruction count,
CI was introduced, and still quite small. The problem magnified when we started
duplicating instructions in bulk for different combinations of CPU features
(flagm, AFP, etc.) that intern multiplies the manual work required to update the
flags by the corresponding constant factor. if it comes down to a choice between
removing this extra coverage and removing the flag, I think we all agree that
removing the flag is the lesser evil.

– The definition of “optimal” is fundamentally problematic. I have often
improved the instruction count of an instruction that was already “optimal”.
This is all kinds of silly, and calls into question whether there’s any value
whatsoever in the existing classifications of the flag. Furthermore, it is often
unknowable, whether an implementation really is optimal. Is it possible to
implement BZHI (with flag calculations) in fewer than eight instructions? We
don’t know, and it’s silly to pretend that we do.

– as a consequence of the problematic definitions , there are so many errors in
both directions that I don’t think there’s much value in preserving the existing
classification at the expense of +progress. Being able to say “32% of
instructions are translated optimally” is neat, but it really doesn’t tell us
anything whatsoever when you dig a little deeper.

So, as the flag is misleading at best and perhaps harmful at worst, let’s remove
it and make the instruction count CI, more useful overall. let’s let the
expected count and the assembly speak for themselves, and cut away the chaff. if
we want a meaningless number to report to management, we can instead calculate
the average blowup factor ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-13 21:14:05 -04:00
Lioncache 47a0f14537 OpcodeDispatcher: Remove unnecessary 128-bit truncating moves from StoreResult
Removes the truncating move that we perform inside the StoreResult
function and instead delegates the responsibility to the instruction
implementations themselves.

This removes a lot of redundant moves that occur on 128-bit variants
of AVX instructions.

Also fixes a weird case where we were handling 128-bit SVE
in VBroadcastFromMem when we already have AdvSIMD instructions
that will perfom the zero-extension behavior for us.
2023-10-17 11:07:04 +02:00
Ryan Houdek cc2eef619c InstCountCI: Update for AFP optimizations
A bunch of random instructions have converted to be optimal in a vacuum.
2023-10-10 03:44:59 -07:00
Ryan Houdek c548625fbe InstCountCI: Update tests for disabling AFP
Doesn't change behaviour yet, just prep work.
2023-10-10 03:17:18 -07:00