mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-09 17:00:19 +02:00
Now the calculation of PF is entirely deferred, by inverting our internal representation of PF. All the (e.g.) logical op needs to do is store the low 8-bits of the result. This is a bit of a mixed bag. Primary ALU ops all save an instruction, by skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky cases are: * Zeroing PF. This now requires writing 1 instead of 0, which may require an extra move for the constant. Some of this will go away when we merge PF+AF into a single register, which is next up on the list. In that case, the two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing x86 view of PF), that will get absorbed into the or, if we also write AF. If we leave AF undefined and need to write a 1, that's a single mov instruction and we couldn't do better anyway if not inverted (since we'd still have a mov wzr even then). So in view of the future work, this isn't something I'm concerned about. * Float comparisons that put Unordered into PF. These require an extra invert to match the new convention. These are already so unnecessarily bloated that I'm not convinced I'm making things materially worse here. But we realistically need multiple destination support in the IR to fix this particular mess. Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>