RCLSE: Optimize redundant store->load operations

The bug that was causing crashes with this was due to inline syscalls.
Now that this is fixed we can re-enable store->load operations.

This allows constant propagation to work significantly better, which
means inline syscalls start working again. This can significantly
improve syscall performance in some cases.

This is most likely to improve performance in dxsetup and vc_redist but
hard to get a real profile.

Additionally this will let us inline cpuid results in the future which
is pretty nice.
This commit is contained in:
Ryan Houdek committed 2023-09-23 06:06:18 -07:00
1 parent 4e9a114858
commit d01b457727
1 file changed
+5 -5
@@ -506,17 +506,17 @@ bool RCLSE::ClassifyContextLoad(FEXCore::IR::IREmitter *IREmit, ContextInfo *Loc
ContextMemberInfo PreviousMemberInfoCopy = *Info;
RecordAccess(Info, Class, Offset, Size, LastAccessType::READ, CodeNode);
if (IsReadAccess(PreviousMemberInfoCopy.Accessed) &&
IsReadAccess(Info->Accessed) &&
PreviousMemberInfoCopy.AccessRegClass == Info->AccessRegClass &&
if (PreviousMemberInfoCopy.AccessRegClass == Info->AccessRegClass &&
PreviousMemberInfoCopy.AccessOffset == Info->AccessOffset &&
PreviousMemberInfoCopy.AccessSize == Size) {
// Optimize the case of redundant reads of the same exact value.
// This optimizes two cases:
// - Previous access was a load, and we have a redundant load of the same value.
// - Previous access was a store, and we are redundantly loading immediately after the store. Eliminating the store.
IREmit->ReplaceAllUsesWithRange(CodeNode, PreviousMemberInfoCopy.ValueNode, IREmit->GetIterator(IREmit->WrapNode(CodeNode)), BlockEnd);
RecordAccess(Info, Class, Offset, Size, LastAccessType::READ, PreviousMemberInfoCopy.ValueNode);
return true;
}
// TODO: Optimize the case of Store->Load.
// TODO: Optimize the case of partial loads.
return false;
}