Commit Graph
9 Commits
Author SHA1 Message Date
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Ryan Houdek bce694ebb5 FEXCore: Moves BitUtils to FHU
No functional change
2023-12-25 06:38:51 -08:00
Ryan Houdek 03f63f99a8 FEXCore: Moves StringUtils to FEXCore headers
Once gdbserver gets moved to the frontend this will need to be in the
includes.
2023-11-02 20:09:12 -07:00
Ryan Houdek fea72ce19c Merge pull request #3120 from Sonicadvance1/more_optimal_x87
FEXCore: Support preserve_all ABI for interpreter fallbacks
2023-09-21 15:35:37 -07:00
Ryan Houdek d588d41ab9 InterpreterFallbacks: Converts X87 and String ops to preserve_all
This improves performance!
2023-09-20 18:51:18 -07:00
Ryan Houdek e85b90c614 FEXCore/Common: Adds SPDX identifier 2023-09-19 17:33:14 -07:00
Ryan Houdek 745729cdc2 SoftFloat-3e: Adds preserve_all attribute to all functions used
This will let FEX's JIT be more optimal
2023-09-18 17:42:48 -07:00
Ryan Houdek 0c5c146fcf FEXCore/JitSymbols: Buffer writes to reduce overhead
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.

The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.

To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
  clang benchmark and seeing how frequently it was still writing.
   - 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
  symbols.
   - In most cases 100ms is enough that you won't notice the blip in
     perf.

One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
2023-09-16 17:52:46 -07:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00