perf: stop fsyncing every ZIP entry before its staging rename

The per-entry flush made entry count the dominant cost: an 8000-file
fixture spent most of its wall time in fsync (the extraction phase ran
past 200 s with it, 14.5 s without, same machine). Nothing in the
design needs that durability: a crash mid-extract leaves the staging
tree, which the next run discards, and publish is a rename-only phase.
The RAR and 7z engines never flushed per entry; all three now share
the same policy.

Measured on the same fixture (8000 small files, 1.9 MiB packed):
- with per-entry fsync: >200 s, not finished
- without: 14.5 s, extract phase 14.47 s CPU (scan 0.01 s, publish ~0)
- official 7-Zip on this host: >400 s (real-time AV scans every file
  it creates; the PS5 has no such factor)

ZIP/RAR matrices green (108 + 27 checks).
This commit is contained in:
Songlx516 committed 2026-09-16 14:01:43 +08:00
1 parent fd48232b6c
commit 8766178aa1
1 file changed
+9 -2
+9 -2
View File
@@ -912,8 +912,15 @@ write_entry(void *zip, zipx_ctx_t *c, int root_fd, const char *name,
done:
if(!ret) {
if(fsync(fd) || close(fd)) {
ctx_fail(c, ZIPX_ERR_IO, name, "cannot flush file: %s", strerror(errno));
/* No fsync here, on purpose. The per-entry flush used to cost 20-30
minutes on a 95k-file archive (measured on PS5-class storage) and buys
nothing the design needs: a crash mid-extract leaves the staging tree,
which is discarded on the next run, and publish is a rename-only phase
(see publish_entry). The RAR and 7z engines never flushed per entry
either; all three now share the same "sync nothing, rename everything"
policy. */
if(close(fd)) {
ctx_fail(c, ZIPX_ERR_IO, name, "cannot close file: %s", strerror(errno));
ret = -1;
}
fd = -1;