Skip to content

Faster runs with a warm result cache - #6627

Merged
ondrejmirtes merged 10 commits into
2.3.xfrom
result-cache-perf
Sep 30, 2026
Merged

ondrejmirtes merged 10 commits into
2.3.xfrom
result-cache-perf

Conversation

@ondrejmirtes

@ondrejmirtes ondrejmirtes commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Runs with a warm result cache spent most of their time on fixed costs that did not depend on how many files changed. Each commit takes one of them away. The first commit fixes a bug in the turbo arena that also affects released versions.

Measured on Drupal core (11,194 analysed files, 28k directories), PHP 8.5, Apple M1 Pro, turbo extension on, fastest of 3 runs. "Before" is 2.3.x at 7c94037, "after" is an earlier version of this PR:

Files reanalysed Before After
0 (nothing changed) 3.41 s 1.58 s
10 6.88 s 2.84 s
1,124 (10%) 10.23 s 6.21 s
2,771 (25%) 13.92 s 9.89 s
5,574 (50%) 20.96 s 16.10 s
7,204 (64%) 23.34 s 19.03 s
8,956 (80%) 26.76 s 21.14 s
11,194 (result cache cleared) 24.12 s 22.08 s
11,185 (every file edited) 50.50 s 27.40 s

The "before" column up to "result cache cleared" is from an earlier session; the machine ran about 5% slower during the "after" session (the unchanged 2.3.x build too), so those gains are if anything understated. The "every file edited" row was measured in one session.

The rework after review (one stat-signature service, one way of caching directory symbols, a batch object, ArenaCache::replace()) does not change these numbers. Before vs. after the rework, medians of 12 interleaved runs each, separate tmpDirs:

Drupal core Before rework After rework
nothing changed 1.568 s 1.567 s
10 files edited (one worker) 2.917 s 2.907 s
250 files edited (forked workers) 4.175 s 4.154 s

WordPress core (full analysis, result cache cleared): 20.0 s -> 19.5 s without turbo, 8.2 s -> 7.9 s with turbo (a small tree - its directory walk and symbol scan were cheap to begin with).

Commits

  1. Let a saved cache entry replace the one in the turbo arena - a bug fix. Cache::load() publishes every entry it reads from disk to the arena, and the first publish of a key won. FileTypeMapper checks its name scope maps after loading them, against the hashes of the files they were created from, so after a file changed the stale map got published, was rejected, and the fresh map could never replace it. Every later lookup in every worker recreated the map: with every file of Drupal core edited, the workers created 44k maps for 11k files. New ArenaCache::replace() in the extension, used by Cache::save(): a load publishes only when the key has nothing yet, a save always wins. Callers checking their entries after loading need to do nothing for it.
  2. Keep directory listings between runs - finding the files walked all 28k directories every run, about a second, and the worker building the symbol index walked them again. A directory is read again only when its mtime, ctime or inode changed. Output identical to Symfony Finder, including order; Finder is still used on Windows and for stream wrappers.
  3. Reuse the file hashes of the result cache while the files are unchanged - sha256 of every analysed file (0.84 s) is skipped while size, mtime, ctime, inode and device match, like git's index.
  4. Store the dependency graph of the result cache with a table of paths - 430k dependency edges, each a path relativized/absolutized on every save/restore (0.5 s). Cache 194 MB -> 161 MB.
  5. Leave the exported nodes in the result cache file until they are needed - 135 MB of exported nodes were decoded on restore and serialized on save; now only the changed files are decoded and the rest is copied as bytes. Peak memory on restore 380 MB -> 126 MB.
  6. Do not parse a changed file when the outcome cannot add anything - when a changed file's dependents all changed too, comparing its exported nodes cannot add anything. With every file edited, restore spent 9 s parsing serially in the main process.
  7. Check stat signatures in one place - FileStatSignatures::begin() takes the time before anything is read, and the reader it returns gives a signature only when it can be trusted, null otherwise (Windows, stream wrappers, a failed stat, a file modified in the second the reading began). DirectoryWalker and ResultCacheManager just compare what they get.
  8. Keep the symbols found in directories between runs, checked by stat - OptimizedDirectorySourceLocatorFactory cached the symbols and checked the cache by hashing every file, except with turbo and forked workers, where every file was scanned natively on every run instead (0.7 s on Drupal) because hashing cost about as much. Now there is one way: the cache entry keeps each file's stat signature next to its hash, and the hash is only computed when the signature cannot vouch for the file. The arena records for the file hashes and symbol maps of workers that are not forked are gone. Files declaring no symbols are cached too; they used to be scanned on every run.
  9. Scan the directory locators of the lazy source locator in one batch - with a single worker, each directory was scanned on its own; now they are batched like PreForkDirectorySymbolScanner does. createBatch() returns an object with createByDirectory(), createByFiles() and scan(), replacing beginBatchedScan()/flushBatchedScan().
  10. Bump expected turbo version - for commit 1.

Commits 3 -> 4 -> 5 -> 6 touch the same code in ResultCacheManager and depend on each other textually. 4 and 5 each bump CACHE_VERSION; if both are kept, one bump is enough. Commit 7 builds on 2 and 3, commit 8 on 7, commit 9 on 8.

Safety

  • The stat-based caches (listings, file hashes, symbols) trust a signature of size, mtime, ctime, inode and device, checked in one place (commit 7). A file or directory modified in the second the reading began is not trusted, because timestamps have a one-second resolution. On Windows nothing is trusted, since ctime there is the creation time; the hashes decide as before.
  • Tried to break them with edits within one second: a fixture project edited right as a second begins, a run started on the cached tmpDir, a second edit at a random moment during that run, then the next cached run compared with a run on an empty tmpDir. The edits keep file sizes (in place), add, remove and rename files, set the mtime back to 2020, or replace a file by renaming a new one over it with the old mtime. 290 iterations with this PR, one worker and four forked workers: no stale result. The same harness against a build with the same-second check removed: 12 stale results in 55 iterations.
  • An older PHPStan reading the new cache file falls back to a full analysis (tested by downgrading and upgrading on Drupal). An empty dependencies section is still written because older versions absolutize it before their cacheVersion check.
  • Incremental results were compared with a full analysis on Drupal for every edit set above (10 files up to every file edited): identical errors, and the same number of reanalysed files as before these changes.
  • Workers spawned instead of forked with turbo on (Windows, or OPcache on) no longer share the symbol maps and directory file hashes through the arena; each builds its maps from the cache entry, which is still shared through the arena by Cache::load().

Tests

  • PHPUnit: all green. New tests for DirectoryWalker (Finder parity including dot files, VCS directories, symlinks, broken symlinks; added/removed/renamed files; replaced directories), FileStatSignatureReader, OptimizedDirectorySourceLocatorFactory (hash fallback, trusted signature, changed file, batches sharing a cache key), CachedExportedNodes, and ArenaCache::replace() in arena-smoke.php (including racing replaces).
  • PHPStan self-analysis and phpcs: clean.
  • Result cache E2E suite, locally on macOS: all pass except bug-14718 (GNU sed -i) and result-cache-ci-notification (needs CI env vars), both pass with those adjusted.
  • E2E suite: the 5 failures locally are environmental (PHP 8.5 vs composer.json constraints, timeout and sudo missing on macOS).

🤖 Generated with Claude Code

https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ

@ondrejmirtes
ondrejmirtes force-pushed the result-cache-perf branch 4 times, most recently from e441809 to 31b90c4 Compare September 30, 2026 11:20
@ondrejmirtes
ondrejmirtes marked this pull request as ready for review September 30, 2026 11:20
@phpstan-bot

Copy link
Copy Markdown
Collaborator

This pull request has been marked as ready for review.

ondrejmirtes and others added 8 commits September 30, 2026 13:30
Cache::load() publishes every entry it reads from disk to the arena, and
arena records were written once - the first publish of a key won. An entry
its caller checks after loading it could therefore get stuck there:
FileTypeMapper checks its name scope maps against the hashes of the files
they were created from, so after a file changed the stale map got
published, was rejected, and the fresh map created instead could never
replace it. Every later lookup in every worker got the stale map back and
created the map again. With every file of Drupal core edited, the workers
created 44k name scope maps for 11k files, parsing the file each time.

ArenaCache::replace() takes over the key's slot instead, and Cache::save()
uses it: a load publishes a copy of the file only when the key has nothing
yet, a save always wins. Callers checking their entries after loading them
need to do nothing for it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
Walking the analysed directories costs a readdir() per directory. On Drupal
core, 28k directories, that is about a second in the main process, and the
worker building the symbol index walks the same directories again.

A directory's listing only changes when an entry in it is added, removed or
renamed, and each of those updates its mtime and ctime. DirectoryWalker now
keeps the listings in tmpDir and reads a directory again only when its
stat changed. A directory modified in the second its listing is read is not
kept, because a second change in that same second would not show in its
stat. The walk yields the same files in the same order as Symfony Finder,
and falls back to Finder on Windows and for stream wrappers or unreadable
directories.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
Every run hashed every analysed file to find the changed ones, on Drupal
core almost a second for 11k files. The result cache now also records each
file's size, mtime, ctime, inode and device, and the recorded hash is
reused while they match - the same check git does against its index.

A signature is only recorded for a file last modified before the second its
hash was taken, and nothing is reused on Windows, where ctime is the
creation time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
Written out with the paths in full, the graph repeats each file in the
dependent list of everything it depends on. On Drupal core that is half a
million paths and 42 MB, and each of them was converted between absolute
and relative on every save and restore - about half a second.

The graph now lists every file once in a path table and refers to it by
position, so only the table is converted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
The exported nodes are most of the result cache - on Drupal core 135 MB of
160 MB. A run needs the decoded nodes of the few files that changed, and
the rest only has to get into the next cache file unchanged. Decoding all
of them on restore and serializing them again on save took about half a
second of every run that re-analysed anything.

The nodes are now stored as an index of files and lengths followed by the
serialized nodes back to back. restore() decodes only the ones it compares,
and save() copies the bytes of the others from the old file. The old file
is closed before the new one is renamed over it, for Windows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
restore() parses every changed file to compare its exported nodes with
the cached ones, which decides whether its dependents, the classes using a
trait it declares, and the files with errors are re-analysed too. When all
of those files changed themselves and are re-analysed anyway, the answer
cannot add anything, so the file is not parsed.

After a branch switch or a formatting run that touched most of the project
that is most of the changed files. On Drupal core with every file edited,
restoring took 9 seconds of parsing in the main process, and the workers
forked from it were slower too. The files re-analysed stay the same.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
DirectoryWalker and ResultCacheManager each built a stat signature, decided
on their own whether it could be trusted - not on Windows, not for a file
modified in the second the reading began - and had to take the time before
reading anything for that to hold.

FileStatSignatures::begin() takes the time, and the reader it returns gives
a signature only when it can be trusted, so the callers just compare and keep
what they get. The directory listings use the same signature as the files now.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
OptimizedDirectorySourceLocatorFactory had two ways of finding the symbols
of a directory. Without the turbo extension, or with workers that are not
forked, they were cached, and the cache was checked by hashing every file.
With the extension and forked workers, every file was scanned natively on
every run instead, because hashing cost about as much as the scan - 0.7s on
Drupal core, paid by every run that analyses at least one file.

There is one way now. The cache entry of a directory keeps each file's stat
signature next to its hash (see FileStatSignatures), and a file whose
signature matches is not even hashed. When the signature cannot vouch for
the file, the hash decides, as before. The arena records that shared the
file hashes and the symbol maps between workers that are not forked are
gone - the cache entry itself is still shared through the arena by
Cache::load().

The cache now also keeps the files that declare no symbols, which used to be
scanned again on every run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
ondrejmirtes and others added 2 commits September 30, 2026 13:52
Before forking, PreForkDirectorySymbolScanner collects every directory
locator and scans them in one go. With a single worker the locators are
created lazily in that worker instead, where each directory was scanned on
its own. The lazy initializer now batches them the same way.

beginBatchedScan() and flushBatchedScan() switched the factory into a
collecting mode that everything calling it in between took part in.
createBatch() returns an object instead: the locators added to it are
scanned by its scan(), and the repository and the Composer locator maker add
to it when they are given one.

A batch can hold two locators with the same cache key (odsl-installed-files
of two Composer projects), and taking the scan lock for the second one does
not wait for the lock the same process holds for the first.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdwtgiEHwkF5BTz73jgHQQ
@ondrejmirtes
ondrejmirtes merged commit ec0d4b6 into 2.3.x Sep 30, 2026
467 of 469 checks passed
@ondrejmirtes
ondrejmirtes deleted the result-cache-perf branch September 30, 2026 11:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants