Repository navigation
Speed up pathlib.Path.glob() by removing redundant regex matching #115060
Description
Activity
- addedperformancePerformance or resource usagePerformance or resource usage
on Feb 6, 2024 Re-opening: there's another optimization possible.
Previous versions of pathlib used
path.exists()orpath.is_dir()to check whether non-wildcard segments exist, rather than callingpath._scandir()on the parent and filtering through a regex. This was dropped when we added the case_sensitive argument.If we can somehow check that the underlying filesystem is case-sensitive, and the user sets
case_sensitive=True(or leaves it at the default ofcase_sensitive=Noneon POSIX), then we can restore this more direct check for existence. And we can do better! For patterns like*/foo/bar.py, we can skip checking whetherfoois a directory, because we're going to checkfoo/bar.pyexists anyway. Likewise forfoo/*, where the_scandir()call failing makes anis_dir()check irrelevant.Alternatively, we could add a normalize_case argument to
glob(), defaulting toTrue. If set toFalse, and case_sensitive is effectivelyFalse, then literal segments in results wouldn't be case-normalized.Alternatively alternatively, we could enable this behaviour specifically when case_sensitive is set to
None(the default), and so users would need to setcase_sensitive=Falseto get fully case-normalized results.I need to think about this more.
Closed because I'm planning to make pathlib use the
globmodule for globbing, see #116392Re-opening because I'm a prize eejit.
- added 4 commits that reference this issue
on Apr 6, 2024 Sorry for the close/open spam, I keep spotting things to do.
Reacted by Erlend E. Aasland
In #104512 we made
pathlib.Path.glob()use a "walk-and-filter" strategy for expanding**wildcards in patterns: when we encounter a**segment, we immediately consume subsequent segments and use them to build a regex that is used to filter results. This saves a bunch ofscandir()calls.However! We actually build a regex for the entire pattern given to
glob(), rather than just the segments following**wildcards. And so when evaluating a pattern likedir*/**/file*, thedir*part is needlessly matched twice against each path. @zooba noted this in a review comment at the time.We should be able to improve performance by building an
re.Patternonly for segments following**wildcards, and not the entireglob()pattern.Linked PRs
pathlib.Path.glob()by removing redundant regex matching #115061pathlib.Path.glob()by skipping directory scanning #116152pathlib.Path.glob()by not scanning literal parts #117732pathlib.Path.glob()by omitting initialstat()#117831