Visitar URL original
Speed up `pathlib.Path.glob()` by removing redundant regex matching · Issue #115060 · python/cpython · GitHub
Skip to content

Speed up pathlib.Path.glob() by removing redundant regex matching #115060

Description

@barneygale

In #104512 we made pathlib.Path.glob() use a "walk-and-filter" strategy for expanding ** wildcards in patterns: when we encounter a ** segment, we immediately consume subsequent segments and use them to build a regex that is used to filter results. This saves a bunch of scandir() calls.

However! We actually build a regex for the entire pattern given to glob(), rather than just the segments following ** wildcards. And so when evaluating a pattern like dir*/**/file*, the dir* part is needlessly matched twice against each path. @zooba noted this in a review comment at the time.

We should be able to improve performance by building an re.Pattern only for segments following ** wildcards, and not the entire glob() pattern.

Linked PRs

Activity

  1. added 2 commits that reference this issue on Feb 6, 2024
  2. added a commit that references this issue on Feb 14, 2024
  3. barneygale commented on Feb 26, 2024

    @barneygale
    ContributorAuthor

    Re-opening: there's another optimization possible.

    Previous versions of pathlib used path.exists() or path.is_dir() to check whether non-wildcard segments exist, rather than calling path._scandir() on the parent and filtering through a regex. This was dropped when we added the case_sensitive argument.

    If we can somehow check that the underlying filesystem is case-sensitive, and the user sets case_sensitive=True (or leaves it at the default of case_sensitive=None on POSIX), then we can restore this more direct check for existence. And we can do better! For patterns like */foo/bar.py, we can skip checking whether foo is a directory, because we're going to check foo/bar.py exists anyway. Likewise for foo/*, where the _scandir() call failing makes an is_dir() check irrelevant.

    Alternatively, we could add a normalize_case argument to glob(), defaulting to True. If set to False, and case_sensitive is effectively False, then literal segments in results wouldn't be case-normalized.

    Alternatively alternatively, we could enable this behaviour specifically when case_sensitive is set to None (the default), and so users would need to set case_sensitive=False to get fully case-normalized results.

    I need to think about this more.

  4. added a commit that references this issue on Feb 29, 2024
  5. barneygale commented on Apr 6, 2024

    @barneygale
    ContributorAuthor

    Closed because I'm planning to make pathlib use the glob module for globbing, see #116392

  6. barneygale commented on Apr 6, 2024

    @barneygale
    ContributorAuthor

    Re-opening because I'm a prize eejit.

  7. added 4 commits that reference this issue on Apr 6, 2024
  8. barneygale commented on Apr 13, 2024

    @barneygale
    ContributorAuthor

    Sorry for the close/open spam, I keep spotting things to do.

  9. added 2 commits that reference this issue on Apr 13, 2024
  10. added 3 commits that reference this issue on Apr 17, 2024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions