Visitar URL original
zipfile.Path is not Path-like · Issue #99818 · python/cpython · GitHub
Skip to content

zipfile.Path is not Path-like #99818

Description

@buhtz

This is about zipfile.Path. The doc says it is compatible to pathlib.Path. But it seems that is not for 100% because it doesn't derive from pathlib.PurePath and can treated as Path-like in all situations.

I was redirected from pandas where I opened an issue about the fact that pandas.read_excel() does accept path-like objects but not zipfile.Path. It was explained to me that zipfile.Path doesn't implement __fspath__ and that is the problem.

Linked PRs

Activity

  1. ericvsmith commented on Nov 27, 2022

    @ericvsmith
    Member

    As explained in pandas-dev/pandas#49906 (comment), I don't think implementing __fspath__ would help your use case. Basically __fspath__ converts the object to a string. There isn't any string that would make sense to be used by any code (for example, the builtin open()) to operate on a file within a zip file.

  2. twoertwein commented on Nov 27, 2022

    @twoertwein

    It might help to hedge the documentation slightly. Maybe something like:

    zipfile.Path implements many methods of pathlib.Path but does not inherit from it nor does zipfile.Path implement os.PathLike.

  3. buhtz commented on Nov 27, 2022

    @buhtz
    Author

    The question is, from your viewpoint as Python core developers, how could the situation be improved.

    As an example csv.reader() can take pathlib.Path and zipfile.Path objects. Works fine without a workaround.

    Maybe zipfile.Path is not intended to act exactly like a pathlib.Path object? If this is the case then you should be more clear and direct about that limitations in the documentation.

    Or do you see a way that 3rd-party-libs like pandas can improve theire code to also accept zipfile.Path objects without explicit checking for that edge case via isinstance(obj, zipfile.Path)? I don't know why csv.reader() works in that case.

  4. AlexWaygood commented on Nov 27, 2022

    @AlexWaygood
    Member

    The doc says it is compatible to pathlib.Path.

    There's some subtlety (probably a little too much subtlety) in the docstring you link to there. The docstring says that zipfile.Path is "pathlib-compatible" (I.e., has a similar interface to pathlib.Path). It doesn't say that ZipFile is "path-like" (and nor do the docs), which is a term that's usually used to mean "an object that conforms to the os.PathLike interface by implementing the __fspath__ method". Somewhat confusingly, a class called zipfile.Path can be "similar to pathlib.Path" without being "path-like".

    Cc. @barneygale, for interest :)

  5. barneygale commented on Nov 27, 2022

    @barneygale
    Contributor

    FWIW, I'm working towards making zipfile.Path a subclass of PurePath (actually, a subclass of a new AbstractPath class that sits between PurePath and Path). It would allow libraries like pandas to work with zip paths like this (pseudocode):

    def read_excel(path):
        if not isinstance(path, pathlib.AbstractPath):
            path = pathlib.Path(path)
        with path.open('rb') as f:
            ...

    Discussion here: https://discuss.python.org/t/make-pathlib-extensible/3428.

  6. barneygale commented on Nov 27, 2022

    @barneygale
    Contributor

    The "pathlib-compatible" bit of the docstring should be changed to "similar to pathlib.Path" IMHO. Saying it's "compatible" is overselling it, particularly as zip paths aren't path-like!

  7. barneygale commented on Nov 27, 2022

    @barneygale
    Contributor

    Apologies for triple-posting, I've been thinking about this :)

    There isn't any string that would make sense to be used by any code (for example, the builtin open()) to operate on a file within a zip file.

    Why not the string representation of the archive member path? The os.PathLike interface definition doesn't say what you can do with the resulting string! Indeed, some functions like os.path.join() and os.path.dirname() call os.fspath() and then apply purely lexical operations on the result. pathlib.PurePath too follows this logic - you can happily open(PureWindowsPath('C:/blah')) from a non-Windows system.

    IMHO this points to an overloading of what "path-like" means, that can be remedied by making it more granular. We could introduce "pure" analogues of __fspath__(), os.fspath() and os.PathLike - something like __purepath__(), os.purepath(), os.PurePathLike.

  8. ericvsmith commented on Nov 27, 2022

    @ericvsmith
    Member

    My point is just that built-in open() is going to call os.fspath(zip_path). There's nothing that zip_path.__fspath__ could return that will make open() work. So maybe __fspath__ isn't what's needed to solve this problem.

  9. fancidev commented on Nov 29, 2022

    @fancidev
    Contributor

    It is probably more accurate to remove the phrase “pathlib-compatible” from the docs or replace it with something like “pathlib.Path-style” or “pathlib.Path-like”.

  10. buhtz commented on Nov 29, 2022

    @buhtz
    Author

    I really appreciate that you folks invest so much energy in that topic. Please let me throw another question into the pit.

    What was the intention about zipfile.Path. When it is not 100% Path-like than this couldn't be the intention to use a zipfile-path as each other path. But this was how I understood it in the first place.

    I describe me use-case here; reading data-files (csv, xlsx) from a path object no matter if it is a regular file (pathlib.Path) or an entry in a zip-archive (zipfile.Path). Are there other use cases you had in mind while designing that zipfile.Path?

  11. jaraco commented on Feb 5, 2023

    @jaraco
    Member

    When designing zipfile.Path, the main use case was for use in abstracting the results from importlib.metadata and importlib.resources when the packages were present in a zipfile. That is, provide access to paths and files of various resources of a package or a distribution's metadata as files in directories. In particular, for importlib.resources, the Traversable protocol was needed. See #88366 for some discussion around Traversable being PathLike and #96870 exploring why Traversable was created. See also python/importlib_resources#232, where @FFY00 and I are discussing the merits and challenges with honoring pathlike objects in importlib.resources.as_file.

    Since zipfile.Path is a Traversable, perhaps all you need to do is use as_file?

    from importlib.resources import as_file
    
    def read_excel(traversable):
        with as_file(traversable) as path:
            with path.open('rb') as f:
               ...

    Improvements I'd like to see:

    • Document the guarantees and limitations of zipfile.Path, perhaps by introducing Traversable.
    • Use type system to declare that zipfile.Path implements Traversable.
    • Possibly highlight as_file as a means to readily get a resource on disk.
  12. jaraco commented on Feb 5, 2023

    @jaraco
    Member

    FWIW, I'm working towards making zipfile.Path a subclass of PurePath (actually, a subclass of a new AbstractPath class that sits between PurePath and Path). It would allow libraries like pandas to work with zip paths like this (pseudocode):

    def read_excel(path):
        if not isinstance(path, pathlib.AbstractPath):
            path = pathlib.Path(path)
        with path.open('rb') as f:
            ...

    Wait - if you're using path.open(), why does it need to be cast to a pathlib.Path? zipfile.Path supports .open without any manipulation. In other words, Pandas should be adapted to accept a traversable, and then it will accept zipfile.Path objects and pathlib.Path objects.

  13. barneygale commented on Feb 5, 2023

    @barneygale
    Contributor

    Bad example on my part. If you're only interested in the interface Traversable provides then sure, you can do:

    def read_excel(path):
        if not isinstance(path, Traversable):
            path = pathlib.Path(path)
        with path.open('rb') as f:
            ...

    If you're looking for an interface more like pathlib's (so including stuff like parents, with_suffix(), exists(), write_text(), glob() and walk()) then pathlib.AbstractPath might be a good option if/when it arrives.

  14. FFY00 commented on Feb 5, 2023

    @FFY00
    Member

    I opened GH-101589 to fix the documentation.

    If the Traversable interface is enough, you can just use it, otherwise, as Jason mentioned, you should use as_file to get a pathlib.Path instance.

    And again, please note that path-like objects are a completely different thing 😅

  15. added a commit that references this issue on Feb 20, 2023
  16. added a commit that references this issue on Feb 25, 2023
  17. hauntsaninja commented on Feb 25, 2023

    @hauntsaninja
    Contributor

    I opened backport PRs #102266 and #102267 ; once merged I think we can close this

  18. added a commit that references this issue on Feb 25, 2023
  19. added 2 commits that reference this issue on Feb 25, 2023
  20. Stannislav commented on Jul 21, 2023

    @Stannislav

    Stumbled upon this discussion while dealing with a similar issue. So, maybe I can contribute a different angle on this.

    file = importlib.resources.files("my.module").joinpath("some-file") is the recommended way of loading package resources.

    Depending on the way the package is distributed, file can be of type pathlib.Path or zipfile.Path, which have non-compatible interfaces, as discussed in this thread. This inconsistency maybe be a stumbling block for those wishing to use importlib.

  21. added 2 commits that reference this issue on Sep 1, 2024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions