PyArrowFileIO.__init__ stores a bound method in an lru_cache:
self.fs_by_scheme: Callable[[str, str | None], FileSystem] = lru_cache(self._initialize_fs)
The cache holds self._initialize_fs, which holds self, so every PyArrowFileIO is in a reference cycle. Reference counting never frees one; it's freed only when the cycle collector runs. Until then its cached pyarrow.fs.S3FileSystem stays open, along with its connection pool.
Reproduction
import gc, sys, weakref
from pyiceberg.io.pyarrow import PyArrowFileIO
gc.disable()
io = PyArrowFileIO({})
io.fs_by_scheme("file", None) # what every read or write does first
ref = weakref.ref(io)
del io
print(sys.version.split()[0], "| freed by refcounting:", ref() is None)
gc.collect()
print(sys.version.split()[0], "| freed after gc.collect():", ref() is None)
3.11.15 | freed by refcounting: False
3.11.15 | freed after gc.collect(): True
3.14.4 | freed by refcounting: False
3.14.4 | freed after gc.collect(): True
Run on pyiceberg 0.11.1 with pyarrow 25.0.1. The same line is on main (pyiceberg/io/pyarrow.py, line 403 at the time of writing).
Why it matters
Catalogs build a new FileIO on every load_table and every commit. SqlCatalog._convert_orm_to_iceberg actually builds two per load: one from load_file_io(...) to read the metadata file, and the table's own from _load_file_io(...). So a process that reloads a table often makes FileIOs, each with its own S3FileSystem and connection pool, faster than they're freed.
On Python ≤ 3.13 the collector frees them quickly enough to hide this. On Python 3.14 the cycle collector is incremental, so they live much longer.
We measured this in litelink, which reloads a table from a reader polling every few milliseconds, against a local rustfs endpoint:
|
3.13 |
3.14 |
| peak open S3 connections |
53 |
1,013 |
| median connection lifetime |
0.53 s |
7.9 s |
On 3.14 that exhausted the object store's 1,024 open-file limit. It began returning 500s and then stopped answering. Calling gc.collect() every 0.2 s brought the peak back down to 13. Full write-up: nhobin219/litelink#137.
Possible fixes
- Break the cycle. For example, keep the filesystem cache in a plain dict on the instance and look it up in
fs_by_scheme, or cache on a module-level function keyed by the FileIO's properties, scheme and netloc, rather than wrapping a bound method.
- Build one FileIO per load, not two.
SqlCatalog._convert_orm_to_iceberg could read the metadata file with the same FileIO it then gives the table.
We worked around it in litelink by sharing one FileIO across loads and commits, through a SqlCatalog subclass: nhobin219/litelink#138. A fix here would remove the need for that.
Investigation done with Claude, reviewed by the litelink maintainer.
PyArrowFileIO.__init__stores a bound method in anlru_cache:The cache holds
self._initialize_fs, which holdsself, so everyPyArrowFileIOis in a reference cycle. Reference counting never frees one; it's freed only when the cycle collector runs. Until then its cachedpyarrow.fs.S3FileSystemstays open, along with its connection pool.Reproduction
Run on pyiceberg 0.11.1 with pyarrow 25.0.1. The same line is on
main(pyiceberg/io/pyarrow.py, line 403 at the time of writing).Why it matters
Catalogs build a new FileIO on every
load_tableand every commit.SqlCatalog._convert_orm_to_icebergactually builds two per load: one fromload_file_io(...)to read the metadata file, and the table's own from_load_file_io(...). So a process that reloads a table often makes FileIOs, each with its ownS3FileSystemand connection pool, faster than they're freed.On Python ≤ 3.13 the collector frees them quickly enough to hide this. On Python 3.14 the cycle collector is incremental, so they live much longer.
We measured this in litelink, which reloads a table from a reader polling every few milliseconds, against a local rustfs endpoint:
On 3.14 that exhausted the object store's 1,024 open-file limit. It began returning 500s and then stopped answering. Calling
gc.collect()every 0.2 s brought the peak back down to 13. Full write-up: nhobin219/litelink#137.Possible fixes
fs_by_scheme, or cache on a module-level function keyed by the FileIO's properties, scheme and netloc, rather than wrapping a bound method.SqlCatalog._convert_orm_to_icebergcould read the metadata file with the same FileIO it then gives the table.We worked around it in litelink by sharing one FileIO across loads and commits, through a
SqlCatalogsubclass: nhobin219/litelink#138. A fix here would remove the need for that.Investigation done with Claude, reviewed by the litelink maintainer.