System Info
accelerate: 1.14.0 (also verified present on `main`)
torch: 2.10.0
python: 3.11.11
platform: macOS-26.5.1-arm64 (also hit on linux/arm64, python 3.11.11, in Docker)
bitsandbytes: NOT installed
Information
Tasks
Reproduction
accelerate contains a circular import between accelerate.utils and accelerate.big_modeling:
accelerate/utils/__init__.py:226 from .bnb import has_4bit_bnb_layers, load_and_quantize_model
accelerate/utils/bnb.py:29 from ..big_modeling import dispatch_model, init_empty_weights
accelerate/big_modeling.py:25 from .hooks import (...)
accelerate/hooks.py:23 from .utils import (...) <-- back to accelerate.utils
Single-threaded this happens to resolve, but only by an ordering accident:
accelerator.py:35 does from accelerate.utils.dataclasses import FP8BackendType
before accelerator.py:37 imports big_modeling, so accelerate.utils is always
fully initialised first and the cycle never has to be broken.
That accident stops protecting you as soon as two threads import accelerate
submodules concurrently: they enter the cycle at different nodes, CPython's import
lock deadlock avoidance lets one of them proceed with a partially initialised
module, and the import fails.
This script fails reliably (10/10 runs here) on a clean install of accelerate 1.14.0.
No GPU, no bitsandbytes, and no other library required:
"""Minimal reproducer: concurrent imports of accelerate submodules fail."""
import importlib
import threading
barrier = threading.Barrier(2)
errors = []
def do_import(name):
barrier.wait()
try:
importlib.import_module(name)
except BaseException as exc:
errors.append(f"{name} -> {type(exc).__name__}: {exc}")
threads = [
threading.Thread(target=do_import, args=(name,))
for name in ("accelerate.utils", "accelerate.big_modeling")
]
for t in threads:
t.start()
for t in threads:
t.join()
print("\n".join(errors) if errors else "ok: both imports succeeded")
Traceback (frozen importlib frames removed for readability):
File "repro.py", line 6, in do_import
importlib.import_module(name)
File ".../importlib/__init__.py", line 126, in import_module
return _bootstrap._gcd_import(name[level:], package, level)
File "<site-packages>/accelerate/utils/__init__.py", line 226, in <module>
from .bnb import has_4bit_bnb_layers, load_and_quantize_model
File "<site-packages>/accelerate/utils/bnb.py", line 29, in <module>
from ..big_modeling import dispatch_model, init_empty_weights
ImportError: cannot import name 'dispatch_model' from partially initialized module
'accelerate.big_modeling' (most likely due to a circular import)
(<site-packages>/accelerate/big_modeling.py)
The two entry points above are just the smallest demonstration — in the wild this is
reached indirectly. Ours came in through transformers, where
transformers/generation/utils.py does from accelerate.hooks import ... while
another thread was entering accelerate.utils through a different transformers
submodule. The symptom in a multi-threaded server (one thread per session/request,
each importing a library that pulls in accelerate) is an intermittent import
failure on a freshly started process, which then works fine on retry — since by
then the modules are in sys.modules.
Depending on which thread wins the race, the same root cause also surfaces as:
ImportError: cannot import name 'AcceleratorState' from partially initialized module 'accelerate.state'
ImportError: cannot import name '_attach_context_parallel_hooks' from partially initialized module 'accelerate.big_modeling'
Expected behavior
Importing accelerate (or any of its submodules) from several threads at once
should not fail.
The cycle can be cut at its weakest link: utils/bnb.py only needs
dispatch_model / init_empty_weights inside two function bodies
(load_and_quantize_model and get_keys_to_not_convert), so the module-level
import is not required. Moving it into those functions makes the reproducer above
pass 6/6, with accelerate.utils.has_4bit_bnb_layers,
accelerate.utils.load_and_quantize_model and accelerate.dispatch_model all
still exported as before:
--- a/src/accelerate/utils/bnb.py
+++ b/src/accelerate/utils/bnb.py
@@ -26,7 +26,6 @@
is_8bit_bnb_available,
)
-from ..big_modeling import dispatch_model, init_empty_weights
from .dataclasses import BnbQuantizationConfig
from .modeling import (
find_tied_parameters,
@@ -85,6 +84,8 @@
`torch.nn.Module`: The quantized model
"""
+ from ..big_modeling import dispatch_model, init_empty_weights
+
load_in_4bit = bnb_quantization_config.load_in_4bit
load_in_8bit = bnb_quantization_config.load_in_8bit
@@ -387,6 +388,8 @@
Input model
"""
# Create a copy of the model
+ from ..big_modeling import init_empty_weights
+
with init_empty_weights():
tied_model = deepcopy(model) # this has 0 cost since it is done inside `init_empty_weights` context manager`
Happy to open a PR with that if it looks like the right shape to you.
One secondary observation while looking at this: utils/__init__.py:226 imports
.bnb unconditionally, so every user pays for the bitsandbytes code path (and is
exposed to this cycle) even without bitsandbytes installed — it is not installed
in the environment above. Guarding it with is_bnb_available(), the way the
deepspeed import a few lines earlier is guarded, would both narrow the blast
radius and skip needless work at import time.
System Info
Information
Tasks
no_trainerscript in theexamplesfolder of thetransformersrepo (such asrun_no_trainer_glue.py)Reproduction
acceleratecontains a circular import betweenaccelerate.utilsandaccelerate.big_modeling:Single-threaded this happens to resolve, but only by an ordering accident:
accelerator.py:35doesfrom accelerate.utils.dataclasses import FP8BackendTypebefore
accelerator.py:37importsbig_modeling, soaccelerate.utilsis alwaysfully initialised first and the cycle never has to be broken.
That accident stops protecting you as soon as two threads import
acceleratesubmodules concurrently: they enter the cycle at different nodes, CPython's import
lock deadlock avoidance lets one of them proceed with a partially initialised
module, and the import fails.
This script fails reliably (10/10 runs here) on a clean install of accelerate 1.14.0.
No GPU, no bitsandbytes, and no other library required:
Traceback (frozen importlib frames removed for readability):
The two entry points above are just the smallest demonstration — in the wild this is
reached indirectly. Ours came in through
transformers, wheretransformers/generation/utils.pydoesfrom accelerate.hooks import ...whileanother thread was entering
accelerate.utilsthrough a differenttransformerssubmodule. The symptom in a multi-threaded server (one thread per session/request,
each importing a library that pulls in accelerate) is an intermittent import
failure on a freshly started process, which then works fine on retry — since by
then the modules are in
sys.modules.Depending on which thread wins the race, the same root cause also surfaces as:
Expected behavior
Importing
accelerate(or any of its submodules) from several threads at onceshould not fail.
The cycle can be cut at its weakest link:
utils/bnb.pyonly needsdispatch_model/init_empty_weightsinside two function bodies(
load_and_quantize_modelandget_keys_to_not_convert), so the module-levelimport is not required. Moving it into those functions makes the reproducer above
pass 6/6, with
accelerate.utils.has_4bit_bnb_layers,accelerate.utils.load_and_quantize_modelandaccelerate.dispatch_modelallstill exported as before:
Happy to open a PR with that if it looks like the right shape to you.
One secondary observation while looking at this:
utils/__init__.py:226imports.bnbunconditionally, so every user pays for the bitsandbytes code path (and isexposed to this cycle) even without
bitsandbytesinstalled — it is not installedin the environment above. Guarding it with
is_bnb_available(), the way thedeepspeedimport a few lines earlier is guarded, would both narrow the blastradius and skip needless work at import time.