Summary
get_max_memory() already treats shared memory correctly for MPS, but the same
reasoning is not applied to integrated CUDA devices, so on those the returned
budget counts the same physical RAM twice.
Current code (src/accelerate/utils/modeling.py, around L822):
# allocate everything in the mps device as the RAM is shared
if is_mps_available():
max_memory["mps"] = psutil.virtual_memory().available
else:
max_memory["cpu"] = psutil.virtual_memory().available
The comment states the principle exactly: when RAM is shared, do not hand out a
second pool. An NVIDIA GB10 (DGX Spark) shares its RAM the same way, but it is
a CUDA device, so it takes the else branch and gets a separate cpu entry
alongside its device entry.
Measured on a GB10:
>>> import torch
>>> from accelerate.utils import get_max_memory
>>> torch.cuda.get_device_properties(0).is_integrated
1
>>> get_max_memory()
{0: 73446760448, 'cpu': 87132405760}
73.4 GB + 87.1 GB = 160.5 GB advertised, on a machine with:
$ grep MemTotal /proc/meminfo
MemTotal: 127535336 kB # 127.5 GB
There is no separate device VRAM on this part. torch.cuda.mem_get_info(0)
reports free=74.1 GB, total=130.6 GB, and those are the same bytes
/proc/meminfo describes, reached over NVLink-C2C rather than PCIe. The two
entries alias one pool: allocating against either shrinks the other, which the
returned dict cannot express.
Why it matters
Everything downstream inherits the number. infer_auto_device_map and
device_map="auto" treat the entries as independent capacities, so a model
sized between the device entry and the sum can be planned to "fit" across a
split whose halves are the same memory.
The hardware is not niche. GB10 / DGX Spark, GH200, GB200, Jetson Thor and Orin
are all integrated-memory parts, and they are specifically the machines people
buy to run large models locally.
Prior art in this file
Suggested fix
Happy to send the PR. Two shapes, and I would rather write the one you want
than guess:
A. Extend the existing shared-RAM rule to integrated CUDA devices. Where
the MPS check happens, also treat an integrated CUDA device as sharing host
RAM, so no independent cpu entry is advertised against it. Smallest change,
and it reuses the reasoning already in the file.
B. Leave get_max_memory() alone and fix the consumer. Teach
infer_auto_device_map that an integrated device and the host are one pool, so
it will not plan a split between two aliases of the same memory.
I lean towards A, since it corrects the value everything else derives from.
Two implementation notes either way:
is_integrated should be read defensively, e.g.
getattr(torch.cuda.get_device_properties(i), "is_integrated", False).
setup.py allows torch>=2.0.0 and I have only confirmed the attribute on
torch 2.10, so a bare access risks breaking older installs.
- This cannot be exercised on your CI, since it needs integrated-memory
hardware. A unit test would have to fake the device property (monkeypatch
get_device_properties) and assert the shape of the returned dict. Happy to
write it that way.
What I am not claiming
I found this while debugging a 4-bit load that failed under device_map="auto"
with "Some modules are dispatched on the CPU or the disk", which
device_map={"": 0} resolved. I could not reproduce that failure with plain
transformers on the same machine and model, so I am not asserting the double
count caused it. The arithmetic above stands on its own.
Environment
accelerate 1.12.0
transformers 5.5.0
torch 2.10.0a0+b558c986e8.nv25.11 (torch.version.cuda 13.0)
device NVIDIA GB10, capability (12, 1), is_integrated=1
driver 580.173.02, CUDA runtime 13.0
platform Linux-6.17.0-1029-nvidia-aarch64-with-glibc2.39, aarch64
python 3.12.3
Summary
get_max_memory()already treats shared memory correctly for MPS, but the samereasoning is not applied to integrated CUDA devices, so on those the returned
budget counts the same physical RAM twice.
Current code (
src/accelerate/utils/modeling.py, around L822):The comment states the principle exactly: when RAM is shared, do not hand out a
second pool. An NVIDIA GB10 (DGX Spark) shares its RAM the same way, but it is
a CUDA device, so it takes the
elsebranch and gets a separatecpuentryalongside its device entry.
Measured on a GB10:
73.4 GB + 87.1 GB = 160.5 GB advertised, on a machine with:
There is no separate device VRAM on this part.
torch.cuda.mem_get_info(0)reports
free=74.1 GB, total=130.6 GB, and those are the same bytes/proc/meminfodescribes, reached over NVLink-C2C rather than PCIe. The twoentries alias one pool: allocating against either shrinks the other, which the
returned dict cannot express.
Why it matters
Everything downstream inherits the number.
infer_auto_device_mapanddevice_map="auto"treat the entries as independent capacities, so a modelsized between the device entry and the sum can be planned to "fit" across a
split whose halves are the same memory.
The hardware is not niche. GB10 / DGX Spark, GH200, GB200, Jetson Thor and Orin
are all integrated-memory parts, and they are specifically the machines people
buy to run large models locally.
Prior art in this file
get_max_memory()returns allocated instead of total devicememory for XPU) is the same class of bug: a device type measured with the
wrong quantity.
Suggested fix
Happy to send the PR. Two shapes, and I would rather write the one you want
than guess:
A. Extend the existing shared-RAM rule to integrated CUDA devices. Where
the MPS check happens, also treat an integrated CUDA device as sharing host
RAM, so no independent
cpuentry is advertised against it. Smallest change,and it reuses the reasoning already in the file.
B. Leave
get_max_memory()alone and fix the consumer. Teachinfer_auto_device_mapthat an integrated device and the host are one pool, soit will not plan a split between two aliases of the same memory.
I lean towards A, since it corrects the value everything else derives from.
Two implementation notes either way:
is_integratedshould be read defensively, e.g.getattr(torch.cuda.get_device_properties(i), "is_integrated", False).setup.py allows
torch>=2.0.0and I have only confirmed the attribute ontorch 2.10, so a bare access risks breaking older installs.
hardware. A unit test would have to fake the device property (monkeypatch
get_device_properties) and assert the shape of the returned dict. Happy towrite it that way.
What I am not claiming
I found this while debugging a 4-bit load that failed under
device_map="auto"with "Some modules are dispatched on the CPU or the disk", which
device_map={"": 0}resolved. I could not reproduce that failure with plaintransformers on the same machine and model, so I am not asserting the double
count caused it. The arithmetic above stands on its own.
Environment