Repository navigation
[draft] Update llama.cpp to b11100 - #1061
Merged
Merged
Conversation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rebases the patch set onto b11100 (659 commits past b10441) and adapts
llamafile's own integration to what upstream moved.
Patches
- Reconcile the 12 patches that drifted. GGML_CALL follows upstream's
reorganisations: the Vulkan backend split into several translation units
(its callbacks are now declared in ggml-vulkan-common.h and its interface
structs live in ggml-vulkan-buffers.cpp), and graph_optimize gained a
ggml_backend_graph_optimize_params whose add_alloc_dep the GPU DSOs call
back into the host with — so that pointer, upstream's lambda behind it
(replaced with a named function, since a lambda cannot carry the
attribute) and the Metal fusion helpers that pass it on all need the
annotation too.
- Drop src_models_{dflash,eagle3}.cpp.patch: upstream now defines
build_arch_graph at the end of both files, which is what they did.
- server-models: upstream replaced the per-model stopping_thread with one
server_monitor thread; move the signal mask there, and give unload_all's
remaining untimed wait the wait_for(30s) treatment. Same for the three
untimed waits server-queue grew.
Build
- New upstream sources: ggml-cpu/iqp.cpp, llama-kv-cache-dsa-iswa,
llama-memory-hybrid-idx, 9 src/models, 2 tools/mtmd/models, common/json{,
-schema}.cpp and the 16 common/parsers chat-template parsers.
- vendor/hash: mtmd-helper now hashes bitmap IDs with hash_sha256_hex().
- Version headers: ggml.c and llama.cpp include ggml-version.h /
llama-version.h, which CMake generates. apply-patches.sh writes both from
the same .in templates into the source tree, where the make build, the GPU
build scripts and metal.c all find them; the -D flags they replace are gone.
- Web UI: upstream deleted tools/ui/embed.cpp for scripts/ui-assets.cmake.
Keep the tool as a llamafile file (upstream's b10441 copy) — it emits the
same interface the CMake templates do, and the cosmocc build has no CMake.
- mtmd_helper_bitmap_init_from_{file,buf} gained an opt argument.
GPU backends
- metal.c: upstream replaced ggml-metal.metal with 20 kernels/*.metal, so
extract those plus their headers and point GGML_METAL_PATH_RESOURCES at
the app dir — the backend flattens the includes itself now, which retires
our shader preprocessor. Adds ggml-metal-{fusion,tuning}.cpp.
- vulkan.sh/.bat: compile every backend .cpp, not just ggml-vulkan.cpp.
Verified on macOS arm64: clean round-trip (reset-repo, setup, clean build,
check) is green, llama-server serves the gzip web UI and the hashed bundles
in router mode, and the Metal dylib builds and runs from scratch. CUDA,
ROCm, Vulkan and Windows still need a run on real hardware.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
llama.cpp b11100 removed --mmap/--no-mmap/--mlock/-dio/-ndio in favour of -lm/--load-mode (#28334); passing them now aborts argument parsing, and two .args examples in creating_llamafiles.md did exactly that. Also corrects the patch README: the built-in tools flag is --tools, not --server-tools (it has been --tools since before b10441). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b11100 deleted upstream's standalone tools/ui/embed.cpp in favour of
scripts/ui-assets.cmake, which renders tools/ui/ui.{cpp,h}.in inside a CMake
build (#28445, to simplify cross-compilation). The cosmocc build has no CMake
step, and the bump initially kept upstream's b10441 embed.cpp as a llamafile
file to fill the gap.
That left us reimplementing the templates' output, which is the part that can
silently drift from what server-http.cpp expects. Replace it with ui-embed.sh,
a translation of that script's emit_files() in the same spirit as BUILD.mk
being a translation of llama.cpp's CMake build. It renders upstream's own
templates, so the generated interface tracks upstream and cannot drift; only
the substitutions (@ASSET_ARRAYS@, @ASSET_TABLE@, @N_ASSETS@, @USE_GZIP@ and
the #cmakedefine) live here. ETags become the SHA-256 upstream computes rather
than the FNV-1a of the old tool.
This also collapses the required-asset list from two copies to one: the
generator embeds whatever it is given, and fetch-ui-assets.sh keeps the single
check it needs anyway to judge a downloaded tarball. A zero-length asset now
drops the UI with a warning instead of aborting the build, matching that
script's fallback.
Two fixes found on the way:
- build_gzip_mirror now passes gzip -n. gzip embeds the *input's* mtime, so
identical assets with different mtimes (a locally built dist/, a repacked
tarball) produced different bytes and different ETags. Upstream pins
SOURCE_DATE_EPOCH=0 for the same reason. Verified: without -n a touch(1)
changes the output, with -n it does not.
- fetch-ui-assets.sh's "KEEP IN SYNC WITH UPSTREAM tools/ui/embed.cpp" banner
pointed at a file that is no longer upstream's; it now names
ui_validate_assets() in scripts/ui-assets.cmake.
Verified: clean round-trip (reset-repo, setup, clean build, check) green;
llama-server serves index.html and the hashed bundles gzip-encoded, with
If-None-Match revalidating 304; the no-asset stub compiles; the empty-asset
guard takes the UI-less path. Generating all 70 assets takes ~2s and only
re-runs when the assets, the templates or the script change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b11100 changed how the FlashAttention vector kernels are selected. fattn.cu now gates every K-V combination on `if constexpr (GGML_CUDA_FA_<K>_<V>)`, and upstream emits one macro per combination from ggml_cuda_fattn_vec_instances() in ggml/cmake/common.cmake. Picking template-instance files is no longer enough: an undefined macro is a hard error, not a skipped branch, so the CUDA build failed with 49 "identifier GGML_CUDA_FA_* is undefined" errors in fattn.cu. collect_gpu_sources() now derives both the instance files and the macros from one list of combinations, so they cannot disagree, and exports them as CUDA_FA_DEFINES for cuda.sh/rocm.sh to append. The default set matches upstream's (q4_0-q4_0, q8_0-q8_0, f16-f16, bf16-bf16) and --fa-all-quants still expands to all 49. The same block goes into the four .bat files. -DGGML_CUDA_FA_ALL_QUANTS is dropped: no source references it any more. Verified on an L40S: ggml-cuda.so builds clean (144 sources, 0 errors) and serves inference both from the CLI and the server -- the latter being what exercises the yield_to_queue GPU-inline patch. The Vulkan module built from these scripts loads and runs too. The .bat files are untested; no Windows box. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
llama-server on Vulkan dies on the *second* request at the default --parallel 4, taking the whole process with it: ggml-backend.cpp: GGML_ASSERT(tensor->data != NULL && "tensor not allocated") ggml_backend_alloc_ctx_tensors_from_buft() splits a context across buffers when a tensor exceeds the backend's max_size (1 GiB on Vulkan). If the tail of the context holds only views, the final alloc_tensor_range() is skipped and those views never get ggml_backend_view_init(), so the persistent KV stream views (layer.k_stream / v_stream) keep data == NULL. The server's state save/restore then reads one and asserts. Slot count decides the KV stream layout, hence the dependence on --parallel: 1 is fine, 4 is not. This is not ours: vanilla b11100 built with Vulkan reproduces it identically on an L40S (only the line number differs, by the offset our free_struct patch adds earlier in the file). Upstream has it as #29221, and PR #25584 is the fix, still in review. Carried here verbatim minus its test, because the failure hits default settings in server mode; drop it at the bump that first includes it. Worth noting for #29221, which concluded the bug was gemma4-specific on an AMD iGPU: it reproduces with dense Qwen3.5-9B on a discrete NVIDIA L40S, so it is neither model- nor vendor-specific. Verified: clean round-trip green; the Vulkan server serves 6/6 requests at the default slot count with the patch, on both vanilla and llamafile, where it previously died on the second. No change on CPU or CUDA, which never hit the split. Integration suite over four runs matches the unpatched baseline (1-2 pre-existing flaky failures, no segfaults), and the multimodal CLI path that segfaulted once during testing is clean 15/15 and did not recur. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BuildMetal() reuses ~/.llamafile/v/<version>/ggml-metal.dylib whenever it exists, before extracting sources or kernels/. With the version left at 0.10.6, a Mac that has run the 0.10.6 release would load that b10441 dylib against the b11100 host. Earlier llama.cpp bumps (#1021, #1043) bumped the version for the same reason. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
setenv() only changes Cosmopolitan's environ; libSystem inside the dlopened dylib never sees it, so the variable was never set for the backend. Kernels are found because [NSBundle bundleForClass:] resolves to the dylib's own directory, where BuildMetal() extracts kernels/. Correct the comment accordingly, and warn when the user's environment sets the variable, since it then overrides the extracted kernels. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
llama.cpp b11100 replaced --mmap/--no-mmap/--mlock/-dio/-ndio with --load-mode and now rejects them, so existing llamafiles whose .args use them would exit at startup. Translate them into one --load-mode (with a warning) placed at the last old flag, so the last flag still wins when mixed with --load-mode. The mode follows the flags' original meaning as independent settings: mmap on by default, --mlock adds locking to it (mmap+mlock), direct I/O takes precedence. Note this differs from b10441's deprecation shim, which mapped --mlock to plain mlock and so silently dropped mmap. Docs: --load-mode is the syntax to use, in .args too; a new "Model Loading Flags" section gives the old-to-new mapping. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
common_log::add() and add_json() (new in b11100) block on an untimed cv_full.wait() when the ring is full. Convert both to wait_for(30s) loops, as done for the other waits, to avoid the XNU untimed-futex expiry. README: move the eagle3/dflash note out of the Bug Fixes table, where it pushed the ggml_src_ggml.c.patch row out of the table. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Bump the llama.cpp submodule to b11100 and reconcile the patch set.
PR Type
Checklist
llama.cpp/,whisper.cpp/, orstable-diffusion.cpp/, I also updated the matching*.patches/files.Tests: