Skip to content

[draft] Update llama.cpp to b11100 - #1061

Merged
aittalam merged 10 commits into
mainfrom
llamacpp-b11100
Sep 30, 2026
Merged

aittalam merged 10 commits into
mainfrom
llamacpp-b11100

Conversation

@aittalam

@aittalam aittalam commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Description

Bump the llama.cpp submodule to b11100 and reconcile the patch set.

PR Type

  • 🦙 sync with upstream llama.cpp

Checklist

  • I understand the code I am submitting.
  • I have run this code locally and verified the change.
  • New and existing tests pass locally, or I have explained why tests were not run.
  • Documentation was updated where necessary.
  • If I changed code in llama.cpp/, whisper.cpp/, or stable-diffusion.cpp/, I also updated the matching *.patches/ files.
  • I have read and followed the contribution guidelines.
  • AI Usage:
    • No AI was used.
    • AI was used in an assistive capacity.
    • This PR includes substantial AI-generated content.

Tests:

  • Smoke test succeeded on Mac (Metal), Linux (CUDA+Vulkan) and Windows (CUDA + Vulkan)
  • Integration tests succeded on Linux (CUDA)

aittalam and others added 4 commits September 22, 2026 11:59
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rebases the patch set onto b11100 (659 commits past b10441) and adapts
llamafile's own integration to what upstream moved.

Patches

- Reconcile the 12 patches that drifted. GGML_CALL follows upstream's
  reorganisations: the Vulkan backend split into several translation units
  (its callbacks are now declared in ggml-vulkan-common.h and its interface
  structs live in ggml-vulkan-buffers.cpp), and graph_optimize gained a
  ggml_backend_graph_optimize_params whose add_alloc_dep the GPU DSOs call
  back into the host with — so that pointer, upstream's lambda behind it
  (replaced with a named function, since a lambda cannot carry the
  attribute) and the Metal fusion helpers that pass it on all need the
  annotation too.
- Drop src_models_{dflash,eagle3}.cpp.patch: upstream now defines
  build_arch_graph at the end of both files, which is what they did.
- server-models: upstream replaced the per-model stopping_thread with one
  server_monitor thread; move the signal mask there, and give unload_all's
  remaining untimed wait the wait_for(30s) treatment. Same for the three
  untimed waits server-queue grew.

Build

- New upstream sources: ggml-cpu/iqp.cpp, llama-kv-cache-dsa-iswa,
  llama-memory-hybrid-idx, 9 src/models, 2 tools/mtmd/models, common/json{,
  -schema}.cpp and the 16 common/parsers chat-template parsers.
- vendor/hash: mtmd-helper now hashes bitmap IDs with hash_sha256_hex().
- Version headers: ggml.c and llama.cpp include ggml-version.h /
  llama-version.h, which CMake generates. apply-patches.sh writes both from
  the same .in templates into the source tree, where the make build, the GPU
  build scripts and metal.c all find them; the -D flags they replace are gone.
- Web UI: upstream deleted tools/ui/embed.cpp for scripts/ui-assets.cmake.
  Keep the tool as a llamafile file (upstream's b10441 copy) — it emits the
  same interface the CMake templates do, and the cosmocc build has no CMake.
- mtmd_helper_bitmap_init_from_{file,buf} gained an opt argument.

GPU backends

- metal.c: upstream replaced ggml-metal.metal with 20 kernels/*.metal, so
  extract those plus their headers and point GGML_METAL_PATH_RESOURCES at
  the app dir — the backend flattens the includes itself now, which retires
  our shader preprocessor. Adds ggml-metal-{fusion,tuning}.cpp.
- vulkan.sh/.bat: compile every backend .cpp, not just ggml-vulkan.cpp.

Verified on macOS arm64: clean round-trip (reset-repo, setup, clean build,
check) is green, llama-server serves the gzip web UI and the hashed bundles
in router mode, and the Metal dylib builds and runs from scratch. CUDA,
ROCm, Vulkan and Windows still need a run on real hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
llama.cpp b11100 removed --mmap/--no-mmap/--mlock/-dio/-ndio in favour of
-lm/--load-mode (#28334); passing them now aborts argument parsing, and two
.args examples in creating_llamafiles.md did exactly that.

Also corrects the patch README: the built-in tools flag is --tools, not
--server-tools (it has been --tools since before b10441).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b11100 deleted upstream's standalone tools/ui/embed.cpp in favour of
scripts/ui-assets.cmake, which renders tools/ui/ui.{cpp,h}.in inside a CMake
build (#28445, to simplify cross-compilation). The cosmocc build has no CMake
step, and the bump initially kept upstream's b10441 embed.cpp as a llamafile
file to fill the gap.

That left us reimplementing the templates' output, which is the part that can
silently drift from what server-http.cpp expects. Replace it with ui-embed.sh,
a translation of that script's emit_files() in the same spirit as BUILD.mk
being a translation of llama.cpp's CMake build. It renders upstream's own
templates, so the generated interface tracks upstream and cannot drift; only
the substitutions (@ASSET_ARRAYS@, @ASSET_TABLE@, @N_ASSETS@, @USE_GZIP@ and
the #cmakedefine) live here. ETags become the SHA-256 upstream computes rather
than the FNV-1a of the old tool.

This also collapses the required-asset list from two copies to one: the
generator embeds whatever it is given, and fetch-ui-assets.sh keeps the single
check it needs anyway to judge a downloaded tarball. A zero-length asset now
drops the UI with a warning instead of aborting the build, matching that
script's fallback.

Two fixes found on the way:

- build_gzip_mirror now passes gzip -n. gzip embeds the *input's* mtime, so
  identical assets with different mtimes (a locally built dist/, a repacked
  tarball) produced different bytes and different ETags. Upstream pins
  SOURCE_DATE_EPOCH=0 for the same reason. Verified: without -n a touch(1)
  changes the output, with -n it does not.
- fetch-ui-assets.sh's "KEEP IN SYNC WITH UPSTREAM tools/ui/embed.cpp" banner
  pointed at a file that is no longer upstream's; it now names
  ui_validate_assets() in scripts/ui-assets.cmake.

Verified: clean round-trip (reset-repo, setup, clean build, check) green;
llama-server serves index.html and the hashed bundles gzip-encoded, with
If-None-Match revalidating 304; the no-asset stub compiles; the empty-asset
guard takes the UI-less path. Generating all 70 assets takes ~2s and only
re-runs when the assets, the templates or the script change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
aittalam and others added 6 commits September 24, 2026 12:34
b11100 changed how the FlashAttention vector kernels are selected. fattn.cu
now gates every K-V combination on `if constexpr (GGML_CUDA_FA_<K>_<V>)`, and
upstream emits one macro per combination from ggml_cuda_fattn_vec_instances()
in ggml/cmake/common.cmake. Picking template-instance files is no longer
enough: an undefined macro is a hard error, not a skipped branch, so the CUDA
build failed with 49 "identifier GGML_CUDA_FA_* is undefined" errors in
fattn.cu.

collect_gpu_sources() now derives both the instance files and the macros from
one list of combinations, so they cannot disagree, and exports them as
CUDA_FA_DEFINES for cuda.sh/rocm.sh to append. The default set matches
upstream's (q4_0-q4_0, q8_0-q8_0, f16-f16, bf16-bf16) and --fa-all-quants
still expands to all 49. The same block goes into the four .bat files.

-DGGML_CUDA_FA_ALL_QUANTS is dropped: no source references it any more.

Verified on an L40S: ggml-cuda.so builds clean (144 sources, 0 errors) and
serves inference both from the CLI and the server -- the latter being what
exercises the yield_to_queue GPU-inline patch. The Vulkan module built from
these scripts loads and runs too. The .bat files are untested; no Windows box.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
llama-server on Vulkan dies on the *second* request at the default
--parallel 4, taking the whole process with it:

  ggml-backend.cpp: GGML_ASSERT(tensor->data != NULL && "tensor not allocated")

ggml_backend_alloc_ctx_tensors_from_buft() splits a context across buffers
when a tensor exceeds the backend's max_size (1 GiB on Vulkan). If the tail of
the context holds only views, the final alloc_tensor_range() is skipped and
those views never get ggml_backend_view_init(), so the persistent KV stream
views (layer.k_stream / v_stream) keep data == NULL. The server's state
save/restore then reads one and asserts. Slot count decides the KV stream
layout, hence the dependence on --parallel: 1 is fine, 4 is not.

This is not ours: vanilla b11100 built with Vulkan reproduces it identically
on an L40S (only the line number differs, by the offset our free_struct patch
adds earlier in the file). Upstream has it as #29221, and PR #25584 is the
fix, still in review. Carried here verbatim minus its test, because the
failure hits default settings in server mode; drop it at the bump that first
includes it.

Worth noting for #29221, which concluded the bug was gemma4-specific on an
AMD iGPU: it reproduces with dense Qwen3.5-9B on a discrete NVIDIA L40S, so
it is neither model- nor vendor-specific.

Verified: clean round-trip green; the Vulkan server serves 6/6 requests at
the default slot count with the patch, on both vanilla and llamafile, where
it previously died on the second. No change on CPU or CUDA, which never hit
the split. Integration suite over four runs matches the unpatched baseline
(1-2 pre-existing flaky failures, no segfaults), and the multimodal CLI path
that segfaulted once during testing is clean 15/15 and did not recur.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BuildMetal() reuses ~/.llamafile/v/<version>/ggml-metal.dylib whenever it
exists, before extracting sources or kernels/. With the version left at
0.10.6, a Mac that has run the 0.10.6 release would load that b10441 dylib
against the b11100 host. Earlier llama.cpp bumps (#1021, #1043) bumped the
version for the same reason.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
setenv() only changes Cosmopolitan's environ; libSystem inside the dlopened
dylib never sees it, so the variable was never set for the backend. Kernels
are found because [NSBundle bundleForClass:] resolves to the dylib's own
directory, where BuildMetal() extracts kernels/. Correct the comment
accordingly, and warn when the user's environment sets the variable, since
it then overrides the extracted kernels.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
llama.cpp b11100 replaced --mmap/--no-mmap/--mlock/-dio/-ndio with
--load-mode and now rejects them, so existing llamafiles whose .args use
them would exit at startup. Translate them into one --load-mode (with a
warning) placed at the last old flag, so the last flag still wins when mixed
with --load-mode.

The mode follows the flags' original meaning as independent settings: mmap
on by default, --mlock adds locking to it (mmap+mlock), direct I/O takes
precedence. Note this differs from b10441's deprecation shim, which mapped
--mlock to plain mlock and so silently dropped mmap.

Docs: --load-mode is the syntax to use, in .args too; a new "Model Loading
Flags" section gives the old-to-new mapping.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
common_log::add() and add_json() (new in b11100) block on an untimed
cv_full.wait() when the ring is full. Convert both to wait_for(30s) loops,
as done for the other waits, to avoid the XNU untimed-futex expiry.

README: move the eagle3/dflash note out of the Bug Fixes table, where it
pushed the ggml_src_ggml.c.patch row out of the table.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@aittalam
aittalam marked this pull request as ready for review September 30, 2026 14:57
@aittalam
aittalam merged commit fef6e44 into main Sep 30, 2026
4 checks passed
@aittalam
aittalam deleted the llamacpp-b11100 branch September 30, 2026 14:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant