Skip to content

perf: vectorize attention value accumulation - #9

Open
Grar00t wants to merge 1 commit into
mainfrom
perf/attention-simd-20261001
Open

Grar00t wants to merge 1 commit into
mainfrom
perf/attention-simd-20261001

Conversation

@Grar00t

@Grar00t Grar00t commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Vectorizes the remaining scalar attention V-accumulation loop using the existing AVX2+FMA and NEON paths, with scalar tails/fallback preserved.

Scope deliberately stays narrow:

  • no matvec rewrite
  • no GQA index change (kvh = (h * n_kv_heads) / n_heads preserved)
  • no ABI/model-format change
  • no proof-format/hash change
  • no quantization, SDOT, SVE, SME, or invented weight path

Observed before change: attention score already routes through SIMD dot_f32, so it was not rewritten.

Verification performed locally from current origin/main (4504604):

  • bash scripts/build.sh --arch generic --smoke PASS
  • bash scripts/build.sh --arch x86_64 --smoke --bench PASS, self-check reports AVX2+FMA
  • bash scripts/build.sh --debug --arch generic --smoke PASS with ASan+UBSan build
  • git diff --check PASS

Ad-hoc long-context attention benchmark on the same machine/config (embed=256, heads=8, kv_heads=2, head_dim=32, ctx=512, one layer) improved median throughput from ~9,971 tok/s to ~15,328 tok/s (~1.54x). This benchmark was temporary and not added to the repository.

AArch64 cross-syntax verification could not be completed because the available clang lacks an AArch64 libc/sysroot; the NEON implementation uses the same existing vfmaq_f32/load/store intrinsics already used elsewhere in this file.

Copilot AI balanced review requested due to automatic review settings October 1, 2026 10:45
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@vercel

vercel Bot commented Oct 1, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
casper Ready Ready Preview, v0 Oct 1, 2026 10:45am UTC
casper-7z Ready Ready Preview, v0 Oct 1, 2026 10:45am UTC
casper-nl Ready Ready Preview, v0 Oct 1, 2026 10:45am UTC

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The SIMD implementations preserve the scalar operation, handle tails correctly, and retain existing architecture fallbacks.

Review effort: Balanced
Findings: None

What changed in this PR

Vectorizes attention value accumulation while preserving scalar fallback and existing model behavior.

Changes:

  • Adds AVX2/FMA and NEON AXPY paths with scalar tails.
  • Uses the helper during attention value accumulation.
File Description
Core_CPP/​niyah_core.c Adds SIMD value accumulation and integrates it into inference.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

This branch was successfully deployed

3 active deployments
Preview – casper-7z — be6eb48c Deployed Oct 1, 2026 by vercel[bot]
Preview – casper — be6eb48c Deployed Oct 1, 2026 by vercel[bot]
Preview – casper-nl — be6eb48c Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants