Skip to content

array: softmax on the CPU spreads its rows across the thread pool #85

array: softmax on the CPU spreads its rows across the thread pool

array: softmax on the CPU spreads its rows across the thread pool #85

Workflow file for this run

name: CI
on:
push:
branches: [main, master]
pull_request:
jobs:
# Static, no build needed: is a GPU op real on the same backends everywhere,
# or did one backend land ahead of the others and never get caught up (that
# is exactly what happened to concat_part/rope before this job existed —
# see tools/check_backend_parity.py's own docstring). Deliberate gaps (the
# M9 LLM decode fast path, still CUDA-only) are recorded in
# tools/backend_parity_allowlist.txt and don't fail this job; an
# unallowlisted asymmetry does.
backend-parity:
name: backend parity (CUDA / Metal / WebGPU)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: python3 tools/check_backend_parity.py
# All three modes — cpu / gpu / auto — on every OS.
# What each hosted runner actually offers:
# - macos-15 (arm64): Metal works through a paravirtual GPU -> a real GPU test
# - ubuntu / windows: no GPU -> gpu/auto exercise the CPU fallback path
test:
name: test / ${{ matrix.os }}
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-15, windows-latest]
steps:
- uses: actions/checkout@v4
- name: Configure
run: cmake -B build -DCMAKE_BUILD_TYPE=Release
- name: Build
run: cmake --build build --config Release
- name: Test (cpu / gpu / auto)
run: ctest --test-dir build -C Release --output-on-failure
# Build verification for the CUDA toolchain. Hosted runners have no NVIDIA
# driver, so the kernels are only compiled and linked. Running the tests here
# doubles as coverage of the dlopen fallback on a driver-less machine, which
# is a permanent test target in its own right.
cuda-build:
name: cuda build+fallback / ${{ matrix.os }}
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
# windows-2022 (not -latest): windows-latest moved to Visual Studio 18,
# which is past nvcc's supported host-compiler window (VS 2017-2022) and
# hard-errors the PTX build (C1189). VS 2022 is in range. The plain
# `test` job below stays on windows-latest — only the CUDA/nvcc path
# needs the pinned toolset.
os: [ubuntu-latest, windows-2022]
steps:
- uses: actions/checkout@v4
- uses: Jimver/cuda-toolkit@v0.2.19
with:
method: network
sub-packages: '["nvcc", "cudart"]'
# nvcc -ptx still preprocesses with the host compiler; on Windows that is
# cl.exe, which msvc-dev-cmd puts on PATH (no-op on Linux via the guard).
- name: Set up MSVC (Windows)
if: runner.os == 'Windows'
uses: ilammy/msvc-dev-cmd@v1
- name: Configure
run: cmake -B build -DCMAKE_BUILD_TYPE=Release -DTENSORLIB_CUDA=ON
- name: Build
run: cmake --build build --config Release
- name: Test (CPU fallback — no driver on hosted runners)
run: ctest --test-dir build -C Release --output-on-failure
# WebGPU (wasm) backend. Unlike cuda-build this does not stop at building —
# it runs the suite.
#
# It can, because what needs verifying is the WGSL and the C++, not Chrome:
# any runtime with navigator.gpu and JSPI will do. Deno has both, and on a
# GPU-less runner it still returns an adapter, through lavapipe (Mesa's
# software Vulkan). Headless Chrome under the same conditions does not expose
# navigator.gpu at all, across three attempts — including with a current
# Chrome and that same lavapipe working. So the obstacle was Chrome's own GPU
# gating, not the runner.
#
# This leaves WebGPU better covered than CUDA rather than worse. nvcc -ptx
# checks CUDA kernels at build time; WGSL is compiled at run time by the
# browser or runtime, so until now a shader error was invisible to the build.
# Running the suite compiles them.
#
# Both gpu and auto run against the ref oracle. auto is included because the
# TENSORLIB_WEBGPU arm of types.h's auto_threshold_() is unreachable from
# native ctest. Note the suite stays green even if every op falls back to
# CPU, so main_wasm.cpp asserts that each kernel family dispatched at least
# once and fails the run otherwise.
wasm:
name: wasm run (webgpu) / ubuntu-latest
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: mymindstorm/setup-emsdk@v14
with:
# Match the machine this was validated on: --use-port=emdawnwebgpu
# pulls a Dawn tied to the emsdk version, so moving this moves
# behaviour.
version: 6.0.3
- uses: denoland/setup-deno@v2
with:
deno-version: v2.x
- name: Install a software Vulkan driver (lavapipe)
run: |
sudo apt-get update -qq
sudo apt-get install -y -qq mesa-vulkan-drivers
# Build the census too, though nothing here runs it (it measures for
# minutes and is a local tool) — that it still compiles is the point.
- name: Build (tests + census)
run: ./test/wasm/build.sh
- name: Run the suite (gpu / auto)
run: deno run --allow-all test/wasm/deno_run.js
# Running the tests on a real CUDA GPU is deliberately out of CI (2026-07):
# paid GPU runners need an Org and bill per minute, and a self-hosted one is
# upkeep we chose not to take on. CUDA's numerical regression is instead a
# manual `ctest --test-dir build` on real NVIDIA hardware.
# To change that, register a runner and set HAS_GPU_RUNNER=true to enable this.
cuda-gpu:
if: vars.HAS_GPU_RUNNER == 'true'
name: cuda gpu / self-hosted
runs-on: [self-hosted, gpu]
steps:
- uses: actions/checkout@v4
- name: Configure
run: cmake -B build -DCMAKE_BUILD_TYPE=Release -DTENSORLIB_CUDA=ON
- name: Build
run: cmake --build build --config Release
- name: Test (cpu / gpu / auto on real GPU)
run: ctest --test-dir build -C Release --output-on-failure