array: softmax on the CPU spreads its rows across the thread pool #85
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| name: CI | |
| on: | |
| push: | |
| branches: [main, master] | |
| pull_request: | |
| jobs: | |
| # Static, no build needed: is a GPU op real on the same backends everywhere, | |
| # or did one backend land ahead of the others and never get caught up (that | |
| # is exactly what happened to concat_part/rope before this job existed — | |
| # see tools/check_backend_parity.py's own docstring). Deliberate gaps (the | |
| # M9 LLM decode fast path, still CUDA-only) are recorded in | |
| # tools/backend_parity_allowlist.txt and don't fail this job; an | |
| # unallowlisted asymmetry does. | |
| backend-parity: | |
| name: backend parity (CUDA / Metal / WebGPU) | |
| runs-on: ubuntu-latest | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - run: python3 tools/check_backend_parity.py | |
| # All three modes — cpu / gpu / auto — on every OS. | |
| # What each hosted runner actually offers: | |
| # - macos-15 (arm64): Metal works through a paravirtual GPU -> a real GPU test | |
| # - ubuntu / windows: no GPU -> gpu/auto exercise the CPU fallback path | |
| test: | |
| name: test / ${{ matrix.os }} | |
| runs-on: ${{ matrix.os }} | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| os: [ubuntu-latest, macos-15, windows-latest] | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - name: Configure | |
| run: cmake -B build -DCMAKE_BUILD_TYPE=Release | |
| - name: Build | |
| run: cmake --build build --config Release | |
| - name: Test (cpu / gpu / auto) | |
| run: ctest --test-dir build -C Release --output-on-failure | |
| # Build verification for the CUDA toolchain. Hosted runners have no NVIDIA | |
| # driver, so the kernels are only compiled and linked. Running the tests here | |
| # doubles as coverage of the dlopen fallback on a driver-less machine, which | |
| # is a permanent test target in its own right. | |
| cuda-build: | |
| name: cuda build+fallback / ${{ matrix.os }} | |
| runs-on: ${{ matrix.os }} | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| # windows-2022 (not -latest): windows-latest moved to Visual Studio 18, | |
| # which is past nvcc's supported host-compiler window (VS 2017-2022) and | |
| # hard-errors the PTX build (C1189). VS 2022 is in range. The plain | |
| # `test` job below stays on windows-latest — only the CUDA/nvcc path | |
| # needs the pinned toolset. | |
| os: [ubuntu-latest, windows-2022] | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - uses: Jimver/cuda-toolkit@v0.2.19 | |
| with: | |
| method: network | |
| sub-packages: '["nvcc", "cudart"]' | |
| # nvcc -ptx still preprocesses with the host compiler; on Windows that is | |
| # cl.exe, which msvc-dev-cmd puts on PATH (no-op on Linux via the guard). | |
| - name: Set up MSVC (Windows) | |
| if: runner.os == 'Windows' | |
| uses: ilammy/msvc-dev-cmd@v1 | |
| - name: Configure | |
| run: cmake -B build -DCMAKE_BUILD_TYPE=Release -DTENSORLIB_CUDA=ON | |
| - name: Build | |
| run: cmake --build build --config Release | |
| - name: Test (CPU fallback — no driver on hosted runners) | |
| run: ctest --test-dir build -C Release --output-on-failure | |
| # WebGPU (wasm) backend. Unlike cuda-build this does not stop at building — | |
| # it runs the suite. | |
| # | |
| # It can, because what needs verifying is the WGSL and the C++, not Chrome: | |
| # any runtime with navigator.gpu and JSPI will do. Deno has both, and on a | |
| # GPU-less runner it still returns an adapter, through lavapipe (Mesa's | |
| # software Vulkan). Headless Chrome under the same conditions does not expose | |
| # navigator.gpu at all, across three attempts — including with a current | |
| # Chrome and that same lavapipe working. So the obstacle was Chrome's own GPU | |
| # gating, not the runner. | |
| # | |
| # This leaves WebGPU better covered than CUDA rather than worse. nvcc -ptx | |
| # checks CUDA kernels at build time; WGSL is compiled at run time by the | |
| # browser or runtime, so until now a shader error was invisible to the build. | |
| # Running the suite compiles them. | |
| # | |
| # Both gpu and auto run against the ref oracle. auto is included because the | |
| # TENSORLIB_WEBGPU arm of types.h's auto_threshold_() is unreachable from | |
| # native ctest. Note the suite stays green even if every op falls back to | |
| # CPU, so main_wasm.cpp asserts that each kernel family dispatched at least | |
| # once and fails the run otherwise. | |
| wasm: | |
| name: wasm run (webgpu) / ubuntu-latest | |
| runs-on: ubuntu-latest | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - uses: mymindstorm/setup-emsdk@v14 | |
| with: | |
| # Match the machine this was validated on: --use-port=emdawnwebgpu | |
| # pulls a Dawn tied to the emsdk version, so moving this moves | |
| # behaviour. | |
| version: 6.0.3 | |
| - uses: denoland/setup-deno@v2 | |
| with: | |
| deno-version: v2.x | |
| - name: Install a software Vulkan driver (lavapipe) | |
| run: | | |
| sudo apt-get update -qq | |
| sudo apt-get install -y -qq mesa-vulkan-drivers | |
| # Build the census too, though nothing here runs it (it measures for | |
| # minutes and is a local tool) — that it still compiles is the point. | |
| - name: Build (tests + census) | |
| run: ./test/wasm/build.sh | |
| - name: Run the suite (gpu / auto) | |
| run: deno run --allow-all test/wasm/deno_run.js | |
| # Running the tests on a real CUDA GPU is deliberately out of CI (2026-07): | |
| # paid GPU runners need an Org and bill per minute, and a self-hosted one is | |
| # upkeep we chose not to take on. CUDA's numerical regression is instead a | |
| # manual `ctest --test-dir build` on real NVIDIA hardware. | |
| # To change that, register a runner and set HAS_GPU_RUNNER=true to enable this. | |
| cuda-gpu: | |
| if: vars.HAS_GPU_RUNNER == 'true' | |
| name: cuda gpu / self-hosted | |
| runs-on: [self-hosted, gpu] | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - name: Configure | |
| run: cmake -B build -DCMAKE_BUILD_TYPE=Release -DTENSORLIB_CUDA=ON | |
| - name: Build | |
| run: cmake --build build --config Release | |
| - name: Test (cpu / gpu / auto on real GPU) | |
| run: ctest --test-dir build -C Release --output-on-failure |