Skip to content

Fix the VMX backend - #10

Open
pkubaj wants to merge 1 commit into
castano:masterfrom
pkubaj:freebsd_ppc64
Open

pkubaj wants to merge 1 commit into
castano:masterfrom
pkubaj:freebsd_ppc64

Conversation

@pkubaj

@pkubaj pkubaj commented Aug 6, 2026

Copy link
Copy Markdown

The ICBC_VMX path had evidently never been compiled: "vectro float", a
vec_ld() call with the wrong arity and no terminating semicolon, vec_div()
against an integer splat (there is no integer vector divide on PowerPC, so
this is ambiguous rather than promoted), and a compound literal that VFloat
has no constructor for. all(), any() and mask() were missing from the
backend entirely; mask() is used by compute_indices3() and
compute_indices4().

vload() now uses vec_xl() where VSX is available. vec_ld() is lvx, which
masks off the low four bits of the address and therefore quietly returns
the containing aligned vector for an unaligned pointer.

mask() follows the NEON implementation, since PowerPC has no movemask:
test each lane's sign bit, AND with {1,2,4,8}, then OR the lanes together.

Verified on FreeBSD/powerpc64le with clang 19: compress_bc1() over 20000
pseudo-random blocks gives byte-identical output and an identical total
error for ICBC_SIMD=ICBC_VMX and ICBC_SIMD=ICBC_SCALAR.

The ICBC_VMX path had evidently never been compiled: "vectro float", a
vec_ld() call with the wrong arity and no terminating semicolon, vec_div()
against an integer splat (there is no integer vector divide on PowerPC, so
this is ambiguous rather than promoted), and a compound literal that VFloat
has no constructor for.  all(), any() and mask() were missing from the
backend entirely; mask() is used by compute_indices3() and
compute_indices4().

vload() now uses vec_xl() where VSX is available.  vec_ld() is lvx, which
masks off the low four bits of the address and therefore quietly returns
the containing aligned vector for an unaligned pointer.

mask() follows the NEON implementation, since PowerPC has no movemask:
test each lane's sign bit, AND with {1,2,4,8}, then OR the lanes together.

Verified on FreeBSD/powerpc64le with clang 19: compress_bc1() over 20000
pseudo-random blocks gives byte-identical output and an identical total
error for ICBC_SIMD=ICBC_VMX and ICBC_SIMD=ICBC_SCALAR.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant