On arm64 the AXWii mix (ramped mix-add, volume envelope, aux returns
and the big-endian bus marshalling) ran the scalar loops, since only an
AVX2 form existed. The Neon forms do eight 16-bit samples per step as
two int32x4 halves and keep the scalar loops for the tail.
The header part of rooklz's 43463622; its other changes are left out.
A new test checks every vector form against the scalar loops at every
tail length and at the ramp and clamp extremes; it passes under
qemu-aarch64 (NEON), with -mavx2 and with neither.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018NKSMYZU43wUfYtm1sxGsC