Speed up affine-reset mm path with factor0 copy and 16-wide SIMD.
Keep scalar column IR so AArch64 can widen MUL_ADD/COPY/SPLAT_STORE to 16 lanes, and lower factor==0 updates to a row copy.
Co-authored-by: Cursor cursoragent@cursor.com
Keep scalar column IR so AArch64 can widen MUL_ADD/COPY/SPLAT_STORE to 16 lanes, and lower factor==0 updates to a row copy.
Co-authored-by: Cursor cursoragent@cursor.com