Skip to content

feat: low-precision float types bf16/float16/float8/float4 and vecf8/vecf4 vector columns (#20567) - #29554

Draft
cpegeric wants to merge 88 commits into
matrixorigin:mainfrom
cpegeric:bug_20567_vector
Draft

cpegeric wants to merge 88 commits into
matrixorigin:mainfrom
cpegeric:bug_20567_vector

Conversation

@cpegeric

@cpegeric cpegeric commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

What type of PR is this?

  • API-change
  • BUG
  • Improvement
  • Documentation
  • Feature
  • Test and CI
  • Code Refactoring

Which issue(s) this PR fixes:

issue #20567

What this PR does / why we need it:

Low-precision float types: scalar bf16/float16/float8/float4, and block-scaled vector columns vecf8(N) (MXFP8) / vecf4(N) (NVFP4).

Draft: the vector_matmul / vector_matmul_merge batch search (CPU, then cuBLASLt GPU) continues on this branch.

Scalar bf16 / float16 / float8 (E4M3) / float4 (E2M1)

Design: docs/design/20260929-low-precision-scalar-floats.md

  • Type ids, parser, container, casts via a float32 bridge, arithmetic/comparison/aggregates, ORDER BY, TAE zonemap and value codec, MySQL protocol, LOAD.
  • Float8/Float4 codecs pinned bit-exactly to CUDA cuda_fp8.h/cuda_fp4.h by a generated golden test.

Vector columns vecf8(N) / vecf4(N)

Design: docs/design/20260930-low-precision-vector-storage.md

  • Cell: 12-byte header (version, format, dim, fp32 global scale g), then block scales (E8M0 per 32 for vecf8; UE4M3 per 16 for vecf4), then elements. Value = g × block scale × element; the layout matches cuBLASLt's block-scaled matmul inputs.
  • DDL: column types in CREATE/ALTER; rejected in primary key, partition key, and every index.
  • Casts: text ↔ vecf8/vecf4, vecf32 ↔ vecf8/vecf4.
  • Arithmetic + - * / promotes to vecf32.
  • Distances: inner_product, l2_distance, l2_distance_sq, l1_distance, cosine_distance, cosine_similarity between vecf8/vecf4/vecf32 in any combination; a text literal binds as vecf32.
  • Also vector_dims, normalize_l2, ANY_VALUE, COUNT, GROUP_CONCAT, ORDER BY / GROUP BY / window.
  • Output, export, CDC/ISCP replication, and LOAD (CSV, Parquet LIST<FLOAT/DOUBLE> and text columns).
  • Distance kernels (pkg/vectorindex/metric/distance_func_vecblock*.go): scalar Go, no SIMD. They are generated (mkvecblock.go), one per operand pair and metric, fully unrolled over 16-element units, with table-lookup decode. 768-d, one row: 200–311 ns for dot/L2, 303–437 ns for L1/cosine.

perf(metric): clear AVX upper state after archsimd kernels

The archsimd distance kernels returned with dirty upper zmm state. The legacy-SSE scalar float code that follows them then carries a false dependency on those bits.

  • Fix: archsimd.ClearAVXUpperBits() after the reduction in 40 kernels. The 4 bf16 AVX2 kernels are excluded (+17% there, no gain after them).
  • Effect, 768-d, Ryzen AI 9 365:
    • vecf8 dequantize ×2 + f32 dot product: 1702 → 364 ns.
    • Scalar loop after an AVX-512 kernel: −7.6%.
    • f16 AVX2 l2sq / l1: −20% / −11%.

Tests

  • Unit tests: codecs (incl. CUDA golden), parse/validation, decode paths, kernels against a float64 reference at tail dims and in both argument orders, SQL functions, casts, frontend/LOAD/CDC paths, zonemap/sort/compare.
  • BVT: dtype/vecblock.sql (65 statements, including flush, prepared statements and rejections), plus the scalar low-precision cases.
  • Regression (mo-tester -m run -g -n), all 0 failed:
Directory Passed Ignored
dtype 7729/7761 32
array 1501/1501 0
vector 2604/2606 2

🤖 Generated with Claude Code

cpegeric and others added 30 commits September 29, 2026 14:38
…rixorigin#20567)

Add the low-precision float value types for the new scalar bf16/float16/float8/
float4 SQL columns. Float8 is OCP/NVIDIA FP8 e4m3 (1-byte, no Inf, max 448); Float4
is OCP MXFP4 e2m1 (1-byte, magnitudes {0,.5,1,1.5,2,3,4,6}, NVIDIA-bit-compatible).
Both convert to/from float32 (round-to-nearest-even, saturating overflow) so all
arithmetic runs on the existing float32 kernels. Reference-value + full round-trip
unit tests included.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rixorigin#20567)

Add T_bf16=73, T_float16=74, T_float8=75, T_float4=76 to the type system: SQL
names, sizes (2/2/1/1 bytes), FixedLength/TypeLen, ToType, String/OidString, and
IsFloat classification. They satisfy FixedSizeT already via ~uint16/~uint8, so
the vector container generics accept them without constraint changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…origin#20567)

Add BF16/FLOAT16/FLOAT8/FLOAT4 keyword tokens and no-length column_type grammar
rules (float4/float8 were reserved-but-UNUSED), regenerate the mysql parser, and
map them in getTypeFromAst (MYSQL_TYPE_FLOAT disambiguated by FamilyString) to
T_bf16/T_float16/T_float8/T_float4. Fixes the issue's 1064 on CREATE TABLE.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rt (matrixorigin#20567)

Start Layer 4 (vector container): GetAny dispatches the four new scalar float type
IDs to their value types. Remaining container switches (Append/Shrink/Shuffle/
String/min-max/compare/sort) still to do -- sort/min-max must compare by float
value, not raw bits.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tches (matrixorigin#20567)

Add the four scalar float type IDs to the value-access (GetAny/AppendAny/String/
RowToString) and movement (Shrink/ShrinkByMask/Shuffle/ShuffleWithBuf) container
switches. Union/const-set and the compare/sort/min-max/sum switches (which must
order by float value, not raw bits) still to do.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…oat4 (matrixorigin#20567)

Contract, format decisions (FP8 e4m3, FP4 e2m1/MXFP4), and invariants (float32
bridge arithmetic, value-order comparison, saturating RNE narrowing). sparsevec
deferred to a separate design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…float4 (matrixorigin#20567)

The four new scalar float types are byte-identical to uint16 (bf16/float16) and
uint8 (float8/float4) for the value-agnostic union-all and const-set copy paths, so
add them to those case labels. Remaining container work: compare/sort/min-max/sum
must order by float value.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…float16/float8/float4 (matrixorigin#20567)

The container ordering and aggregation paths fell through to raw-bit handling for
the low-precision float types, which is wrong: a float's sign bit makes raw uint
order disagree with value order (negatives sort as large positives). Route them
through ToFloat32:

- compareVectorRows: add bf16/float16/float8/float4 via new compareFloatRows.
- InplaceSort / InplaceSortAndCompact: sort (and value-dedup) via ToFloat32; add
  the four types to supportsInplaceSort.
- GetSumValue / GetMinMaxValue: widen via new LowPrecFloatGetSum /
  LowPrecFloatGetMinAndMax; min/max skips NaN like the float32/64 path.
- typeCompatible: recognize the four named types for type-checked column access
  (needed by appendList in the compact path).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…32 bridge (matrixorigin#20567)

Casts to/from the low-precision float types bridge through float32: any source
widens to float32 then rounds to the target format via FromFloat32, and any
low-precision source widens to float32 then reuses the float32 target machinery
(inheriting exact integer-rounding, decimal, and string-formatting semantics).

- func_cast_lowprec_float.go: castToLowPrecFloat (X->newtype), lowPrecFloatToOthers
  (newtype->X), string parse, decimal256 routed to castToDecimal256; init() wires
  the supportedTypeCast gate for all numeric/string sources and targets.
- func_cast.go: route low-precision source/target ahead of the integer/decimal
  fast-paths so a low-precision operand is never misdispatched.
- functionTools.go: NewFunctionResultWrapper builds result vectors for the 4 types
  (was panicking).
- func_testcase.go: test framework builds and compares the 4 types.
- func_cast_test.go: TestCastLowPrecFloat covers float64<->each, a cross-cast, and
  string<->float8.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…16/float16/float8/float4 (matrixorigin#20567)

Binary operators and aggregates treat the low-precision float types as float32,
which holds every bf16/float16/float8/float4 value losslessly:

- type_check.go: fixedTypeCastRule1 (arithmetic + comparison) and fixedTypeCastRule2
  (div) normalize a low-precision operand to float32 before selecting the coercion
  rule, mirroring the existing T_enum->uint16 treatment. float32 has no diagonal
  cast rule (equal types need none), so a low-precision pair that both normalize to
  float32 forces the cast to float32 explicitly.
- list_agg.go: sumAvgTypeCheck and a new minMaxTypeCheck widen a low-precision input
  to float32 (SUM/AVG return float64; MIN/MAX return float32) -- newGenericMinMaxExec
  compares raw uint bits, which do not order floats, so native support is unsafe.

The operand cast itself uses the float32 bridge added for CAST.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t8/float4 (matrixorigin#20567)

The storage/zonemap comparators must order these types by float VALUE, not the raw
uint bits (a negative float's bits exceed a positive's), and several switches panicked
on the unregistered types:

- types.EncodeValue/DecodeValue: register the four types (were panicking in the
  zonemap value path, e.g. UpdateZMAny).
- compute.Compare/CompareGeneric: compare via ToFloat32 (drives zonemap min/max
  maintenance in UpdateZM and Contains/containsBytes pruning).
- zm.getValue: decode the four types (MO_TABLE_COL_MAX read path panicked otherwise).
- zm.SubVecIn: value-aware bound search on a value-sorted column (subVecInLowPrecFloat);
  BatchUpdateZM already feeds value-correct min/max from vector.GetMinMaxValue.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…oat8/float4 (matrixorigin#20567)

Result-set output and CSV/LOAD parsing for the low-precision float types, presenting
them as float32 (they widen losslessly):

- frontend: column type maps to MYSQL_TYPE_FLOAT; the colSlices path widens the column
  into arrFloat32 at build time (lowPrecFloatVecToFloat32Slice) so GetFloat32/GetFloat64
  and the row extractors read it as float32; the legacy any-row and ExtractRowFromEveryVector
  paths return .ToFloat32().
- external LOAD: parse the field as float32 then round to the target format; the empty
  zero-fill and the row-validity predicate accept the four types.

Strict finite/range checking on the string-parse boundary (LOAD + CAST-from-string):
types.RejectNonFiniteNarrowFloat rejects NaN/Inf and out-of-range values instead of
silently saturating, mirroring the narrow-vector element parser (rejectNonFiniteArrayElem).
bf16/float16 overflow is detected as Inf after narrowing; float8/float4 saturate, so their
input magnitude is range-checked (448 / 6). Numeric CAST keeps the saturating constructor,
matching the documented vecint8 precedent (string parse strict, CAST clamps).

Tests: RejectNonFiniteNarrowFloat, Inf conversion (bf16/float16 preserve, float8/float4
saturate), subnormals for all four, and the frontend colSlices output path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…f coverage to ~87%

Add unit tests for the new-code paths that the feature tests did not reach, so the
PR's aggregate diff coverage clears the 75% gate:

- vector: GetAny/AppendAny/String/RowToString/Shrink/ShrinkByMask/Shuffle/ShuffleWithBuf
  and NewFunctionResultWrapper for all four types.
- types: type-registration switches (TypeLen/FixedLength/String/OidString/IsFloat/ToType),
  EncodeValue/DecodeValue, and the slice converters.
- compute: Compare and CompareGeneric order by float value (negative discriminator).
- tae/index: fix TestZonemapLowPrecFloat to use >3 rows so SubVecIn reaches the per-type
  bound search instead of the small-vector early return.
- external: getColData LOAD parse (valid rounds, empty zero-fills, out-of-range/non-finite
  rejected) and appendLoadEmptyNumericZero.
- frontend: getValueFromVector, extractRowFromVector (legacy any-row), and
  convertEngineTypeToMysqlType present the types as FLOAT.
- plan: getTypeFromAst disambiguates the four via FamilyString.
- function: broaden the cast matrix across source types -> bf16 and float8 -> targets.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…float16/float8/float4; add BVT

BVT (test/distributed/cases/dtype/lowprec_float) surfaced two gaps beyond the earlier
per-layer work:

- pkg/sort: ORDER BY did not sort these columns -- IsSupportedType and the comparator
  switch omitted them, and the sortType constraint lacked the scalar slice types, so the
  order operator treated them as equal (stable -> insertion order). Add the four types to
  the constraint, IsSupportedType, and the switch with value-aware comparators (widen via
  ToFloat32, reusing the float32 SQL NaN-aware ordering).
- function: inserting a negative literal (e.g. -2.0) into such a column failed with
  'unary_minus bad value [BF16]'. unaryMinusMatch and UNARY_PLUS (new unaryPlusMatch) now
  cast a low-precision operand to float32, mirroring the binary-operator coercion.

BVT covers DDL, insert (incl. negatives/NULL), value-ordered ORDER BY, comparison,
arithmetic, SUM/AVG/MIN/MAX, CAST both ways, out-of-range reject, and INSERT..SELECT
round-trip; validated 27/27 on a live cluster. Unit tests: TestSortLowPrecFloat and
unary coverage in the ops test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts:
#	pkg/sql/plan/function/list_agg.go
#	pkg/sql/plan/function/type_check.go
…256 cast, non-finite persistence

Three defects found by the self-review gate:

- types/float8.go: Float8FromFloat32 subnormal->normal round-up silently returned
  ZERO. roundMantissaRNE masks its result and signals the boundary overflow only via
  its carry, but the subnormal branch dropped the carry (mant, _ :=) and its mant>=8
  guard was therefore dead. Every value in (0.013671875, 0.015625) narrowed to 0
  instead of the smallest normal 0x08 -- silent data loss. Honor the carry.

- function/func_cast_lowprec_float.go: decimal256 was registered in supportedTypeCast
  but anyToLowPrecFloat lacked the case, so a planned CAST(decimal256 AS bf16/...)
  failed at execution with an internal error. Add the decimal256 case
  (types.Decimal256ToFloat64).

- function/func_cast_lowprec_float.go: the numeric cast path applied no finite check,
  so an overflowing float64/decimal source persisted +Inf for bf16/float16 (and NaN
  for float8), violating the repo-wide never-persist-non-finite invariant (matrixorigin#29084) and
  diverging from the string path. Route the numeric path through
  opUnaryFixedToFixedWithErrorCheck + RejectNonFiniteNarrowFloat so NaN/Inf/out-of-range
  errors uniformly, matching the string cast and matrixorigin#29084's rejectNonFiniteVectorElems.

- tae/index/zm.go: hasNaNBound now covers bf16/float16/float8 for defense-in-depth
  parity with float32/float64 (float4 has no NaN encoding). Latent today since
  GetMinMaxValue skips NaN, but keeps the fail-open guard consistent.

Regression tests: TestFloat8SubnormalRoundUpToNormal, TestCastLowPrecFloatDecimal256AndNonFinite.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…nstraint types.LowPrecFloat

Replace the inline interface{ ToFloat32() float32 } (and the multi-line
FixedSizeTExceptStrType + ToFloat32 form) repeated across the ordering, compare,
sum/min-max, sort, zonemap, frontend, and cast helpers with a single named constraint
types.LowPrecFloat = { FixedSizeTExceptStrType; ToFloat32() float32 }. The union+method
together admit exactly bf16/float16/float8/float4, so the fixed-size storage bound
(DecodeFixed / column accessors) and the value-widening method live in one place.
No behavior change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ecimal256 cases

Cover the self-review fixes end-to-end on a live cluster (47/47):
- out-of-range NUMERIC cast rejected (not just string): CAST(7 AS float4),
  CAST(1000 AS float8), CAST(70000 AS float16), CAST(1e300 AS bf16).
- Inf/NaN cast rejected: CAST('inf'/'nan'/'-inf' ...), and CAST(1e308*100 AS bf16)
  (arithmetic overflow to Inf).
- arithmetic result overflowing the narrow range rejected on cast-back:
  CAST(6*2 AS float4), CAST(a+b AS float4) with 3+4>6.
- FP8 subnormal round-up regression: CAST(0.015 AS float8) = 0.015625 (was silently
  zero); min-subnormal round-trip; underflow to 0; float4 0.5 subnormal.
- decimal256 -> bf16/float8.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…trixorigin#20567)

Design for packed low-precision float vector columns for quantized ML embeddings:
per-32-block E8M0 (MXFP) scaling, one shared byte-exact-MXFP payload format for both
e4m3 (vecf8) and e2m1 (vecf4) elements, and an MO-only header stripped so the GPU path
is a straight memcpy to cuVS/CUTLASS MXFP kernels (no dequant/requant). Contract,
decisions, and invariants only; depends on the scalar float8/float4 element codecs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… P4 dependency (matrixorigin#20567)

- SQL spelling is vecf8/vecf4 to match the existing family (vecf32/vecf64/vecf16/vecbf16),
  not vecfloat8/vecfloat4 (there is no vecfloat16 -- it is vecf16).
- Note that MO's current cuVS integration supports only float/half/int8/uint8 with no MXFP
  ingestion, and the cuVS MXFP contract is not vendored (GPU-box only). CPU storage/distance
  can proceed now; the GPU memcpy path is a forward-looking target pending a cuVS MXFP API.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ot cuVS (matrixorigin#20567)

Rewrite the GPU section: fp8/fp4 are tensor-core GEMM formats, so vecf8/vecf4's GPU value
is brute-force batch similarity (Q x D^T) via cublasLtMatmul block-scaled matmul, NOT a
cuVS index (cuVS accepts only float/half/int8/uint8 + SQ/PQ/binary; no fp8/fp4 path). The
stored 32-block E8M0 MXFP payload is exactly cuBLASLt's CUDA_R_UE8 layout, so transfer is
strip-header + memcpy into the operand/scale tensors. Records hardware (Blackwell sm_100 /
consumer sm_120 incl. RTX 5070), toolchain (CUDA 12.8+/cuBLAS 12.9+), mandatory dataset
tiling, MXFP4-vs-NVFP4 fp4 choice, and a P1-P3 CPU / P4 GPU phasing. New cuBLASLt cgo
integration, separate from cgo/cuvs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…matrixorigin#20567)

Lock the two open decisions:
- vecf4 = NVFP4 (e2m1 4-bit elements + E4M3 unsigned 16-block scale, per-vector-normalized,
  GEMM global=1.0, no stored global). Dictated by cuBLASLt: its FP4 block-scaled matmul
  consumes NVFP4 only (16-block E4M3), not MXFP4 (32-block E8M0), so storing MXFP4 would
  force a lossy re-quantization at GEMM and break the strip-header+memcpy contract. NVFP4 is
  also the more accurate 4-bit format. vecf8 = MXFP8 (e4m3 + E8M0 32-block), which is what
  cuBLASLt's FP8 path takes directly.
- CPU distance accumulate = fp32 (matches the GEMM's fp32 accumulate so CPU==GPU results).

Also records: element codecs are the merged scalar Float8/Float4 (bit-identical to
CUDA_R_8F_E4M3 / CUDA_R_4F_E2M1) so packing is a pure shuffle; only block-scale codecs are
new (E8M0, and reuse Float8 e4m3 for NVFP4's scale); effective bits/element and 1024-dim
storage sizes; the per-vector-independent invariant; and the P4 GPU-box verification items.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…l SQL surface (matrixorigin#20567)

Verified on RTX 5070 (sm_120), CUDA 13.3, cuBLASLt 13.6:
- types.Float8/Float4 match cuda_fp8.h/cuda_fp4.h bit for bit on all
  codes and 4.3M encode inputs (NaN encoding differs; never stored);
  FP4 element 2i is the low nibble.
- Block scales must use cuBLASLt's tiled 128x4 layout; per-row scales give
  wrong results. Elements transfer as stored, scales are re-laid out.
- Dataset must be operand A: queries-as-A with one query returns silently
  wrong results. K is padded to a multiple of 32.
- MXFP8 and NVFP4 (global 1.0) match a CPU reference; recall@10 on
  unit-normalized wiki_all 768-d: MXFP8 0.971, NVFP4 0.918.

Design: cgo/cublaslt dot-product matmul engine; vector_matmul TVF (one
partial per scan pipeline) + vector_matmul_merge aggregate with a JSON
result format; single-CN and multi-CN test plan.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…origin#20567)

Each hit is a two-element array: position 0 the source key as a JSON
string, position 1 the dot product. It halves the result size versus
{"id","score"} objects. Unnest reads '$[0]' / '$[1]'; verified on the
current build, including a 9223372036854775807 id cast back to bigint.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/cuda_fp4.h

Golden data generated on CUDA 13.3: all 256 e4m3 and 16 e2m1 decodes and
1,069 encode inputs (every rounding midpoint +/-1 ulp, subnormals,
saturation, +/-Inf). NaN encoding is excluded (documented difference).
A codec change that breaks GPU bit compatibility now fails a CPU-only UT.
The generator is cgo/cublaslt/test/narrow_float_golden_gen.cu, with its
regeneration command.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
matrixorigin#20567)

All C++/CUDA code lives in cgo/cublaslt (engine + test/), all Go GPU code
in pkg/cublaslt with //go:build gpu on every file and no CPU stub; callers
split *_gpu.go / *_cpu.go like the cuVS table functions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#20567)

Cell = 12-byte header (version, format, reserved, dim, fp32 global g) +
one scale byte per block + packed elements.
- vecf8 (MXFP8): e4m3 elements, E8M0 scale per 32, g = 1.
- vecf4 (NVFP4): e2m1 elements (element 2i in the low nibble), unsigned
  E4M3 scale per 16, per-vector fp32 global g = amax/2688, so any finite
  float32 range fits and block scales stay out of the E4M3 subnormal range.

AppendBlockScaled / ParseBlockScaledCell / Dequantize / string round trip;
the parser rejects malformed cells (length, version, format, reserved and
padding bits, non-finite or negative g, NaN or signed scales, NaN
elements). Scale selection is tested minimal at every E8M0/E4M3 boundary;
nibble order is checked against CUDA's __nv_cvt_float2_to_fp4x2.

Measured with this codec on unit-normalized wiki_all 768-d (50k/200):
recall@10 vecf8 0.971, vecf4 0.922. Design doc updated to the two-level
vecf4 format.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rixorigin#20567)

T_array_float8 (230, vecf8) and T_array_float4 (231, vecf4) are varlena
vector types with Width = dimension. They are deliberately NOT in
IsArrayRelate(), which keeps meaning "Width elements of
GetArrayElementSize bytes": vecf8/vecf4 cells are block-scaled
(vecblock.go) and paths that decode typed slices or call
GetArrayElementSize must not accept them.

- IsBlockScaledVector(): vecf8/vecf4.
- IsVectorType(): IsArrayRelate() || IsBlockScaledVector(), for structural
  checks that apply to every vector column (PK/partition/index rejection,
  dimension guards).
- BlockScaledFormat(): oid -> cell format.
- Type.ArrayCellBytes(): cell byte length for any vector type, the
  replacement for Width*GetArrayElementSize() size estimates.
- Names (VECF8/VECF4, vecf8/vecf4), DescString, OidString, TypeLen,
  FixedLength, ToType, Types map, DecodeValue/EncodeValue (raw bytes).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Parser: VECF8/VECF4 tokens (non-reserved keywords) and ArrayFamily type
  rules; mysql_sql.go regenerated (no conflicts); tree.T prints the width.
- Plan: getTypeFromAst maps every vector SQL name through the new
  types.VectorTypeBySQLName (replaces three hand-maintained six-name lists),
  so vecf8/vecf4 get the default width and the 1..65535 dimension check.
- SHOW CREATE prints VECF8(N)/VECF4(N).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ithmetic/SUM/AVG/inner_product, Parquet LOAD (matrixorigin#20567)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
cpegeric and others added 6 commits October 5, 2026 09:15
…rounding vecf4 quantizer (matrixorigin#20567)

- IF/IFF/CASE/COALESCE keep one bf16/float16/float8/float4/vecf8/vecf4 type,
  so multi-table UPDATE and other DML projections write cells of the column
  width (they stored float32 bytes before).
- != (NOT IN), <=> and BETWEEN round in-range literals to the column type as
  = and IN do; filter simplification no longer keys pre-rounding literals.
- A text literal compared with vecf8/vecf4 is cast to the column dimension.
- types.CompareValue handles the low-precision scalars (data branch diff/merge).
- Data branch SQL text renders dequantized vecf8/vecf4 values.
- vector_matmul scores on the CPU when the GPU engine cannot be created.
- Prepared narrowing accepts integer user variables and decimal-numeral text
  only; IF condition, INTERVAL, CONV and BIT_AND/OR/XOR take low-precision args.
- vecf8/vecf4 comparison decodes without per-row allocation.
- vecf4/vecf8 elements are the exact quotient rounded once (no float32
  intermediate); the design doc states the quantizer's scale choice.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…allback with a GPU (matrixorigin#20567)

- Casts from text to bf16/float16/float8/float4 and to vecf8/vecf4, and from
  vecf8/vecf4 to vecf32, leave rows outside the select list NULL without
  converting them, as the float casts do (a CASE branch no longer errors on an
  unselected invalid value).
- vector_matmul scores on the CPU only with gpu_mode off or no device of
  compute capability 10.0+. With an eligible device enabled, an allocation
  account denial or any engine error fails the query.
- The engine looks up a cuBLASLt algorithm for every tile row bucket
  (128 * 2^i up to the tile capacity) at construction, so a shape the device
  cannot run fails before scoring; tiles run at the cached algorithms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	pkg/defines/const.go
#	pkg/sql/compile/ip_functions_protocol.go
#	pkg/sql/compile/remoterun.go
#	pkg/sql/plan/base_binder.go
#	pkg/sql/plan/function/func_testcase.go
#	pkg/sql/plan/function/operatorSet.go
#	pkg/sql/plan/prepared_binding_test.go
A BLOB of little-endian float32 elements casts to vecf8/vecf4 and quantizes
as the same vecf32 value, matching the binary vector input of vecf32. A
length not a multiple of 4, another dimension or a non-finite value is
rejected; an empty BLOB is NULL. The parity test also covers BLOB shapes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
matrixorigin#20567)

The queries argument may be a BLOB holding nq x N float32 values back to back
(CAST(? AS BLOB) for a client's bytes); text stays a JSON array. The compiler
records the format in the configuration flags; the length must be a non-zero
multiple of 4*N and the values finite. Both forms configure the same queries.
The BLOB-query BVT checks the CPU kernel in cases/ (gpu_mode = 0) and the
cuBLASLt result in gpu_cases/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
matrixorigin#20567)

INSERT/REPLACE VALUES and DEFAULT bound every numeric, hex or bit literal with
the target column type, also inside a CAST, so CAST(X'...' AS BLOB) into a
vector column became CAST(<vector> AS BLOB) and failed (the form a client's
bytes parameter takes). A vector column type is no longer that binding type;
other column types keep it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@qodo-code-review

Copy link
Copy Markdown

Qodo reviews are paused for this user.

Troubleshooting steps vary by plan Learn more →

On a Teams plan?
Reviews resume once this user has a paid seat and their Git account is linked in Qodo.
Link Git account →

Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center?
These require an Enterprise plan - Contact us
Contact us →

@XuPeng-SH XuPeng-SH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decision: REQUEST_CHANGES — value contracts are still not closed

Reviewed d4822f2e307cfc6c76e9bf3792f5714f5af24e4b against base/merge-base 43c37847b16d37cfa8b9eef4235ba4388b5e25ce. The architecture has useful reuse: codecs belong in types, top-k partial/merge/spill use the existing aggregate framework, and GPU execution stays behind the engine boundary. The three inline findings are concrete correctness defects in the new value contracts, not requests for additional compatibility machinery.

Confirmed current-head failures

  1. P1: NVFP4 replication/reload changes already-stored values. A 17-element vector with only elements 0 and 16 nonzero (8.7649145, 5.7432985) stores the latter as 6.2606535. Production CDC emits the decoded text; reloading it stores 6.8867183, approximately 10% additional drift. Both the production CDC extraction/SQL formatter and a fresh single-CN MySQL-protocol insert/select/reinsert reproduce this. This is additional loss during transport, not the permitted initial quantization.
  2. P2: equality disagrees with grouping/distinct. Real SQL says CAST('[1.2031566]' AS vecf4(1)) = CAST('[1.2031565]' AS vecf4(1)) is true, but storing those values produces 2 GROUP BY groups and 2 DISTINCT values. MXFP8 also fails: [447] and [449] both decode to [448], compare equal, and still produce 2 groups. A same-encoding vecf8 neighbour returns 1 group as expected. Deterministic encoding of one input does not establish a canonical encoding for all equal decoded values.
  3. P2: ARM64 fused decoding breaks distance invariants. With normal Go 1.26.4 compilation, a valid vecf4(16) containing sixteen 9.9999994e29 values gives VecBlockL2DistanceSq(v,v) = +Inf; real SQL l2_distance_sq(v,v) returns error 20101. Casting that same stored vector to vecf32 first returns 0. Smaller values also have nonzero self-distance. The new TestVecBlockKernelsSymmetric fails deterministically (3/3); package-local -d=fmahash=n makes both symmetry and the self-distance probe pass. This is a production decode/rounding boundary defect, not merely an overstrict test.

Resolve these at the existing owners: preserve the stored cell/value across transport; extend the shared equality-key contract rather than adding an alternate grouping path; make the kernel generator's decode rounding match the stored-value contract rather than adding self-distance special cases or globally disabling FMA.

Historical review disposition

  • My three old inline threads are now resolved/outdated. The current encoder/parser rejects decoded infinity; host staging/tile admission uses the aggregate account; device selection filters the declared compute-capability baseline. Current finiteness and fake-engine admission/cleanup tests pass.
  • iamlinjunhong's select-list finding is fixed; TestNarrowCastHonorsSelectList passes.
  • The previous NVFP4 K=48/no-algorithm concern is addressed in source by padding K to 32 and preflighting tile algorithms, together with an explicit eligible GPU => execute or fail decision. That decision can be reasonable; it is different from a CPU-fallback promise. The design still says fallback on memory denial at lines 269 and 524, contradicting its final no-fallback decision and implementation. Reconcile those paragraphs and the stale draft/merge references in the PR body. No real GPU execution is claimed here.
  • fengttt's bitwise-format question has an author answer distinguishing CUDA-compatible element/nibble encoding from the custom block/global quantizer. These are not the same assertion; keep that distinction in the final contract. The current CPU types suite, including checked-in CUDA golden verification, passes.

QA, test quality and scope

Fresh CPU native build succeeded. Relevant types/CDC/aggregate/function tests passed; explicit skipped-row CAST passed; the aggregate owning package passed once under -race, and TestVectorMatmulConfigShared passed 100/100 race repetitions. Full types/compare/sort suites passed. The metric suite did not pass, for the symmetry failure above. Reviewer-only overlays and logs were kept outside the clean review worktree; no production code was changed and no subagent was used.

The existing text round-trip check only proves that reparsing succeeds, not that the stored value is preserved. CDC/ISCP fixtures use exact small values, so they miss scale-boundary drift. Add small orthogonal oracles for transport identity, equality-versus-key agreement in both formats, and decoded self-distance with FMA enabled. Do not fix this by deleting the failing symmetry case, widening every tolerance, or piling up repetitions of the exact-grid fixtures.

The 188-file change was mapped by responsibility, with implementation, tests, documentation and generated code considered separately. The mandatory design gate remains blocked by these counterexamples: this is not a blanket acceptance of every remaining hunk, unhappy path, GPU performance claim or device configuration. GPU benchmarks/real-device validation were not rerun on this CPU-only host. Stop-the-cluster upgrades are the scope; mixed-version rolling compatibility is not a blocker, but persisted-value and replication correctness remain required.

Comment thread pkg/cdc/util.go Outdated
Comment thread pkg/sql/plan/function/func_compare.go
Comment thread pkg/vectorindex/metric/distance_func_vecblock_gen_test.go Outdated
cpegeric and others added 6 commits October 5, 2026 13:24
matrixorigin#20567)

options {"metric": "inner_product" | "cosine" | "l2sq"} selects the distance;
the scores are the values of inner_product() (-dot), cosine_distance() (1 with
a zero vector) and l2_distance_sq(), nearest first. Other option keys are
ignored; text that is not a JSON object, or another metric, is an error.

GPU: a row-statistics kernel computes each packed row's squared norm (and the
uint8 element sums, moved off the host) from the values the matmul multiplies;
the fix-up kernel applies the metric, and run uses it too, so the host does no
score arithmetic. The CPU engine computes the same formula with x.x norms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ls (matrixorigin#20567)

The generated kernels decoded an element as an unrounded product (code x
scale) that the compiler could fuse into the following subtract on targets
with fused multiply-add (arm64), so x - y rounded only one side: a vector had
a nonzero self-distance, and +Inf near the float32 limit. Each decoded element
is now an explicit float32 conversion, which rounds it as the stored value is
and prevents the fusion; only the accumulations may still fuse. On arm64 the
L1 kernels drop from 16 fused instructions to 0 and the L2 kernels from 32 to
16 (the d*d accumulation). A self-distance test covers values near the float32
limit and several magnitudes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ixorigin#20567)

Two cells can encode equal values with other block or global scales (vecf8
[447] and [449] both store 448), so = held while GROUP BY, DISTINCT, joins and
set operations keyed the raw cell bytes and kept them apart. The shared
equality-key contract (keycodec AppendCanonicalValue, CanonicalValuesEqual,
CanonicalValueSize and the spill hash) now keys vecf8/vecf4 by their decoded
float32 values, signed zero closed, and every site that routes the other
vector types through it routes vecf8/vecf4 too: the hashmap group/join keys,
COUNT(DISTINCT) and its capacity preflight, approx_count_distinct and its
protocol gate. Runtime filters stay off for vector types and shuffle only
takes integer and string keys. A key-versus-equality oracle covers both
formats; SAMPLE also accepts vecf8/vecf4.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…it (matrixorigin#20567)

CDC/ISCP replication and data branch merge wrote vecf8/vecf4 values as their
decoded text, which the replica quantized again; a vecf4 global scale follows
the decoded maximum, so a stored 6.2606535 replayed as 6.8867183.

The exact text is the cell as stored:
  {"g": g, "b": [{"s": scale, "v": [element values]}, ...]}
("g" only for vecf4; vecf8's global is 1). Every scale and element must be
the value of its code, 32/16 per block and fewer only in the last, so parsing
rebuilds the same cell bytes; anything else is an error, never rounded.
vecblock_json(v) returns it; casting, inserting or LOADing it builds the
cell as written (types.StringToBlockScaled is the one text entry point).
CDC, ISCP and data branch SQL now carry this text; query output and
INTO OUTFILE keep the decoded values. The design doc also drops two stale
CPU-fallback statements.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…matrixorigin#20567)

encoding/json decoded 1000x768 query JSON at 88.5 ms; sonic at 25.8 ms.
Unknown keys and non-numeric values are still rejected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…trixorigin#20567)

INTO OUTFILE and external-table writes emit vecf8/vecf4 cells as the
vecblock JSON text (a quoted CSV field, a JSONL object value), so an
export reloads to the same cells instead of being quantized again.
JSONL LOAD passes an object value of a vecf8/vecf4 column through as
its JSON text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cpegeric

cpegeric commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

fixed

cpegeric and others added 12 commits October 5, 2026 15:37
…norm leaves range (matrixorigin#20567)

A unit's float32 squared norm that overflowed to +Inf or underflowed to 0
gave cosine distance 1 instead of 0, depending on the dimension. Recompute
in float64 as the vecf32 kernels do. A squared L2 beyond the float32 range
is +Inf at every dimension, as for vecf32.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gin#20567)

ORDER BY ties and window PARTITION BY / ORDER BY peers compared cell bytes,
so cells holding the same values with other block scales were distinct
peers while = and GROUP BY treated them as equal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…loat8/float4 (matrixorigin#20567)

e = '0.3' compared in DOUBLE and missed the stored value, as did BETWEEN and
multi-item IN with text items. A prepared IN list with one value outside the
type's range lost the rounding of the others to the list-level numeric
prefix pass; the list now expands to per-item comparisons that bind each
marker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…; Parquet LOAD reads it (matrixorigin#20567)

MODUMP QUERY_RESULT and the checkpointtool SQL-LOAD dump wrote vecf8/vecf4
as decoded values, which reload quantized again (vecf4 values move). Both
now write the vecblock JSON. Parquet string columns accept it as CSV and
JSONL do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rixorigin#20567)

A failed engine creation marked the engine as tried, so a batch retried on
the same executor after a capacity error scored on the CPU. The mark is set
only on success or when no device is in play, and cleared by Free.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… in float64 (matrixorigin#20567)

Follows 8a1ebd1: inner product still errors, cosine distance is 1 and
similarity 0 for the orthogonal pair, as for vecf32.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n#20567)

SET @v from a vecf8/vecf4 value decoded it to vecf32, so writing @v back
quantized it again and vecf4 values could move. The variable now holds the
cell (types.BlockScaledValue) with the column type: it rebuilds the same
bytes, displays the decoded values and survives connection migration.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/float4 (matrixorigin#20567)

A text, float64, decimal or wide integer source was rounded to float32 and
then to the narrow type, so a value just above a tie landed one step low
(cast('1.0039062500001' as bf16) was 1, not 1.0078125). Float32RoundToOdd
keeps the sticky bit, so the narrow rounding is the single correct one, in
casts, INSERT values, CSV and Parquet LOAD.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cell casts back exactly (matrixorigin#20567)

The exact binary form of a vecf8/vecf4 value is its cell (12-byte header,
block scales, element codes). A BLOB of a valid cell of the target casts,
inserts or binds to the same bytes without quantization; its length is
never 4*dim, so a float32 BLOB still quantizes. A BLOB of cell length that
is not a valid cell is an error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…origin#20567)

A Parquet BYTE_ARRAY / FIXED_LEN_BYTE_ARRAY column without a logical type
is binary for a vecf8/vecf4 target, read as a BLOB casts: a value of cell
length is the stored cell, a value of 4*dim bytes is float32 elements.
STRING/JSON columns stay text. The BLOB rule is one helper,
types.BlockScaledFromBinary, shared by the cast and the loader, and a value
of neither length names both. Adds the BVT resource parquet/vecblock.parquet,
written and checked by TestVecBlockParquetResource.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…atrixorigin#20567)

The exact text lists the parts in the cell's order: the global scale (always
written, 1 for vecf8), one scale per block, and the flat element values, so
element i is v[i] * s[i/blockSize] * g and NumPy decodes it in one line.
Smaller than the per-block form, which is no longer accepted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/vecf4 cells (matrixorigin#20567)

With {"query_format":"vecblock"} in the options, the queries of a vecf8/vecf4
column are cells used as given, on the CPU and the GPU: a BLOB of cells back
to back (vecblock_binary) or a JSON array of vecblock JSON objects. A query
taken from a stored vecf4 row is then at distance 0 from it; its decoded
float values quantize to another cell. float32, the default, keeps the
float-value queries.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cpegeric

cpegeric commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the review. All three are fixed, each with a test that fails without the fix.

P1: NVFP4 replication drift. Fixed by carrying the stored cell instead of its decoded values.

  • New exact text form, vecblock_json(v): {"g":g,"s":[block scales],"v":[element values]} (904b890, shape finalized in
    c31f5d0).
    • Every number is a code value of its format: E8M0/UE4M3 scales, E4M3/E2M1 elements.
    • Parsing it rebuilds the same cell bytes with no quantization, and anything that isn't an exact code value is rejected, never
      rounded.
    • CAST, INSERT and LOAD all accept it.
  • Transport now uses it. CDC and ISCP replication SQL, data branch merge, INTO OUTFILE / external writes (CSV and JSONL), MODUMP
    QUERY_RESULT and the checkpointtool dump all emit the exact form (904b890, b03c4e3, 6794455). Your 17-element vector
    replays bit for bit, and element 16 stays at 6.2606535.
  • Exact binary form too, vecblock_binary(v) (836a8ca): a BLOB of a valid cell casts back to the same bytes. Parquet binary
    columns load cells exactly (4ef9ad3), and vector_matmul accepts exact query cells via "query_format":"vecblock"
    (10550e2).
  • Query output still shows the decoded values.
  • Tests: CDC/ISCP/data-branch identity tests with your vector, TestBlockScaledJSONIdentity, and BVT export→LOAD and Parquet round
    trips.

P2: equality vs GROUP BY / DISTINCT. Fixed: every equality consumer now agrees with =.

  • Hash keys (d0064a5): the vecf8/vecf4 key in keycodec is now the decoded float32 values with signed zero unified.
    • This covers GROUP BY, DISTINCT, hash joins, set operations, IN (subquery), COUNT(DISTINCT) and approx_count_distinct.
    • So cells that encode the same values with different scales are one key: vecf8 [447] and [449] both store 448.
    • SAMPLE on vecf8 is fixed as well.
  • Peer groups (5f5fb57): ORDER BY ties and window PARTITION BY / ORDER BY now compare decoded values instead of raw bytes.
  • Tests: TestCanonicalVecBlockKeyMatchesEquality (keys agree with CompareBlockScaledFromBytes), TestBlockScaledPartition, and
    vecblock BVT cases for GROUP BY, DISTINCT, joins, UNION, IN, rank() OVER and PARTITION BY on byte-different, value-equal cells.

P3: arm64 FMA decode. Fixed in 76771e7.

  • Each decoded element is now rounded once to float32 with an explicit conversion (float32(code * scale)), which Go never fuses.
    So the kernels decode exactly as At() and the equality key do, on every architecture.
  • Kernel generator updated and the kernels regenerated.
  • Proof: cross-compiling for arm64, the fused multiply-adds in the decode paths are gone; FMADDS/FMSUBS fell from 560 to 400
    across the kernels. The ones left are in the accumulation, where fusing is the same as for vecf32.
  • Test: TestVecBlockSelfDistance checks that a stored vector is at distance 0 from itself, including values near the float32
    maximum (9.9999994e29).

A self-review on top of these also fixed:

  • Cosine: a float32 norm overflow or underflow is now recomputed in float64, matching vecf32 (8a1ebd1).
  • bf16/float16/float8/float4 text and comparisons:
    • text and float64 values round once (81b54d1);
    • numeral text literals and prepared IN lists round to these types (5c37d1a).
  • User variables keep vecf8/vecf4 cells (7083782).

cpegeric and others added 2 commits October 5, 2026 18:05
… on the same bytes (matrixorigin#20567)

NVFP4 and MXFP8 operands drawn in NVIDIA's layout with one global scale per
operand are scored by a direct cuBLASLt call (alpha = G_x * G_q) and, as MO
cells carrying that global on every row, by the engine. Scores agree to the
float32 rounding of the two runs and the top-10 rows match; at the engine's
tile shape MXFP8 is bit-identical and NVFP4 within 1 ulp (the globals are
applied in double after the GEMM, not as an fp32 alpha).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… alpha (matrixorigin#20567)

The fix-up kernel computes the inner product as acc * fp32(g_row * g_query)
in fp32, as cuBLASLt applies alpha = G_a * G_b with per-tensor global
scales, so NVFP4 scores of rows sharing one global equal a direct NVIDIA
block-scaled GEMM on the same bytes bit for bit (previously within 1 ulp).
Cosine and squared L2 keep the double product, so a row equal to a query
stays at distance 0. MatchesNvidiaBlockScaledGemm now asserts bit identity
at the engine's GEMM shape.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kind/documentation Improvements or additions to documentation kind/enhancement kind/feature size/XXL Denotes a PR that changes 2000+ lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants