feat: low-precision float types bf16/float16/float8/float4 and vecf8/vecf4 vector columns (#20567) - #29554
feat: low-precision float types bf16/float16/float8/float4 and vecf8/vecf4 vector columns (#20567)#29554cpegeric wants to merge 88 commits into
Conversation
…rixorigin#20567) Add the low-precision float value types for the new scalar bf16/float16/float8/ float4 SQL columns. Float8 is OCP/NVIDIA FP8 e4m3 (1-byte, no Inf, max 448); Float4 is OCP MXFP4 e2m1 (1-byte, magnitudes {0,.5,1,1.5,2,3,4,6}, NVIDIA-bit-compatible). Both convert to/from float32 (round-to-nearest-even, saturating overflow) so all arithmetic runs on the existing float32 kernels. Reference-value + full round-trip unit tests included. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rixorigin#20567) Add T_bf16=73, T_float16=74, T_float8=75, T_float4=76 to the type system: SQL names, sizes (2/2/1/1 bytes), FixedLength/TypeLen, ToType, String/OidString, and IsFloat classification. They satisfy FixedSizeT already via ~uint16/~uint8, so the vector container generics accept them without constraint changes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…origin#20567) Add BF16/FLOAT16/FLOAT8/FLOAT4 keyword tokens and no-length column_type grammar rules (float4/float8 were reserved-but-UNUSED), regenerate the mysql parser, and map them in getTypeFromAst (MYSQL_TYPE_FLOAT disambiguated by FamilyString) to T_bf16/T_float16/T_float8/T_float4. Fixes the issue's 1064 on CREATE TABLE. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rt (matrixorigin#20567) Start Layer 4 (vector container): GetAny dispatches the four new scalar float type IDs to their value types. Remaining container switches (Append/Shrink/Shuffle/ String/min-max/compare/sort) still to do -- sort/min-max must compare by float value, not raw bits. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tches (matrixorigin#20567) Add the four scalar float type IDs to the value-access (GetAny/AppendAny/String/ RowToString) and movement (Shrink/ShrinkByMask/Shuffle/ShuffleWithBuf) container switches. Union/const-set and the compare/sort/min-max/sum switches (which must order by float value, not raw bits) still to do. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…oat4 (matrixorigin#20567) Contract, format decisions (FP8 e4m3, FP4 e2m1/MXFP4), and invariants (float32 bridge arithmetic, value-order comparison, saturating RNE narrowing). sparsevec deferred to a separate design. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…float4 (matrixorigin#20567) The four new scalar float types are byte-identical to uint16 (bf16/float16) and uint8 (float8/float4) for the value-agnostic union-all and const-set copy paths, so add them to those case labels. Remaining container work: compare/sort/min-max/sum must order by float value. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…float16/float8/float4 (matrixorigin#20567) The container ordering and aggregation paths fell through to raw-bit handling for the low-precision float types, which is wrong: a float's sign bit makes raw uint order disagree with value order (negatives sort as large positives). Route them through ToFloat32: - compareVectorRows: add bf16/float16/float8/float4 via new compareFloatRows. - InplaceSort / InplaceSortAndCompact: sort (and value-dedup) via ToFloat32; add the four types to supportsInplaceSort. - GetSumValue / GetMinMaxValue: widen via new LowPrecFloatGetSum / LowPrecFloatGetMinAndMax; min/max skips NaN like the float32/64 path. - typeCompatible: recognize the four named types for type-checked column access (needed by appendList in the compact path). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…32 bridge (matrixorigin#20567) Casts to/from the low-precision float types bridge through float32: any source widens to float32 then rounds to the target format via FromFloat32, and any low-precision source widens to float32 then reuses the float32 target machinery (inheriting exact integer-rounding, decimal, and string-formatting semantics). - func_cast_lowprec_float.go: castToLowPrecFloat (X->newtype), lowPrecFloatToOthers (newtype->X), string parse, decimal256 routed to castToDecimal256; init() wires the supportedTypeCast gate for all numeric/string sources and targets. - func_cast.go: route low-precision source/target ahead of the integer/decimal fast-paths so a low-precision operand is never misdispatched. - functionTools.go: NewFunctionResultWrapper builds result vectors for the 4 types (was panicking). - func_testcase.go: test framework builds and compares the 4 types. - func_cast_test.go: TestCastLowPrecFloat covers float64<->each, a cross-cast, and string<->float8. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…16/float16/float8/float4 (matrixorigin#20567) Binary operators and aggregates treat the low-precision float types as float32, which holds every bf16/float16/float8/float4 value losslessly: - type_check.go: fixedTypeCastRule1 (arithmetic + comparison) and fixedTypeCastRule2 (div) normalize a low-precision operand to float32 before selecting the coercion rule, mirroring the existing T_enum->uint16 treatment. float32 has no diagonal cast rule (equal types need none), so a low-precision pair that both normalize to float32 forces the cast to float32 explicitly. - list_agg.go: sumAvgTypeCheck and a new minMaxTypeCheck widen a low-precision input to float32 (SUM/AVG return float64; MIN/MAX return float32) -- newGenericMinMaxExec compares raw uint bits, which do not order floats, so native support is unsafe. The operand cast itself uses the float32 bridge added for CAST. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t8/float4 (matrixorigin#20567) The storage/zonemap comparators must order these types by float VALUE, not the raw uint bits (a negative float's bits exceed a positive's), and several switches panicked on the unregistered types: - types.EncodeValue/DecodeValue: register the four types (were panicking in the zonemap value path, e.g. UpdateZMAny). - compute.Compare/CompareGeneric: compare via ToFloat32 (drives zonemap min/max maintenance in UpdateZM and Contains/containsBytes pruning). - zm.getValue: decode the four types (MO_TABLE_COL_MAX read path panicked otherwise). - zm.SubVecIn: value-aware bound search on a value-sorted column (subVecInLowPrecFloat); BatchUpdateZM already feeds value-correct min/max from vector.GetMinMaxValue. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…oat8/float4 (matrixorigin#20567) Result-set output and CSV/LOAD parsing for the low-precision float types, presenting them as float32 (they widen losslessly): - frontend: column type maps to MYSQL_TYPE_FLOAT; the colSlices path widens the column into arrFloat32 at build time (lowPrecFloatVecToFloat32Slice) so GetFloat32/GetFloat64 and the row extractors read it as float32; the legacy any-row and ExtractRowFromEveryVector paths return .ToFloat32(). - external LOAD: parse the field as float32 then round to the target format; the empty zero-fill and the row-validity predicate accept the four types. Strict finite/range checking on the string-parse boundary (LOAD + CAST-from-string): types.RejectNonFiniteNarrowFloat rejects NaN/Inf and out-of-range values instead of silently saturating, mirroring the narrow-vector element parser (rejectNonFiniteArrayElem). bf16/float16 overflow is detected as Inf after narrowing; float8/float4 saturate, so their input magnitude is range-checked (448 / 6). Numeric CAST keeps the saturating constructor, matching the documented vecint8 precedent (string parse strict, CAST clamps). Tests: RejectNonFiniteNarrowFloat, Inf conversion (bf16/float16 preserve, float8/float4 saturate), subnormals for all four, and the frontend colSlices output path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…f coverage to ~87% Add unit tests for the new-code paths that the feature tests did not reach, so the PR's aggregate diff coverage clears the 75% gate: - vector: GetAny/AppendAny/String/RowToString/Shrink/ShrinkByMask/Shuffle/ShuffleWithBuf and NewFunctionResultWrapper for all four types. - types: type-registration switches (TypeLen/FixedLength/String/OidString/IsFloat/ToType), EncodeValue/DecodeValue, and the slice converters. - compute: Compare and CompareGeneric order by float value (negative discriminator). - tae/index: fix TestZonemapLowPrecFloat to use >3 rows so SubVecIn reaches the per-type bound search instead of the small-vector early return. - external: getColData LOAD parse (valid rounds, empty zero-fills, out-of-range/non-finite rejected) and appendLoadEmptyNumericZero. - frontend: getValueFromVector, extractRowFromVector (legacy any-row), and convertEngineTypeToMysqlType present the types as FLOAT. - plan: getTypeFromAst disambiguates the four via FamilyString. - function: broaden the cast matrix across source types -> bf16 and float8 -> targets. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…float16/float8/float4; add BVT BVT (test/distributed/cases/dtype/lowprec_float) surfaced two gaps beyond the earlier per-layer work: - pkg/sort: ORDER BY did not sort these columns -- IsSupportedType and the comparator switch omitted them, and the sortType constraint lacked the scalar slice types, so the order operator treated them as equal (stable -> insertion order). Add the four types to the constraint, IsSupportedType, and the switch with value-aware comparators (widen via ToFloat32, reusing the float32 SQL NaN-aware ordering). - function: inserting a negative literal (e.g. -2.0) into such a column failed with 'unary_minus bad value [BF16]'. unaryMinusMatch and UNARY_PLUS (new unaryPlusMatch) now cast a low-precision operand to float32, mirroring the binary-operator coercion. BVT covers DDL, insert (incl. negatives/NULL), value-ordered ORDER BY, comparison, arithmetic, SUM/AVG/MIN/MAX, CAST both ways, out-of-range reject, and INSERT..SELECT round-trip; validated 27/27 on a live cluster. Unit tests: TestSortLowPrecFloat and unary coverage in the ops test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts: # pkg/sql/plan/function/list_agg.go # pkg/sql/plan/function/type_check.go
…256 cast, non-finite persistence Three defects found by the self-review gate: - types/float8.go: Float8FromFloat32 subnormal->normal round-up silently returned ZERO. roundMantissaRNE masks its result and signals the boundary overflow only via its carry, but the subnormal branch dropped the carry (mant, _ :=) and its mant>=8 guard was therefore dead. Every value in (0.013671875, 0.015625) narrowed to 0 instead of the smallest normal 0x08 -- silent data loss. Honor the carry. - function/func_cast_lowprec_float.go: decimal256 was registered in supportedTypeCast but anyToLowPrecFloat lacked the case, so a planned CAST(decimal256 AS bf16/...) failed at execution with an internal error. Add the decimal256 case (types.Decimal256ToFloat64). - function/func_cast_lowprec_float.go: the numeric cast path applied no finite check, so an overflowing float64/decimal source persisted +Inf for bf16/float16 (and NaN for float8), violating the repo-wide never-persist-non-finite invariant (matrixorigin#29084) and diverging from the string path. Route the numeric path through opUnaryFixedToFixedWithErrorCheck + RejectNonFiniteNarrowFloat so NaN/Inf/out-of-range errors uniformly, matching the string cast and matrixorigin#29084's rejectNonFiniteVectorElems. - tae/index/zm.go: hasNaNBound now covers bf16/float16/float8 for defense-in-depth parity with float32/float64 (float4 has no NaN encoding). Latent today since GetMinMaxValue skips NaN, but keeps the fail-open guard consistent. Regression tests: TestFloat8SubnormalRoundUpToNormal, TestCastLowPrecFloatDecimal256AndNonFinite. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…nstraint types.LowPrecFloat
Replace the inline interface{ ToFloat32() float32 } (and the multi-line
FixedSizeTExceptStrType + ToFloat32 form) repeated across the ordering, compare,
sum/min-max, sort, zonemap, frontend, and cast helpers with a single named constraint
types.LowPrecFloat = { FixedSizeTExceptStrType; ToFloat32() float32 }. The union+method
together admit exactly bf16/float16/float8/float4, so the fixed-size storage bound
(DecodeFixed / column accessors) and the value-widening method live in one place.
No behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ecimal256 cases
Cover the self-review fixes end-to-end on a live cluster (47/47):
- out-of-range NUMERIC cast rejected (not just string): CAST(7 AS float4),
CAST(1000 AS float8), CAST(70000 AS float16), CAST(1e300 AS bf16).
- Inf/NaN cast rejected: CAST('inf'/'nan'/'-inf' ...), and CAST(1e308*100 AS bf16)
(arithmetic overflow to Inf).
- arithmetic result overflowing the narrow range rejected on cast-back:
CAST(6*2 AS float4), CAST(a+b AS float4) with 3+4>6.
- FP8 subnormal round-up regression: CAST(0.015 AS float8) = 0.015625 (was silently
zero); min-subnormal round-trip; underflow to 0; float4 0.5 subnormal.
- decimal256 -> bf16/float8.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…trixorigin#20567) Design for packed low-precision float vector columns for quantized ML embeddings: per-32-block E8M0 (MXFP) scaling, one shared byte-exact-MXFP payload format for both e4m3 (vecf8) and e2m1 (vecf4) elements, and an MO-only header stripped so the GPU path is a straight memcpy to cuVS/CUTLASS MXFP kernels (no dequant/requant). Contract, decisions, and invariants only; depends on the scalar float8/float4 element codecs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… P4 dependency (matrixorigin#20567) - SQL spelling is vecf8/vecf4 to match the existing family (vecf32/vecf64/vecf16/vecbf16), not vecfloat8/vecfloat4 (there is no vecfloat16 -- it is vecf16). - Note that MO's current cuVS integration supports only float/half/int8/uint8 with no MXFP ingestion, and the cuVS MXFP contract is not vendored (GPU-box only). CPU storage/distance can proceed now; the GPU memcpy path is a forward-looking target pending a cuVS MXFP API. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ot cuVS (matrixorigin#20567) Rewrite the GPU section: fp8/fp4 are tensor-core GEMM formats, so vecf8/vecf4's GPU value is brute-force batch similarity (Q x D^T) via cublasLtMatmul block-scaled matmul, NOT a cuVS index (cuVS accepts only float/half/int8/uint8 + SQ/PQ/binary; no fp8/fp4 path). The stored 32-block E8M0 MXFP payload is exactly cuBLASLt's CUDA_R_UE8 layout, so transfer is strip-header + memcpy into the operand/scale tensors. Records hardware (Blackwell sm_100 / consumer sm_120 incl. RTX 5070), toolchain (CUDA 12.8+/cuBLAS 12.9+), mandatory dataset tiling, MXFP4-vs-NVFP4 fp4 choice, and a P1-P3 CPU / P4 GPU phasing. New cuBLASLt cgo integration, separate from cgo/cuvs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…matrixorigin#20567) Lock the two open decisions: - vecf4 = NVFP4 (e2m1 4-bit elements + E4M3 unsigned 16-block scale, per-vector-normalized, GEMM global=1.0, no stored global). Dictated by cuBLASLt: its FP4 block-scaled matmul consumes NVFP4 only (16-block E4M3), not MXFP4 (32-block E8M0), so storing MXFP4 would force a lossy re-quantization at GEMM and break the strip-header+memcpy contract. NVFP4 is also the more accurate 4-bit format. vecf8 = MXFP8 (e4m3 + E8M0 32-block), which is what cuBLASLt's FP8 path takes directly. - CPU distance accumulate = fp32 (matches the GEMM's fp32 accumulate so CPU==GPU results). Also records: element codecs are the merged scalar Float8/Float4 (bit-identical to CUDA_R_8F_E4M3 / CUDA_R_4F_E2M1) so packing is a pure shuffle; only block-scale codecs are new (E8M0, and reuse Float8 e4m3 for NVFP4's scale); effective bits/element and 1024-dim storage sizes; the per-vector-independent invariant; and the P4 GPU-box verification items. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…l SQL surface (matrixorigin#20567) Verified on RTX 5070 (sm_120), CUDA 13.3, cuBLASLt 13.6: - types.Float8/Float4 match cuda_fp8.h/cuda_fp4.h bit for bit on all codes and 4.3M encode inputs (NaN encoding differs; never stored); FP4 element 2i is the low nibble. - Block scales must use cuBLASLt's tiled 128x4 layout; per-row scales give wrong results. Elements transfer as stored, scales are re-laid out. - Dataset must be operand A: queries-as-A with one query returns silently wrong results. K is padded to a multiple of 32. - MXFP8 and NVFP4 (global 1.0) match a CPU reference; recall@10 on unit-normalized wiki_all 768-d: MXFP8 0.971, NVFP4 0.918. Design: cgo/cublaslt dot-product matmul engine; vector_matmul TVF (one partial per scan pipeline) + vector_matmul_merge aggregate with a JSON result format; single-CN and multi-CN test plan. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…origin#20567) Each hit is a two-element array: position 0 the source key as a JSON string, position 1 the dot product. It halves the result size versus {"id","score"} objects. Unnest reads '$[0]' / '$[1]'; verified on the current build, including a 9223372036854775807 id cast back to bigint. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/cuda_fp4.h Golden data generated on CUDA 13.3: all 256 e4m3 and 16 e2m1 decodes and 1,069 encode inputs (every rounding midpoint +/-1 ulp, subnormals, saturation, +/-Inf). NaN encoding is excluded (documented difference). A codec change that breaks GPU bit compatibility now fails a CPU-only UT. The generator is cgo/cublaslt/test/narrow_float_golden_gen.cu, with its regeneration command. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
matrixorigin#20567) All C++/CUDA code lives in cgo/cublaslt (engine + test/), all Go GPU code in pkg/cublaslt with //go:build gpu on every file and no CPU stub; callers split *_gpu.go / *_cpu.go like the cuVS table functions. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#20567) Cell = 12-byte header (version, format, reserved, dim, fp32 global g) + one scale byte per block + packed elements. - vecf8 (MXFP8): e4m3 elements, E8M0 scale per 32, g = 1. - vecf4 (NVFP4): e2m1 elements (element 2i in the low nibble), unsigned E4M3 scale per 16, per-vector fp32 global g = amax/2688, so any finite float32 range fits and block scales stay out of the E4M3 subnormal range. AppendBlockScaled / ParseBlockScaledCell / Dequantize / string round trip; the parser rejects malformed cells (length, version, format, reserved and padding bits, non-finite or negative g, NaN or signed scales, NaN elements). Scale selection is tested minimal at every E8M0/E4M3 boundary; nibble order is checked against CUDA's __nv_cvt_float2_to_fp4x2. Measured with this codec on unit-normalized wiki_all 768-d (50k/200): recall@10 vecf8 0.971, vecf4 0.922. Design doc updated to the two-level vecf4 format. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rixorigin#20567) T_array_float8 (230, vecf8) and T_array_float4 (231, vecf4) are varlena vector types with Width = dimension. They are deliberately NOT in IsArrayRelate(), which keeps meaning "Width elements of GetArrayElementSize bytes": vecf8/vecf4 cells are block-scaled (vecblock.go) and paths that decode typed slices or call GetArrayElementSize must not accept them. - IsBlockScaledVector(): vecf8/vecf4. - IsVectorType(): IsArrayRelate() || IsBlockScaledVector(), for structural checks that apply to every vector column (PK/partition/index rejection, dimension guards). - BlockScaledFormat(): oid -> cell format. - Type.ArrayCellBytes(): cell byte length for any vector type, the replacement for Width*GetArrayElementSize() size estimates. - Names (VECF8/VECF4, vecf8/vecf4), DescString, OidString, TypeLen, FixedLength, ToType, Types map, DecodeValue/EncodeValue (raw bytes). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Parser: VECF8/VECF4 tokens (non-reserved keywords) and ArrayFamily type rules; mysql_sql.go regenerated (no conflicts); tree.T prints the width. - Plan: getTypeFromAst maps every vector SQL name through the new types.VectorTypeBySQLName (replaces three hand-maintained six-name lists), so vecf8/vecf4 get the default width and the 1..65535 dimension check. - SHOW CREATE prints VECF8(N)/VECF4(N). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ithmetic/SUM/AVG/inner_product, Parquet LOAD (matrixorigin#20567) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rounding vecf4 quantizer (matrixorigin#20567) - IF/IFF/CASE/COALESCE keep one bf16/float16/float8/float4/vecf8/vecf4 type, so multi-table UPDATE and other DML projections write cells of the column width (they stored float32 bytes before). - != (NOT IN), <=> and BETWEEN round in-range literals to the column type as = and IN do; filter simplification no longer keys pre-rounding literals. - A text literal compared with vecf8/vecf4 is cast to the column dimension. - types.CompareValue handles the low-precision scalars (data branch diff/merge). - Data branch SQL text renders dequantized vecf8/vecf4 values. - vector_matmul scores on the CPU when the GPU engine cannot be created. - Prepared narrowing accepts integer user variables and decimal-numeral text only; IF condition, INTERVAL, CONV and BIT_AND/OR/XOR take low-precision args. - vecf8/vecf4 comparison decodes without per-row allocation. - vecf4/vecf8 elements are the exact quotient rounded once (no float32 intermediate); the design doc states the quantizer's scale choice. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…allback with a GPU (matrixorigin#20567) - Casts from text to bf16/float16/float8/float4 and to vecf8/vecf4, and from vecf8/vecf4 to vecf32, leave rows outside the select list NULL without converting them, as the float casts do (a CASE branch no longer errors on an unselected invalid value). - vector_matmul scores on the CPU only with gpu_mode off or no device of compute capability 10.0+. With an eligible device enabled, an allocation account denial or any engine error fails the query. - The engine looks up a cuBLASLt algorithm for every tile row bucket (128 * 2^i up to the tile capacity) at construction, so a shape the device cannot run fails before scoring; tiles run at the cached algorithms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # pkg/defines/const.go # pkg/sql/compile/ip_functions_protocol.go # pkg/sql/compile/remoterun.go # pkg/sql/plan/base_binder.go # pkg/sql/plan/function/func_testcase.go # pkg/sql/plan/function/operatorSet.go # pkg/sql/plan/prepared_binding_test.go
A BLOB of little-endian float32 elements casts to vecf8/vecf4 and quantizes as the same vecf32 value, matching the binary vector input of vecf32. A length not a multiple of 4, another dimension or a non-finite value is rejected; an empty BLOB is NULL. The parity test also covers BLOB shapes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
matrixorigin#20567) The queries argument may be a BLOB holding nq x N float32 values back to back (CAST(? AS BLOB) for a client's bytes); text stays a JSON array. The compiler records the format in the configuration flags; the length must be a non-zero multiple of 4*N and the values finite. Both forms configure the same queries. The BLOB-query BVT checks the CPU kernel in cases/ (gpu_mode = 0) and the cuBLASLt result in gpu_cases/. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
matrixorigin#20567) INSERT/REPLACE VALUES and DEFAULT bound every numeric, hex or bit literal with the target column type, also inside a CAST, so CAST(X'...' AS BLOB) into a vector column became CAST(<vector> AS BLOB) and failed (the form a client's bytes parameter takes). A vector column type is no longer that binding type; other column types keep it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Qodo reviews are paused for this user.Troubleshooting steps vary by plan Learn more → On a Teams plan? Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center? |
XuPeng-SH
left a comment
There was a problem hiding this comment.
Decision: REQUEST_CHANGES — value contracts are still not closed
Reviewed d4822f2e307cfc6c76e9bf3792f5714f5af24e4b against base/merge-base 43c37847b16d37cfa8b9eef4235ba4388b5e25ce. The architecture has useful reuse: codecs belong in types, top-k partial/merge/spill use the existing aggregate framework, and GPU execution stays behind the engine boundary. The three inline findings are concrete correctness defects in the new value contracts, not requests for additional compatibility machinery.
Confirmed current-head failures
- P1: NVFP4 replication/reload changes already-stored values. A 17-element vector with only elements 0 and 16 nonzero (
8.7649145,5.7432985) stores the latter as6.2606535. Production CDC emits the decoded text; reloading it stores6.8867183, approximately 10% additional drift. Both the production CDC extraction/SQL formatter and a fresh single-CN MySQL-protocol insert/select/reinsert reproduce this. This is additional loss during transport, not the permitted initial quantization. - P2: equality disagrees with grouping/distinct. Real SQL says
CAST('[1.2031566]' AS vecf4(1)) = CAST('[1.2031565]' AS vecf4(1))is true, but storing those values produces 2 GROUP BY groups and 2 DISTINCT values. MXFP8 also fails:[447]and[449]both decode to[448], compare equal, and still produce 2 groups. A same-encoding vecf8 neighbour returns 1 group as expected. Deterministic encoding of one input does not establish a canonical encoding for all equal decoded values. - P2: ARM64 fused decoding breaks distance invariants. With normal Go 1.26.4 compilation, a valid
vecf4(16)containing sixteen9.9999994e29values givesVecBlockL2DistanceSq(v,v) = +Inf; real SQLl2_distance_sq(v,v)returns error 20101. Casting that same stored vector to vecf32 first returns 0. Smaller values also have nonzero self-distance. The newTestVecBlockKernelsSymmetricfails deterministically (3/3); package-local-d=fmahash=nmakes both symmetry and the self-distance probe pass. This is a production decode/rounding boundary defect, not merely an overstrict test.
Resolve these at the existing owners: preserve the stored cell/value across transport; extend the shared equality-key contract rather than adding an alternate grouping path; make the kernel generator's decode rounding match the stored-value contract rather than adding self-distance special cases or globally disabling FMA.
Historical review disposition
- My three old inline threads are now resolved/outdated. The current encoder/parser rejects decoded infinity; host staging/tile admission uses the aggregate account; device selection filters the declared compute-capability baseline. Current finiteness and fake-engine admission/cleanup tests pass.
- iamlinjunhong's select-list finding is fixed;
TestNarrowCastHonorsSelectListpasses. - The previous NVFP4 K=48/no-algorithm concern is addressed in source by padding K to 32 and preflighting tile algorithms, together with an explicit eligible GPU => execute or fail decision. That decision can be reasonable; it is different from a CPU-fallback promise. The design still says fallback on memory denial at lines 269 and 524, contradicting its final no-fallback decision and implementation. Reconcile those paragraphs and the stale draft/merge references in the PR body. No real GPU execution is claimed here.
- fengttt's bitwise-format question has an author answer distinguishing CUDA-compatible element/nibble encoding from the custom block/global quantizer. These are not the same assertion; keep that distinction in the final contract. The current CPU types suite, including checked-in CUDA golden verification, passes.
QA, test quality and scope
Fresh CPU native build succeeded. Relevant types/CDC/aggregate/function tests passed; explicit skipped-row CAST passed; the aggregate owning package passed once under -race, and TestVectorMatmulConfigShared passed 100/100 race repetitions. Full types/compare/sort suites passed. The metric suite did not pass, for the symmetry failure above. Reviewer-only overlays and logs were kept outside the clean review worktree; no production code was changed and no subagent was used.
The existing text round-trip check only proves that reparsing succeeds, not that the stored value is preserved. CDC/ISCP fixtures use exact small values, so they miss scale-boundary drift. Add small orthogonal oracles for transport identity, equality-versus-key agreement in both formats, and decoded self-distance with FMA enabled. Do not fix this by deleting the failing symmetry case, widening every tolerance, or piling up repetitions of the exact-grid fixtures.
The 188-file change was mapped by responsibility, with implementation, tests, documentation and generated code considered separately. The mandatory design gate remains blocked by these counterexamples: this is not a blanket acceptance of every remaining hunk, unhappy path, GPU performance claim or device configuration. GPU benchmarks/real-device validation were not rerun on this CPU-only host. Stop-the-cluster upgrades are the scope; mixed-version rolling compatibility is not a blocker, but persisted-value and replication correctness remain required.
matrixorigin#20567) options {"metric": "inner_product" | "cosine" | "l2sq"} selects the distance; the scores are the values of inner_product() (-dot), cosine_distance() (1 with a zero vector) and l2_distance_sq(), nearest first. Other option keys are ignored; text that is not a JSON object, or another metric, is an error. GPU: a row-statistics kernel computes each packed row's squared norm (and the uint8 element sums, moved off the host) from the values the matmul multiplies; the fix-up kernel applies the metric, and run uses it too, so the host does no score arithmetic. The CPU engine computes the same formula with x.x norms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ls (matrixorigin#20567) The generated kernels decoded an element as an unrounded product (code x scale) that the compiler could fuse into the following subtract on targets with fused multiply-add (arm64), so x - y rounded only one side: a vector had a nonzero self-distance, and +Inf near the float32 limit. Each decoded element is now an explicit float32 conversion, which rounds it as the stored value is and prevents the fusion; only the accumulations may still fuse. On arm64 the L1 kernels drop from 16 fused instructions to 0 and the L2 kernels from 32 to 16 (the d*d accumulation). A self-distance test covers values near the float32 limit and several magnitudes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ixorigin#20567) Two cells can encode equal values with other block or global scales (vecf8 [447] and [449] both store 448), so = held while GROUP BY, DISTINCT, joins and set operations keyed the raw cell bytes and kept them apart. The shared equality-key contract (keycodec AppendCanonicalValue, CanonicalValuesEqual, CanonicalValueSize and the spill hash) now keys vecf8/vecf4 by their decoded float32 values, signed zero closed, and every site that routes the other vector types through it routes vecf8/vecf4 too: the hashmap group/join keys, COUNT(DISTINCT) and its capacity preflight, approx_count_distinct and its protocol gate. Runtime filters stay off for vector types and shuffle only takes integer and string keys. A key-versus-equality oracle covers both formats; SAMPLE also accepts vecf8/vecf4. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…it (matrixorigin#20567) CDC/ISCP replication and data branch merge wrote vecf8/vecf4 values as their decoded text, which the replica quantized again; a vecf4 global scale follows the decoded maximum, so a stored 6.2606535 replayed as 6.8867183. The exact text is the cell as stored: {"g": g, "b": [{"s": scale, "v": [element values]}, ...]} ("g" only for vecf4; vecf8's global is 1). Every scale and element must be the value of its code, 32/16 per block and fewer only in the last, so parsing rebuilds the same cell bytes; anything else is an error, never rounded. vecblock_json(v) returns it; casting, inserting or LOADing it builds the cell as written (types.StringToBlockScaled is the one text entry point). CDC, ISCP and data branch SQL now carry this text; query output and INTO OUTFILE keep the decoded values. The design doc also drops two stale CPU-fallback statements. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…matrixorigin#20567) encoding/json decoded 1000x768 query JSON at 88.5 ms; sonic at 25.8 ms. Unknown keys and non-numeric values are still rejected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…trixorigin#20567) INTO OUTFILE and external-table writes emit vecf8/vecf4 cells as the vecblock JSON text (a quoted CSV field, a JSONL object value), so an export reloads to the same cells instead of being quantized again. JSONL LOAD passes an object value of a vecf8/vecf4 column through as its JSON text. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
fixed |
…norm leaves range (matrixorigin#20567) A unit's float32 squared norm that overflowed to +Inf or underflowed to 0 gave cosine distance 1 instead of 0, depending on the dimension. Recompute in float64 as the vecf32 kernels do. A squared L2 beyond the float32 range is +Inf at every dimension, as for vecf32. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gin#20567) ORDER BY ties and window PARTITION BY / ORDER BY peers compared cell bytes, so cells holding the same values with other block scales were distinct peers while = and GROUP BY treated them as equal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…loat8/float4 (matrixorigin#20567) e = '0.3' compared in DOUBLE and missed the stored value, as did BETWEEN and multi-item IN with text items. A prepared IN list with one value outside the type's range lost the rounding of the others to the list-level numeric prefix pass; the list now expands to per-item comparisons that bind each marker. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…; Parquet LOAD reads it (matrixorigin#20567) MODUMP QUERY_RESULT and the checkpointtool SQL-LOAD dump wrote vecf8/vecf4 as decoded values, which reload quantized again (vecf4 values move). Both now write the vecblock JSON. Parquet string columns accept it as CSV and JSONL do. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rixorigin#20567) A failed engine creation marked the engine as tried, so a batch retried on the same executor after a capacity error scored on the CPU. The mark is set only on success or when no device is in play, and cleared by Free. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… in float64 (matrixorigin#20567) Follows 8a1ebd1: inner product still errors, cosine distance is 1 and similarity 0 for the orthogonal pair, as for vecf32. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n#20567) SET @v from a vecf8/vecf4 value decoded it to vecf32, so writing @v back quantized it again and vecf4 values could move. The variable now holds the cell (types.BlockScaledValue) with the column type: it rebuilds the same bytes, displays the decoded values and survives connection migration. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/float4 (matrixorigin#20567) A text, float64, decimal or wide integer source was rounded to float32 and then to the narrow type, so a value just above a tie landed one step low (cast('1.0039062500001' as bf16) was 1, not 1.0078125). Float32RoundToOdd keeps the sticky bit, so the narrow rounding is the single correct one, in casts, INSERT values, CSV and Parquet LOAD. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cell casts back exactly (matrixorigin#20567) The exact binary form of a vecf8/vecf4 value is its cell (12-byte header, block scales, element codes). A BLOB of a valid cell of the target casts, inserts or binds to the same bytes without quantization; its length is never 4*dim, so a float32 BLOB still quantizes. A BLOB of cell length that is not a valid cell is an error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…origin#20567) A Parquet BYTE_ARRAY / FIXED_LEN_BYTE_ARRAY column without a logical type is binary for a vecf8/vecf4 target, read as a BLOB casts: a value of cell length is the stored cell, a value of 4*dim bytes is float32 elements. STRING/JSON columns stay text. The BLOB rule is one helper, types.BlockScaledFromBinary, shared by the cast and the loader, and a value of neither length names both. Adds the BVT resource parquet/vecblock.parquet, written and checked by TestVecBlockParquetResource. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…atrixorigin#20567) The exact text lists the parts in the cell's order: the global scale (always written, 1 for vecf8), one scale per block, and the flat element values, so element i is v[i] * s[i/blockSize] * g and NumPy decodes it in one line. Smaller than the per-block form, which is no longer accepted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/vecf4 cells (matrixorigin#20567) With {"query_format":"vecblock"} in the options, the queries of a vecf8/vecf4 column are cells used as given, on the CPU and the GPU: a BLOB of cells back to back (vecblock_binary) or a JSON array of vecblock JSON objects. A query taken from a stored vecf4 row is then at distance 0 from it; its decoded float values quantize to another cell. float32, the default, keeps the float-value queries. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Thanks for the review. All three are fixed, each with a test that fails without the fix. P1: NVFP4 replication drift. Fixed by carrying the stored cell instead of its decoded values.
P2: equality vs GROUP BY / DISTINCT. Fixed: every equality consumer now agrees with =.
P3: arm64 FMA decode. Fixed in 76771e7.
A self-review on top of these also fixed: |
… on the same bytes (matrixorigin#20567) NVFP4 and MXFP8 operands drawn in NVIDIA's layout with one global scale per operand are scored by a direct cuBLASLt call (alpha = G_x * G_q) and, as MO cells carrying that global on every row, by the engine. Scores agree to the float32 rounding of the two runs and the top-10 rows match; at the engine's tile shape MXFP8 is bit-identical and NVFP4 within 1 ulp (the globals are applied in double after the GEMM, not as an fp32 alpha). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… alpha (matrixorigin#20567) The fix-up kernel computes the inner product as acc * fp32(g_row * g_query) in fp32, as cuBLASLt applies alpha = G_a * G_b with per-tensor global scales, so NVFP4 scores of rows sharing one global equal a direct NVIDIA block-scaled GEMM on the same bytes bit for bit (previously within 1 ulp). Cosine and squared L2 keep the double product, so a row equal to a query stays at distance 0. MatchesNvidiaBlockScaledGemm now asserts bit identity at the engine's GEMM shape. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
What type of PR is this?
Which issue(s) this PR fixes:
issue #20567
What this PR does / why we need it:
Low-precision float types: scalar
bf16/float16/float8/float4, and block-scaled vector columnsvecf8(N)(MXFP8) /vecf4(N)(NVFP4).Draft: the
vector_matmul/vector_matmul_mergebatch search (CPU, then cuBLASLt GPU) continues on this branch.Scalar
bf16/float16/float8(E4M3) /float4(E2M1)Design:
docs/design/20260929-low-precision-scalar-floats.mdFloat8/Float4codecs pinned bit-exactly to CUDAcuda_fp8.h/cuda_fp4.hby a generated golden test.Vector columns
vecf8(N)/vecf4(N)Design:
docs/design/20260930-low-precision-vector-storage.mdg), then block scales (E8M0 per 32 for vecf8; UE4M3 per 16 for vecf4), then elements. Value =g × block scale × element; the layout matches cuBLASLt's block-scaled matmul inputs.+ - * /promotes to vecf32.inner_product,l2_distance,l2_distance_sq,l1_distance,cosine_distance,cosine_similaritybetween vecf8/vecf4/vecf32 in any combination; a text literal binds as vecf32.vector_dims,normalize_l2,ANY_VALUE,COUNT,GROUP_CONCAT, ORDER BY / GROUP BY / window.LIST<FLOAT/DOUBLE>and text columns).pkg/vectorindex/metric/distance_func_vecblock*.go): scalar Go, no SIMD. They are generated (mkvecblock.go), one per operand pair and metric, fully unrolled over 16-element units, with table-lookup decode. 768-d, one row: 200–311 ns for dot/L2, 303–437 ns for L1/cosine.perf(metric): clear AVX upper state after archsimd kernelsThe archsimd distance kernels returned with dirty upper zmm state. The legacy-SSE scalar float code that follows them then carries a false dependency on those bits.
archsimd.ClearAVXUpperBits()after the reduction in 40 kernels. The 4 bf16 AVX2 kernels are excluded (+17% there, no gain after them).Tests
dtype/vecblock.sql(65 statements, including flush, prepared statements and rejections), plus the scalar low-precision cases.mo-tester -m run -g -n), all 0 failed:dtypearrayvector🤖 Generated with Claude Code