fix: Manual float16 conversions to extend platform support - #2773
fix: Manual float16 conversions to extend platform support#2773iwoplaza wants to merge 2 commits into
Conversation
|
pkg.pr.new packages benchmark commit |
Bundle size comparison (
|
| 🟢 Decreased | ➖ Unchanged | 🔴 Increased (max 1.00%) | ❔ Unknown |
|---|---|---|---|
| 0 | 37 | 285 | 0 |
import * as ... in PR vs import * as ... in target (did bundle size increase?):
Click to reveal the results table (152 entries).
| Test | tsdown |
|---|---|
| d_bool.ts | 13.62 kB ( |
| d_f16.ts | 13.62 kB ( |
| d_f32.ts | 13.62 kB ( |
| d_i32.ts | 13.62 kB ( |
| d_u32.ts | 13.62 kB ( |
| d_u16.ts | 13.65 kB ( |
| d_textureDepth2d.ts | 14.07 kB ( |
| d_textureDepthCube.ts | 14.07 kB ( |
| d_texture1d.ts | 14.08 kB ( |
| d_texture2d.ts | 14.08 kB ( |
| d_texture3d.ts | 14.08 kB ( |
| d_textureCube.ts | 14.08 kB ( |
| d_textureDepth2dArray.ts | 14.08 kB ( |
| d_textureDepthCubeArray.ts | 14.09 kB ( |
| d_textureDepthMultisampled2d.ts | 14.09 kB ( |
| d_texture2dArray.ts | 14.09 kB ( |
| d_textureCubeArray.ts | 14.10 kB ( |
| d_textureMultisampled2d.ts | 14.10 kB ( |
| std_discard.ts | 14.92 kB ( |
| std_isBeingTranspiled.ts | 15.01 kB ( |
| std_getTargetShaderLanguage.ts | 15.08 kB ( |
| std_extensionEnabled.ts | 15.13 kB ( |
| std_copy.ts | 15.16 kB ( |
| std_arrayLength.ts | 15.16 kB ( |
| std_range.ts | 15.40 kB ( |
| d_disarrayOf.ts | 15.57 kB ( |
| std_dpdx.ts | 15.88 kB ( |
| std_dpdxCoarse.ts | 15.88 kB ( |
| std_dpdxFine.ts | 15.88 kB ( |
| std_dpdy.ts | 15.88 kB ( |
| std_dpdyCoarse.ts | 15.88 kB ( |
| std_dpdyFine.ts | 15.88 kB ( |
| std_fwidth.ts | 15.88 kB ( |
| std_fwidthCoarse.ts | 15.88 kB ( |
| std_fwidthFine.ts | 15.88 kB ( |
| std_bitcastF32toU32.ts | 47.75 kB ( |
| std_bitcastU32toF32.ts | 47.75 kB ( |
| std_bitcastU32toI32.ts | 47.75 kB ( |
| std_bitcast.ts | 47.76 kB ( |
| std_atomicLoad.ts | 16.68 kB ( |
| std_atomicStore.ts | 16.68 kB ( |
| std_textureBarrier.ts | 16.68 kB ( |
| std_atomicAdd.ts | 16.69 kB ( |
| std_atomicAnd.ts | 16.69 kB ( |
| std_atomicMax.ts | 16.69 kB ( |
| std_atomicMin.ts | 16.69 kB ( |
| std_atomicOr.ts | 16.69 kB ( |
| std_atomicSub.ts | 16.69 kB ( |
| std_atomicXor.ts | 16.69 kB ( |
| std_storageBarrier.ts | 16.69 kB ( |
| std_workgroupBarrier.ts | 16.69 kB ( |
| d_vec2b.ts | 20.07 kB ( |
| d_vec2f.ts | 20.07 kB ( |
| d_vec2h.ts | 20.07 kB ( |
| d_vec2i.ts | 20.07 kB ( |
| d_vec2u.ts | 20.07 kB ( |
| d_vec3b.ts | 20.07 kB ( |
| d_vec3f.ts | 20.07 kB ( |
| d_vec3h.ts | 20.07 kB ( |
| d_vec3i.ts | 20.07 kB ( |
| d_vec3u.ts | 20.07 kB ( |
| d_vec4b.ts | 20.07 kB ( |
| d_vec4f.ts | 20.07 kB ( |
| d_vec4h.ts | 20.07 kB ( |
| d_vec4i.ts | 20.07 kB ( |
| d_vec4u.ts | 20.07 kB ( |
| d_formatToWGSLType.ts | 21.56 kB ( |
| d_uint8.ts | 21.56 kB ( |
| d_float16.ts | 21.57 kB ( |
| d_float16x2.ts | 21.57 kB ( |
| d_float16x4.ts | 21.57 kB ( |
| d_float32.ts | 21.57 kB ( |
| d_float32x2.ts | 21.57 kB ( |
| d_float32x3.ts | 21.57 kB ( |
| d_float32x4.ts | 21.57 kB ( |
| d_sint16.ts | 21.57 kB ( |
| d_sint16x2.ts | 21.57 kB ( |
| d_sint16x4.ts | 21.57 kB ( |
| d_sint32.ts | 21.57 kB ( |
| d_sint32x2.ts | 21.57 kB ( |
| d_sint32x3.ts | 21.57 kB ( |
| d_sint32x4.ts | 21.57 kB ( |
| d_sint8.ts | 21.57 kB ( |
| d_sint8x2.ts | 21.57 kB ( |
| d_sint8x4.ts | 21.57 kB ( |
| d_snorm16.ts | 21.57 kB ( |
| d_snorm16x2.ts | 21.57 kB ( |
| d_snorm16x4.ts | 21.57 kB ( |
| d_snorm8.ts | 21.57 kB ( |
| d_snorm8x2.ts | 21.57 kB ( |
| d_snorm8x4.ts | 21.57 kB ( |
| d_uint16.ts | 21.57 kB ( |
| d_uint16x2.ts | 21.57 kB ( |
| d_uint16x4.ts | 21.57 kB ( |
| d_uint32.ts | 21.57 kB ( |
| d_uint32x2.ts | 21.57 kB ( |
| d_uint32x3.ts | 21.57 kB ( |
| d_uint32x4.ts | 21.57 kB ( |
| d_uint8x2.ts | 21.57 kB ( |
| d_uint8x4.ts | 21.57 kB ( |
| d_unorm10_10_10_2.ts | 21.57 kB ( |
| d_unorm16.ts | 21.57 kB ( |
| d_unorm16x2.ts | 21.57 kB ( |
| d_unorm16x4.ts | 21.57 kB ( |
| d_unorm8.ts | 21.57 kB ( |
| d_unorm8x2.ts | 21.57 kB ( |
| d_unorm8x4.ts | 21.57 kB ( |
| d_unorm8x4_bgra.ts | 21.57 kB ( |
| d_packedFormats.ts | 21.59 kB ( |
| d_isPackedData.ts | 21.63 kB ( |
| d_alignmentOf.ts | 22.52 kB ( |
| std_subgroupAdd.ts | 25.02 kB ( |
| std_subgroupAll.ts | 25.02 kB ( |
| std_subgroupAnd.ts | 25.02 kB ( |
| std_subgroupAny.ts | 25.02 kB ( |
| std_subgroupBallot.ts | 25.02 kB ( |
| std_subgroupBroadcast.ts | 25.02 kB ( |
| std_subgroupBroadcastFirst.ts | 25.02 kB ( |
| std_subgroupElect.ts | 25.02 kB ( |
| std_subgroupExclusiveAdd.ts | 25.02 kB ( |
| std_subgroupExclusiveMul.ts | 25.02 kB ( |
| std_subgroupInclusiveAdd.ts | 25.02 kB ( |
| std_subgroupInclusiveMul.ts | 25.02 kB ( |
| std_subgroupMax.ts | 25.02 kB ( |
| std_subgroupMin.ts | 25.02 kB ( |
| std_subgroupMul.ts | 25.02 kB ( |
| std_subgroupOr.ts | 25.02 kB ( |
| std_subgroupShuffle.ts | 25.02 kB ( |
| std_subgroupShuffleDown.ts | 25.02 kB ( |
| std_subgroupShuffleUp.ts | 25.02 kB ( |
| std_subgroupShuffleXor.ts | 25.02 kB ( |
| std_subgroupXor.ts | 25.02 kB ( |
| d_isBuiltin.ts | 25.25 kB ( |
| d_sizeOf.ts | 25.29 kB ( |
| d_isContiguous.ts | 25.30 kB ( |
| d_getLongestContiguousPrefix.ts | 25.31 kB ( |
| std_textureDimensions.ts | 26.63 kB ( |
| std_textureGather.ts | 26.63 kB ( |
| std_textureLoad.ts | 26.63 kB ( |
| std_textureSample.ts | 26.63 kB ( |
| std_textureSampleBaseClampToEdge.ts | 26.63 kB ( |
| std_textureSampleBias.ts | 26.63 kB ( |
| std_textureSampleCompare.ts | 26.63 kB ( |
| std_textureSampleCompareLevel.ts | 26.63 kB ( |
| std_textureSampleGrad.ts | 26.63 kB ( |
| std_textureSampleLevel.ts | 26.63 kB ( |
| std_textureStore.ts | 26.63 kB ( |
| d_arrayOf.ts | 26.73 kB ( |
| d_size.ts | 26.98 kB ( |
| d_align.ts | 26.98 kB ( |
| d_location.ts | 26.99 kB ( |
| d_interpolate.ts | 26.99 kB ( |
import { ... } in PR vs import * as ... in PR (is the library tree-Shakeable?):
| Test | tsdown |
|---|---|
| tgpu_init.ts | 260.15 kB ( |
| tgpu_initFromDevice.ts | 259.62 kB ( |
| tgpu_resolve.ts | 165.55 kB ( |
| tgpu_resolveWithContext.ts | 165.49 kB ( |
| tgpu_bindGroupLayout.ts | 69.40 kB ( |
| tgpu_mutableAccessor.ts | 66.41 kB ( |
| tgpu_accessor.ts | 66.39 kB ( |
| tgpu_privateVar.ts | 65.74 kB ( |
| tgpu_workgroupVar.ts | 65.74 kB ( |
| tgpu_const.ts | 64.98 kB ( |
| tgpu_fn.ts | 38.59 kB ( |
| tgpu_fragmentFn.ts | 38.59 kB ( |
| tgpu_vertexFn.ts | 38.40 kB ( |
| tgpu_computeFn.ts | 38.11 kB ( |
| tgpu_vertexLayout.ts | 27.22 kB ( |
| tgpu_comptime.ts | 14.91 kB ( |
| tgpu_unroll.ts | 1.66 kB ( |
| tgpu_slot.ts | 1.54 kB ( |
| tgpu_lazy.ts | 1.19 kB ( |
If you wish to run a comparison for other, slower bundlers, run the 'Tree-shake test' from the GitHub Actions menu.
There was a problem hiding this comment.
Pull request overview
This PR replaces Float16Array-dependent float16 read/write paths with a manual IEEE-754 binary16 conversion layer to support runtimes like React Native (Hermes) that don’t implement Float16Array.
Changes:
- Introduces
float16Conversion.tswith helpers to encode/decode float16 values viauint16payloads. - Switches
std/packinganddataIOfloat16 serialization to use the new helpers instead oftyped-binaryfloat16 APIs. - Updates
std/bitcastCPU implementation to avoid constructingFloat16Array, using aUint16Arrayview plus encode/decode.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
| packages/typegpu/src/std/packing.ts | Uses manual float16 read/write wrappers for pack/unpack builtins. |
| packages/typegpu/src/std/bitcast.ts | Removes Float16Array buffer view; adds f16 encode/decode path via Uint16Array. |
| packages/typegpu/src/data/float16Conversion.ts | Adds new float16 encoding/decoding utilities used across the codebase. |
| packages/typegpu/src/data/dataIO.ts | Routes all f16 scalar/vector serialization through the new conversion helpers. |
Comments suppressed due to low confidence (1)
packages/typegpu/src/data/float16Conversion.ts:33
decodeUint16AsFloat16handles subnormals incorrectly: forexponent === 0, IEEE-754 binary16 uses an exponent of -14 (no implicit leading 1). Returningmantissa / 1024makes subnormal values ~16384× too large (e.g. 0x0001).
if (exponent === 0) {
return sign === 0 ? mantissa / 1024 : -mantissa / 1024;
}
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Resolution Time Benchmark---
config:
themeVariables:
xyChart:
plotColorPalette: "#E63946, #3B82F6, #059669"
---
xychart
title "Random Branching (🔴 PR | 🔵 main | 🟢 release)"
x-axis "max depth" [1, 2, 3, 4, 5, 6, 7, 8]
y-axis "time (ms)"
line [0.78, 1.53, 3.27, 5.67, 6.38, 10.78, 18.27, 23.47]
line [0.81, 1.56, 3.50, 5.73, 5.94, 11.18, 18.87, 21.47]
line [0.85, 1.72, 3.58, 5.66, 6.21, 10.12, 18.82, 20.02]
---
config:
themeVariables:
xyChart:
plotColorPalette: "#E63946, #3B82F6, #059669"
---
xychart
title "Linear Recursion (🔴 PR | 🔵 main | 🟢 release)"
x-axis "max depth" [1, 2, 3, 4, 5, 6, 7, 8]
y-axis "time (ms)"
line [0.31, 0.47, 0.59, 0.70, 0.96, 0.99, 1.17, 1.29]
line [0.34, 0.51, 0.63, 0.78, 1.05, 1.02, 1.24, 1.31]
line [0.33, 0.50, 0.61, 0.73, 0.96, 1.00, 1.24, 1.40]
---
config:
themeVariables:
xyChart:
plotColorPalette: "#E63946, #3B82F6, #059669"
---
xychart
title "Full Tree (🔴 PR | 🔵 main | 🟢 release)"
x-axis "max depth" [1, 2, 3, 4, 5, 6, 7, 8]
y-axis "time (ms)"
line [0.82, 1.78, 3.37, 5.72, 10.35, 23.14, 47.39, 96.34]
line [0.82, 1.82, 3.43, 5.58, 10.67, 22.60, 48.38, 98.09]
line [0.77, 1.85, 3.99, 6.40, 11.35, 22.65, 47.92, 98.67]
|
There was a problem hiding this comment.
Caution
The new float16 conversion functions in float16Conversion.ts have algorithmic bugs that corrupt subnormal values — decoded 16384× too large, encoded as garbage bit patterns — and lose negative zero. packages/typegpu/src/data/numeric.ts already exports correct, well-tested toHalfBits/fromHalfBits that do not depend on Float16Array; reusing those would fix all three issues at once.
Reviewed changes
- New
float16Conversion.ts—writeFloat16/readFloat16wrappers for binary I/O plus standaloneencodeFloat16AsUint16/decodeUint16AsFloat16functions implementing IEEE 754 binary16 conversion. dataIO.ts— Alloutput.writeFloat16/input.readFloat16calls replaced with the newwriteFloat16/readFloat16helpers.bitcast.ts—Float16Arraybuffer view replaced withUint16Array; newwriteFloat16ToBuffer/readFloat16FromBufferfunctions handle the f16 bitcast path with manual encode/decode.packing.ts—writeFloat16/readFloat16helpers used inpack2x16float/unpack2x16floatCPU paths.
DeepSeek Pro (free via Pullfrog for OSS) (Kimi K2 not used — the program covers this model; add its provider key to run your pick) | 𝕏
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes
- Removed buggy conversion functions —
encodeFloat16AsUint16/decodeUint16AsFloat16deleted fromfloat16Conversion.ts;writeFloat16/readFloat16now delegate totoHalfBits/fromHalfBitsfromnumeric.ts. - Improved
toHalfBitsrounding — Both the normal and subnormal paths innumeric.tsnow implement proper round-ties-to-even (using round-bit + sticky-bit) instead of the previous add-0.5-and-truncate approach. The normal path also gained an exponent-overflow guard when rounding causes a mantissa carry. - Removed
Float16Arrayfrom bitcast —bufViewsreplacedf16: Float16Arraywithu16: Uint16Array; newwriteFloat16ToBuffer/readFloat16FromBufferfunctions usetoHalfBits/fromHalfBitsfor encode/decode.Float16Arraypurged fromwriteToBuffer/readFromBuffertype signatures. - Added comprehensive tests —
halfBits.test.tsvalidatestoHalfBits/fromHalfBitsagainst nativeFloat16Arrayacross zeros, NaNs, infinities, normals, subnormals, overflow, underflow, and round-ties-to-even (14 tests, all passing).
DeepSeek Pro (free via Pullfrog for OSS) (Kimi K2 not used — the program covers this model; add its provider key to run your pick) | 𝕏
2ae24d2 to
327386b
Compare
327386b to
e7eb8d4
Compare

React Native (Hermes) unfortunately doesn't yet support Float16Array natively, this PR uses the existing
toHalfBitsandfromHalfBitsinstead.I also fixed faulty behavior of
toHalfBitsin some edge cases and added tests validation it behaves like Float16Array would.