3.5 KiB
3.5 KiB
Changelog
One section per release, keyed by the exact tag. verify-version refuses to cut
a release whose tag has no section here, so this gets written before the tag is
pushed — the release notes are generated from it verbatim.
Write it for the person who downloads the build: what they get, what changed for them, what to watch out for. Not a commit dump — the notes already carry the full commit list underneath.
Releases before b10269-1.5.0 predate this file; see the git history.
b10269-1.5.1
Fixed
- Ling-3.0-flash (BailingMoeV3) no longer emits garbage token bursts. The
model is trained with clamped SwiGLU activations in its late layers, and the
per-layer limits live in
config.jsonunderexpert_swiglu_limit_listandshare_expert_swiglu_limit_list. The public HF modeling code ignores those keys and so did this port, which caused deterministic transient logit collapse - output likecount += 1evilledropped into otherwise fine generations. Measured at roughly -20 pass@1 on HumanEval (72.6% -> 93%+ with the fix); the garbage-token repro is eliminated.
Notes
- Re-convert your Ling-3.0-flash GGUF to get the fix. The clamp limits are
written by the converter into two new KVs (
{arch}.swiglu_clamp_expand{arch}.swiglu_clamp_shexp); a GGUF produced before this release does not carry them, and the runtime then defaults to no clamping. Re-download the quant or re-runconversion/bailingmoe.py. - Both KVs are optional and default to zero, so existing GGUFs and every other
architecture are unaffected. The graph needed no change - the SwiGLU clamp
branches in
build_ffn/build_moe_ffnalready trigger on a nonzero per-layer limit, matching the vLLMSwigluStepAndMulsemantics.
b10269-1.5.0
Added
- NVIDIA DGX Spark (GB10) support. New archive
llama-turboquant-linux-arm64-cuda-13.3, built natively for aarch64 with CUDA 13.3 and sm_121 SASS. Other arm64 NVIDIA machines (GH200, GB200, Jetson Thor) run it too, JITing the kernels from PTX on first launch. This is the first Linux arm64 build the fork ships — until now arm64 meant macOS only. - BailingMoeV3 (Ling 3.0) architecture support, including the KDA gate handling.
Changed
- Linux CUDA archives are roughly half the size — 1657 → 956 MB (12.4) and
1879 → 1028 MB (13.3) measured across both the
.zipand.tar.gz. The zips were storinglibcublas.so→.so.13→.so.13.5.1.27as three full copies becausezipfollowed the symlinks. - CUDA 13.3 builds ship Ampere PTX (
80-virtual). A100/H100/B200 were falling back to the Turing PTX floor, which silently disabledcp.asyncand the Ampere MMA path — both gated on__CUDA_ARCH__ >= 800. Those cards get Ampere-class kernels now. No architecture lost support in this release. - Windows CUDA builds got their architecture lists pinned, all runner cores, a
ccache that can actually hold a CUDA build, and 7-Zip instead of
Compress-Archive. Release turnaround drops accordingly.
Notes
- The DGX Spark archive has not yet been validated on real GB10 hardware —
it is built and arch-checked in CI (
cuobjdumpasserts sm_121 SASS is present), but nobody has run it on a Spark yet. Treat this one as beta and report back. - The CUDA 13.3 archives now use
-compress-mode=size. Kernel SASS is unchanged and inference speed is unaffected; the fatbin is decompressed once at module load. It needs a driver from the CUDA 12.4 era or newer.