GLM-5.3-Flash (abliterated EXL3) on 2x NVIDIA DGX Spark with the TensorFold engine: 1.8x faster decode than vLLM, 4x256k concurrent threads, byte-exact speculative decoding. Work in progress.
↑ +3 today
1 tracked repository
Showcase your total open-source impact. This SVG badge updates automatically as our continuous crawler records new stars.
[](https://githubrepo.cloud/developers/jayleaton)
Sorted by stars
GLM-5.3-Flash (abliterated EXL3) on 2x NVIDIA DGX Spark with the TensorFold engine: 1.8x faster decode than vLLM, 4x256k concurrent threads, byte-exact speculative decoding. Work in progress.