llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8, supporting full 262K ctx.
↑ +3 today
1 tracked repository
Showcase your total open-source impact. This SVG badge updates automatically as our continuous crawler records new stars.
[](https://githubrepo.cloud/developers/JakeATX)
Sorted by stars
llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8, supporting full 262K ctx.