Files
rk-llama.cpp/src/models
Georgi Gerganov 1725e316c1 models : optimize qwen3next graph (#19375)
* models : optimizing qwen3next graph

* cont

* wip

* wip

* wip

* wip

* wip

* wip

* wip

* wip

* wip

* wip

* cont : remove redundant q, g chunking

* minor

* minor

* avoid passing masks around

* avoid concats during chunking

* naming + shapes

* update names and use prefix to disable CUDA graphs
2026-02-14 12:57:36 +02:00
..
2026-01-13 23:28:38 +01:00
2025-11-27 16:04:29 +02:00
2025-12-28 17:28:31 +01:00
2025-12-15 18:51:43 +01:00