-
6cbf352851
llm-proxy: propagate client cancellation to the backend (free the slot)
main
williams
2026-07-20 22:14:40 +00:00
-
06f20ae752
llm-proxy: per-backend admission gate (semaphore, bounded queue, 503 on saturation)
claude-noether
2026-07-20 13:50:18 +02:00
-
ac77f054c2
systemd: add llama-server-qwen3-4b-npu.service (NPU Qwen3-4B on :8091)
master
Markus Fritsche
2026-07-19 17:57:39 +02:00
-
-
c121428c58
Initial commit: Boltzmann LLM proxy with compression + classifier middleware
Markus Fritsche
2026-06-15 15:12:25 +02:00