CTX: How far can you reasonably go with Qwen 3.6 27B?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
How far can i stretch the context window with Qwen 3.6 27B (using Q8_0) before it gets too unreliable? I am at 100k right now and i am not quite statisfied.
Other than not quantizing KV cache, is there anything else that can be done to make the model more stable over longer CTX?
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.