Are models with N-Gram tables going to completely change the AI race?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
The news about Qwen 3.8 Flash Next is the first I'm reading about n-gram tables. I may be completely misunderstanding how they work but it seems they could open the door for 1T+ parameter models to be run on a single server with modest GPUs and a ton of system RAM rather than needing a rack of GPU servers connected with something like NVlink.
Could we be looking at shrinking the capability gap between self hosted and flagship models faster than we thought, or am I way off base?
[link] [comments]
More from r/LocalLLaMA
-
Can we reconsider the megathreads?
Aug 26
-
Whoever the fuck predicted we would have gpt 5.5 performance in coding on consumer hardware a couple months ago now, i applaud you
Aug 26
-
[Megathread] GLM-5.3-Flash - former ox-alpha
Aug 26
-
Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?
Aug 26
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.