LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Everything runs through WebGPU, in-browser or in electron/tauri apps. It's fully portable and supports either Nvidia and Apple Silicon (Metal). The actual kernels are optimized for the specific hardware of the device. The Nvidia kernels are aggressively fused into a multi-pass architecture, while the Apple Silicon kernels are created as a fused mega-kernel to minimize the Tile Based Deferred Rendering (TBDR) overhead on WebGPU. Demo: https://warp.sipp.sh
This is still in active development, and I'll be folding this into the Sipp library in the coming weeks. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.