News / #inference Tag Inference 500 articles archived under #inference · RSS Sign in to follow r/LocalLLaMA community 8d ago I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8... TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practical winner so far is a recent CUDA build of llama.cpp with tensor split, Flash Attention, --numa distribute , and large batches. On… 29 r/LocalLLaMA community 8d ago So i tried Remotion with glm 5.3 flash, this mfker is really good. *8bit, vllm, 4x dgx* Prompt: Go download and use Remotion and create a cool 60-second motion graphics video with it. Impress me totally. The motion graphics must be based on stock market visuals. Add many cool, mind-blowing motion graphics and explain what fundamental vs.… 35 r/LocalLLaMA community 9d ago 4x RTX 3090 PCIe 4.0 x16 - advice? Qwen 3.8 Next Flash? We're upgrading our server (Threadripper Pro 5955WX, 128GB 8-channel DDR4) from two RTX 3090 to four cards. Currently we're running Qwen 3.8 27b Q8 with vLLM for a few users. I'm wondering what we should do next: Keep running Qwen 3.8 27b Q8 and enjoy the performance boost and… 33 r/LocalLLaMA community 9d ago Made this motion graphic video via glm 5.3 flash (no vid_gen model used) *8bit, vllm, 4x dgx* Prompt: The given zip file contains the skill.md and all that type of stuff. Your work is to make a amazing infographic video of nearly 40 seconds. It should be based on stock market charts, etc., along with various other kinds of infographics. Please… 11 r/LocalLLaMA community 9d ago Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI access to the project files so it could search through them for bugs and stuff. I already had an old computer that I decided not to… 20 r/LocalLLaMA community 9d ago Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Requirements: M3 or newer, macOS 26.4+, 36 GB Get started with a single command:… 32 r/LocalLLaMA community 9d ago Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings: explicit effort level on every role (unset ones defaulted to xhigh) thinking_token_budget of 7500 maxTokens 8k → 32k (file writes were getting cut off) tool output over 10… 23 r/LocalLLaMA community 9d ago Qwen3.8-Flash-Next (95.5 GiB) on a 64GB Mac at ~27 tok/s, checkpoint + fork I've been running Qwen3.8-Flash-Next as my main local coding model from past few weeks. The file is 95.5 GiB and my Mac has 64 GB . It works because the routed experts stay on SSD and only get read when a token actually routes to them. Finally cleaned it up enough to publish:… 33 r/LocalLLaMA community 9d ago Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark Planning, coding, testing = 8h total. Stack: VSCode Copilot in autopilot mode + SGLang Stats: ∼10k lines generated, ∼800k tokens consumed Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...   submitted by… 20 r/LocalLLaMA community 9d ago Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison Hey r/LocalLLaMA , Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark . We wanted to see… 24 r/LocalLLaMA community 10d ago Still on the Jev waitlist? I hosted OpenJev. It's free, go play with it TypeSafe announced Jev on Tuesday: you give it data plus typed questions (yes/no, pick-one, 0–N scale) and it returns a probability for every option, crazy fast. I signed up and then refreshed my inbox. A lot. Meanwhile Matt Mastracci opened vLLM PR #57250 , which does the same… 5 arXiv — Machine Learning research 10d ago Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems arXiv:2609.19337v1 Announce Type: new Abstract: We present Personalized Federated Hierarchical Gaussian Processes (pFedHGP) for probabilistic regression and classification when data are distributed across heterogeneous clients. Each client's latent function decomposes into (i) a… 7 arXiv — Machine Learning research 10d ago Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems arXiv:2609.19695v1 Announce Type: new Abstract: Federated learning (FL) enables privacy-preserving, on-device training across heterogeneous Internet-of-Things (IoT) deployments such as smart-city water-metering networks, where each smart meter observes a household-specific… 14 arXiv — Machine Learning research 10d ago Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks arXiv:2609.19842v1 Announce Type: new Abstract: Effective EEG decoding requires representations that preserve organization among channels, local waveform dynamics, and long-range temporal context. Existing EEG architectures often capture these structures using separate… 36 arXiv — Machine Learning research 10d ago Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map arXiv:2609.20333v1 Announce Type: new Abstract: We study reconstruction in autoencoders that apply the same forward map before and after setting the observed coordinates to zero. For equal odd input and hidden dimensions $d\geq 3$, among orientation-preserving diffeomorphisms… 29 arXiv — Machine Learning research 10d ago Training Neural Networks to Approach the Optimum Bayes Estimator in Dense Multi-Emitter Localization arXiv:2609.20465v1 Announce Type: new Abstract: We train neural networks on synthesized frames to approach the optimum Bayes estimator for dense emitter localization. The result justifies the future work on training neural networks to achieve high-throughput large-FOV super… 13 arXiv — NLP / Computation & Language research 10d ago AI Should Facilitate Democratic Deliberation at Scale arXiv:2609.20059v1 Announce Type: cross Abstract: AI systems can strengthen democracy by supporting deliberation at scale by addressing cognitive, social, platform-design, and market-driven frictions, while preserving human agency. Unlike proposals such as liquid democracy that… 14 Vercel — AI dev-tools 10d ago GLM 5.3 FlashX now available on AI Gateway GLM 5.3 FlashX is now available on AI Gateway. GLM 5.3 FlashX is a high-speed serving option for Z.ai's multimodal coding model, delivering inference at ~200 tokens per second for faster streamed responses. The higher serving speed is useful for coding agents, tool loops, and… 21 r/LocalLLaMA community 10d ago 600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"? It's not the brightest bulb but it's my new drudgework model for read+find or code tasks I'm willing to let it brute force. Even if it takes 20x more tokens, that's still faster than many local models. Not quite Cerebras but still pretty fun to drive.   submitted by  … 38 r/LocalLLaMA community 10d ago 153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s People kept commenting and asking about single AMD 1xR9700 cards in the comments and discord. Well, I finally had time to do some optimizations for 1xR9700 owners and performance has doubled across the board. You can see the results in BetterBench above if you like visuals or… 19 r/LocalLLaMA community 10d ago Keeping vLLM's Prefix Cache Warm Between Agent Turns   submitted by   /u/bolts98 [link]   [comments] 10 r/LocalLLaMA community 10d ago First M5 Ultra benchmarks just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!   submitted by   /u/Ashefromapex… 8 r/LocalLLaMA community 10d ago We built an open-source GPU profiler you point an AI agent at, instead of reading traces yourself We've been tuning vLLM/SGLang/llama.cpp setups for a long time and got tired of the profiling part: nsys trace, open the GUI, squint, change a flag, repeat. The profilers assume a human is looking at the timeline. These days the thing doing our tuning is usually an agent, and it… 26 Hugging Face Daily Papers research 11d ago The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Abstract Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's… 28 arXiv — Machine Learning research 11d ago FedPGT: Progressive Gradient Transmission for Vehicular Federated Learning over Time-Varying Channels arXiv:2609.18089v1 Announce Type: new Abstract: Vehicular federated learning (VFL) enables privacy-preserving collaborative model training for intelligent transportation systems, where communication resource allocation and gradient sparsification techniques have been explored to… 32 arXiv — Machine Learning research 11d ago Multi-Appliance Non-Intrusive Load Monitoring via Label-Preserving Aggregate Recomposition and Prediction Consistency arXiv:2609.18315v1 Announce Type: new Abstract: Non-intrusive load monitoring (NILM) estimates appliance power sequences from aggregate power, but models trained on source households commonly lose accuracy in unseen households. Aggregate power also contains loads from other… 23 vLLM releases dev-tools 11d ago proto-v0.3.0 Release vllm-proto 0.3.0 25 NVIDIA Developer Blog official-blog 11d ago TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through... 14 r/LocalLLaMA community 11d ago "I'm the one degenerating" -qwen 3.8 27b q4km mtp I was working with my local llm on llama.cpp to get another server up with vllm, but we were running into trouble with the tool calls. When I asked it to investigate it went into a doom loop. Then during the troubleshooting of the doom loop it suddenly became self-aware. Another… 20 r/LocalLLaMA community 11d ago Qwen3.8-27B uncensored Q6_K at 156K context on one RTX 5090, 140-190 tok/s with DFlash2 Setup for one long agentic coding session (tools + vision) on a single 5090: Qwen3.8-27B RVN Heretic (ARA abliterated) at Q6_K, 159744 context , 1 slot, q8_0 K/V, DFlash2 speculative decoding, vision projector on the GPU, Sharp chat template. All numbers measured 2026-09-16.… 18 vLLM releases dev-tools 12d ago proto-v0.2.0: vllm-proto 0.2.0 Validated by PR #56538 CI at fa2a26f . 23 arXiv — Machine Learning research 12d ago Distilling Foundation Models for Agentic What-If Reasoning:Cost, Latency, and Governance in a Hybrid LLM+SLM Architecture arXiv:2609.16091v1 Announce Type: new Abstract: Tabular foundation models deliver strong zero-training predictive performance via in-context learning, but their high inference latency makes them impractical as hot-path decision backends in interactive agentic loops. We distill a… 34 arXiv — Machine Learning research 12d ago LLM Inference in a Flash! arXiv:2609.16161v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown impressive capabilities across a range of natural language processing tasks, and LLM inference has emerged as a critical workload for enabling downstream applications. The demands of serving… 31 arXiv — NLP / Computation & Language research 12d ago Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening arXiv:2609.16517v1 Announce Type: new Abstract: Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change… 15 arXiv — NLP / Computation & Language research 12d ago TIAO: Token Importance-Aware Policy Optimization for Text Summarization arXiv:2609.16748v1 Announce Type: new Abstract: Text summarization requires models to condense content while preserving key qualities such as consistency and coherence. Large language models (LLMs) have shown strong performance on this task and can be further improved through… 34 arXiv — NLP / Computation & Language research 12d ago Cascade: Hierarchical Recoverability Control for Large Language Model Unlearning arXiv:2609.16890v1 Announce Type: new Abstract: Large Language Model (LLM) unlearning is essential for removing sensitive or copyrighted knowledge while preserving general utility. Existing methods often leave residual knowledge in intermediate representations, which can still… 24 NVIDIA Developer Blog official-blog 12d ago Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the... 38 r/LocalLLaMA community 12d ago Qwen3.8-27B-NVFP4 1M context. So far so good. I am a beginner, Took a while to get started, get everything right. This setup is native not container. Still not sure if I did this right, or if I can tune this more. Environment=HF_HUB_OFFLINE=1 Environment=VLLM_LOGGING_LEVEL=INFO Environment=VLLM_ALLOW_LONG_MAX_MODEL_LEN=1… 15 r/MachineLearning community 12d ago I trained a 44M parameter quantized LLM from scratch on 45B tokens. It ships in 19.8 MB and runs at ~1,900 tok/s on CPU. [P] Three weeks back , i posted SHADOW-250M here. It got 360 upvotes, 293 on r/LocalLLaMA and 94 GitHub stars. Thank you. That model was 60 MB, ran around 400 tok/s on CPU and could retrieve records from an archive on disk. What it couldn’t do reliably was reason over what it… 26 arXiv — Machine Learning research 13d ago Discovering and Preserving Category Correlation Knowledge via Adaptive Reciprocal Knowledge Distillation arXiv:2609.13199v1 Announce Type: new Abstract: Knowledge distillation aims to improve the performance of lightweight student models by transferring knowledge from larger and more powerful teacher models. However, a substantial size gap between teacher and student models often… 13 arXiv — Machine Learning research 13d ago Data-Efficient Agentic Graph Domain Adaptation via Reliability-Aware Prototype Learning arXiv:2609.14045v1 Announce Type: new Abstract: Agentic learning systems are often required to adapt after deployment by observing new data and reusing prior knowledge under limited supervision or feedback. For graph-structured prediction, Graph Domain Adaptation (GDA) naturally… 36 arXiv — NLP / Computation & Language research 13d ago DenMark: Robust Semantic Watermarking for Diffusion Language Models arXiv:2609.14257v1 Announce Type: new Abstract: Semantic text watermarks encode signals in meaning rather than surface token choices, offering robustness to paraphrasing and other semantic-preserving edits. Existing semantic watermarking methods are primarily designed for… 15 Hugging Face Daily Papers research 13d ago LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents Abstract A 16.7B-parameter mixture-of-experts diffusion vision-language agent achieves strong multimodal GUI performance while preserving block-parallel decoding efficiency. Generated by thinkingmachines/Inkling-Small Diffusion large language models (dLLMs) achieve high decoding… 37 r/LocalLLaMA community 13d ago Running Qwen3.8-Flash-Next locally on a 12GB VRAM card Now that the dust has settled a bit - here's a write-up on running Qwen3.8-Flash-Next (125B-A6B MoE + 51B n-gram table) on relatively middle-tier hardware (RTX 4070 12GB + 64GB DDR5-5600 + Gen4 NVMe on Linux). I started out with bare 6 tok/s and through latest patches and… 28 r/LocalLLaMA community 14d ago Intern-S2-397B (multimodal, reasoning, coding, and scientific agent capabilities) Model: https://huggingface.co/internlm/Intern-S2-397B Collection: https://huggingface.co/collections/internlm/intern-s2 From Intern Large Models on 𝕏: https://x.com/intern_lm/status/2099425184587370976 vLLM on 𝕏: Day-0 support for Intern Large Models Intern-S2-397B is now… 17 arXiv — Machine Learning research 14d ago On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health arXiv:2609.11961v1 Announce Type: new Abstract: Stress is a pervasive determinant of mental health and a key target for mobile health interventions. On-device language models (ODLMs) offer privacy-preserving inference without cloud dependency, yet their feasibility for health… 35 arXiv — Machine Learning research 14d ago Hidden in Rounds: Predicting the Time Cost of 802.11 Contention in Federated Learning arXiv:2609.12903v1 Announce Type: new Abstract: Federated learning over IEEE~802.11 shares the wireless channel among clients that send model updates. We use ns-3 to measure the frame-delivery ratio and saturation throughput for different client densities and offered loads. A… 9 arXiv — NLP / Computation & Language research 14d ago Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding arXiv:2609.12243v1 Announce Type: new Abstract: Inference-time power sampling via Sequential Monte Carlo (SMC) can substantially improve large language model (LLM) reasoning without requiring post-training. However, many existing SMC approaches rely on equal-weight resampling,… 26 arXiv — NLP / Computation & Language research 14d ago Parameter-Efficient Retrievers for Polish and European Languages arXiv:2609.12913v1 Announce Type: new Abstract: Dense retrieval systems increasingly rely on multi-billion-parameter language models, whose memory and computational requirements make large-scale indexing, frequent corpus updates, and low-latency serving costly. We present a… 12 arXiv — NLP / Computation & Language research 14d ago Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning arXiv:2609.13045v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-linguistic information. However, these models struggle with predicting high-bitrate… 27 Page 2 of 10 · 500 articles ← Newer Older →