News / #inference Tag Inference 500 articles archived under #inference · RSS Sign in to follow r/MachineLearning community 5h ago Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P] AI Engineering from Scratch is an MIT-licensed curriculum: 523 lessons across 20 phases, from linear algebra and backprop to transformers, LLMs, agents, and production serving. The code is stdlib-first, so you see every step instead of calling a library. This month's edition: -… 37 arXiv — NLP / Computation & Language research 7h ago Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge arXiv:2609.31511v1 Announce Type: new Abstract: We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a… 9 r/LocalLLaMA community 7h ago I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2… 36 NVIDIA Developer Blog official-blog 10h ago How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency Every unused watt is capacity left on the table. AI factories are typically provisioned for the unlikely moment when every GPU reaches peak power, creating a... 15 r/LocalLLaMA community 13h ago Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM We built an inference engine for MoE models that don't fit in VRAM + RAM. Most of the model stays on the SSD, and experts are read as tokens need them. This started as a proof of concept, and poc worked, we are getting 9-10 tok/s decode on Qwen3.8-Flash-Next NVFP4 (9.06 on the… 6 r/LocalLLaMA community 18h ago is switching from llama cpp to vllm worth it I have hp z8 g4 with 512 ram and 1x3090 1x5060 16gb. has anyone made the transition from llama cpp to vllm recently? is it worth it? docker under windows or full linux install? I am mainly interested in the model support, it seems that many new local models are supported day 0… 37 r/LocalLLaMA community 1d ago Which of the 16gb VRAM qwen3.8 27b’s is the best? I’m having a hard time finding out which one gives you fastest speed, maximum context with best possible quality. I can run unsloth qwen3.8 27b iq4_xs with 65k q8 kv, context without MTP and vision offloaded to cpu. But also kinda slow for agentic work at like 30ish tok/s )I… 30 r/LocalLLaMA community 1d ago 85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri I've been experimenting with ways to run MoE models that don't fit comfortably in RAM, and I ended up making Overspill, a disk tier for FreeToken . The basic idea came from looking at how Colibri handles experts across disk/RAM/VRAM so I took inspiration from the general… 11 r/LocalLLaMA community 2d ago Custom Models in Oh My Pi: vLLM, llama.cpp, SGLang and More   submitted by   /u/bolts98 [link]   [comments] 27 r/LocalLLaMA community 2d ago Make Volta Fast Again For those who have V100 cards, I wanted to point you to 1Cat-vLLM, a vLLM fork that enables optimized serving for these cards. Showing stats for Qwen3.6-35b comparing a Strix Halo with a hughly optimized llama.cpp fork (pwilkin) and the V100 with 1Cat. It’s not apples to apples,… 22 r/LocalLLaMA community 2d ago How do you use subagents & multiple agent with local models, and how many? Running qwen3.8 27b nvfp4 on vllm at max context only gives around 8 agents with 32k context each. That doesnt seem like much; what use cases do people use multi-agent frameworks and find it helpful for?   submitted by   /u/Ambitious_Fold_2874 [link]   [comments] 15 r/LocalLLaMA community 3d ago VLLM 4x rtx 3060 vs 8x rtx 3060 performance loss Hello! I am currently building my local AI server, I have the Huananzhi H12D-8D EPYC Motherboard with 8x16GB memory sticks at 2666 mhz (waiting for the other components at the moment) I currently have four RTX 3060 12gb gpus and I plan running those at PCIe4 x16 in VLLM. However… 17 r/LocalLLaMA community 3d ago Did anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some I let Codex do some eval on Qwen 3.8 Flash Next IQ4_XS (served via vllm and r9v) and Qwen 3.8 FP8 (served via vllm and radiance). Here are the results: Benchmark Flash-Next IQ4_XS Qwen3.8 27B FP8 Result MMLU-Pro 83.8% 75.0% Flash-Next GPQA Diamond 42.5% 27.5% Flash-Next GSM8K… 8 arXiv — Machine Learning research 3d ago Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning arXiv:2609.29711v1 Announce Type: new Abstract: Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas… 16 arXiv — NLP / Computation & Language research 3d ago TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification arXiv:2609.29633v1 Announce Type: new Abstract: We present TTLab's submission to the AlexandriaX-2026 Subtask~3 on Arabic MT error span detection and classification. Our system frames the task as token-level classification over surface forms, preserving character offsets to… 14 Ollama releases dev-tools 3d ago v0.40.0-rc0: llama-server: prepare to remove compatibility patch Add manifest-list storage so runner-specific manifests can coexist under one tag while preserving existing v1 tags as best-effort downgrade anchors. Show/list/copy/remove/pull/push now understand runner and digest selection and transfer referenced child manifests and layers. Add… 38 r/LocalLLaMA community 3d ago CachyOS Qwen 3.8 27b on 2x 5090 vLLM   submitted by   /u/piddlefaffle12 [link]   [comments] 7 r/LocalLLaMA community 3d ago Qwen 3.8 27b be like... The user is frustrated — I rambled too much and didn't act. Let's just run the test suite and move on. No more forensics. One command, execute, then report. (Original memo is a casual internal monologue in English. Translating faithfully while preserving the informal,… 9 r/LocalLLaMA community 3d ago Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~430 tok/s… 38 r/LocalLLaMA community 4d ago My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context These boards cost me $115 each and I have them connected using llama.cpp with Vulkan and RPC on Bazzite. The boards have roughly 27GB of combined GPU memory and communicate over 1gb Ethernet. For around $300 including psu I’m loving the performance. I have a few more and want to… 22 arXiv — Machine Learning research 4d ago PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation arXiv:2609.26890v1 Announce Type: new Abstract: Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain… 36 arXiv — Machine Learning research 4d ago ZO-COSMO: Index-Free One-Hop Mixing for Decentralized Zeroth-Order Optimization arXiv:2609.27199v1 Announce Type: new Abstract: Sparse communication in decentralized zeroth-order learning requires compatible peer-state coordinates. We characterize this one-hop condition and develop \textsf{ZO-COSMO}, coupling two-query estimation with average-preserving… 11 arXiv — Machine Learning research 4d ago Pheno-GS: Phenoscape-scale Geodesic Sinkhorn arXiv:2609.27633v1 Announce Type: new Abstract: High-throughput single-cell data is now collected across large patient cohorts. Understanding patient-level heterogeneity from cellular-level data motivates phenoscaping: embedding each single-cell distribution as a "datapoint,"… 24 arXiv — NLP / Computation & Language research 4d ago Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition arXiv:2609.27289v1 Announce Type: new Abstract: Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an… 38 arXiv — NLP / Computation & Language research 4d ago MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression arXiv:2609.27380v1 Announce Type: new Abstract: Likelihood-based context compression can account for cross-context redundancy through sequential scoring, but this makes compression outcomes sensitive to context order. We show that different permutations of the same context… 6 NVIDIA Developer Blog official-blog 4d ago How SWE-Serve Exposes the Gap Between Local Tests and Live Serving An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software... 38 arXiv — Machine Learning research 5d ago Federating Quantum and Classical Computing: A Privacy-Preserving Hybrid Approach arXiv:2609.25082v1 Announce Type: new Abstract: Quantum machine learning (QML) is increasingly recognized as one of the most promising near-term applications of quantum computing, viewed as a next-frontier candidate beyond purely classical approaches. Hybrid quantum-classical… 21 arXiv — Machine Learning research 5d ago A Lightweight Plastic-Memory Framework for Graph Few-Shot Class-Incremental Learning arXiv:2609.25781v1 Announce Type: new Abstract: Graph Incremental Learning has garnered increasing attention as dynamic graph data continues to emerge across diverse fields. Conventional approaches primarily address catastrophic forgetting by preserving node-related knowledge… 30 arXiv — Machine Learning research 5d ago GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression arXiv:2609.25963v1 Announce Type: new Abstract: Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation… 28 arXiv — NLP / Computation & Language research 5d ago Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation arXiv:2609.26693v1 Announce Type: new Abstract: A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that… 36 arXiv — NLP / Computation & Language research 5d ago Efficient Iterative Retrieval with Heterogeneous Batching arXiv:2609.25405v1 Announce Type: cross Abstract: Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these… 23 r/LocalLLaMA community 5d ago My local llm when I tell it to do any changes to my vLLM service better make no mistakes I've been running Qwen 3.8 Flash Next and it's a great driver for Hermes and Pi. I told it to add CUDA_DISABLE_PERF_BOOST=1 to reduce my server's idle power draw   submitted by   /u/ZaltyDog [link]   [comments] 38 r/LocalLLaMA community 5d ago M5 ultra AI test results So we have the results: - Prompt processing is up to 4 / 4.5 times faster than M3 ultra depending on how long the context is - tok/s is around 1.5x faster. But: the machine uses twice the power (400w vs 200w), makes more fan noise, and runs much hotter.… 16 r/LocalLLaMA community 5d ago MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap Some people here say MiMo-V2.6 is bad with tools and are going back to GLM-5.3-Flash. I spent today running MiMo-V2.6-Flash-RL as the backend for an agent harness, on 2× DGX Spark with vLLM, using the tonyd2wild recipe. Most of the "tool problems" I hit turned out to be serving… 8 r/LocalLLaMA community 5d ago Fork of FreeToken with DeepSeek-V4.1, vision and speculative decoding (2x3090 numbers inside) I've been running FreeToken on my 2x3090 box for a while and ended up maintaining a fork of it. Posting it in case it's useful to anyone else here. Quick context if you haven't used it: FreeToken is an edge-native MoE serving engine. It offloads experts to host RAM/NVMe and… 31 r/LocalLLaMA community 6d ago Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed with half the context on the same hardware. I benchmarked a couple dozen quants, vllm, sglang, llama.cp and benchmarked settings and… 26 arXiv — NLP / Computation & Language research 6d ago Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs arXiv:2609.22144v1 Announce Type: new Abstract: Preserving safety alignment during large language models fine-tuning is critical, however, recent studies have demonstrated that even benign fine-tuning data may contain safety-degrading samples that silently undermine safety… 14 arXiv — NLP / Computation & Language research 6d ago H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution arXiv:2609.22241v1 Announce Type: new Abstract: We present H2LooP Telecom Model v1, a domain-specialized large language models fine-tuned for the telecommunications industry. We release two domain-adapted model variants serving complementary use cases: a comprehension-focused… 34 arXiv — NLP / Computation & Language research 6d ago DIPLOMAT: Dialogue-Span-Aware Direct Preference Optimization for Polite Persuasive Workplace Negotiation Dialogues arXiv:2609.22256v1 Announce Type: new Abstract: Effective workplace negotiation requires balancing multiple objectives, including achieving task goals, preserving professional relationships, and resolving conflicts constructively. However, misunderstandings, misaligned… 26 arXiv — NLP / Computation & Language research 6d ago Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation arXiv:2609.22603v1 Announce Type: new Abstract: Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while… 27 NVIDIA Developer Blog official-blog 6d ago Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability... 18 r/LocalLLaMA community 6d ago [Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff" https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=webp&s=74765dbd409f4c221640f9f6000a685f6fdbb242 Spent weekend benchmarking the Splash engine (by Incoai) and extending its architecture to native 8-bit on Apple Silicon (M5 Pro, 64 GB unified memory). Splash is… 10 arXiv — Machine Learning research 7d ago MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems arXiv:2609.21533v1 Announce Type: new Abstract: LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs… 15 arXiv — Machine Learning research 7d ago Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability arXiv:2609.21647v1 Announce Type: new Abstract: Hamiltonian normalizing flows are attractive generative models because their phase-space maps are invertible and volume preserving, but most neural constructions are formulated in Euclidean space. We introduce Riemannian Neural… 17 arXiv — Machine Learning research 7d ago Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation arXiv:2609.19122v2 Announce Type: cross Abstract: Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal… 17 arXiv — NLP / Computation & Language research 7d ago Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction arXiv:2609.21231v1 Announce Type: new Abstract: Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving… 17 arXiv — NLP / Computation & Language research 7d ago Consistent Relexicalization of Clinical Documents using Graph-Based Approach arXiv:2609.21387v1 Announce Type: new Abstract: Relexicalization is a pivotal technique in clinical NLP, as it facilitates robust masking of sensitive information while synthesizing datasets that retain high-fidelity, real-world characteristics. However, preserving structural… 28 r/LocalLLaMA community 7d ago Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak. ​ I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster. Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but… 6 r/LocalLLaMA community 8d ago Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0) Hi. I saw some feedback that halogen was degrading at context depth. So I fixed that. Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0: decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter) decode at 258,794:… 18 r/LocalLLaMA community 8d ago Improved TPS of Gemma 4 31B : the journey and also creating custom patches with VLLM fork I improved the TPS of Gemma 4 31B. Improving TPS and performing optimisations requires understanding of the model architecture, and I had to fork VLLM and apply my own patch to break into making a configuration work as per my idea. I wrote a full article so that even beginners… 15 Page 1 of 10 · 500 articles Older →