News / #inference Tag Inference 500 articles archived under #inference · RSS Sign in to follow arXiv — Machine Learning research 28d ago Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms arXiv:2608.28044v1 Announce Type: cross Abstract: Large language model (LLM) inference serving is priced by tokens, but GPU energy is consumed over inference windows. This accounting mismatch makes token-normalized metrics incomplete, since average output-token energy can… 9 arXiv — NLP / Computation & Language research 28d ago Trajectory-Level Speculative Decoding for Diffusion Language Models arXiv:2608.27514v1 Announce Type: new Abstract: Diffusion-based language models (dLLMs) enable parallel token generation through iterative denoising, but existing decoding strategies collapse to single-token generation under low confidence, severely limiting throughput. Unlike… 21 arXiv — NLP / Computation & Language research 28d ago SimpCue: Cue-Based Prompting for Multilingual Text Simplification arXiv:2608.28042v1 Announce Type: new Abstract: Text simplification aims to make complex texts easier to understand while preserving their original meaning. Recent large language models can perform simplification through prompting, but it remains unclear whether adding explicit… 25 arXiv — NLP / Computation & Language research 28d ago A Probabilistic Interpretation of KV Cache Eviction arXiv:2608.28293v1 Announce Type: new Abstract: The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most… 37 arXiv — NLP / Computation & Language research 28d ago ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL arXiv:2608.28476v1 Announce Type: new Abstract: Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously… 27 arXiv — NLP / Computation & Language research 28d ago Semantic Watermarking with Order-Robust Detection over Sub-sentence Units arXiv:2608.27666v1 Announce Type: cross Abstract: Semantic watermarks tie the mark to sentence meaning rather than token choices, promising robustness to content-preserving edits. However, the detector only observes attacker-supplied text, which can be reworded, reordered, or… 34 Hugging Face Daily Papers research 28d ago LMSM: LLM Security Framework Inspired by Linux Security Modules Abstract LMSM applies a modular security framework to LLM serving by separating evidence calibration, policy evaluation, and output gating, enabling flexible interpretability-based enforcement without rebuilding request handling. Generated by thinkingmachines/Inkling-Small Large… 19 r/LocalLLaMA community 28d ago R9V: A designer set of kernels I've been working on for R9700s/RDNA4. Qwen3.8-Flash-Next Unsloth IQ4_XS (w/ TP on 2 R9700s, MTP, SSD n-gram, 128k ctx, vision): TG256 of *78 tok/s* (~3x increase), PP8192 of *1510 tok/s* (~30x increase). TL;DR: Ninfer/DS4 but for RDNA4 Highly custom kernels built for RDNA4, applied to vLLM-Radiance to greatly improve Qwen3.8 Flash Next speeds. This is mostly for dual R9700s with preferably 48GB of RAM or higher, but feel free to tinker. SOTA-Scan/DeepGit report in repo. For… 21 r/LocalLLaMA community 28d ago Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s Qwen 3.8 Flash Next (80gb) now at 3.5 tok/s on 12gb mid range android phone thanks to some optimizations and with a low quantization on dense part. I don't want to promote the project, but simply show that it's possible on a $400–$500 phone   submitted by   /u/dai_app… 9 r/LocalLLaMA community 28d ago Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM If you own 4xR9700 and were waiting for the model to make them shine, then I have some good news for you! It's running at 80-120 tokens/second for generation and 12k token/second prefill for a single request, using tcclaviger's MXFP4-FP8 quant and custom vLLM image… 14 r/LocalLLaMA community 28d ago Don't Sleep on EXL3 Quants I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8\_O KV cache. It's a joy to use a dense 30B model at this size and still get \~30 tok/s on a VRAM-constrained laptop. It's supposed to be only slightly worse than the… 37 r/LocalLLaMA community 28d ago {INTRESTING PAPER BASED ON HBF}2607.10186] FlashAccel: Leveraging High-Bandwidth Flash (HBF) for High-Throughput LLM Inference HBF gives 8x - 16x more capacity than HBM at same cost, and with bandwidth till 3 tb/s.   submitted by   /u/9r4n4y [link]   [comments] 8 r/LocalLLaMA community 29d ago Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s. I used the quantized n-gram to INT4, it's 32 GB, memory-mapped from disk. I confirmed that it works great on 150-160k context, and i was watching all the time my VRAM usage while doing single thread long horizon things - the available VRAM should be enough to push it to over… 12 r/LocalLLaMA community 29d ago I’ve pushed llama.cpp pretty far for Qwen3.8-Flash-Next — is there any reason not to move to vLLM for 200K+ context? I'm currently running Qwen3.8-Flash-Next on a CMP 170HX 64GB + RTX 3090 24GB, with 80GB system RAM. With llama.cpp I've already spent quite a bit of time tuning it: layer split across the two GPUs, PLE on CPU, q8 KV, Flash Attention, detached MTP draft on the 170HX, and some… 29 r/LocalLLaMA community 29d ago Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp) I wanted to share my successful setup for running a Qwen 3.8 27B model with a massive context window on a consumer 16GB GPU (RTX 4070 Ti SUPER). The goal was to fit everything into VRAM without sacrificing quality or speed. 🧠 Key Components Model:… 38 r/LocalLLaMA community 1mo ago use llms to auto annotation your dataset locally hi i make tool for this called llmog it's purpose to make llms free to - auto annotation datasets - reclassification existing yolo datasets running totally local using llama cpp or vllm or use external api you'd rather click than code. 🔗 GitHub: mohamed-em2m/llmog: framework… 36 r/LocalLLaMA community 1mo ago Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks Hey all, and hello fellow DGX Spark-ers! Today I managed some pretty crazy numbers: 181 tok/s aggregate on 2× DGX Spark on Qwen3.8-Flash-Next at 512Kcontext (2.8M kvc) I hit 181 tok/s aggregate today across a multi-agent fleet on a 2-node DGX Spark cluster. Single-stream decode… 22 NVIDIA Developer Blog official-blog 1mo ago Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing,... 22 Hugging Face Daily Papers research 1mo ago Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning Abstract An agentic framework combining LLMs and VLMs enables consistent, multi-instruction editing of long multi-shot videos while preserving spatiotemporal structure. Generated by thinkingmachines/Inkling-Small While generative AI has significantly advanced video editing,… 16 arXiv — Machine Learning research 1mo ago Predicting Quantifiability from Primary Screens to Prioritize Dose-Response Profiling arXiv:2608.26538v1 Announce Type: new Abstract: High-throughput drug screening relies on low-cost primary assays to prioritize compounds for more expensive dose-response profiling, where potency is ultimately quantified. Current screening strategies largely focus on identifying… 20 arXiv — NLP / Computation & Language research 1mo ago Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives arXiv:2608.26372v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something… 23 arXiv — NLP / Computation & Language research 1mo ago Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD arXiv:2608.26735v1 Announce Type: new Abstract: Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities such as reasoning, coding, instruction following, and creative writing. We study this domain--general… 29 r/LocalLLaMA community 1mo ago Ninfer and a 5090 with 3.8 27B is making me cry tears of joy it's so good. Built the latest and I'm getting as much as 220 tokens per second and averaging in the 170s, I can't get over it. If anyone on here is on that project, fuckkkin' chapeau man, really incredible job. I can't believe I was able to like double or more my throughput from llama.cpp… 34 r/LocalLLaMA community 1mo ago GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context Hey all! I'm finally doing some cool stuff with my "thinking heater" (h/t u/-TV-Stand- ). I'm still experimenting with GLM-5.2 (in anticipation of 5.3 coming tomorrow, I hope!) and things are very cool so far. With the release of GLM-5.3-flash, I decided to play with it on the… 25 arXiv — Machine Learning research 1mo ago ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference… 9 arXiv — Machine Learning research 1mo ago NVExplain: Explaining Time Series Forecasting with Latent Trajectory Analysis and Structure-Preserving Surrogates arXiv:2608.25080v1 Announce Type: new Abstract: Time series forecasting models are widely used in high-stakes settings, yet their predictions remain difficult to interpret because existing post-hoc methods often ignore temporal dependence and fail to provide horizon-specific… 20 arXiv — Machine Learning research 1mo ago Transforms for LLM Quantization: The Great Inversion and Format Co-Design arXiv:2608.25188v1 Announce Type: new Abstract: Most competitive 4-bit LLM research pipelines now open the same way: apply a linear, function-preserving transform (rotation, scaling, permutation, non-orthogonal affine) so the outlier mass sits more favorably against the group… 12 arXiv — Machine Learning research 1mo ago Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data arXiv:2608.25794v1 Announce Type: new Abstract: Federated Learning (FL) enables distributed training of machine learning models while preserving data privacy. However, FL struggles with heterogeneous, non-IID client data distributions, resulting in sub-optimal and biased global… 19 arXiv — Machine Learning research 1mo ago When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs arXiv:2608.25941v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study… 17 arXiv — NLP / Computation & Language research 1mo ago Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting arXiv:2608.25115v1 Announce Type: new Abstract: Existing methods for improving Retrieval-Augmented Generation (RAG) efficiency mainly optimize downstream LLM generation, such as context compression or serving optimization. However, RAG is an end-to-end system, and its bottleneck… 29 arXiv — NLP / Computation & Language research 1mo ago TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving arXiv:2608.25523v1 Announce Type: new Abstract: Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memory available… 5 r/LocalLLaMA community 1mo ago Qwen3.8 27B C8 at 972 TG / 5,680 PP on 4x MI100 rig ($6.5k) using my new INT8 vLLM fork Yet another vLLM fork thread here, but this time its for older INT8-centric hardware. This is a complete INT8 serving stack for Qwen3.8 27B based on vLLM, AITER, and a 27B GPTQ INT8 quant w/ DFlash2 . Its not just another vibed autoresearch loop. No, vLLM ships with very little… 9 r/LocalLLaMA community 1mo ago Lemonade end-of-summer project update, now serving 15 engines! Hi everyone, it's been a while since I posted so here's an update on what the Lemonade community has been up to this summer. Our overall mission is to enable local AI builders with everything they need to make great apps and agents, while keeping the stack turnkey, portable, and… 23 r/LocalLLaMA community 1mo ago Open Source Kernel in Qwen3.6-35B-A3B for AMD MI350X: 78,498 output tok/s on 8 GPUs So here's the thing, almost everyone use NVIDIA to run their LLMs, we also do the same, a lot of people we've met use like RTX PRO 6000 or even H100, B300 It seems like everyone eyes is looking at NVIDIA. However we do the math that the raw power alone on AMD GPU MI350X is… 9 arXiv — Machine Learning research 1mo ago Calibration-Preserving Pruning: Compression as a Reliability Contract arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning… 31 arXiv — Machine Learning research 1mo ago PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage arXiv:2608.24040v1 Announce Type: new Abstract: Enterprise AI agents in production often need to be bounded, stateful, observable, and governable rather than fully autonomous. We present PinSieve, a production case study in a large-scale content-quality pipeline. Its deployed… 21 arXiv — Machine Learning research 1mo ago Conditional GraphGANFed: Optimizing Graph-Structured Molecule Generation in Federated Generative Adversarial Networks arXiv:2608.24610v1 Announce Type: new Abstract: Generative adversarial networks (GANs) have garnered considerable attention in molecular discovery for their ability to generate novel and high-quality molecules. To efficiently train a GAN model while preserving data privacy,… 18 TechCrunch — AI news-outlet 1mo ago OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art. 6 Hugging Face Daily Papers research 1mo ago Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization Abstract ERPO replaces action-side policy regularization with input-side query distribution control to stabilize reinforcement learning for language models while preserving response exploration. Generated by thinkingmachines/Inkling-Small Policy optimization (PO) for Large… 4 OpenAI official-blog 1mo ago Jalapeño’s first results show industry-leading speed and efficiency in AI inference Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models. 34 Hugging Face Daily Papers research 1mo ago Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion Abstract Block3D accelerates text-to-3D generation by using block-wise diffusion with confidence-guided correction to reduce inference time while preserving geometric fidelity. Generated by thinkingmachines/Inkling-Small While text-to-3D generation has advanced rapidly,… 13 arXiv — NLP / Computation & Language research 1mo ago Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation arXiv:2608.21606v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whether such information has truly become inaccessible remains challenging. Existing… 19 Hugging Face Daily Papers research 1mo ago TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration Abstract TileMix routes attention score tiles to mixed FP16 or INT8 precision within fused dense attention, recovering long-context accuracy while improving prefill throughput without retraining. Generated by thinkingmachines/Inkling-Small Long-context prefill in large language… 5 r/LocalLLaMA community 1mo ago [2608.16157] FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Source of Claims: https://x.com/Andy_ShuoYang/status/2090856976880472439 Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization! Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s DeepSeek-V4-Flash 284B → RTX 5090… 22 r/LocalLLaMA community 1mo ago Qwen 3.8 27B Aider score I ran the Aider benchmark on Qwen 3.8 27B FP8 with FP8 KV cache 256K context vLLM. The score: 72.9 This matches Gemini 2.5 Pro from 2025-04-12 which also scored 72.9. Beats Claude Opus 4 from 2025-05-25 which scored 72.0. DeepSeek R1 2025-06-06 scored 71.4. It may just be a… 24 arXiv — Machine Learning research 1mo ago FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space arXiv:2608.21096v1 Announce Type: new Abstract: Federated learning enables privacy-preserving collaborative training, but highly heterogeneous client data remain challenging, especially in graph federated learning where clients possess structurally diverse graphs. Existing… 11 arXiv — Machine Learning research 1mo ago Amplifying the imaging power of digital sky surveys with space telescopes data and generative AI arXiv:2608.20666v1 Announce Type: cross Abstract: While Digital sky surveys provide excellent throughput of image data and can cover a large footprint, their imaging power is normally inferior to that of space-based telescopes. Space-based telescopes, on the other hand, provide… 31 arXiv — NLP / Computation & Language research 1mo ago SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields arXiv:2608.20839v1 Announce Type: new Abstract: Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. Existing sampling-based watermarking methods typically inject position-wise i.i.d.… 14 arXiv — NLP / Computation & Language research 1mo ago Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs arXiv:2608.20953v1 Announce Type: new Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding,… 5 arXiv — NLP / Computation & Language research 1mo ago Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI arXiv:2608.20393v1 Announce Type: new Abstract: Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering… 7 Page 5 of 10 · 500 articles ← Newer Older →