News / #inference Tag Inference 500 articles archived under #inference · RSS Sign in to follow arXiv — NLP / Computation & Language research 14d ago From Bench-to-Bedside: A Review of Clinical Trials in Drug Discovery and Development arXiv:2412.09378v4 Announce Type: replace-cross Abstract: Clinical trials bridge basic research and clinical application, serving as essential steps in drug development. This review examines clinical trial phases (Phase I [safety assessment], Phase II [efficacy evaluation],… 38 r/LocalLLaMA community 14d ago R9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG. Pushed out this new update, hopefully decreases the instances of crashes. I torture tested this one for ~12 hours after my fixes and found no instability. Q4 K XL needs more fine tuning, which I will work on in the future. I am simultaneously juggling this + a legitimate… 27 Hacker News — AI on Front Page community 14d ago Why is Google still serving dodgy ads? Article URL: https://www.atomic14.com/2026/09/13/why-is-google-still-serving-dodgy-ads Comments URL: https://news.ycombinator.com/item?id=49686445 Points: 204 # Comments: 92 38 r/LocalLLaMA community 14d ago Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT) VLLM Benchmark: Prefill, Prompt processing - avg, 871.93 tok/s (3 hours constant running xhigh) - 10K prompt, 1000.26 tok/s (16 runs) - 90K prompt, 743,59 tok/s (16 runs) Decode, tok gen - avg, 38.39 tok/s (3 hours constant running xhigh) - 10K, 42.3 tok/s (16 runs) - 90K, 34… 33 r/LocalLLaMA community 15d ago I built a serverless hosting platform for LoRA adapters with vLLM It’s always bothered me that after fine-tuning a model for a project, there isn’t a particularly easy way to host it without either running it locally and keeping a GPU on 24/7 or paying for an entire GPU server. There are managed options for LoRA serving on top of vLLM (AWS),… 32 r/LocalLLaMA community 15d ago Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini I went with the recommended settings for the best generation quality and ran a few benchmarks: temperature: 0.7 top_p: 0.95 top_k: 40 reasoning_effort: high I must… 20 r/LocalLLaMA community 15d ago 2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s What I have: - CPU: EPYC 7551 (32c/64T, Zen 1) - Board: Supermicro H11SSL-i (SP3), Rev 2.0 - RAM: 128 GB DDR4-2133 (all 8 channels full) - GPU: 2x RTX 3090 (48 GB total, PCIe 3.0) - 1500 W PSU What I run: - Qwen3-Flash-Next (177B total / ~6B active MoE, IQ4_XS) on Ilama.cpp.… 38 r/LocalLLaMA community 15d ago This draft model is OP on 16 GB cards for Qwen 3.8 27b https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF I used this draft model with https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with the IQ3_XXS with 128k context and I saw it averaging about 60 tokens per second tg speed on the 16 GB RX 9070 XT. This… 16 Hugging Face Daily Papers research 16d ago Memory as Plans: World-Action Modeling with Memory-Grounded Planning Abstract MaP-WAM improves non-Markovian robotic manipulation by separating memory-grounded planning from plan-conditioned execution, using compact episodic segment records and progress-calibrated action chunks to maintain fixed inference latency. Generated by… 35 OpenAI official-blog 17d ago Rapidly scaling online storage to serve over 1 billion ChatGPT users Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second. 4 vLLM releases dev-tools 17d ago proto-v0.1.0 vllm-proto 0.1.0 16 arXiv — Machine Learning research 17d ago GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models arXiv:2609.10658v1 Announce Type: new Abstract: Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without… 24 arXiv — Machine Learning research 17d ago Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving arXiv:2609.11133v1 Announce Type: new Abstract: Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found… 18 arXiv — NLP / Computation & Language research 17d ago REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving arXiv:2609.11209v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and… 38 arXiv — Machine Learning research 17d ago Optimizing AI Inference Across the Deployment Stack arXiv:2609.10550v1 Announce Type: cross Abstract: AI deployment performance is shaped not by model architecture alone, but by interactions among compression, compiler transformations, and serving policies. Published benchmarks often report latency and throughput under… 22 arXiv — Machine Learning research 17d ago Adaptive Diffusion Freezing: Privacy-preserving Diffusion Models Against Membership Inference Attacks arXiv:2609.10608v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success in generative tasks across various areas, however their training process raises significant privacy concerns, particularly under membership inference attacks (MIAs). Prior studies… 11 arXiv — Machine Learning research 17d ago CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding arXiv:2609.10744v1 Announce Type: cross Abstract: Autoregressive language models can be used to transform a payload text into a stegotext of identical token length by preserving per-position rank information across contexts - a methodology we formalize as Contextual… 11 arXiv — NLP / Computation & Language research 17d ago Automated Identification of Competing Narratives in Political Discourse on Social Media arXiv:2609.11202v1 Announce Type: new Abstract: Social media platforms have become central to shaping political discourse, serving as arenas where narratives form and evolve, influencing public opinion. Identifying and analyzing these narratives, particularly when they compete… 12 arXiv — NLP / Computation & Language research 17d ago Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study arXiv:2609.11450v1 Announce Type: new Abstract: Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical… 8 arXiv — NLP / Computation & Language research 17d ago LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation arXiv:2609.11739v1 Announce Type: new Abstract: Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training… 26 arXiv — NLP / Computation & Language research 17d ago Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing arXiv:2609.11769v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears more neutral. They do not directly show whether a… 5 Hugging Face Daily Papers research 17d ago HyQuant: Hybrid-Precision Quantization for LLM Attention Abstract HyQuant improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead. Generated by thinkingmachines/Inkling-Small Quantization has been… 38 NVIDIA Developer Blog official-blog 17d ago How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as... 6 NVIDIA Developer Blog official-blog 17d ago High-Throughput Structure Prediction with BioNeMo Inference Runtime Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA... 4 Hugging Face Daily Papers research 18d ago Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning Abstract HybridAL adaptively switches from retraining to fine-tuning during active learning based on online stabilization signals, reducing training time while preserving accuracy and improving calibration. Generated by thinkingmachines/Inkling-Small Training strategy, namely… 6 arXiv — Machine Learning research 18d ago Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations arXiv:2609.05658v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy… 35 arXiv — Machine Learning research 18d ago Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression arXiv:2609.05688v1 Announce Type: new Abstract: We study variance-preserving diffusion of the response in mixed linear regression (MLR) with unknown mixing weights. Our analysis separates the statistical guarantees of score matching from the loss geometry and optimization signal… 29 arXiv — Machine Learning research 18d ago LoGIC: Budgeted Context Construction for Node-Level Graph In-Context Learning with Tabular Foundation Models arXiv:2609.05955v1 Announce Type: new Abstract: Tabular foundation models have become powerful graph learners. Systems such as G2T-FM and GraphPFN encode each node as a feature row and make predictions through in-context learning (ICL), with labeled rows serving as the prompt.… 28 arXiv — Machine Learning research 18d ago FANS: Federated Adaptive Network Search Learning for Heterogeneous Devices arXiv:2609.06106v1 Announce Type: new Abstract: Heterogeneous Federated Learning (HFL) aims to train models across devices with diverse resource budgets while preserving data privacy. Existing HFL methods typically bind training to a small predefined menu of model… 31 arXiv — Machine Learning research 18d ago When Retain Constraints Conflict: Mitigating Forget-Retain Interference in Tabular Data arXiv:2609.06786v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of designated training data while preserving model utility, but its behavior on tabular data remains underexplored. This gap is important because tabular prediction is widely used in… 31 arXiv — NLP / Computation & Language research 18d ago Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training arXiv:2609.10052v1 Announce Type: new Abstract: LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study… 13 arXiv — NLP / Computation & Language research 18d ago KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints arXiv:2609.10266v1 Announce Type: new Abstract: LLM serving systems already reuse KV caches, but only when the reused text sits at the very start of the prompt. Two growing workloads break this condition: a retrieval-augmented generation server assembles a different set of… 29 arXiv — NLP / Computation & Language research 18d ago IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier arXiv:2609.10494v1 Announce Type: new Abstract: Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as… 36 arXiv — NLP / Computation & Language research 18d ago Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts arXiv:2609.09241v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection… 16 r/LocalLLaMA community 18d ago Running qwen 3.8 27B iq3 xxs on RTX 3060. Getting anywhere from 10 - 20 tps. Thinking Off . Took about 4 mins and 7 mins. Running on about "IQ3_S - 3.4375 bpw" 37.03.960.932 I slot print_timing: id 0 | task 2665 | prompt processing, n_tokens = 12516, progress = 0.98, t = 37.08 s / 337.55 tokens per second 37.04.777.411… 31 NVIDIA Developer Blog official-blog 18d ago When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill... 27 r/LocalLLaMA community 18d ago Solved: LLM inference on Windows was 2–3x slower when the server window wasn't focused The fix: run the server detached/headless instead of keeping it attached to a console window. RTX 5090, ~27B NVFP4 model via ninfer: Terminal focused: 130–200 tok/s Terminal unfocused: 50–60 tok/s Click the terminal → immediately back to 130+ tok/s At first I thought GPU… 17 r/LocalLLaMA community 18d ago 1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install) mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for a while: a 27B one-bit model answering at up to 30 tokens/s on an RTX 3060 Laptop GPU with 6 GB of VRAM, in Chrome, from a web… 22 r/LocalLLaMA community 19d ago DeepSeek-V4-Flash-Vision-Exp (285B MoE) on 10-12x RTX 3090 — spec decoding, vision Running the full deepseek-ai/DeepSeek-V4-Flash-Vision-Exp on consumer Ampere — 10-12x RTX 3090, SM86-compatible vLLM build. 285B MoE, FP4 experts + FP8 attention, 157 GB weights. Highlights: - **60+ tok/s** decode, DSpark spec (k=3) on 10 GPUs (TP2xPP5), at a 240 W cap - **120+… 17 Hugging Face Daily Papers research 19d ago Reason Through the Latent! Making Latent Visual Reasoning Necessary Abstract CVRR enforces recurrent hidden-state computation as the required image-conditioned pathway for visual reasoning, preserving model competence while distinguishing latent information from actual predictive use. Generated by thinkingmachines/Inkling-Small Latent visual… 35 r/LocalLLaMA community 19d ago Is RX 6800 + 6800 XT a sensible upgrade from 2x RTX 2060 OC 12GB for llama.cpp? I’m currently running llama.cpp on two RTX 2060 12GB cards, so 24GB total VRAM. With Qwen3.8 27B IQ4_XS at 131k context I’m getting around 45 tok/s, which is actually pretty good for this setup. I found a deal on an RX 6800 16GB and an RX 6800 XT 16GB, so I’d be going from 24GB… 8 r/LocalLLaMA community 20d ago For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput I've been making a lot of comments about optimal setup for Strix Halo (gfx1150) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware… 20 r/LocalLLaMA community 20d ago Higher acceptance length, slower prose: Ling’s n=1/2/3 MTP test on one Spark The missing control is visible in sudoingX’s Ling-3.0-flash benchmark graphics. The earlier table leaves Ling’s no-speculation baseline as “not measured.” The later code/prose graphic fills it in: about 23 tok/s without the drafter, against 40.9 on code and 38.7 on prose with… 31 r/LocalLLaMA community 21d ago Qwen 3.8 Next Flash is really really REALLY verbose.. Long time user of 3.6 27b, switched over to Next Flash since it's a logical step up even from 3.8 27b. It's soooo verbose, i'm talking 13 minutes of thinking time on single turn coding requests at approximately 150 tokens per second tg and 7000 tokens per second pp. It's… 24 arXiv — NLP / Computation & Language research 21d ago When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models arXiv:2609.04453v1 Announce Type: cross Abstract: Expert pruning reduces the memory and serving cost of Mixture-of-Experts (MoE) models by removing low-importance experts identified by the router, assuming router probabilities provide a reliable importance signal. We observe… 17 arXiv — Machine Learning research 21d ago RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments arXiv:2609.05403v1 Announce Type: new Abstract: Retail search systems serve diverse geographic regions with distinct query patterns, vocabularies, and product preferences, creating significant data heterogeneity that challenges both privacy-preserving training and model… 22 arXiv — Machine Learning research 21d ago Interface-Induced Trajectory Censoring arXiv:2609.03966v1 Announce Type: cross Abstract: Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's… 38 arXiv — NLP / Computation & Language research 21d ago Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs arXiv:2609.04526v1 Announce Type: new Abstract: Merging a LoRA adapter into its base model is standard deployment practice: it removes the runtime adapter's per-forward overhead and leaves a single standalone checkpoint any serving stack can load. On a native 4-bit microscaling… 27 arXiv — NLP / Computation & Language research 21d ago EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages arXiv:2609.05043v1 Announce Type: new Abstract: Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages, but it can distort task-critical constraints and required outputs, creating corrupted training examples and degrading… 17 r/LocalLLaMA community 21d ago LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM) LayerStoRm: GLM-5.3-Flash UD-Q4_K_XL (186 GiB) at 1M context on 2× RTX 5090 + 2× RTX 5080 (96 GB VRAM total) using RAM for the pinned experts. LayerStoRm is a (still experimental) MIT-licensed continuous expert-streaming inference engine: it runs MoE models far larger than your… 33 Page 3 of 10 · 500 articles ← Newer Older →