News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow r/LocalLLaMA community 1h ago DeepSeek-V4-Flash-Vision Q8 vs Qwen3.8-Flash-Next Q8 I'm using DS-V4-Flash-Vision with Q8_K_XL quantization locally as my everyday engine, and for some time now I've been doing a lot of comparisons with Qwen3.8-Flash-Next, also with Q8_K_XL quantization. It took me quite a while to get Q3.8FN to work reasonably well, and here are… 38 r/LocalLLaMA community 12h ago Villager Simulation Game POC Created with Qwen3.8-27B-UD-Q3_K_XL.gguf - 16GB VRAM https://village-sim-one.vercel.app/ - 16GB VRAM RTX 5070 Ti, fully offloaded - Vision on CPU - Windows, not headless - beellama.cpp - latest version with the kvarn performance enhancements making it as fast as qx_x quants. - MTP n-max = 2 - tg up to 75t/s, pp up to 1700t/s - KV… 14 r/LocalLLaMA community 20h ago M5 Max users: what models are you using & what tk/s are you getting? I was using antirez’s ds4 for a while and getting around 20 tk/s, which worked for my purposes. But I know there have been big advancements between Qwen, the DS4 vision model, and GLM. I’m not sure how the quants affect performance, so what’s the best thing to run right now &… 36 r/LocalLLaMA community 2d ago Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.   submitted by   /u/niacolhealth [link]   [comments] 6 Hugging Face Daily Papers research 2d ago Last Translation Benchmark Abstract The Last Translation Benchmark introduces peer-reviewed, multimodal examples that break leading translation models alongside handcrafted verification rules for reliable, actionable evaluation. Generated by thinkingmachines/Inkling-Small For scientific progress, we need… 22 Hugging Face Daily Papers research 2d ago Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Abstract Puffin-World is a unified multimodal framework that jointly models physics, geometry, and appearance for physically consistent 3D world generation, reconstruction, and closed-loop exploration. Generated by thinkingmachines/Inkling-Small We propose Puffin-World, a… 15 Hugging Face Daily Papers research 2d ago Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding Abstract LatentStream introduces a progressive latent working memory framework that internalizes streaming visual evidence into compact evolving tokens for continuous reasoning. Generated by thinkingmachines/Inkling-Small Streaming video understanding requires multimodal large… 27 arXiv — NLP / Computation & Language research 2d ago Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation arXiv:2609.02998v1 Announce Type: cross Abstract: On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without… 32 arXiv — Machine Learning research 2d ago FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience arXiv:2609.03241v1 Announce Type: new Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or… 36 arXiv — Machine Learning research 2d ago A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds arXiv:2609.03457v1 Announce Type: new Abstract: Accurate cloud resource forecasting is essential for proactive resource provisioning, maintaining Quality of Service (QoS), and reducing operational costs in dynamic cloud environments. The existing forecasting approaches… 38 arXiv — Machine Learning research 2d ago Landmark-Based Discrimination of Injury-Associated Athlete-Sessions from Minute-Resolution Multimodal Football Monitoring Data arXiv:2609.03790v1 Announce Type: new Abstract: Athlete monitoring data may be recorded minute by minute throughout a match or training session, while injury information may only indicate whether the entire session was injury-associated. This creates a modelling problem:… 37 arXiv — NLP / Computation & Language research 2d ago Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards arXiv:2609.03181v1 Announce Type: new Abstract: We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters… 9 arXiv — NLP / Computation & Language research 2d ago What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation arXiv:2609.03254v1 Announce Type: new Abstract: Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must… 29 arXiv — NLP / Computation & Language research 2d ago FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models arXiv:2609.03331v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually… 19 arXiv — NLP / Computation & Language research 2d ago How Far Can Synthetic Data Take Thai OCR? arXiv:2609.03595v1 Announce Type: new Abstract: We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic… 26 arXiv — NLP / Computation & Language research 2d ago KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records arXiv:2609.03597v1 Announce Type: new Abstract: Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten… 24 arXiv — NLP / Computation & Language research 2d ago Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks arXiv:2609.03734v1 Announce Type: new Abstract: BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit… 30 arXiv — NLP / Computation & Language research 2d ago Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR arXiv:2609.04108v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the… 38 arXiv — NLP / Computation & Language research 2d ago Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization arXiv:2609.03158v1 Announce Type: cross Abstract: Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence… 31 arXiv — NLP / Computation & Language research 2d ago MedQA-MM: Shortcuts Behind Medical Visual Reasoning arXiv:2609.03261v1 Announce Type: cross Abstract: A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the… 11 arXiv — NLP / Computation & Language research 2d ago VisCAD: A Foundation Model Suite with Multimodal Industrial CAD Intelligence arXiv:2609.03811v1 Announce Type: cross Abstract: AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level generation maps diverse forms of user intent, including renders, text descriptions, 2D drawings, and real photographs, to… 16 Hugging Face Daily Papers research 2d ago WorldReward: Reward Modeling for Camera-Conditioned World Models Abstract WorldReward is a vision-language reward model that evaluates camera-conditioned world models by aligning video chunks with actions and aggregating preferences for both execution consistency and visual quality. Generated by thinkingmachines/Inkling-Small… 36 Hugging Face Daily Papers research 2d ago Editable Visual Design Abstract A coding agent guided by a vision-language model generates editable layered designs by synthesizing isolated visual assets and iteratively refining native HTML/CSS layouts. Generated by thinkingmachines/Inkling-Small While diffusion base models such as GPT-Image-2 and… 30 Hugging Face Daily Papers research 2d ago LatentPress: Context Compression Beyond Text and Vision Abstract LatentPress compresses conversational and document context into continuous memory tokens read directly by a frozen decoder, achieving high compression with faster inference and improved accuracy over text or OCR methods. Generated by thinkingmachines/Inkling-Small… 5 Hugging Face Daily Papers research 2d ago LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Abstract LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art… 26 r/MachineLearning community 3d ago Mol-JEPA - Multimodal molecular foundation model [R] Hi everyone, I just quickly wanted to share a paper I was working on for around a year now. I created this summary website with key results: https://flogrammer.github.io/moljepa/ TL;DR: its a multimodal JEPA model for molecules. There will be more work to do to improve… 29 Hugging Face official-blog 3d ago NeoMME: an efficient Multimodal-native and Multilingual Encoder Back to Articles a]:hidden"> NeoMME: an efficient Multimodal-native and Multilingual Encoder Team Article Published September 3, 2026 Upvote 8 Tony Wu tonywu71 Hcompany Aurélien Lac h-aurelien-lac Hcompany :last-child]:mb-0"> We introduce NeoMME , a family of 260M and 800M… 8 Hugging Face Daily Papers research 3d ago NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference Abstract NeoMME introduces small bidirectional multimodal encoders pretrained with masked discrete diffusion that achieve strong visual document retrieval and high compression of late-interaction embeddings. Generated by thinkingmachines/Inkling-Small Multimodal models often… 21 Hugging Face Daily Papers research 3d ago FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos Abstract FoldingAgent uses a vision-language model with specialized tools to convert origami videos into executable parametric folding programs via sequential reasoning and physical verification. Generated by thinkingmachines/Inkling-Small We present FoldingAgent, an agentic… 23 vLLM releases dev-tools 3d ago v0.29.0rc2 [Bugfix][Multimodal] Handle prefix-covered items in SHM worker cache … 9 Hugging Face Daily Papers research 3d ago Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering Abstract KBMR uses a multimodal language model to embed images by semantic identity rather than surface appearance, improving retrieval and visual question answering via continuous distillation and hard negative sampling. Generated by thinkingmachines/Inkling-Small… 20 Hugging Face Daily Papers research 3d ago SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions Abstract SnapBench introduces paired corruption benchmarks for mobile snap-and-ask retrieval, revealing that image noise severely degrades multimodal retrieval and proposing an adaptive fusion method to calibrate modality reliability. Generated by thinkingmachines/Inkling-Small… 17 Hugging Face Daily Papers research 3d ago A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss Abstract SimLoss uses embedding-space contrastive supervision to enable single-pass fine-grained image captioning that matches multi-stage quality at much lower latency. Generated by thinkingmachines/Inkling-Small An image may be worth a thousand words, but most captioning… 19 arXiv — Machine Learning research 3d ago DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving arXiv:2609.01609v1 Announce Type: new Abstract: While diffusion models effectively capture multimodal behavioral priors for autonomous driving, offline reinforcement learning (RL) policies remain susceptible to distribution shift, heavy-tailed risk signals, out-of-distribution… 15 arXiv — Machine Learning research 3d ago On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers arXiv:2609.01947v1 Announce Type: new Abstract: Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of examples, constraining supervision to… 13 arXiv — Machine Learning research 3d ago TC-Next: Zero-Shot Multimodal Cyclone Forecasting arXiv:2609.02085v1 Announce Type: new Abstract: We present TropicalCycloneNext (TC-Next), a multimodal deep learning model that forecasts tropical cyclone track and intensity at $6$-$24$ h leads by leveraging a foundation model's forecast fields of atmospheric kinematic and… 38 arXiv — Machine Learning research 3d ago FORGE: Forward-Only Test-Time Adaptation for Integer-Only Vision Models on Microcontrollers arXiv:2609.01683v1 Announce Type: cross Abstract: Vision models deployed on microcontrollers (MCUs) are quantized to integer-only arithmetic and run in inference-only runtimes that do not carry the machinery backpropagation needs: the standard tool for adapting a model to the… 38 arXiv — Machine Learning research 3d ago FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making arXiv:2609.01691v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to make decisions from visual inputs. We introduce FAIRLENS, a benchmark and evaluation framework for measuring both the fairness and the validity of VLM responses in three… 21 arXiv — NLP / Computation & Language research 3d ago MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models arXiv:2609.01772v1 Announce Type: new Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual… 29 arXiv — NLP / Computation & Language research 3d ago Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and Language arXiv:2609.02606v1 Announce Type: new Abstract: Loneliness is a critical public health issue among older adults, linked to higher risks of depression, cognitive decline, and mortality. Scalable, objective methods for its detection remain limited, particularly in natural… 23 arXiv — NLP / Computation & Language research 3d ago Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models arXiv:2609.02082v1 Announce Type: cross Abstract: Visual modality enhances the capabilities of multimodal large language models (MLLMs) but also introduces a safety concern: a benign textual query may convey harmful intent when grounded in a visual image. We term this… 35 arXiv — NLP / Computation & Language research 3d ago EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision arXiv:2609.02133v1 Announce Type: cross Abstract: Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation. We formulate this as response-side affective-orientation control and use… 30 Ollama releases dev-tools 3d ago v0.33.3: gemma4: image and audio input support Safetensors gemma4 imports served by the MLX engine now answer image and audio chats. Images run through both vision architectures: the transformer tower (26B, 31B, e-series) and the 12B's encoder-free unified embedder. Audio arrives through the same intake the ollama API… 22 llama.cpp releases dev-tools 4d ago b10766 model: correctly support input vision for deepseek4 ( #28154 ) model: correctly support input vision for deepseek4 nits Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/44819834 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon… 6 llama.cpp releases dev-tools 4d ago b10762 mtmd: support DeepSeek-V4-Flash-Vision-Exp ( #28133 ) mtmd: support DeepSeek-V4-Flash-Vision-Exp handle min/max token counts from CLI rm debugging use GGML_ROPE_TYPE_VISION nits apply review comments correct token count Website: https://llama.app Attestations:… 29 r/LocalLLaMA community 4d ago Vision support merged for DeepSeek-V4-Flash-Vision-Exp Unsloth GGUFs and Vision support https://huggingface.co/unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF   submitted by   /u/fmillar [link]   [comments] 36 Hugging Face Daily Papers research 4d ago Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry Abstract Credit-addressable reasoning via executable code traces and localized reinforcement learning improves multimodal geometry reasoning by aligning credit assignment with structured reasoning events. Generated by thinkingmachines/Inkling-Small Multimodal geometry reasoning… 5 Hugging Face Daily Papers research 4d ago Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Abstract Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training. Generated by thinkingmachines/Inkling-Small We present Qwen-Drive-1.0,… 29 arXiv — Machine Learning research 4d ago Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy arXiv:2609.00103v1 Announce Type: new Abstract: Memory is widely viewed as an important unsolved problem for LLMs and VLMs, and current benchmarks typically evaluate it by testing accuracy over long text or video. However, accuracy alone misses properties that matter for real… 23 arXiv — Machine Learning research 4d ago CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN arXiv:2609.00590v1 Announce Type: new Abstract: The next generation of mobile networks is envisioned as fully AI-native, with AI-RAN architectures embedding small language models (SLMs) to perform reasoning over real-time telemetry. The state-of-the-art training paradigms for… 26 Page 1 of 10 · 500 articles Older →