News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — NLP / Computation & Language research 17d ago MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions arXiv:2609.11322v1 Announce Type: new Abstract: Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack… 38 arXiv — NLP / Computation & Language research 17d ago ReGround: Grounding Reviewer Comments in Multimodal Evidence arXiv:2609.11460v1 Announce Type: new Abstract: Reviewer comments naturally relate to specific parts of the reviewed paper, yet grounding these comments to the underlying evidence is difficult due to long multimodal documents. Existing benchmarks do not capture this setting and… 8 arXiv — NLP / Computation & Language research 17d ago BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation arXiv:2609.10815v1 Announce Type: cross Abstract: Police body-worn camera (BWC) footage has emerged as a critical aspect of law enforcement that ensures legal transparency, officer accountability, and the protection of civil rights. However, effectively processing this data… 17 arXiv — NLP / Computation & Language research 17d ago New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models arXiv:2609.11022v1 Announce Type: cross Abstract: A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial… 18 arXiv — NLP / Computation & Language research 17d ago A Short Survey of Viewing Large Language Models in Legal Aspect arXiv:2303.09136v2 Announce Type: replace Abstract: Large language models (LLMs) have transformed many fields, including natural language processing, computer vision, and reinforcement learning. These models have also made a significant impact in the field of law, where they are… 27 arXiv — NLP / Computation & Language research 17d ago Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection arXiv:2510.15685v2 Announce Type: replace Abstract: This paper investigates the use of an LLM to generate auxiliary background context for social media posts, and explores four methods to incorporate this context into the input of an SBERT-based Hate Speech Detection (HSD)… 28 arXiv — NLP / Computation & Language research 17d ago Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive Factors arXiv:2511.17036v2 Announce Type: replace Abstract: Visual persuasion uses images to shape cognition, emotion, and behavior, with its effects depending on both visual attributes and semantic context. Despite recent progress, it remains unclear whether Vision-Language Models… 38 Hugging Face Daily Papers research 17d ago SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Abstract Large vision-language models trained on synthetic block-manipulation tasks improve 3D spatial reasoning and generalize to real-world visual tasks. Generated by thinkingmachines/Inkling-Small Large Vision-Language Models (LVLMs) have achieved strong performance on… 25 Hugging Face Daily Papers research 17d ago SenseNova-U1.5: Towards Native Unified Visual Intelligence Abstract SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, expert optimization, and… 30 Hugging Face Daily Papers research 17d ago CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation Abstract A unified vision-language model for coronary angiography uses chain-of-box reasoning and reinforcement learning with verifiable rewards to provide auditable diagnoses and improve zero-shot report generation. Generated by thinkingmachines/Inkling-Small Invasive coronary… 36 llama.cpp releases dev-tools 17d ago b10896 spec: fix failed to decode mtmd chunk with DFlash ( #28587 ) speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter… 34 Hugging Face Daily Papers research 18d ago OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution Abstract OracleZoom improves recursive super-resolution by combining trajectory-based training with cross-scale supervision and a latent prior to reduce hallucinations at extreme magnifications. Generated by thinkingmachines/Inkling-Small Recursive Super-Resolution (SR) extends… 33 r/LocalLLaMA community 18d ago Running Vision Qwen 3.8 27B on a 16GB Card, the config (45tks). I am just sharing my config for Qwen 3.8 27b that fits on a 5060TI, what is cool about this is that you can even get vision! and a 85K context (I have 1.5gb of headroom for more context or a better quant) Model: IQ3_XXS-mtp from… 20 r/LocalLLaMA community 18d ago DeepSeek V4-1 Flash is out Here we go again, DeepSeek is back again with a new model V4-1 Flash A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens Market crash as a service   submitted by   /u/tiguidoio [link]  … 27 r/LocalLLaMA community 18d ago DeepSeek V4.1 Flash: Stronger, Faster, More Accessible Original Source from DeepSeek WeChat Official Account: https://mp.weixin.qq.com/s/qg0NU3NNUbp1co2PdkAPAg Today we're officially releasing the DeepSeek V4.1 Flash model. It is the smallest model in our brand-new model architecture series, with native multimodal visual… 20 arXiv — Machine Learning research 18d ago Grounded and Faithful P&ID Reasoning: Constraining Vision-Language Models with Recovered Evidence Graphs arXiv:2609.05880v1 Announce Type: new Abstract: Piping and Instrumentation Diagrams (P&IDs) are the authoritative maps of process plants: isolation, maintenance, and HAZOP decisions depend on what connects to what. Vision-language models describe these sheets fluently, yet they… 36 arXiv — Machine Learning research 18d ago DART: Distributional Adversarial Recurrent Training for Algorithm Learning arXiv:2609.05988v1 Announce Type: new Abstract: Recurrent reasoning models (RRMs) can solve structured problems, achieving easy-to-hard generalization through iterative computation in hidden space. These models are typically trained with instance-level supervision, which becomes… 29 arXiv — Machine Learning research 18d ago VERPO: Verified Evidence Regularized Policy Optimization arXiv:2609.06100v1 Announce Type: new Abstract: Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by… 33 arXiv — Machine Learning research 18d ago Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier arXiv:2609.06590v1 Announce Type: new Abstract: Mixture of Prompt Experts (MoPE) adapts multimodal transformers through input-dependent prompt composition, while retaining a fixed prompt length. We investigate a layer-wise gating extension in a binary chest X-ray classification… 38 arXiv — Machine Learning research 18d ago Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach arXiv:2609.06668v1 Announce Type: new Abstract: Multimodal graphs couple node attributes in different modalities, such as text and images, with relational structure, enabling topological structure and cross-modality attributes to be modeled jointly. Multimodal graph foundation… 29 arXiv — NLP / Computation & Language research 18d ago SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia arXiv:2609.09672v1 Announce Type: new Abstract: The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA)… 22 arXiv — NLP / Computation & Language research 18d ago $S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants arXiv:2609.09852v1 Announce Type: new Abstract: The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance… 5 arXiv — NLP / Computation & Language research 18d ago Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training arXiv:2609.10052v1 Announce Type: new Abstract: LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study… 13 arXiv — NLP / Computation & Language research 18d ago Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection arXiv:2609.10244v1 Announce Type: new Abstract: We present our system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. A small ($4$B-parameter) VLM is fine-tuned as a per-token classifier reading a two-token feature from its own hidden… 12 arXiv — NLP / Computation & Language research 18d ago On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data arXiv:2609.10321v1 Announce Type: new Abstract: Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from the teacher… 15 arXiv — NLP / Computation & Language research 18d ago MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads arXiv:2609.09206v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention… 19 arXiv — NLP / Computation & Language research 18d ago LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios arXiv:2609.09790v1 Announce Type: cross Abstract: Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities.… 18 arXiv — NLP / Computation & Language research 18d ago From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning arXiv:2609.10335v1 Announce Type: cross Abstract: Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often… 38 Hugging Face Daily Papers research 18d ago Show-Harness: Just a VLM Agent Can Play Robots Abstract Show-Harness links vision-language models to robot control via discrete semantic actions interpreted by embodiment-specific modules, enabling zero-shot and efficient fine-tuned deployment across robots and GUIs. Generated by thinkingmachines/Inkling-Small Foundation… 6 NVIDIA Developer Blog official-blog 18d ago When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill... 27 Hugging Face Daily Papers research 18d ago VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification Abstract VDiff-Bench evaluates multimodal language models on fine-grained image difference identification, revealing major weaknesses in detecting subtle low-level visual changes. Generated by thinkingmachines/Inkling-Small Multimodal Large Language Models (MLLMs) perform… 7 Hugging Face Daily Papers research 18d ago NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting Abstract NOAH is a generative transformer that models full multimodal patient journeys with continuous time dynamics and stochastic latent states, enabling forecasting, zero-shot classification, and counterfactual simulation across diverse clinical data. Generated by… 11 Hugging Face Daily Papers research 18d ago TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model Abstract TANGO is a vision-language framework that predicts whole-body joint actions for humanoid robots to navigate cluttered indoor environments using only simulated training data. Generated by thinkingmachines/Inkling-Small We study the problem of navigating cluttered indoor… 36 Hugging Face Daily Papers research 18d ago Learning 3D Editing without Paired Supervision via Generative Prior Distillation Abstract A feed-forward 3D editing framework distills visual, semantic, and geometric priors from foundation models via differentiable rendering and 3D-aware distribution matching to avoid paired training data. Generated by thinkingmachines/Inkling-Small Instruction-guided 3D… 18 Hugging Face Daily Papers research 18d ago SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation Abstract SynthGait-19k is a large synthetic video dataset for gait analysis that enables benchmarking of video-based gait estimation and shows synthetic supervision transfers to real data. Generated by thinkingmachines/Inkling-Small Accurate estimation of clinically meaningful… 9 r/LocalLLaMA community 19d ago DeepSeek-V4-Flash-Vision-Exp (285B MoE) on 10-12x RTX 3090 — spec decoding, vision Running the full deepseek-ai/DeepSeek-V4-Flash-Vision-Exp on consumer Ampere — 10-12x RTX 3090, SM86-compatible vLLM build. 285B MoE, FP4 experts + FP8 attention, 157 GB weights. Highlights: - **60+ tok/s** decode, DSpark spec (k=3) on 10 GPUs (TP2xPP5), at a 240 W cap - **120+… 17 Hugging Face Daily Papers research 19d ago Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy Abstract Adding Greek to a robot vision-language-action model via machine-translated instructions reveals measurement pitfalls and shows bilingual training improves performance over monolingual baselines, though it remains far below English levels. Generated by… 35 Hugging Face Daily Papers research 19d ago RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? Abstract RoboSPA is a large-scale robotic manipulation benchmark that evaluates vision-language-action models on fine-grained spatial reasoning and long-horizon procedural planning across progressively harder task variants. Generated by thinkingmachines/Inkling-Small… 25 Hugging Face Daily Papers research 19d ago AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Abstract AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.… 37 Hugging Face Daily Papers research 19d ago SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution Abstract SceneMosaic combines learned image priors with vision-language agents to efficiently generate diverse, physically valid indoor scenes by evolving local units and composing them globally. Generated by thinkingmachines/Inkling-Small Diverse and simulation-ready indoor… 36 Hugging Face Daily Papers research 19d ago DriveZero: End-to-End Driving Beyond Human Demonstrations Abstract DriveZero is an end-to-end autonomous driving system that combines a vision foundation model for perception with a closed-loop reinforcement learning action model to learn driving behaviors beyond human demonstrations. Generated by thinkingmachines/Inkling-Small Most… 8 r/LocalLLaMA community 19d ago Best vision models under 6B? Looking into tiny models that have some kind of vision Smart for General stuff + Vision, something like Qwen 3.5 4B or MiniCPM-V-4.6   submitted by   /u/FerLuisxd [link]   [comments] 31 r/LocalLLaMA community 19d ago DeepSeek Flash 4.1 is already being tested via API and rolling out. Translation: "Internal beta testing for an intermediate version of DeepSeek V4.1 Flash is now open; you are welcome to try it out. It adopts a new model architecture featuring native multimodal support, stronger capabilities, faster speeds, and lower costs. Keep your base_url… 25 Hugging Face Daily Papers research 20d ago EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents Abstract EmbodiedSkills proposes a unified framework that validates and verifies robot skill executions through a fixed interface, enabling closed-loop embodied agents with adaptable low-level vision-language-action policies. Generated by thinkingmachines/Inkling-Small… 8 Hugging Face Daily Papers research 20d ago What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation Abstract Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies… 36 r/MachineLearning community 20d ago I reduced image-processing token usage by ~95% compared with GPT-4o direct vision, while maintaining roughly the same accuracy.How significant is that?[P] I'm testing a new approach for reducing the cost of image-based LLM inference. I evaluated it on the MOMA Graph benchmark , using 1,315 questions . Compared with using GPT-4o to process the original images directly, I observed approximately: ~95% lower token usage roughly the… 18 r/LocalLLaMA community 20d ago DeepSeek-V4-Flash-Vision-Exp is amazing at creating game worlds! Model: DeepSeek-V4-Flash-Vision-Exp (local and API when impatient) Time: about one weekend (2 days) of QA and small improvements Full game is here After Qwen3.8-Flash-Next one-shotted a really cool Cat-Hunt game demo, I decided to see what the new DeepSeek vision model can do.… 30 Hugging Face Daily Papers research 21d ago The Attention Triangle in Audio-Video Models Abstract Audio-video diffusion models exhibit bidirectional semantic leakage through cross-modal attention pathways, which can be diagnosed via attention-derived signals and mitigated through inference-time alignment interventions. Generated by thinkingmachines/Inkling-Small… 23 Hugging Face Daily Papers research 21d ago RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Abstract RISE improves language model post-training by recursively generating dense token-level supervision from the model's own reinforcement learning trajectory via self-extrapolation, avoiding external teachers. Generated by thinkingmachines/Inkling-Small On-policy… 9 arXiv — Machine Learning research 21d ago A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision arXiv:2609.04754v1 Announce Type: new Abstract: The Duckworth-Lewis-Stern (DLS) method has been the international standard for revising target scores in rain-interrupted limited-overs cricket since 1999. Despite over two decades of operational use, no large-scale empirical audit… 5 Page 5 of 10 · 500 articles ← Newer Older →