News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — NLP / Computation & Language research 14d ago Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of qualified specialists, limited service reach beyond urban centers, and a… 26 arXiv — NLP / Computation & Language research 14d ago VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation arXiv:2607.28590v1 Announce Type: cross Abstract: Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual… 4 Hugging Face Daily Papers research 14d ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Abstract Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct… 37 Hugging Face Daily Papers research 14d ago Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Abstract GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI… 31 Hugging Face Daily Papers research 14d ago Beacon: Knowing When and How to Perform Agentic Visual Reasoning Abstract The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic… 14 Hugging Face Daily Papers research 14d ago Flux-OPD: On-Policy Distillation with Evolving Contexts Abstract Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student,… 24 r/LocalLLaMA community 14d ago Minimax-H3 video model released, open weights coming in the next few days https://x.com/MiniMax_AI/status/2083006198828417501?s=20 Quote from their article: Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound,… 12 Hugging Face Daily Papers research 14d ago Metis: Memory Foundation Model Abstract Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external… 7 Hugging Face Daily Papers research 14d ago SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them Abstract Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason… 14 r/LocalLLaMA community 14d ago Smallest model (& tips) for intelligent computer use via Hermes? Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive an actual machine via hermes' computer_use tool and cua_driver to click through… 7 r/LocalLLaMA community 15d ago GLM 5.2 with vision on Hugging Face Hi all, I have not seen this model talked about here but it seems like baseten (inference provider on OpenRouter) merged the vision encoder from Kimi k2.6 into GLM 5.2. I think the lack of vision was one of the big complaint when GLM 5.2 came out, I have not tested this model… 34 Hugging Face Daily Papers research 15d ago CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition Abstract Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical… 14 arXiv — NLP / Computation & Language research 15d ago Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement arXiv:2607.26473v1 Announce Type: cross Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic… 14 arXiv — Machine Learning research 15d ago What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations arXiv:2607.27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this?… 12 arXiv — Machine Learning research 15d ago A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment arXiv:2607.26170v1 Announce Type: cross Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background… 21 arXiv — NLP / Computation & Language research 15d ago Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs arXiv:2607.26355v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to perpetuate social biases and reinforce stereotypes. In this study, we… 38 arXiv — NLP / Computation & Language research 15d ago Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation arXiv:2607.26555v1 Announce Type: new Abstract: Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplified by imbalanced data and semantically inconsistent text-image pairs that make… 13 arXiv — NLP / Computation & Language research 15d ago Metis: Memory Foundation Model arXiv:2607.26760v1 Announce Type: new Abstract: Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still… 35 arXiv — NLP / Computation & Language research 15d ago Dual-Path LLM Reasoning for Multimodal Few-Shot Knowledge Graph Completion arXiv:2607.26909v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing facts in knowledge graphs (KGs), thereby improving their completeness and supporting downstream intelligent applications. However, emerging entities and relations in real-world… 36 arXiv — NLP / Computation & Language research 15d ago DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search arXiv:2607.27178v1 Announce Type: new Abstract: State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to… 21 arXiv — NLP / Computation & Language research 15d ago Hearsay: Vision-Language Medical Diagnoses Without an Image arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is… 22 arXiv — NLP / Computation & Language research 15d ago CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators arXiv:2605.08334v2 Announce Type: replace Abstract: We present CustomerSim, an environment and benchmark to evaluate the extent to which Multimodal Large Language Models (MLLMs) can simulate realistic, persona-driven customer behavior in chat-based retail environments. While… 27 Hugging Face Daily Papers research 15d ago TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Abstract Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs… 36 Hugging Face Daily Papers research 15d ago HumanCLAW: Can Vision-Language Models Act Through a Body? Abstract Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller… 19 Vercel — AI dev-tools 15d ago Introducing Enterprise Flexible Commitment for Vercel Marketplace Enterprise customers can now apply a portion of their Flexible Commitment toward eligible resources purchased through the Vercel Marketplace. With Flex Commit support, eligible Marketplace cost can draw directly from your existing commitment making it easier to provision the… 38 Hacker News — AI on Front Page community 15d ago The coolest use for the Vision Pro Article URL: https://christianselig.com/2026/07/vision-pro-house/ Comments URL: https://news.ycombinator.com/item?id=49102774 Points: 221 # Comments: 97 7 MIT News — AI research 15d ago How a medical database developed at MIT evolved into a global standard of data-sharing The visionary PhysioNet platform launched 25 years ago, based on a system developed at MIT in the 1970s. It has become one of the most comprehensive biomedical and clinical data repositories in existence. 17 Hugging Face Daily Papers research 16d ago PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Abstract We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with… 34 r/LocalLLaMA community 16d ago I built a GBNF grammar compiler that makes 8B models reliably call tools - here's how it works (deep dive) I've been building a local agent in Rust (Eris) that runs on llama.cpp and uses an Obsidian-compatible vault as memory. ~50 tools (vault read/write, memory, reminders, web fetch, email, calendar, vision). The biggest pain was getting small models to emit valid tool-calling JSON.… 38 Hugging Face Daily Papers research 16d ago MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Abstract Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are… 31 Hugging Face Daily Papers research 16d ago Pass the Baton: Trajectory-Relayed On-Policy Distillation Abstract On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected… 29 Hugging Face Daily Papers research 16d ago Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking Abstract Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without… 28 arXiv — NLP / Computation & Language research 16d ago MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To… 33 arXiv — NLP / Computation & Language research 16d ago MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice arXiv:2607.25667v1 Announce Type: new Abstract: Psychotherapists need repeated training and supervision by experts; however, scalability is problematic. Here we present MyMentorLLM, a multimodal voice- and text-based simulation environment for deliberate practice, used to… 20 arXiv — NLP / Computation & Language research 16d ago Shieldstral arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety… 19 arXiv — NLP / Computation & Language research 16d ago Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases arXiv:2607.25933v1 Announce Type: new Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating… 32 arXiv — NLP / Computation & Language research 16d ago Pass the Baton: Trajectory-Relayed On-Policy Distillation arXiv:2607.26057v1 Announce Type: new Abstract: On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this… 38 arXiv — NLP / Computation & Language research 16d ago VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents arXiv:2607.24748v1 Announce Type: cross Abstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures. This work studies multimodal… 36 arXiv — NLP / Computation & Language research 16d ago Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model arXiv:2607.24904v1 Announce Type: cross Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an… 6 arXiv — NLP / Computation & Language research 16d ago CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition arXiv:2607.25294v1 Announce Type: cross Abstract: Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus… 4 arXiv — NLP / Computation & Language research 16d ago Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than… 34 arXiv — NLP / Computation & Language research 16d ago Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications arXiv:2607.25642v1 Announce Type: cross Abstract: Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward… 32 arXiv — NLP / Computation & Language research 16d ago A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series arXiv:2607.25947v1 Announce Type: cross Abstract: Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable… 26 arXiv — NLP / Computation & Language research 16d ago VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation arXiv:2510.09733v2 Announce Type: replace Abstract: Visual Retrieval-Augmented Generation (VRAG) has emerged as a promising paradigm for equipping Vision-Language Models (VLMs) with external visual evidence, enabling them to go beyond parametric knowledge when answering visually… 36 Hugging Face Daily Papers research 16d ago Shieldstral Abstract We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content… 14 Hugging Face Daily Papers research 16d ago Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Abstract Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation… 14 Hugging Face Daily Papers research 16d ago FilmBench: A Film-Grade Benchmark for Cinematic Video Generation Abstract Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally,… 17 Hugging Face Daily Papers research 16d ago GNM Head: A Generative aNthropometric Model of the human head Abstract Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight… 22 r/MachineLearning community 16d ago How to deal with text only vector search across multimodal embedding space? [D] My data set is a list of images, each equipped with a a couple sentences of text. A user would search primarily with text only. My default approach is using BM25, but how would I facilitate searching with a vector DB and a model that embeds vectors in a multimodal combined… 37 r/LocalLLaMA community 16d ago microsoft/Mage-VL · Hugging Face - An Efficient Codec-Native Streaming Multimodal Foundation Model Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale. It targets a modern Moravec's paradox of VLMs — strong at complex offline reasoning, yet… 12 Page 6 of 10 · 500 articles ← Newer Older →