News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — NLP / Computation & Language research 10d ago LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning arXiv:2608.01328v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly… 10 arXiv — NLP / Computation & Language research 10d ago Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding arXiv:2608.01560v1 Announce Type: new Abstract: Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding… 29 Hugging Face Daily Papers research 10d ago UEmbed: Unified Sparse and Dense Multimodal Embeddings Abstract Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to… 5 Hugging Face Daily Papers research 10d ago WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Abstract Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or… 24 Hugging Face Daily Papers research 10d ago Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Abstract Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the… 28 Hugging Face Daily Papers research 10d ago VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Abstract Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and… 33 r/MachineLearning community 10d ago I created an autonomous boxing benchmark [D] I created an AI boxing match to test the decision speed, adaptability and strategy. I fed the LLMs with data about the current match and if they have vision, they will get even more data. The match has street rules, anything goes and an AI is not defeated until the ref counts to… 24 Hugging Face Daily Papers research 11d ago N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Abstract We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on… 31 Smol AI News news-outlet 11d ago Qwen 3.8 Max **Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with… 21 arXiv — Machine Learning research 11d ago Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift arXiv:2607.28696v1 Announce Type: new Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. The affected class varies with acquisition protocol and backbone… 9 arXiv — Machine Learning research 11d ago MMFGU: Multimodal Federated Graph Unlearning arXiv:2607.28708v1 Announce Type: new Abstract: Multimodal federated graph learning enables clients to collaboratively train graph models over structural, textual, and visual signals without sharing private local data. However, the presence of heterogeneous multimodal content… 32 arXiv — Machine Learning research 11d ago Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds arXiv:2607.28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to "reflect," yet whether this resembles human revision remains… 34 arXiv — Machine Learning research 11d ago Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation arXiv:2607.29494v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense teacher supervision along student-generated trajectories, but its online rollout process incurs substantial computational cost, particularly when a few long responses delay batch… 6 arXiv — Machine Learning research 11d ago DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs arXiv:2607.28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak. We present DeltaServe, a host-agnostic co-serving design that converts this… 20 arXiv — NLP / Computation & Language research 11d ago ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification arXiv:2607.28637v1 Announce Type: new Abstract: This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subtasks: binary hate speech classification and three-class sentiment analysis. Our… 33 arXiv — NLP / Computation & Language research 11d ago TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal… 6 arXiv — NLP / Computation & Language research 11d ago BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning arXiv:2607.28966v1 Announce Type: new Abstract: Large language models often improve task performance by generating long reasoning traces, but the resulting computation is frequently wasted on redundant verification and revision. Existing probe-based early-exit approaches mainly… 4 arXiv — NLP / Computation & Language research 11d ago Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models arXiv:2607.29079v1 Announce Type: new Abstract: Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time consistency problem on 300 real images, comparing… 8 arXiv — NLP / Computation & Language research 11d ago Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding arXiv:2607.29196v1 Announce Type: new Abstract: Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identifying intended objects or referents, and withholding… 19 arXiv — NLP / Computation & Language research 11d ago Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks arXiv:2607.29585v1 Announce Type: new Abstract: To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; participants who are epistemically vigilant detect when new information conflicts… 8 arXiv — NLP / Computation & Language research 11d ago FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models arXiv:2607.29602v1 Announce Type: new Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic… 5 arXiv — NLP / Computation & Language research 11d ago Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods… 26 arXiv — NLP / Computation & Language research 11d ago WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates… 34 r/LocalLLaMA community 11d ago MiniMax-H3 now on huggingface MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks… 16 Hugging Face Daily Papers research 11d ago N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Abstract We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current… 24 Hugging Face Daily Papers research 11d ago RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models Abstract Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection… 36 Vercel — AI dev-tools 12d ago Qwen 3.8 Max now available on Vercel AI Gateway Qwen 3.8 Max is now available on AI Gateway. Qwen 3.8 Max handles text-only and vision-language work in one model, with 2.4 trillion parameters and a context window of up to 1 million tokens. The model is suited for software engineering and office productivity, along with visual… 13 r/MachineLearning community 13d ago What should we do for EMNLP commitment deadline? [R] We received the reviews, but they don't mention whether we should submit a revised version. Should we prepare one? I also couldn't find anywhere to upload a revision. What exactly is the EMNLP commitment deadline? I had assumed we were supposed to upload an updated version. Do… 22 TechCrunch — AI news-outlet 13d ago Siri AI could come with a paywall for power users Apple CEO Tim Cook envisions users being able to buy more compute for Siri AI via Apple's existing iCloud+ subscriptions. 28 Hugging Face Daily Papers research 14d ago Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Abstract Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes… 14 Hugging Face Daily Papers research 14d ago ReToken: One Token to Improve Vision-Language Models for Visual Retrieval Abstract Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding… 22 Hugging Face Daily Papers research 14d ago See2Think: Do Multimodal Models Really Use Intermediate Visual States? Abstract Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage… 9 arXiv — Machine Learning research 14d ago Regularizing modality contribution drift in multimodal continual learning arXiv:2607.27260v1 Announce Type: new Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge. To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic… 36 arXiv — Machine Learning research 14d ago Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision arXiv:2607.27274v1 Announce Type: new Abstract: EEG-based disease diagnosis requires one prediction per subject, yet common pipelines segment recordings into short instances, inherit the subject label for every instance, and train instance-level classifiers. This assumes that… 10 arXiv — Machine Learning research 14d ago TIER-MoE: Trust-Informed Expert Routing via Conditional Modality Risk for Multimodal Fusion in Biomedical Classification arXiv:2607.27289v1 Announce Type: new Abstract: The promise of multimodal fusion lies in combining complementary sources of evidence, yet more evidence does not always yield a better prediction. Recent multimodal models have advanced fusion through richer cross-modal interaction… 12 arXiv — Machine Learning research 14d ago Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models arXiv:2607.27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral… 14 arXiv — Machine Learning research 14d ago Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective arXiv:2607.27660v1 Announce Type: new Abstract: Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning. In particular, the SCORE framework~\cite{majee2024score} demonstrated that SIMs can serve as… 13 arXiv — Machine Learning research 14d ago FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning arXiv:2607.27665v1 Announce Type: new Abstract: Federated graph learning enables collaborative training over decentralized graph data without sharing raw graph information. As such risks evolve, clients must learn emerging classes from private multimodal graph streams, retain… 26 arXiv — Machine Learning research 14d ago Flux-OPD: On-Policy Distillation with Evolving Contexts arXiv:2607.28022v1 Announce Type: new Abstract: Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision… 13 arXiv — Machine Learning research 14d ago Contrastive Reinforced Policy Optimization via Privileged Self-Distillation arXiv:2607.28026v1 Announce Type: new Abstract: Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it… 12 arXiv — Machine Learning research 14d ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger arXiv:2607.28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate… 18 arXiv — NLP / Computation & Language research 14d ago AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes arXiv:2607.27393v1 Announce Type: new Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has… 4 arXiv — NLP / Computation & Language research 14d ago Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built… 36 arXiv — NLP / Computation & Language research 14d ago Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or… 29 arXiv — NLP / Computation & Language research 14d ago Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis arXiv:2607.27790v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal modalities share a structural isomorphism with natural language, as both can be… 23 arXiv — NLP / Computation & Language research 14d ago AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying… 28 arXiv — NLP / Computation & Language research 14d ago RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning arXiv:2607.28156v1 Announce Type: new Abstract: Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When… 18 arXiv — NLP / Computation & Language research 14d ago Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models arXiv:2607.28166v1 Announce Type: new Abstract: Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide… 28 arXiv — NLP / Computation & Language research 14d ago Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory arXiv:2607.28263v1 Announce Type: new Abstract: Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for prediction. We turn this division of labor into CoMem (Comprehension Memory),… 13 arXiv — NLP / Computation & Language research 14d ago Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the… 35 Page 5 of 10 · 500 articles ← Newer Older →