News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow Hugging Face Daily Papers research 2d ago On-Policy Self-Distillation without Any Supervision Abstract Unsupervised on-policy self-distillation improves large language models by using internal consistency and majority-vote pseudo-solutions to correct confident errors without external supervision. Generated by thinkingmachines/Inkling-Small On-policy (Self-)Distillation… 34 TechCrunch — AI news-outlet 2d ago General Catalyst leads $1.1B round into 2-month-old River AI River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate. 5 r/LocalLLaMA community 2d ago Revision Prompting: Trades slow (decoded) output tokens for cheap (prefilled) input tokens. TL;DR: If you re-run the same prompt whenever the input changes, try sending the old input/output plus a diff of the input, and ask the model for a patch to the output. You generate ~2-10x fewer output tokens, and the untouched parts of the output stay byte-identical. This… 22 Hugging Face Daily Papers research 2d ago MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Abstract MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities. Generated by… 37 Hugging Face Daily Papers research 3d ago Vision-Language Grounding as Bidirectional Concept Correspondence Abstract ConCor-1 treats vision-language grounding as bidirectional concept correspondence, jointly predicting text spans, image segments, and cross-modal matches without prespecified phrases. Generated by thinkingmachines/Inkling-Small Vision-language grounding connects… 30 Hugging Face Daily Papers research 3d ago What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems Abstract A three-stage multimodal framework improves follow-up edit recommendations in image-creation conversations by combining supervised fine-tuning, multi-objective reinforcement learning, and visual verification. Generated by thinkingmachines/Inkling-Small Conversational… 22 arXiv — Machine Learning research 3d ago CONFER: Conflict-Aware Evidence Negotiation for Regime-Calibrated Weak Supervision in Multimodal Emotion Recognition arXiv:2608.07867v1 Announce Type: new Abstract: Multimodal emotion recognition often treats self-reported labels as reliable supervision while overlooking self-report unreliability and cross-modal conflict. We propose \textbf{CONFER}, a graph-based conflict-aware evidence… 9 arXiv — Machine Learning research 3d ago Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees arXiv:2608.08002v1 Announce Type: new Abstract: Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of… 25 arXiv — Machine Learning research 3d ago The Neural Division of Labor: Biologically-Inspired Modular Architectures for Robust Neuromorphic Computing arXiv:2608.08317v1 Announce Type: new Abstract: Biological neural systems achieve high efficiency and robustness through compartmentalized architectures. In contrast, modern artificial neural networks rely on globally entangled structures, which obscure decision logic and suffer… 7 arXiv — Machine Learning research 3d ago Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles arXiv:2608.08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light… 15 arXiv — Machine Learning research 3d ago Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data arXiv:2608.08859v1 Announce Type: new Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under strict computational and memory constraints. A critical yet underexplored challenge… 19 arXiv — NLP / Computation & Language research 3d ago Unified Hallucination Fuzzing for Multimodal Large Language Models arXiv:2608.07525v1 Announce Type: new Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from… 16 arXiv — NLP / Computation & Language research 3d ago Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards arXiv:2608.07531v1 Announce Type: new Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from… 19 arXiv — NLP / Computation & Language research 3d ago Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation arXiv:2608.07763v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which… 19 arXiv — NLP / Computation & Language research 3d ago Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs arXiv:2608.08090v1 Announce Type: new Abstract: Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clearer understanding of the contribution of translated multilingual supervision. We… 17 arXiv — NLP / Computation & Language research 3d ago NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs arXiv:2608.08107v1 Announce Type: new Abstract: Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this phenomenon from the perspective… 8 arXiv — NLP / Computation & Language research 3d ago VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to a 1.04B Spanish/LATAM security decoder via an MLP. To our knowledge, it is the… 9 arXiv — NLP / Computation & Language research 3d ago From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios arXiv:2608.08510v1 Announce Type: new Abstract: Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This "cocktail party" scenario still… 33 arXiv — NLP / Computation & Language research 3d ago OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for… 17 arXiv — NLP / Computation & Language research 3d ago Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue arXiv:2608.08915v1 Announce Type: new Abstract: Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much… 35 arXiv — NLP / Computation & Language research 3d ago PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial… 8 r/LocalLLaMA community 3d ago I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter MoonViT image encoder, then trained a 40.1M-parameter connector between them on… 16 Hugging Face Daily Papers research 3d ago Evidence-RL: Towards Evidence-intensive Visual Reasoning Abstract Counterfactual Evidence Disentanglement improves vision-language model grounding by auditing whether answers causally depend on local visual evidence during reinforcement learning post-training. Generated by thinkingmachines/Inkling-Small Vision-Language Models (VLMs)… 13 TechCrunch — AI news-outlet 3d ago Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision Meta’s new open-weight Muse Glimmer model offers a glimpse of Mark Zuckerberg’s personal superintelligence vision, as well as the emerging divide between AI users can own and access. 36 r/LocalLLaMA community 3d ago Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows Hi r/LocalLLaMA 👋 Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apache 2.0 license. A few specs 30B params, dense Multimodal: interleaved text +… 35 Hugging Face Daily Papers research 3d ago OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction Abstract Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential… 35 r/LocalLLaMA community 3d ago Chat UIs with native audio input for multimodal models? I've been running Gemma 4 E4B with oMLX and I can't find any chat interfaces that directly send the audio file to the model instead of running the audio through a separate STT layer. I can confirm the audio layers work because I ran a couple of requests through Pydantic AI in… 6 Hugging Face Daily Papers research 4d ago Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Abstract Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these… 25 Hugging Face Daily Papers research 4d ago Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning Abstract Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments does not always benefit. We further analyze the limitations in current multimodal environment distributions… 35 r/LocalLLaMA community 4d ago MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates… 10 Smol AI News news-outlet 4d ago not much happened today **Meta** re-enters the open-weight frontier with the release of **Muse Glimmer**, a **30B dense**, multimodal, agent-focused model under **Apache 2.0**, optimized for always-on local agents and consumer hardware. It features **quantization** to keep the model under **20GB**, a… 20 arXiv — Machine Learning research 4d ago CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions arXiv:2608.06516v1 Announce Type: new Abstract: Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the level of task decisions. A connected route can expand cross-modal reach while changing an… 18 arXiv — Machine Learning research 4d ago Recent advances in weakly supervised learning: New supervision paradigms, assumption relaxations, and practical solutions arXiv:2608.06896v1 Announce Type: new Abstract: Deep learning has achieved great success in recent years thanks to the availability of high-quality, well-annotated training data. However, this requirement is often not met in real-world applications. Weakly supervised learning… 30 arXiv — Machine Learning research 4d ago Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning arXiv:2608.06934v1 Announce Type: new Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences. Existing studies, however, often reduce these diverse judgements to… 17 arXiv — Machine Learning research 4d ago Conformal Fusion Under Missing Modalities arXiv:2608.07183v1 Announce Type: new Abstract: Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations. Existing work treats modality… 32 arXiv — Machine Learning research 4d ago An AI4AI Framework for Visual Token Pruning arXiv:2608.07193v1 Announce Type: new Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial and error. As pruning… 11 arXiv — Machine Learning research 4d ago Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery arXiv:2608.06406v1 Announce Type: cross Abstract: Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosystem management. While recent deep learning approaches provide accurate… 38 arXiv — Machine Learning research 4d ago Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection arXiv:2608.06434v1 Announce Type: cross Abstract: Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-system Vision-Language-Action (VLA) architectures combine fast reactive control with slow deliberative reasoning to… 19 arXiv — NLP / Computation & Language research 4d ago TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation arXiv:2608.06396v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relevant experts during downstream adaptation. Yet current approaches have two… 18 arXiv — NLP / Computation & Language research 4d ago Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding arXiv:2608.06532v1 Announce Type: new Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometimes one the model produced without reading the… 21 arXiv — NLP / Computation & Language research 4d ago Simple-OPD: Demystifying Warm-up for On-policy Distillation arXiv:2608.06802v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage before OPD. In this paper, we demystify warm-up for… 37 arXiv — NLP / Computation & Language research 4d ago Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly… 27 arXiv — NLP / Computation & Language research 4d ago ADIAS: Automated Design of Interactive Agentic Systems arXiv:2608.06410v1 Announce Type: cross Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents,… 36 arXiv — NLP / Computation & Language research 4d ago Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier arXiv:2608.06571v1 Announce Type: cross Abstract: Deployed vision-language systems often gate their answers on confidence, making confidence robustness relevant to oversight. We study confidence readouts under white-box, image-only attacks constrained to preserve the generated… 11 arXiv — NLP / Computation & Language research 4d ago SABRE: Scalable and Automated Benchmarking of VLMs under Stress arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and… 22 arXiv — NLP / Computation & Language research 4d ago Kimi K2.5: Visual Agentic Intelligence arXiv:2602.02276v2 Announce Type: replace Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This… 28 Hugging Face Daily Papers research 4d ago When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents Abstract Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments,… 18 Hugging Face Daily Papers research 4d ago Douyin Multimodal Embedding Model Technical Report Abstract Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content,… 36 Hugging Face Daily Papers research 4d ago StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Abstract Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design… 27 Hugging Face official-blog 4d ago Meta is back with Muse Glimmer: local, agentic, multimodal, and open source Back to Articles a]:hidden"> Meta is back with Muse Glimmer: local, agentic, multimodal, and open source! Published August 10, 2026 Update on GitHub Upvote 4 Pedro Cuenca pcuenq merve merve ben burtenshaw burtenshaw Aritra Roy Gosthipaty ariG23498 Great news from the OGs of open… 34 Page 2 of 10 · 500 articles ← Newer Older →