News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow Hugging Face Daily Papers research 25d ago On-Policy Delta Distillation Abstract On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied… 29 r/LocalLLaMA community 26d ago Tool for reproducible management of agent skills Hi LocalLLaMA! I've been building a small CLI tool for managing agent skills. I wanted a quick way to add and switch between different skill sets without manually copying folders around or losing track of which revision was installed (I tend to try different variations of the… 26 Hugging Face Daily Papers research 27d ago On Locality and Length Generalization in Visual Reasoning Abstract A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision models in use today,… 23 Hugging Face Daily Papers research 27d ago Hierarchical Denoising For Multi-Step Visual Reasoning Abstract Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference… 23 Hugging Face Daily Papers research 27d ago RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Abstract Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models… 18 Hugging Face Daily Papers research 28d ago VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance Abstract Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in… 13 arXiv — Machine Learning research 28d ago CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models arXiv:2607.14125v1 Announce Type: new Abstract: Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descriptions, typically formed by inserting a class label (e.g., "cat") into a prompt… 19 arXiv — Machine Learning research 28d ago LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration arXiv:2607.14410v1 Announce Type: new Abstract: Spatially resolved omics studies increasingly combine transcriptomic and epigenomic assays, yet downstream analysis is often still performed using single-modality pipelines. We present LATTICE (Latent Alignment of Tissue-level and… 8 arXiv — Machine Learning research 28d ago Multimodal Semantic-Aware Contrastive Learning For False Negative Mitigation in 3D Medical Imaging arXiv:2607.14995v1 Announce Type: new Abstract: Multimodal Contrastive Learning (CL) has shown significant performance in aligning representations across various data modalities and improving downstream tasks, especially in healthcare. It works by minimizing the distance between… 34 arXiv — NLP / Computation & Language research 28d ago On-Policy Delta Distillation arXiv:2607.15161v1 Announce Type: cross Abstract: On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy… 33 arXiv — Machine Learning research 28d ago A vision foundation model for single-cell biology via spatial gene cartography arXiv:2607.14163v1 Announce Type: cross Abstract: Most single-cell foundation models are adapted from language models, representing each cell as a sequence of gene tokens. This discards the relationships among genes and often the magnitude of their expression. We present… 10 arXiv — Machine Learning research 28d ago Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection arXiv:2607.14236v1 Announce Type: cross Abstract: Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision-driven and can struggle once manipulation enters contact states where the scene is… 27 arXiv — Machine Learning research 28d ago DiMaS: Distribution Matching for Steering Vision-Language-Action Models arXiv:2607.14280v1 Announce Type: cross Abstract: Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a… 22 arXiv — NLP / Computation & Language research 28d ago Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs arXiv:2607.14099v1 Announce Type: new Abstract: Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation… 10 arXiv — NLP / Computation & Language research 28d ago CoEvoT: Co-Evolving Chain-of-Thought Prompting for Graph-LLM Reasoning arXiv:2607.14114v1 Announce Type: new Abstract: Graph learning under distribution shift presents a persistent challenge, where models adapt to new graphs with limited or even no supervision. Recent graph--LLM approaches move toward label-efficient prediction by linearizing… 12 arXiv — NLP / Computation & Language research 28d ago Harnessing LLMs for Reliable Academic Supervision: A Comparative Study arXiv:2607.14707v1 Announce Type: new Abstract: Large language models routinely produce fluent answers to single-shot prompts, yet deploying them as reliable components of a domain decision system is substantially harder. Closing this gap is the work of harness engineering: the… 19 arXiv — NLP / Computation & Language research 28d ago Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence arXiv:2607.15092v1 Announce Type: new Abstract: Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from… 7 arXiv — NLP / Computation & Language research 28d ago TikStance: A Multimodal and Hierarchical Dataset for Multi-target Stance Analysis in TikTok Political Conversations arXiv:2607.15240v1 Announce Type: new Abstract: Political discourse has increasingly moved to short-video platforms, yet computational analysis of such content remains constrained by the scarcity of datasets that jointly preserve audiovisual information and hierarchical… 15 arXiv — NLP / Computation & Language research 28d ago Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA arXiv:2607.15241v1 Announce Type: new Abstract: Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design choices across nine documented… 20 arXiv — NLP / Computation & Language research 28d ago SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions arXiv:2607.15272v1 Announce Type: new Abstract: Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing… 34 arXiv — NLP / Computation & Language research 28d ago MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation arXiv:2607.14264v1 Announce Type: cross Abstract: Automated chest CT report generation remains challenging because clinically faithful reporting requires both whole-volume understanding and accurate description of localized anatomical findings. Here we developed and… 11 arXiv — NLP / Computation & Language research 28d ago SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning arXiv:2607.14333v1 Announce Type: cross Abstract: Vision Language Models (VLMs) demonstrate strong perceptual abilities but remain limited in tasks requiring analytical reasoning across multiple visual states, such as multi-image comparison, change detection, and multi-step… 19 arXiv — NLP / Computation & Language research 28d ago Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction arXiv:2607.14593v1 Announce Type: cross Abstract: As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We present a longitudinal multimodal study of a memory-augmented conversational agent (24… 37 arXiv — NLP / Computation & Language research 28d ago Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment arXiv:2607.14682v1 Announce Type: cross Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge. Current approaches bifurcate into Supervised Fine-Tuning… 14 arXiv — NLP / Computation & Language research 28d ago Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy arXiv:2607.15176v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis).… 20 r/MachineLearning community 28d ago CfP | RTCA @ NeurIPS 2026 [R] Call for Papers and Demos Real-Time Conversational Agents (RTCA): Toward Natural Multimodal Interaction 1st RTCA Workshop [@]() NeurIPS 2026 , Sydney, Australia 11 or 12 December 2026 Website: https://rtcaneurips26.github.io/ We are pleased to share the Call for Papers and Demos… 6 Simon Willison community 29d ago Inkling: Our open-weights model Inkling: Our open-weights model Mira Murati's Thinking Machines Lab just released their first open-weights model. Inkling is "a Mixture-of-Experts transformer with 975B total parameters, 41B active" - an Apache-2.0 licensed multimodal model trained on 45 trillion tokens of text,… 4 Hugging Face Daily Papers research 29d ago Registers Matter for Pixel-Space Diffusion Transformers Abstract Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training,… 15 Latent.Space news-outlet 29d ago [AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B) Thinky's first full LLM release is a banger and bonus: it's open weights! 30 Smol AI News news-outlet 29d ago not much happened today **Moonshot AI** launched **Kimi K3**, a frontier-class open-weights model with **2.8T parameters**, **1M-token context window**, and **native multimodal input**. It features novel **Kimi Delta Attention (KDA)** enabling up to **6.3x faster decoding** and **Attention Residuals**… 18 Hugging Face Daily Papers research 29d ago Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation Abstract We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference,… 10 Hugging Face Daily Papers research 29d ago Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Abstract While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned… 25 arXiv — Machine Learning research 29d ago PQFA: Parallel Quantum Feature Augmentation of Fused Representations for Multimodal Classification arXiv:2607.13466v1 Announce Type: new Abstract: Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid… 31 arXiv — Machine Learning research 29d ago MetaPerch: Learning from metadata for bioacoustics foundation models arXiv:2607.14072v1 Announce Type: new Abstract: Bioacoustic foundation models rely on large-scale citizen science platforms like Xeno-Canto for geographically and ecologically diverse data. Recent work has shown that supervision alone can produce SotA species detection models… 37 arXiv — Machine Learning research 29d ago Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation arXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision, a cost-aware evaluation framework for multimodal AI agents on quantum circuit… 11 arXiv — Machine Learning research 29d ago HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration arXiv:2607.13056v1 Announce Type: cross Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-world collaboration fundamentally requires… 12 arXiv — Machine Learning research 29d ago Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision arXiv:2607.13237v1 Announce Type: cross Abstract: Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge. We propose a human-in-the-loop knowledge acquisition framework that combines active learning with dual-loss optimization… 10 Hugging Face Daily Papers research 29d ago GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch Abstract World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly… 8 Hugging Face Daily Papers research 29d ago Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering Abstract Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception… 30 r/LocalLLaMA community 29d ago Google is updating Gemma 4's chat templates, bringing major fixes to tool calling and reducing "laziness", and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision! Nvm ignore the image links here is the source: https://x.com/googlegemma/status/2077449152062247219 https://huggingface.co/spaces/google/gemma4_vision_token_budget   submitted by   /u/Iwaku_Real [link]   [comments] 16 Hugging Face Daily Papers research 1mo ago Let RGB Be the Language of Vision Abstract This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB images, while general visual tasks can be converted… 25 Hugging Face Daily Papers research 1mo ago SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Abstract Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question… 23 Hugging Face Daily Papers research 1mo ago Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Abstract Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure… 32 r/LocalLLaMA community 1mo ago tencent/Hy-Embodied-RxBrain-1.0 · Hugging Face Introduction RxBrain ( Hy-Embodied-RxBrain-1.0 ) is a unified multimodal foundation model for embodied cognition — a single model that couples language reasoning with visual imagination to deliver three core capabilities: 🤖 Embodied Understanding & Reasoning — question… 35 Smol AI News news-outlet 1mo ago not much happened today **Thinking Machines Lab** launched **Inkling**, its first fully released open-weights foundation model family, featuring **975B parameters** with **41B active parameters** in a **Mixture-of-Experts** architecture. Inkling supports **multimodality** with text, image, and audio… 20 arXiv — Machine Learning research 1mo ago Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning arXiv:2607.12112v1 Announce Type: new Abstract: Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic… 4 arXiv — Machine Learning research 1mo ago Do You Remember? Toward Memory-Centric Multimodal AI arXiv:2607.11919v1 Announce Type: cross Abstract: Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen visual encoder, produce a one-shot text output, and discard internal… 36 arXiv — Machine Learning research 1mo ago SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning arXiv:2607.12042v1 Announce Type: cross Abstract: Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn… 32 arXiv — NLP / Computation & Language research 1mo ago Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings arXiv:2607.12071v1 Announce Type: new Abstract: Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces and neural coding patterns, which severely impedes cross-modal alignment… 34 arXiv — NLP / Computation & Language research 1mo ago WikiSTAR: A System for Shedding Light on the Hidden History of Scientific Wikipedia Articles arXiv:2607.12441v1 Announce Type: new Abstract: Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scientific knowledge evolves over time. Yet scientifically meaningful revisions are… 10 Page 10 of 10 · 500 articles ← Newer