News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — NLP / Computation & Language research 13d ago Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth arXiv:2609.13745v1 Announce Type: new Abstract: Vision--language models (VLMs) can answer chart questions accurately, but output accuracy does not show how they combine the evidence needed to recover an exact value. We study vertical-bar value reading with controlled… 32 arXiv — NLP / Computation & Language research 13d ago SyRHM: Symbolic-Language-Enhanced Reasoning with Associative Retrieval for Zero-shot Harmful Meme Detection arXiv:2609.13794v1 Announce Type: new Abstract: Detecting harmful memes is critical for maintaining safe online communities. However, harmful intent is often implicit, arising from visual-textual incongruity and cultural stereotypes, which challenges existing multimodal… 15 arXiv — NLP / Computation & Language research 13d ago SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity arXiv:2609.13874v1 Announce Type: new Abstract: Multimodal clinical AI typically assumes that the waveform, report, metadata, and downstream predictions attached to a record belong to the same patient. In practice, linkage failures can silently assemble individually plausible… 28 arXiv — NLP / Computation & Language research 13d ago GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning arXiv:2609.14066v1 Announce Type: new Abstract: Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they remain fundamentally limited in reasoning depth and memory structure, suffering from… 32 arXiv — NLP / Computation & Language research 13d ago Learning to Refer from Estimated Listener Gaze arXiv:2609.14207v1 Announce Type: new Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals.… 29 arXiv — NLP / Computation & Language research 13d ago E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning arXiv:2609.14302v1 Announce Type: new Abstract: Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not… 6 arXiv — NLP / Computation & Language research 13d ago Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models arXiv:2609.14779v1 Announce Type: new Abstract: Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the… 30 arXiv — NLP / Computation & Language research 13d ago MedTRACE: Tool-Augmented Multimodal Clinical Reasoning Agents for Evidence-Grounded Decision-Making arXiv:2609.14823v1 Announce Type: new Abstract: Multimodal clinical decision-making requires reliable reasoning over heterogeneous evidence from electronic health records, medical images, and physiological signals. Existing models typically map these inputs directly to diagnoses… 30 Hugging Face Daily Papers research 13d ago PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Abstract PhysBrain 1.5 unifies physical environment understanding, action generation, and future state prediction via joint autoregressive training on discrete vision-language, motion, and visual target sequences, achieving state-of-the-art open-source embodied performance.… 29 Hugging Face Daily Papers research 13d ago Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Abstract The framework dynamically adjusts training prompts via exploration potential scoring and scaffolded rewrites to improve reinforcement learning for multimodal language models. Generated by thinkingmachines/Inkling-Small Training prompts in online reinforcement learning… 8 Hugging Face Daily Papers research 13d ago Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Abstract Discovery Foundation Models enable open-ended scientific discovery through iterative problem formulation, hypothesis testing, and evidence-based revision across dry and wet lab settings. Generated by thinkingmachines/Inkling-Small Foundation models have progressed from… 18 Hugging Face Daily Papers research 13d ago BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender Abstract A benchmark requiring agents to programmatically reconstruct real-world videos in Blender reveals that current models achieve high perceptual similarity but struggle to retain spatiotemporal facts. Generated by thinkingmachines/Inkling-Small Multimodal agents can create… 16 Hugging Face Daily Papers research 13d ago LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents Abstract A 16.7B-parameter mixture-of-experts diffusion vision-language agent achieves strong multimodal GUI performance while preserving block-parallel decoding efficiency. Generated by thinkingmachines/Inkling-Small Diffusion large language models (dLLMs) achieve high decoding… 37 Hugging Face Daily Papers research 13d ago Omni-Streaming Thinking Abstract Omni-Streaming Thinking improves streaming omni-modal reasoning by deferring claims until cross-modal verification, reducing premature commitment and auditory hallucinations. Generated by thinkingmachines/Inkling-Small Streaming omni-modal models must decide what and… 14 r/MachineLearning community 14d ago [P] Built a 100% Client-Side Vision Pipeline for Real-Time Chessboard & Multi-Board Detection (Chrome/Firefox Extension) [P] Hi everyone, Inspired by tools like Chessvision.ai, I wanted to take a different architectural approach and build a browser extension ( ChessInsights AI ) that performs chessboard detection and piece recognition 100% client-side using local inference—with zero image data ever… 23 r/LocalLLaMA community 14d ago Intern-S2-397B (multimodal, reasoning, coding, and scientific agent capabilities) Model: https://huggingface.co/internlm/Intern-S2-397B Collection: https://huggingface.co/collections/internlm/intern-s2 From Intern Large Models on 𝕏: https://x.com/intern_lm/status/2099425184587370976 vLLM on 𝕏: Day-0 support for Intern Large Models Intern-S2-397B is now… 17 r/MachineLearning community 14d ago PhD branding question [R] I'm starting a PhD where I will be doing Graph ML (somewhere along the lines of graph signal processing/ graph deep learning.) My eventual goal is research scientist at big tech, or whichever company has a strong research division, where I can continue similar AI/ML work. I have… 36 arXiv — NLP / Computation & Language research 14d ago Space as an Interventional Invariant: Cross-Modal Predictive Geometry for Stratified Cities and Em-Spaced Intelligence arXiv:2609.11959v1 Announce Type: cross Abstract: Space is a foundational concept across mathematics, physics, spatial cognition, urban science, and embodied intelligence, yet these fields often treat spatial structure either as a shared geometric container or as a collection of… 5 arXiv — Machine Learning research 14d ago On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health arXiv:2609.11961v1 Announce Type: new Abstract: Stress is a pervasive determinant of mental health and a key target for mobile health interventions. On-device language models (ODLMs) offer privacy-preserving inference without cloud dependency, yet their feasibility for health… 35 arXiv — Machine Learning research 14d ago FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences arXiv:2609.11993v1 Announce Type: new Abstract: Machine learning research in financial services is limited by the scarcity of representative open-source datasets. Existing resources are often narrowly focused on a single modality or task and fail to reflect the structured,… 11 arXiv — Machine Learning research 14d ago LatentVerse: A Framework for Understanding Shared and Modality-Specific Information in Multimodal Latent Representations arXiv:2609.12364v1 Announce Type: new Abstract: Latent embeddings have become a central data abstraction in modern machine learning, especially in biomedicine, where foundation models are increasingly used to encode multimodal data like clinical text, medical images, omics, and… 19 arXiv — Machine Learning research 14d ago Beyond the Query: Do Retrieval Signals Improve Adaptive Multimodal RAG Routing? arXiv:2609.12437v1 Announce Type: new Abstract: Adaptive RAG often uses retrieval-time signals to decide whether another retrieval, reranking, or multimodal step should run. We ask whether these signals add routing value once the query itself is already known. Across document,… 16 arXiv — Machine Learning research 14d ago SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation arXiv:2609.12579v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) scores student-generated prefixes with a solution-conditioned self-teacher, yet transfers supervision only through next-token probabilities. We ask whether the aligned final-layer discrepancy… 14 arXiv — Machine Learning research 14d ago ProactiveBench: Can Streaming Video Models Really Interact Like Humans? arXiv:2609.12658v1 Announce Type: new Abstract: Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query a model at a selected timestamp and therefore do… 26 arXiv — Machine Learning research 14d ago Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting arXiv:2609.12890v1 Announce Type: new Abstract: In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make… 5 arXiv — NLP / Computation & Language research 14d ago Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture arXiv:2609.12575v1 Announce Type: new Abstract: Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. From humour to politics to art, people express themselves in words and images that… 11 arXiv — NLP / Computation & Language research 14d ago MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation arXiv:2512.14691v3 Announce Type: replace Abstract: Modern multimodal generative models can synthesize visually compelling images and videos, but it remains unclear whether this visual fluency reflects genuine reasoning: when prompted to generate a solution, can a model preserve… 8 Hugging Face Daily Papers research 14d ago Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models Abstract LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information. Generated by thinkingmachines/Inkling-Small… 20 r/MachineLearning community 14d ago [Upcoming AMA] Waymo AI Team AMA – Drop Your Questions Early! [D] Hi r/MachineLearning , Join our AI leads as they answer your questions on foundation models, simulation, and scaling the Waymo Driver. Our AMA thread is officially open, and you can start dropping your questions now. From multimodality and end-to-end architectures to the… 20 r/LocalLLaMA community 15d ago internlm/Intern-S2 · Hugging Face from internlm: We introduce Intern-S2-397B , our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent… 6 r/LocalLLaMA community 16d ago I am impressed and I owe you one, Qwen 3.8 flash next (vision)! I have enabled the vision for the CIRU Strix UL4 quant of Qwen 3.8 flash next (others quants likely perform very similar) and tried it on a few things, then wanted to show my partner how great it works and she asked it it could identify plants. So I took a photo from a plant… 26 r/LocalLLaMA community 16d ago Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36 I find this new model at HF: "Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and video understanding. Architecture Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta… 26 Latent.Space news-outlet 16d ago [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale We agree with Sebastian: this should have been DeepSeek v5 14 r/LocalLLaMA community 16d ago Got an old slow low vram GPU laying around? Might be worth it to use for Just Vision mmproj llama.cpp For many, Vram is precious, I see many people recommend using --no-mmproj-offload to save gpu vram but it is painfully slow. Especially if you are using it with agentic coding. If possible, add that secondary gpu just for mmproj with --mmdev CUDA1(your gpu). It will be a… 30 Hugging Face Daily Papers research 16d ago Studying Image Tokenizers as Visual Languages in Unified Multimodal Models Abstract Using a controlled autoregressive testbed, the study analyzes task-specific validation losses during multimodal pretraining to evaluate how image tokenizer design affects joint text-image modeling and downstream performance. Generated by thinkingmachines/Inkling-Small… 11 Hugging Face Daily Papers research 16d ago Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking Abstract A reasoning-capable vision-language model that iteratively retrieves and reasons over Wikipedia improves multimodal entity linking for rare entities defined by knowledge-graph structure. Generated by thinkingmachines/Inkling-Small Multimodal entity linking grounds… 29 Hugging Face Daily Papers research 16d ago Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning Abstract Optimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language. Generated by thinkingmachines/Inkling-Small Reasoning… 36 Hugging Face Daily Papers research 16d ago ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Abstract ActReview is a rebuttal-guided post-training framework that generates diagnostic claims and concrete revision suggestions for peer review by leveraging author responses as latent supervision. Generated by thinkingmachines/Inkling-Small As LLMs are increasingly used for… 18 r/MachineLearning community 17d ago Why is TMLR so slow in recent times [D] A final-year PhD student here. A few months back, I submitted a solo-authored paper to TMLR. The reviewers were on time and extremely positive, with some minor revisions. After submitting the revised version, there was absolute silence from the reviewers, with just one… 7 arXiv — Machine Learning research 17d ago M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction arXiv:2609.10559v1 Announce Type: new Abstract: To address the challenges of behavioral multimodality, limited semantic utilization, and long-term error accumulation in vessel trajectory prediction, this paper proposes M3-Former, a multimodal trajectory prediction framework… 28 arXiv — Machine Learning research 17d ago RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty arXiv:2609.10798v1 Announce Type: new Abstract: Image-metadata prediction requires fusing heterogeneous evidence whose reliability can vary across samples and latent factors. Existing representation-level fusion methods typically choose an aggregation architecture, such as… 24 arXiv — Machine Learning research 17d ago EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression arXiv:2609.11058v1 Announce Type: new Abstract: Recent advances in multimodal large language mod- els (MLLMs) have opened new opportunities for edge intelligence by enabling reasoning across heterogeneous sensor modalities, such as vision, text, and telemetry data. However,… 4 arXiv — Machine Learning research 17d ago Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models arXiv:2609.11135v1 Announce Type: new Abstract: Short-term photovoltaic (PV) power and global horizontal irradiance (GHI) forecasts are essential for effective dispatch, reserve scheduling, and grid operations. At these forecasting horizons, errors are predominantly driven by… 21 arXiv — Machine Learning research 17d ago Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting arXiv:2609.10613v1 Announce Type: cross Abstract: In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite… 24 arXiv — Machine Learning research 17d ago Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite Systems arXiv:2609.10746v1 Announce Type: cross Abstract: The growing reliance on Low-Earth Orbit (LEO) satellite communication systems has increased the need for intelligent methods capable of detecting cyberattacks across complex and dynamic space environments. Unlike conventional… 28 arXiv — Machine Learning research 17d ago Meta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy Clustering arXiv:2609.10749v1 Announce Type: cross Abstract: Accurate plant growth estimation is essential for greenhouse monitoring, yet obtaining labeled data remains costly and time-consuming. To address this, we propose a few-shot regression framework that combines Vision Transformer… 16 arXiv — NLP / Computation & Language research 17d ago Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking arXiv:2609.10745v1 Announce Type: new Abstract: Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We… 19 arXiv — NLP / Computation & Language research 17d ago Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction arXiv:2609.10950v1 Announce Type: new Abstract: Recent multimodal sentiment analysis studies increasingly adopt text-centric fusion approaches to exploit the rich sentiment information inherent in the textual modality. However, these approaches often suffer from performance… 17 arXiv — NLP / Computation & Language research 17d ago OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models arXiv:2609.11244v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically… 5 arXiv — NLP / Computation & Language research 17d ago The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based Methods arXiv:2609.11247v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) remains constrained by modality imbalance, yet the field continues to rely on optimization-based balancing methods that promise more than they deliver. We provide three contributions: 1) a… 35 Page 4 of 10 · 500 articles ← Newer Older →