News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — Machine Learning research 3h ago Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness arXiv:2608.12592v1 Announce Type: new Abstract: Continuous physiological time series underpin modern clinical monitoring, yet many of the most informative signals are invasive, expensive, or simply unavailable for a given patient. Conditional generation offers a remedy: an… 6 arXiv — Machine Learning research 3h ago A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings arXiv:2608.12745v1 Announce Type: new Abstract: Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making… 19 arXiv — Machine Learning research 3h ago Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance arXiv:2608.12929v1 Announce Type: new Abstract: The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches. We compared the predictive power of network, environmental, device, and vision… 38 arXiv — Machine Learning research 3h ago CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation arXiv:2608.12944v1 Announce Type: new Abstract: Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the… 23 arXiv — NLP / Computation & Language research 3h ago Vision-Language Models are Fragile Multilingual Associators arXiv:2608.12333v1 Announce Type: new Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark… 37 arXiv — NLP / Computation & Language research 3h ago CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model arXiv:2608.13101v1 Announce Type: new Abstract: Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and… 36 arXiv — NLP / Computation & Language research 3h ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 3h ago Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form… 38 arXiv — NLP / Computation & Language research 3h ago CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation arXiv:2608.13387v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation… 33 arXiv — NLP / Computation & Language research 3h ago Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors arXiv:2608.12746v1 Announce Type: cross Abstract: Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an individual object mention to what the image shows. Most… 4 arXiv — NLP / Computation & Language research 3h ago TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint arXiv:2608.13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce… 30 r/LocalLLaMA community 16h ago SenseNova-Vision: a 7B open model that does segmentation, depth, detection, OCR, and 3D reconstruction with no task-specific heads Stumbled across this new vision model, SenseNova-Vision. It's a 7B MoT model, Apache 2.0 license, which is cool. The main idea is it treats pretty much all computer vision stuff as just one generation problem. Like, instead of needing a bunch of different models for detection,… 17 r/LocalLLaMA community 1d ago I asked DeepSeek-V4-Flash to work with Muse-Glimmer for Vision ability in PI agent and it produced this Same old prompt, just appended a TIP in the end: "Write a single HTML file with a full-page canvas and no libraries. Simulate a realistic side-view of a moving car as the main subject. Keep the car visible in the foreground while the background landscape scrolls continuously to… 33 Hugging Face Daily Papers research 1d ago AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models Abstract AtlasVLA improves embodied AI by replacing reactive control with proactive reasoning via persistent world-ego memory, enabling robust long-horizon manipulation from a single wrist camera. Generated by thinkingmachines/Inkling-Small While Vision-Language-Action (VLA)… 36 arXiv — Machine Learning research 1d ago FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting arXiv:2608.11623v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational… 23 arXiv — NLP / Computation & Language research 1d ago LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection arXiv:2608.11691v1 Announce Type: cross Abstract: Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a… 18 arXiv — Machine Learning research 1d ago REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation arXiv:2608.11698v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond… 8 arXiv — Machine Learning research 1d ago Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision arXiv:2608.12027v1 Announce Type: new Abstract: Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing… 13 arXiv — NLP / Computation & Language research 1d ago CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models arXiv:2608.11534v1 Announce Type: new Abstract: In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that… 10 arXiv — NLP / Computation & Language research 1d ago AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention arXiv:2608.11758v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic… 34 arXiv — NLP / Computation & Language research 1d ago GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation arXiv:2608.11787v1 Announce Type: new Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision… 6 arXiv — NLP / Computation & Language research 1d ago BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model arXiv:2608.11244v1 Announce Type: cross Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support… 9 arXiv — NLP / Computation & Language research 1d ago How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment arXiv:2608.11816v1 Announce Type: cross Abstract: State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced… 25 arXiv — NLP / Computation & Language research 1d ago LookBack: Where and How to Score LVLM Responses via Visual Reference Usage arXiv:2608.11847v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations;… 14 arXiv — NLP / Computation & Language research 1d ago Investigating Learner-Aware Design of LLM-Generated Educational Feedback arXiv:2602.11650v2 Announce Type: replace Abstract: Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and information coverage) to support answer revision and learner acceptance… 12 Hugging Face Daily Papers research 1d ago Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models Abstract Self-Geometry improves vision foundation model predictions by enforcing explicit multi-view geometric constraints via test-time adaptation with LoRA, disentangled losses, and angular neighbor sampling. Generated by thinkingmachines/Inkling-Small Recent Vision Foundation… 21 Hugging Face Daily Papers research 1d ago The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images Abstract Visual tool-use in multimodal LLMs often lacks causal effectiveness, with returned observations frequently failing to influence answers or being used incoherently despite aggregate accuracy improvements. Generated by thinkingmachines/Inkling-Small The… 6 Hugging Face Daily Papers research 1d ago NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs Abstract NeuPAT selectively constrains updates to language-sensitive neurons during multimodal tuning to preserve LLM language capabilities while enabling perceptual adaptation. Generated by thinkingmachines/Inkling-Small Multimodal expansion of large language models (LLMs)… 18 Hugging Face Daily Papers research 1d ago MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Abstract Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only… 21 Hugging Face Daily Papers research 1d ago Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Abstract Spark-to-Paper is a lightweight, composable workflow inside coding assistants that generates research papers by separating planning from reporting, enforcing evidence-based claim revision, and using integrity checks to reduce fabrication. Generated by… 9 r/LocalLLaMA community 1d ago LFM2.5-VL-3B recognizes Steve from Minecraft running locally on an iPhone 17 Liquid AI put out LFM2.5-VL-3B today, which is a 3.1B vision model that weighs roughly 2GB and fits well on a phone Benchmarks are benchmarks so I tried something sillier. Took a photo of a little Steve toy I have, gave it to the model and asked it what it was looking at It… 29 r/LocalLLaMA community 1d ago CohereLabs/North-Micro-Vision-Instruct · Hugging Face North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation for prototyping, task-specific fine-tuning, and specialized multimodal… 8 r/LocalLLaMA community 1d ago LiquidAI/LFM2.5-VL-3B · Hugging Face LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment . It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both text and images, and uses the LFM2.5-2.6B language model as its backbone,… 15 Hugging Face official-blog 1d ago LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge Back to Articles a]:hidden"> LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge Team Article Published August 12, 2026 Upvote - Samuel Stevens samuelstevens LiquidAI Ryan Shubert shubeydoo LiquidAI Sina s-jse LiquidAI Tianshu Yu tianshu-yu LiquidAI Brandon… 24 Hugging Face Daily Papers research 2d ago Articulated Object Reconstruction from Rest-State Observation Abstract A rest-state framework reconstructs articulated objects from a single closed configuration by fusing vision-language outputs into consistent part meshes and validating synthesized motion hypotheses via geometric consistency. Generated by thinkingmachines/Inkling-Small… 29 Hugging Face Daily Papers research 2d ago DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Abstract DistilVDR is a compact 524M vision-document retriever distilled from an 8B teacher using cosine alignment without relevance labels, achieving near-teacher accuracy with far smaller indexes and faster indexing. Generated by thinkingmachines/Inkling-Small Visual document… 33 arXiv — Machine Learning research 2d ago Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory arXiv:2608.09997v1 Announce Type: new Abstract: Transformers have had a profound impact on the world of language processing and computer vision. As efforts to answer the million-dollar question of ``How does a Transformer learn?" have been increasing, existing interpretability… 6 arXiv — Machine Learning research 2d ago ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation arXiv:2608.10905v1 Announce Type: new Abstract: On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight,… 7 arXiv — Machine Learning research 2d ago MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis arXiv:2608.09986v1 Announce Type: cross Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although… 29 arXiv — NLP / Computation & Language research 2d ago Multimodal Item Parameter Estimation using Simulated Response Probabilitie arXiv:2608.10154v1 Announce Type: new Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to… 19 arXiv — NLP / Computation & Language research 2d ago VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise… 14 arXiv — NLP / Computation & Language research 2d ago Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models arXiv:2608.10484v1 Announce Type: cross Abstract: Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2… 21 arXiv — NLP / Computation & Language research 2d ago The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces arXiv:2608.10689v1 Announce Type: cross Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision… 28 arXiv — NLP / Computation & Language research 2d ago Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence arXiv:2608.10720v1 Announce Type: cross Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a… 22 arXiv — NLP / Computation & Language research 2d ago StreamFlow: Dynamic Memory Flows for Streaming Video Understanding arXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded memory. Yet existing paradigms remain limited:… 18 arXiv — NLP / Computation & Language research 2d ago MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment… 16 arXiv — NLP / Computation & Language research 2d ago HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models arXiv:2506.03922v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical… 15 arXiv — NLP / Computation & Language research 2d ago Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pretraining priors. For example, models will confidently assert a five-legged dog has four legs; consequently, on the VLMBias… 13 Hugging Face Daily Papers research 2d ago JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles Abstract A new jigsaw benchmark with interlocking pieces reveals that vision-language models fail at geometric reasoning and suffer a sharp performance drop as puzzle size increases. Generated by thinkingmachines/Inkling-Small Jigsaw puzzle solving requires jointly reasoning… 7 Vercel — AI dev-tools 2d ago Set up coding agents in one command with AI Gateway Using coding agents means setting up multiple accounts, provisioning API keys, and scattering observability and billing. Now, you can route them through AI Gateway to centralize all of this and add controls, with set up in one command: Any of 200+ models in any agent , including… 29 Page 1 of 10 · 500 articles Older →