News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — NLP / Computation & Language research 23d ago Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges arXiv:2607.19011v1 Announce Type: new Abstract: Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description.… 6 arXiv — NLP / Computation & Language research 23d ago DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning arXiv:2607.19088v1 Announce Type: new Abstract: Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We… 19 arXiv — NLP / Computation & Language research 23d ago MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings arXiv:2607.19235v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in… 8 arXiv — NLP / Computation & Language research 23d ago EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation arXiv:2607.18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across… 14 arXiv — NLP / Computation & Language research 23d ago Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks arXiv:2607.18767v1 Announce Type: cross Abstract: The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms of privacy, cost, and scalability. However, SLMs often struggle with complex vision-based tasks, such as grading… 22 Hugging Face Daily Papers research 23d ago ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning Abstract Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely… 27 Simon Willison community 24d ago Nativ: Run AI models locally on your Mac Nativ: Run AI models locally on your Mac Prince Canuma is the developer behind the excellent MLX-VLM Python library for running vision-LLMs using MLX on a Mac. I'm really excited about his new project, which wraps MLX in a full macOS desktop application. It's similar in shape to… 29 Hugging Face Daily Papers research 24d ago TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Abstract Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals… 30 arXiv — Machine Learning research 24d ago Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction arXiv:2607.16233v1 Announce Type: new Abstract: Integrating heterogeneous genomic and clinical modalities for joint cancer subtype classification and survival prediction remains a key challenge in precision oncology. Existing approaches suffer from three limitations: (1) they… 12 arXiv — Machine Learning research 24d ago RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants arXiv:2607.16243v1 Announce Type: new Abstract: Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying. However, multimodal large language models remain impractical for… 4 arXiv — Machine Learning research 24d ago Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation arXiv:2607.16246v1 Announce Type: new Abstract: Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model capabilities interact remains poorly characterized. We studies top-$k$-truncated,… 14 arXiv — Machine Learning research 24d ago Self-Evolving Just-In-Time Memory for Proactive Embodied Safety arXiv:2607.16247v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have empowered embodied agents to execute complex household tasks, they struggle to proactively handle dynamically emerging hazards during closed-loop interactions. Existing safety approaches… 12 arXiv — Machine Learning research 24d ago AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis arXiv:2607.16260v1 Announce Type: new Abstract: Multimodal survival analysis utilizing whole slide images (WSIs) and genomic profiles is fundamental for cancer prognosis. Recently, state-space models like Mamba have emerged as powerful tools for sequence modeling. However,… 8 arXiv — Machine Learning research 24d ago Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision arXiv:2607.16263v1 Announce Type: new Abstract: Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that… 33 arXiv — Machine Learning research 24d ago Multimodal Attention-based Deep Learning for Emergency Triage with Electronic Health Records arXiv:2607.16662v1 Announce Type: new Abstract: Accurate emergency triage decision is critical to avoid clinical deterioration, morbidity, and mortality. Machine learning-based triage system involves acquiring the main presenting complaint in text form and assessing vital signs… 4 arXiv — Machine Learning research 24d ago MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning arXiv:2607.16789v1 Announce Type: new Abstract: Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities. However, training multimodal models faces two main obstacles. First, collecting large-scale, well-aligned… 5 arXiv — Machine Learning research 24d ago ChemFusion: A Multimodal Cross-Attention Network for Reaction Yield Prediction arXiv:2607.17033v1 Announce Type: new Abstract: Forecasting the outcomes of transition-metal-catalyzed reactions is notoriously complex due to the interplay of diverse physical and chemical variables. A persistent computational bottleneck has been effectively merging broad… 13 arXiv — NLP / Computation & Language research 24d ago Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis? arXiv:2607.17262v1 Announce Type: new Abstract: Existing methods for multimodal sentiment analysis (MSA) under missing modalities usually follow a repair-first paradigm. We revisit this assumption and ask: \emph{should every missing modality be repaired?} A per-sample oracle… 21 arXiv — NLP / Computation & Language research 24d ago How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks arXiv:2607.17452v1 Announce Type: new Abstract: Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a… 13 arXiv — NLP / Computation & Language research 24d ago ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG arXiv:2607.16208v1 Announce Type: cross Abstract: Graph-grounded multimodal question answering organizes text, tables, and images in a structured evidence graph, yet end-to-end accuracy depends on which multimodal assets are ranked highly enough to enter downstream reasoning;… 6 arXiv — NLP / Computation & Language research 24d ago One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models arXiv:2607.16442v1 Announce Type: cross Abstract: Machine unlearning is widely used to remove hazardous knowledge from large language models. Modern Vision-Language Models (VLMs), however, process both text and visual inputs, raising a fundamental security question: does… 29 arXiv — NLP / Computation & Language research 24d ago Can Multimodal Large Language Models Understand OCT? arXiv:2607.16609v1 Announce Type: cross Abstract: Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis,… 14 arXiv — NLP / Computation & Language research 24d ago Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing arXiv:2607.16898v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) have recently demonstrated powerful instruction-based image editing capabilities, but they also raise serious concerns about unauthorized manipulation of personal portraits. Existing adversarial… 9 arXiv — NLP / Computation & Language research 24d ago EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation arXiv:2607.17366v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) achieves accurate predictions by integrating multimodal and contextual information in dialogues. While current MERC approaches focus on modeling complex contextual… 34 Hugging Face Daily Papers research 24d ago HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Abstract Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance… 9 Hugging Face Daily Papers research 24d ago ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams Abstract Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest.… 32 Hugging Face Daily Papers research 24d ago FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Abstract Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing… 10 Hugging Face Daily Papers research 24d ago JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models Abstract The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an… 13 Hugging Face Daily Papers research 24d ago SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning Abstract We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. For each query, the model proposes an answer using the same weights, and issues a binary self-verdict (Yes/No).… 7 r/LocalLLaMA community 24d ago According to Agent Arena Kimi K3 ranks at same level as opus thinking My testing with android suggest it is below that. Vision on Opus is the reason. But I don't doubt that for non vision tasks Kimi k3 is on par.   submitted by   /u/Terminator857 [link]   [comments] 4 Hugging Face Daily Papers research 25d ago See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models Abstract Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where… 16 Hugging Face Daily Papers research 25d ago RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Abstract Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial videos and other multimodal… 24 r/LocalLLaMA community 25d ago MiniCPM-Robot model series - MiniCPM-RobotManip & MiniCPM-RobotTrack 🚀 MiniCPM enters the physical world — enabling robots to understand, remember, and act. We open-source MiniCPM-Robot, our first embodied AI model series, including: 🤖 MiniCPM-RobotManip — a 1.5B general-purpose Vision-Language-Action (VLA) model for robotic manipulation. 🐕… 23 Hugging Face Daily Papers research 25d ago S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation Abstract We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model… 15 arXiv — Machine Learning research 25d ago Diffusion models recover accurate mixture weights despite score function insensitivity arXiv:2607.15485v1 Announce Type: new Abstract: Score-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture… 23 arXiv — Machine Learning research 25d ago Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework arXiv:2607.15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer… 6 arXiv — Machine Learning research 25d ago Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment arXiv:2607.15928v1 Announce Type: new Abstract: Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce. Existing… 38 arXiv — Machine Learning research 25d ago An Empirical Study of Handcrafted Feature Learning and Convolutional Neural Networks for Facial Expression Recognition arXiv:2607.15288v1 Announce Type: cross Abstract: Facial expression recognition is an important computer vision task with applications in human--computer interaction, mental health monitoring, driver alert systems, and behavioral analysis. While convolutional neural networks… 5 arXiv — Machine Learning research 25d ago AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning arXiv:2607.15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of… 31 arXiv — Machine Learning research 25d ago MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation arXiv:2607.15299v1 Announce Type: cross Abstract: In this paper, we propose MLLM-DataEngine, a novel closed-loop system that bridges data generation, model training, and evaluation. Within each loop iteration, the MLLM-DataEngine first analyzes the weakness of the model based on… 29 arXiv — Machine Learning research 25d ago Intentional Electromagnetic Interference Attacks on Facial Recognition arXiv:2607.15512v1 Announce Type: cross Abstract: Attacks on general computer vision algorithms are often relegated to the digital domain, with the optimization performed purely in the digital world and then translated to physical mediums for implementation. In the field of… 7 arXiv — Machine Learning research 25d ago Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models arXiv:2607.15565v1 Announce Type: cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering… 12 arXiv — Machine Learning research 25d ago Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents arXiv:2607.15657v1 Announce Type: cross Abstract: Multimodal AI agents increasingly rely on persistent long-term memory to ground generation in past visual and textual episodes. We show that unconditional trust in visual data creates a critical vulnerability. We propose Lucid, a… 36 arXiv — NLP / Computation & Language research 25d ago Large Language Models as Unified Multimodal Learners for Clinical Prediction arXiv:2607.15380v1 Announce Type: new Abstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion… 22 arXiv — NLP / Computation & Language research 25d ago ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning arXiv:2607.16131v1 Announce Type: new Abstract: Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail… 35 arXiv — NLP / Computation & Language research 25d ago HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection arXiv:2607.16076v1 Announce Type: cross Abstract: Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal… 35 arXiv — NLP / Computation & Language research 25d ago An Exam for Active Observers arXiv:2607.16165v1 Announce Type: cross Abstract: Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a… 18 arXiv — NLP / Computation & Language research 25d ago Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing arXiv:2606.07636v2 Announce Type: replace-cross Abstract: Long-form video editing over heterogeneous footage requires agents to coordinate source selection, multimodal analysis, timeline construction, narration and subtitle alignment, rendering, and revision while exposing… 37 Hugging Face Daily Papers research 25d ago Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Abstract We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel… 21 Hugging Face Daily Papers research 25d ago On-Policy Delta Distillation Abstract On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied… 29 Page 9 of 10 · 500 articles ← Newer Older →