News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow llama.cpp releases dev-tools 9d ago b11049 test-llama-archs : generate dummy test vocab ( #29084 ) Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48623939 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64,… 37 arXiv — NLP / Computation & Language research 10d ago Learn Your Own Thoughts: Abstract Token Curriculum arXiv:2609.19717v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on… 29 arXiv — NLP / Computation & Language research 10d ago Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning arXiv:2609.19878v1 Announce Type: cross Abstract: Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the… 6 arXiv — Machine Learning research 10d ago QoS-Aware Federated Learning for Multimodal In-Cabin Interaction in Smart Vehicles arXiv:2609.20123v1 Announce Type: new Abstract: Modern smart vehicles leverage multimodal sensors, ranging from high-bandwidth vision systems to low-rate physiological monitors, to provide personalized in-cabin services. However, integrating high-fidelity multimodal fusion with… 34 arXiv — Machine Learning research 10d ago Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment arXiv:2609.20199v1 Announce Type: new Abstract: Cognitive workload reflects the mental effort required during task performance and is central to the design of adaptive human-machine systems. The use of biosignals to measure cognitive workload has been extensively researched and… 30 arXiv — Machine Learning research 10d ago SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption arXiv:2609.20302v1 Announce Type: new Abstract: Multimodal gradient balancing methods modulate encoder gradients with a shared scalar per modality, implicitly assuming that corruption is uniform across the training batch. In practice, corruption is sample-heterogeneous: within a… 30 arXiv — NLP / Computation & Language research 10d ago Modality Discrepancy Transformer for Ambivalence and Hesitancy Recognition arXiv:2609.19148v1 Announce Type: new Abstract: Ambivalence and hesitancy (A/H) are affective states in which individuals express contradictory signals across facial, vocal, and linguistic channels. Automatically recognising A/H in clinical videos requires detecting cross-modal… 5 arXiv — NLP / Computation & Language research 10d ago To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives arXiv:2609.19167v1 Announce Type: new Abstract: As AI systems evolve into personalized digital companions, a central capability is reasoning over a user's long-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences.… 5 arXiv — NLP / Computation & Language research 10d ago Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion arXiv:2609.19417v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) is widely used for multimodal, cross-document question answering. However, building corpus-level graphs is expensive, slow to query, and difficult to maintain. We present TrioRAG, a… 18 arXiv — NLP / Computation & Language research 10d ago Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection arXiv:2609.19778v1 Announce Type: new Abstract: Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for… 24 arXiv — NLP / Computation & Language research 10d ago Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning arXiv:2609.20252v1 Announce Type: new Abstract: High-quality representations are essential for a wide range of downstream tasks. Dedicated embedding models are explicitly optimized for representation learning, yet their training data are often more limited in scale and diversity… 34 arXiv — NLP / Computation & Language research 10d ago RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning arXiv:2609.20784v1 Announce Type: new Abstract: Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged… 8 arXiv — NLP / Computation & Language research 10d ago Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models arXiv:2609.19384v1 Announce Type: cross Abstract: Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude… 5 arXiv — NLP / Computation & Language research 10d ago From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning arXiv:2609.19445v1 Announce Type: cross Abstract: The rapid expansion of multimodal models has surfaced formidable bottlenecks in computation, memory, and deployment, catalyzing the rise of Efficient Multimodal Learning (EML) as a pivotal research frontier. Despite intensive… 30 arXiv — NLP / Computation & Language research 10d ago Geopolitical Divisions Across Languages in Large Language Models arXiv:2609.20005v1 Announce Type: cross Abstract: People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how… 17 Vercel — AI dev-tools 10d ago GLM 5.3 FlashX now available on AI Gateway GLM 5.3 FlashX is now available on AI Gateway. GLM 5.3 FlashX is a high-speed serving option for Z.ai's multimodal coding model, delivering inference at ~200 tokens per second for faster streamed responses. The higher serving speed is useful for coding agents, tool loops, and… 21 r/MachineLearning community 10d ago Future of general LLM work (interp/inference/alignment) vs agentic/physical AI (VLA, multimodal) for career [D] I'm at a crossroads with two grad school options that would take me in somewhat different research directions, and I wanted some general advice on these fields, their growth, and industry alignment. I'm leaving out the specifics of the programs since I'm tryna compare the… 5 Hugging Face Daily Papers research 11d ago PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection Abstract Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains… 19 Hugging Face Daily Papers research 11d ago ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Abstract Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction… 21 Hugging Face Daily Papers research 11d ago HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses Abstract Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different… 12 arXiv — Machine Learning research 11d ago Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives arXiv:2609.17572v1 Announce Type: new Abstract: Auditing vision-language models (VLMs) for societal bias requires distinguishing direct algorithmic valuation disparities from confounders embedded within archival metadata. In this study, we audit Contrastive Language-Image… 24 arXiv — Machine Learning research 11d ago Walking the Score Manifold: Continuous-time Generative Dynamics on Learned Data Manifolds arXiv:2609.17901v1 Announce Type: new Abstract: Generative modeling of time-dependent data is typically formulated on a discrete temporal grid, restricting supervision to the observed timestamps in the training data. We instead frame generation as continuous-time evolution on a… 32 arXiv — Machine Learning research 11d ago Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers arXiv:2609.18321v1 Announce Type: new Abstract: Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization. The same reuse makes imperfect supervision persistent. Since even strong… 22 arXiv — Machine Learning research 11d ago Provable Guarantees for Spectral Structured Prediction arXiv:2609.18527v1 Announce Type: new Abstract: Structured prediction is the simultaneous prediction of multiple labels, and is widely used in various fields, such as natural language processing and computer vision. In this paper, we study binary node label recovery on signed… 8 arXiv — NLP / Computation & Language research 11d ago Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts arXiv:2609.18385v1 Announce Type: new Abstract: Emotions are an essential aspect of human communication, particularly on social media, where authors frequently combine text and images to convey their emotions. Yet prior work on emotion analysis of social media posts has… 28 arXiv — NLP / Computation & Language research 11d ago ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts arXiv:2609.18844v1 Announce Type: new Abstract: Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often… 21 arXiv — NLP / Computation & Language research 11d ago MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents arXiv:2609.19059v1 Announce Type: new Abstract: Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrieve and use earlier evidence across dialogue, files, and… 38 arXiv — NLP / Computation & Language research 11d ago NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation arXiv:2609.17699v1 Announce Type: cross Abstract: We present NeMo Data Designer (NDD), an open-source, general-purpose framework for multi-modal synthetic data generation (SDG). Designed to be intuitive to use, NDD provides a declarative configuration format in which human… 37 arXiv — NLP / Computation & Language research 11d ago G-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement arXiv:2609.18009v1 Announce Type: cross Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense… 23 r/LocalLLaMA community 11d ago Qwen3.8-27B uncensored Q6_K at 156K context on one RTX 5090, 140-190 tok/s with DFlash2 Setup for one long agentic coding session (tools + vision) on a single 5090: Qwen3.8-27B RVN Heretic (ARA abliterated) at Q6_K, 159744 context , 1 slot, q8_0 K/V, DFlash2 speculative decoding, vision projector on the GPU, Sharp chat template. All numbers measured 2026-09-16.… 18 arXiv — Machine Learning research 12d ago OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning arXiv:2609.16057v1 Announce Type: new Abstract: Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three limitations remain. (1) Existing methods often distill task-specific experience with limited generalizability.… 31 arXiv — Machine Learning research 12d ago A Decision-Support Audit Protocol for Supervision Drift in Proxy-Labeled Credit-Risk Prediction arXiv:2609.16102v1 Announce Type: new Abstract: Credit-risk models are trained on proxy labels and deployed under temporal and segment change, yet no single transfer metric separates base-rate shift, probability-scale shift, and feature-label relationship change. We contribute a… 32 arXiv — Machine Learning research 12d ago Robust Fault Detection in Mechanical Multimodal Time Series via Self-Supervised Cross-Modal Reconstruction arXiv:2609.16314v1 Announce Type: new Abstract: Fault detection is essential in industrial systems, enabling early identification of abnormal behaviour and improving safety, reliability, and operational efficiency. Modern systems increasingly rely on heterogeneous sensing… 35 arXiv — Machine Learning research 12d ago OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation arXiv:2609.16459v1 Announce Type: new Abstract: Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories while conditioning on the same… 26 arXiv — Machine Learning research 12d ago AsyncCouple-Flow: Asynchronous Cross-Modal Coupling and Flow Matching for Spatio-Temporal Forecasting arXiv:2609.16573v1 Announce Type: new Abstract: Multi-modal spatio-temporal forecasting (MM-STF) supports weather nowcasting, traffic prediction, and earth-system modeling by combining heterogeneous sources such as physical fields, satellite imagery, and in-situ sensors. Three… 27 arXiv — Machine Learning research 12d ago Distributed JEPA: A Self-Supervised Framework for Energy Forecasting arXiv:2609.17029v1 Announce Type: new Abstract: Traditional energy forecasting solutions rely on task-specific supervision and energy asset representations, limiting transferability and the ability to capture general temporal dynamics across heterogeneous assets. We address this… 25 arXiv — Machine Learning research 12d ago ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation arXiv:2609.17499v1 Announce Type: new Abstract: Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decisions. As one of the most advanced… 24 arXiv — NLP / Computation & Language research 12d ago Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation arXiv:2609.16059v1 Announce Type: new Abstract: Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fine-Tuning (SFT), which often leads to surface-level pattern matching and degrades… 22 arXiv — NLP / Computation & Language research 12d ago Efficient Multimodal Generative Recommendation with Latent Narrative Reasoning arXiv:2609.16070v1 Announce Type: new Abstract: Generative recommendation reformulates item prediction as semantic identifier generation, yet episodic content introduces a fundamentally different setting where the target is determined by narrative evolution rather than user… 4 arXiv — NLP / Computation & Language research 12d ago Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue arXiv:2609.16396v1 Announce Type: new Abstract: Negation is typically modeled through its linguistic realization, although spoken interaction is accompanied by tightly coordinated nonverbal behavior. We ask whether contexts centered on spoken negation cues contain measurable… 37 arXiv — NLP / Computation & Language research 12d ago Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs arXiv:2609.17327v1 Announce Type: new Abstract: This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages. We employ several fine-tuned vision-language… 25 arXiv — NLP / Computation & Language research 12d ago VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs arXiv:2609.16722v1 Announce Type: cross Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly… 28 arXiv — NLP / Computation & Language research 12d ago LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture arXiv:2609.16730v1 Announce Type: cross Abstract: Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation… 33 r/LocalLLaMA community 13d ago jinfer: An open-source AI inference engine for the JVM. Finally, AI in jar. For years, the JVM has watched the AI revolution from the bench. Every model, AI framework, every breakthrough, built with/for Python. jinfer is an inference engine built for the JVM from first principles: chat, vision, audio transcription, embeddings, reranking, and TTS. No… 7 Hugging Face Daily Papers research 13d ago HazardAuditor: From Executable Threats to Safer Computer-Use Agents Abstract HazardAuditor provides execution-grounded safety supervision for computer-use agents and introduces Guard Policy Optimization to align generative guard training with sequence-level safety outcomes. Generated by thinkingmachines/Inkling-Small Computer-use agents… 8 arXiv — Machine Learning research 13d ago ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation arXiv:2609.13857v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (ERC) requires adapting to the instance-dependent reliability of different evidence sources. Lexical content may be decisive, vocal expression may provide complementary cues, or… 17 arXiv — Machine Learning research 13d ago Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision arXiv:2609.13947v1 Announce Type: new Abstract: In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor. However, the logic chip integrated with a CMOS image sensor (CIS) is tightly constrained in… 12 arXiv — Machine Learning research 13d ago Data-Efficient Agentic Graph Domain Adaptation via Reliability-Aware Prototype Learning arXiv:2609.14045v1 Announce Type: new Abstract: Agentic learning systems are often required to adapt after deployment by observing new data and reusing prior knowledge under limited supervision or feedback. For graph-structured prediction, Graph Domain Adaptation (GDA) naturally… 36 arXiv — Machine Learning research 13d ago Multimodal deep learning from spectra for small-molecule structure identification: enhancing robustness with mixed-condition training arXiv:2609.14360v1 Announce Type: new Abstract: In practical molecular characterization, small-molecule structure identification benefits from complementary spectroscopic evidence, but missing, degraded, or mismatched spectra challenge multimodal models. Herein, we incorporate… 17 arXiv — NLP / Computation & Language research 13d ago TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams arXiv:2609.13158v1 Announce Type: new Abstract: Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize… 13 Page 3 of 10 · 500 articles ← Newer Older →