News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow Hugging Face Daily Papers research 16d ago UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Abstract Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision… 11 Hugging Face Daily Papers research 16d ago WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Abstract Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves… 35 r/MachineLearning community 16d ago Agent Mini: a minimal, local-first AI agent you can actually read, understand, and extend. [P] It is about ~3k lines of Python, uses Ollama by default, and ships with practical tools for shell, files, web search, memory, and vision. No LangChain. No LiteLLM. No big framework stack. Just asyncio , httpx , and a small ReAct loop tuned for local/smaller models. An agent that… 18 MIT Technology Review — AI news-outlet 17d ago Samsung’s chip workers are jumping ship to rival SK Hynix Lee, an engineer at Samsung’s semiconductor division, clocks out when his shift ends. He used to work longer hours, going the extra mile to excel at his projects. But lately, he’s been coming straight home to work on his job application for the chipmaker’s South Korean rival SK… 21 arXiv — Machine Learning research 17d ago Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control arXiv:2607.22779v1 Announce Type: new Abstract: Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control. In this field, deep learning approaches have become the gold standard. However, current architectures struggle to scale; model… 29 arXiv — NLP / Computation & Language research 17d ago Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training arXiv:2607.22794v1 Announce Type: cross Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability. To address this critical issue, we present the first… 22 arXiv — Machine Learning research 17d ago Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation arXiv:2607.23125v1 Announce Type: new Abstract: Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models,… 20 arXiv — NLP / Computation & Language research 17d ago Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh arXiv:2607.23446v1 Announce Type: new Abstract: A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples containing relevant law improves later use of retrieved law. We curate 2{,}165 bilingual QA… 31 arXiv — NLP / Computation & Language research 17d ago Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking arXiv:2607.23514v1 Announce Type: new Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using… 28 arXiv — NLP / Computation & Language research 17d ago StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting arXiv:2607.24191v1 Announce Type: new Abstract: Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly… 36 arXiv — NLP / Computation & Language research 17d ago Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair arXiv:2607.24604v1 Announce Type: new Abstract: Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed… 9 arXiv — NLP / Computation & Language research 17d ago Kimi K3: Open Frontier Intelligence arXiv:2607.24653v1 Announce Type: new Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention… 20 arXiv — NLP / Computation & Language research 17d ago MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models arXiv:2607.22586v1 Announce Type: cross Abstract: Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual… 4 arXiv — NLP / Computation & Language research 17d ago Revitalizing Public Urban Places through Cultural and Political Memory: A Technological Approach with LLMs and Augmented Reality arXiv:2607.22613v1 Announce Type: cross Abstract: This paper explores the intersection of memory, place, and identity, examining how new technologies, particularly Apple Vision Pro, can illuminate this nexus. Leveraging digital twins and virtual reality, it investigates how… 23 arXiv — NLP / Computation & Language research 17d ago RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus arXiv:2607.22709v1 Announce Type: cross Abstract: The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Memes often rely on a semantic clash between visual and textual modalities, where hateful… 17 arXiv — NLP / Computation & Language research 17d ago Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction arXiv:2607.22729v1 Announce Type: cross Abstract: Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text. This omits crucial visual cues, such as facial… 18 Hugging Face Daily Papers research 17d ago Kimi K3: Open Frontier Intelligence Abstract We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow… 7 Hugging Face Daily Papers research 17d ago JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Abstract Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative… 4 Hugging Face Daily Papers research 17d ago DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes Abstract While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a… 37 Hugging Face Daily Papers research 17d ago From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Abstract Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply… 15 Hugging Face Daily Papers research 17d ago OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Abstract Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging… 12 Hugging Face Daily Papers research 17d ago ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding Abstract Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and… 23 Hugging Face Daily Papers research 17d ago Data Pyramid for Embodied Manipulation Abstract Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by… 18 Hugging Face Daily Papers research 17d ago DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification Abstract Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may… 29 NVIDIA Developer Blog official-blog 17d ago NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they... 9 Hugging Face Daily Papers research 18d ago SceneActBench: Can Agents Act on the 3D Scenes They See? Abstract Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present… 36 Hugging Face Daily Papers research 18d ago Scaling Native Multimodal Pre-Training From Scratch Abstract Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on… 28 Hugging Face Daily Papers research 18d ago Multimodal Speaker Verification as a Threat to Speaker Anonymization Abstract Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic,… 27 Hugging Face Daily Papers research 18d ago Three-Body Scattering for Generative Modeling Abstract Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressive factorization. Instead, we show that a proper distributional energy can induce sample-level motion and provide direct regression supervision for a… 9 Smol AI News news-outlet 18d ago not much happened today **Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with… 33 arXiv — Machine Learning research 18d ago RED-PIM: Reducing Data Movement for Transformers using Processing-in-Memory arXiv:2607.21731v1 Announce Type: new Abstract: Transformers are widely used across many domains, including natural language processing, computer vision, web search, and DNA sequence analysis. Given their broad applicability, improving the performance of transformer models is… 16 arXiv — NLP / Computation & Language research 18d ago LeAct: Learning to Reason from Expert Actions arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines,… 17 arXiv — Machine Learning research 18d ago DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection arXiv:2607.22035v1 Announce Type: new Abstract: Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain… 25 arXiv — Machine Learning research 18d ago Autoregressive EHR Foundation Models with Multimodal Inputs arXiv:2607.22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a… 33 arXiv — Machine Learning research 18d ago IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data arXiv:2607.22351v1 Announce Type: new Abstract: The speed of sound in tissue is a prerequisite for well-focused imaging and has diagnostic value, but recovering it from raw pulse-echo channel data is fundamentally a nonlinear inverse problem. Learned solvers are fast yet label… 25 arXiv — Machine Learning research 18d ago LunarFM: A Shared Multimodal Representation of the Moon's Surface arXiv:2607.22408v1 Announce Type: new Abstract: The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar… 25 arXiv — Machine Learning research 18d ago FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs arXiv:2607.21624v1 Announce Type: cross Abstract: Transformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains… 20 arXiv — Machine Learning research 18d ago Computer Vision Based Neurology Brain Activity Rejection Architecture and Implementation arXiv:2607.21654v1 Announce Type: cross Abstract: The electroencephalogram (EEG) is a valuable and widely applied tool for investigating brain disorders and behavioral changes. It offers a minimally restrictive and non-invasive method. However, challenges in using EEG for… 36 arXiv — Machine Learning research 18d ago Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives arXiv:2607.21660v1 Announce Type: cross Abstract: Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in… 15 arXiv — NLP / Computation & Language research 18d ago Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks are often inconsistent and show unsatisfying… 23 arXiv — NLP / Computation & Language research 18d ago Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms arXiv:2607.21780v1 Announce Type: new Abstract: Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into their constituent documents is difficult, especially for low-resource languages.… 31 arXiv — NLP / Computation & Language research 18d ago Scaling Native Multimodal Pre-Training From Scratch arXiv:2607.22043v1 Announce Type: new Abstract: Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this… 31 arXiv — NLP / Computation & Language research 18d ago Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination arXiv:2607.22067v1 Announce Type: new Abstract: The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on… 26 arXiv — NLP / Computation & Language research 18d ago Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models arXiv:2607.21617v1 Announce Type: cross Abstract: Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful transcribers: when text is imperfect, they often tend… 6 arXiv — NLP / Computation & Language research 18d ago Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level… 15 arXiv — NLP / Computation & Language research 18d ago Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation arXiv:2607.22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting. In such settings, a reliable uncertainty signal matters more than raw… 4 arXiv — NLP / Computation & Language research 18d ago Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine-tuning, requiring large, task-specific speech… 17 llama.cpp releases dev-tools 18d ago b10142 mtmd: Add Vision Support for Minimax-M3 ( #25113 ) Add preliminary MiniMax-M3 support Text-only port that re-uses existing components: MiniMax-M2 style GQA with per-head QK-norm and partial rotary, DeepSeek-V3 style leading-dense and routed/shared experts, and swigluoai… 18 r/LocalLLaMA community 18d ago Vision Support for Minimax-M3 has been merged into llama.cpp   submitted by   /u/Time_Reaper [link]   [comments] 14 Hugging Face Daily Papers research 19d ago VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression Abstract Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated extensive research on visual token compression. While training-free strategies rely on heuristic metrics and suffer… 12 Page 7 of 10 · 500 articles ← Newer Older →