News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow r/LocalLLaMA community 4d ago Space Bunny is the new stealth model in OpenCode. Free to try, multimodal I’m really curious to know which company released it.   submitted by   /u/Greney_Yunan [link]   [comments] 31 Latent.Space news-outlet 4d ago 🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics) Radical Numerics is using biological chain-of-thought and multimodal perception to keep up with the bio-defense arms race, design new genomes and gain insights into biology itself. 12 arXiv — Machine Learning research 5d ago Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow arXiv:2609.25131v1 Announce Type: new Abstract: Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also exposes correct intermediate predictions to later errors. An experiment on Sudoku… 20 arXiv — Machine Learning research 5d ago A JEPA Recipe for Tabular Foundation Models arXiv:2609.25541v1 Announce Type: new Abstract: Tabular foundation models learn to predict cell values in context, whereas world-model self-supervision asks for prediction in representation space (LeCun, 2022; Assran et al., 2023). On a tabular foundation-model prior, the latent… 22 arXiv — Machine Learning research 5d ago Margin-Drop Coordinates for Cross-Budget Robustness Evaluation arXiv:2609.26081v1 Announce Type: new Abstract: Fixed-budget robustness evaluation can select the wrong frozen vision encoder. An encoder that survives a shallow attack may lose most of that robustness when the same evaluation is strengthened. We ask whether the shallow… 21 arXiv — Machine Learning research 5d ago Beyond Imitation: Auditing the Recoverability of Reasoning in Distilled Models arXiv:2609.26216v1 Announce Type: new Abstract: A correct teacher solution becomes useful supervision when the receiving student can continue its reasoning. We measure this compatibility with prefix recovery: after revealing 25%, 50%, or 75% of a verified solution, we test… 19 arXiv — NLP / Computation & Language research 5d ago ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains arXiv:2609.25055v1 Announce Type: new Abstract: In this report we present results of the ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains. This competition aimed to advance research in document understanding through the task of Visual Question… 14 arXiv — NLP / Computation & Language research 5d ago Qwen3.8-Omni: Towards Native Omni-Modal Agents arXiv:2609.25611v1 Announce Type: new Abstract: We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash… 30 arXiv — NLP / Computation & Language research 5d ago Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance arXiv:2609.26035v1 Announce Type: new Abstract: Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and explicit belief revision can be implemented as a behavior layer over a fixed… 15 arXiv — NLP / Computation & Language research 5d ago SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children arXiv:2609.26090v1 Announce Type: new Abstract: Language is the target of most early intervention for autistic children. Because the goal and the method change from child to child, the work falls to a teacher who takes one child at a time and judges each scene as it unfolds.… 18 arXiv — NLP / Computation & Language research 5d ago One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data arXiv:2609.26097v1 Announce Type: new Abstract: Remote-sensing (RS) multimodal large language models (MLLMs) are trained and evaluated only in English, while text-only instruction data covers over 100 languages. We propose MODL (Mutually Orthogonal Domain-Language composition),… 19 arXiv — NLP / Computation & Language research 5d ago Modality-Gated Deep Adapters: Adding a Modality to a Frozen Embedding Model with Exact Preservation arXiv:2609.26182v1 Announce Type: new Abstract: Multimodal embedding models are deployed at scale: retrieval indices, benchmark results, and behavioral audits all depend on the base model's exact outputs. Extending such a model to a new modality with existing parameter-efficient… 29 arXiv — NLP / Computation & Language research 5d ago Beyond Static Charts: Can Language and Vision Language Models Generate Interactive Data Visualization Interfaces? arXiv:2609.26208v1 Announce Type: new Abstract: Data visualization is central to analytical reasoning, but real-world analysis increasingly requires language-driven interactive interfaces rather than static charts. Although recent large language and vision language models… 25 arXiv — NLP / Computation & Language research 5d ago Same Chart, Different Story: Bias in Vision-Language Chart Interpretation arXiv:2609.26210v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to interpret charts and generate natural-language explanations for socially consequential data. However, they may produce different narratives for the same chart when only the… 11 arXiv — NLP / Computation & Language research 5d ago Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion arXiv:2609.26381v1 Announce Type: new Abstract: Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems. Recent vision-based parsers improve accuracy, but need GPUs and may introduce noise into the extracted text.… 34 arXiv — NLP / Computation & Language research 5d ago Combining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding arXiv:2609.26399v1 Announce Type: new Abstract: Scene safety understanding plays a life-or-death role in situational awareness in various critical domains. Traditional methods that rely on learning direct mappings between scenes and safety levels often lack interpretability,… 14 arXiv — NLP / Computation & Language research 5d ago Knowledge Pull Requests for Continual Document Authoring arXiv:2609.26634v1 Announce Type: new Abstract: We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times,… 33 arXiv — NLP / Computation & Language research 5d ago Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding arXiv:2609.26638v1 Announce Type: new Abstract: Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed. Unlike open-ended text generation,… 5 r/LocalLLaMA community 5d ago Fork of FreeToken with DeepSeek-V4.1, vision and speculative decoding (2x3090 numbers inside) I've been running FreeToken on my 2x3090 box for a while and ended up maintaining a fork of it. Posting it in case it's useful to anyone else here. Quick context if you haven't used it: FreeToken is an edge-native MoE serving engine. It offloads experts to host RAM/NVMe and… 31 arXiv — Machine Learning research 6d ago Generalized Multimodal Foundation Model arXiv:2609.22107v1 Announce Type: new Abstract: Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single tasks, making it… 31 arXiv — Machine Learning research 6d ago DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction arXiv:2609.22184v1 Announce Type: new Abstract: Drug-target relation prediction supports candidate screening, drug repositioning, and mechanism analysis. Existing models often use incomplete drug or protein representations, model cross-modal interactions shallowly, or train… 26 arXiv — Machine Learning research 6d ago Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation arXiv:2609.22254v1 Announce Type: new Abstract: On-policy distillation (OPD) is a promising approach for transferring knowledge between language models, where a student receives dense token-level supervision along its own generated trajectories. However, teacher supervision can… 16 arXiv — Machine Learning research 6d ago Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization arXiv:2609.22879v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, but fine-tuning them with reinforcement learning (RL) remains constrained by the cost of real-world robot interaction. Model-based… 38 arXiv — NLP / Computation & Language research 6d ago Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation arXiv:2609.22094v1 Announce Type: new Abstract: Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining for every policy change and suffering from label scarcity since multimedia cannot be… 32 arXiv — NLP / Computation & Language research 6d ago Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization arXiv:2609.22149v1 Announce Type: new Abstract: Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after tabular--text pairings are disrupted. We present a projection-based evaluator for… 20 arXiv — NLP / Computation & Language research 6d ago Beyond Raw Context Transfer: Representation-based Federated Retrieval-Augmented Generation arXiv:2609.22162v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves the factuality of large language models (LLMs) and vision-language models (VLMs) by grounding generation in external knowledge. However, most existing RAG frameworks assume a… 22 arXiv — NLP / Computation & Language research 6d ago Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models arXiv:2609.22206v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance across a wide range of multimodal tasks, yet understanding and quantifying their predictive uncertainty remains underexplored despite being central for… 36 arXiv — NLP / Computation & Language research 6d ago SALSA: Semi-Autonomous Literature Summarization Assistant arXiv:2609.22210v1 Announce Type: new Abstract: SALSA (Semi-Autonomous Literature Summarization Assistant) is an open- source, human-in-the-loop platform for extracting structured scientific datasets from multimodal literature sources. The software combines document parsing,… 36 arXiv — NLP / Computation & Language research 6d ago Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models arXiv:2609.22234v1 Announce Type: new Abstract: Instruction hierarchy (IH) alignment teaches language models to prioritize higher-level instructions when inputs conflict. While studied primarily in text-only settings, vision-language models (VLMs) introduce new challenges for… 26 arXiv — NLP / Computation & Language research 6d ago MIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills Assessment arXiv:2609.22778v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require expert judgment remains unclear. We investigate this challenge in the context of… 15 arXiv — NLP / Computation & Language research 6d ago Automatic multimodal UX improvement recommendations from LLM agent user simulations arXiv:2609.22971v1 Announce Type: new Abstract: Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However,… 22 llama.cpp releases dev-tools 6d ago b11081 test-llama-archs : make tensor data stdev configurable and improve help ( #29133 ) test-llama-archs : make tensor data stdev configurable Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp test-llama-archs : expand usage and add examples Assisted-by:… 19 llama.cpp releases dev-tools 6d ago b11075 ggml-metal : simplify fusion pattern op list declaration ( #29206 ) ggml-metal : derive non-empty fusion ops from ops_all Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp ggml-metal : drop _all suffix from fusion op pattern vectors Assisted-by:… 4 arXiv — Machine Learning research 7d ago Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces arXiv:2609.20904v1 Announce Type: new Abstract: Hybrid motor-imagery brain-computer interfaces (MI-BCIs) combining EEG and fNIRS can outperform EEG-only systems by exploiting complementary electrophysiological and hemodynamic information. To obtain such hybrid information when… 7 arXiv — Machine Learning research 7d ago From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities arXiv:2609.20991v1 Announce Type: new Abstract: Physiological emotion recognition using wearable sensors has important applications in mental health monitoring, affective computing, and human-computer interaction. However, existing studies typically evaluate a single model,… 10 arXiv — Machine Learning research 7d ago M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection arXiv:2609.21164v1 Announce Type: new Abstract: Integrating diverse data modalities --- such as clinical notes, laboratory results, and medical imaging --- is essential for advancing clinical decision-making. While Large Language Models (LLMs) have shown remarkable performance… 25 arXiv — Machine Learning research 7d ago On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation arXiv:2609.21561v1 Announce Type: new Abstract: On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into the model. However, privileged information can change… 19 arXiv — Machine Learning research 7d ago Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise arXiv:2609.22053v1 Announce Type: new Abstract: Graph Convolutional Networks (GCNs) are highly sensitive to label noise, since corrupted supervision can propagate through the graph and degrade learned node representations. This work proposes PCC+GCN, a hybrid framework that uses… 25 arXiv — NLP / Computation & Language research 7d ago VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering arXiv:2609.20843v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) enables models to answer natural-language questions through structured graph reasoning and has achieved substantial progress across many benchmarks and applications. Recently, multimodal… 28 arXiv — NLP / Computation & Language research 7d ago Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders arXiv:2609.20845v1 Announce Type: new Abstract: A decoder that turns video or audio into text conventionally consumes the entire input before emitting a word. Offline this is merely more than the task requires; live it is impossible, since a caption cannot wait for a match to… 21 arXiv — Machine Learning research 7d ago Fragment-Aware Vision Transformers for Fresco-Fragment Style Classification arXiv:2609.21012v1 Announce Type: cross Abstract: Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconographic structure. In archaeological settings, however, artworks often survive… 15 arXiv — NLP / Computation & Language research 7d ago Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions arXiv:2609.20830v1 Announce Type: new Abstract: Revision-capable generation is appealing because it can insert or revise earlier content, but many non-autoregressive and edit-based approaches obtain this flexibility through repeated sequence-level computation. We propose… 11 arXiv — NLP / Computation & Language research 7d ago MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs arXiv:2609.20850v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related… 29 arXiv — NLP / Computation & Language research 7d ago Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction arXiv:2609.21392v1 Announce Type: new Abstract: Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of… 28 arXiv — NLP / Computation & Language research 7d ago DiaVLo: Diagnosing Behaviours of Vision-Language Models arXiv:2609.22008v1 Announce Type: new Abstract: Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable… 12 arXiv — NLP / Computation & Language research 7d ago Offline Multimodal Large Language Models for Decision Support in Air Operations arXiv:2609.21390v1 Announce Type: cross Abstract: Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often… 29 arXiv — NLP / Computation & Language research 7d ago Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis arXiv:2609.21651v1 Announce Type: cross Abstract: Farmer.Chat is Digital Green's farm advisory service for smallholder farmers. When something looks wrong with a crop, the farmer takes a photograph and sends it, and that photograph is the whole question: no symptom described, no… 31 arXiv — NLP / Computation & Language research 7d ago The Spoken Wikipedia Presentation Corpus arXiv:2609.21676v1 Announce Type: cross Abstract: We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that… 23 arXiv — NLP / Computation & Language research 7d ago Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model arXiv:2601.06848v2 Announce Type: replace Abstract: Multimodal aspect-based sentiment analysis (MABSA) aims to identify aspect-level sentiments by jointly modeling textual and visual information, which is essential for fine-grained opinion understanding in social media. Existing… 14 r/LocalLLaMA community 7d ago One more 'you should try ExllamaV3/exl3 for flash next' appreciation post After seeing a few posts on here about it, I finally tried exl3 3bpw and exllamav3 for running flash next - with amazing results. On 3x3090s, 128GB DDR4: 1500 prefill, 80 tps decode On 1x5090, 128G. DDR4: 1500 prefill, 29 tps decode Both at 262k context, both with vision/spec… 9 Page 2 of 10 · 500 articles ← Newer Older →