News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow llama.cpp releases dev-tools 1h ago b11227 context : do not re-reserve the scheduler when toggling causal_attn ( #28751 ) context : do not re-reserve the scheduler when toggling causal_attn llama_context::set_causal_attn() marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this… 27 arXiv — NLP / Computation & Language research 7h ago Recursive Self-Improvement via On-Policy Distillation for Reasoning arXiv:2609.30652v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the… 15 arXiv — NLP / Computation & Language research 7h ago SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages arXiv:2609.30739v1 Announce Type: new Abstract: Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing… 10 arXiv — NLP / Computation & Language research 7h ago Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting arXiv:2609.31169v1 Announce Type: new Abstract: Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the… 16 arXiv — NLP / Computation & Language research 7h ago Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers arXiv:2609.31403v1 Announce Type: new Abstract: Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the… 24 arXiv — NLP / Computation & Language research 7h ago ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs arXiv:2609.31448v1 Announce Type: new Abstract: Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from… 29 arXiv — NLP / Computation & Language research 7h ago What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study arXiv:2609.30402v1 Announce Type: cross Abstract: Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices… 4 arXiv — NLP / Computation & Language research 7h ago Affective Flow Language Model for Emotional Support Conversation arXiv:2602.08826v3 Announce Type: replace Abstract: Large language models (LLMs) have advanced emotional support conversation, but existing alignment methods rely mainly on sparse preferences at the response level or outcomes at the dialogue level, providing limited supervision… 36 NVIDIA Developer Blog official-blog 10h ago How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency Every unused watt is capacity left on the table. AI factories are typically provisioned for the unlikely moment when every GPU reaches peak power, creating a... 15 r/LocalLLaMA community 1d ago Which of the 16gb VRAM qwen3.8 27b’s is the best? I’m having a hard time finding out which one gives you fastest speed, maximum context with best possible quality. I can run unsloth qwen3.8 27b iq4_xs with 65k q8 kv, context without MTP and vision offloaded to cpu. But also kinda slow for agentic work at like 30ish tok/s )I… 30 r/LocalLLaMA community 2d ago internlm/Intern-Decision 4B and 0.8B https://huggingface.co/internlm/Intern-Decision-0.8B Intern-Decision-4B Demo | Model Weights | GitHub Intern-Decision-4B is a multimodal structured decision model fine-tuned from Qwen3.5-4B . It accepts a shared state, a schema of named questions, and optional images, and… 6 r/MachineLearning community 2d ago What are people building in computer vision, and what's still painful? [D] I've built a lot of ML systems over the years, mainly computer vision models optimised to run on mobile phones. For example, my previous company built the food recognition model for MyFitnessPal. I'm interested in what people are actually deploying in industry now. Are edge… 17 llama.cpp releases dev-tools 2d ago b11183 metal : split fa kernels into per-dtype libraries ( #29329 ) metal : split fa kernels into per-dtype libraries Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp cont : minor fix comment Website: https://llama.app Attestations:… 11 arXiv — Machine Learning research 3d ago Time-Series Foundation Models That Understand Data Revisions arXiv:2609.28576v1 Announce Type: new Abstract: Historical observations are not always fixed: statistical agencies revise previously published values as new evidence arrives. Forecasting from a contemporary download can therefore expose a model to information unavailable at the… 13 arXiv — Machine Learning research 3d ago UO-FIE: Combining Exact-Label Supervision with Graded Utility for Factivity Inference arXiv:2609.28605v1 Announce Type: new Abstract: The Factivity Inference Evaluation 2026 (FIE2026) classifies Chinese context-hypothesis pairs into nine ordered factivity intervals. Its evaluation metric rewards both exact predictions and proximity to the correct interval, while… 19 arXiv — NLP / Computation & Language research 3d ago LastOPD: Taming Collapse in Latent On-Policy Distillation arXiv:2609.28845v1 Announce Type: cross Abstract: On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises… 28 arXiv — Machine Learning research 3d ago Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD arXiv:2609.29142v1 Announce Type: new Abstract: Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense… 13 arXiv — Machine Learning research 3d ago ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models arXiv:2609.29398v1 Announce Type: new Abstract: Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation… 7 arXiv — NLP / Computation & Language research 3d ago Reasoning Instructions Can Break Answer Decoding in Vision--Language Models arXiv:2609.29278v1 Announce Type: new Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On… 7 arXiv — NLP / Computation & Language research 3d ago Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams arXiv:2609.29333v1 Announce Type: new Abstract: One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded… 24 arXiv — NLP / Computation & Language research 3d ago ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts arXiv:2609.29349v1 Announce Type: new Abstract: ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for… 29 arXiv — NLP / Computation & Language research 3d ago Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions arXiv:2609.29496v1 Announce Type: new Abstract: Natural language explanation generation serves as a key mechanism for exposing and evaluating vision-language reasoning. Prior work on explanation-driven vision-language models predominantly follows a post-hoc (answer-first)… 11 arXiv — NLP / Computation & Language research 3d ago LLMersion: A Local-First AI Agent Framework for Low-Cost Home Language Learning toward Educational Equity arXiv:2609.29672v1 Announce Type: new Abstract: Artificial intelligence helps education most where an essential provision has been rationed by cost. For language learners that provision is a teacher's voice, which binds listening, reading, speaking, and writing into one act.… 8 arXiv — NLP / Computation & Language research 3d ago An Empirical Study of VLM Pipelines for Long-Document QA arXiv:2609.29933v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document to the model,… 11 arXiv — NLP / Computation & Language research 3d ago Return or Revise? Learning When Revision Helps Retrieval-Augmented QA arXiv:2609.30087v1 Announce Type: new Abstract: We consider the decision of whether to return an existing draft answer or revise it using retrieved evidence, as in answer-revision systems. Draft confidence estimates whether the current answer is correct, but the decision… 10 arXiv — NLP / Computation & Language research 3d ago SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data arXiv:2609.30238v1 Announce Type: new Abstract: Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information… 27 arXiv — NLP / Computation & Language research 3d ago Small yet Assistive: Spatially-Aware Post-Training for Low Vision arXiv:2609.28757v1 Announce Type: cross Abstract: An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate… 4 llama.cpp releases dev-tools 3d ago b11172 metal : optimize sparse FA + clean-up ( #29377 ) metal : cache sparse FA indices in shared memory Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp metal : simplify shared memory size calculation Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp pi : update general… 7 The Information — AI news-outlet 3d ago Former OpenAI Data Center Chief Is Now At Nvidia Chris Malone, OpenAI’s former head of data centers, joined Nvidia as the vice president of Nvidia’s DSX Platform this month, according to his LinkedIn profile. DSX is the Nvidia division that helps customers design and build AI data centers according to Nvidia’s specifications.… 4 Hugging Face official-blog 3d ago Accelerating vision-language models with LFM2.5-VL-DSpark Back to Articles a]:hidden"> Accelerating vision-language models with LFM2.5-VL-DSpark Team Article Published September 24, 2026 Upvote 4 xx tugot17 LiquidAI Yuri Khrustalev ykhrustalev LiquidAI Leonie Monigatti iamleonie LiquidAI Viviana Márquez vivianamarquez LiquidAI… 19 arXiv — Machine Learning research 4d ago A Leakage-Aware Multimodal Evaluation Framework for Early Intraoperative Acute Kidney Injury Prediction arXiv:2609.26848v1 Announce Type: new Abstract: Postoperative acute kidney injury (AKI) after major non-cardiac surgery carries substantial morbidity, yet early intraoperative risk stratification remains difficult. In this retrospective cohort study, we propose SynerT, a… 35 arXiv — Machine Learning research 4d ago CORE-STACK+: Meta-Learning for Deep Stacked Generalization arXiv:2609.26905v1 Announce Type: new Abstract: Stacking heterogeneous vision backbones (CNNs, ViTs, and hybrids) is the de facto recipe for accuracy, calibration, and robustness, yet two coupled pathologies limit its returns. Prediction-space multicollinearity ill-conditions… 9 arXiv — Machine Learning research 4d ago A Scaling Study for fMRI Foundation Models arXiv:2609.27232v1 Announce Type: new Abstract: Scaling laws have guided large-model development in computer vision and natural language processing, but the relationships among data, model size, and compute remain unclear for functional magnetic resonance imaging (fMRI)… 26 arXiv — Machine Learning research 4d ago PhyMo: A Physical-Field Modality for Multimodal AI4Physics arXiv:2609.27554v1 Announce Type: new Abstract: Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However,… 4 arXiv — Machine Learning research 4d ago VCMM: Variance-Calibrated Momentum for Multimodal Learning arXiv:2609.27577v1 Announce Type: new Abstract: Multimodal joint training often suffers from modality imbalance, where a dominant modality suppresses the optimization of others. Existing methods mainly balance modality learning by modulating gradient magnitudes or directions,… 28 arXiv — Machine Learning research 4d ago NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers arXiv:2609.27735v1 Announce Type: new Abstract: Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to… 29 arXiv — Machine Learning research 4d ago LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations arXiv:2609.28086v1 Announce Type: new Abstract: We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or… 37 arXiv — NLP / Computation & Language research 4d ago Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation arXiv:2609.26926v1 Announce Type: new Abstract: Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this… 31 arXiv — NLP / Computation & Language research 4d ago Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition arXiv:2609.27289v1 Announce Type: new Abstract: Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an… 38 arXiv — NLP / Computation & Language research 4d ago Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents arXiv:2609.27353v1 Announce Type: new Abstract: Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena… 5 arXiv — NLP / Computation & Language research 4d ago PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models arXiv:2609.27395v1 Announce Type: new Abstract: Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number,… 15 arXiv — NLP / Computation & Language research 4d ago Exact Feedback Is Not Control: Evaluating Text-based Closed-Loop Revision in LLMs arXiv:2609.28150v1 Announce Type: new Abstract: Closed-loop revision is increasingly used in large language model (LLM) applications, but failures may reflect incomplete feedback or ineffective responses to correct feedback. We introduce a fixed-budget revision protocol with… 32 arXiv — NLP / Computation & Language research 4d ago Small Cues, Big Consequences: Learning Pivotal Cues for Multimodal Meme Classification arXiv:2609.26907v1 Announce Type: cross Abstract: Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal classifiers can miss such evidence when relying mainly on global image-text… 21 arXiv — NLP / Computation & Language research 4d ago ContraVis: Evidence-Grounded Visual Analytics for Contradiction Review in Legal Contracts arXiv:2609.27014v1 Announce Type: cross Abstract: Legal contracts are structurally complex documents in which contradictions may emerge across distant and interconnected provisions. Although large language models (LLMs) improve legal language understanding, contradiction… 12 arXiv — NLP / Computation & Language research 4d ago Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges arXiv:2609.27110v1 Announce Type: cross Abstract: Vision-language models (VLMs) are deployed as zero-shot judges of image aesthetics, and panels of several models are recommended, on thin evidence, as the way to make such judges reliable. On two human-rated datasets, EVA and… 38 arXiv — NLP / Computation & Language research 4d ago What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit arXiv:2609.27408v1 Announce Type: cross Abstract: Benchmarks for vision-language models offer their answer choices in some convention: a letter, a color name, a pixel coordinate. That convention is treated as neutral. We find it is not, and that the limits a benchmark reports… 26 The Information — AI news-outlet 4d ago Google Nears Release of Flagship Gemini 4 AI Model Google is nearing the release of its newest flagship model, Gemini 4, the head of its DeepMind division said Wednesday, a long-awaited development after the tech giant fell behind rivals Anthropic and OpenAI in the AI model race. Speaking at The Information’s AI Agenda Live… 13 llama.cpp releases dev-tools 4d ago b11151: convert : allow vision target for DFlash/Dspark (#29339) Resolve the target arch with get_model_architecture so vision targets (e.g. Lfm2VlForConditionalGeneration) map to their text model for the vocab. Fix double rope reorder for LFM2/LFM2.5 DSpark drafters 17 The Information — AI news-outlet 4d ago Patreon Co-Founder Sam Yam Joins OpenAI to Lead New Creator Division OpenAI has hired Patreon co-founder Sam Yam to lead a new creator product division, Creator Product, Yam announced Wednesday on X , a sign of renewed interest in attracting consumers. Two other former Patreon executives, head of product Drew Rowny and head of engineering Shannon… 19 r/LocalLLaMA community 4d ago apple/LensVLM-9B · Hugging Face https://huggingface.co/bartowski/LensVLM-9B-GGUF LensVLM-9B LensVLM is a 9B Vision Language Model (VLM) that scans compressed images of text, then selectively expands only the relevant pages to their uncompressed form via learned tools. Paper: LensVLM: Selective Context… 29 Page 1 of 10 · 500 articles Older →