News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow r/LocalLLaMA community 19d ago Best chat model that fits in 128gb I'm looking for a model to chat with, reasoning, maybe get some career or life coaching. I don't care at all about multimodal or coding ability Just it's intelligence in remembering context in a conversation or a specific topic, thinking out of the box, etc. Must fit in 128gb,… 23 r/LocalLLaMA community 19d ago Benchmarks: TensorSharp vs. llama.cpp Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like Gemma4, DiffusionGemma, Qwen3.6 with multi-modal (image, vision, audio), Qwen… 38 r/MachineLearning community 20d ago Neurips Position Track Rebuttal and Reviews [R] Hello! This is my first time submitting an actual conference paper (only done workshops so far). Got a 3/3/5/7 for the Position Paper Track. Reviews all seem quite addressable. Meta review also seemed kinda positive? Included wording such as "a revision should include..."… 10 r/LocalLLaMA community 21d ago FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3   submitted by   /u/pmttyji [link]   [comments] 23 r/LocalLLaMA community 21d ago swiss-ai/Apertus-v1.5 70B/8B https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multilingual, multimodal, fully open, and transparent AI. The models support a wide… 23 Latent.Space news-outlet 21d ago [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model A HUGE win for BFL! 22 arXiv — Machine Learning research 21d ago Multimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs arXiv:2607.20517v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and the need for multi-hop reasoning across dispersed evidence. We present… 14 arXiv — Machine Learning research 21d ago ReliableTableQA:How Much Supervision Does Reliability Annotation Need? arXiv:2607.20537v1 Announce Type: new Abstract: We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether the query is answerable, but whether the computed answer is statistically meaningful. In real… 34 arXiv — Machine Learning research 21d ago Monkey King Bang: A Unified Scientific Multimodal Foundation Model arXiv:2607.20557v1 Announce Type: new Abstract: Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition. Existing systems are either specialised for individual domains or unify scientific… 30 arXiv — Machine Learning research 21d ago Adaptive Confidence-weighted Expansion for Trustworthy Multi-Omics Multimodal Fusion arXiv:2607.20742v1 Announce Type: new Abstract: Multimodal learning is a robust approach to improve predictive performance in applications such as medical prognosis. However, the clinical applicability of models that use multimodal learning is hampered by their poor performance… 22 arXiv — Machine Learning research 21d ago Best-of-Evidence: Best-of-N Selection under Partial Verification arXiv:2607.20950v1 Announce Type: new Abstract: BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reliably. Many vision-language tasks instead provide only partial… 7 arXiv — Machine Learning research 21d ago Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification arXiv:2607.21068v1 Announce Type: new Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reducing dependency on specialist-only assessment, for which neural network-based deep… 8 arXiv — Machine Learning research 21d ago Spectral Transformation for Layer-wise Global Rank Discovery in Federated LoRA for Vision Transformers arXiv:2607.21074v1 Announce Type: new Abstract: Fine-tuning Vision Transformers (ViTs) with low-rank adapters (LoRA) promises better communication efficiency under federated setup, yet existing aggregation strategies face fundamental limitations. Independently averaging these… 28 arXiv — Machine Learning research 21d ago The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works arXiv:2607.21273v1 Announce Type: new Abstract: Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow. We show that under group-normalized RL (GRPO), this… 28 arXiv — Machine Learning research 21d ago Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning arXiv:2607.21290v1 Announce Type: new Abstract: Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as modern game telemetry provides multiple related supervision signals from the same structured observations. We study… 18 arXiv — Machine Learning research 21d ago M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data arXiv:2607.21343v1 Announce Type: new Abstract: Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains… 16 Hugging Face Daily Papers research 21d ago ReferTrack: Referring Then Tracking for Embodied Visual Tracking Abstract Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their… 31 r/MachineLearning community 21d ago GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R] The interesting finding from a new [arXiv paper]( https://arxiv.org/abs/2607.16165 ) isn't that a frontier vision model failed a new benchmark, that happens weekly, but the specific shape of the failure and the fact that the models cannot patch it by writing their own code. The… 25 Hugging Face Daily Papers research 22d ago SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments Abstract Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to… 37 Hugging Face Daily Papers research 22d ago An Exam for Active Observers Abstract Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's… 4 Hugging Face Daily Papers research 22d ago Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations Abstract Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change… 9 Hugging Face Daily Papers research 22d ago Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Abstract Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible.… 36 arXiv — Machine Learning research 22d ago Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models arXiv:2607.19399v1 Announce Type: new Abstract: It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger… 34 arXiv — Machine Learning research 22d ago Trustworthy Privacy-Preserving Multimodal Federated Learning for Personalised Breast Cancer Prediction arXiv:2607.19532v1 Announce Type: new Abstract: Federated learning has emerged as a potential solution to privacy concerns associated with using sensitive health data for training predictive models, particularly in personalised cancer care. This research investigates whether… 22 arXiv — Machine Learning research 22d ago Time Series Network Utilization KPI Forecasting Using Advanced AI/ML Models arXiv:2607.19974v1 Announce Type: new Abstract: The rapid proliferation of data-intensive applications, cloud infrastructure, and IoT ecosystems has made proactive resource provisioning critical for maintaining optimal network performance. However, network administrators face a… 33 arXiv — NLP / Computation & Language research 22d ago VizRAG: Enhancing Retrieval-Augmented Generation with Hypergraph Visualization arXiv:2607.19830v1 Announce Type: new Abstract: Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic facts among entities, rather than relying solely on binary relationships. Despite the advancements in multimodal large… 4 arXiv — NLP / Computation & Language research 22d ago ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models arXiv:2607.20092v1 Announce Type: cross Abstract: Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a… 23 arXiv — NLP / Computation & Language research 22d ago Self-supervision drives representational convergence in medical foundation models more than clinical supervision arXiv:2607.20274v1 Announce Type: cross Abstract: Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure. Whether this convergence… 8 arXiv — NLP / Computation & Language research 22d ago Test-Time Training for Modality Order Consistency in Vision-Language Models arXiv:2607.20351v1 Announce Type: cross Abstract: We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently… 18 arXiv — NLP / Computation & Language research 22d ago Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations arXiv:2607.20379v1 Announce Type: cross Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false… 33 arXiv — NLP / Computation & Language research 22d ago Missing-by-Design: Certifiable Modality Deletion for Revocable Multimodal Sentiment Analysis arXiv:2602.16144v4 Announce Type: replace Abstract: As multimodal systems increasingly process sensitive personal data, the ability to selectively revoke specific data modalities has become a critical requirement for privacy compliance and user autonomy. We present… 37 arXiv — NLP / Computation & Language research 22d ago Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution arXiv:2604.03472v4 Announce Type: replace Abstract: Co-evolutionary self-play, where one language model generates problems and another solves them, promises curriculum learning without human supervision. The promise breaks down early in practice. The proposer converges to a… 10 arXiv — NLP / Computation & Language research 22d ago Emotion Collider: Dual Hyperbolic Mirror Manifolds for Sentiment Recovery via Anti Emotion Reflection arXiv:2602.16161v4 Announce Type: replace-cross Abstract: Emotional expression underpins natural communication and effective human-computer interaction. We present Emotion Collider (EC-Net), a hyperbolic hypergraph framework for multimodal emotion and sentiment modeling. EC-Net… 27 Hugging Face Daily Papers research 22d ago Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Abstract Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support… 31 Vercel — AI dev-tools 22d ago Inspect feature flag history with Vercel CLI Vercel Flags version history can now be inspected from the Vercel CLI with the new vercel flags versions command. Run vercel flags versions to print the full revision history for a flag, with each revision's author, message, timestamp, and changed environments. Filter to a… 13 r/LocalLLaMA community 22d ago microsoft/Fara1.5-27B · Hugging Face Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers . It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to… 37 llama.cpp releases dev-tools 23d ago b10085 mtmd : use align_corners for qwen3vl vision position embedding interpolation ( #25781 ) The Qwen3-VL learned position embedding is interpolated to the runtime patch grid with the default bilinear+antialias (align_corners=False) sampling, while the transformers reference uses… 32 Hugging Face Daily Papers research 23d ago Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges Abstract Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor… 15 Hugging Face Daily Papers research 23d ago Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Abstract Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers… 20 Hugging Face Daily Papers research 23d ago Delineate Anything v2: A Global Foundation Model for Field Delineation Abstract Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chain transparency, and carbon accounting. While vision foundation models like SAM show remarkable zero-shot capabilities, they frequently fail in geospatial… 17 Hugging Face Daily Papers research 23d ago EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration Abstract Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be… 5 Hugging Face Daily Papers research 23d ago HPD-Parsing: Hierarchical Parallel Document Parsing Abstract Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through… 26 arXiv — Machine Learning research 23d ago AHEAD: Advancing Multi-Class Label Aggregation with Interpretable Cross-Annotator Modeling arXiv:2607.18465v1 Announce Type: new Abstract: Crowdsourced labeling provides valuable labeled data for domains across natural language processing, computer vision, and video. Label aggregation aims to infer latent true labels from noisy and biased annotations, with the key… 4 arXiv — Machine Learning research 23d ago Robust Multi-View Classification under Noisy Supervision via Global Anchor Consensus arXiv:2607.18561v1 Announce Type: new Abstract: In recent years, multi-view learning has attracted increasing attention, as it integrates the complementary information of heterogeneous views. Most existing multi-view classification methods rely on accurate annotations to… 34 arXiv — Machine Learning research 23d ago KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale arXiv:2607.18885v1 Announce Type: new Abstract: Kernel-based alignment of CLIP toward a vision centric teacher such as DINOv2 (KUEA) improves CLIP's visual representations while preserving text-encoder compatibility, using a fixed trade-off weight tuned on curated ImageNet-1K.… 35 arXiv — Machine Learning research 23d ago One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models arXiv:2607.19128v1 Announce Type: new Abstract: Vision-language models (VLMs) provide a unified representation space for textual and visual information, yet their potential as general-purpose backbones for graph-structured data remains largely unexplored. In practice, attributed… 29 arXiv — NLP / Computation & Language research 23d ago PathReportEval: A Systematic Benchmark for Pathology Report Generation arXiv:2607.18448v1 Announce Type: new Abstract: Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure because existing studies use heterogeneous datasets, model settings, visual encoders,… 19 arXiv — NLP / Computation & Language research 23d ago Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning arXiv:2607.18615v1 Announce Type: new Abstract: Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combine a language backbone with visual components, which makes unlearning more complex. There is a surprising phenomenon when… 24 arXiv — NLP / Computation & Language research 23d ago Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv:2607.18666v1 Announce Type: new Abstract: A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio… 36 arXiv — NLP / Computation & Language research 23d ago HPD-Parsing: Hierarchical Parallel Document Parsing arXiv:2607.18839v1 Announce Type: new Abstract: Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page… 11 Page 8 of 10 · 500 articles ← Newer Older →