Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 10d ago
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
Abstract Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent…
20 -
Hugging Face Daily Papers research 11d ago
Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
Abstract When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present Fathom, a key scan…
36 -
Hugging Face Daily Papers research 11d ago
In-Context Robot Learning with VLM Agents
Abstract Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment…
24 -
Hugging Face Daily Papers research 11d ago
Assessing nnU-Net Generalization across Brain Tumor Populations in BraTS-GoAT 2026
Abstract BraTS-GoAT evaluates tumor segmentation across heterogeneous populations. We trained a conventional 3D nnU-Net on 1,351 labeled cases using five-fold cross-validation and 1,000 epochs per fold. The final predictor averaged all folds and applied test-time mirroring. On…
11 -
Hugging Face Daily Papers research 11d ago
The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
Abstract Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's…
28 -
Hugging Face Daily Papers research 11d ago
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
Abstract Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains…
19 -
Hugging Face Daily Papers research 11d ago
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
Abstract Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint. Four are left unbounded by the parallelism plans in…
5 -
Hugging Face Daily Papers research 11d ago
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
Abstract In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value…
4 -
Hugging Face Daily Papers research 11d ago
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
Abstract Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators,…
16 -
Hugging Face Daily Papers research 11d ago
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
Abstract Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking…
17 -
Hugging Face Daily Papers research 11d ago
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
Abstract Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction…
21 -
Hugging Face Daily Papers research 11d ago
A Zeroth-Order Paradigm for LLM Preference Alignment
Abstract Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference…
34 -
Hugging Face Daily Papers research 11d ago
HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
Abstract Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different…
12 -
Hugging Face Daily Papers research 11d ago
Agora: Git as Shared Memory for Collective AutoResearch
Abstract Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratch, so more agents tend to mean more duplicated search rather than more discovery. Agora is a shared…
9 -
Hugging Face Daily Papers research 11d ago
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention
Abstract Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization…
4 -
Hugging Face Daily Papers research 11d ago
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
Abstract Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather…
33 -
Hugging Face Daily Papers research 11d ago
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
Abstract Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning…
19 -
Hugging Face Daily Papers research 11d ago
EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset
Abstract 3D hand mesh reconstruction is a challenging yet essential task for downstream applications, including human-robot interaction and AR/VR. Although conventional cameras have been widely adopted for this task, methods that rely on them struggle in low-light environments…
22 -
Hugging Face Daily Papers research 11d ago
Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
Abstract In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson…
13 -
Hugging Face Daily Papers research 11d ago
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
Abstract We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three…
17 -
Hugging Face Daily Papers research 11d ago
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Abstract Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce…
36 -
Hugging Face Daily Papers research 11d ago
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
Abstract We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling…
28 -
Hugging Face Daily Papers research 13d ago
Enabling Creative Exploration for Vibe Design Agents
Abstract Separating design exploration from code generation via structured intermediate specifications enables diverse UI alternatives without altering downstream generation settings. Generated by thinkingmachines/Inkling-Small Vibe design agents turn natural-language briefs…
7 -
Hugging Face Daily Papers research 13d ago
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Abstract AlayaVista decouples panoramic scene evolution from perspective video synthesis to enable efficient, high-fidelity interactive world modeling, supported by the MUGEN dataset. Generated by thinkingmachines/Inkling-Small Interactive video world models must maintain broad…
30 -
Hugging Face Daily Papers research 13d ago
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Abstract RSIAgent is a training-free multi-agent framework that enables recursive self-improvement via autonomous memory construction and broad-then-deep exploration to adapt digital agents to new environments. Generated by thinkingmachines/Inkling-Small Digital agents must…
6 -
Hugging Face Daily Papers research 13d ago
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Abstract HazardAuditor provides execution-grounded safety supervision for computer-use agents and introduces Guard Policy Optimization to align generative guard training with sequence-level safety outcomes. Generated by thinkingmachines/Inkling-Small Computer-use agents…
8 -
Hugging Face Daily Papers research 13d ago
Agent as Policy for Robotic Manipulation
Abstract A general-purpose agent directly controls a physical robot by interpreting visuals, writing executable programs, and revising actions based on physical feedback across diverse manipulation tasks. Generated by thinkingmachines/Inkling-Small We demonstrate that a…
36 -
Hugging Face Daily Papers research 13d ago
Realtime-Venus: A full-duplex interaction system with asynchronous delegation
Abstract Realtime-Venus is a proactive full-duplex system with separate audio-visual and audio models that integrate continuous perception, conversational control, and native speech generation via a shared causal timeline and dual-loop runtime. Generated by…
23 -
Hugging Face Daily Papers research 13d ago
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Abstract Backdoor vulnerability in fine-tuned language models varies drastically with poison selection, and a learned set-scoring method improves worst-case attack success by identifying high-impact poisoned examples. Generated by thinkingmachines/Inkling-Small Backdoor…
15 -
Hugging Face Daily Papers research 13d ago
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Abstract Elo-per-token analysis reveals that LLM agents initially scale faster than independent sampling but eventually slow, while parallel short sessions improve performance over single long runs. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents…
31 -
Hugging Face Daily Papers research 13d ago
Atria Dawn: The Dawn of Agentic Superintelligence
Abstract Atria Dawn Preview is a foundation agentic language model trained through verified tool interactions that achieves strong benchmark results and demonstrates a shift toward human-AI project-level collaboration in scientific research. Generated by…
9 -
Hugging Face Daily Papers research 13d ago
Kaininja: Extending Native 3D Generators to the Part Level
Abstract KaiNinja extends a native 3D generator to part-level outputs using a dual-volume representation that resolves interface conflicts, improving both part and whole-object fidelity without segmentation. Generated by thinkingmachines/Inkling-Small Native 3D generators turn…
4 -
Hugging Face Daily Papers research 13d ago
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Abstract ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency. Generated by…
26 -
Hugging Face Daily Papers research 13d ago
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
Abstract PhysBrain 1.5 unifies physical environment understanding, action generation, and future state prediction via joint autoregressive training on discrete vision-language, motion, and visual target sequences, achieving state-of-the-art open-source embodied performance.…
29 -
Hugging Face Daily Papers research 13d ago
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
Abstract The framework dynamically adjusts training prompts via exploration potential scoring and scaffolded rewrites to improve reinforcement learning for multimodal language models. Generated by thinkingmachines/Inkling-Small Training prompts in online reinforcement learning…
8 -
Hugging Face Daily Papers research 13d ago
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
Abstract Vidu S2 introduces real-time interactive avatar and video editing models that support high-resolution spatial video generation and dynamic reference updates. Generated by thinkingmachines/Inkling-Small We present Vidu S2, which comprises Vidu S2-Avatar, a real-time…
25 -
Hugging Face Daily Papers research 13d ago
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence
Abstract Discovery Foundation Models enable open-ended scientific discovery through iterative problem formulation, hypothesis testing, and evidence-based revision across dry and wet lab settings. Generated by thinkingmachines/Inkling-Small Foundation models have progressed from…
18 -
Hugging Face Daily Papers research 13d ago
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Abstract Attention-DP3 improves 3D diffusion policies by injecting object-level geometric cues via attention to stabilize performance under heavy clutter. Generated by thinkingmachines/Inkling-Small 3D point-cloud observations are inherently ambiguous in complex, cluttered…
27 -
Hugging Face Daily Papers research 13d ago
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Abstract A benchmark requiring agents to programmatically reconstruct real-world videos in Blender reveals that current models achieve high perceptual similarity but struggle to retain spatiotemporal facts. Generated by thinkingmachines/Inkling-Small Multimodal agents can create…
16 -
Hugging Face Daily Papers research 13d ago
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Abstract A 16.7B-parameter mixture-of-experts diffusion vision-language agent achieves strong multimodal GUI performance while preserving block-parallel decoding efficiency. Generated by thinkingmachines/Inkling-Small Diffusion large language models (dLLMs) achieve high decoding…
37 -
Hugging Face Daily Papers research 13d ago
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Abstract Dream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations. Generated by thinkingmachines/Inkling-Small Recursive self-improvement is becoming increasingly vital…
25 -
Hugging Face Daily Papers research 13d ago
Omni-Streaming Thinking
Abstract Omni-Streaming Thinking improves streaming omni-modal reasoning by deferring claims until cross-modal verification, reducing premature commitment and auditory hallucinations. Generated by thinkingmachines/Inkling-Small Streaming omni-modal models must decide what and…
14 -
Hugging Face Daily Papers research 13d ago
Studying Without a Syllabus: Task-Agnostic Environment Preprocessing
Abstract An agent can explore unfamiliar environments without task-specific guidance to build reusable artifacts that reduce later inference costs, though larger study budgets do not always improve results. Generated by thinkingmachines/Inkling-Small Before an LLM agent tackles…
33 -
Hugging Face Daily Papers research 13d ago
TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
Abstract TRACE is a training-free framework that ranks visual evidence by future utility and diversity, reserves native tokens for spatial coverage, and contracts retired frames to reduce latency and memory in GUI agents. Generated by thinkingmachines/Inkling-Small GUI agents…
8 -
Hugging Face Daily Papers research 13d ago
ActionSplice: In-Flight Action Editing for Interactive World Models
Abstract ActionSplice introduces counterfactual state transport to splice revised actions into chunk-autoregressive video world models without replaying completed evaluations, improving fidelity and speed. Generated by thinkingmachines/Inkling-Small Chunk-autoregressive video…
38 -
Hugging Face Daily Papers research 13d ago
DataFlex-RL: An Evaluation Platform for RLVR Data Policies
Abstract DataFlex-RL evaluates reinforcement learning data policies and finds that uniform sampling matches or exceeds adaptive rollout selection, reweighting, and domain mixing across math, logic, and science benchmarks. Generated by thinkingmachines/Inkling-Small Data policies…
8 -
Hugging Face Daily Papers research 13d ago
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Abstract A compact Thai text-to-speech model is distilled from a large voice-cloning teacher using synthetic data from brief voice references, achieving strong on-device accuracy and prosody with minimal errors. Generated by thinkingmachines/Inkling-Small In low-resource…
33 -
Hugging Face Daily Papers research 13d ago
How Far Can Synthetic Data Take Thai OCR?
Abstract A Thai OCR model trained solely on synthetic documents achieves strong real-world performance by isolating key transfer factors like typography, layout, and handwriting glyphs. Generated by thinkingmachines/Inkling-Small We investigate what makes synthetic OCR…
8 -
Hugging Face Daily Papers research 14d ago
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Abstract SAS improves sparse attention by training a selector end-to-end with language modeling loss via continuous gating inside attention softmax, yielding better context ranking under tight budgets. Generated by thinkingmachines/Inkling-Small Post-training attention…
21 -
Hugging Face Daily Papers research 14d ago
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
Abstract Diverse Skill Routing improves LLM agent skill selection by balancing relevance with non-redundancy via a determinantal point process, boosting multi-skill coverage. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents increasingly rely on…
27