Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 10d ago
Progressive Agent Skill Generation via Reinforcement Learning
Abstract Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence sources. In contrast, learning-based approaches offer a more unified way to model skill generation across heterogeneous…
24 -
Hugging Face Daily Papers research 10d ago
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Abstract Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use…
11 -
Hugging Face Daily Papers research 10d ago
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Abstract Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from…
6 -
Hugging Face Daily Papers research 10d ago
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Abstract Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to…
38 -
Hugging Face Daily Papers research 10d ago
UEmbed: Unified Sparse and Dense Multimodal Embeddings
Abstract Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to…
5 -
Hugging Face Daily Papers research 10d ago
Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
Abstract Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations,…
38 -
Hugging Face Daily Papers research 10d ago
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Abstract Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural…
4 -
Hugging Face Daily Papers research 10d ago
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
Abstract Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or…
24 -
Hugging Face Daily Papers research 10d ago
StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field
Abstract Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local…
32 -
Hugging Face Daily Papers research 10d ago
ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
Abstract To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expose semantic tool schemas in static…
24 -
Hugging Face Daily Papers research 10d ago
Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis
Abstract Existing indoor layout generators produce globally plausible layouts yet may retain local violations such as collisions, out-of-bounds placements, obstructed openings, and blocked circulation. Most prior work focuses on full-scene synthesis or scene-level optimization,…
8 -
Hugging Face Daily Papers research 10d ago
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Abstract Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the…
28 -
Hugging Face Daily Papers research 10d ago
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Abstract Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and…
33 -
Hugging Face Daily Papers research 10d ago
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
Abstract The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over…
14 -
Hugging Face Daily Papers research 11d ago
Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
Abstract Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use…
31 -
Hugging Face Daily Papers research 11d ago
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Abstract Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages…
8 -
Hugging Face Daily Papers research 11d ago
Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
Abstract Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their…
17 -
Hugging Face Daily Papers research 11d ago
SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing
Abstract Autonomous multi-vehicle racing requires real-time planning of diverse competitive behaviors in intense interactions. Existing planners often struggle to balance strategic diversity and computational efficiency. To address this challenge, we propose Sampling-based…
32 -
Hugging Face Daily Papers research 11d ago
In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
Abstract Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks well-established…
18 -
Hugging Face Daily Papers research 11d ago
Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark
Abstract Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improved enhancement, yet most methods depend on carefully curated training data pairs, with limited robustness under…
33 -
Hugging Face Daily Papers research 11d ago
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
Abstract AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request…
15 -
Hugging Face Daily Papers research 11d ago
Mental World Modeling
Abstract World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends,…
33 -
Hugging Face Daily Papers research 11d ago
One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
Abstract Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding…
29 -
Hugging Face Daily Papers research 11d ago
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
Abstract Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a…
4 -
Hugging Face Daily Papers research 11d ago
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Abstract While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for…
5 -
Hugging Face Daily Papers research 11d ago
Meshy T2: Fast Native Mesh Generation with Flow Matching
Abstract Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is essential for film, gaming, and interactive 3D applications. Mainstream approaches serialize a mesh into a token sequence…
19 -
Hugging Face Daily Papers research 11d ago
ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
Abstract In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting significant inefficiency in capturing the dynamics of…
30 -
Hugging Face Daily Papers research 11d ago
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
Abstract As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term…
18 -
Hugging Face Daily Papers research 11d ago
N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
Abstract We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on…
31 -
-
Hugging Face Daily Papers research 11d ago
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
Abstract Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and…
8 -
Hugging Face Daily Papers research 11d ago
Enhancing Rubric-based RL via Self-Distillation
Abstract Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no optimization signal. Recent…
8 -
Hugging Face Daily Papers research 11d ago
Scaling Properties of Text Conditioning in Visual Generation
Abstract We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion…
23 -
Hugging Face Daily Papers research 11d ago
QQWorld: Quantile-Quantile Matching for World Model Regularization
Abstract Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic…
23 -
Hugging Face Daily Papers research 11d ago
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
Abstract We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current…
24 -
Hugging Face Daily Papers research 11d ago
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Abstract Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where…
38 -
Hugging Face Daily Papers research 11d ago
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Abstract System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability…
20 -
Hugging Face Daily Papers research 11d ago
RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
Abstract Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection…
36 -
Hugging Face Daily Papers research 13d ago
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
Abstract Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different…
27 -
Hugging Face Daily Papers research 13d ago
β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
Abstract On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD…
27 -
Hugging Face Daily Papers research 14d ago
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
Abstract Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate…
37 -
Hugging Face Daily Papers research 14d ago
Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Abstract This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method, this work focuses on causal…
37 -
Hugging Face Daily Papers research 14d ago
PhiZero: A World Model Built Around Physical Language
Abstract We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics…
26 -
Hugging Face Daily Papers research 14d ago
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Abstract Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes…
14 -
Hugging Face Daily Papers research 14d ago
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Abstract Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding…
22 -
Hugging Face Daily Papers research 14d ago
Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
Abstract Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rather than modeling which agents can be trusted and under what conditions. This limitation is particularly important in multi-agent systems, where a central…
18 -
Hugging Face Daily Papers research 14d ago
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Abstract Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage…
9 -
Hugging Face Daily Papers research 14d ago
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability
Abstract Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools. Yet research has largely passed over this medium: prior systems design…
37 -
Hugging Face Daily Papers research 14d ago
Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing
Abstract In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive,…
6 -
Hugging Face Daily Papers research 14d ago
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
Abstract Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, yet their reliability in open information environments remains underexplored. A key concern is whether apparently…
11