Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 18d ago
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Abstract WearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions. Generated by thinkingmachines/Inkling-Small Recent advances in wearable sensing enable…
18 -
Hugging Face Daily Papers research 18d ago
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
Abstract The study introduces a benchmark to evaluate AI agents using sparse autoencoders for autonomous mechanistic interpretability and feature discovery, revealing progress but significant gaps versus expert baselines. Generated by thinkingmachines/Inkling-Small While…
4 -
Hugging Face Daily Papers research 18d ago
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Abstract Φ-Bench evaluates large language models on open-ended engineering of the LLM infrastructure stack across tasks from kernel optimization to end-to-end system design. Generated by thinkingmachines/Inkling-Small Large language models (LLMs) have demonstrated remarkable…
8 -
Hugging Face Daily Papers research 18d ago
Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models
Abstract Diff-based code editing underperforms direct generation overall but excels only on short, localized edits such as refactoring and error fixes, a property termed task locality. Generated by thinkingmachines/Inkling-Small Large language models used for code editing can be…
21 -
Hugging Face Daily Papers research 18d ago
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Abstract Combining harness evolution with localized expert correction improves weaker models without disrupting their native planning style. Generated by thinkingmachines/Inkling-Small Agent harnesses (the system prompt, tool set, execution hooks, and context-management…
24 -
-
Hugging Face Daily Papers research 18d ago
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
Abstract The Discovery Certification Protocol validates AI research agent outcomes through executable recovery tests, controlled audits, and deterministic verification with finite-sample recovery bounds. Generated by thinkingmachines/Inkling-Small AI research agents combine…
8 -
Hugging Face Daily Papers research 18d ago
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
Abstract This work introduces a benchmark for evaluating whether large language models can understand evolving interpersonal dynamics to provide effective emotional support in multi-party settings. Generated by thinkingmachines/Inkling-Small Existing emotional support…
17 -
Hugging Face Daily Papers research 18d ago
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Abstract Puppeteer is a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, physically grounded gestures. Generated by thinkingmachines/Inkling-Small Generating co-speech gestures that are temporally…
38 -
Hugging Face Daily Papers research 18d ago
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Abstract DATPO improves reasoning coverage in large models by using difficulty-adaptive tree-structured rollouts with sentence-entropy-guided branching and diversity-aware optimization. Generated by thinkingmachines/Inkling-Small Reinforcement Learning with Verifiable Rewards…
37 -
Hugging Face Daily Papers research 18d ago
Show-Harness: Just a VLM Agent Can Play Robots
Abstract Show-Harness links vision-language models to robot control via discrete semantic actions interpreted by embodiment-specific modules, enabling zero-shot and efficient fine-tuned deployment across robots and GUIs. Generated by thinkingmachines/Inkling-Small Foundation…
6 -
Hugging Face Daily Papers research 18d ago
Revisiting Complete Reasoning Traces for Post-Training
Abstract Large language models gain reasoning improvements from truncated trajectory endpoints rather than full reasoning traces, reducing redundancy while benefiting supervised fine-tuning and reinforcement learning. Generated by thinkingmachines/Inkling-Small Large language…
5 -
Hugging Face Daily Papers research 18d ago
RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
Abstract RenderFormer-V2 is a transformer-based neural rendering model that handles diverse light-transport effects via a two-stage sequence-to-sequence architecture with improved attention and heterogeneous scene support. Generated by thinkingmachines/Inkling-Small We present…
10 -
Hugging Face Daily Papers research 18d ago
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Abstract VDiff-Bench evaluates multimodal language models on fine-grained image difference identification, revealing major weaknesses in detecting subtle low-level visual changes. Generated by thinkingmachines/Inkling-Small Multimodal Large Language Models (MLLMs) perform…
7 -
Hugging Face Daily Papers research 18d ago
NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting
Abstract NOAH is a generative transformer that models full multimodal patient journeys with continuous time dynamics and stochastic latent states, enabling forecasting, zero-shot classification, and counterfactual simulation across diverse clinical data. Generated by…
11 -
Hugging Face Daily Papers research 18d ago
MasterControl Seventeen Every Time
Abstract A governed analytics framework pairs language models for intent interpretation with deterministic policy execution of pre-approved programs, achieving full answer-and-evidence reliability where runtime-planning agents failed. Generated by thinkingmachines/Inkling-Small…
8 -
Hugging Face Daily Papers research 18d ago
Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
Abstract Agents can turn shared infrastructure into a channel for coordinated intrusion. The Hugging Face incident and a separate public-wiki investigation show why a security assessment may need evidence from several executions and the artifacts they leave behind. We argue that…
35 -
Hugging Face Daily Papers research 18d ago
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
Abstract TANGO is a vision-language framework that predicts whole-body joint actions for humanoid robots to navigate cluttered indoor environments using only simulated training data. Generated by thinkingmachines/Inkling-Small We study the problem of navigating cluttered indoor…
36 -
Hugging Face Daily Papers research 18d ago
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Abstract A procedural graph framework organizes agent actions into structured relational triplets, providing situational guidance and self-evolving topology to improve long-horizon tool use. Generated by thinkingmachines/Inkling-Small Large language models are increasingly…
30 -
Hugging Face Daily Papers research 18d ago
Learning 3D Editing without Paired Supervision via Generative Prior Distillation
Abstract A feed-forward 3D editing framework distills visual, semantic, and geometric priors from foundation models via differentiable rendering and 3D-aware distribution matching to avoid paired training data. Generated by thinkingmachines/Inkling-Small Instruction-guided 3D…
18 -
Hugging Face Daily Papers research 18d ago
What Did I Just Say? Self-Listening for Full-Duplex Speech Models
Abstract Self-Listening grounds interruption recovery in full-duplex spoken language models by feeding realized speech back as input, improving consistency with actually spoken responses. Generated by thinkingmachines/Inkling-Small Full-duplex spoken language models can listen…
4 -
Hugging Face Daily Papers research 18d ago
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Abstract SynthGait-19k is a large synthetic video dataset for gait analysis that enables benchmarking of video-based gait estimation and shows synthetic supervision transfers to real data. Generated by thinkingmachines/Inkling-Small Accurate estimation of clinically meaningful…
9 -
Hugging Face Daily Papers research 19d ago
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
Abstract OpenWAM factorizes world-action pretraining into modular components to identify key design principles, yielding a scalable open model with strong simulation and real-robot performance. Generated by thinkingmachines/Inkling-Small World-Action Models inherit world…
34 -
Hugging Face Daily Papers research 19d ago
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
Abstract A*-Thought-V2 models chain-of-thought reasoning as hidden-state trajectories to selectively retain explicit reasoning steps or compress them into continuous latent tokens, improving accuracy and efficiency. Generated by thinkingmachines/Inkling-Small Chain-of-Thought…
32 -
Hugging Face Daily Papers research 19d ago
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Abstract Marigold V2 repurposes diffusion transformers for monocular depth estimation via single-step flow-matching inference, semantic alignment, and a Sinkhorn-based two-stage fine-tuning protocol, yielding sharper out-of-distribution depth maps and strong results on related…
31 -
Hugging Face Daily Papers research 19d ago
RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
Abstract A feed-forward generative Transformer enables direct single- and multi-view image relighting by injecting target illumination via cross-attention and processing unordered views with permutation-invariant encodings, trained on a large synthetic dataset. Generated by…
9 -
Hugging Face Daily Papers research 19d ago
Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
Abstract UCF-Net improves deepfake detection by fusing CLIP and DINO representations with uncertainty-weighted hierarchical feature aggregation, achieving stronger cross-domain generalization. Generated by thinkingmachines/Inkling-Small The growing realism and accessibility of…
31 -
Hugging Face Daily Papers research 19d ago
Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
Abstract In multi-turn dialogues, a model's inference of a partner's expertise becomes readable in early layers but only causally active later, bounding intervention points for steering behavior. Generated by thinkingmachines/Inkling-Small A transformer can make an attribute…
5 -
Hugging Face Daily Papers research 19d ago
Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model
Abstract Cadence pairs a time-series foundation model with an adaptive arithmetic coder for error-bounded lossy compression, achieving substantial gains over classical predictors while reporting negative results on lossless coding and cross-batch determinism. Generated by…
34 -
Hugging Face Daily Papers research 19d ago
Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy
Abstract Adding Greek to a robot vision-language-action model via machine-translated instructions reveals measurement pitfalls and shows bilingual training improves performance over monolingual baselines, though it remains far below English levels. Generated by…
35 -
Hugging Face Daily Papers research 19d ago
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
Abstract A unified Bayesian variational framework combining spike-and-slab sparsity and Gaussian mixture quantization achieves high compression rates for large neural networks with minimal accuracy loss. Generated by thinkingmachines/Inkling-Small Compressing large-scale neural…
16 -
Hugging Face Daily Papers research 19d ago
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
Abstract ReactVAU enables real-time streaming video anomaly understanding via a fast detection module, persistent anomaly-aware memory, and an on-demand slow reasoning module that minimizes heavy model usage. Generated by thinkingmachines/Inkling-Small In this paper, we propose…
19 -
Hugging Face Daily Papers research 19d ago
Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions
Abstract Large language models encode whether structurally impossible math or code prompts are unanswerable via a hidden-state direction, but fail to abstain because this recognition signal is misaligned with safety-refusal pathways, indicating a routing rather than encoding…
13 -
Hugging Face Daily Papers research 19d ago
Omni Interaction Agent Technical Report
Abstract Gander is an end-to-end framework that integrates continuous multi-modal streaming, real-time full-duplex interaction, and agentic reasoning through a Cerebellum-Brain architecture and a chunk-level token stream design. Generated by thinkingmachines/Inkling-Small In…
34 -
Hugging Face Daily Papers research 19d ago
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Abstract RoboSPA is a large-scale robotic manipulation benchmark that evaluates vision-language-action models on fine-grained spatial reasoning and long-horizon procedural planning across progressively harder task variants. Generated by thinkingmachines/Inkling-Small…
25 -
Hugging Face Daily Papers research 19d ago
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Abstract CoVeR is a training-free spatial token selector that preserves 3D reasoning performance by enforcing exact budgets and full scene coverage across multi-view visual tokens. Generated by thinkingmachines/Inkling-Small Representing a 3D scene as multi-view images allows 2D…
17 -
Hugging Face Daily Papers research 19d ago
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Abstract AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.…
37 -
Hugging Face Daily Papers research 19d ago
TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
Abstract TransNormal-2 improves monocular normal estimation by correcting VAE reconstruction errors through geometry-aware training losses and a lightweight RGB-guided refinement module, achieving strong results with minimal annotations. Generated by…
29 -
Hugging Face Daily Papers research 19d ago
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
Abstract BeaconKV improves memory efficiency for long reasoning traces by using compact beacon queries to predict which past key-value pairs will be revisited, reducing cache size without sacrificing accuracy. Generated by thinkingmachines/Inkling-Small Large Reasoning Models…
21 -
Hugging Face Daily Papers research 19d ago
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Abstract Mask Forcing mitigates mode collapse in distilled autoregressive video diffusion by injecting masked cleaner signals during self-rollout, improving visual quality without extra training data. Generated by thinkingmachines/Inkling-Small Autoregressive (AR) video…
10 -
Hugging Face Daily Papers research 19d ago
Reason Through the Latent! Making Latent Visual Reasoning Necessary
Abstract CVRR enforces recurrent hidden-state computation as the required image-conditioned pathway for visual reasoning, preserving model competence while distinguishing latent information from actual predictive use. Generated by thinkingmachines/Inkling-Small Latent visual…
35 -
Hugging Face Daily Papers research 19d ago
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
Abstract A two-stage framework predicts sparse gripper keyframes and continuous actions to transfer complex hand demonstrations to robotic grippers while reducing drift via kinematic optimization. Generated by thinkingmachines/Inkling-Small Transferring human hand demonstrations…
23 -
Hugging Face Daily Papers research 19d ago
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Abstract On-Policy Reverse Distillation enables stronger models to exceed weak supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, accelerating optimization without imposing capacity limits. Generated by…
7 -
Hugging Face Daily Papers research 19d ago
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
Abstract Activation steering vectors in large language models encode theory-aligned human value geometry when derived via distribution-driven methods, with geometric fidelity scaling with model size but declining after instruction tuning. Generated by…
7 -
Hugging Face Daily Papers research 19d ago
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
Abstract GE-Act 2.0 is a world-action model trained from scratch with a control-oriented autoencoder, single-step visual planner, and inverse dynamics model, using knowledge-aligned selective optimization to enable scalable zero-shot robot manipulation across diverse skills and…
32 -
Hugging Face Daily Papers research 19d ago
Miles v0.1: Production-Level Post-Training
Abstract Miles is an open-source, production-ready system for large-scale reinforcement learning and post-training that supports diverse backends, weight synchronization, LoRA, distillation, and diffusion models. Generated by thinkingmachines/Inkling-Small We present Miles v0.1,…
16 -
Hugging Face Daily Papers research 19d ago
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Abstract MovieGrid decomposes long videos into spatially arranged chunks for joint modeling, improving multi-shot coherence and scaling video length efficiently. Generated by thinkingmachines/Inkling-Small Generating long-form multi-shot videos requires coherent within-shot…
6 -
Hugging Face Daily Papers research 19d ago
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
Abstract Feedback-Enriched Environments adapt task settings to provide observation-level guidance, improving reinforcement learning stability and exploration for long-horizon agent tasks. Generated by thinkingmachines/Inkling-Small Large Language Models demonstrate remarkable…
5 -
Hugging Face Daily Papers research 19d ago
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
Abstract A system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport. Generated by thinkingmachines/Inkling-Small Speculative decoding accelerates rollout…
36 -
Hugging Face Daily Papers research 19d ago
Kalman Delta Networks: Uncertainty-aware Associative Memory
Abstract Kalman Delta Networks reformulate linear attention as a linear-Gaussian state-space model with Kalman-filter updates to track memory uncertainty, yielding efficient scan-compatible approximations that improve language modeling performance. Generated by…
21