Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 14d ago
Online Learning with LLM Experts from Limited Feedback
Abstract Adaptive routing of prompts to LLM experts is formulated as a contextual bandit problem with limited feedback, yielding algorithms with sublinear regret bounds and effective routing strategies. Generated by thinkingmachines/Inkling-Small We study adaptive routing of…
35 -
Hugging Face Daily Papers research 14d ago
StepAudio 3 Gen Technical Report
Abstract StepAudio 3 Gen is a discrete autoregressive audio generation model using residual vector quantization tokens to unify text-to-speech, voice design, sound effects, and music within a single framework. Generated by thinkingmachines/Inkling-Small We introduce StepAudio 3…
20 -
Hugging Face Daily Papers research 14d ago
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Abstract A data-centric framework with specialized systems for reasoning analysis, cost reduction, and execution verification enables small teams to train open-weight cyber agents that achieve top-tier performance on benchmark suites. Generated by thinkingmachines/Inkling-Small…
17 -
Hugging Face Daily Papers research 14d ago
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Abstract Benchmark Radar is a searchable living database and discovery engine for AI evaluation benchmarks that aggregates sources, score histories, and evidence to support benchmark selection and comparison. Generated by thinkingmachines/Inkling-Small Benchmark researchers and…
7 -
Hugging Face Daily Papers research 14d ago
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
Abstract LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information. Generated by thinkingmachines/Inkling-Small…
20 -
Hugging Face Daily Papers research 14d ago
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Abstract COBRA-Skills improves LLM agent skill optimization by using contextual-bandit prioritization and evidence-based evolution to cut evaluation costs while maintaining high performance. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents can…
21 -
Hugging Face Daily Papers research 14d ago
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
Abstract Posterior Label Correction DPO improves preference optimization by routing noisy pairwise labels into clean, flipped, or tied cases using calibrated policy-reference margins. Generated by thinkingmachines/Inkling-Small Direct Preference Optimization (DPO) simplifies…
24 -
Hugging Face Daily Papers research 14d ago
SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image
Abstract A physics-guided framework improves part-aware 3D generation by resolving inter-part penetration, recovering contact graphs, and refining parameterized connectors via simulation feedback to ensure stable, physically valid assemblies. Generated by…
25 -
Hugging Face Daily Papers research 16d ago
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Abstract Using a controlled autoregressive testbed, the study analyzes task-specific validation losses during multimodal pretraining to evaluate how image tokenizer design affects joint text-image modeling and downstream performance. Generated by thinkingmachines/Inkling-Small…
11 -
Hugging Face Daily Papers research 16d ago
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking
Abstract A reasoning-capable vision-language model that iteratively retrieves and reasons over Wikipedia improves multimodal entity linking for rare entities defined by knowledge-graph structure. Generated by thinkingmachines/Inkling-Small Multimodal entity linking grounds…
29 -
Hugging Face Daily Papers research 16d ago
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
Abstract Optimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language. Generated by thinkingmachines/Inkling-Small Reasoning…
36 -
Hugging Face Daily Papers research 16d ago
Adaptive Bridge: A Proxy-Based Decoupling Layer for Mitigating DDS Backpressure in ROS 2
Abstract A proxy layer isolates critical ROS 2 subscribers from degraded ones via topic splitting and dynamic rate control to eliminate DDS backpressure. Generated by thinkingmachines/Inkling-Small In systems built on Robot Operating System 2 (ROS 2) and using Data Distribution…
26 -
Hugging Face Daily Papers research 16d ago
Beyond Solver Verdicts: Generative Reward Models for Autoformalization
Abstract Neurosymbolic reasoning is vulnerable to incorrect but verdict-matching formal translations, which are addressed by a generative verification method that scores reference equivalence without an oracle and improves downstream accuracy. Generated by…
6 -
Hugging Face Daily Papers research 16d ago
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
Abstract ActReview is a rebuttal-guided post-training framework that generates diagnostic claims and concrete revision suggestions for peer review by leveraging author responses as latent supervision. Generated by thinkingmachines/Inkling-Small As LLMs are increasingly used for…
18 -
Hugging Face Daily Papers research 16d ago
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
Abstract The study introduces a benchmark to evaluate whether research methods are specified clearly enough for implementation, finding that identifying missing details is the primary challenge for language models. Generated by thinkingmachines/Inkling-Small A research idea may…
29 -
Hugging Face Daily Papers research 16d ago
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
Abstract A natural-language proof-generation pipeline using post-trained Nemotron 3 Ultra checkpoints achieves gold-medal performance on IMO 2026 through iterative verification and refinement without external tools. Generated by thinkingmachines/Inkling-Small We study how model…
9 -
Hugging Face Daily Papers research 16d ago
Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs
Abstract RCWM recursively builds complex 3D worlds as executable code from a single image by alternating global and local reconstruction with shared camera alignment. Generated by thinkingmachines/Inkling-Small Code world models represent worlds as executable programs, but this…
7 -
Hugging Face Daily Papers research 16d ago
World in World: Explore the World with World Models
Abstract A training-free interface enables flexible camera and time control in frozen autoregressive video world models by routing heterogeneous visual evidence through native self-attention with correspondence-guided queries and per-channel attention guidance. Generated by…
11 -
Hugging Face Daily Papers research 16d ago
Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Abstract MaP-WAM improves non-Markovian robotic manipulation by separating memory-grounded planning from plan-conditioned execution, using compact episodic segment records and progress-calibrated action chunks to maintain fixed inference latency. Generated by…
35 -
Hugging Face Daily Papers research 17d ago
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
Abstract Negative Self-Distillation improves large language model reasoning by pushing models away from self-generated flawed reasoning via a dynamic gating mechanism that protects linguistic capabilities. Generated by thinkingmachines/Inkling-Small On-Policy Self-Distillation…
16 -
Hugging Face Daily Papers research 17d ago
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
Abstract FreeFlow is a hierarchical transformer for optical flow that eliminates task-specific inductive biases and achieves state-of-the-art accuracy using window, shifted-window, and global attention. Generated by thinkingmachines/Inkling-Small Optical flow methods typically…
32 -
Hugging Face Daily Papers research 17d ago
Mi-Ripple: Restoring Images Degraded by Iterative AI Editing
Abstract Mi-Ripple reduces digital ripple artifacts in edited images by separating lattice artifacts from texture and applying targeted spectral filtering and reference cleaning. Generated by thinkingmachines/Inkling-Small Iterative reference-conditioned image editing can…
36 -
Hugging Face Daily Papers research 17d ago
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
Abstract A hierarchical multi-agent reinforcement learning framework combining graph attention and dynamic role assignment improves tactical coordination and win rates in air combat. Generated by thinkingmachines/Inkling-Small Multi-Agent Reinforcement Learning (MARL) has…
25 -
Hugging Face Daily Papers research 17d ago
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
Abstract A benchmark for transit-kiosk policy reasoning shows that small parameter-efficiently fine-tuned language models can match larger models on structured tool-use and fare-quoting tasks. Generated by thinkingmachines/Inkling-Small We introduce MetroLLM-Bench, a 955-case…
9 -
Hugging Face Daily Papers research 17d ago
UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
Abstract UniH3 unifies hierarchical homogeneity and heterogeneity for universal medical image restoration via memory-based homogeneity priors and a heterogeneity balancer. Generated by thinkingmachines/Inkling-Small All-in-One medical image restoration (MedIR) aims to address…
13 -
Hugging Face Daily Papers research 17d ago
TempCloze: Can Video-LLMs Identify the Missing Middle?
Abstract TempCloze evaluates visual temporal reasoning in Video-LLMs by requiring identification of missing video segments from distractors targeting semantics, alignment, and progression. Generated by thinkingmachines/Inkling-Small Temporal reasoning benchmarks for Video-LLMs…
29 -
Hugging Face Daily Papers research 17d ago
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
Abstract X-AuT progressively prunes audio-encoder layers in speech large language models and restores accuracy via behavioral probes, representation alignment, cross-scale distillation, and LoRA adaptation. Generated by thinkingmachines/Inkling-Small Reducing audio-encoder depth…
28 -
Hugging Face Daily Papers research 17d ago
Generative Late-Interaction Embeddings For Visual Document Retrieval
Abstract Generative Late-Interaction Embeddings compress visual document retrieval vectors by learning a small basis set that regenerates full embeddings on demand, improving accuracy under tight storage limits without retraining the encoder. Generated by…
28 -
Hugging Face Daily Papers research 17d ago
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Abstract EvoSafeHarness optimizes deployable safety harnesses by jointly searching natural-language policies and executable logic tailored to a frozen model and target domain, improving safety-utility trade-offs across agent benchmarks. Generated by…
35 -
Hugging Face Daily Papers research 17d ago
HyQuant: Hybrid-Precision Quantization for LLM Attention
Abstract HyQuant improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead. Generated by thinkingmachines/Inkling-Small Quantization has been…
38 -
Hugging Face Daily Papers research 17d ago
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
Abstract NCP-ArchPreview is a large latent-space language model that jointly trains next-token and next-concept prediction to improve pretraining efficiency and downstream performance. Generated by thinkingmachines/Inkling-Small We introduce NCP-ArchPreview, a latent-space…
33 -
Hugging Face Daily Papers research 17d ago
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
Abstract Large vision-language models trained on synthetic block-manipulation tasks improve 3D spatial reasoning and generalize to real-world visual tasks. Generated by thinkingmachines/Inkling-Small Large Vision-Language Models (LVLMs) have achieved strong performance on…
25 -
Hugging Face Daily Papers research 17d ago
SenseNova-U1.5: Towards Native Unified Visual Intelligence
Abstract SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, expert optimization, and…
30 -
Hugging Face Daily Papers research 17d ago
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation
Abstract A unified vision-language model for coronary angiography uses chain-of-box reasoning and reinforcement learning with verifiable rewards to provide auditable diagnoses and improve zero-shot report generation. Generated by thinkingmachines/Inkling-Small Invasive coronary…
36 -
Hugging Face Daily Papers research 17d ago
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
Abstract The study proposes a reference-based hidden-state auditing method that measures representational bias shift across model variants to detect internal bias changes with minimal compute. Generated by thinkingmachines/Inkling-Small Existing bias auditing methods typically…
32 -
Hugging Face Daily Papers research 17d ago
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
Abstract Researchers introduce a 400-scenario benchmark and evidence-based monitor to study how instrumental goals, oversight, and strategic hints drive covert misaligned behavior in LLM agents. Generated by thinkingmachines/Inkling-Small We study scheming in LLM agents, in…
33 -
Hugging Face Daily Papers research 17d ago
The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements
Abstract For sparse binary signals, sufficient sample sizes for maximum-likelihood support recovery are identified in high-SNR regimes, revealing an information-theoretic threshold and trade-offs between measurement sparsity and computational cost, with analysis also covering…
6 -
Hugging Face Daily Papers research 17d ago
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
Abstract Influence-guided response rewriting of selected training examples produces stronger and more persistent behavioral shifts in language models than conventional reweighting, highlighting the broader intervention leverage of influential data. Generated by…
24 -
Hugging Face Daily Papers research 17d ago
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Abstract StochBench introduces 450 graduate-level stochastic processes problems in Lean 4 to benchmark formal theorem proving on domain-specific applied mathematics. Generated by thinkingmachines/Inkling-Small Leading benchmarks for formal theorem proving with large language…
37 -
Hugging Face Daily Papers research 17d ago
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
Abstract A non-invasive brain decoding approach maps MEG responses to semantic embeddings to reconstruct sentence-level text without requiring word-level alignment. Generated by thinkingmachines/Inkling-Small Non-invasive speech decoding remains constrained by the low…
34 -
Hugging Face Daily Papers research 18d ago
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Abstract AgentGrad improves multi-agent prompt optimization by identifying target agents through sequential intervention and clustering gradients semantically to avoid mixing unrelated errors. Generated by thinkingmachines/Inkling-Small Large language model (LLM)-based…
28 -
Hugging Face Daily Papers research 18d ago
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
Abstract OracleZoom improves recursive super-resolution by combining trajectory-based training with cross-scale supervision and a latent prior to reduce hallucinations at extreme magnifications. Generated by thinkingmachines/Inkling-Small Recursive Super-Resolution (SR) extends…
33 -
Hugging Face Daily Papers research 18d ago
PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
Abstract PlannerForge is an LLM-agent framework that unifies all stages of scenario-based autonomous driving testing and improves generation, selection, modification, and planning performance across commercial and open-source models. Generated by thinkingmachines/Inkling-Small…
7 -
Hugging Face Daily Papers research 18d ago
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
Abstract This survey examines inference-efficiency techniques for video large language models, analyzing cost reductions across frame sampling, encoding, token compression, and language model stages while identifying evaluation gaps. Generated by thinkingmachines/Inkling-Small…
25 -
Hugging Face Daily Papers research 18d ago
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Abstract SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of…
20 -
Hugging Face Daily Papers research 18d ago
Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning
Abstract HybridAL adaptively switches from retraining to fine-tuning during active learning based on online stabilization signals, reducing training time while preserving accuracy and improving calibration. Generated by thinkingmachines/Inkling-Small Training strategy, namely…
6 -
Hugging Face Daily Papers research 18d ago
Programmable World Model
Abstract A programmable world model separates explicit state evolution from video generation using executable rules and 3D bounding boxes to maintain persistent, controllable environments. Generated by thinkingmachines/Inkling-Small Recent video world models generate…
7 -
Hugging Face Daily Papers research 18d ago
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Abstract AgenticGen improves advertising video generation by decomposing it into strategy selection and draft generation stages supervised by online business feedback and human quality rewards. Generated by thinkingmachines/Inkling-Small Advertising video generation is not only…
34 -
Hugging Face Daily Papers research 18d ago
SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
Abstract SyncWorld is an action-conditioned world model that uses visual calibration episodes to learn environment-specific action-to-visual mappings, enabling zero-shot simulation and test-time policy improvement across unseen robotic settings. Generated by…
6 -
Hugging Face Daily Papers research 18d ago
DF26: We Cannot Tell Fake From Real Anymore
Abstract A new benchmark for AI-generated public-speaking videos reveals that both humans and current detectors perform near chance, underscoring the need for robustness to modern generative distribution shifts. Generated by thinkingmachines/Inkling-Small We introduce DF26, a…
13