Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 14h ago
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Abstract Latent Dynamics Reasoning integrates kinematic dynamics in structured latent space to enable video world models that extrapolate physical laws far beyond training distributions with far fewer parameters and faster inference. Generated by thinkingmachines/Inkling-Small…
35 -
Hugging Face Daily Papers research 20h ago
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Abstract StateFlow introduces a persistent 3D world state to enable iterative, controllable previsualization for film and game design by constructing, evolving, and accessing structured scene and camera representations. Generated by thinkingmachines/Inkling-Small…
31 -
Hugging Face Daily Papers research 20h ago
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
Abstract The benchmark evaluates autonomous coding agents on open-ended world-model research by having them iteratively improve a starter model across game environments using a shared structured-state format. Generated by thinkingmachines/Inkling-Small World modeling is an…
20 -
Hugging Face Daily Papers research 20h ago
AVA-Encoder: Towards Agent-Native Video Representation Learning
Abstract AVA-Encoder learns structured video representations via agentic auto-encoding using knowledge graphs to enable cinematic video generation and reasoning with reduced token usage. Generated by thinkingmachines/Inkling-Small Creative agents still lack an effective way to…
30 -
Hugging Face Daily Papers research 22h ago
Parameter Exploration for RLVR via Variational Learning
Abstract Parameter-space exploration via perturbed policy sampling improves LLM reinforcement learning by diversifying rollouts and reducing training failures compared to action-space methods. Generated by thinkingmachines/Inkling-Small Exploration has been a focus of…
20 -
Hugging Face Daily Papers research 23h ago
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Abstract Mechanist is an autonomous agentic system that uses AI to discover and control the mechanisms underlying model intelligence, generating hypotheses, performing causal interventions, and improving safety and performance. Generated by thinkingmachines/Inkling-Small AI…
34 -
Hugging Face Daily Papers research 23h ago
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
Abstract Two studies define measurable GPU control gates for LLM-agent services by analyzing concurrent cohort scheduling and on-device routing versus host redispatch. Generated by thinkingmachines/Inkling-Small LLM-agent services repeatedly execute small deterministic…
25 -
Hugging Face Daily Papers research 1d ago
Self-Evolving Embodied Agents via Skill-Harness Evolution
Abstract SHAPER is a train-free framework that improves embodied agents by evolving reusable skills and a context-code harness around a frozen foundation model through environment rollouts. Generated by thinkingmachines/Inkling-Small Embodied agents are increasingly built as…
26 -
Hugging Face Daily Papers research 1d ago
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
Abstract Replacing individual LLM agents with low-parameter surrogates fitted from cheap queries enables scalable society simulations, with validity predicted by an interaction-order and memory taxonomy. Generated by thinkingmachines/Inkling-Small Simulating societies of many…
5 -
Hugging Face Daily Papers research 1d ago
Simplex Relaxation for Discrete Diffusion
Abstract Simplax enriches uniform discrete diffusion via Dirichlet-categorical augmentation to improve reverse sampling and generative quality on text and Sudoku tasks. Generated by thinkingmachines/Inkling-Small Discrete diffusion models for categorical generation are defined…
30 -
Hugging Face Daily Papers research 1d ago
Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands
Abstract This work introduces a dedicated model for per-joint hand visibility estimation and demonstrates its benefit for multi-view 3D hand pose annotation. Generated by thinkingmachines/Inkling-Small Hand Pose Estimation (HPE) is a fundamental technology for various…
4 -
Hugging Face Daily Papers research 1d ago
Persistent Recursive Worlds Enable Autonomous Software Evolution
Abstract Genesis organizes long-horizon software development around a persistent project rather than persistent agents, enabling multi-day compiler construction and numerical module reimplementation with low cost and high performance. Generated by thinkingmachines/Inkling-Small…
22 -
Hugging Face Daily Papers research 1d ago
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Abstract OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows. Generated by thinkingmachines/Inkling-Small AI agents…
28 -
Hugging Face Daily Papers research 1d ago
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
Abstract AtlasVLA improves embodied AI by replacing reactive control with proactive reasoning via persistent world-ego memory, enabling robust long-horizon manipulation from a single wrist camera. Generated by thinkingmachines/Inkling-Small While Vision-Language-Action (VLA)…
36 -
Hugging Face Daily Papers research 1d ago
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
Abstract Self-Geometry improves vision foundation model predictions by enforcing explicit multi-view geometric constraints via test-time adaptation with LoRA, disentangled losses, and angular neighbor sampling. Generated by thinkingmachines/Inkling-Small Recent Vision Foundation…
21 -
Hugging Face Daily Papers research 1d ago
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
Abstract Visual tool-use in multimodal LLMs often lacks causal effectiveness, with returned observations frequently failing to influence answers or being used incoherently despite aggregate accuracy improvements. Generated by thinkingmachines/Inkling-Small The…
6 -
Hugging Face Daily Papers research 1d ago
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
Abstract NeuPAT selectively constrains updates to language-sensitive neurons during multimodal tuning to preserve LLM language capabilities while enabling perceptual adaptation. Generated by thinkingmachines/Inkling-Small Multimodal expansion of large language models (LLMs)…
18 -
Hugging Face Daily Papers research 1d ago
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Abstract Stronger models can build inference-time harnesses that substantially improve weaker models' task performance without parameter updates by offloading reasoning into structured code and routing. Generated by thinkingmachines/Inkling-Small Recent work on distillation…
36 -
Hugging Face Daily Papers research 1d ago
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Abstract A closed-loop framework combining physics-based video synthesis, diffusion-based video dereflection, and a new benchmark achieves state-of-the-art video reflection removal with fast inference. Generated by thinkingmachines/Inkling-Small Videos captured through glass…
29 -
Hugging Face Daily Papers research 1d ago
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Abstract Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only…
21 -
Hugging Face Daily Papers research 1d ago
CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
Abstract CoinRAG improves retrieval-augmented generation efficiency and accuracy by reusing fine-grained semantic nugget caches instead of full chunks. Generated by thinkingmachines/Inkling-Small Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited…
4 -
Hugging Face Daily Papers research 1d ago
Agent Safety Should Be a Runtime Contract
Abstract Agent safety should be enforced at runtime through preventive controls and verifiable evidence rather than relying solely on training-time alignment methods. Generated by thinkingmachines/Inkling-Small The dominant paradigm treats AI safety as a property to be instilled…
24 -
Hugging Face Daily Papers research 1d ago
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Abstract Spark-to-Paper is a lightweight, composable workflow inside coding assistants that generates research papers by separating planning from reporting, enforcing evidence-based claim revision, and using integrity checks to reduce fabrication. Generated by…
9 -
Hugging Face Daily Papers research 1d ago
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Abstract ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents…
12 -
Hugging Face Daily Papers research 1d ago
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Abstract DSAgentBench evaluates autonomous agents on complete, multi-tool data-science workflows in real computing environments and reveals major performance gaps. Generated by thinkingmachines/Inkling-Small Real-world data science involves long-horizon workflows that span data…
33 -
Hugging Face Daily Papers research 1d ago
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
Abstract SkillZip compresses self-evolving agent skills by finding a minimal faithful structural explanation that shares repeated rules and procedures while preserving rare exceptions, without requiring evaluation rollouts. Generated by thinkingmachines/Inkling-Small…
26 -
Hugging Face Daily Papers research 1d ago
InSight-doc: Agentic Visual Perception for Long-Document Understanding
Abstract InSight-doc adaptively allocates visual resolution during reasoning to improve long-document understanding while reducing latency and hallucinations. Generated by thinkingmachines/Inkling-Small Long-document understanding often requires reasoning over many visually rich…
27 -
Hugging Face Daily Papers research 1d ago
Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
Abstract A new attention mechanism replaces fixed scaled dot-product attention with a learned power-law bilinear operator, with verified architecture, measured stability, and machine-checked proofs. Generated by thinkingmachines/Inkling-Small The Large Language Model from Power…
10 -
Hugging Face Daily Papers research 2d ago
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
Abstract A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning. Generated by thinkingmachines/Inkling-Small We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of…
8 -
Hugging Face Daily Papers research 2d ago
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
Abstract Decoding-Level Taboo is a runtime logit-space stress test that reveals how large language models handle off-nominal generation paths, showing that robustness depends on scale and instruction alignment. Generated by thinkingmachines/Inkling-Small Large language model…
32 -
Hugging Face Daily Papers research 2d ago
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
Abstract UniMoMo compresses trained recommendation mixture-of-experts models into smaller standard MoE checkpoints via functional similarity grouping and layer-adaptive protection, preserving accuracy while accelerating inference. Generated by thinkingmachines/Inkling-Small…
27 -
Hugging Face Daily Papers research 2d ago
Beyond Pixels: From Video Priors to 4D Worlds
Abstract Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining. Generated by thinkingmachines/Inkling-Small 4D generation synthesizes…
21 -
Hugging Face Daily Papers research 2d ago
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
Abstract The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models. Generated by thinkingmachines/Inkling-Small The rapid advancement of Large Language…
16 -
Hugging Face Daily Papers research 2d ago
Articulated Object Reconstruction from Rest-State Observation
Abstract A rest-state framework reconstructs articulated objects from a single closed configuration by fusing vision-language outputs into consistent part meshes and validating synthesized motion hypotheses via geometric consistency. Generated by thinkingmachines/Inkling-Small…
29 -
Hugging Face Daily Papers research 2d ago
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Abstract DistilVDR is a compact 524M vision-document retriever distilled from an 8B teacher using cosine alignment without relevance labels, achieving near-teacher accuracy with far smaller indexes and faster indexing. Generated by thinkingmachines/Inkling-Small Visual document…
33 -
Hugging Face Daily Papers research 2d ago
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Abstract Adversarial Fréchet Distance improves generator post-training by adding a learnable adversarial feature space to static Fréchet losses, with whitening to stabilize optimization. Generated by thinkingmachines/Inkling-Small Fréchet distance has recently emerged as an…
10 -
Hugging Face Daily Papers research 2d ago
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
Abstract Pruning strategies applied at different pipeline stages reduce token usage and latency in long-horizon research agents, with early pruning yielding the greatest efficiency gains. Generated by thinkingmachines/Inkling-Small Long-horizon research agents solve open-ended…
10 -
Hugging Face Daily Papers research 2d ago
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
Abstract Open multilingual translation models are improved via group relative policy optimization with reference-free quality rewards and checkpoint interpolation, surpassing strong open and proprietary baselines. Generated by thinkingmachines/Inkling-Small We study…
11 -
Hugging Face Daily Papers research 2d ago
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Abstract Combodied Agents integrate digital and embodied tools into a closed-loop framework that models individual human-state trajectories over time to provide proportionate, consent-aware support. Generated by thinkingmachines/Inkling-Small After an older adult misses a…
30 -
Hugging Face Daily Papers research 2d ago
iFAN: Inference-Aware Learning for Plain Mask Transformers
Abstract A training framework called iFAN improves mask transformers by aligning query ranking with mask quality and distilling stronger intermediate predictions to the final layer. Generated by thinkingmachines/Inkling-Small Query-based mask transformers assemble segmentation…
25 -
Hugging Face Daily Papers research 2d ago
TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity
Abstract A unified toolbox enables reproducible comparison and extension of time-series dataset similarity methods for forecasting, classification, and generation tasks. Generated by thinkingmachines/Inkling-Small The rapid advancement of artificial intelligence (AI) has…
10 -
Hugging Face Daily Papers research 2d ago
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Abstract A new jigsaw benchmark with interlocking pieces reveals that vision-language models fail at geometric reasoning and suffer a sharp performance drop as puzzle size increases. Generated by thinkingmachines/Inkling-Small Jigsaw puzzle solving requires jointly reasoning…
7 -
Hugging Face Daily Papers research 2d ago
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Abstract Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator. Generated by thinkingmachines/Inkling-Small Omni-modal dialogue models can understand…
20 -
Hugging Face Daily Papers research 2d ago
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Abstract Agentic systems can achieve open-ended improvement through multi-component co-evolution that progressively removes fixed human constraints across agents, environments, and evolution mechanisms. Generated by thinkingmachines/Inkling-Small Agentic systems are increasingly…
11 -
Hugging Face Daily Papers research 2d ago
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Abstract SPIEval benchmarks mobile assistant LLMs on scattered personal data tasks, revealing major gaps in information retrieval and verification. Generated by thinkingmachines/Inkling-Small Large language models (LLMs) are increasingly deployed as mobile assistants, where a…
36 -
Hugging Face Daily Papers research 2d ago
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Abstract A new benchmark called VibeLifeBench evaluates long-horizon proactive agents across simulated multi-week everyday tasks, revealing that current frontier models perform poorly. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents are increasingly…
22 -
Hugging Face Daily Papers research 2d ago
The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents
Abstract Gated Hindsight Distillation improves GUI agent training by using future screenshots as privileged evidence to recover correct reasoning when standard imitation fails. Generated by thinkingmachines/Inkling-Small GUI agents are commonly trained offline from successful…
21 -
Hugging Face Daily Papers research 2d ago
Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers
Abstract SiPE integrates a lightweight syntactic prior from dependency parses into positional embeddings across transformer architectures, improving syntactic generalization and language understanding without altering self-attention or increasing inference cost. Generated by…
15 -
Hugging Face Daily Papers research 2d ago
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
Abstract Atelier improves artist-grounded image generation by translating vague artistic intent into explicit control states that separate scene content from style, reducing reliance on stereotypical shortcuts. Generated by thinkingmachines/Inkling-Small Artist-grounded image…
20 -
Hugging Face Daily Papers research 2d ago
On-Policy Self-Distillation without Any Supervision
Abstract Unsupervised on-policy self-distillation improves large language models by using internal consistency and majority-vote pseudo-solutions to correct confident errors without external supervision. Generated by thinkingmachines/Inkling-Small On-policy (Self-)Distillation…
34