huggingface-daily-papers-2026-05-21
AI · · 1 min read
HuggingFace Daily Papers
Fetched: 2026-05-21T11:09:39Z
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
- Upvotes: 0
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
- Upvotes: 0
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
- Upvotes: 0
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
- Upvotes: 0
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
- Upvotes: 0
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
- Upvotes: 0
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
- Upvotes: 0
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
- Upvotes: 0
OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
- Upvotes: 0
Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency
- Upvotes: 0
Stitched Value Model for Diffusion Alignment
- Upvotes: 0
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
- Upvotes: 0
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
- Upvotes: 0
The Unlearnability Phenomenon in RLVR for Language Models
- Upvotes: 0
OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation
- Upvotes: 0