arxiv-2026-06-08
AI · · 6 min read
arXiv Papers - 2026-06-08
Total papers: 15
How reliable are LLMs when it comes to playing dice?
- Authors: Luca Avena, Gianmarco Bet, Bernardo Busoni
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07515v1
- Summary: We investigate the probabilistic reasoning capabilities of large language models through a controlled benchmarking study on discrete probability problems. We constructed two datasets, respectively a set of standard exercises and a set of counterintuitive exercises, designed to trigger heuristic reasoning, and evaluated 8 state-of-the-art models, each tested with and without Chain-of-Thought prompting. Models achieve an average accuracy of 0.96 on standard problems but only 0.59 on counterintuiti
Agentopia: Long-Term Life Simulation and Learning in Agent Societies
- Authors: Xintao Wang, Sirui Zheng, Hongqiu Wu et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07513v1
- Summary: Humans learn from social life. Simulating this process with LLM-powered agents represents a promising research direction, raising a natural question: whether LLMs can learn from such simulated social experience to better understand and replicate human behavior. However, prior agent society simulations typically operate at the scale of days, limiting the depth of social interactions and long-term growth. In this paper, we study long-term life simulation and LLM learning in agent societies, with t
MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism
- Authors: Cong Chen, Guo Gan, Kaixiang Ji et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07512v1
- Summary: Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduce MemDreamer to decouple perception and reasoning, shifting long-video understanding into an agentic exploration process. As a plug-and-play framework, it incrementally streams videos to construct a Hierarchical Graph Memory, a top-down three-tier architecture for semantic abstraction, anchored by a
Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
- Authors: Songhao Wu, Zhongxin Chen, Yuxuan Liu et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07502v1
- Summary: Large language models exhibit impressive zero-shot capabilities across a wide range of downstream tasks. However, they struggle to function as off-the-shelf embedding models, leading to suboptimal performance on massive text embedding benchmarks. In this paper, we identify a potential cause underlying this deficiency. Our motivation stems from an unexpected observation: text embeddings tend to align with frequent but uninformative tokens when projected onto the vocabulary space. We argue that th
Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning
- Authors: Fatema Siddika, Md Anwar Hossen, Tanwi Mallick et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07500v1
- Summary: Continual learning in Large Language Models (LLMs) is hindered by the plasticity-stability dilemma, where acquiring new capabilities often leads to catastrophic forgetting of previous knowledge. Existing methods typically treat parameters uniformly, failing to distinguish between specific task knowledge and shared capabilities. We introduce Mixture of Sparse Experts for Task Agnostic Continual Learning (SETA), a framework that resolves the plasticity-stability conflict through adaptive sparse su
Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization
- Authors: Ming Sun, Kun Yuan
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07496v1
- Summary: Decentralized stochastic optimization is a fundamental paradigm for large-scale learning over networks, where agents communicate only with their neighbors and no central coordinator is required. For strongly convex problems, communication efficiency is mainly determined by the condition number (κ=L/μ) and the network spectral gap (1-β). Although deterministic decentralized methods can simultaneously achieve accelerated (\sqrtκ) and (1/\sqrt{1-β}) dependences, no existing stochastic metho
Second-Order Path Kernel Interpolation Formulas in Machine Learning
- Authors: Jin Guo, Roy Y. He, Jean-Michel Morel
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07495v1
- Summary: Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gradient descent. It expresses the model's prediction as an integral, along the optimization path, of a data-dependent kernel that aligns the model's gradients at the test and training data. Such a first-order characterization remains valid for models trained with batch-based stoch
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
- Authors: Ekaterina Grishina, Stepan Kuznetsov, Askar Tsyganov et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07492v1
- Summary: The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale. This drives a demand for a proper methodology for fair comparison between algorithms. Naive aggregation of performance metrics (e.g., averaging NDCG over benchmarks) can yield misleading rankings, undermining practical selection. To address this problem, we introduce a novel, data-driven ranking methodology based on B
Twelve quick tips for designing AI-driven HPC workflows
- Authors: Jamie J. Alnasir
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07491v1
- Summary: High-performance computing (HPC) clusters remain the backbone of large-scale scientific computation, traditionally executing deterministic, linear pipelines optimised for predictable performance. However, the pervasive integration of artificial intelligence (AI) and foundation models into scientific research has introduced a fundamentally new computational paradigm. AI-driven workflows are characteristically iterative, data-driven, and probabilistic, introducing unique challenges regarding data
How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope
- Authors: Jeremy Yang, Kate Zyskowski, Noah Yonack et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07489v1
- Summary: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end. Using production data from Perplexity's Search and Computer products, we study this transition by examining how AI agents accelerate and reshape knowledge work. Three key empirical findings emerge. First, using sessions with near-identical initial query pairs as natural experiments for the same underlying task attempted with both
CoMetaPNS: Continually Meta-learning Personalized Neural Surrogates for Cardiac Electrophysiology Simulations
- Authors: Ryan Missel, Xiajun Jiang, Linwei Wang
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07488v1
- Summary: Personalized virtual heart simulations face challenges in model personalization and computational cost. While neural surrogates offer state-of-the-art solutions, they typically address either efficient personalization or training generalizable models. Recent work reframes this by learning the process of personalizing a surrogate using limited subject-specific context data, through few-shot generative modeling with set-conditioned surrogates and meta-learned amortized inference. These methods, ho
Network Recovery from Cascade Data: A Debiased Jacobian-Based Machine Learning Approach
- Authors: Lei Huang
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07483v1
- Summary: Many important outcomes unfold as dynamic cascades, including product adoption, disease spread, financial distress, and information diffusion. A central challenge is to recover the hidden influence network behind these cascades. Existing methods typically assume a specific diffusion model, and their performance degrades substantially when that assumption is misspecified. We propose CascadeNet, a Jacobian-based machine learning framework for network recovery that does not require specifying a dif
Drifting Models for Surrogate Flow Modeling
- Authors: Chris R. Jung, Markus Dörr, Natalie Jüngling et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07481v1
- Summary: While Computational Fluid Dynamics (CFD) provides high-fidelity flow fields for optimizing indoor environments, its computational cost limits rapid exploration. To solve this problem generative surrogates offer better distribution modeling than deterministic networks, but iterative sampling is slow. To enable high-quality, single-pass generation, we adapt the novel generative drifting framework to fluid mechanics. We introduce a conditional architecture that performs drifting in a learned VAE la
Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification
- Authors: Sercan Karakaş, Yusuf Şimşek
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07479v1
- Summary: Turkish idiomatic light verb constructions (LVCs) are challenging for multiword expression processing because they often share the same surface form as fully literal verb-object combinations while functioning as a single, partially idiomatic predicate. We frame Turkish LVC detection as a binary classification task (literal meaning vs. idiomatic meaning) and evaluate on a manually created controlled set (N=147) with matched negatives: out-of-domain random sentences and in-domain literal controls
Unsupervised Continual Clustering via Forward-Backward Knowledge Distillation
- Authors: Mohammadreza Sadeghi, Sareh Soleimani, Zihan Wang et al.
- Published: 2026-06-05
- arXiv ID: http://arxiv.org/abs/2606.07474v1
- Summary: Unsupervised Continual Learning (UCL) aims to enable neural networks to learn sequential tasks without labels or access to past data. A major challenge in this setting is Catastrophic Forgetting, where models forget previously learned tasks upon learning new ones. This challenge is amplified in UCL due to the absence of labels to guide learning and memory retention. Existing mitigation strategies, such as knowledge distillation and replay buffers, often raise memory and privacy concerns. Moreove