02
Embedded agency, continual skill evolution, and the mechanics of test-time reasoning
Fresh results probe how foundation-model agents cooperate via similarity inference, consolidate experience into transferable skills, and allocate inference compute, alongside formal treatments of causality, ablation, and pretraining inductive biases.
- paperarXivA game theory for foundation models shows new paths to rational cooperation through similarity inference↗
Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus Hutter, James Manyika, Rif A. Saurous, João Sacramento, Blaise Agüera y Arcas
Foundation model agents in social dilemmas converge to stable cooperation under optimal planning, contradicting classical Nash predictions of mutual defection. The authors introduce the embedded Bayesian agent, which maintains uncertainty about its own decision algorithms and treats its deliberation as evidence of similarity to partners, formalizing an embedded equilibrium as a new solution concept.
game theoryembedded agencymulti-agentfoundation modelscooperation - paperarXivReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning↗
Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Failed expert trajectories, previously discarded, provide valuable signals when treated as objects of reflection rather than imitation. ReflectRL elicits reflective reasoning from these golden negatives then transitions the policy to direct reasoning, consistently improving on-policy training across nine benchmarks and multiple backbones with low overhead.
reinforcement learningreasoningon-policyLLMsreflection - paperarXivTest-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility↗
Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary
The paper formalizes test-time scaling as budgeted inference over the autoregressive prefix tree and distinguishes single-trajectory, leaf-level, and prefix-level regimes. It supplies evaluation profiles that separate system performance from candidate diagnostics, protocol-matched compute reporting, and reproducibility requirements, releasing over two billion reasoning traces.
test-time scalingreasoninginferenceevaluationreproducibility - paperarXivLogic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility↗
Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino
Pre-pretraining on formal logical derivations accelerates linguistic skill acquisition, reaching 80 percent accuracy with 36 billion fewer tokens than standard initialization at 100B-token scale. The resulting lower-rank, spectrally concentrated representations also enable pruning to roughly 33 percent sparsity while matching dense baseline performance.
pretraininglogicinductive biascompressibilityLLMs - paperarXivCross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse↗
Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani
A closed-form ridge regression mapper transfers KV caches across same-family models that share head dimensions, stripping RoPE for position-free reuse. It retains 73-98 percent of standalone-prefill accuracy on multiple pairs while running 2.7-25x faster than re-prefill and remaining stable across multi-turn handoffs.
KV cacheefficiencymodel familiesinferencesystems - paperarXivPAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents↗
Shuhan Xue, Zixin Ding, Yichen Shen, Yinjie Wang, Zhenfei Yin, Yingcheng Wu, Yuxin Chen, Mengdi Wang, Ling Yang
PAST-Bench isolates whether retained experience (preferences, histories, skills) produces later-task gains via the intended save-retrieve-update pathway across 26 scenarios. Gains exist but are uneven; Hermes+ interventions raise average improvement and pathway evidence, especially on state-update tasks, though results remain model- and capability-dependent.
agentsself-improvementmemorybenchmarkpersonal agents - paperarXivContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?↗
Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang
The benchmark evaluates in-context continual skill learning across five domains of interconnected subtasks ordered by difficulty. Sequential execution improves performance, yet much of the gain arises from context adaptation rather than reusable skill abstraction; explicit skills help selectively while weaker models accumulate fragmented task-specific collections.
agentscontinual learningskillsin-context learningbenchmark - paperarXivComputing Actual Causes for Neural Network Predictions under Structured Causal Inputs↗
Jannick Strobel, Muqsit Azeem, Stefan Leue
Explanations are formalized as Halpern-Pearl actual causes under Boolean SCMs that capture input dependencies, computed via bound propagation and branch-and-bound with completeness and minimality guarantees. The method scales to search spaces of 10^13 candidates and shows that ignoring dependencies produces 14.9 percent spurious causes.
causalityexplainabilityneural networksactual causationSCM - paperarXivSensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling↗
Nathan Labiosa, David Buff, Ena Nayak, Erica Donno
Layer maps for representational sensitivity, causal recovery via activation restoration, and adapter repair capacity dissociate under surface perturbations, with anti-correlation between sensitivity and causality. Two propagation regimes appear, late-accumulation strengthens with scale, and cascade disruption explains why causally flagged early layers are often the worst adapter sites.
interpretabilityrobustnesslayersperturbationscaling - paperarXivA Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation↗
Abdallah Khemais
Exact results characterize when weight-space ablation of residual carriers collapses conditional computation, contrasting the effects of patching (contrast) versus ablation (absolute level). An attention-head interaction formula is derived with second-order remainder, and synthetic transformers confirm the predictions with strong rank correlation to measured interactions.
mechanistic interpretabilityablationtheoryresidual streamcausality