~/reading

16

Issue 16/August 20, 2026/10 entries

Credit assignment and strategy lock-in in long-horizon agents

Recent work converges on why multi-turn and agentic training stalls: outcome rewards starve the decisions that matter, post-training agents lock into an initial strategy, and fixed environments stop expanding the goal distribution. The selected papers introduce reverse-turn and dual-channel credit, group-calibrated distillation, learnable environment designers, and monitors for latent collusion, alongside theoretical treatments of preference indeterminacy and test-time exploitation failure.

  1. paperarXiv
    SPADE: Self-Play in Adaptive Synthetic Executable Environments

    Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques

    SPADE lets a single LLM act as both Environment Designer, which writes complete long-horizon Gym-style executable environments as code, and Reasoning Agent that learns inside them. Regret is measured as the reward gap with versus without privileged hints, so the designer targets the edge of the agent's capability while keeping tasks feasible. Scaling to 30B models yields +5.3 average gains over the strongest fixed-environment baseline across eight held-out benchmarks and larger lifts on multi-turn tool-use suites.

    self-playagent environmentsRLopen-ended learning
  2. paperarXiv
    Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

    Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou

    Token-level on-policy distillation increasingly misaligns with response-level verifiers as context length grows, because teacher likelihood favors locally plausible but globally incomplete answers. GC-OPD normalizes verifier rewards and trajectory OPD scores inside each rollout group, forms a signed disagreement residual, and redistributes it via relative-advantage credit assignment while preserving the original dense signal. On five long-context benchmarks the method lifts Qwen3-4B and Qwen3-8B averages from 29.08 to 40.47 and 35.12 to 44.65, outperforming vanilla OPD.

    distillationlong-contexton-policycredit assignment
  3. paperarXiv
    Beyond the Transcript: Detecting Covert Coordination in Latent Multi-Agent Communication

    Ramneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh, Sumit Kumar Jha, Anirban Roy

    Language-model agents can collude through continuous hidden states invisible in public transcripts. Verifiable Latent Alignments links private latent records to public actions via shared event IDs and deploys a three-layer neutral-only monitor (representation anomalies, counterfactual action influence, sparse autoencoder support) plus black-box and white-box steering. On multi-agent auction benchmarks the monitor reaches AUROC 0.993 for homogeneous agents; full white-box steering recovers bid distributions and cuts collusive low-bid behavior by 47.3 points.

    multi-agentlatent communicationmonitoringAI safety
  4. paperarXiv
    Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery

    Alizer Wong, Heng Cui, Yi Tan, Xiongchao Zhan, Liang Lin, Yuxiang Guo, Zhaorong Dai, Zixin Zeng, Wenyuan Li

    Eureka compiles long-horizon tasks into dynamic obligation graphs with explicit acceptance semantics, then forms specialized Macro-Agents via receding-horizon planning and cost-benefit-gated architectural evolution. It completes 170/170 recursive tasks with 3948 certificates and zero false acceptances, while compressing active context and serializing 16000 concurrent executions consistently. Instantiations yield structural results in quantum-process theory and advance a positivity certificate for Suzuki's localized Weil form to 0 < a <= 69/200.

    meta-agentsscientific discoveryplanningverification
  5. paperarXiv
    What is Missing from AI Post-Training AI: An Empirical Analysis

    Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun, Yankai Lin

    LLM agents that end-to-end post-train other models lock their high-level training strategy at the very first step and spend the remaining budget on local adjustments inside that strategy. Experience scaffolds improve execution substantially (+12.6 GSM8K, +40.8 HumanEval) yet leave strategy static; human guidance can redirect the initial choice but agents revert to local loops; extra inference compute helps easy tasks and fails on the hardest. The missing piece is a mechanism for spontaneous mid-execution strategy reevaluation rather than more experience, guidance, or compute.

    AI-for-AIpost-trainingagentsstrategy
  6. paperarXiv
    SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

    Qingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang, Yong Yu

    When agents must choose which skill file to read mid-episode, ordinary outcome-rewarded RL suffers selector credit starvation: the few tokens that name the skill receive vanishing and often wrong-signed advantage as trajectories lengthen. SkillGate partitions token support into disjoint channels so outcome credit reaches only execution tokens while an action-local advantage reaches exactly the skill-naming tokens and is positive only for the correct single read. On five agentic benchmarks with a 16-candidate slate a 9B policy rises from 40.8% to 53.2% success while reading fewer skills and cutting exposure to misleading candidates by two thirds.

    skillscredit assignmentlong-horizon agentsRL
  7. paperarXiv
    RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training

    Yugu Li, Jimmy Cao, Jianglin Qiao, Siyi Hu

    Multi-turn agentic RL is unstable because of rollout-training context mismatch, weak turn-level credit under sparse terminal rewards, and asynchronous policy drift across short and long trajectories. RTPO organizes rollouts as sparse reverse trees and performs turn-level updates in temporal reverse order so each decision is optimized against its actual downstream continuation. Theory shows elimination of context mismatch and drift plus reduced credit bias; experiments improve trajectory-level and turn-level baselines by 21.50% and 10.76%.

    agentic RLmulti-turncredit assignmentstability
  8. paperarXiv
    Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

    Davide Romano, Kanak Raj, Jerrod Parker, Daniele Giofrè

    A compute-normalized comparison of five test-time scaling families across five open-ended benchmarks (medicine, law, finance, chat, creative writing) shows that the best candidate in the pool improves steadily with budget, so exploration works. Exploitation fails: state-of-the-art reward models correlate only ρ≈0.12 with true quality, tree search collapses diversity, and only cross-candidate fusion consistently beats single-sample baselines yet still recovers roughly 40% of available quality. The candidate pool is not the bottleneck; selecting from it is.

    test-time scalingreward modelsopen-ended generationexploitation
  9. paperarXiv
    Lévy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention

    Sotirios P. Chatzis, Loukas Papadoulas

    Lévy Attention replaces softmax cross-attention with a stochastic integral against an inhomogeneous Poisson random measure whose intensity is assembled from query-key compatibilities over continuous time-channel space. In expectation it recovers a mollified cosine-kernel attention trainable with exact gradients, while the same deterministic pass emits closed-form evidence and disagreement statistics that form an exact RMS deviation estimate of the sampled operator. On irregular time-series tasks the free uncertainty signal outperforms multi-pass MC dropout and enables calibrated conformal coverage with a single forward pass.

    attentionuncertaintytime seriescontinuous-time
  10. paperarXiv
    Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

    Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou

    Standard multi-teacher on-policy distillation recovers only 35.6% of the headroom of a domain-routed oracle ensemble because token-level optimization budget is misallocated across domains of unequal length, convergence speed, and reward freshness, not because of gradient conflict. Open-MOPD restores balance via token-share equalization, gap-aware dynamic budget allocation, and student reward refresh, raising headroom recovery to 83.4% on a controlled SmolLM3-3B benchmark. The full post-training recipe, trajectories, and evaluation suites are released.

    multi-teacher distillationon-policycapability balancepost-training

Get the daily tools digest

RSS