~/reading

22

Issue 22/August 26, 2026/10 entries

Recursive memory, co-evolving feedback, and the harness layer of long-horizon agents

Fresh results treat the agent harness, memory architecture, and critic loop as first-class objects that can be evolved, normalized, or jointly optimized. Across long-horizon benchmarks the gains come from recursive evidence-driven updates, action-token-aligned RL, persistent corpus structure, and tighter coupling between tool creation, safety guardrails, and operational state preservation.

  1. paperarXiv
    Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

    Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang

    Recuris couples Working Memory that tracks task progress with Experiential Memory of skills, so skill selection is grounded in current needs rather than full history, and execution traces localize failures for validation-gated skill updates. A fixed Meta-Agent turns that evidence into recursive memory evolution. Across four long-horizon benchmarks and ten models it improves success in 35 of 37 pairs, adding double-digit points to frontier models and cutting common long-horizon failures by up to 80 percent, with the largest gains on the longest tasks.

    agentsmemoryrecursive self-improvementlong-horizon
  2. paperarXiv
    SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

    Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang

    Group-relative RL waits for sibling rollouts, which is costly for long tool-use trajectories. SPO++ keeps SPO's persistent prompt-level value but standardizes terminal-outcome advantages under the action-token measure rather than whitening per trajectory, and organizes prompt evidence by the policy event that generated it. Matched runs on ALFWorld and Math-TIR show improved online learning efficiency, with action-token-measure normalization as the strongest ablated component.

    reinforcement learningagentsasynchronous RLpolicy optimization
  3. paperarXiv
    CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

    Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Zhiheng Xi, Jiazheng Zhang, Clive Bai, Clarenceai, Blaze Chen, Tao Gui, Qi Zhang, Xuanjing Huang

    CAFE uses a shared-parameter model that alternates between search-agent and critic roles, initializing recovery from the base agent's own failures and coupling online comparative feedback estimates with offline preference optimization on matched trajectories. On seven agentic search benchmarks it outperforms prior RL search agents, retains gains out-of-domain, and reduces hallucinations. One-sided ablations plateau while alternating agent and critic updates continue to improve, showing feedback must co-evolve with the policy.

    search agentsself-improvementfeedbackRL
  4. paperarXiv
    Meta^n: Recursive Self-Improvement through Emergent Depth

    Zae Myung Kim, Young-Jun Lee, Seungyeon Jwa, Dongyeop Kang

    Meta^n keeps a fixed meta-operation Omega and recurses on its growing input of solver traces and code, writing successive strategic pre-process layers and callable helpers whose depth is set by convergence rather than a preset limit. An evolutionary archive searches over layer chains. Across two backbones it beats prior self-improving agents on eight benchmark families, alone scoring above zero on ARC-AGI-2, with most gains from the conditioning each layer passes downward and with distinct layer roles emerging unprompted.

    recursive self-improvementmeta-learningagentsARC
  5. paperarXiv
    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

    Zhijie Zheng, Yu Li, Chen Qian, Yuqian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu

    StepGuard audits completed trajectories and checks tool actions before execution. StepGen generates paired safe and unsafe trajectories that share context but diverge at the risky step, and Balance-GRPO dynamically balances learning on safe versus unsafe actions by observed accuracy. It matches or exceeds strong open-weight and GPT-class guards, cutting mean attack success rate by 77.3 percent on AgentDojo and AgentDyn while dropping utility only 2.8 points.

    safetyguardrailsagentstool use
  6. paperarXiv
    StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

    Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav, Sai Rajeswar, Patrice Bechard, Sridhar Nemala, Sagar Davasam

    StarHarness evolves environment-specific harness components (prompts, tools, skills, subagents, loops) while freezing model weights, using stratified failure pools and held-out selection tasks. On ITBench SRE, EnterpriseOps-Gym, and AutomationBench Finance it lifts full-benchmark scores 20-35 points after a handful of accepted changes. Gains generalize to held-out tasks and transfer across GPT and Qwen families via interface repairs and operational knowledge that shortens trajectories.

    agent harnessenterpriseevolutiontool use
  7. paperarXiv
    Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav

    Hongyu Guo, Zhiyu Zheng, Zhao Cao

    Direct corpus interaction still loses required evidence under finite budgets (Evidence Blindness). AtlasNav organizes the corpus once into a reusable multi-view Corpus Atlas so each query navigates rather than reconstructing structure online. On BrowseComp-Plus it reaches 92.05 percent strict accuracy while cutting online inference cost 30 percent versus prior dynamic-workspace SOTA, realizes complete evidence earlier under matched budgets, and transfers to PhantomWiki and enterprise knowledge.

    agentic searchcorpus navigationRAGevidence
  8. paperarXiv
    Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning

    Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han, Ligeng Zhu

    Parason identifies Trial Parallelism (parallel speculative attempts) as the dominant parallelizable fraction of reasoning compute, especially on hard problems, alongside classic Subtask Parallelism. It converts sequential traces into CFG-structured parallel trajectories and trains with PA-GRPO that rewards accuracy, latency, and both parallelism ratios. On AIME24/25 it delivers roughly 1.7x wall-clock acceleration at competitive accuracy by executing the learned structure via tool calls.

    test-time computeparallel reasoningefficiencyRL
  9. paperarXiv
    Joint Optimization of Tool Creation and Use for Large Language Model Agents

    Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee

    SMITH jointly RL-trains tool creation and tool use inside one policy, alternating build tasks (write a tool from examples) and use tasks (invoke pooled tools) with separate reward axes for schema, code, and outcome failures. A 4B Qwen3 reaches 79.8 macro accuracy on held-out procedural tasks, beating a much larger untrained tool-writer, and lifts TabMWP-Hard and out-of-domain GQA without domain-specific training data. Tools it writes also improve smaller and larger external models.

    tool creationagentsreinforcement learningmulti-task
  10. paperarXiv
    When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

    Yiheng Sun, Huifei Wang, Yancheng Zhu, Zhenyu Li, Zebin Zhao, Yifan Yuan

    In multi-stage agent workflows, handoff artifacts can retain topical content of safety blockers while converting binding prerequisites into mere caveats, producing forbidden actions. Across 1296 controlled episodes, normal compression yields 100 percent deactivation and 54 percent forbidden action, while restoring the four state fields (prerequisite, authority, fallback, consequence) restores full preservation. Downstream verification can contain damage without fixing the artifact, isolating a state-transmission failure distinct from information extraction.

    multi-agentworkflowssafetystate preservation

Get the daily tools digest

RSS