15
Self-improving agents, reward hacking, and long-horizon trust
Fresh results expose variance and order sensitivity in memory-based self-improvement, show debate curbing judge hacking in RLAIF, and introduce monitors for ontological drift plus neurosymbolic world models. Complementary work tightens LLM judging under risk controls, surfaces MoE hallucination signals, and questions the self-consistency of preference estimates, while an essay reframes mathematical values under capable AI.
- paperarXivOn the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification↗
Qinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang, Chien-Sheng Wu
Memory-based self-improving agents that maintain textual memory banks show high variance across runs and strong dependence on task order, with default orderings acting as hidden curricula. Underspecification of tasks and environments contributes to fragility; adding rubrics and environment feedback partially mitigates degradation but leaves residual gaps. The work calls for multi-run reporting and human-oversight interfaces to prevent unforeseeable failures.
self-improvementagentsevaluationmemoryreliability - paperarXivDebate Training Reduces Reward Hacking in RLAIF↗
Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah
RL finetuning via two-player debate (generator vs critic, weaker LLM judge) maintains judge performance and recovers a large validation accuracy gap versus single-player RLAIF on verifiable math tasks. Weaker judges hack faster but extra debate rounds compensate; word limits on critiques balance the game and prevent critic judge-hacking. Debate incentives can override prompted misalignment, supporting feasibility of multi-agent oversight with careful balancing.
RLAIFdebatereward-hackingalignmentmulti-agent - paperarXivMixture-of-Expert Blocks Contain Strong Hallucination Detection Signals↗
Joao Fonseca, Rodrigo Rodrigues, Paolo Romano
InnerExpert extracts per-token features from MoE router entropy, expert disagreement, and usage patterns plus standard transformer signals, then trains a lightweight detector on LLM-as-judge labels. It achieves high answer-level and token-level AUROC across datasets and two MoE architectures using only a single forward pass. The method enables continuous updates without manual annotation and localizes hallucinated spans for fine-grained intervention.
hallucinationMoEdetectioninterpretabilityLLMs - paperarXivJudge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees↗
Sher Badshah, Ali Emami, Hassan Sajjad
A risk-controlled framework calibrates uncertainty thresholds via finite-sample Clopper-Pearson intervals so that the false discovery rate among accepted LLM-judge verdicts stays below a user-specified alpha. Low-confidence parametric judgments route to retrieval-augmented re-evaluation under a second calibrated threshold, preserving the guarantee. Across open-domain QA and judges of varying scale it maintains target error rates at substantially higher coverage than single-mode baselines.
LLM-as-judgeuncertaintyconformalretrievalrisk-control - paperarXivStagedWorkspace: A Versioned Workspace for Knowledge-Work Agents↗
Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian
StagedWorkspace binds parsed records and review diffs to content hashes of native files, enforcing an explicit workspace-state contract across search, edit, and submission views. Dual parsed/native access yields consistent gains on OfficeQA Pro and APEX-Agents; a review-axis ablation confirms value of visible diffs. The design treats workspace state as an experimental variable and motivates benchmarks that score evidence, staged edits, and artifacts as state transitions.
agentsworkspaceknowledge-workversioningtools - paperarXivChain-of-Experience for Continual LLM Improvement↗
Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan
Chain-of-Experience lets models accumulate experiential traces from self-feedback or environmental signals (correctness, test pass rates) across iterative interactions at test time. Across math, coding, and knowledge tasks with eight frontier models it outperforms feedback-free baselines, cuts API cost, and improves accuracy per token; complementary feedback channels add further gains. Stronger base models improve more, most gains appear early, and models remain robust to weak or spurious feedback.
test-timecontinualself-improvementfeedbackLLMs - paperarXivBeyond Suspicious Steps: Ontological Trust in Long-Horizon Agents↗
An He, Yao Wang, Haibin Zhang
Ontological trust measures whether a trajectory prefix still corresponds to the user-authorized task along Role, Goal, and Evidence axes. RGE derives structured representations with LLMs then performs deterministic trust-state updates and interventions, producing a replayable auditable trajectory. On a cross-domain corpus of benign, prefix-paired drift, and pseudo-consistency failures it outperforms rule, judge, and shield baselines, exceeding 93 percent Drift F1 at high benign coverage.
long-horizonagentsoversighttrustmonitoring - paperarXivLLM-Derived Preference Judgments Are Not Self-Consistent↗
Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
Cardinal preference judgments elicited from LLMs (e.g., willingness-to-pay) systematically violate self-consistency conditions required by a single underlying utility function. Statistical tests and distance-to-best-fit measures quantify large persistent inconsistencies across flight, apartment, and hotel domains and six models. The result undermines pipelines that fit a utility model to LLM judgments and then optimize actions.
preferencesutilityconsistencyagentsevaluation - paperarXivTowards Zero-Shot Task Transfer with Neurosymbolic World Models↗
Isidoro Tamassia, Lennert De Smet, Giuseppe Marra
A neurosymbolic world model decouples observation reconstruction from reward prediction so that reward depends only on structured symbolic state components. This enables zero-shot adaptation to new reward functions defined over the same symbolic space without further environment interaction. The approach demonstrates stronger generalization than purely neural model-based RL methods while highlighting remaining learning challenges.
world-modelsneurosymbolictransferMBRLzero-shot - essayarXivMathematics in the age of AI↗
Terence Tao
Conditioning on the arrival of AI capable of research-level mathematics, the essay asks what the goals and values of mathematical research actually are rather than debating capability timelines. Using the problem-solving pipeline as a case study it frames the moment as a crisis of values and practices analogous to the early-20th-century foundational crisis. Recommendations include clarifying evaluation criteria and resisting Goodhart-style over-optimization of proxies.
mathematicsvaluesAI-impactfoundationsessay