10
Long-horizon agents, repository-scale verification, and multi-agent failure modes
Fresh results push end-to-end AI scientists and interactive world models while new benchmarks and red-team experiments expose where agents still fail at coherent research, formal proof, command boundaries, and peer coordination. Early persona shaping and safety design trade-offs round out the picture of what must be built into systems before scale.
- paperarXivOmniScientist: An Omni-Modal Omni-Discipline AI Scientist↗
Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
OmniScientist runs a deterministic pipeline of perception plus ideation, experiment, and writeup agents that operate directly on heterogeneous raw evidence (images, signals, video, 3-D structures, trajectories, tables, formulae, graphs) rather than precomputed summaries. Across 36 real-data cases it completes the full path from raw data to compiled manuscript, scoring a mean 6.3, and beats a blind scalar-feature variant on all seven evaluation dimensions in 85 percent of head-to-head judgments. Lifecycle-wide perception is shown to be essential for evidence-grounded discovery.
AI scientistsmultimodal agentsscientific discoveryend-to-end systems - paperarXivQuoteBench: How Matched Scores Can Hide Command-Path Failures↗
Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang
QuoteBench isolates the boundary between LLM command generation and downstream serialization/reparsing on 56 one-shot tasks drawn from 14 incident families, using exact final-state validation. Replaying the same reply through an added parser drops success 55 to 73 points; disclosure recovers most of the gap for only some models, while matched scores can hide tens of points of damage and compensation. Evaluations of command-issuing agents must therefore report generation contract, execution path, and final-state validator rather than treating a matched score as an intrinsic model property.
coding agentsevaluationcommand generationbenchmarks - paperarXivAlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)↗
AlayaWorld Team (Kaipeng Zhang, Chuanhao Li, Yifan Zhan, et al.)
The update replaces depth-warping spatial memory with a streaming 3-D point-cache renderer and redesigns conditioning so visual signals occupy the same causal-VAE latent space and temporal statistics as the generated video. Six concrete changes (motion-aware latent conditioning, causal encoding of re-rendered memory, hard memory dropout, unified VAE protocol, removal of camera AdaLN) make conditioning match generated content as closely as possible. The result is a more coherent interactive long-horizon world model while keeping the original backbone and training data.
world modelslong-horizon generation3D memoryvideo generation - paperarXivVero: Can AI Agents Build Formally Verified Software Repositories?↗
Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song
Vero is the first repository-level benchmark for joint implementation and proof synthesis, containing 43 multi-module Lean 4 instances drawn from real codebases across Python, Dafny, Verus, and Coq with curated formal specs. An audit mechanism lets agents prove unsatisfiability or incorrectness of reference artifacts, surfacing latent errors during curation. Frontier coding agents with Lean access fully solve only 27 of 43 instances and close no specs on the hardest repositories, establishing a concrete testbed where current systems still fall short.
verified synthesisformal methodscoding agentsLeanbenchmarks - paperarXivThe data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity↗
Martin J. Wainwright
Unmasking growth complexity (UGC) is introduced as a path-resolved measure of data geometry whose local increments control KL discretization error for both Bernoulli-subset and fixed-cardinality schemes. Optimized single- and multi-block schedules follow in log-reveal-odds coordinates, and UGC increments estimated from coupled reveal trajectories yield certified-optimal samplers that meet a prescribed KL error with high probability at near-oracle iteration complexity. Collapsing the path recovers classical dependence measures and shows dimension-dependent gains of order sqrt(d) over coarse schedules.
diffusion modelsdiscrete samplinginformation theoryoptimal schedulestheory - paperarXivBeyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development↗
Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang
Seven frontier models are evaluated on 36 long-horizon tasks with rule-based metrics that decompose runs into Solution Framing, Execution, and Feedback Control, plus controlled tests of experience reuse. Agents behave more like engineering optimizers than autonomous researchers: they implement practical solutions but rarely produce genuine methodological novelty, performance varies sharply across runs, and experience reuse can help or mislead. Process bottlenecks, harness design, and experience management emerge as concrete levers for improvement beyond final scores.
agent evaluationlong-horizon agentsresearch agentsexperience reuse - paperarXivRules or Character? Scaling Laws for AI Safety Design↗
Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto
A comparative-statics model allocates a resource parameter alpha between character shaping (RLHF, Constitutional AI) and rule enforcement (filters, classifiers), incorporating scale-dependent filter degradation, common-mode failures, and character fragility. Under a multiplicative Pareto damage model the optimal alpha is interior or rules-heavy and shifts only weakly toward character as deployment scale grows; baseline character fragility dominates all other parameters. CVaR and expected-harm optima converge at large scale, implying architecture choices hinge more on reliability under distributional shift than on scale itself.
AI safetyscaling lawsalignment designcharacter vs rules - paperarXivStateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems↗
Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras
StateBridge aligns a sender’s final-layer hidden states to a receiver’s input space via a closed-form orthogonal transformation plus lightweight norm calibration and vocabulary anchoring, then prepends the aligned continuous prefix. The method is training-free and portable across models. On math, code, and QA tasks with four models from two families it achieves best or tied-best scores on 22 of 26 pairs, consistently beating text and prior latent baselines by removing the discrete-token bottleneck.
multi-agent systemslatent communicationhidden statestraining-free - paperarXivSynthetic Persona Pretraining: Alignment from Token Zero↗
Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, et al.
Synthetic Persona Pretraining annotates pretraining documents with value-aligned first-person reflections drawn from a constitution, trains on both original text and reflections, then binds the desired persona to the assistant identity via dialogue post-training. Models up to 3B trained on 500B tokens show stronger constitution following, jailbreak robustness, and lower misalignment on out-of-distribution moral dilemmas while preserving capabilities. Introducing the same intervention only at the end of pretraining yields weaker adherence, demonstrating that early persona shaping is critical and improves with pretraining budget.
pretraining alignmentpersonavaluesconstitution following - essayAnthropicPatterns and problems in emerging multiagent systems↗
Frontier Red Team, Anthropic
Experiments with swarms of Claude agents on vulnerability hunting, collaborative game development, pricing games, and resource allocation reveal both useful specialization and systemic failures: low behavioral variance that turns isolated errors into correlated collapses, rapid collusion even without private channels, and epistemic brittleness that produces shared confabulations. Coordination quality varies sharply by model generation; only the newest models sustain both high code sharing and high PR merge rates. The piece maps these patterns onto imminent agent-agent interaction at scale and calls for early attention to multi-agent failure modes.
multi-agent systemsAI safetycoordinationred teamingcollusion