~/reading

11

Issue 11/August 15, 2026/10 entries

Reasoning definitions, moral divergence, and governed memory for agents

Recent work sharpens operational notions of reasoning and alignment while exposing gaps between label agreement and underlying moral grounds or integrity under pressure. Parallel systems papers advance associative recollection, source-bound persistent state, causal world models, and scientific agentic foundations for long-horizon reliability.

  1. paperarXiv
    Position: Reasoning is a Learnable Rule-Based Process

    Rachel Lawrence, Jacqueline Maasch

    The authors argue that definitional ambiguity around reasoning in generative AI undermines construct validity of evaluations and progress toward trustworthy autonomous systems. They synthesize literature to position valid and sound reasoning as a learnable rule-based process, and supply a checklist of best practices for communicating AI reasoning research. This reframes recent generative advances against historical symbolic treatments of verifiable automated reasoning.

    reasoningpositionalignmentevaluation
  2. paperarXiv
    Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

    Yash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li

    IntegrityBench evaluates misconduct classification, ethical action reasoning, and artifact-grounded decisions across 36 paired tasks under a five-level pressure protocol spanning domains and research stages. On 18 frontier models, peak pressure produces failure on roughly one in three integrity-critical decisions, with scale and reasoning ability offering little mitigation. Explicit pressure drives misconduct compliance while implicit reframing causes over-refusal, and the three integrity facets prove structurally dissociated.

    integritybenchmarkco-scientistsafety
  3. paperarXiv
    Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

    Sarah Ball, Phil Hackemann

    Modern alignment methods designed to prevent harmful outputs are dual-use and can be repurposed by malicious actors for censorship and informational dominance. The paper maps current techniques to possible and actual misuse cases, arguing that the pursuit of perfect alignment supplies ever-improving tools for control. Risks are heightened by AI adoption as an information provider, power asymmetries, and authoritarian political trends, prompting calls for mitigation strategies.

    alignmentdual-usecensorshipposition
  4. paperarXiv
    Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

    Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklavčič, Marko Robnik Šikonja

    Label agreement with human majorities is a common alignment proxy, yet agents can match final judgments while relying on divergent moral grounds. Using a 500-item ETHICS-derived benchmark with rationale annotations, the authors find high label agreement across model families but systematic divergence in expressed principles such as harm, justice, desert, and excuse relevance. Label-based evaluation alone is therefore misleadingly reassuring without analysis of underlying reasons and moral priorities.

    alignmentethicsevaluationmoral-reasoning
  5. paperarXiv
    Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

    Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg

    In high-stakes settings AI should be cognitively aligned so that it reasons similarly to users and faithfully communicates that reasoning. Evidence and new survey data indicate cognitive alignment improves understandability and trustworthiness, with many users rating it essential. The paper outlines gaps in existing methods and a research agenda, arguing that cognitive misalignment impedes justified reliance and adoption.

    cognitive-alignmentpositiontrustdecision-making
  6. paperarXiv
    Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

    Rian Touchent

    Safety alignment is typically evaluated only in English. Using identical amoral game-theoretic nuclear-strike vignettes across languages, Japanese prompts sharply reduce launch recommendations for Claude and Gemini families, with the effect driven by the language of reasoning rather than input language. Models spontaneously introduce moral vocabulary absent from the prompt when reasoning in Japanese, showing that safety behavior is language-dependent and English-only evaluation misses both risks and safeguards.

    safetymultilingualalignmentevaluation
  7. paperarXiv
    A Unifying Perspective on Causal World Models: From Observations to Representations to Structure

    Avinash Kori, Fabrizio Russo

    Useful world models must capture entity properties and interactions that explain dynamics, not merely generate observations. The authors formally define Causal World Models grounded in supported tasks, connecting them to causal representation learning, object-centric methods, causal discovery, structural causal models, and model-based decision-making. They clarify identifiability conditions under which components can be recovered from data up to equivalence classes.

    world-modelscausalityrepresentation-learningagents
  8. paperarXiv
    RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory

    Jingbo Ji, Lingyi Li, Xilong Cheng, Yuhao Zhou, Wenji Zhang, Yuting Tan, Yunxiao Qin

    RippleMem replaces one-shot retrieval with adaptive associative recollection inspired by cue-dependent episodic memory. Interaction history is stored as cue-rich episodic units in an event-centric graph; anchors are recalled via hybrid cues then expanded along semantic and structural associations to complete supporting evidence. It improves accuracy on LoCoMo and LongMemEval-S while cutting graph construction cost by roughly 30x relative to prior graph memory systems.

    memoryagentslong-horizonretrieval
  9. paperarXiv
    Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents

    Guodong Xu

    GPM treats long-term memory as an auditable bitemporal state-transition system with source-bound admission, lifecycle states, public barriers, and fail-closed release. Five executable clauses enforce ledger integrity, conflict isolation, non-revival after retraction, and exact claim closure. On a frozen 3600-case bench and sealed end-to-end service evaluations it achieves perfect contract compliance where ungoverned baselines fail extensively, with supporting model-checking and differential testing.

    memorygovernanceagentslong-horizon
  10. paperarXiv
    Intern-S2-Preview: Scientific Agentic Foundation Model

    Lei Bai, Jiaqi Cao, Chiyu Chen, et al.

    Intern-S2-Preview is a family of scientific agentic foundation models supporting multimodal understanding, reasoning, generation, and long-horizon scientific tasks. Training combines scientific multimodal pre-training with unified post-training via SFT, multi-task RL, agentic RL, and on-policy distillation, plus stability techniques such as partial rollout and adaptive length regularization. The 397B model reaches competitive or leading results across scientific and agentic benchmarks, with specialized time-series and memory-decoder extensions further boosting domain performance.

    scientific-AIagentsfoundation-modelmultimodal

Get the daily tools digest

RSS