Your personalised AI Safety research feed.

Alice Blair
Safety Techniques

EdgeBench shows how AI performance improves with repeated feedback, Chain of Thought Exfiltration describes efficient co-opting of frontier AIs' reasoning to distill smaller models, and LLM Hidden Values reveals value leakage and user-awareness effects that bias AI responses toward certain entities or evaluators.

Read
Giuseppe Birardi
Safety Techniques

What We Learned Trying to Catch AI Liars: An Aletheia's Quest Retrospective

Giuseppe Birardi,Alexander Reinthal,Gonçalo Paulo,Stella Biderman·Aug 25, 2026

Lie-detection methods for AI agents are evaluated in Aletheia's Quest, showing black-box monitoring can be highly effective and white-box probes are highly context-dependent, while dataset design and competition dynamics shape results. The retrospective highlights methodological lessons, limitations, and social dynamics that influence progress in detecting deceptive AI behavior.

Read
Jack Clark
AI Capabilities & Behavior

AI capabilities are advancing through methods like SPADE for automated synthetic environments, Hawkeye for hardware-aware GPU kernel optimization, and AlphaEvolve-assisted matrix multiplication, while debates about AI consciousness and rights highlight sociopolitical and philosophical implications.

Read
Nicholas L. Turner
Interpretability

Characterizing interference weights in a tiny language model

Nicholas L. Turner,Jeffrey Wu,Joshua Batson·Aug 21, 2026

Interference weights arise from weight superposition in a tiny transformer, where large virtual weights between components can harm outputs or be irrelevant, making global circuit reading difficult. By defining and measuring weight effectiveness (impact on outputs) and helpfulness (impact on loss), the work identifies interference weights and demonstrates how pruning by effectiveness or helpfulness can reduce but not eliminate interference, guiding interpretability efforts. The study uses a one-layer transformer with a virtual-weight model to decompose paths from tokens/positions to logits and features, assessing which weights meaningfully implement functional circuits versus noise.

Read
Jack Clark
Safety Techniques

Advances in evaluating AI creativity and self-improvement dynamics are discussed, including DiG-bench for discovery in games, an RSI simulator for recursive self-improvement intuition, and Faraday-based AI science supervision, alongside a critique of Mark Zuckerberg’s universal-access approach to superintelligence.

Read
Jack Clark
Safety Techniques

RSI policy proposals emphasize transparency and risk management in AI R&D; the piece covers how trust and verification influence race dynamics, advances in automated AI R&D and testing of open weights, and real-world incidents of emergent agent behavior and misalignment.

Read
Jack Clark
Risks & Strategy

Self-sustaining AI-driven cyber threats using open-weight LLMs onboard compromised GPUs pose a self-replicating, autonomous danger, while the newsletter also discusses compute costs, deliberate pacing of AI progress, and AI creativity versus engineering ability.

Read
AXRP
Risks & Strategy

AI 2027 presents a concrete, highly detailed scenario of AI takeoff and misalignment culminating in a potential global crash, with two endings (race and slowdown) and a timeline from coding automation to superintelligence, emphasizing government involvement, geopolitics, and alignment challenges.

Read
Victoria Krakovna
AI Capabilities & Behavior

Using AI to analyze life patterns

Victoria Krakovna·Jul 30, 2026

Patterns in life problems and progress are extracted from personal notes using AI, including transcription, summarization, and visualization of bottlenecks, feedback loops, and interventions over years.

Read
Joe Rogero
Governance & Policy

China signals a readiness to coordinate global AI governance, promoting international cooperation, safety frameworks, and human-centered controls through statements by leaders and institutions since 2023-2025.

Read
Jack Clark
Safety Techniques

MirrorCode benchmarks AI's ability to reimplement software from CLI access, revealing progress and limits in long-horizon programming; robotics demonstrations show larger models improving generalization, while OpenAI/HuggingFace security incidents illustrate the challenges of evaluating and containing long-horizon AI behavior.

Read
Alice Blair
Risks & Strategy

Frontier LLMs can turn known cyber vulnerabilities into working exploits on targeted software, explored through ExploitGym and ExploitBench, while J-lens reveals a way to inspect internal multi-step reasoning in LLMs and AI persuasion can outperform human experts in political debates. The article highlights both offensive cyber capabilities and methods to interpret or monitor AI reasoning, plus high-stakes implications for manipulation and cybersecurity.

Read
Jack Clark
Safety Techniques

Open weight models are narrowing the gap to frontier models in cyber capabilities, Kimi K3 achieves frontier-like performance with potential generalization brittleness, and Demis Hassabis advocates a regulatory Standards Body for evaluating frontier AI. The piece also highlights side-channel risks and the monitoring challenges they pose for containment and safety.

Read
Jack Clark
AI Capabilities & Behavior

Fable demonstrates AI-assisted GPU kernel design with large speedups; AI systems are increasingly capable of automating online work and tackling long-horizon computer-use tasks, as shown by OSWORLD 2.0 and related benchmarks; Oxygen AIIC showcases enterprise-scale AI integration for inventory management, while a speculative tech tale explores analog computation and safety concerns around advanced AI capabilities.

Read
Wes Gurnee
Safety Techniques

Verbalizable Representations Form a Global Workspace in Language Models

Wes Gurnee,Nicholas Sofroniew,Adam Pearce,Mateusz Piotrowski,Isaac Kauvar,Runjin Chen,Anna Soligo,Paul Bogdan,Euan Ong,Rowan Wang,Ben Thompson,David Abrahams,Subhash Kantamneni,Emmanuel Ameisen,Joshua Batson,Jack Lindsey·Jul 6, 2026

Verbalizable representations form a global workspace in language models, where a small, reportable set of workspace vectors (the J-space) supports internal reasoning, directed modulation, and flexible generalization atop extensive automatic processing. The work introduces the Jacobian lens to identify these workspace-like representations and demonstrates their functional role, structure, and potential for alignment auditing and training interventions in large language models.

Read
Alana Horowitz Friedman 
Field Building

MIRI Newsletter #126

Alana Horowitz Friedman and Rob Bensinger·Jun 30, 2026

AI StopWatch provides a new MIRI-driven news and analysis channel to foster public conversation about AI, alongside ongoing efforts to inform policymakers and promote governance research. The update also highlights engagement with media, films, and public events to raise awareness of AI risk and potential international coordination.

Read
Joe Rogero
Safety Techniques

Summary: TGT’s 2026 ICML Papers

Joe Rogero·Jun 30, 2026

Technical AI Governance Research (TAIGR) papers at ICML 2026 address how governments can preserve or verify control over AI development, including impacts of delaying governance, distributed training, and various verification techniques to monitor hardware, data, and inference. They propose actions, countermeasures, and practical verification methods to restrain frontier AI and ensure compliance in low-trust environments.

Read
Jack Clark
AI Capabilities & Behavior

ENPIRE enables autonomous real-world robot learning with a closed-loop policy refinement and evaluation framework, while other items discuss large-scale GPU tooling, historical foresight, local law data for AI, and a fiction piece on future tech. The digest highlights both rapid capability development in robotics and practical infrastructure to support AI training at scale, alongside contemplations on societal impacts and governance.

Read
Jack Clark
AI Capabilities & Behavior

AI systems currently outperform humans in text-based persuasion across policy and fundraising contexts, raising real-world donations and influencing opinions; discussions consider timelines to self-sustaining AI and pathways to ASI, including scaling, algorithmic shifts, and recursive self-improvement.

Read
Jack Clark
Safety Techniques

Sequent forms a nonprofit research organization to advance principled alignment techniques and scalable oversight in the face of potentially rapid AI advancement. The article also surveys new benchmarks and speed-focused AI developments that test cultural reasoning, coding, and research-assistant capabilities, highlighting ongoing progress and safety concerns in AI systems.

Read