Skip to content
HN On Hacker News ↗

Anthropic Risk August 2026 [pdf]

▲ 56 points 58 comments by artninja1988 2w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

1 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,518
PEAK AI % 1% · §1
Analyzed
Aug 14
backend: pangram/v3.3
Segments scanned
1 windows
avg 1518 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,518 words · 1 segments analyzed

Human AI-generated
§1 Human · 1%

Risk Report: August 2026 anthropic.com 1 Introduction and executive summary 7 1.1 Structure of the report 8 1.2 Executive summary of findings 9 1.3 Changes to our RSP since the most recent Risk Report 13 1.3.1 Updated threshold for automation of AI R&D 13 1.3.2 Updated threshold for development of novel biological and chemical weapons 14 1.3.3 Coverage dates of risk reports 14 1.3.4 Redaction scope and transparency 15 1.3.5 Governance and review changes around Risk Reports 15 1.4 Notes on coverage of unreleased models 15 2 Autonomy threat model 1: Misalignment in high-stakes settings 17 2.1 Overview 17 2.2 Threat model 18 2.2.1 Specific pathways 20 2.3 Relevant AI models 20 2.4 Summary and methodology 20 2.5 Definitions 22 2.6 Claims and core argument 25 2.7 Claim 1: Models are unlikely to have strong covert capabilities 27 2.8 Claim 2: Expected harm from known misalignment is low 36 2.9 Claim 3: Unknown severe pervasive misalignment is very unlikely 39 2.9.1 Claim 3.1: Experience with prior Anthropic models suggests that unknown severe pervasive misalignment is unlikely in covered models 39 2.9.2 Claim 3.2: Our experience with adversarially-designed training processes suggests that unknown severe pervasive misalignment is somewhat unlikely in covered models 40 2.9.3 Claim 3.3: It is somewhat unlikely that the specific training processes used to train the covered models produced unknown pervasive misalignment 41 2.9.4 Claim 3.4: The presence of unknown severe pervasive misalignment is unlikely given our alignment assessment 45 2.10 Claim 4: Expected harm from unknown severe context-dependent misalignment is low 51 2.11 Claim 5: Catastrophic harm from known misalignment and unknown context-dependent misalignment is likely to be mitigated 53 2.11.1 Claim 5.1: Catastrophic harm from reversible threats other than diffuse sandbagging, emerging from internal usage, and caused by known or context-dependent misalignment, is likely to be mitigated 54 2.11.2 Claim 5.2: Catastrophic harm from diffuse sandbagging caused by known or context-dependent misalignment is likely to be mitigated 60 2.11.3 Claim 5.3: Catastrophic harm from irreversible threats, emerging from internal usage, and caused by known or context-dependent misalignment is likely to be mitigated 60 2.11.4 Claim 5.4: Catastrophic harm from external usage caused by known or context-dependent misalignment is unlikely compared with internal usage, and somewhat likely to be mitigated 62 2.12 Claim 6: Catastrophic harm from unknown pervasive misalignment is somewhat likely to be 2 mitigated 63 2.13 Claim 7: Threat modeling is sufficient 63 2.14 Claim 8: Risk from deliberately engineered misalignment is low 65 2.15 Pathway-specific risk assessments 65 2.16 Limitations 67 2.16.1 Models might have stronger covert capabilities than we believe 67 2.16.2 Relevant capabilities and propensities might change suddenly with scale 68 2.16.3 Evaluation awareness may be more prevalent than we believe 68 2.16.4 There may be forms of context-dependent misalignment that are harder to find and more harmful than we expect 68 2.17 Considerations outside the scope of this section 69 2.18 Pre-internal-deployment review process 70 2.19 Overall assessment of risk 71 2.20 Claude’s review of this risk assessment 72 2.21 Looking forward 74 2.22 Models other than Mythos 5 and Model 2 75 2.23 Appendix: Detailed description of internal usage monitoring 76 2.23.1 Asynchronous monitoring 76 2.23.1.1 Automated offline monitoring overview 76 2.23.1.2 Automated offline monitoring evaluations 78 2.23.1.3 Offline monitoring during training and evaluations 81 2.23.1.4 Other forms of asynchronous monitoring 81 2.23.2 Blocking interventions 82 2.23.2.1 Model weight security 82 2.23.2.2 Sandboxing in training and evaluations 82 2.23.2.3 PR reviews 83 2.23.2.4 Sandboxing and blocking classifiers in internal deployments 83 2.23.3 Important limitations of our model-external risk mitigations 83 2.23.4 Changes to risk mitigations since our previous risk report 84 2.24 Appendix: Details of power seeking environment evaluation 84 2.25 Appendix: Active research into reward hacking generalization 86 2.25.1 Setup 87 2.25.2 Evaluation behavior 87 3 Autonomy threat model 2: Risks from automated R&D 93 3.1 Overview 93 3.2 Threat model 94 3.3 Relevant AI models 96 3.4 Substitution of our AI models for Anthropic researchers 97 3.4.1 Failure modes from internal use 98 3.4.2 Researcher survey 99 3.4.3 CoBench 99 3 3.5 Acceleration of AI progress due to automated AI R&D 101 3.5.1 AECI capability trajectory 102 3.5.2 Conclusions on the overall degree of AI R&D acceleration 103 3.6 Automated R&D in other domains 103 3.6.1 Introduction 103 3.6.2 Robotics 105 3.6.3 Biotechnology 105 3.6.4 Energy 106 3.6.5 Semiconductors 108 3.6.6 Weapons development 108 3.6.7 Neurotechnology 109 3.6.8 Nanotechnology 110 3.7 Our risk mitigations 110 3.7.1 Changes to risk mitigations since our previous risk report 110 3.8 Overall assessment of risk 112 3.9 Looking forward 112 3.10 Connection to our recommendations for industry-wide safety 113 4 Chemical and biological weapons production 115 4.1 Overview 115 4.2 Threat models 116 4.2.1 CB-1 threat model 116 4.2.2 CB-2 threat model 118 4.2.3 Timescale and scope of uplift 121 4.3 Relevant AI models 122 4.4 Current state of model capabilities 123 4.4.1 Notes on how we weigh evidence 123 4.4.2 CB-1 evidence 123 4.4.2.1 General update from a randomized controlled trial on uplift from AI models 124 4.4.3 CB-2 evidence for Mythos Preview, Fable 5, and Mythos 5 125 4.4.4 CB-2 evidence for Opus 4.8 and other Opus and Sonnet models 128 4.5 Our risk mitigations 130 4.5.1 Robustness levels: Level 1, Level 2, Level 3 132 4.5.2 Coverage levels 133 4.5.2.1 Extension in coverage since our prior Risk Report 133 4.5.2.2 Wider coverage of CB classifiers for Fable 5 134 4.5.3 Evidence about robustness of classifiers 135 4.5.3.1 Ongoing bug bounty results 135 4.5.3.2 Notes on jailbreak methods that remain viable on at least some models in some cases 136 4.5.3.2.1 Boundary-point jailbreaking 136 4.5.3.2.2 UK AISI-sourced jailbreak 139 4 4.5.3.2.3 Coverage gap identified by both Anthropic and UK AISI 139 4.5.3.2.4 Bug-bounty-sourced jailbreak 139 4.5.4 Offline monitoring 139 4.5.5 Bioclassifier exemptions 140 4.5.5.1 Exemptions for Mythos Preview, Mythos 5, and Fable 5 140 4.5.5.2 Exemptions for all other sub-Mythos-class commercial models 141 4.5.5.3 Access controls for other models not available for general access 142 4.5.5.3.1 Helpful-only models 142 4.5.5.3.2 Claude Gov 143 4.5.5.3.3 Pre-deployment access to unreleased models 143 4.5.5.4 Malicious access via prompt injection 144 4.5.6 Threat intelligence 145 4.5.7 Model weight security 146 4.5.8 Changes to our risk mitigations since our last Risk Report 146 4.5.8.1 Slight relaxations to our model weight storage criteria 146 4.5.8.2 Incidents related to our classifiers and access controls 147 4.5.8.2.1 Unauthorized Mythos Preview access on our human feedback platforms 147 4.5.8.2.2 All human feedback vendor traffic run without blocking biological classifiers 148 4.6 Overall assessment of risk 150 4.6.1 Risks from the CB-1 threat model 150 4.6.2 Risks from the CB-2 threat model 151 4.7 Looking forward 153 4.8 Connection to our recommendations for industry-wide safety 155 5 Cross-cutting content 158 5.1 Acceleration dynamics 158 5.1.1 Distillation 159 5.1.1.1 Introduction 159 5.1.1.2 Thinking and connector text summarization 160 5.1.1.2.1 Mitigations for deliberate bypasses 163 5.1.2 Our overall state on acceleration dynamics 163 5.2 Safety process failures 163 5.2.1 Introduction 163 5.2.2 Partial refusals on safety work undermining stress-testing research 164 5.2.3 Exposing chain-of-thought reasoning to grading pressure 165 5.2.4 Directly training on misaligned behavior during a production training run 166 5.2.5 An instance of unmonitored unrestricted agents with access to sensitive resources 167 5.2.6 Repeated training on alignment-faking transcript datasets 168 5.3 Benefits from Anthropic’s operating as a frontier AI company 169 5.3.1 Key examples of beneficial Anthropic activities and choices since our last Risk Report169 5.4 Overall risk assessment and risk-benefit determination 173 5 5.5 Looking back on roadmap progress and AI development and deployment decisions 174 5.5.1 Decisions to develop and deploy increasingly capable models 174 5.5.2 Updates on our Frontier Safety Roadmap 174 6 Appendices 175 6.1 Threat model criteria 175 6.2 [Appendix redacted] 176 6.3 [Appendix redacted] 176 6.4 Model weight security 176 6.4.1 Overview 176 6.4.2 Notable security controls 177 6.5 List of minor incidents related to our CB risk mitigations 178 6.6 Our AI models 181 6.6.1 List of all Claude models, sorted by release date 185 6 1 Introduction and executive summary This report evaluates the degree to which Anthropic’s AI systems pose catastrophic risk in several categories, in light of what we know about both their capabilities and the measures we have in place for mitigating risk. We focus on a small number of particularly salient catastrophic risks, for reasons discussed below. This report is part of our implementation of version 3.4 of our Responsible Scaling Policy (see also our Frontier Safety Roadmap, which describes our goals and specific plans for safety mitigations). Our system cards, which are published each time we release a model, provide analysis on some dimensions of risk—in particular, assessing our AI models for capabilities that may have dangerous as well as beneficial uses. This report, which we aim to publish every 3–6 months (see the RSP for timing details), goes beyond the analysis presented in those documents in several ways: 1. We discuss not only properties of our AI models that are relevant to catastrophic risk, but also properties of our risk mitigations, including security controls and deployment safeguards. By considering the whole picture of