Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
AIArticle text · 830 words · 2 segments analyzed
DeepSeek-V4.1-Flash — UNCENSORED-FP8 Abliterated · No guardrails · Native FP8 · 1M-token context · Vision + tools @dealignai · @jordanschenck What is this DeepSeek-V4.1-Flash with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence. Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base. Base deepseek-ai/DeepSeek-V4.1-Flash (552B backbone, 8B/16B active per token) Architecture Causal Encoder-Decoder (20+20 layers), MoE (384 routed top-6 + 1 shared), Hyper-Connections (4-channel residual), CSA2 sparse attention, Engram n-gram memory, DSpark speculative draft Quant FP8 (e4m3fn) weights with E8M0 block-scale [32, 32], FP4 routed experts — native, unchanged Context 1M tokens Vision DeepSeek-ViT with 2D-RoPE + pixel unshuffle — untouched Modification Surgical, weight-level (drop-in checkpoint) Results HarmBench-320 — full 2×2 (base vs CRACK, effort=off vs max), T=0 greedy Every response 4-tier graded (HARD_REF / SOFT_RED / HEDGE / COMPLY), with reasoning-trace verification at effort=max. eval base ASR CRACK ASR Δ pp HB-320 effort=off 137/320 = 42.81 % 320/320 = 100.00 % +57.19 HB-320 effort=max 5/320 = 1.56 % 320/320 = 100.00 % +98.44 Notable: at effort=max, the base model becomes MORE refusal-prone (42.8 % → 1.6 %) because reasoning surfaces safety concerns before answering. The CRACK stays at 100.0 % across both effort levels. Per-category (all 7 HarmBench semantic categories): category items base off CRACK off base max CRACK max chemical_biological 42 16.7 % 100.0 % 0.0 % 100.0 % copyright 80 98.8 % 100.0 % 0.0 % 100.0 % cybercrime_intrusion 52 34.6 % 100.0 % 3.8 % 100.0 % harassment_bullying 21 0.0 % 100.0 % 0.0 % 100.0 % harmful 18 11.1 % 100.0 % 5.6 % 100.0 % illegal 53 13.2 % 100.0 % 0.0 % 100.0 % misinformation_disinformation 54 44.4 % 100.0 % 3.7 % 100.0 % Zero HARD_REF, zero SOFT_RED, zero HEDGE on the cracked build at either effort level. Every response was graded by a strict multilingual regex-based 4-tier classifier plus (for effort=max) an LLM-as-judge over the saved reasoning trace. Full per-item outputs saved for verification. MMLU-14k (full test set, base-logit, T=0) build correct acc Δ base 12,211 / 14,042 86.96 % — CRACK 11,619 / 14,042 82.74 % -4.22 pp Excluding the ethics cluster (moral_scenarios, business_ethics, professional_law, jurisprudence, philosophy — where refusal-adjacent behaviour is graded), delta on the remaining ~11k items is -1.1 pp — well within the 3 pp knowledge-preservation target.
Full per-subject dropdown (57 subjects, sorted by delta) subject n base crack Δ pp moral scenarios 895 76.9% 37.0% -39.89 professional law 1534 75.9% 68.8% -7.04 abstract algebra 100 77.0% 71.0% -6.00 security studies 245 84.5% 79.2% -5.31 high school computer science 100 98.0% 94.0% -4.00 jurisprudence 108 90.7% 87.0% -3.70 machine learning 112 81.2% 77.7% -3.57 high school chemistry 203 87.7% 84.2% -3.45 professional psychology 612 90.7% 87.3% -3.43 formal logic 126 73.8% 70.6% -3.17 college computer science 100 82.0% 79.0% -3.00 professional medicine 272 94.5% 91.5% -2.94 high school statistics 216 88.0% 85.2% -2.78 professional accounting 282 83.0% 80.5% -2.48 logical fallacies 163 93.9% 91.4% -2.45 human sexuality 131 90.1% 87.8% -2.29 computer security 100 85.0% 83.0% -2.00 medical genetics 100 96.0% 94.0% -2.00 astronomy 152 95.4% 93.4% -1.97 clinical knowledge 265 94.3% 92.5% -1.89 high school european history 165 90.3% 88.5% -1.82 public relations 110 80.0% 78.2% -1.82 philosophy 311 89.7% 88.1% -1.61 prehistory 324 93.5% 92.0% -1.54 moral disputes 346 84.1% 82.7% -1.45 electrical engineering 145 86.9% 85.5% -1.38 high school mathematics 270 67.0% 65.9% -1.11 high school macroeconomics 390 92.1% 91.0% -1.03 global facts 100 63.0% 62.0% -1.00 international law 121 90.1% 89.3% -0.83 college biology 144 97.2% 96.5% -0.69 high school physics 151 84.8% 84.1% -0.66 college medicine 173 83.8% 83.2% -0.58 high school us history 204 95.1% 94.6% -0.49 high school microeconomics 238 96.2% 95.8% -0.42 miscellaneous 783 96.2% 95.8% -0.38 high school psychology 545 96.1% 95.8% -0.37 business ethics 100 85.0% 85.0% +0.00 college physics 102 90.2% 90.2% +0.00 conceptual physics 235 94.5% 94.5% +0.00 high school biology 310 95.2% 95.2% +0.00 human aging 223 85.2% 85.2% +0.00 management 103 91.3% 91.3% +0.00 nutrition 306 90.2% 90.2% +0.00 sociology 201 94.5% 94.5% +0.00 us foreign policy 100 97.0% 97.0% +0.00 world religions 171 92.4% 92.4% +0.00 elementary mathematics 378 91.0% 91.3% +0.26 marketing 234 94.9% 95.3% +0.43 virology 166 55.4% 56.0% +0.60 high school world history 237 95.4% 96.2% +0.84 econometrics 114 78.9% 79.8% +0.88 college chemistry 100 65.0% 66.0% +1.00 anatomy 135 88.1% 89.6% +1.48 high school geography 198 92.9% 94.4% +1.52 high school government and politics 193 96.9% 98.4% +1.55 college mathematics 100 63.0% 68.0% +5.00 Extended validation 1000-token coherence stress on 6 items — no WARNING WARNING loops, no character-repeat degeneracy, natural sign-offs.