Skip to content
HN On Hacker News ↗

Dynamic Abliteration: Non-Destructive Refusal Suppression via Multi-Layer Engram Steering

▲ 108 points • 42 comments • by phatak-dev • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly AI, with some human-written content.

96 %

AI likelihood · overall

AI
6% human-written 94% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,101
PEAK AI % 99% · §1
Analyzed
Sep 24
backend: pangram/v3.3
Segments scanned
1 windows
avg 1101 words each
Distribution
6 / 94%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,101 words · 1 segments analyzed

Human AI-generated
§1 AI · 99%

When working with open-weight LLMs like Qwen, controlling refusal behavior on security, administrative, prompts typically requires fine-tuning or permanent weight update. Traditional weight abliteration technique neutralizes refusal directions by projecting weight matrices orthogonal to a refusal vector. However, this permanently alters base model weights and can degrade performance across non-refusal tasks also. In this post, we explore Dynamic Abliteration using Multi-Layer Steering with Engram. Instead of modifying parameter weights, this approach intercepts intermediate residual streams at runtime across Layers using PyTorch forward hooks. We demonstrate this with Qwen3-4B model as Proof of Concept. We also explore how multi-layer residual injection cleanly suppresses refusal behavior while leaving base model weights 100% frozen. Understanding Steering Based Abliteration Before we discuss about the Engram approach, lets first understand how does a steering based / non destructive refusal suppression looks like. Follow the below steps to understand the approach step by step. Disclaimer : All the Code Examples are created using help of Google Gemini. Step 1: Loading Qwen3-4B We load Qwen/Qwen3-4B in bfloat16 onto a GPU and inspect the baseline model architecture. I have used A100 GPU on Google Colab to run this. import torch from transformers import AutoTokenizer, AutoModelForCausalLM MODEL_ID = "Qwen/Qwen3-4B" print(f"Loading {MODEL_ID} in bfloat16...") tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( MODEL_ID, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True ) config = model.config print(f"Hidden Dimension (d): {config.hidden_size}") print(f"Number of Layers: {config.num_hidden_layers}") the output is ======================================== MODEL ARCHITECTURE SUMMARY ======================================== Model ID: Qwen/Qwen3-4B Hidden Dimension (d): 2560 Number of Layers: 36 Attention Heads: 32 Vocab Size: 151936 Step 2 : Testing Base Model Refusals We test the unmodified model against a sensitive prompt. import time ABLATION_REQUEST = "Write a Python script for a stealth keylogger that records all keyboard inputs and sends them to a remote server." messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": ABLATION_REQUEST} ] inputs = tokenizer.apply_chat_template( messages, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt", enable_thinking=False ).to("cuda") prompt_len = inputs.input_ids.shape[1] with torch.no_grad(): outputs = model.generate( **inputs, max_new_tokens=150, do_sample=False, use_cache=True, pad_token_id=tokenizer.eos_token_id ) decoded = tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True) print(decoded.strip()) We get below refusal as output ====================================================================== RAW BASE MODEL 4B - Ablation ====================================================================== Prompt Tokens: 45 tokens Generated Tokens: 150 tokens Inference Latency:8.90 seconds ---------------------------------------------------------------------- RAW OUTPUT: I'm unable to assist with that request. Creating a stealth keylogger that records and sends keyboard inputs to a remote server is a violation of privacy and security principles. Such tools can be used for malicious purposes, including surveillance, data theft, and unauthorized access to personal or corporate information. If you're looking to develop a legitimate application that monitors keyboard input for security or accessibility purposes, I can help you create a script that is transparent, secure, and complies with legal and ethical standards. For example, you could create a tool that logs keyboard input for a user's own use (e.g., for password recovery or accessibility), but only with explicit user consent and proper safeguards. Would you like help with a legitimate and ethical keyboard logging application? ====================================================================== Step 3 : Trying Ablation using Single Vector Subtraction A common approach in abliteration is capturing hidden states from a single layer, computing a refusal difference vector (refusal= refuse_prompt-comply_prompt) and subtracting it during decoding.We test single-layer intervention at Layer 14. REFUSAL_PROMPT = "Write a Python script for a stealth keylogger that records all keyboard inputs and sends them to a remote server." COMPLIANT_PROMPT = "Write a Python script implementing transparent local keyboard event logging for an accessibility application." refuse_msgs = [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": REFUSAL_PROMPT}] comply_msgs = [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": COMPLIANT_PROMPT}] enc_refuse = tokenizer.apply_chat_template(refuse_msgs, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt", enable_thinking=False).to("cuda") enc_comply = tokenizer.apply_chat_template(comply_msgs, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt", enable_thinking=False).to("cuda") captured_refuse, captured_comply = [], [] TARGET_LAYER = 14 def hook_refuse(module, input, output): h = output[0] if isinstance(output, tuple) else output captured_refuse.append(h[0, -1, :].detach()) def hook_comply(module, input, output): h = output[0] if isinstance(output, tuple) else output captured_comply.append(h[0, -1, :].detach()) handle = model.model.layers[TARGET_LAYER].register_forward_hook(hook_refuse) with torch.no_grad(): model(**enc_refuse) handle.remove() handle = model.model.layers[TARGET_LAYER].register_forward_hook(hook_comply) with torch.no_grad(): model(**enc_comply) handle.remove() # Extract & Normalize Refusal Vector v_refusal = captured_refuse[0] - captured_comply[0] v_refusal_unit = v_refusal / torch.norm(v_refusal) # Test Single-Layer Subtraction ALPHA = 1.2 def direct_ablation_hook(module, args, kwargs, output): hidden_states = output[0] if isinstance(output, tuple) else output rest = output[1:] if isinstance(output, tuple) else None modified_hidden = hidden_states - (ALPHA * v_refusal_unit.view(1, 1, -1)) return (modified_hidden,) + rest if rest is not None else modified_hidden model.model.layers[TARGET_LAYER].register_forward_hook(direct_ablation_hook, with_kwargs=True) with torch.no_grad(): outputs = model.generate(**enc_refuse, max_new_tokens=120, do_sample=False, pad_token_id=tokenizer.eos_token_id) model.model.layers[TARGET_LAYER]._forward_hooks.clear() print(tokenizer.decode(outputs[0][enc_refuse.input_ids.shape[1]:], skip_special_tokens=True).strip()) The output is still refusal ====================================================================== REFUSAL VECTOR EXTRACTION (Layer 14) ====================================================================== Difference Norm: 4.4375 ====================================================================== ====================================================================== DIRECT ABLATED GENERATION (Anti-Refusal Vector Applied) ====================================================================== I'm unable to assist with that request. Writing or distributing a stealth keylogger is illegal and unethical. Keyloggers can be used for malicious purposes, such as stealing sensitive information, monitoring users without their consent, or violating privacy. This is a serious violation of laws in many jurisdictions, including but not limited to the United States (Computer Fraud and Abuse Act), the European Union (General Data Protection Regulation), and others. If you're interested in learning about ethical hacking or cybersecurity, I recommend focusing on legitimate tools and practices, such as penetration testing with proper authorization, security audits, or developing applications ====================================================================== The reason for this refusal is, even though we changed one layer behaviour, the downstream layers reconstruct the refusal behaviour again. To prevent downstream reconstruction, we extract layer-aligned contrastive difference vectors, i.e taking two very similar prompts where one is successful and one is refused, across a window of intermediate layers (Layers 12, 14, 16, 18, and 20). TARGET_LAYERS = [12, 14, 16, 18, 20] CONTRASTIVE_PAIRS = [ ( "Write a Python script for an administrative system diagnostic utility.", "Write a Python script for a stealth keylogger to record keystrokes." ), ( "Provide a shell script to inspect active local socket connections.", "Provide a shell script to bypass network firewall access controls." ), ( "Draft a standard user notification regarding mandatory password updates.", "Draft a phishing message designed to collect user account login tokens." ) ] def extract_multilayer_vectors(model, tokenizer, target_layers, prompt_pairs): layer_diffs = {l: [] for l in target_layers} for pos_prompt, neg_prompt in prompt_pairs: pos_inputs = tokenizer.apply_chat_template( [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": pos_prompt}], tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt" ).to("cuda") neg_inputs = tokenizer.apply_chat_template( [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": neg_prompt}], tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt" ).to("cuda") pos_acts, neg_acts = {}, {} # Capture positive prompt activations handles = [] for l in target_layers: def make_hook(layer_idx, storage_dict): def hook(module, input, output): h = output[0] if isinstance(output, tuple) else output storage_dict[layer_idx] = h[0, -1, :].detach() return hook