Skip to content
HN On Hacker News ↗

GitHub - Sparticle62ops/pssa: A custom AI architecture being developed in rust

▲ 90 points • 38 comments • by sparticle62 • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

99 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,643
PEAK AI % 99% · §1
Analyzed
Sep 30
backend: pangram/v3.3
Segments scanned
1 windows
avg 1643 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,643 words · 1 segments analyzed

Human AI-generated
§1 AI · 99%

PSSA: a plastic state-space architecture PSSA is a small language model that is not a transformer. It reads text one token at a time through a recurrent state-space layer, keeps a bank of episodic memories it can look things up in, and rewrites part of its own weights while it runs. It is written in Rust from scratch, with no PyTorch, no TensorFlow, and no ML framework of any kind underneath it. At matched parameters and on the same corpus, it learns faster than a transformer and generates text about twelve times quicker on the same CPU. Why Rust, and why that is not the point Not for speed points, and not because the language makes the architecture better. PSSA needed per-token weight updates, a memory bank written during the forward pass, and a scalar reference path that every batched kernel could be differentiated against. Expressing that inside an autograd framework meant fighting the framework at every step, so the linear algebra is written directly instead. That made the plastic parts straightforward and the gradients checkable against a reference to around 3e-8. The architecture is the claim here. The implementation language is a detail, and a Python port is welcome. How it differs from a transformer A transformer scores every pair of tokens in the context, so its cost per step grows with the square of the sequence length and the whole context is re-read at every step. PSSA carries one fixed-size state along the sequence in a single left-to-right pass, and looks things up in a memory bank instead of re-reading the context, so cost grows linearly with length. The model Every token goes through one PSSA layer: a selective state-space recurrence, a bounded read from an episodic memory bank in hyperbolic space, a learned gate that decides how much of that read reaches the residual stream, and a SiLU MLP. The defaults are d_m = 256 channels, d_s = 16 states per channel, and a rank-16 adapter. The recurrence Write x for the layer-normalized token embedding. Three projections are read off the token itself, which is what makes the recurrence selective rather than fixed: delta = softplus(W_delta x) per-channel step size, delta in R^d_m B = W_B x input map, B in R^d_s C = W_C x output map, C in R^d_s The transition is diagonal, one rate per (channel, state) pair, kept negative by construction so the recurrence cannot blow up: A = -softplus(A_raw) A in R^(d_m x d_s) Discretizing that continuous system with step delta gives the per-token update. h carries across tokens and across chunk boundaries during training: Abar_ij = exp(delta_i * A_ij) Bbar_ij = delta_i * B_j h_ij <- Abar_ij * h_ij + Bbar_ij * x_i y_i = sum_j C_j * h_ij A_raw is initialized so each channel's 16 rates sit on log-spaced timescales tau from 1.5 to 200 tokens, in the spirit of the HiPPO initialization. A single channel therefore starts out holding the last two tokens and the last two hundred at the same time, and training moves those horizons rather than discovering them from scratch. This half of the layer is a selective diagonal SSM and claims no novelty; it is the same family as S4 and Mamba, written out scalar-first so the backward pass can be checked term by term. The memory read The part that is specific to PSSA is what happens to y. A query is formed from both the current token and the current state, so retrieval is conditioned on where the recurrence has got to and not only on the token in hand: q = W_qx x + W_qh y qh = proj(q) diffeomorphic map into the Poincare ball, |qh| < 1 The read is bounded at four slots, weighted by a softmax over hyperbolic distance at temperature tau_mem: w = softmax(-d_H(qh, k_s) / tau_mem) over the 4 nearest slots m = sum_k w_k * v_k Hyperbolic distance grows toward the boundary of the ball, so slots holding general context and slots holding one specific episode stay separable without widening the read. Four slots is a fixed cost per token regardless of how much the bank holds. Gate, adapter, MLP The read does not join the stream unconditionally. A learned per-channel gate decides how much of it lands, alongside a low-rank SiLU adapter that carries targeted updates: g = sigmoid(W_gate x) z = s * y + g (elementwise) W_proj m + adapter(x) u = W_2 silu(W_1 z) z_out = z + u The write path Writes are the reason the architecture is called plastic. A slot is inserted when the incoming state is novel against what the bank already holds, each slot carries a refractory counter that rate-limits how often it can be overwritten, and fast plastic updates are folded back into the base transition matrix by a closed-form ridge regression rather than living in the external store forever: A_base <- A_base + (H^T H + lambda I)^-1 H^T dH The refractory counter is what keeps a stream of contradictory updates from erasing a slot that repeated evidence has already stabilized, and consolidation is what stops the bank from being the only place long-range structure is stored. What is and is not new here The recurrence is standard selective-SSM machinery. The claims are the hyperbolic bounded read conditioned on the recurrent state, the novelty and refractory rules on writes, and the ridge consolidation step from fast weights into the transition matrix. Everything is implemented against a scalar reference path that the batched and parallel implementations are differentiated against on every commit, currently agreeing to a maximum gradient error around 3e-8 (cargo run --release --example twin_check). The result Two models, same corpus, same tokenizer, same optimizer schedule, same seed, same number of parameters. One is PSSA, one is a standard transformer. Over 12.7M tokens of cleaned WikiText-103: PSSA finished at 3.98 training cross-entropy, the transformer at 4.43. That is a gap of 0.45 nats, perplexity 53.7 against 83.7. The transformer spent its entire 12.7M-token budget to reach a loss PSSA had already passed around 2M tokens in. The two curves never cross, and they never touch: It holds on text neither model has seen Training loss only says a model fit the stream it was fed. So both checkpoints were scored on a 198,939-token slice cut from a part of the corpus neither run ever touched: Every checkpoint of both runs, 64 PSSA links and 43 transformer links, scored on a bounded 9,934-token window of that unseen slice. The curves never cross: PSSA is ahead from the first link and finishes 0.51 nats lower. The table below is the final checkpoint of each run on the full slice. Held-out slice, 198,939 unseen tokens PSSA Transformer Cross-entropy 3.997 4.429 Perplexity 54.4 83.8 Next-token accuracy 24.1% 18.0% The held-out gap, 0.43 nats, is essentially the training gap. PSSA is not memorizing harder, it is generalizing better. And it is much faster to run Generating 200 tokens on the same CPU, same prompt, same sampler: PSSA Transformer 200 tokens 226 ms 2,735 ms Relative 12x faster baseline A recurrent model carries a fixed-size state, so the cost of each new token does not grow with the length of what came before. A transformer re-reads its whole context every step. What is actually different about it A recurrent state-space core. Learned continuous state matrices carry information forward in a fixed-size state, instead of attention over the full context window. An episodic memory bank. 512 slots with hyperbolic (Poincare-style) retrieval and bounded top-4 search, written to and read from during the run. Plastic weights. Fast updates reinforce what works, novelty drives growth, and a refractory gate rate-limits overwrites so repeated contradictory input does less damage. Closed-form consolidation. A ridge-regression step folds the fast plastic updates back into the base transition matrix, the way sleep consolidates a day's learning. No framework. Hand-written linear algebra in Rust, with a CUDA path for training and a scalar CPU reference that every gradient is checked against (max gradient difference 2.98e-8). What this is not Being straight about the scale, because the numbers above are easy to over-read: These are 1.5M-parameter models on 12.7M tokens. That is a research prototype, not a competitor to anything you have heard of. Text quality at this scale is poor for both models. PSSA emits "a barget of the Prian Academy", the transformer "a material circulation of the United States". The comparison is about learning efficiency, not fluency. The speed comparison is CPU-to-CPU, which is fair. The training throughput numbers further down are not hardware-matched and should not be read as an architecture result. Two experiments are still unmeasured: retention of earlier skills after a corpus switch, and whether ablating the memory bank changes the loss. Try it git clone https://github.com/Sparticle62ops/pssa.git cd pssa cargo build --release ./target/release/oxide_ai_pssa Running it with no arguments gives you a home screen listing every command plus any checkpoint and corpus it finds in the working directory. Where the project needs help Compute The whole result above was trained on a free hosted notebook with a single entry-level GPU, in 200,000-token links, because a session gets cut after a few hours. Every interesting question left, whether the gap holds at 10x or 100x these parameters, whether the memory bank matters at scale, how it does against a modern recurrent baseline, needs one thing: a GPU with real VRAM and allocations measured in days instead of hours. Anything meaningfully above the entry-level card this ran on changes what can be asked. If you have compute to grant, or you work somewhere that does, that is the single highest-leverage thing anyone can offer this project. Sponsorship Sponsorship funds compute and nothing else. In return you get named here and in the write-up of any result your hardware made possible. Get in touch before sending anything so the details can be agreed.