Skip to content
HN On Hacker News ↗

GitHub - volotat/mini-AGI: Continual learning model trained from scratch on 8GB VRAM laptop with batch-1 stream of data.

▲ 277 points • 80 comments • by volotat • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly AI, with some human-written content.

91 %

AI likelihood · overall

AI
6% human-written 94% AI-generated
SEGMENTS · HUMAN 2 of 5
SEGMENTS · AI 2 of 5
WORD COUNT 1,203
PEAK AI % 95% · §4
Analyzed
Sep 21
backend: pangram/v3.3
Segments scanned
5 windows
avg 241 words each
Distribution
6 / 94%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,203 words · 5 segments analyzed

Human AI-generated
§1 Human · 19%

mini-AGI - is a continual learning byte-level language model that assembles its own architecture, trains from scratch on a single 8 GB VRAM GPU, and keeps learning from everything it reads.

§2 AI · 86%

It stores its weights as ordinary files on disk and pages them onto the card as it needs them, so the parameter count is bounded by free disk space rather than by VRAM. It grows new capacity while training when it runs short, prunes what nothing asks for, and reads through exactly the same code path it serves on.

§3 Human · 6%

Targeted at a PC or laptop with at least an 8 GB VRAM GPU on the board. NOTE: as of now this is a small toy-level model. Do not expect a frontier level capabilities. This is rather a small experiment to show, that continual learning from the single stream of data without catastrophic forgetting is possible. Furthermore it is possible on a modest hardware. Which means that almost everyone could train their own version of the model (or simply continue training this one) exactly as they see it fit. And the capabilities would be bounded by the actual hardware, scale and quality of the data available and the amount of time one willing to spend on training the model. Here is how min-run dashboard looks like. The model is pointed to the corpus to constantly read and learn from. History - here is the samples from the whole training run history so far. You can inspect them yourself to see how the model improved over the course of training/reading the corpus. The weights are not published yet. The run is still reading its first pass over the corpus, the weights go up once it has been through all of it, which is a couple of weeks away at the current rate.

§4 AI · 95%

Motivation Every language model you can actually own today is a model somebody else trained and then froze. You can fine-tune around the edges of it, but you cannot train one from scratch on your own hardware, and you cannot keep training it on what you do day to day - the moment you try, it forgets what it knew before. The result is that a personal model is always somebody else's model with a thin layer of you on top, and it stops learning the day it ships. mini-AGI model has small enough GPU footprint that it is possible to train end-to-end on one consumer card, and it is built so that training never has to stop. It reads a stream of characters one chunk at a time, takes a gradient step on each, and the same path serves generation. There is no separate fine-tuning regime and no frozen base: reading and being trained are the same event. Three constraints shape everything else in the design: It has to fit on 8 GB. Not with quantisation - training needs gradients and optimiser state, which is roughly three times the weights again. So the weights live on disk and only the working set is resident. It has to not forget. A model that learns continually and overwrites itself is worse than one that does not learn at all. It has to be able to read anything. The alphabet is the 256 byte values, so there is no tokenizer to fit and no data type that needs a new vocabulary. The model is genuinely yours: trained on your hardware, on your data, that keeps learning from every conversation you have with it, and that nobody else can take it away or switch it off. How the architecture works Characters (bytes) does not pass through a fixed stack of layers as it would be in a traditional LLM. Instead, it passes through two dense prelude blocks and then through one recurrent block applied up to 24 times, each application choosing its own experts from a shared pool. The latent state between applications is never decoded - it is merged with the embedded input by an adapter each time round, so the loop cannot drift away from the text it is reading. Three distinct blocks, up to 26 block-applications per character. Adaptive depth. A halting head scores every character at every row, and the character stops as soon as another row would not change the answer. Easy characters take one row, hard ones take many. This is the PonderNet recipe: while training, every depth is computed and weighted by its halting probability, so the halting head learns through those weights. Routing per block-application, not per character. Each of the 26 applications picks its own top-8 experts, so one character touches far more of the pool than "top-8" suggests, and the same expert can be selected several times at different depths. What varies is which eight at each point. No expert is assigned a subject. There are no labels anywhere. Soft top-k routing distributes capability across the pool by itself, and a character can combine fragments from several experts. The cost is that capabilities share parameters and so can interfere. This is the architecture assembling itself, one character at a time, captured from the live model - nothing here is drawn by hand. Each tile on the left is one expert; colour is expert identity and stays the same for the whole clip. A row is one application of the recurrent block, and the eight tiles in it are the eight experts that row actually ran. The stack grows downward as the model keeps going, and the amber line is where halting stopped it - the grey rows below are computation the model declined to spend. The trace on the right is how many rows each character took. It moves constantly between 4 and 14 against a ceiling of 24, and the caret under the text shows which character is being read. Positions are rotary and carry no learned parameters, which is why the context window can be extended by continued training rather than by re-initialising anything. ...and the same thing while it writes The clip above is the model reading - every character is held-out text it is being shown. This one is the model writing: it was primed with 2,500 characters of a held-out story and then continued on its own, so the grey text is what it was given and the green text is entirely its own. Greedy decoding, no sampling anywhere - run it twice and you get the same sentence. Two things are worth watching. The stack behaves the same way, because generating and reading are the same forward pass in this model - the only difference is whether the next character comes from a file or from the model's own argmax. And writing costs more depth than reading: about 9.9 rows a character against 8.0 on the same subject. The dotted lines mark where the working set was re-chosen, which happens every 64 characters; in this clip nothing swapped, because the prompt had already pulled the right experts onto the card.

§5 Mixed · 34%

What it produced, continuing a story about a cherry tree: They worked together and saw their favorite shore. One day, they wanted to play with their favorite shore. They wanted to play with it, but Grammatically correct and on-topic.