Skip to content
HN On Hacker News ↗

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

▲ 158 points • 43 comments • by Betelbuddy • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

90 %

AI likelihood · overall

AI
15% human-written 85% AI-generated
SEGMENTS · HUMAN 0 of 2
SEGMENTS · AI 1 of 2
WORD COUNT 355
PEAK AI % 97% · §1
Analyzed
Sep 17
backend: pangram/v3.3
Segments scanned
2 windows
avg 178 words each
Distribution
15 / 85%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 355 words · 2 segments analyzed

Human AI-generated
§1 AI · 97%

View PDF HTML (experimental) Abstract:The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only a fraction of an enormous stored parameter bank for each token. That success is built on static pretraining data. A deployed model faces a different world, where much of the data that would make it more useful is not in its training set but in the live interaction it is currently handling, such as the facts a user supplies or the corrections they give. A conventional model cannot learn from this data, because its weights are frozen after training. Instead, the knowledge and behaviour supplied at run time are placed in the prompt, by retrieval or instruction, and re-read on every request only to be discarded once the request ends. We ask how an architecture could learn from live interaction by writing it into its weights. Taking inspiration from MoE, we propose the \textbf{Infinite-Parameter LLM}. A compact hypernetwork turns the data given at run time into a low-rank modulation of a shared base network, so the feed-forward weights are generated from live data rather than stored in a fixed bank. Where prior weight generators read the context once and freeze, we carry a Bayesian belief over the generator's latent code and update it online, so the effective weight is re-derived from that evolving belief as the session proceeds rather than fixed after one read. The stored footprint stays fixed, yet the weights the model can compile are effectively infinite. For the knowledge and behaviour supplied at run time, carrying them in the weights rather than the prompt is amortized in compute, frees the context window, persists across turns, and can generalise better than in-context use. We specify an evaluation protocol that tests exactly this against in-context learning and retrieval.

§2 Mixed · 47%

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2609.18842 [cs.AI] (or arXiv:2609.18842v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.18842 arXiv-issued DOI via DataCite (pending registration) Submission history From: Jinli Hu Dr [view email] [v1] Wed, 16 Sep 2026 15:49:34 UTC (50 KB)