Skip to content
HN On Hacker News ↗

GitHub - MakazhanAlpamys/Soup: Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.

▲ 139 points 28 comments by MakazhanAlpamys 3w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

96 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,480
PEAK AI % 96% · §1
Analyzed
Aug 4
backend: pangram/v3.3
Segments scanned
1 windows
avg 1480 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,480 words · 1 segments analyzed

Human AI-generated
§1 AI · 96%

Fine-tune and post-train LLMs in one command. No SSH, no config hell. Website · Quick Start · Config · Docs · Commands · Models · Discord Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done. pip install "soup-cli[train]" # add [train] to fine-tune; bare `soup-cli` is the light CLI soup init --template chat soup train Why Soup? Training LLMs is still painful. Even experienced teams spend 30-50% of their time fighting infrastructure instead of improving models. Soup fixes that. Zero SSH. Never SSH into a broken GPU box again. One config. A simple YAML file is all you need. Auto everything. Batch size, GPU detection, quantization — handled. Works locally. Train on your own GPU with QLoRA. No cloud required. What's New v0.72.4 — align on a laptop: DPO, ORPO, SimPO and KTO over layer streaming. Layer streaming keeps the frozen base out of VRAM and feeds it to the GPU one decoder layer at a time. It used to support supervised fine-tuning only; now it runs the preference losses too. DPO's reference model is free. DPO needs a reference to compare against, and a second copy of the model would double memory and defeat the whole point. Soup uses the same streamed base with its adapters switched off — one set of weights, one stream. Measured on an RTX 3050 4 GB: streamed DPO peaked at 0.914× the supervised-fine-tuning peak. Forcing a real second model in the same test cost +730 MB — exactly one copy of the weights. KTO is not reference-free, however it is usually described: it picks its reference the same way DPO does, so it gets the same treatment. ORPO and SimPO genuinely are. Bit-exact against a normal, non-streamed run of the same loss — 0.0 difference, the bar every release in this series has to clear. The VRAM pre-flight knows a paired loss is twice the rows, because chosen and rejected go through the model as one tensor. Honest cost: the reference is free in memory, not in time — DPO reads the layer stack 1.52× as often per step as supervised fine-tuning does. grpo / ppo stay excluded on purpose: generation re-reads every layer per token, which is exactly what streaming cannot amortise. Still BETA. # soup.yaml — then just `soup train --config soup.yaml` training: stream_layers: true # base streams out of VRAM; only the adapter trains quantization: 4bit # NF4 — ~4x smaller store, so 8B fits a 4 GB card batch_size: 4 # v0.72.3: bigger batches amortise the weight read stream_source: auto # RAM when it fits, NVMe disk when it does not Trained with stream_layers: true on v0.72.0? That adapter is inert — its tensors were saved under keys with an extra .inner. segment, so every loader returned the untuned base. Fixed in v0.72.1; re-run or re-save. Check with: python -c "from safetensors.torch import load_file; print([k for k in load_file('adapter_model.safetensors') if '.inner.' in k][:3])" Previous release — v0.71.40, soup reward synth (generate a reward verifier from your data) Point soup reward synth at a JSONL of reference outputs and it infers a deterministic verifier, writes a readable / committable .py reward function, and — the part nobody else does — refuses to emit one that can't tell your references from bad answers (four families: numeric / json_schema / regex / tool_call; a mandatory calibration report is the moat). Reward ensembles (reward_fn: "accuracy,format") also train now. (#311) soup reward synth references.jsonl -o reward.py --output-report calib.json Previous release — v0.71.39, CI for weights not prompts (emit + provenance-bind the ship verdict) soup ship's verdict became emittable, committable, and provenance-bound: --emit-evidence makes a run replay into an identical verdict, eval.ship in soup.yaml + --config makes the gate policy reviewable, and --config binds evidence to the exact recipe that produced it (stale evidence → exit 3). soup ship --push owner/repo#N posts the SHIP / DON'T-SHIP card on the PR. Previous release — v0.71.38, The gate grows teeth (real leg-2 regression gate) soup ship's regression leg became real: a fixed, extraction-based scorer over seven bundled, offline suites (MCQ · arithmetic · tool-calling · JSON validity · safety/refusal). A tune that wins your task but quietly breaks tool-calling now gets a DON'T SHIP. Zero new deps. soup ship --base ./base --adapter ./my-lora --task-eval my_task.jsonl # exit 0 = SHIP · 2 = DON'T SHIP · 3 = bad flags · 1 = runtime error Previous release — v0.71.33, soup draft (measure speculative decoding) soup draft measure reports a draft model's acceptance rate + real plain-vs-assisted tok/s (exit 0/2/1 for CI); soup draft distill distils your target into a dense tiny draft, auto-wired into soup serve --auto-spec. The honest result on a small same-family pair: distillation didn't move acceptance (69.3% → 69.3%) and assisted decoding was a net slowdown — which is exactly the number you want before shipping speculative decoding. soup draft measure --target ./my-tuned-model --draft HuggingFaceTB/SmolLM2-135M-Instruct \ --prompts prod-prompts.jsonl # -> acceptance %, real tok/s, ship-or-not Full history: CHANGELOG.md · GitHub Releases. Quick Start 1. Install # Light core: CLI + config + data tools, no PyTorch pip install soup-cli # Add the training stack (torch, transformers, peft, trl, datasets, …) pip install "soup-cli[train]" # Everything (train + serve + ui + data) in one shot pip install "soup-cli[all]" # Or from GitHub (latest dev) pip install git+https://github.com/MakazhanAlpamys/Soup.git The full extras table (fast, mlx, serve, eval, ui, vision, audio, …) lives in docs/models.md. Use double quotes around the extra. They are the only spelling that works in every shell — cmd.exe, PowerShell, bash, and zsh. Older tutorials and videos (including some of ours) show the single-quoted pip install 'soup-cli[train]'. That is bash / zsh / PowerShell syntax, and it fails on Windows cmd.exe, which has no single-quote quoting and hands the quotes straight to pip: ERROR: Invalid requirement: "'soup-cli[train]'": Expected package name at the start of dependency specifier If you hit that, swap the ' for " — pip is rejecting a literal quote character, nothing is wrong with the package. (Dropping the quotes entirely works on Windows too, but zsh then reads [train] as a glob and fails.) soup init, soup data …, and the other data/inspection commands work on the light install. Fine-tuning (soup train) needs the [train] extra. 2. Create a config soup init # interactive wizard soup init --template chat # or start from a template Templates: chat, code, tool-calling, medical, reasoning, vision, kto, orpo, simpo, ipo, bco, rlhf, pretrain, moe, longcontext, embedding, audio. 3. Train, test, ship soup train --config soup.yaml # LoRA, quantization, batching — all handled soup chat --model ./output # talk to your model soup push --model ./output --repo you/my-model soup merge --adapter ./output # merge LoRA into the base soup export --model ./output --format gguf --quant q4_k_m # GGUF for Ollama / llama.cpp More export targets (ONNX, TensorRT, AWQ, GPTQ, BitNet) and deployment options live in docs/serving-and-export.md. Configuration A complete soup.yaml: base: meta-llama/Llama-3.1-8B-Instruct task: sft # backend: unsloth # 2-5x faster, pip install "soup-cli[fast]" data: train: ./data/train.jsonl format: alpaca val_split: 0.1 training: epochs: 3 lr: 2e-5 batch_size: auto lora: r: 64 alpha: 16 quantization: 4bit output: ./output config/schema.py is the single source of truth for every field. Advanced data, training, and PEFT options are documented under Documentation. Documentation The full feature reference lives in docs/. Start here: Guide Covers Training tasks & methods SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, tool-calling, PRM, pre-training, distillation, classification, vision/audio/TTS, unlearning, RAFT/RA-DIT, loop-hardening detectors PEFT, long context & efficiency DoRA, LoRA+, rsLoRA, VeRA, OLoRA, NEFTune, PiSSA, ReLoRA, optimizer & PEFT zoo, LLaMA Pro, GaLore, YaRN/LongLoRA, packing, curriculum, auto-tuning Performance & quantization QAT, FP8, Quant Menu (I + II), KV-cache, NVFP4, save formats, Cut Cross-Entropy, gradient checkpointing, kernels, activation offloading, layer streaming, multi-GPU / DeepSpeed / FSDP Data engineering Formats, the Axolotl/LF-parity pipeline, data tools, synthetic generation & forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs Evaluation & probes Eval design/gate, eval-gated training, benchmarks, NLG metrics, calibration, Elo arena, diagnose, post-train X-ray probes, A/B, drift, tunability, soup advise Serving & export OpenAI-compatible server, batch inference, benchmarking, merge/export, Anthropic Messages endpoint, speculative decoding (train + measure your own draft), deploy autopilot, Web UI, Agent Forge Adapters, registry & governance Adapter lifecycle/management, model registry, Soup Cans, the data flywheel (soup loop), knowledge editing, steering, supply-chain controls (scan/sign/BOM/attest/audit/airgap) Compliance & governance quickstart HIPAA/SOC2/EU-AI-Act/SR-11-7 init templates, provenance (BOM/attest/repro-receipt), audit log, air-gap, model-card autogen (soup card), CI gate (soup ci init) Backends, platform & ops MLX/Unsloth backends, alternative hubs, HF Hub integration, autopilot, experiment tracking, plan/apply, env lockfiles, hardware-fit, completions, plugins, utility commands Command reference The full soup command list Supported models & extras Recommended model families, the VRAM size guide, the pip extras matrix Data Formats All formats are auto-detected from JSONL, JSON, CSV, Parquet, or TXT: alpaca — {"instruction": ..., "input": ..., "output": ...} sharegpt — {"conversations": [{"from": "human", "value": ...}, ...]} chatml — {"messages": [{"role": "user", "content": ...}, ...]}