Skip to content

Train your own

Your conversations → your private dataset → one GPU job → a model that is yours, back on the phone. The app does the plumbing; Hugging Face does the compute and bills your account per minute.

Models › Train your own. Needs a Hugging Face write token in Settings › Hugging Face and at least one dataset you shared from History.

The confirmation card names the exact POST, the GPU, the cap and the estimated cost

Nothing starts before you confirm the exact job and its cost.

Running: step 14 of 24, loss, dollars spent so far, the stage ladder

Live: step, loss, dollars so far, each stage with its time.

Done: held-out loss before and after, tool-call accuracy, add to Models

Held-out loss before → after. One tap adds the 4-bit model.

What happens

  1. Dataset — pick one you shared (History › Share to Hugging Face). The app reads the shards and counts the conversations.
  2. Recipe — base Qwen3.5 0.8B / 2B / 4B (suggested by dataset size: < 300 → 0.8B, ≤ 2 000 → 2B), LoRA or full fine-tune, GPU, epochs, and a hard stop in hours — Hugging Face kills the job there, so the cap is also the most it can cost.
  3. Estimate — steps × measured seconds-per-step on that GPU, plus install, scoring and the two conversions. Honest, not a quote; it leans long.
  4. Confirm — the same approval card the agent shows for an http call: POST huggingface.co/api/jobs/<you>, the flavor, the cap, the script arguments. Deny and nothing is sent.
  5. Job — the app commits train/loopl_sft.py into your dataset repo and launches it with uv run. The script trains, merges, scores on held-out conversations, pushes <you>/<name> (weights), -4bit (MLX, for the phone) and -GGUF — all private.
  6. Result — eval loss before → after, how often the tuned model picked the right tool, and Add to Models. The repos also appear under Models › Mine.

Keep going — rounds, epochs, stitching

A tuned model is a base too. Three controls on the Recipe card make training a thing you keep doing rather than do once:

  • Continue a model. Pick one of your Mine models and the next job trains on top of it — round 2, round 3… The badge tells you the round; the model card carries the whole lineage (Qwen3.5-0.8B → loopl-app-sim → loopl-app-sim-r2). The script trains from the bf16 twin of the 4-bit model you run on the phone — training from 4-bit weights is refused, and a model without a bf16 twin is named as "can't continue". Round ≥ 2 defaults to LoRA (r32, lr 5e‑5) and replays every shard of the dataset so round 1 is not forgotten; turn Include all previous shards off to train on the new ones only.
  • Epochs: Auto (recommended). Up to 3 passes, scored on held-out conversations after each; the checkpoint with the lowest loss is kept and training stops the moment loss rises. The guide the footnote shows is the one we measured: under 200 conversations 3–4 · 200–1 000 → 2 · over 1 000 → 1–2. Pick 1–4 by hand when you know better. The result card draws held-out loss per epoch and marks the epoch that was kept.
  • Stitch related conversations. By day joins the chats you shared on the same day — oldest first — into longer multi-turn samples (several per day when a day is long), so the model learns to carry context across chats; By topic joins chats about the same thing. Tool call / result pairs always travel together; the originals train too. The footnote previews how many stitched samples your dataset yields.

Measured, round 2 from the app: base loopl-app-sim (itself trained from the app), 74 conversations stitched by day, epochs auto → 3 trained, best epoch 2 — A10G, 334 s, ≈ $0.10, held-out loss 0.58 → 0.22, tool calls 12/12, vision kept. The result lists in Mine as round 2 and can be continued again.

Cost guide

base GPU 200 conversations 1 000 5 000
0.8B LoRA A10G $1.00/h ≈ $0.25 · 13 min ≈ $0.50 · 30 min ≈ $1.90 · 1.9 h
2B full A100 $2.50/h ≈ $0.80 · 20 min ≈ $2.30 · 55 min ≈ $10 · 4 h
4B LoRA A100 $2.50/h ≈ $1.00 · 25 min ≈ $2.80 · 1.1 h ≈ $12 · 5 h

Measured: 191 conversations, 0.8B LoRA, 2 epochs on an A10G — 8 min, $0.13, held-out loss 2.38 → 2.15. With pictures: 74 conversations (14 with images), 0.8B LoRA on an A10G — 6 min, ≈ $0.10, held-out loss 0.75 → 0.29; the result shows up in Mine with its vision badge and describes photos on the phone.

Vision is kept, not tuned. The export keeps the base model's vision tower, so a model trained from conversations with pictures still sees; the training itself learns from the text of those turns.

For SDK users

TrainingPlan → TrainingEstimate; TrainingLaunch.jobSpec builds the job body; TrainingSession uploads the script, creates the job through HFHubClient and polls TrainingProgress (the script's [loopl] stage=… lines). The token travels only as the job's HF_TOKEN secret. The training script lives in loopl-train (train/loopl_sft.py ≥ 1.1.0 — the MLX export keeps the vision tower), bundled with the app byte-for-byte.