Hugging Face published TRL v1.15.0 on 8 October 2026 at 19:26 UTC. The tag is not a prerelease. The notes’ main change is a fused language-model head, on by default, for SFT, DPO, KTO, GRPO, RLOO, and Distillation: “there is nothing to turn on.”
What the head was materializing
The LM head is the last layer: one score per vocabulary token. For a batch those scores are a tensor shaped [batch, sequence, vocab], a number at every position for every vocab entry. On a small model with a large vocabulary, that grid is the memory cost. The notes say a Triton kernel projects hidden states through the head in tiles and reduces them to per-token log-probs and entropy, so the full tensor is never built. The gain, they say, is largest for small models with large vocabularies. SFT’s previous default, chunked_nll, already avoided full logits, so that path is unchanged.
Hugging Face’s table
Every figure below is Hugging Face’s, from the release notes. The table header: gemma-3-1b, 262k vocabulary, 79 GiB of GPU memory, bf16, batch size 1 except where noted, gradient checkpointing, sdpa. The details say one B300, allocator capped at 79 GiB (“about an H100 80GB’s usable memory”), google/gemma-3-1b-pt from its config with random bf16 weights. Not measured: H100 hardware itself, flash-attention, real generation, multi-GPU. GRPO and RLOO use fixed random completions, so only scoring is measured.
Max trainable sequence length, v1.14.2 to v1.15.0:
| Trainer | v1.14.2 | v1.15.0 | Change |
|---|---|---|---|
| DPO | 10,240 | 59,392 | 5.80 times |
| KTO (batch size 2) | 9,216 | 63,488 | 6.89 times |
| GRPO (scoring only) | 28,672 | 114,688 | 4.00 times |
| RLOO (scoring only) | 23,552 | 100,352 | 4.26 times |
SFT (loss_type="nll") |
20,480 | 107,520 | 5.25 times |
SFT (default chunked_nll) |
107,520 | 107,520 | 1.00 times |
The notes’ headline says up to 6.9 times longer sequences, matching the KTO row. Distillation uses the fused head and has no row in this table.
At 8,192 tokens the notes say peak memory drops 52% to 82% and steps are 2.3% to 10.9% faster. The table under that sentence prints DPO 48.97 to 12.00 GiB (down 75.5%, 1.035 times the tokens per second), KTO 64.91 to 11.98 (down 81.5%, 1.045 times), GRPO 24.15 to 10.30 (down 57.4%, 1.023 times), RLOO 29.28 to 14.02 (down 52.1%, 1.029 times), and SFT with nll loss 30.33 to 8.87 (down 70.7%, 1.029 times). That tokens-per-second column runs from 1.023 times to 1.045 times. Scoring needs Triton on a GPU: Linux with CUDA, ROCm, or XPU.
Hugging Face’s 9 October repost of a post by Quentin Gallouédec rounds the same work to “up to 82% less peak VRAM” and “7x longer sequences,” with DPO at 10k then 59k tokens and GRPO at 29k then 115k. The release table is finer. Gallouédec’s post was created at 08:02 UTC, and the repost carries that text.
What to check before upgrading
- A PEFT adapter on
lm_headnow raises. Usemodules_to_save=["lm_head"]. use_liger_kernel=Trueis deprecated in DPO, KTO, GRPO, and RLOO, and removed in v2.0.0. Liger’s layer kernels still apply.fused_linear_cross_entropyis forced off, and setting it explicitly raises. Usemodel_init_kwargs={"use_kernels": True}.- In DPO and KTO,
compute_metricsandcompute_loss(..., return_outputs=True)still get full logits from an extra forward, only when used. That path is deprecated and removed in v2.0.0. - The full-logits path, the
use_liger_kernelchunked path, and_forward_redirectionare gone from GRPO, RLOO, DPO, and KTO. - Python 3.10 is dropped.
nn.DataParallelis refused. The experimental MiniLLM trainer is removed. vLLM 0.20.0, 0.20.1, and 0.20.2 are dropped. - Scoring needs Triton on a GPU (CUDA, ROCm, or XPU). The notes do not describe a CPU path. The sequence and memory multiples above are Hugging Face’s B300 numbers.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.