Skip to content

Archive

Training Objectives

1 articles
Artificial Intelligence 24 Sep 2026 5 min read

Multi-Token Prediction Adds Parallel Future-Token Losses to a Shared Model Trunk

A next-token language model normally applies one predictive objective at position t: the hidden representation at that position is used to score token x_(t+1). Multi-token prediction changes that training boundary. One shared model trunk produces the representation, while multiple output heads predict several subsequent tokens from that position. The mechanism adds supervision at multiple future offsets without requiring a separate transformer trunk for every offset. It is therefore a change to the training objective and prediction heads, not a claim that ordinary autoregressive generation can emit several unchecked tokens as one exact step.