Post-Training Lens Reframes Capabilities
Visual status: no verified article image is available. The reporting remains text-first.
A new theory published on arXiv reframes how we think about post-training for large language models, insisting there are two separate outcomes: elicitation and creation. The authors argue that the key question is not whether a training step is imitation or discovery, but whether it reweights behaviors the model could already reach or actually expands its reachable abilities. In practical terms, this changes how teams should evaluate SFT and RL in the wild. The paper formalizes this idea with a free energy perspective and a concept called accessible support, the set of behaviors a model can actually produce under finite budgets.
The core distinction rests on four words researchers say matter most in post-training. If updates stay near the pretraining distribution and simply tilt probabilities toward desirable outputs, they are capability elicitation, not new ground. If updates broaden what the model can practically do, via search, tool use, or new information, the authors call that capability creation. The framework unifies supervised fine tuning and reinforcement learning as two flavors of reweighting a reference distribution, with the difference lying in whether the update expands or just shifts the existing behavioral landscape.
For practitioners, the framework translates into concrete evaluation questions. How much of the improvement is just a reweighting within the model’s current reach, and how much is actually enabling the model to access new behaviors under realistic compute budgets? The paper uses the term accessible support to capture this nuance, emphasizing that the same algorithm can appear to improve capabilities while only nudging the same set of behaviors, or, alternatively, unlocks new ones when external signals or tools push the model into previously unreachable territory. This reframing matters for product bets and claims about performance gains.
From a product standpoint, the lens suggests a sharper focus on how teams deploy post-training. If most gains are elicitation, teams should temper expectations about radical new behaviors emerging from SFT or RL alone and instead invest in scaffolds that widen the practical reach, such as retrieval, tools, or dynamic information integration. If a project genuinely expands accessible behavior, the payoff can be substantial, but it also raises questions about compute and data requirements, because expanding the support is typically more demanding than reweighting. The paper’s free-energy view makes these tradeoffs explicit.
A vivid analogy helps: think of a model as a musician with a set of songs they can play now. Elicitation is like fine tuning the tempo or emphasis on existing tunes, while creation is teaching the musician new genres and instruments they could actually perform given new training and tools. The framework says you can get impressive shade and nuance by nudging the same repertoire, but real leaps come when the musician gains access to new pieces and orchestral support that broaden what concerts are even possible.
What to watch next for developers shipping this quarter? First, ensure that claimed improvements are measured against the model’s existing reachable set, not just a different failure mode or synthetic benchmark. Second, quantify whether new capabilities rely on external signals like retrieval or tools or hinge on a broader information flow that increases the behavioral space. Third, prepare to articulate compute and data budgets in terms of accessible support rather than raw metric gains, so stakeholders understand where growth actually comes from. The authors’ framework puts these questions front and center, moving the field toward clearer, more honest reporting of what post-training actually accomplishes.
Limitations remain. The proposal is theoretical and calls for practical methods to measure accessible support in real systems, something not yet standardized at scale. Without concrete benchmarks, teams may struggle to separate elegant theory from actionable engineering. Still, the guiding insight is robust: distinguish gains that reweight the familiar from those that genuinely extend a model’s reach, and design post-training programs with that split in mind. The rest will follow, as practitioners align incentives, metrics, and tooling around this clearer map of capability growth.
- On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspectivearxiv.org / Primary source / Published MAY 12, 2026 / Accessed MAY 13, 2026