Tasks and evaluation interfaces [ftip-001M]

A model does not have a benchmark score without a task distribution, an inference procedure, and an evaluator. Keeping these objects separate is essential for post-training: a training change and an evaluation-time search change can raise the same reported score while answering different research questions.