Systems that improve on purpose, not by luck
Training can produce a better system without anyone being able to say why, or repeat it. This planned paper is about the reproducibility of improvement itself: whether a gain came from the change we made, and whether making that change again would produce it again.
A training run that ends better than it started is not, by itself, progress you can build on. If you cannot say which change caused the improvement, and cannot get it back by making that change again, then what you have is a fortunate outcome rather than a method. The difference only becomes visible when you try to do it twice.
We are interested in treating improvement as something that should reproduce, the way a result in any other discipline is expected to. That means separating the effect of a deliberate change from the noise of initialisation, ordering, and the many small choices that go unrecorded. It is slower than chasing the best single run, and we think it is the only way the gains compound instead of evaporating.
What we intend to measure
- When a system improves, how do we attribute the gain to the change we made rather than to chance in the run?
- If we repeat the same change from a fresh start, how consistently does the improvement come back?
- How much of what looks like method is really variance we have not accounted for?
- What would have to be recorded for someone else to reproduce an improvement without guessing?
Open questions
- Can reproducibility of improvement be made routine without making experimentation prohibitively slow?
- What is a fair standard for calling a training change real, as opposed to a run that happened to land well?