A frontier release landed in the middle of a live engagement last month. The client runs a document-heavy intake loop: claims arrive, get classified, extracted, and routed, with exceptions queued for a human. The new model was measurably better at the extraction step, which is the step that drives everything downstream.
Here is the entire account of the migration. An engineer pointed the capability boundary at the new model in a staging copy of the loop. The standing evaluation suite replayed three weeks of production traffic against it overnight: extraction accuracy up, exception rate down about a fifth, no contract violations. The next morning the configuration change went to production. Elapsed human effort, roughly forty minutes. Nobody scheduled a meeting. The Tuesday model became the Wednesday model.
The same month, a firm we assessed in the spring scoped the identical upgrade at two quarters. Not because their engineers are worse. Because the old model's output dialect had leaked into eleven downstream consumers, none of them under contract, and every one of them a place where last year's frontier had been quietly hard-coded.
Frontier capability arrives on roughly the same schedule for everyone. What differs is the delivery dock. A loop that speaks to capability through a versioned contract, and re-evaluates continuously against live outcomes, takes delivery of each release as routine input. A loop that lets one model's dialect soak into its consumers pays a re-integration project per release, and the releases are speeding up. Both firms will run the same model eventually. Only one of them was designed to receive it.
END