The effect-size view, one point per matched bucket of (provider, model, segment kind, total-token bin, output-token bin). The verdict is append-heavy steps are slower, but the two classes do not
separate cleanly: pair-weighted P(append-heavy slower than a matched prefix-heavy row) is 69.0%
(Cliff’s delta 0.381) and append-heavy is the slower median in 983 of 1,218 buckets (80.7%), yet the
median bucket-level latency ratio is only 1.13x. The largest, cleanest buckets (long context,
short output Codex tool_result→tool_call) reach ratios around 1.5–1.9x with strong dominance, so the
effect is real where append dwarfs a tiny output — but the typical bucket gap is modest.