ICML2026
Capability Traps in DPO
Marco Pollanen
Abstract
Direct Preference Optimization (DPO) is often tuned by selecting checkpoints with high preference margins. We show that this selection rule can fail. Across dense sweeps in three 7B open-weight families under controlled DPO recipes, the DPO preference margin can strongly anticorrelate with capability probes, with the strongest observed case in Llama logic probes (Pearson , ). In Mistral, matched-duration controls show that transient exposure to higher produces persistent arithmetic (, ) and format (, ) degradation relative to a duration-matched constant- run, while logic and sycophancy show no significant effect. An expanded 14-probe logic suite further shows that aggregate capability scores can hide stable opposition between probe clusters: direct-inference probes are consistently negative while fallacy-detection probes are consistently positive, with 13 of 14 probes sign-stable across 3 seeds and 3 values. Sensitivity profiles differ across architectures: Mistral shows elevated seed variance near , Llama shows capability-specific rigidity, and Qwen trades off smoothly. These findings motivate probe-resolved sweeps rather than margin-based checkpoint selection.