The experiments evaluate the DNO algorithm (specifically, DNO-Prct) using an iterative training process that combines GPT-4-Turbo scores with curated pairwise comparisons. UltraFeedback forms the core dataset, with additional large-scale trials. Evaluation is performed using AlpacaEval 2.0, MT-Bench, and the OpenLLM Leaderboard. The results highlight how DNO approaches state-of-the-art performance through efficient and scalable preference modeling.




