In this section, we present detailed theoretical proofs supporting the Direct Nash Optimization (DNO) framework. The proof of Theorem 2 involves a two-step procedure, starting with regression using logarithmic loss and leading to a squared error limit. The definitions and assumptions rely heavily on the concentrability of reinforcement learning theory (specifically in the works of Xie et al., 2021, 2023). While the section simplifies some concepts for clarity, a comprehensive theoretical analysis is beyond the scope of the paper. The proofs also leverage standard results from regression theory, with additional references provided for deeper understanding.





