In this section, we provide detailed theoretical proofs that support the Direct Nash Optimization (DNO) framework. The proof of Theorem 2 involves a two-step procedure, starting with a regression using logarithmic loss and leading to a squared error limit. The definitions and assumptions rely heavily on the concentrability of reinforcement learning theory (specifically in the works of Xie et al., 2021, 2023). Although the section simplifies some concepts for clarity, a complete theoretical analysis is beyond the scope of the paper. The proofs also make use of standard results from regression theory, with additional references provided for deeper understanding.





