This article presents the Batched Prompting approach used to efficiently label preferences with GPT-4, where all candidate responses are evaluated in a single batch rather than individually. The cost analysis of scaling the experiment to 600k training entries reveals that sampling was the most time-consuming and costly stage, with an approximate cost of $6,000 per iteration. Annotation was the highest cost, amounting to $34,000 per iteration due to token volumes. Training was relatively economical, requiring only 12-24 hours.



