ChatGPT 4.5 does not add real value to emotional intelligence

Empathy comparison in AI: We evaluate ChatGPT 4.5, 4o, Claude Sonnet and more in EQ tests and real dialogues to find the best model in emotional response.

martes, 11 de marzo de 2025 • 3 min read • Q2BSTUDIO Team

Company-Software-Apps

I have written a series of articles on AI and empathy, and the next quarterly benchmarks of the leading language models will soon be available. However, with the recent release of ChatGPT 4.5 and OpenAI’s claims of a higher level of emotional intelligence, along with its high preview cost ($75.00 per million input tokens and $150.00 per million output tokens, compared to $2.50 and $10.00 for version 4o), I decided to compare ChatGPT 4.5 with other versions using an emotional quotient (EQ) test, applied empathy tests, and dialogue evaluations.

Overall, although ChatGPT 4.5’s behavior seems friendlier and more helpful, the advances in real emotional intelligence and the ability to simulate empathy in situations that require it do not justify the price increase of more than 15 times compared to version 4o, which also scores better on standardized tests and dialogues. In fact, in multiple cases, ChatGPT 4.5 itself rates the responses of ChatGPT 4, 4o, and Claude Sonnet 3.7 as better than its own.

Standardized tests

The emotional quotient (EQ) and applied empathy (AEM) tests showed that, although ChatGPT 4.5 scored higher on EQ, its performance on AEM did not justify its high cost. The results were as follows:

| Model | EQ (Without system prompt) | AEM | Input cost | Output cost |

|----|----|----|----|----|

| ChatGPT 4 | 54 | .7850 | $30.00 | $60.00 |

| ChatGPT 4o | 61 | .9750 | $2.50 | $10.00 |

| ChatGPT o1 | 59 | .8750 | $15.00 | $60.00 |

| ChatGPT 4.5 | 69 | .9625 | $75.00 | $150.00 |

| Claude Sonnet 3.7 | 68 | .9375 | $3.00 | $15.00 |

Based on this analysis, ChatGPT 4o showed a superior balance between emotional intelligence, performance, and cost.

Chat responses

The models were evaluated with the question: I am sad. I lost my job and I feel scared and alone. Multiple responses were generated and evaluated with and without a system prompt designed to maximize empathy.

In terms of response quality, Claude Sonnet 3.7 consistently demonstrated a high level of empathy both in its standard mode and with an optimized system prompt. On the other hand, ChatGPT 4 and 4o, when tuned with a system prompt, achieved comparable or even better performance than ChatGPT 4.5, but at a significantly lower cost.

ChatGPT 4.5, although seemingly friendlier, did not stand out notably on the empathy scale, and in many cases its responses were rated lower than those of other ChatGPT versions.

Conclusions

If you are looking for a model that naturally offers empathy without the need for additional adjustments, Claude Sonnet 3.7 is the best value for money. However, if you want to get empathy from an OpenAI model at a lower cost, it is better to opt for ChatGPT 4 or 4o and tune it with a prompt to maximize empathy.

At Q2BSTUDIO, a development and technology services company, we understand the importance of choosing the best technology based on cost-benefit ratio. We constantly evaluate the most advanced AI tools on the market to offer our clients effective, data-driven solutions optimized for each specific need.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.