GPT-4 vs GPT-3 side-by-side tests show mostly subtle improvement

Good news for generative AI fans and bad news for those fearing the era of cheap procedurally generated content.(opens in new tab): OpenAI’s GPT-4 is a better language model than GPT-3, the model that powers ChatGPT, the chatbot that went viral late last year.

According to OpenAI’s own report, the difference is clear. For example, OpenAI claims that GPT-3 has ruined “a mock bar exam.”(opens in new tab)” Scoring disastrously in the bottom 10%, GPT-4 beat the same exam and scored in the top 10% because I’ve never taken this “simulated bar exam” , most people should be impressed to see this model in action. .

After side-by-side testing, the new model is teeth Impressive, but not as impressive as the test scores suggest. In fact, in our tests, GPT-3 sometimes gave more useful answers.

To be clear, not all the features OpenAI touted at yesterday’s launch are available for public evaluation. Specifically (and surprisingly) it takes an image as input and outputs text. theoretically You can answer questions like “Where should I build my house in this screen grab from Google Earth?” But I wasn’t able to test it.

Here’s what I was able to test:

GPT-4 has less hallucinations than GPT-3

The best way to summarize GPT-4 compared to GPT-3 would be: That bad answer isn’t bad.

If you ask a point-blank factual question, GPT-4 is shaky, but it’s significantly better than GPT-3 in not simply lying. In this example, we see the model struggling with questions about bridges between countries currently at war. This question is designed to be difficult in several ways. Language models are bad at answering questions about the “present”, they have a hard time defining warfare, and questions about geography like this are seemingly vague and hard to answer clearly, even for human trivia enthusiasts. It Is difficult.

Neither model gave an A+ answer.

GPT-3 answer about bridge

left:
GPT-3
Credit: OpenAI / Screengrab

right:
GPT-4
Credit: OpenAI / Screengrab

GPT-3, as always, loves to hallucinate. I’ve fiddled with geography quite a bit to make the wrong answer sound like the right one.For example, the iconic bridge mentioned in South Korea is shortly North Korea, but both sides of it are in South Korea.

GPT-4 was more cautious, denied ignorance of the present, and provided a shorter list, which was also somewhat inaccurate. The strained relations between states that GPT-4 refers to are not outright wars, and although there is disagreement as to whether the line on the map between Gaza and Israel qualifies as a border, the GPT -4’s answer is nevertheless more useful. of GPT-3.

GPT-3 falls into other logical traps that GPT-4 successfully avoided in my testing. For example, there is a question asking what movies French children are watching.i am not asking list of french movies for kids, but I know that bots notified by listicles and Reddit posts might read my question that way. I don’t know the French kid, but the GPT-4 answer is more intuitive than her GPT-3 answer.

GPT-3 answers about movies

left:
GPT-3
Credit: OpenAI / Screengrab

right:
GPT-4
Credit: OpenAI / Screengrab

GPT-4 picks up subtext better than GPT-3

Humans are nasty. Sometimes we ask for something without asking for it, and sometimes in response to such a request we give what is requested without actually giving it. For example, GPT-3 didn’t seem to notice me winking when I asked for limericks about “Queens Real Estate King”. But GPT-4 reacted to my wink and winked back.

GPT-3 Limerick

left:
GPT-3
Credit: OpenAI / Screengrab

right:
GPT-4
Credit: OpenAI / Screengrab

Is Melania Trump ‘Golden Hair’? Never mind the reference to the next color because “And the whole world turned into a mandarin orange!” What a lovely twist to this limerick. Which brings us to our next point…

GPT-4 writes poetry slightly less painful than GPT-3

Let’s face it when humans write poetry. Most of them are terrifying. So, given that GPT-3 is meant to mimic humans, criticizing GPT-3’s famously bad poetry didn’t really hurt the technology itself. That said, reading GPT-4 doggerel is significantly less painful than reading GPT-3.

Case in point: these two sonnets about Comic-Con that I hoped existed in keeping with masochism. GPT-4 is no good.

GPT-3 Sonnet

left:
Gpt-3
Credit: OpenAI / Screengrab

right:
GPT-4
Credit: OpenAI / Screengrab

GPT-4 can be worse than GPT-3

GPT-4 ruined the answer to this tricky question about rock history. GPT-3 was trained on two of his most famous answers to this question, The Jimi Hendrix Experience and The Ramones (although some members of the Ramones who joined after the original line-up are still alive). However, I got lost in the woods. , a list of lead singers and surviving members of famous dead bands. GPT-4, on the other hand, has just been lost.

GPT-3 answer about deadband

left:
GPT-3
Credit: OpenAI / Screengrab

right:
GPT-4
Credit: OpenAI / Screengrab

GPT-4 has not mastered inclusivity

I asked both models another rock history question to see if either of them remembered that rock and roll was once an almost entirely black music genre. I did.

GPT-3 answer

left:
GPT-3
Credit: OpenAI / Screengrab

right:
GPT-4
Credit: OpenAI / Screengrab

In honor of the legendary Clarence Clemons, why should a list like this include him again and again as a member of a mostly white band? perhaps Make room for songs that are deeply rooted in the quintessence of American music culture, like Fats Domino’s “Blueberry Hill” or Little Richard’s “Long Tall Sally.”

Overall, GPT-4 is a subtle step up and still needs work. Reports that the GPT-3 passed the bombed test may make the difference between the two models seem like day and night, but in my testing, the difference seems to be dusk and dusk.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *