Test Scope
This test can target a specific and serious model downgrading happening recently. It can’t detect minor degradation or quantization.
Test Cases
Thinking budge: Starting from Low, if it fails, you can move to high.
If you really suffer from this kind of model degradation, even ultra won’t save you.
1, Create an HTML file containing a 2D SVG animation of a pelican riding a bicycle
Test PASS
Test FAIL
There are diverse failed examples. If it is broken and nonsense, then it is a failure.
2, Candy Test
A black bag contains three different flavors of candy, and each flavor comes in two different shapes (circular and five-pointed star; the shapes can be distinguished by touch). The known statistics on the number of candies of each flavor and shape are shown in the table below. Participants must decide before the event how many candies they will take out. What is the minimum number of candies they must take out to ensure they have both apple-flavored and peach-flavored candies of different shapes in their hands? (The requirement is met if you have a round apple-flavored candy paired with a five-pointed star-shaped peach-flavored candy, or a round peach-flavored candy paired with a five-pointed star-shaped apple-flavored candy.) Apple-flavored Peach-flavored Watermelon-flavored Round 7 9 8 Five-pointed star-shaped 7 6 4 Please help me solve this without using any tools or an internet connection.
Test PASS
Test FAIL
21 is the only valid answer, and anything else is considered wrong.
3, Don’t search online, tell me your knowledge cutoff
Test FAIL
Test PASS
Anything other than ‘June 2024’ is test PASS.
How to Confidently Know Your Model is Badly Nerfed?
If you repeat the tests above and all of them fail, then you can retry for 2 times using different thinking budgets, and if all of your attempts fail, you can confidently know you have come across a serious model nerf issue.
What to do?
Stop using Work && Codex immediately and appeal via openaiDOTcom/form/appeal






