
Wenn die KI diesen Test besteht, sollten Sie aufpassen | Die Macher eines neuen Tests namens „Humanity’s Last Exam“ argumentieren, dass wir bald nicht mehr in der Lage sein werden, Tests zu erstellen, die hart genug für KI-Modelle sind.
https://www.nytimes.com/2025/01/23/technology/ai-test-humanitys-last-exam.html
4 Kommentare
„If you’re looking for a new reason to be nervous about artificial intelligence, try this: Some of the smartest humans in the world are struggling to create tests that A.I. systems can’t pass.
For years, A.I. systems were measured by giving new models a variety of standardized benchmark tests. Many of these tests consisted of challenging, S.A.T.-caliber problems in areas like math, science and logic. Comparing the models’ scores over time served as a rough measure of A.I. progress.
But A.I. systems eventually got too good at those tests, so new, harder tests were created — often with the types of questions graduate students might encounter on their exams.
Those tests aren’t in good shape, either. New models from companies like OpenAI, Google and Anthropic have been getting high scores on many Ph.D.-level challenges, limiting those tests’ usefulness and leading to a chilling question: Are A.I. systems getting too smart for us to measure?
This week, researchers at the Center for AI Safety and Scale AI are releasing a possible answer to that question: A new evaluation, called “Humanity’s Last Exam,” that they claim is the hardest test ever administered to A.I. systems.“
This is so stupid. Tech companies are just going to build a model that is going that is capable of answering these questions and people that have no clue will claim it is over AGI has been achieved.
Got news for Humans. We are outmoded and no longer necessary to keep society moving. We’ll just eat shit, and billionaires will live with robot butlers
So I know this sounds like a really dumb question, but based on my knowledge of AI and learning models, it would be an easy feat for an AI to solve any problem that exists as long as that problem has a solution which also exists as a key.
Do these tests, ask AI to conceive of something novel based upon minimal input?
If not, then it’s just rote recall on a masive scale.