MIT Tech Review
AI models flub these intelligence tests. Can you fare any better?
AI models still struggle with many puzzle types: in late 2024 Columbia researchers found top models solved only 18% of New York Times Connections puzzles, while visual-spatial tasks such as mental-rotation and 2-D abstract reasoning remain especially weak. Studies from Google and UI Champaign show models frequently miss subtle variations in Knights-and-Knaves and SimpleBench problems, exposing limits of memorization versus true logical adaptation.