I tested 80 models on two simple grid-based problems, asking them to locate 2026 and compute the sum of the neighbors when the numbers are placed in a spiral on the grid. The results surprised me, as models performed better on the problem I thought to be harder, but also I got to see the models cheating. Read more at https://mihai.page/ai-2026-1/
Remote
Mihai Maruseac
@mihaimaruseac@universeodon.com
Building AGI with Privacy and Security at OpenAI.
Previously: ML Supply chain security @ Google Open Sourse Security Team (GOSST, released model signing & GUAC).
Previously: TensorFlow Security & OSS @ Google Research.
Previously: Haskell+differential privacy+ML @ LeapYear.
0 Followers
0 Following
2 Posts
Joined November 20, 2022
Blog:
Open post
I've been using LLMs to answer questions and learn in the past weeks. I'm summarizing how Deep Research and Study Mode (Guided Learning / Study and Learn) from Gemini and ChatGPT can be used for this, without going into too many details. https://mihai.page/deep-research-study-mode/
0
0
0
0