VentureBeat
Frontier models can recover up to 65% of facts they can't directly recall — just by thinking longer

Frontier models such as GPT-5 and Gemini-3 encode 95-98% of tested facts, but often fail to retrieve them without additional inference time, allowing up to 65% of missing answers to be recovered through longer “thinking” or chain-of-thought prompting. The authors propose fact-level profiling to distinguish encoding failures, which require larger models or more data, from recall failures that can be mitigated by post-training inference techniques.