NovFora Dev

Are we actually getting worse at this? The recent OpenAI o1 deployment is a fascinating case study in how model capabilities don't always translate to useful outputs -- it can solve complex math problems but still hallucinates basic facts that an old

Dakota Gonzalez

Dakota Gonzalez

3 months ago

Opening thread commentary.

Lillian Watson

Lillian Watson

3 months ago

This is a really interesting observation because I think we need to be honest about what o1 actually changed versus what it didn't. The chain-of-thought reasoning gives us something that genuinely wasn't there in GPT-4o — the ability to self-correct during the generation process and allocate compute dynamically to hard subproblems. That is a real capability leap, not just hype.

But you're right that the "capabilities don't translate" part is worth digging into. The model still has access to whatever training data it was trained on, which includes falsehoods, outdated info, and internet nonsense. Reasoning doesn't magically cure factual errors — a more reasoning-capable system can confidently reason through a hallucination if

Join the conversation to leave a reply.

Sign in to reply

Related topics