Discussion about this post

User's avatar
Srimallya Maitra's avatar

Llms are trained to generate the next token given the context. And the context length becomes the short term memory. It seems to me that every inference step we recompile the context. We are actively generating the context along the way.

That’s thinking.

Deep thinking is recompiling the context many time.

There might be a learn gating mechanism which decides to turn on the thinking. And that’s active memory generation.

It can be bert like separate encoder or a diffusion based system.

Mykyta Storozhenko's avatar

They also ran it w/ same compute allotment between reasoning and not reasoning models which is dumb because reasoning models tend to perform better when they have more compute since more compute = more tokens to think with. Shit paper, and ofc it’s Apple that publishes it, being the biggest losers when it comes to AI. real shocker. Anyway, obviously llms, reasoning or not, differ from shape rotator type thinking. They’re wordcels and can’t do shape rotator thinking, but most people can’t rotate shapes that well anyway, so make of that what you will.

5 more comments...

No posts

Ready for more?