Thinking fast and slow in AI: The role of metacognition (2021)

arxiv.org

169 points by teleforce 19 hours ago


gchamonlive - 11 hours ago

There was this post a few days ago https://news.ycombinator.com/item?id=49797323

It had this to say in the linked post:

  This led to the natural question: can gzip do language modeling? (...). Here’s some real, unedited output after priming it on tiny Shakespeare:

  gzipt --corpus data/tinyshakespeare.txt --prompt $'MENENIUS:\n' --length 200

  MENENIUS:
  'Though all at once canq

  MARCIUS:
  Pray now, nocamest thou to a morsel.

  LARTIUS:
  Hence, and
  I' the end admire, where G
  again; and after it ag .
Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.

The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.

red75prime - 16 hours ago

It reminds me of "At the time we drew boxes labeled 'perception', 'cognition' with arrows between them." An imprecise quote that I can't place.

I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.

hoppp - 8 hours ago

Most humans have weak meta-cognition, a large percentage doesn't have verbal thoughts.

Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.

In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.

arbirk - 10 hours ago

Interestingly that is not what we got, but maybe we should loop at architectures like this again? The JEPA loop is interesting, but might fail for the in-flexibility of the component ordering

creativeSlumber - 14 hours ago

How relevant is this fast/slow thinking thing with regards to current frontier models?

I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.