Autoregressive Language Model on the 6502 Processor

mattbeton.com

136 points by nmstoker 7 days ago


derefr - 4 days ago

> The model weights and inference code need to be contained within 25KB of user-space memory

Wouldn’t it be era-appropriate to allow relying on banked memory? You’d still need to hold the inference code, but you could effectively stream(ing page) the weights as you compute on them.

actionfromafar - 4 days ago

The 6502 is notoriously unfit for a C compiler, so probably there is room for more performance in the future. :)

bmc7505 - 4 days ago

Cool to think this demo would have been possible over fifty years ago. I wonder what someone from 1975 would have said if you had shown this to them back then.

tyromaniac - 4 days ago

This is super cool! As someone who's worked a little with NES programming and tried out cc65, I'm surprised he didn't just hand write some assembly, he likely couldve saved a lot of space if I had to guess.

vintermann - 4 days ago

From my experience, modern machine learning models don't scale well down at all, and it's almost certainly better to just use a simple Markov chain variant of some sort - like Niall on the Amiga, or whatever Terry Pratchett used to come up with Foul Ol' Ron's catchphrase, "Millennium hand and shrimp".

torment-nexus - 4 days ago

The biggest win for AI dev efficiency is cutting down what gets loaded into context. Semantically matching tasks to the top tools helps a lot.

toplinesoftsys - 4 days ago

This is amazing project! I hope it will result in real miniaturization of AI - for example, edge LLM inside of glasses. That will be awesome.

- 4 days ago
[deleted]
aghilmort - 4 days ago

really great work

jkwang - 4 days ago

[flagged]

- 4 days ago
[deleted]
leonmeng - 4 days ago

[flagged]