I'm not paying $20 for ChatGPT or Claude because a free local LLM does

xda-developers.com

29 points by hsnewman 2 hours ago


smcleod - 3 minutes ago

If you were only paying $20 for LLMs previously then you were getting very low usage / little done.

reddit_clone - 5 minutes ago

Ok. What can I do with an M4 MBP, with 48G memory?

mrandish - 30 minutes ago

Thanks for writing and sharing. Since I also have a 4070 Ti Super, I'm always interested in hearing people's experiences with local models that fit 'middle-ground' hardware. I kind of feel left out because most posts I see seem to either be about clever ways of making older, lower-end cards usable, leveraging mega-CPU RAM (128GB) or how aweseome high-end GPUs like 4090/5090 can be.

kristianp - 26 minutes ago

> llama.cpp serves it over localhost at speeds that stop being a complaint after a few minutes of use.

"That stop being a complaint"? That's a strange construction.

I wonder if qwen 4 will make these smaller cards more viable by allowing the ngram storage to be hosted on CPU RAM.

braggerxyz - an hour ago

20$/month vs. 1200$ gpu. That's a lot of months you can pay the subscription

kittikitti - an hour ago

This is a great article! Thank you for sharing. I also recommend opencode that has built in features for local model hosting.

WarmWash - an hour ago

"I spent $1100 on a graphics card so I can save $20/mo running a mid/low intelligence model at 33tk/s with 16k context"