Why your local LLM feels dumber than it is

forum.level1techs.com

75 points by felineflock 4 hours ago


jonplackett - 2 hours ago

I just got qwen 3.8 27b mlx running on my Macbook Pro and honestly I’m pretty blown away by how not-dumb it is.

catlifeonmars - 8 minutes ago

[delayed]

JacobJack - 38 minutes ago

> And the comparisons in this post are not going to be running some 2.58-bit-gguf-in-ollama with a couple test prompts.

Genuine question : is there something fundamentally wrong with Ollama ?

I use Ollama because it is easy to set up and manage (and also because VLLM is not super Windows friendly).

I thought the main advantage of VLLM was better concurrency management (better batching).

But if the quality of the interference itself is an issue, then maybe I should reconsider my choice.

anotherCodder - 2 hours ago

[flagged]