WebLLM: high-performance in-browser LLM inference engine

github.com

101 points by saikatsg 16 hours ago


TekMol - 14 hours ago

This seems to be the demo:

https://chat.webllm.ai/

I am getting:

    WebGPUNotAvailableError: WebGPU is not supported in
    your current environment, but it is necessary to
    run the WebLLM engine.
On both, FireFox and Chromium on Linux.
mandeepj - 8 hours ago

It is kinda obvious, but maybe that's why it's not stated anywhere: each browser session will result in a download of 500 MB to ~1 GB, depending on your model selection. So, it's better to add a disclaimer if you end up using WebLLM in a customer-facing site.

MarioMan - 10 hours ago

I really enjoy this engine. I’ve used it for personal projects, but it hasn’t been updated since Gemma 2. I suggest using Transformers.js instead these days.

refulgentis - 15 hours ago

Project is de facto dead, used it for many years and had to rip it out 6 months ago, don't waste your time.

init0 - 11 hours ago

You might like webml-kit https://npm.im/webml-kit

conceptme - 12 hours ago

Bake me a cake

responds with

> Error: Cannot initialize runtime because of requested maxStorageBuffersPerShaderStage exceeds limit. requested=10, limit=9.

adastra22 - 11 hours ago

A WebX technology that actually involves browsers!

paidx - 4 hours ago

[flagged]

nnevatie - 14 hours ago

[flagged]

- 7 hours ago
[deleted]
asyncze - 13 hours ago

[dead]

ycCantCode - 12 hours ago

[dead]