vitalik.eth
vitalik.eth|Sep 17, 2026 01:01
Qwen 3.8 flash is truly impressive, and llama.cpp has been rapidly getting better and better at processing it columns are: pre-existing prompt, new prompt, generated, input tok/s, output tok/s This is on my laptop (strix halo). I think we're very close to the point where you can just use local models for a large share of tasks, and for anything more advanced, workflows like "use your local model to orchestrate queries to powerful models so your queries don't leak your personal information" actually become viable.
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads