this post was submitted on 31 Jul 2026
39 points (100.0% liked)
chat
8640 readers
183 users here now
Chat is a text only community for casual conversation, please keep shitposting to the absolute minimum. This is intended to be a separate space from c/chapotraphouse or the daily megathread. Chat does this by being a long-form community where topics will remain from day to day unlike the megathread, and it is distinct from c/chapotraphouse in that we ask you to engage in this community in a genuine way. Please keep shitposting, bits, and irony to a minimum.
As with all communities posts need to abide by the code of conduct, additionally moderators will remove any posts or comments deemed to be inappropriate.
Thank you and happy chatting!
founded 5 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this model can be compressed to as low as 12 gigs (though I wouldn't recommend going that low due to a huge quality drop). And not all of this kind of model has to reside in vram - you can have most of it in ram and the rest gets cached in vram
I'm curious, can the model fill your whole VRAM or do you have to leave extra space for the actual, you know, whatever it's doing? Like, could you actually do anything with the 12gb model if you only have 12 GB of VRAM?
Yes, you need to leave some space for the context (kv cache) - a lot of space if you need a lot of context. Large context just allows you to generate more text and run longer tasks.
12 gigs is actually a good amount to run this class of model. As I said, not all of the model has to live in the vram
here is an example of a guy running a 35 billion parameter model on gtx 1060 (6 gigs of vram) at 36 tokens a second with max context size, a good speed for most tasks. Point is, these days you can run fairly capable models on ancient hardware if you know what you're doing
Oh damn. Nice find. Hurts a little that my rig is barely above his floor tho. lol.
Okay, that makes sense. The only thing I do with my GPU is play video games. So you're pretty much good to just fill VRAM pretty much all the way, depending on settings.
I found a YouTube link in your comment. Here are links to the same video on alternative frontends that protect your privacy: