this post was submitted on 31 Jul 2026
39 points (100.0% liked)
chat
8642 readers
181 users here now
Chat is a text only community for casual conversation, please keep shitposting to the absolute minimum. This is intended to be a separate space from c/chapotraphouse or the daily megathread. Chat does this by being a long-form community where topics will remain from day to day unlike the megathread, and it is distinct from c/chapotraphouse in that we ask you to engage in this community in a genuine way. Please keep shitposting, bits, and irony to a minimum.
As with all communities posts need to abide by the code of conduct, additionally moderators will remove any posts or comments deemed to be inappropriate.
Thank you and happy chatting!
founded 5 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Yes, you need to leave some space for the context (kv cache) - a lot of space if you need a lot of context. Large context just allows you to generate more text and run longer tasks.
12 gigs is actually a good amount to run this class of model. As I said, not all of the model has to live in the vram
here is an example of a guy running a 35 billion parameter model on gtx 1060 (6 gigs of vram) at 36 tokens a second with max context size, a good speed for most tasks. Point is, these days you can run fairly capable models on ancient hardware if you know what you're doing
Oh damn. Nice find. Hurts a little that my rig is barely above his floor tho. lol.
Okay, that makes sense. The only thing I do with my GPU is play video games. So you're pretty much good to just fill VRAM pretty much all the way, depending on settings.
I found a YouTube link in your comment. Here are links to the same video on alternative frontends that protect your privacy: