this post was submitted on 21 Sep 2026
44 points (97.8% liked)
Linux
15034 readers
287 users here now
A community for everything relating to the GNU/Linux operating system (except the memes!)
Also, check out:
Original icon base courtesy of lewing@isc.tamu.edu and The GIMP
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
i mean, you can just do the math.
my 7900XTX consumes 300W at full load. my 5700X3D uses 95. say 500W for the entire system working at maximum. at 24GB of VRAM, i can run 20b parameter models at varying speeds, on average about 5-10 tokens per second. in the most generous case, then, it takes 110ish seconds at 500W to generate 1024 tokens which is the limit i've configured. that's 15 Wh per prompt, or 70 prompts per kWh, which isn't really possible to do on that machine since each prompt takes more than a minute.
when playing a game, the hardware isn't all maxed out, call it 350W average, which means a bit less than three hours of gaming per kWh.
if we scale all this up, bearing in mind that specialised hardware is more efficient, a prompt that results in 10 seconds of work for a 2kW server, that's 5Wh, which means 200 prompts per kWh. but these things are running constantly, and with 10-second responses you can fit 360 replies per hour, for a total of 1.8kWh per hour, ideally.
so no, those numbers make no sense at all.
I find it kinda hard to believe that a prompt takes up an entire server for a full 10 seconds. That does not seem in any way scalable.
context-switching is expensive, but you are right that 10 seconds may be too long. i just pulled something.
then again they've optimised the size of the machines to the point that a single rack can fit more than 100 blade servers, and a dc can have a thousand racks.