The article talks a lot trash about AMD and ROCm but vulkan works fine too. In fact from a datacenter GPU standpoint there is an AMD option called the V620 available on US EBay that I was able to haggle to $350, with 32GB VRAM, 512GB/s bandwidth, and runs the same Qwen-3.6-27b at about 20t/s. It requires a few of the same fan shenanigans this guy did but there is no need to pull specific past software versions to make it usable in Linux
Technology
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
This is nice to know. Thank you.
Honestly even for the prices of around 500$ that I'm seeing it for it looks like a pretty good value to get 32gb of vram. I see it says 300w on AMD's product page for it does it have any way of power limiting the card to get more efficiency/less heat?
I spent a lot of time researching and testing different methods for that, the only thing that worked was LACT in Linux. Using that I was able to undervolt 100mV and drop GPU power usage about 10%. On my B450 ITX board with a Ryzen 2400GE CPU it pulls 30w idle from the wall, and about 300w inferencing with VRAM filled. (330w before LACT) My fan solution ended up being to buy a 3d-printed shroud off ebay, the fan that came with it was super loud so I switched to an arctic p8 Max, and control it with the motherboard targeting a t-sensor header with the probe attached to the backplate.
Here I was expecting a graphics demo to blow our collective minds but instead I got a story about a local LLM for cheap. It is . I should have known better.
Were I the author / tech cobbler here, I'd be concerned that too much time with an LLM, local or otherwise, might erode or dull my apparently fairly sharp reasoning and tech skills. (Clarification: Not my sharpness, theirs. I'm a potato.)
Other thoughts: For a minute I thought this whole thing was a tribute to, or a troll in the manner of, that one Redditor that always spun their stories around to being about their dad beating them with jumper cables.
Also, my old PC developed an issue like the warm reboot problem, except with the network interface. I couldn't just restart, I had to power off and back on. I never did bother to find out whether it was early signs of hardware failure or whether it was an old hardware / newer kernel mismatch.
They cannot run graphics
The change of the meaning of the G in GPU from "graphics" to "general" is even less well documented and used than the "V" of DVD changing from "video" to "versatile".
Indeed it only occurred to me what it must have changed to and to go looking to confirm after seeing your comment.
And frankly they ought to have changed the name to something like "MPPU" if they wanted it to stick (massively parallel).
Back in my day, we memorized logarithm tables and we liked it!
Seems like an awful lot of trouble to save $100 not buying a 5060 Ti that also has 16GB.
I had no idea you could get a 16gb card this cheap! TIL.
It's pretty low performance, though... My card from 8 years ago nearly matches the passmark rating.
Sure it's got a lot of VRAM, but the 4080 has five times the compute power.
That's fine I just need to display pictures of your mom (they are very large) (/s)
The main thing that itches me with the V100 is the fact that given that pascal is about to be EOL, a 2017 card is probably soon next
The ampere class is soon to fall out of relevance as well.
Why tho?
You're in the wrong community if you're asking questions like that.
Ever watched bringus studios? Man played games on a drive through computer. As the old saying goes all hardware is good hardware if you know what to do.
82db? Wild. I guess they don't really care that much about noise in data centers though.
If it doesn't sound like a jet plane taking off, is it really a server at all?
Hey, my home server netbook can play jet_plane_taking_off.opus.
They mostly don't, but this is also not how the datacenters cool them.
A datacenter will either have an open water loop, or an all in one taking heat to a more advantagous place for a radiator to be, or at the very least better managed airflow with bigger fans and more specific air baffles.
This thing has no such luxury and has a small area and unknown broader thermal context, so screaming it is to make up for the limitations of the scenario.
No they really dont. Big ass fans running 24/7 to help the small fans running 24/7. It all blends into an easily ignored drone though just dont try to have a conversation in there.
eBay has some rad Chinese mezzanine boards for these guys too. Nvlink works and everything lol

Cool mezzanine boards, but can we talk about your dope af custom jig for offset mounting arbitrary boards?
The last time I saw a passively cooled GPU was the '90s.
I am using PTM sheets and they idle at decent temps, though I did ziptie some high CFM fans behind them lol

Yeah but the 0 point doing this unless you want to run AI models for some reason. These GPUs can't do video game graphics so this isn't a solution to the GPU shortage.
This is a bit like me writing an article about NASCAR, now I can turn left whenever I want. But I haven't magically acquired a functional vehicle for a fraction of its value. I've purchased a second hand specialist product that is usually useless outside of that environment.
Not an AI guy, but I do like using niche hardware wrong to get results cheap. Can anyone tell me what this would be like for gaming or general computing? My 1660 super was a budget pick when I got it back in '18.
This isn't the smart way, though.
What the homelabbers do (at least before the RAM crisis) is buy Xeon/TR/EPYC boards on the cheap, and then run gaming GPUs for hybrid inference.
This is what I do. I run MiMo 2.5 at 8-10t/s on a 7800X3D/RTX 3090/128GB CPU RAM, more with Dflash. That's a 300B model: it's not even in the same class as Qwen 27B, which is what the dev in OP's article is trying to run.
And this is small-time: setups with 4-8 memory channels can run stuff like Kimi or Deepseek Pro, even faster. Or they can run smaller LLMs with quantization types that are very fast on CPUs, and get crazy speeds.
...And besides, Qwen 27B can run fine on a 4080, with the right framework. It will fit in 16GB as an exl3.
Not that this isn't a cool hardware hacking project.
...But it's kind of the wrong approach. It's about 2 years out of date, as MoEs are king in LLM land now. RAM is horrendously expensive, yes, but so are most used V100s, or used 3090s.
