this post was submitted on 19 Jul 2026
175 points (88.5% liked)

Technology

86893 readers
3309 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
top 50 comments
sorted by: hot top controversial new old
[–] mierdabird@lemmy.dbzer0.com 66 points 2 weeks ago (2 children)

The article talks a lot trash about AMD and ROCm but vulkan works fine too. In fact from a datacenter GPU standpoint there is an AMD option called the V620 available on US EBay that I was able to haggle to $350, with 32GB VRAM, 512GB/s bandwidth, and runs the same Qwen-3.6-27b at about 20t/s. It requires a few of the same fan shenanigans this guy did but there is no need to pull specific past software versions to make it usable in Linux

[–] ranzispa@mander.xyz 10 points 2 weeks ago

This is nice to know. Thank you.

[–] JohnWorks@sh.itjust.works 6 points 2 weeks ago (1 children)

Honestly even for the prices of around 500$ that I'm seeing it for it looks like a pretty good value to get 32gb of vram. I see it says 300w on AMD's product page for it does it have any way of power limiting the card to get more efficiency/less heat?

[–] mierdabird@lemmy.dbzer0.com 5 points 2 weeks ago

I spent a lot of time researching and testing different methods for that, the only thing that worked was LACT in Linux. Using that I was able to undervolt 100mV and drop GPU power usage about 10%. On my B450 ITX board with a Ryzen 2400GE CPU it pulls 30w idle from the wall, and about 300w inferencing with VRAM filled. (330w before LACT) My fan solution ended up being to buy a 3d-printed shroud off ebay, the fan that came with it was super loud so I switched to an arctic p8 Max, and control it with the motherboard targeting a t-sensor header with the probe attached to the backplate.

[–] palordrolap@fedia.io 43 points 2 weeks ago (3 children)

Here I was expecting a graphics demo to blow our collective minds but instead I got a story about a local LLM for cheap. It is . I should have known better.

Were I the author / tech cobbler here, I'd be concerned that too much time with an LLM, local or otherwise, might erode or dull my apparently fairly sharp reasoning and tech skills. (Clarification: Not my sharpness, theirs. I'm a potato.)

Other thoughts: For a minute I thought this whole thing was a tribute to, or a troll in the manner of, that one Redditor that always spun their stories around to being about their dad beating them with jumper cables.

Also, my old PC developed an issue like the warm reboot problem, except with the network interface. I couldn't just restart, I had to power off and back on. I never did bother to find out whether it was early signs of hardware failure or whether it was an old hardware / newer kernel mismatch.

[–] Evotech@lemmy.world 4 points 2 weeks ago (1 children)
[–] palordrolap@fedia.io 6 points 2 weeks ago

The change of the meaning of the G in GPU from "graphics" to "general" is even less well documented and used than the "V" of DVD changing from "video" to "versatile".

Indeed it only occurred to me what it must have changed to and to go looking to confirm after seeing your comment.

And frankly they ought to have changed the name to something like "MPPU" if they wanted it to stick (massively parallel).

[–] Artisian@lemmy.world 4 points 2 weeks ago

Back in my day, we memorized logarithm tables and we liked it!

load more comments (1 replies)
[–] melfie@lemmy.zip 20 points 2 weeks ago (3 children)

Seems like an awful lot of trouble to save $100 not buying a 5060 Ti that also has 16GB.

[–] ranzispa@mander.xyz 16 points 2 weeks ago (1 children)

448 GB/s memory vs 900 GB/s and at a cheaper price.

load more comments (1 replies)
[–] Lucidlethargy@sh.itjust.works 4 points 2 weeks ago (1 children)

I had no idea you could get a 16gb card this cheap! TIL.

It's pretty low performance, though... My card from 8 years ago nearly matches the passmark rating.

load more comments (1 replies)
load more comments (1 replies)
[–] frongt@lemmy.zip 14 points 2 weeks ago (2 children)

Sure it's got a lot of VRAM, but the 4080 has five times the compute power.

[–] plz1@sh.itjust.works 8 points 2 weeks ago (3 children)

You're not getting a 4080 for $200 though...

load more comments (3 replies)
[–] LodeMike@lemmy.today 6 points 2 weeks ago

That's fine I just need to display pictures of your mom (they are very large) (/s)

[–] First_Thunder@lemmy.zip 14 points 2 weeks ago (2 children)

The main thing that itches me with the V100 is the fact that given that pascal is about to be EOL, a 2017 card is probably soon next

[–] pigup@lemmy.world 6 points 2 weeks ago

The ampere class is soon to fall out of relevance as well.

load more comments (1 replies)
[–] Malyca@lemmy.zip 13 points 2 weeks ago (3 children)
[–] scrubbles@poptalk.scrubbles.tech 17 points 2 weeks ago

You're in the wrong community if you're asking questions like that.

[–] vaultdweller013@sh.itjust.works 8 points 2 weeks ago (1 children)

Ever watched bringus studios? Man played games on a drive through computer. As the old saying goes all hardware is good hardware if you know what to do.

load more comments (1 replies)
[–] DudeImMacGyver@kbin.earth 11 points 2 weeks ago (3 children)

82db? Wild. I guess they don't really care that much about noise in data centers though.

[–] notabot@piefed.social 29 points 2 weeks ago (1 children)

If it doesn't sound like a jet plane taking off, is it really a server at all?

[–] muhyb@programming.dev 3 points 2 weeks ago

Hey, my home server netbook can play jet_plane_taking_off.opus.

[–] jj4211@lemmy.world 8 points 2 weeks ago

They mostly don't, but this is also not how the datacenters cool them.

A datacenter will either have an open water loop, or an all in one taking heat to a more advantagous place for a radiator to be, or at the very least better managed airflow with bigger fans and more specific air baffles.

This thing has no such luxury and has a small area and unknown broader thermal context, so screaming it is to make up for the limitations of the scenario.

[–] MorningWood@anarchist.nexus 5 points 2 weeks ago (1 children)

No they really dont. Big ass fans running 24/7 to help the small fans running 24/7. It all blends into an easily ignored drone though just dont try to have a conversation in there.

load more comments (1 replies)
[–] pech@lemmy.world 10 points 2 weeks ago (2 children)

eBay has some rad Chinese mezzanine boards for these guys too. Nvlink works and everything lol 3566 file-QvfRnBhkoKQBxmqtmrGBLV

[–] Septimaeus@infosec.pub 4 points 2 weeks ago (2 children)

Cool mezzanine boards, but can we talk about your dope af custom jig for offset mounting arbitrary boards?

load more comments (2 replies)
[–] VindictiveJudge@lemmy.world 4 points 2 weeks ago (1 children)

The last time I saw a passively cooled GPU was the '90s.

[–] pech@lemmy.world 8 points 2 weeks ago

I am using PTM sheets and they idle at decent temps, though I did ziptie some high CFM fans behind them lol 3581

[–] echodot@feddit.uk 9 points 2 weeks ago

Yeah but the 0 point doing this unless you want to run AI models for some reason. These GPUs can't do video game graphics so this isn't a solution to the GPU shortage.

This is a bit like me writing an article about NASCAR, now I can turn left whenever I want. But I haven't magically acquired a functional vehicle for a fraction of its value. I've purchased a second hand specialist product that is usually useless outside of that environment.

[–] Postmortal_Pop@lemmy.world 8 points 2 weeks ago (4 children)

Not an AI guy, but I do like using niche hardware wrong to get results cheap. Can anyone tell me what this would be like for gaming or general computing? My 1660 super was a budget pick when I got it back in '18.

load more comments (4 replies)
[–] brucethemoose@lemmy.world 4 points 2 weeks ago* (last edited 2 weeks ago)

This isn't the smart way, though.

What the homelabbers do (at least before the RAM crisis) is buy Xeon/TR/EPYC boards on the cheap, and then run gaming GPUs for hybrid inference.

This is what I do. I run MiMo 2.5 at 8-10t/s on a 7800X3D/RTX 3090/128GB CPU RAM, more with Dflash. That's a 300B model: it's not even in the same class as Qwen 27B, which is what the dev in OP's article is trying to run.

And this is small-time: setups with 4-8 memory channels can run stuff like Kimi or Deepseek Pro, even faster. Or they can run smaller LLMs with quantization types that are very fast on CPUs, and get crazy speeds.

...And besides, Qwen 27B can run fine on a 4080, with the right framework. It will fit in 16GB as an exl3.


Not that this isn't a cool hardware hacking project.

...But it's kind of the wrong approach. It's about 2 years out of date, as MoEs are king in LLM land now. RAM is horrendously expensive, yes, but so are most used V100s, or used 3090s.

load more comments
view more: next ›