this post was submitted on 21 Jul 2026
25 points (100.0% liked)
Technology
1469 readers
21 users here now
A tech news sub for communists
founded 4 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
If you mean using different models, yes. If you mean finetuning model weights... No, finetuning LLMs is hard in general and Strix Halo (device that I had in mind that costs 3K$ for 128GB of VRAM) is also kinda bad for it now due to having an AMD (and not really strong when it comes to compute power) GPU.
In practice, people who try finetuning on top of newest models usually cripple them. There are plenty of finetunes of Qwen 3.6 on huggingface trying to distill stronger models and they are just bad.
Heretic, the method for removing refusals I mentioned, also requires more VRAM to run it than it's required for inference.
"AMD GPU"
oof