this post was submitted on 19 Jul 2026
271 points (89.0% liked)

Technology

86485 readers
3177 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
(page 2) 50 comments
sorted by: hot top controversial new old
[–] DudeImMacGyver@kbin.earth 16 points 1 day ago (4 children)

82db? Wild. I guess they don't really care that much about noise in data centers though.

[–] MorningWood@anarchist.nexus 8 points 1 day ago (1 children)

No they really dont. Big ass fans running 24/7 to help the small fans running 24/7. It all blends into an easily ignored drone though just dont try to have a conversation in there.

[–] DudeImMacGyver@kbin.earth 5 points 1 day ago

What?

Sorry, I think I might have hearing damage from all this constant noise.

load more comments (3 replies)
[–] First_Thunder@lemmy.zip 16 points 1 day ago (3 children)

The main thing that itches me with the V100 is the fact that given that pascal is about to be EOL, a 2017 card is probably soon next

[–] deltapi@lemmy.world 4 points 1 day ago

About to be? Pascal cards no longer work in TrueNAS scale due to Nvidia dropping driver support

[–] pigup@lemmy.world 6 points 1 day ago

The ampere class is soon to fall out of relevance as well.

[–] brucethemoose@lemmy.world 3 points 1 day ago* (last edited 1 day ago)

This isn't the smart way, though.

What the homelabbers do (at least before the RAM crisis) is buy Xeon/TR/EPYC boards on the cheap, and then run gaming GPUs for hybrid inference.

This is what I do. I run MiMo 2.5 at 8-10t/s on a 7800X3D/RTX 3090/128GB CPU RAM, more with Dflash. That's a 300B model: it's not even in the same class as Qwen 27B, which is what the dev in OP's article is trying to run.

And this is small-time: setups with 4-8 memory channels can run stuff like Kimi or Deepseek Pro, even faster. Or they can run smaller LLMs with quantization types that are very fast on CPUs, and get crazy speeds.

...And besides, Qwen 27B can run fine on a 4080, with the right framework. It will fit in 16GB as an exl3.


Not that this isn't a cool hardware hacking project.

...But it's kind of the wrong approach. It's about 2 years out of date, as MoEs are king in LLM land now. RAM is horrendously expensive, yes, but so are most used V100s, or used 3090s.

load more comments
view more: ‹ prev next ›