Decreasing the power limit of the 5090 to 480W yields negligible inference slowdown, per test with Qwen 3.6-27b.
Read the original at old.reddit.com→I run my inference machine in the living room, so noise and heat output are a significant concern. Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% less t/s in...
Original headline: "Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible."
Coverage timeline
- Aug 4, 15:39 UTC r/LocalLLaMA lead source Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible.