QFN llama.cpp: user asks for config tips for running a squished ISTA model on dual RTX 5060 Ti and 32 GB RAM
Read the original at www.reddit.com→https://imgur.com/a/Ef2xyNu Using this squished down ISTA model on 2x 5060ti 16gb and 32gb ddr4 ram I'm wondering if my settings are correct as I cant really find much consistent feedback for this model on this...
Original headline: "QFN llama.cpp Any juice left to squeeze?"
Coverage timeline
- Oct 4, 20:35 UTC r/LocalLLaMA lead source QFN llama.cpp Any juice left to squeeze?