Tested Nemotron 3.5 Lightning locally on coding; Hermes Agent shows strong tool-calling capabilities but overall code quality below expectations for its size
Read the original at old.reddit.com→Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quality of the code...
Original headline: "Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work"
Coverage timeline
- Aug 12, 10:15 UTC r/LocalLLaMA lead source Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work