TPUs for inference discussed; user experiments and scale considerations highlighted
Read the original at old.reddit.com→I discovered recently that Google uses their own TPUs, like tiny ASIC cards like the toy ones that existed for bitcoin. And while it sounds inefficient the fact they use thousands of them because...they can...means...
Original headline: "Has anyone here fiddled with TPUs for inference ?"
Coverage timeline
- Aug 8, 10:45 UTC r/LocalLLaMA lead source Has anyone here fiddled with TPUs for inference ?