r/LocalLLaMA · · 1 min read

Has anyone here fiddled with TPUs for inference ?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I discovered recently that Google uses their own TPUs, like tiny ASIC cards like the toy ones that existed for bitcoin. And while it sounds inefficient the fact they use thousands of them because...they can...means at scale they aren't so bad.

Has no one here given them a try? I see web search results of tiny ones that can be purchased and look like nvme adapted where I search them for \~58 euros. Not sure what 40 TOPS translates to compared to my Nvidia 5060.

But not just that, but the user experience with them, are they a nightmare to use ?

submitted by /u/misanthrophiccunt
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA