Skip to content

localmodel.run

About localmodel.run

I kept Googling "can I run Llama 3 70B on a 16 GB Mac" and getting Reddit threads with conflicting answers and no math shown. So I built a calculator that does the actual memory math, cites its sources, and gives a straight yes or no.

The site covers Mac, Windows, Linux, iOS, and Android. It recommends the right tool per platform because the wrong one cuts your throughput in half.

The browser became a real target in 2026, so the site now measures it the same way: 54 models that run in the tab over WebGPU with Transformers.js, every download size read from the HuggingFace file tree, and six pages where you can run them live with one click. Nothing downloads until you press Run, and nothing you type, say, or drop in ever leaves the tab.


Data sources

Text model figures are validated against real GGUF files on the Ollama library and HuggingFace repos. Image, video, and audio figures are sourced from vendor specs and benchmarks. The formula is public. You can check it on the methodology page.


Builder

Built by Ansuman Shah, an indie dev in Bengaluru. A decade of building web and mobile apps at scale, and I was still Googling whether a model would fit in my RAM. So I built the tool that answers that and shows every source. I make small, focused tools under the NoodleApps name. Find me on X or GitHub.

Data last validated 2026-08-03. Figures are estimates, not guarantees. Always verify before a large download.