Skip to content

Device profile · Windows

Best local LLMs for AMD Ryzen AI Halo (128GB)

AMD Ryzen AI Halo (128GB) has ~96 GB usable for model weights and runs 118 of 135 popular models. Best tool: LM Studio.

117 run well 1 tight fit 17 too heavy
Usable memory
~96 GB
Models run
118
Too large
17
Top pick
30.5B
Top pick Q4_K_M

Runs at Q4_K_M using ~20.7 GB of ~96 GB usable. You have room for FP16 for higher quality.

Runs on AMD Ryzen AI Halo (128GB)

Compatible models 118 total

Best way to run models on Windows

Runtime guide Windows

Beginner: LM Studio, Best GUI on Windows, auto-detects CUDA/Vulkan backends.

Power user: Ollama (CUDA), Scriptable server; CUDA path is fastest on NVIDIA.

AMD GPUs run via Vulkan/ROCm at roughly half CUDA throughput. NVIDIA is the smooth path on Windows.

Full Windows tool guide →

FAQ

What is the best local LLM for AMD Ryzen AI Halo (128GB)?

Qwen3 30B-A3B is the strongest model that runs comfortably, using ~20.7 GB at Q4_K_M of the ~96 GB usable on AMD Ryzen AI Halo (128GB).

How much of AMD Ryzen AI Halo (128GB)'s memory can I use for a model?

About 96 GB. On a CPU-only machine, leave headroom for the OS and apps.

Which tool should I use on Windows?

LM Studio (Best GUI on Windows, auto-detects CUDA/Vulkan backends.) or Ollama (CUDA) for speed. AMD GPUs run via Vulkan/ROCm at roughly half CUDA throughput. NVIDIA is the smooth path on Windows.

Compare nearby devices

Hardware with a similar usable pool, and how many of the 135 tracked models each runs next to AMD Ryzen AI Halo (128GB)'s 118.

Sources

Memory figures are estimates. See methodology.