Skip to content
logo AIFinitee

All things AI & LLM

  • All Things AI & LLM
  • Archive
  • Strix Halo
  • Apple Silicon
  • Distributed Inference
Subscribe

Quantization

Home » Quantization
GLM 5.3 Flash running on a 128GB PC — featured banner
Posted inAMD Distributed Inference

GLM 5.3 Flash on 128GB: Two Easiest Ways to Run It — and When Two Strix Halos Beat One

GLM 5.3 Flash on one 128GB URAM machine, or even Two for higher quality Here's the headline that should make you sit up: a 320-billion-parameter multimodal model — one that…
Tags: ds4, GLM, local LLM, Quantization, Strix Halo
GLM 5.2 at 4-bit with MTP on Mac Studio — featured banner
Posted inApple Silicon

GLM 5.2 at 4-Bit With MTP: The Easiest Way In

And the 512GB Mac Studio that Apple just stopped selling. Apple discontinued the 512GB Mac Studio. Not the chip — the M3 Ultra is still on the shelf — but…
Tags: GLM, LLM Inference, Local AI, Mac Studio, Quantization
DeepSeek V4 Flash Q2 running on Strix Halo 128GB — featured banner
Posted inAMD Distributed Inference

Running DeepSeek V4 Flash Q2 on Strix Halo 128GB

AIfinitee believes in local compute as an alternative to cloud and that the most capable open models should be runnable on hardware you actually own. This post is the recipe for…

Tags: Cluster Computing, DeepSeek, LLM Inference, Local AI, Quantization, Strix Halo
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top