Skip to content
logo AIFinitee

All things AI & LLM

  • All Things AI & LLM
  • Archive
  • Strix Halo
  • Apple Silicon
  • Distributed Inference
Subscribe

local LLM

Home » local LLM
Aimee, the AIFinitee chibi mascot, reads a fairy-tale storybook and says Once Upon... while a cheerful chibi onion wearing round glasses replies A Time - a playful nod to engram lookup memory.
Posted inPrimer

Once Upon a Time: DeepSeek V4.1-Flash and Qwen3.8-Flash-Next – What the hekk is an Engram?!

Two of the most interesting open-model releases of the quarter just landed, and they share a spec-sheet line that has a lot to do with benchmarks - but ignores the…
Tags: DeepSeek, Engram, Local AI, local LLM, Qwen, Strix Halo
Banner: two purple-haired chibi anime girls affectionately hugging two AMD mini-PCs with glowing cyan accents, dark purple background
Posted inAMD Distributed Inference

Repeat after me: Good things come in TP=2 (DeepSeek V4 Flash Q4 on 2x Strix Halos RDMA)

Models continue to improve in quality for a similar memory size. Sometimes, though, one box just isn't enough room for those that want to run higher quality/intelligent models. DeepSeek V4…
Tags: DeepSeek, ds4, local LLM, RDMA, Strix Halo, tensor parallelism
GLM 5.3 Flash running on a 128GB PC — featured banner
Posted inAMD Distributed Inference

GLM 5.3 Flash on 128GB: Two Easiest Ways to Run It — and When Two Strix Halos Beat One

GLM 5.3 Flash on one 128GB URAM machine, or even Two for higher quality Here's the headline that should make you sit up: a 320-billion-parameter multimodal model — one that…
Tags: ds4, GLM, local LLM, Quantization, Strix Halo
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top