Skip to content
logo

AIFinitee

All things AI & LLM

  • All Things AI & LLM
  • Archive
  • Strix Halo
  • Apple Silicon
  • Distributed Inference
Subscribe
Top Stories
Chibi Aimee with purple cat-clip pigtails and yellow eyes rides a dark mini-PC like a horse at full gallop down a floodlit stadium track, a huge plume of dust and speed lines trailing behind at golden hour
Strix Halo: Value Laggard becomes competent Qwen 3.8 Flash Next with Halogen (450+ prefill & 35+ tps)
Crying chibi Aimee with purple pigtails wails with her mouth wide open as glowing data and folder icons stream out of her mini-PC in a cyan arc toward the cloud, dark purple background
PSA: ZCode caught uploading Git repos without permission! The Case for Open-Source Harnesses & MiniMax Code Release
Aimee, the AIFinitee chibi mascot, reads a fairy-tale storybook and says Once Upon... while a cheerful chibi onion wearing round glasses replies A Time - a playful nod to engram lookup memory.
Once Upon a Time: DeepSeek V4.1-Flash and Qwen3.8-Flash-Next – What the hekk is an Engram?!
Banner: two purple-haired chibi anime girls affectionately hugging two AMD mini-PCs with glowing cyan accents, dark purple background
Repeat after me: Good things come in TP=2 (DeepSeek V4 Flash Q4 on 2x Strix Halos RDMA)
GLM 5.3 Flash running on a 128GB PC — featured banner
GLM 5.3 Flash on 128GB: Two Easiest Ways to Run It — and When Two Strix Halos Beat One
MiniMax H3 prompting guide — featured banner
MiniMax H3: Your guide to proper prompting for better videos
GLM 5.2 at 4-bit with MTP on Mac Studio — featured banner
GLM 5.2 at 4-Bit With MTP: The Easiest Way In
DeepSeek V4 Flash Q2 running on Strix Halo 128GB — featured banner
Running DeepSeek V4 Flash Q2 on Strix Halo 128GB
ComfyUI artwork generated on AMD Ryzen AI MAX 96GB — featured banner
ComfyUI on AMD Ryzen AI MAX: 96 GB Unified Memory vs 16 GB NVIDIA
Speed versus smarts: choosing the right local model size — featured banner
Speed vs. Smarts: When Bigger Models Win for Local AI Coding
Linux vs Windows for LLM inference — featured banner
Linux vs Windows for LLM Inference: More Tokens/Second on Linux
Two-node AMD Ryzen AI MAX Strix Halo cluster — featured banner
Strix Halo Cluster
Chibi Aimee with purple cat-clip pigtails and yellow eyes rides a dark mini-PC like a horse at full gallop down a floodlit stadium track, a huge plume of dust and speed lines trailing behind at golden hour
Posted inAMD Distributed Inference

Strix Halo: Value Laggard becomes competent Qwen 3.8 Flash Next with Halogen (450+ prefill & 35+ tps)

We have wanted to write about Qwen 3.8 Flash Next for quite a while, and yet, we haven't been doing that — because nothing we ran it on made it…
Continue Reading
Posted by
Crying chibi Aimee with purple pigtails wails with her mouth wide open as glowing data and folder icons stream out of her mini-PC in a cyan arc toward the cloud, dark purple background
Posted inMiniMax

PSA: ZCode caught uploading Git repos without permission! The Case for Open-Source Harnesses & MiniMax Code Release

ZCode silently packaged Git histories and uploaded them to Aliyun OSS with a server-held key. What we verified, how to block it, and the open-source fix.
Continue Reading
Posted by
Aimee, the AIFinitee chibi mascot, reads a fairy-tale storybook and says Once Upon... while a cheerful chibi onion wearing round glasses replies A Time - a playful nod to engram lookup memory.
Posted inPrimer

Once Upon a Time: DeepSeek V4.1-Flash and Qwen3.8-Flash-Next – What the hekk is an Engram?!

Two of the most interesting open-model releases of the quarter just landed, and they share a spec-sheet line that has a lot to do with benchmarks - but ignores the…
Continue Reading
Posted by
Banner: two purple-haired chibi anime girls affectionately hugging two AMD mini-PCs with glowing cyan accents, dark purple background
Posted inAMD Distributed Inference

Repeat after me: Good things come in TP=2 (DeepSeek V4 Flash Q4 on 2x Strix Halos RDMA)

Models continue to improve in quality for a similar memory size. Sometimes, though, one box just isn't enough room for those that want to run higher quality/intelligent models. DeepSeek V4…
Continue Reading
Posted by
GLM 5.3 Flash running on a 128GB PC — featured banner
Posted inAMD Distributed Inference

GLM 5.3 Flash on 128GB: Two Easiest Ways to Run It — and When Two Strix Halos Beat One

GLM 5.3 Flash on one 128GB URAM machine, or even Two for higher quality Here's the headline that should make you sit up: a 320-billion-parameter multimodal model — one that…
Continue Reading
Posted by
MiniMax H3 prompting guide — featured banner
Posted inMiniMax

MiniMax H3: Your guide to proper prompting for better videos

As much as cloud remains a dominant option for the casual consumer, it is worth to remember that AI (LLM or units of compute) are also units of Intellect. Ownership…
Continue Reading
Posted by
You May Have Missed
Chibi Aimee with purple cat-clip pigtails and yellow eyes rides a dark mini-PC like a horse at full gallop down a floodlit stadium track, a huge plume of dust and speed lines trailing behind at golden hour
Posted inAMD Distributed Inference

Strix Halo: Value Laggard becomes competent Qwen 3.8 Flash Next with Halogen (450+ prefill & 35+ tps)

Posted by
Crying chibi Aimee with purple pigtails wails with her mouth wide open as glowing data and folder icons stream out of her mini-PC in a cyan arc toward the cloud, dark purple background
Posted inMiniMax

PSA: ZCode caught uploading Git repos without permission! The Case for Open-Source Harnesses & MiniMax Code Release

Posted by
Aimee, the AIFinitee chibi mascot, reads a fairy-tale storybook and says Once Upon... while a cheerful chibi onion wearing round glasses replies A Time - a playful nod to engram lookup memory.
Posted inPrimer

Once Upon a Time: DeepSeek V4.1-Flash and Qwen3.8-Flash-Next – What the hekk is an Engram?!

Posted by
Banner: two purple-haired chibi anime girls affectionately hugging two AMD mini-PCs with glowing cyan accents, dark purple background
Posted inAMD Distributed Inference

Repeat after me: Good things come in TP=2 (DeepSeek V4 Flash Q4 on 2x Strix Halos RDMA)

Posted by
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top