Skip to content
logo AIFinitee

All things AI & LLM

  • All Things AI & LLM
  • Archive
  • Strix Halo
  • Apple Silicon
  • Distributed Inference
Subscribe

Strix Halo

Home » Strix Halo
Aimee, the AIFinitee chibi mascot, reads a fairy-tale storybook and says Once Upon... while a cheerful chibi onion wearing round glasses replies A Time - a playful nod to engram lookup memory.
Posted inPrimer

Once Upon a Time: DeepSeek V4.1-Flash and Qwen3.8-Flash-Next – What the hekk is an Engram?!

Two of the most interesting open-model releases of the quarter just landed, and they share a spec-sheet line that has a lot to do with benchmarks - but ignores the…
Tags: DeepSeek, Engram, Local AI, local LLM, Qwen, Strix Halo
Banner: two purple-haired chibi anime girls affectionately hugging two AMD mini-PCs with glowing cyan accents, dark purple background
Posted inAMD Distributed Inference

Repeat after me: Good things come in TP=2 (DeepSeek V4 Flash Q4 on 2x Strix Halos RDMA)

Models continue to improve in quality for a similar memory size. Sometimes, though, one box just isn't enough room for those that want to run higher quality/intelligent models. DeepSeek V4…
Tags: DeepSeek, ds4, local LLM, RDMA, Strix Halo, tensor parallelism
GLM 5.3 Flash running on a 128GB PC — featured banner
Posted inAMD Distributed Inference

GLM 5.3 Flash on 128GB: Two Easiest Ways to Run It — and When Two Strix Halos Beat One

GLM 5.3 Flash on one 128GB URAM machine, or even Two for higher quality Here's the headline that should make you sit up: a 320-billion-parameter multimodal model — one that…
Tags: ds4, GLM, local LLM, Quantization, Strix Halo
MiniMax H3 prompting guide — featured banner
Posted inMiniMax

MiniMax H3: Your guide to proper prompting for better videos

As much as cloud remains a dominant option for the casual consumer, it is worth to remember that AI (LLM or units of compute) are also units of Intellect. Ownership…
Tags: LLM Inference, Local AI, Strix Halo
DeepSeek V4 Flash Q2 running on Strix Halo 128GB — featured banner
Posted inAMD Distributed Inference

Running DeepSeek V4 Flash Q2 on Strix Halo 128GB

AIfinitee believes in local compute as an alternative to cloud and that the most capable open models should be runnable on hardware you actually own. This post is the recipe for…

Tags: Cluster Computing, DeepSeek, LLM Inference, Local AI, Quantization, Strix Halo
ComfyUI artwork generated on AMD Ryzen AI MAX 96GB — featured banner
Posted inAMD Distributed Inference

ComfyUI on AMD Ryzen AI MAX: 96 GB Unified Memory vs 16 GB NVIDIA

At AIFinitee, we have spent months benchmarking LLM inference on our dual-node AMD Ryzen AI MAX 395+ cluster. We have measured tokens per second across MiniMax-M2 and Qwen3.5-397B. We have…
Tags: ComfyUI, LLM Inference, Local AI, Strix Halo
Speed versus smarts: choosing the right local model size — featured banner
Posted inAMD Distributed Inference

Speed vs. Smarts: When Bigger Models Win for Local AI Coding

At AIFinitee, we've spent months chasing tokens per second. Our two-node AMD Ryzen AI MAX cluster hits 17-20 tok/s with MiniMax-M2. Our Linux-vs-Windows benchmarks showed how your OS quietly taxes…
Tags: Benchmarking, LLM Inference, Local AI, Strix Halo
Linux vs Windows for LLM inference — featured banner
Posted inAMD Distributed Inference

Linux vs Windows for LLM Inference: More Tokens/Second on Linux

Same GPU, different OS = different performance. See why Linux delivers 5-30% more tokens/sec than Windows across NVIDIA, AMD, and Intel GPUs.
Tags: Benchmarking, Linux, LLM Inference, Local AI, Strix Halo
Two-node AMD Ryzen AI MAX Strix Halo cluster — featured banner
Posted inAMD Distributed Inference MiniMax

Strix Halo Cluster

One of the beliefs we hold at AIfinitee is that LLMs will only get bigger. Bigger LLMs will mean more memory for usage in either cloud compute or local compute.…
Tags: Benchmarking, Cluster Computing, LLM Inference, Local AI, Strix Halo
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top