Skip to content
logo AIFinitee

All things AI & LLM

  • All Things AI & LLM
  • Archive
  • Strix Halo
  • Apple Silicon
  • Distributed Inference
Subscribe

Cluster Computing

Home » Cluster Computing
DeepSeek V4 Flash Q2 running on Strix Halo 128GB — featured banner
Posted inAMD Distributed Inference

Running DeepSeek V4 Flash Q2 on Strix Halo 128GB

AIfinitee believes in local compute as an alternative to cloud and that the most capable open models should be runnable on hardware you actually own. This post is the recipe for…

Tags: Cluster Computing, DeepSeek, LLM Inference, Local AI, Quantization, Strix Halo
Two-node AMD Ryzen AI MAX Strix Halo cluster — featured banner
Posted inAMD Distributed Inference MiniMax

Strix Halo Cluster

One of the beliefs we hold at AIfinitee is that LLMs will only get bigger. Bigger LLMs will mean more memory for usage in either cloud compute or local compute.…
Tags: Benchmarking, Cluster Computing, LLM Inference, Local AI, Strix Halo
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top