Skip to content
logo AIFinitee

All things AI & LLM

  • Home
  • Articles
  • Blog
Subscribe

Cluster Computing

Home ยป Cluster Computing
Running DeepSeek V4 Flash Q2 on Strix Halo 128GB
Posted inAMD Distributed Inference

Running DeepSeek V4 Flash Q2 on Strix Halo 128GB

AIfinitee belives local compute as an alternative to cloud and that the most capable open models should be runnable on hardware you actually own. This post is the recipe for…
Tags: AMD, Cluster Computing, DeepSeek, LLM Inference, Local AI, Quantization, Strix Halo
AMD Ryzen AI MAX Two-Node Cluster Guide | LLM Setup (17-20 tok/s)
Posted inAMD Distributed Inference MiniMax

AMD Ryzen AI MAX Two-Node Cluster Guide | LLM Setup (17-20 tok/s)

One of the beliefs we hold at AIfinitee is that LLMs will only get bigger. Bigger LLMs will mean more memory for usage in either cloud compute or local compute.…
Tags: AMD, Benchmarking, Cluster Computing, LLM Inference, Local AI, Strix Halo
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top