Skip to content
logo AIFinitee

All things AI & LLM

  • All Things AI & LLM
  • Archive
  • Strix Halo
  • Apple Silicon
  • Distributed Inference
Subscribe

Benchmarking

Home ยป Benchmarking
Speed versus smarts: choosing the right local model size โ€” featured banner
Posted inAMD Distributed Inference

Speed vs. Smarts: When Bigger Models Win for Local AI Coding

At AIFinitee, we've spent months chasing tokens per second. Our two-node AMD Ryzen AI MAX cluster hits 17-20 tok/s with MiniMax-M2. Our Linux-vs-Windows benchmarks showed how your OS quietly taxes…
Tags: Benchmarking, LLM Inference, Local AI, Strix Halo
Linux vs Windows for LLM inference โ€” featured banner
Posted inAMD Distributed Inference

Linux vs Windows for LLM Inference: More Tokens/Second on Linux

Same GPU, different OS = different performance. See why Linux delivers 5-30% more tokens/sec than Windows across NVIDIA, AMD, and Intel GPUs.
Tags: Benchmarking, Linux, LLM Inference, Local AI, Strix Halo
Two-node AMD Ryzen AI MAX Strix Halo cluster โ€” featured banner
Posted inAMD Distributed Inference MiniMax

Strix Halo Cluster

One of the beliefs we hold at AIfinitee is that LLMs will only get bigger. Bigger LLMs will mean more memory for usage in either cloud compute or local compute.…
Tags: Benchmarking, Cluster Computing, LLM Inference, Local AI, Strix Halo
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top