And the 512GB Mac Studio that Apple just stopped selling. Apple discontinued the 512GB Mac Studio. Not the chip — the M3 Ultra is still on the shelf — but…
AIfinitee belives local compute as an alternative to cloud and that the most capable open models should be runnable on hardware you actually own. This post is the recipe for…
At AIFinitee, we have spent months benchmarking LLM inference on our dual-node AMD Ryzen AI MAX 395+ cluster. We have measured tokens per second across MiniMax-M2 and Qwen3.5-397B. We have…
At AIFinitee, we've spent months chasing tokens per second. Our two-node AMD Ryzen AI MAX cluster hits 17-20 tok/s with MiniMax-M2. Our Linux-vs-Windows benchmarks showed how your OS quietly taxes…
One of the beliefs we hold at AIfinitee is that LLMs will only get bigger. Bigger LLMs will mean more memory for usage in either cloud compute or local compute.…