Posted inAMD Distributed Inference
Repeat after me: Good things come in TP=2 (DeepSeek V4 Flash Q4 on 2x Strix Halos RDMA)
Models continue to improve in quality for a similar memory size. Sometimes, though, one box just isn't enough room for those that want to run higher quality/intelligent models. DeepSeek V4…
