Rendered at 01:57:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Tsiklon 10 hours ago [-]
This looks like a very cool high end x86 workstation. CPU alone is $12000 retail.
Theres a few potential “problems” I see with it - main system RAM is about 30-40x slower than that of the GPU’s and you will have PCI-e as a bottleneck getting working data to or between the GPUs.
In practice is the 400-500GB/s main memory limiting some use cases, or is the PCI-E bus the main issue with these types of systems?
kimixa 7 hours ago [-]
Realistically the models will always be resident in the GPU during processing - the CPU and it's memory would probably more be used for related development tasks (as this is a workstation not a cloud server).
If it's just running the same model all day that CPU and it's RAM would be pretty wasted.
happyPersonR 2 hours ago [-]
lol $12k is basically the main problem
5 hours ago [-]
Neywiny 11 hours ago [-]
So there's a path to 4 in order to run trillion parameter models? But you can't buy them or fit them in the case? Maybe in the Halo 2...
That all said, running frontier models locally offline would be very nice
ilaksh 11 hours ago [-]
It's over $100k.
Neywiny 9 hours ago [-]
Cheaper than an entry level human. But I've never used a model that size for development so idk if they could replace it. The work I do with even nemotron 550b, it's so dumb it's infuriating. Same with the free Google one (for anything sanitized/generic)
ilaksh 5 hours ago [-]
Try Qwen 3.8 27b or the antirez ds4 0731 quants
Can run on MacBook m5 or new m5 studio or m5 studio ultra or rtx 6000 pro.
Neywiny 3 hours ago [-]
I have it on my 7900xtx but haven't compared it just yet. I might though. Honestly what I need is a lack of knowledge trained into the model. Maybe 27b will be better. I need it to want ground truth data, not hallucinate based on quantized distilled documents from training.
karmakaze 9 hours ago [-]
The previous 'personal' AI Station I had pictured was the a16z one[0].
Each MI350P[1] in the TR Halo Station has 144GB VRAM and with 4.6 PFLOPs peak MXFP6 performance.
Theres a few potential “problems” I see with it - main system RAM is about 30-40x slower than that of the GPU’s and you will have PCI-e as a bottleneck getting working data to or between the GPUs.
In practice is the 400-500GB/s main memory limiting some use cases, or is the PCI-E bus the main issue with these types of systems?
If it's just running the same model all day that CPU and it's RAM would be pretty wasted.
That all said, running frontier models locally offline would be very nice
Can run on MacBook m5 or new m5 studio or m5 studio ultra or rtx 6000 pro.
Each MI350P[1] in the TR Halo Station has 144GB VRAM and with 4.6 PFLOPs peak MXFP6 performance.
Four liquid cooled? Yes please.
[0] https://a16z.com/building-a16zs-personal-ai-workstation-with...
[1] https://www.amd.com/en/products/accelerators/instinct/mi350/...