Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
aappleby
7 months ago
|
parent
|
context
|
favorite
| on:
Show HN: Run Qwen3-Next-80B on 8GB GPU at 1tok/2s ...
Why even bother with the GPU at that point? CPU would be just as fast if you're bottlenecked on SSD bandwidth.
anuarsh
7 months ago
[–]
CPU is much slower than GPU. You can actually use both by offloading some layers to CPU as o.offload_layers_to_cpu(layers_num=12). It is faster to load from RAM than from SSD.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: