input, 2048-token context steps, and 128 greedy generation tokens at every
frontier. Each prefill number is for the next 2048-token chunk. The complete
sweeps are in m5_max.csv and
gb10.csv.
[td]Machine[/td]
[td]Backend[/td]
[td]Context[/td]
[td]Prefill[/td]
[td]Generation[/td]
MacBook Pro M5 Max, 128 GB
Metal
2048
790.18 t/s
39.35 t/s
MacBook Pro M5 Max, 128 GB
Metal
16384
572.53 t/s
36.14 t/s
MacBook Pro M5 Max, 128 GB
Metal
32768
557.04 t/s
34.36 t/s
MacBook Pro M5 Max, 128 GB
Metal
65536
398.50 t/s
27.64 t/s
DGX Spark GB10, 128 GB
CUDA
2048
825.76 t/s
18.05 t/s
DGX Spark GB10, 128 GB
CUDA
16384
872.44 t/s
15.10 t/s
DGX Spark GB10, 128 GB
CUDA
32768
855.94 t/s
14.43 t/s
DGX Spark GB10, 128 GB
CUDA
65536
822.98 t/s
13.84 t/s
Older measurements for machines and model variants not rerun in this pass are
kept for reference. They used the earlier CLI prompt procedure and are not
directly comparable with the table above.
[td]Machine[/td]
[td]Quant[/td]
[td]Prompt[/td]
[td]Prefill[/td]
[td]Generation[/td]
MacBook Pro M3 Max, 128 GB
q2
short
58.52 t/s
26.68 t/s
MacBook Pro M3 Max, 128 GB
q2
11709 tokens
250.11 t/s
21.47 t/s
Mac Studio M3 Ultra, 512 GB
q2
short
84.43 t/s
36.86 t/s
Mac Studio M3 Ultra, 512 GB
q2
11709 tokens
468.03 t/s
27.39 t/s
Mac Studio M3 Ultra, 512 GB
q4
short
78.95 t/s
35.50 t/s
Mac Studio M3 Ultra, 512 GB
q4
12018 tokens
448.82 t/s
26.62 t/s
Mac Studio M3 Ultra, 512 GB
PRO q2
32768 tokens
138.82 t/s
9.56 t/s
以上是看 https://github.com/antirez/ds4 这个仓库的数据。
最近 dsv4 涨价太猛了,所以尝试找找本地部署的方案,找到这个方案,由于手里也没有设备同时也不太懂这个 速度的问题,也无法测试具体的实践效果如何,不知道有相关设备的朋友,能不能帮忙看看或者测试一下。
目前看了咸鱼 M3 MAX 128G 的大概 2.5 万左右,感觉按照 dsv4 flash 的价格的话可以接受本地部署了。M5 MAX 苹果官网看了一下要 5 万多,暂时没有这么多预算。
看了大家的大部分人的回复,本地部署暂时不考虑,只能继续上云 api 了。

