如果 2 - 3 万人民币设备,部署本地 dsv4flash , token 基本够用,你是会选择本地部署还是接云 api?

查看 74|回复 9
作者:foryou2023   
The current q2 results use ds4-bench with the standard Promessi sposi
input, 2048-token context steps, and 128 greedy generation tokens at every
frontier. Each prefill number is for the next 2048-token chunk. The complete
sweeps are in m5_max.csv and
gb10.csv.
[td]Machine[/td]
[td]Backend[/td]
[td]Context[/td]
[td]Prefill[/td]
[td]Generation[/td]
MacBook Pro M5 Max, 128 GB
Metal
2048
790.18 t/s
39.35 t/s
MacBook Pro M5 Max, 128 GB
Metal
16384
572.53 t/s
36.14 t/s
MacBook Pro M5 Max, 128 GB
Metal
32768
557.04 t/s
34.36 t/s
MacBook Pro M5 Max, 128 GB
Metal
65536
398.50 t/s
27.64 t/s
DGX Spark GB10, 128 GB
CUDA
2048
825.76 t/s
18.05 t/s
DGX Spark GB10, 128 GB
CUDA
16384
872.44 t/s
15.10 t/s
DGX Spark GB10, 128 GB
CUDA
32768
855.94 t/s
14.43 t/s
DGX Spark GB10, 128 GB
CUDA
65536
822.98 t/s
13.84 t/s
Older measurements for machines and model variants not rerun in this pass are
kept for reference. They used the earlier CLI prompt procedure and are not
directly comparable with the table above.
[td]Machine[/td]
[td]Quant[/td]
[td]Prompt[/td]
[td]Prefill[/td]
[td]Generation[/td]
MacBook Pro M3 Max, 128 GB
q2
short
58.52 t/s
26.68 t/s
MacBook Pro M3 Max, 128 GB
q2
11709 tokens
250.11 t/s
21.47 t/s
Mac Studio M3 Ultra, 512 GB
q2
short
84.43 t/s
36.86 t/s
Mac Studio M3 Ultra, 512 GB
q2
11709 tokens
468.03 t/s
27.39 t/s
Mac Studio M3 Ultra, 512 GB
q4
short
78.95 t/s
35.50 t/s
Mac Studio M3 Ultra, 512 GB
q4
12018 tokens
448.82 t/s
26.62 t/s
Mac Studio M3 Ultra, 512 GB
PRO q2
32768 tokens
138.82 t/s
9.56 t/s
以上是看 https://github.com/antirez/ds4 这个仓库的数据。
最近 dsv4 涨价太猛了,所以尝试找找本地部署的方案,找到这个方案,由于手里也没有设备同时也不太懂这个 速度的问题,也无法测试具体的实践效果如何,不知道有相关设备的朋友,能不能帮忙看看或者测试一下。
目前看了咸鱼 M3 MAX 128G 的大概 2.5 万左右,感觉按照 dsv4 flash 的价格的话可以接受本地部署了。M5 MAX 苹果官网看了一下要 5 万多,暂时没有这么多预算。
看了大家的大部分人的回复,本地部署暂时不考虑,只能继续上云 api 了。
        
   
   
   

部署, 本地, 性能

donaldturinglee   
看了大家的大部分人的回复,本地部署暂时不考虑,只能继续上云 api 了。
cvooc   
2.5 万部署的能用在工作吗?生产环境还是直接上 API 吧
FakerLeung   
现在迭代这么快, 本地部署 每秒几十 token 只能个人同一时间跑单任务用, 除非极度缺乏安全感,
真不如 api 跑批量来的爽.
xing7673   
我看双 DGX 好像能干到 4 人同时 50t/s ?还是 FP8 的精度?
darksword21   
你这又是 q2 又是几十 k 的上下文,部署起来也基本没啥用
crocoBaby   
没有意义,发这种 bench 的人也是闲的
op351   
很多人早算过这笔帐,肯定是云 api 划算的
ff521   
现阶段肯定是直接买云服务商的算力  
云服务商都还在算力设备采购难,到货周期长,回本周期长的阶段,就不用想本地部署省钱这码事了  
现在本地部署的好处就一个 安全合规  
不存在本地部署吊打云服务商的算力和质量还能省钱的
donaldturinglee   
hx170 这个破解的显卡有人玩过吗
您需要登录后才可以回帖 登录 | 立即注册

返回顶部