sparkle/benchmarks/reports/8bit-g64-baseline-2026-07-25.md
fiser_jun 4745f264b2
2026-08-04 14:34:00 +08:00

62 lines
3.0 KiB
Markdown

# Qwen3-Next-80B-A3B 8bit(g64)基线 — M2 Max 64GB
日期:2026-07-25 · 机器:MacBook Pro M2 Max / 64GB / 内置 SSD
模型:`models/Qwen3-Next-80B-A3B-Instruct-MLX-8bit`(**bits=8, group_size=64**,79GB)
数据:`models/experts_8bit_g64/blobs`(48 层,77GB,v1 affine,stride 3,342,336B,page-aligned)+ `models/qn_mtp_weights.safetensors`(3.3GB)
## 性能
| 场景 | tok/s | accept_len | 备注 |
|---|---|---|---|
| 冷启动首轮(32 tok) | 5.0 | 1.889 | 含 kernel 编译 + 池冷填充 |
| 稳态(128 tok,64 槽) | **9.4** | 2.246 | 符合 8-12 tok/s 预期 |
| server 短请求(33 tok,含 prefill) | 6.6 | — | OpenAI 链路实测 |
- 常驻内存:RSS ≈ **23.5 GB**(`--expert-slots 64`,池 ~10GB + dense + 缓冲;64GB 机器余量充足,可上探 96-128 槽)
- 加载时间:约 60-90s(含预热)
## 正确性(三层证据链)
1. **kernel 单元等价**(`verify_8bit_equiv.py`):L0/L15/L47 blob fused kernel vs Python 参考 **diff = 0**
2. **端到端对拍**(`verify_e2e_spec.py`):贪心 vs 投机 token 17 确定性分叉 → 触发对照
3. **KV 量化对照**(`verify_teacher_forced.py`,seq=1 vs seq=2 生产原语):
- `KV_QUANT=0`:**24/24 token 全一致**,cosine 0.99935 → 8bit 权重路径数值干净
- `KV_QUANT=1`(生产默认):23/24(95.8%),cosine 0.99870,翻转处 margin=0 近 tie → 达质量门(≥95%/≥0.99)
- 结论:分叉全部归因于 KV 量化(KV 量化既有有损特性),与权重量化位宽无关
## 已知事项
- `VirtualPool.acquire_host` latent bug:单次前向单层唯一专家数超池容量时 fetch 回退会 AttributeError;生产 prefill 分块 chunk=2 永不触发。勿用整段 seq=N 一次性前向做诊断。
- server `/admin` 调参页池内存估算已按 g64 修正为 ~160MB/槽。
- server 模型 id 取目录名:`Qwen3-Next-80B-A3B-Instruct-MLX-8bit`。
## 复现命令
```bash
# CLI
.venv/bin/sparkle --model models/Qwen3-Next-80B-A3B-Instruct-MLX-8bit \
--expert-dir models/experts_8bit_g64 \
--qn-config models/Qwen3-Next-80B-A3B-Instruct-MLX-8bit/config.json \
--mtp-out models/qn_mtp_weights.safetensors --expert-slots 64 --stats
# OpenAI server(前端 Sparkle 源对接,端口 8317)
.venv/bin/python -m mlx_streaming.server --port 8317 \
--model models/Qwen3-Next-80B-A3B-Instruct-MLX-8bit \
--expert-dir models/experts_8bit_g64 \
--qn-config models/Qwen3-Next-80B-A3B-Instruct-MLX-8bit/config.json \
--mtp-out models/qn_mtp_weights.safetensors --expert-slots 64
```
## 追记(2026-07-25 晚):AUTOPIN + 调优后
| 配置 | decode tok/s | 备注 |
|---|---|---|
| 64 槽 + AUTOPIN | ~8.0-9.4 | 稳定,RSS ~24GB |
| **96 槽 + AUTOPIN(生产)** | **9.3-10.3** | 甜点位,RSS ~32GB;CLI 128tok: 9.3(MTP4)/9.7(MTP8) |
| 160 槽 + AUTOPIN | 不可用 | RSS ~34GB 系统耗尽卡死(AUTOPIN eager 预热多占 ~5GB) |
- AUTOPIN(热度持久化+启动 pin 定死):命中率 72.2%→98.6%,pin 3439 专家/6.5s
- PREFILL_CHUNK=8:676 tok prompt 2m49s→33s
- MTP_BITS 4→8 A/B:accept_len 2.246→2.207,无效,保持 4
- 槽位调整改进程级重启(进程内重建在新 C++ pin 状态下会把引擎拖到 0.24 tok/s)