基于 Niko1221/Strata v0.1.35 的优化分支,让 Qwen3.8-Flash-Next 在个人电脑上更合理地使用显存与系统内存。 我们重点保留经过本机验证的内存规划改动,并沿用上游的计算内核、FP8 ngram、服务接口与会话缓存。
| 分支 | 定位 | 实测设备 |
|---|---|---|
main |
单卡内存规划与上游 RAM 多会话缓存配合 | Windows,RTX 5090 32 GiB 显存,48 GB 系统内存 |
strata-2080tix2 |
双卡权重分配、RAM 专家常驻与动态交换 | Linux,两张扩容至 22 GiB 的 RTX 2080 Ti,32 GB 系统内存 |
双卡测试使用扩容卡,普通 11 GiB RTX 2080 Ti 的容量与性能需要另行验证。 两个分支的优化和测试范围各自独立。
主分支保留两项改动:
- 独立 prefill 缓冲选择。 填满专家缓存后,根据剩余显存估算独立预填充缓冲。
保留配置的显存余量,且批量不小于上游借用缓冲方案时才启用。无需移出专家来腾空间;
显存不足时自动沿用上游借用路径。正式配置仍使用
--expert-cache auto。 - 专家与会话的 RAM 预算配合。 启用 RAM 常驻专家和多会话缓存时,专家加载至少为 会话缓存预算、最低空闲 RAM 和额外 256 MiB 留出空间,避免两者分别按同一份空闲内存做预算。 已有更大的专家 headroom 会被保留;会话缓存关闭时,维持上游预算。
本机部署沿用上游的 RAM 多会话保存、恢复、隔离与淘汰,配置为 2 GiB、最多 2 个缓存槽, 最低空闲 RAM 2560 MiB;对应至少 4.75 GiB 的专家 headroom。配置不是安装器的全局默认值, 缓存可容纳的会话数量还取决于长度。前缀检查点与多会话缓存均来自上游。
双卡分支的改动: dense 权重按所属 GPU 的层加载;在 RAM 常驻两张 GPU 的专家补集; 动态交换在确认两张卡上传完成后提交专家归属。该分支配合上游的专用 prefill 缓冲与按层状态分配。 详见 双卡实现与测试记录。
主分支:2026-10-02,RTX 5090 / 48 GB RAM,IQ3_S、FP8 ngram、INT8 KV,配置上下文 262144。 以下是顺序对照,文件缓存与运行状态会影响结果。
| 测试 | 对照 | 优化版 | 含义 |
|---|---|---|---|
| 17944-token 完整 prefill,自动专家缓存、暖文件 | 上游 2925 tok/s | 2919 tok/s | 默认配置预填充速度接近 |
| 17944-token 完整 prefill,较小专家缓存、同一二进制 | 借用缓冲 2312 tok/s | 独立缓冲 2902 tok/s | 独立缓冲在有显存余量时值得保留 |
| 切回相同的 17944-token 提示 | — | 复用 17937 tokens,只重读 7 tokens | 上游多会话恢复有效 |
独立缓冲对照将专家缓存字节预算设为 5400 MiB,实际分配 7056 槽、13.39 GiB; 两组 prefill 批量均为 8192,独立缓冲约 3805 MiB,显存余量 1536 MiB。 这是一次顺序测试,不代表默认配置或所有提示都有同等收益。 解码复测波动较大,尚未证明主分支有稳定的解码提速。
双卡分支在上述扩容双卡上,单轮同提示测试中,13K prefill 从旧改造版的 1040.5 提升至 1207.4 tok/s,1024-token decode 从 40.3 提升至 51.5 tok/s。 没有清空 OS 文件缓存;详细条件与限制见双卡文档。
主分支已通过 11 项内存规划边界检查、30 项上游会话缓存测试,以及真实模型 A→B→A 状态一致性检查。 完整服务包含 GPU 视觉编码器的启动与 HTTP 会话恢复也已验证。 更多参数、测量和范围见 主分支优化说明。
服务可以接收多个客户端请求并排队,同一时刻执行 1 条推理。 RAM 多会话缓存加速不同历史之间的切换,不增加同时推理数量。 当前没有应用层队列长度上限或推理限流;已验证排队、断连取消与请求交接,尚未做高并发容量评测。 SSD 会话缓存暂不实现,RAM 缓存随引擎退出而丢失。
git clone https://github.com/spideytznn/Strata.git
cd Strata使用自定义改动需要从源码编译:Windows 运行 START-HERE.bat --build,Linux 运行
./setup.sh --build。已有安装调整构建时可加 --setup。安装器准备依赖与模型,并按设备构建引擎;
仅使用上游预编译引擎不会包含本仓库的 C++ 改动。
主分支自定义改动的实测平台为 Windows / CUDA;Linux 双卡使用对应分支。
本机已验证配置使用 IQ3_S、FP8 ngram、INT8 KV、MTP4、GPU 视觉与 262144 上下文上限;
上下文上限是配置值,不代表已填满 262K 实测。硬件、后台内存占用与模型大小都会影响可运行配置。
保留自己的模型路径,并在生成配置的 args 中按需加入:
--conversation-cache-mib 2048
--conversation-cache-slots 2
--conversation-cache-min-free-mib 2560
STRATA_PREFILL_OWN_AUTO=0 可关闭本地独立缓冲选择,用于对照。
安装步骤见 AI_SETUP,模型选择见 MODELS,
服务接口与完整参数见 DETAILS。模型、机器专用配置与编译产物不随源码上传。
原作者 Niko1221/Strata 提供基础引擎、计算内核、服务层、 前缀检查点与 RAM 多会话缓存。FP8 ngram 使用上游读取与打包实现,本仓库没有把这些功能标为自研。 模型来自 Qwen,量化版本来自 ISTA-DASLab; 底层还使用 llama.cpp / ggml。 源码遵循 MIT License,模型和依赖遵循各自许可。
An optimized fork of Niko1221/Strata v0.1.35 for running Qwen3.8-Flash-Next on personal computers with a practical balance of VRAM and system RAM. We retain measured memory-planning changes while using upstream compute kernels, FP8 ngram support, service APIs, and conversation caching.
| Branch | Focus | Tested hardware |
|---|---|---|
main |
Single-GPU memory planning with upstream RAM conversation caching | Windows, RTX 5090 with 32 GiB VRAM, 48 GB system RAM |
strata-2080tix2 |
Dual-GPU weight placement, resident RAM experts, and dynamic exchange | Linux, two RTX 2080 Ti cards modified to 22 GiB each, 32 GB system RAM |
The dual-GPU measurements use modified cards. Capacity and performance on ordinary 11 GiB RTX 2080 Ti cards require separate validation. Each branch has its own implementation and test scope.
Main retains two changes:
- Independent prefill buffer selection. After filling the expert cache, the planner prices independent
buffers against remaining free VRAM. It preserves the configured VRAM reserve and selects them only if
the batch is at least as large as upstream's borrowed-buffer option. It does not evict experts to make
room. If free VRAM is insufficient, upstream borrowing remains active. Production keeps
--expert-cache auto. - Coordinated expert and conversation RAM budgets. With resident RAM experts and conversation caching enabled, expert loading reserves at least the conversation budget plus the minimum free-RAM floor plus 256 MiB, so both allocations do not budget against the same free memory independently. Larger existing expert headroom is respected. Disabling conversation caching preserves upstream headroom.
Our local deployment uses upstream RAM conversation parking, restoration, isolation, and eviction with a 2 GiB budget and up to 2 slots, plus a 2560 MiB free-RAM floor. This implies at least 4.75 GiB of expert headroom. These are deployment settings, not installer-wide defaults. The number of conversations that fit depends on their lengths. Prefix checkpoints and conversation caching are upstream features.
The dual-GPU branch loads dense weights for each GPU's assigned layers, keeps the combined expert complement resident in RAM, and commits dynamic-exchange ownership only after both GPUs confirm uploads. It uses upstream dedicated prefill buffers and layer-owned state allocation. See the dual-GPU implementation and measurements.
Main: 2026-10-02, RTX 5090 / 48 GB RAM, IQ3_S, FP8 ngram, INT8 KV, configured context limit 262144. These are sequential comparisons affected by file caching and runtime conditions.
| Test | Reference | Optimized | Interpretation |
|---|---|---|---|
| Full 17944-token prefill, auto expert cache, warm files | Upstream 2925 tok/s | 2919 tok/s | Similar default prefill speed |
| Full 17944-token prefill, smaller expert cache, same binary | Borrowed buffers 2312 tok/s | Independent buffers 2902 tok/s | The independent option is useful when free VRAM permits |
| Return to the same 17944-token prompt | — | 17937 tokens reused, only 7 reread | Upstream conversation restoration works |
The independent-buffer comparison used a 5400 MiB expert-cache byte budget, yielding 7056 sized slots and 13.39 GiB of actual allocation. Both arms used 8192-token batches; independent buffers used about 3805 MiB, preserving a 1536 MiB VRAM reserve. This single sequential pair does not establish the same improvement for default settings or all prompts. Repeated decode measurements varied substantially; a stable main-branch decode speedup has not been established.
On the modified dual-GPU machine, one matched-prompt run improved 13K prefill from 1040.5 in the old custom version to 1207.4 tok/s, and 1024-token decode from 40.3 to 51.5 tok/s. OS file caches were not cleared. Conditions and limitations are documented on that branch.
Main passes 11 planner boundary checks, 30 upstream conversation-cache tests, and real-model A→B→A state parity checks. Full-service startup with the GPU vision encoder and HTTP conversation restoration also pass. See main-branch implementation and validation for detailed parameters and scope.
Multiple clients can submit requests and wait in the queue; one inference request runs at a time. RAM conversation caching speeds up switching histories without adding parallel inference slots. There is currently no application-level queue-length cap or inference rate limit. Queueing, disconnect cancellation, and request handover are checked; high-concurrency capacity has not been measured. SSD session caching is deferred. RAM snapshots are lost when the engine exits.
git clone https://github.com/spideytznn/Strata.git
cd StrataBuild from source to use the custom changes: run START-HERE.bat --build on Windows or ./setup.sh --build
on Linux. Add --setup when changing an existing installation's build/settings. Setup prepares dependencies
and the model and builds for your GPU. An upstream prebuilt engine does not contain this fork's C++ changes.
Main's custom changes are tested on Windows / CUDA; use the separate branch for the tested Linux dual-GPU setup.
The local validated deployment uses IQ3_S, FP8 ngram, INT8 KV, MTP4, GPU vision, and a 262144 context limit.
The context limit is configured, not a full-262K input measurement. Hardware, background RAM use, and model
size determine what fits. Keep your own model paths and optionally add these engine args to the generated config:
--conversation-cache-mib 2048
--conversation-cache-slots 2
--conversation-cache-min-free-mib 2560
STRATA_PREFILL_OWN_AUTO=0 disables the local independent-buffer choice for comparison.
See AI_SETUP for installation, MODELS for model choices, and
DETAILS for service APIs and all options. Models, machine-specific configs, and build
artifacts are not included in the source repository.
Niko1221/Strata provides the base engine, compute kernels, service layer, prefix checkpoints, and RAM conversation cache. FP8 ngram uses upstream readers and packing tools; these are credited to upstream rather than presented as our additions. The model is from Qwen, with quantized versions from ISTA-DASLab. Strata also uses llama.cpp / ggml. Source is under the MIT License; models and dependencies retain their own licenses.