DSH Plugins
返回列表
🤖

dsh-llama-model-manager

模型推理 更新于 2026.09.12

在终端中运行以下命令:

dsh plugin install DoctorxPriestess/dsh-llama-model-manager

将以下提示词粘贴到 DeepSeek Harness 对话框中:

在 DeepSeek Harness 终端中执行 dsh plugin install DoctorxPriestess/dsh-llama-model-manager 即可安装本插件,源码仓库地址为 https://github.com/DoctorxPriestess/dsh-llama-model-manager 。

插件介绍

Running local llama.cpp GGUF models inside DSH works fine until you want to switch to a different one. Stop the server, edit the provider config, restart, and hope no request was mid-flight. On Windows the situation gets worse: Node's child.kill(SIGINT) silently compiles down to TerminateProcess, so llama.cpp never gets a chance to free its model and you lose a dozen gigabytes of VRAM every single time. dsh-llama-model-manager absorbs the entire lifecycle-start, stop, switch, recover-behind one stable OpenAI-compatible gateway, so DSH always talks to a single fixed URL while the model underneath changes freely.

Model switching is guarded by a serialization gate: in-flight inference holds a shared ticket, a switch request acquires an exclusive one and waits for every outstanding request to drain before the connection is torn down. Stopping a model delivers a genuine Ctrl+C console control event through a hidden PowerShell helper that performs the P/Invoke dance, letting llama.cpp run its own cleanup and call llama_model_free. Measured end-to-end on a 27B model: clean exit code zero, full VRAM returned, no window ever flashes. If DSH crashes mid-session, the plugin cross-checks the recorded pid, executable name, bound port, and reported model path before it will touch a leftover process, and it refuses to act on any pid it cannot positively attribute.

Built for Windows users who run local inference through DSH and switch between two or three GGUF models on a regular basis. It never reads or writes DSH's settings.yaml, ships zero npm dependencies and no build step, and exposes a live settings page where you can see the current model, tail logs, and start, stop, switch, or restart with one click. All you need is a llama-server.exe build and your .gguf files.

使用场景

  • 在 DSH 中一键切换多个本地 GGUF 模型而无需修改 provider 配置
  • 安全停止 llama-server 并完整释放显存,避免进程残留占用资源
  • 为 DSH 提供稳定的 OpenAI 兼容端点,后端模型可随时更换

适合人员

  • 在 Windows 上使用 DSH 运行本地 llama.cpp 推理的用户
  • 需要频繁切换多个 GGUF 模型的本地部署用户
  • 希望避免显存泄漏和进程残留的开发者