dsh-model-jury
在终端中运行以下命令:
dsh plugin install gjjkbssg/dsh-model-jury
将以下提示词粘贴到 DeepSeek Harness 对话框中:
在 DeepSeek Harness 终端执行 dsh plugin install gjjkbssg/dsh-model-jury 完成安装,源码位于 https://github.com/gjjkbssg/dsh-model-jury 。
插件介绍
Most model-council setups share a blind spot: several models give independent answers, then another model is asked to judge them—so the final verdict still rests on a single LLM's subjective take. dsh-model-jury removes that bottleneck by making code the judge. Three seats answer blind, critique each other anonymously, revise, and a deterministic aggregator computes quorum, vote state, and keeps dissent and critical risks visible to the user at all times.
The protocol runs in three rounds. Round 1 dispatches the same question, instructions, and JSON schema to all seats concurrently; no seat sees another response. A random per-run mapping then assigns responses to P1, P2, P3 position labels. Round 2 redistributes the anonymized positions after scrubbing provider, model, and product identity terms, and asks each model to point out strongest points, weakest points, missing evidence, and actual disagreement—without rewarding consensus. Round 3 lets each surviving model revise, merge, stay undecided, or preserve dissent. The aggregator accepts only P1, P2, P3, hybrid, or undecided, computes vote state and quorum deterministically, and surfaces any critical-risk flag prominently. The plugin is strictly deliberation-only: it never edits files, installs dependencies, commits, deploys, or performs network mutations. GLM and DeepSeek seats receive no tools; the Codex seat runs with permissionMode: never.
The reference deployment seats GPT/Codex, GLM, and DeepSeek, but the protocol is provider-neutral—any compatible model can be routed through the same blind, anonymous, structured state machine via a CouncilSeat implementation. It suits developers and teams working in DeepSeek Harness who want structured multi-perspective evaluation for architecture reviews, strategy decisions, or complex technical questions without a single model acting as the final authority. Every run writes owner-only trace files containing structured responses, safe call metadata, prompts, aggregation data, and the final report—never environment snapshots, request headers, API keys, OAuth state, or hidden chain-of-thought.
使用场景
- 架构或技术方案评审时,希望多个模型独立作答并匿名互评,最终由代码计算投票状态而非依赖单一裁判模型
- 策略决策或复杂技术问答中,需要保留异议和关键风险标记,避免共识偏差掩盖少数派的有效批评
- 多供应商(GPT/Codex、GLM、DeepSeek)混合环境下,用统一的盲评协议对同一问题做多轮交叉评审
- scenarios_en need to be 3 items, let me fix
适合人员
- 已在使用 DeepSeek Harness 并希望引入多模型交叉验证的开发者
- 需要透明、可复现的多模型评审流程而不依赖单一裁判 LLM 的团队
- 对模型输出持审慎态度、希望保留异议和降级状态可见性的技术负责人