DSH Plugins
返回列表
🤖

dsh-codex-media

模型推理 更新于 2026.08.25

在终端中运行以下命令:

dsh plugin install binsarjr/dsh-codex-media

将以下提示词粘贴到 DeepSeek Harness 对话框中:

在终端执行 dsh plugin install binsarjr/dsh-codex-media 即可将本插件(源码地址 https://github.com/binsarjr/dsh-codex-media)安装到当前 DeepSeek Harness。

插件介绍

Text-only models hit an awkward wall: DeepSeek Harness rejects image messages outright, and a naive document pipeline tends to dump the entire extracted text into the agent's context, blowing the token budget in one turn. dsh-codex-media takes a practical stance: instead of asking a language model to see, it hands the model a local file path and lets a local OpenAI Codex CLI or a compatible endpoint do the heavy lifting, returning only a concise answer to the agent. The plugin ships with zero runtime dependencies (Node 22+ built-ins only), deliberately omits any upload machinery, and pairs naturally with dsh-drop-to-path, which owns the drop-and-save UX.

Three capabilities sit behind one engine. analyze_image handles PNG, JPEG, WebP, and GIF for freeform description or targeted questions. analyze_document covers PDF, Office, RTF, and common text formats, and the engine extracts only the final assistant text so binary documents never flood the context window. generate_image turns a text prompt into a local file; by default it rides the Hermes Agent oneshot authenticated through an existing ChatGPT/Codex login, so no separate API key is required. Four transports are available on a single call interface, every request passes extension allow-lists, size caps, and a PDF magic-byte check, and a per-call timeout kills the process tree or aborts the HTTP request. Document contents are always treated as untrusted data, and the analysis prompt instructs the model to ignore any instructions embedded inside the file.

If you are building agent workflows on DeepSeek or another text-only model and occasionally need to glance at an image or read through a PDF, or if you want image generation to run entirely on a locally authenticated Codex CLI without juggling another API key, this plugin fills that gap. It does not change which model you use; it simply extends the model's reach to the files it cannot see natively.

使用场景

  • 在文本模型中让 Agent 查看并描述本地图片
  • 让 Agent 读取 PDF 或 Office 文档并仅返回针对性回答
  • 无需额外 API Key,通过 Codex CLI 本地生成图像文件

适合人员

  • 用 DeepSeek 等文本模型搭建 Agent 工作流的开发者
  • 希望推理链完全本地化、避免多套 API Key 的团队
  • 需要为 Agent 扩展视觉理解与文档阅读能力的集成者