vision-exp-tile
在终端中运行以下命令:
dsh plugin install Nicholaskin/vision-exp-tile
将以下提示词粘贴到 DeepSeek Harness 对话框中:
在 DeepSeek Harness 终端中运行 dsh plugin install Nicholaskin/vision-exp-tile 即可安装,插件源码地址为 https://github.com/Nicholaskin/vision-exp-tile
插件介绍
When DeepSeek released the v4-flash-vision-exp model, any image above 800×800 pixels gets automatically downsampled before it reaches the model. Fine table text, dense charts, long-document details simply blur out. vision-exp-tile solves this by slicing large images into lossless 800×800 tiles before inference, feeding each tile to the vision model at full resolution, and then auto-aggregating the per-tile answers into a single structured result using coordinate annotations. In other words, it turns the official downscaling rule from a hard limit into a sweet spot: every tile bypasses resampling, stays within a predictable token budget, and the final output reads like you fed the full image in one shot.
The plugin ships three strategies out of the box. Smart mode lets the model inspect the whole image first, ask clarifying questions, then target specific regions. Pipeline mode runs the entire flow hands-free: pre-check, local OCR via Paddle, Rapid, or Windows OCR (auto-fall-through to the vision API if none are installed), pixel-grid text transcription, proportional region cropping, and layered aggregation. Full mode tiles the entire canvas on a uniform grid for maximum coverage. It depends on no third-party DSH plugin; OCR engines are optional user-side environments, and the pipeline degrades gracefully at every step. Temp files older than 24 hours are cleaned up automatically, and all billing goes straight to the DeepSeek official API dashboard.
This plugin is for DeepSeek Harness (DSH) users who regularly hit the "image too big, details lost" wall: screenshot analysis, long-document OCR, dense table and chart reading, UI review, or any workflow where large images carry critical detail. No coding required, no manual cropping. If you already run picturereader for small images, vision-exp-tile detects it and routes work automatically—large images here, small documents there—without touching the other plugin at all.
使用场景
- 上传超大截图或长文档,让视觉模型无损识别其中的文字与细节
- 对密集表格或图表执行全自动 OCR 流水线,自动汇总结构化结果
- 使用智能模式让模型先了解需求,再精读指定坐标区域
适合人员
- 经常用 DeepSeek 视觉模型但被大图降采样卡住的用户
- 需要批量识别截图、文档、图表细节的 AI 应用开发者
- 已部署 DSH 并希望零额外插件依赖扩展图像能力的团队