xby-recog
在终端中运行以下命令:
dsh plugin install xby-skill/xby-recog
将以下提示词粘贴到 DeepSeek Harness 对话框中:
在终端执行 dsh plugin install xby-skill/xby-recog 即可从 GitHub(https://github.com/xby-skill/xby-recog)安装该插件到 DeepSeek Harness。
插件介绍
Extracting structured data from images is a routine yet tedious task in production—scanned documents, handwritten forms, ID cards, passports, bank cards, license plates—where manual entry is slow and error-prone, and generic OCR cannot reliably parse field-level details on specific documents. xby-recog bridges that gap by bringing the xby recognition API into DeepSeek Harness as a plugin, so professional-grade image recognition can be invoked directly within a conversational workflow.
The plugin spans three categories of recognition. First, general text OCR in two tiers: a speed-optimised mode and a high-accuracy mode, both suited to documents, signage, and screenshots. Second, handwriting-specific OCR with dedicated text-line detection for notes, signatures, and hand-filled forms. Third, structured document and ticket recognition covering license plates (colour and bounding box), national ID cards (auto front/back detection and ID validation), passports and HK/MO/TW travel permits (with MRZ machine-readable zone parsing), bank cards (Luhn checksum validation), business licences, driving licences, and vehicle registration certificates. Every tool accepts three input formats—image URL, Base64 string, or local file path—and the API key persists automatically across restarts after a one-time setup.
It is well suited to developers building automated document review, data extraction, or form-digitisation pipelines, as well as application engineers handling bulk image-to-text workloads in customer service, finance, logistics, or government workflows. Because each capability maps to a single named tool, teams can compose multi-step recognition workflows without writing bespoke API wrappers, keeping the whole pipeline inside the Harness conversation loop.
使用场景
- 在对话流程中直接调用车牌、身份证、护照等证件识别,免写额外 API 封装。
- 批量处理扫描件与手写表单,自动提取结构化字段并校验有效性。
- 将多步图像识别(OCR + 证件解析 + 机读码)组合成端到端数据提取流水线。
适合人员
- 需要自动化文档审核与数据录入的开发者。
- 处理证件、票据批量识别的应用工程师。
- 在 Harness 中将视觉 API 工作流化调用的团队。