Best · Curated
Best DSH Vision & OCR Plugins
Vision plugins give text-only coding agents image understanding, OCR and structured evidence extraction.
Editor's pick
Editor's pick
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
Ranking
Tools & Capabilities plugin ranking
019834,561 downloads/mo02745,314 downloads/mo03792,974 downloads/mo04692,974 downloads/mo05742,630 downloads/mo06692,533 downloads/mo07712,359 downloads/mo08642,293 downloads/mo09672,141 downloads/mo10721,876 downloads/mo11671,656 downloads/mo12761,608 downloads/mo13691,554 downloads/mo14641,234 downloads/mo15691,234 downloads/mo16671,158 downloads/mo1768839 downloads/mo1866759 downloads/mo1969753 downloads/mo2064719 downloads/mo2166709 downloads/mo2261605 downloads/mo2366585 downloads/mo2466540 downloads/mo2566526 downloads/mo2659502 downloads/mo2766495 downloads/mo2859451 downloads/mo2961440 downloads/mo3061400 downloads/mo3161390 downloads/mo3266275 downloads/mo3364259 downloads/mo3466227 downloads/mo3558161 downloads/mo36570 downloads/mo37100 downloads/mo38520 downloads/mo39650 downloads/mo40570 downloads/mo41500 downloads/mo42100 downloads/mo43520 downloads/mo44100 downloads/mo45570 downloads/mo46550 downloads/mo47100 downloads/mo48100 downloads/mo49570 downloads/mo50100 downloads/mo51100 downloads/mo52570 downloads/mo53590 downloads/mo54270 downloads/mo55570 downloads/mo56620 downloads/mo57600 downloads/mo58620 downloads/mo59550 downloads/mo60570 downloads/mo61300 downloads/mo62570 downloads/mo63400 downloads/mo64100 downloads/mo65100 downloads/mo66570 downloads/mo67550 downloads/mo68100 downloads/mo69570 downloads/mo70570 downloads/mo
ysr666/dsh-vision-router
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
Scorp1o117/dsh-plugin-marketplace
Browse and search GitHub dsh-plugin repositories from Web UI Settings, sort by stars, and install plugins through the DSH CLI.
988hj7tczd-oss/dsh-computer-use
Cross-platform Computer Use for DeepSeek Harness: virtual-mouse human-like operation (screen_observe + computer_click + type + key + scroll + drag, 11 tools), AX-tree zero-vision-cost mode with free GLM-4V-Flash vision fallback, safety guards (snapshot TTL / app whitelist / dangerous-action approval / password-box protection).
qphotoai/dsh-computer-use-windows
Windows-ready DSH computer-use plugin with singleton host dependencies, safe UIA input, and CUA Driver vision fallback.
maxwell-feng/dsh-windows-ocr
Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
Aik358/dsh-cua-pre
Windows desktop automation for DeepSeek Harness: accessibility-first observe/act loop with 30 standard tools (element/coordinate targets, auto/a11y/event strategy routing), superseded-observations and no-replay safety semantics, persistent kill switch, RuntimeId/rect-drift stale protection, chat cards and a floating panel (live ops, frame wall, env auto-detect settings), tiled vision describe for low-res models, JSONL audit trail and a pid allowlist. Requires Windows 10+ and Python 3.9+ with uiautomation and pillow (installable from the settings page).
maxwell-feng/dsh-tesseract-ocr
Local OCR for attached images via Tesseract: only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
dsh-plugins/dsh-auxiliary
Auxiliary models for DeepSeek Harness: vision understanding and context compression through dedicated model routes. DeepSeek Harness 辅助模型插件:为视觉理解、上下文压缩、审批审查、子代理、会话标题与图片生成提供独立的模型路由、工具与系统提示,全程不触碰主对话模型。
Flyvhidbwo/dsh-vision-proxy
DeepSeek brain + automatic image transcription: attach images in the GUI and each one is transcribed via the official deepseek-v4-flash-vision-exp by default (a pure-text V4-Pro brain can see images), with any OpenAI-compatible VLM or local Ollama as alternatives.
Linjiangxian0203/dsh-remote-tunnel
Remote host tunnel manager: runs dsh web on a remote Linux server (systemd supervision, no root needed) behind an auto-reconnecting SSH tunnel; allocates and registers remote ports on the server, multi-user safe, with an audit view.
siegfly/dsh-deepseek-vision
A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.
lujianjun19/dsh-llm-github-copilot
GitHub Copilot LLM adapter: OAuth device-flow sign-in, live model discovery from the Copilot API, vision support for image-capable models (gpt-4.1, gpt-4o), and two wire protocols (Chat Completions and Responses API) with automatic endpoint routing per model.
bug-huntter/dsh-vision-plugin
Configurable image recognition for text-only DSH models: image messages are first transcribed by an OpenAI-compatible vision model (Base URL, model ID and API key set in a Settings section) and then passed to the main model as text, while image-input support is advertised. The API key auth scheme is selectable — OpenAI, Anthropic, Gemini or Azure style request headers — and a missing key is reported before any request is sent.
zzy6-a/vision-use
Desktop computer-use for DSH: captures the Windows screen into the agent vision channel and drives mouse/keyboard through a Codex-style overlay with Esc abort; auto-detects Windows native or WSL hosts.
xiaoshihou514/dsh-vision
Native vision capability extension, using either Zhipu (free) or Qwen-VL (local).
mokuyoaxis/dsh-iris
Media and vision workspace for DeepSeek Harness: image, video and speech generation, image Q&A and element locating, long-image OCR, pixel diff, HTML-screenshot verification and video summarization, with DashScope and OpenAI-compatible providers, model pools and a workbench client.
GodD6366/dsh-sub2api
Connect a sub2api gateway to DeepSeek Harness: OpenAI-compatible multi-provider routes (OpenAI / Claude / Grok / Gemini) behind one base URL, with per-key model discovery, usage lookup, and global vision/image tools.
maxmilian/dsh-grafana-query
Read-only Grafana tools that query metrics through the data source proxy: instance health, data source listing, instant and range PromQL queries, current alert state, and provisioned alert rules. Range queries downsample to a point budget; an explicitly supplied step that would exceed it is rejected rather than silently changed, so the model never receives data at a resolution it did not ask for.
sunxin-ai/dsh-design-qa
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Viger1/dsh-preview
Headless-browser verification tools so the agent opens the page it just built, reads the rendered DOM and computed styles, checks the console, and screenshots the result; ships a frontend-verify skill.
AngelosZou/dsh-pdf-reader
A content-aware PDF reading plugin for vision models: profiles each page for figures (vector and raster), tables, formula risk and double-column layout, then applies content-aware hybrid extraction, rendering figure/table/formula pages as high-DPI region crops. Provides a low-resolution preview to understand the page layout, and renders a specified region at high resolution. Packaged as multiple tools for agents.
moon09300731/dsh-vision-tools
Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.
ruby1304/dsh-vision-subagent
Vision for any DSH route: paste images in the Web composer with intent-aware auto-analysis, delegate workspace image reads to a Kimi/MiniMax vision subagent, and materialize pasted originals for editing.
SPYQWER1/dsh-codex-tools
Codex-backed `web_search`, `image_gen`, and `image_vision` tools for DeepSeek Harness, reusing ChatGPT OAuth login state.
fourzkw/dsh-drawai
An editable diagram canvas in the DSH sidebar plus two agent tools (diagram_read, diagram_apply) that read and edit the workspace in place, using native .drawio files: the canvas opens and saves them losslessly — cells it does not understand are preserved byte-for-byte — and a file-fingerprint revision powers optimistic locking against concurrent edits.
Elohia/dsh-plugin-image-input
Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).
niyongsheng/free-vision-skill
Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.
ld-1101/dsh-vision-plugin
Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).
wang-bool/visual-review
Renders pasted/uploaded images inline in the DSH Web chat and gives text-only models vision: the model-invokable visual_review tool calls any OpenAI-compatible multimodal API first, falling back to a local Qwen3-VL worker.
Elohia/dsh-plugin-mm-vision
Synesthesia Encoder for DSH: a vision model translates images into compact structured spatial text (canvas/elements/percentage coordinates), giving text-only LLMs pixel-level image understanding via the `mm_vision` tool.
Ck-epsilon/aura-vision
Aura Vision - free vision OCR plugin for DeepSeek Harness web profile (permanent bundle, GLM-4V-Flash free tier, tile-based long-document recognition)
XMoon/dsh-profile-settings
Per-profile settings overlay for DeepSeek Harness: global settings.yaml stays the baseline while each profile overrides any namespace through its own profiles/<name>/settings.patch.yml — object sections merge recursively, scalars and arrays replace wholesale, and !unset masks an inherited value. The overlay is transparent to existing plugins (they keep reading ctx.settings unchanged), writes land in the profile overlay only, and official schema validation, revision semantics, expectedRevision conflicts, watchers and events are untouched. Ships a settings command family (get/set/unset/mask/unmask/promote/demote/migrate/diff/layers) plus a Profile Settings section in the web Settings panel over a loopback RPC channel.
Harzva/dsh-maclens
Apple on-device Vision tools for text-only dsh models: local OCR (zh-Hans + 30 langs), image classification, face detection, document layout, and a combined describe — 100% offline, no API key, tall-screenshot slicing.
linenxi-ctrl/dsh-vision
External vision plugin for DeepSeek Harness: whale-button config panel, image recognition with auto-reply, and agent screenshot/recognize tools.
Leeminjing/dsh-eyes
Give text-only DeepSeek models on-demand vision: upload images, DeepSeek answers by calling a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).
linkingoscar/dsh-attachment-formats
Codex-style attachment formats for the DeepSeek Harness Web GUI: PDF text-layer extraction, Office text extraction, scanned-PDF OCR, long-document spill + index cards, image-to-PNG.
GOU-GEE/deepseek-vision#plugins/dsh-plugin-deepseek-vision
Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.
Kalospacer/dsh-model-capabilities
Adds a settings page for declaring per-model reasoning efforts and vision support on pi-ai provider routes.
StvLi/dsh-ros2
ROS2 debugging toolset and robot-state vision analysis for DeepSeek Harness: node/topic/service/action/interface/TF enumeration, whole-graph topology JSON, rosdep checks, approval-gated builds and custom message scaffolding, GUI screenshots and multimodal vision observation, plus headless RViz2 offscreen rendering (low-poly meshes, direct pixel read, GPU passthrough - motion rendering at 30Hz) with parallel VLM realtime analysis.
jmxsxwyzjdwl/dsh-mmroute
Transparent multimodal routing for text-only models: every image in every model call is fully transcribed (verbatim OCR, data, uncertainty zones, injection-hardened) by your own multimodal understander, with focused re-look via vision_relook and automatic retry on image-related failures. No bundled endpoints, no borrowed logins.
Renji004/dsh-omni-vision
Local eyes for text-only models: eyes_render draws text/shapes/Mermaid onto a canvas in the Web GUI, eyes_paste captures pasted images, eyes_ocr reads text via the built-in Windows OCR (offline), and eyes_analyze inspects pixels as structured data — no vision model required.
tinqiao-oss/clawtouch-mcp#dsh-clawtouch
Computer use through an external USB HID device: name a target in plain language, a vision model locates it in a cropped window screenshot, and a Raspberry Pi Pico 2 moves the real mouse and types on the real keyboard. 7 tools. Windows has window listing, per-window cropping and automatic window raising; macOS needs pyobjc and has neither; Linux is unsupported.
shinjiyu/dsh-plugin-multimodal
Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.
littleblakew/msds-chain-mcp#dsh
Chemical safety intelligence through the hosted MSDS Chain MCP endpoint: 23 tools for compatibility and mixing order, GHS hazards, PPE, storage, waste disposal, exposure limits, transport classification, multi-region regulatory compliance, SDS lookup and version diffs, and signed audit reports, each answer citing the supplier SDS and revision date it is grounded in.
dami9527/dsh-image-pathify
Lets text-only models handle pasted chat images, with a native vision experience, batch image viewing, and a built-in OpenAI-compatible analyze_image tool; vision-capable models are unaffected.
Junkrat9527/dsh-autovision
dsh-autovision: paste an image into a text-only model composer and a configured multimodal model transcribes it to text automatically. Twin-provider auto-routing + agent-callable read-image tool. No built-in keys, no relay.
DamonKoy/dsh-web-ui#dsh-tool-describe-image
Gives a text-only model image understanding via a vision-language model, exposed as a `describe_image` tool.
xing666173/dsh-vision-hub#file-drop
Drag-and-drop file upload for PDF, Word, Excel and images: dropped files are saved to a local directory and referenced by path, no base64 bloat in the chat.
NagasakiSoyo-ui/dsh-llm-deepseek-vision
Vision-augmented DeepSeek adapter plugin for DeepSeek Harness: a vision-capable model describes image input, then a text-only DeepSeek model reasons over the description
314857493/dsh-vision#vision-route
Registers a `deepseek-vision` provider route: the Web GUI accepts pasted images and transcribes them to text via the free Zhipu GLM vision API before delegating to the DeepSeek adapter.
314857493/dsh-vision#vision-tool
Model-facing `vision` tool for DeepSeek Harness: describe and OCR image files by calling the free Zhipu GLM vision API directly (glm-4v-flash fallback chain), no external CLI required.
boheastill/phone-eye
Let your AI agent see and operate a real Android phone: phone_look (vision + UI-tree fusion), tap/swipe/type, screenshot — over adb, for any MCP client.
xie129716/computer-user-vision
Windows computer use forked from computer-user: 13 computer_* tools that read the screen and drive the mouse and keyboard. Image-capable routes get the screenshot as a real image with an exact image-to-screen mapping, so no external OCR; elements return as UI Automation refs so a click lands on the exact control rectangle; Ctrl+Alt+Esc stops every call.
xiaoyuink/dsh-image-vision
Image understanding for any DSH model: vision, OCR, grounding, and crop tools with domain presets for histopathology, cell biology, anatomy, clinical images and scientific figures.
NOirBRight/dsh-llm-opencode-go
OpenCode Go subscription chat for DSH with per-model Completions, Responses, or Anthropic protocol selection, model discovery, context, vision and thinking metadata, effort defaults, and usage meters.
CZX2244/dsh-bilibili
Bilibili video analysis: metadata, transcript (ASR fallback via Bijian/sherpa-onnx/whisper.cpp), comments, danmaku, and sharp keyframes with optional local vision descriptions.
gugu123a/dsh-tool-see-image
Provides the see_image tool for DSH: send an image file to a configurable OpenAI-compatible vision model and relay its description back to a text-only model.
jyh20030112/dsh-visual-plugin
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
azwosile/dsh-highres-vision
For the DeepSeek Harness native vision model deepseek-v4-flash-vision-exp, raises image admission limits to 32 MiB / 8192 px / 600 images and adds a highres_read tool that tiles large images, then returns the whole image plus 800x800 tiles through the host read_image tool.
JohnXu22786/model-catalog
dsh plugin: model catalog auto-discovery - fetch model listings, pricing and capabilities from OpenAI-compatible API hosts, normalized into ready-to-use config
ZI-LV68/dsh-deepseek-model-router
Auto-switch DeepSeek models: routes every request to the vision model when images are present, to the pro model for complex tasks, and to the fast model otherwise; includes a switch_model tool for manual overrides.
drscrewdriver/dsh-llm-openai-completions
OpenAI-completions-compatible adapter for custom gateways (vLLM / LM Studio / self-hosted proxies): always role:"system", thinking driven by the model config, Qwen-style response split, and vision-model image input (single / multiple).
br1nosense/dsh-vision-solution
Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.
TZHR-invest/dsh-plugins#dsh-vision-tool
Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.
pearjelly/deep-blend#bundle
Render and iterate on Blender scenes from DSH: 16 tools for scene specs, previews, visual review and delivery renders, with immutable revisions and an approval gate.
Isanti2016/dsh-quicksight
Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).
nyantused-cpun/gewu-tools
Model-agnostic visual-inspection pipeline for text-only agents: page-by-page HTML screenshots plus a ready-made vision-subagent briefing contract (gewu_prep), then source-code truth verification of every finding (gewu_locate); validated on mimo-v2.5 & qwen3.7-plus.
ankye/dsh-client-vision#tool-vision
Screen capture and external vision recognition: take_screenshot, list_windows, analyze_image and view_image tools with a configurable GPT vision channel (gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra), API key via the credentials service, and a settings card; view_image shows the screenshot in the Web UI while the model context keeps text only.
kaixinbaba/dsh-vision-recognizer
Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.
poiuyjie/dsh-vision-opencode
Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.
View all Tools & Capabilities plugins Tools & Capabilities,Back to leaderboard dshplugins.cc。