证据快照复核于 2026-09-05GitHub 数据核对日期: 2026-08-21
来源已审查独立 Skill搜索、视觉与数据agents-working-with-images Profileui-reconstruction Profilescreenshot-analysis Profile

vision-skills

用本地命令行工具完成图片 OCR、元素定位、像素级对比、裁剪、SVG 描摹和 HTML 截图。

快速了解

它能做什么

用本地命令行工具完成图片 OCR、元素定位、像素级对比、裁剪、SVG 描摹和 HTML 截图。

本站提供的是中文说明,不代表该项目或 Plugin 自身提供中文界面;语言支持请以上游文档为准。

能力
搜索、视觉与数据视觉image-understandingUI

选择前先看

视觉技能工具包让文本型智能体通过专用命令检查图片文件。可用 glance 做描述或 OCR,用 ground 和 detect 获取坐标,用 trace 和 pixel_diff 获取基于像素的几何信息或变更区域,并提供长截图、颜色、前景提取和 HTML 渲染脚本。

适合谁

需要分析截图、图表、示意图、图标或 UI 参考图,并希望采用可复用命令行流程的智能体和开发者。

常见任务

  • 读取截图文字,或处理长滚动截图并检查分段边界。
  • 定位指定 UI 控件,放大后检查其状态。
  • 先精确找出两张截图的变化区域,再解释变化内容。
  • 借助裁剪、描摹、颜色分析和 HTML 截图,从参考图复原 UI 或图形。

权限与数据

使用本地图片路径和共享的视觉服务配置。

权限
  • 读取支持的 PNG、JPEG、GIF 和 WebP 图片文件。
  • 在本地写入裁剪图、SVG、OCR 文档和截图等输出。
  • HTML 截图渲染时可能调用本地 Chrome 系浏览器。
数据处理
  • 依赖 API 的命令可能会把图片或选定裁剪区域发送至已配置的视觉服务。
  • trace、crop、pixel_diff 和 dominant_colors 等本地工具在文档中说明不会使用视觉 API。
外部服务
  • 通过 VISION_BASE_URL 等 VISION 配置指定的视觉 API 服务。
  • html_shot 所需的本地 Chrome、Chromium 或 Edge。
凭据
  • 使用依赖 API 的命令时,需要为已配置视觉服务提供 VISION_API_KEY。

局限

  • ground 和 detect 的坐标框由 0–1000 网格缩放而来,不适合像素级精确判断;需要精确数值时应使用 trace 或像素工具。
  • 高密度界面的检测数量可能波动;文档建议按布局区域分别检查。
  • 部分命令依赖可选工具或软件包;若命令不可用,应按文档提示报告缺失情况。
  • 仅支持 PNG、JPEG、GIF 和 WebP 输入。

DSHub 已核对

  • 源提交已固定。
  • 已采集完整的技能文档。
  • 仓库许可证为 MIT。

DSHub 未核对

  • DSHub 未安装或执行该技能。
  • 未验证可选本地命令、Python 软件包、浏览器及视觉服务配置是否可用。

固定版本安装

主要操作

这个独立 Skill没有 DSH Plugin 安装操作,请根据源码文档使用真实交付方式。

访问源码项目

维护者原文

Skill 使用说明

查看 commit bf68366 对应的 SKILL.md
维护者编写的上游内容原文于 2026/9/5skills/vision-skills/SKILL.md 获取,正文和仓库相对媒体固定到 commit bf68366d2a22,内容哈希为 d902eab31823。以下是未经 DSHub 翻译的上游原文,语言可能与当前页面不同;第三方托管的 badge 可能独立更新。

name: vision-skills description: >- Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and scripts/html_shot.py (HTML file to image). Use for any task involving an image — questions, text, splitting and transcribing long screenshots or chat histories, locating elements, comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading values off a chart, operating a GUI from screenshots — and to re-check an image yourself when a description you were given lacks a detail.

vision-skills

Five local CLIs that give a text-only agent eyes. They read one shared vision config (VISION_API_KEY / VISION_BASE_URL / VISION_MODEL / LANG), plus the optional Python-client settings VISION_API_PROTOCOL, VISION_REASONING_EFFORT, and VISION_USER_AGENT — no extra credentials.

Pick the tool by the question you are answering:

Question Tool
"What does this image show / say?" glance
"Where is X?" — a thing you can name ground
"Where are all the Xs?" — every instance of a kind detect
"What is its exact shape, size, offset?" trace
"Cut this box out as its own image file" crop
"OCR this long screenshot / scrolling page / chat history" scripts/long_screenshot_ocr.py
"Extract the icon/logo foreground as transparent PNG — manual region or auto (cropped+scaled screenshots)" scripts/extract_fg.py
"Turn this HTML file into a viewport or full-page screenshot" scripts/html_shot.py
"Which colours dominate a region, and which palette value fits it?" scripts/dominant_colors.py
A relation none of them return — a gap, a distance between two located things code over the pixels (Pillow)

glance answers what something is; ground and detect answer where. You give ground a description of a particular thing; you give detect a kind and it enumerates the instances.

Both give real coordinates, but they are not pixel-exact: the box arrives on a 0-1000 grid and is scaled to your image, so the last pixel or few are not reliable. That is accurate enough to crop with, to click, to compare positions against. When a number has to be exact, trace derives it from the actual pixels — offsets, sizes, shapes.

Use the provided tools before hand-rolled pixels

Everything this toolkit ships a tool for, call the tool — do not rewrite it with Pillow in the middle of a task. The CLIs exist so the same pixel work is not hand-coded differently every time:

  • cut a box out of an image → crop, not Image.open(...).crop(...)
  • sample a region's palette → scripts/dominant_colors.py
  • compare two images → scripts/pixel_diff.py
  • vectorize to SVG → trace
  • locate / inventory elements → ground / detect
  • describe / OCR an image → glance
  • safely split, OCR, and merge a long screenshot → scripts/long_screenshot_ocr.py
  • HTML file to a viewport or full-page screenshot → scripts/html_shot.py

Hand-written Pillow is only for what none of them return: a relation between two things you already located (a gap, a distance), a resize or overlay, drawing. If you catch yourself writing .crop(), .convert(), or histogram code where one of the tools above fits, replace it with the tool call — same coordinates, same box format, and the output feeds the next tool directly.

glance — ask about an image

glance <image>                                 # detailed description
glance <image> -q "<question>"                 # targeted question (qualitative only)
glance <image> --ocr                           # verbatim OCR
glance <image> --region X1,Y1,X2,Y2 -q "..."   # zoom into a crop
glance <img1> <img2> -q "..."                  # compare in ONE call

When you do compare with glance, pass all paths to one call — separate calls cannot see both images, so two descriptions compared afterwards are two hallucination surfaces, not a comparison. --region uploads only the crop, so small text and icons become readable.

But "what changed between these two?" is not a glance question. A one-word badge or a small shift is a rounding error to a vision model and exact to scripts/pixel_diff.py. Diff first to get the box, then glance --region that box to read what the change actually is.

For a tall scrolling screenshot, do not send the whole image through one OCR call and accept the model's downscaling loss. Run the long-screenshot workflow, which finds low-content cut bands, invokes glance on each chunk, uses structured extraction for chat histories, merges only duplicated overlap, and writes a boundary audit:

python3 scripts/long_screenshot_ocr.py work/page.png -o work/page.ocr.md
python3 scripts/long_screenshot_ocr.py work/chat.png --mode chat --resume -o work/chat.ocr.md

Read references/long-screenshot-ocr.md before using it. It defines the verification pass for unsafe cuts and chat-message boundaries.

ground — locate a named target

ground <image> "<target description>"
ground <image> "<target>" --region X1,Y1,X2,Y2

Output: x1: .., y1: .., x2: .., y2: .. in original-image pixels — with --region too (crop hits are mapped back).

Provider-native 0-1000 boxes do not all use the same array order: Gemini uses [y0, x0, y1, x1], while Qwen3-VL, Qwen3.5, and Qwen3.6 use [x0, y0, x1, y1]. Grounding code must select the order by model family (or an explicit override) before scaling to pixels; never parse every provider as Gemini-style yxyx.

If several boxes come back numbered, your description matched more than one element rather than picking out a single thing. Narrow it with what distinguishes the one you mean — its text, its position, the block it sits in — and ask again.

The box is a handle, not just an answer — it feeds the next call:

$ ground screenshot.png "the send button"
x1: 1067, y1: 841, x2: 1108, y2: 881
$ glance screenshot.png --region 1067,841,1108,881 -q "is it enabled or greyed out?"

That two-step is how you inspect anything too small to survive a full-image pass.

detect — find every instance of a kind

detect <image>                        # every UI element
detect <image> "buttons"              # one kind only
detect <image> --region X1,Y1,X2,Y2   # inside one box

You name a particular thing for ground; you name a kind for detect and it enumerates the instances. Output is a numbered list with each item's visible text and box. A full-screen pass is a fast first draft — counts vary run to run on dense screens. For completeness, detect the layout blocks first, then detect --region each block.

trace — exact shape geometry (local, no vision API)

trace <image>                                  # b/w spline SVG to stdout
trace <image> --polygon                        # boxy diagrams/wireframes
trace <image> --region X1,Y1,X2,Y2 -o out.svg  # crop first

Coordinates come from the actual pixels, not a model's estimate. Flat, high-contrast graphics only; text becomes curves (pair with --ocr when the text matters). Small images are upscaled automatically before tracing, so a 30px icon traces as readily as a screenshot — size is not a reason to skip the tool. Before shipping or reusing a traced SVG, read references/restore-graphic.md — it holds the reuse traps and the ship-vs-hand-write call.

crop — cut a pixel box out of an image (local, no vision API)

crop <image> --region X1,Y1,X2,Y2             # writes <image-stem>.crop.png next to the input
crop <image> --region X1,Y1,X2,Y2 -o out.png
crop <image> --region X1,Y1,X2,Y2 --scale 4   # upscale the cut-out 4x (LANCZOS) first

The same X1,Y1,X2,Y2 pixel boxes ground/detect print, clamped to the image bounds. Once a box is worth keeping — the same crop is about to feed pixel_diff, dominant_colors, and trace in turn — cut it to a file once and reuse it, instead of re-cropping in memory on every call. --scale N upscales the cut-out before writing (default output name becomes <image-stem>.crop@Nx.png): for icons too small for ground/trace to see clearly, crop with --scale 4, then run ground/trace on the upscaled file — coordinates it returns are in the upscaled grid, divide by N to map back to the original image. Requires the optional pillow.

extract_fg — icon foreground as transparent PNG: manual region or auto (local, no vision API)

# manual: you know the region (and optionally the background colour)
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 -o icon.png
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 --mode dark          # grey/black line logos
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 --exclude-color '#E6E6E6'
# auto: `crop --scale` cut-outs with the icon centred — no region needed
crop shot.png --region X1,Y1,X2,Y2 --scale 4 -o d/icon1.png
python3 scripts/extract_fg.py d/icon1.png d/icon2.png       # writes <stem>.clean.png next to each input
python3 scripts/extract_fg.py d/icon1.png --disc-radius 60
python3 scripts/extract_fg.py d/icon1.png --boxes "101,84,184,171"

Manual mode keeps every sufficiently large connected component of the region (separate logo sub-shapes stay together; specks drop out). Auto mode takes a crop --scale cut-out with the icon centred (disc + glyph): the disc centre is the image centre, the disc radius defaults to min(w,h)/2 * 0.6, and the disc colour is sampled from a ring around the centre; that colour is excluded and the glyph is picked as the most saturated among the three largest coloured components (white rings, ripples, and text fall away), output as a 1:1 transparent PNG. When auto inference fails, override the radius with --disc-radius, or pass a ground box (in the upscaled grid) as --boxes to recentre and re-filter by overlap. Multiple images may be passed at once (auto mode). Requires the optional pillow (and numpy for auto mode).

html_shot — render an HTML file to an image (local, needs a Chrome-family browser)

python3 scripts/html_shot.py page.html                      # writes page.png, 1280x800
python3 scripts/html_shot.py page.html --width 1440 --height 900 -o page.png
python3 scripts/html_shot.py page.html --scale 2            # 2x pixels: small text stays readable
python3 scripts/html_shot.py page.html --full-page           # complete scroll height, same layout viewport
python3 scripts/html_shot.py page.html --full-page --max-pixels 40000000

The visual-alignment loop: write HTML, screenshot it at the reference viewport, then compare it with the design. Use pixel_diff to locate material differences, not to chase a zero-difference score. Rendering happens in headless Chrome/Chromium/Edge — no Python dependencies. The default captures only the viewport. Use --full-page for the complete document while keeping --width and --height as the layout viewport, so vh/svh and responsive breakpoints do not change. Add --max-pixels N when the page height is untrusted. --wait-ms N pauses for fonts, images, or animation before capturing. Paths are relative to this skill's own directory.

pixel_diff — where two images differ (local, no vision API)

python3 scripts/pixel_diff.py <a> <b>      # path is relative to this skill dir

Prints an overall difference percentage plus the worst regions as x1: .. boxes you can feed straight into glance --region. Exact where a vision model rounds off.

dominant_colors — a region's palette, and the exact value among candidates (local, no vision API)

python3 scripts/dominant_colors.py <image> --region X1,Y1,X2,Y2          # top colour clusters + shares
python3 scripts/dominant_colors.py <image> --region X1,Y1,X2,Y2 \
  --candidates '#F9FAFA,#F5F5F5,#F3F3F3,#EDEDED'                        # pick the best candidate

A vision model names a colour ("light gray") but not its value. The first mode downsamples, quantizes, and merges near-duplicates to list the region's significant colours with the share each owns — the histogram shows which colour is the background and which is the accent. Given the candidate palette your label implies, the second mode scores each candidate by how close the region's pixels are to it and prints the winner. Take the value from here, never from glance's prose. Paths are relative to this skill's own directory.

Work from a copy, not a temp path

If the image lives in a temp directory, before your first tool call on one, copy it somewhere durable and run everything against the copy — that is what keeps the image reachable later:

cp "<the temp path>" work/shot.png
glance work/shot.png -q "..."

Exception: the user asked for the image to stay in a temp folder.

When you have a description instead of the image

If an image reached you only as text — a description written by a person, a tool, or another model — and the image's file path is visible in the conversation, do not reason past a missing detail. Look again yourself:

  1. glance <path> -q "<the specific detail>" — one qualitative follow-up.
  2. ground <path> "<target>" then glance <path> --region <that box> -q "..." — locate, then zoom. The reliable way to inspect one element closely.

If the file no longer exists, say so instead of guessing.

Coarse to fine — the method behind every task above

For a single question about an image, glance is the whole answer. For anything multi-step, work outside-in:

  1. One full-image pass (glance, or a description you already have) for the layout and an inventory of what is where.
  2. For any element that matters, ground it, then zoom with glance --region <box> -q "...". Full-image passes routinely miss small text and icons; a crop puts all the pixels on one detail, so the model sees it at effectively higher resolution. When the same box will be checked more than once, cut it to a file first with crop.
  3. Never take a prose answer for a pixel-level fact — exact colors, small offsets, sizes. Vision models confidently report styling that is not there: coloured syntax highlighting in a monochrome code block, a border that does not exist. Get the number from trace, from a ground box, or from pixel_diff; sample the pixels yourself only for what those cannot return.

Use cases

Each file below is one job, start to finish: when it applies, the call sequence, and how to tell you got it right.

The job Read
OCR a long screenshot, scrolling page, or chat history without losing text at chunk boundaries references/long-screenshot-ocr.md
Rebuild a page or component as HTML/CSS, including a roughly three-minute fast approximation mode, or align an existing UI with its reference image references/restore-ui.md
Extract or rebuild an icon, logo, illustration, or other isolated graphic as transparent PNG/SVG references/restore-graphic.md
Turn a sketch, diagram, or whiteboard into Mermaid, Graphviz, or another structured representation references/restore-structure.md
Operate a GUI from screenshots — locate, act, verify each step references/gui.md

Notes

  • Only PNG / JPEG / GIF / WebP images are supported.
  • If a command is not found, the optional tools were not installed — report this to the user instead of improvising a replacement.
  • If the vision API fails, relay the error faithfully; never fabricate image content.

Source repository: https://github.com/Anionex/agent-vision-toolkit

Installation guide: https://github.com/Anionex/agent-vision-toolkit/blob/main/AGENT_INSTALL.md

有意识地管理

安装与管理

前置条件与目标 Profile

目标 agents-working-with-images Profile, ui-reconstruction Profile, screenshot-analysis Profile

交付方式 Skill 文件 — https://raw.githubusercontent.com/Anionex/agent-vision-toolkit/bf68366d2a2250691ef42f3ca1b464d2eba1ab36/skills/vision-skills/SKILL.md

兼容性与访问范围

Requires local vision CLIs; some functions need optional dependencies Not declared in supplied evidence

检查兼容性证据

风险事实

external-image-processing

May send input images to the vision API configured through VISION_API_KEY, VISION_BASE_URL, and VISION_MODEL.

证据
local-file-access

Reads image files and can write crops, SVGs, OCR output, and rendered screenshots to local paths.

证据
证据与编辑审查Manifest、Bundle patch、分发与新鲜度

不可变证据

审查状态与源码活动

人工已批准

在核对来源内容和不可变发布记录后,已由人工批准发布。AI 参与了内容草稿生成,最终发布决定由人工完成。

人工审查于 2026/9/5 UTC 17:23GitHub 事实核对日期: 2026/9/5 UTC 16:31

自当前证据基线以来,没有记录到重要源码变化。

下一步

比较生态 Artifact 类型

订阅重要变化: vision-skills