Claude Skill MartBrowse skillsQuick linesLearn by videoTerminal guideWhat is a Skill?
← Back to list

Vision — External Vision Model Image Analysis

Send screenshots, UI mockups, or diagrams to external vision models (Doubao, Qwen, DeepSeek, OpenAI, Claude) and get a text description back.

UtilitiesIntermediate★ 171⑂ 8AI score 9/10Last updated: Aug 25, 2026

What it does

  • Runs vision.py with an image path plus a prompt and returns a textual description of the image.
  • Supports png/jpg/webp/gif; provider resolution order is --provider flag → VISION_PROVIDER env → first API key found.
  • Beyond the built-ins (doubao, qwen, deepseek, openai, anthropic), any OpenAI-compatible endpoint (vLLM, Ollama, LiteLLM, OpenRouter, Azure, self-hosted proxies) can be wired in via {NAME}_API_KEY, {NAME}_BASE_URL, {NAME}_MODEL — no code changes.
  • Fine-tune output with VISION_MODEL, VISION_TEMPERATURE, and VISION_MAX_TOKENS.
  • Includes a routing guard: run python vision.py --check-routing; if the session is native, skip the tool and analyze the image directly to avoid needless API cost.

Who it's for

  • Developers doing screenshot-based debugging or UI review in environments without native image understanding.
  • Teams wanting to plug Chinese-market models (Doubao, Qwen) or self-hosted vision models into a Claude Code workflow.
  • Anyone benchmarking several vision APIs for cost and quality.

Usage examples

  1. python vision.py "screenshot.png" "Describe the page layout and any visible UI issues." — quick UI sanity check with auto-detected provider.
  2. python vision.py --provider qwen "mockup.png" "List all components, colors, and spacing patterns." — design mockup breakdown.
  3. python vision.py -p openai "after.png" "Compare with app design spec, flag differences." — visual regression review.

· · · Install guide · · ·

Try it now, no install

Paste this into Claude to use the skill without installing anything.

Read the instructions in this file and follow them to help me:
https://raw.githubusercontent.com/xiincs/claude-code-vision-skill/HEAD/vision/SKILL.md

What I want: (describe your task here)

If Claude can't open the link, open it yourself and paste the contents instead.

↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.

Install in the Claude app (no terminal)
  1. Download the ZIP with the button below.
  2. In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
  3. Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
↓ Download ZIP
Install in Claude Code

Let Claude do it — paste this into Claude Code

Install the skill I found on Claude Skill Mart.
Copy the vision folder from the GitHub repo xiincs/claude-code-vision-skill into my ~/.claude/skills/vision/.
When it's done, tell me in one line what this skill can do.

Install with a command instead

git clone https://github.com/xiincs/claude-code-vision-skill.git /tmp/claude-code-vision-skill && cp -r /tmp/claude-code-vision-skill/vision ~/.claude/skills/vision

⚠ This is a third-party skill. Check the source repository before installing.

  1. Confirm Python 3 and pip are available: python --version.
  2. Clone the repo: git clone https://github.com/xiincs/claude-code-vision-skill.git
  3. Copy the skill folder: mkdir -p ~/.claude/skills && cp -r claude-code-vision-skill/vision ~/.claude/skills/vision
  4. Export the API key for your chosen provider, e.g. export DASHSCOPE_API_KEY="sk-..." (or DOUBAO_API_KEY, OPENAI_API_KEY, DEEPSEEK_API_KEY, ANTHROPIC_API_KEY).
  5. Optionally pin a default with export VISION_PROVIDER=qwen. If you use the anthropic provider, also run pip install anthropic.
  6. Reload your shell (or source ~/.zshrc) and smoke-test: python ~/.claude/skills/vision/vision.py "test.png" "Describe this image".
  7. Restart Claude Code and ask something like "analyze this screenshot" — the skill will be triggered automatically.