Vision — External Vision Model Image Analysis
Send screenshots, UI mockups, or diagrams to external vision models (Doubao, Qwen, DeepSeek, OpenAI, Claude) and get a text description back.
UtilitiesIntermediate★ 171⑂ 8AI score 9/10Last updated: Aug 25, 2026
What it does
- Runs
vision.pywith an image path plus a prompt and returns a textual description of the image. - Supports png/jpg/webp/gif; provider resolution order is
--providerflag →VISION_PROVIDERenv → first API key found. - Beyond the built-ins (doubao, qwen, deepseek, openai, anthropic), any OpenAI-compatible endpoint (vLLM, Ollama, LiteLLM, OpenRouter, Azure, self-hosted proxies) can be wired in via
{NAME}_API_KEY,{NAME}_BASE_URL,{NAME}_MODEL— no code changes. - Fine-tune output with
VISION_MODEL,VISION_TEMPERATURE, andVISION_MAX_TOKENS. - Includes a routing guard: run
python vision.py --check-routing; if the session isnative, skip the tool and analyze the image directly to avoid needless API cost.
Who it's for
- Developers doing screenshot-based debugging or UI review in environments without native image understanding.
- Teams wanting to plug Chinese-market models (Doubao, Qwen) or self-hosted vision models into a Claude Code workflow.
- Anyone benchmarking several vision APIs for cost and quality.
Usage examples
python vision.py "screenshot.png" "Describe the page layout and any visible UI issues."— quick UI sanity check with auto-detected provider.python vision.py --provider qwen "mockup.png" "List all components, colors, and spacing patterns."— design mockup breakdown.python vision.py -p openai "after.png" "Compare with app design spec, flag differences."— visual regression review.
· · · Install guide · · ·
Try it now, no install
Paste this into Claude to use the skill without installing anything.
Read the instructions in this file and follow them to help me: https://raw.githubusercontent.com/xiincs/claude-code-vision-skill/HEAD/vision/SKILL.md What I want: (describe your task here)
If Claude can't open the link, open it yourself and paste the contents instead.
↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.
Install in the Claude app (no terminal)
- Download the ZIP with the button below.
- In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
- Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
Install in Claude Code
Let Claude do it — paste this into Claude Code
Install the skill I found on Claude Skill Mart. Copy the vision folder from the GitHub repo xiincs/claude-code-vision-skill into my ~/.claude/skills/vision/. When it's done, tell me in one line what this skill can do.
Install with a command instead
git clone https://github.com/xiincs/claude-code-vision-skill.git /tmp/claude-code-vision-skill && cp -r /tmp/claude-code-vision-skill/vision ~/.claude/skills/vision⚠ This is a third-party skill. Check the source repository before installing.
- Confirm Python 3 and pip are available:
python --version. - Clone the repo:
git clone https://github.com/xiincs/claude-code-vision-skill.git - Copy the skill folder:
mkdir -p ~/.claude/skills && cp -r claude-code-vision-skill/vision ~/.claude/skills/vision - Export the API key for your chosen provider, e.g.
export DASHSCOPE_API_KEY="sk-..."(orDOUBAO_API_KEY,OPENAI_API_KEY,DEEPSEEK_API_KEY,ANTHROPIC_API_KEY). - Optionally pin a default with
export VISION_PROVIDER=qwen. If you use the anthropic provider, also runpip install anthropic. - Reload your shell (or
source ~/.zshrc) and smoke-test:python ~/.claude/skills/vision/vision.py "test.png" "Describe this image". - Restart Claude Code and ask something like "analyze this screenshot" — the skill will be triggered automatically.
View source on GitHub ↗License: MIT