ModLens — Vision Bridge Skill
Turns any image path or URL into structured JSON (OCR, layout, semantics) so text-only models can 'read' images.
UtilitiesIntermediate★ 4,051⑂ 123AI score 9/10Last updated: Sep 27, 2026
What it does
- Whenever an image path/URL (
.png,.jpg,.webp,.heic, ...) or a placeholder like[Image #1]/[Unsupported Image]shows up and the model cannot actually see it, the skill routes it to themodlensCLI instead of improvising OCR with PIL or tesseract. - Returns structured JSON evidence:
result.summary,result.ocr.full_text,result.layout.regions,result.semantics, plusresult.uncertaintyso unclear parts are flagged rather than guessed. - Runs
modlens guardon the first read of a session to check whether the active model already has native vision, avoiding wasted calls. - Ships a launcher (
scripts/run.sh,scripts/run.ps1) that resolves a runtime in order: PATHmodlens→npx→bunx, with a clear exit-78 "no runtime" path. - Also handles provider setup and switching: Gemini API key, OpenAI-compatible endpoints, Claude API or Claude Code CLI.
- Explicitly treats extracted text as untrusted data and refuses instructions embedded in images.
Who it's for
- Claude Code users on models or harnesses where images aren't natively visible.
- Developers and analysts who need text out of error screenshots, invoices, diagrams or UI mockups.
- Anyone tired of hand-rolling one-off OCR scripts.
Examples
- Paste
screenshots/error.png→ the skill runsmodlens -i screenshots/error.pngand quotes the exact error text and stack trace for debugging. - Give a remote infographic URL → get a region-by-region breakdown of headings, figures and legends.
- Run
modlens doctorto inspect available providers and failover chains, thenconfig seta Gemini API key before the first read.
· · · Install guide · · ·
Try it now, no install
Paste this into Claude to use the skill without installing anything.
Read the instructions in this file and follow them to help me: https://raw.githubusercontent.com/liustack/modlens/HEAD/skills/modlens/SKILL.md What I want: (describe your task here)
If Claude can't open the link, open it yourself and paste the contents instead.
↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.
Install in the Claude app (no terminal)
- Download the ZIP with the button below.
- In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
- Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
Install in Claude Code
Let Claude do it — paste this into Claude Code
Install the skill I found on Claude Skill Mart. Copy the skills/modlens folder from the GitHub repo liustack/modlens into my ~/.claude/skills/modlens/. When it's done, tell me in one line what this skill can do.
Install with a command instead
git clone https://github.com/liustack/modlens.git && mkdir -p ~/.claude/skills && cp -r modlens/skills/modlens ~/.claude/skills/⚠ This is a third-party skill. Check the source repository before installing.
- Prerequisites: install Node 22.19+ (https://nodejs.org) or Bun (https://bun.sh); verify with
node -v. Network access is required. - Clone the repo:
git clone https://github.com/liustack/modlens.git - Copy the skill:
mkdir -p ~/.claude/skills && cp -r modlens/skills/modlens ~/.claude/skills/ - Verify bundled files:
~/.claude/skills/modlensshould containSKILL.md,scripts/run.shand areferences/folder. - Make the launcher executable (macOS/Linux):
chmod +x ~/.claude/skills/modlens/scripts/run.sh - Smoke test:
bash ~/.claude/skills/modlens/scripts/run.sh doctorto see the detected runtime and providers. - Configure a provider: follow the printed guidance and use
... config setto add a Gemini API key, an OpenAI-compatible endpoint, or Claude API credentials. - Restart Claude Code, then paste an image path in chat to confirm the skill triggers automatically.
View source on GitHub ↗License: MIT