Generate and edit images across 30+ models (Gemini, Seedream, Recraft, GPT-Image, Riverflow) through one OpenRouter Image API call.
Official skill for generating and validating animated GIFs that meet Slack's emoji and message specs.
Describe a scientific diagram in plain language and get a publication-oriented PNG that an AI reviewer scores against your document type's quality bar.
Describe your content and get a publication-quality infographic that auto-refines until it passes a quality threshold.
Turn a list of shell commands into polished animated terminal GIFs using VHS, ready to drop into your README.
Turns audio and video into speaker-labeled, timestamped transcripts — and owns ASR audio preprocessing as a first-class job.
Turns a topic or article into a full educational comic — storyboard, consistent characters, generated pages, and a merged PDF.
Designs and generates article cover images using a 5-dimension system of type, palette, rendering style, text level, and mood.
Analyzes your article, decides where illustrations belong, and generates a visually consistent set of images via a Type × Style × Palette system.
Turns any content into a publication-ready infographic by combining 21 layout types with 22 visual styles.
One CLI-backed skill that generates images through 10+ providers — OpenAI, Google, DashScope, MiniMax, Replicate and more.
Turns any article or idea into a 1–10 image card series for social media, with 12 styles, 8 layouts and 3 palettes.
Generate emotionally expressive Chinese and Japanese speech with StepFun's stepaudio-2.5-tts using natural-language instructions and inline prosody directives.
Transcribe up to 30 minutes of audio in a single API call using StepFun's stepaudio-2.5-asr SSE endpoint.
Reliably download YouTube videos, audio, and HLS (m3u8) streams in high quality using yt-dlp and ffmpeg.
Compares two videos and generates a self-contained interactive HTML report with PSNR/SSIM metrics and frame-by-frame visuals.
A CLI skill that generates text and images through a reverse-engineered Gemini Web API, with vision input and multi-turn sessions.