codex-vision
Claude Code skill: send images to (and generate/edit images from) OpenAI Codex CLI. Wraps codex exec -i with three modes + opt-in tmux session.
メディアタスク向けの 11,743 件のツールとスキル
Claude Code skill: send images to (and generate/edit images from) OpenAI Codex CLI. Wraps codex exec -i with three modes + opt-in tmux session.
"Generate images via Google Nano Banana (Gemini Image API). Use when: user asks to generate, create, or make an image. NOT for: video generation, image analysis/understanding."
Agentic motion videos built with Remotion
Codex skill for turning Bilibili videos into Obsidian study notes with subtitle download, keyframe capture, formula correction, and supplemental explanations.
Video analysis skill for Claude Code — frame extraction, scene detection, transcription, parallel subagent analysis
Standalone TTS CLI for Apple Silicon using Qwen3-TTS and MLX. No server needed.
"Local PDF extraction skill for OpenClaw using OpenDataLoader CLI. Use when: extracting text from PDFs, improving table/column parsing, preparing data for RAG/embeddings, processing scanned PDFs with OCR, describing images/charts."
Skill for content creators designing educational media: plan objectives, outline, and scripts that are platform-aware and easy to publish.
Comprehensive portfolio review for photography. Extract EXIF metadata, analyze composition/technical execution using vision models, generate ratings and critiques. Use when reviewing multiple photos for portfolio curation, getting feedback on photo quality, or identifying strongest work.
播客音频转逐字稿 - OpenClaw Skill
Claude Code skill — transcribe & summarize Apple Voice Memos on-device using Apple Silicon GPU (MPS). M1/M2/M3/M4 only.
Documentary-style video editing skill for Claude — sound bite selection, narrative assembly, FCPXML output
OpenClaw skill for local text-to-speech using Edge TTS
Use when creating or revising technical presentations, engineering report decks, architecture reviews, research summaries, project updates, demo-day talks, or slide-style deliverables that may need HTML, PDF, PPTX, diagrams, flowcharts, sequence diagrams, mind maps, or visual technical summary cards.
OpenClaw skill for joining meetings and transcription
一个AI精准口播剪辑应用,做自媒体最好用的口播剪辑助手!
ContentStudio is a tool to schedule social-media posts across Facebook, LinkedIn, Twitter/X, Instagram, YouTube, TikTok, Pinterest, and Google Business Profile. Use when the user wants to list/create/delete/approve posts, manage media, or audit workspaces, accounts, campaigns, labels, categories, or team-members on their ContentStudio account.
OpenCode skill for deep server diagnostics and professional PDF report generation (Unraid, Linux, NAS, homelab)
Use this skill to generate well-branded interfaces and assets for Therapist Resources, either for production or throwaway prototypes/mocks. Contains essential design guidelines, colors, type, fonts, voice DNA, logos, and UI kit components.
一款基于 Qwen2-VL-2B-Instruct 构建的、可用于生产的图像转文本 (OCR) 工具,能够识别屏幕截图、文档、收据、名片、表格和复杂布局中的文本,并支持中文、英文、日文、韩文和混合语言内容。
Build and maintain a personal or team knowledge base ("second brain") inside an Obsidian vault. Use this skill whenever the user wants to ingest sources (meeting notes, articles, transcripts, brain dumps, recipes, research, project docs, supplier conversations) into a linked wiki, query their notes, save a conversation into their vault, capture a quick thought, log daily work progress, or generate a weekly summary. Triggers on words like "second brain", "knowledge base", "vault", "wiki", "Obsidian", "my notes", "ingest this", "save this to my notes", "remember this", "capture this", "log this", "what did I work on", "weekly summary", and on any request to organize, extract concepts from, or link content across saved material. Works for any source type and any domain — professional, personal, or both.
Codex skill for AI image generation — supports built-in tool, custom endpoints (9Router, etc), auto-reads auth.json
All-in-one: Torrents + Subtitles + Player - Streamlined media experience
Use when the task is to capture reusable screenshots from a local or remote webpage and insert those images into a Markdown article or note. Best for blog posts, Zhihu/公众号 drafts, documentation pages, product pages, or local static-site previews where the user wants selected sections cropped by CSS selector and turned into Markdown image blocks.
Batch OCR for scanned PDFs using vision-language models. Supports customizable system prompts for context-aware extraction of blurry, multilingual, and special-character documents.
Generate a PaperBench-compatible rubric.json from an ML paper folder containing a PDF and optional paper.md/addendum, using checkpointed multi-pass analysis, top-down rubric construction, and validator-driven repair.
A skill for anchoring style consistency across batch-generated images and text, with drift control and imagegen-ready texture workflows.
"Converts a PDF (local path or URL) into structured markdown. Handles text-based PDFs, image-based/scanned PDFs (OCR), and mixed PDFs. Outputs research-library markdown with frontmatter, section headings, preserved tables, key statistics, and source attribution. Use when a PDF needs to be read, converted, transcribed, or added to a research library."
Generate ready-to-paste prompts for Suno AI to create Thai-language songs (Thai-only or Thai+English code-switching). Use this skill whenever the user mentions Suno, AI music generation, แต่งเพลงด้วย AI, สร้างเพลงไทย, generating song lyrics for AI, T-pop/Thai pop/Thai rock songwriting, or wants to produce the full Lyrics + Styles + Title + slider settings (Vocal Gender, Weirdness, Style Influence) for Suno v4.5/v4.5+/v4.5-all. Trigger this skill on any phrase like "ช่วยแต่งเพลง suno", "อยากทำเพลงไทย AI", "เขียน prompt สำหรับ suno", "Thai song lyrics for Suno", or any request that combines Thai songwriting with an AI music generator — even if Suno isn't named explicitly. The skill produces output formatted for direct paste into Suno's Custom Mode fields, follows Thai pop/rock conventions (tonal-melodic alignment, ฉันทลักษณ์, common chord-progression-friendly phrasing), warns about likely tone-mispronunciation (เพี้ยน) risk, and gives iteration tips.
Unified NVIDIA NIM skill for OCR, layout, tables, charts, and reranking across Codex, Claude, and OpenClaw
"Build or update interactive JSX lesson apps in a workspace that follows the `<workspace_root>/<course>/claude_lessons/<slug>/` layout. Supports two modes: (1) new mode — build a lesson from scratch via a 6-phase multi-agent pipeline (scoping → content analysis → plan + approval gate → execution → review → deploy); (2) update mode — modify an existing lesson in place (content, media, structure). Trigger when the user asks to create, build, make, write, or add a lesson, OR when the user asks to update, rework, revise, improve, refresh, modify, tweak, fix, or enhance an existing lesson, OR references an existing lesson by course and slug. Replaces the legacy jsx-lesson skill for new builds and all updates."
Transform YouTube videos into structured knowledge — TL;DR, key takeaways, timestamped claims, topic timeline, and notable quotes. A Claude Code / Codex / Gemini skill.
一个 Agent Skill,根据 B 站视频字幕或者音频 STT 转写结果生成视频内容总结 prompt
PDF skill for AI coding agents — generate polished PDFs from topics or Markdown files
A development skill for documentation-grounded Unitree G1 SDK, ROS2, DDS, service-interface, and RealSense camera workflows.
"AI-native vertical video engine with niche intelligence. Takes a one-line topic and a niche profile, and outputs a finished YouTube Short/Reel/TikTok with AI-generated b-roll, voiceover, burned-in captions, background music, and thumbnail. Supports multiple LLM providers (Claude, Gemini, GPT, Ollama), TTS providers (Edge TTS, ElevenLabs), and 15+ content niches."
Resolve a paper from title, DOI, paper page URL, or direct PDF URL; fetch available source text; use the latest available frontier subagents for deep method and experiment analysis; then save a comprehensive Korean Markdown summary.
Given a DOI, returns the PDF download URL from Sci-Hub. That's it.
A creative OS for kids where OpenClaw turns ideas into games, stories, music, art, science apps, and more.
"Use when the user wants to run API-backed NeMo-Skills generation on NERSC Perlmutter with sshproxy MFA, podman-hpc image bootstrap, remote preflight checks, job submission, polling, and result retrieval. Require an absolute env file, an absolute local input.jsonl path, an absolute local prompt.yaml path, and a supported model."
Conversation-driven podcast editing skill for Claude Code and Codex with Groq Whisper, ffmpeg, Gemini cover generation, subtitles, reels, and publishing assets
A Claude Code skill that generates AI podcast scripts and audio from source content. NotebookLM-style two-person conversations via Podcastfy.
GPT-Image-2 + Seedance 2.0 AI Video Production Pipeline — automated image generation, video creation, narration, and post-production
Hermes agent skill that does blob tracking videos
跨平台媒体分析报告生成技能
从视频中提取字幕或转录音频,输出 SRT/TXT/VTT/JSON/TSV 格式。支持 Bilibili 和 YouTube。用户提供视频 URL 并要求提取字幕、转录文本时使用。
Provide a URL to download a report, simultaneously saving both a PDF and a Markdown file locally. The former is for your LLM; the latter is for your colleagues. Your ultimate research tool.
Skill from ZhouFung/ocr-skill
Claude Code skill for the Kommodo API. Search recordings, read AI summaries, download transcripts, rename/retag, and publish recordings as pages.
Claude Code skill: Download all images from an esa post for local viewing