url-pdf-download-ocr
Provide a URL to download a report, simultaneously saving both a PDF and a Markdown file locally. The former is for your LLM; the latter is for your colleagues. Your ultimate research tool.
11,743 个媒体任务的工具和技能
Provide a URL to download a report, simultaneously saving both a PDF and a Markdown file locally. The former is for your LLM; the latter is for your colleagues. Your ultimate research tool.
OpenClaw skill for Smallest AI — sub-100ms TTS (Lightning v3.1) and 64ms STT (Pulse) with 30+ languages
OpenClaw skill for organizing scanned documents using local AI models (OCR + classification)
MiniMax multimodal skill — TTS, voice cloning, image generation for OpenClaw
Codex skill for Volcengine Jimeng AI image and video generation
Claude Code skill that animates a static image into a 9:16 video using Google Veo 3.1. Claude reads the image and writes the cinematic prompt directly, no separate vision-model step.
Given a DOI, returns the PDF download URL from Sci-Hub. That's it.
Claude Code skill: send images to (and generate/edit images from) OpenAI Codex CLI. Wraps codex exec -i with three modes + opt-in tmux session.
App Store & Google Play submission readiness audit skill for Claude Code — checks iOS Privacy Manifest, Android Data Safety, permissions, screenshots, age rating, and metadata limits
Fine-tune Qwen3-TTS for text-to-speech with custom voices
A Claude Code skill that generates AI podcast scripts and audio from source content. NotebookLM-style two-person conversations via Podcastfy.
Unified NVIDIA NIM skill for OCR, layout, tables, charts, and reranking across Codex, Claude, and OpenClaw
播客音频转逐字稿 - OpenClaw Skill
Claude Code skill: transcribe & summarize any YouTube video. Two-tier (captions → Whisper). Zero-dep CLI.
Build and maintain a personal or team knowledge base ("second brain") inside an Obsidian vault. Use this skill whenever the user wants to ingest sources (meeting notes, articles, transcripts, brain dumps, recipes, research, project docs, supplier conversations) into a linked wiki, query their notes, save a conversation into their vault, capture a quick thought, log daily work progress, or generate a weekly summary. Triggers on words like "second brain", "knowledge base", "vault", "wiki", "Obsidian", "my notes", "ingest this", "save this to my notes", "remember this", "capture this", "log this", "what did I work on", "weekly summary", and on any request to organize, extract concepts from, or link content across saved material. Works for any source type and any domain — professional, personal, or both.
Hermes agent skill that does blob tracking videos
Codex skill for previewable HTML long images with native Canvas PNG export
Generate diagram/animation videos from natural language using Manim. Agent skill for Claude Code, Codex, and other AI coding agents.
"Converts a PDF (local path or URL) into structured markdown. Handles text-based PDFs, image-based/scanned PDFs (OCR), and mixed PDFs. Outputs research-library markdown with frontmatter, section headings, preserved tables, key statistics, and source attribution. Use when a PDF needs to be read, converted, transcribed, or added to a research library."
Agentic motion videos built with Remotion
跨平台媒体分析报告生成技能
Video analysis skill for Claude Code — frame extraction, scene detection, transcription, parallel subagent analysis
让agent能分析视频内容,需要LLM支持分析图片
Audit a mobile app's MMP/SDK integration via mitmproxy: capture traffic, classify by SDK (Meta, AppsFlyer, Adjust, Branch, TikTok, RevenueCat, Superwall, etc.), generate a PDF audit with platform-specific code fixes.
image-creation skills
Nano Banana - Claude Code skill for AI image generation with Gemini. Optimizes Chinese prompts to English and generates images via Gemini API.
"Generate AI images using Pollinations API. Use when user wants to generate images, create AI art, edit images, or needs visuals for docs/presentations. Triggers on 'generate image', 'create picture', 'make AI art', 'draw something'."
Copilot CLI skill for automated Azure portal screenshot capture with PII redaction
Extract Formulas From PDF
Use when creating or revising technical presentations, engineering report decks, architecture reviews, research summaries, project updates, demo-day talks, or slide-style deliverables that may need HTML, PDF, PPTX, diagrams, flowcharts, sequence diagrams, mind maps, or visual technical summary cards.
"Local PDF extraction skill for OpenClaw using OpenDataLoader CLI. Use when: extracting text from PDFs, improving table/column parsing, preparing data for RAG/embeddings, processing scanned PDFs with OCR, describing images/charts."
Use when the task is to capture reusable screenshots from a local or remote webpage and insert those images into a Markdown article or note. Best for blog posts, Zhihu/公众号 drafts, documentation pages, product pages, or local static-site previews where the user wants selected sections cropped by CSS selector and turned into Markdown image blocks.
Skill from ZhouFung/ocr-skill
The video AI notes tool is provided by Baidu. Based on the video download address provided by the user, it downloads and parses the video, and finally generates AI notes corresponding to the video (a total of three types of notes can be generated: document notes, outline notes, and image-text notes).
Use whenever the user gives a Xiaohongshu note link + a KOC/KOL account name (or mentions /xhs, /xhs-image-gen, "换皮生图", "改成 xx 账号风格", "把这条小红书改成我自己账号的样子"), or asks for native 3:4 XHS posters written into Feishu Base. Any agent (Claude, Codex, or others) can run this skill — image generation uses the agent's built-in image_gen tool when available (Path A, e.g. Codex), or shells out to a `codex exec` sidecar when not (Path B, e.g. Claude). Title and content are also rewritten — title via Skill(dbs-xhs-title) Top-1 pick, content by the running agent itself following dbs-content's five-dimension principles. All artifacts (original + rewritten title/content + generated images) are written back to a Feishu Base task record. Triggers via natural language match, /xhs alias, or /xhs-image-gen slash command.
Skill from james-town77/claude-screenshot-skill
Generate, edit, batch-generate, or troubleshoot images through Azure OpenAI or Microsoft Foundry image model deployments using env vars or an env-file. Use for GPT Image, gpt-image-2, gpt-image-1.5, gpt-image-1, Azure OpenAI image generation, deployment-based image APIs, base64 image outputs, or an Azure-backed alternative to image_gen.
Generate, edit, and restore images using Google Gemini -- from any AI agent.
AI Agents skill that gives agents the capability to render PlantUML diagrams as images
Conversation-driven podcast editing skill for Claude Code and Codex with Groq Whisper, ffmpeg, Gemini cover generation, subtitles, reels, and publishing assets
Generate a PaperBench-compatible rubric.json from an ML paper folder containing a PDF and optional paper.md/addendum, using checkpointed multi-pass analysis, top-down rubric construction, and validator-driven repair.
Minimal runnable secretary skill for TTS playback, audio export, Hermes media output, and daily fixed-time reminders.
视频自动化生成系统 - video-exporter 技能
Generate a virtual tattoo placement preview using the TryInk API. Browse designs by style or body part, then render a tattoo onto any body photo URL.
OpenCode skill for deep server diagnostics and professional PDF report generation (Unraid, Linux, NAS, homelab)
Download YouTube video transcripts with automatic translation to Vietnamese and AI summarization. Perfect for learning English from YouTube, managing knowledge base, and extracting insights from video content. Triggers when user wants to extract, translate, or summarize YouTube video content.
AI agent skill for Plex Media Server — search, watchlist, libraries, sessions via CLI
Give your AI agents memory across sessions. Lightweight session transcript search for OpenClaw.
Stable Diffusion web UI
AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, GitHub Copilot CLI, OpenClaw, Factory Droid, Trae, Google Antigravity). Turn any folder of code, docs, papers, images, or videos into a queryable knowledge graph