Claude KVM
MCP server — control remote desktops via VNC with a native Swift daemon and Apple Vision OCR
Agent-S - Autonomous GUI Agent
Agent-S is a powerful autonomous agent that can control your computer's graphical interface to complete complex tasks. It combines vision and action understanding to interact with any GUI element.
Unblind
Route images to vision API. Never pretend to see. Never Read/Edit settings.json.
Piranha Vision & DevTools
BJJ video analysis — YOLO pose detection, AI technique analysis, and highlight reels.
Kaito Query Service
AI LLM with Gemini, MiniMax, Replicate, OpenRouter. Vision, search, code review. USDC on Base.
Photographi Mcp
Visual Intelligence Command Center: A Local Computer Vision Engine for Photo Libraries
vision-tools
Five local CLIs that give a text-only agent eyes. They read one shared vision config (`VISION_API_KEY` / `VISION_BASE_URL` / `VISION_MODEL` / `LANG`) â no extra credentials.
Io.Github.PyJudge/Pdf4vllm
PDF reader for vision LLMs. Auto-detects text corruption and switches to image mode.
XRay-Vision
AI-powered codebase analysis — call graphs, security, dead code, complexity. 150+ tools.
Io.Github.AdonaiVera/Fiftyone Mcp Server
Control FiftyOne computer vision datasets through AI assistants using 80+ operators.
AI Analysis Guide
> AI/ML is the technology for extracting value from data. This skill systematically covers all aspects of AI analysis â from machine learning fundamentals, deep learning, natural language processing, and computer vision to practical model development workflows.
game-art-direction
Direct game art and creative vision. Turn GAME_DESIGN into a production-level ART_DIRECTION defining a recognizable visual style, camera and composition, world and character grammar, functional colour/light/material, HUD feedback, motion and transition specs, audio direction, and a signature moment for every screen and mode. Use for what should the game look like, set the art direction, define the visual style. 游æç¾æ¯ä¸åææ¹åãæ GAME_DESIGN 转æçå级 ART_DIRECTIONï¼å®ä¹å¯è¾¨è¯è§è§é£æ ¼ãé头æå¾ãä¸çä¸è§è²è¯æ³ãåè½æ§è²å æè´¨ãçé¢åé¦ãè¿å¨ä¸è½¬åºè§æ ¼ã声鳿¹ååæ¯ä¸ªçé¢/模å¼çç¾åæ¸¸ææ¶å»ãç¨äºå¤ææ¸¸æåºè¯¥é¿ä»ä¹æ ·ãå¶å®æ¸¸æç¾æ¯æ¹åçéæ±ã
Image Recongnition Mcp
MCP server for AI-powered image recognition and description using OpenAI vision models.
aesthetic
Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Integrate localized specialized skills (chrome-devtools, ImageMagick) with native vision intelligence to achieve premium aesthetic standards.
Mobile Device Mcp
AI control of Android/iOS devices. Screenshots, UI tree, AI vision, Flutter, video. 49 tools.
modlens
Plug-in vision for text-only models. Hard rule: when a file path or URL with an image extension (.png, .jpg, .jpeg, .webp, .gif, .heic, .heif) appears anywhere in the conversation (typed by the user, injected as a `[Image: source: <path>]` line, or inside a tag) and you cannot see that image's content, run this skill on it before any other approach: no self-built OCR, no PIL, no tesseract. Also triggers on pasted-image placeholders such as `[Image #1]` and `[Unsupported Image]`. If you can actually see the image, do not use this skill. When unsure, run `modlens guard` before the first read of a session: a deny verdict means the active model has native vision and must read the image itself. Runs the modlens CLI to convert the image into structured JSON evidence: every word transcribed, layout regions, semantics, visual clues. Also use when the user asks how to install, configure, or switch modlens providers (Gemini API key, OpenAI-compatible endpoints, Claude API or Claude Code CLI).
arch-check
Run the charter architecture fitness check â Checkpoint 7 (Light Verify), alongside `crap-analyzer`. Turns the charter's architectural vision from prose into an objective gate: `dae_arch.py` reads the manifest's `architecture:` rules and reports violations.
Mulmocast Vision
Easy and stylish presentation slide generator.
citation-check-skill
Vision-enabled verification gate with web search. Use when users want to (1) verify slides/reports/PDFs/images against authoritative online sources, (2) validate that citations actually exist and say what's claimed, (3) check charts/graphs/tables for accuracy, (4) audit AI-generated content in doc-only mode (no external knowledge). Two modes - search mode validates against web, doc-only mode ensures everything traces to provided documents. Supports content in any language.
Vision for a text-only agent
Your model cannot receive image content directly. When the user pastes an image, Claude Code shows it as an `[Unsupported Image]` placeholder â a text marker with **no usable pixel data**. You cannot read it. Instead, obtain the image through a route you can access, then send it through the `analy
Nature Vision Mcp
Identifies biological species, returning Latin names with confidence scores.
mechanical-engineering-research
Research, write, code, analyze, present, and develop proposals for thermal-fluid mechanical engineering work with source-aware rigor. Use for heat transfer, fluid mechanics, thermodynamics, HVAC, energy systems, turbomachinery, pumps, piping, CFD, experiments, correlations, standards, datasheets, papers, patents, AI/ML tools, computer vision, sequence regression, surrogate modeling, research coding, Overleaf, VS Code, GitHub, git, federal grant proposals, DOE/NSF/NASA-style narratives, invention disclosure, provisional patent support, commercialization, and trade studies. Produces technical briefs, critical literature reviews, proposal narratives, review-criteria responses, manuscript sections, methods, results discussions, data-analysis plans, plots, presentations, design comparisons, calculation plans, reproducible code, repository workflows, disclosure drafts, patent-support packets, and research roadmaps.
Image Recognition Mcp
MCP server for AI-powered image recognition and description using OpenAI vision models.
architect
Designs system architecture and selects technology stack based on vision analysis. Use after vision analysis for technical decisions. Triggers on: design architecture, select tech stack, choose framework.
Snapgrab
URL to screenshot with metadata. Python MCP server. Claude Vision optimized.