Skills

All Skills

ocr

Skills tagged with #ocr

@ARAS-Workspace
MCP

Claude KVM

MCP server — control remote desktops via VNC with a native Swift daemon and Apple Vision OCR

mcpgithub
ARAS-Workspace/claude-kvm
5mo ago
0
@lukisch
MCP

BACH FileCommander

43 tools for filesystem, process management, sessions, search, OCR, ZIP, and PDF export.

mcpgithubsearchfile
lukisch/bach-filecommander-mcp
5mo ago
0
@alvinxx1123

experience-cleaning-skill

Clean OCR and manually entered interview experiences before storage so the RAG corpus has less noise, fewer duplicates, and more consistent company, department, and question fields.

alvinxx1123/sspOffer-interview-assistant+4 more
5mo ago
440
@RECERQA
MCP

Io.Github.RECERQA/Rq Scan

MCP Server for RQ-SCAN - AI-powered document OCR and data extraction platform

mcpgithubai
RECERQA/rq-scan
5mo ago
0
@BitYoungjae

chalkak-ocr-models-release

Release and sync the `chalkak-ocr-models` AUR package by bumping `aur/chalkak-ocr-models/PKGBUILD` and `.SRCINFO`, refreshing checksums against the GitHub asset `ocr-models-vN`, and pushing packaging-only metadata to AUR. Use when publishing OCR model files to AUR, updating model package version/checksum, or fixing `chalkak-ocr-models` AUR metadata.

BitYoungjae/ChalKak+1 more
5mo ago
390
@SamMorrowDrums
MCP

reMarkable MCP Server

Access your reMarkable tablet - read documents, browse files, extract text and OCR

mcpgithubfile
SamMorrowDrums/remarkable-mcp+1 more
5mo ago
0
@ChrBoebel
MCP

Optical Context MCP

Compress OCR-heavy PDFs into dense packed images so agents can work with long visual documents.

mcpgithub
ChrBoebel/optical-context-mcp
5mo ago
0
@filegraph
MCP

Document Processing

Extract text from documents, manipulate PDFs, and perform OCR on images.

mcpaifile
filegraph/docconvert
5mo ago
0
@spencermarx

ocr

AI-powered multi-agent code review. Simulates a team of Principal Engineers reviewing code from different perspectives. Use when asked to review code, check a PR, analyze changes, or perform code review.

spencermarx/open-code-review
5mo ago
300
@sh3ll3x3c
MCP

Native DevTools

MCP server for desktop, browser (CDP), and Android automation. Screenshot, OCR, click, type.

mcpgithubbrowser
sh3ll3x3c/native-devtools-mcp
5mo ago
0
@zai-org

glmocr

Extract text from images using GLM-OCR API. Supports images and PDFs with high accuracy OCR, table recognition, formula extraction, and handwriting recognition. Use this skill whenever the user wants to extract text from images, perform OCR on pictures, scan documents, convert images to text, or process any image files to get their textual content.

zai-org/GLM-OCR+4 more
5mo ago
2.0K0
@linxule
MCP

MinerU

MinerU document parsing API — PDFs, images, DOCX, PPTX with OCR and batch processing.

mcpgithubapi
linxule/mineru-mcp
5mo ago
0
@Deesmo
MCP

Io.Github.Deesmo/Arch Tools Mcp

53 AI tools for agents: web, crypto, AI generation, OCR, and more. Pay with Stripe or USDC.

mcpgithubaiweb
Deesmo/Arch-AI-Tools
5mo ago
0
@liustack

modlens

Plug-in vision for text-only models. Hard rule: when a file path or URL with an image extension (.png, .jpg, .jpeg, .webp, .gif, .heic, .heif) appears anywhere in the conversation (typed by the user, injected as a `[Image: source: <path>]` line, or inside a tag) and you cannot see that image's content, run this skill on it before any other approach: no self-built OCR, no PIL, no tesseract. Also triggers on pasted-image placeholders such as `[Image #1]` and `[Unsupported Image]`. If you can actually see the image, do not use this skill. When unsure, run `modlens guard` before the first read of a session: a deny verdict means the active model has native vision and must read the image itself. Runs the modlens CLI to convert the image into structured JSON evidence: every word transcribed, layout regions, semantics, visual clues. Also use when the user asks how to install, configure, or switch modlens providers (Gemini API key, OpenAI-compatible endpoints, Claude API or Claude Code CLI).

liustack/modlens
1mo ago
3.9K0
@TaewoooPark

answer-processing

Use whenever the user uploads a hand-written or scanned answer PDF to be graded against a reference solution. Converts answer PDFs in `answers/*.pdf` to markdown in `answers/converted/*.md` using the pdf skill (OCR as needed), then performs strategy-based grading against `converted/solutions/*.md` or `quizzes/*_answers.md`. Invoked by `/grade`.

TaewoooPark/PAIDEIA+1 more
4mo ago
380
@QuartzUnit
MCP

Docpick

Schema-driven document extraction with local OCR + LLM. Document in, Structured JSON out.

mcpgithubllm
QuartzUnit/docpick
5mo ago
0
@deapi-ai
MCP

deAPI MCP Server

33 AI tools: transcription, image/video generation, TTS, music, OCR, embeddings via deAPI

mcpgithubapiai
deapi-ai/mcp-server-deapi
5mo ago
0
@wavyrai
MCP

reMarkable MCP Server

Access your reMarkable tablet - read documents, browse files, extract text and OCR handwritten notes

mcpgithubaifile
wavyrai/rm-mcp
5mo ago
0
@AI-Riksarkivet

htr-transcription

Guide for using HTRflow MCP tools to transcribe handwritten documents. Use when: transcribe handwriting, HTR, handwritten document, OCR historical document, read old handwriting, digitize manuscript, transcribe old letters, recognize handwritten text.

AI-Riksarkivet/htrflow_app
5mo ago
390
@damionrashford

rival-search-mcp

Deterministic deep research via RivalSearchMCP. 9 tools: 5-engine web search (DuckDuckGo/Bing/Yahoo/Mojeek/Wikipedia), 9-platform social search (Reddit/HN/StackOverflow/Dev.to/Medium/ProductHunt/Bluesky/Lobste.rs/Lemmy), 5-source news (Google/Bing/Guardian/GDELT/DDG), 5 academic DBs (OpenAlex/CrossRef/arXiv/PubMed/EuropePMC), GitHub search, website mapping, content extraction with OCR, and research topic synthesis. No API keys required. Use when the user needs web research, competitive analysis, content discovery, or academic paper search.

damionrashford/RivalSearchMCP
4mo ago
920