Vision-primitives MCP server for text-only LLMs: describe, locate (coordinates), OCR with bbox, annotate, crop, zoom, automated anomaly scanning, computer use — powered by Xiaomi MiMo V2.5. Single-fil
Vision Primitives MCP — 给纯文本 LLM 装上眼睛的视觉原语 MCP 服务器
让纯文本模型(DeepSeek / Codex / 任意 MCP 客户端)通过 28 个 MCP 工具获得完整视觉能力:描述 → 定位(坐标)→ OCR → 标注 → 裁切/放大 → 异常扫描 → 电脑操控。
项目定位:通用视觉推理桥接层——文本模型 + 任意 VLM 的通用视觉工作流,本地隐私 + 单文件轻量。不与 UI-TARS / CogAgent 等端到端 GUI 模型竞争(它们有专门训练);核心能力是无 grounding 模型的兜底定位:somlocate(编号递归)+ cvlocate(颜色/模板)让 MiMo 这类无 grounding 训练的模型达到像素级定位(实测 0-4px,接近甚至超过 grounding 模型)。视觉后端可切换(小米 MiMo V2.5 云端 / LM Studio 本地 Qwen3-VL 等),单文件 Python,核心仅依赖 Pillow;numpy 可选(模板匹配 285x 加速);YOLO 检测器可选(models/icondetect.pt 放置后自动启用)。
测试图(900×600,元素位置为已知真值):红圆中心 (150,140)、绿三角中心 (740,417):
整图定位对比(MiMo V2.5 vs 本地 Qwen3-VL-8B,各模式偏差)
模型 locate 红圆 locate 绿三角 som 红圆 som 绿三角 单次调用 uiparse 全流程 uirefine 全流程 ------------------------ MiMo V2.5(云端) 10-64px(波动) 79-97px(波动) 33-82px(波动) 13-123px(波动) 15-25s 21.5s — MiMo + som-cv(兜底管线) 0px 4px — — 10-12s — — Qwen3-VL-8B(本地) 90px 164px 77px 123px 10-30s 29.4s 357s(超时) Qwen2.5-VL-7B(本地) 28px 27px 15px 123px 1.3-1.7s 12.4s 12.6s
要点:Qwen2.5-VL-7B 的 grounding 专才(RefCOCO 93.7%)+ 非思考型架构,定位精度 3-6 倍、速度 10-20 倍于 Qwen3-VL-8B;绿三角 som 两种模型都锁死(123px,首轮选偏),用 final="cv" 或裁切定位可解。
模型 describe OCR(prompt 修复后) som-cv(颜色目标) uilocate 文本锚定 scratch 论文推理 ------------------ MiMo V2.5(云端) 4.7s ✓ 8.9s,4/4 块 0px(3.2s) 8.6s ✓ 45.7s 3轮(图趋势漏答) Qwen2.5-VL-7B(本地) 1.1s ✓ 3.5-4.6s,4/4 块 0px 12.8s ✓ 9.6s 1轮全对 Qwen3-VL-8B(本地) ~10s 23s,4/4 块 0px — —
要点:OCR 对 prompt 长度敏感(简化指令后 Qwen2.5-VL 召回 1→4 块、链路 10 倍加速);scratch 多轮推理依赖模型的"自主细看"决策(Qwen2.5-VL 一轮全对,MiMo 漏答图趋势);som-cv 的颜色分割部分为纯本地确定性,模型只需选对格子。
From the project README.
Add the radar badge to your README — it shows your project was picked up by MCP Radar and links to this page:
[](https://mcp.liqiwa.com/s/zouyuanqing--vision-primitives-mcp.html)
Pulse — Hermes-style self-improving AI agent. Reliability-first rebuild with evaluated skill self-evolution, multi-agent team orchestration, dialectic user modeling, and fully self-hosted default stac
FFZackFair92/unreal-engine-mcpDrive Unreal Engine 5 from an AI agent — MCP server over the Remote Control API and in-editor Python. No C++ plugin to compile.
russeell/jobfindsme🔥 AI 求职雷达|一句话同时搜 BOSS直聘 + 猎聘,找到匹配你简历的岗位。
MuhittinYilmazer/akanaSelf-hosted, local-first AI assistant server: swappable LLM providers, review-gated memory, an encrypted vault, and wake-word voice. No accounts, no telemetry.
clementrx/Performance-agentOpen-source AI Strength & Conditioning coach for your terminal. Runs in Claude Code / Gemini CLI / Codex — the LLM narrates, a deterministic engine calculates, and it won't lie about unrealistic goals
sai-chaithanya-navuluri/content-coreproduction AI platform kernel: LLM abstraction, workflow engine, RAG, MCP server, FastAPI, evaluation framework, closed-loop feedback
The top new MCP servers of the week, every Monday. No spam, unsubscribe anytime.