</>MCP Agents Market
Plugin

modlens

by liustack3.8kTypeScriptUpdated 2026-08-30

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

DeepSeek HarnessClaude CodeCodexOpenCodePi

ModLens is a vision plugin designed for DeepSeek Harness and other coding agent frameworks that enables text-only AI models to process and analyze images. It allows users to paste images directly into chat interfaces without saving files first, automatically converting visual content into structured JSON evidence including OCR text, layout regions, and semantic metadata. The plugin supports multiple vision engines through a failover chain, including free options like Gemini API and Antigravity CLI, as well as reusing existing credentials from Claude Code, Codex, OpenCode, and Pi. It installs with a single command and requires zero configuration changes to harness settings, working seamlessly with DeepSeek-V4, GLM, and other text-only coding models.

Key Features

Direct image pasting into chat interfaces without file system operations, with automatic tool triggering
Structured JSON output containing OCR transcription, reading-order layout regions, and entity/relation extraction
Ten vision engine sources: six built-in API providers (Gemini, OpenAI-compatible, Anthropic, Antigravity CLI, Claude CLI, Kimi CLI) with automatic failover
Zero-configuration installation that reuses existing credentials from Claude Code, Codex, OpenCode, and Pi
OpenAI-compatible universal socket supporting qwen-vl, GLM, SiliconFlow, OpenRouter, vLLM, Ollama, and custom gateways
Comma-separated API key rotation on authentication, rate-limit, or quota failures
One-command DeepSeek Harness installation via npx with automatic model variant wrapping (e.g., DeepSeek-V4-Flash (modlens vision))
Non-invasive architecture requiring no hooks, wrappers, or configuration file modifications

Use Cases

  • 01Enabling DeepSeek-V4, GLM-5.3, and other text-only coding models to read screenshots, diagrams, and UI mockups during development
  • 02Analyzing dense charts, scatter plots, and data visualizations with 128+ elements for technical documentation
  • 03Reading tweet screenshots including author, caption, engagement metrics, and embedded images
  • 04Processing presentation slides, reading titles, layout structure, and content for meeting summaries
  • 05Batch processing multiple images in a single conversation to compare visual families or design iterations
  • 06Extracting structured data from receipts, forms, and documents via OCR with layout preservation

Related Plugins

View more

modlens — FAQ

What is ModLens and what does it do?+

ModLens is a vision plugin that adds image-reading capabilities to text-only AI coding agents like DeepSeek-V4 and GLM-5.3. When you paste an image or provide a file path, it automatically sends the image to a vision engine and returns structured JSON evidence (OCR text, layout regions, semantic metadata) that the text model can use to answer questions about the image.

How do I install ModLens in DeepSeek Harness?+

Run the command `npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.25.2` to install instantly. After installation, run `modlens doctor` to check which vision engines are available, and optionally configure a free Gemini API key or reuse existing credentials from other agent CLIs on your machine.

Which AI clients and harnesses does ModLens work with?+

ModLens is verified to work with DeepSeek Harness (dsh), Claude Code, Codex, OpenCode, and Pi. The installation process is one command for DeepSeek Harness and follows skill installation procedures for the other frameworks.

Do I need an API key to use ModLens?+

Not necessarily. ModLens offers multiple options: you can use the free Antigravity CLI (no key required, just browser sign-in), a free Gemini API key from Google AI Studio (no credit card), or reuse existing credentials from Claude Code, Codex, OpenCode, or Pi if already signed in. Paid API keys from OpenAI, Anthropic, or any OpenAI-compatible provider also work.

Is ModLens free to use?+

Yes, the ModLens plugin itself is free and open-source under the MIT License. Vision engine costs depend on your chosen provider: Antigravity CLI and Gemini API both offer free tiers, while reusing credentials from your existing agent subscriptions (Claude Code, Kimi Code, etc.) consumes those respective quotas.

How does the failover chain work when multiple vision engines are configured?+

ModLens tries configured vision engines in order: fast API providers (Gemini, OpenAI-compatible, Anthropic) attempt first in 5-10 seconds, then agent CLIs (Claude, Codex, OpenCode, Pi) serve as backup in 15-45 seconds. The first successful result wins, and meta.attempts logs every attempt so fallbacks are never silent.

How do I install modlens?+

Open the source repository on GitHub and follow its README. modlens is a plugin — MCP Agents Market links you directly to the official repo.

Is modlens free?+

modlens is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches