← Back to directory
A

AI Vision MCP Server

Community
AI-powered image and video analysis MCP server using Gemini and Vertex AI
Category
Other #42 of 230
Stars
★ 80 Popular
Transport
stdio (local process)
Runtime
Node.js 18+
Credentials
API key / credential required
License
MIT
Last commit
⚠
Tools
5
64FMRS · C

The AI Vision MCP server provides comprehensive image and video analysis tools with multiple providers and file sources, designed for the MCP ecosystem. Its strengths include flexibility, validation, and error handling, but it requires external API keys and attention to data privacy and cost.

Strongest · Documentation 15/20 Weakest · Reliability 10/20

Reliability
10/20
Security and permissions
14/20
Maintenance
12/20
Documentation
15/20
Setup experience
13/20
Why each score
Reliability 10/20
Evidence: The source structure is well-organized (src with providers, services, storage, etc.), and the declared tools match the file structure. Zod validation and error handling with retry and circuit breakers are implemented. However, no tests or CI workflows are present in the repository, so we cannot verify that the server actually starts and completes the MCP handshake as described. Deduction: lack of test evidence; static review cannot confirm behavior.
Security and permissions 14/20
Evidence: Credentials are passed via environment variables, and the README examples do not contain real secrets. Temporary file handling for annotated images is implemented. However, there is no evidence of least-privilege design, confirmation for dangerous operations (e.g., writing files, external network calls), or explicit data-flow disclosure for GCS uploads. No red-line issues like malware or credential theft were found. Deduction: missing documentation and confirmation mechanisms for security-sensitive operations.
Maintenance 12/20
Evidence: The repository is active (71 stars, 5 open issues), has an MIT license, and includes contribution guidelines and development scripts. However, there are no release tags, dependency update policy, or security response channel mentioned. Deduction: governance and versioning gaps.
Documentation 15/20
Evidence: The README is extensive, covering multiple client configurations (Claude Desktop, Claude Code, Cursor, Cline), an environment variable guide (60+ variables), tool examples with JSON, and troubleshooting. Architecture and error handling are described. Deduction: some limitations (e.g., video only YouTube) are noted, but cost, rate limits, and full data flow are not fully disclosed.
Setup experience 13/20
Evidence: Installation steps are clear, with support for npx and source build, and specific configuration code for several main clients. Timeout adjustments are recommended. However, without CI or tests, we cannot verify that the setup works out-of-the-box, and some platform-specific issues (e.g., Windows paths) are only partially addressed. Deduction: no execution evidence and incomplete verification of all environment variable combinations.

Static review · not runListed 2026-08-07

Read the FMRS scoring method →

Fit and risk

What it can accessReads local filesWrites / deletes local filesUses the network

Best for

  • Developers looking to quickly integrate AI vision capabilities, especially within Google's AI ecosystem.
  • Design teams needing UI/UX compliance and accessibility audits.
  • Workflows requiring automated image and video analysis.

Not for

  • Not suitable for scenarios requiring offline or on-premise AI models.
  • Not for integrations with non-Google cloud resources (e.g., AWS).
  • May not be ideal for complex video editing or real-time streaming needs.

Required permissions

  • Requires access to Google AI Studio API or Vertex AI API credentials.
  • For local file analysis, needs read access to the local filesystem.
  • If using Vertex AI, needs access to a Google Cloud Storage bucket.
  • For video analysis, needs access to YouTube or local video files.

Risks and side effects

  • API key leakage: environment variables contain sensitive credentials; store them securely.
  • Data privacy: images/videos sent to Google Cloud may contain sensitive information; ensure compliance.
  • Cost: API usage may incur costs based on volume.
  • Potential timeout or transport errors if not configured properly.

Setup

Before you start

Runtime:Node.js 18+

IMAGE_PROVIDER required Provider for image analysis, google or vertex_ai; always required
VIDEO_PROVIDER required Provider for video analysis, google or vertex_ai; always required
GEMINI_API_KEY optionalsecret Google AI Studio API key, required for the google provider; get it at aistudio.google.com
VERTEX_PRIVATE_KEY optionalsecret Vertex AI service account private key (PEM), required for the vertex_ai provider
Other optional settings (9)
VERTEX_CLIENT_EMAIL optional Vertex AI service account email, required for the vertex_ai provider
VERTEX_PROJECT_ID optional GCP project ID, required for the vertex_ai provider
GCS_BUCKET_NAME optional Google Cloud Storage bucket name, needed for Vertex AI video upload
MCP_TIMEOUT optional MCP startup timeout in ms; Claude Code docs suggest 60000
MCP_TOOL_TIMEOUT optional MCP tool execution timeout in ms; docs suggest about 300000
TEMPERATURE_FOR_DETECT_OBJECTS_IN_IMAGE optional Temperature for detect_objects_in_image, default 0.0 for deterministic output
TOP_P_FOR_DETECT_OBJECTS_IN_IMAGE optional Nucleus sampling top_p for detect_objects_in_image, default 0.95
TOP_K_FOR_DETECT_OBJECTS_IN_IMAGE optional Top-k vocabulary selection for detect_objects_in_image, default 30
MAX_TOKENS_FOR_DETECT_OBJECTS_IN_IMAGE optional Max token limit for detect_objects_in_image, default 8192
  1. Install Node.js 18+ and npm.
  2. Configure environment variables: either for Google AI Studio (set IMAGE_PROVIDER=google, VIDEO_PROVIDER=google, GEMINI_API_KEY) or Vertex AI (set IMAGE_PROVIDER=vertex_ai, VIDEO_PROVIDER=vertex_ai, VERTEX_CLIENT_EMAIL, VERTEX_PRIVATE_KEY, VERTEX_PROJECT_ID, GCS_BUCKET_NAME).
  3. Add the server to your MCP client: For Claude Desktop, add a configuration like { "mcpServers": { "ai-vision-mcp": { "command": "npx", "args": ["ai-vision-mcp"], "env": { ... } } } } to your config file. For Claude Code, use claude mcp add ai-vision-mcp -e ... -- npx ai-vision-mcp. For Cursor, add to ~/.cursor/mcp.json. For Cline, add to cline_mcp_settings.json.
  4. Increase MCP client timeout to >1 minute and tool execution timeout to ~5 minutes.

Check that it works

After install, the client's tool list should show analyze_image, compare_images, detect_objects_in_image, audit_design, and analyze_video; run analyze_image with a question like 'What is this image about?' to confirm the connection works.

Troubleshooting

  1. If 'Transport closed' appears, ensure no logs are written to stdout; use console.error instead.
  2. Check that the imagescript dependency loads correctly: run npm run doctor or npm run check:imagescript.
  3. Verify environment variables are set correctly, especially API keys and provider selection.
  4. Increase MCP client timeout settings to avoid timeouts during long analyses.

Things to try

Once connected, you can ask your AI assistant things like:

  • Analyze this image and describe it in detail: https://example.com/image.jpg
  • Compare these two images and tell me which has the best lighting quality
  • Detect all objects in this image and generate an annotated image with bounding boxes
  • What is this YouTube video about: https://www.youtube.com/watch?v=9hE5-98ZeCg

Tools 5

analyze_image read-only
Analyzes an image and returns a detailed description, with modes for general, palette, hierarchy, and components.
compare_images read-only
Compares multiple images (2-4) and returns a detailed comparison analysis.
detect_objects_in_image writes
Detects objects in an image and generates annotated images with bounding boxes, returning coordinates and a summary.
audit_design read-only
Audits UI/UX design compliance with pixel analysis, WCAG contrast checks, and AI critique.
analyze_video read-only
Analyzes a video and returns a detailed description, supporting YouTube URLs, GCS URIs, or local file paths.

Use cases

Describe image content for accessibility or content moderation.
Compare design mockups or screenshots.
Automate object detection for inventory or quality control.
Audit UI designs for accessibility compliance (WCAG).
Analyze video content for summaries or insights.

Supported clients

Claude Desktop
Claude Code
Cursor
Cline

Listed from the project's documentation, not tested by this site.

Overview

The AI Vision MCP server is a Model Context Protocol (MCP) server that provides AI-powered image and video analysis using Google Gemini and Vertex AI models. It supports multiple file sources (URLs, local files, base64), offers tools for image analysis, image comparison, object detection with bounding boxes, UI/UX design auditing, and video analysis. The server features dual provider support (Google and Vertex AI), built-in Google Cloud Storage integration, Zod-based validation, and robust error handling with retries and circuit breakers.

Similar servers

BioMCP 75 · B

One binary. One grammar. Evidence from the biomedical sources you already trust.

★ 653 · Tools 11 Compare with this →

Source revision ad9acf02bf44 Data synced 2026-10-11 Read the FMRS scoring method