← Back to directory
H

Human MCP

Community
Bringing Human Capabilities to AI Agents
Category
Dev Tools #295 of 438
Stars
★ 292 Popular
Transport
stdio (local process) · Streamable HTTP
Runtime
Node.js 22+ · Bun 1.2+
Credentials
API key / credential required
License
MIT
Last commit
⚠
Tools
31
46FMRS · D

Human MCP is a feature-rich MCP server that offers a wide range of human-like capabilities. It is well-structured and supports multiple providers, making it flexible. However, it is not official, requires external API keys, and may not be suitable for text-only workflows.

Strongest · Documentation 13/20 Weakest · Reliability 4/20

Reliability
4/20
Security and permissions
8/20
Maintenance
9/20
Documentation
13/20
Setup experience
12/20
Why each score
Reliability 4/20
Only README/manifest supplied; no package.json, source, tests, or CI. README claims 29 tools but also states 27 production-ready tools; Brain section lists four tool names under a '3 tools' heading. No way to verify startup, MCP handshake, or tool-to-behavior mapping. Deducted heavily for missing execution evidence and internal inconsistencies; capped at 4.
Security and permissions 8/20
No red-line issues found: no real secrets in examples, no malware or covert exfiltration evidence; API keys are placeholders and R2 upload is disclosed. However, no source code to audit credential storage, least privilege, or confirmation flows. HTTP transport auto-uploads local files to Cloudflare R2 without per-file confirmation; CORS/DNS/rate-limit settings are only documented, not verified. Main risks visible but implementation and scoping unverified; score 8.
Maintenance 9/20
Evidence: MIT license, repo not archived, 3 open issues. Missing: commit history, release cadence, dependency updates, issue response, security-response channel; README version claim v2.16.0 is unverifiable. Scores as 'active but governance/versioning evidence gap'; 9.
Documentation 13/20
README is very extensive: API key setup, multiple client configs (Claude Desktop/Code, OpenCode, Cline, Cursor, Windsurf), env vars, tool JSON examples, formats, troubleshooting, architecture. Deductions: tool-count contradiction (29 vs 27), 'Brain' heading says 3 tools but lists 4, clone URL points to human-mcp/human-mcp instead of mrgoonie/human-mcp, package name not verifiable, cost and limits only partially addressed, no reviewable evidence. Score 13.
Setup experience 12/20
Setup instructions are clear and multi-client, with env var and config JSON examples; Node 22+/Bun prerequisites stated. However, no CI or tests in supplied files, so per static calibration setup cannot exceed 15. Deductions: npx package name (@goonnguyen/human-mcp) not verifiable from repo, README clone URL inconsistent, some CLI commands (claude mcp test) are unverified and may not exist. Score 12.

Static review · not runListed 2026-08-07

Read the FMRS scoring method →

Fit and risk

What it can accessReads local filesUses the networkControls a browser

Best for

  • AI coding agents needing visual debugging capabilities
  • Developers who need automated screenshots and visual analysis
  • Users interested in content generation (images, videos, music, speech)
  • Users who need advanced reasoning capabilities

Not for

  • Users who only need text-based functionality (overkill)
  • Users seeking a fully local solution without API keys (relies on external AI providers)
  • Users needing audio processing (Ears not yet implemented)

Required permissions

  • Requires Google Gemini API key (mandatory)
  • Optional API keys: Minimax, ZhipuAI, ElevenLabs
  • Optional Cloudflare R2 credentials for HTTP transport file uploads
  • Network access to external AI services
  • Local file access to read input files

Risks and side effects

  • API key exposure: keys are stored in config, may be accidentally shared
  • Cost: generative AI usage may incur charges, especially in production
  • Data privacy: files may be uploaded to third-party services (Cloudflare R2) for processing
  • Dependency on external services: relies on availability of Google, Minimax, etc.
  • Configuration errors: invalid API keys or network issues causing failures

Setup

Before you start

Runtime:Node.js 22+ · Bun 1.2+

GOOGLE_GEMINI_API_KEY requiredsecret Google Gemini API key, required for core visual analysis; create it at Google AI Studio (aistudio.google.com).
MINIMAX_API_KEY optionalsecret Minimax platform key for speech (Speech 2.6), music (Music 2.5), and video (Hailuo 2.3).
ZHIPUAI_API_KEY optionalsecret ZhipuAI key for GLM-4.6V vision, CogView-4 images, and CogVideoX-3 video.
ELEVENLABS_API_KEY optionalsecret ElevenLabs key for TTS, music generation, and sound effects.
HTTP_SECRET optionalsecret Auth secret for HTTP transport mode.
CLOUDFLARE_CDN_ACCESS_KEY optionalsecret Cloudflare R2 access key with R2:Object:Write permission.
CLOUDFLARE_CDN_SECRET_KEY optionalsecret Cloudflare R2 secret key.
Other optional settings (16)
USE_VERTEX optional Set to 1 to switch to Vertex AI auth instead of a Gemini API key.
VERTEX_PROJECT_ID optional GCP project ID for Vertex AI, used in production deployments.
VERTEX_LOCATION optional Vertex AI region, defaults to us-central1.
GOOGLE_APPLICATION_CREDENTIALS optional Path to a service account JSON for Vertex AI production auth.
SPEECH_PROVIDER optional Default speech provider (gemini/minimax/elevenlabs).
VIDEO_PROVIDER optional Default video provider (gemini/minimax/zhipuai).
VISION_PROVIDER optional Default vision provider (gemini/zhipuai).
IMAGE_PROVIDER optional Default image provider (gemini/zhipuai).
TRANSPORT_TYPE optional Transport mode: stdio, http, or both; stdio is default.
HTTP_PORT optional HTTP transport listen port, defaults to 3000.
HTTP_HOST optional HTTP server bind address.
LOG_LEVEL optional Log verbosity, e.g. info.
MCP_TIMEOUT optional MCP request timeout in milliseconds, shown in Claude Code config examples.
CLOUDFLARE_CDN_BUCKET_NAME optional Cloudflare R2 bucket name for automatic local-file uploads in HTTP mode.
CLOUDFLARE_CDN_ENDPOINT_URL optional R2 S3-compatible endpoint URL.
CLOUDFLARE_CDN_BASE_URL optional Public CDN base URL for R2-hosted files.
  1. Obtain a Google Gemini API key from Google AI Studio.
  2. Set the environment variable via 'export GOOGLE_GEMINI_API_KEY="your_api_key"' or directly provide it in client config.
  3. For Claude Desktop, edit the config file (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json) to add the mcpServers entry using the npx command.
  4. For Claude Code, use claude mcp add --scope user human-mcp npx @goonnguyen/human-mcp --env GOOGLE_GEMINI_API_KEY=your_key.
  5. For other clients, refer to specific configurations in the README.
claude_desktop_config.json
{
  "mcpServers": {
    "human-mcp": {
      "command": "npx",
      "args": [
        "@goonnguyen/human-mcp"
      ],
      "env": {
        "GOOGLE_GEMINI_API_KEY": "your_gemini_api_key_here"
      }
    }
  }
}

Shown for Claude Desktop. Other clients may use a different file or key (VS Code uses "servers") — the configurator below converts it.

.vscode/mcp.json
{
  "servers": {
    "human-mcp": {
      "command": "npx",
      "args": [
        "@goonnguyen/human-mcp"
      ],
      "env": {
        "GOOGLE_GEMINI_API_KEY": "your_gemini_api_key_here"
      }
    }
  }
}

Goes in your project's .vscode/mcp.json (VS Code uses a "servers" key).

Terminal
claude mcp add human-mcp -e GOOGLE_GEMINI_API_KEY=your_gemini_api_key_here -- npx @goonnguyen/human-mcp

Run it in a terminal; replace any <…> placeholders with your own values first.

Check that it works

The client's tool list should show the 29 human-mcp tools (e.g. eyes_analyze, gemini_gen_image, mouth_speak); send 'use eyes_analyze on this test screenshot' to confirm the API connection works.

Troubleshooting

  1. Connection failed: check Node.js/npm or Bun installation, verify API key, check network connectivity
  2. Tool not found: restart MCP client, check server logs
  3. API errors: validate Google Gemini API key, check quota and usage limits, review network/firewall
  4. Permission errors: check npm global installation permissions, use npx instead
  5. Vertex AI authentication: verify GCP project ID, ADC status, and IAM permissions

Things to try

Once connected, you can ask your AI assistant things like:

  • Use eyes_analyze to check this UI screenshot for layout bugs and accessibility issues
  • Use gemini_gen_image to generate a modern dark-theme login page mockup at 16:9
  • Use jimp_crop_image to crop this image to a centered 800x600 area, then rmbg_remove_background to remove the background
  • Use playwright_screenshot_fullpage to capture the full page at https://example.com/dashboard

Tools 31

eyes_analyze read-only
Analyze images, videos, and GIFs for UI bugs, errors, and accessibility
eyes_compare read-only
Compare two images to find visual differences
eyes_read_document read-only
Extract text and data from PDF, DOCX, XLSX, PPTX, and more
eyes_summarize_document read-only
Generate summaries and insights from documents
gemini_gen_image writes
Generate images from text using Imagen API
gemini_gen_video writes
Generate videos from text with Veo 3.0
gemini_image_to_video writes
Animate images into videos
minimax_gen_music writes
Generate music with vocals using Minimax Music 2.5
Show 23 more tools
elevenlabs_gen_music writes
Generate music using ElevenLabs Music API
elevenlabs_gen_sfx writes
Generate sound effects from text descriptions
gemini_edit_image writes
Comprehensive AI image editing (inpainting, outpainting, style transfer, etc.)
gemini_inpaint_image writes
Add/modify areas in images without mask
gemini_outpaint_image writes
Expand image borders
gemini_style_transfer_image writes
Apply artistic styles to images
gemini_compose_images writes
Combine multiple images
jimp_crop_image writes
Crop images (manual, center, etc.)
jimp_resize_image writes
Resize images with multiple algorithms
jimp_rotate_image writes
Rotate images
jimp_mask_image writes
Apply grayscale masks
rmbg_remove_background writes
AI background removal with three quality levels
playwright_screenshot_fullpage read-only
Capture full page screenshot including scrollable content
playwright_screenshot_viewport read-only
Capture visible viewport screenshot
playwright_screenshot_element read-only
Capture screenshot of a specific element
mouth_speak writes
Convert text to speech with 30+ voices and 24 languages
mouth_narrate writes
Long-form content narration with chapter breaks
mouth_explain writes
Generate spoken code explanations with technical analysis
mouth_customize writes
Test and compare different voices and styles
mcp__reasoning__sequentialthinking read-only
Native sequential thinking with thought revision
brain_analyze_simple read-only
Fast pattern-based analysis (problem solving, root cause, SWOT, etc.)
brain_patterns_info read-only
List available reasoning patterns and frameworks
brain_reflect_enhanced read-only
AI-powered meta-cognitive reflection for complex analysis

Use cases

Debugging UI issues: analyze screenshots to find layout problems
Error investigation: analyze screen recordings to detect errors
Accessibility audit: check UI for accessibility compliance
Image generation for design prototypes
Video generation for marketing materials
Code explanation audio: generate spoken explanations for code reviews
Documentation narration: convert technical docs to audio
Advanced problem solving: analyze complex technical issues with multi-step reasoning
Automated web screenshots for documentation and testing

Supported clients

Claude Desktop
Claude Code
OpenCode
Cline
Cursor
Windsurf

Listed from the project's documentation, not tested by this site.

Overview

Human MCP is a comprehensive Model Context Protocol server that provides AI coding agents with human-like capabilities including visual analysis, document processing, speech generation, content creation, image editing, browser automation, and advanced reasoning. It supports multiple AI providers (Google Gemini, Minimax, ZhipuAI, ElevenLabs) and offers 29 tools organized into four categories: Eyes (vision), Hands (content generation), Mouth (speech), and Brain (reasoning). The server is installed via npx and requires API keys for configuration.

Similar servers

Context7 80 · B

Upstash's official server providing up-to-date third-party library docs for AI coding assistants

★ 62.9k · Tools 2 Compare with this →

Source revision e65a87855811 Data synced 2026-10-11 Read the FMRS scoring method