← Back to directory
S

Supadata MCP Server

Official
Video transcripts and web scraping for LLM clients via Supadata
Category
Dev Tools #199 of 438
Stars
★ 63 Popular
Runtime
Node.js
Credentials
API key / credential required
License
MIT
Last commit
Tools
9
54FMRS · D

Supadata MCP Server is an officially maintained server covering video transcripts, web scraping/crawling, and AI-powered structured extraction, with built-in retry and rate-limiting. It suits developers needing multi-platform content ingestion, but relies on a third-party Supadata API key and service, making it unsuitable for use cases that can't tolerate sending data off-device or that have strict compliance requirements.

Strongest · Documentation 13/20 Weakest · Reliability 8/20

Reliability
8/20
Security and permissions
12/20
Maintenance
11/20
Documentation
13/20
Setup experience
10/20
Why each score
Reliability 8/20
No manifest was cached, and the prompt provides no source code, package.json, CI configuration, or test file contents. The README only lists 9 tools (transcript, metadata, extract, scrape, map, crawl plus three status-check tools) with CLI-style usage snippets, but there is no way to verify these examples match the actual implementation. 'npm test' is mentioned but no test cases or coverage are shown, and no CI workflow evidence exists. Per the static calibration rule, reliability cannot exceed 12 without verifiable execution evidence (real CI workflows plus committed tests covering key paths); given the near-total absence of source/manifest evidence, the happy path is only plausible on paper while error handling, dependency control, and edge cases are entirely unverifiable, warranting a score of 8.
Security and permissions 12/20
The server requires a SUPADATA_API_KEY and uses it to call a third-party Supadata API that fetches/scrapes arbitrary user-supplied URLs (YouTube/TikTok/Instagram/Twitter/generic web pages) — a clear external network and third-party data-flow dependency. The README does not disclose what data is sent to Supadata, retention policy, or any local logging/caching exposure. The crawl tool's default of up to 100 pages, while not destructive (no delete/pay/remote-command defaults), lacks any stated confirmation step for broad scraping/crawling, and there is no mention of least-privilege scoping such as domain allowlisting. No fabricated identity, malware, secret leakage, or known critical unmitigated vulnerability was found, so the 0-4 red-line does not apply, but incomplete permission-scoping and data-boundary disclosure caps the score at 12.
Maintenance 11/20
The repository is not archived, uses the MIT license, and shows 62 stars with 4 open issues, indicating some activity, but the provided material contains no commit history, release/version evidence, dependency update cadence, or security-response channel. The Contributing section is a generic template (fork, branch, npm test, PR) without stated maintainer response times or versioning policy. Evidence is insufficient to confirm sustained maintenance quality, resulting in a score of 11 — basic governance signals present but largely unverified.
Documentation 13/20
The README is well organized, covering feature highlights, environment variable configuration, individual sections for all 9 tools with CLI usage examples, and a helpful tool-selection comparison table. However, notable gaps remain: actual client setup/connection steps are entirely offloaded to an external integration guide (docs.supadata.ai) with no in-repo JSON config example for Claude/Cursor/etc.; there is no discussion of API rate limits, cost, error codes, or troubleshooting; and the retry/rate-limit config values are listed without explaining their real-world triggers or user impact. This warrants a score of 13 — usable but with hidden assumptions and troubleshooting gaps.
Setup experience 10/20
Installation nominally follows standard npm steps (install/build/test), but the actual MCP client connection configuration (e.g., a Claude Desktop mcpServers JSON snippet) is not included in this repository's README — it is deferred to an external integration guide whose accuracy and step count cannot be verified from the given material. Per the static calibration rule, setup cannot exceed 15 without CI-plus-test execution evidence; given the missing in-repo connection example and reliance on only an env-var instruction, the path from repo to working client connection is not demonstrably few-step or self-contained, yielding a score of 10.

Static review · not runListed 2026-08-26

Read the FMRS scoring method →

Fit and risk

What it can accessUses the network

Best for

  • Developers integrating video transcripts into content pipelines
  • AI application builders needing combined web and video data for research or summarization
  • Automation workflows requiring site-wide URL discovery and multi-page crawling

Not for

  • General-purpose browser automation requiring clicks, logins, or form interaction
  • Scraping targets with strict terms-of-service constraints that haven't been reviewed
  • Privacy-sensitive scenarios where sending target URLs or content to a third-party API is unacceptable

Required permissions

  • Requires a valid SUPADATA_API_KEY to call the Supadata API
  • Requires outbound network access to reach target video platforms and web URLs

Risks and side effects

  • Target URLs and scraped content are sent to and processed by the third-party Supadata service
  • A misconfigured or leaked API key could be misused and incur unexpected usage costs
  • Scraping or crawling target sites may violate their terms of service or copyright restrictions
  • Async jobs (transcript, extract, crawl) depend on Supadata's service availability and rate limits

Setup

Before you start

Runtime:Node.js

SUPADATA_API_KEY requiredsecret API key for the Supadata service; sign up at https://supadata.ai to obtain one.
  1. Get a Supadata API key and set it as the SUPADATA_API_KEY environment variable.
  2. Follow Supadata's integration guide (docs.supadata.ai/integrations/mcp) to configure the MCP server for Claude, ChatGPT, Cursor, Windsurf, VS Code, or other clients.
  3. For local development, clone the repo and run npm install and npm run build.

Check that it works

Confirm that supadata_transcript and related tools appear in your client's tool list, then ask it to scrape a simple page (e.g. https://example.com); returning content proves the connection works.

Troubleshooting

  1. Confirm the SUPADATA_API_KEY environment variable is set correctly
  2. Transcript, extract, and crawl operations are asynchronous — poll the matching check_status tool with the returned job ID
  3. Verify the target URL belongs to a supported platform (YouTube, TikTok, Instagram, Twitter) or is an accessible file/web URL
  4. On failures, check network connectivity and the retry/rate-limit configuration (maxAttempts, initialDelay, maxDelay, backoffFactor)

Things to try

Once connected, you can ask your AI assistant things like:

  • Extract the English transcript from this YouTube video: https://youtube.com/watch?v=example
  • Get metadata for this video, including title and author
  • Use AI to extract the main topics discussed in this video
  • Scrape https://example.com and return the page content as markdown

Tools 9

supadata_transcript read-only
Extracts transcripts from supported video platforms (YouTube, TikTok, Instagram, Twitter) and file URLs, with an optional language parameter.
supadata_check_transcript_status read-only
Checks the progress of a transcript extraction job using its job ID.
supadata_metadata read-only
Fetches metadata for a media URL, including platform info, title, description, author details, engagement stats, media details, tags, and creation date.
supadata_extract read-only
Uses AI to extract structured data from a video URL; accepts a prompt and/or JSON Schema for the output format, returning a job ID for async processing.
supadata_check_extract_status read-only
Checks the progress of an AI extraction job using its job ID.
supadata_scrape read-only
Extracts content from a single URL with advanced options such as language.
supadata_map read-only
Discovers all indexed URLs on a website to find relevant pages before scraping.
supadata_crawl read-only
Starts an asynchronous crawl job that extracts content from multiple related pages on a site.
Show 1 more tools
supadata_check_crawl_status read-only
Checks the progress of a crawl job using its job ID.

Use cases

Bulk-extracting transcripts from YouTube, TikTok, Instagram, or Twitter videos
Scraping a single web page for summarization or knowledge-base ingestion
Crawling a blog or documentation site across multiple related pages
Discovering the full set of URLs on a site before scraping
Fetching video metadata (title, author, engagement stats, etc.)
Extracting AI-driven structured data from video content

Supported clients

Claude
ChatGPT
Cursor
Windsurf
VS Code

Listed from the project's documentation, not tested by this site.

Overview

Supadata MCP Server is the official Model Context Protocol server for Supadata, adding video transcript extraction, web scraping, crawling, and site discovery to Cursor, Claude, and other LLM clients. It extracts transcripts from YouTube, TikTok, Instagram, Twitter, and file URLs, retrieves media metadata (platform info, title, description, author details, engagement stats, tags, creation date), and uses AI to pull structured data from video content. On the web side it offers single-page scraping, multi-page crawling, and site-wide URL discovery. Built-in configurable retry and rate-limiting logic handles transient failures. Requires a SUPADATA_API_KEY environment variable to authenticate against Supadata's backend API.

Similar servers

Context7 80 · B

Upstash's official server providing up-to-date third-party library docs for AI coding assistants

★ 62.9k · Tools 2 Compare with this →

Source revision 40c868a3f820 Data synced 2026-10-11 Read the FMRS scoring method