Cover image
Try Now
2025-04-06

Servidor MCP que extrae contenido web usando Readability.js

3 years

Works with Finder

1

Github Watches

0

Github Forks

0

Github Stars

MCP Web Extractor

A Model Context Protocol (MCP) server that extracts web content using Readability.js. This tool fetches web pages and extracts the main content, making it ideal for saving clean, readable versions of articles to Obsidian notes.

Features

  • Extracts readable content from any URL
  • Removes ads, sidebars, and other distractions
  • Returns clean text along with metadata (title, excerpt, etc.)
  • Easy integration with Obsidian via MCP

Installation

# Clone the repository
git clone https://github.com/iemong/mcp-web-extractor.git
cd mcp-web-extractor

# Install dependencies
npm install

# Build the project
npm run build

# Start the server
npm start

The server will start on http://localhost:3000 with the MCP endpoint at http://localhost:3000/mcp.

Usage

As a standalone service

You can use the included client example to extract content from a URL:

ts-node-esm client-example.ts

With Obsidian

The obsidian-integration.ts file provides an example of how to integrate this MCP server with Obsidian. You can use it as a starting point for creating an Obsidian plugin that extracts web content.

API

The MCP server provides the following capability:

  • extract-content: Extracts readable content from a given URL
    • Parameters: { url: string }
    • Returns: { title, content, textContent, excerpt, siteName }

License

MIT

相关推荐

  • Joshua Armstrong
  • Confidential guide on numerology and astrology, based of GG33 Public information

  • https://suefel.com
  • Latest advice and best practices for custom GPT development.

  • Emmet Halm
  • Converts Figma frames into front-end code for various mobile frameworks.

  • Elijah Ng Shi Yi
  • Advanced software engineer GPT that excels through nailing the basics.

  • https://maiplestudio.com
  • Find Exhibitors, Speakers and more

  • lumpenspace
  • Take an adjectivised noun, and create images making it progressively more adjective!

  • https://appia.in
  • Siri Shortcut Finder – your go-to place for discovering amazing Siri Shortcuts with ease

  • Carlos Ferrin
  • Encuentra películas y series en plataformas de streaming.

  • Yusuf Emre Yeşilyurt
  • I find academic articles and books for research and literature reviews.

  • tomoyoshi hirata
  • Sony α7IIIマニュアルアシスタント

  • apappascs
  • Descubra la colección más completa y actualizada de servidores MCP en el mercado. Este repositorio sirve como un centro centralizado, que ofrece un extenso catálogo de servidores MCP de código abierto y propietarios, completos con características, enlaces de documentación y colaboradores.

  • ShrimpingIt
  • Manipulación basada en Micrypthon I2C del expansor GPIO de la serie MCP, derivada de AdaFruit_MCP230xx

  • jae-jae
  • Servidor MCP para obtener contenido de la página web con el navegador sin cabeza de dramaturgo.

  • ravitemer
  • Un poderoso complemento Neovim para administrar servidores MCP (protocolo de contexto del modelo)

  • patruff
  • Puente entre los servidores Ollama y MCP, lo que permite a LLM locales utilizar herramientas de protocolo de contexto del modelo

  • pontusab
  • La comunidad de cursor y windsurf, encontrar reglas y MCP

  • JackKuo666
  • 🔍 Habilitar asistentes de IA para buscar y acceder a la información del paquete PYPI a través de una interfaz MCP simple.

  • av
  • Ejecute sin esfuerzo LLM Backends, API, frontends y servicios con un solo comando.

  • WangRongsheng
  • 🧑‍🚀 全世界最好的 llM 资料总结(数据处理、模型训练、模型部署、 O1 模型、 MCP 、小语言模型、视觉语言模型) | Resumen de los mejores recursos del mundo.

    Reviews

    3 (1)
    Avatar
    user_ln2muBCY
    2025-04-17

    The mcp-web-extractor by iemong is truly a game-changer. It offers a seamless experience for extracting data from any web page. The comprehensive documentation on its GitHub page is incredibly helpful and the tool's efficiency is unmatched. Highly recommend for anyone in need of a robust web extraction solution!