Back to rankings

sqzw-x/mdcx

Python

Movie metadata scraper

crawlerembyjav-scraperjellyfinmetadatamovie-crawlermovie-metadatamovie-scrappermoviespythonscraper
Star Growth
Stars
3.7k
Forks
480
Weekly Growth
Issues
11
1k2k3k
Nov 2023Sep 2024Aug 2025Jul 2026
ArtifactsPyPIpip install mdcx
README

MDCx

python

上游项目

  • yoshiko2/Movie_Data_Capture: CLI 工具, 开源版本现已不活跃, 新版本已闭源商业化.
  • moyy996/AVDC: 上述项目早期的一个 Fork, 使用 PyQt 实现了图形界面, 已停止维护
  • @Hermit/MDCx: AVDC 的 Fork, 一度在 anyabc/something 分发源代码及可执行文件.
  • 2023-11-3 @anyabc 因未知原因销号删库, 其分发的最后一个版本号为 20231014.
  • 本项目基于 @Hermit/MDCx, 对代码进行了大幅的重构与拆分, 以提高可维护性

向相关开发者表示敬意.

构建

一般情况请勿自行构建, 至 Release 下载最新版

Windows 7

即将放弃对 Windows 7 的支持. #494

Windows 7 上需使用 Python 3.8 构建, 代码及依赖均兼容, 可在本地自行构建. 也可使用 GitHub Actions 构建:

  1. fork 本仓库, 在仓库设置中启用 Actions
  2. 参考 为存储库创建配置变量, 设置 BUILD_FOR_WINDOWS_LEGACY 变量, 值非空即可
  3. 在 Actions 中手动运行 Build and Release

macOS

低版本 macOS: 需注意 opencv 兼容性问题, 参考 issue #82. 也可使用 GitHub Actions 构建, 步骤同上, 需设置 BUILD_FOR_MACOS_LEGACY 变量, 值非空即可; 以及 MACOS_LEGACY_CV_VERSION 变量, 值为兼容的 opencv-contrib-python-headless 版本

授权许可

本插件项目在 GPLv3 许可授权下发行。此外,如果使用本项目表明还额外接受以下条款:

  • 本项目仅供学习以及技术交流使用
  • 请勿在公共社交平台上宣传此项目
  • 使用本软件时请遵守当地法律法规
  • 法律及使用后果由使用者自己承担
  • 禁止将本软件用于任何的商业用途
Related repositories
firecrawl/firecrawl

The API to search, scrape, and interact with the web at scale. 🔥

TypeScriptnpmGNU Affero General Public License v3.0aicrawler
firecrawl.dev
154.1k8.8k
D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

PythonPyPIBSD 3-Clause "New" or "Revised" Licensecrawlercrawling
scrapling.readthedocs.io/en/latest/
70.6k7k
scrapy/scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

PythonPyPIBSD 3-Clause "New" or "Revised" Licensepythonscraping
scrapy.org
63.3k11.8k
NaiboWang/EasySpider

A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。

JavaScriptnpmGNU Affero General Public License v3.0code-freecrawler
easyspider.net
44.3k5.4k
iawia002/lux

👾 Fast and simple video download library and CLI tool written in Go

GoGo ModulesMIT Licensedownloadergo
31.6k3.3k
mendableai/firecrawl

🔥 Turn entire websites into LLM-ready markdown or structured data. Scrape, crawl and extract with a single API.

TypeScriptnpmGNU Affero General Public License v3.0aicrawler
firecrawl.dev
29.5k2.5k
ScrapeGraphAI/Scrapegraph-ai

Python scraper based on AI

PythonPyPIMIT Licensescrapingscraping-python
scrapegraphai.com
28.5k2.8k
gocolly/colly

Elegant Scraper and Crawler Framework for Golang

GoGo ModulesApache License 2.0golangscraper
go-colly.org
25.4k1.9k
apify/crawlee

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

TypeScriptnpmApache License 2.0web-scrapingweb-crawling
crawlee.dev
24.8k1.6k
jhao104/proxy_pool

Python ProxyPool for web spider

PythonPyPIMIT Licensecrawlerproxy
jhao104.github.io/proxy_pool/
23.5k5.4k
Evil0ctal/Douyin_TikTok_Download_API

🚀「Douyin_TikTok_Download_API」是一个开箱即用的高性能异步抖音、快手、TikTok、Bilibili数据爬取工具,支持API调用,在线批量解析及下载。

PythonPyPIApache License 2.0pythonpywebio
douyin.wtf
18.9k2.7k
projectdiscovery/katana

A next-generation crawling and spidering framework.

GoGo ModulesMIT Licensecrawlerweb-spider
17.2k1.2k