How to Run a Fast Website Parser for LLMs on 14 Megabytes of RAM
A lightweight Rust-based web scraper that uses just 14MB of RAM and outputs clean Markdown for LLMs.
Language
HomeLanguages
Sections
A lightweight Rust-based web scraper that uses just 14MB of RAM and outputs clean Markdown for LLMs.
A Chrome extension that detects anti-bot systems, CAPTCHAs, and 21 fingerprinting techniques in real time, helping developers and security researchers identify what protection is running on any website.
Krawl is a deception server that pretends to be a vulnerable web app, feeding scanners infinite trap pages and wasting their resources.
When Scrapy hits its limits and Python crawls, Crawly on the BEAM VM offers a powerful alternative for large-scale web scraping with familiar concepts and built-in management tools.
Spider- flow turns web scraping into visual programming. With over 11k GitHub stars, this Java-based platform lets you build parsers by drawing flowcharts in your browser.