Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
-
Updated
Aug 7, 2026 - Rust
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Use LLMs to robustly extract web data
Fully automated and hands-free, accurately extracting and understanding web content — powered by machine learning agents.
Windows desktop app for collecting, reviewing, ranking, and exporting Google Scholar search results.
Low-Cost Cross-Domain Web Structured Information Extraction using specialized LoRA adapters.
Replayable Browser Agent
基于Scala Akka的分布式主题网络爬虫
A swipeable shortlist of live San Francisco apartment listings, powered by Context.dev.
Self-hosted web scraping and Markdown extraction for AI agents
Automatic extraction of the information on local event from a webpage with Machine Learning
Hermes plugin to improve web search with SearXNG for higher-quality search results and Crawl4AI for LLM-optimized webpage extraction.
Free local web search/extraction router for AI agents. Go CLI + MCP, BYOK/free-first routing, keyless DDGS/Scrapling fallback, setup writers and client guides.
A powerful and lightweight web scraping library with LLM extraction capabilities. This library combines web scraping with AI-powered content extraction using either OpenAI or OpenRouter APIs.
Predicting product recommendation score using the data available on the website of the client
Fast local Tavily-compatible web search and extraction backend for Hermes
Programming assignments for Web Information Extraction and Retrieval, FRI UL, 2021. PA1: standalone webcrawler of .gov.si web sites, PA2: approaches of the structured web data extraction, PA3: Data processing and indexing and Data retrieval.
Glasses Web Reader for Even Realities G2 — three-layer browser (sources → articles → reader) using r.jina.ai for clean URL extraction.
Validate public URLs and turn them into reviewable AI-ready Markdown or JSON.
Standalone Crawl4AI web extraction and bounded crawling plugin for Hermes Agent
Add a description, image, and links to the web-extraction topic page so that developers can more easily learn about it.
To associate your repository with the web-extraction topic, visit your repo's landing page and select "manage topics."