Verified project record
ScrapeGraphAI/Scrapegraph-ai
A Python web scraping library that uses LLMs and direct graph logic to create scraping pipelines for websites and local documents via simple prompts. It supports various LLMs through APIs or local models via Ollama.
Project overview
The library uses graph-based logic to handle web scraping tasks using large language models, reducing the need to write and maintain custom parsers for individual sites or document formats.
- Project type
- RAG · AI Search · Data Processing
- Use cases
- Knowledge Q&A · Search & Research
- Deployment
- Refer to project documentation
- License
- MIT
Best for
- Audiences including developers, researchers, and data teams who need to create scraping pipelines using Python and are equipped to manage LLM configurations.
Key capabilities
- A single-page scraper that extracts information given a user prompt and an input source.
- A multi-page scraper that extracts information from the top n search results of a search engine.
- A single-page scraper that extracts information from a website and generates an audio file.
- A single-page scraper that extracts information from a website and generates a Python script.
- Extracts information from multiple pages given a single prompt and a list of sources using parallel LLM calls.
- Allows using different LLMs through APIs (OpenAI, Groq, Azure, Gemini, MiniMax) or local models via Ollama.
- Creates scraping pipelines for local documents (XML, HTML, JSON, Markdown, etc.).
- Collects anonymous usage metrics which can be disabled by setting an environment variable.
Limitations and risks
- Requires managing browser rendering, proxies, and scaling when self-hosted.
- The library collects anonymous usage metrics by default, which can be opted out of.
Getting started
- Setup involves a medium difficulty process starting with 'pip install scrapegraphai', followed by 'playwright install'. Users then configure LLM API keys or a local model runtime (Ollama) in the graph_config and run SmartScraperGraph. Playwright browser is required for fetching.
Alternatives and comparisons
- Provides shared retrieval infrastructure handling authentication, ingestion, syncing, indexing, and retrieval so developers do not have to rebuild pipelines.
- A full-stack RAG tutorial combining systematic theory with hands-on projects to build production-ready intelligent QA and knowledge retrieval systems.
- Backup and retrieve personal Telegram messages with local tokenization and vector semantic search, alongside AI-powered features.
Project comparisons
Evidence and sources
- README: [ScrapeGraphAI](https://scrapegraphai.com) is a *web scraping* python library that uses LLM and direct graph logic to create scraping pipelines for websites and local documents (X…
- README: The reference page for Scrapegraph-ai is available on the official page of PyPI: [pypi](https://pypi.org/project/scrapegraphai/).
- README: It is possible to use different LLM through APIs, such as **OpenAI**, **Groq**, **Azure**, **Gemini**, **MiniMax** and more, or local models using **Ollama**.
- README: For OpenAI and other models you just need to change the llm config! > ```python >graph_config = { > "llm": { > "api_key": "YOUR_OPENAI_API_KEY", > "model": "openai/gpt-4o-mini", >…
- README: graph_config = { "llm": { "model": "ollama/llama3.2", "model_tokens": 8192, "format": "json", }, "verbose": True, "headless": False, }
AI Search
Find projects, verify facts, compare options, or turn a complex need into an actionable plan
Try a searchA click only fills the search box; you stay in control
Project Details
0