AI Guides › Workbench

Three Open-Source Scrapers: ScrapeGraphAI, Scrapling and Agent Reach

By Nigel Guy · 8 min read

Most people who want data off a website do one of two things: copy and paste by hand until they give up, or install whichever scraper a video recommended and then discover it solves a different problem from theirs. The three tools here are free and open source, but they are not interchangeable. One reads a page the way you would and hands back the fields you named, one survives a site redesign, and one routes your AI agent to platforms that normally want a login. Pick by the job, not by the star count.

The rule: decide what kind of page you are scraping and how often before you install anything, and give every tool its own virtual environment.

Five words you need first

Word What it means here
Terminal The text window where you type commands (Terminal on a Mac, PowerShell on Windows).
pip Python's installer. pip install x downloads package x.
Virtual environment A private folder of packages for one project, so tools do not break each other.
Selector A short address for part of a page, such as .product for every element with the class "product".
MCP server A small program that lets Claude Code, Cursor or Codex call a tool directly from chat.

The kit at a glance

Tool What it does Cost at time of writing Best for The catch
ScrapeGraphAI You describe the data in plain English; a language model reads the page and returns JSON Library free (MIT). The model costs whatever your provider charges, or nothing if you run one locally with Ollama. An optional hosted API has a free tier of 500 one-off credits, then paid plans from about £15 a month (billed in US dollars) One-off pulls from messy pages where writing selectors would take longer than the job Needs Python 3.12 or newer. Every page goes through a model, so it is slower and costs tokens; small local models make mistakes
Scrapling A fast scraping library whose "adaptive" mode re-finds elements after a site changes its layout Free (BSD-3-Clause) Pages you scrape repeatedly, where breakage is the real cost You still write selectors. Adaptive mode needs a first successful run to learn from
Agent Reach A command-line router that installs and health-checks upstream tools so an AI agent can read YouTube, RSS, GitHub, Reddit, X and others Free (MIT), no API fees Giving Claude Code or Codex reach into platforms beyond plain web pages Login-gated platforms use your browser cookies, which risks the account. Much of the documentation is in Chinese

What you need

Make one folder per tool and a virtual environment inside it:

mkdir scrape-test && cd scrape-test
python3 -m venv .venv
source .venv/bin/activate      # Windows: .venv\Scripts\Activate.ps1

Fill in nothing; run as is. Your prompt line should now start with (.venv).

ScrapeGraphAI: describe the data, get JSON

Install it and the browser it drives:

pip install scrapegraphai
playwright install

If you are using Ollama, download a model first with ollama pull llama3.2. Then save this as extract.py, replacing the bracketed parts:

import json
from scrapegraphai.graphs import SmartScraperGraph

config = {
    "llm": {"model": "ollama/llama3.2", "model_tokens": 8192, "format": "json"},
    "headless": True,
}

job = SmartScraperGraph(
    prompt="[WHAT_TO_EXTRACT, e.g. every event name, date and ticket price]",
    source="[PAGE_URL]",
    config=config,
)
print(json.dumps(job.run(), indent=2))

Run python extract.py. To use a hosted model instead, the project's README shows "openai/gpt-4o-mini" with an api_key entry; treat that model name as an example and check your provider's current list.

The prompt field is where most people lose accuracy. Write it like a brief, not a wish:

Extract every [ITEM_TYPE] listed on this page.
Return a JSON list. Each entry has exactly these keys: [KEY_1], [KEY_2], [KEY_3].
Use null when a value is not shown on the page. Never infer or invent a value.
Ignore navigation, adverts, cookie banners and "related" sections.
If the page holds no [ITEM_TYPE], return an empty list.

Fill in the item type (say, "job vacancy") and the exact field names you want back.

Scrapling: selectors that survive a redesign

Install it with its fetchers, then let it download its browsers:

pip install "scrapling[fetchers]"
scrapling install

Save this as watch.py:

from scrapling.fetchers import StealthyFetcher

StealthyFetcher.adaptive = True
page = StealthyFetcher.fetch("[PAGE_URL]", headless=True, network_idle=True)

# First run: auto_save=True remembers what these elements look like.
cards = page.css("[CSS_SELECTOR]", auto_save=True)
# After the site changes, swap the line above for:
# cards = page.css("[CSS_SELECTOR]", adaptive=True)
for card in cards:
    print(card.text)

Replace the URL and the selector (right-click an item in your browser, choose Inspect, and read its class name). The mechanism is the two flags: auto_save=True stores a fingerprint of each matched element on the first run; adaptive=True uses that fingerprint to relocate them when the old selector stops matching.

Agent Reach: wider reach for your agent

The trap here is the name. pip install agent-reach on PyPI installs a different project by a different author. Install the one in this guide from its GitHub archive, ideally with pipx, which keeps command-line tools isolated:

pipx install https://github.com/Panniantong/agent-reach/archive/main.zip
agent-reach install --env=auto
agent-reach doctor

The second command is a read-only check by default; it only changes your system if you rerun it with --system, which you should do only after reading what it plans to install. doctor lists each platform and whether it is ready. Web pages, YouTube, RSS and public GitHub work without a login; Reddit, Facebook, Instagram and XiaoHongShu need one, and X is only partly available without it.

Running them from Claude Code, Cursor or Codex

Scrapling ships an MCP server. Install pip install "scrapling[ai]" in its environment, find the full path with which scrapling-mcp (Windows: where scrapling-mcp), then:

ScrapeGraphAI and Agent Reach have no MCP step in this guide; ask the agent to run your script or the upstream command instead. A prompt that keeps the agent honest:

You are helping me collect public data with [TOOL_NAME], already installed in [FOLDER_PATH].
Goal: get [FIELDS] for [ITEM_TYPE] from [PAGE_URL_OR_PLATFORM], saved as [OUTPUT_FILE].
Steps:
1. Check robots.txt for that site and tell me if this path is disallowed. Stop if it is.
2. Test on a single page first and show me the first three records.
3. Wait for my "go" before running across more pages, and keep to one request every few seconds.
Rules: do not log in, use cookies, or install anything system-wide without asking me.
If a field is missing on the page, leave it empty; never fill gaps with guesses.
If the URL, fields or output format are unclear, ask me before starting.
Before replying, check that every record you show came from the page you fetched.

Fill in the tool, folder, fields, source and output file name.

How you know it worked

The usual snags

Symptom Likely cause
pip: command not found Try pip3, or python3 -m pip
ScrapeGraphAI refuses to install Python older than 3.12
Browser errors on first run You skipped playwright install or scrapling install
Plausible-looking but wrong JSON Small local model; try a larger one and tighten the prompt
A tool vanished after closing the terminal Virtual environment not activated; rerun the activate line

How to choose

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo