AI Guides › Workbench

ScrapeGraphAI: The Free Local Route and the Paid Hosted Route

By Nigel Guy · 7 min read

The usual way people meet ScrapeGraphAI is a short clip promising a free tool that "scrapes anything". So they paste two install lines into a terminal, hit a Python version error, connect the hosted version to Claude Code instead, and only later notice they are spending credits. There are really two products with one name. One is a free, open-source Python library you run on your own machine. The other is a paid cloud API that your AI assistant calls for you. Both are useful, but they cost and fail in different ways.

The rule: decide which route you are on before you install anything. The local library is free but you have to run it yourself. The hosted MCP server needs no setup but spends credits. In both cases, check the output against the page.

The words you need first

Word What it means here
Scraping Reading a web page with a program and pulling out the parts you want
Python package A bundle of code you install with pip, Python's installer
Virtual environment A private folder of packages for one project, so installs don't clash
Playwright A tool that drives a real, invisible browser so pages that need JavaScript load properly
Ollama A free app that runs AI models on your own computer
MCP server A plug-in that gives an assistant such as Claude Code new tools, in this case scraping
Credits ScrapeGraphAI's cloud currency. Each cloud request uses some

The routes compared

Route What it does Cost at time of writing Best for The catch
Library + Ollama (local) You write a short Python file. A model on your machine reads the page and returns structured data Free. MIT-licensed library, free Ollama, free Llama 3.2 model (about 2 GB download) One-off lists, learning, pages you'd rather not send to a cloud service Needs Python 3.12 or newer. Small local models miss or invent fields more often
Library + a cloud model key Same script, but the reading is done by a paid model such as OpenAI's Library free. You pay that provider per use Messier pages where a small local model struggles Two bills to watch, and the page text goes to the model provider
Hosted MCP server (sgai) Claude Code, Cursor or Codex calls ScrapeGraphAI's cloud for you. No Python needed 500 one-off free credits, then paid plans billed in US dollars (Starter is $20 a month at time of writing). The £ figure depends on your card's exchange rate People who already work in an AI coding assistant A basic extraction costs 5 credits, and 10 with the anti-bot "stealth" option. Tool calls time out after 60 seconds

Route 1: the free local route, by hand

What you need: Python 3.12 or later (PyPI lists the latest release, 2.3.0, as needing at least 3.12), a terminal, a few gigabytes of free disk space, and ideally 8 GB of RAM or more. RAM needs vary by machine, so treat that as a guide, not a guarantee.

Step 1 — check Python. Run python3 --version (on Windows, python --version). If it shows 3.11 or older, install a current version from python.org before going any further.

Step 2 — make a project folder and a virtual environment. The package's own page recommends this.

mkdir scraper && cd scraper
python3 -m venv .venv
source .venv/bin/activate      # Windows: .venv\Scripts\activate

Step 3 — install the library and a browser.

pip install scrapegraphai
playwright install chromium

Plain playwright install downloads every browser Playwright supports. Adding chromium gets only the one you need.

Step 4 — install Ollama and fetch a model. Download Ollama from ollama.com, open it, then run:

ollama pull llama3.2

That fetches the default 3-billion-parameter version, which Ollama lists at about 2 GB. Keep the Ollama app running while you scrape.

Step 5 — write the script. Save the following as scrape.py. Fill in TARGET_URL, and replace the field list in WHAT_TO_EXTRACT with what you actually want.

import json
from scrapegraphai.graphs import SmartScraperGraph

TARGET_URL = "[PAGE_URL]"
WHAT_TO_EXTRACT = (
    "You are extracting data from one web page for a spreadsheet. "
    "Return a JSON object with a key 'items', holding a list. "
    "Each entry has these fields: [FIELD_1], [FIELD_2], [FIELD_3]. "
    "Copy values exactly as written on the page. "
    "If a field is missing for an entry, use null. Never guess or fill gaps. "
    "Skip navigation, adverts and footer links."
)

settings = {
    "llm": {
        "model": "ollama/llama3.2",
        "model_tokens": 8192,
        "format": "json",
    },
    "headless": True,
}

scraper = SmartScraperGraph(prompt=WHAT_TO_EXTRACT, source=TARGET_URL, config=settings)
result = scraper.run()

with open("result.json", "w", encoding="utf-8") as f:
    json.dump(result, f, indent=2, ensure_ascii=False)
print(json.dumps(result, indent=2, ensure_ascii=False))

Step 6 — run it. In the same activated terminal, run python scrape.py. The results print on screen and are saved to result.json.

Route 2: the hosted route in Claude Code, Cursor or Codex

This route skips Python entirely, but every request runs on ScrapeGraphAI's cloud and uses credits. Make an account at scrapegraphai.com first, then add the server for your tool:

Tool How to add it (from ScrapeGraphAI's docs)
Claude Code claude mcp add --transport http sgai https://mcp.scrapegraphai.com/mcp. Add --scope user to make it available in every project. Sign in through /mcp inside Claude Code, or add --header "Authorization: Bearer $SGAI_API_KEY" to use an API key instead
Cursor Add an sgai entry with "url": "https://mcp.scrapegraphai.com/mcp" to ~/.cursor/mcp.json, restart Cursor, and log in from its MCP settings
Codex Add [mcp_servers.sgai] with url = "https://mcp.scrapegraphai.com/mcp" to ~/.codex/config.toml, then run codex mcp login sgai

To check it worked, ask the assistant to run the credits tool. If it shows your balance, the connection works. Then give it a bounded job. Fill in the bracketed parts:

You have the ScrapeGraphAI (sgai) tools. I need structured data from one page.

Page: [PAGE_URL]
Fields I want for each item: [FIELD_LIST]
Credit limit for this task: [MAX_CREDITS]

Steps:
1. Check my credit balance first and tell me what this task should cost before you spend anything.
2. Use one extraction call on that single page only. Do not crawl, search, set up monitors or turn on stealth mode unless I say so.
3. Return a table with one row per item and one column per field. Use "not on page" where a value is missing. Never estimate one.
4. Below the table, list any fields that were ambiguous and how many items you found.

If the page URL or the field list is missing or unclear, ask me before calling any tool.
Before you answer, check that every value in the table came from the tool result and not from your own knowledge.

When it goes wrong

Symptom Likely cause Fix
pip refuses: requires Python >=3.12 Old Python Install 3.12 or newer, then recreate the virtual environment
Connection refused on port 11434 Ollama isn't running Open the Ollama app, then run the script again
Empty list from a page you can see Content loads late, sits behind a login, or the site blocks bots Try another page, or switch off headless and watch what loads. Don't try to get round a login
Plausible values that aren't on the page Small model filling gaps Ask for nulls, cut the field list down, or try a larger model
playwright not found Virtual environment not activated Run the activate line again from Step 2

How to choose

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo