Scrapling: Adaptive Web Scraping Framework for Python
What Scrapling is
Scrapling is a Python web scraping framework that, in the README's words, handles everything from a single request to a full-scale crawl. It bundles three things that are usually separate libraries: a fast HTML parser with CSS, XPath and BeautifulSoup-style selection; a set of fetchers for plain HTTP requests and full browser automation; and a spider framework modeled on Scrapy for concurrent, multi-session crawls. It is written and maintained by Karim Shoair and released under the BSD-3-Clause license.
Its headline idea is adaptive scraping: the parser can remember the elements you selected and, when a site's markup changes, relocate them using similarity matching instead of failing on a stale selector. That targets a familiar maintenance cost of scraping jobs, where a redesign silently breaks extraction. The project is aimed at developers doing data extraction for research, monitoring, RAG corpus building or analytics, as well as people who want to pull a page into Markdown from the terminal without writing code.
How it works
Scrapling is organized in layers. At the bottom is the parser (scrapling.parser.Selector), which can be used on its own against any HTML string. Above it are the fetchers in scrapling.fetchers, each returning a response object you query with the same selector API:
Fetcher/FetcherSession: HTTP requests that can mimic a browser's TLS fingerprint and headers, with HTTP/3 support.DynamicFetcher/DynamicSession: full browser automation using Playwright's Chromium or Google Chrome for JavaScript-heavy sites.StealthyFetcher/StealthySession: a headless browser mode with fingerprint spoofing. The README lists handling Cloudflare Turnstile and interstitial challenges among its capabilities.
On top sits scrapling.spiders, a Scrapy-like crawling framework with start_urls, async parse callbacks and Request/Response objects. A single spider can register several sessions (for example a fast HTTP session and a browser session) and route individual requests between them by session ID. Crawls checkpoint to a directory so they can be paused with Ctrl+C and resumed later.
Key features
- Adaptive element tracking: save selections with
auto_save=Trueand later passadaptive=Trueto relocate elements after the page structure changes. - Flexible selection: CSS, XPath, BeautifulSoup-style
find_all, text and regex search, plusfind_similar()to locate elements like one you already found. - Concurrent crawling: configurable concurrency limits, per-domain throttling and download delays.
- AutoThrottle: per-domain delays tuned from response times, doubled (or set from
Retry-After) when a site starts rate-limiting, and relaxed once it stops. - Robots.txt compliance: an optional
robots_txt_obeyflag that respectsDisallow,Crawl-delayandRequest-ratedirectives. - Development mode: caches responses to disk on the first run and replays them, so you can iterate on
parse()without re-hitting target servers. - Spider templates:
CrawlSpider,SitemapSpider,XMLFeedSpider,CSVFeedSpider,ShopifySpiderandSiteToMarkdownSpider. - Streaming and export: stream items with
spider.stream(), or export to JSON, JSONL, CSV or XML. - Background API capture:
capture_xhrcollects matching XHR/fetch responses a page makes while loading. - Session and proxy management: persistent sessions for cookies and state, plus a built-in
ProxyRotator. - Remote browsers: connect to an existing browser over CDP with
cdp_url. - AI tooling: an MCP server for Claude, Cursor and other clients that narrows pages with CSS selectors and strips prompt-injection content, an Agent Skill for coding agents, and
page.markdown()for LLM-ready Markdown. - Scrapy integration: decorate an existing Scrapy callback with
scrapling_responseto use Scrapling's parser without a rewrite. - CLI and shell: an IPython-based scraping shell and an
extractcommand that saves a page as text, Markdown or HTML.
Getting started
Scrapling requires Python 3.10 or higher. The base package includes only the parser:
pip install scraplingImporting from scrapling.fetchers or scrapling.spiders requires the fetcher dependencies and browsers:
pip install "scrapling[fetchers]"
scrapling install # normal install
scrapling install --force # force reinstallExtras exist for the MCP server (scrapling[ai]), RAG helpers (scrapling[rag]), the shell (scrapling[shell]) and everything (scrapling[all]). A Docker image with all extras and browsers is also published:
docker pull pyd4vinci/scraplingA basic HTTP fetch against the scraping practice site used in the README:
from scrapling.fetchers import Fetcher, FetcherSession
with FetcherSession(impersonate='chrome') as session: # Use latest version of Chrome's TLS fingerprint
page = session.get('https://quotes.toscrape.com/', stealthy_headers=True)
quotes = page.css('.quote .text::text').getall()
# Or use one-off requests
page = Fetcher.get('https://quotes.toscrape.com/')
quotes = page.css('.quote .text::text').getall()And a spider that follows pagination and exports the results:
from scrapling.spiders import Spider, Request, Response
class QuotesSpider(Spider):
name = "quotes"
start_urls = ["https://quotes.toscrape.com/"]
concurrent_requests = 10
async def parse(self, response: Response):
for quote in response.css('.quote'):
yield {
"text": quote.css('.text::text').get(),
"author": quote.css('.author::text').get(),
}
next_page = response.css('.next a')
if next_page:
yield response.follow(next_page[0].attrib['href'])
result = QuotesSpider().start()
print(f"Scraped {len(result.items)} quotes")
result.items.to_json("quotes.json")Passing crawldir="./crawl_data" to the spider enables checkpointed pause and resume.
Use cases
- Price and catalog monitoring: track product data on retail sites over time, relying on adaptive selectors to survive layout changes; the
ShopifySpidertemplate reads a store's catalog through its JSON API. - RAG corpus building: crawl documentation or a knowledge site into Markdown with
SiteToMarkdownSpider, with no LLM in the loop. - Research datasets: collect public data for academic or journalistic analysis, using AutoThrottle and robots.txt compliance to stay polite.
- Agent web access: give an MCP-capable assistant a scraping tool that returns trimmed, sanitized page content.
- Incremental Scrapy migration: keep an existing Scrapy project and swap in Scrapling's parser per callback.
- Feeds and sitemaps: ingest RSS, XML or CSV feeds and sitemap-driven crawls with the bundled templates.
How it compares
The README positions Scrapling's API as familiar to Scrapy and BeautifulSoup users, reusing the pseudo-elements from Scrapy/Parsel, and it adapts Parsel code for its selector translator. It publishes parser benchmarks from its own benchmarks.py: on a 5,000-nested-element text extraction test it reports 1.99 ms for Scrapling versus 2.06 ms for Parsel/Scrapy and 2.56 ms for raw lxml, with BeautifulSoup variants far slower, and it claims its similarity search is about 5.5x faster than AutoScraper. These are project-reported numbers averaged over 100+ runs. Compared with Scrapy, the main differences are the built-in browser fetchers, adaptive selectors and multi-session routing in one package; Scrapy remains the more established ecosystem with its own middleware and extensions.
Things to know before adopting
- Legal and ethical use: the README's disclaimer says the library is for educational and research purposes, that users must comply with local and international scraping and privacy laws, and to always respect websites' terms of service and robots.txt files. Enabling
robots_txt_obeyand AutoThrottle is a sensible default for any crawl. - Access controls: the stealth and anti-bot features exist, but a site's bot protection is often an expression of its terms. Check whether a site offers an official API or data export, and whether scraping it is permitted, before reaching for browser-based fetchers.
- Install footprint: the base install is parser-only; fetchers pull in browsers and fingerprinting dependencies via
scrapling install. - License: BSD-3-Clause, permissive for commercial use with attribution; the translator submodule is adapted from BSD-licensed Parsel.
- Quality signals: the README reports 92% test coverage, full type hints checked with PyRight and MyPy, and a Docker image built on each release.
- Sponsor listings: the top of the README carries sponsor entries for proxy and scraping service vendors; those are third-party commercial services, not part of the open-source project.
Project activity
As of October 2026 the repository has about 85,100 stars on GitHub. It was created on 2024-10-13, is written in Python, and is released under the BSD-3-Clause license. It is published on PyPI as scrapling. Source code is at github.com/D4Vinci/Scrapling and documentation is at scrapling.readthedocs.io.
Enjoying this project?
Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.
Repository:https://github.com/D4Vinci/Scrapling
GitHub - D4Vinci/Scrapling: Scrapling: Adaptive Web Scraping Framework for Python
Scrapling is a BSD-3-Clause Python web scraping framework with an adaptive parser, HTTP and browser fetchers, and a Scrapy-like spider API, for developers build...
github - d4vinci/scrapling

