Scrapling: Adaptive Web Scraping Framework for Python
GitHub Repo
BSD-3-Clause
October 2, 2026 at 09:19 AM
0 views

Scrapling: Adaptive Web Scraping Framework for Python

@D4VinciProject Author

What Scrapling is

Scrapling is a Python web scraping framework that, in the README's words, handles everything from a single request to a full-scale crawl. It bundles three things that are usually separate libraries: a fast HTML parser with CSS, XPath and BeautifulSoup-style selection; a set of fetchers for plain HTTP requests and full browser automation; and a spider framework modeled on Scrapy for concurrent, multi-session crawls. It is written and maintained by Karim Shoair and released under the BSD-3-Clause license.

Its headline idea is adaptive scraping: the parser can remember the elements you selected and, when a site's markup changes, relocate them using similarity matching instead of failing on a stale selector. That targets a familiar maintenance cost of scraping jobs, where a redesign silently breaks extraction. The project is aimed at developers doing data extraction for research, monitoring, RAG corpus building or analytics, as well as people who want to pull a page into Markdown from the terminal without writing code.

How it works

Scrapling is organized in layers. At the bottom is the parser (scrapling.parser.Selector), which can be used on its own against any HTML string. Above it are the fetchers in scrapling.fetchers, each returning a response object you query with the same selector API:

  • Fetcher / FetcherSession: HTTP requests that can mimic a browser's TLS fingerprint and headers, with HTTP/3 support.
  • DynamicFetcher / DynamicSession: full browser automation using Playwright's Chromium or Google Chrome for JavaScript-heavy sites.
  • StealthyFetcher / StealthySession: a headless browser mode with fingerprint spoofing. The README lists handling Cloudflare Turnstile and interstitial challenges among its capabilities.

On top sits scrapling.spiders, a Scrapy-like crawling framework with start_urls, async parse callbacks and Request/Response objects. A single spider can register several sessions (for example a fast HTTP session and a browser session) and route individual requests between them by session ID. Crawls checkpoint to a directory so they can be paused with Ctrl+C and resumed later.

Key features

  • Adaptive element tracking: save selections with auto_save=True and later pass adaptive=True to relocate elements after the page structure changes.
  • Flexible selection: CSS, XPath, BeautifulSoup-style find_all, text and regex search, plus find_similar() to locate elements like one you already found.
  • Concurrent crawling: configurable concurrency limits, per-domain throttling and download delays.
  • AutoThrottle: per-domain delays tuned from response times, doubled (or set from Retry-After) when a site starts rate-limiting, and relaxed once it stops.
  • Robots.txt compliance: an optional robots_txt_obey flag that respects Disallow, Crawl-delay and Request-rate directives.
  • Development mode: caches responses to disk on the first run and replays them, so you can iterate on parse() without re-hitting target servers.
  • Spider templates: CrawlSpider, SitemapSpider, XMLFeedSpider, CSVFeedSpider, ShopifySpider and SiteToMarkdownSpider.
  • Streaming and export: stream items with spider.stream(), or export to JSON, JSONL, CSV or XML.
  • Background API capture: capture_xhr collects matching XHR/fetch responses a page makes while loading.
  • Session and proxy management: persistent sessions for cookies and state, plus a built-in ProxyRotator.
  • Remote browsers: connect to an existing browser over CDP with cdp_url.
  • AI tooling: an MCP server for Claude, Cursor and other clients that narrows pages with CSS selectors and strips prompt-injection content, an Agent Skill for coding agents, and page.markdown() for LLM-ready Markdown.
  • Scrapy integration: decorate an existing Scrapy callback with scrapling_response to use Scrapling's parser without a rewrite.
  • CLI and shell: an IPython-based scraping shell and an extract command that saves a page as text, Markdown or HTML.

Getting started

Scrapling requires Python 3.10 or higher. The base package includes only the parser:

pip install scrapling

Importing from scrapling.fetchers or scrapling.spiders requires the fetcher dependencies and browsers:

pip install "scrapling[fetchers]"

scrapling install           # normal install
scrapling install  --force  # force reinstall

Extras exist for the MCP server (scrapling[ai]), RAG helpers (scrapling[rag]), the shell (scrapling[shell]) and everything (scrapling[all]). A Docker image with all extras and browsers is also published:

docker pull pyd4vinci/scrapling

A basic HTTP fetch against the scraping practice site used in the README:

from scrapling.fetchers import Fetcher, FetcherSession

with FetcherSession(impersonate='chrome') as session:  # Use latest version of Chrome's TLS fingerprint
    page = session.get('https://quotes.toscrape.com/', stealthy_headers=True)
    quotes = page.css('.quote .text::text').getall()

# Or use one-off requests
page = Fetcher.get('https://quotes.toscrape.com/')
quotes = page.css('.quote .text::text').getall()

And a spider that follows pagination and exports the results:

from scrapling.spiders import Spider, Request, Response

class QuotesSpider(Spider):
    name = "quotes"
    start_urls = ["https://quotes.toscrape.com/"]
    concurrent_requests = 10
    
    async def parse(self, response: Response):
        for quote in response.css('.quote'):
            yield {
                "text": quote.css('.text::text').get(),
                "author": quote.css('.author::text').get(),
            }
            
        next_page = response.css('.next a')
        if next_page:
            yield response.follow(next_page[0].attrib['href'])

result = QuotesSpider().start()
print(f"Scraped {len(result.items)} quotes")
result.items.to_json("quotes.json")

Passing crawldir="./crawl_data" to the spider enables checkpointed pause and resume.

Use cases

  • Price and catalog monitoring: track product data on retail sites over time, relying on adaptive selectors to survive layout changes; the ShopifySpider template reads a store's catalog through its JSON API.
  • RAG corpus building: crawl documentation or a knowledge site into Markdown with SiteToMarkdownSpider, with no LLM in the loop.
  • Research datasets: collect public data for academic or journalistic analysis, using AutoThrottle and robots.txt compliance to stay polite.
  • Agent web access: give an MCP-capable assistant a scraping tool that returns trimmed, sanitized page content.
  • Incremental Scrapy migration: keep an existing Scrapy project and swap in Scrapling's parser per callback.
  • Feeds and sitemaps: ingest RSS, XML or CSV feeds and sitemap-driven crawls with the bundled templates.

How it compares

The README positions Scrapling's API as familiar to Scrapy and BeautifulSoup users, reusing the pseudo-elements from Scrapy/Parsel, and it adapts Parsel code for its selector translator. It publishes parser benchmarks from its own benchmarks.py: on a 5,000-nested-element text extraction test it reports 1.99 ms for Scrapling versus 2.06 ms for Parsel/Scrapy and 2.56 ms for raw lxml, with BeautifulSoup variants far slower, and it claims its similarity search is about 5.5x faster than AutoScraper. These are project-reported numbers averaged over 100+ runs. Compared with Scrapy, the main differences are the built-in browser fetchers, adaptive selectors and multi-session routing in one package; Scrapy remains the more established ecosystem with its own middleware and extensions.

Things to know before adopting

  • Legal and ethical use: the README's disclaimer says the library is for educational and research purposes, that users must comply with local and international scraping and privacy laws, and to always respect websites' terms of service and robots.txt files. Enabling robots_txt_obey and AutoThrottle is a sensible default for any crawl.
  • Access controls: the stealth and anti-bot features exist, but a site's bot protection is often an expression of its terms. Check whether a site offers an official API or data export, and whether scraping it is permitted, before reaching for browser-based fetchers.
  • Install footprint: the base install is parser-only; fetchers pull in browsers and fingerprinting dependencies via scrapling install.
  • License: BSD-3-Clause, permissive for commercial use with attribution; the translator submodule is adapted from BSD-licensed Parsel.
  • Quality signals: the README reports 92% test coverage, full type hints checked with PyRight and MyPy, and a Docker image built on each release.
  • Sponsor listings: the top of the README carries sponsor entries for proxy and scraping service vendors; those are third-party commercial services, not part of the open-source project.

Project activity

As of October 2026 the repository has about 85,100 stars on GitHub. It was created on 2024-10-13, is written in Python, and is released under the BSD-3-Clause license. It is published on PyPI as scrapling. Source code is at github.com/D4Vinci/Scrapling and documentation is at scrapling.readthedocs.io.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
scrapling
Created
October 2
Last Updated
October 2, 2026 at 09:19 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.