scrapfly/scrapfly-scrapers
quality grade B, 69 out of 100Scalable Python web scraping scripts for +40 popular domains
- stars
- 1.0k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
Scalable Python web scraping scripts for +40 popular domains
Write web scrapers in Ruby using a clean, AI-assisted DSL. Kimurai uses AI to figure out where the data lives, then caches the selectors and scrapes with pure Ruby. Get the intelligence of an LLM without the per-request latency or token costs.
Faster requests on Python 3
Go HTTP client with browser-identical TLS/HTTP2 fingerprinting. Bypass bot detection by perfectly mimicking Chrome, Firefox, and Safari at the cryptographic level (JA3/JA4, Akamai fingerprint, header order). Supports HTTP/1.1, HTTP/2, HTTP/3, sessions, cookies, and proxies.
A JavaScript library for generating random user agents with data that's updated daily.
HTTP(S)/SOCKS5 rotating residential proxies - code examples & general information.
Use Web Scraper API to extract data from Google Finance, including stock titles, pricing, and price changes in percentages.
AI Map is an AI-powered website mapping tool by Oxylabs AI Studio that uses natural language prompts to intelligently discover and extract relevant URLs from any website.
AgentQL is a suite of tools for connecting your AI to the web. Featuring a query language and Playwright integrations for interacting with elements and extracting data quickly, precisely, and at scale. Includes REST API, Python and JavaScript SDKs, browser debugger.
The complete web scraping toolkit for PHP.
In this tutorial, we showcase how to scrape public Google data with Python and Oxylabs API.
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
Simple web scraping for R
Finviz analysis python library.
Learn how to build your own Google Jobs scraper that simultaneously scrapes Google Jobs for multiple search queries and geo-locations with Python and Oxylabs’ Google Jobs Scraper API. https://oxylabs.io/blog/how-to-scrape-google-jobs
Scrape tweets, profiles, followers and following from Twitter/X, no API key needed. Python library with smart multi-account pooling, proxy support and async.
A guide for extracting titles, authors, and citations from Google Scholar using Python and Oxylabs SERP Scraper API.
A code for extracting best-selling items, search results, and currently available deals from Amazon using Python and Oxylabs E-Commerce Scraper API.
Web Scraping Framework
Learn step-by-step how to scrape Google Trends data and make a result comparison using Python and Oxylabs SERP API. Extract keywords, their popularity, breakdown by region, related queries, and more.
The process of extracting product data from Amazon using Python, including titles, ratings, prices, images, and descriptions.
Structured data gathering from any website using AI-powered scraper, crawler, and browser automation. Scraping and crawling with natural language prompts. Equip your LLM agents with fresh data. AI Studio python SDK for intelligent web data gathering.
Free Trial Amazon Scraper API for extracting search, product, offer listing, reviews, question and answers, best sellers and sellers data.
A Smart, Automatic, Fast and Lightweight Web Scraper for Python
24,523 repositories in the index in total.