Bookmarks tagged with #scraping.
Show all
Show all
tamnd/kage: Shadow any website for offline viewing, with the JavaScript stripped out
Shadow any website for offline viewing, with the JavaScript stripped out - tamnd/kage
Saved
on: 2026-06-15
lwthiker/curl-impersonate
curl-impersonate: A special build of curl that can impersonate Chrome & Firefox - lwthiker/curl-impersonate
Saved
on: 2025-01-18
browserbase/stagehand: An AI web browsing framework focused on simplicity and extensibility.
The AI Browser Automation Framework
Saved
on: 2025-01-09
steel-dev/steel-browser
🔥 Open Source Browser API for AI Agents & Apps. Steel Browser is a batteries-included browser instance that lets you automate the web without worrying about infrastructure. - steel-dev/steel-br...
Saved
on: 2024-12-02
SiteOne Crawler
A very useful and free website analyzer you'll ♥ as a Dev/DevOps, QA engineer, SEO or Security specialist, website owner or consultant. It performs in-depth analyzes of your website, generates an offline or markdown version of the website, provides a detailed HTML audit report and works on all popular platforms - Windows, macOS and Linux (x64 and arm64 too).
Saved
on: 2024-10-04
Crawlee · Build reliable crawlers. Fast. | Crawlee
Crawlee helps you build and maintain your crawlers. It's open source, but built by developers who scrape millions of pages every day for a living.
Saved
on: 2022-08-23
Reversing private APIs, Safeway, and not-so-extreme couponing
Saved
on: 2019-10-15
Headless Chrome support in Cloud Functions and App Engine | Hacker News
Saved
on: 2018-08-20
emadehsan/thal: Getting started with Puppeteer and Chrome Headless for Web
Getting started with Puppeteer and Chrome Headless for Web Scraping - emadehsan/thal
Saved
on: 2017-08-29
Advanced Web Scraping: Bypassing "403 Forbidden," captchas, and more | sang
Saved
on: 2017-03-16
cantino/huginn
Create agents that monitor and act on your behalf. Your agents are standing by! - huginn/huginn
Saved
on: 2016-02-02
Embed/README.md at master · php-embed/Embed
Get info from any web service or page
Saved
on: 2016-01-28
Web Scraping With Node.js — Smashing Magazine
As the volume of data on the web has increased, web scraping has become increasingly widespread, and a number of powerful services have emerged to simplify it. You can use Node.js to create a powerful web scraper that is both extremely versatile and completely free. A basic understanding of Node.js is recommended for this article; so, if you haven’t already, check it out before continuing. Also, web scraping may violate the terms of service for some websites, so just make sure you’re in the clear there before doing any heavy scraping.
Saved
on: 2015-11-21
How To Scrape a Website Using Node.js and Puppeteer
In this tutorial, you will build a web scraping application using Node.js and Puppeteer. Your app will grow in complexity as you progress. First, you will code your app to open Chromium and load a special website designed as a web-scraping sandbox: books.toscrape.com . In the next two steps, you will scrape all the books on a single page of books.toscrape and then all the books across multiple pages. Then you will filter your scraping by category and save your data as JSON.
Saved
on: 2015-11-21
DiDOM/README.md at master · Imangazaliev/DiDOM
Simple and fast HTML and XML parser
Saved
on: 2015-10-24
awesome-web-scraping/php.md at master · lorien/awesome-web-scraping
List of libraries, tools and APIs for web scraping and data processing. - lorien/awesome-web-scraping
Saved
on: 2015-08-18
search-script-scrape/README.md at master · stanfordjournalism/search-script-scrape
101 real world web scraping exercises in Python 3 for data journalists - stanfordjournalism/search-script-scrape
Saved
on: 2015-08-17
Welcome to dryscrape’s documentation! — dryscrape 1.0.1 documentation
Saved
on: 2015-06-03
PHP Based Scraper: How much should I pay someone/freelancer : PHP
Saved
on: 2014-10-12
ParseHub | Free web scraping - The most powerful web scraper
Saved
on: 2014-09-24
scraperjs/README.md at master · ruipgil/scraperjs
A complete and versatile web scraper
Saved
on: 2014-08-19
http://jakeaustwick.me/python-web-scraping-resource/?mc_list=python
Saved
on: 2014-08-05
https://github.com/hickford/MechanicalSoup/blob/master/README.md
Saved
on: 2014-07-10
mingcheng/php-readability: Back the fun of reading - PHP Port for Arc90′s Readability
Back the fun of reading - PHP Port for Arc90′s Readability - mingcheng/php-readability
Saved
on: 2014-06-26