magnASCII.dev Simone Magnaschi
Senior Full Stack Web Dev
Bookmarks tagged with #scraping.
Show all

tamnd/kage: Shadow any website for offline viewing, with the JavaScript stripped out

Shadow any website for offline viewing, with the JavaScript stripped out - tamnd/kage
Saved on: 2026-06-15

lwthiker/curl-impersonate

curl-impersonate: A special build of curl that can impersonate Chrome & Firefox - lwthiker/curl-impersonate
Saved on: 2025-01-18

steel-dev/steel-browser

🔥 Open Source Browser API for AI Agents & Apps. Steel Browser is a batteries-included browser instance that lets you automate the web without worrying about infrastructure. - steel-dev/steel-br...
Saved on: 2024-12-02

SiteOne Crawler

A very useful and free website analyzer you'll ♥ as a Dev/DevOps, QA engineer, SEO or Security specialist, website owner or consultant. It performs in-depth analyzes of your website, generates an offline or markdown version of the website, provides a detailed HTML audit report and works on all popular platforms - Windows, macOS and Linux (x64 and arm64 too).
Saved on: 2024-10-04

Crawlee · Build reliable crawlers. Fast. | Crawlee

Crawlee helps you build and maintain your crawlers. It's open source, but built by developers who scrape millions of pages every day for a living.
Saved on: 2022-08-23

emadehsan/thal: Getting started with Puppeteer and Chrome Headless for Web

Getting started with Puppeteer and Chrome Headless for Web Scraping - emadehsan/thal
Saved on: 2017-08-29

Francis Kim

Software Engineer
Saved on: 2016-08-23

cantino/huginn

Create agents that monitor and act on your behalf. Your agents are standing by! - huginn/huginn
Saved on: 2016-02-02

Embed/README.md at master · php-embed/Embed

Get info from any web service or page
Saved on: 2016-01-28

Web Scraping With Node.js — Smashing Magazine

As the volume of data on the web has increased, web scraping has become increasingly widespread, and a number of powerful services have emerged to simplify it. You can use Node.js to create a powerful web scraper that is both extremely versatile and completely free. A basic understanding of Node.js is recommended for this article; so, if you haven’t already, check it out before continuing. Also, web scraping may violate the terms of service for some websites, so just make sure you’re in the clear there before doing any heavy scraping.
Saved on: 2015-11-21

How To Scrape a Website Using Node.js and Puppeteer

In this tutorial, you will build a web scraping application using Node.js and Puppeteer. Your app will grow in complexity as you progress. First, you will code your app to open Chromium and load a special website designed as a web-scraping sandbox: books.toscrape.com . In the next two steps, you will scrape all the books on a single page of books.toscrape and then all the books across multiple pages. Then you will filter your scraping by category and save your data as JSON.
Saved on: 2015-11-21

DiDOM/README.md at master · Imangazaliev/DiDOM

Simple and fast HTML and XML parser
Saved on: 2015-10-24

awesome-web-scraping/php.md at master · lorien/awesome-web-scraping

List of libraries, tools and APIs for web scraping and data processing. - lorien/awesome-web-scraping
Saved on: 2015-08-18

search-script-scrape/README.md at master · stanfordjournalism/search-script-scrape

101 real world web scraping exercises in Python 3 for data journalists - stanfordjournalism/search-script-scrape
Saved on: 2015-08-17

scrapinghub/portia

Visual scraping for Scrapy
Saved on: 2014-11-17

binux/pyspider

A Powerful Spider(Web Crawler) System in Python
Saved on: 2014-11-17

scraperjs/README.md at master · ruipgil/scraperjs

A complete and versatile web scraper
Saved on: 2014-08-19

mingcheng/php-readability: Back the fun of reading - PHP Port for Arc90′s Readability

Back the fun of reading - PHP Port for Arc90′s Readability - mingcheng/php-readability
Saved on: 2014-06-26
❤️
</>
2026