Some of the most valuable data on the web has no API. Prices, job listings, research datasets, product catalogs — all sitting in HTML, waiting. Web scraping turns that public data into structured datasets you can analyze. Done respectfully, it is a legitimate engineering skill; done carelessly, it gets you blocked or worse.
This guide covers the full ladder: polite crawling rules, static scraping with BeautifulSoup, JavaScript-heavy sites with Playwright, and storing results properly. All examples use Python 3.
1. The Polite Crawler's Checklist
Before writing a single request, do this — it keeps you ethical and unblocked:
- Read robots.txt (
https://site.com/robots.txt) and honorDisallowrules. - Check the Terms of Service — some sites prohibit scraping explicitly.
- Identify yourself with a real User-Agent and contact email.
- Rate-limit: 1 request per 2–5 seconds is polite; bursts get you banned.
- Prefer official APIs when they exist — scraping is plan B.
2. Static Pages: requests + BeautifulSoup
Most blogs, docs, and listings are server-rendered HTML. requests fetches, BeautifulSoup parses. Always set a timeout and a User-Agent:
import requests
from bs4 import BeautifulSoup
HEADERS = {"User-Agent": "DevInsightsBot/1.0 (contact@codingtutorials.site)"}
resp = requests.get("https://example-blog.com/articles", headers=HEADERS, timeout=15)
resp.raise_for_status()
soup = BeautifulSoup(resp.text, "html.parser")
for card in soup.select("article.post-card"):
title = card.select_one("h2").get_text(strip=True)
link = card.select_one("a")["href"]
print(title, "->", link)
3. Pagination and Resilient Fetching
Real crawlers handle failures. Wrap requests in a retry loop with exponential backoff, and walk paginated listings until no “next” link remains:
import time
import requests
def fetch(url, tries=4):
for attempt in range(tries):
try:
r = requests.get(url, headers=HEADERS, timeout=15)
if r.status_code == 429: # rate limited: back off
time.sleep(2 ** attempt * 5)
continue
r.raise_for_status()
return r.text
except requests.RequestException:
time.sleep(2 ** attempt) # 1s, 2s, 4s...
return None
# walk pages politely
page, all_titles = 1, []
while True:
html = fetch(f"https://example-blog.com/articles?page={page}")
if not html:
break
titles = parse_titles(html) # your BeautifulSoup logic
if not titles:
break
all_titles.extend(titles)
page += 1
time.sleep(3) # be nice to the server
4. JavaScript-Rendered Pages: Playwright
When content loads via JavaScript (React/Vue apps), HTTP fetching returns an empty shell. Playwright drives a real headless browser, waits for content, then hands you the rendered HTML:
# pip install playwright && playwright install chromium
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(user_agent=HEADERS["User-Agent"])
page.goto("https://example-shop.com/products", wait_until="networkidle")
page.wait_for_selector(".product-card") # wait for JS render
html = page.content() # now parse with BeautifulSoup
browser.close()
Playwright is heavier than requests — reserve it for pages that truly need JavaScript. For everything else, stay with BeautifulSoup.
5. Store It Right: SQLite
Do not accumulate CSVs in a folder. SQLite is a zero-setup database in a single file — perfect for scraped datasets you will query later:
import sqlite3
db = sqlite3.connect("articles.db")
db.execute("""CREATE TABLE IF NOT EXISTS articles
(id INTEGER PRIMARY KEY, title TEXT UNIQUE, url TEXT,
scraped_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP)""")
for title, url in all_titles:
db.execute("INSERT OR IGNORE INTO articles (title, url) VALUES (?, ?)",
(title, url))
db.commit()
print(db.execute("SELECT COUNT(*) FROM articles").fetchone()[0], "rows stored")
6. Develop Faster: Cache Responses Locally
While building your scraper you will re-run it dozens of times. Hitting the live server every run is slow and rude. Cache HTML to disk during development so repeat runs cost the target site nothing:
import hashlib, os
import requests
CACHE_DIR = ".cache"
os.makedirs(CACHE_DIR, exist_ok=True)
def fetch_cached(url):
key = hashlib.sha256(url.encode()).hexdigest()
path = os.path.join(CACHE_DIR, key + ".html")
if os.path.exists(path):
return open(path, encoding="utf-8").read()
html = fetch(url) # your polite fetch() from section 3
if html:
open(path, "w", encoding="utf-8").write(html)
return html
Delete the cache folder when you want a fresh crawl. This one habit makes scraper development dramatically faster — and keeps you a good citizen of the web.
Frequently Asked Questions (FAQ)
Q: Is web scraping legal?
Scraping publicly accessible data is generally allowed in many jurisdictions, but Terms of Service, copyright, and anti-circumvention rules vary. Always check robots.txt and ToS, never bypass login walls or CAPTCHAs, and consult a lawyer for commercial use.
Q: BeautifulSoup or Scrapy?
BeautifulSoup for scripts and one-off jobs. Scrapy when you need a full crawling framework: link following, middleware, pipelines, and concurrency out of the box.
Q: How do I avoid getting IP-blocked?
Respectful rate limiting (a few seconds between requests), a real User-Agent, honoring robots.txt, and scraping off-peak. If you need scale, that is a sign to look for an official API or data partnership.
Conclusion
Start simple: requests plus BeautifulSoup handles most of the web. Reach for Playwright only when JavaScript rendering forces your hand, always crawl politely, and store results in SQLite so your data stays queryable. Respectful scraping is a superpower — abusive scraping ruins it for everyone.
💡 Engineering Key Takeaway
Check robots.txt first, identify your bot honestly, back off on 429s, and prefer APIs over scraping. The best scraper is the one nobody notices.