Some of the most valuable data on the web has no API. Prices, job listings, research datasets, product catalogs — all sitting in HTML, waiting. Web scraping turns that public data into structured datasets you can analyze. Done respectfully, it is a legitimate engineering skill; done carelessly, it gets you blocked or worse.

This guide covers the full ladder: polite crawling rules, static scraping with BeautifulSoup, JavaScript-heavy sites with Playwright, and storing results properly. All examples use Python 3.

1. The Polite Crawler's Checklist

Before writing a single request, do this — it keeps you ethical and unblocked:

2. Static Pages: requests + BeautifulSoup

Most blogs, docs, and listings are server-rendered HTML. requests fetches, BeautifulSoup parses. Always set a timeout and a User-Agent:

Python (BeautifulSoup Basics)
import requests
from bs4 import BeautifulSoup

HEADERS = {"User-Agent": "DevInsightsBot/1.0 (contact@codingtutorials.site)"}

resp = requests.get("https://example-blog.com/articles", headers=HEADERS, timeout=15)
resp.raise_for_status()

soup = BeautifulSoup(resp.text, "html.parser")
for card in soup.select("article.post-card"):
    title = card.select_one("h2").get_text(strip=True)
    link = card.select_one("a")["href"]
    print(title, "->", link)

3. Pagination and Resilient Fetching

Real crawlers handle failures. Wrap requests in a retry loop with exponential backoff, and walk paginated listings until no “next” link remains:

Python (Retries + Pagination)
import time
import requests

def fetch(url, tries=4):
    for attempt in range(tries):
        try:
            r = requests.get(url, headers=HEADERS, timeout=15)
            if r.status_code == 429:  # rate limited: back off
                time.sleep(2 ** attempt * 5)
                continue
            r.raise_for_status()
            return r.text
        except requests.RequestException:
            time.sleep(2 ** attempt)  # 1s, 2s, 4s...
    return None

# walk pages politely
page, all_titles = 1, []
while True:
    html = fetch(f"https://example-blog.com/articles?page={page}")
    if not html:
        break
    titles = parse_titles(html)  # your BeautifulSoup logic
    if not titles:
        break
    all_titles.extend(titles)
    page += 1
    time.sleep(3)  # be nice to the server

4. JavaScript-Rendered Pages: Playwright

When content loads via JavaScript (React/Vue apps), HTTP fetching returns an empty shell. Playwright drives a real headless browser, waits for content, then hands you the rendered HTML:

Python (Playwright)
# pip install playwright && playwright install chromium
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(user_agent=HEADERS["User-Agent"])
    page.goto("https://example-shop.com/products", wait_until="networkidle")
    page.wait_for_selector(".product-card")  # wait for JS render
    html = page.content()          # now parse with BeautifulSoup
    browser.close()

Playwright is heavier than requests — reserve it for pages that truly need JavaScript. For everything else, stay with BeautifulSoup.

5. Store It Right: SQLite

Do not accumulate CSVs in a folder. SQLite is a zero-setup database in a single file — perfect for scraped datasets you will query later:

Python (SQLite Storage)
import sqlite3

db = sqlite3.connect("articles.db")
db.execute("""CREATE TABLE IF NOT EXISTS articles
              (id INTEGER PRIMARY KEY, title TEXT UNIQUE, url TEXT,
               scraped_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP)""")

for title, url in all_titles:
    db.execute("INSERT OR IGNORE INTO articles (title, url) VALUES (?, ?)",
               (title, url))
db.commit()
print(db.execute("SELECT COUNT(*) FROM articles").fetchone()[0], "rows stored")

6. Develop Faster: Cache Responses Locally

While building your scraper you will re-run it dozens of times. Hitting the live server every run is slow and rude. Cache HTML to disk during development so repeat runs cost the target site nothing:

Python (Response Cache)
import hashlib, os
import requests

CACHE_DIR = ".cache"
os.makedirs(CACHE_DIR, exist_ok=True)

def fetch_cached(url):
    key = hashlib.sha256(url.encode()).hexdigest()
    path = os.path.join(CACHE_DIR, key + ".html")
    if os.path.exists(path):
        return open(path, encoding="utf-8").read()
    html = fetch(url)          # your polite fetch() from section 3
    if html:
        open(path, "w", encoding="utf-8").write(html)
    return html

Delete the cache folder when you want a fresh crawl. This one habit makes scraper development dramatically faster — and keeps you a good citizen of the web.

Frequently Asked Questions (FAQ)

Q: Is web scraping legal?

Scraping publicly accessible data is generally allowed in many jurisdictions, but Terms of Service, copyright, and anti-circumvention rules vary. Always check robots.txt and ToS, never bypass login walls or CAPTCHAs, and consult a lawyer for commercial use.

Q: BeautifulSoup or Scrapy?

BeautifulSoup for scripts and one-off jobs. Scrapy when you need a full crawling framework: link following, middleware, pipelines, and concurrency out of the box.

Q: How do I avoid getting IP-blocked?

Respectful rate limiting (a few seconds between requests), a real User-Agent, honoring robots.txt, and scraping off-peak. If you need scale, that is a sign to look for an official API or data partnership.

Conclusion

Start simple: requests plus BeautifulSoup handles most of the web. Reach for Playwright only when JavaScript rendering forces your hand, always crawl politely, and store results in SQLite so your data stays queryable. Respectful scraping is a superpower — abusive scraping ruins it for everyone.

💡 Engineering Key Takeaway

Check robots.txt first, identify your bot honestly, back off on 429s, and prefer APIs over scraping. The best scraper is the one nobody notices.

SK

Written by Sajid Khan

Principal Software Engineer & Author

Sajid is a full-stack engineer and tech writer passionate about web performance, resilient backend architectures, and developer mentorship. He authors in-depth tutorials on modern JavaScript, React, and systems engineering.