Python Scraping for OSINT: Choose the Access Method First
Collect publicly available information responsibly using Python and standard libraries
The parsing is the short part. What decides whether your collection is usable is which rung of the access ladder you took, what you did with personal data afterwards, and whether you can still say where each fact came from.

Most scraping tutorials hand you requests, BeautifulSoup and a loop, and stop there. For OSINT work that is the least important part. The parts that decide whether your collection is usable — and whether it is lawful — are which access method you chose, what you did with personal data afterwards, and whether you can still say where each fact came from three weeks later.
This covers those, with the parsing as the short section it deserves to be.
The legal position, which is more nuanced than either camp says
Two things are commonly stated and both are wrong.
"Scraping public data is illegal" is wrong. In hiQ Labs v. LinkedIn, the Ninth Circuit found that scraping data a site publishes without authentication is unlikely to constitute unauthorised access under the Computer Fraud and Abuse Act. Public means public, and the CFAA is about circumventing access controls.
"Scraping public data is therefore fine" is also wrong, and the same case shows why: hiQ ultimately lost on breach of contract. The terms of service were enforceable even though the CFAA claim was not. So the exposure moves from criminal statute to contract, and contract is where most scraping disputes actually live.
Then there is the part OSINT practitioners cannot skip. GDPR applies to personal data whether or not it was public. A name being on a public profile does not remove it from scope. If you collect personal data about people in the UK or EU from a source other than the person themselves, Article 14 requires you to tell them — including what you collected, why, and how long you will keep it — subject to a set of exemptions you should read rather than assume apply to you.
The practical consequences for how you work:
- Authentication is the bright line. Scraping behind a login you agreed terms to access is a different legal posture from scraping an unauthenticated page. Treat the login as the boundary.
- robots.txt is not law, but it is evidence. Ignoring it does not by itself create liability. It does make it much harder to argue your access was authorised, and it will be the first exhibit if anyone asks.
- Purpose limitation is a real constraint. Collecting for one investigation and retaining for general-purpose enrichment is the step that turns a defensible activity into an indefensible one.
None of this is legal advice, and jurisdictions differ substantially. If the work is professional, the conversation with your legal team should happen before the first request, not after the first complaint.
Do not write a scraper until you have been down this ladder
Every rung below is cheaper, more robust and less contentious than the one under it. Most people start at the bottom.

The rung people miss is the third one, and it is the single most useful technique in this article. Almost every modern site renders from its own internal API. The HTML you would painstakingly parse is assembled in the browser from JSON that you can request directly.
Open developer tools, go to the Network tab, filter to Fetch/XHR, reload the page and watch what it calls. You will usually find a clean, paginated, documented-in-all-but-name endpoint returning exactly the fields you wanted — and it will not break the next time someone changes a CSS class.
# what the page itself fetches, copied straight out of devtools
curl -s 'https://example.com/api/v2/listings?page=1&per_page=50' \
-H 'Accept: application/json' | jq '.results[] | {id, title, updated_at}'A stack that holds up
requests and BeautifulSoup work, and for a hundred pages they are fine. For collection at any scale, two substitutions are worth making.
import asyncio, httpx
from selectolax.parser import HTMLParser
HEADERS = {
# Identify yourself. This is what gets you rate-limited instead of banned.
"User-Agent": "LearnCybers-Research/1.0 (+https://learncybers.com/contact)",
}
async def fetch(client, url, sem):
async with sem: # a real ceiling on concurrency
r = await client.get(url, headers=HEADERS, timeout=20)
r.raise_for_status()
await asyncio.sleep(1.0) # politeness, inside the semaphore
return r.text
async def main(urls):
sem = asyncio.Semaphore(4) # four in flight, not four hundred
async with httpx.AsyncClient(follow_redirects=True) as client:
pages = await asyncio.gather(*(fetch(client, u, sem) for u in urls),
return_exceptions=True)
for url, html in zip(urls, pages):
if isinstance(html, Exception):
print(f"FAILED {url}: {html}")
continue
tree = HTMLParser(html)
title = tree.css_first("h1")
print(url, title.text(strip=True) if title else None)httpx gives you async and HTTP/2, which matters once you are fetching thousands of pages. selectolax parses HTML substantially faster than BeautifulSoup because it wraps a C parser, and the selector syntax is close enough that switching costs an afternoon.
Two details in that snippet do more work than they look like. The semaphore is a real concurrency ceiling rather than a polite suggestion — without it, gather will happily open every connection at once. And the sleep sits inside the semaphore, so it actually paces requests instead of running in parallel and achieving nothing.
The User-Agent matters more than people expect. A string that identifies who you are and how to reach you is the difference between an administrator throttling you and an administrator blocking your entire range. Anyone who has run a site has seen both kinds of crawler.
Headless browsers, and when to give in
Playwright is the last rung because it is the expensive one: a real browser per worker, hundreds of megabytes of memory, and an order of magnitude slower than an HTTP request.
Use it when the content genuinely does not exist without JavaScript execution, when the site signs its API requests in a way you cannot reproduce, or when interaction is required to reach the data. Do not use it because the HTTP approach took twenty minutes to figure out.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
# Block what you do not need — this alone roughly halves the time
page.route("**/*.{png,jpg,jpeg,webp,gif,svg,woff,woff2,css}",
lambda route: route.abort())
page.goto(url, wait_until="networkidle")
html = page.content()
browser.close()One thing worth knowing: if a site deploys a commercial anti-bot service, working around it is where "collecting public information" starts to look like "circumventing an access control". That is the boundary the CFAA cares about, and it is a good place to stop and ask whether you need this source.
The part that makes it OSINT rather than a database
Collection is the easy half. The failure mode of open-source intelligence is not missing data, it is confident wrongness — and scraping produces confident wrongness at scale.
- Record provenance per field, not per record. Which URL, fetched at what time, and what the raw response was. A conclusion you cannot re-derive is a rumour with a timestamp.
- Keep the raw response. Store the HTML or JSON you actually received, not only your parse of it. When your extraction logic turns out to be wrong — and it will — you can reprocess rather than re-collect.
- Two independent sources, or say so. Name collisions are the single largest source of false OSINT conclusions. Two profiles with the same name are two people until something ties them together.
- Timestamp everything, including absence. "No result on 14 September" is a finding. "No result" is not.
- Expect planted data. If your collection is about someone who might be aware of it, some of what you find may exist because you were expected to find it.
Minimisation, which is both ethics and compliance
The instinct is to collect everything now in case it is useful later. Resist it, for three reasons that all point the same way.
It is the step that breaks purpose limitation under GDPR. It creates a dossier that is itself a breach risk — a scraped dataset of personal information is exactly the kind of thing that ends up in a public bucket. And it makes analysis worse, because signal density falls as you add fields nobody asked about.
Write down the question before you write the scraper. Collect the fields that answer it. Set a deletion date and honour it. If the question changes, collect again — that is cheaper than defending a dataset you cannot justify.
Scope and sources
The hiQ Labs v. LinkedIn outcome is summarised above from the published decisions; the case moved through several stages and the CFAA and contract questions were answered differently, which is precisely the point. GDPR Article 14 has exemptions, including where notification is impossible or involves disproportionate effort, and whether they apply to your work is a question for a lawyer rather than an article.
Nothing here is legal advice. Laws differ by jurisdiction and change; the technical methods are ours and are intended for collection from sources you are authorised to access. Do not use them against systems you do not have permission to query.


