System · 2026 · active
An async crawler that walks a site, checks every link it finds, and reports what is broken — written to be well-behaved rather than fast. The interesting constraints are all limits: how many requests at once, how close together per host, what to retry, and what the site has asked you not to touch.
A crawler is the standard first async project because the naive version is twenty lines. The naive version is also indistinguishable from a denial-of-service attempt: unbounded concurrency, no spacing, no robots.txt, retries that hammer whatever just failed.
So the exercise I actually wanted was the opposite one — build the limits first and let throughput be whatever politeness leaves over.
One pooled httpx.AsyncClient for the whole run, with explicit connection limits. Creating a client per request is the common error: it throws away connection reuse and quietly opens as many sockets as you have URLs. Reusing one makes the ceiling a number someone chose.
Retries use exponential backoff with jitter, and the policy is data rather than control flow. Without jitter, everything that failed together retries together, and the retry storm is often worse than the original failure. Passing it as data means the retry behaviour can be tested without waiting out real delays.
Per-host spacing and robots.txt are enforced in the fetch path, not left to the caller — a politeness rule that depends on every call site remembering it is not a rule.
Fetching sits behind a Fetcher Protocol. Test doubles satisfy it structurally, so the retry path is exercised against a real failing-then-succeeding fetcher with no mocking library involved. Structural typing is what makes that possible: the double never imports or inherits from the thing it stands in for, so the test cannot accidentally depend on the implementation.
asyncio rather than threads. Crawling is I/O-bound almost end to end, so the event loop holds far more in-flight requests per unit of memory. The cost is that any CPU-bound analysis has to leave the loop deliberately — which is why page analysis is designed for a process pool rather than being awaited inline.
mypy --strict from the first commit rather than added later. Retrofitting types onto async code is materially harder than writing them: the awaitable boundaries are exactly where the annotations matter, and exactly where they are painful to reconstruct afterwards.
The network layer is done. The analysis stage — CPU-bound page checks in a process pool — is in progress, and the report output is not written yet. Listed here as active for that reason, rather than presented as finished.