Build a Googlebot Crawler Lab
Build a local crawler with deduplication, robots rules, retries, and metrics.
Introduction
30 Second Summary
Automated visitors can waste a website's capacity by requesting the same pages repeatedly. A brief server problem can also stop useful work before the site recovers.
In this project, you will build a controlled local web crawler in Python that explores a safe website running on your computer. You will compare a naive crawl with a Googlebot-inspired design built around URL deduplication, robots.txt rules, per-host pacing, and failure recovery.
What You'll Build
Your terminal turns crawler design into a visible experiment, showing a naive crawl fail before the improved crawler finishes the same local site safely.
By the end of this project, you'll have:
- A working local crawl you can launch with one Python program and watch move through a graph of linked pages.
- A side-by-side terminal comparison that makes duplicate requests, blocked paths, host waits, retries, and recovery visible.
- A saved crawl report and architecture map in crawl_report.txt and architecture.md that you can use to explain how the local design becomes a distributed system.
- Secret Mission: Design deterministic host ownership and fenced worker failover without allowing two workers to crawl the same host at once.
Are there any prerequisites?
Basic Python and system design knowledge will help. You need Python 3.7 or newer plus Visual Studio Code.
The lab uses Python's standard library on your computer, so it needs no cloud account or third-party packages.
Before We Start
Before the hands-on work begins, this checkpoint captures your commitment to a controlled, loopback-only crawler lab. Its safeguards prevent duplicate work, disallowed-path requests, excessive request rates, and uncontrolled retries from wasting capacity or causing harmful load at web scale.
Set Up the Local Crawler Workspace
A verified local workspace keeps cloud setup from hiding crawler behavior.
This lab uses the Python standard library. No third-party package is required.
In this step, get ready to:
- Open a local folder in Visual Studio Code.
- Verify a supported Python 3 runtime.
- Run a readiness script from the integrated terminal.
Create the local workspace
Visual Studio Code keeps the editor beside the integrated terminal. A Desktop folder gives the lab one easy-to-find home.
Why use the standard library?
A crawler framework could hide the decisions this lab needs to expose. Python's built-in tools keep those decisions visible.
- Press Cmd+Space (macOS) or the Windows key (Windows) to open your search bar.
- Type Visual Studio Code into the search bar.
- Press Enter to open Visual Studio Code.
- Click File in the top menu bar.
- Click Open Folder.
- Select Desktop in the folder picker.
- Click New Folder.
- Enter googlebot-crawler-lab as the folder name.
The folder picker now shows your named lab folder on the Desktop.
- Click Create to make the folder.
- Select googlebot-crawler-lab in the folder picker.
- Click Open to use the folder as your workspace.
- Confirm that you trust the folder if Visual Studio Code displays a trust prompt.
- Click the New File icon in the Explorer toolbar.
- Enter crawler_lab.py as the file name.
- Press Enter to create the file.
You will see crawler_lab.py inside googlebot-crawler-lab in the Explorer sidebar.
Verify a supported Python version
The lab later uses ThreadingHTTPServer. That API requires Python 3.7 or newer.
- Click Terminal in the Visual Studio Code top menu.
- Click New Terminal to open the integrated terminal.
- Check your installed Python version by running this command:
python3 --version
What does this command check?
- The python3 command selects your installed Python 3 runtime.
- The --version flag prints its version number.
✔️ My version is supported
Your Python runtime is ready for the local server used later in the lab.
ⓧ I see an older version
Your current runtime is too old for ThreadingHTTPServer.
- Visit the Python macOS downloads page.
- Download the current macOS installer.
- Run the downloaded installer.
- Accept the standard installation.
- Complete the installer's Install Certificates.command step.
- Switch back to the Visual Studio Code integrated terminal.
- Check the upgraded version by running this command:
python3 --version
What confirms the upgrade?
The command must report Python version 3.7 or newer.
ⓧ Command not found
Python 3 is not available through the integrated terminal yet.
- Visit the Python macOS downloads page.
- Download the current macOS installer.
- Run the downloaded installer.
- Accept the standard installation.
- Complete the installer's Install Certificates.command step.
- Switch back to the Visual Studio Code integrated terminal.
- Check the installed version by running this command:
python3 --version
What confirms the installation?
The command must report Python version 3.7 or newer.
- Keep the integrated terminal open for the readiness run.
Add and run the readiness check
A fixed terminal message gives you a visible checkpoint before the crawler code grows. It proves the editor-to-terminal workflow is working.
- Select crawler_lab.py in the Explorer sidebar.
- Add the readiness check to crawler_lab.py by pasting this code:
print("Crawler lab ready")
What does this code do?
- The line writes a fixed readiness message to the terminal.
- That output proves Python executed the file saved by Visual Studio Code.
- Save crawler_lab.py by pressing Cmd+S (macOS) or Ctrl+S (Windows).
Before you run it, what message do you expect the terminal to print?
- Run the readiness script from the integrated terminal by running this command:
python3 crawler_lab.py
What should I see?
You will see Crawler lab ready in the integrated terminal.
That is your first working checkpoint. The same editor-to-terminal loop now carries the crawler through every later experiment.
Readiness message missing?
Confirm that the integrated terminal is inside googlebot-crawler-lab. Check that crawler_lab.py is saved before running it.
Compare your file with the full-code tab below. Help me troubleshoot why crawler_lab.py does not print its readiness message.
✔️ Awesome, I've got everything!
Great. Double-check that you saved crawler_lab.py.
ⓧ I'd like to double check the full code
print("Crawler lab ready")
Your local workspace is ready. Next, you will build a loopback website that makes the naive crawler's failures visible.
Build the Demo Web and Naive Crawler
Your verified Python workflow is ready inside VS Code. Now the lab needs a controlled website that your crawler can visit safely.
Real pages can link back to each other. A basic URL frontier can keep admitting those repeated discoveries until it wastes requests or encounters a server failure.
In this step, get ready to:
- Build a local website with cyclic links and a temporary failure.
- Extract links into an ordinary list frontier.
- Run the naive crawler to expose repeated requests and unsafe behavior.
Build the controlled demo website
A loopback HTTP server keeps every request on your computer. Its pages form a small HTML graph with a private route and one simulated overload response.
- In crawler_lab.py from earlier, select the existing print("Crawler lab ready") line.
- Replace the selected line with these imports and the crawler identity:
from html.parser import HTMLParser
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from time import sleep
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit, urlunsplit
from urllib.request import Request, urlopen
USER_AGENT = "MiniGooglebot"
What do these foundations provide?
- The standard-library imports provide the server, parser, URL tools, request client, thread, delay, and network exceptions used throughout the lab.
- The USER_AGENT value gives every crawler request a consistent identity.
- Save crawler_lab.py.
- Confirm the imports load by running this command in the integrated terminal from earlier:
python3 crawler_lab.py
What does this check prove?
Running the script asks Python to load every import and create USER_AGENT. A clean return to the terminal prompt confirms this foundation is valid.
Your terminal returns to the prompt without a traceback. The standard-library foundation is ready.
Seeing an import or syntax error?
Confirm that every import starts at the far left of crawler_lab.py. Check that MiniGooglebot remains inside matching double quotes.
Compare the first nine lines with the code block above. Help me diagnose this Python import error.
The server handler maps request paths to response bodies. Start its page map with the seed page that exposes every designed crawler problem.
- Place your cursor two lines below the USER_AGENT value.
- Add the handler and seed page by pasting this code:
class DemoSiteHandler(BaseHTTPRequestHandler):
flaky_hits = 0
pages = {
"/": """<!doctype html>
<html><body>
<h1>Crawler Lab</h1>
<ul>
<li><a href="/page-a">Page A</a></li>
<li><a href="/page-a#team">Page A fragment alias</a></li>
<li><a href="/private/notes">Private notes</a></li>
<li><a href="/flaky">Flaky page</a></li>
</ul>
</body></html>""",
}
What does this page introduce?
- The pages dictionary stores the response body for each route.
- The root page links to a normal page and a fragment alias of that page.
- The root page also advertises a private path and the route that simulates overload.
- Save crawler_lab.py.
- Check the handler syntax by running the script again:
python3 crawler_lab.py
What does this handler check prove?
Python now parses the handler class and its multiline page body. Returning to the prompt confirms the indentation and quotes are balanced.
The script returns cleanly. Your seed page is now represented inside the handler.
Seeing an indentation or string error?
Confirm that pages is indented four spaces inside DemoSiteHandler. Keep each page body between its original triple quotes.
Check that the dictionary ends with an indented closing brace. Help me fix this handler definition.
The remaining pages create a cycle between the home page, Page A, and Page B. Separate private and flaky routes make policy and recovery gaps observable.
- Inside the pages dictionary, place your cursor before its closing brace.
- Add the remaining route bodies by pasting this code:
"/page-a": """<!doctype html>
<html><body>
<h1>Page A</h1>
<a href="/">Home</a>
<a href="/page-b">Page B</a>
</body></html>""",
"/page-b": """<!doctype html>
<html><body>
<h1>Page B</h1>
<a href="/page-a">Page A again</a>
<a href="/flaky">Flaky page again</a>
</body></html>""",
"/private/notes": """<!doctype html>
<html><body><h1>Private crawler notes</h1></body></html>""",
"/flaky": """<!doctype html>
<html><body>
<h1>Recovered page</h1>
<a href="/">Home</a>
</body></html>""",
How does the page graph behave?
- Page A links back to the seed page and forward to Page B.
- Page B links back to Page A and repeats the flaky-page discovery.
- The private route exists as a normal page even though the server later publishes a rule against crawling it.
- Save crawler_lab.py.
- Confirm the expanded page map loads by running the script:
python3 crawler_lab.py
What does this page-map check prove?
Python loads every multiline page into the pages dictionary. A clean return confirms the route entries are separated correctly.
The terminal returns to the prompt without a traceback. The complete local page graph is ready.
Does the page map fail to load?
Confirm that each route entry ends with a comma. Check that the new entries remain inside the pages dictionary.
Make sure each opening triple quote has a matching closing triple quote. Help me inspect this page map.
A request handler must translate each response into bytes and send the correct headers. Suppressing the default request log keeps the crawler's own output readable.
- Place your cursor after the closing brace of the pages dictionary.
- Add the response helpers inside DemoSiteHandler by pasting this code:
def log_message(self, format, *args):
return
def send_body(self, status, body, content_type="text/html; charset=utf-8"):
payload = body.encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
What do these helpers do?
- The log_message() override prevents the server's default access log from mixing with crawler events.
- The send_body() helper encodes response text before writing it to the network stream.
- The response headers tell the client what content arrived and how many bytes it contains.
- Save crawler_lab.py.
- Confirm the response helpers load by running the script:
python3 crawler_lab.py
What does this helper check prove?
Python parses both methods as members of DemoSiteHandler. A clean return confirms their indentation keeps them inside the class.
The terminal returns without a traceback. The handler can now construct HTTP responses.
Are the helper methods outside the class?
Confirm that both def lines begin with four spaces. Their method bodies must begin with eight spaces.
Check the closing brace above these methods for four spaces of indentation. Help me correct this class indentation.
The request method controls every designed server response. It publishes the crawler rule and makes the first request to the flaky route return a server error.
- Place your cursor after the send_body() method.
- Add the request-routing method inside DemoSiteHandler by pasting this code:
def do_GET(self):
path = urlsplit(self.path).path
if path == "/robots.txt":
rules = "User-agent: MiniGooglebot\nDisallow: /private/\n"
self.send_body(200, rules, "text/plain; charset=utf-8")
return
if path == "/flaky":
DemoSiteHandler.flaky_hits += 1
if DemoSiteHandler.flaky_hits == 1:
self.send_body(500, "Temporary overload", "text/plain; charset=utf-8")
return
body = self.pages.get(path)
if body is None:
self.send_body(404, "Not found", "text/plain; charset=utf-8")
return
self.send_body(200, body)
How does request routing work?
- The request path selects a rule file, the flaky response, a known page, or the missing-page response.
- The flaky_hits counter makes only the first request to /flaky return status 500.
- The /robots.txt response disallows the private directory for the crawler identity.
- Save crawler_lab.py.
- Confirm the request method loads by running the script:
python3 crawler_lab.py
What does this routing check prove?
Python now parses every branch in do_GET(). Returning to the prompt confirms the nested conditions are valid.
The script returns cleanly. The handler can now serve pages, crawler rules, missing paths, and the temporary overload response.
Seeing an error inside do_GET()?
Check that each return remains inside its matching condition. Confirm that the final self.send_body(200, body) line aligns with the outer conditions.
Keep both newline escapes inside the crawler-rule string. Help me fix this request-routing method.
The server launcher binds to the loopback address with an ephemeral port. A background thread lets the crawler make requests while the main script continues.
- Place your cursor after the completed DemoSiteHandler class.
- Add the server launcher at the left margin by pasting this code:
def start_demo_site():
DemoSiteHandler.flaky_hits = 0
server = ThreadingHTTPServer(("127.0.0.1", 0), DemoSiteHandler)
thread = Thread(target=server.serve_forever, daemon=True)
thread.start()
host, port = server.server_address
return server, f"http://{host}:{port}"
How does the server start safely?
- The loopback address keeps the demo site local to your computer.
- Port 0 asks the operating system to choose an available ephemeral port.
- The returned base URL gives the crawler the exact host and port assigned to this run.
- Save crawler_lab.py.
- Confirm the launcher definition loads by running the script:
python3 crawler_lab.py
What does this launcher check prove?
Python creates the start_demo_site() function without starting it yet. A clean return confirms the server and thread construction is syntactically valid.
The terminal returns to the prompt without a traceback. Your controlled website is fully defined.
Does the launcher definition fail?
Confirm that start_demo_site() begins at the far left of the file. Check that all seven lines in its body use four spaces.
Keep the loopback host and port inside one tuple. Help me inspect the server launcher.
Add link discovery and the naive frontier
The crawler needs to read anchor elements before it can discover more pages. URL normalization resolves relative paths and removes fragments so each discovery has a consistent URL form.
- Place your cursor after start_demo_site().
- Add the link parser by pasting this code:
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag != "a":
return
for name, value in attrs:
if name == "href" and value:
self.links.append(value)
How does the parser discover links?
- The links list stores every discovered anchor destination.
- The start-tag callback ignores elements that are not anchors.
- Each non-empty href value is appended in discovery order.
- Save crawler_lab.py.
- Confirm the parser definition loads by running the script:
python3 crawler_lab.py
What does this parser check prove?
Python loads LinkParser as a valid parser subclass. Returning to the prompt confirms both methods are nested correctly.
The script returns cleanly. Link discovery is ready for the crawler.
Does LinkParser fail to load?
Confirm that both methods are indented inside LinkParser. Keep the for loop inside handle_starttag().
Check that self.links uses the same spelling in both methods. Help me fix this parser class.
Relative destinations need the current page as context. Removing fragments makes the two Page A links resolve to the same normalized URL, even though this first crawler still admits both discoveries.
- Place your cursor after the LinkParser class.
- Add URL normalization and link extraction by pasting this code:
def normalize_url(base_url, href):
absolute_url = urljoin(base_url, href)
defragmented_url, _ = urldefrag(absolute_url)
parts = urlsplit(defragmented_url)
if parts.scheme not in {"http", "https"} or not parts.netloc:
return None
path = parts.path or "/"
return urlunsplit(
(parts.scheme.lower(), parts.netloc.lower(), path, parts.query, "")
)
def extract_links(page_url, body):
parser = LinkParser()
parser.feed(body)
return [normalize_url(page_url, href) for href in parser.links]
What happens to each discovered URL?
- The base page turns relative destinations into absolute URLs.
- Fragment removal collapses destinations such as /page-a#team to the Page A resource.
- The scheme and network location are converted to lowercase before the URL is rebuilt.
- The extraction helper parses a response body and normalizes every discovered link.
- Save crawler_lab.py.
- Confirm both URL helpers load by running the script:
python3 crawler_lab.py
What does this URL check prove?
Python now resolves the references used by normalize_url() and extract_links(). A clean return confirms their expressions and indentation are valid.
The terminal returns without a traceback. The crawler can now turn page links into normalized destinations.
Seeing an error in the URL helpers?
Confirm that the tuple passed to urlunsplit() stays inside its parentheses. Check that extract_links() begins at the far left.
Keep the list comprehension inside the return statement. Help me diagnose these URL helpers.
The naive crawler uses a plain list as its frontier. Every same-host discovery is appended without checking whether that normalized URL was already scheduled.
- Place your cursor after extract_links().
- Add the naive crawl loop by pasting this code:
def naive_crawl(seed_url, max_fetches=8):
allowed_netloc = urlsplit(seed_url).netloc
frontier = [seed_url]
attempts = 0
while frontier and attempts < max_fetches:
url = frontier.pop(0)
attempts += 1
path = urlsplit(url).path
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
try:
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
print(f"NAIVE FETCH {response.getcode()} {path}")
except HTTPError as error:
print(f"NAIVE ERROR {error.code} {path}")
continue
except URLError as error:
print(f"NAIVE NETWORK ERROR {path}: {error.reason}")
continue
for discovered_url in extract_links(url, body):
if discovered_url and urlsplit(discovered_url).netloc == allowed_netloc:
frontier.append(discovered_url)
How does the naive crawler work?
- The frontier starts with the seed URL and removes one item from the front for each attempt.
- The network-location check prevents discovered links from leaving the local demo host.
- Successful responses are decoded before their links are extracted.
- HTTP and network failures are logged without any retry.
- Every same-host discovery is appended because this crawler has no seen set.
- Save crawler_lab.py.
- Confirm the crawl loop loads by running the script:
python3 crawler_lab.py
What does this crawl-loop check prove?
Python parses the request, exception, and discovery branches inside naive_crawl(). Returning to the prompt confirms the function is ready to call.
The terminal returns without a traceback. Your ordinary list frontier is ready for its first crawl.
Does naive_crawl fail to load?
Confirm that both except blocks align with try. Keep the discovery loop inside the while loop.
Check that frontier.append(discovered_url) remains inside the same-host condition. Help me fix this crawl loop.
Run the naive crawl and inspect the failure
The entry point starts the local server before handing its seed URL to the crawler. A finally block shuts down the listener even when the crawl encounters an error.
- Place your cursor after naive_crawl().
- Add the server lifecycle and script entry point by pasting this code:
def main():
server, base_url = start_demo_site()
sleep(0.05)
try:
print("NAIVE CRAWLER")
naive_crawl(f"{base_url}/")
finally:
server.shutdown()
server.server_close()
if __name__ == "__main__":
main()
How does the lab control its server?
- The short delay gives the background server thread time to begin serving requests.
- The generated base URL supplies the seed for the naive crawl.
- The finally block stops the server and closes its listener after every normal run or raised exception.
- The module guard calls main() when you run the file as a script.
- Save crawler_lab.py.
Before you run the finished slice, do you expect the ordinary list frontier to visit each path once or revisit paths it has already discovered?
- Start the demo server and naive crawler by running this command:
python3 crawler_lab.py
What should you see?
- The output begins with NAIVE CRAWLER.
- Repeated NAIVE FETCH lines show that the frontier admits URLs it has already scheduled.
- The /private/notes path is fetched even though the site publishes a rule against it.
- The first flaky request logs NAIVE ERROR 500 /flaky without recovering.
This shortfall is intentional. The run exposes duplicate work, ignored publisher rules, and missing failure recovery on the main path.
You have the failure baseline working: one local command now makes the crawler's design gaps visible in the terminal.
Missing the expected naive crawl events?
Confirm that main() calls naive_crawl() inside its try block. Check that max_fetches still defaults to 8.
If the terminal reports an address or connection problem, stop any previous run before trying again. Help me diagnose the missing crawler events.
✔️ Awesome, I've got everything!
Great. Double-check that you saved crawler_lab.py after the successful naive crawl.
ⓧ I'd like to double check the full code
from html.parser import HTMLParser
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from time import sleep
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit, urlunsplit
from urllib.request import Request, urlopen
USER_AGENT = "MiniGooglebot"
class DemoSiteHandler(BaseHTTPRequestHandler):
flaky_hits = 0
pages = {
"/": """<!doctype html>
<html><body>
<h1>Crawler Lab</h1>
<ul>
<li><a href="/page-a">Page A</a></li>
<li><a href="/page-a#team">Page A fragment alias</a></li>
<li><a href="/private/notes">Private notes</a></li>
<li><a href="/flaky">Flaky page</a></li>
</ul>
</body></html>""",
"/page-a": """<!doctype html>
<html><body>
<h1>Page A</h1>
<a href="/">Home</a>
<a href="/page-b">Page B</a>
</body></html>""",
"/page-b": """<!doctype html>
<html><body>
<h1>Page B</h1>
<a href="/page-a">Page A again</a>
<a href="/flaky">Flaky page again</a>
</body></html>""",
"/private/notes": """<!doctype html>
<html><body><h1>Private crawler notes</h1></body></html>""",
"/flaky": """<!doctype html>
<html><body>
<h1>Recovered page</h1>
<a href="/">Home</a>
</body></html>""",
}
def log_message(self, format, *args):
return
def send_body(self, status, body, content_type="text/html; charset=utf-8"):
payload = body.encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
def do_GET(self):
path = urlsplit(self.path).path
if path == "/robots.txt":
rules = "User-agent: MiniGooglebot\nDisallow: /private/\n"
self.send_body(200, rules, "text/plain; charset=utf-8")
return
if path == "/flaky":
DemoSiteHandler.flaky_hits += 1
if DemoSiteHandler.flaky_hits == 1:
self.send_body(500, "Temporary overload", "text/plain; charset=utf-8")
return
body = self.pages.get(path)
if body is None:
self.send_body(404, "Not found", "text/plain; charset=utf-8")
return
self.send_body(200, body)
def start_demo_site():
DemoSiteHandler.flaky_hits = 0
server = ThreadingHTTPServer(("127.0.0.1", 0), DemoSiteHandler)
thread = Thread(target=server.serve_forever, daemon=True)
thread.start()
host, port = server.server_address
return server, f"http://{host}:{port}"
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag != "a":
return
for name, value in attrs:
if name == "href" and value:
self.links.append(value)
def normalize_url(base_url, href):
absolute_url = urljoin(base_url, href)
defragmented_url, _ = urldefrag(absolute_url)
parts = urlsplit(defragmented_url)
if parts.scheme not in {"http", "https"} or not parts.netloc:
return None
path = parts.path or "/"
return urlunsplit(
(parts.scheme.lower(), parts.netloc.lower(), path, parts.query, "")
)
def extract_links(page_url, body):
parser = LinkParser()
parser.feed(body)
return [normalize_url(page_url, href) for href in parser.links]
def naive_crawl(seed_url, max_fetches=8):
allowed_netloc = urlsplit(seed_url).netloc
frontier = [seed_url]
attempts = 0
while frontier and attempts < max_fetches:
url = frontier.pop(0)
attempts += 1
path = urlsplit(url).path
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
try:
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
print(f"NAIVE FETCH {response.getcode()} {path}")
except HTTPError as error:
print(f"NAIVE ERROR {error.code} {path}")
continue
except URLError as error:
print(f"NAIVE NETWORK ERROR {path}: {error.reason}")
continue
for discovered_url in extract_links(url, body):
if discovered_url and urlsplit(discovered_url).netloc == allowed_netloc:
frontier.append(discovered_url)
def main():
server, base_url = start_demo_site()
sleep(0.05)
try:
print("NAIVE CRAWLER")
naive_crawl(f"{base_url}/")
finally:
server.shutdown()
server.server_close()
if __name__ == "__main__":
main()
Your local lab now reproduces the waste and failure modes behind a naive crawler. Next, you will replace the ordinary list with a time-ordered frontier that rejects duplicate URLs before they consume fetch capacity.
Add a Deduplicated URL Frontier
Your naive crawler exposed how cyclic links can fill a queue with repeated work. Every duplicate consumes fetch capacity that could have reached a new page.
A URL frontier controls which page the crawler fetches next. This version combines normalized URL keys with schedule-time deduplication while leaving the robots rule and temporary server failure visible.
In this step, get ready to:
- Build a time-ordered frontier with normalized same-host URLs.
- Reject duplicate URLs before they enter the frontier.
- Compare the naive crawler with the deduplicated crawler.
Build the time-ordered frontier
A min-heap keeps the next eligible URL at the front. Sequence numbers give URLs with equal ready times a stable order.
- Replace the import section at the top of crawler_lab.py with these imports:
from collections import Counter
from heapq import heappop, heappush
from html.parser import HTMLParser
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from time import monotonic, sleep
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit, urlunsplit
from urllib.request import Request, urlopen
What do these imports add?
- Counter stores event totals without requiring every metric key to be initialized.
- heappush inserts scheduled work into the frontier.
- heappop removes the frontier entry with the smallest ready time.
- monotonic provides the scheduler clock that the next step uses for request pacing.
- Save crawler_lab.py.
- Confirm the updated imports still parse by running the script in the integrated terminal from earlier:
python3 crawler_lab.py
What does this check prove?
The naive crawl log confirms that Python loaded every import successfully. The local server also completes its usual shutdown.
Seeing an import or syntax error?
Check that each import appears once at the top of crawler_lab.py. Make sure no earlier import line remains partially edited.
Use the reported line number to find the mismatch. Help me check these Python imports.
The crawler needs its frontier state before it can admit a seed URL. The scheduled set records normalized URLs at admission time.
- Find the end of naive_crawl() in crawler_lab.py.
- Add the class below naive_crawl() by pasting this code:
class Crawler:
def __init__(self, seed_url):
self.allowed_netloc = urlsplit(seed_url).netloc
self.frontier = []
self.sequence = 0
self.scheduled = set()
self.host_ready = {}
self.robots_cache = {}
self.metrics = Counter()
self.schedule(seed_url)
What does this state represent?
- allowed_netloc limits the crawl to the loopback website that supplied the seed.
- frontier holds heap entries for pending URLs.
- scheduled remembers every normalized URL admitted to the frontier.
- metrics records duplicate and scope decisions for later reporting.
- Add the scheduling method directly below __init__() inside Crawler by pasting this code:
def schedule(self, url, ready_at=0.0, attempt=0, is_retry=False):
normalized_url = normalize_url(url, url)
if normalized_url is None:
self.metrics["out_of_scope"] += 1
return
if urlsplit(normalized_url).netloc != self.allowed_netloc:
self.metrics["out_of_scope"] += 1
return
if not is_retry:
if normalized_url in self.scheduled:
self.metrics["duplicates"] += 1
print(f"DEDUP {urlsplit(normalized_url).path}")
return
self.scheduled.add(normalized_url)
self.sequence += 1
heappush(
self.frontier,
(ready_at, self.sequence, normalized_url, attempt),
)
How does admission deduplication work?
- normalize_url() turns fragment aliases into the same canonical URL key.
- allowed_netloc rejects discoveries outside the local seed host.
- scheduled blocks a normalized URL before another frontier entry can be created.
- heappush() stores the accepted URL with its ready time and sequence number.
- Save crawler_lab.py.
- Check that the new class parses by running the script again:
python3 crawler_lab.py
What does this run confirm?
The familiar naive log confirms that Python parsed the new class and its heap operations. The class is ready for a crawl loop to use it.
Seeing an indentation error?
Align schedule() with __init__() inside Crawler. Keep the method body indented one additional level.
Check that the closing parenthesis for heappush() lines up with the call. Help me fix the Crawler class indentation.
Fetch each scheduled URL once
The frontier now controls admission. A fetch method gives the crawl loop one place to perform each local HTTP request.
- Add fetch() directly below schedule() inside Crawler by pasting this code:
def fetch(self, url):
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
return response.getcode(), body
What does the fetch method return?
- Request attaches the crawler's user agent to the local request.
- urlopen() fetches the scheduled URL with a bounded timeout.
- body holds decoded HTML for link extraction.
- response.getcode() supplies the status printed by the crawl loop.
- Save crawler_lab.py.
- Confirm the fetch method contains valid Python by running the current script:
python3 crawler_lab.py
What does this check confirm?
The naive section still completes without a traceback. Python has successfully loaded the new request method.
Script failing before the crawl log?
Check that fetch() is indented inside Crawler. Confirm that its request header matches the existing USER_AGENT identifier.
Use the traceback line to check quotes and parentheses. Help me debug the fetch method.
The crawl loop repeatedly removes the earliest frontier entry. Every successful page sends its discovered links back through the same admission check.
- Add crawl() directly below fetch() inside Crawler by pasting this code:
def crawl(self, max_successful_pages=10):
while self.frontier and self.metrics["fetched"] < max_successful_pages:
ready_at, _, url, attempt = heappop(self.frontier)
parts = urlsplit(url)
path = parts.path
self.metrics["fetch_attempts"] += 1
try:
status, body = self.fetch(url)
except HTTPError as error:
self.metrics["errors"] += 1
print(f"ERROR {error.code} {path}")
continue
except URLError as error:
self.metrics["errors"] += 1
print(f"NETWORK ERROR {path}: {error.reason}")
continue
self.metrics["fetched"] += 1
print(f"FETCH {status} {path}")
for discovered_url in extract_links(url, body):
if discovered_url:
self.schedule(discovered_url)
return self.metrics
How does the crawl loop use the frontier?
- heappop() removes the entry with the earliest ready time.
- fetch() returns the HTTP status and page body for that URL.
- extract_links() discovers the next candidate URLs from successful HTML.
- schedule() filters every discovery before it can enlarge the frontier.
- Save crawler_lab.py.
Compare both crawler runs
Both crawlers need the same simulated failure state for a fair comparison. Resetting flaky_hits gives the improved crawler its own first-request HTTP 500.
- Replace the existing main() function with this version:
def main():
server, base_url = start_demo_site()
sleep(0.05)
try:
print("NAIVE CRAWLER")
naive_crawl(f"{base_url}/")
DemoSiteHandler.flaky_hits = 0
print("\nIMPROVED CRAWLER")
crawler = Crawler(f"{base_url}/")
crawler.crawl()
finally:
server.shutdown()
server.server_close()
What changes in the comparison?
- naive_crawl() runs first with its ordinary list frontier.
- DemoSiteHandler.flaky_hits resets the temporary error before the second run.
- Crawler starts with the same seed URL and schedules normalized discoveries.
- finally still closes the loopback server after both runs.
- Save crawler_lab.py.
Before you run the comparison, which crawler do you expect to print fewer fetch lines for cyclic links?
- Run both crawler versions from the integrated terminal with this command:
python3 crawler_lab.py
What should you see?
- The naive section prints repeated NAIVE FETCH paths because it admits every discovery.
- The improved section prints DEDUP for the fragment alias and cyclic links.
- The improved section still fetches /private/notes because robots rules are not enforced in this version.
- The improved section prints ERROR 500 /flaky because retry recovery is not present yet.
That is the frontier doing its job. Each normalized URL now receives at most one initial fetch attempt.
Missing the DEDUP lines?
Confirm that main() creates Crawler after the naive run. Check that schedule() adds each normalized URL to scheduled before pushing it onto the heap.
Make sure the fragment alias still exists in the demo homepage. Help me diagnose the missing DEDUP output.
✔️ Awesome, I've got everything!
Your normalized frontier is rejecting repeated discoveries. Double-check that crawler_lab.py is saved.
ⓧ I'd like to double check the full code
from collections import Counter
from heapq import heappop, heappush
from html.parser import HTMLParser
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from time import monotonic, sleep
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit, urlunsplit
from urllib.request import Request, urlopen
USER_AGENT = "MiniGooglebot"
class DemoSiteHandler(BaseHTTPRequestHandler):
flaky_hits = 0
pages = {
"/": """<!doctype html>
<html><body>
<h1>Crawler Lab</h1>
<ul>
<li><a href="/page-a">Page A</a></li>
<li><a href="/page-a#team">Page A fragment alias</a></li>
<li><a href="/private/notes">Private notes</a></li>
<li><a href="/flaky">Flaky page</a></li>
</ul>
</body></html>""",
"/page-a": """<!doctype html>
<html><body>
<h1>Page A</h1>
<a href="/">Home</a>
<a href="/page-b">Page B</a>
</body></html>""",
"/page-b": """<!doctype html>
<html><body>
<h1>Page B</h1>
<a href="/page-a">Page A again</a>
<a href="/flaky">Flaky page again</a>
</body></html>""",
"/private/notes": """<!doctype html>
<html><body><h1>Private crawler notes</h1></body></html>""",
"/flaky": """<!doctype html>
<html><body>
<h1>Recovered page</h1>
<a href="/">Home</a>
</body></html>""",
}
def log_message(self, format, *args):
return
def send_body(self, status, body, content_type="text/html; charset=utf-8"):
payload = body.encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
def do_GET(self):
path = urlsplit(self.path).path
if path == "/robots.txt":
rules = "User-agent: MiniGooglebot\nDisallow: /private/\n"
self.send_body(200, rules, "text/plain; charset=utf-8")
return
if path == "/flaky":
DemoSiteHandler.flaky_hits += 1
if DemoSiteHandler.flaky_hits == 1:
self.send_body(500, "Temporary overload", "text/plain; charset=utf-8")
return
body = self.pages.get(path)
if body is None:
self.send_body(404, "Not found", "text/plain; charset=utf-8")
return
self.send_body(200, body)
def start_demo_site():
DemoSiteHandler.flaky_hits = 0
server = ThreadingHTTPServer(("127.0.0.1", 0), DemoSiteHandler)
thread = Thread(target=server.serve_forever, daemon=True)
thread.start()
host, port = server.server_address
return server, f"http://{host}:{port}"
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag != "a":
return
for name, value in attrs:
if name == "href" and value:
self.links.append(value)
def normalize_url(base_url, href):
absolute_url = urljoin(base_url, href)
defragmented_url, _ = urldefrag(absolute_url)
parts = urlsplit(defragmented_url)
if parts.scheme not in {"http", "https"} or not parts.netloc:
return None
path = parts.path or "/"
return urlunsplit(
(parts.scheme.lower(), parts.netloc.lower(), path, parts.query, "")
)
def extract_links(page_url, body):
parser = LinkParser()
parser.feed(body)
return [normalize_url(page_url, href) for href in parser.links]
def naive_crawl(seed_url, max_fetches=8):
allowed_netloc = urlsplit(seed_url).netloc
frontier = [seed_url]
attempts = 0
while frontier and attempts < max_fetches:
url = frontier.pop(0)
attempts += 1
path = urlsplit(url).path
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
try:
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
print(f"NAIVE FETCH {response.getcode()} {path}")
except HTTPError as error:
print(f"NAIVE ERROR {error.code} {path}")
continue
except URLError as error:
print(f"NAIVE NETWORK ERROR {path}: {error.reason}")
continue
for discovered_url in extract_links(url, body):
if discovered_url and urlsplit(discovered_url).netloc == allowed_netloc:
frontier.append(discovered_url)
class Crawler:
def __init__(self, seed_url):
self.allowed_netloc = urlsplit(seed_url).netloc
self.frontier = []
self.sequence = 0
self.scheduled = set()
self.host_ready = {}
self.robots_cache = {}
self.metrics = Counter()
self.schedule(seed_url)
def schedule(self, url, ready_at=0.0, attempt=0, is_retry=False):
normalized_url = normalize_url(url, url)
if normalized_url is None:
self.metrics["out_of_scope"] += 1
return
if urlsplit(normalized_url).netloc != self.allowed_netloc:
self.metrics["out_of_scope"] += 1
return
if not is_retry:
if normalized_url in self.scheduled:
self.metrics["duplicates"] += 1
print(f"DEDUP {urlsplit(normalized_url).path}")
return
self.scheduled.add(normalized_url)
self.sequence += 1
heappush(
self.frontier,
(ready_at, self.sequence, normalized_url, attempt),
)
def fetch(self, url):
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
return response.getcode(), body
def crawl(self, max_successful_pages=10):
while self.frontier and self.metrics["fetched"] < max_successful_pages:
ready_at, _, url, attempt = heappop(self.frontier)
parts = urlsplit(url)
path = parts.path
self.metrics["fetch_attempts"] += 1
try:
status, body = self.fetch(url)
except HTTPError as error:
self.metrics["errors"] += 1
print(f"ERROR {error.code} {path}")
continue
except URLError as error:
self.metrics["errors"] += 1
print(f"NETWORK ERROR {path}: {error.reason}")
continue
self.metrics["fetched"] += 1
print(f"FETCH {status} {path}")
for discovered_url in extract_links(url, body):
if discovered_url:
self.schedule(discovered_url)
return self.metrics
def main():
server, base_url = start_demo_site()
sleep(0.05)
try:
print("NAIVE CRAWLER")
naive_crawl(f"{base_url}/")
DemoSiteHandler.flaky_hits = 0
print("\nIMPROVED CRAWLER")
crawler = Crawler(f"{base_url}/")
crawler.crawl()
finally:
server.shutdown()
server.server_close()
if __name__ == "__main__":
main()
Your crawler now prevents repeated discoveries from wasting frontier space. Next, you will enforce publisher rules and recover from temporary overload.
Enforce Politeness and Recover from Overload
Your deduplicated URL frontier now stops repeated discoveries from wasting fetch capacity. The remaining run still enters /private/notes.
A temporary HTTP 500 also ends /flaky's journey through the frontier. The crawler needs publisher rules plus a clock for each host.
You will add a robots cache and a host scheduler. You will also use bounded exponential backoff to turn temporary overload into controlled recovery.
In this step, get ready to:
- Block URLs that the local robots.txt rules disallow.
- Pace requests with a separate readiness time for each host.
- Retry temporary server failures with a bounded delay.
Load and enforce robots.txt rules
A crawler checks /robots.txt before requesting a page. An unreachable rules file triggers a complete block in this lab so the crawler fails closed.
- In crawler_lab.py, update the existing time import to include sleep by using this finished line:
from time import monotonic, sleep
Why add sleep?
- The existing monotonic() clock measures when a host becomes eligible.
- The sleep() function pauses the crawler until that eligibility time arrives.
- Add the robots parser import below the existing urllib imports by using this line:
from urllib.robotparser import RobotFileParser
What does this import provide?
The RobotFileParser class loads rules from /robots.txt. Its can_fetch() method answers whether the crawler may request a specific URL.
- Save crawler_lab.py.
- Verify that the new imports load from the integrated terminal by running this command:
python3 crawler_lab.py
What does this check prove?
You should see NAIVE CRAWLER followed by IMPROVED CRAWLER. Reaching both sections confirms that Python loaded the new standard-library imports.
Does the script stop at an import?
Check that RobotFileParser uses the same capitalization as the finished import line. Confirm that robotparser is lowercase in the module path.
Save the file before running it again. Help me troubleshoot the crawler imports.
The crawler also needs fixed limits for request spacing and retries. Keeping these values near USER_AGENT makes the crawl policy visible in one place.
- Replace the existing USER_AGENT line with this policy group:
USER_AGENT = "MiniGooglebot"
POLITENESS_SECONDS = 0.15
MAX_RETRIES = 2
RETRYABLE_STATUS_CODES = {429, 500}
What do these policies control?
- POLITENESS_SECONDS sets the minimum delay after each request attempt to one host.
- MAX_RETRIES limits how many times a retryable URL can return to the frontier.
- RETRYABLE_STATUS_CODES identifies overload responses that may succeed later.
A per-origin cache prevents the crawler from loading the same rules before every page. Each cached parser becomes the policy source for its origin.
- Inside the Crawler class, add robots_for() between schedule() and fetch() by pasting this method:
def robots_for(self, url):
parts = urlsplit(url)
origin = urlunsplit((parts.scheme, parts.netloc, "", "", ""))
if origin in self.robots_cache:
return self.robots_cache[origin]
parser = RobotFileParser()
parser.set_url(urljoin(origin, "/robots.txt"))
try:
parser.read()
print(f"ROBOTS loaded {parts.netloc}")
except URLError:
parser.parse(["User-agent: *", "Disallow: /"])
print(f"ROBOTS unreachable, fail closed {parts.netloc}")
self.metrics["robots_fetches"] += 1
self.robots_cache[origin] = parser
return parser
What does this method do?
- The origin key groups pages that share the same scheme plus network location.
- The cache returns an existing parser when the crawler has already loaded rules for that origin.
- The exception path parses a rule that disallows every URL when the rules file is unreachable.
- The metric records each rules retrieval attempt.
- Save crawler_lab.py.
- Confirm that Python can define the new method by running the lab again:
python3 crawler_lab.py
What should you see?
The program should reach the improved crawl without a syntax error. This confirms that robots_for() belongs inside the Crawler class.
- Inside crawl(), add this rules check immediately before self.metrics["fetch_attempts"] += 1:
robots = self.robots_for(url)
if not robots.can_fetch(USER_AGENT, url):
self.metrics["blocked"] += 1
print(f"BLOCKED {path}")
continue
How does the check protect the site?
- The crawler asks the cached parser whether USER_AGENT may fetch the current URL.
- A blocked URL increments the blocked metric.
- The continue statement skips the network request for that URL.
- Save crawler_lab.py.
Before the run, predict whether /private/notes will still reach the fetcher.
- Test the robots decision from the integrated terminal by running:
python3 crawler_lab.py
What changed in the crawl?
You should see ROBOTS loaded once for the local origin. You should also see BLOCKED /private/notes instead of a successful fetch for that path.
That private path is now protected. The improved crawler checks publisher rules before spending a page request.
Still fetching the private path?
Confirm that the rules check appears before the call to self.fetch(url). Check that USER_AGENT matches the local rules.
Make sure the blocked branch ends with continue. Help me debug the robots check.
Pace requests with the host clock
Deduplication controls which URLs enter the frontier. A host clock controls when an admitted URL may contact its server.
The heap already stores a ready time with each URL. The crawler can combine that time with host_ready to preserve both retry delays and host pacing.
- In crawl(), replace the lines that pop a frontier item and derive path with this scheduling block:
ready_at, _, url, attempt = heappop(self.frontier)
parts = urlsplit(url)
host = parts.netloc
path = parts.path
allowed_at = max(ready_at, self.host_ready.get(host, 0.0))
now = monotonic()
if allowed_at > now:
wait_seconds = allowed_at - now
self.metrics["waits"] += 1
print(f"WAIT {wait_seconds:.2f}s {host}")
sleep(wait_seconds)
How does the host clock work?
- allowed_at chooses the later time from the frontier entry and the host clock.
- monotonic() gives the scheduler a clock that does not move backward.
- A future eligibility time records a wait before pausing the crawler.
- Replace the existing successful-fetch metric and log lines with this block:
self.metrics["fetched"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
print(f"FETCH {status} {path}")
Why update the clock after a fetch?
A completed request moves that host's next eligible time forward by POLITENESS_SECONDS. URLs for the same host must wait for that clock before reaching the fetcher.
- Save crawler_lab.py.
Before the run, predict whether pages on the same local host will still be fetched back to back.
- Test the host clock from the integrated terminal by running:
python3 crawler_lab.py
What proves the pacing works?
You should see at least one line beginning with WAIT in the improved section. Each wait shows that a URL remained behind the host's next eligible time.
No wait lines in the improved crawl?
Check that self.host_ready[host] is updated immediately after a successful fetch. Confirm that allowed_at uses the larger of the frontier time and the host time.
Make sure sleep(wait_seconds) remains inside the eligibility condition. Help me debug the host scheduler.
Retry temporary overload
A temporary server error should delay work instead of deleting it. The retry entry keeps the same normalized URL plus an incremented attempt number.
The delay starts at 0.25 seconds. Each later attempt doubles it until MAX_RETRIES stops further rescheduling.
- Inside crawl(), replace the existing HTTPError and URLError handlers with this recovery block:
except HTTPError as error:
self.metrics["errors"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
if error.code in RETRYABLE_STATUS_CODES and attempt < MAX_RETRIES:
backoff = 0.25 * (2 ** attempt)
self.metrics["retries"] += 1
print(f"RETRY {error.code} {path} in {backoff:.2f}s")
self.schedule(
url,
ready_at=monotonic() + backoff,
attempt=attempt + 1,
is_retry=True,
)
else:
print(f"ERROR {error.code} {path}")
continue
except URLError as error:
self.metrics["errors"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
print(f"NETWORK ERROR {path}: {error.reason}")
continue
How does recovery stay controlled?
- Every failed attempt advances the host clock before the crawler considers more work.
- Only status codes in RETRYABLE_STATUS_CODES return to the frontier.
- The attempt number makes the backoff grow while enforcing the retry limit.
- is_retry=True bypasses admission deduplication for the known URL.
- Save crawler_lab.py.
Before the final run, predict whether /flaky will disappear after its first failure or return through the frontier.
- Run the completed polite crawler from the integrated terminal:
python3 crawler_lab.py
What should the finished run show?
The improved section should print ROBOTS loaded. It should also print BLOCKED /private/notes plus at least one WAIT line.
You should then see RETRY 500 /flaky before FETCH 200 /flaky. The first line records overload while the second proves recovery.
You have closed all three gaps from the earlier run. The crawler now respects publisher rules, paces its host, and recovers from temporary overload.
Does the flaky page stay failed?
Confirm that 500 appears in RETRYABLE_STATUS_CODES. Check that the retry call passes is_retry=True.
Verify that the retry uses attempt + 1 and a future ready_at value. Help me debug the retry path.
✔️ Awesome, I've got everything!
Great. Double-check that you saved crawler_lab.py after the successful recovery run.
ⓧ I'd like to double check the full code
from collections import Counter
from heapq import heappop, heappush
from html.parser import HTMLParser
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from time import monotonic, sleep
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit, urlunsplit
from urllib.request import Request, urlopen
from urllib.robotparser import RobotFileParser
USER_AGENT = "MiniGooglebot"
POLITENESS_SECONDS = 0.15
MAX_RETRIES = 2
RETRYABLE_STATUS_CODES = {429, 500}
class DemoSiteHandler(BaseHTTPRequestHandler):
flaky_hits = 0
pages = {
"/": """<!doctype html>
<html><body>
<h1>Crawler Lab</h1>
<ul>
<li><a href="/page-a">Page A</a></li>
<li><a href="/page-a#team">Page A fragment alias</a></li>
<li><a href="/private/notes">Private notes</a></li>
<li><a href="/flaky">Flaky page</a></li>
</ul>
</body></html>""",
"/page-a": """<!doctype html>
<html><body>
<h1>Page A</h1>
<a href="/">Home</a>
<a href="/page-b">Page B</a>
</body></html>""",
"/page-b": """<!doctype html>
<html><body>
<h1>Page B</h1>
<a href="/page-a">Page A again</a>
<a href="/flaky">Flaky page again</a>
</body></html>""",
"/private/notes": """<!doctype html>
<html><body><h1>Private crawler notes</h1></body></html>""",
"/flaky": """<!doctype html>
<html><body>
<h1>Recovered page</h1>
<a href="/">Home</a>
</body></html>""",
}
def log_message(self, format, *args):
return
def send_body(self, status, body, content_type="text/html; charset=utf-8"):
payload = body.encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
def do_GET(self):
path = urlsplit(self.path).path
if path == "/robots.txt":
rules = "User-agent: MiniGooglebot\nDisallow: /private/\n"
self.send_body(200, rules, "text/plain; charset=utf-8")
return
if path == "/flaky":
DemoSiteHandler.flaky_hits += 1
if DemoSiteHandler.flaky_hits == 1:
self.send_body(500, "Temporary overload", "text/plain; charset=utf-8")
return
body = self.pages.get(path)
if body is None:
self.send_body(404, "Not found", "text/plain; charset=utf-8")
return
self.send_body(200, body)
def start_demo_site():
DemoSiteHandler.flaky_hits = 0
server = ThreadingHTTPServer(("127.0.0.1", 0), DemoSiteHandler)
thread = Thread(target=server.serve_forever, daemon=True)
thread.start()
host, port = server.server_address
return server, f"http://{host}:{port}"
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag != "a":
return
for name, value in attrs:
if name == "href" and value:
self.links.append(value)
def normalize_url(base_url, href):
absolute_url = urljoin(base_url, href)
defragmented_url, _ = urldefrag(absolute_url)
parts = urlsplit(defragmented_url)
if parts.scheme not in {"http", "https"} or not parts.netloc:
return None
path = parts.path or "/"
return urlunsplit(
(parts.scheme.lower(), parts.netloc.lower(), path, parts.query, "")
)
def extract_links(page_url, body):
parser = LinkParser()
parser.feed(body)
return [normalize_url(page_url, href) for href in parser.links]
def naive_crawl(seed_url, max_fetches=8):
allowed_netloc = urlsplit(seed_url).netloc
frontier = [seed_url]
attempts = 0
while frontier and attempts < max_fetches:
url = frontier.pop(0)
attempts += 1
path = urlsplit(url).path
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
try:
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
print(f"NAIVE FETCH {response.getcode()} {path}")
except HTTPError as error:
print(f"NAIVE ERROR {error.code} {path}")
continue
except URLError as error:
print(f"NAIVE NETWORK ERROR {path}: {error.reason}")
continue
for discovered_url in extract_links(url, body):
if discovered_url and urlsplit(discovered_url).netloc == allowed_netloc:
frontier.append(discovered_url)
class Crawler:
def __init__(self, seed_url):
self.allowed_netloc = urlsplit(seed_url).netloc
self.frontier = []
self.sequence = 0
self.scheduled = set()
self.host_ready = {}
self.robots_cache = {}
self.metrics = Counter()
self.schedule(seed_url)
def schedule(self, url, ready_at=0.0, attempt=0, is_retry=False):
normalized_url = normalize_url(url, url)
if normalized_url is None:
self.metrics["out_of_scope"] += 1
return
if urlsplit(normalized_url).netloc != self.allowed_netloc:
self.metrics["out_of_scope"] += 1
return
if not is_retry:
if normalized_url in self.scheduled:
self.metrics["duplicates"] += 1
print(f"DEDUP {urlsplit(normalized_url).path}")
return
self.scheduled.add(normalized_url)
self.sequence += 1
heappush(
self.frontier,
(ready_at, self.sequence, normalized_url, attempt),
)
def robots_for(self, url):
parts = urlsplit(url)
origin = urlunsplit((parts.scheme, parts.netloc, "", "", ""))
if origin in self.robots_cache:
return self.robots_cache[origin]
parser = RobotFileParser()
parser.set_url(urljoin(origin, "/robots.txt"))
try:
parser.read()
print(f"ROBOTS loaded {parts.netloc}")
except URLError:
parser.parse(["User-agent: *", "Disallow: /"])
print(f"ROBOTS unreachable, fail closed {parts.netloc}")
self.metrics["robots_fetches"] += 1
self.robots_cache[origin] = parser
return parser
def fetch(self, url):
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
return response.getcode(), body
def crawl(self, max_successful_pages=10):
while self.frontier and self.metrics["fetched"] < max_successful_pages:
ready_at, _, url, attempt = heappop(self.frontier)
parts = urlsplit(url)
host = parts.netloc
path = parts.path
allowed_at = max(ready_at, self.host_ready.get(host, 0.0))
now = monotonic()
if allowed_at > now:
wait_seconds = allowed_at - now
self.metrics["waits"] += 1
print(f"WAIT {wait_seconds:.2f}s {host}")
sleep(wait_seconds)
robots = self.robots_for(url)
if not robots.can_fetch(USER_AGENT, url):
self.metrics["blocked"] += 1
print(f"BLOCKED {path}")
continue
self.metrics["fetch_attempts"] += 1
try:
status, body = self.fetch(url)
except HTTPError as error:
self.metrics["errors"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
if error.code in RETRYABLE_STATUS_CODES and attempt < MAX_RETRIES:
backoff = 0.25 * (2 ** attempt)
self.metrics["retries"] += 1
print(f"RETRY {error.code} {path} in {backoff:.2f}s")
self.schedule(
url,
ready_at=monotonic() + backoff,
attempt=attempt + 1,
is_retry=True,
)
else:
print(f"ERROR {error.code} {path}")
continue
except URLError as error:
self.metrics["errors"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
print(f"NETWORK ERROR {path}: {error.reason}")
continue
self.metrics["fetched"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
print(f"FETCH {status} {path}")
for discovered_url in extract_links(url, body):
if discovered_url:
self.schedule(discovered_url)
return self.metrics
def main():
server, base_url = start_demo_site()
sleep(0.05)
try:
print("NAIVE CRAWLER")
naive_crawl(f"{base_url}/")
DemoSiteHandler.flaky_hits = 0
print("\nIMPROVED CRAWLER")
crawler = Crawler(f"{base_url}/")
crawler.crawl()
finally:
server.shutdown()
server.server_close()
if __name__ == "__main__":
main()
Your crawler can now protect a host while surviving temporary overload. Next, you will measure those decisions and map each local component to a distributed system boundary.
Measure the Crawl and Map It to Web Scale
Your improved crawler now respects publisher rules. It also controls request timing before recovering from temporary overload.
A working crawler becomes a defensible system design when its behavior is measurable. Metrics expose each decision while a scale map shows where local state belongs in a distributed design.
In this step, get ready to:
- Record crawler events in a saved report.
- Map local components to distributed replacements.
- Compare the naive crawler with the completed crawler.
Write the crawl report
The crawler already updates a Counter as events occur. A report preserves those measurements after the process closes.
- In crawler_lab.py, insert this import above the existing heapq import:
from collections import Counter
What does this import provide?
The Counter class stores each metric under a name. Missing metric names return zero when the report reads them.
- Find the end of the Crawler class in crawler_lab.py.
- Insert the first part of write_report() above main() by pasting this code:
def write_report(metrics, filename="crawl_report.txt"):
metric_names = [
"fetched",
"blocked",
"duplicates",
"retries",
"errors",
"robots_fetches",
"waits",
"fetch_attempts",
"out_of_scope",
]
lines = ["Googlebot-Inspired Crawler Lab Report", ""]
lines.extend(f"{name}: {metrics[name]}" for name in metric_names)
lines.extend(
[
"",
"Scale-out component map",
"URL frontier -> durable partitioned priority queue",
"scheduled set -> sharded URL-seen key-value store",
"host_ready map -> host scheduler owned by one shard",
"fetch method -> stateless fetcher worker pool",
"LinkParser -> independently scalable parser workers",
"Counter metrics -> centralized observability pipeline",
]
)
What does this code prepare?
- The metric_names list fixes the order of the nine measurements.
- The first lines.extend() call converts every metric into a labeled report line.
- The second lines.extend() call adds six production replacements for local crawler components.
- Save crawler_lab.py.
Before you run the lab, do you expect the existing crawl behavior to remain intact after these additions?
- Check that the updated file still runs from the integrated terminal by running this command:
python3 crawler_lab.py
What does this check prove?
You should see the naive section followed by the improved section. Reaching the end proves the new import and report preparation compile successfully.
- Place your cursor on a new line after the closing parenthesis in write_report().
- Finish the function by pasting this code:
with open(filename, "w", encoding="utf-8") as report:
report.write("\n".join(lines) + "\n")
print("\nCrawl summary")
for name in metric_names:
print(f"{name}: {metrics[name]}")
print(f"Report written to {filename}")
How does the report get saved?
- The file block writes the assembled lines to crawl_report.txt using UTF-8 encoding.
- The final loop prints the same nine measurements in the terminal.
- The last message names the report file created by the run.
- In main(), find this existing line:
crawler.crawl()
What happens to the returned metrics?
This line runs the improved crawler. Its returned Counter currently has no variable to hold it.
- Replace the existing line with these two lines:
metrics = crawler.crawl()
write_report(metrics)
What changes in main?
- The metrics variable keeps the completed crawl's event counts.
- The write_report(metrics) call saves those counts after the improved crawl finishes.
- Save crawler_lab.py.
Before you run the updated lab, which crawler decisions do you expect the final summary to count?
- Generate the report from the integrated terminal by running this command:
python3 crawler_lab.py
What should the run produce?
The terminal should end with Crawl summary followed by nine metric values. The final line should confirm that the program wrote crawl_report.txt.
- Select crawl_report.txt in the Visual Studio Code Explorer sidebar.
- Confirm the report starts with the nine named metrics.
- Confirm the six scale-out mappings appear below the metrics.
You now have durable evidence of the crawler's decisions. The report preserves successful fetches alongside blocks, duplicates, retries, errors, robots fetches, waits, fetch attempts, and scope rejections.
Report file missing?
Confirm that write_report(metrics) is inside the try block. It must appear immediately after metrics = crawler.crawl().
Check that the integrated terminal is still using the googlebot-crawler-lab folder. Help me diagnose why crawl_report.txt was not created.
✔️ Awesome, I've got everything!
The crawler now prints its summary and saves crawl_report.txt after every completed run.
ⓧ I'd like to double check the full code
from collections import Counter
from heapq import heappop, heappush
from html.parser import HTMLParser
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Thread
from time import monotonic, sleep
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit, urlunsplit
from urllib.request import Request, urlopen
from urllib.robotparser import RobotFileParser
USER_AGENT = "MiniGooglebot"
POLITENESS_SECONDS = 0.15
MAX_RETRIES = 2
RETRYABLE_STATUS_CODES = {429, 500}
class DemoSiteHandler(BaseHTTPRequestHandler):
flaky_hits = 0
pages = {
"/": """<!doctype html>
<html><body>
<h1>Crawler Lab</h1>
<ul>
<li><a href="/page-a">Page A</a></li>
<li><a href="/page-a#team">Page A fragment alias</a></li>
<li><a href="/private/notes">Private notes</a></li>
<li><a href="/flaky">Flaky page</a></li>
</ul>
</body></html>""",
"/page-a": """<!doctype html>
<html><body>
<h1>Page A</h1>
<a href="/">Home</a>
<a href="/page-b">Page B</a>
</body></html>""",
"/page-b": """<!doctype html>
<html><body>
<h1>Page B</h1>
<a href="/page-a">Page A again</a>
<a href="/flaky">Flaky page again</a>
</body></html>""",
"/private/notes": """<!doctype html>
<html><body><h1>Private crawler notes</h1></body></html>""",
"/flaky": """<!doctype html>
<html><body>
<h1>Recovered page</h1>
<a href="/">Home</a>
</body></html>""",
}
def log_message(self, format, *args):
return
def send_body(self, status, body, content_type="text/html; charset=utf-8"):
payload = body.encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
def do_GET(self):
path = urlsplit(self.path).path
if path == "/robots.txt":
rules = "User-agent: MiniGooglebot\nDisallow: /private/\n"
self.send_body(200, rules, "text/plain; charset=utf-8")
return
if path == "/flaky":
DemoSiteHandler.flaky_hits += 1
if DemoSiteHandler.flaky_hits == 1:
self.send_body(500, "Temporary overload", "text/plain; charset=utf-8")
return
body = self.pages.get(path)
if body is None:
self.send_body(404, "Not found", "text/plain; charset=utf-8")
return
self.send_body(200, body)
def start_demo_site():
DemoSiteHandler.flaky_hits = 0
server = ThreadingHTTPServer(("127.0.0.1", 0), DemoSiteHandler)
thread = Thread(target=server.serve_forever, daemon=True)
thread.start()
host, port = server.server_address
return server, f"http://{host}:{port}"
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag != "a":
return
for name, value in attrs:
if name == "href" and value:
self.links.append(value)
def normalize_url(base_url, href):
absolute_url = urljoin(base_url, href)
defragmented_url, _ = urldefrag(absolute_url)
parts = urlsplit(defragmented_url)
if parts.scheme not in {"http", "https"} or not parts.netloc:
return None
path = parts.path or "/"
return urlunsplit(
(parts.scheme.lower(), parts.netloc.lower(), path, parts.query, "")
)
def extract_links(page_url, body):
parser = LinkParser()
parser.feed(body)
return [normalize_url(page_url, href) for href in parser.links]
def naive_crawl(seed_url, max_fetches=8):
allowed_netloc = urlsplit(seed_url).netloc
frontier = [seed_url]
attempts = 0
while frontier and attempts < max_fetches:
url = frontier.pop(0)
attempts += 1
path = urlsplit(url).path
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
try:
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
print(f"NAIVE FETCH {response.getcode()} {path}")
except HTTPError as error:
print(f"NAIVE ERROR {error.code} {path}")
continue
except URLError as error:
print(f"NAIVE NETWORK ERROR {path}: {error.reason}")
continue
for discovered_url in extract_links(url, body):
if discovered_url and urlsplit(discovered_url).netloc == allowed_netloc:
frontier.append(discovered_url)
class Crawler:
def __init__(self, seed_url):
self.allowed_netloc = urlsplit(seed_url).netloc
self.frontier = []
self.sequence = 0
self.scheduled = set()
self.host_ready = {}
self.robots_cache = {}
self.metrics = Counter()
self.schedule(seed_url)
def schedule(self, url, ready_at=0.0, attempt=0, is_retry=False):
normalized_url = normalize_url(url, url)
if normalized_url is None:
self.metrics["out_of_scope"] += 1
return
if urlsplit(normalized_url).netloc != self.allowed_netloc:
self.metrics["out_of_scope"] += 1
return
if not is_retry:
if normalized_url in self.scheduled:
self.metrics["duplicates"] += 1
print(f"DEDUP {urlsplit(normalized_url).path}")
return
self.scheduled.add(normalized_url)
self.sequence += 1
heappush(
self.frontier,
(ready_at, self.sequence, normalized_url, attempt),
)
def robots_for(self, url):
parts = urlsplit(url)
origin = urlunsplit((parts.scheme, parts.netloc, "", "", ""))
if origin in self.robots_cache:
return self.robots_cache[origin]
parser = RobotFileParser()
parser.set_url(urljoin(origin, "/robots.txt"))
try:
parser.read()
print(f"ROBOTS loaded {parts.netloc}")
except URLError:
parser.parse(["User-agent: *", "Disallow: /"])
print(f"ROBOTS unreachable, fail closed {parts.netloc}")
self.metrics["robots_fetches"] += 1
self.robots_cache[origin] = parser
return parser
def fetch(self, url):
request = Request(url, headers={"User-Agent": f"{USER_AGENT}/1.0"})
with urlopen(request, timeout=2) as response:
body = response.read().decode("utf-8", errors="replace")
return response.getcode(), body
def crawl(self, max_successful_pages=10):
while self.frontier and self.metrics["fetched"] < max_successful_pages:
ready_at, _, url, attempt = heappop(self.frontier)
parts = urlsplit(url)
host = parts.netloc
path = parts.path
allowed_at = max(ready_at, self.host_ready.get(host, 0.0))
now = monotonic()
if allowed_at > now:
wait_seconds = allowed_at - now
self.metrics["waits"] += 1
print(f"WAIT {wait_seconds:.2f}s {host}")
sleep(wait_seconds)
robots = self.robots_for(url)
if not robots.can_fetch(USER_AGENT, url):
self.metrics["blocked"] += 1
print(f"BLOCKED {path}")
continue
self.metrics["fetch_attempts"] += 1
try:
status, body = self.fetch(url)
except HTTPError as error:
self.metrics["errors"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
if error.code in RETRYABLE_STATUS_CODES and attempt < MAX_RETRIES:
backoff = 0.25 * (2 ** attempt)
self.metrics["retries"] += 1
print(f"RETRY {error.code} {path} in {backoff:.2f}s")
self.schedule(
url,
ready_at=monotonic() + backoff,
attempt=attempt + 1,
is_retry=True,
)
else:
print(f"ERROR {error.code} {path}")
continue
except URLError as error:
self.metrics["errors"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
print(f"NETWORK ERROR {path}: {error.reason}")
continue
self.metrics["fetched"] += 1
self.host_ready[host] = monotonic() + POLITENESS_SECONDS
print(f"FETCH {status} {path}")
for discovered_url in extract_links(url, body):
if discovered_url:
self.schedule(discovered_url)
return self.metrics
def write_report(metrics, filename="crawl_report.txt"):
metric_names = [
"fetched",
"blocked",
"duplicates",
"retries",
"errors",
"robots_fetches",
"waits",
"fetch_attempts",
"out_of_scope",
]
lines = ["Googlebot-Inspired Crawler Lab Report", ""]
lines.extend(f"{name}: {metrics[name]}" for name in metric_names)
lines.extend(
[
"",
"Scale-out component map",
"URL frontier -> durable partitioned priority queue",
"scheduled set -> sharded URL-seen key-value store",
"host_ready map -> host scheduler owned by one shard",
"fetch method -> stateless fetcher worker pool",
"LinkParser -> independently scalable parser workers",
"Counter metrics -> centralized observability pipeline",
]
)
with open(filename, "w", encoding="utf-8") as report:
report.write("\n".join(lines) + "\n")
print("\nCrawl summary")
for name in metric_names:
print(f"{name}: {metrics[name]}")
print(f"Report written to {filename}")
def main():
server, base_url = start_demo_site()
sleep(0.05)
try:
print("NAIVE CRAWLER")
naive_crawl(f"{base_url}/")
DemoSiteHandler.flaky_hits = 0
print("\nIMPROVED CRAWLER")
crawler = Crawler(f"{base_url}/")
metrics = crawler.crawl()
write_report(metrics)
finally:
server.shutdown()
server.server_close()
if __name__ == "__main__":
main()
Document the scale-out design
Counts describe what happened in one process. The architecture document assigns each local structure a production boundary that can preserve state or scale independently.
- Select the googlebot-crawler-lab folder in the Explorer sidebar.
- Click the New File icon in the Explorer toolbar.
- Enter architecture.md as the file name.
- Add the scope, pipeline, and component boundaries by pasting this content:
# Googlebot-Inspired Crawler Scale Map
## Scope
This project models public crawler behaviors and general scalable-crawler patterns. It does not reproduce Google's proprietary implementation.
## Local pipeline
Seed URL -> URL frontier -> robots and host scheduler -> fetcher -> HTML parser -> URL normalization and seen set -> URL frontier
The demo site and crawler share one process only to keep the lab safe and observable.
## Production component boundaries
| Local component | Distributed replacement | Reason to separate it |
| --- | --- | --- |
| `frontier` heap | Durable partitioned priority queue | Survive restarts and scale queued work beyond one process |
| `scheduled` set | Sharded URL-seen key-value store | Make URL admission idempotent across workers |
| `robots_cache` | Replicated cache backed by durable policy records | Reuse rules without fetching robots.txt before every page |
| `host_ready` map | Host scheduler partitioned by host ownership | Keep one authoritative next-allowed time per host |
| `fetch` method | Stateless fetcher pool | Scale network-bound work independently |
| `LinkParser` | Parser worker pool reading stored responses | Replay parsing without refetching the source website |
| `Counter` metrics | Metrics and event pipeline | Observe throughput, errors, blocks, retries, and queue growth |
How do these boundaries scale?
- The pipeline follows a URL from discovery back into the frontier.
- The frontier and seen set become durable shared state.
- Robots rules and host clocks become reusable scheduling records.
- Fetching and parsing become separate worker pools.
- Save architecture.md.
- Confirm the document shows the local pipeline above a seven-row component table.
Component boundaries define where state lives. Core decisions define the rules that every boundary must preserve.
- Append the core decisions and failure boundaries to architecture.md by pasting this content:
## Core decisions
1. Partition frontier ownership by host, not by full URL, so all politeness state for one host has one owner.
2. Deduplicate when admitting a URL to the frontier so repeated discoveries do not enlarge the queue.
3. Treat fetch delivery as retryable and processing as idempotent. A replayed URL must not create duplicate downstream records.
4. Put a durable checkpoint between fetching and parsing at web scale so parser recovery does not spend another network request.
5. Cache robots rules, but refresh them according to the applicable protocol and crawler policy.
## Failure boundaries
- Frontier failure: restore queued URLs, attempt counts, and ready times from durable state.
- Fetcher failure: retry the leased URL after its visibility or lease timeout.
- Parser failure: replay the stored response without contacting the host again.
- Host-owner failure: fence the old owner before assigning the host and its clock to another worker.
Why do these rules matter?
Host-based ownership preserves one authoritative politeness clock for each host. Schedule-time deduplication keeps repeated discoveries out of the queue.
Durable checkpoints make recovery possible after a component failure. Idempotency prevents replayed work from creating duplicate downstream records.
- Save architecture.md.
- Confirm the document names recovery behavior for four failure boundaries.
A useful system design also states its limits. Explicit omissions keep the lab from implying that it reproduces a complete search crawler.
- Finish architecture.md by appending this content:
## Deliberate omissions
- JavaScript rendering
- DNS caching and resolution infrastructure
- Content fingerprinting and canonical clustering
- Persistent page storage and indexing
- Multi-region coordination
These belong in follow-on projects after the frontier, politeness, and retry mental model is solid.
Why name the omissions?
The omissions separate the lab's tested behaviors from the capabilities of a web-scale search system. They also define clear boundaries for future designs.
- Save architecture.md.
- Confirm the final section lists five deliberate omissions.
Architecture section missing?
Check that Core decisions follows the component table. Confirm that Deliberate omissions follows the failure boundaries.
Save the file before comparing it with the reference. Help me compare my architecture document with the required scale map.
✔️ Awesome, I've got everything!
The architecture document now connects the local crawler to durable state, worker pools, recovery rules, and explicit design limits.
ⓧ I'd like to double check the full code
# Googlebot-Inspired Crawler Scale Map
## Scope
This project models public crawler behaviors and general scalable-crawler patterns. It does not reproduce Google's proprietary implementation.
## Local pipeline
Seed URL -> URL frontier -> robots and host scheduler -> fetcher -> HTML parser -> URL normalization and seen set -> URL frontier
The demo site and crawler share one process only to keep the lab safe and observable.
## Production component boundaries
| Local component | Distributed replacement | Reason to separate it |
| --- | --- | --- |
| `frontier` heap | Durable partitioned priority queue | Survive restarts and scale queued work beyond one process |
| `scheduled` set | Sharded URL-seen key-value store | Make URL admission idempotent across workers |
| `robots_cache` | Replicated cache backed by durable policy records | Reuse rules without fetching robots.txt before every page |
| `host_ready` map | Host scheduler partitioned by host ownership | Keep one authoritative next-allowed time per host |
| `fetch` method | Stateless fetcher pool | Scale network-bound work independently |
| `LinkParser` | Parser worker pool reading stored responses | Replay parsing without refetching the source website |
| `Counter` metrics | Metrics and event pipeline | Observe throughput, errors, blocks, retries, and queue growth |
## Core decisions
1. Partition frontier ownership by host, not by full URL, so all politeness state for one host has one owner.
2. Deduplicate when admitting a URL to the frontier so repeated discoveries do not enlarge the queue.
3. Treat fetch delivery as retryable and processing as idempotent. A replayed URL must not create duplicate downstream records.
4. Put a durable checkpoint between fetching and parsing at web scale so parser recovery does not spend another network request.
5. Cache robots rules, but refresh them according to the applicable protocol and crawler policy.
## Failure boundaries
- Frontier failure: restore queued URLs, attempt counts, and ready times from durable state.
- Fetcher failure: retry the leased URL after its visibility or lease timeout.
- Parser failure: replay the stored response without contacting the host again.
- Host-owner failure: fence the old owner before assigning the host and its clock to another worker.
## Deliberate omissions
- JavaScript rendering
- DNS caching and resolution infrastructure
- Content fingerprinting and canonical clustering
- Persistent page storage and indexing
- Multi-region coordination
These belong in follow-on projects after the frontier, politeness, and retry mental model is solid.
Run and explain the finished lab
The finished run places the flawed crawler beside the improved crawler in one terminal history. That comparison makes the seen set, host clock, and retry frontier observable.
Before you run the final check, which naive failures do you expect the improved section to prevent or recover from?
- Run the completed crawler lab from the integrated terminal with this command:
python3 crawler_lab.py
What does the final run prove?
- The naive section contains repeated NAIVE FETCH paths.
- The naive section includes /private/notes before logging NAIVE ERROR 500 /flaky.
- The improved section prints DEDUP for repeated discoveries.
- The improved section prints BLOCKED /private/notes after applying the robots rule.
- The improved section prints RETRY 500 /flaky before later printing FETCH 200 /flaky.
- The terminal ends with the crawl summary and report confirmation.
That completes the crawler loop. It discovers pages, rejects duplicate work, protects the host, recovers from overload, and records the outcome.
Final run stops early?
Read the first traceback line that points to crawler_lab.py. Compare the nearby indentation with the full-file reference.
Confirm the previous run finished before starting another run. Help me diagnose why the crawler stops before the crawl summary.
- Select crawl_report.txt in the Explorer sidebar.
- Confirm the file records all nine crawler measurements.
- Confirm the file contains six scale-out component mappings.
- Select architecture.md in the Explorer sidebar.
- Trace the local pipeline from the seed URL back to the URL frontier.
- Match scheduled to the sharded URL-seen store.
- Match host_ready to the host scheduler.
How do the safeguards protect the crawl?
The scheduled set rejects normalized URLs before duplicate work enters the frontier. This preserves fetch capacity.
The host_ready clock gives each host one next-allowed request time. This prevents request bursts.
Retry entries preserve the URL, attempt number, and future ready time. Temporary overload returns to the time-ordered frontier without bypassing host pacing.
You have closed the loop from crawler behavior to system design. Your lab now demonstrates the local safeguards and the boundaries needed to scale them.
Secret mission
Preserve Politeness During a Worker Failure
Extend your crawler architecture with deterministic host ownership across three frontier workers. Design a fenced failover sequence that restores crawl state before another worker resumes requests.
Clean Up Your Resources
Clean Up Your Resources
Choose whether to keep your resources available, pause the crawler for later, or delete the lab entirely. This local-only project has no ongoing costs.
Resources you used:
- A local googlebot-crawler-lab folder containing crawler_lab.py, architecture.md, and crawl_report.txt.
- An in-process loopback HTTP listener that exists only while crawler_lab.py runs.
Keep everything running
No action is needed while you are actively building. Your files remain available for another crawl or architecture review.
- Leave the googlebot-crawler-lab folder on your Mac.
- Return to the VS Code workspace whenever you want to rerun crawler_lab.py.
Each completed run reaches the finally block. That cleanup closes the local listener automatically.
Pause - I'll come back to this later
Pausing keeps every file while ensuring the local listener is stopped. A completed run already reaches this state through its finally block.
- Return to the VS Code workspace from earlier.
- Select the integrated terminal from earlier.
- Stop the active Python process if crawler_lab.py is still running.
- Confirm that the terminal prompt has returned.
- Keep the googlebot-crawler-lab folder for your next session.
Your crawler is now stopped. Every lab file remains ready for your return.
Delete - I don't want to use this again
Complete removal deletes the local lab so you can start fresh. No cloud resource or external database is affected.
- Return to the googlebot-crawler-lab workspace in VS Code.
- Stop the current Python process in the integrated terminal if the crawler is still running.
- Confirm that the terminal prompt has returned.
- Close the googlebot-crawler-lab workspace in VS Code.
- Locate the googlebot-crawler-lab folder in Finder.
- Move the googlebot-crawler-lab folder to the Trash.
- Empty the Trash.
- Confirm that the googlebot-crawler-lab folder no longer appears in Finder.
The local crawler code is now removed. Its report and fenced-failover design are removed with the same folder.
Nice Work!
Nice Work!
You did it! You built a local web crawler lab that reveals scheduling failures before resolving them with explicit safeguards.
You've learned how to:
- Run a controlled crawler demo against a local web graph. Observe repeated fetches, private-path access, and temporary server failure.
- Build a normalized URL frontier that rejects duplicates before scheduling. Enforce robots.txt rules, per-host pacing, and bounded retries.
- Capture crawl observability in crawl_report.txt. Map each local component to a distributed production boundary in architecture.md.
- Secret Mission: Design deterministic host sharding with one fenced lease owner per host. Restore durable crawl state before a replacement worker resumes fetching.
Ready to quiz yourself?