Do you want to scrape links?

Scripting is where Python really pays off. The Python Scripting Essentials Bundle has the kind of small automation projects you would build for work.

Scraping links is the first thing you do with a page you did not write. You get the html, find every a tag and read the href value out of it.

That gives you the menu, the footer, every article link and every mail address, in the order they appear in the page.

Scraping links is the first thing you do when you look at a page you did not write. You get the html, find every a tag and read the href value.

That gives you the menu, the footer, every article link and every mail address on the page, in the order they appear.

The module urllib2 can be used to download webpage data. Webpage data is always formatted in HTML format.

To cope with the HTML format data, we use a Python module named BeautifulSoup. BeautifulSoup is a Python module for parsing webpages (HTML).

All of the links will be returned as a list, like so:

['//slashdot.org/faq/slashmeta.shtml', ... ,'mailto:[email protected]', '#', '//slashdot.org/blog', '#', '#', '//slashdot.org']

We scrape a webpage with these steps:

  • download webpage data (html)
  • create beautifulsoup object and parse webpage data
  • use soups method findAll to find all links by the a tag
  • store all links in list

To get all links from a webpage:

 
from bs4 import BeautifulSoup
from urllib.request import Request, urlopen
import re

req = Request("http://slashdot.org")
html_page = urlopen(req)

soup = BeautifulSoup(html_page, "lxml")

links = []
for link in soup.findAll('a'):
links.append(link.get('href'))

print(links)

How does it work?

This line downloads the webpage data (which is surrounded by HTML tags):

req = Request("http://slashdot.org")
html_page = urlopen(req)

The next line loads it into a BeautifulSoup object:

soup = BeautifulSoup(html_page, "lxml")

The link codeblock will then get all links using .findAll(‘a’), where ‘a’ is the indicator for links in html.

links = []
for link in soup.findAll('a'):
links.append(link.get('href'))

Finally we show the list of links:

print(links)

Download network examples

Getting the links right

  • soup.findAll('a') returns every link on the page.
  • Filter on href, some anchors are used for jumping inside the page only.
  • Skip values that start with #, those are page anchors.
  • Links can start with //, that means http or https. Add the scheme yourself.

The result is a normal Python list, so you can sort it, remove duplicates with set() or loop over it.

Getting the links right

  • soup.findAll('a') returns every link on the page.
  • Filter on href, some anchors are only used to jump inside the page.
  • Skip the values that start with #, those are page anchors.
  • A value that starts with // means http or https. Add the scheme yourself.
  • urljoin(page_url, href) turns a relative link into a full address.

The result is a normal Python list, so you can sort it, remove the duplicates with set() or loop over it. Two passes are usually enough: one for the links, one for the pages behind them.

Add a User-Agent header when you download the page, a lot of sites send a 403 error to scripts that do not send one. And put a small delay between requests, otherwise you will get blocked.

Reading helps, writing fixes it. PyChallenge has exercises on this and you can try them right now.