We use the module urllib2 to download webpage data. Any webpage is formatted using a markup language known as HTML.

Scripting is where Python really pays off. The Python Scripting Essentials Bundle has the kind of small automation projects you would build for work.

Every picture on a webpage has an address behind it. The img tag holds that address in its src attribute, so if you can read the tag you know the image url. That is all you need to build your own gallery, save the files to a folder or check which images a site uses.

The example below scrapes a page and prints every image url it finds. BeautifulSoup does the reading of the html for you, so there is no regex to write.

Extracting image links:
To extract all image links use:

from BeautifulSoup import BeautifulSoup
import urllib2
import re

html_page = urllib2.urlopen("http://imgur.com")
soup = BeautifulSoup(html_page)
images = []
for img in soup.findAll('img'):
images.append(img.get('src'))

print(images)

Explanation
First we import the required modules:

from BeautifulSoup import BeautifulSoup
import urllib2
import re
`

We get the webpage data using:

html_page = urllib2.urlopen("http://imgur.com")

Then we extract all image links using:

images = []
for img in soup.findAll('img'):
images.append(img.get('src'))

Finally we print the links:

print(links)

How it works

  • urllib.request.urlopen(url) downloads the html of the page.
  • BeautifulSoup(html_page) turns that html into something you can search.
  • soup.findAll('img') gives you every image tag on the page.
  • img.get('src') returns the url of that image.

Relative image urls are common, a page may say /img/logo.png instead of the full address. You can join them with the url of the page using urllib.parse.urljoin(page_url, img_url).

Always add a User-Agent header when you download a page. A lot of sites send a 403 error to scripts that do not send one.

If you want to try this yourself, the exercises on PyChallenge turn the same idea into a few short problems with instant feedback.