← thecodex.expert · The Codex Family of Knowledge
Tier 2 · Intermediate · Go Project

Web Scraper

Fetch a web page and extract specific information from its HTML. Learn to pull structured data out of the messy web.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).

Where this shows up: price monitoring, news aggregation, research data collection, search engines, market analysis. When a site offers no API, scraping is how you get its data — carefully and respectfully.

2 How to Think About It

Think about fetch-then-extract, before any code:

The plan — in plain English
1. Download the page’s HTML. → 2. Search the raw text for the pattern you want. → 3. Extract the piece you need from each match. → Always: check the site allows it and do not hammer it with requests.

Download page HTML

Parse the HTML

Find target elements

Extract text or links

Use the data

3 The Build — explained part by part

Here is the complete scraper. It uses only Go’s standard library — net/http to fetch, regexp to find links — nothing to install. Each part is explained below.

Goscraper.go
package main

import (
	"fmt"
	"io"
	"net/http"
	"regexp"
)

// hrefPattern finds <a ...href="...">  opening tags. A real project would use
// a proper HTML parser (golang.org/x/net/html); this simple pattern keeps the
// project to the standard library only, same as the Python version's choice
// of html.parser over installing BeautifulSoup.
var hrefPattern = regexp.MustCompile(`(?i)<a\s+[^>]*?href\s*=\s*"([^"]*)"`)

// fetch downloads a page's HTML. A polite scraper identifies itself with a
// User-Agent.
func fetch(url string) (string, error) {
	req, err := http.NewRequest("GET", url, nil)
	if err != nil {
		return "", err
	}
	req.Header.Set("User-Agent", "CodexBot/1.0")

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		return "", err
	}
	defer resp.Body.Close()

	body, err := io.ReadAll(resp.Body)
	if err != nil {
		return "", err
	}
	return string(body), nil
}

// extractLinks finds every <a href="..."> in a block of HTML.
func extractLinks(html string) []string {
	matches := hrefPattern.FindAllStringSubmatch(html, -1)
	links := make([]string, 0, len(matches))
	for _, m := range matches {
		links = append(links, m[1])
	}
	return links
}

func main() {
	html, err := fetch("https://example.com")
	if err != nil {
		fmt.Println("Could not fetch the page:", err)
		return
	}
	for _, link := range extractLinks(html) {
		fmt.Println(link)
	}
}
⚠ No in-browser playground here
Go compiles to a real binary, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once Go is installed.
What each part does — in plain words
regexp.MustCompile(...) — Go has no HTML parser in its standard library (Python’s html.parser has no direct equivalent here), so this project uses a regular expression to find <a href="..."> tags instead. It is compiled once, at package level, so the pattern is not re-parsed on every call — a small but real Go performance habit.

http.NewRequest + req.Header.Set — build the request by hand (rather than the shorter http.Get) specifically so we can attach a User-Agent header, identifying the scraper politely — the same courtesy the Python version practices.

http.DefaultClient.Do(req) — actually send the request and get a response back. defer resp.Body.Close() — guarantees the connection is released even if something below panics; defer is one of Go’s signature features, and this is its most common use.

io.ReadAll(resp.Body) — read the whole response body into memory as bytes, then convert to a string.

FindAllStringSubmatch(html, -1) — run the pattern against the whole page; -1 means “find every match, not just the first”. Each match is a slice where index 0 is the whole tag and index 1 is the captured href value, which is what extractLinks collects.
Common mistakes — and how to avoid them
✗ Forgetting defer resp.Body.Close() — each unclosed response leaks a connection, and a long-running scraper will eventually run out of sockets.
✓ Always defer the close immediately after checking the error, before doing anything else with the response.
✗ Treating a regex HTML match as a real HTML parser — it will misfire on nested quotes, commented-out tags, or minified HTML with unusual spacing.
✓ For anything beyond a learning project, use golang.org/x/net/html, a real (if not standard-library) HTML tokenizer.
✗ Skipping the User-Agent header — many sites block requests that look like they come from no real browser or bot identity.
✓ Always identify your scraper, as the code above does.

4 Test & Prove Each Part

We test the extraction logic on a fixed HTML string — no network needed, fully predictable.

Links are extracted from anchor tags
A page with no links returns an empty result
Multiple links are all collected, in order
Goscraper_test.go
package main

import (
	"reflect"
	"testing"
)

func TestExtractsLink(t *testing.T) {
	got := extractLinks(`<a href="/page">Link</a>`)
	want := []string{"/page"}
	if !reflect.DeepEqual(got, want) {
		t.Errorf("extractLinks(...) = %v; want %v", got, want)
	}
}

func TestNoLinks(t *testing.T) {
	got := extractLinks(`<p>No links here</p>`)
	if len(got) != 0 {
		t.Errorf("extractLinks(...) = %v; want empty", got)
	}
}

func TestMultipleLinks(t *testing.T) {
	got := extractLinks(`<a href="/a">A</a><a href="/b">B</a>`)
	want := []string{"/a", "/b"}
	if !reflect.DeepEqual(got, want) {
		t.Errorf("extractLinks(...) = %v; want %v", got, want)
	}
}

Run with go test -v ./.... We test extraction on fixed HTML strings, never the live web — instant, reliable, and they pass offline. The network-fetching half (fetch) is separate on purpose, so it never gets in the way of testing the parsing logic.

5 The Interface

INPUTa URLthe page to scrape
What it expects
fetch("https://example.com")
OUTPUTextracted linksevery href on the page
What it returns
/one
/two

6 Run It & Automate It

Save the code as scraper.go and run it with go run scraper.go — Go compiles and executes in one step, no separate build needed while you are experimenting.

Run it locally
go run scraper.go
Fetches the URL in main and lists its links. Scrape responsibly: check a site’s robots.txt and terms, and do not send rapid repeated requests.

A CI tool like Jenkins runs go test automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
/one
/two
If it breaks — how to fix it
🚨 Get "https://...": dial tcp: ... no such host
Check the URL is correct and reachable, and that you have network access from where the program runs.
🚨 No links found on a real page.
The page may load its content with JavaScript, which this simple fetch-and-regex approach never runs. That needs a headless browser, a much heavier tool.
🚨 golang.org/x/net/html: no such file or directory (if you try the "try this next" upgrade)
That package is not in the standard library. Run go get golang.org/x/net/html first to add it to your module.
GroovyJenkinsfile
// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
    agent any                                  // run on any available machine

    stages {
        stage('Get the code') {
            steps { checkout scm }             // download the latest code
        }
        stage('Set up Go') {
            steps {
                sh 'go version'                                 // confirm Go is installed
                sh 'test -f go.mod || go mod init web_scraper'  // create a module if none exists
            }
        }
        stage('Run the tests') {
            steps {
                sh 'go vet ./...'                    // catch obvious mistakes before running
                sh 'go test -v ./...'                // run every test, show each result
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed — look above.' }
    }
}
🎯 Try this next — make it yours
  1. Extract other things. Write a second regex for headings or image sources. (Teaches: targeting more tags.)
  2. Use a real parser. Swap the regex for golang.org/x/net/html for robust extraction. (Teaches: adding a dependency with go get.)
  3. Follow links. Scrape linked pages too, politely, with a delay between requests. (Teaches: crawling.)
What you learned
You learned to fetch a web page with net/http, identify your scraper politely, and extract data from raw HTML with regexp — plus defer for reliable cleanup, one of Go’s most-used features. Related: Standard Library, Error Handling & the Standard Library.