1 The Problem
We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).
2 How to Think About It
Think about fetch-then-extract, before any code:
3 The Build — explained part by part
Here is the complete scraper. It uses only Go’s standard library — net/http to fetch, regexp to find links — nothing to install. Each part is explained below.
package main
import (
"fmt"
"io"
"net/http"
"regexp"
)
// hrefPattern finds <a ...href="..."> opening tags. A real project would use
// a proper HTML parser (golang.org/x/net/html); this simple pattern keeps the
// project to the standard library only, same as the Python version's choice
// of html.parser over installing BeautifulSoup.
var hrefPattern = regexp.MustCompile(`(?i)<a\s+[^>]*?href\s*=\s*"([^"]*)"`)
// fetch downloads a page's HTML. A polite scraper identifies itself with a
// User-Agent.
func fetch(url string) (string, error) {
req, err := http.NewRequest("GET", url, nil)
if err != nil {
return "", err
}
req.Header.Set("User-Agent", "CodexBot/1.0")
resp, err := http.DefaultClient.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
body, err := io.ReadAll(resp.Body)
if err != nil {
return "", err
}
return string(body), nil
}
// extractLinks finds every <a href="..."> in a block of HTML.
func extractLinks(html string) []string {
matches := hrefPattern.FindAllStringSubmatch(html, -1)
links := make([]string, 0, len(matches))
for _, m := range matches {
links = append(links, m[1])
}
return links
}
func main() {
html, err := fetch("https://example.com")
if err != nil {
fmt.Println("Could not fetch the page:", err)
return
}
for _, link := range extractLinks(html) {
fmt.Println(link)
}
}
html.parser has no direct equivalent here), so this project uses a regular expression to find <a href="..."> tags instead. It is compiled once, at package level, so the pattern is not re-parsed on every call — a small but real Go performance habit.http.NewRequest + req.Header.Set — build the request by hand (rather than the shorter
http.Get) specifically so we can attach a User-Agent header, identifying the scraper politely — the same courtesy the Python version practices.http.DefaultClient.Do(req) — actually send the request and get a response back. defer resp.Body.Close() — guarantees the connection is released even if something below panics;
defer is one of Go’s signature features, and this is its most common use.io.ReadAll(resp.Body) — read the whole response body into memory as bytes, then convert to a
string.FindAllStringSubmatch(html, -1) — run the pattern against the whole page;
-1 means “find every match, not just the first”. Each match is a slice where index 0 is the whole tag and index 1 is the captured href value, which is what extractLinks collects.
defer resp.Body.Close() — each unclosed response leaks a connection, and a long-running scraper will eventually run out of sockets.defer the close immediately after checking the error, before doing anything else with the response.golang.org/x/net/html, a real (if not standard-library) HTML tokenizer.User-Agent header — many sites block requests that look like they come from no real browser or bot identity.4 Test & Prove Each Part
We test the extraction logic on a fixed HTML string — no network needed, fully predictable.
package main
import (
"reflect"
"testing"
)
func TestExtractsLink(t *testing.T) {
got := extractLinks(`<a href="/page">Link</a>`)
want := []string{"/page"}
if !reflect.DeepEqual(got, want) {
t.Errorf("extractLinks(...) = %v; want %v", got, want)
}
}
func TestNoLinks(t *testing.T) {
got := extractLinks(`<p>No links here</p>`)
if len(got) != 0 {
t.Errorf("extractLinks(...) = %v; want empty", got)
}
}
func TestMultipleLinks(t *testing.T) {
got := extractLinks(`<a href="/a">A</a><a href="/b">B</a>`)
want := []string{"/a", "/b"}
if !reflect.DeepEqual(got, want) {
t.Errorf("extractLinks(...) = %v; want %v", got, want)
}
}
Run with go test -v ./.... We test extraction on fixed HTML strings, never the live web — instant, reliable, and they pass offline. The network-fetching half (fetch) is separate on purpose, so it never gets in the way of testing the parsing logic.
5 The Interface
What it expects
fetch("https://example.com")What it returns
/one
/two6 Run It & Automate It
Save the code as scraper.go and run it with go run scraper.go — Go compiles and executes in one step, no separate build needed while you are experimenting.
go run scraper.goFetches the URL in
main and lists its links. Scrape responsibly: check a site’s robots.txt and terms, and do not send rapid repeated requests.A CI tool like Jenkins runs go test automatically whenever the code changes — every line below has a plain explanation.
/one
/twogo get golang.org/x/net/html first to add it to your module.// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up Go') {
steps {
sh 'go version' // confirm Go is installed
sh 'test -f go.mod || go mod init web_scraper' // create a module if none exists
}
}
stage('Run the tests') {
steps {
sh 'go vet ./...' // catch obvious mistakes before running
sh 'go test -v ./...' // run every test, show each result
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
- Extract other things. Write a second regex for headings or image sources. (Teaches: targeting more tags.)
- Use a real parser. Swap the regex for
golang.org/x/net/htmlfor robust extraction. (Teaches: adding a dependency withgo get.) - Follow links. Scrape linked pages too, politely, with a delay between requests. (Teaches: crawling.)
net/http, identify your scraper politely, and extract data from raw HTML with regexp — plus defer for reliable cleanup, one of Go’s most-used features. Related: Standard Library, Error Handling & the Standard Library.