← thecodex.expert · The Codex Family of Knowledge
Tier 3 · Upper-Intermediate · Java Project

Web Scraper

Fetch a real page over HTTP with java.net.http and extract every link from its HTML. Teaches HttpClient, a hand-written link extractor, and testing against a real local server rather than the live internet.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).

Where this shows up: price monitoring, news aggregation, research data collection, search engines, market analysis. When a site offers no API, scraping is how you get its data — carefully and respectfully.

2 How to Think About It

Two separate jobs, kept in two separate methods: getting the raw HTML over the network, and pulling links out of text you already have. Only the first one needs a network at all.

The plan — in plain English
1. Fetch the page with HttpClient. → 2. Check the status code before trusting the body. → 3. Extract every href with a regular expression. → 4. Print what was found.

Download page HTML

Parse the HTML

Find target elements

Extract text or links

Use the data

3 The Build — explained part by part

Here is the complete scraper. java.net.http, standard since Java 11, needs no external HTTP library at all — a genuine advantage over languages whose standard library ships no HTTP client.

JavaWebScraper.java
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

/**
 * Web Scraper: fetches a page over real HTTP with java.net.http and extracts
 * every link from its HTML.
 */
public class WebScraper {

    private static final Pattern HREF = Pattern.compile("href=\"([^\"]+)\"");

    static String fetch(HttpClient client, String url) throws Exception {
        HttpRequest request = HttpRequest.newBuilder(URI.create(url)).GET().build();
        HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
        if (response.statusCode() != 200) {
            throw new RuntimeException("Unexpected status: " + response.statusCode());
        }
        return response.body();
    }

    static List<String> extractLinks(String html) {
        List<String> links = new ArrayList<>();
        Matcher m = HREF.matcher(html);
        while (m.find()) {
            links.add(m.group(1));
        }
        return links;
    }

    public static void main(String[] args) throws Exception {
        if (args.length != 1) {
            System.out.println("Usage: java WebScraper <url>");
            return;
        }
        HttpClient client = HttpClient.newHttpClient();
        String html = fetch(client, args[0]);
        List<String> links = extractLinks(html);
        System.out.println("Found " + links.size() + " links:");
        links.forEach(link -> System.out.println("  " + link));
    }
}
⚠ No in-browser playground here
Java compiles to JVM bytecode and needs a real JDK to run, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once a JDK is installed.
What each part does — in plain words
HttpClient.newHttpClient() and HttpRequest.newBuilder(...).GET().build() — java.net.http, part of the standard library since Java 11, so this project needs no third-party HTTP library at all, unlike languages whose standard library ships no HTTP client of its own.

response.statusCode() != 200 checked before returning the body — a 404 or 500 response still arrives as a normal, successful network exchange with no exception thrown; checking the status code explicitly is the only way to know the page actually loaded.

Pattern.compile("href=\"([^\"]+)\"") — a hand-written regular expression rather than a full HTML parser, which is a reasonable, honest trade for a learning project: it works for simple well-formed pages but is not a substitute for a real parser like Jsoup on messy real-world HTML.

fetch and extractLinks are separate, independently testable methods — the test suite proves the extraction logic works on plain strings with zero network involved, and separately proves the network call works against a real local server, rather than one large method doing both at once.
Common mistakes — and how to avoid them
✗ Assuming any response means success — a 404 page still returns a body, just not the one you wanted.
✓ Always check statusCode() before trusting body(), as fetch does.
✗ Testing against a real website over the live internet — the test becomes flaky, slow, and breaks the moment that page’s content changes.
✓ Spin up a real local HttpServer inside the test itself, as fetchesAndExtractsLinksFromARealLocalServer does — still a genuine network round trip, just not a dependency on the outside world.

4 Test & Prove Each Part

We test link extraction on plain strings, and separately prove the network call works, against a real server this test starts and stops itself.

Extracting links from simple HTML returns them in order
HTML with no links returns an empty list, not null
A real local HTTP server is fetched and its links extracted correctly
A non-200 response throws instead of returning a misleading empty result
JavaWebScraperTest.java
import com.sun.net.httpserver.HttpServer;
import org.junit.Test;
import java.net.InetSocketAddress;
import java.net.http.HttpClient;
import java.util.List;
import static org.junit.Assert.assertEquals;
import static org.junit.Assert.assertTrue;

public class WebScraperTest {

    @Test
    public void extractsHrefsFromHtml() {
        String html = "<a href=\"/one\">One</a><a href=\"/two\">Two</a>";
        List<String> links = WebScraper.extractLinks(html);
        assertEquals(List.of("/one", "/two"), links);
    }

    @Test
    public void noLinksMeansEmptyListNotNull() {
        assertTrue(WebScraper.extractLinks("<p>no links here</p>").isEmpty());
    }

    @Test
    public void fetchesAndExtractsLinksFromARealLocalServer() throws Exception {
        HttpServer server = HttpServer.create(new InetSocketAddress("localhost", 0), 0);
        server.createContext("/", exchange -> {
            byte[] body = "<a href=\"/about\">About</a><a href=\"/contact\">Contact</a>"
                    .getBytes();
            exchange.sendResponseHeaders(200, body.length);
            exchange.getResponseBody().write(body);
            exchange.close();
        });
        server.start();
        try {
            int port = server.getAddress().getPort();
            HttpClient client = HttpClient.newHttpClient();
            String html = WebScraper.fetch(client, "http://localhost:" + port + "/");
            List<String> links = WebScraper.extractLinks(html);
            assertEquals(List.of("/about", "/contact"), links);
        } finally {
            server.stop(0);
        }
    }

    @Test(expected = RuntimeException.class)
    public void nonOkStatusThrows() throws Exception {
        HttpServer server = HttpServer.create(new InetSocketAddress("localhost", 0), 0);
        server.createContext("/", exchange -> {
            exchange.sendResponseHeaders(404, -1);
            exchange.close();
        });
        server.start();
        try {
            int port = server.getAddress().getPort();
            HttpClient client = HttpClient.newHttpClient();
            WebScraper.fetch(client, "http://localhost:" + port + "/");
        } finally {
            server.stop(0);
        }
    }
}

Compile and run with javac -cp junit-4.13.2.jar and hamcrest-core-1.3.jar WebScraper.java WebScraperTest.java then java -cp .:junit-4.13.2.jar:hamcrest-core-1.3.jar org.junit.runner.JUnitCore WebScraperTest. The last two tests use com.sun.net.httpserver.HttpServer — the same JDK-bundled server class the REST API project builds on — to start a real, ephemeral-port server inside the test itself, so the network call is genuine but the test never depends on any address outside localhost.

5 The Interface

INPUTINPUTa URL
What it expects
java WebScraper http://localhost:8099/index.html
OUTPUTOUTPUTevery link found on the page
What it returns
Found 3 links:
  /about
  /contact
  https://example.com

6 Run It & Automate It

Save the code as WebScraper.java and compile it with javac — that turns your source into .class bytecode files, which java then runs on the JVM. No separate install step: any real JDK ships both tools.

Run it locally
javac WebScraper.java && java WebScraper http://localhost:8099/index.html
Point it at any page you are allowed to fetch, local or on the real internet.

A CI tool like Jenkins runs the same compile-then-test steps automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ python3 -m http.server 8099 &
$ java WebScraper http://localhost:8099/index.html
Found 3 links:
  /about
  /contact
  https://example.com
If it breaks — how to fix it
🚨 java.net.ConnectException: Connection refused
Nothing is listening at that host and port. If you are testing locally, make sure a server is actually running there first.
🚨 java.lang.RuntimeException: Unexpected status: 404
This is fetch working as designed — the URL you gave it does not exist on that server. Double-check the path.
GroovyJenkinsfile
// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
    agent any                          // run on any available machine
    environment {
        CP = 'junit-4.13.2.jar:hamcrest-core-1.3.jar'   // JUnit + its one dependency
    }

    stages {
        stage('Get the code') {
            steps { checkout scm }     // download the latest code
        }
        stage('Set up JDK') {
            steps {
                sh 'java -version'           // confirm a JDK is installed
                sh 'javac -cp "$CP" *.java'   // compile the program and its tests together
            }
        }
        stage('Run the tests') {
            steps {
                sh 'java -cp ".:$CP" org.junit.runner.JUnitCore WebScraperTest'
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed — look above.' }
    }
}
🎯 Try this next — make it yours
  1. Follow links one level deep. Fetch every extracted link and report how many are reachable. (Teaches: recursive or iterative crawling, and cycle avoidance.)
  2. Extract more than links. Pull out <title> and <img src> too. (Teaches: generalizing a regex-based extractor, and where it starts to strain.)
  3. Add a real HTML parser. If you have network access, swap the regex for Jsoup. (Teaches: what a purpose-built parser handles that a regex cannot — nested tags, malformed HTML.)
What you learned
You learned java.net.http’s HttpClient/HttpRequest/HttpResponse trio built into the standard library, why a status code must be checked explicitly, and how to test real network code against a real local server instead of the live internet. Related: Standard Library, IO and NIO.