← thecodex.expert · The Codex Family of Knowledge
Tier 2 · Intermediate · Kotlin Project

Web Scraper

Fetch a web page and extract specific information from its HTML. Learn to pull structured data out of the messy web.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).

Where this shows up: price monitoring, news aggregation, research data collection, search engines, market analysis. When a site offers no API, scraping is how you get its data — carefully and respectfully.

2 How to Think About It

Think about fetch-then-extract, before any code:

The plan — in plain English
1. Download the page’s HTML. → 2. Search the raw text for the pattern you want. → 3. Extract the piece you need from each match. → Always: check the site allows it and do not hammer it with requests.

Download page HTML

Parse the HTML

Find target elements

Extract text or links

Use the data

3 The Build — explained part by part

Here is the complete scraper. Read each part’s note below — you should understand the whole thing from the notes alone.

KotlinwebScraper.kt
import java.net.URI
import java.net.http.HttpClient
import java.net.http.HttpRequest
import java.net.http.HttpResponse

private val HREF = Regex("href=\"([^\"]+)\"")

/** Fetches [url] over real HTTP and returns the response body as text. */
fun fetch(client: HttpClient, url: String): String {
    val request = HttpRequest.newBuilder(URI.create(url)).GET().build()
    val response = client.send(request, HttpResponse.BodyHandlers.ofString())
    check(response.statusCode() == 200) { "Unexpected status: ${response.statusCode()}" }
    return response.body()
}

/** Every href="..." target found in [html], in the order they appear. */
fun extractLinks(html: String): List<String> =
    HREF.findAll(html).map { it.groupValues[1] }.toList()

fun main(args: Array<String>) {
    if (args.size != 1) {
        println("Usage: kotlin WebScraperKt <url>")
        return
    }
    val client = HttpClient.newHttpClient()
    val html = fetch(client, args[0])
    val links = extractLinks(html)
    println("Found ${links.size} links:")
    links.forEach { println("  $it") }
}
⚠ No in-browser playground here
Kotlin compiles to real JVM bytecode, not something a browser can run directly — running it live would need either a server-side compiler or a third-party embed, the same kind of external dependency this site avoids relying on for a core teaching example. Copy the code below and run it with a real kotlinc on your own machine instead; the “Run It” section explains exactly how.
What each part does — in plain words
HttpClient.newHttpClient() — the JDK's own built-in HTTP client (since Java 11, available to Kotlin with no extra dependency at all): no separate library to add for a feature this common.

HttpRequest.newBuilder(URI.create(url)).GET().build() — the client is built once; each request is its own small, immutable object describing exactly one call.

check(response.statusCode() == 200) { "..." } — Kotlin's check throws an IllegalStateException with the given message if the condition is false, a concise one-line alternative to a multi-line if (...) throw ....

val HREF = Regex("href=\"([^\"]+)\"") — a simple, pragmatic pattern, not a full HTML parser; HREF.findAll(html).map { it.groupValues[1] } walks every match and pulls out just the captured URL.
Common mistakes — and how to avoid them
✗ Assuming every HTTP response succeeded just because the request didn't throw.
✓ A 404 or 500 response is still a normal, non-throwing result from HttpClient.send; always check response.statusCode() before trusting the body.
✗ Reaching for a regex to parse arbitrary, deeply nested HTML in general.
✓ A regex is a reasonable, pragmatic choice for one narrow pattern like href="..." (as here), but a real crawler handling arbitrary pages should use a proper HTML parser — a regex cannot reliably handle nested tags or attributes split across lines.
✗ Forgetting that HttpClient.send throws on a connection failure (DNS, refused connection, timeout).
✓ These are a different failure mode from a non-200 status code; a real program should handle both, typically with a try/catch around the call and a status check right after it.

4 Test & Prove Each Part

How do we know this works? We pull the real logic into small, plain functions and check each one against cases we already know the answer to.

Every href="..." target is found, in order
A page with no links returns an empty list
Other attributes like src="..." are correctly ignored
KotlinwebScraperTest.kt
import kotlin.test.Test
import kotlin.test.assertEquals

class WebScraperTest {
    @Test
    fun findsEveryLink() {
        val html = """<a href="https://a.com">A</a> <a href="/b">B</a>"""
        assertEquals(listOf("https://a.com", "/b"), extractLinks(html))
    }

    @Test
    fun noLinksIsAnEmptyList() {
        assertEquals(emptyList(), extractLinks("<p>nothing here</p>"))
    }

    @Test
    fun ignoresOtherAttributes() {
        val html = """<img src="pic.png"> <a href="/ok">ok</a>"""
        assertEquals(listOf("/ok"), extractLinks(html))
    }
}

Compile with kotlinc webScraper.kt webScraperTest.kt -include-runtime -d webScraper.jar and run with JUnit's own runner. extractLinks takes a plain String of HTML, so the tests check known snippets directly — no real network call or running server involved.

5 The Interface

INPUTa URLthe page to scrape
What it expects
fetch("https://example.com")
OUTPUTextracted linksevery href on the page
What it returns
/one
/two

6 Run It & Automate It

Save the code as webScraper.kt and compile it with kotlinc webScraper.kt -include-runtime -d webScraper.jar. It takes the URL to fetch as a command-line argument.

Run it locally
kotlinc webScraper.kt -include-runtime -d webScraper.jar && java -jar webScraper.jar https://example.com
Fetches the given URL and prints every link found on the page.

A CI tool like Jenkins compiles and tests automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
Found 3 links:
  https://example.com
  /about
  /contact
If it breaks — how to fix it
🚨 Usage: kotlin WebScraperKt <url>
No URL was given on the command line; run it as java -jar webScraper.jar https://example.com, with the full URL including https://.
🚨 java.net.ConnectException or the program hangs for a long time
Check the URL is reachable from wherever this is running — a sandboxed or offline environment may block outbound requests entirely; try it against a server you control, as this page's own verification did.
GroovyJenkinsfile
// Jenkinsfile &mdash; compiles and tests automatically every time the code changes.
pipeline {
    agent any                                  // run on any available machine

    stages {
        stage('Get the code') {
            steps { checkout scm }             // download the latest code
        }
        stage('Set up Kotlin') {
            steps {
                sh 'kotlinc -version'                             // confirm the compiler is installed
            }
        }
        stage('Compile and test') {
            steps {
                sh 'kotlinc webScraper.kt webScraperTest.kt -include-runtime -d build.jar'  // one real JVM jar, no build tool required
                sh 'java -cp build.jar:kotlin-test-junit.jar:junit.jar org.junit.runner.JUnitCore WebScraperTest'
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed &mdash; look above.' }
    }
}
🎯 Try this next — make it yours

You have a working web scraper. Extend it:

  1. Follow the links. Fetch each discovered link too, up to some depth. (Teaches: recursion, or an explicit work queue.)
  2. Filter to one domain. Use java.net.URI to resolve relative links and ignore external ones. (Teaches: resolving a relative URL against a base.)
  3. Extract more than links. Pull out every <img src="..."> too, with a second pattern. (Teaches: generalising the extraction approach.)
  4. Respect robots.txt. Fetch and check it before scraping anything else. (Teaches: a real-world scraping courtesy, and one more HTTP call.)
What you learned
You learned the JDK’s built-in HttpClient (no extra dependency needed), Kotlin’s check for a concise invariant check, and a pragmatic regex-based alternative to a full HTML parser for simple link-scraping. Related reference: Exception Handling in Kotlin, Kotlin & Java Interop.