1 The Problem
We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).
2 How to Think About It
Strip away the crate that usually hides this: an HTTP GET is a socket connection, a few lines of text sent, and a response read back.
href="..." occurrences.
3 The Build — explained part by part
Here is the complete scraper. This is the one project on this site where Rust’s sandbox constraint (no crates.io) turns into a genuine teaching opportunity rather than just a workaround: writing an HTTP client by hand shows you exactly what reqwest normally does for you.
use std::env;
use std::io::{Read, Write};
use std::net::TcpStream;
/// A minimal, hand-written HTTP/1.1 GET client over a raw `TcpStream`.
/// Idiomatic real-world Rust reaches for the `reqwest` crate here, but this
/// build environment cannot fetch crates.io, so this project does the thing
/// `reqwest` normally hides: open a socket, write the request line and
/// headers ourselves, and read the raw response back. It is more code, but
/// it is also a genuinely good look at what an HTTP client actually is
/// underneath — very much in the spirit of Rust's systems-programming roots.
fn fetch(host: &str, port: u16, path: &str) -> std::io::Result<String> {
let mut stream = TcpStream::connect((host, port))?;
let request = format!(
"GET {path} HTTP/1.1\r\nHost: {host}\r\nUser-Agent: codex-scraper/1.0\r\nConnection: close\r\n\r\n"
);
stream.write_all(request.as_bytes())?;
let mut raw = String::new();
stream.read_to_string(&mut raw)?;
Ok(raw)
}
/// Splits a raw HTTP/1.1 response into (status_line, headers, body). The
/// blank line (`\r\n\r\n`) is the boundary HTTP itself defines between
/// headers and body.
fn split_response(raw: &str) -> (&str, &str) {
match raw.split_once("\r\n\r\n") {
Some((head, body)) => (head, body),
None => (raw, ""),
}
}
/// Rust's standard library has no regex engine (that is the `regex` crate,
/// also unreachable here), so link extraction is a small hand-written
/// scanner: find every `href="..."` occurrence and pull out the quoted
/// text. It is less flexible than a real regex, but it is dependency-free,
/// entirely readable, and correct for well-formed HTML.
fn extract_links(html: &str) -> Vec<String> {
let mut links = Vec::new();
let mut rest = html;
while let Some(start) = rest.find("href=\"") {
rest = &rest[start + "href=\"".len()..];
if let Some(end) = rest.find('"') {
links.push(rest[..end].to_string());
rest = &rest[end + 1..];
} else {
break;
}
}
links
}
fn main() {
let url = env::args().nth(1).unwrap_or_else(|| "http://127.0.0.1:8080/".to_string());
let (host, port, path) = match parse_url(&url) {
Some(parts) => parts,
None => {
eprintln!("Could not parse URL: {url} (expected http://host[:port]/path)");
return;
}
};
match fetch(&host, port, &path) {
Ok(raw) => {
let (head, body) = split_response(&raw);
println!("{}", head.lines().next().unwrap_or(""));
let links = extract_links(body);
println!("Found {} link(s):", links.len());
for link in &links {
println!(" {link}");
}
}
Err(e) => eprintln!("Request failed: {e}"),
}
}
/// A tiny hand-written URL splitter — just enough for `http://host:port/path`.
/// A real project would use the `url` crate for this.
fn parse_url(url: &str) -> Option<(String, u16, String)> {
let rest = url.strip_prefix("http://")?;
let (authority, path) = match rest.find('/') {
Some(i) => (&rest[..i], &rest[i..]),
None => (rest, "/"),
};
let (host, port) = match authority.split_once(':') {
Some((h, p)) => (h.to_string(), p.parse().ok()?),
None => (authority.to_string(), 80),
};
Some((host, port, path.to_string()))
}
#[cfg(test)]
mod tests {
use super::*;
use std::io::BufReader;
use std::net::TcpListener;
use std::thread;
#[test]
fn extracts_links_from_html() {
let html = r#"<a href="/about">About</a><a href="https://example.com">Ex</a>"#;
let links = extract_links(html);
assert_eq!(links, vec!["/about", "https://example.com"]);
}
#[test]
fn returns_no_links_for_plain_text() {
assert!(extract_links("just some text, no tags here").is_empty());
}
#[test]
fn splits_headers_from_body_on_the_blank_line() {
let raw = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\n\r\n<html>hi</html>";
let (head, body) = split_response(raw);
assert!(head.starts_with("HTTP/1.1 200 OK"));
assert_eq!(body, "<html>hi</html>");
}
#[test]
fn parses_a_simple_url() {
let (host, port, path) = parse_url("http://example.com:9090/links").unwrap();
assert_eq!(host, "example.com");
assert_eq!(port, 9090);
assert_eq!(path, "/links");
}
/// End-to-end: starts a real local TCP server on an OS-assigned port,
/// serves one canned HTML response, and confirms `fetch` + `extract_links`
/// pull the right links out of an actual network round trip, not just a
/// hand-fed string.
#[test]
fn fetches_and_extracts_links_from_a_real_local_server() {
let listener = TcpListener::bind("127.0.0.1:0").unwrap();
let port = listener.local_addr().unwrap().port();
let handle = thread::spawn(move || {
let (stream, _) = listener.accept().unwrap();
let mut reader = BufReader::new(stream.try_clone().unwrap());
let mut request_line = String::new();
std::io::BufRead::read_line(&mut reader, &mut request_line).unwrap();
let body = r#"<html><body><a href="/one">One</a><a href="/two">Two</a></body></html>"#;
let response = format!(
"HTTP/1.1 200 OK\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{}",
body.len(),
body
);
let mut stream = stream;
stream.write_all(response.as_bytes()).unwrap();
});
let raw = fetch("127.0.0.1", port, "/").unwrap();
handle.join().unwrap();
let (_, body) = split_response(&raw);
let links = extract_links(body);
assert_eq!(links, vec!["/one", "/two"]);
}
}
rustup) is installed.format!("GET {path} HTTP/1.1\r\nHost: {host}\r\n...\r\n\r\n") — a valid HTTP/1.1 request is a plain text protocol: a request line, headers each ending in
\r\n, and a blank line marking the end of headers. Writing it out by hand is the entire “request” a crate like reqwest builds for you.raw.split_once("\r\n\r\n") — that same blank line is how you find where headers end and the body begins in the response — HTTP defines this boundary explicitly, so no guessing is involved.
extract_links — Rust’s standard library ships no regex engine, so this is a hand-written scanner: repeatedly find
href=", then find the closing quote, and slice out what is between them. Less flexible than a real regex, but dependency-free and fully readable.the local-server test — rather than only testing
extract_links on a hand-written string, one test starts a real TcpListener on an OS-assigned port, serves one canned response, and confirms fetch genuinely round-trips over a real socket — the same rigor Go’s version of this project used.
Connection: close in the request headers — without it, a real server may keep the connection open waiting for another request, and read_to_string would then block forever waiting for the stream to end.Connection: close for a one-shot client like this one, as the code above does.href="..." only ever uses double quotes — real-world HTML sometimes uses single quotes.regex/scraper crates, if reachable) would handle both; this project’s scanner is deliberately simple and documented as such.4 Test & Prove Each Part
We test link extraction on known strings, the header/body split, the URL parser, and — the important one — a real fetch against a real local server.
#[cfg(test)]
mod tests {
use super::*;
use std::io::BufReader;
use std::net::TcpListener;
use std::thread;
#[test]
fn extracts_links_from_html() {
let html = r#"<a href="/about">About</a><a href="https://example.com">Ex</a>"#;
let links = extract_links(html);
assert_eq!(links, vec!["/about", "https://example.com"]);
}
#[test]
fn returns_no_links_for_plain_text() {
assert!(extract_links("just some text, no tags here").is_empty());
}
#[test]
fn splits_headers_from_body_on_the_blank_line() {
let raw = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\n\r\n<html>hi</html>";
let (head, body) = split_response(raw);
assert!(head.starts_with("HTTP/1.1 200 OK"));
assert_eq!(body, "<html>hi</html>");
}
#[test]
fn parses_a_simple_url() {
let (host, port, path) = parse_url("http://example.com:9090/links").unwrap();
assert_eq!(host, "example.com");
assert_eq!(port, 9090);
assert_eq!(path, "/links");
}
/// End-to-end: starts a real local TCP server on an OS-assigned port,
/// serves one canned HTML response, and confirms `fetch` + `extract_links`
/// pull the right links out of an actual network round trip, not just a
/// hand-fed string.
#[test]
fn fetches_and_extracts_links_from_a_real_local_server() {
let listener = TcpListener::bind("127.0.0.1:0").unwrap();
let port = listener.local_addr().unwrap().port();
let handle = thread::spawn(move || {
let (stream, _) = listener.accept().unwrap();
let mut reader = BufReader::new(stream.try_clone().unwrap());
let mut request_line = String::new();
std::io::BufRead::read_line(&mut reader, &mut request_line).unwrap();
let body = r#"<html><body><a href="/one">One</a><a href="/two">Two</a></body></html>"#;
let response = format!(
"HTTP/1.1 200 OK\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{}",
body.len(),
body
);
let mut stream = stream;
stream.write_all(response.as_bytes()).unwrap();
});
let raw = fetch("127.0.0.1", port, "/").unwrap();
handle.join().unwrap();
let (_, body) = split_response(&raw);
let links = extract_links(body);
assert_eq!(links, vec!["/one", "/two"]);
}
}
Run with cargo test. The last test is the one worth reading closely: it spawns a thread running a real TcpListener, serves one hand-built HTTP response, and lets fetch connect to it exactly as it would to a real website — proving the socket code works over an actual network round trip, not just against a string.
5 The Interface
What it expects
http://127.0.0.1:8099/index.htmlWhat it returns
HTTP/1.0 200 OK
Found 2 link(s):
/about
https://example.com6 Run It & Automate It
Save the code as src/main.rs inside a Cargo project's src/ folder and run it with cargo run — Cargo compiles and executes in one step while you are experimenting, then cargo build --release gives you an optimized binary once you are done.
cargo run -- http://127.0.0.1:8099/index.htmlPoint it at any plain HTTP (not HTTPS — this client has no TLS) server, including one you started locally with
python3 -m http.server.A CI tool like Jenkins runs cargo test automatically whenever the code changes — every line below has a plain explanation.
$ python3 -m http.server 8099 &
$ cargo run -- http://127.0.0.1:8099/index.html
HTTP/1.0 200 OK
Found 2 link(s):
/about
https://example.comhttp:// URLs. An https:// URL needs TLS, which this hand-rolled client deliberately does not implement.Connection: close — without it, read_to_string can block forever waiting for a server that keeps the connection open.// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up Rust') {
steps {
sh 'rustc --version' // confirm Rust is installed
sh 'cargo build' // compile, downloading any crates
}
}
stage('Run the tests') {
steps {
sh 'cargo clippy -- -D warnings' // catch obvious mistakes before running
sh 'cargo test' // run every test, show each result
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
- Follow redirects. Check for a 3xx status and a
Locationheader, then fetch again. (Teaches: reading response headers, not just the body.) - Use the real
reqwestcrate. If you have network access, compare how much of this code disappears. (Teaches: what a well-designed dependency buys you.) - Add a timeout. Use
TcpStream::set_read_timeoutso a hung server cannot block forever. (Teaches: socket timeouts.)