1 The Problem
We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).
2 How to Think About It
Strip away the library that usually hides this: an HTTP GET is a socket connection, a few lines of text sent, and a response read back. The two design questions worth pausing on are who closes the socket, and how links get pulled out of the HTML.
getaddrinfo and connect a TCP socket to it, wrapped in an RAII Socket. → 2. Write a valid HTTP/1.1 request line and headers, ending in a blank line. → 3. Read the raw response bytes back with recv. → 4. Split it into a status line and body on the blank line HTTP itself defines. → 5. Match every href="..." in the body with std::regex.
3 The Build — explained part by part
Here is the complete scraper. C++ has no standard-library HTTP client either (unlike Go or Java) so writing one by hand over BSD sockets is still what an HTTP client looks like underneath any library — but two things below are genuinely nicer than the C version.
#pragma once
#include <optional>
#include <string>
#include <vector>
struct Url {
std::string host;
int port;
std::string path;
};
// A tiny hand-written URL splitter -- just enough for
// "http://host[:port]/path". Returns std::nullopt on a malformed URL.
std::optional<Url> parse_url(const std::string &url);
// A small RAII wrapper around a raw BSD socket file descriptor. The
// destructor closes the socket automatically -- there is no equivalent in
// C, which must remember to call close() on every exit path by hand,
// including every early "return -1 on failure" branch. Move-only, since a
// socket handle should never be silently duplicated and closed twice.
class Socket {
public:
Socket();
explicit Socket(int fd);
~Socket();
Socket(const Socket &) = delete;
Socket &operator=(const Socket &) = delete;
Socket(Socket &&other) noexcept;
Socket &operator=(Socket &&other) noexcept;
int fd() const { return fd_; }
bool valid() const { return fd_ >= 0; }
private:
int fd_ = -1;
};
// A minimal, hand-written HTTP/1.1 GET client over a raw socket. C++ has no
// standard-library HTTP client either (unlike Go's net/http or Java's
// java.net.http), so this project does by hand what any HTTP library does
// underneath: open a socket, write the request line and headers ourselves,
// and read the raw response back. Returns std::nullopt on failure.
std::optional<std::string> fetch(const std::string &host, int port, const std::string &path);
struct SplitResponse {
std::string status_line;
std::string body;
};
// Splits a raw HTTP/1.1 response into a status line and a body, at the
// blank line ("\r\n\r\n") HTTP itself defines as the boundary. Returns
// std::nullopt if no such boundary was found.
std::optional<SplitResponse> split_response(const std::string &raw);
// Unlike C, C++'s standard library does ship a real regex engine
// (<regex>), so link extraction can be a genuine pattern match for every
// href="..." occurrence instead of a hand-written character scanner.
std::vector<std::string> extract_links(const std::string &html);
#define _POSIX_C_SOURCE 200809L
#include "WebScraper.hpp"
#include <cstring>
#include <netdb.h>
#include <regex>
#include <sstream>
#include <sys/socket.h>
#include <unistd.h>
#include <utility>
// ---------- Socket ----------
Socket::Socket() : fd_(-1) {}
Socket::Socket(int fd) : fd_(fd) {}
Socket::~Socket() {
if (fd_ >= 0) close(fd_);
}
Socket::Socket(Socket &&other) noexcept : fd_(other.fd_) {
other.fd_ = -1;
}
Socket &Socket::operator=(Socket &&other) noexcept {
if (this != &other) {
if (fd_ >= 0) close(fd_);
fd_ = other.fd_;
other.fd_ = -1;
}
return *this;
}
// ---------- parse_url ----------
std::optional<Url> parse_url(const std::string &url) {
std::string rest = url;
const std::string prefix = "http://";
if (rest.rfind(prefix, 0) != 0) return std::nullopt;
rest = rest.substr(prefix.size());
auto slash = rest.find('/');
std::string host_port = slash == std::string::npos ? rest : rest.substr(0, slash);
std::string path = slash == std::string::npos ? "/" : rest.substr(slash);
if (host_port.empty()) return std::nullopt;
int port = 80;
std::string host = host_port;
auto colon = host_port.find(':');
if (colon != std::string::npos) {
host = host_port.substr(0, colon);
try {
port = std::stoi(host_port.substr(colon + 1));
} catch (const std::exception &) {
return std::nullopt;
}
}
if (host.empty()) return std::nullopt;
return Url{host, port, path};
}
// ---------- fetch ----------
std::optional<std::string> fetch(const std::string &host, int port, const std::string &path) {
addrinfo hints{};
hints.ai_family = AF_INET;
hints.ai_socktype = SOCK_STREAM;
addrinfo *result = nullptr;
std::string port_str = std::to_string(port);
if (getaddrinfo(host.c_str(), port_str.c_str(), &hints, &result) != 0) {
return std::nullopt;
}
Socket sock(socket(result->ai_family, result->ai_socktype, result->ai_protocol));
if (!sock.valid()) {
freeaddrinfo(result);
return std::nullopt;
}
if (connect(sock.fd(), result->ai_addr, result->ai_addrlen) < 0) {
freeaddrinfo(result);
return std::nullopt;
}
freeaddrinfo(result);
std::ostringstream request;
request << "GET " << path << " HTTP/1.1\r\n"
<< "Host: " << host << "\r\n"
<< "Connection: close\r\n"
<< "\r\n";
std::string req = request.str();
if (send(sock.fd(), req.c_str(), req.size(), 0) < 0) return std::nullopt;
std::string response;
char buf[4096];
ssize_t n;
while ((n = recv(sock.fd(), buf, sizeof(buf), 0)) > 0) {
response.append(buf, static_cast<size_t>(n));
}
return response;
}
// ---------- split_response ----------
std::optional<SplitResponse> split_response(const std::string &raw) {
auto boundary = raw.find("\r\n\r\n");
if (boundary == std::string::npos) return std::nullopt;
auto line_end = raw.find("\r\n");
std::string status_line = raw.substr(0, line_end);
std::string body = raw.substr(boundary + 4);
return SplitResponse{status_line, body};
}
// ---------- extract_links ----------
std::vector<std::string> extract_links(const std::string &html) {
static const std::regex href_re(R"re(href="([^"]*)")re");
std::vector<std::string> links;
auto begin = std::sregex_iterator(html.begin(), html.end(), href_re);
auto end = std::sregex_iterator();
for (auto it = begin; it != end; ++it) {
links.push_back((*it)[1].str());
}
return links;
}
#include "WebScraper.hpp"
#include <iostream>
int main(int argc, char **argv) {
if (argc < 2) {
std::cerr << "Usage: web_scraper <url>\n";
return 1;
}
auto url = parse_url(argv[1]);
if (!url) {
std::cerr << "Could not parse URL (expected http://host[:port]/path)\n";
return 1;
}
auto raw = fetch(url->host, url->port, url->path);
if (!raw) {
std::cerr << "Could not fetch " << argv[1] << "\n";
return 1;
}
auto split = split_response(*raw);
if (!split) {
std::cerr << "Malformed HTTP response\n";
return 1;
}
std::cout << split->status_line << "\n";
auto links = extract_links(split->body);
std::cout << "Found " << links.size() << " link(s):\n";
for (const auto &link : links) {
std::cout << " " << link << "\n";
}
return 0;
}
close() automatically, on every exit path, including an early return std::nullopt from fetch. C’s version of this project has to remember to call close(fd) by hand on every one of those same early-return branches — miss one, and that is a leaked file descriptor. Socket is move-only (copy is deleted) since a socket handle should never be silently duplicated and closed twice.std::regex href_re(R"re(href=\"([^\"]*)\")re"); — unlike C, C++’s standard library ships a real regex engine in
<regex>. extract_links can match every href="..." occurrence and capture the quoted text directly, instead of hand-writing a strstr/strchr scanner the way the C version had to (even though C’s own libc separately ships POSIX <regex.h>, this project’s C version chose not to use it, for reasons explained on that page).getaddrinfo(host.c_str(), port_str.c_str(), &hints, &result) — the same modern, IPv4/IPv6-agnostic hostname resolution C uses, called identically from C++ since it is a POSIX C API with no C++ standard-library equivalent.
raw.find("\r\n\r\n") —
std::string::find locates the same blank-line boundary HTTP itself defines between headers and body, the C++ equivalent of C’s strstr.
Socket be copied instead of moved — two copies with the same fd would both try to close it, the second call operating on an already-closed (or worse, reused) descriptor.R"(href="([^"]*)")" for a pattern that itself ends in )" — the raw string terminates at the first )" it finds, silently truncating the pattern and breaking the build with a confusing error far from the real cause.R"re(...)re", whenever the pattern itself might contain )".4 Test & Prove Each Part
Eight checks, including a real end-to-end test that starts an actual HTTP server on a background std::thread bound to an OS-assigned port, then has fetch() talk to it over a real socket — no mocking, the same discipline the C version used with a background pthread.
#define _POSIX_C_SOURCE 200809L
#include "WebScraper.hpp"
#include <arpa/inet.h>
#include <cassert>
#include <cstring>
#include <iostream>
#include <netinet/in.h>
#include <sys/socket.h>
#include <thread>
#include <unistd.h>
#define RUN(name) do { name(); std::cout << "PASS: " << #name << "\n"; } while (0)
static void parse_url_splits_host_port_and_path() {
auto u = parse_url("http://example.com:8080/path/to/page");
assert(u.has_value());
assert(u->host == "example.com");
assert(u->port == 8080);
assert(u->path == "/path/to/page");
}
static void parse_url_defaults_to_port_80_and_root_path() {
auto u = parse_url("http://example.com");
assert(u.has_value());
assert(u->port == 80);
assert(u->path == "/");
}
static void parse_url_rejects_a_non_http_scheme() {
assert(!parse_url("ftp://example.com/").has_value());
assert(!parse_url("not a url at all").has_value());
}
static void split_response_finds_the_blank_line_boundary() {
std::string raw = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\n\r\n<html>body</html>";
auto s = split_response(raw);
assert(s.has_value());
assert(s->status_line == "HTTP/1.1 200 OK");
assert(s->body == "<html>body</html>");
}
static void split_response_rejects_a_response_with_no_blank_line() {
assert(!split_response("HTTP/1.1 200 OK\r\nno blank line here").has_value());
}
static void extract_links_finds_every_href() {
std::string html = R"(<a href="/one">One</a><a href="https://two.example/">Two</a>)";
auto links = extract_links(html);
assert(links.size() == 2);
assert(links[0] == "/one");
assert(links[1] == "https://two.example/");
}
static void extract_links_returns_empty_for_html_with_no_links() {
assert(extract_links("<html><body>No links here</body></html>").empty());
}
// A real end-to-end test: start a tiny HTTP server on a background thread
// bound to an OS-assigned port, then have fetch() talk to it over an
// actual socket -- no mocking, the same discipline the C version of this
// project used with a background pthread.
static void fetch_and_extract_links_work_against_a_real_local_server() {
int server_fd = socket(AF_INET, SOCK_STREAM, 0);
assert(server_fd >= 0);
int opt = 1;
setsockopt(server_fd, SOL_SOCKET, SO_REUSEADDR, &opt, sizeof(opt));
sockaddr_in addr{};
addr.sin_family = AF_INET;
addr.sin_addr.s_addr = INADDR_ANY;
addr.sin_port = 0; // let the OS pick a free port
assert(bind(server_fd, reinterpret_cast<sockaddr *>(&addr), sizeof(addr)) == 0);
assert(listen(server_fd, 1) == 0);
socklen_t len = sizeof(addr);
getsockname(server_fd, reinterpret_cast<sockaddr *>(&addr), &len);
int port = ntohs(addr.sin_port);
std::thread server([server_fd]() {
int client = accept(server_fd, nullptr, nullptr);
if (client < 0) return;
char buf[1024];
recv(client, buf, sizeof(buf), 0); // drain the request, ignore its contents
std::string body = "<html><body><a href=\"/page1\">One</a><a href=\"/page2\">Two</a></body></html>";
std::string response = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\nContent-Length: " +
std::to_string(body.size()) + "\r\n\r\n" + body;
send(client, response.c_str(), response.size(), 0);
close(client);
});
auto raw = fetch("127.0.0.1", port, "/");
server.join();
close(server_fd);
assert(raw.has_value());
auto split = split_response(*raw);
assert(split.has_value());
assert(split->status_line == "HTTP/1.1 200 OK");
auto links = extract_links(split->body);
assert(links.size() == 2);
assert(links[0] == "/page1");
assert(links[1] == "/page2");
}
int main() {
RUN(parse_url_splits_host_port_and_path);
RUN(parse_url_defaults_to_port_80_and_root_path);
RUN(parse_url_rejects_a_non_http_scheme);
RUN(split_response_finds_the_blank_line_boundary);
RUN(split_response_rejects_a_response_with_no_blank_line);
RUN(extract_links_finds_every_href);
RUN(extract_links_returns_empty_for_html_with_no_links);
RUN(fetch_and_extract_links_work_against_a_real_local_server);
std::cout << "All tests passed.\n";
return 0;
}
Compile and run with g++ -std=c++20 -Wall -Wextra -Wpedantic -pthread -o test_run WebScraper.cpp test_WebScraper.cpp && ./test_run. The -pthread flag is required because the end-to-end test spins up a real server on a std::thread.
5 The Interface
What it expects
$ ./scraper http://example.com/page.htmlWhat it returns
HTTP/1.1 200 OK
Found 2 link(s):
/foo
/bar6 Run It & Automate It
Save the code as WebScraper.hpp / WebScraper.cpp / main.cpp and compile it with g++ — that turns your source directly into a native executable for your machine. No separate runtime needed: the compiled binary runs on its own.
g++ -std=c++20 -pthread -o scraper main.cpp WebScraper.cpp && ./scraper http://example.com/Point it at any plain HTTP (not HTTPS — this project does not implement TLS) URL.
A CI tool like Jenkins runs the same compile-then-test-then-check-for-leaks steps automatically whenever the code changes — every line below has a plain explanation.
$ ./scraper http://127.0.0.1:8199/index.html
HTTP/1.0 200 OK
Found 2 link(s):
/foo
/barhttps:// (almost every real site now does) will fail to connect on port 80 the way this expects.split->body before calling extract_links. Many sites gzip-compress their response, and this project does not decompress Content-Encoding: gzip.// Jenkinsfile — compiles, tests, and checks for leaks on every change.
pipeline {
agent any
stages {
stage('Get the code') {
// download the latest code
steps { checkout scm }
}
stage('Compile') {
steps {
// confirm a compiler is installed
sh 'g++ --version'
// compile with strict warnings on
sh 'g++ -std=c++20 -Wall -Wextra -o app *.cpp -pthread'
}
}
stage('Run the tests') {
steps {
// prints PASS/FAIL, exits non-zero on failure
sh './app'
}
}
stage('Check for memory leaks') {
steps {
// fails the build on any leak or invalid access
sh 'valgrind --error-exitcode=1 --leak-check=full ./app'
}
}
}
post {
success { echo 'All tests passed, no leaks found.' }
failure { echo 'A test or Valgrind check failed — see above.' }
}
}
- Follow redirects. A 301/302 response includes a
Locationheader — fetch that URL next. (Teaches: recursive or looped fetching, and a depth limit to avoid an infinite redirect loop.) - Extract image sources too. Match
src="..."on<img>tags with a secondstd::regex. (Teaches: composing more than one pattern over the same text.) - Resolve relative links to absolute URLs.
/foofound onhttp://example.com/pageshould becomehttp://example.com/foo. (Teaches: URL-joining logic most HTTP libraries hide from you.)
std::regex for a genuine pattern match instead of a hand-written character scanner — both direct upgrades over this project’s C version. Related: Classes and RAII, STL Algorithms.