← thecodex.expert · The Codex Family of Knowledge
Tier 0 · Absolute Beginner · C++ Project

Word Counter

Count lines, words, and characters in a text file, the same three numbers the real wc command reports, then find its most frequent words. The interesting part is not the counting — it is what C++’s standard library gives you for free that C does not.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a tool that takes some text — typed in or read from a file — and reports how many words, characters, and lines it has. It teaches the core string operations for breaking text into pieces and measuring them.

Where this shows up: word-count limits on forms and essays, reading-time estimates, search indexing, text analysis, validating input length. Measuring and slicing text is one of the most common jobs in software.

2 How to Think About It

Three separate jobs, each one small: count basics, tally word frequency, and pick the top few. The design question worth pausing on is what data structure holds the tally.

The plan — in plain English
1. Read the whole file into a string. → 2. Count lines, words, and characters in one pass. → 3. Tally how many times each lowercase word appears. → 4. Sort the tally and keep the top few.

Get the text

Characters = length

Words = split on spaces, count

Lines = split on newlines, count

Show all three

3 The Build — explained part by part

Here is the complete counter. Read the note on word_frequency first — it is the one place this project genuinely improves on the C version, not just translates it.

C++WordCounter.hpp / WordCounter.cpp / main.cpp
#pragma once
#include <string>
#include <utility>
#include <vector>

struct Counts {
    int lines = 0;
    int words = 0;
    int chars = 0;
};

// Counts lines, words, and characters in `text`, the same three numbers the
// real `wc` command reports.
Counts count_all(const std::string &text);

// Counts how often each lowercase word appears in `text`, case-insensitive.
// Unlike C, C++'s standard library has a real hash map -- std::unordered_map
// -- so this needs no hand-rolled linear-search table.
std::vector<std::pair<std::string, int>> word_frequency(const std::string &text);

// Returns the top `n` entries from `counts`, sorted by count descending
// (ties broken alphabetically for a stable, reproducible order).
std::vector<std::pair<std::string, int>> top_n(
    const std::vector<std::pair<std::string, int>> &counts, int n);

#include "WordCounter.hpp"
#include <algorithm>
#include <cctype>
#include <sstream>
#include <unordered_map>

Counts count_all(const std::string &text) {
    Counts c;
    c.chars = static_cast<int>(text.size());
    bool in_word = false;
    for (char ch : text) {
        if (ch == '\n') c.lines++;
        if (std::isspace(static_cast<unsigned char>(ch))) {
            in_word = false;
        } else if (!in_word) {
            in_word = true;
            c.words++;
        }
    }
    if (!text.empty() && text.back() != '\n') c.lines++; // count a trailing partial line
    return c;
}

static std::string to_lower(const std::string &s) {
    std::string out = s;
    std::transform(out.begin(), out.end(), out.begin(),
                    [](unsigned char ch) { return std::tolower(ch); });
    return out;
}

std::vector<std::pair<std::string, int>> word_frequency(const std::string &text) {
    std::unordered_map<std::string, int> counts;
    std::istringstream stream(text);
    std::string word;
    while (stream >> word) {
        counts[to_lower(word)]++;
    }
    return std::vector<std::pair<std::string, int>>(counts.begin(), counts.end());
}

std::vector<std::pair<std::string, int>> top_n(
    const std::vector<std::pair<std::string, int>> &counts, int n) {
    std::vector<std::pair<std::string, int>> sorted = counts;
    std::sort(sorted.begin(), sorted.end(), [](const auto &a, const auto &b) {
        if (a.second != b.second) return a.second > b.second;
        return a.first < b.first;
    });
    if (static_cast<int>(sorted.size()) > n) sorted.resize(n);
    return sorted;
}

#include "WordCounter.hpp"
#include <fstream>
#include <iostream>
#include <sstream>

int main(int argc, char **argv) {
    if (argc < 2) {
        std::cerr << "Usage: word_counter <file>\n";
        return 1;
    }
    std::ifstream file(argv[1]);
    if (!file) {
        std::cerr << "Could not open " << argv[1] << "\n";
        return 1;
    }
    std::stringstream buffer;
    buffer << file.rdbuf();
    std::string text = buffer.str();

    Counts c = count_all(text);
    std::cout << "Lines: " << c.lines << "\n";
    std::cout << "Words: " << c.words << "\n";
    std::cout << "Characters: " << c.chars << "\n";

    auto freq = word_frequency(text);
    auto top = top_n(freq, 3);
    std::cout << "Top words:\n";
    for (const auto &[word, count] : top) {
        std::cout << "  " << word << ": " << count << "\n";
    }
    return 0;
}
⚠ No in-browser playground here
C++ compiles to a real, native binary, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once a C++17-or-newer compiler like g++ or clang++ is installed.
What each part does — in plain words
std::unordered_map<std::string, int> counts; — this is the whole reason this project reads differently from its C counterpart. C has no hash map in its standard library, so the C version of this project hand-rolled a linear-search array of (word, count) pairs — correct, but O(n) per lookup. std::unordered_map is a real hash table, built in, giving average O(1) insert and lookup with no extra code.

std::vector<std::pair<std::string, int>> — once counting is done, the map is copied out into a vector of pairs so it can be sorted; std::unordered_map has no defined order, so anything that needs a stable, sorted result has to leave the map and land in a container that std::sort understands.

std::sort(sorted.begin(), sorted.end(), [](const auto &a, const auto &b) { ... }) — a lambda as the comparison: count descending first, then alphabetical for a stable tie-break, so the top-N result is reproducible run to run instead of depending on hash-table iteration order.

std::stringstream buffer; buffer << file.rdbuf(); — the whole-file-into-a-string idiom in main.cpp: rdbuf() gives direct access to the stream’s underlying buffer, letting one line slurp an entire file rather than a manual read loop.
Common mistakes — and how to avoid them
✗ Expecting std::unordered_map iteration to come out in insertion or alphabetical order — it comes out in hash bucket order, which looks random and can even change between runs.
✓ Copy the entries you need into a std::vector and sort that, as top_n does here, whenever order matters.
✗ Comparing words without normalising case first — "The" and "the" would then count as two different words.
✓ Lower-case every word with std::transform and std::tolower before it goes into the map, as word_frequency does here.

4 Test & Prove Each Part

Five small checks, each proving one rule: the raw counts, case-insensitive tallying, descending sort order, the empty-input edge case, and alphabetical tie-breaking. Same hand-written assert() harness as the rest of this project — the package registry needed for Catch2 or GoogleTest is not reachable in this environment either.

Line, word, and character counts match a hand-counted example
Word frequency is case-insensitive ("The", "the", "THE" all count as one word)
top_n sorts by count descending
Empty text produces zero counts and an empty frequency list, not a crash
Ties in count are broken alphabetically, for a reproducible order
C++test_WordCounter.cpp
#include "WordCounter.hpp"
#include <algorithm>
#include <cassert>
#include <iostream>

#define RUN(name) do { name(); std::cout << "PASS: " << #name << "\n"; } while (0)

static void counts_lines_words_and_chars_correctly() {
    Counts c = count_all("the cat sat\non the mat\n");
    assert(c.lines == 2);
    assert(c.words == 6);
}

static void frequency_is_case_insensitive() {
    auto freq = word_frequency("The the THE cat");
    assert(freq.size() == 2);
    auto it = std::find_if(freq.begin(), freq.end(),
                            [](const auto &p) { return p.first == "the"; });
    assert(it != freq.end());
    assert(it->second == 3);
}

static void top_n_sorts_by_count_descending() {
    auto freq = word_frequency("a a a b b c");
    auto top = top_n(freq, 2);
    assert(top.size() == 2);
    assert(top[0].first == "a" && top[0].second == 3);
    assert(top[1].first == "b" && top[1].second == 2);
}

static void empty_text_counts_zero_everything_that_matters() {
    Counts c = count_all("");
    assert(c.words == 0);
    assert(c.lines == 0);
    auto freq = word_frequency("");
    assert(freq.empty());
}

static void ties_in_top_n_break_alphabetically() {
    auto freq = word_frequency("zebra zebra apple apple");
    auto top = top_n(freq, 2);
    assert(top[0].first == "apple");
    assert(top[1].first == "zebra");
}

int main() {
    RUN(counts_lines_words_and_chars_correctly);
    RUN(frequency_is_case_insensitive);
    RUN(top_n_sorts_by_count_descending);
    RUN(empty_text_counts_zero_everything_that_matters);
    RUN(ties_in_top_n_break_alphabetically);
    std::cout << "All tests passed.\n";
    return 0;
}

Compile and run with g++ -std=c++20 -o test_run WordCounter.cpp test_WordCounter.cpp && ./test_run.

5 The Interface

INPUTINPUTa path to a text file, given as a command-line argument
What it expects
$ ./wc notes.txt
OUTPUTOUTPUTline/word/character counts, then the top 3 words
What it returns
Lines: 2
Words: 6
Characters: 24
Top words:
  the: 2
  cat: 1
  mat: 1

6 Run It & Automate It

Save the code as WordCounter.hpp / WordCounter.cpp / main.cpp and compile it with g++ — that turns your source directly into a native executable for your machine. No separate runtime needed: the compiled binary runs on its own.

Run it locally
g++ -std=c++20 -o wc main.cpp WordCounter.cpp && ./wc notes.txt
Point it at any text file you have lying around.

A CI tool like Jenkins runs the same compile-then-test-then-check-for-leaks steps automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ ./wc notes.txt
Lines: 2
Words: 6
Characters: 24
Top words:
  the: 2
  cat: 1
  mat: 1
If it breaks — how to fix it
🚨 Could not open notes.txt
The path is relative to the current working directory the program was launched from, not the source file’s location — check with pwd or pass an absolute path.
🚨 terminate called after throwing an instance of ‘std::bad_alloc’
Almost always means argv[1] was read without first checking argc < 2 — the program tried to open a null or garbage path.
GroovyJenkinsfile
// Jenkinsfile — compiles, tests, and checks for leaks on every change.
pipeline {
    agent any

    stages {
        stage('Get the code') {
            // download the latest code
            steps { checkout scm }
        }
        stage('Compile') {
            steps {
                // confirm a compiler is installed
                sh 'g++ --version'
                // compile with strict warnings on
                sh 'g++ -std=c++20 -Wall -Wextra -o app *.cpp'
            }
        }
        stage('Run the tests') {
            steps {
                // prints PASS/FAIL, exits non-zero on failure
                sh './app'
            }
        }
        stage('Check for memory leaks') {
            steps {
                // fails the build on any leak or invalid access
                sh 'valgrind --error-exitcode=1 --leak-check=full ./app'
            }
        }
    }

    post {
        success { echo 'All tests passed, no leaks found.' }
        failure { echo 'A test or Valgrind check failed — see above.' }
    }
}
🎯 Try this next — make it yours
  1. Ignore punctuation. Strip leading/trailing punctuation from each word before counting. (Teaches: character classification with std::ispunct.)
  2. Read from stdin when no file is given. Fall back to std::cin like the real wc does. (Teaches: treating std::cin and an ifstream polymorphically through std::istream&.)
  3. Report the longest word. Track it alongside the frequency tally. (Teaches: a running max during a single pass.)
What you learned
You learned to reach for std::unordered_map as C++’s built-in hash table instead of hand-rolling a linear-search table the way C has to, why an unordered map still needs a vector-plus-sort step whenever order matters, and how rdbuf() slurps a whole file in one line. Related: STL Containers, Standard Library.