1 The Problem
We want a tool that takes some text — typed in or read from a file — and reports how many words, characters, and lines it has. It teaches the core string operations for breaking text into pieces and measuring them.
2 How to Think About It
Three separate jobs, each one small: count basics, tally word frequency, and pick the top few. The design question worth pausing on is what data structure holds the tally.
3 The Build — explained part by part
Here is the complete counter. Read the note on word_frequency first — it is the one place this project genuinely improves on the C version, not just translates it.
#pragma once
#include <string>
#include <utility>
#include <vector>
struct Counts {
int lines = 0;
int words = 0;
int chars = 0;
};
// Counts lines, words, and characters in `text`, the same three numbers the
// real `wc` command reports.
Counts count_all(const std::string &text);
// Counts how often each lowercase word appears in `text`, case-insensitive.
// Unlike C, C++'s standard library has a real hash map -- std::unordered_map
// -- so this needs no hand-rolled linear-search table.
std::vector<std::pair<std::string, int>> word_frequency(const std::string &text);
// Returns the top `n` entries from `counts`, sorted by count descending
// (ties broken alphabetically for a stable, reproducible order).
std::vector<std::pair<std::string, int>> top_n(
const std::vector<std::pair<std::string, int>> &counts, int n);
#include "WordCounter.hpp"
#include <algorithm>
#include <cctype>
#include <sstream>
#include <unordered_map>
Counts count_all(const std::string &text) {
Counts c;
c.chars = static_cast<int>(text.size());
bool in_word = false;
for (char ch : text) {
if (ch == '\n') c.lines++;
if (std::isspace(static_cast<unsigned char>(ch))) {
in_word = false;
} else if (!in_word) {
in_word = true;
c.words++;
}
}
if (!text.empty() && text.back() != '\n') c.lines++; // count a trailing partial line
return c;
}
static std::string to_lower(const std::string &s) {
std::string out = s;
std::transform(out.begin(), out.end(), out.begin(),
[](unsigned char ch) { return std::tolower(ch); });
return out;
}
std::vector<std::pair<std::string, int>> word_frequency(const std::string &text) {
std::unordered_map<std::string, int> counts;
std::istringstream stream(text);
std::string word;
while (stream >> word) {
counts[to_lower(word)]++;
}
return std::vector<std::pair<std::string, int>>(counts.begin(), counts.end());
}
std::vector<std::pair<std::string, int>> top_n(
const std::vector<std::pair<std::string, int>> &counts, int n) {
std::vector<std::pair<std::string, int>> sorted = counts;
std::sort(sorted.begin(), sorted.end(), [](const auto &a, const auto &b) {
if (a.second != b.second) return a.second > b.second;
return a.first < b.first;
});
if (static_cast<int>(sorted.size()) > n) sorted.resize(n);
return sorted;
}
#include "WordCounter.hpp"
#include <fstream>
#include <iostream>
#include <sstream>
int main(int argc, char **argv) {
if (argc < 2) {
std::cerr << "Usage: word_counter <file>\n";
return 1;
}
std::ifstream file(argv[1]);
if (!file) {
std::cerr << "Could not open " << argv[1] << "\n";
return 1;
}
std::stringstream buffer;
buffer << file.rdbuf();
std::string text = buffer.str();
Counts c = count_all(text);
std::cout << "Lines: " << c.lines << "\n";
std::cout << "Words: " << c.words << "\n";
std::cout << "Characters: " << c.chars << "\n";
auto freq = word_frequency(text);
auto top = top_n(freq, 3);
std::cout << "Top words:\n";
for (const auto &[word, count] : top) {
std::cout << " " << word << ": " << count << "\n";
}
return 0;
}
std::unordered_map is a real hash table, built in, giving average O(1) insert and lookup with no extra code.std::vector<std::pair<std::string, int>> — once counting is done, the map is copied out into a vector of pairs so it can be sorted;
std::unordered_map has no defined order, so anything that needs a stable, sorted result has to leave the map and land in a container that std::sort understands.std::sort(sorted.begin(), sorted.end(), [](const auto &a, const auto &b) { ... }) — a lambda as the comparison: count descending first, then alphabetical for a stable tie-break, so the top-N result is reproducible run to run instead of depending on hash-table iteration order.
std::stringstream buffer; buffer << file.rdbuf(); — the whole-file-into-a-string idiom in
main.cpp: rdbuf() gives direct access to the stream’s underlying buffer, letting one line slurp an entire file rather than a manual read loop.
std::unordered_map iteration to come out in insertion or alphabetical order — it comes out in hash bucket order, which looks random and can even change between runs.std::vector and sort that, as top_n does here, whenever order matters."The" and "the" would then count as two different words.std::transform and std::tolower before it goes into the map, as word_frequency does here.4 Test & Prove Each Part
Five small checks, each proving one rule: the raw counts, case-insensitive tallying, descending sort order, the empty-input edge case, and alphabetical tie-breaking. Same hand-written assert() harness as the rest of this project — the package registry needed for Catch2 or GoogleTest is not reachable in this environment either.
#include "WordCounter.hpp"
#include <algorithm>
#include <cassert>
#include <iostream>
#define RUN(name) do { name(); std::cout << "PASS: " << #name << "\n"; } while (0)
static void counts_lines_words_and_chars_correctly() {
Counts c = count_all("the cat sat\non the mat\n");
assert(c.lines == 2);
assert(c.words == 6);
}
static void frequency_is_case_insensitive() {
auto freq = word_frequency("The the THE cat");
assert(freq.size() == 2);
auto it = std::find_if(freq.begin(), freq.end(),
[](const auto &p) { return p.first == "the"; });
assert(it != freq.end());
assert(it->second == 3);
}
static void top_n_sorts_by_count_descending() {
auto freq = word_frequency("a a a b b c");
auto top = top_n(freq, 2);
assert(top.size() == 2);
assert(top[0].first == "a" && top[0].second == 3);
assert(top[1].first == "b" && top[1].second == 2);
}
static void empty_text_counts_zero_everything_that_matters() {
Counts c = count_all("");
assert(c.words == 0);
assert(c.lines == 0);
auto freq = word_frequency("");
assert(freq.empty());
}
static void ties_in_top_n_break_alphabetically() {
auto freq = word_frequency("zebra zebra apple apple");
auto top = top_n(freq, 2);
assert(top[0].first == "apple");
assert(top[1].first == "zebra");
}
int main() {
RUN(counts_lines_words_and_chars_correctly);
RUN(frequency_is_case_insensitive);
RUN(top_n_sorts_by_count_descending);
RUN(empty_text_counts_zero_everything_that_matters);
RUN(ties_in_top_n_break_alphabetically);
std::cout << "All tests passed.\n";
return 0;
}
Compile and run with g++ -std=c++20 -o test_run WordCounter.cpp test_WordCounter.cpp && ./test_run.
5 The Interface
What it expects
$ ./wc notes.txtWhat it returns
Lines: 2
Words: 6
Characters: 24
Top words:
the: 2
cat: 1
mat: 16 Run It & Automate It
Save the code as WordCounter.hpp / WordCounter.cpp / main.cpp and compile it with g++ — that turns your source directly into a native executable for your machine. No separate runtime needed: the compiled binary runs on its own.
g++ -std=c++20 -o wc main.cpp WordCounter.cpp && ./wc notes.txtPoint it at any text file you have lying around.
A CI tool like Jenkins runs the same compile-then-test-then-check-for-leaks steps automatically whenever the code changes — every line below has a plain explanation.
$ ./wc notes.txt
Lines: 2
Words: 6
Characters: 24
Top words:
the: 2
cat: 1
mat: 1pwd or pass an absolute path.argv[1] was read without first checking argc < 2 — the program tried to open a null or garbage path.// Jenkinsfile — compiles, tests, and checks for leaks on every change.
pipeline {
agent any
stages {
stage('Get the code') {
// download the latest code
steps { checkout scm }
}
stage('Compile') {
steps {
// confirm a compiler is installed
sh 'g++ --version'
// compile with strict warnings on
sh 'g++ -std=c++20 -Wall -Wextra -o app *.cpp'
}
}
stage('Run the tests') {
steps {
// prints PASS/FAIL, exits non-zero on failure
sh './app'
}
}
stage('Check for memory leaks') {
steps {
// fails the build on any leak or invalid access
sh 'valgrind --error-exitcode=1 --leak-check=full ./app'
}
}
}
post {
success { echo 'All tests passed, no leaks found.' }
failure { echo 'A test or Valgrind check failed — see above.' }
}
}
- Ignore punctuation. Strip leading/trailing punctuation from each word before counting. (Teaches: character classification with
std::ispunct.) - Read from stdin when no file is given. Fall back to
std::cinlike the realwcdoes. (Teaches: treatingstd::cinand anifstreampolymorphically throughstd::istream&.) - Report the longest word. Track it alongside the frequency tally. (Teaches: a running max during a single pass.)
std::unordered_map as C++’s built-in hash table instead of hand-rolling a linear-search table the way C has to, why an unordered map still needs a vector-plus-sort step whenever order matters, and how rdbuf() slurps a whole file in one line. Related: STL Containers, Standard Library.