1 The Problem
We want a tool that takes some text — typed in or read from a file — and reports how many words, characters, and lines it has. It teaches the core string operations for breaking text into pieces and measuring them.
2 How to Think About It
Think of the file as one long string to slice up three different ways: by whitespace (words), by length (characters), and by line.
String with Files.readString. → 2. Split on whitespace to count words and lines. → 3. Build a frequency table with a HashMap. → 4. Report the counts and the top words.
3 The Build — explained part by part
Here is the complete counter. A small record bundles the four counts together so the method returns one clear value instead of four loose numbers.
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
/**
* Word Counter: reads a text file and reports line, word, and character
* counts, plus the most frequent words.
*/
public class WordCounter {
record Counts(int lines, int words, int chars, int bytes) {}
static Counts countAll(String text) {
int chars = text.length();
int bytes = text.getBytes().length;
String[] lines = text.isEmpty() ? new String[0] : text.split("\n", -1);
int lineCount = text.isEmpty() ? 0 : lines.length - (text.endsWith("\n") ? 1 : 0);
String[] words = text.trim().isEmpty() ? new String[0] : text.trim().split("\\s+");
return new Counts(lineCount, words.length, chars, bytes);
}
static Map<String, Integer> wordFrequency(String text) {
Map<String, Integer> freq = new LinkedHashMap<>();
for (String raw : text.toLowerCase().split("\\s+")) {
String word = raw.replaceAll("[^a-z0-9']", "");
if (word.isEmpty()) continue;
freq.merge(word, 1, Integer::sum);
}
return freq;
}
static List<Map.Entry<String, Integer>> topN(Map<String, Integer> freq, int n) {
return freq.entrySet().stream()
.sorted((a, b) -> b.getValue() - a.getValue())
.limit(n)
.toList();
}
public static void main(String[] args) throws IOException {
if (args.length != 1) {
System.out.println("Usage: java WordCounter <file>");
return;
}
String text = Files.readString(Path.of(args[0]));
Counts c = countAll(text);
System.out.printf("Lines: %d Words: %d Characters: %d Bytes: %d%n",
c.lines(), c.words(), c.chars(), c.bytes());
System.out.println("Top words:");
for (var entry : topN(wordFrequency(text), 5)) {
System.out.printf(" %-12s %d%n", entry.getKey(), entry.getValue());
}
}
}
equals, and toString for free, which is exactly enough structure for a method that needs to return four related numbers at once.Files.readString(Path.of(args[0])) — the modern
java.nio.file API from the course’s IO and NIO lesson; it reads an entire file into a String in one call, replacing the old BufferedReader-in-a-loop pattern for anything that comfortably fits in memory.text.length() vs text.getBytes().length — a Java
String is UTF-16 internally, so .length() counts chars (UTF-16 code units, close enough to characters for this project) while .getBytes() re-encodes to UTF-8 bytes — the two numbers only match for plain ASCII text, the same byte-vs-character trap every other language on this site documents.freq.merge(word, 1, Integer::sum) —
merge inserts 1 for a new key or adds 1 to an existing value via the given function, all in one call — Java’s answer to Python’s dict.get(k, 0) + 1 or Rust’s entry().or_insert(0).
text.length() as “the word count” — that counts every character in the file, not the number of words.text.trim().split("\\s+")) and count the resulting array’s length."The" and "the" would be tallied as two different words..toLowerCase() before splitting, as wordFrequency does.4 Test & Prove Each Part
We test the counting logic directly against in-memory strings, without touching a real file for most of the checks.
import org.junit.Test;
import java.util.List;
import java.util.Map;
import static org.junit.Assert.assertEquals;
public class WordCounterTest {
@Test
public void countsLinesWordsAndChars() {
WordCounter.Counts c = WordCounter.countAll("hello world\nsecond line\n");
assertEquals(2, c.lines());
assertEquals(4, c.words());
}
@Test
public void emptyTextCountsAsZero() {
WordCounter.Counts c = WordCounter.countAll("");
assertEquals(0, c.lines());
assertEquals(0, c.words());
}
@Test
public void wordFrequencyIsCaseInsensitiveAndStripsPunctuation() {
Map<String, Integer> freq = WordCounter.wordFrequency("The fox. THE Fox! the fox");
assertEquals(Integer.valueOf(3), freq.get("the"));
assertEquals(Integer.valueOf(3), freq.get("fox"));
}
@Test
public void topNReturnsMostFrequentFirst() {
Map<String, Integer> freq = WordCounter.wordFrequency("a a a b b c");
List<Map.Entry<String, Integer>> top = WordCounter.topN(freq, 2);
assertEquals("a", top.get(0).getKey());
assertEquals("b", top.get(1).getKey());
}
}
Compile and run with javac -cp junit-4.13.2.jar and hamcrest-core-1.3.jar WordCounter.java WordCounterTest.java then java -cp .:junit-4.13.2.jar:hamcrest-core-1.3.jar org.junit.runner.JUnitCore WordCounterTest. Notice these tests never touch the filesystem — they pass plain strings straight to countAll and wordFrequency, which is why splitting file-reading out of the counting logic mattered.
5 The Interface
What it expects
java WordCounter sample.txtWhat it returns
Lines: 2 Words: 10 Characters: 48 Bytes: 48
Top words:
the 3
fox 26 Run It & Automate It
Save the code as WordCounter.java and compile it with javac — that turns your source into .class bytecode files, which java then runs on the JVM. No separate install step: any real JDK ships both tools.
javac WordCounter.java && java WordCounter sample.txtCreate a sample.txt with a few lines of text first, in the same folder.
A CI tool like Jenkins runs the same compile-then-test steps automatically whenever the code changes — every line below has a plain explanation.
$ printf "the quick brown fox\nthe lazy dog the fox jumped\n" > sample.txt
$ java WordCounter sample.txt
Lines: 2 Words: 10 Characters: 48 Bytes: 48
Top words:
the 3
fox 2
quick 1
brown 1
lazy 1java from. Check your current directory, or pass a full path as the argument.split("\\s+") handles runs of whitespace correctly, but a single literal-space split would not.// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
environment {
CP = 'junit-4.13.2.jar:hamcrest-core-1.3.jar' // JUnit + its one dependency
}
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up JDK') {
steps {
sh 'java -version' // confirm a JDK is installed
sh 'javac -cp "$CP" *.java' // compile the program and its tests together
}
}
stage('Run the tests') {
steps {
sh 'java -cp ".:$CP" org.junit.runner.JUnitCore WordCounterTest'
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
- Ignore common stop words. Skip "the", "a", "and" from the top-words list. (Teaches: filtering a stream before collecting.)
- Read from stdin too. Support piping text in when no filename is given. (Teaches:
System.inas a fallback input source.) - Count sentences. Split on
.,!, and?. (Teaches: regex alternation.)
Files.readString for one-call file reading, a record for bundling related results, the UTF-16-chars-vs-UTF-8-bytes distinction in Java strings, and the merge method for one-line frequency counting. Related: IO and NIO, Collections and Generics.