← thecodex.expert · The Codex Family of Knowledge
Tier 3 · Upper-Intermediate · Java Project

Log Analyser

Parse semi-structured log lines and report counts by level and the most frequent messages. Teaches Optional for a parse that might fail, and a generic top-N counter reusable for any frequency table.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a log analyser: read a server log file, count how many entries are errors versus normal, find which hour had the most traffic, and list the most frequent error messages. It teaches parsing semi-structured text at scale and aggregating it into useful insight — a daily task in operations.

Where this shows up: monitoring and observability, debugging production issues, security analysis, performance tuning. When something breaks at 3am, the person who can analyse the logs is the one who fixes it.

2 How to Think About It

Real logs have garbage lines mixed in. Design for that from the start: parsing a single line returns an Optional, never throws, and the rest of the program only ever sees successfully parsed entries.

The plan — in plain English
1. Parse each line, keeping only the ones that succeed. → 2. Count entries by level. → 3. Count entries by message text. → 4. Report the level breakdown and the top few messages.

Read log lines

Parse each: time, level, message

Count errors

Count entries per hour

Count message frequency

Report insights

3 The Build — explained part by part

Here is the complete analyser. parseLine returning Optional<Entry> instead of throwing is the load-bearing decision — it is what makes a malformed line an ordinary, expected outcome instead of a crash.

JavaLogAnalyser.java
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;

/**
 * Log Analyser: parses "YYYY-MM-DD HH:MM:SS LEVEL message" log lines and
 * reports level counts and the most frequent messages.
 */
public class LogAnalyser {

    record Entry(String date, String time, String level, String message) {}

    /** Returns empty for a line that doesn't structurally look like a log line, rather than guessing. */
    static Optional<Entry> parseLine(String line) {
        String[] parts = line.split(" ", 4);
        if (parts.length < 4) return Optional.empty();
        String date = parts[0], time = parts[1], level = parts[2], message = parts[3];
        if (!date.contains("-") || !time.contains(":")) return Optional.empty();
        return Optional.of(new Entry(date, time, level, message));
    }

    static List<Entry> parseAll(List<String> lines) {
        return lines.stream().flatMap(l -> parseLine(l).stream()).toList();
    }

    static Map<String, Long> countByLevel(List<Entry> entries) {
        Map<String, Long> counts = new LinkedHashMap<>();
        for (Entry e : entries) {
            counts.merge(e.level(), 1L, Long::sum);
        }
        return counts;
    }

    /** Reusable for any string-keyed frequency table, not just log messages. */
    static <T> List<Map.Entry<T, Long>> topN(Map<T, Long> counts, int n) {
        return counts.entrySet().stream()
                .sorted((a, b) -> Long.compare(b.getValue(), a.getValue()))
                .limit(n)
                .toList();
    }

    static Map<String, Long> countByMessage(List<Entry> entries) {
        Map<String, Long> counts = new LinkedHashMap<>();
        for (Entry e : entries) {
            counts.merge(e.message(), 1L, Long::sum);
        }
        return counts;
    }

    public static void main(String[] args) throws Exception {
        if (args.length != 1) {
            System.out.println("Usage: java LogAnalyser <file>");
            return;
        }
        List<Entry> entries = parseAll(Files.readAllLines(Path.of(args[0])));
        System.out.println("Parsed " + entries.size() + " entries.");
        System.out.println("By level:");
        countByLevel(entries).forEach((level, count) -> System.out.printf("  %-6s %d%n", level, count));
        System.out.println("Top messages:");
        for (var entry : topN(countByMessage(entries), 3)) {
            System.out.printf("  %-30s %d%n", entry.getKey(), entry.getValue());
        }
    }
}
⚠ No in-browser playground here
Java compiles to JVM bytecode and needs a real JDK to run, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once a JDK is installed.
What each part does — in plain words
static Optional<Entry> parseLine(String line) — from the course’s Optional and Null Safety lesson: a genuinely malformed line is not an error condition worth an exception, it is an expected possibility the caller must handle. Checking date.contains("-") and time.contains(":") catches lines that split into 4 parts by accident but are not actually log lines — the exact gap a naive splitn-based check would miss.

lines.stream().flatMap(l -> parseLine(l).stream()) — an Optional has a .stream() method that yields zero or one element, which is exactly what flatMap needs to turn “a stream of maybe-parsed lines” into “a stream of only the entries that parsed” in one line, with no explicit filtering step.

static <T> List<Map.Entry<T, Long>> topN(Map<T, Long> counts, int n) — a generic method from the course’s Collections and Generics lesson: it works identically whether counting log levels, messages, or anything else keyed by any type T, because nothing in its body assumes what T actually is.

counts.merge(key, 1L, Long::sum) — the same one-call insert-or-increment pattern the word-counter project used, here counting by Long instead of Integer since a log file can have far more lines than a word-counted text file.
Common mistakes — and how to avoid them
✗ Letting parseLine throw on a malformed line — a single bad line anywhere in a large log file would crash the entire analysis.
✓ Return Optional.empty() for anything that does not structurally look right, as parseLine does, and skip it downstream.
✗ Checking only parts.length == 4 to validate a line — "not a real log line" also splits into exactly 4 parts by split(" ", 4), so length alone lets garbage through.
✓ Also check that the date and time fields actually look like a date and a time, as parseLine does.

4 Test & Prove Each Part

We test the parser on both well-formed and deliberately malformed input, and the counting logic in isolation.

A well-formed log line parses into the right level and message
A line that does not look like a log line returns Optional.empty(), not a thrown exception
A malformed line mixed into good ones is skipped, not fatal
Counting by level tallies correctly
The top-N helper returns the most frequent entries first, for any key type
JavaLogAnalyserTest.java
import org.junit.Test;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import static org.junit.Assert.assertEquals;
import static org.junit.Assert.assertTrue;

public class LogAnalyserTest {

    @Test
    public void parsesAWellFormedLine() {
        Optional<LogAnalyser.Entry> e = LogAnalyser.parseLine("2026-09-28 10:00:01 INFO Server started");
        assertTrue(e.isPresent());
        assertEquals("INFO", e.get().level());
        assertEquals("Server started", e.get().message());
    }

    @Test
    public void rejectsALineThatDoesNotLookLikeALogLine() {
        assertTrue(LogAnalyser.parseLine("not a real log line").isEmpty());
        assertTrue(LogAnalyser.parseLine("too short").isEmpty());
    }

    @Test
    public void malformedLinesAreSkippedNotCrashed() {
        List<String> lines = List.of(
                "2026-09-28 10:00:01 INFO ok",
                "garbage",
                "2026-09-28 10:00:02 ERROR also ok");
        List<LogAnalyser.Entry> entries = LogAnalyser.parseAll(lines);
        assertEquals(2, entries.size());
    }

    @Test
    public void countsByLevelCorrectly() {
        List<LogAnalyser.Entry> entries = LogAnalyser.parseAll(List.of(
                "2026-01-01 00:00:00 INFO a",
                "2026-01-01 00:00:01 INFO b",
                "2026-01-01 00:00:02 ERROR c"));
        Map<String, Long> counts = LogAnalyser.countByLevel(entries);
        assertEquals(Long.valueOf(2), counts.get("INFO"));
        assertEquals(Long.valueOf(1), counts.get("ERROR"));
    }

    @Test
    public void topNReturnsMostFrequentFirst() {
        Map<String, Long> counts = Map.of("a", 5L, "b", 1L, "c", 3L);
        List<Map.Entry<String, Long>> top = LogAnalyser.topN(counts, 2);
        assertEquals("a", top.get(0).getKey());
        assertEquals("c", top.get(1).getKey());
    }
}

Compile and run with javac -cp junit-4.13.2.jar and hamcrest-core-1.3.jar LogAnalyser.java LogAnalyserTest.java then java -cp .:junit-4.13.2.jar:hamcrest-core-1.3.jar org.junit.runner.JUnitCore LogAnalyserTest. The malformed-line test is the important one: it feeds a genuinely garbage line into parseAll alongside good ones and checks the good ones still come through, rather than only testing parseLine in isolation.

5 The Interface

INPUTINPUTa log file path
What it expects
java LogAnalyser server.log
OUTPUTOUTPUTlevel counts and top messages
What it returns
Parsed 6 entries.
By level:
  INFO   3
  ERROR  2
Top messages:
  Request received  2

6 Run It & Automate It

Save the code as LogAnalyser.java and compile it with javac — that turns your source into .class bytecode files, which java then runs on the JVM. No separate install step: any real JDK ships both tools.

Run it locally
javac LogAnalyser.java && java LogAnalyser sample.log
Create a sample.log with a few log lines first, in the same folder.

A CI tool like Jenkins runs the same compile-then-test steps automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ java LogAnalyser sample.log
Parsed 6 entries.
By level:
  INFO   3
  ERROR  2
  WARN   1
Top messages:
  Request received               2
  Database timeout               2
  Server started                 1
If it breaks — how to fix it
🚨 java.nio.file.NoSuchFileException
The log file was not found relative to the folder you ran java from. Check your current directory or pass a full path.
🚨 The entry count looks lower than the number of lines in the file.
That usually means some lines genuinely failed the date/time shape check in parseLine — open the file and check whether a line is missing its timestamp or level.
GroovyJenkinsfile
// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
    agent any                          // run on any available machine
    environment {
        CP = 'junit-4.13.2.jar:hamcrest-core-1.3.jar'   // JUnit + its one dependency
    }

    stages {
        stage('Get the code') {
            steps { checkout scm }     // download the latest code
        }
        stage('Set up JDK') {
            steps {
                sh 'java -version'           // confirm a JDK is installed
                sh 'javac -cp "$CP" *.java'   // compile the program and its tests together
            }
        }
        stage('Run the tests') {
            steps {
                sh 'java -cp ".:$CP" org.junit.runner.JUnitCore LogAnalyserTest'
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed — look above.' }
    }
}
🎯 Try this next — make it yours
  1. Filter by time range. Only count entries between two timestamps. (Teaches: parsing the date/time fields into real LocalDateTime values.)
  2. Export to CSV. Write the level counts to a file instead of stdout. (Teaches: reusing the Files API for output, not just input.)
  3. Detect error bursts. Flag when 3+ ERROR lines happen within any 60-second window. (Teaches: a sliding-window algorithm over parsed timestamps.)
What you learned
You learned returning Optional for a parse that might legitimately fail, combining it with flatMap to filter a stream in one step, and writing a generic method that works for any key type without knowing what that type is. Related: Optional and Null Safety, Streams and Lambdas.