← thecodex.expert · The Codex Family of Knowledge
Tier 3 · Upper-Intermediate · Kotlin Project

Log Analyser

Parse server log files to count errors, find the busiest hours, and spot the top offenders. Turn raw logs into insight.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a log analyser: read a server log file, count how many entries are errors versus normal, find which hour had the most traffic, and list the most frequent error messages. It teaches parsing semi-structured text at scale and aggregating it into useful insight — a daily task in operations.

Where this shows up: monitoring and observability, debugging production issues, security analysis, performance tuning. When something breaks at 3am, the person who can analyse the logs is the one who fixes it.

2 How to Think About It

Think about turning lines into counts, before any code:

The plan — in plain English
1. Read the log file line by line. → 2. Parse each line into its parts: timestamp, level (INFO/ERROR), message. → 3. Count as you go: errors, entries per hour, message frequencies. → 4. Report the totals and the top items.

Read log lines

Parse each: time, level, message

Count errors

Count entries per hour

Count message frequency

Report insights

3 The Build — explained part by part

Here is the complete analyser. Read each part’s note below — you should understand the whole thing from the notes alone.

KotlinlogAnalyser.kt
import java.io.File

/** One parsed line: "YYYY-MM-DD HH:MM:SS LEVEL message". */
data class Entry(val date: String, val time: String, val level: String, val message: String)

/** Returns null for a line that doesn't structurally look like a log line, rather than guessing. */
fun parseLine(line: String): Entry? {
    val parts = line.split(" ", limit = 4)
    if (parts.size < 4) return null
    val (date, time, level, message) = parts
    if (!date.contains("-") || !time.contains(":")) return null
    return Entry(date, time, level, message)
}

fun parseAll(lines: List<String>): List<Entry> = lines.mapNotNull { parseLine(it) }

fun countByLevel(entries: List<Entry>): Map<String, Int> =
    entries.groupingBy { it.level }.eachCount()

fun countByMessage(entries: List<Entry>): Map<String, Int> =
    entries.groupingBy { it.message }.eachCount()

/** Reusable for any string-keyed frequency table, not just log messages. */
fun <T> topN(counts: Map<T, Int>, n: Int): List<Map.Entry<T, Int>> =
    counts.entries.sortedByDescending { it.value }.take(n)

fun main(args: Array<String>) {
    if (args.size != 1) {
        println("Usage: kotlin LogAnalyserKt <file>")
        return
    }
    val entries = parseAll(File(args[0]).readLines())
    println("Parsed ${entries.size} entries.")
    println("By level:")
    countByLevel(entries).forEach { (level, count) -> println("  ${level.padEnd(6)} $count") }
    println("Top messages:")
    for (entry in topN(countByMessage(entries), 3)) {
        println("  ${entry.key.padEnd(30)} ${entry.value}")
    }
}
⚠ No in-browser playground here
Kotlin compiles to real JVM bytecode, not something a browser can run directly — running it live would need either a server-side compiler or a third-party embed, the same kind of external dependency this site avoids relying on for a core teaching example. Copy the code below and run it with a real kotlinc on your own machine instead; the “Run It” section explains exactly how.
What each part does — in plain words
data class Entry(val date: String, val time: String, val level: String, val message: String) — every parsed line becomes one of these; the compiler will not let any function hand back something with a missing or mistyped field.

fun parseLine(line: String): Entry? — returns null for a line that doesn't structurally look like a log line, rather than guessing or throwing; parseAll then uses mapNotNull to quietly drop every line that failed to parse.

entries.groupingBy { it.level }.eachCount() — Kotlin's standard-library answer to Python's collections.Counter: group by a key selector, then count each group's size, in one chained call with no explicit loop or mutable counter.

fun <T> topN(counts: Map<T, Int>, n: Int): List<Map.Entry<T, Int>> — a generic function, reusable for any string-keyed (or otherwise typed) frequency table, not just log messages; sortedByDescending { it.value }.take(n) does the sorting and limiting in one readable chain.
Common mistakes — and how to avoid them
✗ Assuming every line has exactly four space-separated parts.
✓ Check parts.size < 4 and return null for malformed lines, the same guard the Python version uses, so one bad line does not crash the whole report.
✗ Reaching for a mutable HashMap<String, Int> and a hand-written increment loop as a counter.
✓ groupingBy { ... }.eachCount() does the exact same job with no mutable state of your own to get wrong.
✗ Checking date.contains("-") and time.contains(":") and believing that fully validates the format.
✓ These are deliberately loose, structural checks — good enough to reject obviously-wrong input (as the tests confirm) without the complexity of a real date/time parser, which a production log analyser would eventually want.

4 Test & Prove Each Part

How do we know this works? We pull the real logic into small, plain functions and check each one against cases we already know the answer to.

A well-formed line parses into its four fields correctly
A line that doesn't look like a log line returns null, not a crash
Each level is counted correctly across several lines
The most frequent entries come back first, in order
KotlinlogAnalyserTest.kt
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNull

class LogAnalyserTest {
    @Test
    fun parsesAWellFormedLine() {
        val entry = parseLine("2026-01-01 10:00:00 ERROR disk full")
        assertEquals(Entry("2026-01-01", "10:00:00", "ERROR", "disk full"), entry)
    }

    @Test
    fun rejectsALineThatDoesNotLookLikeALog() {
        assertNull(parseLine("just some text"))
    }

    @Test
    fun countsEachLevel() {
        val entries = parseAll(
            listOf(
                "2026-01-01 10:00:00 ERROR disk full",
                "2026-01-01 10:00:01 INFO ok",
                "2026-01-01 10:00:02 ERROR disk full",
            )
        )
        assertEquals(mapOf("ERROR" to 2, "INFO" to 1), countByLevel(entries))
    }

    @Test
    fun topNReturnsTheMostFrequentFirst() {
        val counts = mapOf("a" to 1, "b" to 5, "c" to 3)
        assertEquals(listOf("b", "c"), topN(counts, 2).map { it.key })
    }
}

Compile with kotlinc logAnalyser.kt logAnalyserTest.kt -include-runtime -d logAnalyser.jar and run with JUnit's own runner. parseLine and the counting functions take plain strings and collections, so the tests check known log lines directly, without touching a real log file.

5 The Interface

INPUTINPUTlog file
What it expects
2026-06-24 14:30:00 ERROR Database timeout
OUTPUTOUTPUTreport
What it returns
Total: 6  Errors: 3
Busiest hour: ("14", 3)
Top errors: [("Database timeout", 2), ...]

6 Run It & Automate It

Save the code as logAnalyser.kt and compile it with kotlinc logAnalyser.kt -include-runtime -d logAnalyser.jar. It takes the log file's path as a command-line argument.

Run it locally
kotlinc logAnalyser.kt -include-runtime -d logAnalyser.jar && java -jar logAnalyser.jar server.log
Reads the given file and prints level counts and the most frequent messages.

A CI tool like Jenkins compiles and tests automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
Parsed 5 entries.
By level:
  INFO   2
  ERROR  2
  WARN   1
Top messages:
  disk full                      2
  server started                 1
  request ok                     1
If it breaks — how to fix it
🚨 Usage: kotlin LogAnalyserKt <file>
No file path was given on the command line; run it as java -jar logAnalyser.jar server.log with a real file in that location.
🚨 Parsed 0 entries, even though the file clearly has log lines
Check the log format matches DATE TIME LEVEL MESSAGE with single spaces between the first three fields, and that the date contains a - and the time contains a : — parseLine relies on exactly that shape.
GroovyJenkinsfile
// Jenkinsfile &mdash; compiles and tests automatically every time the code changes.
pipeline {
    agent any                                  // run on any available machine

    stages {
        stage('Get the code') {
            steps { checkout scm }             // download the latest code
        }
        stage('Set up Kotlin') {
            steps {
                sh 'kotlinc -version'                             // confirm the compiler is installed
            }
        }
        stage('Compile and test') {
            steps {
                sh 'kotlinc logAnalyser.kt logAnalyserTest.kt -include-runtime -d build.jar'  // one real JVM jar, no build tool required
                sh 'java -cp build.jar:kotlin-test-junit.jar:junit.jar org.junit.runner.JUnitCore LogAnalyserTest'
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed &mdash; look above.' }
    }
}
🎯 Try this next — make it yours

You have a working log analyser. Extend it:

  1. Stream instead of loading the whole file. Use File(...).useLines { ... } for logs too large to fit comfortably in memory. (Teaches: lazy, sequence-based line processing.)
  2. Parse real timestamps. Use java.time.LocalDateTime instead of treating date and time as plain strings. (Teaches: real date/time parsing and its pitfalls.)
  3. Filter by a time range. Only count entries between two given timestamps. (Teaches: comparing parsed LocalDateTime values.)
  4. Output JSON. Serialize the report instead of printing plain text. (Teaches: structuring output for another program to consume.)
What you learned
You learned data class for a strongly-typed parsed record, mapNotNull for quietly dropping what fails to parse, and groupingBy { }.eachCount() as Kotlin’s built-in answer to a frequency counter. Related reference: Data Classes, Collections in Kotlin.