← thecodex.expert · The Codex Family of Knowledge
Tier 3 · Upper-Intermediate · Rust Project

Log Analyser

Parse server log files to count errors, find the busiest hours, and spot the top offenders. Turn raw logs into insight.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a log analyser: read a server log file, count how many entries are errors versus normal, find which hour had the most traffic, and list the most frequent error messages. It teaches parsing semi-structured text at scale and aggregating it into useful insight — a daily task in operations.

Where this shows up: monitoring and observability, debugging production issues, security analysis, performance tuning. When something breaks at 3am, the person who can analyse the logs is the one who fixes it.

2 How to Think About It

Think about turning lines into counts, before any code:

The plan — in plain English
1. Read the log file line by line. → 2. Parse each line into its parts: hour, level (INFO/ERROR), message. → 3. Count as you go: errors, entries per hour, message frequencies. → 4. Report the totals and the top items.

Read log lines

Parse each: time, level, message

Count errors

Count entries per hour

Count message frequency

Report insights

3 The Build — explained part by part

Here is the complete analyser. Rust has no built-in equivalent of Python’s collections.Counter or a stdlib regex engine, so this project builds its counting and “top N” logic explicitly with a plain HashMap and a hand-written line parser.

Rustsrc/main.rs
use std::collections::HashMap;
use std::env;
use std::fs::File;
use std::io::{BufRead, BufReader};

#[derive(Debug, PartialEq)]
struct Entry {
    hour: String,
    level: String,
    message: String,
}

/// Parses one log line shaped like `2026-06-24 14:30:00 ERROR Database
/// timeout`. Splits on whitespace but caps it at 4 pieces, so the message
/// (which may itself contain spaces) stays whole as the last piece instead
/// of being split further. Lines that do not look like a date + time +
/// level + message are rejected rather than half-parsed.
fn parse_line(line: &str) -> Option<Entry> {
    let parts: Vec<&str> = line.splitn(4, ' ').collect();
    if parts.len() != 4 {
        return None;
    }
    let [date, time, level, message] = [parts[0], parts[1], parts[2], parts[3]];
    if !date.contains('-') || !time.contains(':') {
        return None; // does not look like "YYYY-MM-DD HH:MM:SS LEVEL message"
    }
    let hour = time.split(':').next()?.to_string();
    Some(Entry {
        hour,
        level: level.to_string(),
        message: message.to_string(),
    })
}

#[derive(Debug, PartialEq)]
struct Report {
    total: usize,
    errors: usize,
    busiest_hour: Option<(String, u32)>,
    top_errors: Vec<(String, u32)>,
}

/// Rust's standard library has no `Counter`-style type, so this hand-builds
/// one with a `HashMap<String, u32>` and then sorts a copy of it into a
/// ranked `Vec` — exactly the pattern Go's version of this project uses,
/// since Go's standard library has no built-in "most common" helper either.
fn top_n(counts: &HashMap<String, u32>, n: usize) -> Vec<(String, u32)> {
    let mut pairs: Vec<(String, u32)> = counts.iter().map(|(k, v)| (k.clone(), *v)).collect();
    pairs.sort_by(|a, b| b.1.cmp(&a.1).then(a.0.cmp(&b.0)));
    pairs.truncate(n);
    pairs
}

fn analyse(lines: &[String]) -> Report {
    let mut total = 0;
    let mut errors = 0;
    let mut by_hour: HashMap<String, u32> = HashMap::new();
    let mut by_message: HashMap<String, u32> = HashMap::new();

    for line in lines {
        let entry = match parse_line(line) {
            Some(e) => e,
            None => continue,
        };
        total += 1;
        *by_hour.entry(entry.hour).or_insert(0) += 1;
        if entry.level == "ERROR" {
            errors += 1;
            *by_message.entry(entry.message).or_insert(0) += 1;
        }
    }

    let busiest_hour = top_n(&by_hour, 1).into_iter().next();
    let top_errors = top_n(&by_message, 3);

    Report { total, errors, busiest_hour, top_errors }
}

fn main() {
    let path = env::args().nth(1).unwrap_or_else(|| "server.log".to_string());
    let file = match File::open(&path) {
        Ok(f) => f,
        Err(e) => {
            eprintln!("Could not open {path}: {e}");
            return;
        }
    };

    let lines: Vec<String> = BufReader::new(file).lines().map_while(Result::ok).collect();
    let report = analyse(&lines);

    println!("Total entries: {}", report.total);
    println!("Errors: {}", report.errors);
    if let Some((hour, count)) = &report.busiest_hour {
        println!("Busiest hour: (\"{hour}\", {count})");
    }
    print!("Top errors: [");
    let rendered: Vec<String> = report
        .top_errors
        .iter()
        .map(|(msg, count)| format!("(\"{msg}\", {count})"))
        .collect();
    println!("{}]", rendered.join(", "));
}

#[cfg(test)]
mod tests {
    use super::*;

    fn lines(text: &str) -> Vec<String> {
        text.lines().map(String::from).collect()
    }

    #[test]
    fn a_line_parses_into_hour_level_and_message() {
        let entry = parse_line("2026-06-24 14:30:00 ERROR Database timeout").unwrap();
        assert_eq!(entry.hour, "14");
        assert_eq!(entry.level, "ERROR");
        assert_eq!(entry.message, "Database timeout");
    }

    #[test]
    fn errors_are_counted_correctly() {
        let log = lines(
            "2026-06-24 14:30:00 INFO Server started\n\
             2026-06-24 14:31:00 ERROR Database timeout\n\
             2026-06-24 14:32:00 ERROR Disk full\n",
        );
        let report = analyse(&log);
        assert_eq!(report.total, 3);
        assert_eq!(report.errors, 2);
    }

    #[test]
    fn the_busiest_hour_is_identified() {
        let log = lines(
            "2026-06-24 09:00:00 INFO a\n\
             2026-06-24 14:00:00 INFO b\n\
             2026-06-24 14:05:00 INFO c\n\
             2026-06-24 14:10:00 INFO d\n",
        );
        let report = analyse(&log);
        assert_eq!(report.busiest_hour, Some(("14".to_string(), 3)));
    }

    #[test]
    fn a_malformed_line_is_skipped_not_a_crash() {
        let log = lines("not a real log line\n2026-06-24 14:00:00 INFO fine\n");
        let report = analyse(&log);
        assert_eq!(report.total, 1);
    }
}
⚠ No in-browser playground here
Rust compiles to a real binary, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once Rust (via rustup) is installed.
What each part does — in plain words
line.splitn(4, ' ') — split on spaces, but stop after 4 pieces, so the message (which may itself contain spaces, like “Database timeout”) stays whole as the last piece instead of being split further.

if !date.contains('-') || !time.contains(':') { return None; } — a cheap sanity check that the line actually looks like a log line before trusting its shape further, since Rust has no regex in its standard library to validate the format more precisely.

HashMap<String, u32> used as a counter, with *by_hour.entry(hour).or_insert(0) += 1 — the entry API looks up a key and inserts a default first if missing, all in one expression.

fn top_n(counts: &HashMap<String, u32>, n: usize) -> Vec<(String, u32)> — since a HashMap has no defined iteration order and no built-in “most common” method, this copies the counts into a Vec of pairs and sorts it explicitly with sort_by, comparing by count descending. Exactly what Python’s Counter.most_common(n) does automatically — here it is one small function, written once, reused for both the busiest hour and the top errors.

BufReader::new(file).lines().map_while(Result::ok) — reads the file line by line without loading the whole thing into memory at once, the right approach for a log file that could be gigabytes long in a real system. map_while rather than filter_map stops at the first read error instead of potentially looping forever on a broken stream — clippy actually caught this exact distinction during development.
Common mistakes — and how to avoid them
✗ Splitting with plain line.split(' ') instead of splitn(4, ' ') — a multi-word error message gets chopped into extra pieces and the field count check breaks.
✓ Always cap the split count when the last field can contain the separator character.
✗ Assuming a HashMap iterates in the order you inserted keys — for (k, v) in &counts visits entries in an unspecified order every run.
✓ If order matters (like “top 3”), always sort explicitly afterward, never rely on iteration order.
✗ Using .filter_map(Result::ok) on a fallible line iterator — clippy flags this because it can loop forever if the underlying read keeps failing rather than reaching end-of-file.
✓ Use .map_while(Result::ok) instead, which stops at the first error.

4 Test & Prove Each Part

We test parsing a log line and the aggregation, using a few known lines.

A line parses into hour, level, and message
Errors are counted correctly
The busiest hour is identified
A malformed line is skipped, not treated as a crash
Rustsrc/main.rs (tests module)
#[cfg(test)]
mod tests {
    use super::*;

    fn lines(text: &str) -> Vec<String> {
        text.lines().map(String::from).collect()
    }

    #[test]
    fn a_line_parses_into_hour_level_and_message() {
        let entry = parse_line("2026-06-24 14:30:00 ERROR Database timeout").unwrap();
        assert_eq!(entry.hour, "14");
        assert_eq!(entry.level, "ERROR");
        assert_eq!(entry.message, "Database timeout");
    }

    #[test]
    fn errors_are_counted_correctly() {
        let log = lines(
            "2026-06-24 14:30:00 INFO Server started\n\
             2026-06-24 14:31:00 ERROR Database timeout\n\
             2026-06-24 14:32:00 ERROR Disk full\n",
        );
        let report = analyse(&log);
        assert_eq!(report.total, 3);
        assert_eq!(report.errors, 2);
    }

    #[test]
    fn the_busiest_hour_is_identified() {
        let log = lines(
            "2026-06-24 09:00:00 INFO a\n\
             2026-06-24 14:00:00 INFO b\n\
             2026-06-24 14:05:00 INFO c\n\
             2026-06-24 14:10:00 INFO d\n",
        );
        let report = analyse(&log);
        assert_eq!(report.busiest_hour, Some(("14".to_string(), 3)));
    }

    #[test]
    fn a_malformed_line_is_skipped_not_a_crash() {
        let log = lines("not a real log line\n2026-06-24 14:00:00 INFO fine\n");
        let report = analyse(&log);
        assert_eq!(report.total, 1);
    }
}

Run with cargo test. We feed analyse a few known log lines so every count can be checked by hand. This is how you trust an analyser before running it on millions of real lines.

5 The Interface

INPUTINPUTlog file
What it expects
2026-06-24 14:15:44 ERROR Database timeout
OUTPUTOUTPUTreport
What it returns
Total entries: 6
Errors: 3
Busiest hour: ("14", 4)
Top errors: [("Database timeout", 2), ...]

6 Run It & Automate It

Save the code as src/main.rs inside a Cargo project's src/ folder and run it with cargo run — Cargo compiles and executes in one step while you are experimenting, then cargo build --release gives you an optimized binary once you are done.

Run it locally
cargo run -- server.log
Point it at a server.log file to get a full report; defaults to server.log in the current directory if no argument is given.

A CI tool like Jenkins runs cargo test automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ cargo run -- server.log
Total entries: 6
Errors: 3
Busiest hour: ("14", 4)
Top errors: [("Database timeout", 2), ("Disk full", 1)]
If it breaks — how to fix it
🚨 Could not open server.log: No such file or directory (os error 2)
Create a server.log file in the same directory, or pass a path as the first argument.
🚨 Counts look wrong or entries are silently skipped.
Check your log format matches the parser exactly — the date field must contain a - and the time field a :, or parse_line rejects the whole line.
GroovyJenkinsfile
// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
    agent any                                  // run on any available machine

    stages {
        stage('Get the code') {
            steps { checkout scm }             // download the latest code
        }
        stage('Set up Rust') {
            steps {
                sh 'rustc --version'                // confirm Rust is installed
                sh 'cargo build'                     // compile, downloading any crates
            }
        }
        stage('Run the tests') {
            steps {
                sh 'cargo clippy -- -D warnings'     // catch obvious mistakes before running
                sh 'cargo test'                       // run every test, show each result
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed — look above.' }
    }
}
🎯 Try this next — make it yours
  1. Date filtering. Only include entries from one specific day. (Teaches: string comparison on the date field.)
  2. Use the real regex crate. If you have network access, replace the hand-written parser with a real regular expression. (Teaches: what a dependency buys you over hand-rolled parsing.)
  3. Live tail. Keep reading a log file as new lines arrive, like tail -f. (Teaches: following a growing file with std::thread::sleep and re-reading.)
What you learned
You learned to parse semi-structured text with hand-written checks instead of a regex crate, aggregate it with a HashMap used as a counter, and find top items with your own sort_by-based “most common” helper, since Rust has no built-in Counter either. Turning raw logs into insight is a vital operations skill in any language. Related: Collections, The Standard Library Tour.